Multi-mode guided controllable creative font generation method and system
Through the multimodal-guided controllable creative font generation method, the problems of insufficient artistic expression and readability, precise control and multilingual adaptability in the prior art are solved, and the balance between creativity, readability and adaptability is achieved, and the typesetting of non-Latin letters is supported, and higher artistic expression and design control are provided.
Patent Information
- Application Number
- CN202510234257.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The existing art glyph generation technology has shortcomings in artistic expression and readability, precise control, and multilingual adaptability, making it difficult to achieve a balance between cultural background and linguistic diversity.
The multi-modal guided controllable creative font generation method is adopted to obtain target text, target fonts and multi-modal data, and generate conversion prompts and artistic design prompts, and use the multi-mask guided diffusion process to perform glyph transformation and artistic glyph image generation.
A balance between creativity, readability and adaptability is achieved, supporting non-Latin letter typesetting, providing higher artistic expression and design controls, ensuring an organic integration of creativity and readability.
Smart Images

Figure CN120107416A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of artistic word generation, and in particular to a multi-modal guided controllable creative font generation method and system. Background Art
[0002] Artistic font generation is the process of transforming characters into visually expressive forms, designed to enhance artistry or convey specific meanings. It combines ordinary characters with stylistic and decorative elements to create font designs that are both readable and attractive. This fusion of vision and meaning can effectively attract the audience's attention and strengthen the delivery of information. In recent years, artistic font generation has been widely used in advertising design, typesetting, brand image and other fields, and has achieved remarkable results.
[0003] With the rise of multimodal large language models, researchers have tried to combine the advantages of large language models and large visual models to solve more complex multimodal tasks. Large language models are good at language understanding and reasoning, while large visual models perform well in visual tasks (such as image segmentation), especially under the guidance of language cues, they can effectively generate images that meet visual needs. Multimodal information refers to input that contains multiple different types of data at the same time, such as text, images, audio, video, etc. In this context, through multimodal art font generation, these different information sources are integrated, glyph transformation and style transfer are performed, and artistic, creative and expressive text graphics are created.
[0004] However, although existing technologies can achieve glyph transformation and style transfer, they still face some challenges:
[0005] (1) Balance between artistic expression and readability
[0006] Most existing methods are limited by fixed frameworks and concepts, lacking sufficient flexibility and adaptability, especially when dealing with cultural backgrounds and language diversity, and it is difficult to achieve a good balance between artistic expression and readability. This results in the lack of visual appeal of the generated results while sacrificing readability.
[0007] (2) Lack of precise control, including background and regional texture
[0008] Current end-to-end generation methods usually generate characters directly without providing enough flexibility for designers to participate in personalized adjustments. This limits the background control and fine management of regional textures in artistic glyph generation. In addition, many methods (such as WordArtDesigner and MetaDesigner) introduce complex backgrounds and artifacts during the generation process, resulting in increased complexity of fonts and applications.
[0009] (3) Language restrictions
[0010] The current methods for generating artistic glyphs mainly focus on the generation of Latin letters. When processing non-Latin characters (such as Chinese, Japanese, and Korean), many methods still have significant difficulties, such as TextDiffuser, TextDiffuser2, AnyText, and FontDiffuser. This limits the application of these methods in multilingual and multicultural environments and cannot meet the needs of users.
[0011] Therefore, existing artistic font generation technologies are insufficient in terms of artistic expression, readability, precise control, and cross-language adaptability. Summary of the invention
[0012] In order to solve the above problems, the present disclosure proposes a controllable creative font generation method and system with multimodal guidance, which solves the limitations of existing methods in artistic font generation by using creative texture generation of diffusion model through multimodal guidance, and provides better solutions in terms of artistic expression and readability, precise control and multi-language adaptability.
[0013] According to some embodiments, the present disclosure adopts the following technical solutions:
[0014] A multi-modal guided controllable creative font generation method, comprising:
[0015] Obtain target text, target font, and multimodal data for representing user intent;
[0016] Generate transformation hints and art design hints based on multimodal data;
[0017] Extract multiple paths from the target text and the target font, select the path with the highest similarity to perform glyph transformation according to the transformation prompt, and obtain the transformed image;
[0018] A multi-mask guided diffusion process is exploited to generate the final artistic glyph image for the transformed image using artistic design cues as hints.
[0019] According to some embodiments, the present disclosure adopts the following technical solutions:
[0020] A multi-modal guided controllable creative font generation system, comprising:
[0021] A data acquisition module is configured to: acquire a target text, a target font, and multimodal data for representing user intention;
[0022] The prompt generation module is configured to: generate conversion prompts and art design prompts based on the multimodal data;
[0023] The glyph transformation module is configured to: extract multiple paths from the target text and the target font, select the path with the highest similarity to perform glyph transformation according to the transformation prompt, and obtain a transformed image;
[0024] The image generation module is configured to generate a final artistic glyph image for the transformed image using artistic design cues as hints using a multi-mask guided diffusion process.
[0025] According to some embodiments, the present disclosure adopts the following technical solutions:
[0026] A computer program product comprises a computer program, wherein when the computer program is executed by a processor, the computer program implements the multi-modal guided controllable creative font generation method.
[0027] According to some embodiments, the present disclosure adopts the following technical solutions:
[0028] A non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the multi-modal guided controllable creative font generation method is implemented.
[0029] According to some embodiments, the present disclosure adopts the following technical solutions:
[0030] An electronic device comprises: a processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the multi-modal guided controllable creative font generation method.
[0031] Compared with the prior art, the present invention has the following beneficial effects:
[0032] The present invention provides a multi-modal guided, multi-level controlled art typesetting framework, which combines an iterative feedback mechanism to solve the key challenge of balancing creativity, readability and adaptability, while supporting typesetting systems with non-Latin letters; it is worth noting that the framework does not rely on training and is user-driven; while improving artistic expression, it combines design principles with generative models to simplify the generation process into a series of sequential steps. The framework transforms the input through iterative feedback, continuously optimizes the generation results, and ensures the organic integration of creativity and readability.
[0033] This paper adopts a chain of thought (CoT) method based on a multimodal large language model (MLLM). By extracting design features from multimodal input and integrating user intentions, abstract concepts are transformed into specific and actionable prompts to improve the effect of artistic expression.
[0034] In order to achieve multi-level control, the present invention extracts vector paths aligned with multi-modal inputs under multi-level control to ensure the uniformity of the output. In addition, a texture diffusion process that maintains background stability is implemented, which can accurately control the changes of different shapes, colors and textures and ensure readability, while solving the limitations of language through the diffusion process.
[0035] In order to provide better control and consistent output, the present invention designs a feedback module, which includes a scoring mechanism, label-based feedback, and an automated feedback mechanism. The module supports iterative optimization to ensure that the generated output meets high-quality aesthetic standards. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings constituting a part of the present disclosure are used to provide a further understanding of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation on the present disclosure.
[0037] Figure 1 This is a schematic diagram of the model structure of Example 1. DETAILED DESCRIPTION
[0038] The present disclosure is further described below in conjunction with the accompanying drawings and embodiments.
[0039] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present disclosure belongs.
[0040] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0041] Example 1
[0042] In one embodiment of the present disclosure, a method for generating a controllable creative font guided by multimodality is provided, comprising:
[0043] Step S1: obtaining a target text, a target font, and multimodal data for representing user intention;
[0044] Step S2: generating conversion prompts and art design prompts based on multimodal data;
[0045] Step S3: extract multiple paths from the target text and the target font, select the path with the highest similarity to perform glyph transformation according to the transformation prompt, and obtain the transformed image;
[0046] Step S4: Generate a final artistic glyph image for the transformed image using artistic design cues as hints using a multi-mask guided diffusion process.
[0047] As an embodiment, a multimodal guided controllable creative font generation method disclosed in the present invention uses a generation model to generate controllable creative fonts for input target text, target font and multimodal data used to represent user intentions, and finally obtains an artistic font image that meets the user's expectations. Through multimodal guidance, creative texture generation using a diffusion model is used to solve the limitations of existing methods in artistic font generation, especially to provide better solutions in terms of artistic expression and readability, precise control and multilingual adaptability. The generation model here, such as Figure 1 As shown, it includes four modules: multimodal intention extraction module, automatic path matching and conversion module, background-preserving texture generation module and iterative prompt feedback optimization module. Taking the target text as "Guilin" and the target font as "fontname.ttf" as an example, the four modules are explained.
[0048] 1. Multimodal Intent Extraction Module
[0049] Multimodal data is data provided by users to represent user intent. Text, images, audio, and video are all centered around the target text and target font. For example, the text and audio are an introduction to Guilin: "Karst peaks of various shapes stand in the distance, covered with lush vegetation. The sky is filled with brilliant clouds, with blue-purple and pink-orange colors interweaving to create a dreamy atmosphere. The calm water is like a mirror, perfectly reflecting the sky and mountains, making the scenery more ethereal. Some houses can be seen on the shore, adding a touch of life to the natural scenery, giving the whole a feeling of tranquility, remoteness, and overwhelming beauty."
[0050] This module is used to extract user intent from multimodal data and transform abstract concepts into specific generation prompts. In the input multimodal data (such as text, images, audio, video), the system can analyze and generate specific artistic glyph design prompts. For specific concepts in multimodal data (such as "peach", "maple tree"), the system can effectively generate matching artistic glyphs. However, for abstract concepts in multimodal data (such as "baroque" and "cold"), it is often difficult to provide detailed descriptions, so innovative prompt analysis and generation strategies are needed to obtain more detailed descriptions.
[0051] To analyze and extract specific prompts, a prompt analysis and generation strategy based on large language models is designed. The chain of thought (CoT) method of large language models is adopted, and specific prompts are extracted and generated through task-specific templates (i.e., the prompt engineering Prompt of large language models).
[0052] Specifically, receive the user's multimodal input, which is used to characterize the user's intention. Specifically, first receive the user's multimodal input U p ∈{text,image,audio,video}, and input it into a multimodal intention extraction template containing multiple design dimensions. Using large language models and prompt engineering, gradually generate transformation prompts P trans related to the user's intention and artistic design prompts P art .
[0053] For the transformation prompt P trans , by identifying keywords in the input (such as "mountain"), use it as the initial transformation prompt. Subsequent iterative optimization can adjust the transformation prompt to precisely control the transformation of glyphs and ensure alignment with the expected output.
[0054] For P art , a multi-step sub-prompt generation chain is designed. Each sub-prompt corresponds to a specific design dimension (such as "semantic theme", "font style", "layout and structure", "color and color palette", "shape and structure", etc.). Through these dimensions, comprehensively guide the generation process of artistic glyphs; by decomposing the generation process, ensure coverage of all relevant design elements, and support users to perform real-time fine-tuning through interaction to adjust each design dimension according to the user's personalized needs, such as color, shape, or layout, etc.
[0055] II. Automatic Path Matching and Transformation Module
[0056] This module is mainly responsible for extracting paths from the input target text and target font, and automatically selecting the most suitable path for glyph transformation according to the transformation prompt P trans . Specifically:[[]]
[0057] 1. Automatic Level Path Matching
[0058] By analyzing the transformation prompt P trans , select the most suitable path from multiple levels (such as strokes, paths, characters, and words). This process is divided into two stages:
[0059] (1) Path Preprocessing
[0060] Firstly, the open source font engine library FreeType is used to extract glyph outlines from user-input text and TrueType fonts, and each outline is converted into a cubic Bezier curve. Then, the Bezier curve is standardized through the vectorization command. Finally, the path of each character is segmented to generate the original path set W, and it is ensured that non-Latin characters can maintain shape consistency during the conversion process.
[0061] Specifically, based on the user's target text U w and the target font U fn , get the image I of the target text ori .
[0062] Decompose the path into the set of primitive paths W:
[0063] First, use the outline extraction function in the open source font engine library FreeType Extract the outline of each character and process it uniformly through the function Convert the outline to a cubic Bezier curve. The Bezier curve is a commonly used mathematical representation used to describe the font outline. There are three main types of curves in the glyph outline: linear, quadratic Bezier curves, and cubic Bezier curves. For convenience and unified processing, both linear Bezier curves and quadratic Bezier curves are converted to cubic Bezier curves.
[0064] Then, through the function Convert the Bezier chain to vector format, i.e. path commands in SVG format ( <path>Tags); these commands use a combination of commands such as M (move to point) and C (cubic Bezier curve) with control points to describe the character outline, and then add the necessary content to form an SVG file. If the input is a string, multiple characters will be converted to SVG files, and each character will generate an SVG file separately. Then, the font is scaled and aligned, and the layout of the entire string is standardized to ensure that the characters can be correctly centered within the given canvas size.
[0065] After that, the text and characters are segmented from the perspective of vectors (Path Dec.), with each path as the minimum segmentation unit, and the path segmentation is performed through the "M" command to form the basic segmentation unit L i For non-Latin characters, the number of control points that make up the character (string) is pre-calculated to ensure smoother transformations and better shape rendering consistency;
[0066] Finally, the original path set W is obtained:
[0067]
[0068] (2) Automatic path selection
[0069] Using CLIP encoder to convert the prompt P trans Each path is compared in a high-dimensional feature space, the similarity between them is calculated, and the path with the highest similarity is selected for conversion; this process ensures that the selected path can visually match the expected conversion hint while retaining the readability of the glyph.
[0070] Specifically, first, use the differentiable rasterizer DiffVG The original path set W is rasterized. Then, the encoder ε of the contrastive language-image pre-trained model CLIP is used to transform the prompt P trans And each rasterized path L in the original path set W i Encode them, map them into the same high-dimensional feature space, and calculate all combinable and P trans At the same time, in order to ensure that the selected path has sufficient deformation ability, the number of commands of the path is pre-calculated to evaluate the complexity of each path, and small-area or low-confidence paths are filtered out. If the number of commands of a path is greater than or equal to a certain threshold, the path is considered to have higher complexity and can present richer shape changes, which can better meet user expectations. Finally, a group of paths with the highest similarity score above this threshold are selected for transformation. This local transformation method can maintain the readability of the output.
[0071] The optimal path set L is selected as follows:
[0072]
[0073] Among them, F m (·) represents the automatic level path matching function, and cos represents the calculation of the cosine similarity score.
[0074] 2. Font transformation
[0075] Transform the optimal path set L into the transformation hint P trans Aligned deformation path sets
[0076] First, the optimal path set L generates the corresponding grating image I. Then, the image I is optimized and cropped to generate a random enhanced image Next, convert the prompt P trans and The input is fed into the frozen visual-language model, and the best path set L is transformed into the transition hint P by calculating the shape diffusion similarity loss SDS Loss. trans Align to a set of deformation paths under the guidance of Deformation Path Set The transformation is done through the following process:
[0077]
[0078] Among them, F t (·) represents a glyph transformation function.
[0079] Set the deformation path Align with the original path set W (Path Ali.) and rasterize Combine them (Path Com.) to generate the deformed image I mod In order to maintain the structure of the characters, by calculating L and The center point and size difference between them are adjusted using the bounding box; finally, the deformed image I mod Generate as follows:
[0080]
[0081] 3. Texture and background preservation module
[0082] In order to ensure the stability of the background and the precise control of the texture during the generation of artistic fonts, this module uses the U-Net model, which involves two stages: denoising and denoising. The denoising stage uses a diffusion process based on multi-mask guidance. In this process, the transformed image is used. As input, multiple masks are used to add noise to ensure accurate control of the area and maintain the stability of the background; in the denoising stage, noise is gradually removed during the reverse diffusion process, and finally an artistic font that meets the user's expectations is generated. During this denoising process, the feedback module repeatedly optimizes the artistic design hint P art , specifically:
[0083] 1. Add noise
[0084] Image I mod Convert to latent space representation z 0 , and use the region mask M reg To constrain the noise area, the area mask M reg It helps artistic expression, ensures a stable background, accurate texture, and controls the area, so as to accurately control the positioning of artistic fonts, specifically:
[0085] Image I mod Transformed into latent space representation via variational autoencoder (VAE) ε Then, use the region mask to constrain the noise area, thereby precisely controlling the position of artistic typography.
[0086] To extract the mask, two methods are designed:
[0087] (1) Automatic path alignment: From the aligned path set The contour information is extracted from the image and downsampled to fit the dimension of the latent space.
[0088] (2) User interaction: Use the Segment Anything Model to extract accurate masks and filter out masks that are too large or too small. An interactive mode is also provided, allowing users to generate M by clicking on the selection. reg , enhancing the accuracy and personalization of the process.
[0089] By element-wise multiplication, the region mask M reg Applied to the latent space representation z 0 In the forward process, it is the potential representation z 0 Add Gaussian noise of time step t to obtain the initial noisy potential representation x in :
[0090]
[0091] In order to keep the background stable during the generation process, an inverse mask is defined M reg ′ =1-M reg ; Then, M reg ′ With z 0 Perform element-wise multiplication to obtain the noise-free latent representation x.
[0092] At each time step t∈{1,...,T}, the Gaussian noise corresponding to time step t is added to the latent representation x, resulting in a noisy latent representation x in the background bg :x bg =q(x t |x).
[0093] Next, x in With x bg Concatenate to get the noise potential representation x at time step t t .
[0094] 2. Denoising
[0095] Introducing global masks Ensure clear edges of the background by:
[0096] Use Art Design Tips art As a hint, and introduce the global mask As the conditional information c, in order to keep the clear edge of the background, the predicted latent representation is obtained through the following denoising process
[0097]
[0098] Among them, ∈ θ (·,·) represents the noise predicted by the U-Net network.
[0099] This denoising process is repeated at each time step t∈{ε,...,1} until the predicted latent representation x is obtained. 0 , and then through the decoder Convert it into output image I output ; It should be noted that x bg Obtained through the forward process, and x in Optimized through a denoising process.
[0100] 4. Iterative Feedback Optimization Module
[0101] In order to further optimize the output results, this embodiment introduces a feedback module that combines user feedback and automated feedback. Based on user feedback and label feedback, this module adjusts the template in the multimodal intent extraction module (i.e., the prompt project Prompt of the large language model), and works together through a weighted adjustment mechanism to gradually refine and continuously optimize the generated results.
[0102] Specifically, it includes user feedback and label feedback. User feedback is the user guiding the system to adjust relevant design elements, such as color, shape, texture, etc. through rating and label feedback; for example, if the user gives a low rating to the color, the weight of color-related prompts will be enhanced; the automatic feedback part uses tools such as the CLIP model to automatically score the output and adjust the generation prompts, thereby continuously improving the quality of the model output; user feedback and automatic feedback interact with each other to ensure that the generation process can not only adapt to the user's personalized needs, but also improve the overall design quality. Through this mechanism, the generation process of artistic fonts can be precisely controlled, and more customized and high-quality output can be provided according to user needs.
[0103] The user feedback is divided into rating feedback and label feedback. The following is an explanation of the three types of feedback:
[0104] (1) Rating feedback: Users rate various aspects of the generation, prompting the model to adjust relevant elements based on the feedback. For example, if the "color" aspect scores low, the model will pay more attention to color in the subsequent generation process and prioritize the low-scoring parts through a weighted system, thereby optimizing the output effect.
[0105] (2) Label feedback: Users can annotate dimensions that they are dissatisfied with (e.g., material, shape, etc.) to guide the model to focus on these aspects. This directional feedback helps the model adjust content according to the user’s specific preferences, improving personalization and accuracy.
[0106] (3) Automated feedback: Use automatic scoring tools (such as CLIP, EasyOCR, etc.) to evaluate the output results and automatically adjust the generation prompts; automated feedback complements user feedback input and can make more extensive quality adjustments to the generated results, thereby ensuring continuous optimization of the output.
[0107] These feedback mechanisms are interconnected, with user feedback providing specific guidance and automated feedback ensuring overall quality adjustments, resulting in more accurate and personalized output results.
[0108] In order to verify the performance of the method in this embodiment, the method in this embodiment is compared with FontDiffuser, AnyText, TextDiffuser, TextDiffuser-2, DALLE 3 in ChatGPT-4o, and Stable Diffusion. The four indicators of CLIP score, DINO score, OCR confidence, and Hausdorff Distence (HD) are used to quantitatively evaluate the generated images. CLIP score and DINO score are used to evaluate the style consistency and similarity in the generated output. OCR confidence and Hausdorff Distence (HD) are used to evaluate the contour similarity. Table 1 shows the comparison results of these indicators.
[0109] Table 1
[0110]
[0111] It can be seen that AnyText and TextDiffuser performed poorly in the various indicators evaluated, especially in CLIP score and Hausdorff distance (HD), and both failed to effectively balance style and readability. This is mainly due to the limitations of their training data, which makes it difficult for them to generate diverse and accurate styles, especially when dealing with non-Latin characters. In contrast, DALLE 3 and Stable Diffusion performed well in CLIP and DINO scores and can better capture style consistency, but scored low in OCR confidence and HD indicators, indicating that they performed well in style alignment, but had problems with the accuracy of glyph structure, which easily led to distortion of text shape.
[0112] The method of this embodiment outperforms other models in all indicators and performs excellently, ensuring the consistency of style, richness of semantics, and accuracy of glyph structure, thereby generating artistic typography that is both beautiful and readable, and can provide more precise and diverse artistic font effects.
[0113] Example 2
[0114] In one embodiment of the present disclosure, a multi-modal guided controllable creative font generation system is provided, comprising:
[0115] A data acquisition module is configured to: acquire a target text, a target font, and multimodal data for representing user intention;
[0116] The prompt generation module is configured to: generate conversion prompts and art design prompts based on the multimodal data;
[0117] The glyph transformation module is configured to: extract multiple paths from the target text and the target font, select the path with the highest similarity to perform glyph transformation according to the transformation prompt, and obtain a transformed image;
[0118] The image generation module is configured to generate a final artistic glyph image for the transformed image using artistic design cues as hints using a multi-mask guided diffusion process.
[0119] Example 3
[0120] In one embodiment of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the computer program implements the multi-modal guided controllable creative font generation method.
[0121] Example 4
[0122] In one embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the multi-modal guided controllable creative font generation method is implemented.
[0123] Example 5
[0124] In one embodiment of the present disclosure, an electronic device is provided, including: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the multi-modal guided controllable creative font generation method.
[0125] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0126] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0127] Although the above describes the specific implementation methods of the present disclosure in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present disclosure. Technical personnel in the relevant field should understand that on the basis of the technical solution of the present disclosure, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present disclosure.< / path>
Claims
1. A multi-modal guided controllable creative font generation method, characterized in that: include: Obtain target text, target font, and multimodal data for representing user intent; Generate transformation hints and art design hints based on multimodal data; Extract multiple paths from the target text and the target font, select the path with the highest similarity to perform glyph transformation according to the transformation prompt, and obtain the transformed image; A multi-mask guided diffusion process is exploited to generate the final artistic glyph image for the transformed image using artistic design cues as hints.
2. A method for generating a controllable creative font guided by multimodality as claimed in claim 1, characterized in that: The generation of the conversion prompt is to identify keywords in the multimodal data and generate the conversion prompt; The generation of the art design prompts utilizes a sub-prompt generation chain to generate sub-prompts under different design dimensions.
3. A method for generating a controllable creative font guided by multimodality as claimed in claim 1, characterized in that: The multiple paths are extracted from the target text and the target font, specifically: Extracting glyph outlines from the target text and the target font, and converting each outline into a cubic Bezier curve; Standardize Bezier curves through vectorization commands; The path of each character is segmented to generate a set of original paths.
4. A method for generating a controllable creative font guided by multi-modality as claimed in claim 3, characterized in that: The method of selecting the path with the highest similarity according to the conversion prompt is to use the CLIP encoder to compare the conversion prompt and each path in the original path set in a high-dimensional feature space, calculate the similarity between them, select the path with the highest similarity, and form the best path set.
5. A method for generating a controllable creative font guided by multi-modality as claimed in claim 4, characterized in that: The glyph transformation is to transform the best path set into modified paths aligned with the transformation hint, align the modified paths with the original path set, and combine them by rasterization to generate a transformed image.
6. The method for generating a controllable creative font guided by multi-modality according to claim 1, characterized in that: The multi-mask guided diffusion process is specifically as follows: Convert the transformed image into a latent space representation and use a region mask to constrain the noisy region; Introduce a global mask to ensure clear edges of the background; Through the reverse process, denoising is gradually performed, and finally an artistic font image that meets the user's expectations is generated.
7. A multi-modal guided controllable creative font generation system, characterized in that: include: A data acquisition module is configured to: acquire a target text, a target font, and multimodal data for representing user intention; The prompt generation module is configured to: generate conversion prompts and art design prompts based on the multimodal data; The glyph transformation module is configured to: extract multiple paths from the target text and the target font, select the path with the highest similarity to perform glyph transformation according to the transformation prompt, and obtain a transformed image; The image generation module is configured to generate a final artistic glyph image for the transformed image using artistic design cues as hints using a multi-mask guided diffusion process.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating a controllable creative font guided by multimodality as described in any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by the processor, the method for generating controllable creative fonts guided by multimodality as described in any one of claims 1 to 6 is implemented.
10. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement a multimodal guided controllable creative font generation method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Complex scene text editing method and system based on stroke-level guide diffusion model
CN117475035A
Word art picture generation method based on text guidance
CN118228684A
Non-paired cross-modal medical image conversion method based on contour guide path regularization
CN119444893A
Artistic font generation method based on semantic input
CN119444925A
Object Search in Digital Images
US20210224312A1