Font style generation method and device, computer equipment and storage medium
By adopting a font generation model based on diffusion model in the font generation of few samples and fine-tuning of the LoRA module, combined with the fusion of multiple LoRA modules, the shortcomings of structural quality and style expressiveness in font generation are solved, and high-quality and diverse font generation are achieved.
Patent Information
- Application Number
- CN202510108483.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-13
AI Technical Summary
The existing few-sample font generation algorithms are susceptible to pattern crashes during the generation process, resulting in poor quality of the generated font structure and difficulty in effectively capturing fine-grained style characteristics, resulting in insufficient style expressiveness of the generated fonts.
A font generation model based on diffusion model is adopted, and the model is fine-tuned through the LoRA module, combined with the fusion of multiple LoRA modules, an independent style space is built to improve the diversity and style refinement of font generation.
By better capturing complex font structures, the generated target font images are more coherent and high-quality, reducing the demand for computing resources and improving the diversity of font generation and style refinement consistency.
Smart Images

Figure CN119991850A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of computer vision technology, and in particular, relates to a font style generation method, device, computer equipment and storage medium. Background Art
[0002] Few-Shot Font Generation (FFG) is a technology that generates new font styles using limited font samples. Its goal is to generate new fonts with diverse and fine-grained features by using a small number of reference font images, which can significantly reduce the time and cost required for manually designing fonts and organizing font datasets.
[0003] At present, few-shot font generation algorithms usually follow the paradigm of style-content decoupling. First, the content (character shape, outline, etc.) of the font is separated from the style (strokes, style, etc.), and the conversion process from the source font to the target font is learned by using a generative adversarial network (GAN) or other generative models. The key to few-shot font generation is to be able to effectively extract features from limited samples and generate new fonts through the learning ability of the model. When using a generative adversarial network (GAN) for font generation, since GAN is easily affected by mode collapse during the generation process, the generated fonts often fail to achieve ideal results in terms of structural quality, especially in terms of pen coherence and details. In addition, existing GAN methods often cannot effectively capture fine-grained style features in font style learning, resulting in insufficient style expression of the generated fonts. Therefore, it is necessary to use a fine-tuning algorithm to fine-tune the GAN method. However, traditional fine-tuning algorithms consume a lot of computing resources, especially when dealing with complex models, which may lead to extended model training time and resource waste. Summary of the invention
[0004] The present application provides a font style generation method, apparatus, computer device and storage medium, aiming to solve at least one of the above-mentioned technical problems in the prior art to a certain extent.
[0005] In order to solve the above problems, this application provides the following technical solutions:
[0006] A font style generating method, comprising:
[0007] Construct a font generation model based on the diffusion model;
[0008] The LoRA module is used to fine-tune the font generation model based on the diffusion model to obtain an optimized font generation model;
[0009] A source font image and a reference font image are acquired, the source font image and the reference font image are input into the optimized font generation model, and a target font image is generated by the optimized font generation model.
[0010] The technical solution adopted by the embodiment of the present application also includes: the construction of a font generation model based on a diffusion model is specifically:
[0011] The font generation model based on the diffusion model adopts the UNet network and is built on the basis of LDM. The cross attention layer is introduced into the font generation model through LDM, and the DDIM sampling strategy is used for gradual denoising.
[0012] The technical solution adopted in the embodiment of the present application also includes: the LoRA module is used to fine-tune the font generation model based on the diffusion model to obtain an optimized font generation model, specifically:
[0013] The LoRA module is injected into the cross-attention layer of the UNet network, and the cross-attention layer includes a linear layer. The linear layer calculates the attention mechanism of keys, queries and values to fine-tune the parameters of the font generation model based on the diffusion model.
[0014] The technical solution adopted in the embodiment of the present application also includes: the LoRA module is used to fine-tune the font generation model based on the diffusion model to obtain an optimized font generation model, including:
[0015] Collecting fonts to be trained, classifying the fonts to be trained, and using a clustering algorithm to cluster and generate M basic style fonts;
[0016] The font to be trained is used to train a font generation model based on a diffusion model, and in the training stage, a similarity weight between the font to be trained and M basic style fonts is calculated;
[0017] The corresponding LoRA modules are trained for the M basic style fonts respectively, and the LoRA module parameters in the training process are adjusted using the similarity weights;
[0018] During the sampling process, the basic style fonts are used as prior knowledge of the fonts to be trained, and all LoRA modules are combined according to the similarity weights between the font to be trained and each basic style font to generate a font generation model for learning fine-grained font styles.
[0019] The technical solution adopted in the embodiment of the present application also includes: the clustering algorithm is used to cluster and generate M basic style fonts, specifically:
[0020] The K-Means unsupervised clustering algorithm is used to dynamically cluster the fonts to be trained according to their font styles, and the fonts to be trained with similar styles are classified into the same cluster to obtain M basic style fonts.
[0021] The technical solution adopted by the embodiment of the present application also includes: the similarity weight between the font to be trained and the M-type basic style fonts is calculated as follows:
[0022]
[0023] Among them, S i The font F to be trained i The style characteristics of Basic style font The style characteristics of The font F to be trained i And basic style fonts The L1 distance between them, τ represents the temperature parameter, and the normalized exponential function is used to convert Convert to weights and get the font F to be trained i With base font The similarity weight between
[0024] The technical solution adopted by the embodiment of the present application also includes: obtaining a source font image and a reference font image, inputting the source font image and the reference font image into the optimized font generation model, and generating a target font image through the optimized font generation model, specifically:
[0025] Get the source font image I that provides character information respectively c and a reference font image I providing style information s ;
[0026] The source font image I is respectively encoded by a content encoder and a style encoder. c and reference font image I s Perform encoding processing to obtain content feature F c and style features F s ;
[0027] The content feature F c and style features F s After dimension compression, the two are merged and used as the condition y in the diffusion model as input;
[0028] From the source font image I through the VAE encoder ε c The compressed latent space is extracted, a forward diffusion process is performed in the latent space, and during the forward diffusion process, the condition y is injected into the UNet network to guide the generation, and the latent space feature z is obtained after t steps.0 ;
[0029] The latent space feature z is decoded by VAE decoder D 0 Decode and generate the target font image I g .
[0030] Another technical solution adopted by the embodiment of the present application is: a font style generating device, comprising:
[0031] Model building module: used to build a font generation model based on the diffusion model;
[0032] Model fine-tuning module: used to fine-tune the font generation model based on the diffusion model using the LoRA module to obtain an optimized font generation model;
[0033] Font generation module: used for acquiring a source font image and a reference font image, inputting the source font image and the reference font image into the optimized font generation model, and generating a target font image through the optimized font generation model.
[0034] Another technical solution adopted by the embodiment of the present application is: a computer device, the computer device includes a processor and a memory coupled to the processor, wherein:
[0035] The memory stores program instructions for implementing the font style generation method;
[0036] The processor is used to execute the program instructions stored in the memory to control the font style generation method.
[0037] Another technical solution adopted by the embodiment of the present application is: a storage medium storing program instructions executable by a processor, wherein the program instructions are used to execute the font style generating method.
[0038] Compared with the prior art, the beneficial effects of the embodiments of the present application are as follows: the font style generation method, device, computer equipment and storage medium of the embodiments of the present application construct a font generation model by introducing a diffusion model, and fine-tuning the font generation model using the LoRA module, which can better capture the complex font structure during the font generation process, making the generated target font image more coherent and high-quality. At the same time, LoRA significantly reduces the amount of parameters that need to be fine-tuned through low-rank decomposition technology, thereby reducing the demand for computing resources and being able to perform micro-model adjustments more efficiently under conditions of few samples. And through the fusion of multiple LoRAs, more sophisticated font style learning is achieved to construct an independent style space, which not only improves the diversity of font generation, but also more accurately captures the subtle style differences between different fonts, so that the model can be adaptively adjusted for different styles, achieving higher style refinement and consistency, thereby improving the quality and diversity of font generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a flowchart of a font style generating method according to an embodiment of the present application;
[0040] Figure 2 A schematic diagram of a font generation model structure based on a diffusion model according to an embodiment of the present application;
[0041] Figure 3 This is a schematic diagram of the LoRA module structure of an embodiment of the present application;
[0042] Figure 4 This is a schematic diagram of the integration of the LoRA modules in an embodiment of the present application;
[0043] Figure 5 This is a schematic diagram of the structure of a font style generating device according to an embodiment of the present application;
[0044] Figure 6 A schematic diagram of the computer device structure of an embodiment of the present application;
[0045] Figure 7 A schematic diagram of the structure of a storage medium according to an embodiment of the present application. DETAILED DESCRIPTION
[0046] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0047] The terms "first", "second", "third" in this application are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Thus, the features defined as "first", "second", "third" can expressly or implicitly include at least one of the features. In the description of this application, the meaning of "multiple" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. In the embodiments of this application, all directional indications (such as up, down, left, right, front, back...) are only used to explain the relative position relationship, movement, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication also changes accordingly. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or computer device that includes a series of steps or units is not limited to the steps or units listed, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or computer devices.
[0048] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0049] Specifically, see Figure 1 , is a flow chart of the font style generation method of the embodiment of the present application. The font style generation method of the embodiment of the present application comprises the following steps:
[0050] S100: Construct a font generation model based on the diffusion model;
[0051] In this step, if Figure 2As shown, it is a schematic diagram of the font generation model structure based on the diffusion model of the embodiment of the present application. The font generation model based on the diffusion model adopts the UNet network and is built on the basis of the LDM (Latent Diffusion Model). The LDM introduces a cross-attention layer in the font generation model and adopts the DDIM (Denoising Diffusion Implicit Models) sampling strategy. During the sampling process, DDIM uses a deterministic sampling method and continuous noise prediction to gradually remove noise from the image. In each step of the noise process, the low-resolution semantic information captured by the encoder is used, and the high-resolution detail information of the decoder is combined to generate the final high-quality target image, which can achieve a higher quality generation effect with fewer processing steps.
[0052] It can be understood that the embodiments of the present application may also use a variational autoencoder (VAE) or a generative adversarial network (GAN) as a substitute for the LDM model, and a similar generation effect can be achieved by performing specific network adjustments in combination with a style encoder, etc.
[0053] S110: using the LoRA (Low-Rank Adaptation) module to fine-tune the font generation model based on the diffusion model to obtain an optimized font generation model;
[0054] In this step, LoRA is a technology used to reduce the amount of fine-tuning parameters for large models. In large-scale pre-trained models, the weights of the model are usually represented by large matrices, which contain a large number of parameters. LoRA reduces the amount of parameters by decomposing the model weight updates into two low-rank matrices, and significantly reduces the computational overhead in the fine-tuning process. Specifically, if the update amount of the weight matrix W is denoted as ΔW, LoRA represents ΔW as two low-rank matrices A and B through low-rank decomposition, that is, ΔW = A × B. By only training these two small-scale low-rank matrices, LoRA can achieve task-specific adaptive adjustments while keeping the pre-training weights W unchanged, thereby reducing the amount of parameters and computational overhead, while maintaining model performance and significantly improving fine-tuning efficiency. Specifically, Figure 3As shown, it is a schematic diagram of the LoRA module structure of an embodiment of the present application. The LoRA module includes a cross-attention module, a residual block and a spatial transformation station. When the LoRA module is used to fine-tune the font generation model based on the diffusion model, the LoRA module is injected into the cross-attention layer of the UNet network. The cross-attention layer includes multiple Linear layers (linear layers). The attention mechanism of the key (Key), query (Query) and value (Value) is calculated through multiple Linear layers, and the parameters of the font generation model based on the diffusion model are fine-tuned, so that the fine-tuned font generation model can use the basic font style to represent other font styles.
[0055] Specifically, fine-tuning the font generation model based on the diffusion model using the LoRA module includes the following steps:
[0056] S111: Collect the fonts to be trained, classify the fonts to be trained, use a clustering algorithm to cluster and generate M basic style fonts, and use the M basic style fonts as the core objects of the LoRA module fine-tuning font generation model;
[0057] Among them, the embodiment of the present application adopts the K-Means unsupervised clustering algorithm to dynamically cluster the training fonts according to the font style, classifies the training fonts of similar styles into the same cluster, and obtains M types of basic style fonts, thereby effectively establishing the connection between styles and improving the consistency of styles.
[0058] S112: training a font generation model based on a diffusion model using the font to be trained, and calculating a similarity weight between the font to be trained and the M-type basic style fonts during the training phase;
[0059] There are different degrees of similarity between the style features of the font to be trained and each type of basic style font. The embodiment of the present application quantifies the similarity by calculating the distance between them. The similarity weight calculation method is specifically as follows:
[0060]
[0061] Among them, the font to be trained F i Corresponding style feature S i , basic style font Corresponding style characteristics Calculate the font F to be trained i And basic style fonts The L1 distance between Then the distance is normalized by the exponential function Converted into weights to measure the similarity between them. Where τ represents the temperature parameter, which is used to adjust the impact of distance on weight distribution. Parameter Indicates the font to be trained Fi With base font The similarity weight between Contains M values, each value corresponds to the i-th font F i Style characteristics and basic style fonts The similarities between the style features.
[0062] S113: During the training process, the corresponding LoRA module is trained for each category of basic style fonts respectively, and the parameters of the LoRA module during the training process are adjusted using the similarity weight, so that the LoRA module can fit the style of the target font more accurately;
[0063] S114: In the sampling process, the basic style font is used as the prior knowledge of the font to be trained, and the similarity weight between the font to be trained and each basic style font is calculated. Combine multiple LoRA modules to generate a font generation model that learns fine-grained font styles;
[0064]
[0065] Each font has a set of correlation coefficients with the M basic style fonts. Multiple LoRA modules are used to establish the association between each font and multiple LoRAs. Formula (2) is used to combine these M LoRA modules and inject them into the model for joint training to generate a font generation model that learns fine-grained font styles. Taking the kth basic style font as an example, its corresponding LoRA module is A k and B k , A k and B k The corresponding similarity weight Multiply, and then add the weighted results of all LoRA modules to the pre-trained weight W 0 Add them together to get the final weight W′.
[0066] Among them, Figure 4 As shown, this is a schematic diagram of the LoRA module fusion of an embodiment of the present application. The embodiment of the present application introduces the LoRA module during the fine-tuning process, which greatly reduces the consumption of computing resources and can achieve efficient font generation with less time and cost. And by dynamically adjusting the parameters of LoRA by fusing the parameters of multiple LoRA modules, flexible adaptation of fonts of different styles is achieved, so that the basic style font can fully learn and represent the style characteristics of other fonts, thereby constructing an independent and flexible adaptive parameter space for the target font to be generated, which not only improves the diversity of font generation, but also allows users to customize personalized styles according to their own needs to meet the personalized needs of different users.
[0067] S120: Acquire a source font image and a reference font image, input the source font image and the reference font image into an optimized font generation model, and generate a target font image through the font generation model;
[0068] In this step, if Figure 2 As shown, the target font image generation process of the optimized font generation model includes the following steps:
[0069] S121: Obtain source font images I that provide character information respectively c and a reference font image I providing style information s ;
[0070] S122: The source font image I is respectively encoded by the content encoder and the style encoder. c and reference font image I s Perform encoding processing to obtain content feature F c and style features F s ;
[0071] S123: Content feature F c and style features F s After dimension compression, merge them and use them as the conditional y in the input diffusion model;
[0072] S124: From the source font image I through the VAE encoder ε c The compressed latent space is extracted from the latent space, and the forward diffusion process is performed in the latent space. During the forward diffusion process, the condition y is injected into the U-Net network to guide the generation, and the latent space feature z is obtained after t steps. 0 ;
[0073] S125: Through VAE decoder D, the latent space feature z 0 Decode and finally generate the target font image I g .
[0074] Based on the above, the font style generation method of the embodiment of the present application constructs a font generation model by introducing a diffusion model, and uses the LoRA module to fine-tune the font generation model. It can better capture the complex font structure during the font generation process, making the generated target font image more coherent and high-quality. At the same time, LoRA significantly reduces the amount of parameters that need to be fine-tuned through low-rank decomposition technology, thereby reducing the demand for computing resources and being able to perform micro-model adjustments more efficiently under small sample conditions. And through the fusion of multiple LoRAs, more sophisticated font style learning is achieved to construct an independent style space, which not only improves the diversity of font generation, but also more accurately captures the subtle style differences between different fonts, so that the model can be adaptively adjusted for different styles, achieving higher style refinement and consistency, thereby improving the quality and diversity of font generation.
[0075] See also Figure 5 , is a schematic diagram of the structure of a font style generating device according to an embodiment of the present application. The font style generating method device 40 according to an embodiment of the present application comprises:
[0076] Model construction module 41: used to construct a font generation model based on a diffusion model; wherein, the font generation model based on the diffusion model adopts a UNet network and is constructed on the basis of an LDM (Latent Diffusion Model). A cross-attention layer is introduced into the font generation model through the LDM, and a DDIM (Denoising Diffusion Implicit Models) sampling strategy is adopted. During the sampling process, DDIM uses a deterministic sampling method and continuous noise prediction to gradually remove noise from the image. In each step of the noise process, the low-resolution semantic information captured by the encoder is used, and the high-resolution detail information of the decoder is combined to generate the final high-quality target image, which can achieve a higher quality generation effect with fewer processing steps.
[0077] Model fine-tuning module 42: used to use the LoRA module to fine-tune the diffusion model-based font generation model to obtain an optimized font generation model; wherein, LoRA is a technology used to reduce the amount of fine-tuning parameters of large models. In large-scale pre-trained models, the weights of the model are usually represented by large matrices, which contain a large number of parameters. LoRA reduces the amount of parameters by decomposing the model weight update into two low-rank matrices, and significantly reduces the computational overhead in the fine-tuning process. Specifically, if the update amount of the weight matrix W is denoted as ΔW, LoRA represents ΔW as two low-rank matrices A and B through low-rank decomposition, that is, ΔW=A×B. LoRA can achieve task-specific adaptive adjustments while keeping the pre-training weights W unchanged by only training these two small-scale low-rank matrices, thereby reducing the amount of parameters and computational overhead, while maintaining model performance and significantly improving fine-tuning efficiency. Specifically, Figure 3 As shown, it is a schematic diagram of the LoRA module structure of an embodiment of the present application. The LoRA module includes a cross-attention module, a residual block and a spatial transformation station. When the LoRA module is used to fine-tune the font generation model based on the diffusion model, the LoRA module is injected into the cross-attention layer of the UNet network. The cross-attention layer includes multiple Linear layers (linear layers). The attention mechanism of the key (Key), query (Query) and value (Value) is calculated through multiple Linear layers, and the parameters of the font generation model based on the diffusion model are fine-tuned, so that the fine-tuned font generation model can use the basic font style to represent other font styles.
[0078] Specifically, the LoRA module is used to fine-tune the font generation model based on the diffusion model, including:
[0079] Collect the fonts to be trained, classify them, use a clustering algorithm to cluster and generate M basic style fonts, and use the M basic style fonts as the core objects of the LoRA module fine-tuning font generation model; wherein, the embodiment of the present application uses the K-Means unsupervised clustering algorithm to dynamically cluster the training fonts according to the font style, and classifies the training fonts of similar styles into the same cluster to obtain M basic style fonts, thereby effectively establishing the connection between the styles and improving the style consistency.
[0080] The font to be trained is used to train the font generation model based on the diffusion model, and in the training stage, the similarity weight between the font to be trained and the M-type basic style fonts is calculated;
[0081] The corresponding LoRA modules are trained for each category of basic style fonts, and the LoRA module parameters in the training process are adjusted using similarity weights, so that the LoRA modules can fit the style of the target font more accurately; specifically, when training the LoRA modules, the embodiment of the present application uses multiple low-rank adaptation matrices (LoRA) to construct an independent parameter space based on style clustering. Each LoRA module represents a specific font style or style change through an independent low-rank matrix, so that each LoRA module can be selected and weighted combined according to the input style requirements during the generation process, so that the subtle differences between different font styles can be flexibly handled, so that the model can be adaptively adjusted for different styles to achieve higher style refinement and consistency, thereby improving the quality and diversity of font generation.
[0082] In the inference stage, the parameters of M LoRA modules are fused according to the similarity weights between the target font and the basic style font to obtain a font generation model that can generate a custom font style. Figure 4 As shown, this is a schematic diagram of the LoRA module fusion of an embodiment of the present application. The embodiment of the present application introduces the LoRA module during the fine-tuning process, which greatly reduces the consumption of computing resources and can achieve efficient font generation with less time and cost. And by dynamically adjusting the parameters of LoRA by fusing the parameters of multiple LoRA modules, flexible adaptation of fonts of different styles is achieved, so that the basic style font can fully learn and represent the style characteristics of other fonts, thereby constructing an independent and flexible adaptive parameter space for the target font to be generated, which not only improves the diversity of font generation, but also allows users to customize personalized styles according to their own needs to meet the personalized needs of different users.
[0083] The font generation module 43 is used to obtain a source font image and a reference font image, input the source font image and the reference font image into the optimized font generation model, and generate a target font image through the optimized font generation model; specifically, the target font image generation process includes:
[0084] Get the source font image I that provides character information respectively c and a reference font image I providing style information s ;
[0085] The source font image I is encoded by the content encoder and the style encoder respectively. c and reference font image I s Perform encoding processing to obtain content feature F c and style features F s ;
[0086] The content feature F c and style features F s After dimension compression, merge them and use them as the conditional y in the input diffusion model;
[0087] From the source font image I through the VAE encoder ε c The compressed latent space is extracted from the latent space, and the forward diffusion process is performed in the latent space. During the forward diffusion process, the condition y is injected into the U-Net network to guide the generation, and the latent space feature z is obtained after t steps. 0 ;
[0088] Through the VAE decoder D, the latent space feature z 0 Decode and finally generate the target font image I g .
[0089] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0090] The device provided in the embodiment of the present application can be applied in the aforementioned method embodiment. For details, please refer to the description of the aforementioned method embodiment, which will not be repeated here.
[0091] See also Figure 6 , is a schematic diagram of the computer device structure of an embodiment of the present application. The computer device 50 includes:
[0092] A memory 51 storing executable program instructions;
[0093] A processor 52 connected to the memory 51;
[0094] The processor 52 is used to call the executable program instructions stored in the memory 51 and perform the following steps: construct a font generation model based on a diffusion model; use the LoRA module to fine-tune the font generation model based on the diffusion model to obtain an optimized font generation model; obtain a source font image and a reference font image, input the source font image and the reference font image into the optimized font generation model, and generate a target font image through the optimized font generation model.
[0095] The processor 52 may also be referred to as a CPU (Central Processing Unit). The processor 52 may be an integrated circuit chip having signal processing capabilities. The processor 52 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0096] See also Figure 7 , which is a schematic diagram of the structure of the storage medium of the embodiment of the present application. The storage medium of the embodiment of the present application stores program instructions 61 capable of implementing the following steps: constructing a font generation model based on a diffusion model; using a LoRA module to fine-tune the font generation model based on the diffusion model to obtain an optimized font generation model; obtaining a source font image and a reference font image, inputting the source font image and the reference font image into the optimized font generation model, and generating a target font image through the optimized font generation model.
[0097] Among them, the program instruction 61 can be stored in the above-mentioned storage medium in the form of a software product, including several instructions for a computer device (which can be a personal computer, a server, or a network computer device, etc.) or a processor (processor) to execute all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), disk or optical disk and other media that can store program instructions, or terminal computer devices such as computers, servers, mobile phones, tablets, etc. Among them, the server can be an independent server, or it can be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms.
[0098] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the system embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0099] In addition, each functional unit in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional units. The above is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the description and drawings of this application, or directly or indirectly used in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. A font style generation method, characterized in that: include: Construct a font generation model based on the diffusion model; The LoRA module is used to fine-tune the font generation model based on the diffusion model to obtain an optimized font generation model; A source font image and a reference font image are acquired, the source font image and the reference font image are input into the optimized font generation model, and a target font image is generated by the optimized font generation model.
2. The font style generation method according to claim 1, characterized in that: The construction of the font generation model based on the diffusion model is specifically as follows: The font generation model based on the diffusion model adopts the UNet network and is built on the basis of LDM. A cross attention layer is introduced into the font generation model through LDM, and a DDIM sampling strategy is used for gradual denoising.
3. The font style generation method according to claim 2, characterized in that: The LoRA module is used to fine-tune the font generation model based on the diffusion model to obtain an optimized font generation model, specifically: The LoRA module is injected into the cross-attention layer of the UNet network, and the cross-attention layer includes a linear layer. The linear layer calculates the attention mechanism of keys, queries and values to fine-tune the parameters of the font generation model based on the diffusion model.
4. The font style generation method according to claim 3, characterized in that: The LoRA module is used to fine-tune the font generation model based on the diffusion model to obtain an optimized font generation model, including: Collecting fonts to be trained, classifying the fonts to be trained, and using a clustering algorithm to cluster and generate M basic style fonts; The font to be trained is used to train a font generation model based on a diffusion model, and in the training stage, a similarity weight between the font to be trained and M basic style fonts is calculated; The corresponding LoRA modules are trained for the M basic style fonts respectively, and the LoRA module parameters in the training process are adjusted using the similarity weights; During the sampling process, the basic style fonts are used as prior knowledge of the fonts to be trained, and all LoRA modules are combined according to the similarity weights between the font to be trained and each basic style font to generate a font generation model for learning fine-grained font styles.
5. The font style generation method according to claim 4, characterized in that: The clustering algorithm is used to cluster and generate M basic style fonts, specifically: The K-Means unsupervised clustering algorithm is used to dynamically cluster the fonts to be trained according to their font styles, and the fonts to be trained with similar styles are classified into the same cluster to obtain M basic style fonts.
6. The font style generation method according to claim 4, characterized in that: The calculating of the similarity weight between the to-be-trained font and the M basic style fonts is specifically as follows: Among them, S i The font F to be trained i The style characteristics of Basic style font The style characteristics of The font F to be trained i And basic style fonts The L1 distance between them, τ represents the temperature parameter, and the normalized exponential function is used to convert Convert to weights and get the font F to be trained i With base font The similarity weight between 7. The font style generation method according to any one of claims 1 to 6, characterized in that: The step of acquiring a source font image and a reference font image, inputting the source font image and the reference font image into the optimized font generation model, and generating a target font image by using the optimized font generation model is specifically as follows: Get the source font image I that provides character information respectively c and a reference font image I providing style information s ; The source font image I is respectively encoded by a content encoder and a style encoder. c and reference font image I s Perform encoding processing to obtain content feature F c and style features F s ; The content feature F c and style features F s After dimension compression, the two are merged and used as the condition y in the diffusion model as input; From the source font image I through the VAE encoder ε c Extract the compressed latent space, perform a forward diffusion process in the latent space, and in the forward diffusion process, inject the condition y into the UNet network to guide the generation, and obtain the latent space feature z0 after t steps; The latent space feature z0 is decoded by the VAE decoder D to generate the target font image I g .
8. A font style generating device, characterized in that: include: Model building module: used to build a font generation model based on the diffusion model; Model fine-tuning module: used to fine-tune the font generation model based on the diffusion model using the LoRA module to obtain an optimized font generation model; Font generation module: used for acquiring a source font image and a reference font image, inputting the source font image and the reference font image into the optimized font generation model, and generating a target font image through the optimized font generation model.
9. A computer device, characterized in that: The computer device includes a processor and a memory coupled to the processor, wherein: The memory stores program instructions for implementing the font style generation method according to any one of claims 1 to 7; The processor is used to execute the program instructions stored in the memory to control the font style generation method.
10. A storage medium, characterized in that: Program instructions executable by a processor are stored, and the program instructions are used to execute the font style generation method according to any one of claims 1 to 7.