Stylized font generation method based on diffusion model
The method addresses font generation challenges by integrating pre-trained encoders and a U-Net diffusion model for precise font generation, improving structural accuracy and artistic expression while reducing computational costs.
Patent Information
- Application Number
- CN202510502113.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-15
AI Technical Summary
Existing font generation methods are difficult to accurately capture the subtle stroke changes and complex structural characteristics of Chinese characters, resulting in the generated images being prone to confusing strokes and imbalanced structures, and the style transfer effect is limited, making it difficult to meet the requirements of high-end design and personalized customization.
The style font generation method based on diffusion model is adopted, combined with pre-trained style encoder, content encoder and stroke structure encoder, the deep integration of multimodal information is achieved through adaptive feature fusion, and iterative denoising generation is used to design a special loss function optimization generation process.
It significantly improves the structural accuracy and style expressiveness of font generation, reduces the calculation cost and training complexity, improves the robustness and generalization ability of the model, and generates high-quality and personalized Chinese fonts.
Smart Images

Figure CN120318367A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and image generation, and particularly relates to a method for generating stylized fonts of diffusion models. Background Art
[0002] Currently, with the continuous innovation of digital media and information dissemination methods, fonts, as an important part of visual communication, have an increasing demand for artistry and personalization. High-quality and personalized fonts can not only enhance the brand image and user experience, but also be widely used in fields such as digital publishing, advertising design, interface customization, and cultural and creative industries. Automatically generating fonts with rich styles and fine structures using advanced generation technologies not only helps to build a large font library, but also provides new creative tools for designers, promoting the deep integration of traditional font design and modern computer vision technologies, thereby improving the visual communication effect while meeting the market's urgent need for diverse and innovative fonts.
[0003] Currently, mainstream font generation methods mostly adopt generative adversarial networks (GANs) or variational autoencoders (VAEs). Although these methods have achieved font style conversion and generation to a certain extent, there are still deficiencies in the balance between the overall structure and local details. Traditional methods often have difficulty accurately capturing the subtle stroke changes and complex structural features in fonts, resulting in problems such as chaotic strokes and unbalanced structures in the generated images. In addition, prior technologies usually rely on fixed feature extractors and cannot effectively distinguish the detail differences of the same Chinese character under different styles. The style transfer effect is limited, and the generated fonts often lack uniqueness in artistic expressiveness and are difficult to meet the requirements of high-end design and personalized customization. These defects limit the wide application of traditional technologies in font library construction and digital creative design. Summary of the Invention
[0004] This method proposes a style font generation scheme based on the diffusion model. The core lies in combining a pre-trained style encoder, a content encoder, and a stroke structure encoder, and achieving deep integration of multi-modal information through adaptive feature fusion. During the generation process, first, the pre-trained style encoder is used to extract style features from the reference font image, and its parameters are frozen during the generation stage to accelerate the calculation speed and maintain the stable transmission of style information. At the same time, a content encoder is constructed to extract the basic structural features of the font image, then a Chinese character stroke structure vector is constructed, and a specially designed stroke structure encoder is used to obtain its features. Subsequently, using the adaptive feature fusion mechanism, the content, stroke structure, and style features are deeply fused to generate conditional features, which are then used as guidance information and input into the diffusion model based on U-Net. Through multi-step denoising, the target font image is gradually restored and generated. This method achieves accurate restoration of font details and overall structure during the generation process, significantly improving the structural accuracy and style expressiveness of the font, providing new technical support and application value for font library construction and font design.
[0005] To achieve the above object, the technical solution of the present invention is as follows: A style font generation method based on a diffusion model, the method comprising the following steps:
[0006] S1: Pre-train a style encoder, construct classification labels required for various font style images and font style classification tasks, build a style classification model network for extracting input font style features and predicting font style labels, and continuously train to obtain a pre-trained style encoder;
[0007] S2: Construct a common Chinese character stroke order structure vector required for the font generation task, encode it into a fixed-dimensional representation in combination with the stroke order and structural features of Chinese characters. At the same time, gradually superimpose Gaussian noise on the input image to generate a training data set from clear to pure noise;
[0008] S3: Take the standard font sample as the content image input, and use the content encoder (extract the content feature E c , this feature E c represents the core structural information of the standard font sample, such as the font outline and layout; at the same time, take the reference style font sample as the style image input, and use the pre-trained style encoder to extract E s , this feature E s captures the artistic style information of the reference style font sample, such as the stroke thickness; in addition, use the stroke order structure encoder to additionally extract the stroke order structure feature E sc of the standard font, this feature E sc describes the font structure and stroke order information, providing additional content feature guidance for subsequent fusion.
[0009] S4: Use the feature fusion module F (a custom fusion module that combines multiple feature vectors through feature weighting and attention mechanism) to fuse the extracted features, generating the conditional feature c = F(E c , E s , E sc ), where the conditional feature c integrates the content feature E c , the style feature E s , and the stroke structure feature E sc ; Subsequently, add this conditional feature c as a conditional input to the U-Net network and generate a stylized font image through an iterative denoising process;
[0010] S5: During the font image generation process, design a dedicated loss function to continuously optimize the generation model, making the generated font image gradually approach the target style.
[0011] Among them, the pre-training process of the style encoder extracts the style features of the reference image, and its purpose is decoupled from the font generation task, thus effectively reducing the computational cost; during this process, first construct a font style image and perform enhancement processing on the image, including cropping, flipping, and adding random noise; at the same time, prepare various font style labels, including running script and cursive script of various calligrapher styles; in the subsequent font generation task, the parameters of the style encoder will be frozen, mainly used to extract style features, and reduce the computational cost in font generation;
[0012] The style encoder consists of a Res2Net backbone network, a Transformer encoder, and multiple multi-layer perceptrons (MLPs). Finally, obtain the classification label with the highest probability through the Softmax function and calculate the loss through the cross-entropy loss function to optimize the network parameters;
[0013] Among them, the expression of the cross-entropy loss function is:
[0014]
[0015] Among them, L represents the cross-entropy loss, C represents the total number of categories of font style classification, y i represents the true label of the sample in the i-th class of font style. When the sample belongs to the i-th class of font style, y i = 1, otherwise y i = 0, represents the predicted probability of the i-th class of font style calculated through the Softmax function, that is, the probability value that the model predicts the sample belongs to the i-th class of font style.
[0016] Among them, the stroke order structure vector in S2 is a 36-dimensional vector, which is used to describe the stroke order and structure information of Chinese characters in the content image during the font generation stage, so as to guide the model to better understand the structure (such as upper-lower structure, left-right structure, etc.) and stroke order (such as dot, horizontal, left-falling stroke, etc.) information of Chinese characters.
[0017] The stroke order structure encoder in S2 consists of multiple convolutional operations, batch normalization layers (Batch Normalization), ReLU activation functions, and pooling layers. Finally, the extracted features are mapped to a 128-dimensional vector through a fully connected layer. Specifically, first, the 36-dimensional structure stroke order vector is reshaped into a two-dimensional tensor with a shape of (1, 6, 6). Subsequently, the input passes through the first convolutional layer (convolution kernel size 3×3, output channels 32), followed by batch normalization and ReLU activation, and then the feature map size is halved to (32, 3, 3) through a 2×2 max pooling layer. Then, it passes through the second convolutional layer (convolution kernel size 3×3, output channels 64), followed by batch normalization and ReLU activation, and the feature map size is further compressed to (64, 1, 1) through a 2×2 max pooling layer. Subsequently, it passes through the third convolutional layer (convolution kernel size 3×3, output channels 128), followed by batch normalization and ReLU activation, and the feature map is compressed into a 128-dimensional vector through a global average pooling layer. Finally, the feature vector is mapped to a 128-dimensional stroke order structure feature vector E through a fully connected layer sc 。
[0018] The content encoder in S2 receives a font content image with a shape of (1, 64, 64) as input and extracts features layer by layer using multiple convolutional modules. First, the input passes through the first convolutional module, including two 3×3 convolutional layers (channel number 32). After each convolutional layer, there is batch normalization and a Leaky ReLU activation function, and the feature size is reduced to (32, 32, 32) through 2×2 max pooling. Subsequently, it enters the second convolutional module, including two 3×3 convolutional layers (channel number 64). After each layer, there is still batch normalization and ReLU activation, and it is further compressed to (64, 16, 16) through 2×2 average pooling. Then, the third convolutional module consists of three 3×3 convolutional layers (channel number 128), and batch normalization and SiLU activation are added after each convolutional layer. Finally, the feature map size is reduced to (128, 1, 1) through global average pooling. Finally, the features are flattened and mapped to a 128-dimensional content feature E through a fully connected layer c 。
[0019] The feature fusion module in S4 includes feature weighting, cross-attention mechanism, and self-attention mechanism. First, perform feature weighted fusion on the content feature and the stroke structure encoding feature; then, fuse the weighted result with the style encoding feature through the cross-attention mechanism; finally, further optimize the fused feature representation through the self-attention mechanism to obtain the final conditional vector c. The specific fusion process can be described by the following expressions. The content feature E c and the stroke structure feature E sc The fusion part is:
[0020] F cw = α·E c + β·E sc ,
[0021] where E c represents the content encoding feature, E sc represents the stroke structure feature, α, β are learnable weighting parameters used to adjust the contribution ratio of the two features. Here, α and β are set to 0.6 and 0.4 respectively. F cw represents the weighted result of the content feature and the stroke feature,
[0022] The weighted result F cw and the style feature E s The part of cross-attention fusion is:
[0023]
[0024] where Q is the query vector (Query) generated by the feature weighted result F cw , K = V is the value vector (Value) generated by the style feature,
[0025] Finally, perform self-attention fusion on the cross-attention fusion feature:
[0026]
[0027] where, Q = K = V are the query, key, and value vectors generated by the cross-attention fusion result F ca .
[0028] Through the adaptive fusion module, the model can fully integrate the content feature, the stroke structure feature, and the style feature to obtain the conditional feature c and capture the dependency relationship between the features.
[0029] In step S2, by performing multiple noise addition operations on the style font image, a pure noise image is gradually obtained to meet the diffusion training requirements and form sample data. The noise addition process is modeled using a forward Markov chain, and its specific expression is:
[0030]
[0031] wherein represents a scaled version of the font image at the previous time step, used to retain the main features of the image; β t represents the noise addition coefficient at time step t, controlling the intensity of the noise; I is the identity matrix, indicating that the noise is an isotropic Gaussian distribution;
[0032] The corresponding step-by-step noise addition formula is:
[0033]
[0034] where x t represents the image obtained after the t-th noise addition operation, t ∈ {1, 2,..., T}; x t-1 represents the output of the previous noise addition operation or the initial font image; β t is the noise addition coefficient at the t-th step, used to control the noise intensity; ε represents the random noise sampled from a Gaussian distribution with zero mean and unit variance
[0035] In step S4, for the noise image x t at time step t, a step-by-step denoising operation is performed through the U-Net model, aiming to gradually restore the final target font image x0. This denoising process is modeled as a reverse Markov chain, assuming that the conditional probability distribution of generating x t-1 depends only on the noise image x t at the current time step, time step t, and the conditional feature c. This conditional feature c combines the content feature E c , the stroke order structure feature E sc , and the style feature E s , providing rich context information in the font generation task, thereby guiding the U-Net to gradually approach the target font style during the reverse generation process; specifically, the conditional probability distribution in the denoising conditional process can be expressed as:
[0036]
[0037] where μ θ (x t , t, c) represents the denoising mean predicted by the model, is the fixed noise variance for each step, controlling the uncertainty of generation; c represents the conditional feature, combined by the content feature E c , the stroke order structure feature E sc , and the style feature E s ;
[0038] The mean calculation formula for the denoising process is:
[0039]
[0040] Among them, x t represents the noisy image at the current time step, t represents the time step, and c represents the conditional feature, which is obtained by fusing the content feature E c , the stroke structure feature E sc and the style feature E s . ∈ θ (x t , t, c) represents the predicted noise component by the U-Net model. α t = 1 - β t represents the noise addition coefficient at time step t, where β t is a predefined noise scheduling parameter. represents the cumulative noise addition coefficient from time step 1 to t. μ θ (x t , t, c) represents the denoising mean predicted by the model and is used to generate the image x t-1 at the next time step. In step S5, in order to make the generated font image close to the target font in terms of style consistency, structural details, and noise restoration, this method constructs two types of loss functions to optimize the generation process. On the one hand, a loss function based on contrastive learning is designed. The image semantic features are extracted from the image through a pre-trained VGG network and mapped into feature vectors, denoted as
[0041] v = f VGG (x),
[0042] where x is the font image, and the obtained v is the feature extracted from the font image in the VGG network. During the contrastive learning process, the feature distance between the generated font and the target font (i.e., the positive sample) is made as close as possible, while the feature distance between the font images of the same Chinese character but different styles (i.e., the negative samples) is made as far as possible. The specific contrast loss formula is:
[0043]
[0044] where N is the number of batch samples, v i represents the feature vector extracted from the i-th generated image through the VGG network, represents the feature vector of the target font image corresponding to the i-th generated image, represents the feature vector of the font image from different styles but belonging to the same Chinese character, that is, the negative sample. sim(a, b) is defined as the cosine similarity between vector a and vector b, τ is the temperature parameter used to smooth the similarity distribution, and M represents the number of negative samples selected in the contrastive learning.
[0045] In addition, to ensure the noise prediction accuracy of the diffusion model during the denoising process, this method also uses the mean squared error (MSE) loss to measure the difference between the noise predicted by the model and the actually added noise. The formula is as follows:
[0046]
[0047] where N is the number of batch samples, ∈ i represents the Gaussian noise actually added to the i-th sample at time step t, satisfying ∈ i ~N(0, I), represents the noise component predicted by the U-Net model at time step t, is the noise image of the i-th sample at time step t, c is the conditional feature obtained through adaptive feature fusion, which fuses the content feature E c , the stroke order structure feature E sc and the style feature E s ;
[0048] The total loss function is composed of a weighted combination of the contrast loss and the MSE loss:
[0049]
[0050] where, in the total loss function, the contribution ratio of the MSE loss is 1.0, and the contribution ratio of the contrast loss is 0.5. Therefore, during the optimization process, more emphasis is placed on noise correction, while also ensuring the effectiveness of the style consistency constraint.
[0051] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the described method for generating a style font based on a diffusion model.
[0052] A computer-readable storage medium stores computer instructions thereon, characterized in that: when the computer instructions are executed by a processor, they implement the described method for generating a style font based on a diffusion model.
[0053] Compared with the prior art, the advantages of the present invention are as follows:
[0054] 1. Significantly reduce the computational cost and training complexity:
[0055] By completely decoupling the pre-training process of the style encoder from the font generation task, the present invention avoids the redundant calculations in the end-to-end training of traditional diffusion models. For example, in the prior art, such as unconditional diffusion models (DDPM) or GAN-based font generation methods, it is often necessary to repeatedly train the entire model in each generation task, resulting in waste of computing resources and extension of training time. Through the pre-training strategy of the present invention, the style encoder is first independently trained to capture general font style features, and then its parameters are frozen in the generation stage and only used for feature extraction, which theoretically reduces the amount of model parameter updates.
[0056] 2. Improve the accuracy of Chinese character generation and the understanding of fine-grained structures:
[0057] Compared with existing font generation technologies (such as simple CNN-based or unconditional diffusion models), the present invention introduces a 36-dimensional stroke order structure vector to simultaneously encode the stroke order and structural information of Chinese characters, which significantly improves the model's ability to capture the unique attributes of Chinese characters. From a theoretical perspective, traditional methods often ignore the stroke order dynamics, resulting in distortion in the stroke connection and structural transition of the generated fonts (such as blurred boundaries or incorrect orders); the present invention not only adds an additional stroke order vector, but also inserts a structural separator (such as "0" indicating a structural transition) in the stroke vector, achieving fine-grained encoding. Theoretical derivation shows that this encoding method can reduce the Chinese character structure reconstruction error by 15% (simulation calculation based on the mean square error MSE metric), thereby guiding the U-Net model to more accurately fuse conditional features (E c 、E sc and E s ) during the denoising stage, and finally generating high-quality font images.
[0058] 3. Improve the robustness and generalization ability of the model:
[0059] When dealing with diverse font datasets, the prior art often suffers from poor generalization due to insufficient data augmentation or incomplete feature extraction (such as a decline in performance on new font styles); in the pre-training stage of the present invention, through data augmentation (such as cropping, flipping, and adding noise) combined with rich font labels (such as regular script, Song typeface, and celebrity fonts), and integrating global structure information into the stroke order structure encoding, this theoretically enhances the model's anti-noise ability and adaptability, reducing the risk of overfitting. Brief Description of the Drawings
[0060] Figure 1 is the overall flowchart of the present invention, schematically showing the steps of the style font generation method based on the diffusion model.
[0061] Figure 2 is the architecture diagram of the pre-trained style encoder of the present invention, detailing the structure and pre-training process of the style encoder.
[0062] Figure 3 This is the overall generation architecture diagram of the present invention, showing the reverse diffusion mechanism and conditional feature application in the font generation process. Detailed implementation manners
[0063] The present invention discloses a method. To deepen the understanding of the present invention, the following will introduce the solution in detail in combination with the accompanying drawings and implementation manners.
[0064] Example 1: Refer to Figures 1 - 3 , a method for generating stylized fonts based on a diffusion model, the method comprising the following steps:
[0065] S1: Pre-train a style encoder, construct classification labels required for various font style images and font style classification tasks, build a style classification model network for extracting input font style features and predicting font style labels, and continuously train to obtain a pre-trained style encoder;
[0066] S2: Construct a common Chinese character stroke order structure vector required for the font generation task, encode it into a fixed-dimensional representation in combination with the stroke order and structural characteristics of Chinese characters. At the same time, gradually superimpose Gaussian noise on the input image to generate a training data set from clear to pure noise;
[0067] S3: Input a standard font sample as the content image, use a content encoder (extract content feature E c , this feature E c represents the core structure information of the standard font sample, such as font outline and layout; at the same time, input a reference stylized font sample as the style image, and use the pre-trained style encoder to extract E s , this feature E s captures the artistic style information of the reference stylized font sample, such as stroke thickness; in addition, use a stroke order structure encoder to additionally extract the stroke order structure feature E sc of the standard font, this feature E sc describes the font structure and stroke order information, providing additional content feature guidance for subsequent fusion,
[0068] S4: Use a feature fusion module F (a custom fusion module that combines multiple feature vectors through feature weighting and attention mechanism) to fuse the extracted features to generate a conditional feature c = F(E c , E s , E sc ), where the conditional feature c integrates the content feature E c , the style feature E s and the stroke order structure feature E sc ; subsequently, use this conditional feature c as a conditional input and add it to the U-Net network to generate a stylized font image through an iterative denoising process;
[0069] S5: During the font image generation process, a dedicated loss function is designed to continuously optimize the generation model, making the generated font images gradually approach the target style.
[0070] The style encoder adopts a pre-training strategy to extract the style features of the reference image. This pre-training process is completely decoupled from the font generation task, which means that the style encoder is first independently trained to capture general font style information, thus avoiding repeated training of the model in subsequent font generation tasks, significantly reducing the computational cost and improving the overall system efficiency. Specifically, in the pre-training stage, first, a diverse dataset of font style images is constructed, including various reference font samples. Subsequently, these images are enhanced through data augmentation techniques such as cropping, flipping, and adding random noise to enhance the robustness and generalization ability of the model, ensuring that the style encoder can accurately extract style features under different variants. At the same time, rich font style labels are prepared, such as different fonts, including regular script, Song typeface, and the fonts of different calligraphers and famous people. These labels help train an encoder that can finely capture the details of artistic styles. After pre-training, in the font generation task, the parameters of the style encoder will be frozen and used only as a feature extraction module. This not only simplifies the generation process but also significantly reduces the consumption of computational resources. For example, compared with the end-to-end training method, the present invention can shorten the training time by more than 30% and improve the quality and consistency of the generated images. Through this pre-training decoupling mechanism, the present invention realizes efficient reuse of style features, solves the computationally intensive problem in font generation of traditional diffusion models, and highlights the innovative value of the pre-trained encoder.
[0071] Specifically, for the described style encoder, the encoder is composed of a Res2Net backbone network, a Transformer encoder, and multiple multi-layer perceptrons (MLPs). Finally, the classification label with the highest probability is obtained through the Softmax function, and the loss is calculated through the cross-entropy loss function to optimize the network parameters. The expression of its cross-entropy loss function is:
[0072]
[0073] where L represents the cross-entropy loss, C represents the total number of categories for font style classification, y i represents the true label of the sample in the i-th class of font style. When the sample belongs to the i-th class of font style, y i = 1, otherwise y i = 0, represents the predicted probability of the i-th class of font style calculated through the Softmax function, that is, the probability value that the model predicts the sample belongs to the i-th class of font style.
[0074] The stroke order structure vector is a 36-dimensional vector. The specific stroke order structure connection vector encoding method is as follows: First, each stroke type is encoded into a corresponding number according to the stroke sequence number commonly used in Chinese characters, for example, dot represents 1, horizontal represents 2, vertical represents 3, left-falling represents 4, right-falling represents 5, right-falling represents 6, and so on, to ensure that each stroke has a unique sequence number. Taking the character "明" as an example, its stroke order information is [丨,乛,一,一,丿,,一,一], then the initial stroke order vector is [3,23,2,2,4,16,2,2,0,0..];
[0075] In order to enable the model to not only learn the stroke information in the content image, but also better understand Chinese characters by introducing structural information; common Chinese character structures such as single-character structure, upper-lower structure, left-right structure and surrounding structure are encoded by inserting the number "0" in the stroke order vector as a structure separator; for example, the word "Ming" belongs to the left-right structure, so "0" is inserted in the stroke order vector to indicate the transition from the left structure to the right structure. The final stroke order structure connection vector is [3,23,2,2,0,4,16,2,2], where "0" indicates the conversion of the structure. The constructed stroke order structure connection vector is input into the stroke order structure encoder to extract features, and the feature E is obtained. sc This encoding method not only clarifies the order of strokes, but also effectively conveys the stroke connection information of Chinese characters in the structural span, so that the style font generation method can fully understand and reproduce the unique style of Chinese characters at the fine-grained and global structural levels, and achieve high-quality and personalized font generation effects.
[0076] The stroke order structure encoder is composed of multiple convolution operations, batch normalization layers (Batch Normalization), ReLU activation function and pooling layer, and finally the extracted features are mapped to a 128-dimensional vector through a fully connected layer; specifically, the 36-dimensional structural stroke order vector is first reshaped into a two-dimensional tensor with a shape of (1,6,6); then, the input passes through the first convolutional layer (convolution kernel size 3×3, output channel number 32), followed by batch normalization and ReLU activation, and then the feature map size is halved to (32,3,3) through a 2×2 maximum pooling layer; then, it passes through the second convolutional layer (convolution kernel size 3×3, output channel number 64), followed by batch normalization and ReLU activation, and the feature map size is further compressed to (64,1,1) through a 2×2 maximum pooling layer; then, it passes through the third convolutional layer (convolution kernel size 3×3, output channel number 128), followed by batch normalization and ReLU activation, and the feature map is compressed to a 128-dimensional vector through a global average pooling layer; finally, the feature vector is mapped to a 128-dimensional feature vector E through a fully connected layer. sc .
[0077] Next, the content encoder receives a font content image with a shape of (1, 64, 64) as input and uses multiple convolutional modules to extract features layer by layer. First, the input passes through the first convolutional module, which includes two 3×3 convolutional layers (with 32 channels). After each convolutional layer, batch normalization and the Leaky ReLU activation function are immediately followed, and the feature size is reduced to (32, 32, 32) through 2×2 max pooling; then it enters the second convolutional module, which includes two 3×3 convolutional layers (with 64 channels). Batch normalization and the ReLU activation are still attached after each layer, and it is further compressed to (64, 16, 16) through 2×2 average pooling; then, the third convolutional module consists of three 3×3 convolutional layers (with 128 channels), and batch normalization and the SiLU activation are added after each convolutional layer. Finally, the feature map size is reduced to (128, 1, 1) through global average pooling; finally, the features are flattened and mapped to a 128-dimensional feature vector E c ;
[0078] Subsequently, the feature fusion module includes feature weighting, cross-attention mechanism, and self-attention mechanism. First, feature weighting fusion is performed on the content feature and the stroke order structure encoding feature; then, the weighted result and the style encoding feature are fused through the cross-attention mechanism; finally, the fused feature representation is further optimized through the self-attention mechanism to obtain the final conditional feature vector c. The specific fusion process can be described by the following expression. The content feature E c and the stroke order structure feature E sc The fusion part is:
[0079] F cw = α·E c + β·E sc ,
[0080] where E c represents the content encoding feature, E sc represents the stroke order structure feature, and α, β are learnable weighting parameters used to adjust the contribution ratio of the two features. Here, α and β are set to 0.6 and 0.4 respectively, and F cw represents the weighted result of the content feature and the stroke order feature.
[0081] The weighted result F cw and the style feature E s The cross-attention fusion part is:
[0082]
[0083] where Q is obtained from the feature weighting result F cwThe generated query vector (Query), where K = V is the value vector generated from the style features. The cross-attention fusion features are finally fused with a self-attention:
[0084]
[0085] Among them, Q = K = V are the query, key, and value vectors generated from the cross-attention fusion result.
[0086] Through the adaptive fusion module, the model can fully integrate the content features, stroke structure features, and style features, and capture the dependencies between features.
[0087] Furthermore, the noise addition process is modeled using a forward Markov chain, and its specific expression is:
[0088]
[0089] Where represents the scaled version of the font image at the previous time step, used to retain the main features of the image; β t represents the noise addition coefficient at time step t, controlling the intensity of the noise; I is the identity matrix, indicating that the noise is an isotropic Gaussian distribution; correspondingly, the corresponding step-by-step noise addition formula is:
[0090]
[0091] Where x t represents the image obtained after the t-th noise addition operation, t ∈ {1, 2,..., T}; x t-1 represents the output of the previous noise addition operation or the initial font image; β t is the noise addition coefficient at the t-th step, used to control the noise intensity; ε represents the random noise sampled from a Gaussian distribution with zero mean and unit variance Sampled random noise.
[0092] The noise addition process is modeled as a reverse Markov chain. Assuming that the conditional probability distribution of generating x t-1 at each step depends only on the noise image x t at the current time step, the time step information t, and the conditional feature c.
[0093] This conditional feature c combines the content feature E c , the stroke structure feature E sc and the style feature E s , providing rich context information in the font generation task, thus guiding the U-Net to gradually approximate the target font style during the reverse generation process; specifically, the conditional probability distribution in the denoising conditional process can be expressed as:
[0094]
[0095] Among them, μ θ (x t , t, c) represents the denoising mean predicted by the model, is the fixed noise variance for each step, controlling the uncertainty of generation; c represents the conditional feature, which is composed of the content feature E c , the stroke structure feature E sc and the style feature E s fused together;
[0096] The mean calculation formula for the denoising process is:
[0097]
[0098] Among them, x t represents the noisy image at the current time step, t represents the noisy image at the current time step, c represents the conditional feature, which is composed of the content feature E c , the stroke structure feature E sc and the style feature E s fused to obtain; ∈ θ (x t , t, c) represents the noise component predicted by the U-Net model, α t = 1 - β t represents the noise addition coefficient at time step t, represents the cumulative noise addition coefficient from 0 to t, μ θ (x t , t, c) represents the denoising mean at the current step t predicted by the model, used to generate the image x t-1 .
[0099] In step S5, in order to make the generated font image close to the target font in terms of style consistency, structural details, and noise restoration, this method constructs a reverse Markov chain, assuming that the conditional probability distribution of generating x t-1 only depends on the noisy image x t at the current time step, the time step information t, and the conditional feature c. This conditional feature c combines the content feature E c , the stroke structure feature E sc and the style feature E s , providing rich context information in the font generation task, thereby guiding the U-Net to gradually approach the target font style during the reverse generation process; specifically, the conditional probability distribution in the denoising conditional process can be expressed as:
[0100]
[0101] Among them, μ θ (xt , t, c) represents the denoising mean predicted by the model, is the fixed noise variance for each step, controlling the uncertainty of generation; c represents the conditional feature, which is composed of the content feature E c , the stroke structure feature E sc and the style feature E s fused together;
[0102] The mean calculation formula for the denoising process is:
[0103]
[0104] where, x t represents the noisy image at the current time step, t represents the noisy image at the current time step, c represents the conditional feature, which is composed of the content feature E c , the stroke structure feature E sc and the style feature E s fused to obtain; ∈ θ (x t , t, c) represents the noise component predicted by the U-Net model, α t = 1 - β t represents the noise addition coefficient at time step t, represents the cumulative noise addition coefficient from 0 to t, μ θ (x t , t, c) represents the denoising mean of the current step t predicted by the model, used to generate the image x t-1 .
[0105] This method constructs two types of loss functions to optimize the generation process; on the one hand, a loss function based on contrastive learning is designed. The image semantic features are extracted from the image through a pre-trained VGG network and mapped into feature vectors, denoted as
[0106] v = f VGG (x),
[0107] where, x is the font image, and the obtained v is the feature extracted from the font image in the VGG network; during the contrastive learning process, the feature distance between the generated font and the target font (i.e., the positive sample) is made as close as possible, while the feature distance between the font images of the same Chinese character but different styles (i.e., the negative samples) is made as far as possible; the specific contrast loss formula is:
[0108]
[0109] where, N is the number of batch samples, v i represents the feature vector extracted from the i-th generated image through the VGG network, denotes the feature vector of the target font image corresponding to the i-th generated image, denotes the feature vectors of font images from different styles but belonging to the same Chinese character, i.e., negative samples. sim(a, b) is defined as the cosine similarity between vector a and vector b, τ is the temperature parameter used to smooth the similarity distribution, and M represents the number of negative samples selected in the contrastive learning;
[0110] In addition, to ensure the noise prediction accuracy of the diffusion model during the denoising process, this method also uses the mean squared error (MSE) loss to measure the difference between the noise predicted by the model and the actually added noise. The formula is:
[0111]
[0112] where N is the number of batch samples, ∈ i denotes the Gaussian noise actually added to the i-th sample at time step t, satisfying ∈ i ~N(0, I), denotes the noise component predicted by the U-Net model at time step t, is the noise image of the i-th sample at time step t, c is the conditional feature obtained through adaptive feature fusion, which fuses the content feature E c and the stroke order structure feature E sc as well as the style feature E s ;
[0113] The total loss function is composed of a weighted combination of the contrastive loss and the MSE loss:
[0114]
[0115] Among them, in the calculation of the total loss function of this method, the contribution ratio of the MSE loss is 1.0, and the contribution ratio of the contrastive loss is 0.5. Therefore, during the optimization process, more emphasis is placed on noise correction, while also ensuring the effectiveness of the style consistency constraint.
[0116] It should be noted that the above embodiments are not used to limit the protection scope of the present invention. Any equivalent transformation or substitution made on the basis of the above technical solutions falls within the protection scope of the claims of the present invention.
Claims
1. A method for generating stylized fonts based on diffusion models, characterized in that, The method includes the following steps: S1: Pre-train a style encoder, construct various font style images and classification labels required for the font style classification task, build a style classification model network for extracting input font style features and predicting font style labels, and continuously train to obtain a pre-trained style encoder; S2: Construct the vector of the stroke order structure of common Chinese characters required for the font generation task, encode it into a fixed-dimensional representation by combining the stroke order and structural features of Chinese characters. At the same time, gradually superimpose Gaussian noise on the input image to generate a training data set from clear to pure noise; S3: Input the standard font sample as the content image and use the content encoder to extract the content feature E c , where the feature E c represents the core structure information of the standard font sample. At the same time, input the reference style font sample as the style image and use the pre-trained style encoder to extract E s , where the feature E s captures the artistic style information of the reference style font sample. Use the stroke order structure encoder to additionally extract the stroke order structure feature E sc , where the feature E sc describes the font structure and stroke order information, providing additional content feature guidance for subsequent fusion. S4: Use the feature fusion module F to fuse the extracted features to generate the conditional feature c = F(E c , E s , E sc ), where the conditional feature c integrates the content feature E c , the style feature E s and the stroke structure feature E sc ; Subsequently, take the conditional feature c as the conditional input and add it to the U-Net network to generate a stylized font image through an iterative denoising process; S5: During the font image generation process, design a special loss function to continuously optimize the generation model, so that the generated font image gradually approaches the target style.
2. The method for generating a stylized font based on a diffusion model according to claim 1, wherein, The style encoder adopts a pre-training strategy to extract the style features of the reference image. This pre-training process is completely decoupled from the font generation task. Specifically, in the pre-training stage, first construct a diverse font style image data set, including various reference font samples; Subsequently, perform enhancement processing on these images, such as data enhancement techniques such as cropping, flipping, and adding random noise, to enhance the robustness and generalization ability of the model, and ensure that the style encoder can accurately extract style features under different variants; at the same time, prepare rich font style labels, including regular script, Song typeface, and the fonts of different calligraphers and celebrities. After the pre-training is completed, in the font generation task, the parameters of the style encoder will be frozen and only used as a feature extraction module; The style encoder consists of a Res2Net backbone network, a Transformer encoder, and multiple multi-layer perceptrons (MLPs). Finally, the classification label with the highest probability is obtained through the Softmax function, and the loss is calculated through the cross-entropy loss function to optimize the network parameters; The expression of the cross-entropy loss function is: Among them, L represents the cross-entropy loss, C represents the total number of categories for font style classification, and y i represents the true label of the sample in the i-th font style. When the sample belongs to the i-th font style, y i = 1; otherwise, y i = 0, represents the predicted probability of the i-th font style calculated by the Softmax function, that is, the probability value that the model predicts the sample belongs to the i-th font style.
3. A method for generating a stylized font based on a diffusion model according to claim 1, characterized in that, The stroke order structure vector in S2 is a 36-dimensional vector, which is used to describe the stroke order and structural information of Chinese characters in the content image during the font generation stage, so as to guide the model to better understand the structure and stroke order information of Chinese characters. The specific encoding method is: first, encode each stroke as a corresponding number according to the commonly used stroke types of Chinese characters, and use the method of inserting the number "0" in the stroke order vector as a structure separator for encoding; Subsequently, the constructed stroke order structure connection vector is input into the stroke order structure encoder to extract the feature E sc .
4. A method for generating a stylized font based on a diffusion model according to claim 2, characterized in that, In S2, the stroke order structure encoder consists of multiple convolutional operations, a batch normalization layer (Batch Normalization), a ReLU activation function, and a pooling layer. Finally, the extracted features are mapped to a 128-dimensional vector through a fully connected layer. Specifically, first, the 36-dimensional structural stroke order vector is reshaped into a two-dimensional tensor with a shape of (1, 6, 6). Subsequently, the input passes through the first convolutional layer (with a convolutional kernel size of 3×3 and 32 output channels), followed by batch normalization and ReLU activation, and then the feature map size is halved to (32, 3, 3) through a 2×2 max pooling layer. Then, it passes through the second convolutional layer (with a convolutional kernel size of 3×3 and 64 output channels), followed by batch normalization and ReLU activation, and the feature map size is further compressed to (64, 1, 1) through a 2×2 max pooling layer. Subsequently, it passes through the third convolutional layer (with a convolutional kernel size of 3×3 and 128 output channels), followed by batch normalization and ReLU activation, and the feature map is compressed into a 128-dimensional vector through the global average pooling layer; finally, the feature vector is mapped to the stroke order structure feature E of 128 dimensions through the fully connected layer sc 。 5. A method for generating a stylized font based on a diffusion model according to claim 2, characterized in that, In S2, the content encoder receives a font content image with a shape of (1, 64, 64) as input and extracts features layer by layer using multiple convolutional modules. First, the input passes through the first convolutional module, which includes two 3×3 convolutional layers (with 32 channels). After each convolutional layer, batch normalization and the Leaky ReLU activation function are immediately applied, and the feature size is reduced to (32, 32, 32) through 2×2 max pooling. Subsequently, it enters the second convolutional module, which consists of two 3×3 convolutional layers (with 64 channels). Batch normalization and ReLU activation are still attached after each layer, and further compression is performed to (64, 16, 16) through 2×2 average pooling. Then, the third convolutional module is composed of three 3×3 convolutional layers (with 128 channels), and batch normalization and SiLU activation are added after each convolutional layer. Finally, the feature map size is reduced to (128, 1, 1) through global average pooling. Ultimately, the features are flattened and mapped to 128-dimensional content features E c .
6. The method for generating a stylized font based on a diffusion model according to claim 1, wherein The feature fusion module in S4 includes feature weighting, cross-attention mechanism, and self-attention mechanism. First, the content feature and the stroke structure encoding feature are fused with feature weighting; then, the weighted result and the style encoding feature are fused through the cross-attention mechanism; finally, the fused feature representation is further optimized through the self-attention mechanism to obtain the final conditional feature vector. The specific fusion process can be described by the following expression, where the content feature E c and the stroke structure feature E sc The fusion part is: F cw = α·E c + β·E sc , Among them, E c represents the content feature, and E sc represents the stroke order structure feature. α and β are learnable weighting parameters used to adjust the contribution ratio of the two features. Here, α and β are set to 0.6 and 0.4 respectively. F cw represents the weighted result of the content feature and the stroke order feature. Weighted result F cw and style feature E s The part for cross-attention fusion is as follows: where Q is the query vector generated by the feature weighted result F cw and K = V is the value vector generated by the style feature E s Finally, self-attention fusion is performed on the cross-attention fusion features: Among them, Q = K = V are the query, key, and value vectors generated from the cross-attention fusion result F ca Through the adaptive fusion module, the model can fully integrate the content features, stroke order structure features, and style features to obtain the conditional feature c and capture the dependencies between the features.
7. A method for generating a stylized font based on a diffusion model according to claim 1, characterized in that: In step S2, by performing multiple noise addition operations on the style font image, a pure noise image is gradually obtained to meet the requirements of diffusion training and form sample data. The noise addition process is modeled using a forward Markov chain, and its specific expression is: Among them represents a scaled version of the font image at the previous time step, used to retain the main features of the image; β t represents the noise addition coefficient at time step t, controlling the intensity of the noise; I is the identity matrix, indicating that the noise is an isotropic Gaussian distribution; The corresponding step-by-step noise addition formula is: where x t represents the image obtained after the t-th noise addition operation, where t ∈ {1, 2, …, T}; x t-1 represents the output of the previous noise addition operation or the initial font image; β t is the noise addition coefficient at the t-th step, used to control the noise intensity; ε represents the random noise sampled from a Gaussian distribution with zero mean and unit variance and In step S4, for the noisy image x at time step t t , a step-by-step denoising operation is performed through the U-Net model, aiming to gradually restore the final target font image x0. This denoising process is modeled as a reverse Markov chain, assuming that the conditional probability distribution of generating x t-1 depends only on the noisy image x at the current time step t , time step t, and the conditional feature c, which combines the content feature E c , stroke structure feature E sc and style feature E s , providing rich context information in the font generation task, thereby guiding the U-Net to gradually approach the target font style during the reverse generation process; specifically, the conditional probability distribution in the denoising conditional process can be expressed as: Among them, μ θ (x t , t, c) represents the denoising mean predicted by the model, is the fixed noise variance at each step, controlling the uncertainty of generation; c represents the conditional feature, which is fused by the content feature E c , the stroke order structure feature E sc and the style feature E s ; The mean calculation formula for the denoising process is: Among them, x t represents the noisy image at the current time step, t represents the current time step, and c represents the conditional feature, which is obtained by fusing the content feature E c , the stroke structure feature E sc and the style feature E s ; ∈ θ (x t , t, c) represents the noise component predicted by the U-Net model, and α t = 1 - β t represents the noise addition coefficient at time step t, represents the cumulative noise addition coefficient from 0 to t, and μ θ (x t , t, c) represents the denoising mean at the current step t predicted by the model, which is used to generate the image x t-1 .
8. A method for generating a stylized font based on a diffusion model according to claim 1, characterized in that: In step S5, based on the loss function of contrastive learning, the image semantic features are extracted from the image through the pre-trained VGG network and mapped into feature vectors, denoted as v = f VGG (x), Among them, x is the font image, and the obtained v is the feature extracted from the font image in the VGG network. During the contrastive learning process, the feature distance between the generated font and the target font (i.e., the positive sample) is made as close as possible, while the feature distance between font images of the same Chinese character but different styles (i.e., the negative samples) is made as far as possible. The specific contrastive loss formula is: where N is the number of batch samples, v i represents the feature vector extracted from the i-th generated image through the VGG network, represents the feature vector of the target font image corresponding to the i-th generated image, represents the feature vectors of font images from different styles but belonging to the same Chinese character, that is, negative samples. sim(a, b) is defined as the cosine similarity between vector a and vector b, τ is the temperature parameter used to smooth the similarity distribution, and M represents the number of negative samples selected in the contrastive learning; The mean squared error (MSE) loss is used to measure the difference between the noise predicted by the model and the actually added noise, and its formula is: where N is the number of batch samples, ∈ i denotes the Gaussian noise actually added to the i-th sample at time step t, satisfying ∈ i ~N(0, I), denotes the noise component predicted by the U-Net model at time step t, is the noise image of the i-th sample at time step t, c is the conditional feature obtained by adaptive feature fusion, which fuses the content feature E c , the stroke order structure feature E sc and the style feature E s ; The total loss function is composed of a weighted combination of the contrastive loss and the MSE loss: Among them, in the total loss function, the contribution ratio of the MSE loss is 1.0, and the contribution ratio of the contrastive loss is 0.5, so as to focus more on noise correction during the optimization process and ensure the effectiveness of the style consistency constraint at the same time.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: When the processor executes the program, it implements a style font generation method based on a diffusion model as described in any one of claims 1 to 8 above.
10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instruction is executed by the processor, it implements a style font generation method based on a diffusion model as described in any one of claims 1-8.
Citation Information
Cited By
Condition-controllable image sample expansion method and system for surface defect detection
CN120525000A
Vectorization Chinese character graph generation method based on large model
CN121010668A
A large model-based vectorized Chinese character pattern generation method
CN121010668B
Hair body generation method and device based on diffusion model
CN121304849A
Multi-source style and content interaction generation method for few-sample font generation
CN121564127A