Controllable Chinese character new font generation method and system based on diffusion model
Through the diffusion model-based method, using the skeleton layout and stroke style feature guidance, a new font that meets expectations is generated, which solves the problems of uncontrollable generation results and unsmooth strokes in the existing technology, and achieves diversified Chinese font generation.
Patent Information
- Application Number
- CN202510610417.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-26
AI Technical Summary
The prior art has problems in the generation of Chinese characters in Chinese fonts, which are difficult to customize according to the specific needs of users or designers, and the strokes are not smooth.
Using a diffusion model-based method, the skeleton layout information and stroke style features of the character image are used as conditional guidance, and the U-Net network of the diffusion model is used to perform the feature-guided diffusion process, generate a new font character image that meets expectations, and optimize the denoising process through the mean square error loss function.
It realizes the generation of diverse new fonts, maintains high stroke fluency, and can generate expected new fonts based on the specified target skeleton structure and stroke style.
Smart Images

Figure CN120543697A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer image processing, and in particular relates to a method and system for generating a controllable new Chinese character font based on a diffusion model. Background Art
[0002] Font design is widely used in print, digital media, and brand identity. In print, font design enhances text readability and imparts a unique style. In digital media, it infuses content with aesthetic appeal and ensures optimal rendering across various screens. In brand identity, font design conveys a company's personality and establishes visual identity. Therefore, font design is not only a vehicle for textual expression but also a crucial component of visual communication and brand building. However, compared to English and Arabic letters, Chinese characters have complex glyph structures and are numerous. Designing a complete font requires individually designing each character to ensure a consistent style, a complex and time-consuming process. With the advent of the artificial intelligence era, cutting-edge technologies such as deep learning are continuously developing, leading to breakthroughs in font generation technology. Current research on new font generation based on style transfer aims to streamline designers' workflows, but requires designers to first design sample fonts, which can lead to a lack of inspiration. While font generation methods based on style fusion can automatically generate new font styles without the need for reference samples, their development has been relatively slow and presents several challenges and shortcomings. Fonts generated through interpolation and blending often suffer from choppy strokes, and the generated results are uncontrollable, making them difficult to customize to the specific needs of users or designers. Summary of the Invention
[0003] The purpose of the present invention is to provide a controllable Chinese character new font generation method and system based on a diffusion model. By taking the skeleton layout information and stroke style characteristics of the character image as the conditional guidance of the diffusion model, learning its unraveling representation, and realizing the reorganization of the skeleton layout and stroke style of different fonts, a variety of new fonts are generated.
[0004] To achieve the above object, the technical solution of the present invention is: a method for generating a controllable new Chinese character font based on a diffusion model, comprising:
[0005] S1. Input the target stroke style character image and the target skeleton character image into the stroke style extraction network and the skeleton information extraction network respectively through two branches;
[0006] S2, extracting multi-scale skeleton features through the skeleton information extraction network;
[0007] S3, extracting stroke style features through a stroke style extraction network;
[0008] S4, through the feature-guided diffusion model, the denoising operation prediction of each step of the U-Net network in the diffusion process is used to generate a new font character image that meets the expectations;
[0009] S5. Optimize each step of the denoising process by using the mean square error loss function.
[0010] Furthermore, in S2, the skeleton information extraction network includes a skeleton extraction algorithm for extracting a skeleton image of a target skeleton character image, and a skeleton encoder for extracting multi-scale skeleton features of the skeleton image.
[0011] Furthermore, the skeleton extraction algorithm includes binarization processing, skeleton extraction, grayscale conversion, and dimension expansion.
[0012] Furthermore, through the skeleton encoder E k Extracting multi-scale skeleton features F from skeleton images k ={F k1 ,F k2}, where F k1 is the skeleton feature of the first layer of the input U-Net network, F k2 To input the skeleton features of the second layer of the UNet network, specifically:
[0013] F k =E k (I k )
[0014] Where, I k Represents the skeleton image, skeleton encoder E k It is a neural network composed of convolutional layers.
[0015] Furthermore, in S3, the stroke style extraction network includes a stroke style encoder for extracting style features of the target stroke style character image and a multi-layer perceptron MLP for projecting the style features into the intermediate feature space of the diffusion model.
[0016] Furthermore, through the stroke style encoder E s Extract the style feature F of the target stroke style character image s , specifically:
[0017] F s =E s (I s )
[0018] Where, I s Represents the target stroke style character image, the stroke style encoder E s It is a neural network composed of the first eight layers of VGG11.
[0019] Furthermore, in S4, the denoising operation prediction of each step of the U-Net network during the diffusion process of the feature-guided diffusion model is implemented as follows:
[0020]
[0021] Where, ω k (t) is the skeleton feature weight associated with time step t, T is the total number of steps in the diffusion process of the diffusion model, and λ is a hyperparameter that controls the decay rate; Represents the output features of the i-th layer encoder of the U-Net network at time step t; represents the output features of the encoder in the i-th layer of the U-Net network, MLP(·) represents the feature mapping performed by the multi-layer perceptron MLP, and L represents the set of layers in the U-Net network into which the style features are injected.
[0022] Furthermore, in S4, a new font character image that meets expectations is generated, which is specifically implemented as follows:
[0023]
[0024] p θ (x t-1 |x t ,F k ,F s )=Ν(x t-1 ;μ θ (x t ,t,F k ,F s ),∑ θ (x t ,t,F k ,F s ))
[0025] Where p θ () is the conditional probability distribution, p(x T ) is x T The marginal probability of x0 is the initial data distribution, x t is the data of time step t with noise, T is the time step in the diffusion process, N() is the probability density function of Gaussian distribution, μ θ (x t ,t,F k ,F s ) is the denoised mean value predicted by the U-Net network during the diffusion process of the diffusion model, ∑ θ (x t ,t,F k ,F s ) is the denoised variance of the U-Net network prediction during the diffusion process of the diffusion model.
[0026] Furthermore, S5 is implemented as follows:
[0027]
[0028] Where, L MSE is the mean square error loss function, ε is the true noise, ε θ is the noise predicted by the U-Net network, is the joint expectation of the original distribution x0, the noise ε, and the time step t.
[0029] The present invention also provides a controllable Chinese character new font generation system based on a diffusion model, including a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, it can implement any of the method steps described above.
[0030] Compared with the existing technology, the present invention has the following beneficial effects: the present invention draws on the idea of structural borrowing in font design and proposes a controllable Chinese character new font generation method based on a diffusion model. By decomposing the character image into representations of skeleton features and stroke features, the skeleton and stroke features of different fonts are extracted and recombined to generate a variety of new fonts. The proposed algorithm can generate the expected new font according to the specified target skeleton structure and stroke style, and maintain a high level of stroke fluency. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 Schematic diagram of the method system flow of an embodiment of the present invention.
[0032] Figure 2 Schematic diagram of a skeleton encoder according to an embodiment of the present invention.
[0033] Figure 3 Schematic diagram of a stroke style encoder according to an embodiment of the present invention. DETAILED DESCRIPTION
[0034] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.
[0035] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present application belongs.
[0036] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0037] The present invention provides a method for generating a controllable new Chinese character font based on a diffusion model, comprising:
[0038] S1. Input the target stroke style character image and the target skeleton character image into the stroke style extraction network and the skeleton information extraction network respectively through two branches;
[0039] S2, extracting multi-scale skeleton features through the skeleton information extraction network;
[0040] S3, extracting stroke style features through a stroke style extraction network;
[0041] S4, through the feature-guided diffusion model, the denoising operation prediction of each step of the U-Net network in the diffusion process is used to generate a new font character image that meets the expectations;
[0042] S5. Optimize each step of the denoising process by using the mean square error loss function.
[0043] The following is a specific implementation process of the present invention.
[0044] like Figure 1 As shown, the present invention provides a controllable new Chinese character font generation method based on a diffusion model, which mainly includes the following steps:
[0045] S1, input the target stroke style character image and the target skeleton character image into the network through two branches respectively;
[0046] S2, extracting multi-scale skeleton features through the skeleton information extraction network;
[0047] S3, extracting stroke style features through a stroke style extraction network;
[0048] S4, using the features extracted in S2 and S3 to guide the denoising operation prediction of each step of the diffusion model in the U-net network, to generate a new font character image that meets the expectations;
[0049] S5. Optimize each step of the denoising process by using the mean square error loss function.
[0050] Furthermore, in S1, the target skeleton character image is input into the skeleton information extraction network branch; the target stroke style character image is input into the stroke style extraction network branch.
[0051] Furthermore, in S2, the skeleton information extraction network includes a skeleton extraction algorithm for extracting a skeleton image of a target skeleton character image, and a skeleton encoder for extracting multi-scale skeleton features of the skeleton image.
[0052] Furthermore, the skeleton extraction algorithm includes binarization processing, skeleton extraction, grayscale conversion, and dimension expansion.
[0053] Further, such as Figure 2 As shown, through the skeleton encoder E k Extracting multi-scale skeleton features F from skeleton images k ={F k1 ,F k2}, where F k1 is the skeleton feature of the first layer of the input U-Net network, F k2 To input the skeleton features of the second layer of the UNet network, specifically:
[0054] F k =E k (I k )
[0055] Where, I k Represents the skeleton image, skeleton encoder E k It is a neural network composed of convolutional layers.
[0056] Furthermore, in S3, the stroke style extraction network includes a stroke style encoder for extracting style features of the target stroke style character image and a multi-layer perceptron MLP for projecting the style features into the intermediate feature space of the diffusion model.
[0057] Further, such as Figure 3 As shown, through the stroke style encoder E s Extract the style feature F of the target stroke style character image s , specifically:
[0058] F s =E s (I s )
[0059] Where, I s Represents the target stroke style character image, the stroke style encoder E s It is a neural network composed of the first eight layers of VGG11.
[0060] Furthermore, in S4, the denoising operation prediction of each step of the U-Net network during the diffusion process of the feature-guided diffusion model is implemented as follows:
[0061]
[0062] Where, ω k (t) is the skeleton feature weight associated with time step t, T is the total number of steps in the diffusion process of the diffusion model, and λ is a hyperparameter that controls the decay rate; Represents the output features of the i-th layer encoder of the U-Net network at time step t; represents the output features of the encoder in the i-th layer of the U-Net network, MLP(·) represents the feature mapping performed by the multi-layer perceptron MLP, and L represents the set of layers in the U-Net network into which the style features are injected.
[0063] Furthermore, in S4, a new font character image that meets expectations is generated, which is specifically implemented as follows:
[0064]
[0065] p θ (x t-1 |x t ,F k ,F s )=Ν(x t-1 ;μ θ (x t ,t,F k ,F s ),∑ θ (x t ,t,F k ,F s ))
[0066] Where p θ () is the conditional probability distribution, p(x T ) is x T The marginal probability of x0 is the initial data distribution, x t is the data of time step t with noise, T is the time step in the diffusion process, N() is the probability density function of Gaussian distribution, μ θ (x t ,t,F k ,F s ) is the denoised mean value predicted by the U-Net network during the diffusion process of the diffusion model, ∑ θ (x t ,t,F k ,F s ) is the denoised variance of the U-Net network prediction during the diffusion process of the diffusion model.
[0067] Furthermore, S5 is implemented as follows:
[0068]
[0069] Where, L MSE is the mean square error loss function, ε is the true noise, ε θ is the noise predicted by the U-Net network, is the joint expectation of the original distribution x0, the noise ε, and the time step t.
[0070] The present invention also provides a controllable Chinese character new font generation system based on a diffusion model, including a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, it can implement any of the method steps described above.
[0071] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0072] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0073] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0074] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0075] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other manner. Any person skilled in the art may utilize the above-disclosed technical content to modify or modify the present invention into equivalent embodiments. However, any simple modifications, equivalent variations, and modifications to the above embodiments that do not depart from the technical content of the present invention and are based on the technical essence of the present invention remain within the scope of protection of the present invention.
Claims
1. A method for generating a controllable new Chinese character font based on a diffusion model, characterized in that: include: S1. Input the target stroke style character image and the target skeleton character image into the stroke style extraction network and the skeleton information extraction network respectively through two branches; S2, extracting multi-scale skeleton features through the skeleton information extraction network; S3, extracting stroke style features through a stroke style extraction network; S4, through the feature-guided diffusion model, the denoising operation prediction of each step of the U-Net network in the diffusion process is used to generate a new font character image that meets the expectations; S5. Optimize each step of the denoising process by using the mean square error loss function.
2. The method for generating a controllable new Chinese character font based on a diffusion model according to claim 1, characterized in that: In S2, the skeleton information extraction network includes a skeleton extraction algorithm for extracting a skeleton image of a target skeleton character image and a skeleton encoder for extracting multi-scale skeleton features of the skeleton image.
3. The method for generating a controllable new Chinese character font based on a diffusion model according to claim 2, wherein: The skeleton extraction algorithm includes binarization processing, skeleton extraction, grayscale conversion, and dimension expansion.
4. The method for generating a controllable new Chinese character font based on a diffusion model according to claim 2, wherein: Through the skeleton encoder E k Extracting multi-scale skeleton features F from skeleton images k ={F k1 ,F k2 }, where F k1 is the skeleton feature of the first layer of the input U-Net network, F k2 To input the skeleton features of the second layer of the UNet network, specifically: F k =E k (I k ) Where, I k Represents the skeleton image, skeleton encoder E k It is a neural network composed of convolutional layers.
5. The method for generating a controllable new Chinese character font based on a diffusion model according to claim 4, characterized in that: In S3, the stroke style extraction network includes a stroke style encoder for extracting style features of the target stroke style character image and a multi-layer perceptron MLP for projecting the style features into the intermediate feature space of the diffusion model.
6. The method for generating a controllable new Chinese character font based on a diffusion model according to claim 5, characterized in that: Through the stroke style encoder E s Extract the style feature F of the target stroke style character image s , specifically: F s =E s (I s ) Where, I s Represents the target stroke style character image, the stroke style encoder E s It is a neural network composed of the first eight layers of VGG11.
7. The method for generating a controllable new Chinese character font based on a diffusion model according to claim 6, characterized in that: In S4, the denoising operation prediction of each step of the U-Net network in the diffusion process of the feature-guided diffusion model is implemented as follows: Where, ω k (t) is the skeleton feature weight associated with time step t, T is the total number of steps in the diffusion process of the diffusion model, and λ is a hyperparameter that controls the decay rate; Represents the output features of the i-th layer encoder of the U-Net network at time step t; represents the output features of the encoder in the i-th layer of the U-Net network, MLP(·) represents the feature mapping performed by the multi-layer perceptron MLP, and L represents the set of layers in the U-Net network into which the style features are injected.
8. The method for generating a controllable new Chinese character font based on a diffusion model according to claim 7, characterized in that: In S4, a new font character image that meets expectations is generated. The specific implementation is as follows: p θ (x t-1 |x t ,F k ,F s )=Ν(x t-1 ;μ θ (x t ,t,F k ,F s ),∑ θ (x t ,t,F k ,F s )) Where p θ () is the conditional probability distribution, p(x T ) is x T The marginal probability of x0 is the initial data distribution, x t is the data of time step t with noise, T is the time step in the diffusion process, N() is the probability density function of Gaussian distribution, μ θ (x t ,t,F k ,F s ) is the denoised mean value predicted by the U-Net network during the diffusion process of the diffusion model, ∑ θ (x t ,t,F k ,F s ) is the denoised variance of the U-Net network prediction during the diffusion process of the diffusion model.
9. The method for generating a controllable new Chinese character font based on a diffusion model according to claim 8, characterized in that: S5 is specifically implemented as follows: Where, L MSE is the mean square error loss function, ε is the true noise, ε θ is the noise predicted by the U-Net network, is the joint expectation of the original distribution x0, the noise ε, and the time step t.
10. A controllable Chinese character new font generation system based on diffusion model, characterized in that: The method comprises a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the method steps according to any one of claims 1 to 9 can be implemented.
Citation Information
Cited By
Vectorization Chinese character graph generation method based on large model
CN121010668A
A large model-based vectorized Chinese character pattern generation method
CN121010668B