A method for generating Chinese character styles with natural writing characteristics

CN116935406BActive Publication Date: 2026-08-14TONGJI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-18
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0006]虽然现有方法很好地对风格迁移时的风格特征以及需迁移汉字的汉字骨架进行了约束,但仍具有以下不足之处:1.所有风格迁移的结果都建立在单字之上,即整个端到端的生成过程不仅如此,这些方法单字内生成的笔画固定,没有考虑字内笔画间以及字与字之间的连笔构造,生成结果较为单一,缺少变化;2.参考汉字必须是训练集中包含的宋体等标准字体,在实际应用场景中必须提供对应的正书书体,缺少鲁棒性,有较大局限

Benefits of technology

[0015]This invention constructs a dataset of commonly used Chinese characters' cursive strokes and proposes a generative adversarial network with an attention mechanism focused on cursive strokes. Supplemented by a groundbreaking cursive consistency loss function, it trains the mapping process from the cursive Chinese character skeleton dataset to styled Chinese characters, achieving style transfer from Chinese character skeletons to arbitrary style fonts while maintaining natural handwriting characteristics. Utilizing a skeleton extraction refinement algorithm, style transfer between Chinese characters of arbitrary styles is achieved. Compared to existing style transfer methods, it offers a significant improvement in terms of natural handwriting characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116935406B_ABST
    Figure CN116935406B_ABST
Patent Text Reader

Abstract

This invention proposes a method for generating Chinese character styles with natural handwriting characteristics, comprising the following steps: S1 Dataset construction; S2 Structural design and training of the generation model; S3 Inputting the trained and optimized model to generate single-line style Chinese characters with natural handwriting characteristics. This invention constructs a dataset of 8876 commonly used and less commonly used Chinese characters with fully connected strokes, along with a dataset of style characters. Through training an adversarial generative network, it achieves a mapping from Chinese character skeletons to style characters, while simultaneously realizing true "connected meaning despite broken strokes, and continuous flow of energy despite broken characters." This method has broad application value in the construction of datasets in the fields of art design and optical character recognition, and plays a significant role in advancing the field of generating natural handwriting style Chinese characters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for generating Chinese character styles with natural handwriting characteristics. Background Technology

[0002] As the carrier of Chinese civilization, Chinese characters possess a unique artistic beauty through their diverse writing styles. In contemporary visual design, typography is a fundamental design element and an indispensable tool for expression. For typography users, a well-chosen font can often fully embody the designer's concept, thus maintaining stylistic unity.

[0003] With the development of generative models in computer vision, style transfer and generative models have gradually become a major means of replacing traditional design methods. These models extract style features from a given dataset of style fonts and perform style transfer on other previously undesigned Chinese characters. However, existing Chinese character generation methods are based on generating components of reference characters. The generation results of identical character components are fixed and rigid, like movable type printing, lacking the interplay between strokes and components. Secondly, the application of character components is limited to single characters, resulting in rigid characters that lack the interplay between characters, thus failing to match the quality of masterpieces. Furthermore, as a state-of-the-art image generation model, diffusion models cannot adequately fit the skeleton of Chinese characters and cannot be applied to the field of Chinese character generation.

[0004] Chinese patent application CN202010333081.2 discloses a method and system for Chinese character style transfer based on a multi-task adversarial learning network, including: acquiring images of Chinese characters to be style transferred; inputting the images of Chinese characters to be style transferred into a trained multi-task adversarial learning network; and outputting multiple font images after style transfer from the trained multi-task adversarial learning network. Using a unified encoder to learn the universal visual patterns of reference fonts that are important for all target fonts maximizes the propagation of feature-level information across tasks while preserving task-specific features in their respective network channels, this multi-task training strategy makes the training of the Chinese character style transfer network more stable, improves the network's generalization ability, and results in font styles that are more consistent with the target fonts with clear stroke boundaries.

[0005] Chinese patent application number CN202011564611.0 proposes a Chinese font style transfer method. Based on the original generative adversarial network, it innovatively adds two auxiliary networks to the generator composed of a recurrent generative adversarial network. The first is to extract structural features from the original image and the image generated by the generator through a Chinese character classification and recognition residual network. The second is to extract the style features generated by the generator using a style encoder to ensure style consistency.

[0006] While existing methods effectively constrain the stylistic features and the skeletal structure of the characters to be transferred during style transfer, they still have the following shortcomings: 1. All style transfer results are based on individual characters, meaning the entire end-to-end generation process is not only limited to this, but the strokes generated within a single character are fixed, without considering the connection between strokes within a character or between characters, resulting in relatively monotonous and unchanging results; 2. The reference characters must be standard fonts such as Songti included in the training set, and corresponding regular script fonts must be provided in practical application scenarios, lacking robustness and having significant limitations. Summary of the Invention

[0007] To address the shortcomings of existing style transfer methods, which produce rigid and inflexible results, and to achieve the natural handwriting quality of Chinese characters generated through style transfer, thus realizing the effect of "broken strokes but connected meaning, and broken characters but connected spirit," this invention proposes a Chinese character style generation method with natural handwriting quality. This method has wide application value in the construction of datasets in the fields of art and design as well as optical character recognition.

[0008] Technical solution:

[0009] A method for generating Chinese character styles with natural handwriting characteristics is proposed. In the data preprocessing stage, Bézier curves of 8876 regular script Chinese characters are processed and fully connected to the strokes, constructing a dataset of paired Chinese character skeleton lines and their corresponding style images. In the style transfer stage, a neural network with a recurrent adversarial generative network as the base model is trained, and an attention mechanism is added. The encoded stroke information is then concatenated with the encoding of the embedding layer. Simultaneously, a stroke consistency loss function is creatively designed to measure natural handwriting characteristics, thus focusing the generator on generating natural handwriting. In summary, this invention provides a method for generating Chinese character styles with natural handwriting characteristics, realizing the mapping process from fully connected skeleton Chinese characters to styled Chinese characters, thereby generating single-line styled Chinese characters with natural handwriting characteristics.

[0010] A method for generating Chinese character styles with natural handwriting characteristics includes:

[0011] S1 dataset construction;

[0012] Structural design and training of the S2 generative model;

[0013] The model trained and optimized using S3 input is used to generate single-line style Chinese characters with natural handwriting characteristics.

[0014] Advantages of the technical solution of this invention:

[0015] This invention constructs a dataset of commonly used Chinese characters' cursive strokes and proposes a generative adversarial network with an attention mechanism focused on cursive strokes. Supplemented by a groundbreaking cursive consistency loss function, it trains the mapping process from the cursive Chinese character skeleton dataset to styled Chinese characters, achieving style transfer from Chinese character skeletons to arbitrary style fonts while maintaining natural handwriting characteristics. Utilizing a skeleton extraction refinement algorithm, style transfer between Chinese characters of arbitrary styles is achieved. Compared to existing style transfer methods, it offers a significant improvement in terms of natural handwriting characteristics.

[0016] This invention constructs a dataset of 8876 commonly used and less commonly used Chinese characters, consisting of complete stroke connections, along with a dataset of stylistic Chinese characters. By training an adversarial generative network, it achieves a true mapping from character skeletons to stylistic characters, realizing "connected meaning even when strokes are broken, and connected spirit even when characters are broken." This invention has broad application value in the construction of datasets in the fields of art design and optical character recognition. It also significantly advances the field of generating naturally written Chinese character styles. Attached Figure Description

[0017] Figure 1 This is the implementation process of the present invention;

[0018] Figure 2 This is an example of a paired dataset titled "Chinese Character Skeleton - Zhao Mengfu's Running Script".

[0019] Figure 3 This is the model architecture for the example implementation;

[0020] Figure 4 This example demonstrates the generation of a Su Shi-style font (in paired characters, the left side represents the generated result, and the right side represents the target).

[0021] Figure 5 This example demonstrates the generation of the Shu Tong style font (left is the generated result, right is the target).

[0022] Figure 6 This example demonstrates the generation of a Zhao Mengfu-style font (left is the generated result, right is the target). Detailed Implementation

[0023] The technical solutions provided in this application will be further described below with reference to specific embodiments and accompanying drawings. The advantages and features of this application will become clearer from the following description.

[0024] The overall process is as follows Figure 1 As shown.

[0025] S1 dataset construction:

[0026] This invention constructs a fully connected Chinese character skeleton dataset by editing the closed curves in the Bézier curves and then rendering them.

[0027] Bézier curves are classified into first-order, second-order, and third-order curves according to their order, denoted as:

[0028] M(a,b)L(c,d) (a first-order Bézier curve from point (a,b) to point (c,d);

[0029] M(a,b)C(x,y,c,d) (a second-order Bézier curve starting from point (a,b) and ending at point (c,d), with control point coordinates (x,y).

[0030] M(a,b)Q(x1,y1,x2,y2,c,d) (a third-order Bézier curve starting from point (a,b) and ending at point (c,d), with control points at coordinates (x1,y1) and (x2,y2).

[0031] In the dataset of this invention, the skeleton of a Chinese character is composed of several lines, where the number of lines is the number of strokes in the Chinese character. Each line is formed by connecting the three types of curves mentioned above.

[0032] By connecting the end of each stroke of a Bézier curve to the beginning of the next stroke, a fully connected dataset of Chinese character skeletons is constructed. The addition method involves adding a line segment between each stroke.<path stroke-width="3"stroke="#000"d="M a b L c d"fill="none"stroke-linecap="round"> "Where, a, b, c, and d represent the endpoint coordinates of the previous line segment and the starting coordinates of the next line segment, respectively. Next, by rendering styled Chinese character images from TTF font files and mapping them one-to-one with the skeleton dataset based on their Unicode encoding, a pair of skeleton-style Chinese character datasets is constructed, such as..." Figure 2 As shown.

[0033] Structural design and training of the S2 generative model

[0034] S2.1 Constructing the Model Structure

[0035] First, to ensure style consistency, this invention uses a convolutional neural network to construct a style encoder. A series of Chinese characters of the same style are input into the encoder, and the extracted high-dimensional style features are reduced in dimensionality using PCA to obtain style feature vectors.

[0036] Next, to ensure that the generated Chinese characters retain the original character skeleton and focus on the correspondence between strokes, the Bezier curves of the Chinese character skeleton are connected end to end to generate a fully connected Chinese character skeleton, which is then input into the structure encoder. The resulting structural features are also reduced in dimensionality by PCA to obtain structural feature vectors.

[0037] To ensure the model focuses on generating cursive strokes, this invention creatively proposes a module focused on cursive stroke information during training. A cursive stroke encoding attention matrix is ​​established and input into the model along with the dataset annotations. Specifically, the Chinese character skeleton rendered as an image is divided into 8*8 sub-images. If the original Bézier curve contains cursive strokes, the sub-image is set to 1; otherwise, it is set to 0, resulting in a 1024-dimensional one-hot encoding matrix. The attention encoding matrix is ​​flattened and concatenated with the two feature vectors mentioned above, and finally input into the generator. During generator training, the generation quality of the generated results at cursive strokes is measured against the training set, and a cursive stroke consistency loss function is defined as the training direction.

[0038] The discriminator aims to distinguish between the generator's output and the paired datasets provided in the training set, classifying the former as false and the latter as true. Backpropagation, using a comprehensive loss function, continuously trains and optimizes both the discriminator and the generator.

[0039] The overall architecture of the model is as follows Figure 3 As shown.

[0040] Training the S2.2 generative model

[0041] S2.2.1 Define the skeleton image of Chinese characters as c g The style of the Chinese character image is c f The Chinese character skeleton to be generated is c p , and c f Other complete Chinese character images with the same style are c s c g and c f The image vector obtained by concatenating along the channel dimension is c. v c extracted using a style encoder p The style code is s p .

[0042] Additionally, the character c x Style coding uses s x This indicates that the cursive coding uses T. x This indicates that x can take any character, as described above.

[0043] The S2.2.2 comprehensive loss function consists of the following four parts:

[0044] 1) Adversarial loss, used to measure the distance between the generator's output and paired datasets;

[0045]

[0046] Where D s(c) represents the discriminator, which determines whether an image c with style s was generated by the generator. G(c,s) indicates that this is the result generated by the generator after adding style based on the Chinese character skeleton. f For style Chinese character images c f Style coding.

[0047] During training, the adversarial loss function optimizes both the generator and the discriminator. The generator aims to produce images so that the discriminator cannot distinguish between the original input and the generated image. The discriminator, on the other hand, aims to identify as accurately as possible which image was generated by the generator and which was the original image.

[0048] 2) L1 loss, used to measure the difference between the generated result and the standard result at the pixel level;

[0049] Target mean absolute error loss function:

[0050] Where c v c f s s For the style encoding of the Chinese character image to be transferred, G(c v ,s s ) is C v and S s The result after using the generator.

[0051] The target mean absolute error loss function is used to optimize the generator. The more similar the generated Chinese character image is to the target image, the more ideal the generation is considered. The similarity between the two images is determined by calculating the pixel-wise mean absolute error loss function, which is then optimized to approach zero. This loss function ensures that the generator's output is visually similar to the target output.

[0052] 3) Style coding loss, used to evaluate the coding effect of the style encoder;

[0053] Style coding loss function:

[0054] Where c v s s s f The meaning is the same as described above.

[0055] The style encoding loss function is used to optimize both the generator and the style encoder. Its workflow is as follows: Noise is randomly sampled and passed through a mapping network to obtain a style code corresponding to a novel, specific style. This style, along with the skeleton, is then fed into the generator to produce a target image G. The style extracted from the current image using the style encoder is compared with the style generated from the noise. The mean absolute error loss is calculated, with the goal of minimizing it to zero. This process ensures that the extracted style remains consistent with the original style after passing through the generator. This optimization improves both the accuracy of the style encoder's style code extraction and the generator's utilization of the style code.

[0056] 4) Stroke consistency loss, used to focus on the consistency between the strokes of the generated Chinese character and the natural handwriting of the target Chinese character;

[0057]

[0058] Let S(c) represent the extraction of text structure from text image c using a ligature encoding extractor, outputting it as an embedding vector. The original image and the target font image generated by the generator are fed into a residual neural network. The mean squared error loss of the output embedding vector is calculated. Through training, the loss function is minimized, resulting in a final effect where the ligature encodings of the original image and the target font are consistent across the embedding layers. This loss function-optimized generator can preserve the ligature information of the text during the generation process, increasing its naturalness.

[0059] A comprehensive loss function is used to reconcile the above four loss functions. The comprehensive loss function of the model can be expressed as follows:

[0060]

[0061] Where λ gan , λ L1 , λ enc , λ stroke Let λ represent the parameters used to combine the loss functions. This formula represents the parameters used during network training and subsequent experiments. gan The value of λ is set to 10. L1 , λ enc and λ stroke Set to 1, the target expression used. Subscripts G, E, F, and D represent the generator, style encoder, structure encoder, and discriminator, respectively.

[0062] To verify the model's ability to generate natural-looking Chinese characters, several fonts were selected for example verification. Figure 4 , Figure 5 , Figure 6The results of natural handwriting style transfer for three representative example fonts are shown. Su Shi's font, Shu Tong's font, and Zhao Mengfu's font were selected as experimental examples. The image on the left is the transfer result of the model after inputting the Chinese character skeleton, and the image on the right is the target image in the test set.

[0063] The above description is merely a description of preferred embodiments of this application and is not intended to limit the scope of this application in any way. Any changes or modifications made by those skilled in the art based on the above-disclosed technical content should be considered as equivalent and valid embodiments and fall within the scope of protection of the technical solution of this application.

Claims

1. A method for generating Chinese character styles with natural handwriting characteristics, characterized in that, Including the following steps: S1 dataset construction; Structural design and training of the S2 generative model; S3 inputs are used to train and optimize the model to generate single-line style Chinese characters with natural handwriting; S1 Dataset Construction: A fully connected Chinese character skeleton dataset was constructed by editing the closed curves in the Bézier curves and then rendering them. Specifically: Bézier curves are classified into first-order, second-order, and third-order curves according to their order, denoted as: M(a,b)L(c,d) represents a first-order Bézier curve from point (a,b) to point (c,d); M(a,b)C(x,y,c,d) represents a second-order Bézier curve that starts from point (a,b) and ends at point (c,d), with the control point having coordinates (x,y). M(a,b)Q(x1,y1,x2,y2,c,d) represents a third-order Bézier curve starting from point (a,b) and ending at point (c,d), with control points at coordinates (x1,y1) and (x2,y2). In the dataset, the skeleton of a Chinese character is composed of several lines, the number of which is the number of strokes in the character; each line is formed by connecting the three types of curves mentioned above. By connecting the end of each stroke of a Bézier curve to the beginning of the next stroke, a fully connected dataset of Chinese character skeletons is constructed. Its features are, S2 specifically refers to: S2.1 Constructing the Model Structure First, to ensure style consistency, a style encoder is constructed using a convolutional neural network. A series of Chinese characters of the same style are input into the encoder, and the extracted high-dimensional style features are reduced in dimensionality using PCA to obtain style feature vectors. Next, to ensure that the generated Chinese characters retain the original skeleton and focus on the correspondence between strokes, the Bezier curves of the Chinese character skeleton are connected end to end to generate a fully connected Chinese character skeleton and input into the structure encoder. The resulting structural features are also reduced in dimensionality by PCA to obtain the structural feature vector. Finally, to ensure the model can focus on the generation of cursive strokes, a module focusing on cursive stroke information was proposed during training; a cursive stroke encoding attention matrix was established and input into the model as a label of the dataset. The goal of the discriminator is to distinguish between the generator's generated results and the paired datasets provided in the training set, classifying the former as false and the latter as true; the discriminator and generator are continuously trained and optimized through backpropagation using a comprehensive loss function. S2.2 Training of the generative model S2.2.1 Defines the skeleton image of Chinese characters as follows: The style of Chinese character images is The Chinese character skeleton to be generated is ,and Other complete Chinese character images with the same style are , and The image vector obtained by concatenating along the channel dimension is: Extracted using a style encoder The style code is ; S2.2.2 The comprehensive loss function consists of the following four parts: 1) Adversarial loss, used to measure the distance between the generator's output and paired datasets; (1) in The discriminator determines whether an image c with style s was generated by the generator. This indicates that the generator added styles based on the Chinese character skeleton. Style encoding for styled Chinese character images; 2) L1 loss, used to measure the difference between the generated result and the standard result at the pixel level; Target mean absolute error loss function: (2) in, Style encoding for Chinese character images whose style needs to be transferred. C v and S s The result after using the generator; 3) Style coding loss, used to evaluate the coding effect of the style encoder; Style coding loss function: (3) 4) Ligature consistency loss, used to focus on the consistency between the ligatures of the generated Chinese characters and the natural handwriting of the target Chinese characters; (4) The model's overall loss function: Among them gan , L1 , enc , stroke These are the parameters used to combine the loss functions; subscripts. , , , These represent the generator, style encoder, structure encoder, and discriminator, respectively.

2. The method for generating Chinese character styles with natural handwriting characteristics as described in claim 1, characterized in that, The method involves connecting the end of each stroke of a Bézier curve to the beginning of the next stroke, thus constructing a fully connected dataset of Chinese character skeletons. The addition method is as follows: Add a line segment between each stroke. <path stroke-width="3" stroke="#000" d="Ma b L c d" fill="none" stroke-linecap="round">< / path> "Where, a, b, c, and d are the coordinates of the endpoint of the previous line segment and the starting coordinates of the next line segment, respectively." Next, by rendering style Chinese character images from TTF font files and matching them one-to-one with the skeleton dataset according to Unicode encoding, a pair of skeleton-style Chinese character datasets are constructed.

3. The method for generating Chinese character styles with natural handwriting characteristics as described in claim 1, characterized in that, A connection-encoding attention matrix is ​​established and input into the model as a label for the dataset. Specifically, the Chinese character skeleton rendered as an image is divided into 8*8 sub-images. If the original Bézier curve contains connection strokes, the sub-image is set to 1; otherwise, it is set to 0, resulting in a 1024-dimensional one-hot encoding matrix. The attention encoding matrix is ​​flattened and concatenated with the two feature vectors mentioned above, and finally input into the generator. During the generator training process, the generation quality of the generated result at connection strokes is measured against the training set, and a connection stroke consistency loss function is defined as the training direction.

Citation Information

Patent Citations

  • A Chinese Character Style Transfer Method and System Based on Multi-Task Adversarial Learning Networks

    CN111553246B

  • Chinese font style migration method

    CN112633430A

  • Method and device for beautifying cursive style of handwritten Chinese characters

    CN102013109A

  • Chinese character font generation method

    CN116152374A