Chinese character style migration method and migration system thereof
By introducing channel Transformer and spatial Transformer structures into the Chinese character generation network, combined with AdaIN technology, the problems of missing outline details of Chinese character generation fonts and single font types in the existing technology are solved, and a higher quality and diverse Chinese character style transfer effect is achieved.
Patent Information
- Application Number
- CN202311820682.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-06-27
AI Technical Summary
The existing Chinese character generation network model has lost the font outline details when generating font outline details, the font types are single, and paired data sets are required.
The Trans-StarGANv2 network is proposed. By introducing channel Transformer and spatial Transformer structures, combined with adaptive instance normalization (AdaIN) technology, a variety of target style fonts are generated, which solves the problem of loss of font stroke details and styles, and improves the training intensity of the model through perceived loss.
It has better results in FID and LPIPS indicators, and is easier to be recognized in the subjective vision of the human eye, the generated font structure is clearer, the details are richer, and the overall quality is higher.
Smart Images

Figure CN120220168A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of Chinese character image processing, and particularly relates to a Chinese character style transfer method and a transfer system thereof. Background Art
[0003] In recent years, with the popularity of social software and the emergence of short videos, the communication between people has gradually been replaced by voice and video from the original letters, and even the situation of forgetting how to write a character when picking up a pen occurs. In the era of data information, Chinese characters also need to be emphasized, so the research on Chinese characters by humans is becoming more and more extensive. At present, a large number of studies mainly focus on the recognition and classification of Chinese characters, and there are not many studies on the generation of Chinese characters, especially in the generation of Chinese characters with multiple styles. Common Chinese character generation methods include two categories, traditional Chinese character generation methods and Chinese character generation methods based on deep learning.
[0004] In view of the problems that the current Chinese character generation network model loses details in the generated font outline part, the generated font types are single, and paired data sets are required, the present application hereby proposes a new Chinese character style transfer method and a transfer system thereof. Summary of the Invention
[0005] The purpose of the present invention is to propose the Trans-StarGANv2 network for the problems that the current Chinese character generation network model loses details in the generated font outline part, the generated font types are single, and paired data sets are required. This network generates multiple target style fonts through a single generator based on the StarGANv2 network. First, to solve the problem of losing the style of the generated font stroke details, a channel Transformer and a spatial Transformer structure are introduced into the generator. Compared with the pure convolutional font generation network, this structure can capture the global feature relationship, while the convolutional neural network has problems such as losing the dependence relationship of long-distance features. Second, perceptual loss is introduced to improve the training intensity of the overall model generator and discriminator. The proposed network model is experimented on a data set containing three fonts. The experimental results show that: compared with other font generation network models such as the original StarGANv2 model, the proposed network not only shows better results in terms of FID and LPIPS metrics, but also is more easily recognized by the subjective vision of the human eye.
[0006] The present invention realizes the above purpose through the following technical solutions:
[0007] In the first aspect of the present invention, a Chinese character style transfer method is provided, and the method includes,
[0008] Obtaining Chinese character image data to be style transferred;
[0009] Based on the Chinese character image data, it is input into the trained neural network model;
[0010] Obtain the Chinese character image data after style transfer;
[0011] The neural network model includes a generator module, a discriminator module, a mapping network module, and a StyleEncoder style encoder module. The generator module extracts spatial features from the Chinese character image data and performs normalization processing, and combines with the StyleEncoder style encoder module to generate the required style picture data. The discriminator module discriminates the authenticity of the picture data.
[0012] As a further optimized solution of the present invention, the generator module is a U-shaped structure, including a CBT feature extraction module, an SBT feature extraction module, and a CBTV2 feature fusion module. The CBT feature extraction module extracts features based on Self-Attention in the channel dimension, and improves and reduces the spatial resolution size based on the upsampling module and the downsampling module. After extracting the channel dimension, it passes through the SBT module to extract spatial dimension features, and then through upsampling and using the adaptive instance normalization of CBTV2 to combine with the style vector extracted by the StyleEncoder style encoder in the target font to generate the required style picture.
[0013] As a further optimized solution of the present invention, the CBT feature extraction module includes a LayerNorm layer, an MDTA layer, and a GDFN layer. Among them, the LayerNorm layer is a normalization layer. The vector will be normalized in each dimension after passing through the LayerNorm layer, so that the features remain the same in scale;
[0014] The calculation process of the MDTA layer is
[0015]
[0016] where X is the input: W d is a 3×3 depth convolution, and W P is a 1×1 point convolution;
[0017] Perform a Reshape operation on the Q and K projections and then perform a dot product to generate a channel attention map with a size of ; After that,
[0018] In the formula and are the results obtained after Reshape of Q, K, and V respectively; α is a learnable scaling parameter;
[0019] The calculation process of the GDFN layer is
[0020]
[0021] Among them, ⊙ in the formula represents element-wise multiplication; φ represents the GELU non-linear function; W d 1 and W d 2 respectively represent the first linear projection layer and the second linear projection layer after 3×3 depth convolution; W p 1 and W p 2 respectively represent the first linear projection layer and the second linear projection layer after 1×1 dimension expansion convolution; W p 0 represents 1×1 dimension reduction convolution.
[0022] As a further optimization scheme of the present invention, the calculation of the SBT feature extraction module is as follows:
[0023] Q = W l Q X;
[0024] K = W l K X;
[0025] V = W l V X;
[0026] Among them, X is the input; W represents a fully connected layer; next, a reshaping operation, that is, Reshape, is performed on the Q and K projections to enable their dot product interaction to generate a spatial attention map of size ; after that,
[0027]
[0028] Among them, and are respectively the results after reshaping Q, K, and V, that is, the results obtained after the Reshape operation; α is a learnable scaling parameter; after that,
[0029]
[0030] Among them, ⊙ is element-wise multiplication; φ represents the GELU non-linear function; W l 1 represents the linear mapping of the first linear projection layer; W l 2 represents the linear mapping of the second linear projection layer; W l 0 represents the fusion linear mapping.
[0031] As a further optimization solution of the present invention, the CBTV2 feature fusion module includes an AdaIN layer, and the formula of the AdaIN layer is:
[0032] where s is the style vector and x is the input content.
[0033] As a further optimization solution of the present invention, preprocess the Chinese character image data to be style-transferred.
[0034] The second aspect of the present invention provides a Chinese character style transfer system, including a memory and a processor. The memory includes a Chinese character style transfer method program, and when the Chinese character style transfer method program is executed by the processor, the above steps are implemented.
[0035] The beneficial effects of the present invention are as follows: The present invention proposes a Channel-Based Transformer feature extraction module, a Spatial-Based Transformer feature extraction module, and a Channel-Based Transformer V2 feature fusion module combined with adaptive instance normalization AdaIN, which enhances the model's ability to collect features and improves the accuracy of the generated picture style;
[0036] The present invention introduces a style-aware loss in the font generation task. The perceptual loss is calculated based on the feature representation of the deep neural network and is more suitable for evaluating the quality of the images generated by the generative adversarial network than the traditional pixel-level loss.
[0037] The present invention proposes a new font generation model, the Trans-StarGAN V2 model, which has a clearer font structure, richer detail styles, and higher overall quality compared to the output of the original model. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a schematic structural diagram of the steps of the Chinese character style transfer method of the present invention;
[0039] Figure 2 is the network structure of the Trans-StarGAN v2 of the present invention;
[0040] Figure 3 is the structural diagram of the CBT module of the present invention;
[0041] Figure 4 is the structural diagram of the SBT module of the present invention;
[0042] Figure 5 is the structural diagram of the CBTV2 module of the present invention;
[0043] Figure 6These are the example font generation results of the Trans-StarGAN v2 and StarGAN v2 of the present invention. Detailed implementation manners
[0044] The following further describes the present application in detail with reference to the accompanying drawings. It is necessary to point out here that the following specific implementation manners are only used to further illustrate the present application and cannot be understood as limiting the protection scope of the present application. Those skilled in the art can make some non-essential improvements and adjustments to the present application according to the above application content.
[0045] Refer to Figures 1 to 6 As shown, a Chinese character style transfer method includes:
[0046] Step S102: Obtain Chinese character image data to be style-transferred;
[0047] Step S104: Input the Chinese character image data into the trained neural network model;
[0048] Step S106: Obtain the Chinese character image data after style transfer;
[0049] The neural network model includes a generator module, a discriminator module, a mapping network module, and a StyleEncoder style encoder module. The generator module extracts spatial features from the Chinese character image data and performs normalization processing, and combines with the StyleEncoder style encoder module to generate the required style picture data. The discriminator module discriminates the authenticity of the picture data.
[0050] In this embodiment, the Trans-StarGANv2 network is adopted, and the overall loss function of this network is:
[0051]
[0052] Where is the adversarial loss, as shown in Equation (2), where
[0053]
[0054] Where is the style reconstruction loss, as shown in the formula:
[0055]
[0056] Where is the style diversity loss, as shown in the formula:
[0057]
[0058] Where The cyclic consistency loss is shown in the formula:
[0059]
[0060] In this paper, in addition to the above losses, a new style-aware loss is introduced. The pre-trained model VGG19 is used to extract and calculate the feature representations of the generated image and the real image. By calculating the differences and adding them to the overall training loss function, the style difference between the model-generated image and the target style image is further reduced, thereby improving the overall performance of the network. The feature maps obtained from the first five convolutions of VGG19 are used to calculate the perceptual loss, which are conv1, conv2, conv3, conv4, and conv5 respectively. In the above formula, represents the feature maps extracted by each convolutional layer in the pre-trained VGG network, I out and I gt are the output image and the original image respectively:
[0061]
[0062] In this embodiment, the generator module has a U-shaped structure and mainly consists of three parts: the Channel-Based Transformer (CBT) feature extraction module, the Spatial-Based Transformer (SBT) feature extraction module, and the Channel-Based Transformer V2 (CBTV2) feature fusion module. The Channel-Based Transformer and the Spatial-Based Transformer are used to extract features in both the channel and spatial dimensions simultaneously, making the generated results of our model more realistic and closer to the desired results. The Channel-Based Transformer module extracts features by using Self-Attention in the channel dimension. At the same time, the upsampling module and the downsampling module are used to increase and decrease the spatial resolution respectively. After extracting the channel dimension features, the SBT module extracts the spatial dimension features, and then through upsampling and using the adaptive instance normalization of CBTV2, the style vector extracted from the target font by the StyleEncoder is combined to generate the required style image. The discriminator is used to distinguish the authenticity of the images generated by the generator, and the adversarial training between the generator and the discriminator is carried out to improve the performance of the entire generation network.
[0063] Among them, the characteristics of the CBT feature extraction module are as follows:
[0064] Currently, the generator of the generative adversarial network mainly uses a convolutional neural network to extract features. However, as a new neural network structure, Transformer enables feature extraction to no longer rely solely on convolutional neural networks. The Transformer model solves the shortcomings of CNN such as limited receptive fields, and the unique self-attention mechanism in Transformer helps the model obtain information not limited to adjacent regions but any position. Therefore, this paper proposes a feature extraction module based on the Transformer structure, called the Channel-Based Transformer feature extraction module (CBT), as Figure 3 shown. Each CBT module consists of a LayerNorm layer, a Multi-Dconv Head Transposed Attention (MDTA) layer, and a Gated-Dconv Feed-Forward Network (GDFN) layer. Among them, the LayerNorm layer is the normalization layer. After passing through the LayerNorm layer, the vector will be normalized in each dimension, thereby making the features remain the same in scale and increasing the overall generalization ability of the model. Using LayerNorm in Transformer can also increase the model's processing ability for high-resolution images. The MDTA layer does not calculate attention in the pixel dimension but in the channel dimension. Therefore, there is no need to perform interactive calculations in the pixel dimension, but to calculate the covariance in the feature channel dimension to extract the feature map. It first performs cross-channel pixel aggregation through a 1×1 convolution and uses a 3×3 depth convolution to achieve channel-level aggregation of local context. This structure has two advantages:
[0065] One is the introduction of depth convolution, which emphasizes local context before calculating the feature covariance to generate a global attention map; the other is to calculate the cross-channel cross-covariance to generate an attention map that implicitly encodes the global context. The specific calculation process is as follows
[0066]
[0067] In the formula: X is the input: W d is a 3×3 depth convolution, and W P is a 1×1 point convolution. Next, a Reshape operation is performed on the Q and K projections and then a dot product is performed to generate a channel attention map of size . The subsequent calculations are similar to those of a conventional Transformer, as shown specifically below:
[0068]
[0069] In the formula and They are the results obtained after reshaping Q, K, and V respectively; α is a learnable scaling parameter. The GDFN layer is a gated feed-forward network based on local content fusion. It emphasizes spatial context and uses 1×1 convolutions to increase the dimension and then 3×3 convolutions to extract features, followed by gating using the GELU activation function. The gating mechanism in GDFN determines and controls whether complementary features need to be passed forward, and allows subsequent layers in the network hierarchy to focus specifically on finer image attributes, thus producing high-quality outputs. The calculation process is shown in the following formula:
[0070]
[0071] In the formula, ⊙ represents element-wise multiplication; φ represents the GELU non-linear function; W d 1 and W d 2 represent the 3×3 depthwise convolutions of the first linear projection layer and the second linear projection layer respectively; W p 1 and W p 2 represent the 1×1 dimensionality expansion convolutions of the first linear projection layer and the second linear projection layer respectively; W p 0 represents the 1×1 dimensionality reduction convolution.
[0072] The SBT feature extraction module makes the finally generated picture optimal by performing global interaction on the spatial features after the CBT module. Since the SBT module is connected after the CBT and has undergone multiple downsampling processes, it is equivalent to the features being reduced multiple times in the spatial dimension. Therefore, it mainly focuses on the channel level and can well adapt to the spatial-based Self-Attention operation. The SBT module is improved based on the FFN layer of the conventional Transformer module. The specific structure of SBT is as Figure 4 shown. The SBT module replaces the MDTA layer in the CBT with the multi-head self-attention mechanism (MHSA), and at the same time replaces the GDFN in the CBT with the Gated Feed-Forward Network (GFFN). The generator passes through the CBT feature extraction module and downsampling, which is equivalent to the features being reduced by 16 times in the spatial dimension, only 16×16 in size, while the channel dimension is expanded to 512. After being processed by the CBT feature extraction module, the channel dimension has undergone comprehensive global interaction and does not require Self-Attention in the channel dimension. Instead, Self-Attention operation is performed in the spatial dimension to make up for the far lack of global interaction in the 1×1 and 3×3 convolution operations, so that the extracted features have better global representation ability;
[0073] The calculation of the SBT feature extraction module is as follows:
[0074] Q = W l Q X;
[0075] K = W l K X;
[0076] V = W l V X;
[0077] where X is the input; W represents the fully connected layer; next, a reshaping operation, i.e., Reshape, is performed on the projections of Q and K to generate a spatial attention map of size through their dot product interaction; then, the calculation is the same as that of the conventional Transformer
[0078]
[0079] where and are the results after reshaping Q, K, and V, respectively, i.e., the results obtained after the Reshape operation; α is a learnable scaling parameter; the GFFN layer is still designed as the element-wise product of two linear projection layers, and one of them is non-linearly activated by GELU. The proposed gated feed-forward network also emphasizes the spatial context based on local content mixing (similar to the MDTA module). The difference is that GFFN performs a linear mapping on the channels and emphasizes the fusion of channel features more. The gating mechanism, like MDTA, controls which complementary features should flow forward and allows subsequent layers in the network hierarchy to focus specifically on finer image attributes, thus generating high-quality outputs; where
[0080]
[0081] where ⊙ is the element-wise multiplication; φ represents the GELU non-linear function; W l 1 represents the linear mapping of the first linear projection layer; W l 2 represents the linear mapping of the second linear projection layer; W l 0 represents the fusion linear mapping.
[0082] The main difference between the CBTV2 feature fusion module and the CBT module is that there is an additional feature fusion layer AdaIN. The AdaIN layer has two inputs, namely the content input X and the style input S. The mean and variance of the content feature are aligned with the mean and variance of the style feature. CBT V2 is used in the generator upsampling process to generate the target font we need through the AdaIN layer combined with the style vector. The formula of the AdaIN layer is:
[0083] Among them, s is the style vector and x is the input content.
[0084] experiment:
[0085] Dataset and experimental environment
[0086] This experiment uses a self-built dataset, which contains three different styles of fonts: Fang Zheng Shuti, Lishu and Xingkai. The three fonts are used as the three classes of the model, and the dataset is divided into a training set and a validation set. Each font training set contains 10,000 different font images, and each font validation set contains 2,000 different font images. The dataset is preprocessed, each font image is resized to 256×256, and the larger blank parts of the image are cropped to maximize the font filling ratio.
[0087] Experimental parameter setting and evaluation indicators
[0088] The experimental environment of this paper is built based on the Pytorch framework using pycharm. The code is written based on Python 3.6.7 under the Windows system, and Nvidia A4000 GPU is used for training, with a video memory size of 16G. In the training experiment, the Adam optimizer is selected to optimize the network, and β1=0, β2=0.99, and the batch-size during training is set to 4; the iteration is 100000 times; the initial learning rate is 1e-4, and we set λ sty =1,λ ds =1,λ cyc =1 and λ reg = 1, the learning rates of G, D, E and F are set to 10 -4 .
[0089] Three different metrics are used to evaluate the accuracy of the generated font styles of the network model in this paper. (1) Fréchet Inception Distance (FID), which measures the distance between two multivariate normal distributions; (2) Learned Perceptual Image Patch Similarity (LPIPS), which is a learnable perceptual image patch similarity used to measure the difference between two images, which is more consistent with human perception.
[0090] Experimental results analysis
[0091] In order to fully verify the effectiveness of Trans-StarGanV2 in Chinese character generation, we will conduct a comprehensive comparison with other font generation models in terms of objective measurement indicators and human subjective visual effects. All comparisons are based on the same data set to ensure fairness.
[0092] We are Figure 6 The figure shows the comparison of the results of the Trans-StarGAN v2 model and the original network model StarGAN v2 using the same input fonts and target fonts on our three font datasets. Tables 1 and 2 are the font style transfer FID indicators corresponding to the Trans-StarGAN v2 model and StarGAN v2 respectively.
[0093] The result example figures are the result figures of the proposed model and the original model, where the first row is the input font and the first column is the target font. It can be seen from the figure that compared with the original model, the proposed model retains the style of the target font in the details of the font, especially the complex font. Take the word "Lu" as an example. When converting from regular script to relaxed style, the proposed model converts the vertical fold and dot of the bottom stroke of "Lu" into the vertical fold and dot of relaxed style, retaining the stroke style of the target font, while the original model splits the vertical fold into dot, horizontal stroke and dot for style transfer.
[0094] When converting from regular script to regular script, the model in this paper converts the vertical fold and dot at the bottom of the character "陆" into the vertical fold and dot of regular script, retaining the stroke style of the target font. The original model splits the vertical fold into dots, horizontal strokes and dots before performing style transfer. The same is true for other fonts. Although the overall style transfer is performed, the model in this paper, thanks to the Transformer structure, can better extract detailed features in terms of stroke details, making the generated font style closer to the target style. The above is the result of subjective observation. Then we further judge the generation effect of the model by calculating the FID indicator of the model.
[0095] Table 1 FID index of Trans-StarGAN v2 generation results on the dataset
[0096]
[0097] Table 2 FID Metrics of the Generated Results of StarGAN v2 on the Dataset
[0098]
[0099] By observing the FID metrics of the two models, it can be seen that the change range of the model in this paper is the largest when converting from Founder Shuti to Xingkai, and the metric decreases by 7 points. It can also be clearly seen from the result graph that the effect is much better than the original model StarGAN v2. There is also a significant improvement in the FID metrics for other style conversions.
[0100] Table 3 Metrics of the Generated Results of Each Font Generation Model on the Dataset
[0101]
[0102]
[0103] Table 3 shows the metrics of the training results of various font generation models on the dataset in this paper. It can be found that our model is superior to other models in terms of metrics. Compared with zi2zi, HCCG-CycleGAN, and StyleGAN, our model has obvious advantages in the evaluation metrics of FID and LPIPS. Our method only needs to construct the target style dataset and then train immediately compared with models such as zi2zi and HCCG-CycleGAN, and endow the input original image with the style features we need. After using the Transformer structure, it has better performance in aspects such as overall font detail extraction and fusion. Experiments prove that the Channel-Based Transformer module and the Spatial-Based Transformer module can fully explore the feature relationships of spatial channels and have a significant effect on the improvement of various evaluation metrics for font style generation and fusion.
[0104] Ablation Experiment
[0105] To verify the effectiveness of the proposed network model, we conduct ablation experiments on the model. First, our model uses Transformer and convolution in combination. To verify the effectiveness of the Transformer structure, we conduct three groups of ablation experiments, respectively replacing the CBT module and the SBT module with the ResBLK of the original model, and replacing CBTV2 with the AdaINResBLK of the original model, and then verifying the effectiveness through the FID and LPIPS metrics.
[0106] Table 4 Results of the Ablation Experiment on the Generator Network Structure
[0107]
[0108]
[0109] It is found through experimental results that when only channel or spatial feature extraction is performed, the FID index has increased by 30 - 40, indicating that the obtained images are quite different from the target style. This shows that compared with convolution, the Transformer structure is no longer limited to small receptive fields in feature extraction but extracts features from a global perspective, which is more advantageous than convolution. However, using CBTV2 and AdaINResBLK for upsampling style fusion has a relatively small improvement effect. This is because AdaIN is used to combine the target style, and both calculate the mean and variance for alignment. This is also the reason for choosing AdaIN as the style transfer module, which has good effects in terms of speed and flexibility.
[0110] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A Chinese character style transfer method, characterized in that, The method includes: Obtaining Chinese character image data to be style-transferred; Based on the Chinese character image data, inputting it into the trained neural network model; Obtaining the Chinese character image data after style transfer; The neural network model includes a generator module, a discriminator module, a mapping network module, and a StyleEncoder style encoder module. The generator module extracts spatial features from the Chinese character image data and performs normalization processing, and combines with the StyleEncoder style encoder module to generate the required style picture data. The discriminator module discriminates the authenticity of the picture data.
2. The Chinese character style transfer method according to claim 1, characterized in that: The generator module is of a U-shaped structure, including a CBT feature extraction module, an SBT feature extraction module, and a CBTV2 feature fusion module. The CBT feature extraction module extracts features based on Self-Attention in the channel dimension, and improves and reduces the spatial resolution size based on the upsampling module and the downsampling module. After extracting the channel dimension, it extracts spatial dimension features through the SBT module, then performs upsampling and uses the adaptive instance normalization of CBTV2 to combine with the style vector extracted from the target font by the StyleEncoder style encoder to generate the required style picture.
3. A Chinese character style transfer method according to claim 2, characterized in that: The CBT feature extraction module includes a LayerNorm layer, an MDTA layer, and a GDFN layer. The LayerNorm layer is a normalization layer. When a vector passes through the LayerNorm layer, it will be normalized in each dimension, so that the features remain the same in scale. The calculation process of the MDTA layer is: Among them, X is the input: W d is a 3×3 depth convolution, W P is a 1×1 point convolution; Perform a Reshape operation on the Q and K projections and then perform a dot product to generate a channel attention map of size ; After that, where and are the results obtained after reshaping Q, K, and V respectively; α is a learnable scaling parameter; The calculation process of the GDFN layer is: where ⊙ is element-wise multiplication; φ represents the GELU non-linear function; W d 1 and W d 2 respectively represent the first and second linear projection layers after 3×3 depthwise convolution; W p 1 and W p 2 respectively represent the first and second linear projection layers after 1×1 expansion convolution; W p 0 represents 1×1 dimensionality reduction convolution.
4. A Chinese character style transfer method according to claim 3, characterized in that: The calculation of the SBT feature extraction module is as follows: Q = W l Q X; K = W l K X; V = W l V X; Among them, X is the input; W represents a fully connected layer; next, a reshaping operation, that is, Reshape, is performed on the projections of Q and K to enable their dot product interaction to generate a spatial attention map of size ; after that, Among them, and are respectively the results after reshaping Q, K, and V, that is, the results obtained after the Reshape operation; α is a learnable scaling parameter; then, where ⊙ is element-wise multiplication; φ represents the GELU non-linear function; W l 1 represents the linear mapping of the first linear projection layer; W l 2 represents the linear mapping of the second linear projection layer; W l 0 represents the fused linear mapping.
5. A Chinese character style transfer method according to claim 4, characterized in that: The CBTV2 feature fusion module includes an AdaIN layer, and the formula of the AdaIN layer is: Among them, s is the style vector and x is the input content.
6. A Chinese character style transfer method according to claim 1, characterized in that: Preprocessing the Chinese character image data to be style-transferred.
7. A Chinese character style transfer system, characterized in that, It includes a memory and a processor. The memory includes a Chinese character style transfer method program. When the Chinese character style transfer method program is executed by the processor, it implements the steps described in claim 1 above.
Citation Information
Cited By
Vectorization Chinese character graph generation method based on large model
CN121010668A
A large model-based vectorized Chinese character pattern generation method
CN121010668B