Self-supervised font generation method for enhancing font style extraction by introducing dynamic convolution kernel

The self-supervised Chinese character generation method enhanced by dynamic convolution kernels solves the problem of inaccurate extraction by fixed convolution kernels, and achieves high-precision feature extraction and edge detail capture for Chinese characters of different styles. The generated fonts are more natural and suitable for calligraphy and ancient book digitization.

CN121505069APending Publication Date: 2026-02-10HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511883863.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing deep learning models face a contradiction between the static nature of fixed convolutional kernels and the diversity of Chinese character styles when extracting stylistic features. This leads to inaccurate extraction when dealing with fonts of special styles and patterns, making it impossible to distinguish between inherent stylistic features and sample defects, thus affecting the accuracy and completeness of the generated fonts.

Method used

A self-supervised Chinese character generation method enhanced by dynamic convolutional kernels is adopted. By constructing a dynamic convolutional kernel generation module, an edge information extraction module, and a self-supervised network, combined with a CBAM attention weighting enhancement module and multi-layer residual blocks, adaptive style extraction and edge constraints of dynamic convolutional kernels are realized. A total adversarial loss function is constructed for training.

Benefits of technology

It significantly improves the adaptability and accuracy of feature extraction for Chinese characters of different styles, and the generated fonts have more natural edge details. It is suitable for scenarios such as calligraphy font generation and ancient book digitization, and overcomes the problems of poor cross-style generalization and neglect of edge details in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121505069A_ABST
    Figure CN121505069A_ABST
Patent Text Reader

Abstract

The invention discloses a self-supervised Chinese character generation method introducing dynamic convolution to enhance font style extraction, which is suitable for the fields of computer vision, font design and calligraphy education digitization, and comprises the following steps: constructing a dynamic convolution enhanced self-supervised network model which introduces a dynamic convolution kernel into the font style extraction field, and the convolution kernel parameters are dynamically adjusted to adapt to the edge and structure characteristics of the Chinese characters of different styles, so that the accurate capture of the style information is realized. According to the method, font style information extraction is enhanced through dynamic convolution, and the problems that in a traditional font generation method, a fixed convolution kernel is depended, the kernel scale and weight need to be manually adjusted, and the fixed kernel and a style extraction target are obviously separated, so that the model adjusting and optimizing process is tedious, and the cross-style generalization ability is poor are solved; therefore, the feature extraction adaptability and accuracy of different styles of Chinese characters are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and deep learning technology, specifically a Chinese character generation method based on dynamic convolution kernels and self-supervised learning. Background Technology

[0002] In the fields of multi-style font generation and intelligent calligraphy scoring, deep learning models can utilize multimodal font data (such as reference fonts of different styles, images of calligraphy works, and stroke structure annotations) to achieve font generation. Chinese character style extraction is a core technical step in multi-style font generation, intelligent calligraphy scoring, and the digitization of ancient books. Its goal is to accurately capture "macro-style types" (such as regular script, running script, and clerical script) and "fine-grained style features" (such as stroke pauses, stroke curvature, stroke spacing, and ink density) from reference samples, providing a core basis for subsequent generation or scoring.

[0003] Existing methods for extracting font style information using deep learning typically employ shallow convolutional layers in models like VGG and ResNet to extract the font's framework and style features, and fixed convolutional kernels and attention enhancement mechanisms to extract edge information. However, the static nature of fixed convolutional kernels fundamentally contradicts the diversity of Chinese character styles, thus affecting the extraction of styles for fonts with special characteristics and styles. In practical applications, reference fonts often exhibit anomalies such as intentional ink breaks (artistic expression), scanning noise, and incomplete strokes; some artistic fonts even have significantly different frameworks from the original text. In these situations, current style extraction techniques lack feature discrimination capabilities: they may misjudge "false edges" caused by ink breaks as genuine style features, leading to large areas of blank space in the subsequently generated font; they cannot distinguish between inherent style features and sample defects, such as mistakenly extracting stains from ancient book scans as a certain style, affecting extraction accuracy. Summary of the Invention

[0004] The present invention addresses the shortcomings of the existing technology by proposing a self-supervised font generation method that incorporates dynamic convolution kernels to enhance font style extraction. This method aims to extract accurate font styles more effectively, thereby improving the accuracy and completeness of font generation and better assisting designers in font design.

[0005] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: The self-supervised Chinese character generation method of the present invention, which introduces dynamic convolution to enhance font style extraction, is characterized by the following steps: Step 1: Collect Chinese character samples in N styles, and obtain M reference Chinese character images for each style to form a style reference dataset. Perform preprocessing to obtain the preprocessed style Chinese character sample set. ,in, Indicates the first preprocessed step A set of Chinese character samples in various styles, and , express The first in Let N be the number of Chinese character images, M be the number of styles, H be the height of the Chinese character images, and W be the width of the Chinese character images; The true style tag is recorded as ; Collection of standard reference font image sets ;in, Indicates the first A standard reference font image; let The real content tag is recorded as ; Step 2: Construct a self-supervised network with dynamic convolutional kernel enhancement, including: a basic style encoder module, a CBAM attention weighting enhancement module, a dynamic convolutional kernel generation module, an edge information extraction module, and a dynamic convolutional feature extraction module, and then... Processing yields the first... The first style A pure edge image and the The first style One final style feature ; Step 3: Construct the generative network, including: a content encoder, a style-content blending module, a T-layer residual block, and a U-layer transposed convolutional block, and then... and Processing yields the first... The first style A preliminary font image is generated. ; Step 4: Build and optimize the network, including: content discriminator Style discriminator and to and Processing is performed to construct the first Total combat loss of each style ; Step 5: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full The input is placed into a self-supervised network enhanced by dynamically generated convolutional kernels, and the edge information extraction module outputs the first... The first style Initial generated edge map Therefore, using equation (1) to construct the first The first style pixel-level self-supervised loss : (1) In equation (1), i is the row index of the pixel, with a value range from 1 to H, and j is the column index of the pixel, with a value range from 1 to W. Step 6: Construct the first equation using equation (2) The first style Total loss function : (2) In equation (2), These are weight parameters; Step 7: Minimize the total loss using the Adam optimizer The total network consisting of the self-supervised network, the generative network, and the tuning network is iteratively trained until the maximum number of iterations is reached, thereby obtaining the initial trained total model. Step 8: Utilize the generative network in the initially trained overall model to... and Process the data and output the m-th high-quality multi-style font image of the n-th style. ; Step 9: Construct the direct comparison loss using equation (3) The system is used to train the overall model after initial training, thereby obtaining the optimal overall model. The generator network in the optimal overall model is then used as a self-supervised Chinese character generator to generate multi-style font images. (3) In equation (3), c is the index of the RGB channel.

[0006] The self-supervised Chinese character generation method for enhancing font style extraction by introducing dynamic convolution as described in this invention is also characterized in that step 2 includes: Step 2.1: The basic style encoder module... Process and output the first... The first style Preliminary stylistic features Where h represents the height of the preliminary style feature, and w represents the width of the preliminary style feature. The number of output channels for initial style features; Step 2.2: The CBAM attention weighting enhancement module utilizes equation (4) to... Process and output the first... The first style Attention-weighted features ; (4) In equation (4), For element-wise multiplication, ChannelAtt represents channel attention operation, and SpatialAtt represents spatial attention operation; Step 2.3: The dynamic convolution kernel generation module... Processing is performed to obtain the first... The first style A sequence of dynamic convolutional kernels ,in, For the number of dynamic kernels, The size of the dynamic core; Step 2.4: The edge information extraction module uses... right Perform convolution processing to obtain the first... The first style The first edge feature map; then the edge continuity is checked using the Canny operator. The first style The edge feature map is processed to obtain the first edge feature map. The first style A pure edge image ; Step 2.5: The dynamic convolution feature extraction module, based on... Using equation (5) Perform dynamic convolution operations and output the first... The first style One final style feature : (5) In equation (5), For convolution operations, for The k-th dynamic convolution kernel.

[0007] Furthermore, step 2.3 includes: Step 2.3.1: Use equation (6) to... Compress the number of channels and output the first... The first style Style characteristics after channel compression : (6) In equation (6), To compress the convolution kernel, To compress the bias term, This represents the number of channels after compression, and satisfies... , This indicates a convolution operation with a 1×1 kernel; Step 2.3.2: Use equation (7) to... Perform global average pooling operation and output the first... The first style Global style encoding vectors : (7) In equation (7), express The Middle Line number The full-channel feature vector corresponding to the column space location; Step 2.3.3: The three-layer MLP network utilizes equation (8) to... Processing yields the first... The first style The original parameter matrix of the dynamic convolution kernel : (8) In equation (8), , , These are the weight matrices for the first, second, and third layers of the MLP network, respectively. , , These are the bias terms for the first, second, and third layers of the MLP network, respectively. Indicates the activation function; Step 2.3.4: Use equation (9) to... Perform the dimension reconstruction operation to obtain the first dimension. The first style Dynamic convolution kernels : (9) In equation (9), Reshape represents the dimension reconstruction operation.

[0008] Furthermore, step 3 includes: Step 3.1: The content encoder uses equation (10) to... Perform product operations and bias processing, then output. Content features : (10) In equation (10), For content convolution kernel, For bias terms, This represents the convolution operation of the content encoder; Step 3.2: The style-content hybrid module will... The number of channels and After adjusting the number of channels to be consistent, we obtain the first... The first style Style features after channel alignment ; Step 3.3: The residual blocks of layer T are sequentially processed... and Perform residual fusion and output the T-th layer residual block. The first style One fusion feature ; Step 3.4: The transposed convolutional blocks of layer U are sequentially processed... Perform an upsampling operation and output the first... The first style A preliminary font image is generated. .

[0009] Furthermore, step 4 includes: Step 4.1: Content Discriminator Using equations (11) and (12) respectively and After processing, the probability of reasonableness of the m-th content in the n-th style is obtained. The probability of reasonableness of the m-th content in the n-th style : (11) (12) In equations (11) and (12), , , They are respectively The three convolutional layers, , , for The weight kernels corresponding to the three convolutional layers, , , for The bias terms corresponding to the three convolutional layers, and BN for batch normalization. Use the Sigmoid activation function; Step 4.2: Construct a content discriminator using equation (13) The loss of style : (13) Step 4.3: Style Discriminator Using equations (14) and (15) respectively and Perform binary classification to obtain the style consistency probability of the m-th style for the n-th style. The probability of consistency with the m-th style of the n-th style : (14) (15) In equations (14) and (15), , , They are respectively The three convolutional layers, , , for The weight kernels corresponding to the three convolutional layers, , , for The bias terms of the three convolutional layers; Step 4.4: Construct a style discriminator using equation (16) The nth style loss : (16) Step 4.5: Construct the first step of the optimization network using equation (17). Total combat loss of each style : (17) In equation (17), , There are two weight parameters.

[0010] The present invention provides an electronic device comprising a processor and a memory, wherein the memory stores a computer program as described in any one of claims 1-5, and the processor executes the computer program to implement the method described therein.

[0011] The present invention discloses a computer-readable storage medium storing a computer program, characterized in that the computer program is executed by a processor to perform the steps of the method.

[0012] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The adaptive style extraction module based on dynamic convolution enhancement in this invention. This module uses a dynamic convolution kernel generator to automatically generate suitable convolution kernel parameters according to the style differences of the input Chinese characters (such as "thin horizontal strokes and thick vertical strokes" in regular script and "flowing and elegant strokes" in running script). This overcomes the problems of traditional font generation methods, which rely on fixed convolution kernels, require manual adjustment of kernel scale and weights, and have a significant separation between the fixed kernel and the style extraction target, resulting in a cumbersome model tuning process and poor cross-style generalization ability (such as generating running script but retaining the brushstrokes of regular script). It significantly improves the adaptability and accuracy of feature extraction for Chinese characters of different styles, while shortening the model tuning time for new styles.

[0013] 2. The adaptive style extraction module of this invention also achieves deep collaboration between dynamic convolution kernels and self-supervised edge constraints. The generated dynamic convolution kernels, while adapting to the target style (such as the symmetrical structure of seal script and the bold strokes of cursive script), can accurately focus on the key areas of Chinese character edges, forming a closed-loop optimization logic of "dynamic extraction - self-supervised verification". After the edge features extracted by the dynamic kernel are verified by the pure edge map, the kernel parameters are further fine-tuned through loss backpropagation, so that the kernel can not only adapt to the overall style features, but also accurately capture fine-grained details such as edge transitions and stroke connections. This effectively solves the defect of traditional dynamic convolution methods that only focus on the overall style adaptation and ignore the accuracy of edge details (such as broken strokes and edge distortion in the generated font). This collaborative design makes the style adaptability and detail capture capability of the dynamic kernel complementary, so that the extracted style features have both global style consistency and maintain the authenticity of edge details. The edge details of the generated font are more restored than those of the single dynamic convolution method, and the stroke connections are more natural. It is especially suitable for scenarios such as calligraphy font generation and ancient book digitization where edge details are strictly required. Attached Figure Description

[0014] Figure 1 This is a flowchart of the self-supervised font generation method of the present invention; Figure 2 This is a model structure diagram of the style extraction part of the self-supervised font generation method of the present invention. Detailed Implementation

[0015] In this implementation, a self-supervised Chinese character generation method that introduces dynamic convolution to enhance font style extraction is applied to multi-style Chinese character generation scenarios, such as digital preservation of traditional calligraphy, personalized font customization, and development of calligraphy education aids. Through a multi-stage strategy of "data preprocessing → dynamic style extraction → deep style-content fusion → adversarial training optimization → direct comparison loss tuning," it effectively addresses the shortcomings of existing technologies, such as poor style adaptation with fixed convolution kernels, loss of generated details, and weak robustness to broken ink samples. Specifically, for example... Figure 1 As shown, it includes the following steps: Step 1: Collect Chinese character samples in N styles, including 266 fonts such as Canglang Xingkai, Classical Textbook Song, Baige Tianxing, Artistic Handwriting, and Songti. For each style, obtain M reference Chinese character images. For each font, collect 787 Chinese character images. Take 6 images at a time to construct the sample set. In the actual use case, M=6. This forms the style reference dataset, which is then preprocessed to obtain the preprocessed style Chinese character sample set. ,in, Indicates the first preprocessed step A set of Chinese character samples in various styles, and , express The first in Let N be the number of Chinese character images, M be the number of styles, H be the height of the Chinese character images, and W be the width of the Chinese character images; The true style tag is recorded as H=W=128.

[0016] Collection of standard reference font image sets ;in, Indicates the first A standard reference font image; the actual use cases include and cover all Chinese characters in the style sample set. The real content tag is recorded as ; In a specific implementation, the style sample reference image is preprocessed to obtain a style Chinese character sample set. The process can be divided into two scenarios: When the font source is a downloaded font file: a 128×128 black and white binary image is generated directly without any additional processing; When the font source is handwritten data captured by a camera: 1) Grayscale conversion (weighted average method) ); 2) Binarization (Otsu algorithm, (Values ​​range from 135 to 160); 3) Denoising and Normalization (Gaussian Filtering) The size is uniformly 128×128, and the pixel values ​​are normalized to [0,1]).

[0017] Step 2: As Figure 2 As shown, a self-supervised network with dynamic convolutional kernel enhancement is constructed, including: a basic style encoder module, a CBAM attention weighting enhancement module, a dynamic convolutional kernel generation module, an edge information extraction module, and a dynamic convolutional feature extraction module. Furthermore, [the network is then used for...]. Processing yields the first... The first style A pure edge image and the The first style One final style feature .

[0018] Step 2.1: Basic Style Encoder Module Process and output the first... The first style Preliminary stylistic features Where h represents the height of the preliminary style feature, and w represents the width of the preliminary style feature. The number of output channels for preliminary style features; in specific use cases The image resolution is consistent with the input image, and the 512 channels can fully accommodate the multi-dimensional features of different styles, such as the regularity of regular script and the smoothness of running script.

[0019] In a specific implementation, the basic style encoder has a 7-layer structure: Layer 1 is a 7×7 convolution (outputting 64 channels), Layers 2-5 are 4×4 convolutions (outputting 128, 256, 256, and 256 channels respectively), Layer 6 is global average pooling, and Layer 7 is a 1×1 convolution (outputting 512 channels); the convolution kernels are initialized using a He normal distribution (He). The bias term is initialized to 0 and is used to capture global textures and local details of the style.

[0020] Step 2.2: The CBAM attention-weighted enhancement module utilizes equation (4) to... Process and output the first... The first style Attention-weighted features ; (4) In equation (4), For element-wise multiplication, ChannelAtt represents channel attention operation, and SpatialAtt represents spatial attention operation.

[0021] In a specific implementation: Channel attention: through global average pooling (GAP) Input a two-layer MLP (input 512 → hide 128 → output 512, Sigmoid activated) to enhance the weight of key style channels; Spatial attention: The GAP is concatenated with the result of global max pooling (GMP), and then a 1×1 convolution (outputting 1 channel, sigmoid activation) is used to locate style key regions; Element-level multiplication ( This enables the coordinated enhancement of channel and spatial features.

[0022] Step 2.3: The dynamic convolution kernel generation module... Processing yields the first... The first style A dynamic convolution kernel sequence ,in, For the number of dynamic kernels, The size of the dynamic core; Step 2.3.1: Use equation (6) to... Compress the number of channels and output the first... The first style Style characteristics after channel compression : (6) In equation (6), To compress the convolution kernel, To compress the bias term, This represents the number of channels after compression, and satisfies... , This indicates a convolution operation with a 1×1 kernel; in a specific instance, a 1×1 convolution is used to... Channel number compressed to (512 / 8).

[0023] Step 2.3.2: Use equation (7) to... Perform global average pooling operation and output the first... The first style Global style encoding vectors Global style encoding vector ; (7) In equation (7), express The Middle Line number The full-channel feature vector corresponding to the column space location.

[0024] Step 2.3.3: The three-layer MLP network utilizes equation (8) to... Processing yields the first... The first style The original parameter matrix of the dynamic convolution kernel : (8) In equation (8), , , These are the weight matrices for the first, second, and third layers of the MLP network, respectively. , , These are the bias terms for the first, second, and third layers of the MLP network, respectively. This represents the activation function; in a specific example, z is input into a 3-layer MLP (input 64 → hidden 256 → hidden 128 → output 1152), generating 128 3×3 convolutional kernel parameters. .

[0025] Step 2.3.4: Use equation (9) to... Perform the dimension reconstruction operation to obtain the first dimension. The first style Dynamic convolution kernels : (9) In equation (9), Reshape represents the dimension reconstruction operation. Dimension reconstruction logic: ... Reorganized into Where "512" is the number of input channels (and ). (The number of channels is consistent), satisfying the basic rule that "the number of input channels of the convolution kernel = the number of feature channels".

[0026] Step 2.4: Edge information extraction module uses right Perform convolution processing to obtain the first... The first style The edge feature map is then compared with the edge continuity check using the Canny operator (low threshold 0.1, high threshold 0.3, determined after optimization on the validation set). The first style The edge feature map is processed to obtain the first edge feature map. The first style A pure edge image ; Step 2.5: The dynamic convolution feature extraction module extracts features based on... Using equation (5) Perform dynamic convolution operations and output the first... The first style One final style feature : (5) In equation (5), For convolution operations, for The k-th dynamic convolution kernel.

[0027] Convolutional logic: 128 dynamic kernels (each 512×3×3) are respectively connected to... Perform independent convolution on (512×128×128). Each kernel outputs a single-channel style detail map of 1×128×128 (such as connected strokes and starting stroke details), and then stacks them by channels into a 128-channel feature map.

[0028] The actual meaning of "summation": in the weight book It is not pixel summation, but the traversal and integration of the convolution operations of 128 kernels. The engineering implementation is "per-kernel convolution + channel stacking" to ensure that style details are not lost.

[0029] Output feature value: The 128 channels correspond to 128 types of style details respectively, accurately representing fine-grained features such as stroke strength and ink shade, laying a foundation for subsequent fusion.

[0030] Step 3: Construct a generation network, including: content encoder, style-content hybrid module, T-layer residual block, U-layer transposed convolution block, and perform and processing to obtain the th <00级别的初步生成字体图像个初步生成字体图像 ; Step 3.1: The content encoder performs product operation and bias processing on using Equation (10), and outputs the content feature of : (10) In Equation (10), is the content convolution kernel, is the bias term, represents the convolution operation of the content encoder; specifically, "two-dimensional convolution + batch normalization" is used to extract stable Chinese character structure features from .

[0031] Content encoder structure: 6-layer network design, the first layer is 7×7 convolution (64 channels, IN normalization, ReLU, maintaining the input size); the second to fourth layers are 4×4 convolution (128, 256, 512 channels); the fifth to sixth layers are residual blocks (to avoid gradient disappearance), and the output is .

[0032] Characteristics of content features: Focus on structural information such as stroke layout and radical position, independent of style. For example, the content features of the character "wood" remain the same in regular script and running script styles.

[0033] Function of batch normalization: Adjust the feature value range to make the distribution of content features and style features consistent, providing compatibility for residual fusion.

[0034] Step 3.2: The style-content hybrid module will... The number of channels and After adjusting the number of channels to be consistent, we obtain the first... The first style Style features after channel alignment ; Alignment: The number of channels is 512. The number of channels is 128, and it is processed by 1×1 convolution (kernel dimension (512×128×1×1)). The number of channels was adjusted to 512, resulting in... .

[0035] Step 3.3: The residual blocks of layer T are sequentially processed... and Perform residual fusion and output the T-th layer residual block. The first style One fusion feature Choose T=6 residual blocks (to balance fusion effect and computational cost), each layer contains a "BN+ReLU+3×3 convolution" structure, padding=1, and keep the feature map size unchanged.

[0036] Step 3.4: The transposed convolutional blocks of layer U are sequentially processed... Perform an upsampling operation and output the first... The first style A preliminary font image is generated. ,Right now .

[0037] Step 4: Build and optimize the network, including: content discriminator Style discriminator and to and Processing is performed to construct the first Total combat loss of each style .

[0038] Step 4.1: Content Discriminator Using equations (11) and (12) respectively and After processing, the probability of reasonableness of the m-th content in the n-th style is obtained. The probability of reasonableness of the m-th content in the n-th style : (11) (12) In equations (11) and (12), , , They are respectively The three convolutional layers, , , for The weight kernels corresponding to the three convolutional layers, , , for The bias terms corresponding to the three convolutional layers, and BN for batch normalization. It is a Sigmoid activation function; in practice, the weight kernel... , , Initialized using a He normal distribution, bias term , , Initialize to 0; the LeakyReLU activation function avoids gradient sparsity, making discriminator training more stable.

[0039] Step 4.2: Construct a content discriminator using equation (13) The loss of style : (13) Step 4.3: Style Discriminator Using equations (14) and (15) respectively and Perform binary classification to obtain the style consistency probability of the m-th style for the n-th style. The probability of consistency with the m-th style of the n-th style : (14) (15) In equations (14) and (15), , , They are respectively The three convolutional layers, , , for The weight kernels corresponding to the three convolutional layers, , , for The bias terms of the three convolutional layers; using The three-layer convolutional structure, weight kernel , , Independently initialized using the He normal distribution, bias term , , Initialize to 0 to ensure the independence of style feature recognition.

[0040] Step 4.4: Construct a style discriminator using equation (16) The nth style loss : (16) Step 4.5: Construct the first step of the optimization network using equation (17). Total combat loss of each style : (17) In equation (17), , There are two weighting parameters. In this embodiment, the weighting parameters... =1.0、 =1.0 (determined after 5 iterations of optimization on the validation set), ensuring that structure discrimination and style discrimination are equally important, and avoiding optimization bias; the total adversarial loss directly reflects the discriminator's comprehensive ability to distinguish between "structure + style".

[0041] Step 5: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] The input is placed into a self-supervised network enhanced by dynamically generated convolutional kernels, and the edge information extraction module outputs the first... The first style Initial generated edge map Using the same Canny operator as in step 2.4 (low threshold 0.1, high threshold 0.3), ensures consistency between the supervision benchmark and the extraction logic of generated edges, resulting in... Therefore, equation (1) is used to construct the first... The first style pixel-level self-supervised loss : (1) In equation (1), i is the row index of the pixel, with a value range from 1 to H, and j is the column index of the pixel, with a value range from 1 to W. Step 6: Construct the first step using equation (2) The first style Total loss function : (2) In equation (2), These are the weighting parameters. Weighting parameters: =0.1 (determined by 5 validation set optimizations), balancing pixel loss and adversarial loss to avoid generating stiff or uncontrolled images.

[0042] Step 7: Minimize the total loss using the Adam optimizer The total network consisting of the self-supervised network, the generative network, and the tuning network is iteratively trained until the maximum number of iterations is reached, thereby obtaining the initial trained total model. Step 8: Utilize the generative network in the initially trained overall model to... and Process the data and output the m-th high-quality multi-style font image of the n-th style. ; In practice, the training strategy was as follows: ≥50,000 iterations, batch size of 8, training on an NVIDIA RTX 4060 (8GB VRAM) based on the PyTorch framework, with each iteration taking 0.3 seconds and the total training time being approximately 4.2 hours.

[0043] Convergence criterion: The loss on the validation set decreases by less than 1e-5 for 5 consecutive validation set iterations to avoid overfitting. Generate the result after convergence. (3-channel RGB image).

[0044] Step 9: Construct the direct comparison loss using equation (3) The system is used to train the overall model after initial training, thereby obtaining the optimal overall model. The generator network in the optimal overall model is then used as a self-supervised Chinese character generator to generate multi-style font images. (3) In equation (3), c is the index of the RGB channel.

[0045] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.

[0046] In this embodiment, a computer-readable storage medium stores a computer program, which, when executed by a processor, performs the steps of the above-described method. The processor's operating system is compatible with Windows 11 / Ubuntu 20.04, the deep learning framework is PyTorch 1.18, the programming language is Python 3.7, and the dependent libraries include OpenCV (image preprocessing), NumPy (numerical computation), and Matplotlib (result visualization).

[0047] In summary, this invention solves the problem of poor style adaptation with fixed convolution kernels by using a style extraction module enhanced by dynamic convolution; it filters out false edges with broken ink, reducing the blank rate of generated fonts; and it improves the accuracy of font generation tasks, significantly enhancing cross-style generalization ability.

Claims

1. A self-supervised Chinese character generation method that incorporates dynamic convolution to enhance font style extraction, characterized in that, Includes the following steps: Step 1: Collect Chinese character samples in N styles, and obtain M reference Chinese character images for each style to form a style reference dataset. Perform preprocessing to obtain the preprocessed style Chinese character sample set. ,in, Indicates the first preprocessed step A set of Chinese character samples in various styles, and , express The first in Let N be the number of Chinese character images, M be the number of styles, H be the height of the Chinese character images, and W be the width of the Chinese character images; The true style tag is recorded as ; Collection of standard reference font image sets ;in, Indicates the first A standard reference font image; let The real content tag is recorded as ; Step 2: Construct a self-supervised network with dynamic convolutional kernel enhancement, including: a basic style encoder module, a CBAM attention weighting enhancement module, a dynamic convolutional kernel generation module, an edge information extraction module, and a dynamic convolutional feature extraction module, and then... Processing is performed to obtain the first... The first style A pure edge image and the The first style One final style feature ; Step 3: Construct the generative network, including: a content encoder, a style-content blending module, a T-layer residual block, and a U-layer transposed convolutional block, and then... and Processing is performed to obtain the first... The first style A preliminary font image is generated. ; Step 4: Build and optimize the network, including: content discriminator Style discriminator and to and Processing is performed to construct the first Total combat loss of each style ; Step 5: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] The input is placed into a self-supervised network enhanced by dynamically generated convolutional kernels, and the edge information extraction module outputs the first... The first style Initial generated edge map Therefore, using equation (1) to construct the first The first style pixel-level self-supervised loss : (1) In equation (1), i is the row index of the pixel, with a value range from 1 to H, and j is the column index of the pixel, with a value range from 1 to W. Step 6: Construct the first equation using equation (2) The first style Total loss function : (2) In equation (2), These are weight parameters; Step 7: Minimize the total loss using the Adam optimizer The total network consisting of the self-supervised network, the generative network, and the tuning network is iteratively trained until the maximum number of iterations is reached, thereby obtaining the initial trained total model. Step 8: Utilize the generative network in the initially trained overall model to... and Process the data and output the m-th high-quality multi-style font image of the n-th style. ; Step 9: Construct the direct comparison loss using equation (3) The system is used to train the overall model after initial training, thereby obtaining the optimal overall model. The generator network in the optimal overall model is then used as a self-supervised Chinese character generator to generate multi-style font images. (3) In equation (3), c is the index of the RGB channel.

2. The self-supervised Chinese character generation method based on dynamic convolution to enhance font style extraction as described in claim 1, characterized in that, Step 2 includes: Step 2.1: The basic style encoder module... Process and output the first... The first style Preliminary stylistic features Where h represents the height of the preliminary style feature, and w represents the width of the preliminary style feature. The number of output channels for initial style features; Step 2.2: The CBAM attention weighting enhancement module utilizes equation (4) to... Process and output the first... The first style Attention-weighted features ; (4) In equation (4), For element-wise multiplication, ChannelAtt represents channel attention operation, and SpatialAtt represents spatial attention operation; Step 2.3: The dynamic convolution kernel generation module... Processing is performed to obtain the first... The first style A dynamic convolution kernel sequence ,in, For the number of dynamic kernels, The size of the dynamic core; Step 2.4: The edge information extraction module uses... right Perform convolution processing to obtain the first... The first style The first edge feature map; then the edge continuity is checked using the Canny operator. The first style The edge feature map is processed to obtain the first edge feature map. The first style A pure edge image ; Step 2.5: The dynamic convolution feature extraction module, based on... Using equation (5) Perform dynamic convolution operations and output the first... The first style One final style feature : (5) In equation (5), For convolution operations, for The k-th dynamic convolution kernel.

3. The self-supervised Chinese character generation method based on dynamic convolution to enhance font style extraction as described in claim 2, characterized in that, Step 2.3 includes: Step 2.3.1: Use equation (6) to... Compress the number of channels and output the first... The first style Style characteristics after channel compression : (6) In equation (6), To compress the convolution kernel, To compress the bias term, This represents the number of channels after compression, and satisfies... , This indicates a convolution operation with a 1×1 kernel; Step 2.3.2: Use equation (7) to... Perform global average pooling operation and output the first... The first style Global style encoding vectors : (7) In equation (7), express The Middle Line number The full-channel feature vector corresponding to the column space location; Step 2.3.3: The three-layer MLP network utilizes equation (8) to... Processing is performed to obtain the first... The first style The original parameter matrix of the dynamic convolution kernel : (8) In equation (8), , , These are the weight matrices for the first, second, and third layers of the MLP network, respectively. , , These are the bias terms for the first, second, and third layers of the MLP network, respectively. Indicates the activation function; Step 2.3.4: Use equation (9) to... Perform the dimension reconstruction operation to obtain the first dimension. The first style Dynamic convolution kernels : (9) In equation (9), Reshape represents the dimension reconstruction operation.

4. The self-supervised Chinese character generation method based on dynamic convolution to enhance font style extraction as described in claim 1, characterized in that, Step 3 includes: Step 3.1: The content encoder uses equation (10) to... Perform product operations and bias processing, then output. Content features : (10) In equation (10), For content convolution kernel, For bias terms, This represents the convolution operation of the content encoder; Step 3.2: The style-content hybrid module will... The number of channels and After adjusting the number of channels to be consistent, we obtain the first... The first style Style features after channel alignment ; Step 3.3: The residual blocks of layer T are sequentially processed... and Perform residual fusion and output the T-th layer residual block. The first style One fusion feature ; Step 3.4: The transposed convolutional blocks of layer U are sequentially processed... Perform an upsampling operation and output the first... The first style A preliminary font image is generated. .

5. The self-supervised Chinese character generation method based on dynamic convolution to enhance font style extraction as described in claim 1, characterized in that, Step 4 includes: Step 4.1: Content Discriminator Using equations (11) and (12) respectively and After processing, the probability of reasonableness of the m-th content in the n-th style is obtained. The probability of reasonableness of the m-th content in the n-th style : (11) (12) In equations (11) and (12), , , They are respectively The three convolutional layers, , , for The weight kernels corresponding to the three convolutional layers, , , for The bias terms corresponding to the three convolutional layers, and BN for batch normalization. Use the Sigmoid activation function; Step 4.2: Construct a content discriminator using equation (13) The loss of style : (13) Step 4.3: Style Discriminator Using equations (14) and (15) respectively and Perform binary classification to obtain the style consistency probability of the m-th style for the n-th style. The probability of consistency with the m-th style of the n-th style : (14) (15) In equations (14) and (15), , , They are respectively The three convolutional layers, , , for The weight kernels corresponding to the three convolutional layers, , , for The bias terms of the three convolutional layers; Step 4.4: Construct a style discriminator using equation (16) The nth style loss : (16) Step 4.5: Construct the first step of the optimization network using equation (17). Total combat loss of each style : (17) In equation (17), , There are two weight parameters.

6. An electronic device includes a processor and a memory, the memory storing a computer program, wherein the processor executes the computer program to implement the steps of the method as described in any one of claims 1-5.

7. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1-5.