A Chinese character font attribute control method and system based on hyperplane modeling and latent code editing

CN122549367APending Publication Date: 2026-08-11FUZHOU UNIV ZHICHENG COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-23
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,将现有潜码编辑方法(如InterFaceGAN)简单转用于汉字字体领域存在以下本质困难:

Benefits of technology

[0024]首先,无需依赖大量成对标注数据,通过超平面建模方法即可在生成器潜在空间中学习到与目标属性对应的编辑方向,显著降低了数据标注成本;其次,通过基于同一字体多字符样本的层级响应强度统计平均分配差异化编辑权重,能够有效避免对汉字结构控制层的过度干预,在属性编辑过程中保持汉字字形结构的稳定性与笔画的协调性,解决了现有方法易出现笔画断裂、粘连或拓扑关系破坏的问题;再次,所获得的属性编辑方向具备跨字体复用性,无需针对新字体或新字符进行额外训练,大幅提升了方法的实用性与泛化能力;最后,通过字体潜码映射器与生成器的联合训练以及潜在编码的迭代优化,能够保证生成字体图像的质量与语义一致性,满足精细化字体设计的需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122549367A_ABST
    Figure CN122549367A_ABST
Patent Text Reader

Abstract

This invention provides a method and system for controlling Chinese character font attributes based on hyperplane modeling and latent code editing, belonging to the field of deep learning generative model technology. Based on the latent codes of multiple font samples with attribute scores in the generator's hierarchical latent space, a linear support vector machine is trained to obtain a separating hyperplane. The normal vector of the separating hyperplane is determined as the latent direction corresponding to the target attribute. The response intensity of the latent codes of multiple character samples under the same font at each level of the latent direction is calculated. The average response intensity of each level is statistically averaged to obtain the average response intensity of each level, and differentiated editing weights are assigned to different levels. The latent code of the Chinese character font image to be edited is obtained in the generator's hierarchical latent space. Based on the latent direction and the differentiated editing weights, the latent code to be edited is hierarchically weighted to obtain the edited latent code. The edited latent code is input into the generator to generate the attribute-edited Chinese character font image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of deep learning generative model technology, specifically relating to a method and system for controlling Chinese character font attributes based on hyperplane modeling and latent code editing. Background Technology

[0002] Chinese characters, as information carriers, are characterized by their vast number and complex structure, leading to long design cycles and high costs in traditional font design. With the development of deep learning technology, significant progress has been made in automatic Chinese font generation methods based on generative models. However, existing research largely focuses on font style transfer or overall style conversion, with less attention paid to the continuous and controllable editing of individual font attributes.

[0003] Existing methods for editing font attributes typically rely on large amounts of paired annotation data, making it difficult to establish a stable correspondence between user subjective perception and font visual attributes. More importantly, Chinese characters have strict structural constraints; adjusting a particular attribute (such as weight, height, width, or flatness) often affects the overall structural balance, leading to problems such as uneven strokes, localized overlap, or disruption of topological relationships. For example, when adjusting the weight attribute, existing methods easily result in an imbalance in the thickness ratio of horizontal and vertical strokes, deformation of turning edges, or unnatural connections between strokes.

[0004] Latent code editing technology, by finding editing directions corresponding to semantic attributes in the latent space of a generative model and adjusting the directionality of the latent encoding, can achieve continuous and controllable changes in attributes while maintaining semantic consistency. This technology has been proven to have good semantic separability and structure preservation capabilities in fields such as face image editing. However, simply applying existing latent code editing methods (such as InterFaceGAN) to the field of Chinese character fonts presents the following fundamental difficulties:

[0005] First, human face images are natural RGB images with a relatively loose structure and a high tolerance for local deformation; while Chinese characters are black and white binary images with strict stroke topological relationships and strong structural constraints, and simple directional perturbations can easily lead to stroke breakage or adhesion.

[0006] Second, there is a large body of research on the mapping relationship between facial attributes (such as smile and age) and latent coding; however, the stylistic attributes of Chinese fonts (such as font weight and height) are highly subjective and semantically ambiguous, and there is a lack of effective means to establish a correspondence between users' subjective feelings and visual attributes.

[0007] Third, existing latent code editing methods typically apply a uniform amplitude of directional perturbation to all levels of the latent code; however, the structural information of Chinese fonts is mainly controlled by low-resolution layers, while stroke details are mainly controlled by high-resolution layers. Uniform perturbation can easily lead to excessive intervention in the structural layers, causing local morphological distortion.

[0008] The aforementioned difficulties indicate that existing latent code editing frameworks for human faces cannot be directly transferred to the field of Chinese character fonts. Systematic adaptation and innovation are urgently needed to address the unique structural and semantic characteristics of Chinese characters. Therefore, exploring latent code editing methods that can achieve continuous attribute adjustment while maintaining the structural stability of Chinese characters is of great significance for improving the controllability and practicality of the font generation process. Summary of the Invention

[0009] To address the shortcomings and deficiencies of existing technologies, this invention provides a method and system for controlling Chinese character font attributes based on hyperplane modeling and latent code editing. First, based on the latent codes of multiple font samples with attribute scores in the generator's hierarchical latent space, a linear support vector machine is trained to obtain a separating hyperplane, and its normal vector is determined as the latent direction corresponding to the target attribute. Then, the response intensity of the latent codes of multiple character samples under the same font at each level of this latent direction is calculated, and the average response intensity of each level is obtained through statistical averaging. Differential editing weights are then assigned to different levels accordingly. For the Chinese character font image to be edited, the initial latent code is extracted through a font latent code mapper and iteratively optimized to obtain the latent code to be edited. Then, hierarchical weighted editing is performed based on the aforementioned latent direction and differential editing weights. Finally, the attribute-edited Chinese character font image is generated by the generator. This invention establishes a correspondence between visual attributes and subjective feelings by introducing a font personality dimension, obtains semantically stable attribute editing directions using a hyperplane modeling method, and achieves continuous and controllable adjustment of font attributes while maintaining the structural stability of Chinese characters through a hierarchical weighted editing strategy, exhibiting good cross-font generalization ability.

[0010] The specific technical solution adopted by this invention to solve its technical problem is as follows:

[0011] A method for controlling Chinese character font attributes based on hyperplane modeling and latent code editing, specifically including the following steps:

[0012] Based on the latent encodings of multiple font samples with attribute scores in the generator's hierarchical latent space, a linear support vector machine (SVM) is trained to obtain a separating hyperplane. The normal vector of this separating hyperplane is then determined as the latent direction corresponding to the target attribute. Specifically, the SVM learns the separating hyperplane by maximizing the attribute class margin. This hyperplane defines the optimal decision boundary between different attribute levels in the original high-dimensional hierarchical latent space, and its normal vector is the principal direction of change of the target attribute in the latent space, referred to as the latent direction. This latent direction possesses stable semantic representation capabilities, can characterize the continuous change trend of attributes in the latent space, and exhibits cross-font generalization across different characters.

[0013] The latent encoding of multiple character samples under the same font is calculated, and the response intensity at each level in the latent direction is obtained by statistically averaging the response intensity at each level. Differential editing weights are then assigned to different levels based on the average response intensity. Here, the response intensity reflects the sensitivity of each level in the hierarchical latent space to changes in the target attribute. By statistically averaging multiple character samples under the same font, the influence of random fluctuations of a single character on the hierarchical weight allocation can be eliminated, improving the statistical robustness and cross-character generalization ability of the weight allocation. Differential editing weights are assigned based on a positive correlation between the response intensity of each level. Smaller editing weights are applied to structural control levels that control the overall proportion and spatial layout of the glyphs, while larger editing weights are applied to non-structural control levels that control stroke shape and detail expression. Inhibitory weights are applied to levels with excessively low response intensity and insignificant correlation, thus concentrating the perturbation energy of subsequent editing operations on levels that are meaningful to attribute changes and have low structural risk.

[0014] The latent encoding of the Chinese character font image to be edited is obtained in the hierarchical latent space of the generator. This latent encoding serves as the starting point for subsequent editing operations and contains a complete style and structural representation of the font image to be edited in the hierarchical latent space.

[0015] Based on the latent direction and the differentiated editing weights, the latent code to be edited is subjected to hierarchical weighted editing to obtain the edited latent code. During editing, the product of the corresponding differentiated editing weight and the global editing intensity parameter is applied to each level of latent code as the editing amplitude, and a linear displacement is performed along the latent direction to achieve hierarchical control of attribute changes. Among them, the global editing intensity parameter controls the overall attribute change amplitude and determines the strength of attribute adjustment; the differentiated editing weights control the distribution ratio of editing disturbances at different levels and determine the control precision of the attribute expression region. Through the coordinated adjustment of the two, continuous and controllable editing of the target attribute is achieved.

[0016] The edited latent encoding is input into the generator to generate a Chinese character font image with edited attributes.

[0017] Furthermore, the average response intensity of each layer is obtained by calculating the arithmetic mean of the absolute values ​​of the inner products of the corresponding layer latent vectors and the latent directions of multiple character samples under the same font. The absolute value of the inner product measures the projection magnitude of the latent vector of that layer onto the latent direction, i.e., the magnitude of the response of that layer to attribute changes.

[0018] Further, obtaining the latent encoding of the Chinese character font image to be edited includes: inputting the Chinese character font image to be edited into a font latent code mapper to extract an initial latent encoding, and using the initial latent encoding as a starting point to obtain the latent encoding to be edited through iterative optimization. The iterative optimization aims to minimize the difference between the input image and the image reconstructed by the generator, while constraining the latent encoding to not deviate from the generator's effective latent distribution. The font latent code mapper is a multi-scale convolutional coding network, composed of a convolutional backbone and a skip mapping module, used to extract font features at different resolutions, projecting each scale feature into a sub-latent encoding of the corresponding level, and then concatenating them to form a complete latent encoding matching the number of levels in the generator.

[0019] Furthermore, the generator is a StyleGAN2 generation network adapted to adjust for Chinese character font characteristics, including assigning independent latent vectors to each resolution layer to form W. + A hierarchical latent space and weighted demodulation are used instead of adaptive instance normalization. The design of assigning independent latent vectors to each resolution layer allows different levels of the generator to independently control visual features at different scales, from overall structure to stroke details, providing a structural foundation for hierarchical weighted editing. The introduction of weighted demodulation can eliminate the droplet-like artifacts generated during the generation process by traditional adaptive instance normalization, improving the generation quality of Chinese character font images.

[0020] Furthermore, the separating hyperplane satisfies the following relation: Where d is the normal vector of the separating hyperplane, corresponding to the latent direction, and b is the bias term. For potential encoding.

[0021] Furthermore, the multiple font samples with attribute scores are obtained through the following method: A hierarchical representation system of font visual attributes is established by introducing a font personality dimension; continuous attribute level annotations are performed on the font samples along the target attribute dimension; and an attribute classifier is used to predict the attributes of the annotated font images to obtain attribute scores. The font personality dimension includes visual attribute representations at the structural, stroke, and detail levels. The structural level includes the character's width-to-height ratio, center-of-gravity distribution, and overall expansion direction; the stroke level includes stroke thickness, visual weight, and curvature; and the detail level includes stroke endpoint shapes, turning methods, and local decorative treatments. By establishing the above hierarchical representation system, the user's subjective aesthetic perception of font style is systematically linked to quantifiable multi-scale visual features, providing a semantic foundation for the accurate acquisition of subsequent attribute scores and effective modeling of latent directions.

[0022] The present invention also provides a Chinese character font attribute control system, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the methods described above.

[0023] Compared with the prior art, the present invention and its preferred embodiments have at least the following beneficial effects:

[0024] First, without relying on a large amount of paired labeled data, the hyperplane modeling method can learn the editing direction corresponding to the target attribute in the generator's latent space, significantly reducing the cost of data annotation. Second, by statistically averaging the hierarchical response intensity based on multi-character samples of the same font, differentiated editing weights can be effectively allocated, avoiding excessive intervention in the Chinese character structure control layer. This maintains the stability of the Chinese character's shape structure and the coordination of strokes during attribute editing, solving the problems of stroke breakage, adhesion, or topological relationship destruction that are common in existing methods. Third, the obtained attribute editing direction is reusable across fonts, eliminating the need for additional training for new fonts or characters, greatly improving the method's practicality and generalization ability. Finally, through joint training of the font latent code mapper and the generator, as well as iterative optimization of the latent encoding, the quality and semantic consistency of the generated font images can be guaranteed, meeting the needs of refined font design. Attached Figure Description

[0025] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:

[0026] Figure 1 This is a schematic diagram of the latent direction modeling and hierarchical weight determination process in an embodiment of the present invention.

[0027] Figure 2 This is a schematic diagram illustrating the process of editing and generating Chinese character font image attributes according to an embodiment of the present invention.

[0028] Figure 3 This is a schematic diagram illustrating the multi-level influencing factors of font personality in an embodiment of the present invention.

[0029] Figure 4 This is a schematic diagram of the font latent code mapper structure according to an embodiment of the present invention.

[0030] Figure 5 This is a schematic diagram of the generator structure according to an embodiment of the present invention.

[0031] Figure 6 This is a visualization distribution diagram of the potential encoding t-SNE dimensionality reduction in an embodiment of the present invention.

[0032] Figure 7 This is a diagram showing the average response intensity of each layer in an embodiment of the present invention. Detailed Implementation

[0033] To make the features and advantages of the present invention more apparent and understandable, specific embodiments are described below in detail:

[0034] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0035] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0036] This invention aims to solve the following technical problems existing in the prior art: existing font attribute editing methods rely on a large amount of paired annotation data, making it difficult to achieve the correspondence between the user's subjective feeling and the font's visual attributes; the stability of the font structure and the coordination of strokes are easily damaged during the single attribute editing process; and there is a lack of systematic latent directional modeling and hierarchical editing schemes suitable for the characteristics of Chinese fonts.

[0037] To address this, this invention discloses a method and system for controlling Chinese character font attributes based on hyperplane modeling and latent code editing. The method includes: constructing a font latent code mapper and jointly training it with a generator; mapping attribute-annotated font images to a latent space to obtain latent codes; obtaining attribute scores for font samples through an attribute classifier; training a linear support vector machine based on the latent codes with attribute scores to obtain a separating hyperplane, using its normal vector as the latent direction for attribute editing; calculating the response intensity of the latent code at each level of the latent direction, and applying differentiated weights to different levels for hierarchical weighted editing; and obtaining the attribute-edited Chinese character font image through a generator after mapping, encoding, and hierarchical weighted editing of the font image to be edited. This invention establishes a correlation between visual attributes and semantics by introducing a font personality dimension, and utilizes joint training and hierarchical weighting strategies to achieve continuous and controllable adjustment of font attributes while maintaining the structural stability of Chinese characters, exhibiting cross-font generalization capability.

[0038] Compared with existing technologies, this invention establishes a semantic bridge between user subjective perception and font visual attributes by introducing a font personality dimension, reducing the reliance on a large amount of paired labeled data. Based on this, a jointly trained font latent code mapper is used to adapt the generator's latent space to the structural and stylistic characteristics of Chinese characters. Furthermore, latent coding regularization constraints are introduced to limit the degree to which the encoder output deviates from the average latent coding distribution, thereby improving the stability and generation quality of the latent coding and providing a good foundation for subsequent hyperplane learning. Then, a separating hyperplane modeling method based on linear support vector machines is adopted to learn attribute latent directions with stable semantic expressions in the original high-dimensional latent space, achieving continuous, smooth, and controllable adjustment of font attributes. The obtained latent directions are reusable across fonts, requiring no additional training for new characters. Finally, a hierarchical weighted editing strategy is used, allocating differentiated weights based on the response intensity of the latent directions at each level, concentrating attribute editing on the main response levels while suppressing excessive intervention in the structural control layer. This allows attribute editing to be completed while maintaining the stability of the Chinese character structure and the coordination of strokes, effectively avoiding structural distortion and stroke adhesion.

[0039] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0040] like Figure 1 and Figure 2 As shown, this embodiment provides a method for controlling Chinese character font attributes based on hyperplane modeling and latent code editing, specifically including the following steps:

[0041] S1. Construct and label a dataset of Chinese character font attributes, as follows:

[0042] S1.1 Introducing the dimension of font personality, establishing a hierarchical representation system of font visual attributes from the structural, stroke, and detail levels. For example... Figure 3 As shown, the structural level includes the width-to-height ratio of the characters, the distribution of the center of gravity, and the overall direction of development; the stroke level includes the thickness of the strokes, visual weight, and the shape of the strokes; and the detail level includes the shape of the stroke endpoints, the way the strokes turn, and the local decorative treatment.

[0043] S1.2. Based on the continuous variation range of the target attribute, select multiple font samples covering the entire attribute spectrum. The target attribute includes font weight, height / slimness, or width / flatness. Taking font weight as an example, select 100 font samples from Founder Type Library and Hanyi Type Library with continuous thickness variations from thin to thick to ensure that the data has a continuous distribution in the font weight dimension.

[0044] S1.3 Render the font file into a character image of a uniform format, and perform character alignment and resolution standardization processing. In this embodiment, the TrueType Font font file is rendered into a character image of a uniform format, all font images are subjected to resolution standardization processing, uniformly adjusted to 128×128 resolution, and the characters are center aligned.

[0045] S1.4. Based on the visual features of the font in the target attribute dimension, the samples are continuously labeled with attribute levels. Taking font weight as an example, the font samples are divided into multiple weight levels according to the thickness of the strokes (such as thin, regular, medium, thick, extra thick, etc.) and assigned category labels to form a training dataset with attribute labels.

[0046] S2. Construct a font latent code mapper and jointly train it with the generator on the dataset; map the font image labeled with target attributes to the generator through the trained font latent code mapper. Latent space, to obtain the corresponding latent encoding The details are as follows:

[0047] S2.1 Construct a font latent code mapper, such as Figure 4 As shown, the font latent code mapper is a multi-scale convolutional coding network, consisting of a convolutional backbone and a skip mapping module. After extracting font features at different resolutions, it projects the features at each scale into sub-latent codes of the corresponding layers, and finally concatenates them to form a complete latent code corresponding to the number of layers in the generator. Latent encoding;

[0048] S2.2, Jointly train the font latent code mapper and the generator on the dataset, such as... Figure 5 As shown, the generator is a generative network based on StyleGAN2 with adaptive adjustments for Chinese character font characteristics. This includes: adapting the generator input resolution to match the size of the Chinese character font image (e.g., 128×128); assigning independent latent vectors to each resolution layer of the generator, forming a multi-layer structure. A hierarchical latent space (e.g., 12 layers) is used; weight demodulation is employed instead of adaptive instance normalization to eliminate the teardrop-shaped generation artifacts introduced by the adaptive instance normalization operation in traditional StyleGAN; the discriminator reduces the number of high-resolution layers and deep convolutional modules, and the network enters discrimination immediately after the input resolution is quickly downsampled to an 8×8 feature scale. Since Chinese character font images have relatively regular structures, simple backgrounds, and concentrated semantics, the discriminator can easily capture generation defects. Therefore, simplifying the discriminator structure can match its discriminative ability to the actual needs of the font task and alleviate the training imbalance problem.

[0049] The overall loss function for joint training is:

[0050]

[0051] In the formula, , , , and Let these represent the adversarial loss, reconstruction loss, latent space regularization loss, path regularization loss, and R1 regularization loss, respectively. , , , and These represent the weights of each loss.

[0052] In this embodiment, the loss weights are preferably set as follows: =1、 =1、 =0.03、 =0.5 and =10. Training uses the Adam optimizer, with a batch size of 8 and a learning rate of 2×10. -4 .

[0053] S2.3 Map the font image labeled with target attributes to the generator through the trained font latent code mapper. Latent space, to obtain the corresponding latent encoding .

[0054] S3. Construct an attribute classifier and process the latent encoding. The corresponding font images are used for attribute prediction to obtain the score of each font sample on the target attribute, as follows:

[0055] S3.1 Construct a deep convolutional neural network as an attribute classifier to establish a mapping relationship between font images and attribute scores. In this embodiment, the font attribute classifier uses ResNet50 as the backbone network, extracts high-level structural features of the font image through residual connection structures, and establishes a mapping relationship between font images and attribute scores.

[0056] S3.2. Input the attribute-annotated font image dataset into the attribute classifier for training to obtain the trained attribute classifier. Specifically, the output layer design of the attribute classifier uses different modeling methods depending on the attribute type, as follows:

[0057] For discrete category attributes, the output layer uses a fully connected layer combined with the Softmax activation function, with cross-entropy as the optimization objective;

[0058] For prediction of continuous attributes (such as word weight), a linear output header is used and the mean squared error is used as the loss function;

[0059] When modeling multiple attributes simultaneously, a multi-task learning structure is adopted, sharing the feature extraction layer and setting multiple output branches.

[0060] S3.3 Input the font image corresponding to the latent encoding in S2 into the attribute classifier to obtain the attribute prediction score.

[0061] S4. Based on the latent encoded sample set with attribute scores, train a linear support vector machine in the latent space. Obtain the separating hyperplane by maximizing the attribute class margin. Determine the normal vector of the separating hyperplane as the latent direction corresponding to the attribute editing, as follows:

[0062] S4.1 Obtain the latent encoded sample set with attribute scores, and perform dimensionality reduction on the latent encoding using the t-SNE method. Dimensionality reduction visualization is performed to verify that the target attribute values ​​exhibit a continuous gradient change trend in the embedding space, thus obtaining the attribute distribution characteristics. For example... Figure 6 As shown, the word weight attribute value exhibits a clear and continuous gradient change trend in the two-dimensional embedding space, verifying the existence of a main change trend in the latent space that is highly correlated with the word weight attribute.

[0063] S4.2, in the original high dimension In the latent space, a linear support vector machine is trained based on latent encoded samples with attribute scores. The separating hyperplane is obtained by maximizing the class margin, and this hyperplane is represented as:

[0064]

[0065] in, This refers to the hyperplane normal vector, i.e., the latent direction of the attribute. This is a bias term.

[0066] S4.3 Calculate the normal vector of the separating hyperplane as the principal direction of change of the target attribute in the latent space to obtain the latent direction corresponding to the font attribute. .

[0067] S5. Calculate the latent encoding of multiple character samples under the same font. The response intensity at each level in the latent direction is statistically averaged to obtain the average response intensity at each level. Based on the average response intensity at each level, differentiated weighting coefficients are assigned to different levels. Get global editing intensity parameters Based on weighting coefficients and For latent encoding A hierarchical weighted editing strategy is constructed to apply differentiated editing magnitudes to each level of the potential encoding. Specifically:

[0068] S5.1 Obtain the latent encoding of multiple character samples under the same font. The response intensity of each latent component is calculated, and the response intensity of multiple samples is statistically averaged to obtain the average response intensity of each layer.

[0069]

[0070] in, Let be the average response intensity of the i-th layer. For the k-th character sample Layer latent vectors, The latent direction is obtained in S4, and N is the number of samples, which is 100 in this embodiment;

[0071] S5.2. Based on the average response intensity of each layer, design differentiated weighting coefficients for different levels of the latent coding. In this approach, higher weights are assigned to levels with higher response intensity to enhance stroke shape variations; lower weights are applied to structural layers to maintain overall proportional stability; and disturbances are moderately suppressed for levels with lower response intensity to reduce ineffective intervention. For example... Figure 7 As shown, the word weight latent direction exhibits higher response strength in the middle layer and some higher layers, while the response is relatively weaker in the lower-level structural control layer. Accordingly, the weights of the five layers with the strongest response are reset to 1, the weights of the three layers with the weakest response are reset to 0.1, and the weights of the remaining layers are reset to 0.

[0072] S5.3 Obtain global editing intensity parameters and weighting coefficients for each level The product of the weighting coefficient and the global editing intensity is applied to the latent encoding of each layer as the layer editing amplitude, and a linear shift is performed along the latent direction to obtain the edited latent encoding:

[0073]

[0074] in, This is the edited latent vector of the i-th layer. Given the original latent vector of layer i, globally edit the intensity parameters. Used to control the overall magnitude of change and determine the strength of attribute adjustment, i.e. The larger the value, the more significant the change in character weight; the weight coefficient of the i-th layer. The distribution of control disturbances at different levels is achieved by adjusting... The value of can be used to control the area where attributes are expressed.

[0075] S6. Use an image of a Chinese character font of any style. Input to the font latent code mapper to extract the initial latent code. The final latent code is obtained through iterative optimization. Based on the potential direction and hierarchical weight coefficients, for Perform hierarchical weighted editing to obtain the edited potential code. ; through the generator Generate Chinese character font images with edited attributes The specific implementation process of this step is as follows:

[0076] S6.1. Take an image of a Chinese character font of any style. Input to the font latent code mapper to extract the initial latent code. ;

[0077] S6.2, with initial latent encoding To optimize the starting point, a reconstructed image is generated using a StyleGAN2-based generator. To minimize the input image With reconstructed images The difference between them is the optimization objective of the latent encoding, while constraining the latent encoding not to deviate from the effective latent distribution of the generator, specifically:

[0078]

[0079] In the formula w + L represents the optimized latent encoding. gt Represents the input Chinese character font image; L opt The joint optimization objective is as follows:

[0080]

[0081] In the formula The reconstruction loss is expressed in L2 norm form and is used to measure the direct difference between the generated image and the target font image in pixel space. The perceptual loss is calculated based on features extracted from a pre-trained VGG network and is used to constrain the semantic consistency of the font in terms of structural layout, stroke relationships, and overall style. The p-norm regularization loss is used to explicitly constrain the update magnitude of the latent vectors at each layer, so that the optimized latent encoding remains within a reasonable manifold neighborhood of the generator's latent space. By measuring the L2 distance between the current latent vector and the initial predicted latent vector layer by layer, the overall offset of the latent encoding is penalized, which effectively prevents excessive offset of single-layer or local latent vectors during the optimization process and avoids font structure destruction or style semantic distortion. This represents the edge-preserving loss, which improves the clarity of stroke contours by constraining the consistency between the generated image and the target image in the gradient domain. , , and These represent the weights of the reconstruction loss, perception loss, p-norm regularization loss, and edge preservation loss, respectively. Preferably, in this embodiment, the weights of each loss are set as follows: =10, =1, =0.001, =0.5, and the optimization iteration rounds are set to 300 rounds. After a fixed number of iterations, the final latent code w is obtained. + .

[0082] S6.3. Based on the latent direction and hierarchical weight coefficients, for w + Perform hierarchical weighted editing to obtain the edited potential code. :

[0083]

[0084] By adjusting the value of the global editing intensity parameter α, the font weight attribute can be continuously, smoothly, and controllably adjusted. For example, setting α=0.3 generates a font with slightly thicker strokes; setting α=0.7 generates a font with significantly thicker strokes; and setting α=1.0 generates a font with very heavy strokes.

[0085] S6.4, Using the generator to... Generate Chinese character font image after attribute editing I edit .

[0086] Following the process described in S1 to S6, by constructing labeled datasets with either a height-slimness attribute or a width-flatness attribute, and training the corresponding separating hyperplanes, the latent directions for the height-slimness attribute and the width-flatness attribute can be obtained, enabling continuous and controllable editing of the font's aspect ratio.

[0087] Based on the above method design, this embodiment of the invention also provides a Chinese character font attribute control system based on hyperplane modeling and latent code editing, including a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, it can implement any of the above method steps.

[0088] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0089] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0090] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0091] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0092] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0093] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A Chinese character font attribute control method based on hyperplane modeling and latent code editing, characterized in that, Includes the following steps: Based on the latent encoding of multiple font samples with attribute scores in the generator's hierarchical latent space, a linear support vector machine is trained to obtain a separating hyperplane, and the normal vector of the separating hyperplane is determined as the latent direction corresponding to the target attribute. Calculate the response intensity of the latent encoding of multiple character samples under the same font at each level in the latent direction, and obtain the average response intensity of each level by statistical averaging. Based on the average response intensity of each level, assign differentiated editing weights to different levels. Obtain the potential encoding of the Chinese character font image to be edited in the generator's hierarchical latent space; Based on the latent direction and the differentiated editing weight, the latent code to be edited is subjected to hierarchical weighted editing to obtain the edited latent code; The edited latent encoding is input into the generator to generate a Chinese character font image with edited attributes.

2. The Chinese character font attribute control method based on hyperplane modeling and latent code editing according to claim 1, characterized in that: The hierarchical weighted editing specifically involves applying the product of the differentiated editing weight and the global editing intensity parameter to each level of potential encoding as the editing amplitude, and then performing a linear displacement along the potential direction.

3. The Chinese character font attribute control method based on hyperplane modeling and latent code editing according to claim 1, characterized in that: The average response intensity of each layer is obtained by calculating the arithmetic mean of the absolute values ​​of the inner products of the corresponding layer latent vectors and the latent directions of multiple character samples under the same font.

4. The Chinese character font attribute control method based on hyperplane modeling and latent code editing according to claim 1, characterized in that: The allocation rule for the differentiated editing weight is as follows: the editing weight is positively correlated with the average response intensity of each level, the editing weight is less than that of the non-structure control level for the structure control level, and the editing weight is no greater than the preset upper limit for the level with the average response intensity lower than the preset threshold.

5. The Chinese character font attribute control method based on hyperplane modeling and latent code editing according to claim 1, characterized in that: The process of obtaining the potential encoding of the Chinese character font image to be edited includes: inputting the Chinese character font image to be edited into a font latent code mapper to extract the initial potential encoding, and obtaining the potential encoding to be edited through iterative optimization starting from the initial potential encoding.

6. The Chinese character font attribute control method based on hyperplane modeling and latent code editing according to claim 5, characterized in that: The font latent code mapper is a multi-scale convolutional coding network, consisting of a convolutional backbone and a skip mapping module. It is used to extract font features at different resolutions and concatenate them to form a complete latent code that matches the number of generator layers.

7. The method for controlling Chinese character font attributes based on hyperplane modeling and latent code editing according to claim 1, characterized in that: The generator is a StyleGAN2 generation network adapted to the characteristics of Chinese character fonts, including independent latent vectors allocated to each resolution layer to form W + Hierarchical latent space, using weight demodulation instead of adaptive instance normalization.

8. The Chinese character font attribute control method based on hyperplane modeling and latent code editing according to claim 1, characterized in that: The separating hyperplane satisfies the following relationship: Where d is the normal vector of the separating hyperplane, corresponding to the latent direction, and b is the bias term. For potential encoding.

9. The Chinese character font attribute control method based on superplane modeling and latent code editing according to claim 1, characterized in that: The multiple font samples with attribute scores are obtained by introducing a font personality dimension to establish a hierarchical representation system of font visual attributes, labeling the font samples with continuous attribute levels on the target attribute dimension, and using an attribute classifier to predict the attributes of the labeled font images to obtain attribute scores; the font personality dimension includes visual attribute representations at the structural level, stroke level, and detail level.

10. A Chinese character font attribute control system, characterized by comprising: The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method of any one of claims 1 to 9.