A font library generation method, device, equipment and storage medium

By extracting quantifiable font features and using font knowledge graphs for multi-level evaluation, the problem of character quality consistency in the generation of ultra-large font libraries is solved, and the rapid and efficient generation and quality control of ultra-large font libraries are achieved.

CN122200672BActive Publication Date: 2026-07-21BEIJING HANYI KEYIN INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING HANYI KEYIN INFORMATION TECH
Filing Date
2026-05-13
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing methods for generating ultra-large character sets rely on manual design or AI-assisted generation, which cannot achieve consistent character quality assessment. This results in a large gap between the generated results and industry standards, making it difficult to meet the requirements of ultra-large character sets for consistent character quality.

Method used

By generating initial glyph images based on sample glyphs, extracting quantifiable font features, and using a pre-built font knowledge graph for multi-level evaluation, including single-character integrity, family consistency, and cluster consistency evaluation, correction instructions are generated for optimization, and finally converted into a font library file.

Benefits of technology

It achieves consistent character quality across ultra-large font libraries, shortens the production cycle, and ensures component uniformity and style consistency within the character cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122200672B_ABST
    Figure CN122200672B_ABST
Patent Text Reader

Abstract

The application provides a font library generation method, device and equipment and a storage medium, and relates to the field of font library generation. The method comprises the following steps: generating initial glyph images of a target character set based on sample glyphs; extracting quantifiable font features from the initial glyph images, and performing multi-level evaluation based on a pre-constructed font knowledge graph. The multi-level evaluation at least comprises cluster consistency evaluation on glyphs belonging to the same character cluster. Correction instructions are generated according to the evaluation results, and the correction instructions are fed back to the initial glyph image generation stage for optimization. The initial glyph images that pass the multi-level evaluation are converted into font library files. By extracting quantifiable font features and performing multi-level evaluation based on the pre-constructed knowledge graph, the problem of uncontrolled deformation of the same component in different characters can be found and corrected at the cluster level. In combination with the evaluation results, correction instructions are generated and fed back for optimization, so that the final font library meets the requirements of character quality consistency of the super-large font library.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of font generation, specifically to a font generation method, apparatus, device, and storage medium. Background Technology

[0002] Current methods for generating ultra-large font libraries primarily rely on manual design or early AI-assisted generation. Manual design requires designers to draw each character individually, resulting in a production cycle of months or even years for a single ultra-large font library, leading to high costs and difficulty in ensuring stylistic consistency. While large-model-based font generation methods can achieve batch generation of single-character images, they lack the ability to automatically and comprehensively evaluate the quality of the generated characters. This results in significant discrepancies between the generated results and industry standards, failing to meet the consistent character quality requirements of ultra-large font libraries. Summary of the Invention

[0003] In view of the above problems, embodiments of this application provide a font library generation method, apparatus, device and storage medium, which overcomes or at least partially solves the problem that the prior art is unable to meet the character quality consistency requirements of ultra-large font libraries.

[0004] The first aspect of this application provides a font library generation method, which includes: generating an initial glyph image of a target character set based on sample glyphs; extracting quantifiable font features from the initial glyph image and performing multi-level evaluation based on a pre-built font knowledge graph, wherein the multi-level evaluation includes at least a cluster consistency evaluation of glyphs belonging to the same character cluster; generating correction instructions based on the evaluation results and feeding the correction instructions back to the initial glyph image generation stage for generation optimization; and converting the initial glyph image that has passed the multi-level evaluation into a font library file.

[0005] In this embodiment, through an intelligent closed-loop architecture of generation, evaluation, and feedback, only a small number of samples are needed to drive the batch generation of the target character set. At the same time, by extracting quantifiable font features and performing multi-level evaluation based on a pre-built knowledge graph, especially by performing cluster consistency analysis on the glyphs of the same character cluster, it is possible to discover and correct the problem of uncontrolled deformation of the same component in different characters at the cluster level. Then, combined with the evaluation results, correction instructions are generated and feedback is fed back for optimization, ensuring that the final font library meets the requirements of ultra-large font library for character quality consistency and significantly shortening the production cycle of ultra-large font library.

[0006] In one alternative approach, the quantifiable font features include at least one or more of the following: stroke skeleton features, contour geometric features, component structure features, and topological features. Stroke skeleton features include the number of strokes, the type and number of intersections; contour geometric features include the character frame ratio, centroid coordinates, and the mean and variance of stroke thickness; component structure features include the component spatial relationship matrix, the angle of the main stroke, and the tightness of the central area; and topological features include the number of closed regions.

[0007] In one alternative approach, quantifiable font features are extracted from the initial glyph image, including: stroke skeleton extraction, contour vectorization, and contour point set generation. Stroke skeleton extraction includes: extracting the skeleton point set, stroke segmentation information, and intersection information of the glyph; calculating the number of strokes, intersection type and quantity, main stroke angle, and central tightness based on the skeleton point set, stroke segmentation information, and intersection information. Contour vectorization includes: converting the glyph image into vector contours to obtain a closed contour set and component segmentation information; the closed contour set is composed of Bézier curve segments, and the outer and inner contour types are labeled; based on the closed contour set and component segmentation information, a contour subset corresponding to each component is obtained. Contour point set generation includes: discretely sampling from the Bézier curves of the closed contour set to generate a contour point set. Based on the contour point set, closed contour set, and contour subset, the character frame ratio, centroid coordinates, mean and variance of stroke thickness, number of closed regions, and component spatial relationship matrix are calculated.

[0008] This embodiment deconstructs the character shape into structured data such as skeleton point sets and contour subsets, and calculates font features such as the number of strokes, character frame ratio, and component spatial relationship matrix based on this data, providing a quantifiable data foundation for subsequent multi-level evaluation. By segmenting components to obtain the contour subsets corresponding to each component, cross-character comparisons of the same components in different characters can be performed, providing key support for cluster consistency analysis.

[0009] In one optional approach, the pre-built font knowledge graph includes general character structure rules, specific font style rules, and character cluster association rules. The general character structure rules store the basic structural rules for different character systems, used to constrain the number of strokes, the number of closed regions, and the component spatial relationship matrix. Specific font style rules are bound to the target font style, storing the font's style characteristics, used to constrain the mean and variance of stroke thickness, the angle of the main stroke, the tightness of the central area, the character frame proportion, and the center of gravity coordinates. The character cluster association rules cluster characters according to radicals, structural types, and visual similarity, storing the hierarchical relationships among members within each character cluster. This is used to constrain the cross-character cluster consistency of the component spatial relationship matrix. The cluster consistency constraints include: statistical parameters of the spatial relationship matrices of each glyph component within the same character cluster, serving as a criterion for determining component proportion consistency; and a morphological similarity threshold for the contour subsets of the same component within the same character cluster, serving as a criterion for determining component morphological consistency.

[0010] In this embodiment, a font knowledge graph containing general text structure rules, specific font style rules, and character cluster association rules is constructed to systematically and structurally store professional font design knowledge as executable quantitative standards. The general text structure rules and specific font style rules provide single-character-level rationality judgment criteria for features such as stroke count, character frame ratio, and component spatial relationship matrix. The character cluster association rules, by storing statistical parameters of the component spatial relationship matrix within a cluster and morphological similarity thresholds for subsets of identical component outlines, provide a quantitative judgment benchmark for cross-character component consistency analysis. This enables the statistical discovery and correction of uncontrolled deformation of the same component in different characters, fundamentally ensuring component uniformity and style consistency across a massive font library.

[0011] In one optional approach, multi-level evaluation includes single-character integrity evaluation, family consistency evaluation, and cluster consistency evaluation. Single-character integrity evaluation, based on general character structure rules and specific font style rules, examines quantifiable font features item by item to assess the physical structural rationality and basic style conformity of individual glyphs. Family consistency evaluation, based on specific font style rules and a benchmark glyph database that has passed evaluation, compares the quantifiable font features of each stroke type in the current glyph with the statistical distribution of features of the same stroke type in the benchmark glyph database to assess style consistency. The benchmark glyph database that has passed evaluation refers to the set of glyphs that has passed evaluation. Cluster consistency evaluation, based on character cluster association rules, performs cluster consistency analysis on all glyphs within the same character cluster. The analysis includes: statistically analyzing the component spatial relationship matrix of all glyphs within the character cluster, calculating the deviation of the current glyph from statistical parameters, and marking glyphs with deviations exceeding a preset multiple as having abnormal component proportions; calculating the morphological similarity of the contour subsets of the same components of all glyphs within the character cluster, and marking glyphs with similarity below a morphological similarity threshold as having abnormal component morphology.

[0012] In this embodiment, the single-character integrity assessment checks the physical structure and basic style of each glyph item by item based on general character structure rules and specific font style rules to ensure the rationality of the structure of a single character. The family consistency assessment compares the glyphs with the benchmark glyph database that has passed the assessment to ensure the style uniformity of the entire font library in terms of stroke shape and visual weight. The cluster consistency assessment is based on character cluster association rules and performs cross-character joint analysis on glyphs with the same radical or structural type from two dimensions: component proportion and component shape. This can accurately identify and locate the deformation out-of-control problem of the same component in different characters. The three levels work together to discover defects in single characters, monitor the overall style, and ensure component-level consistency from a statistical perspective, thus achieving professional and automated control of the quality of ultra-large font libraries.

[0013] In one alternative approach, the correction instructions include one or more of the following: local geometric correction instructions, style optimization instructions, and cluster alignment correction instructions. Local geometric correction instructions are generated based on problem features and their deviation values ​​identified in single-character integrity assessments, and are used to correct structural defects in individual glyphs. Style optimization instructions are generated based on style deviations identified in family consistency assessments, and are used to correct style deviations between the glyph and existing glyphs. Cluster alignment correction instructions are generated based on component proportion anomalies or component shape anomalies identified in cluster consistency assessments, and are used to correct statistical deviations between the glyph and its associated character cluster.

[0014] In one alternative approach, converting the initial glyph image, which has undergone multi-level evaluation, into a font file includes: converting the initial glyph image into a vector contour, performing curve smoothing and node simplification optimization to obtain an optimized vector contour; and then encapsulating the optimized vector contour according to the target font format to generate a font file.

[0015] A second aspect of this application provides a font generation apparatus, comprising: a generation module for generating initial glyph images of a target character set based on sample glyphs; an extraction and evaluation module for extracting quantifiable font features from the initial glyph images and performing multi-level evaluation based on a pre-built font knowledge graph, wherein the multi-level evaluation includes at least a cluster consistency evaluation of glyphs belonging to the same character cluster; a correction module for generating correction instructions based on the evaluation results and feeding the correction instructions back to the initial glyph image generation stage for generation optimization; and an output module for converting the initial glyph images that have passed the multi-level evaluation into font files.

[0016] The font generation device provided in this embodiment can ensure that the final font meets the requirements of character quality consistency for ultra-large font libraries, and significantly shortens the production cycle of ultra-large font libraries.

[0017] A third aspect of this application provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the character library generation method provided in the first aspect of this application.

[0018] The computer equipment provided in this embodiment can ensure that the final font library meets the requirements of character quality consistency for ultra-large font libraries, and significantly shorten the production cycle of ultra-large font libraries.

[0019] A fourth aspect of this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the character library generation method provided in the first aspect of this application.

[0020] The computer-readable storage medium provided in this embodiment can ensure that the final font library meets the requirements of character quality consistency for ultra-large font libraries, and significantly shorten the production cycle of ultra-large font libraries.

[0021] The above description is merely an overview of the technical solutions of the embodiments of this application. In order to better understand the technical means of the embodiments of this application and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of this application more obvious and understandable, specific implementation methods of this application are described below. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart illustrating a font generation method provided for some embodiments of this application.

[0024] Figure 2 This is a schematic diagram of the structure of a font generation device provided in some embodiments of this application. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.

[0027] The terms "comprising" and "having," and any variations thereof, used in the specification, claims, and drawings of this application are intended to cover without excluding other meanings. The words "a" or "an" do not exclude the presence of multiples.

[0028] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of the phrase "embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0029] Furthermore, the terms "first," "second," etc., in the specification and claims of this application or in the aforementioned drawings are used to distinguish different objects rather than to describe a specific order, and may explicitly or implicitly include one or more of the features.

[0030] In the description of this application, unless otherwise stated, "multiple" means two or more (including two), and similarly, "multiple groups" means two or more (including two groups).

[0031] In the description of this application, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linkage" should be interpreted broadly. For example, "connection" or "linkage" in mechanical structures can refer to a physical connection, such as a fixed connection, for example, a connection fixed by fasteners, such as a connection fixed by screws, bolts, or other fasteners; a physical connection can also be a detachable connection, such as a snap-fit ​​or interlocking connection; a physical connection can also be an integral connection, such as a connection formed by welding, bonding, or integral molding. In circuit structures, "connection" or "linkage" can refer not only to a physical connection but also to an electrical connection or a signal connection. For example, it can be a direct connection, i.e., a physical connection, or an indirect connection through at least one intermediate component, as long as the circuit is connected; it can also refer to the internal connection of two components. Signal connection can refer not only to signal connection through a circuit but also to signal connection through a media, such as radio waves. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0032] This application provides a method for generating a font library. Please refer to [link / reference]. Figure 1 , Figure 1 This is a flowchart illustrating a font generation method provided for some embodiments of this application.

[0033] like Figure 1 As shown, the font generation method provided in this application embodiment includes the following steps 101 to 104: Step 101: Generate an initial glyph image of the target character set based on the sample glyphs.

[0034] Among them, sample glyphs refer to a small number of reference character images provided by the user. For example, there can be 50-300 reference character images, which are used to define the style characteristics of the target font.

[0035] The target character set refers to the encoding set of all characters that need to be generated. For example, the target character set can be the 88,000 Chinese characters defined by the GB18030-2022 standard. Using this target character set as the production task list, the corresponding initial glyph image is generated for each character one by one.

[0036] Step 102: Extract quantifiable font features from the initial glyph image and perform multi-level evaluation based on the pre-built font knowledge graph.

[0037] Among them, multi-level evaluation includes at least cluster consistency evaluation of glyphs belonging to the same character cluster.

[0038] Quantifiable font features include at least one or more of the following: stroke skeleton features, outline geometric features, component structural features, and topological features. Stroke skeleton features may include the number of strokes and the type and number of intersections. Outline geometric features may include the character frame ratio, centroid coordinates, and mean and variance of stroke thickness. Component structural features may include the component spatial relationship matrix, the angle of the main stroke, and the tightness of the central area. Topological features include the number of closed regions.

[0039] Step 103: Generate correction instructions based on the evaluation results, and feed the correction instructions back to the initial glyph image generation stage for generation optimization.

[0040] In practical applications, the optimized initial glyph image can be continuously evaluated at multiple levels until the evaluation is passed.

[0041] Step 104: Convert the initial glyph images that have passed the multi-level evaluation into a font file.

[0042] Specifically, converting the initial glyph image that has passed multi-level evaluation into a font file includes: converting the initial glyph image that has passed multi-level evaluation into a vector outline, performing curve smoothing and node simplification optimization to obtain an optimized vector outline, and encapsulating the optimized vector outline according to the target font format to generate a font library. For example, it can be encapsulated in TrueType or OpenType format to output the final font library.

[0043] In this embodiment, an intelligent closed-loop architecture of generation, evaluation, and feedback allows for the batch generation of target character sets with only a small number of samples. Simultaneously, by extracting quantifiable font features and performing multi-level evaluation based on a pre-built knowledge graph, particularly by performing cluster consistency analysis on glyphs of the same character cluster, it is possible to identify and correct uncontrolled deformation issues of the same component in different characters at the cluster level. Furthermore, by combining the evaluation results to generate correction instructions and provide feedback for optimization, the final font library meets the character quality consistency requirements of ultra-large font libraries and significantly shortens the production cycle of ultra-large font libraries.

[0044] In some embodiments, quantifiable font features are extracted from the initial glyph image, including: stroke skeleton extraction, contour vectorization, and contour point set generation. Stroke skeleton extraction includes: extracting the skeleton point set, stroke segmentation information, and intersection information of the glyph; calculating the number of strokes, the type and number of intersection points, the angle of the main stroke, and the tightness of the central area based on the skeleton point set, stroke segmentation information, and intersection information. Contour vectorization includes: converting the glyph image into vector contours to obtain a closed contour set and component segmentation information; the closed contour set is composed of Bézier curve segments, and the outer and inner contour types are marked; based on the closed contour set and component segmentation information, a contour subset corresponding to each component is obtained. Contour point set generation includes: discretely sampling from the Bézier curves of the closed contour set to generate a contour point set. Based on the contour point set, the closed contour set, and the contour subset, the character frame ratio, centroid coordinates, the mean and variance of stroke thickness, the number of closed regions, and the component spatial relationship matrix are calculated.

[0045] Specifically, the number of strokes can be obtained by counting the number of independent strokes in the stroke segmentation information. The type and number of intersection points can be obtained by identifying the intersection positions and connection relationships of strokes based on the intersection point information. The angle of the main stroke can be obtained by identifying the main stroke based on the stroke segmentation information, extracting the skeleton point set corresponding to the main stroke, performing straight line fitting, and calculating the angle between the fitted line and the horizontal line. The tightness of the central grid can be obtained by dividing the character frame into nine equal parts based on the skeleton point set, counting the number of skeleton points falling into the central grid, and calculating the skeleton point density within the central grid. The proportion of the character frame can be obtained by calculating the minimum bounding rectangle of all points based on the contour point set, obtaining the width W and height H, and calculating the proportion W / H. The centroid coordinates can be obtained by calculating the average coordinates of all points based on the contour point set as the centroid. The mean and variance of stroke thickness can be obtained by drawing a straight line along the normal direction of the skeleton points based on the skeleton point set and the contour point set, intersecting the left and right contours determined by the contour point set, calculating the distance of the intersection point as the stroke thickness, and counting the mean and variance of all measurement points. The number of closed regions can be determined based on the set of closed contours and the contour type, where each outer contour and its contained inner contours constitute a closed region. The component spatial relationship matrix can be obtained by obtaining a subset of contour points corresponding to each component based on component segmentation information, calculating the bounding rectangle of each component, and establishing the proportional relationship or relative position matrix between components.

[0046] This embodiment deconstructs the character shape into structured data such as skeleton point sets and contour subsets, and calculates font features such as the number of strokes, character frame ratio, and component spatial relationship matrix based on this data, providing a quantifiable data foundation for subsequent multi-level evaluation. By segmenting components to obtain the contour subsets corresponding to each component, cross-character comparisons of the same components in different characters can be performed, providing key support for cluster consistency analysis.

[0047] In some embodiments, the pre-built font knowledge graph includes general text structure rules, font-specific style rules, and character cluster association rules.

[0048] The general character structure rules store the basic structural rules for different character systems, which are used to constrain the number of strokes, the number of closed regions, and the component spatial relationship matrix. For example, the number of strokes specifies the range of strokes for different characters; the number of closed regions specifies the number of closed regions for different characters; and the component spatial relationship matrix specifies the component ratio range for common structures such as left-right structures and top-bottom structures.

[0049] Specific font style rules are bound to the target font style, storing the font's style characteristics to constrain the mean and variance of stroke thickness, main stroke angle, central grid tightness, character frame ratio, and center of gravity coordinates. For example, the mean and variance of stroke thickness specifies the proportional range of thin horizontal strokes and thick vertical strokes, and the allowable fluctuation range of stroke thickness; the main stroke angle specifies the range of the main stroke's tilt angle; the central grid tightness specifies the allowable range of the density of points in the central grid skeleton; the character frame ratio specifies the allowable range of the character's width-to-height ratio; and the center of gravity coordinates specify the allowable range of center of gravity offset.

[0050] The character cluster association rule clusters characters according to radical, structural type, and visual similarity to form character clusters. It stores the subordinate relationships of each member within a character cluster and uses them to constrain the cross-character cluster consistency of the component spatial relationship matrix. The cluster consistency constraints include: statistical parameters of the spatial relationship matrix of each glyph component within the same character cluster, which serve as the criterion for determining the consistency of component proportions; and morphological similarity thresholds of the contour subsets of the same component within the same character cluster, which serve as the criterion for determining the consistency of component morphology.

[0051] In this embodiment, a font knowledge graph containing general text structure rules, specific font style rules, and character cluster association rules is constructed to systematically and structurally store professional font design knowledge as executable quantitative standards. The general text structure rules and specific font style rules provide single-character-level rationality judgment criteria for features such as stroke count, character frame ratio, and component spatial relationship matrix. The character cluster association rules, by storing statistical parameters of the component spatial relationship matrix within a cluster and morphological similarity thresholds for subsets of identical component outlines, provide a quantitative judgment benchmark for cross-character component consistency analysis. This enables the statistical discovery and correction of uncontrolled deformation of the same component in different characters, fundamentally ensuring component uniformity and style consistency across a massive font library.

[0052] In some embodiments, the multi-level assessment includes single-word integrity assessment, family consistency assessment, and cluster consistency assessment.

[0053] The single-character integrity assessment is based on general character structure rules and specific font style rules. It checks each quantifiable font feature and evaluates the rationality of the physical structure of a single character and its conformity with the basic style.

[0054] Family consistency assessment is based on specific font style rules and a benchmark glyph database that has passed the assessment. It compares the quantifiable font features of each stroke type in the current glyph with the statistical distribution of features of the same stroke type in the benchmark glyph database to assess style consistency. The benchmark glyph database that has passed the assessment refers to the set of glyphs that have passed the assessment; the same stroke type is, for example, all of them are dots, horizontals, verticals, etc.

[0055] Cluster consistency evaluation is based on character cluster association rules. It conducts cluster consistency analysis on all glyphs within the same character cluster. The analysis includes: statistically analyzing the component spatial relationship matrix of all glyphs within the character cluster, calculating the deviation degree of the current glyph relative to the statistical parameters, and marking the glyphs with a deviation exceeding a preset multiple as having abnormal component ratios. For example, the statistical analysis can be calculating the mean and standard deviation, and the preset multiple can be 2 standard deviations. Calculate the morphological similarity of the contour subsets of the same components of all glyphs within the character cluster, and mark the glyphs with a similarity lower than the morphological similarity threshold as having abnormal component morphologies.

[0056] For example, based on the character cluster association rules, incorporate the current glyph into its所属 character cluster, such as the "three dots of water" cluster, extract the component spatial relationship matrix of all glyphs within the character cluster that have passed the evaluation, such as the proportion of the width of the left component, and calculate its mean μ and standard deviation σ. Compare the proportion of the width of the left component of the current glyph with the range of μ±2σ. If it exceeds this range, mark it as having an abnormal component ratio.

[0057] For example, extract the contour subset of a certain component in the current glyph, such as the contour subset of "氵", and calculate the morphological similarity with the contour subsets of the same components in other glyphs within the character cluster. The similarity calculation method can use contour matching algorithms, such as Hausdorff distance, shape context, etc. If the similarity is lower than the morphological similarity threshold, for example, the morphological similarity threshold is 0.85, mark it as having an abnormal component morphology.

[0058] When both dimensions of component ratio and component morphology pass the evaluation, the cluster consistency evaluation is considered passed; if either dimension is abnormal, it triggers the generation of corresponding correction instructions.

[0059] In this embodiment, the evaluation of single - character integrity逐项检查 the physical structure and basic style of each glyph based on general text structure rules and specific font style rules to ensure the rationality of the structure of a single character; the family consistency evaluation statistically compares with the database of benchmark glyphs that have passed the evaluation to ensure the style unity of the entire font library in terms of stroke morphology and visual weight; the cluster consistency evaluation is based on character cluster association rules, and conducts cross - character joint analysis on glyphs of the same radical or structure type from two dimensions of component ratio and component morphology, which can accurately identify and locate the problem of out - of - control deformation of the same component in different characters. The three levels cooperate with each other, which can not only detect single - character defects, but also monitor the overall style, and more importantly, ensure component - level consistency from a statistical perspective, achieving professional and automated control of the quality of a super - large font library.

[0060] In some embodiments, the correction instructions include one or more of local geometry correction instructions, style optimization instructions, and cluster alignment correction instructions. The local geometry correction instructions are generated based on the problem features and their deviation values found by the single-character integrity assessment. Specifically, they are generated based on the problem features such as stroke breakage and center-of-gravity offset marked by the single-character integrity assessment and their deviation values. For example, the local geometry correction instructions include the problem position coordinates, the current values of the problem features, and geometric adjustment parameters, and are used to correct the structural defects of a single glyph. The style optimization instructions are generated based on the style deviations found by the family consistency assessment. Specifically, they are generated based on the feature deviations such as stroke thickness and main stroke angle found by the family consistency assessment. For example, the style optimization instructions include the reference feature values, the current feature values, and style adjustment parameters, and are used to correct the style deviations between the glyphs and the generated font library. The cluster alignment correction instructions are generated based on the abnormal component ratios or abnormal component morphologies found by the cluster consistency assessment. For example, the cluster alignment correction instructions include cluster statistical parameters or morphological similarity thresholds, abnormal individual deviation values, and alignment adjustment parameters, and are used to correct the statistical deviations between the glyphs and their affiliated clusters.

[0061] In practical applications, the correction instructions can be directly input into the stage of generating the initial glyph image, and the generation parameters are automatically adjusted for iterative optimization. The correction instructions can also push the glyphs that fail the assessment and the correction instructions to the manual proofreading interface, and after receiving manual confirmation or adjustment, they are returned to the stage of generating the initial glyph image. The pre-constructed font knowledge graph has the ability to be dynamically updated and can automatically adjust the threshold parameters and cluster division criteria in the rule library according to the results of the cluster consistency assessment and the feedback of manual correction.

[0062] The following uses a specific example to illustrate a font library generation method provided by this application.

[0063] In this embodiment, taking the generation of an extra-large font library in the "KaiTi" style that complies with the GB18030-2022 standard as an example, the specific implementation process of this method is described in detail. The target character set contains more than 88,000 Chinese characters, and the sample glyphs are 300 reference glyphs in the "KaiTi" style provided by the user.

[0064] First, preprocess and analyze the features of the input 300 "KaiTi" sample glyphs. Through operations such as stroke skeleton extraction and contour vectorization, the quantifiable font features of each sample character are extracted, including the number of strokes, the ratio of the literal box, the center-of-gravity coordinates, the mean and variance of stroke thickness, the component spatial relationship matrix, the main stroke angle, the tightness of the middle palace, and the number of closed regions. Statistically aggregate the feature values of the 300 samples to generate a 32-dimensional style vector representing the "KaiTi" style.

[0065] Automatically disassemble and recognize the basic stroke forms and their variants from sample glyphs, and establish a stroke component library containing 15 basic strokes such as horizontal, vertical, left-falling, right-falling, etc. and 214 common radicals. Each component records its typical outline, skeleton features and combination rules.

[0066] Taking the target character set (88,000 character encoding of GB18030 standard), 32-dimensional style vector and stroke component library as inputs, batch generation is carried out using Tongyi large model and glyph control adapter architecture. The glyph control adapter converts the character encoding, style vector and component parameters into an input format understandable by the large model, and controls the generated results to conform to the design logic of "KaiTi". The generation module supports multi-GPU parallel processing, divides the character set into multiple subsets for simultaneous generation, and outputs the initial glyph images of each character. For example, the initial glyph image is in bitmap format.

[0067] Extract quantifiable font features from the generated initial glyph images, specifically including: Stroke skeleton extraction: For example, a hybrid method combining parallel thinning algorithm and convolutional neural network is used to extract the skeleton point set, stroke segmentation information and intersection information of the character "海".

[0068] Calculate the following quantifiable font features based on the above information: Number of strokes: Count the stroke segmentation information to get that the character "海" has 10 strokes; Type and number of intersection points: Based on the intersection information, 3 stroke intersection points are identified; Angle of the main stroke: Identify the main stroke as the long horizontal stroke in the right part "每" of "海", and perform linear fitting on the skeleton points of this horizontal stroke to get an angle of 3°; Tightness of the central part: Divide the literal box into nine equal parts, and count the density of skeleton points in the central grid as 0.14.

[0069] Contour vectorization: Convert the image of the character "海" into a vector contour, and obtain a closed contour set and component segmentation information. The closed contour set is composed of Bezier curve segments, marking the types of outer and inner contours. Based on the closed contour set and component segmentation information, the character "海" is segmented into the left component "氵" and the right component "每", and the corresponding contour subsets are obtained respectively.

[0070] Contour Point Set Generation: Discretely sample from the Bezier curves of the closed contour set to generate a contour point set. Calculate the following quantifiable font features based on the contour point set, closed contour set, and contour subsets: Literal Box Ratio: Calculate the minimum bounding rectangle based on the contour point set, with width W = 122 pixels, height H = 100 pixels, and ratio W / H = 1.22; Centroid Coordinates: Calculate the centroid based on the contour point set, with coordinates (61, 51); Mean and Variance of Stroke Thickness: Measure the stroke width along the normal direction of the skeleton points, obtaining a mean of 2.6 pixels and a variance of 0.32; Number of Closed Regions: Based on the closed contour set, it is determined that the character "海" has no closed regions, with a value of 0; Component Spatial Relationship Matrix: Based on the contour subsets of the left component "氵" and the right component "每", calculate the bounding rectangles respectively. The width ratio of the left component is 29%, and the width ratio of the right component is 71%.

[0071] Load the pre-built "KaiTi" font knowledge graph, which contains the following rules: (1) General Character Structure Rules: Rule G1: The standard number of strokes of the Chinese character "海" is 10; Rule G2: Among left-right structure Chinese characters, the width ratio of the left component usually ranges from 25% to 35%; Rule G3: The number of closed regions should be consistent with the standard structure of the character. "海" should have no closed regions. (2) Specific Font Style Rules ("KaiTi" Style): Rule S1: The mean stroke thickness should be between 2.4 and 2.8 pixels, and the variance should not exceed 0.4; Rule S2: The angle of the main stroke should be within the range of 0° ± 5°; Rule S3: The tightness of the middle palace should be between 0.12 and 0.18; Rule S4: The literal box ratio should be between 1.18 and 1.26; Rule S5: The vertical coordinate of the centroid should be within ±4 pixels of the vertical center of the literal box. (3) Character Cluster Association Rules: Cluster the characters by radicals to form a "three-point water" cluster, which includes characters such as "江, 河, 湖, 海, 洋". Cluster Rule Storage: Statistical Parameters: Based on the characters "江, 河, 湖" that have passed the evaluation, the mean width ratio of the left component "氵" is 28%, and the standard deviation is 2%; Morphological Similarity Threshold: The similarity threshold of the contour subsets of the same component "氵" is 0.85.

[0072] Extract the quantifiable font features and compare them with the loaded pre-built knowledge graph of the "KaiTi" font. Perform three levels of evaluation: (1) Single-character integrity evaluation: Check the number of strokes: The feature value 10 is consistent with rule G1, passed; Check the number of enclosed areas: The feature value 0 is consistent with rule G3, passed; Check the mean and variance of stroke thickness: The mean 2.6 is within the range of rule S1, and the variance 0.32 does not exceed 0.4, passed; Check the main stroke angle: 3° is within the range of rule S2, passed; Check the tightness of the middle palace: 0.14 is within the range of rule S3, passed; Check the ratio of the literal box: 1.22 is within the range of rule S4, passed; Check the center of gravity coordinates: The vertical coordinate is 51, the vertical center of the literal box is 50, and the deviation of 1 pixel is within the range of rule S5, passed. Conclusion of single-character integrity evaluation: The basic structure and style of this glyph meet the requirements and have no serious defects. (2) Family consistency evaluation: Compare the features of the character "海" with the benchmark glyph database that has passed the evaluation (including other characters in the generated "KaiTi" font library). The mean thickness of the "dot" stroke in the benchmark database is 2.5 pixels, and the thickness value of the "dot" stroke in the character "海" is 2.6, with a deviation of 0.1 within the allowable range; The mean main stroke angle of the "horizontal" stroke in the database is 3.5°, and the main stroke angle of the character "海" is 3°, with a deviation of 0.5° passed. The comparison of other features is within the allowable deviation range. Conclusion of family consistency evaluation: The style of this glyph is basically consistent with the generated font family. (3) Cluster consistency evaluation: Incorporate the character "海" into the "三点水" cluster for analysis: Analysis of component ratio consistency: In the cluster, there are already characters such as "江", "河", and "湖". The width ratios of the left component "氵" are 27%, 29%, and 28% respectively, with a mean of 28% and a standard deviation of 2%. The width ratio of the left component of the character "海" is 29%. Compared with the cluster mean of 28%, the deviation is 0.5 times the standard deviation (less than the preset 2-fold standard deviation threshold) and is not marked as an abnormal component ratio. Analysis of component form consistency: Extract the contour subset of the "氵" component in the character "海" and calculate the morphological similarity with the contour subsets of the "氵" components in the characters "江", "河", and "湖" in the cluster. The average similarity is 0.91, which is higher than the preset threshold of 0.85 and is not marked as an abnormal component form. Conclusion of cluster consistency evaluation: This character performs normally in its所属 cluster and has no abnormalities. Since all three levels of evaluation have passed, the system generates a "passed" conclusion for the character "海" and does not need to generate correction instructions.

[0073] Suppose in another scenario, the width ratio of the left component "氵" of the generated character "海" is 35%, exceeding the cluster average by 28% and reaching 3.5 times the standard deviation (exceeding the 2 - standard - deviation threshold). The cluster consistency evaluation will mark it as an abnormal component ratio and generate the following cluster alignment correction instructions. Instruction type: Cluster alignment correction instruction; Character: "海"; Abnormal type: Component ratio abnormality; Problem component: "氵"; Current value: 35%; Mean in cluster statistical parameters is 28%; Standard deviation is 2%; Deviation multiple: 3.5; Suggested operation: Shrink the horizontal scale of the left component "氵" to 80% of its original width. The alignment adjustment parameters are as follows: Component ID: "氵"; Scaling ratio: 0.8; Scaling direction: Horizontal.

[0074] This correction instruction is fed back to the generation stage. The glyph control adapter in the generation module analyzes the adjustment parameters in the instruction and adjusts the generation weight of the left component "氵" to 80% of its original horizontal scale when regenerating the character "海". After 1 - 2 rounds of iterative optimization, the width ratio of the left component of the newly generated character "海" drops to 29% and passes the cluster consistency evaluation.

[0075] For the glyph images of the character "海" that pass through the multi - level evaluation, perform the output steps: (1) Vectorization: Convert the "海" character image into a vector contour, perform curve smoothing and node reduction optimization on the contour, reduce the average number of nodes from more than 1500 to 350, and at the same time retain the pen - tip features of "KaiTi". (2) Encapsulation: Encapsulate the optimized vector contour in the OpenType font format to generate a part of the font library file. Traverse all 88,000 characters of the GB18030 standard, and repeat the above process for each glyph that passes the evaluation. Finally, generate a complete "KaiTi" super - large font library file, and the super - large font library file can be in the OTF format.

[0076] Through the above process, this embodiment realizes the automatic generation from 300 sample glyphs to an 88,000 - character super - large font library. The generated "KaiTi" font library passes professional quality inspection, and the stroke standardization, weight uniformity, and component consistency all meet industry standards. Especially through the cluster consistency evaluation mechanism, it effectively ensures the morphological unity of common components such as "氵", "辶", and "木" in different characters, and solves the problem of out - of - control component deformation commonly found in existing generation methods.

[0077] In some embodiments, when a font library generation method is applied to the generation of ancient - book reproduction fonts, the sample glyphs are scanned ancient - book sample character images, and the method further includes: Before extracting quantifiable font features, perform background color removal, skew correction, and blur enhancement on the sample images. In the evaluation step, add verification of block - printed features, including knife - carving trace detection, weathering effect consistency evaluation, stone - slab texture simulation, etc.

[0078] Another embodiment of this application also provides a font library generation device, refer to Figure 2 , Figure 2 This is a schematic diagram of the structure of a font generation device provided in some embodiments of this application.

[0079] like Figure 2 As shown, the font generation device 2 of this embodiment may include: a generation module 21, an extraction and evaluation module 22, a correction module 23, and an output module 24.

[0080] The system comprises the following modules: Generation module 21 generates initial glyph images of the target character set based on sample glyphs; Extraction and evaluation module 22 extracts quantifiable font features from the initial glyph images and performs multi-level evaluation based on a pre-built font knowledge graph, including at least a cluster consistency evaluation of glyphs belonging to the same character cluster; Correction module 23 generates correction instructions based on the evaluation results and feeds these instructions back to the initial glyph image generation stage for optimization; and Output module 24 converts the initial glyph images that have passed the multi-level evaluation into font library files.

[0081] Another embodiment of this application provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the font generation method provided in the foregoing embodiments of this application.

[0082] Another embodiment of this application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the font generation method provided in the foregoing embodiments of this application.

[0083] The font generation apparatus, computer equipment, and computer-readable storage medium provided in the embodiments of this application can be used to perform the above-described... Figure 1 The technical solutions of the method embodiments shown are similar in principle and in effect, and will not be described again here.

[0084] Those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments, combinations of features from different embodiments are intended to be within the scope of this application and form different embodiments. For example, in the claims, any of the claimed embodiments can be used in any combination.

[0085] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for generating a font library, characterized in that, The method includes: Generate initial glyph images of the target character set based on sample glyphs; Quantifiable font features are extracted from the initial glyph image, and multi-level evaluation is performed based on a pre-constructed font knowledge graph. The multi-level evaluation includes at least a cluster consistency evaluation of glyphs belonging to the same character cluster. Based on the evaluation results, correction instructions are generated, and these correction instructions are fed back to the stage of generating the initial glyph image for optimization. The initial glyph image that has passed the multi-level evaluation is converted into a font file; The pre-built font knowledge graph includes general text structure rules, specific font style rules, and character cluster association rules; The general text structure rules store the basic structural rules of different text systems, which are used to constrain the number of strokes, the number of closed regions, and the component spatial relationship matrix. The specific font style rules are bound to the target font style, storing the style features of the font, and are used to constrain the mean and variance of stroke thickness, the angle of the main stroke, the tightness of the central part, the proportion of the character frame, and the coordinates of the center of gravity. The character cluster association rule clusters characters according to radical, structural type, and visual similarity to form character clusters, and stores the membership relationship of each member within the character cluster. This is used to constrain the cross-character cluster consistency of the component spatial relationship matrix. The cluster consistency constraint includes: statistical parameters of the component spatial relationship matrix of each glyph within the same character cluster, as a criterion for determining component proportion consistency; and morphological similarity threshold of the contour subset of the same component within the same character cluster, as a criterion for determining component morphological consistency. The multi-level assessment includes single-word integrity assessment, family consistency assessment, and cluster consistency assessment. The single-character integrity assessment is based on the general character structure rules and the specific font style rules. It checks the quantifiable font features item by item and evaluates the physical structure rationality and basic style conformity of a single character. The family consistency assessment is based on the specific font style rules and the benchmark glyph database that has passed the assessment. It compares the quantifiable font features of each stroke type in the current glyph with the feature statistical distribution of the same stroke type in the benchmark glyph database to assess style consistency. The benchmark glyph database that has passed the assessment refers to the set of glyphs that has passed the assessment. The cluster consistency assessment is based on the character cluster association rules and performs cluster consistency analysis on all glyphs within the same character cluster. The analysis includes: performing statistical analysis on the component spatial relationship matrix of all glyphs within the character cluster, calculating the deviation of the current glyph from the statistical parameters, and marking glyphs with deviations exceeding a preset multiple as having abnormal component proportions; and calculating the morphological similarity of the contour subsets of the same components of all glyphs within the character cluster, and marking glyphs with similarity scores below the morphological similarity threshold as having abnormal component morphology.

2. The method according to claim 1, characterized in that, The quantifiable font features include at least one or more of the following: stroke skeleton features, outline geometric features, component structure features, and topological features; wherein, the stroke skeleton features include the number of strokes, the type and number of intersection points; the outline geometric features include the character frame ratio, the center of gravity coordinates, the mean and variance of stroke thickness; the component structure features include the component spatial relationship matrix, the angle of the main stroke, and the tightness of the central area; and the topological features include the number of closed regions.

3. The method according to claim 2, characterized in that, The extraction of quantifiable font features from the initial glyph image includes: stroke skeleton extraction, contour vectorization, and contour point set generation; The stroke skeleton extraction includes: extracting the skeleton point set, stroke segmentation information, and intersection information of the character shape; and calculating the number of strokes, the type and number of intersections, the angle of the main stroke, and the tightness of the central part based on the skeleton point set, the stroke segmentation information, and the intersection information. The contour vectorization includes: converting the glyph image into a vector contour to obtain a closed contour set and component segmentation information; the closed contour set is composed of Bézier curve segments and the outer contour and inner contour types are marked; based on the closed contour set and the component segmentation information, a contour subset corresponding to each component is obtained; The contour point set generation includes: discrete sampling from the Bézier curve of the closed contour set to generate the contour point set; Based on the set of contour points, the set of closed contours, and the subset of contours, calculate the character frame ratio, centroid coordinates, mean and variance of stroke thickness, number of closed regions, and component spatial relationship matrix.

4. The method according to claim 1, characterized in that, The correction instructions include one or more of the following: local geometry correction instructions, style tuning instructions, and cluster alignment correction instructions; The local geometric correction instruction is generated based on the problem features and deviation values ​​found in the single character integrity assessment, and is used to correct the structural defects of a single character. The style tuning instruction is generated based on the style deviations found in the family consistency assessment and is used to correct the style deviations between the glyphs and the generated glyphs. The cluster alignment correction instruction is generated based on the abnormal component ratio or abnormal component shape found in the cluster consistency assessment, and is used to correct the statistical deviation between the glyph and the character cluster to which it belongs.

5. The method according to claim 1, characterized in that, The process of converting the initial glyph image, which has undergone the multi-level evaluation, into a font file includes: The initial glyph image, which has undergone multi-level evaluation, is converted into a vector contour. Curve smoothing and node simplification optimization are then performed to obtain the optimized vector contour. The optimized vector outline is encapsulated according to the target font format to generate a font library file.

6. A font generation device, characterized in that, The device includes: The generation module is used to generate initial glyph images of the target character set based on sample glyphs; An extraction and evaluation module is used to extract quantifiable font features from the initial glyph image and perform multi-level evaluation based on a pre-built font knowledge graph. The multi-level evaluation includes at least a cluster consistency evaluation of glyphs belonging to the same character cluster. The pre-built font knowledge graph includes general text structure rules, specific font style rules, and character cluster association rules; The general text structure rules store the basic structural rules of different text systems, which are used to constrain the number of strokes, the number of closed regions, and the component spatial relationship matrix. The specific font style rules are bound to the target font style, storing the style features of the font, and are used to constrain the mean and variance of stroke thickness, the angle of the main stroke, the tightness of the central part, the proportion of the character frame, and the coordinates of the center of gravity. The character cluster association rule clusters characters according to radical, structural type, and visual similarity to form character clusters, and stores the membership relationship of each member within the character cluster. This is used to constrain the cross-character cluster consistency of the component spatial relationship matrix. The cluster consistency constraint includes: statistical parameters of the component spatial relationship matrix of each character within the same character cluster, as a criterion for determining component proportion consistency; and morphological similarity threshold of the contour subset of the same component within the same character cluster, as a criterion for determining component morphological consistency. The multi-level assessment includes single-word integrity assessment, family consistency assessment, and cluster consistency assessment. The single-character integrity assessment is based on the general character structure rules and the specific font style rules. It checks the quantifiable font features item by item and evaluates the physical structure rationality and basic style conformity of a single character. The family consistency assessment is based on the specific font style rules and the benchmark glyph database that has passed the assessment. It compares the quantifiable font features of each stroke type in the current glyph with the feature statistical distribution of the same stroke type in the benchmark glyph database to assess style consistency. The benchmark glyph database that has passed the assessment refers to the set of glyphs that has passed the assessment. The cluster consistency assessment is based on the character cluster association rules and performs cluster consistency analysis on all glyphs within the same character cluster. The analysis includes: performing statistical analysis on the component spatial relationship matrix of all glyphs within the character cluster, calculating the deviation of the current glyph from the statistical parameters, and marking glyphs with deviations exceeding a preset multiple as having abnormal component proportions; and calculating the morphological similarity of the contour subsets of the same components of all glyphs within the character cluster, and marking glyphs with similarity scores below the morphological similarity threshold as having abnormal component morphology. The correction module is used to generate correction instructions based on the evaluation results and feed the correction instructions back to the initial glyph image generation stage for generation optimization. The output module is used to convert the initial glyph image obtained through the multi-level evaluation into a font file.

7. A computer device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the font generation method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the character library generation method as described in any one of claims 1 to 5.