Font generation method and device, equipment and storage medium
By extracting and fusing features at multiple scales and combining them with an adaptive transformation function, the problems of local geometric distortion and global style inconsistency in font generation are solved, achieving efficient adaptation and accurate generation for different writing styles.
Patent Information
- Application Number
- CN202511058345.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-18
AI Technical Summary
Existing font generation technologies in the healthcare and fintech fields suffer from insufficient feature decoupling, lack of cross-scale interaction, and weak dynamic adaptability, resulting in local geometric distortion and global style inconsistency, making it difficult to meet the specific needs of these fields.
Multiple deep separable convolutional branches with different dilation rates are used to extract features at multiple scales. Features are dynamically fused through a spatial attention mechanism, a hierarchical feature pyramid is constructed, and cross-scale fusion is performed through multi-scale residual connections. Style embedding vectors are extracted by combining convolutional neural networks and transformed into distribution parameters in the latent space. New font images are generated using an adaptive transformation function.
It significantly improves the local geometric accuracy and global style consistency of font generation, enhances the model's ability to generalize to different writing styles, and meets the needs of multi-scale feature interaction and dynamic adaptability in the fields of healthcare and fintech.
Smart Images

Figure CN120976360A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology and can be applied to fields such as medical health and financial technology. In particular, it relates to a font generation method, apparatus, device and storage medium. Background Technology
[0002] In recent years, deep learning-based font generation technology has demonstrated significant application potential in the healthcare and fintech sectors. In healthcare, font generation is widely used in standardized medical reports, drug label design, and patient record management, requiring a high degree of consistency between local stroke details (such as the standardized writing of drug names) and overall layout style (such as table alignment). In fintech, font generation technology is applied to automated financial statements, contract template generation, and transaction voucher design, ensuring the accurate representation of numbers, symbols, and text to avoid compliance risks caused by font distortion. However, existing technologies still face significant challenges in addressing these needs.
[0003] While traditional Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs) can generate diverse font styles, their generation process is often accompanied by pattern collapse and training instability, leading to local geometric distortions (such as broken strokes in the "mg" unit in medical reports) or global style inconsistencies (such as mixing title and body text fonts in financial contracts). Furthermore, component-based methods rely on predefined font libraries, making it difficult to meet the dynamic needs of the healthcare field for special symbols (such as ECG waveform annotations) or the fintech field for multilingual mixed typesetting (such as combinations of Chinese and English monetary figures).
[0004] In healthcare settings, fonts need to simultaneously optimize both micro-level details (such as the clarity of drug dosage numbers) and macro-level layout (such as column alignment in medical record tables). In fintech, transaction documents need to balance the accuracy of character strokes (such as tamper-proof design for monetary amounts) and the overall visual consistency of the document (such as consistent headers and footers in multi-page contracts). However, existing flow models generally employ a single-scale coupling layer, which cannot effectively model multi-scale feature interactions at the stroke, component, and overall style levels, leading to an imbalance between local details and global structure in the generated results.
[0005] In the healthcare field, font generation requires dynamic adjustment of font styles based on patient diagnostic report types (such as CT image annotations and surgical records); in the fintech field, font parameters need to be flexibly switched according to contract types (such as loan agreements and investment prospectuses) or regulatory requirements (such as financial statement formats in different countries). However, the transformation function of traditional normalized flow is fixed, making it difficult to dynamically modulate feature fusion strategies through input style embedding vectors (such as reference fonts), thus limiting the responsiveness of the generative model to domain-specific needs.
[0006] Normalizing flows offer a novel approach to addressing the aforementioned problems due to their accurate probability density modeling capabilities and reversible transformation properties. Their reversibility ensures the traceability of the generation process (e.g., version control of medical documents), while accurate probabilistic modeling helps quantify the uncertainty of font generation (e.g., anti-counterfeiting design of financial instruments). However, the application of existing normalizing flow models in healthcare and fintech still faces the following bottlenecks:
[0007] 1. Lack of explicit decoupling mechanisms for stroke details, component structure, and global style in medical and financial documents;
[0008] 2. It is difficult to drive real-time adjustment of model parameters through style embedding vectors;
[0009] 3. Existing methods are mostly designed for general font generation and do not fully consider the domain-specific needs of healthcare and fintech scenarios. Summary of the Invention
[0010] The purpose of this invention is to provide a font generation method, apparatus, device, and storage medium, aiming to solve problems such as insufficient feature decoupling, lack of cross-scale interaction, and weak dynamic adaptability in existing font generation technologies.
[0011] In a first aspect, embodiments of the present invention provide a font generation method, including:
[0012] Obtain a font image and a reference font, and preprocess the font image to obtain a feature map;
[0013] Multiple scale features are extracted from the feature map by using multiple depthwise separable convolutional branches with different dilation rates;
[0014] Multiple scale features are dynamically fused using a spatial attention mechanism to obtain fused features;
[0015] A hierarchical feature pyramid is constructed, and low-level high-resolution features and high-level semantic features in the feature map are fused across scales through multi-scale residual connections to obtain constrained features.
[0016] Style embedding vectors are obtained by extracting styles from the reference font using a convolutional neural network.
[0017] The style embedding vector is transformed into distribution parameters within the latent space;
[0018] The fused features, the distribution parameters, and the constraint features are input into an adaptive transformation function to obtain hierarchical features;
[0019] The hierarchical features are converted into a new font image using a decoder.
[0020] Secondly, embodiments of the present invention provide a font generation apparatus, comprising:
[0021] An acquisition unit is used to acquire a font image and a reference font, and to preprocess the font image to obtain a feature map;
[0022] The feature extraction unit is used to extract multiple scale features from the feature map through multiple depthwise separable convolutional branches with different dilation rates;
[0023] The feature fusion unit is used to dynamically fuse features at multiple scales through a spatial attention mechanism to obtain fused features;
[0024] The construction unit is used to construct a hierarchical feature pyramid. Through multi-scale residual connections, the low-level high-resolution features and high-level semantic features in the feature map are fused across scales to obtain constrained features.
[0025] The style extraction unit is used to extract style from the reference font through a convolutional neural network to obtain a style embedding vector;
[0026] A transformation unit is used to transform the style embedding vector into distribution parameters within the latent space;
[0027] A transformation unit is used to input the fused features, the distribution parameters, and the constraint features into an adaptive transformation function to obtain hierarchical features;
[0028] A conversion unit is used to convert the hierarchical features into a new font image via a decoder.
[0029] Thirdly, embodiments of the present invention provide a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the font generation method described in the first aspect.
[0030] Fourthly, embodiments of the present invention also provide a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program that, when executed by a processor, implements the font generation method described in the first aspect.
[0031] This invention discloses a font generation method, apparatus, device, and storage medium, comprising: acquiring a font image and a reference font, and preprocessing the font image to obtain a feature map; extracting multiple scale features from the feature map through multiple depthwise separable convolutional branches with different dilation rates; dynamically fusing the multiple scale features through a spatial attention mechanism to obtain fused features; constructing a hierarchical feature pyramid, and fusing low-level high-resolution features and high-level semantic features in the feature map across scales through multi-scale residual connections to obtain constrained features; extracting style from the reference font through a convolutional neural network to obtain a style embedding vector; transforming the style embedding vector into distribution parameters in a latent space; inputting the fused features, the distribution parameters, and the constrained features into an adaptive transformation function to obtain hierarchical features; and converting the hierarchical features into a new font image through a decoder. This invention explicitly separates stroke-level, component-level, and style-level features through parallel dilated convolutional branches, and dynamically fuses multi-scale information using a spatial attention mechanism, effectively solving the problems of local geometric distortion and global style inconsistency. Meanwhile, an adaptive transformation function is introduced, enabling the coupling layer parameters to dynamically adjust according to the input font style, significantly improving the model's generalization ability to different writing styles. Furthermore, multi-scale residual connections aggregate low-level high-resolution features with high-level semantic features, mitigating the detail loss problem in deep network processing. This invention also provides a font generation device, a computer-readable storage medium, and a computer device, all possessing the aforementioned beneficial effects, which will not be elaborated further here. Attached Figure Description
[0032] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0033] Figure 1 This is a schematic diagram of an application environment for a font generation method according to an embodiment of the present invention;
[0034] Figure 2 This is a flowchart illustrating the font generation method.
[0035] Figure 3 This is another flowchart illustrating the font generation method;
[0036] Figure 4 This is a schematic diagram of the sub-processes of the font generation method;
[0037] Figure 5 A schematic block diagram of a font generation device;
[0038] Figure 6 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0039] Figure 7 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0041] It should be understood that, when used in this specification and the appended claims, the terms “comprising” and “including” indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more of its features, integrals, steps, operations, elements, components and / or collections thereof.
[0042] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0043] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0044] The font generation method provided in this embodiment of the invention can be applied to, for example, Figure 1In this application environment, the client communicates with the server via a network. The server can obtain a font image and a reference font from the client, and preprocess the font image to obtain a feature map. Multiple scale features are extracted from the feature map using multiple depthwise separable convolutional branches with different dilation rates. These multiple scale features are dynamically fused using a spatial attention mechanism to obtain fused features. A hierarchical feature pyramid is constructed, and low-level high-resolution features and high-level semantic features within the feature map are fused across scales using multi-scale residual connections to obtain constrained features. Style is extracted from the reference font using a convolutional neural network to obtain a style embedding vector. The style embedding vector is transformed into distribution parameters in the latent space. The fused features, distribution parameters, and constrained features are input into an adaptive transformation function to obtain hierarchical features. The hierarchical features are converted into a new font image using a decoder. The new font image is fed back to the client. In this invention, stroke-level, component-level, and style-level features are explicitly separated by parallel dilated convolutional branches, and multi-scale information is dynamically fused using a spatial attention mechanism, effectively solving the problems of local geometric distortion and global style inconsistency. Simultaneously, an adaptive transformation function is introduced, enabling the coupling layer parameters to dynamically adjust according to the input font style, significantly improving the model's generalization ability to different writing styles. Furthermore, multi-scale residual connections are used to aggregate low-level high-resolution features with high-level semantic features, mitigating the detail loss problem in deep network processing. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, AR devices, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0045] Please see Figures 2-4 This embodiment provides a font generation method, including:
[0046] S101: Obtain a font image and a reference font, and preprocess the font image to obtain a feature map;
[0047] In healthcare scenarios, font images can be obtained from electronic medical record systems and medical science poster design platforms, such as the fonts for diagnostic text in medical records and the title fonts for science posters. Reference fonts can also be obtained from standard font libraries in medical brand promotional materials and professional medical journals, such as the exclusive fonts for hospital names and the standardized fonts for medical papers. In the fintech field, font images can be obtained from financial contract template systems and financial APP interface design documents, such as the fonts for electronic contract terms and the title fonts for APP wealth management modules. Reference fonts can also be obtained from financial institution brand VI materials and standard font libraries for financial statements, such as the unified fonts for bank annual reports and the compliant fonts for securities APPs. The obtained font images are first normalized to a uniform 256×256 pixels to adapt to the model input; then grayscale conversion is performed to eliminate color interference and highlight the font structure; finally, Gaussian filtering is used for noise suppression, such as removing noise introduced during medical image scanning or financial interface screenshots, to obtain clear and regular feature maps, laying the foundation for subsequent multi-scale coupling layer processing, ensuring that font generation accurately adapts to scenario requirements in healthcare information display and fintech business documents.
[0048] S102: Extract multiple scale features from the feature map using multiple depthwise separable convolutional branches with different dilation rates;
[0049] In this embodiment, the coupling layer is the core component of the normalized flow, and it has been extended at multiple scales. Given the input feature map of the l-th layer... H represents the height of the image, W represents the width, and C represents the number of channels. The processing flow is as follows:
[0050] Multi-scale features are extracted using K depthwise separable convolutional branches with different dilation rates:
[0051]
[0052] in, The expansion rate is d k 3×3 depthwise separable convolution The output feature of the k-th branch. The level increase of the dilated convolution ensures that the receptive field covers multi-level features from stroke details (d=1) to the whole character (d=8).
[0053] The parallel branch structure enables the synchronous extraction of features at multiple scales, avoiding the omission or confusion of feature information at a single scale. It provides a hierarchical and complete foundation for subsequent feature fusion, and is adapted to the accuracy and standardization requirements of font generation in the medical and financial fields.
[0054] In some embodiments, extracting multiple scale features from a feature map using multiple depthwise separable convolutional branches with different dilation rates includes:
[0055] Extract stroke detail features from the feature map through the depthwise separable convolution branch with the first dilation rate;
[0056] Extract component features from the feature map through the depthwise separable convolution branch with the second dilation rate;
[0057] Extract character local structure features from the feature map through the depthwise separable convolution branch with the third dilation rate;
[0058] Extract overall style features from the feature map through the depthwise separable convolution branch with the fourth dilation rate.
[0059] For the input feature map, the depthwise separable convolution branch with the first dilation rate is mainly responsible for extracting stroke detail features. This branch uses 3×3 depthwise separable convolution. Since the dilation rate is small, the receptive field is concentrated in the local area, and it can accurately capture the subtle features such as the thickness change of the strokes and the stroke direction in the font. For example, when processing the regular script character '永', it can clearly extract the staccato details at the end of the right-falling stroke and the starting and ending features of the horizontal stroke. The depthwise separable convolution branch with the second dilation rate focuses on extracting component features. At this dilation rate, the receptive field of the convolution is moderately enlarged, and it can cover the components composed of multiple strokes in the font, such as the independent components like 'horizontal fold hook' and 'left-falling and right-falling strokes' in the character '永'. This branch can effectively capture the morphological structure of these components and the connection relationship between components, ensuring the integrity of component features. The depthwise separable convolution branch with the third dilation rate is used to extract character local structure features. As the dilation rate further increases, the receptive field coverage is wider, and it can focus on the combination method and local layout between different components in the character. Taking the character '永' as an example, this branch can extract local structure information such as the relative position between the dot and the horizontal stroke and the interpenetrating layout between the left-falling and right-falling strokes, providing a structural basis for subsequent feature fusion and style transformation. The depthwise separable convolution branch with the fourth dilation rate is dedicated to extracting overall style features. This largest dilation rate enables the receptive field of the convolution to cover the entire character image, thereby capturing the overall style attributes of the font, such as the stretching and smoothness of regular script and the dignity and regularity of Song typeface, and at the same time, it can also extract the overall typesetting tendency of the character, such as macroscopic features like the density and the center of gravity position.
[0060] Through these four depthwise separable convolution branches with different dilation rates, the model can comprehensively extract multi-scale features from the feature map from stroke details to overall style, laying a foundation for subsequent attention fusion and conditional transformation, and further ensuring the performance of the generated font in terms of detail accuracy and style consistency.
[0061] It should be noted that the first dilation rate, the second dilation rate, the third dilation rate, and the fourth dilation rate can be actually adjusted according to the type of the input font, the complexity, and the requirements of the generation task. Preferably, extracting multiple-scale features from the feature map through multiple depthwise separable convolution branches with different dilation rates includes:
[0062] Extracting stroke detail features from the feature map through a depthwise separable convolution branch with a dilation rate of 1;
[0063] Extracting component features from the feature map through a depthwise separable convolution branch with a dilation rate of 2;
[0064] Extracting character local structure features from the feature map through a depthwise separable convolution branch with a dilation rate of 4;
[0065] Extracting overall style features from the feature map through a depthwise separable convolution branch with a dilation rate of 8.
[0066] In this embodiment, the branches with different dilation rates can accurately match the feature requirements of different levels of the font: the branch with a dilation rate of 1 focuses on the local area and can carefully capture details such as the thickness change of the stroke and the direction of the pen tip, such as the connected details of the font in medical prescriptions or the stroke pauses in financial contracts; the branch with a dilation rate of 2 can effectively extract the combination relationship between components, such as the connection features between radicals like "氵" and "讠" and the main body in Chinese characters, ensuring the accuracy of the glyph structure of professional terms in medical and health documents; the branch with a dilation rate of 4 can capture the arrangement of the local structure of characters, such as the local typesetting features of the combination of numbers and words in financial statements; the branch with a dilation rate of 8 can grasp the overall style and ensure the consistency of the overall style of medical publicity materials or financial brand fonts.
[0067] At the same time, depthwise separable convolution can reduce the parameter calculation amount when extracting features, improve the processing efficiency, and is suitable for scenarios with real-time requirements such as real-time font generation in online consultations in the medical and health field and dynamic font updates on fintech platforms. The parallel branch structure not only realizes the synchronous extraction of multi-scale features but also avoids the problem of feature confusion under a single scale, providing a clear and complete information basis for subsequent feature fusion and font generation.
[0068] In the medical and health field, this solution takes the generation of standardized drug label fonts as an example. Extracting features through the parallel dilation convolution branches of the multi-scale coupling layer:
[0069] The depthwise separable convolution branch with a dilation rate of 1: Capturing stroke detail features, such as the pause at the end of the stroke of the lowercase "mg" character in the drug name, to ensure the standard writing of the dosage unit;
[0070] Depthwise separable convolution branch with dilation rate 2: Extract component features, such as the hooked horizontal and vertical structure of the character "司" in the Chinese drug name "aspirin", and maintain the spatial relationship between components;
[0071] Depthwise separable convolution branch with dilation rate 4: Model local structural features of characters, such as the alignment of drug names and dosage numbers ("500mg"), to avoid reading obstacles caused by font deformation;
[0072] Depthwise separable convolution branch with dilation rate 8: Extract overall style features, such as the style consistency between the title text and the main text in medical labels (e.g., the title is in bold Song typeface and the main text is in standard Kai typeface).
[0073] In the fintech field, taking the generation of multilingual financial contract fonts as an example:
[0074] Depthwise separable convolution branch with dilation rate 1: Precisely extract the stroke turning details of English amount numbers (e.g., "$10,000.00") to ensure the clarity of symbols and numbers;
[0075] Depthwise separable convolution branch with dilation rate 2: Analyze Chinese character components (such as the structures of "万" and "圆") to maintain the integrity of components in multilingual mixed layout;
[0076] Depthwise separable convolution branch with dilation rate 4: Model the local structures of Chinese and English characters (such as the spacing between amount numbers and currency symbols) to avoid format disorders;
[0077] Depthwise separable convolution branch with dilation rate 8: Extract overall style features, such as the unified font style of the contract header and footer (such as the modern sense of sans-serif fonts), to ensure the visual consistency of the document.
[0078] Through the above multi-scale feature extraction mechanism, the model realizes the standardized writing of drug labels in the medical and health field, ensures the typesetting accuracy of multilingual financial documents in the fintech field, and simultaneously meets the collaborative optimization requirements of local details and global styles.
[0079] S103: Dynamically fuse multiple-scale features through a spatial attention mechanism to obtain fused features;
[0080] Dynamically fuse multiple-scale features through a spatial attention mechanism, and the obtained fused features include:
[0081] Successively perform concatenation, convolution, and activation function calculations on multiple-scale features to obtain an attention weight map;
[0082] Perform weighted summation on the attention weight map and multiple-scale features to obtain fused features.
[0083] The approach of dynamically fusing features at multiple scales through spatial attention mechanisms has significant advantages. In terms of fusion accuracy, first concatenating, convolving, and calculating activation functions on multi-scale features to obtain an attention weight map allows the model to automatically identify key features at different locations. For example, in healthcare scenarios, it can enhance stroke detail features for key diagnostic terms in medical record fonts and strengthen overall style features for the title area of promotional posters; in fintech scenarios, it can emphasize component features for core contract clauses to ensure structural clarity and emphasize style features for brand logo areas to enhance recognizability.
[0084] By using a weighted summation of weighted graphs and multi-scale features to obtain fused features, dynamic feature adaptation is achieved. This avoids feature redundancy or loss of key information caused by fixed fusion methods, and allows for flexible adjustment of the contribution of features at each scale according to specific content requirements. Furthermore, the entire process requires no manual intervention in the feature fusion strategy. While ensuring fusion efficiency, it provides a high-quality feature foundation for subsequent font generation that balances local detail accuracy with global style consistency, thus meeting the stringent requirements and scenario adaptability of font generation in fields such as healthcare and finance.
[0085] In some embodiments, multiple scale features are sequentially concatenated, convolved, and activation function calculated to obtain an attention weight map, including:
[0086] Dynamically fuse multi-scale features using spatial attention mechanisms:
[0087]
[0088] in This is the attention weight map, where σ is the sigmoid activation function; Conv 1×1 It is a 1×1 convolution.
[0089] In this embodiment, the stroke details, component features, local structural features, and overall style features extracted by depth-separable convolutional branches with dilation rates of 1, 2, 4, and 8 are concatenated in the channel dimension to form a high-dimensional feature tensor (e.g., input size of 512×512×64).
[0090] Next, a 1×1 convolutional layer is used to compress the number of feature channels after splicing to 1 / 4 (e.g., from 64 channels to 16 channels), which reduces computational complexity and extracts cross-scale correlation information.
[0091] Then, the Sigmoid activation function is applied to the dimensionality-reduced features to generate an attention weight map A∈[0,1]H×WA∈[0,1]H×W, where high-value regions correspond to key features (such as the intersection of strokes in the drug dosage unit "mg").
[0092] The 1×1 convolution realizes feature channel fusion without changing the spatial dimension, which not only compresses redundant information but also preserves the spatial position relationship of key features, and is suitable for scenarios with strict requirements for the font spatial structure such as medical prescriptions and financial contracts. The sigmoid activation function maps the weight values to the interval of 0-1, enabling the model to clearly distinguish the importance of features at each scale. For example, in the generation of medical diagnosis term fonts, it strengthens the weight of stroke details, and in the generation of financial brand fonts, it enhances the weight of the overall style.
[0093] In some embodiments, the attention weight map and multiple scale features are weighted and summed to obtain the fused feature, including:
[0094]
[0095] Where, is the attention weight map, is the scale feature, and ⊙ represents element-wise multiplication. This operation enables the model to adaptively strengthen important scale features.
[0096] In the field of healthcare, taking the generation of standardized drug dosage label fonts as an example, the fused feature extraction is achieved through the weighted sum of the attention weight map and multi-scale features:
[0097] The element-wise multiplication A⊙F is performed on the attention weight map and multi-scale features corresponding to stroke details, components, local structures, and overall styles respectively k , dynamically strengthening the features of key regions. For example, for the drug dosage unit "mg", the attention weight map will assign high weights to the hook stroke of the "m" character (local details of the d = 1 branch) and the horizontal fold structure of the "g" character (component features of the d = 2 branch), suppressing redundant background noise and ensuring the standard writing of dosage units.
[0098] Then, through weighted summation, different scale features are integrated to form the fused feature Ffusion. This operation strengthens the starting stroke turning point of the number "5" (local structure of the d = 4 branch) and the overall style feature (regular script pen tip of the d = 8 branch) in the character "500mg", avoiding the risk of dosage misreading caused by font deformation.
[0099] In the field of fintech, taking the generation of multi-language financial contract amount fonts as an example:
[0100] For the mixed layout of English currency symbols and the Chinese character "wan", the attention weight map will dynamically strengthen the diagonal stroke of the symbol (stroke details of the d = 1 branch) and the radical of the character "wan" (component features of the d = 2 branch), while suppressing irrelevant background interference.
[0101] Integrate features of different scales through element-wise multiplication and weighted summation. For example, in "¥10,000.00 Ten thousand yuan", strengthen the precise outline of the amount number "10,000.00" (local structure with d = 4 branches) and the sans-serif font style of the contract title (overall style with d = 8 branches), and ensure the alignment of the spacing between Chinese and English characters and visual consistency.
[0102] The above weighted summation operation dynamically adjusts the contribution degrees of multi-scale features through a spatial attention mechanism, solves the problems of the standardization of drug dosage writing in the field of medical health and the multi-language typesetting accuracy in the field of fintech, and at the same time improves the geometric accuracy and style diversity of the generated results.
[0103] S104: Construct a hierarchical feature pyramid, and perform cross-scale fusion of the low-level high-resolution features and high-level semantic features in the feature map through multi-scale residual connections to obtain constrained features;
[0104] Specifically, to maintain the cross-layer feature consistency, design multi-scale residual connections:
[0105]
[0106] Among them, Downsample uses average pooling with a stride of 2, and M = 3 represents the backtracking depth; h l-m represents the feature map of the l - m layer; r l represents the constrained feature.
[0107] In this embodiment, downsampling is performed through average pooling with a stride of 2, which can retain key structural information while compressing the feature dimension. When generating medical record fonts in the medical health scenario, it can stably transmit basic stroke features such as "horizontal and vertical". When generating contract fonts in the field of fintech, it can ensure the structural coherence of numbers such as "0 - 9". Secondly, the setting of the backtracking depth M = 3 enables the features of the current layer to establish associations with the features of the previous 3 layers, forming a cross-layer feature chain, and avoiding feature loss or distortion in the deep network. For example, in the generation of medical prescription fonts, it can ensure that the "艹" head component feature of the character "药" remains consistent from the shallow layer to the deep layer; when generating financial statement fonts, it can maintain the stroke features of the symbol "¥" from being weakened.
[0108] In the field of medical health, taking the generation of drug dosage label fonts as an example, this solution maintains cross-layer feature consistency through multi-scale residual connections:
[0109] In the encoder-decoder structure, set the backtracking depth M = 3, backtrack from the current layer l three layers forward (l - 1, l - 2, l - 3), and extract low-level high-resolution features (such as the stroke details of the character "mg" in the drug name). Adjust F through average pooling AvgPool with a stride of 2 s=2 Adjust F l-mThe resolution (e.g., reduced from 512×512 to 256×256) is adjusted to match the current layer feature F. l Alignment.
[0110] The three layers of features F are backtracked. l-1 F l-2 F l-3 With the current layer feature F l The weighted fusion formula is as follows:
[0111]
[0112] Among them, low-level features (such as F) l-3 Preserve the stroke transition details of the character "mg", and high-level features (such as F) l Model the overall layout structure of dosage units to ensure consistency between the generated results in local geometry (such as the hook of "mg") and global layout (such as the alignment of "500mg").
[0113] S105: Extract style from the reference font using a convolutional neural network to obtain a style embedding vector;
[0114] Specifically, obtain the image data of the reference font. The reference font is a font image containing the characteristics of the target style. It can be a Chinese character, number, letter or symbol image of a specific style, and the image format is a bitmap format that meets the requirements of subsequent processing.
[0115] Next, preprocessing operations are performed on the reference font image. The preprocessing includes: adjusting the image size to a preset pixel specification (e.g., 256×256 pixels) to fit the input size requirements of the convolutional neural network.
[0116] Then, the adjusted image is converted to grayscale, transforming the color image into a single-channel grayscale image to eliminate the interference of color information on style feature extraction;
[0117] Then, the grayscale image is subjected to noise suppression processing using the Gaussian filtering algorithm to remove noise such as spots and spikes in the image, resulting in a standardized input image;
[0118] The preprocessed input image is then fed into a pre-defined convolutional neural network. This network consists of a shallow convolutional module, a mid-level convolutional module, a deep convolutional module, pooling layers, and fully connected layers connected sequentially. Specifically: the shallow convolutional module contains 2-3 convolutional layers with 3×3 kernels, extracting stroke edge details of the reference font through 16-32 convolutional channels, such as the sharpness or roundness of the strokes, the thickness transitions of the strokes, and the shape of the turning points; the mid-level convolutional module contains 2-3 convolutional layers with 3×3 kernels, extracting component structural features of the reference font through 64-128 convolutional channels, such as the proportion of radicals to the main body, and the relationships between components. Spacing, overall structural compactness, etc.; Deep convolutional modules contain 2-3 convolutional layers, using 3×3 convolutional kernels, and extract the overall style attributes of the reference font through 256-512 convolutional channels, such as the uprightness of the font, whether it has decorative elements, and the overall visual style (such as solemn, lively, simple, etc.); Pooling layers use max pooling with a stride of 2 to reduce the dimensionality of the feature maps output by each convolutional module, reducing the amount of data while retaining key feature information; Fully connected layers contain 2-3 neurons, which integrate and map the pooled features, converting high-dimensional feature vectors into vectors of a preset dimension (such as 256-dimensional or 512-dimensional) through linear transformation.
[0119] Next, the output layer of the convolutional neural network outputs a style embedding vector. This style embedding vector can fully encode all style features of the reference font, from stroke details to component structure to overall style. It can be used for style transfer and style consistency control of the target font in the subsequent font generation process.
[0120] S106: Transform the style embedding vector into distribution parameters within the latent space;
[0121] In this embodiment, the potential space adopts a mixture of Gaussian distributions;
[0122]
[0123] Where, p Z (z) represents the probability density value of the mixture Gaussian distribution at point z, i.e., the probability density of sample z occurring; π j This represents the mixing coefficient of the j-th Gaussian component; μ represents the probability density function of the i-th Gaussian component. ij This represents the mean; Variance, μ ij and These are the distribution parameters.
[0124] Where the mixing coefficient π j The distribution parameters are predicted by the hypernetwork:
[0125] {π j ,μ ij ,σ ij} = HyperNet(e style )
[0126] e style The 256-dimensional style embedding vector extracted from the reference font using a CNN:
[0127] e style =CNN(h l+1 )
[0128] The mixing coefficients and parameters such as the mean and variance of the Gaussian components are not fixed, but are predicted and generated by the hypernetwork based on the style embedding vector of the reference font. This design creates a strong bond between the distribution parameters of the latent space and the style of the reference font: the hypernetwork can learn to transform the stroke features (such as the sharpness of the strokes) and structural features (such as the compactness of components) contained in the style embedding vector into corresponding distribution parameters. For example, if the style embedding vector shows that the reference font is "flowing cursive script", the hypernetwork can predict the "high variance" component (corresponding to the stroke extension) and the appropriate mixing coefficient (highlighting the weight of the "connected strokes" related components). This correlation ensures that the samples generated in the latent space (such as the latent vectors used for font generation) always revolve around the core features of the reference style, avoiding style shift.
[0129] S107: Input the fused features, the distribution parameters, and the constraint features into the adaptive transformation function to obtain hierarchical features;
[0130] Specifically, the fused features and constrained features are input into the adaptive transformation function to obtain hierarchical features, including:
[0131] The fused features and distribution parameters are input into the adaptive transformation function to generate the first scale parameters and the second scale parameters.
[0132] The hierarchical features are obtained by fusing the first scale parameter, the second scale parameter, and the constraint features.
[0133] The adaptive transformation function can dynamically adapt the first and second scale parameters generated by the fusion feature processing to the feature attributes. In the healthcare scenario, when generating electronic medical records, it can adjust the parameters for the font of diagnostic terms and strengthen the features related to stroke clarity. In the fintech field, when generating contract text, it can adapt the font of core clauses and optimize the component structure stability parameters.
[0134] Secondly, by fusing scale parameters with constraint features, the hierarchical features retain effective multi-scale information from the fused features (such as stroke details in medical prescriptions and the overall style of financial statements) while also being standardized by the constraint features, ensuring consistency across layers (such as uniform overall font layout in medical records and compliant font structure in contracts). This mechanism avoids the separation between local and global aspects in feature transformation, generating hierarchical features that combine detailed accuracy with overall consistency, providing a reliable foundation for subsequent font generation and meeting the rigorous font generation and scenario adaptation requirements of the medical and financial fields.
[0135] Furthermore, the fused features and distribution parameters are input into the adaptive transformation function to generate the first scale parameters and the second scale parameters, including:
[0136] The first scale parameter is generated according to the following formula:
[0137]
[0138] The second-scale parameters are generated according to the following formula:
[0139]
[0140] The AdaIN operation is defined as follows:
[0141]
[0142] in, γ represents the first scale parameter; MLP represents a two-layer perceptron; s β s γ t and β t All represent learnable parameters; F l Indicates fusion characteristics; Let F represent the second scale parameter; μ(F) represent the mean of the fusion feature F; and σ(F) represent the standard deviation of the fusion feature F, where the distribution parameters are μ(F) and σ(F).
[0143] The AdaIN operation standardizes the fused feature F by applying the mean μ(F) and standard deviation σ(F), and adjusts them by combining learnable parameters γ, β and γ, β. This effectively separates style information from content information in the features. In healthcare scenarios, when generating electronic medical records, it can separate font styles such as "SimSun" from glyph content such as "diagnostic terms," ensuring that the font style is consistent and the content is clear across different medical records. In the fintech field, when generating contract text, it can distinguish between standardized styles such as "Bold" and core content such as "clause text," ensuring the compliance of the contract font while not losing the text structure features.
[0144] Secondly, the two-layer perceptron (MLP) processing of the AdaIN output maps features to scale parameters adapted to subsequent computations. This preserves key information in the fused features (such as stroke details in medical prescriptions and component structures in financial statements) while dynamically adjusting through learnable parameters, ensuring that the first and second scale parameters accurately adapt to different scenario requirements. This mechanism provides a flexible and controllable parameter foundation for subsequent hierarchical feature generation, meeting the stylistic stability and content accuracy requirements of font generation in the medical and financial fields.
[0145] Where, γ s and β s As parameters for the AdaIN operation, the fused features F are first processed. l Perform normalization adjustment, γ s Controls the scaling of features (such as enhancing or weakening the contrast of stroke details), β s The baseline offset of the control features (such as adjusting overall brightness or color tendency) is then processed by an MLP to generate the first scale parameters. The first scale parameter is used to control the "scaling intensity" adjustment of the feature.
[0146] γ t and β t With γ s β s Similarly, both adjust the fused feature F through the AdaIN operation. l However, the adjustment objective is to generate second-scale parameters. The second scale parameter is used to control the adjustment of the feature's "offset direction and magnitude". Among them, γ t The sensitivity to control offset (e.g., more sensitive to features with large style differences), β t The basic direction of the offset is controlled (e.g., offset towards structural features of the target style). The adjusted features are processed by MLP to generate second-scale parameters.
[0147] Furthermore, the first scale parameter, the second scale parameter, and the constraint features are fused and calculated to obtain hierarchical features, including:
[0148] Hierarchical features are generated according to the following formula:
[0149]
[0150] Among them, h l+1 This represents the hierarchical features output by the (l+1)th coupled layer; This represents a slice of the feature map along the channel dimension, specifically the first d channels. This represents a slice of the feature map along the channel dimension; ⊙ represents element-wise multiplication; exp represents exponential operation. represents the first scale parameter; represents the second scale parameter; r l represents the constraint feature.
[0151] This structure can form a feature pyramid, effectively alleviating the problem of detail loss in deep networks. At the same time, by slicing the feature map along the channel dimension (the first d channels and the remaining channels), and performing element-wise multiplication ⊙ and addition operations respectively in combination with the first scale parameter s and the second scale parameter t, fine-tuning of the features is achieved. In the medical and health scenario, when generating electronic medical records, the stroke channel features of diagnostic terms can be optimized specifically (such as enhancing the stroke clarity of the character "cancer"), while maintaining the stability of the overall layout channel features; in the fintech field, when generating contract texts, the component channel features of core terms can be adjusted mainly (such as standardizing the structure of the two characters "guarantee"), taking into account the unity of the full-text style channel features.
[0152] Secondly, the direct superposition of constraint features further strengthens the consistency of cross-layer features, avoiding feature distortion caused by scale parameter adjustment. For example, in the generation of medical prescription fonts, the stroke features of numbers related to "dosage" can be ensured to be coherent from shallow to deep layers; when generating financial statement fonts, the font styles of key data such as "yield rate" can be maintained without deviation. This mechanism enables hierarchical features to have both the flexibility of dynamic adjustment and the stability of structure, providing a reliable basis for subsequent high-quality font generation and adapting to the accuracy and standardization requirements of font generation in the medical and financial fields.
[0153] S108: Convert the hierarchical features into a new font image through a decoder.
[0154] In the medical and health field, taking the generation of standardized drug dosage label fonts as an example, the hierarchical features H l+1 are converted into a new font image through a decoder:
[0155] The hierarchical features H l+1 (including the stroke details, component structures, and overall style features of the character "500mg") are input into the decoder. This feature has been modeled by a mixture of Gaussian distributions, integrating the pen style of handwritten fonts (such as the hook of "m") and the geometric accuracy of dosage numbers (such as the starting turning point of "5").
[0156] Then, through multiple layers of transposed convolutions, the low-resolution features (such as 32×32) are gradually upsampled to the target resolution (such as 512×512) to restore the complete glyph structure of the drug label. For example, in "500mg", the transposed convolution layer gradually refines the staccato details of the characters "mg" and the stroke closed curves of the dosage numbers "500".
[0157] Then, introduce residual connections in the middle layer of the decoder (such as the 3rd layer), and add the hierarchical feature H l+1 element-wise to the upsampled feature to strengthen the key stroke details (such as the hook end of the "m" character).
[0158] The font image output by the decoder needs to meet the medical industry specifications. For example, the typesetting alignment of the generated result is constrained by a predefined L1 loss function (such as the spacing between the dosage unit "mg" and the number "500" in "500mg") to ensure the readability and compliance of drug labels.
[0159] In the field of fintech, take the generation of multi-language financial contract amount fonts as an example:
[0160] Input the hierarchical feature H l+1 (including the slash stroke with the English symbol "$" and the surrounding radical of the Chinese character "wan") into the decoder. This feature has been modeled by a mixture of Gaussian distributions to ensure the joint probability density characteristics of Chinese and English characters.
[0161] Then, gradually upsample the low-resolution feature (such as 64×64) to 1024×1024 through transposed convolutional layers to restore the complete glyphs of the amount symbol and Chinese characters. For example, in "10,000.00壹万圆整", the transposed convolutional layers gradually refine the slash stroke of the symbol and the horizontal hook structure of the character "wan".
[0162] Next, introduce a cross-attention mechanism in the high layer of the decoder (such as the 5th layer) to dynamically adjust the local structures of Chinese and English characters (such as the precise contour of the "$" symbol and the surrounding radical of the character "yuan") to ensure the visual consistency of cross-language typesetting.
[0163] The font image output by the decoder needs to meet the financial industry standards. For example, the anti-counterfeiting features of the generated result are constrained by the discriminator in adversarial training (such as the tamper-proof design of amount numbers), and at the same time, ensure that the spacing and font style of Chinese and English characters (such as the modern sense of sans-serif fonts) meet the requirements of the contract template.
[0164] The above decoding process solves the problems of the standardization of drug dosage writing in the field of medical and health and the accuracy of multi-language typesetting in the field of fintech through multi-stage upsampling and style consistency constraints, and at the same time realizes the collaborative optimization of local details and global styles through the conditional fusion of hierarchical features.
[0165] Take the conversion from Chinese regular script to Song typeface as an example to illustrate the workflow:
[0166] 1. Input data:
[0167] Reference font: Regular script character "Yong" (512×512 pixels), including obvious pen strokes and pauses;
[0168] Target font: Song typeface character 'Yong' (low resolution 256×256 pixels), with straight strokes and no decorations.
[0169] 2. Feature extraction:
[0170] The encoder captures the pen tip details of regular script through the d = 1 branch (such as the thickness variation of the right-falling stroke);
[0171] The d = 8 branch extracts the overall structural features (such as the interpenetrating layout of the character 'Yong');
[0172] Attention map a l Strengthen the d = 4 branch features at the stroke intersections.
[0173] 3. Style transfer:
[0174] The decoder adjusts the AdaIN parameters according to the Song typeface style embedding, and converts the pen tips of regular script into straight strokes of Song typeface;
[0175] Multi-scale residuals ensure that the horizontal hook component in the character 'Yong' maintains topological correctness.
[0176] 4. Output result:
[0177] Generate a high-resolution Song typeface 'Yong' of 512×512, with trumpet-shaped decorations formed at the ends of the strokes (Song typeface features).
[0178] In this embodiment, the stroke-level, component-level, and style-level features are explicitly separated through parallel dilated convolutional branches, and the multi-scale information is dynamically fused by combining the spatial attention mechanism, effectively solving the problems of local geometric distortion and global style inconsistency. At the same time, an adaptive transformation function is introduced to enable the parameters of the coupling layer to be dynamically adjusted according to the input font style, significantly enhancing the generalization ability of the model to different writing styles. In addition, the low-level high-resolution features and high-level semantic features are aggregated through multi-scale residual connections, alleviating the problem of detail loss in deep network processing.
[0179] Please refer to Figure 5 , this embodiment provides a font generation device 500, including:
[0180] An acquisition unit 501, configured to acquire a font image and a reference font, and preprocess the font image to obtain a feature map;
[0181] A feature extraction unit 502, configured to extract multi-scale features from the feature map through multiple depthwise separable convolutional branches with different dilation rates;
[0182] A feature fusion unit 503, configured to dynamically fuse the multi-scale features through a spatial attention mechanism to obtain a fused feature;
[0183] Construction unit 504 is used to construct a hierarchical feature pyramid and fuse low-level high-resolution features and high-level semantic features in the feature map across scales through multi-scale residual connections to obtain constrained features.
[0184] The style extraction unit 505 is used to extract style from the reference font through a convolutional neural network to obtain a style embedding vector;
[0185] Transformation unit 506 is used to transform the style embedding vector into distribution parameters within the latent space;
[0186] Transformation unit 507 is used to input the fused features, the distribution parameters and the constraint features into an adaptive transformation function to obtain hierarchical features;
[0187] The conversion unit 508 is used to convert the hierarchical features into a new font image through a decoder.
[0188] Furthermore, the feature extraction unit 502 includes:
[0189] A stroke detail extraction subunit is used to extract stroke detail features from the feature map through a depthwise separable convolutional branch with a first dilation rate;
[0190] A component feature extraction subunit is used to extract component features from the feature map through a depthwise separable convolutional branch with a second dilation rate;
[0191] A character local extraction subunit is used to extract local structural features of characters from the feature map through a depthwise separable convolution branch with a third dilation rate;
[0192] A global style extraction subunit is used to extract global style features from the feature map through a depthwise separable convolutional branch with a fourth dilation rate.
[0193] Furthermore, the feature fusion unit 503 includes:
[0194] The activation function calculation subunit is used to sequentially concatenate, convolve, and calculate activation functions on features at multiple scales to obtain an attention weight map.
[0195] The weighted summation subunit is used to perform a weighted summation of the attention weight map and multiple scale features to obtain fused features.
[0196] Furthermore, the transformation unit 507 includes:
[0197] The parameter generation subunit is used to input the fused features and the distribution parameters into the adaptive transformation function to generate the first scale parameter and the second scale parameter;
[0198] The fusion calculation subunit is used to fuse the first scale parameter, the second scale parameter and the constraint feature to obtain hierarchical features.
[0199] Furthermore, the parameter generation subunit includes:
[0200] The first-scale parameter generation sub-unit is used to generate the first-scale parameters according to the following formula:
[0201]
[0202] The second-scale parameter generation sub-unit is used to generate the second-scale parameters according to the following formula:
[0203]
[0204] The AdaIN operation is defined as follows:
[0205]
[0206] in, γ represents the first scale parameter; MLP represents a two-layer perceptron; s β s γ t and β t All represent learnable parameters; F l Indicates fusion characteristics; denoted by σ(F), representing the second scale parameter; μ(F) represents the mean of the fusion feature F; σ(F) represents the standard deviation of the fusion feature F.
[0207] Furthermore, the fusion computing subunit includes:
[0208] Hierarchical feature generation subunits are used to generate hierarchical features according to the following formula:
[0209]
[0210] Among them, h l+1 This represents the hierarchical features output by the (l+1)th coupled layer; This represents a slice of the feature map along the channel dimension, specifically the first d channels. This represents a slice of the feature map along the channel dimension; ⊙ represents element-wise multiplication; exp represents exponential operation. Indicates the first scale parameter; Represents the second scale parameter; r l Represents constraint features.
[0211] Furthermore, the potential space adopts a mixture Gaussian distribution.
[0212] This invention provides a font generation device. First, it acquires a font image and a reference font, and preprocesses the font image to obtain a feature map. Then, it extracts multiple scale features from the feature map using multiple depthwise separable convolutional branches with different dilation rates. Next, it dynamically fuses these multiple scale features using a spatial attention mechanism to obtain fused features. A hierarchical feature pyramid is constructed, and low-level high-resolution features and high-level semantic features within the feature map are fused across scales using multi-scale residual connections to obtain constrained features. A convolutional neural network extracts style from the reference font to obtain style embedding vectors. These style embedding vectors are then transformed into distribution parameters in the latent space. The fused features, distribution parameters, and constrained features are input into an adaptive transformation function to obtain hierarchical features. Finally, a decoder converts these hierarchical features into a new font image. By explicitly separating stroke-level, component-level, and style-level features through parallel dilated convolutional branches and dynamically fusing multi-scale information using a spatial attention mechanism, the device effectively solves the problems of local geometric distortion and global style inconsistency. Furthermore, the introduction of an adaptive transformation function allows the coupling layer parameters to be dynamically adjusted according to the input font style, significantly improving the model's generalization ability to different writing styles. Furthermore, by aggregating low-level high-resolution features with high-level semantic features through multi-scale residual connections, the problem of detail loss in deep network processing is alleviated.
[0213] For specific limitations regarding the font generation device, please refer to the limitations on the font generation method above, which will not be repeated here. Each unit in the aforementioned font generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These units can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each unit.
[0214] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements the functions or steps of a font generation method on the server side.
[0215] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 7As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the client-side functions or steps of a font generation method.
[0216] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0217] Obtain a font image and a reference font, and preprocess the font image to obtain a feature map;
[0218] Multiple scale features are extracted from the feature map by using multiple depthwise separable convolutional branches with different dilation rates;
[0219] Multiple scale features are dynamically fused using a spatial attention mechanism to obtain fused features;
[0220] A hierarchical feature pyramid is constructed, and low-level high-resolution features and high-level semantic features in the feature map are fused across scales through multi-scale residual connections to obtain constrained features.
[0221] Style embedding vectors are obtained by extracting styles from the reference font using a convolutional neural network.
[0222] The style embedding vector is transformed into distribution parameters within the latent space;
[0223] The fused features, the distribution parameters, and the constraint features are input into an adaptive transformation function to obtain hierarchical features;
[0224] The hierarchical features are converted into a new font image using a decoder.
[0225] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0226] Obtain a font image and a reference font, and preprocess the font image to obtain a feature map;
[0227] Multiple scale features are extracted from the feature map by using multiple depthwise separable convolutional branches with different dilation rates;
[0228] Multiple scale features are dynamically fused using a spatial attention mechanism to obtain fused features;
[0229] A hierarchical feature pyramid is constructed, and low-level high-resolution features and high-level semantic features in the feature map are fused across scales through multi-scale residual connections to obtain constrained features.
[0230] Style embedding vectors are obtained by extracting styles from the reference font using a convolutional neural network.
[0231] The style embedding vector is transformed into distribution parameters within the latent space;
[0232] The fused features, the distribution parameters, and the constraint features are input into an adaptive transformation function to obtain hierarchical features;
[0233] The hierarchical features are converted into a new font image using a decoder.
[0234] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0235] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0236] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0237] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A font generation method, characterized in that, include: Obtain a font image and a reference font, and preprocess the font image to obtain a feature map; Multiple scale features are extracted from the feature map by using multiple depthwise separable convolutional branches with different dilation rates; Multiple scale features are dynamically fused using a spatial attention mechanism to obtain fused features; A hierarchical feature pyramid is constructed, and low-level high-resolution features and high-level semantic features in the feature map are fused across scales through multi-scale residual connections to obtain constrained features. Style embedding vectors are obtained by extracting styles from the reference font using a convolutional neural network. The style embedding vector is transformed into distribution parameters within the latent space; The fused features, the distribution parameters, and the constraint features are input into an adaptive transformation function to obtain hierarchical features; The hierarchical features are converted into a new font image using a decoder.
2. The font generation method according to claim 1, characterized in that, The extraction of multiple scale features from the feature map through multiple depthwise separable convolutional branches with different dilation rates includes: Stroke detail features are extracted from the feature map through a depth-separable convolutional branch with a first dilation rate; Component features are extracted from the feature map using a depth-separable convolutional branch with a second dilation rate; Character local structural features are extracted from the feature map through a depthwise separable convolutional branch with a third dilation rate; The overall style features are extracted from the feature map by a depthwise separable convolutional branch with a fourth dilation rate.
3. The font generation method according to claim 1, characterized in that, The method of dynamically fusing multiple scale features through spatial attention mechanism to obtain fused features includes: Multiple scale features are sequentially concatenated, convolved, and activation functions are calculated to obtain an attention weight map. The attention weight map and multiple scale features are weighted and summed to obtain the fused features.
4. The font generation method according to claim 1, characterized in that, The step of inputting the fused features, the distribution parameters, and the constraint features into the adaptive transformation function to obtain hierarchical features includes: The fused features and the distribution parameters are input into the adaptive transformation function to generate the first scale parameter and the second scale parameter; The first scale parameter, the second scale parameter, and the constraint feature are fused and calculated to obtain hierarchical features.
5. The font generation method according to claim 4, characterized in that, The step of inputting the fused features and the distribution parameters into the adaptive transformation function to generate the first scale parameters and the second scale parameters includes: The first scale parameter is generated according to the following formula: The second-scale parameters are generated according to the following formula: The AdaIN operation is defined as follows: in, γ represents the first scale parameter; MLP represents a two-layer perceptron; s β s γ t and β t All represent learnable parameters; F l Indicates fusion characteristics; denoted by σ(F), representing the second scale parameter; μ(F) represents the mean of the fusion feature F; σ(F) represents the standard deviation of the fusion feature F.
6. The font generation method according to claim 4, characterized in that, The step of fusing the first scale parameter, the second scale parameter, and the constraint feature to obtain hierarchical features includes: Hierarchical features are generated according to the following formula: Among them, h l+1 This represents the hierarchical features output by the (l+1)th coupled layer; This represents a slice of the feature map along the channel dimension, specifically the first d channels. This represents a slice of the feature map along the channel dimension; ⊙ represents element-wise multiplication; exp represents exponential operation. Indicates the first scale parameter; Represents the second scale parameter; r l Represents constraint features.
7. The font generation method according to claim 1, characterized in that, The potential space adopts a mixture Gaussian distribution.
8. A font generation device, characterized in that, include: An acquisition unit is used to acquire a font image and a reference font, and to preprocess the font image to obtain a feature map; The feature extraction unit is used to extract multiple scale features from the feature map through multiple depthwise separable convolutional branches with different dilation rates; The feature fusion unit is used to dynamically fuse features at multiple scales through a spatial attention mechanism to obtain fused features; The construction unit is used to construct a hierarchical feature pyramid. Through multi-scale residual connections, the low-level high-resolution features and high-level semantic features in the feature map are fused across scales to obtain constrained features. The style extraction unit is used to extract style from the reference font through a convolutional neural network to obtain a style embedding vector; A transformation unit is used to transform the style embedding vector into distribution parameters within the latent space; A transformation unit is used to input the fused features, the distribution parameters, and the constraint features into an adaptive transformation function to obtain hierarchical features; A conversion unit is used to convert the hierarchical features into a new font image via a decoder.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the font generation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the font generation method as described in any one of claims 1 to 7.