Artificial intelligence-based font data generation method and device, equipment and medium
By working together with a cross-domain feature fusion module and a hybrid generator, the problem of non-standard structure in Chinese character generation is solved, achieving structural standardization and visual aesthetics in the generated results, which is applicable to Chinese character generation in the financial and medical fields.
Patent Information
- Application Number
- CN202510680001.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-10-28
AI Technical Summary
In existing technologies, Chinese character glyph generation methods based on CycleGAN have the problem of non-standard structure of the generated results, which makes it difficult to meet the needs of calligraphy standards and practical applications. In particular, in the electronic contract signing in the financial field, it may cause the signature to fail manual or automated verification, increasing operational risks and forgery risks.
A character generation model employing a cross-domain feature fusion module, a hybrid generator, and a multi-target discriminator is used. The generation process is constrained by a preset total loss function. Combined with cross-domain feature extraction and cross-attention fusion strategies, the generated Chinese character images are ensured to simultaneously meet structural standardization and visual aesthetics.
It achieves full-process control from semantic understanding to visual generation, and the generated Chinese character images simultaneously meet the requirements of structural standardization and visual aesthetics, improving the standardization and reliability of the generation results. It is suitable for Chinese character generation scenarios in the financial and medical fields.
Smart Images

Figure CN120853199A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and can be applied to fields such as fintech and digital healthcare. In particular, it relates to a glyph data generation method, device, computer device, and storage medium based on artificial intelligence. Background Art
[0002] In the field of Chinese character glyph generation, deep learning-based generative adversarial network (GAN) technology, especially CycleGAN, has been used to achieve the automated generation of realistic Chinese character glyphs. However, the prior art generally faces the problem of unregulated generation results, such as structural distortion. Such defects stem from the lack of constraints on Chinese character structures in traditional CycleGAN models, resulting in the generated Chinese character glyphs being difficult to meet the requirements of calligraphy norms or practical application needs.
[0003] For example, in the scenario of signing electronic contracts in the financial field, the generation of handwritten Chinese characters needs to strictly follow the structural standards of name signatures to ensure legal effect and anti-counterfeiting ability. Signatures generated by traditional CycleGAN may have problems such as笔画粘连 or component misalignment (e.g., the disproportion of the "弓" and "长" in the character "张"), resulting in the signature being unable to pass manual or automated verification, and thus triggering disputes over the validity of the contract. Such problems not only increase the operational risks of financial institutions but may also damage the rights and interests of customers due to the risk of signature forgery.
[0004] Therefore, there is an urgent need to provide a Chinese character glyph generation method with the ability to constrain structures to improve the standardization of generation results to meet the application requirements of high-precision demand scenarios. Summary of the Invention
[0005] The purpose of the embodiments of this application is to propose a glyph data generation method, device, computer device, and storage medium based on artificial intelligence to solve the technical problem that the existing Chinese character glyph generation methods have unregulated generation results.
[0006] In a first aspect, a glyph data generation method based on artificial intelligence is provided, including:
[0007] Receiving an input Chinese character encoding and the corresponding low-resolution glyph image;
[0008] Invoking a pre-constructed glyph generation model; wherein, the glyph generation model includes a cross-domain feature fusion module, a hybrid generator, and a multi-object discriminator, and the multi-object discriminator is used to constrain the glyph generation process of the hybrid generator through a preset total loss function during the model training stage;
[0009] Extracting features from the Chinese character encoding based on the first feature extraction model in the cross-domain feature fusion module to obtain the corresponding semantic features;
[0010] Based on the second feature extraction model in the cross-domain feature fusion module, feature extraction is performed on the character image to obtain the corresponding visual features;
[0011] Based on the cross-domain feature fusion module, a preset cross-attention fusion strategy is used to perform feature fusion processing on the semantic features and the visual features to obtain the corresponding fused features;
[0012] The fusion features are converted into corresponding Chinese character images based on the fusion generator.
[0013] The Chinese character image is then processed for output.
[0014] Secondly, an artificial intelligence-based glyph data generation device is provided, comprising:
[0015] The receiving module is used to receive the input Chinese character encoding and the corresponding low-resolution character image;
[0016] The first calling module is used to call a pre-built character generation model; wherein, the character generation model includes a cross-domain feature fusion module, a hybrid generator and a multi-objective discriminator, and the multi-objective discriminator is used to constrain the character generation process of the hybrid generator through a preset total loss function during the model training phase;
[0017] The first extraction module is used to extract features from the Chinese character encoding based on the first feature extraction model in the cross-domain feature fusion module to obtain the corresponding semantic features;
[0018] The second extraction module is used to extract features from the character image based on the second feature extraction model in the cross-domain feature fusion module to obtain the corresponding visual features.
[0019] The fusion module is used to perform feature fusion processing on the semantic features and the visual features based on the cross-domain feature fusion module, using a preset cross-attention fusion strategy, to obtain the corresponding fused features;
[0020] The conversion module is used to convert the fused features into corresponding Chinese character images based on the fusion generator;
[0021] The output module is used to process the output of the Chinese character image.
[0022] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described artificial intelligence-based character data generation method.
[0023] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described artificial intelligence-based glyph data generation method.
[0024] In the above-mentioned scheme implemented by the AI-based character shape data generation method, device, computer equipment, and storage medium, the input Chinese character encoding and corresponding low-resolution character shape image are first received; then, a pre-constructed character shape generation model is invoked; wherein, the character shape generation model includes a cross-domain feature fusion module, a hybrid generator, and a multi-target discriminator, the multi-target discriminator being used to constrain the character shape generation process of the hybrid generator through a preset total loss function during the model training phase; then, based on the first feature extraction model in the cross-domain feature fusion module, features are extracted from the Chinese character encoding to obtain corresponding semantic features; and based on the second feature extraction model in the cross-domain feature fusion module, features are extracted from the character shape image to obtain corresponding visual features; and based on the cross-domain feature fusion module, a preset cross-attention fusion strategy is used to perform feature fusion processing on the semantic features and the visual features to obtain corresponding fused features; subsequently, based on the hybrid generator, the fused features are converted into the corresponding Chinese character image; finally, the Chinese character image is output. Upon receiving the input Chinese character encoding and the corresponding low-resolution character image, this application invokes a character generation model comprising a cross-domain feature fusion module, a hybrid generator, and a multi-target discriminator. Semantic features are extracted from the Chinese character encoding based on the first feature extraction model in the cross-domain feature fusion module, and visual features are extracted from the character image based on the second feature extraction model in the same module. Then, based on the cross-domain feature fusion module, a pre-cross-attention fusion strategy is used to fuse the semantic features and the visual features to obtain fused features. Finally, the hybrid generator converts the fused features into the corresponding Chinese character image and outputs it. This application, through the use of a character generation model and the collaborative work of the cross-domain feature fusion module and the hybrid generator, performs character generation processing corresponding to the input Chinese character encoding and the corresponding low-resolution character image. This enables end-to-end control from semantic understanding to visual generation, effectively ensuring that the generated Chinese character image simultaneously meets structural requirements and visual aesthetics. Attached Figure Description
[0025] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is an exemplary system architecture diagram to which this application can be applied;
[0027] Figure 2 This is a flowchart of an embodiment of the AI-based glyph data generation method according to this application;
[0028] Figure 3 This is a schematic diagram of the structure of an embodiment of the AI-based character data generation apparatus according to this application;
[0029] Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0031] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0032] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0033] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0034] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0035] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.
[0036] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0037] It should be noted that the AI-based character data generation method provided in this application is generally executed by a server / terminal device, and correspondingly, the AI-based character data generation device is generally located in the server / terminal device.
[0038] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0039] Continue to refer to Figure 2 This document illustrates a flowchart of an embodiment of the AI-based character shape data generation method according to this application. The order of steps in the flowchart can be changed, and some steps can be omitted, depending on different requirements. The AI-based character shape data generation method provided in this application can be applied to any scenario requiring character shape data generation, and thus can be applied to products in these scenarios, such as Chinese character shape generation scenarios in the financial and medical fields. The AI-based character shape data generation method includes the following steps:
[0040] Step S201: Receive the input Chinese character encoding and the corresponding low-resolution character image.
[0041] In this embodiment, the AI-based glyph data generation method runs on an electronic device (e.g., Figure 1 The server / terminal device shown can obtain the input Chinese character encoding and the corresponding low-resolution glyph images through wired connection or wireless connection. It should be noted that the above wireless connection methods can include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultrawideband) connections, and other currently known or future-developed wireless connection methods. The execution subject of this application is specifically a glyph generation system, or simply referred to as the system. Among them, the above Chinese character encoding is specifically the Chinese character Unicode encoding (e.g., U+548C represents "and"). The above low-resolution glyph image refers to the low-quality glyph image of the Chinese character corresponding to the above Chinese character encoding.
[0042] This application can be applied to the business scenarios of Chinese character glyph generation in the financial insurance field and the digital medical field. Exemplarily, in the financial insurance field, the business scenarios of Chinese character glyph generation can include intelligent policy customization and anti-fraud. Scenario example: personalized policy generation and signature verification. Requirement background: The insurance industry needs to process a large number of standardized policies, but different customers (such as enterprises, individuals) have personalized requirements for the display form of terms, signature styles, etc. At the same time, the legal verification of electronic policies needs to ensure the uniqueness and non-tamperability of signatures or key terms. Application of Chinese character glyph generation: Dynamic clause generation: By generating Chinese character glyphs in a specific style (such as regular script, Song typeface), key information such as the customer name, amount, and clause number in the policy terms is embedded in the template in "calligraphy-level" fonts to enhance professionalism and readability. Example: Generate a "Policy effective date: January 1, 2024" with a calligraphy style for high-end customers, replacing the mechanical font. Anti-counterfeiting signature generation: Based on the customer's historical signature data, a "handwritten style" Chinese character signature exclusive to the customer is generated through a generative adversarial network (GAN) and embedded in the electronic policy. Combined with blockchain technology, signature uniqueness verification is achieved. Example: The signature "Zhang San" of customer Zhang San is generated in a unique cursive style and compared with the historical signatures in the database for structural similarity to prevent forgery. Visualization of risk clauses: Glyphs of risk keywords (such as "deductible", "exclusions") in complex clauses are strengthened (such as bold, color change, special font) to enhance the customer's reading experience. Value manifestation: Enhance the customer's trust in the policy (personalization and anti-counterfeiting). Reduce the risk of disputes caused by ambiguous clauses.
[0043] In the field of digital healthcare, the business scenarios of Chinese character glyph generation can include assisting in diagnosis and communicating with patients. Scenario examples: medical report generation and multilingual adaptation. Requirement background: In medical scenarios, doctors need to quickly generate reports containing information such as patient names, diagnosis results, and medication instructions, and need to meet the reading needs of different patient groups (such as elderly patients and foreign patients). In addition, rare diseases or complex cases require visual glyphs to assist doctors in understanding. Applications of Chinese character glyph generation: Elderly-friendly report generation: To address the vision decline problem of elderly patients, generate Chinese character glyphs with large fonts and high contrast, and adjust the stroke thickness (such as bolding keywords like "hypertension" and "diabetes") to improve readability. Example: Generate the "160" in "Systolic blood pressure: 160 mmHg" as bold red font to distinguish it from "Diastolic blood pressure". Multilingual mixed layout: When foreign patients seek medical treatment, mix Chinese characters with pinyin and English terms in the layout. Ensure the unified font style of Chinese characters and pinyin / English through glyph generation technology (such as using sans-serif fonts uniformly). Example: Generate the mixed layout of "Headache (Tóutòng, Headache)" to avoid reading obstacles caused by mismatched fonts. Symbolic annotation for rare diseases: Generate glyphs with symbols or decorative strokes for rare disease names or special symptoms (such as "Marfan syndrome") to assist doctors in quickly identifying them. Example: Generate a "Fan" character with a wavy line next to "Marfan syndrome" to prompt doctors to pay attention to genetic characteristics. Value manifestation: Improve the readability of medical reports and patient compliance. Assist doctors in quickly locating key information and reduce the risk of misdiagnosis.
[0044] Step S202, call the pre-constructed glyph generation model; wherein, the glyph generation model includes a cross-domain feature fusion module, a hybrid generator, and a multi-object discriminator, and the multi-object discriminator is used to constrain the glyph generation process of the hybrid generator through a preset total loss function during the model training phase.
[0045] In this embodiment, the above-mentioned character generation model consists of three main modules: a cross-domain feature fusion module, a hybrid generator, and a multi-objective discriminator. These modules work together to achieve full-process control from semantic understanding to visual generation, ensuring that the generated result (Chinese character image) simultaneously meets structural standardization and visual aesthetics. The main functions of the cross-domain feature fusion module are: (1) extracting the semantic embedding representation of Chinese characters using the BERT model; (2) extracting the visual features of the input image using the ResNet-50 model; and (3) dynamically aligning the semantic and visual feature spaces using a cross-attention mechanism. The hybrid generator is responsible for converting the fused features generated based on the cross-domain feature fusion module into high-quality Chinese character images. This module adopts an improved U-Net structure. To ensure the high quality of the generated Chinese characters, the above-mentioned multi-objective discriminator is designed to constrain the generation process through various loss functions. These loss functions work together to optimize the generation result. The construction process of the above-mentioned character generation model will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0046] Step S203: Based on the first feature extraction model in the cross-domain feature fusion module, feature extraction is performed on the Chinese character encoding to obtain the corresponding semantic features.
[0047] In this embodiment, the specific implementation process of extracting features from the Chinese character encoding based on the first feature extraction model in the cross-domain feature fusion module to obtain the corresponding semantic features will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0048] Step S204: Based on the second feature extraction model in the cross-domain feature fusion module, feature extraction is performed on the character image to obtain the corresponding visual features.
[0049] In this embodiment, the specific implementation process of extracting features from the character image based on the second feature extraction model in the cross-domain feature fusion module to obtain the corresponding visual features will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0050] Step S205: Based on the cross-domain feature fusion module, a preset cross-attention fusion strategy is used to perform feature fusion processing on the semantic features and the visual features to obtain the corresponding fused features.
[0051] In this embodiment, the specific implementation process of using the cross-domain feature fusion module to perform feature fusion processing on the semantic features and the visual features to obtain the corresponding fused features using a preset cross-attention fusion strategy will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0052] Step S206: Based on the hybrid generator, the fusion features are converted into corresponding Chinese character images.
[0053] In this embodiment, the specific implementation process of converting the fusion features into corresponding Chinese character images based on the hybrid generator will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0054] Step S207: Output processing is performed on the Chinese character image.
[0055] In this embodiment, the specific implementation process of outputting the Chinese character image described above will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0056] Upon receiving the input Chinese character encoding and the corresponding low-resolution character image, this application invokes a character generation model comprising a cross-domain feature fusion module, a hybrid generator, and a multi-target discriminator. Semantic features are extracted from the Chinese character encoding based on the first feature extraction model in the cross-domain feature fusion module, and visual features are extracted from the character image based on the second feature extraction model in the same module. Then, based on the cross-domain feature fusion module, a pre-cross-attention fusion strategy is used to fuse the semantic features and the visual features to obtain fused features. Finally, the hybrid generator converts the fused features into the corresponding Chinese character image and outputs it. This application, through the use of a character generation model and the collaborative work of the cross-domain feature fusion module and the hybrid generator, performs character generation processing corresponding to the input Chinese character encoding and the corresponding low-resolution character image. This enables end-to-end control from semantic understanding to visual generation, effectively ensuring that the generated Chinese character image simultaneously meets structural requirements and visual aesthetics.
[0057] In some alternative implementations, step S203 includes the following steps:
[0058] The Chinese character encoding is parsed into the corresponding Chinese character based on a preset standard library.
[0059] In this embodiment, the aforementioned standard library is a library with the function of parsing Unicode encoding. The received Chinese character encoding can be parsed using this standard library to obtain the corresponding Chinese character.
[0060] Perform standardization processing on the Chinese character to obtain the corresponding target Chinese character.
[0061] In this embodiment, the above standardization processing includes converting the Chinese character into a standard regular script or Song typeface glyph to avoid interference of font style differences on subsequent processing.
[0062] Convert the target Chinese character into the corresponding stroke sequence.
[0063] In this embodiment, by using a predefined Chinese character stroke database (such as a Chinese character stroke database), a Chinese character can be disassembled into a stroke sequence (e.g., "he" → ["丿","一","丨",...]). Among them, each stroke is mapped to a one-hot encoding, forming a matrix with a sequence length of 8 and a dimension of 20 (assuming a total of 20 basic strokes), that is, a stroke sequence.
[0064] Based on the first feature extraction model in the cross-domain feature fusion module, perform encoding processing on the stroke sequence to obtain the corresponding semantic embedding representation.
[0065] In this embodiment, the above first feature extraction model can specifically adopt the BERT model. In the input layer of the BERT model, the stroke sequence is regarded as a "pseudo-word sequence", and [CLS] and [SEP] tags are added. And through the 12-layer Transformer encoding of the BERT model, the 768-dimensional output at the [CLS] position is taken as the semantic embedding E_BERT.
[0066] Among them, after disassembling the target Chinese character C into a stroke sequence S = {s1,…,s n}, the semantic embedding representation of the Chinese character is extracted through the BERT model, that is, the stroke-level context embedding corresponding to the Chinese character is obtained: E BERT = BERT("[CLS]" + s1 + … + s n + "[SEP]"). Among them, "+" represents a concatenation operation, and the output dimension is n×768. For example, the character "he" is decomposed into 8 strokes ["丿","一","丨","丿","丶","丨", "一"]. The BERT model can capture the topological relationship and relative position constraints between strokes.
[0067] Regard the semantic embedding representation as the semantic feature corresponding to the Chinese character encoding.
[0068] This application uses a standard library to parse Chinese character encoding into Chinese characters, standardizes these characters to obtain target Chinese characters, then converts them into corresponding stroke sequences. The stroke sequences are then encoded using the first feature extraction model in the cross-domain feature fusion module, and the resulting semantic embedding representation is used as the semantic feature corresponding to the Chinese character encoding. This allows for efficient and accurate extraction of semantic features from Chinese character encoding, improving extraction efficiency and ensuring the accuracy of the obtained semantic features.
[0069] In some optional implementations of this embodiment, step S204 includes the following steps:
[0070] Obtain the preset preprocessing strategy.
[0071] In this embodiment, the preprocessing strategy includes denoising, binarization, and correction. Specifically, denoising includes applying non-local means or Gaussian filtering to eliminate noise. Binarization includes converting the image into a black-and-white binary image using an adaptive thresholding method (such as the Otsu algorithm) to enhance stroke contours. Correction includes correcting slanted characters through affine transformation to ensure horizontal / vertical alignment of stroke directions.
[0072] The character image is processed based on the preprocessing strategy to obtain the corresponding specified image.
[0073] In this embodiment, the above-mentioned character image can be preprocessed according to the strategy content of the above preprocessing strategy to obtain the preprocessed image, which is then used as the corresponding designated image.
[0074] Based on the second feature extraction model in the cross-domain feature fusion module, feature extraction is performed on the specified image to obtain the corresponding specified features.
[0075] In this embodiment, the second feature extraction model described above can specifically employ the ResNet-50 model. The process of visual feature extraction using the ResNet-50 model includes: extracting multi-level features of the image, focusing on stroke outlines, intersections, and overall layout, to obtain visual features F. CNN Specifically, the input image I is processed using ResNet-50. c ∈R H×W×3 (∈ indicates belonging to):
[0076]
[0077] Among them, here This indicates the spatial downsampling rate, and 1024 represents the number of channels.
[0078] The specified feature is used as the visual feature corresponding to the glyph image.
[0079] This application processes character images using a preprocessing strategy to obtain corresponding specified images. Then, it extracts features from the specified images based on the second feature extraction model in the cross-domain feature fusion module and uses the obtained specified features as the visual features corresponding to the character images. This enables efficient and accurate extraction of visual features from character images, improves the efficiency of visual feature extraction, and ensures the accuracy of the obtained visual features.
[0080] In some optional implementations, the step of fusing the semantic features and visual features using a preset cross-attention fusion strategy based on the cross-domain feature fusion module to obtain the corresponding fused features includes: after obtaining the semantic features and visual features, establishing the correspondence between these two heterogeneous features by performing cross-attention fusion. The specific calculation process is as follows:
[0081] Q = E BERT ·W Q K = flatten(F) CNN )·W K
[0082]
[0083] in, d is a learnable parameter. k =64 represents the attention head dimension, · represents the dot product operator, the flatten operation flattens the spatial dimension, and Softmax represents the activation function, which can be understood as normalizing the values in the matrix. The final generated fused feature is:
[0084] F fuSed =LayerNorm(E BERT +A cross ·W O )
[0085] Among them, W O ∈R 768×768 To output the projection matrix, LayerNorm represents the layer normalization operation. In this way, the glyph generation model can dynamically associate the semantic structure and visual representation of Chinese characters, such as associating the semantic concept of "horizontal stroke" with its proper position and shape in the image.
[0086] In some optional implementations, the hybrid generator includes a multilayer perceptron and a target convolutional neural network model, wherein the target convolutional neural network model includes an upsampling module and a downsampling module; step S206 includes the following steps:
[0087] Obtain the preset random noise vector.
[0088] In this embodiment, the network architecture of the hybrid generator includes five downsampling modules and five corresponding upsampling modules. Each module contains residual connections to ensure effective transmission of deep information. During downsampling, the number of feature channels increases layer by layer from 64 to 1024; during upsampling, it decreases back to 64 layer by layer, and finally outputs a 3-channel RGB image through a 1×1 convolution. A random noise vector z∈R^100 conforming to a standard normal distribution can be generated according to actual needs. The random noise vector is injected into the generator through an MLP to balance determinism and diversity.
[0089] Based on the multilayer perceptron in the hybrid generator, the random noise vector and the fused feature are concatenated to obtain the corresponding first feature.
[0090] In this embodiment, after obtaining the fused features, the hybrid generator is responsible for converting these features into high-quality Chinese character images. This module adopts an improved U-Net structure (i.e., a target convolutional neural network model), and the specific operation process is as follows:
[0091] First, a latent representation is constructed by fusing the random noise vector z with the fused feature F. fused The splicing provides initial conditions for the generation process:
[0092] h0 = MLP([z + avgpool(F)) fused )])
[0093] Here, MLP stands for Multilayer Perceptron, used to map initial features; avgpool represents the average pooling operation, compressing the fused features to a dimension that matches the noise vector; and Z represents the random noise vector. This step introduces randomness into the Chinese character generation process, ensuring that even for the same Chinese character, diverse styles of results can be generated.
[0094] The first feature is downsampled based on the downsampling module in the target convolutional neural network model to obtain the corresponding second feature.
[0095] In this embodiment, the first feature can be progressively downsampled using the downsampling module in the target convolutional neural network model, increasing the number of channels from 64 to 1024, and multi-scale features can be extracted to obtain the corresponding second feature.
[0096] Based on the upsampling module in the target convolutional neural network model, the second feature is subjected to progressive upsampling processing to obtain the corresponding output image.
[0097] In this embodiment, the above-mentioned progressive upsampling process includes: performing a transposed convolution operation on the l-th layer in the target convolutional neural network model:
[0098] h l = AdaIN(ConvTranspose(h l-1 ), γ l ·F fused + β l )
[0099] where γ l , β l are affine parameters learned from F fused for controlling the generated style; AdaIN represents adaptive instance normalization, which is responsible for injecting the information of the fused features into the generation process; ConvTranspose represents the transposed convolution operation for implementing the upsampling and refinement of the feature map.
[0100] To handle the long-range dependencies between strokes in Chinese characters, the system introduces a self-attention mechanism at the key layer:
[0101]
[0102] where W q , W k , W v ∈ R C×d are projection matrices, C represents the channel dimension of the input feature h l , is the attention head dimension after dimensionality reduction, and · T represents matrix transpose. The self-attention mechanism enables the model to capture the relationships between distant pixels, ensuring the overall structural coordination of the generated glyphs, such as the symmetric balance relationship between the left-falling stroke and the right-falling stroke in the character "永".
[0103] Take the output image as the Chinese character image corresponding to the fused feature.
[0104] In this application, by obtaining a preset random noise vector, then based on the use of the multi-layer perceptron in the hybrid generator, the random noise vector and the fused feature are concatenated to obtain a first feature, and then based on the use of the downsampling module in the target convolutional neural network model, the first feature is downsampled to obtain a second feature, and further based on the upsampling module in the target convolutional neural network model, the second feature is progressively upsampled to obtain an output image and used as the Chinese character image corresponding to the fused feature, thereby enabling the efficient and accurate generation of Chinese character images that meet the structural standardization and visual aesthetics.
[0105] In some optional implementation manners, step S207 includes the following steps:
[0106] The structural accuracy of the Chinese character image is verified.
[0107] In this embodiment, the structural accuracy verification includes stroke structure matching, symmetry evaluation, and balance evaluation. Specifically, the stroke structure matching includes: constructing a standard template library containing standard stroke sequences of relevant Chinese characters and corresponding spatial positional relationships; then using image processing techniques (such as edge detection and morphological operations) to extract the stroke contours of the generated image, and spatially aligning and topologically matching the extracted strokes with the strokes in the standard template library. Subsequently, the matching degree (such as Dice coefficient, IoU) between the generated image strokes and the standard template is calculated. A threshold (such as Dice coefficient ≥ 0.85) is set to determine whether the structural accuracy meets the standard. The symmetry evaluation includes: detecting the symmetry (such as angle difference, length ratio) of the left-falling stroke and right-falling stroke in the target Chinese character corresponding to the above-mentioned Chinese character image. Then, the symmetry deviation (such as angle deviation ≤ 5°, length deviation ≤ 10%) is calculated. The balance evaluation includes: calculating the centroid position of the character and comparing it with the standard centroid position. A centroid offset ≤ 2 pixels is considered balanced.
[0108] The system pre-sets thresholds corresponding to stroke structure matching, symmetry evaluation, and balance evaluation (e.g., Di ce coefficient ≥ 0.85, symmetry deviation ≤ 5°). Only when the Chinese character image is detected to simultaneously meet the threshold requirements of stroke structure matching, symmetry evaluation, and balance evaluation will the Chinese character image be determined to have passed the structural accuracy verification; otherwise, the Chinese character image will be determined to have failed the structural accuracy verification.
[0109] If the Chinese character image passes the structural accuracy verification, then the visual authenticity verification of the Chinese character image is performed.
[0110] In this embodiment, the visual authenticity verification includes FID (Frechet Inception Distance) and IS (Inception Score) evaluation. FID evaluation involves extracting feature vectors from the generated and real images using a pre-trained Inception-v3 model. The Frechet distance between the feature distributions is calculated; a smaller distance indicates higher visual quality. IS evaluation involves calculating the entropy of the class probability distribution and the KL divergence of the conditional probability distribution of the generated image. A higher IS value indicates better image diversity and clarity. Furthermore, thresholds for FID and IS are pre-set, for example, FID ≤ 50 and IS ≥ 3.0. If a Chinese character image simultaneously meets the threshold requirements for both FID and IS, it is considered to have passed the visual authenticity verification; otherwise, it is considered to have failed.
[0111] If the Chinese character image passes the visual authenticity verification, then the preset image output method is obtained.
[0112] In this embodiment, the selection of the above-mentioned image output method is not specifically limited and can be determined according to the user's actual needs. For example, any one of the methods such as email sending, interface display, and message sending can be used. Here, the "user" refers to the user who inputs the above-mentioned Chinese character encoding and the corresponding low-resolution character image.
[0113] The Chinese character image is output based on the image output method described above.
[0114] In this embodiment, the Chinese character image can be sent to the user according to the selected image output method to complete the output processing of the Chinese character image.
[0115] After generating Chinese character images, this application intelligently verifies both structural accuracy and visual realism, thereby efficiently and objectively evaluating the structural accuracy and visual realism of the generated images. Once the image passes both structural accuracy and visual realism verifications, it automatically processes the output based on the acquired image output method, effectively ensuring that the output image meets expected standards and thus improving the user experience.
[0116] In some optional implementations of this embodiment, before step S202, the electronic device may further perform the following steps:
[0117] Obtain pre-collected sample data.
[0118] In this embodiment, a standard stroke sequence dataset containing a first number of Chinese characters is prepared in advance (for semantic feature extraction), along with high-quality images corresponding to the Chinese characters (for visual feature extraction and discriminator training), and low-quality images corresponding to the Chinese characters (for testing the robustness of visual feature extraction). All prepared data are then integrated to obtain corresponding sample data.
[0119] Call the pre-built original model containing the target structure.
[0120] In this embodiment, the original model is a pre-built model that includes a cross-domain feature fusion module, a hybrid generator, and a multi-target discriminator.
[0121] Obtain the preset end-to-end training strategy.
[0122] In this embodiment, the end-to-end training strategy includes a forward propagation phase, a backpropagation and optimization phase, and an iterative training phase. Specifically, the forward propagation phase includes: Input processing: Converting Unicode encoding U+6C38 into a stroke sequence and extracting semantic embeddings using a BERT model. If the input contains low-quality Chinese character images, visual features are extracted using ResNet-50. Feature fusion: Using a cross-attention mechanism, the semantic embeddings and visual features are fused to obtain fused features. Latent representation construction: Generating a random noise vector, concatenating it with the fused features, and mapping it to an initial latent representation using a multilayer perceptron. Image generation: Inputting the initial latent representation into the U-Net generator, and generating high-resolution images of the specified Chinese characters through progressive upsampling.
[0123] The backpropagation and optimization phase includes: Discriminator evaluation: The generated specified Chinese character image G(z) and a real high-quality Chinese character image are input into the multi-objective discriminator. The discriminator evaluates the quality of the generated image from the perspectives of adversarial loss, feature matching loss, cross-domain similarity loss, and stroke constraint loss. Loss calculation: The adversarial loss is calculated to measure the distribution difference between the generated image and the real image. The feature matching loss is calculated to ensure the consistency of features between the generated image and the real image in the intermediate layers of the discriminator. The cross-domain similarity loss is calculated to ensure the consistency between the visual features and semantic embeddings of the generated image. The stroke constraint loss is calculated to ensure that the stroke structure of the generated image conforms to the standard. Combining loss functions: The various loss functions are combined according to their weights to obtain the total loss function. Optimizing the generator and discriminator: The parameters of the generator and discriminator are updated using a gradient descent algorithm (such as Adam) based on the total loss function. Generator parameter update direction: Minimize the total loss function to generate more realistic and structurally accurate Chinese character images. Discriminator parameter update direction: Maximize the adversarial loss (while minimizing other losses) to improve the discriminative ability.
[0124] The iterative training phase includes repeating the forward and backward propagation and optimization phases described above until the training termination conditions are met (such as reaching the maximum number of iterations, stable image quality, etc.), and a well-trained model is obtained. Monitoring and tuning: Periodically evaluate the quality of the generated images (e.g., using metrics such as FID and IS). Adjust hyperparameters such as loss function weights and learning rate based on the evaluation results.
[0125] Based on the end-to-end training strategy and the total loss function, the original model is trained using the sample data to obtain a target model that meets the construction requirements.
[0126] In this embodiment, the sample data can be trained using the total loss function and sample data according to the strategy content of the end-to-end training strategy, so as to obtain a target model that meets the construction requirements and serve as the final character generation model.
[0127] The target model is used as the glyph generation model.
[0128] This application acquires pre-collected sample data; then calls a pre-built original model containing the target structure; subsequently acquires a preset end-to-end training strategy; and then, based on the end-to-end training strategy and the total loss function, trains the original model using the sample data to obtain a target model that meets the construction requirements; and uses the target model as the glyph generation model. This application, by acquiring pre-collected sample data and calling a pre-built original model containing the target structure, and then using the sample data to train the original data based on the combination of the end-to-end training strategy and the total loss function, can efficiently and accurately construct a glyph generation model that meets the requirements, improving the model construction efficiency of the glyph generation model and ensuring the model processing effect of the obtained glyph generation model.
[0129] In some optional implementations of this embodiment, before step S202, the electronic device may further perform the following steps:
[0130] Obtain the preset adversarial loss function, feature matching loss function, cross-domain similarity loss function, and stroke constraint loss function.
[0131] In this embodiment, the aforementioned adversarial loss function realizes the game process between the generator and the discriminator:
[0132] L adv =E[logD] {patch} (I real )]+E[log(1-D patch (G(z)))]
[0133] Among them, D patch Let represent the PatchGAN discriminator, employing a multi-scale discrimination strategy (16×16, 64×64, 256×256 pixel blocks); E represents the expected value; log represents the logarithmic operation; G(z) represents the generator output; I real Represents real Chinese character images. Adversarial loss continuously improves the realism of the generated Chinese characters, making them closer to the real samples.
[0134] The feature matching loss function described above ensures that the generated image and the real image have the same distribution in the feature space:
[0135]
[0136] Where, φ i ∑ represents the features of the i-th intermediate layer of the discriminator; ∑ represents the summation operation; ||·||1 represents the L1 norm. Feature matching loss helps stabilize the training process and prevent mode collapse.
[0137] The above cross-domain similarity loss function is used to establish the connection between semantic understanding and visual generation:
[0138]
[0139] Among them, ψ is the VGG-19 feature extractor, and CosSim represents the calculation of cosine similarity. This loss function ensures that the generated Chinese characters are consistent with the input semantic features, for example, ensuring that the structure of the two strokes and one stroke of the character "mountain" is correctly presented.
[0140] The above stroke constraint loss function is used to ensure the standardization of the generated glyphs:
[0141] L stroke = ||M stroke ·G(z) - M stroke ·I ref ||1
[0142] Among them, M stroke is the stroke mask matrix, which is pre-generated by existing OCR technology; I ref is the reference standard glyph. This loss ensures that the key strokes in the generated glyph follow the standard writing.
[0143] Obtain the preset weight coefficients.
[0144] In this embodiment, the generation of the above weight coefficients is not specifically limited and can be determined according to actual business requirements. For example, the weight coefficients can be determined by grid search or validation set tuning: λ1 = 1.0, λ2 = 10.0, λ3 = 5.0, λ4 = 2.0, where λ1, λ2, λ3, λ4 are the weight coefficients corresponding to the above adversarial loss function, feature matching loss function, cross-domain similarity loss function, and stroke constraint loss function, respectively.
[0145] Based on the weight coefficients, use the preset combination formula to combine the adversarial loss function, the feature matching loss function, the cross-domain similarity loss function, and the stroke constraint loss function to obtain the corresponding combined loss function.
[0146] In this embodiment, the above combination formula specifically includes:
[0147]
[0148] Among them, L total is the total loss function, and ‖w‖ represents the parameter regularization term to prevent overfitting.
[0149] Use the combined loss function as the total loss function.
[0150] This application combines multiple loss functions—adversarial loss function, feature matching loss function, cross-domain similarity loss function, and stroke constraint loss function—based on the obtained weight coefficients and combined formulas. This allows for the efficient and accurate construction of the total loss function, ensuring its accuracy. This facilitates subsequent use of the total loss function to guide the optimization direction of the hybrid generator, thereby optimizing the character generation results and ensuring that the generated Chinese characters are of high quality.
[0151] In some alternative implementations, the user information obtained is subject to user consent and complies with relevant laws and policies.
[0152] Furthermore, any software tools or components not belonging to our company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
[0153] Furthermore, existing methods for generating Chinese character glyphs have three major limitations: First, characters generated by single-modal methods (such as pure GANs) often exhibit misaligned radicals or overlapping strokes; second, existing cross-domain fusion methods mostly employ simple feature splicing, failing to establish a multi-dimensional mapping relationship between the form, meaning, and sound of Chinese characters; and third, the loss function design does not fully consider the topological constraints of Chinese character writing, resulting in generated results that do not conform to standards such as the "Standard for Stroke Order of Commonly Used Modern Chinese Characters." In addition, traditional methods perform poorly in small-sample learning scenarios, often requiring a large amount of training data to generate high-quality glyphs, which severely limits their application in areas such as rare characters and ancient scripts.
[0154] This application overcomes the above defects through the following innovations: (1) Designing a cross-attention feature fusion module to dynamically align the semantic graph encoded by BERT with the visual features extracted by CNN; (2) Introducing a cross-domain similarity loss function to force the geometric constraints of key elements such as radicals and strokes to be maintained in the latent space; (3) Combining adversarial training and transfer learning to improve the generation robustness in small sample scenarios by utilizing the prior knowledge of the pre-trained model.
[0155] This application achieves a seamless integration of semantic understanding and visual generation, solving the form-meaning separation problem in traditional methods. The system can simultaneously meet the requirements of structural standardization and visual aesthetics. Furthermore, the multi-target discriminator implements triple supervision at the stroke level, radical level, and whole character level to ensure that the generated characters conform to professional calligraphy standards. It also significantly improves the small sample learning ability, requiring only 500 samples to complete the adaptation of new fonts, reducing data requirements by 98% compared to traditional methods.
[0156] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0157] It should be emphasized that, in order to further ensure the privacy and security of the aforementioned Chinese character images, the images can also be stored in a blockchain node.
[0158] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0159] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0160] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0161] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0162] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of an artificial intelligence-based character data generation device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0163] like Figure 3 As shown, the character data generation device 300 based on artificial intelligence described in this embodiment includes: a receiving module 301, a first calling module 302, a first extraction module 303, a second extraction module 304, a fusion module 305, a conversion module 306, and an output module 307.
[0164] in:
[0165] The receiving module 301 is used to receive the input Chinese character encoding and the corresponding low-resolution character image;
[0166] The first calling module 302 is used to call a pre-built character generation model; wherein, the character generation model includes a cross-domain feature fusion module, a hybrid generator and a multi-objective discriminator, and the multi-objective discriminator is used to constrain the character generation process of the hybrid generator through a preset total loss function during the model training phase;
[0167] The first extraction module 303 is used to extract features from the Chinese character encoding based on the first feature extraction model in the cross-domain feature fusion module to obtain the corresponding semantic features.
[0168] The second extraction module 304 is used to extract features from the character image based on the second feature extraction model in the cross-domain feature fusion module to obtain the corresponding visual features.
[0169] The fusion module 305 is used to perform feature fusion processing on the semantic features and the visual features based on the cross-domain feature fusion module using a preset cross-attention fusion strategy to obtain the corresponding fused features;
[0170] The conversion module 306 is used to convert the fusion features into corresponding Chinese character images based on the fusion generator;
[0171] The output module 307 is used to perform output processing on the Chinese character image.
[0172] In some optional implementations of this embodiment, the first extraction module 303 includes:
[0173] The parsing submodule is used to parse the Chinese character encoding into corresponding Chinese characters based on a preset standard library;
[0174] The standardization submodule is used to standardize the Chinese characters to obtain the corresponding target Chinese characters.
[0175] The conversion submodule is used to convert the target Chinese character into a corresponding stroke sequence;
[0176] The encoding submodule is used to encode the stroke sequence based on the first feature extraction model in the cross-domain feature fusion module to obtain the corresponding semantic embedding representation;
[0177] The first determining submodule is used to treat the semantic embedding representation as a semantic feature corresponding to the Chinese character encoding.
[0178] In some optional implementations of this embodiment, the second extraction module 304 includes:
[0179] The first acquisition submodule is used to acquire the preset preprocessing strategy;
[0180] The first processing submodule is used to process the glyph image based on the preprocessing strategy to obtain the corresponding specified image;
[0181] The extraction submodule is used to extract features from the specified image based on the second feature extraction model in the cross-domain feature fusion module to obtain the corresponding specified features.
[0182] The second determining submodule is used to use the specified feature as the visual feature corresponding to the glyph image.
[0183] In some optional implementations of this embodiment, the hybrid generator includes a multilayer perceptron and a target convolutional neural network model, the target convolutional neural network model including an upsampling module and a downsampling module; the transformation module 306 includes:
[0184] The second acquisition submodule is used to acquire a preset random noise vector;
[0185] The splicing submodule is used to splice the random noise vector and the fused feature based on the multilayer perceptron in the hybrid generator to obtain the corresponding first feature.
[0186] The second processing submodule is used to perform downsampling processing on the first feature based on the downsampling module in the target convolutional neural network model to obtain the corresponding second feature;
[0187] The third processing submodule is used to perform progressive upsampling processing on the second feature based on the upsampling module in the target convolutional neural network model to obtain the corresponding output image.
[0188] The second determining submodule is used to treat the output image as a Chinese character image corresponding to the fused feature.
[0189] In some optional implementations of this embodiment, the output module 307 includes:
[0190] The first verification submodule is used to verify the structural accuracy of the Chinese character image;
[0191] The second verification submodule is used to perform visual authenticity verification on the Chinese character image if the Chinese character image passes the structural accuracy verification.
[0192] The third acquisition submodule is used to acquire a preset image output mode if the Chinese character image passes the visual authenticity verification.
[0193] The output submodule is used to process the Chinese character image output based on the image output method.
[0194] In some optional implementations of this embodiment, the AI-based glyph data generation device further includes:
[0195] The first acquisition module is used to acquire pre-collected sample data;
[0196] The second calling module is used to call the pre-built original model containing the target structure;
[0197] The second acquisition module is used to acquire the preset end-to-end training strategy;
[0198] The training module is used to train the original model using the sample data based on the end-to-end training strategy and the total loss function to obtain a target model that meets the construction requirements.
[0199] The first determining module is used to use the target model as the glyph generation model.
[0200] In some optional implementations of this embodiment, the AI-based glyph data generation device further includes:
[0201] The third acquisition module is used to acquire preset adversarial loss functions, feature matching loss functions, cross-domain similarity loss functions, and stroke constraint loss functions;
[0202] The fourth acquisition module is used to acquire preset weight coefficients;
[0203] The combination module is used to combine the adversarial loss function, the feature matching loss function, the cross-domain similarity loss function, and the stroke constraint loss function based on the weight coefficients using a preset combination formula to obtain the corresponding combined loss function.
[0204] The second determining module is used to use the combined loss function as the total loss function.
[0205] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed] for details. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0206] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0207] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0208] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for character data generation methods based on artificial intelligence. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.
[0209] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions of the artificial intelligence-based character data generation method.
[0210] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.
[0211] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the artificial intelligence-based glyph data generation method described above.
[0212] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
Claims
1. A method for generating glyph data based on artificial intelligence, characterized in that, Includes the following steps: Receive the input Chinese character encoding and the corresponding low-resolution character image; The pre-built character generation model is invoked; wherein, the character generation model includes a cross-domain feature fusion module, a hybrid generator and a multi-objective discriminator, and the multi-objective discriminator is used to constrain the character generation process of the hybrid generator through a preset total loss function during the model training phase; Based on the first feature extraction model in the cross-domain feature fusion module, feature extraction is performed on the Chinese character encoding to obtain the corresponding semantic features; Based on the second feature extraction model in the cross-domain feature fusion module, feature extraction is performed on the character image to obtain the corresponding visual features; Based on the cross-domain feature fusion module, a preset cross-attention fusion strategy is used to perform feature fusion processing on the semantic features and the visual features to obtain the corresponding fused features; The fusion features are converted into corresponding Chinese character images based on the fusion generator. The Chinese character image is then processed for output.
2. The method for generating glyph data based on artificial intelligence according to claim 1, characterized in that, The step of extracting features from the Chinese character encoding based on the first feature extraction model in the cross-domain feature fusion module to obtain the corresponding semantic features specifically includes: Based on a preset standard library, the Chinese character encoding is parsed into the corresponding Chinese character characters; The Chinese characters are standardized to obtain the corresponding target Chinese characters; Convert the target Chinese character into its corresponding stroke sequence; The stroke sequence is encoded based on the first feature extraction model in the cross-domain feature fusion module to obtain the corresponding semantic embedding representation; The semantic embedding is represented as a semantic feature corresponding to the Chinese character encoding.
3. The method for generating glyph data based on artificial intelligence according to claim 1, characterized in that, The step of extracting features from the glyph image based on the second feature extraction model in the cross-domain feature fusion module to obtain the corresponding visual features specifically includes: Obtain the preset preprocessing strategy; The character image is processed based on the preprocessing strategy to obtain the corresponding specified image; Based on the second feature extraction model in the cross-domain feature fusion module, feature extraction is performed on the specified image to obtain the corresponding specified features; The specified feature is used as the visual feature corresponding to the glyph image.
4. The method for generating glyph data based on artificial intelligence according to claim 1, characterized in that, The hybrid generator includes a multilayer perceptron and a target convolutional neural network model, wherein the target convolutional neural network model includes an upsampling module and a downsampling module; the step of converting the fused features into corresponding Chinese character images based on the hybrid generator specifically includes: Obtain a preset random noise vector; Based on the multilayer perceptron in the hybrid generator, the random noise vector and the fused feature are concatenated to obtain the corresponding first feature; The first feature is downsampled based on the downsampling module in the target convolutional neural network model to obtain the corresponding second feature. Based on the upsampling module in the target convolutional neural network model, the second feature is subjected to progressive upsampling processing to obtain the corresponding output image; The output image is used as the Chinese character image corresponding to the fused feature.
5. The method for generating glyph data based on artificial intelligence according to claim 1, characterized in that, The step of outputting the Chinese character image specifically includes: The structural accuracy of the Chinese character image is verified. If the Chinese character image passes the structural accuracy verification, then the visual authenticity verification of the Chinese character image is performed. If the Chinese character image passes the visual authenticity verification, then the preset image output method is obtained; The Chinese character image is output based on the image output method described above.
6. The method for generating glyph data based on artificial intelligence according to claim 1, characterized in that, Prior to the step of invoking the pre-built glyph generation model, the following is also included: Obtain pre-collected sample data; Call the pre-built original model containing the target structure; Obtain the preset end-to-end training strategy; Based on the end-to-end training strategy and the total loss function, the original model is trained using the sample data to obtain a target model that meets the construction requirements. The target model is used as the glyph generation model.
7. The method for generating glyph data based on artificial intelligence according to claim 1, characterized in that, Prior to the step of invoking the pre-built glyph generation model, the following is also included: Obtain the preset adversarial loss function, feature matching loss function, cross-domain similarity loss function, and stroke constraint loss function; Obtain the preset weight coefficients; Based on the weight coefficients, the adversarial loss function, the feature matching loss function, the cross-domain similarity loss function, and the stroke constraint loss function are combined using a preset combination formula to obtain the corresponding combined loss function; The combined loss function is used as the total loss function.
8. A character data generation device based on artificial intelligence, characterized in that, include: The receiving module is used to receive the input Chinese character encoding and the corresponding low-resolution character image; The first calling module is used to call a pre-built character generation model; wherein, the character generation model includes a cross-domain feature fusion module, a hybrid generator and a multi-objective discriminator, and the multi-objective discriminator is used to constrain the character generation process of the hybrid generator through a preset total loss function during the model training phase; The first extraction module is used to extract features from the Chinese character encoding based on the first feature extraction model in the cross-domain feature fusion module to obtain the corresponding semantic features; The second extraction module is used to extract features from the character image based on the second feature extraction model in the cross-domain feature fusion module to obtain the corresponding visual features. The fusion module is used to perform feature fusion processing on the semantic features and the visual features based on the cross-domain feature fusion module, using a preset cross-attention fusion strategy, to obtain the corresponding fused features; The conversion module is used to convert the fused features into corresponding Chinese character images based on the fusion generator; The output module is used to process the output of the Chinese character image.
9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the artificial intelligence-based glyph data generation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the artificial intelligence-based glyph data generation method as described in any one of claims 1 to 7.