Font generation model training method and related device

Through training font generating models, the description information of reference fonts and target fonts automatically generates new types of fonts that conform to the description information, solve the problem of low font design efficiency in existing technology, and achieve rapid and efficient font design.

WO2024235314A9PCT designated stage expired Publication Date: 2025-05-08LABORATORY FOR ARTIFICIAL INTELLIGENCE IN DESIGN LIMITED +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/093922
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-05-17
Filing Date
2024-05-17
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

In the prior art, designing and creating new font styles requires the participation of professional typography designers and is less efficient.

Method used

By training the font to generate models, the description information of the sample text and the target font of the reference font will automatically generate a new type of font that conforms to the description information. This method includes obtaining the description information and target text images of the sample text image and target font of the reference font. Enter this information to the font generating model to be trained, and use the confrontation generation network for training to obtain a new type of font that can automatically generate the description information. The font generating model completed by training.

Benefits of technology

The efficiency of font design is improved, so that designers can quickly generate new types of fonts that meet specific descriptions, reducing the dependence of artificial design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024093922_08052025_PF_FP_ABST
    Figure CN2024093922_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a font generation model training method and a related device. The method comprises: obtaining a sample text image of a reference font, description information of a target font and a target text image of the target font, wherein the sample text image and the target text image comprise the same text; inputting the sample text image of the reference font and the description information of the target font into a font generation model to be trained, so as to obtain a predicted text image of the target font; and according to the predicted text image of the target font and the target text image of the target font, training a generative adversarial network comprising the font generation model and a discrimination network, so as to obtain a trained font generation model.
Need to check novelty before this filing date? Find Prior Art

Description

Training method and related equipment for font generation model

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This disclosure claims priority to Chinese patent application number 202310558049.8, filed on May 17, 2023, entitled “Training method and related equipment for font generation model”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to the field of computer technology, and in particular to a font generation model training method, a font generation model training device, a font generation method, a font generation device, an electronic device, and a storage medium. Background Art

[0004] In the related art, designing and creating a new font style requires the participation of professional typesetting designers. Designers need to combine their own experience, expertise and creativity to design a font that suits given needs. This font design method is inefficient.

[0005] It should be noted that the information disclosed in the above background section is intended only to enhance understanding of the background of this disclosure and, therefore, may include information that does not constitute prior art known to persons of ordinary skill in the art. In the related art, designing and creating a new font style requires the involvement of a professional typographer. Designers must combine their experience, expertise, and creativity to design a font that suits a given requirement, resulting in a relatively inefficient font design method.

[0006] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field.

[0007] Summary of the Invention

[0008] The embodiments of the present disclosure provide a training device for a font generation model, a font generation method, a font generation device, an electronic device, and a storage medium. The font generation model trained by the method can automatically generate a new font that conforms to the font description information based on the input font description information, thereby improving the efficiency of font design.

[0009] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.

[0010] An embodiment of the present disclosure provides a method for training a font generation model, comprising: obtaining a sample text image of a reference font, descriptive information of a target font, and a target text image of the target font, wherein the sample text image and the target text image include the same text; inputting the sample text image of the reference font and the descriptive information of the target font into the font generation model to be trained to obtain a predicted text image of the target font; and training a generative adversarial network including the font generation model and a discriminant network based on the predicted text image of the target font and the target text image of the target font to obtain a trained font generation model.

[0011] In some exemplary embodiments of the present disclosure, a generative adversarial network including the font generation model and the discriminant network is trained based on the predicted text image of the target font and the target text image of the target font to obtain a trained font generation model, including: inputting the predicted text image of the target font and the target text image of the target font as input images to the discriminant network respectively to obtain the degree of truth or falsehood of the input images; wherein the degree of truth or falsehood is used to indicate the degree of truth or falsehood of the input image predicted by the discriminant network as the target text image; and training the discriminant network based on the degree of truth or falsehood of the input images.

[0012] In some exemplary embodiments of the present disclosure, a generative adversarial network including the font generation model and the discriminant network is trained based on the predicted text image of the target font and the target text image of the target font to obtain a trained font generation model, including: determining a first loss based on the difference between the pixel value of the predicted text image of the target font and the pixel value of the target text image of the target font; determining a second loss based on the difference between the feature value of the predicted text image of the target font and the feature value of the target text image of the target font; determining a third loss based on the similarity between the description information of the target font and the predicted text image of the target font; using the discriminant network to discriminate the degree of truth or falsehood of the predicted text image of the target font, and determining a fourth loss based on the degree of truth or falsehood; determining a first font attribute value corresponding to the predicted text image of the target font and a second font attribute value corresponding to the description information of the target font, and determining a fifth loss based on the first font attribute value and the second font attribute value; training the font generation model based on the first loss, the second loss, the third loss, the fourth loss and the fifth loss.

[0013] In some exemplary embodiments of the present disclosure, the second loss is determined based on the difference between the feature values ​​of the predicted text image of the target font and the feature values ​​of the target text image of the target font, including: inputting the predicted text image of the target font into a feature extraction model to obtain the feature values ​​of the predicted text image of the target font; inputting the target text image of the target font into the feature extraction model to obtain the feature values ​​of the target text image of the target font; and determining the second loss based on the difference between the feature values ​​of the predicted text image of the target font and the feature values ​​of the target text image of the target font.

[0014] In some exemplary embodiments of the present disclosure, the third loss is determined based on the similarity between the description information of the target font and the predicted text image of the target font, including: inputting the description information of the target font and the predicted text image of the target font into a trained image-text comparison model, obtaining the similarity between the description information of the target font and the predicted text image of the target font, and determining the third loss based on the similarity.

[0015] In some exemplary embodiments of the present disclosure, determining a first font attribute value corresponding to the predicted text image of the target font and a second font attribute value corresponding to the description information of the target font includes: inputting the predicted text image of the target font into a font image attribute prediction model to obtain the first font attribute value; and inputting the description information into a font description attribute prediction model to obtain the second font attribute value.

[0016] In some exemplary embodiments of the present disclosure, the font generation model includes a style description encoder, a font structure encoder and a style font generation module; wherein, the sample text image of the reference font and the description information of the target font are input into the font generation model to be trained to obtain the predicted text image of the target font, including: inputting the sample text image of the reference font into the font structure encoder for processing to obtain the image structure feature vector of the sample text image; inputting the description information of the target font into the style description encoder for processing to obtain the font style feature vector of the description information; splicing the image structure feature vector and the font style feature vector to obtain a fused feature vector; inputting the fused feature vector into the style font generation module for processing to obtain the predicted text image of the target font.

[0017] In some exemplary embodiments of the present disclosure, obtaining a sample text image of a reference font, descriptive information of a target font, and a target text image of the target font includes: obtaining font files of multiple initial fonts from a font library; determining one of the initial fonts as the reference font, determining initial fonts other than the reference font as the target font, and obtaining descriptive information of the target font; generating a sample text image of the reference font based on the font file of the reference font, and generating a target text image of the target font based on the font file of the target font.

[0018] An embodiment of the present disclosure provides a font generation method, comprising: obtaining a sample text image of a reference font and description information of a predicted font; inputting the sample text image of the reference font and the description information of the predicted font into a trained font generation model to obtain a predicted text image of the predicted font; wherein the font generation model is trained according to any of the above methods.

[0019] In some exemplary embodiments of the present disclosure, the above method also includes: determining a third font attribute value corresponding to the predicted text image of the predicted font; obtaining a fourth font attribute value obtained by adjusting the third font attribute value; inputting the predicted text image of the predicted font and the fourth font attribute value into a font attribute adjustment model to obtain the adjusted predicted text image of the predicted font.

[0020] An embodiment of the present disclosure provides a training device for a font generation model, comprising: an acquisition module for acquiring a sample text image of a reference font, descriptive information of a target font, and a target text image of the target font, wherein the sample text image and the target text image include the same text; an acquisition module for inputting the sample text image of the reference font and the descriptive information of the target font into a font generation model to be trained to obtain a predicted text image of the target font; and a training module for training a generative adversarial network including the font generation model and a discriminant network based on the predicted text image of the target font and the target text image of the target font to obtain a trained font generation model.

[0021] An embodiment of the present disclosure provides a font generation device, comprising: an acquisition module for acquiring image-text groups of multiple objects, each image-text group comprising an object image and text information corresponding to the object image; an acquisition module for inputting the image-text groups of each object into a multimodal data representation model trained by any of the above-described methods to obtain multimodal feature data of each object.

[0022] An embodiment of the present disclosure provides an electronic device, comprising: at least one processor; and a storage device for storing at least one program, wherein when the at least one program is executed by the at least one processor, the at least one processor implements any of the above-mentioned font generation model training methods or font generation methods.

[0023] An embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements any of the above-mentioned font generation model training methods or font generation methods.

[0024] The training method of the font generation model provided by the embodiment of the present disclosure inputs the sample text image of the reference font and the description information of the target font into the font generation model to be trained to obtain the predicted text image of the target font; the adversarial generative network including the font generation model and the discriminant network is trained based on the predicted text image of the target font and the target text image of the target font, and the trained font generation model obtained in this way can automatically generate a new font that meets the font description information based on the input font description information, thereby improving the efficiency of font design; in addition, the sample text image of the reference font and the target text image of the target font include the same text, which can enable the trained model to convert the font while ensuring that the text content remains unchanged, thereby improving the accuracy of the font generation model.

[0025] It should be understood that the general description above and the detailed descriptions below are merely exemplary and explanatory and do not limit the present disclosure. The present disclosure provides a font generation model training device, a font generation method, a font generation device, an electronic device, and a storage medium. The font generation model trained by this method can automatically generate new fonts that conform to the font description information based on the input font description information, thereby improving the efficiency of font design.

[0026] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.

[0027] An embodiment of the present disclosure provides a method for training a font generation model, comprising: obtaining a sample text image of a reference font, descriptive information of a target font, and a target text image of the target font, wherein the sample text image and the target text image include the same text; inputting the sample text image of the reference font and the descriptive information of the target font into the font generation model to be trained to obtain a predicted text image of the target font; and training a generative adversarial network including the font generation model and a discriminant network based on the predicted text image of the target font and the target text image of the target font to obtain a trained font generation model.

[0028] In some exemplary embodiments of the present disclosure, a generative adversarial network including the font generation model and the discriminant network is trained based on the predicted text image of the target font and the target text image of the target font to obtain a trained font generation model, including: inputting the predicted text image of the target font and the target text image of the target font as input images to the discriminant network respectively to obtain the degree of truth or falsehood of the input images; wherein the degree of truth or falsehood is used to indicate the degree of truth or falsehood of the input image predicted by the discriminant network as the target text image; and training the discriminant network based on the degree of truth or falsehood of the input images.

[0029] In some exemplary embodiments of the present disclosure, a generative adversarial network including the font generation model and the discriminant network is trained based on the predicted text image of the target font and the target text image of the target font to obtain a trained font generation model, including: determining a first loss based on the difference between the pixel value of the predicted text image of the target font and the pixel value of the target text image of the target font; determining a second loss based on the difference between the feature value of the predicted text image of the target font and the feature value of the target text image of the target font; determining a third loss based on the similarity between the description information of the target font and the predicted text image of the target font; using the discriminant network to discriminate the degree of truth or falsehood of the predicted text image of the target font, and determining a fourth loss based on the degree of truth or falsehood; determining a first font attribute value corresponding to the predicted text image of the target font and a second font attribute value corresponding to the description information of the target font, and determining a fifth loss based on the first font attribute value and the second font attribute value; training the font generation model based on the first loss, the second loss, the third loss, the fourth loss and the fifth loss.

[0030] In some exemplary embodiments of the present disclosure, the second loss is determined based on the difference between the feature values ​​of the predicted text image of the target font and the feature values ​​of the target text image of the target font, including: inputting the predicted text image of the target font into a feature extraction model to obtain the feature values ​​of the predicted text image of the target font; inputting the target text image of the target font into the feature extraction model to obtain the feature values ​​of the target text image of the target font; and determining the second loss based on the difference between the feature values ​​of the predicted text image of the target font and the feature values ​​of the target text image of the target font.

[0031] In some exemplary embodiments of the present disclosure, the third loss is determined based on the similarity between the description information of the target font and the predicted text image of the target font, including: inputting the description information of the target font and the predicted text image of the target font into a trained image-text comparison model, obtaining the similarity between the description information of the target font and the predicted text image of the target font, and determining the third loss based on the similarity.

[0032] In some exemplary embodiments of the present disclosure, determining a first font attribute value corresponding to the predicted text image of the target font and a second font attribute value corresponding to the description information of the target font includes: inputting the predicted text image of the target font into a font image attribute prediction model to obtain the first font attribute value; and inputting the description information into a font description attribute prediction model to obtain the second font attribute value.

[0033] In some exemplary embodiments of the present disclosure, the font generation model includes a style description encoder, a font structure encoder and a style font generation module; wherein, the sample text image of the reference font and the description information of the target font are input into the font generation model to be trained to obtain the predicted text image of the target font, including: inputting the sample text image of the reference font into the font structure encoder for processing to obtain the image structure feature vector of the sample text image; inputting the description information of the target font into the style description encoder for processing to obtain the font style feature vector of the description information; splicing the image structure feature vector and the font style feature vector to obtain a fused feature vector; inputting the fused feature vector into the style font generation module for processing to obtain the predicted text image of the target font.

[0034] In some exemplary embodiments of the present disclosure, obtaining a sample text image of a reference font, descriptive information of a target font, and a target text image of the target font includes: obtaining font files of multiple initial fonts from a font library; determining one of the initial fonts as the reference font, determining initial fonts other than the reference font as the target font, and obtaining descriptive information of the target font; generating a sample text image of the reference font based on the font file of the reference font, and generating a target text image of the target font based on the font file of the target font.

[0035] An embodiment of the present disclosure provides a font generation method, comprising: obtaining a sample text image of a reference font and description information of a predicted font; inputting the sample text image of the reference font and the description information of the predicted font into a trained font generation model to obtain a predicted text image of the predicted font; wherein the font generation model is trained according to any of the above methods.

[0036] In some exemplary embodiments of the present disclosure, the above method also includes: determining a third font attribute value corresponding to the predicted text image of the predicted font; obtaining a fourth font attribute value obtained by adjusting the third font attribute value; inputting the predicted text image of the predicted font and the fourth font attribute value into a font attribute adjustment model to obtain the adjusted predicted text image of the predicted font.

[0037] An embodiment of the present disclosure provides a training device for a font generation model, comprising: an acquisition module for acquiring a sample text image of a reference font, descriptive information of a target font, and a target text image of the target font, wherein the sample text image and the target text image include the same text; an acquisition module for inputting the sample text image of the reference font and the descriptive information of the target font into a font generation model to be trained to obtain a predicted text image of the target font; and a training module for training a generative adversarial network including the font generation model and a discriminant network based on the predicted text image of the target font and the target text image of the target font to obtain a trained font generation model.

[0038] An embodiment of the present disclosure provides a font generation device, comprising: an acquisition module for acquiring image-text groups of multiple objects, each image-text group comprising an object image and text information corresponding to the object image; an acquisition module for inputting the image-text groups of each object into a multimodal data representation model trained by any of the above-described methods to obtain multimodal feature data of each object.

[0039] An embodiment of the present disclosure provides an electronic device, comprising: at least one processor; and a storage device for storing at least one program, wherein when the at least one program is executed by the at least one processor, the at least one processor implements any of the above-mentioned font generation model training methods or font generation methods.

[0040] An embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements any of the above-mentioned font generation model training methods or font generation methods.

[0041] The training method of the font generation model provided by the embodiment of the present disclosure inputs the sample text image of the reference font and the description information of the target font into the font generation model to be trained to obtain the predicted text image of the target font; the adversarial generative network including the font generation model and the discriminant network is trained based on the predicted text image of the target font and the target text image of the target font, and the trained font generation model obtained in this way can automatically generate a new font that meets the font description information based on the input font description information, thereby improving the efficiency of font design; in addition, the sample text image of the reference font and the target text image of the target font include the same text, which can enable the trained model to convert the font while ensuring that the text content remains unchanged, thereby improving the accuracy of the font generation model.

[0042] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0044] FIG1 is a schematic diagram showing an exemplary system architecture to which a font generation model training method or a font generation method according to an embodiment of the present disclosure can be applied.

[0045] FIG2 is a flowchart of a method for training a font generation model in an exemplary embodiment of the present disclosure.

[0046] FIG3 is a schematic diagram of a font generation model and a discriminant network in an exemplary embodiment of the present disclosure.

[0047] FIG4 is a schematic diagram of a training image-text comparison model in an exemplary embodiment of the present disclosure.

[0048] FIG5 is a flowchart of a font generation method in another exemplary embodiment of the present disclosure.

[0049] FIG6 is a block diagram of a training apparatus for a font generation model in an exemplary embodiment of the present disclosure.

[0050] FIG. 7 is a block diagram of a font generating apparatus in an exemplary embodiment of the present disclosure.

[0051] FIG8 is a schematic structural diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0052] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. Like reference numerals in the drawings represent like or similar parts, and thus repetitive description thereof will be omitted.

[0053] The features, structures or characteristics described in the present disclosure may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0054] The accompanying drawings are merely schematic illustrations of the present disclosure. Identical reference numerals in the drawings denote identical or similar components, and thus their repeated descriptions will be omitted. Some of the block diagrams shown in the accompanying drawings do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in at least one hardware module or integrated circuit, or in different networks and / or processor devices and / or microcontroller devices.

[0055] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all content and steps, nor must they be executed in the order described. For example, some steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.

[0056] In addition, in the description of the present disclosure, the terms "a", "an", "the", "said" and "at least one" are used to indicate the presence of at least one element or component; the terms "comprising", "including" and "having" are used to express open-ended inclusion and mean that additional elements or components may exist in addition to the listed elements or components; the terms "first", "second" and "third" etc. are used only as labels and are not intended to limit the quantity of their objects.

[0057] FIG1 is a schematic diagram showing an exemplary system architecture to which a font generation model training method or a font generation method according to an embodiment of the present disclosure can be applied.

[0058] As shown in Figure 1, the system architecture may include a server 101, a network 102, a terminal device 103, a terminal device 104, and a terminal device 105. The network 102 is used as a medium for providing a communication link between the terminal device 103, the terminal device 104, or the terminal device 105 and the server 101. The network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0059] Server 101 may be a server that provides various services, such as a background management server that supports devices operated by users using terminal device 103, terminal device 104, or terminal device 105. The background management server may analyze and process received data such as requests, and feed back the processing results to terminal device 103, terminal device 104, or terminal device 105.

[0060] Terminal device 103, terminal device 104 and terminal device 105 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a wearable smart device, a virtual reality device, an augmented reality device, etc., but are not limited thereto.

[0061] In an embodiment of the present disclosure, the server 101 can: obtain a sample text image of a reference font, description information of a target font, and a target text image of the target font, wherein the sample text image and the target text image include the same text; input the sample text image of the reference font and the description information of the target font into a font generation model to be trained to obtain a predicted text image of the target font; train a generative adversarial network including a font generation model and a discriminant network based on the predicted text image of the target font and the target text image of the target font to obtain a trained font generation model.

[0062] In an embodiment of the present disclosure, the server 101 can: obtain a sample text image of a reference font and description information of a predicted font; input the sample text image of the reference font and description information of the predicted font into a trained font generation model to obtain a predicted text image of the predicted font; wherein the font generation model is trained according to the above method.

[0063] It should be understood that the number of terminal devices 103, terminal devices 104, terminal devices 105, networks 102 and servers 101 in Figure 1 are merely schematic. Server 101 can be a physical server, a server cluster composed of multiple servers, or a cloud server. Depending on actual needs, it can have any number of terminal devices, networks and servers.

[0064] Below, each step of the method for training a font generation model in an exemplary embodiment of the present disclosure will be described in more detail with reference to the accompanying drawings and embodiments.

[0065] Figure 2 is a flow chart of a method for training a font generation model in an exemplary embodiment of the present disclosure. The method provided in the embodiment of the present disclosure can be executed by any electronic device with computing processing capabilities, such as the server or terminal device shown in Figure 1, but the present disclosure is not limited thereto.

[0066] As shown in FIG2 , the training method of the font generation model provided by the embodiment of the present disclosure may include the following steps.

[0067] In step S202 , a sample text image of a reference font, description information of a target font, and a target text image of the target font are obtained, wherein the sample text image and the target text image include the same text.

[0068] In the disclosed embodiment, text refers to Chinese characters expressed in Chinese, such as "love", "you", "I", "he", etc.; font refers to the style of text, such as Songti, Heiti, custom font, etc. The reference font and target font are both a type of font, and there can be multiple target fonts; those skilled in the art can determine the reference font and target font according to actual needs, for example, the reference font is Songti, and the target font is Chaozhiti; the description information of the target font refers to the text information used to describe the font features of the target font, for example, referring to Figure 3, the description information 302 of the target font can be "The Chaozhiti font has a square shape with an oblique stroke, and the strokes are clear and concise..."

[0069] In the embodiment of the present disclosure, the sample text image and the target text image include the same text. Referring to Figure 3, for example, the sample text image 301 of the reference font is the image of the word "love" in Song font, and the target text image 303 of the target font is the image of the word "love" in Chaozhi font.

[0070] In an exemplary embodiment, obtaining a sample text image of a reference font, descriptive information of a target font, and a target text image of the target font includes: obtaining font files of multiple initial fonts from a font library; determining one of the initial fonts as a reference font, determining initial fonts other than the reference font as target fonts, and obtaining descriptive information of the target font; generating a sample text image of the reference font based on the font file of the reference font, and generating a target text image of the target font based on the font file of the target font.

[0071] For example, 400 existing TTF (True Type Font) font files of initial fonts and the Chinese description of font features corresponding to each initial font can be obtained from the font library; Songti is determined as the reference font, and the initial fonts other than Songti are determined as the target fonts; the top 200 most commonly used Songti Chinese character images are automatically generated based on the TTF font file of Songti, and the top 200 most commonly used target font Chinese character images are automatically generated based on the TTF font file of the target font.

[0072] In the embodiment of the present disclosure, after the sample text image and the target text image are generated, the size of the image may be adjusted, the image may be grayscaled, normalized, and the like.

[0073] In step S204 , the sample text image of the reference font and the description information of the target font are input into the font generation model to be trained to obtain a predicted text image of the target font.

[0074] In the embodiment of the present disclosure, a generative adversarial network (GAN) can be established. The generative adversarial network includes a generator and a discriminator. The generator network can process the input data to obtain an output result. The output result will try to imitate the target text image of the target font. The input of the adversarial network is the target text image of the target font or the output result of the generator network (i.e., the predicted text image of the target font). Its purpose is to separate the target text image and the predicted text image output by the generator network as much as possible. The generator network needs to output an output result similar to the target text image, so that the output result is not distinguished by the discriminator network as much as possible. The generator network and the adversarial network compete with each other, and the parameters of the two are constantly adjusted, which ultimately makes it impossible for the discriminator network to determine whether the output result of the generator network is true. At this time, the output result of the generator network is close to the target text image.

[0075] In the disclosed embodiment, the font generation model may be a generation network; referring to FIG3 , the sample text image 301 of the reference font and the description information 302 of the target font may be input into the font generation model to be trained for processing. The font generation model to be trained may adjust the sample text image 301 of the reference font according to the description information 302 of the target font, and output a predicted text image 304 of the target font.

[0076] In an exemplary embodiment, the font generation model may include a style description encoder 307, a font structure encoder 305 and a style font generation module 306; wherein, inputting the sample text image of the reference font and the description information of the target font into the font generation model to be trained to obtain the predicted text image of the target font may include: inputting the sample text image 301 of the reference font into the font structure encoder 305 for processing to obtain the image structure feature vector 308 of the sample text image; inputting the description information 302 of the target font into the style description encoder 307 for processing to obtain the font style feature vector 309 of the description information; splicing the image structure feature vector 308 and the font style feature vector 309 to obtain a fused feature vector 310; inputting the fused feature vector 310 into the style font generation module 306 for processing to obtain the predicted text image 304 of the target font.

[0077] In the embodiment of the present disclosure, a language pre-training model based on a transformer can be pre-trained to learn the mapping relationship between font style description information and font style feature vectors, and the trained pre-training model can be used as the style description encoder 307 in the font generation model; the style description encoder 307 can convert the input target font description information 302 into a font style feature vector 309 (also called a style description vector).

[0078] In the embodiment of the present disclosure, the font structure encoder 305 can be a downsampling structure composed of multiple convolution modules, which is used to reduce the dimension of the input reference font sample text image 301 to the feature vector space to obtain the image structure feature vector 308; after the image structure feature vector 308 and the font style feature vector 309 are spliced, the dimension is reduced through a fully connected network to generate a fixed-length fusion feature vector 310; the style font generation module 306 can be composed of multiple deconvolution modules, which is used to restore the feature vector to an image and generate an image with the same pixel size as the sample text image of the reference font; in the upsampling process, the feature information on the corresponding scale can be introduced into the upsampling or deconvolution process through skip connection to ensure information transmission in the model, make the model easier to train, and reduce the possibility of overfitting.

[0079] In step S206, a generative adversarial network including a font generation model and a discriminant network is trained based on the predicted text image of the target font and the target text image of the target font to obtain a trained font generation model.

[0080] In the embodiment of the present disclosure, a generative adversarial network including a font generation model and a discriminant network can be trained based on the predicted text image 304 of the target font and the target text image 303 of the target font; specifically, the discriminant network can be first trained based on the predicted text image 304 of the target font and the target text image 303 of the target font, so that the discriminant network can identify whether the input image is a predicted text image category (which can be understood as "false") or a target text image category (which can be understood as "true"). It can be understood that the predicted text image 304 of the target font belongs to the predicted text image category, and the target text image 303 of the target font belongs to the target text image category; then, the font generation model is trained based on the predicted text image 304 of the target font and the predicted category output by the discriminant network, so that the predicted text image 304 generated by the font generation model cannot be identified as a predicted text image category by the discriminant network (it can be understood that the predicted text image generated by the font generation model can be indistinguishable from the real thing), thereby obtaining a trained font generation model.

[0081] In an exemplary embodiment, training a generative adversarial network including a font generation model and a discriminant network based on a predicted text image of a target font and a target text image of a target font to obtain a trained font generation model may include: inputting the predicted text image of the target font and the target text image of the target font as input images to the discriminant network respectively to obtain the degree of truth or falsehood of the input images; wherein the degree of truth or falsehood is used to indicate the degree of truth or falsehood of the input image predicted by the discriminant network as the target text image, for example, if the discriminant network predicts the input image as the target text image category, it is considered that the discriminant network predicts the input image as "true"; for example, if the discriminant network predicts the input image as the predicted text image category, it is considered that the discriminant network predicts the input image as "false"; and training the discriminant network according to the degree of truth or falsehood of the input image.

[0082] In the disclosed embodiment, the discriminant network may be a structure composed of multiple convolutional layers, which is used to judge the authenticity of the images generated by the generative network.

[0083] Specifically, when the predicted text image 304 of the target font generated by the font generation model is input into the discriminant network as an input image, the predicted category of the discriminant network for the predicted text image 304 is obtained. At this time, the predicted category may be the predicted text image category (which can be understood as "false") or the target text image category (which can be understood as "true"). At this time, the actual category of the predicted text image 304 is the predicted text image category (which can be understood as "false"); if the predicted category predicted by the discriminant network is the predicted text image category, it means that the discriminant network has recognized it correctly and there is no need to adjust the parameters of the discriminant network; if the predicted category predicted by the discriminant network is the target text image category, it means that the discriminant network has recognized it incorrectly, and the parameters of the discriminant network are adjusted until the discriminant network can correctly recognize the category of the input image.

[0084] Similarly, when the target text image 303 of the target font is input as an input image to the discriminant network, the predicted category of the target text image 303 by the discriminant network is obtained. At this time, the predicted category may be the predicted text image category (which can be understood as "false") or the target text image category (which can be understood as "true"). At this time, the actual category of the target text image 303 is the target text image category (which can be understood as "true"); if the predicted category predicted by the discriminant network is the target text image category, it means that the discriminant network has recognized it correctly and there is no need to adjust the parameters of the discriminant network; if the predicted category predicted by the discriminant network is the predicted text image category, it means that the discriminant network has recognized it incorrectly, and the parameters of the discriminant network are adjusted until the discriminant network can correctly recognize the category of the input image.

[0085] In an embodiment of the present disclosure, when training the discriminant network, a first adversarial loss function (for example, a cross entropy loss function) can be determined based on the predicted category of the input image and the actual category of the input image, and the discriminant network is trained using the Adam optimizer based on the first adversarial loss function.

[0086] In an exemplary embodiment, a generative adversarial network including a font generation model and a discriminant network is trained based on a predicted text image of a target font and a target text image of a target font to obtain a trained font generation model, including: determining a first loss based on a difference between a pixel value of the predicted text image of the target font and a pixel value of the target text image of the target font; determining a second loss based on a difference between a feature value of the predicted text image of the target font and a feature value of the target text image of the target font; determining a third loss based on a similarity between the description information of the target font and the predicted text image of the target font; using a discriminant network to discriminate the degree of truth or falsehood of the predicted text image of the target font, and determining a fourth loss based on the degree of truth or falsehood; determining a first font attribute value corresponding to the predicted text image of the target font and a second font attribute value corresponding to the description information of the target font, and determining a fifth loss based on the first font attribute value and the second font attribute value; and training the font generation model based on the first loss, the second loss, the third loss, the fourth loss, and the fifth loss.

[0087] Specifically, the first loss may be, for example, the difference between the pixel value of the predicted text image of the target font and the pixel value of the target text image of the target font calculated by the L1 loss function.

[0088] Specifically, the second loss can be, for example, the difference between the feature values ​​of the predicted text image of the target font and the feature values ​​of the target text image of the target font calculated by the contextual loss function, wherein the feature values ​​of the predicted text image of the target font and the feature values ​​of the target text image of the target font can be respectively extracted by a pre-trained feature extraction model.

[0089] In an exemplary embodiment, the second loss is determined based on the difference between the feature values ​​of the predicted text image of the target font and the feature values ​​of the target text image of the target font, including: inputting the predicted text image of the target font into a feature extraction model to obtain the feature values ​​of the predicted text image of the target font; inputting the target text image of the target font into a feature extraction model to obtain the feature values ​​of the target text image of the target font; and determining the second loss based on the difference between the feature values ​​of the predicted text image of the target font and the feature values ​​of the target text image of the target font.

[0090] The feature extraction model may be a pre-trained VGG (Visual Geometry Group Network) model, but the present disclosure is not limited thereto.

[0091] Specifically, the third loss can be determined, for example, by a trained image-text comparison model. The image-text comparison model can be, for example, a CLIP (Contrastive Language-Image Pretraining) model. The training process of the image-text comparison model is introduced below.

[0092] In an exemplary embodiment, the third loss is determined based on the similarity between the description information of the target font and the predicted text image of the target font, including: inputting the description information of the target font and the predicted text image of the target font into a trained image-text comparison model, obtaining the similarity between the description information of the target font and the predicted text image of the target font, and determining the third loss based on the similarity.

[0093] FIG4 is a schematic diagram of a training image-text comparison model in an exemplary embodiment of the present disclosure.

[0094] Specifically, referring to FIG4 , the image-text contrast model (e.g., CLIP model) may include a text encoder 401 and an image encoder 402. The image encoder 402 uses a multi-layer convolutional neural network to downsample the image information, and the output vector is passed through a fully connected layer to generate an image vector (T1, T2, T3, ..., T N), the text encoder 401 uses a transformer-based language pre-training model (such as MacBERT (Masked Language Modeling as Correction Bidirectional Encoder Representation from Transformers, a bidirectional conversion encoder based on an error-correcting masked language model) model), and the output vector is passed through a fully connected layer to generate a text vector (I1, I2, I3, ..., I N ); performing a dot product operation on the image vector and the text vector to generate a result matrix 403.

[0095] Specifically, during the training phase of the CLIP model, the MSE (mean square) loss function can be used to calculate the loss value between the generated result matrix and a diagonal matrix whose diagonals are all 1. Based on the loss value, the Adam optimizer is used to update the weights of the text encoder 401 and the image encoder 402 until the loss value no longer decreases or reaches the expected loss value. After the CLIP model training is completed, the trained CLIP model can be used as a loss function of a generative adversarial network. The main task is to determine the similarity between the description information of the target font and the similarity of the predicted text image of the target font. The larger the value of the similarity, the more dissimilar it is, and the smaller the value, the more similar it is.

[0096] Specifically, the third loss is determined according to the similarity. The similarity may be directly used as the third loss, or the third loss may be obtained by multiplying the similarity with a weight.

[0097] Specifically, the fourth loss can be determined, for example, based on the degree of truth or falsehood of the predicted text image of the target font determined by the discriminant network; for example, the predicted text image of the target font is input into the discriminant network to obtain the predicted category of the predicted text image. For example, the discriminant network predicts the input image as the target text image category, that is, it is considered that the discriminant network predicts the input image as "true"; for example, the discriminant network predicts the input image as the predicted text image category, that is, it is considered that the discriminant network predicts the input image as "false"; the fourth loss is determined based on the difference between the predicted category of the predicted text image and the target text image category.

[0098] Specifically, after the stage training of the discriminant network is completed, the predicted text image 304 of the target font generated by the font generation model is input into the discriminant network to obtain the predicted category of the predicted text image 304 predicted by the discriminant network. If the predicted category is the predicted text image category, it means that the predicted text image generated by the font generation model has not been able to deceive the discriminant model. At this time, the model parameters of the font generation model need to be adjusted until the predicted text image 304 generated by the adjusted font generation model is input into the discriminant network. The predicted category output by the discriminant network is the target text image category, which makes the discriminant model unable to correctly identify the predicted text image, indicating that the predicted text image generated by the font generation model is more accurate.

[0099] In an embodiment of the present disclosure, when training a font generation model, a second adversarial loss function can be determined based on the predicted category of the predicted text image and the target text image category, and the discriminant network can be trained using the second adversarial loss function.

[0100] Specifically, the fifth loss may be determined, for example, based on a font attribute value corresponding to the predicted text image of the target font (hereinafter referred to as a “first font attribute value”) and a font attribute value corresponding to the description information of the target font (hereinafter referred to as a “second font attribute value”);

[0101] In an exemplary embodiment, determining a first font attribute value corresponding to a predicted text image of a target font and a second font attribute value corresponding to description information of the target font includes: inputting the predicted text image of the target font into a font image attribute prediction model to obtain the first font attribute value; and inputting the description information into a font description attribute prediction model to obtain the second font attribute value.

[0102] In the embodiment of the present disclosure, the font image attribute prediction model and the font description attribute prediction model can both be trained neural network models, wherein the font image attribute prediction model can predict the attribute values ​​of each preset attribute for the input image, and the font description attribute prediction model can predict the attribute values ​​of each preset attribute for the input text.

[0103] The preset attributes refer to font attributes, such as line thickness, italics, artistic sense, realism, complexity, formality, happiness, freshness, etc. Each preset attribute corresponds to an attribute value, and the attribute value can range from 0 to 100.

[0104] The following introduces the training process of the font image attribute prediction model and the font description attribute prediction model.

[0105] For the font image attribute prediction model, for example, the font-attributes dataset can be used as a training dataset, which includes multiple fonts and multiple attributes corresponding to each labeled font and the attribute values ​​of each attribute; during the training process, the font image attribute prediction model uses a multi-layer convolutional neural network to downsample the input character image, output the attribute feature vector corresponding to each preset attribute, and process each dimension of the attribute feature vector through the softmax (activation) layer to obtain the predicted attribute value corresponding to each preset attribute; the loss is calculated based on the predicted attribute value and the labeled attribute value, and the parameters of the font image attribute prediction model are updated based on the loss to complete the training of the font image attribute prediction model.

[0106] For the font description attribute prediction model, for example, a font dataset including tff files of multiple fonts and their corresponding Chinese descriptions can be used as a training dataset, and a preset number (for example, 200) of commonly used Chinese character images are generated using the tff files of each font, and these images are respectively input into the trained font image attribute prediction model, and the attribute values ​​of the preset attributes of the font of each image are predicted, and the average of the predicted attribute values ​​is taken as the attribute value of this font, and the attribute value of this font is used as the label of the font description attribute prediction model; the Chinese description corresponding to the font is input into the font description attribute prediction model to obtain the attribute values ​​of each preset attribute, and the loss value between the attribute value of each preset attribute obtained by the font description attribute prediction model and the attribute value of this font is calculated, and the weight of the font description attribute prediction model is updated according to the loss value to complete the training of the font description attribute prediction model.

[0107] The font description attribute prediction model in the embodiment of the present disclosure may use the MacBERT model, but the present disclosure is not limited thereto.

[0108] After the font image attribute prediction model and the font description attribute prediction model are trained, the predicted text image of the target font is input into the font image attribute prediction model to obtain a first font attribute value; the description information is input into the font description attribute prediction model to obtain a second font attribute value; and the difference between the first font attribute value and the second font attribute value is used as the fifth loss.

[0109] In the embodiment of the present disclosure, the font generation model is trained according to the first loss, the second loss, the third loss, the fourth loss and the fifth loss. The sum of the first loss, the second loss, the third loss, the fourth loss and the fifth loss can be taken as the total loss to train the font generation model; or the first loss, the second loss, the third loss, the fourth loss and the fifth loss can be multiplied by their respective weights to obtain the total loss to train the font generation model; in the process of training the font generation model, the model parameters of the font generation model are adjusted so that the total loss meets the preset conditions to obtain a trained font generation model.

[0110] The training method of the font generation model provided by the embodiment of the present disclosure inputs the sample text image of the reference font and the description information of the target font into the font generation model to be trained to obtain the predicted text image of the target font; the adversarial generative network including the font generation model and the discriminant network is trained based on the predicted text image of the target font and the target text image of the target font, and the trained font generation model obtained in this way can automatically generate a new font that meets the font description information based on the input font description information, thereby improving the efficiency of font design; in addition, the sample text image of the reference font and the target text image of the target font include the same text, which can enable the trained model to convert the font while ensuring that the text content remains unchanged, thereby improving the accuracy of the font generation model.

[0111] In addition, during the model training process, both the reference font and the target font can be derived from the font library. Using the target text image of the target font as the label for model training can reduce the labeling of training data during the model training process.

[0112] Figure 5 is a flowchart of a font generation method in an exemplary embodiment of the present disclosure. The method provided in the embodiment of the present disclosure can be executed by any electronic device with computing processing capabilities, such as the server or terminal device shown in Figure 1, but the present disclosure is not limited thereto.

[0113] As shown in FIG5 , the font generation method provided by the embodiment of the present disclosure may include the following steps.

[0114] In step S502, a sample character image of a reference font and description information of a predicted font are obtained.

[0115] In actual applications, the reference font can be consistent with the reference font during model training. For example, if Songti is used as the reference font during model training, then Songti should also be used as the reference font in actual applications.

[0116] In practical applications, the predicted font refers to a font that does not exist in the existing font library and that the user wants to create, and the description information of the predicted font refers to text information input by the user that describes the font features of the predicted font.

[0117] In step S504, the sample text image of the reference font and the description information of the predicted font are input into the trained font generation model to obtain the predicted text image of the predicted font; wherein the font generation model is trained according to any of the above embodiments.

[0118] In an embodiment of the present disclosure, a sample text image of a reference font and description information of a predicted font are input into a trained font generation model. The font generation model adjusts the sample text image of the reference font according to the description information of the predicted font and outputs a predicted text image of the predicted font.

[0119] The font generation method provided by the embodiment of the present disclosure automatically generates a new font that meets the description information based on the sample text image of the reference font and the description information of the predicted font, thereby improving the efficiency of font design.

[0120] The disclosed embodiments may also provide a visual interactive interface, which may include a first input box, a second input box, and a generate button. The user may input the Chinese characters to be generated through the first input box, input the Chinese description of the font to be generated (i.e., the predicted font) through the second input box, and then click the generate button to generate a picture of the Chinese characters in the predicted font.

[0121] The visual interactive interface provided by the embodiment of the present disclosure can also display the attribute values ​​of various preset attributes of the generated predicted font, so that the user can make more detailed adjustments to the generated predicted font, thereby improving the user experience.

[0122] In an exemplary embodiment, the method may further include: determining a third font attribute value corresponding to the predicted text image of the predicted font; obtaining a fourth font attribute value obtained by adjusting the third font attribute value; inputting the predicted text image of the predicted font and the fourth font attribute value into a font attribute adjustment model to obtain an adjusted predicted text image of the predicted font.

[0123] Specifically, the predicted text image of the predicted font can be input into the above-mentioned trained font image attribute prediction model to obtain the attribute values ​​of the predicted text image of the predicted font under each preset attribute (called the "third font attribute value"), and the obtained third font attribute value is displayed in the visual interactive interface; the user can modify the unsatisfactory part according to needs, obtain the font attribute value obtained by adjusting the third font attribute value (called the "fourth font attribute value"), input the predicted text image of the predicted font and the fourth font attribute value into the font attribute adjustment model, and obtain the adjusted predicted text image of the predicted font.

[0124] For example, if the user is not satisfied with the font thickness of the generated predicted text image, the attribute value of the font thickness of the currently generated predicted text image is 63, and the user can manually adjust the attribute value to 89; the attribute value of the adjusted font thickness, the attribute values ​​of other unadjusted attributes, and the currently generated predicted text image are input into the font attribute adjustment model to obtain the predicted text image of the adjusted predicted font.

[0125] The font attribute adjustment model in the embodiment of the present disclosure may include an attribute embedding module, a Chinese character embedding module, a fully connected layer and a generator.

[0126] Among them, the attribute embedding module can embed each preset attribute into a high-dimensional space to obtain an attribute vector; the attribute value range of each preset attribute can be, for example, between 0 and 100, and the attribute value can be normalized to between 0 and 1 as the weight of the attribute vector; the Chinese character embedding module can map Chinese characters to a high-dimensional space to obtain a character vector, which is used as a conditional vector for generating predicted character images to control the Chinese characters that the model needs to generate; the attribute vectors of all preset attributes and their weights are spliced ​​together after calculation as the input of the fully connected layer, and the multi-layer fully connected layer reduces the dimension of the input vector and outputs the latent space vector. The latent space vector and the character vector are spliced ​​together to form a conditional latent space vector. The conditional latent space vector will pass through a generator composed of multiple layers of deconvolution to generate a predicted text image corresponding to the attribute value and the Chinese character. The discriminator consists of multiple convolutional layers to determine whether the input image is a real image or a generated image.

[0127] Among them, the generator part can be trained using a loss function consisting of a GAN (Generative Adversarial Networks) loss function, an MSE loss function, a Contextual loss function, an attribute prediction loss function and a character classification loss function, wherein the GAN loss function is used to calculate the generative adversarial network loss, the MSE loss function is used to calculate the pixel value difference between the character image generated by the generator and the target image, the Contextual loss function is used to calculate the difference between the feature values ​​of the character image generated by the generator and the target image, the attribute prediction loss function is used to calculate the loss between the font attribute value corresponding to the character image generated by the generator and the font attribute value corresponding to the description information, and the character classification loss function is used to calculate the similarity between the character category to which the character image generated by the generator belongs and the target character category; the discriminator part can be trained using the GAN loss function, and the trained attribute embedding module, Chinese character embedding module, fully connected layer and generator are used as the above-mentioned font image attribute adjustment model.

[0128] The font generation method provided by the embodiments of the present disclosure can be applied to font design scenarios, such as poster font design, advertising font design, film and television font design, and outer packaging font design. Using the font generation method provided by the embodiments of the present disclosure, designers only need to use Chinese to describe the current font design requirements and confirm the Chinese characters that need to be generated. The font generation model can then generate Chinese character images that meet the designer's font requirements. After the Chinese character images are generated, the designer only needs to add the images to the required locations to complete the design of the poster text. This not only reduces the designer's work pressure and working time, but also, because the fonts generated by the model are not subject to copyright restrictions, it can also reduce the cost of purchasing font copyright fees and reduce the company's operating costs.

[0129] It should also be understood that the above is merely intended to help those skilled in the art better understand the embodiments of the present disclosure, and is not intended to limit the scope of the embodiments of the present disclosure. Based on the above examples, those skilled in the art can obviously make various equivalent modifications or variations. For example, certain steps in the above method may be unnecessary, or certain new steps may be added. Or any combination of any two or more of the above embodiments. Such modifications, variations, or combinations also fall within the scope of the embodiments of the present disclosure.

[0130] It should also be understood that the above description of the embodiments of the present disclosure focuses on emphasizing the differences between the various embodiments. The same or similar points that are not mentioned can be referenced to each other. For the sake of brevity, they will not be repeated here.

[0131] It should also be understood that the size of the sequence numbers of the above processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present disclosure.

[0132] It should also be understood that in the various embodiments of the present disclosure, unless otherwise specified or there is a logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other, and the technical features in different embodiments can be combined to form new embodiments based on their internal logical relationships.

[0133] The above describes in detail an example of a training method for a font generation model provided by the present disclosure. It is understandable that, in order to implement the above functions, the computer device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present disclosure.

[0134] The following are embodiments of the apparatus disclosed herein, which can be used to implement the method embodiments disclosed herein. For details not disclosed in the apparatus embodiments disclosed herein, please refer to the method embodiments disclosed herein.

[0135] FIG6 is a block diagram of a training apparatus for a font generation model in an exemplary embodiment of the present disclosure.

[0136] As shown in FIG6 , the training device 600 for the font generation model may include: an acquisition module 602 , an obtaining module 604 and a training module 606 .

[0137] The acquisition module 602 is configured to obtain a sample text image of a reference font, description information of a target font, and a target text image of the target font, wherein the sample text image and the target text image include the same text. The acquisition module 604 is configured to input the sample text image of the reference font and the description information of the target font into the font generation model to be trained to obtain a predicted text image of the target font. The training module 606 is configured to train a generative adversarial network comprising the font generation model and a discriminant network based on the predicted text image of the target font and the target text image of the target font, thereby obtaining a trained font generation model.

[0138] In some exemplary embodiments of the present disclosure, the training module 606 is used to: input the predicted text image of the target font and the target text image of the target font as input images to the discriminant network respectively to obtain the degree of truth or falsehood of the input images; wherein the degree of truth or falsehood is used to indicate the degree of truth or falsehood of the input image predicted by the discriminant network as the target text image; and train the discriminant network according to the degree of truth or falsehood of the input images.

[0139] In some exemplary embodiments of the present disclosure, the training module 606 is used to: determine a first loss based on the difference between the pixel values ​​of the predicted text image of the target font and the pixel values ​​of the target text image of the target font; determine a second loss based on the difference between the feature values ​​of the predicted text image of the target font and the feature values ​​of the target text image of the target font; determine a third loss based on the similarity between the description information of the target font and the predicted text image of the target font; use the discriminant network to discriminate the degree of truth or falsehood of the predicted text image of the target font, and determine a fourth loss based on the degree of truth or falsehood; determine a first font attribute value corresponding to the predicted text image of the target font and a second font attribute value corresponding to the description information of the target font, and determine a fifth loss based on the first font attribute value and the second font attribute value; train the font generation model based on the first loss, the second loss, the third loss, the fourth loss and the fifth loss.

[0140] In some exemplary embodiments of the present disclosure, the training module 606 is used to: input the predicted text image of the target font into the feature extraction model to obtain the feature value of the predicted text image of the target font; input the target text image of the target font into the feature extraction model to obtain the feature value of the target text image of the target font; and determine the second loss based on the difference between the feature value of the predicted text image of the target font and the feature value of the target text image of the target font.

[0141] In some exemplary embodiments of the present disclosure, the training module 606 is used to: input the description information of the target font and the predicted text image of the target font into the trained image-text comparison model, obtain the similarity between the description information of the target font and the predicted text image of the target font, and determine the third loss based on the similarity.

[0142] In some exemplary embodiments of the present disclosure, the training module 606 is used to: input the predicted text image of the target font into the font image attribute prediction model to obtain the first font attribute value; input the description information into the font description attribute prediction model to obtain the second font attribute value.

[0143] In some exemplary embodiments of the present disclosure, the font generation model includes a style description encoder, a font structure encoder, and a style font generation module; wherein the acquisition module 604 is used to: input a sample text image of the reference font into the font structure encoder for processing to obtain an image structure feature vector of the sample text image; input the description information of the target font into the style description encoder for processing to obtain a font style feature vector of the description information; concatenate the image structure feature vector and the font style feature vector to obtain a fused feature vector; input the fused feature vector into the style font generation module for processing to obtain a predicted text image of the target font. In some exemplary embodiments of the present disclosure, the acquisition module 602 is used to: obtain font files of multiple initial fonts from a font library; determine one of the initial fonts as the reference font, determine initial fonts other than the reference font as the target font, and obtain description information of the target font; generate a sample text image of the reference font based on the font file of the reference font, and generate a target text image of the target font based on the font file of the target font.

[0144] FIG. 7 is a block diagram of a font generating apparatus in an exemplary embodiment of the present disclosure.

[0145] As shown in FIG. 7 , the font generation device 700 may include an acquisition module 702 and an obtaining module 704 .

[0146] Among them, the acquisition module 702 is used to obtain image-text groups of multiple objects, each image-text group includes an object image and text information corresponding to the object image; the acquisition module 704 is used to input the image-text group of each object into the multimodal data representation model trained by any of the above methods to obtain multimodal feature data of each object.

[0147] In some exemplary embodiments of the present disclosure, the above-mentioned device also includes: a determination module, used to determine a third font attribute value corresponding to the predicted text image of the predicted font; an acquisition module 702 is used to obtain a fourth font attribute value obtained by adjusting the third font attribute value; the acquisition module 704 inputs the predicted text image of the predicted font and the fourth font attribute value into a font attribute adjustment model to obtain the adjusted predicted text image of the predicted font.

[0148] It should be noted that the block diagrams shown in the above figures are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor terminal devices and / or microcontroller terminal devices.

[0149] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0150] Figure 8 is a schematic diagram showing the structure of an electronic device suitable for implementing the exemplary embodiments of the present disclosure according to an exemplary embodiment. It should be noted that the electronic device shown in Figure 8 is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0151] As shown in Figure 8, electronic device 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to the program stored in read-only memory (ROM) 802 or the program loaded from storage section 808 into random access memory (RAM) 803. Various programs and data required for the operation of system 800 are also stored in RAM 803. CPU 801, ROM 802 and RAM 803 are connected to each other via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0152] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, and the like; an output section 807 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 808 including a hard disk; and a communication section 809 including a network interface card such as a LAN card or a modem. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 810 as needed, so that computer programs read therefrom can be installed into the storage section 808 as needed.

[0153] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from a removable medium 811. When the computer program is executed by the central processing unit (CPU) 801, the above-mentioned functions defined in the system of the present disclosure are performed.

[0154] It should be noted that the computer-readable medium described in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.

[0155] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0156] The units involved in the embodiments described in the present disclosure may be implemented by software or by hardware. The described units may also be provided in a processor. For example, they may be described as follows: a processor includes a sending unit, an acquisition unit, a determination unit, and a first processing unit. The names of these units do not, in some cases, constitute a limitation on the units themselves. For example, the sending unit may also be described as a "unit that sends a picture acquisition request to the connected server."

[0157] As another aspect, the present disclosure further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist independently and not be incorporated into the electronic device. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed by an electronic device, the electronic device implements the method described in the above embodiments. For example, the electronic device may implement the steps shown in Figure 2.

[0158] According to one aspect of the present disclosure, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various optional implementations of the above-described embodiments.

[0159] It should be understood that any number of elements in the drawings of the present disclosure is for illustration only and not for limitation, and any naming is for distinction only and does not have any limiting meaning.

[0160] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0161] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A training method for a font generation model, comprising: Acquire a sample text image of a reference font, description information of a target font, and a target text image of the target font, wherein the sample text image and the target text image include the same text; Inputting the sample text image of the reference font and the description information of the target font into the font generation model to be trained to obtain the predicted text image of the target font; The adversarial generative network including the font generation model and the discriminant network is trained according to the predicted text image of the target font and the target text image of the target font to obtain a trained font generation model.

2. The method according to claim 1, training a generative adversarial network including the font generation model and the discriminant network according to the predicted text image of the target font and the target text image of the target font to obtain a trained font generation model, comprising: Inputting the predicted text image of the target font and the target text image of the target font as input images to the discrimination network respectively, and obtaining the degree of truth or falsehood of the input images; The degree of truth or falsehood is used to indicate the degree of truth or falsehood of the input image predicted by the discriminant network as a target text image; The discriminant network is trained according to the authenticity of the input image.

3. The method according to claim 1 or 2, training a generative adversarial network including the font generation model and the discriminant network according to the predicted text image of the target font and the target text image of the target font to obtain a trained font generation model, comprising: determining a first loss according to a difference between a pixel value of the predicted text image of the target font and a pixel value of the target text image of the target font; determining a second loss according to a difference between a feature value of the predicted text image of the target font and a feature value of the target text image of the target font; determining a third loss according to a similarity between the description information of the target font and the predicted text image of the target font; Using the discriminant network to discriminate the authenticity of the predicted text image of the target font, and determining a fourth loss according to the authenticity; Determining a first font attribute value corresponding to the predicted text image of the target font and a second font attribute value corresponding to the description information of the target font, and determining a fifth loss according to the first font attribute value and the second font attribute value; The font generation model is trained according to the first loss, the second loss, the third loss, the fourth loss and the fifth loss.

4. The method according to claim 3, determining the second loss according to the difference between the feature value of the predicted text image of the target font and the feature value of the target text image of the target font, comprising: Inputting the predicted text image of the target font into a feature extraction model to obtain a feature value of the predicted text image of the target font; Inputting the target text image of the target font into the feature extraction model to obtain the feature value of the target text image of the target font; The second loss is determined according to a difference between a feature value of the predicted text image of the target font and a feature value of the target text image of the target font.

5. The method according to claim 3, determining the third loss according to the similarity between the description information of the target font and the predicted text image of the target font, comprising: The description information of the target font and the predicted text image of the target font are input into the trained image-text comparison model to obtain the similarity between the description information of the target font and the predicted text image of the target font, and the third loss is determined according to the similarity.

6. The method according to claim 3, determining a first font attribute value corresponding to the predicted text image of the target font and a second font attribute value corresponding to the description information of the target font, comprising: Inputting the predicted text image of the target font into a font image attribute prediction model to obtain the first font attribute value; The description information is input into a font description attribute prediction model to obtain the second font attribute value.

7. The method according to claim 1, wherein the font generation model comprises a style description encoder, a font structure encoder and a style font generation module; in, Inputting the sample text image of the reference font and the description information of the target font into the font generation model to be trained to obtain the predicted text image of the target font, including: Inputting a sample character image of the reference font into the font structure encoder for processing to obtain an image structure feature vector of the sample character image; Inputting the description information of the target font into the style description encoder for processing to obtain a font style feature vector of the description information; Concatenate the image structure feature vector and the font style feature vector to obtain a fused feature vector; The fused feature vector is input into the style font generation module for processing to obtain a predicted text image of the target font.

8. The method according to claim 1, obtaining a sample text image of a reference font, description information of a target font, and a target text image of the target font, comprises: Get font files of multiple initial fonts from the font library; Determine one of the initial fonts as the reference font, determine the initial fonts other than the reference font as the target font, and obtain description information of the target font; A sample character image of the reference font is generated according to the font file of the reference font, and a target character image of the target font is generated according to the font file of the target font.

9. A font generation method, comprising: Obtain sample text images of reference fonts and description information of predicted fonts; The sample text image of the reference font and the description information of the predicted font are input into a trained font generation model to obtain a predicted text image of the predicted font; wherein the font generation model is trained by the method according to any one of claims 1-8.

10. The method according to claim 9, further comprising: Determining a third font attribute value corresponding to the predicted text image of the predicted font; Acquire a fourth font attribute value obtained by adjusting the third font attribute value; The predicted text image of the predicted font and the fourth font attribute value are input into a font attribute adjustment model to obtain an adjusted predicted text image of the predicted font.

11. A training device for a font generation model, comprising: An acquisition module, configured to acquire a sample text image of a reference font, description information of a target font, and a target text image of the target font, wherein the sample text image and the target text image include the same text; An acquisition module, used for inputting the sample text image of the reference font and the description information of the target font into the font generation model to be trained, so as to obtain the predicted text image of the target font; A training module is used to train a generative adversarial network including the font generation model and the discriminant network according to the predicted text image of the target font and the target text image of the target font to obtain a trained font generation model.

12. A font generation device, comprising: An acquisition module, used to acquire sample text images of reference fonts and description information of predicted fonts; An acquisition module is used to input the sample text image of the reference font and the description information of the predicted font into a trained font generation model to obtain the predicted text image of the predicted font; wherein the font generation model is trained and obtained by the method according to any one of claims 1-8.

13. An electronic device comprising: at least one processor; A storage device for storing at least one program, which, when executed by the at least one processor, enables the at least one processor to implement the method according to any one of claims 1 to 10.

14. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the executable instructions are executed by a processor to implement the method according to any one of claims 1 to 10.