Style font generation method and device, equipment and medium
By building a fine-tuning dataset and adding LoRA branches to the diffusion model for training, the target style font generation model is generated, which solves the problem of inefficient style font generation and achieves the effect of quickly generating specific style fonts, which is suitable for marketing and advertising copywriting.
Patent Information
- Application Number
- CN202510616636.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-15
AI Technical Summary
In the prior art, the generation of style fonts is inefficient, and it is necessary to frequently iterate the input of propt to generate the required style fonts, resulting in inefficient generation and limited applicable scenarios.
By receiving several font images, building a fine-tuning dataset, adding the LoRA branch to the depth of the Unet network in the pre-trained diffusion model, performing fine-tuning training, generating the target style font generation model, and directly inputting text information can generate a style font picture.
It improves the efficiency of style font generation, without frequent iteration of the input of propt, can quickly generate the required style fonts, suitable for copywriting such as marketing, promotion and advertising.
Smart Images

Figure CN120495451A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence technology and financial technology, and in particular to a method, device, equipment and medium for generating a style font. Background Art
[0002] Stylish fonts are fonts with specific design styles and visual characteristics. The style of a font not only affects the readability of information but also shapes the atmosphere and tone of the work. Stylish fonts can be used in marketing, promotional, and advertising copy. For example, in the marketing activities of various business units (BUs), various marketing copy is designed and produced, and the content often relies on artistic fonts with various styles. These artistic fonts echo the marketing theme. Traditionally, designers create stylish fonts using various design software, such as Adobe Photoshop, but this is time-consuming and labor-intensive, and difficult to reuse. As the image generation capabilities of Diffusion (also known as diffusion model) continue to improve, more and more businesses are embracing and adopting image generation to create artistic fonts of various styles. However, because the results generated by Diffusion models are often highly random, even with the assistance of ControlNet (a neural network technology used to enhance image generation), frequent iterations of the prompts fed into the model are required to generate the desired stylish fonts. However, achieving high-quality prompts is difficult, resulting in inefficient stylish font generation and significantly limiting its application scenarios. Summary of the Invention
[0003] The present invention provides a style font generation method, device, computer equipment and medium to solve the technical problem of low style font generation efficiency in the prior art.
[0004] In a first aspect, a method for generating a style font is provided, comprising:
[0005] receiving a plurality of font images, and constructing a fine-tuning dataset based on the plurality of font images;
[0006] Add the preset LoRA branch to the deep layer of the Unet network in the pre-trained diffusion model to obtain the initial style font generation model;
[0007] Fine-tune the initial style font generation model based on the fine-tuning dataset to obtain a target style font generation model;
[0008] If text information is received, the text information is input into the target style font generation model to generate an image, thereby obtaining a target style font image.
[0009] In a second aspect, a device for generating a style font is provided, comprising:
[0010] a fine-tuning data construction unit, configured to receive a plurality of font images and construct a fine-tuning data set based on the plurality of font images;
[0011] The initial model generation unit is used to add the preset LoRA branch to the deep layer of the Unet network in the pre-trained diffusion model to obtain the initial style font generation model;
[0012] A model fine-tuning training unit, configured to perform fine-tuning training on the initial style font generation model based on the fine-tuning dataset to obtain a target style font generation model;
[0013] The target style font generation unit is configured to, upon receiving text information, input the text information into the target style font generation model to generate an image, thereby obtaining a target style font image.
[0014] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned method for generating a style font when executing the computer program.
[0015] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned style font generation method are implemented.
[0016] In the solution implemented by the above-mentioned style font generation method, device, computer equipment and storage medium, a number of font images are received and a fine-tuning data set is constructed based on the number of font images; a preset LoRA branch is added to the deep layer of the Unet network in the pre-trained diffusion model to obtain an initial style font generation model; the initial style font generation model is fine-tuned and trained based on the fine-tuning data set to obtain a target style font generation model; if text information is received, the text information is input into the target style font generation model for image generation to obtain a target style font image. The required style font can be generated without frequently iterating the prompt input into the model, thereby improving the efficiency of style font generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 This is a flow chart of a method for generating a style font according to an embodiment of the present invention;
[0019] Figure 2 yes Figure 1 A schematic flow chart of a specific implementation of step S1;
[0020] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step S11;
[0021] Figure 4 yes Figure 1 A schematic flow chart of a specific implementation of step S3;
[0022] Figure 5 1 is a schematic structural diagram of a style font generation device provided by an embodiment of the present invention;
[0023] Figure 6 is a structural diagram of a computer device in one embodiment of the present invention;
[0024] Figure 7 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0026] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0027] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0028] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0029] See also Figure 1 , Figure 1 This is a flow chart of a method for generating a style font provided by an embodiment of the present invention. The method for generating a style font is applied in a server.
[0030] like Figure 1 As shown, the method includes the following steps S1 to S4.
[0031] S1. Receive a number of font images, and construct a fine-tuning dataset based on the font images.
[0032] In this embodiment, a user inputs several font images on a display terminal and constructs a fine-tuning dataset from these images. This dataset serves as a training set for subsequent model training, and the fine-tuning dataset includes font images of the desired style. These font images may include 3 to 5 images, requiring only a small number of font images to be input by the user.
[0033] In one embodiment, see Figure 2 , step S1 includes:
[0034] S11, obtaining a target font image set based on the plurality of font images and a preset image screening strategy;
[0035] S12, obtaining prompt words corresponding to each target font image in the target font image set; wherein the prompt words include a plurality of style description words;
[0036] S13, obtaining style description words common to all target font images in the target font image set as target style description words, and forming a target style description word set;
[0037] S14. Construct the fine-tuning dataset based on the target font image set and the target style description word set.
[0038] In this embodiment, first, due to limitations in user knowledge and experience, the styles of several input font images may differ or even be incorrect. Therefore, an image screening strategy is used to select target font images with the same style from the multiple font images. These selected target font images are then grouped into a target font image set, i.e., each target font image in the target font image set has the same style. Next, a prompt word corresponding to each target font image is obtained. The prompt word includes several style descriptors. A tool capable of generating image-related prompt words can be used to reversely calculate the prompt word corresponding to each target font image. Then, from the several style descriptors corresponding to each target font image, a style descriptor common to all target font images is selected as the target style descriptor. These selected target style descriptors are then grouped into a target style descriptor set. Finally, for each target font image in the target font image set, the target font image and all target style descriptors in the target style descriptor set are combined into a set of fine-tuning data. The fine-tuning data corresponding to each target font image is then grouped into a fine-tuning dataset. The amount of fine-tuning data in the fine-tuning dataset depends on the number of target font images.
[0039] For example, a target font picture set obtained by screening from several font pictures input by the user includes three target font pictures, namely target font picture a, target font picture b and target font picture c; and the style description words corresponding to target font picture a are "glossy", "smooth edges" and "soft colors", the style description words corresponding to target font picture b are "glossy" and "soft colors", and the style description words corresponding to target font picture c are "glossy", "dreamy and romantic" and "soft colors", so the screened target style description words are "glossy" and "soft colors"; finally, a fine-tuning dataset is constructed from target font picture a, target font picture b, target font picture c, "glossy" and "soft colors", and the fine-tuning dataset is used as a training set for subsequent training models, so that the subsequently trained target style font generation model can generate font pictures with "glossy" and "soft colors" styles. Such font pictures with "glossy" and "soft colors" styles can be embedded in marketing, promotion or advertising copy in commercial or financial businesses.
[0040] In one embodiment, see Figure 3 , step S11 includes:
[0041] S111, combining two of the font pictures in the plurality of font pictures into font picture groups to obtain a plurality of initial font picture groups;
[0042] S112, calculating the similarity between two font images included in each initial font image group to obtain an initial image similarity set;
[0043] S113, clustering the plurality of font images according to the initial image similarity set to obtain at least one font image subset;
[0044] S114: If there is a unique font image subset whose corresponding initial image similarity is greater than a preset similarity threshold, the font image subset greater than the similarity threshold is used as the target font image set;
[0045] S115. If there are multiple font image subsets whose corresponding initial image similarities are greater than the similarity threshold, the font image subsets greater than the similarity threshold are used as candidate font image subsets, and the candidate font image subset containing the largest number of font images is selected as the target font image set.
[0046] In this embodiment, several font images input by the user are combined in pairs to obtain several initial font image groups. The similarity between two font images included in each initial font image group is calculated, and the similarity between the two font images included in the initial font image group is used as the initial image similarity. The initial image similarities corresponding to all initial font image groups are then combined into an initial image similarity set. The several font images are then clustered using the initial image similarity set to obtain at least one font image subset, wherein the initial image similarity corresponding to any two font images in the font image subset is the same, that is, the initial image similarity corresponding to a font image subset is the initial image similarity corresponding to any two font images in the font image subset.
[0047] After obtaining at least one font picture subset, it is necessary to determine whether there is a font picture subset in the at least one font picture subset whose initial picture similarity is greater than a preset similarity threshold, so as to screen out a qualified font picture subset in the at least one font picture subset as a target font picture set. Preferably, the similarity threshold can be set to 0.7. Specifically, if only the initial picture similarity corresponding to a unique font picture subset is greater than the similarity threshold, then the unique font picture subset greater than the similarity threshold is used as the target font picture set. If the initial picture similarity corresponding to multiple font picture subsets is greater than the similarity threshold, then these font picture subsets greater than the similarity threshold are used as candidate font picture subsets, and then the candidate font picture subset containing the largest number of font pictures is selected as the target font picture set, that is, the font picture subset greater than the similarity threshold and containing the largest number of font pictures is used as the target font picture set, thereby ensuring that the target font pictures in the target font picture set have the same style.
[0048] In one embodiment, step S112 includes:
[0049] Input each initial font image group into the pre-trained style transfer model in sequence to obtain two font image features corresponding to each initial font image group;
[0050] For each initial font picture group, the similarity between two font picture features corresponding to the initial font picture group is calculated to obtain the initial picture similarity set.
[0051] In this embodiment, a pre-trained style transfer model is used to extract font image features of a font image. Preferably, the style transfer model can be a StyTr2 model. The StyTr2 model can better handle long-range dependencies of image features and avoid the loss of content and style details. Specifically, for each initial font image group, the two font images contained in the initial font image group are sequentially input into the pre-trained style transfer model. The style transfer model outputs two font image features corresponding to the initial font image group. By calculating the similarity between the two font image features, the initial image similarity corresponding to the initial font image group is obtained, and the initial image similarities corresponding to all initial font image groups are formed into an initial image similarity set.
[0052] In one embodiment, after step S113, the method further includes:
[0053] If the initial image similarities corresponding to each font image subset are less than or equal to the similarity threshold, then obtain several new font images, use the several new font images to update the several font images, and return to execute the step of combining the font images in the several font images into font image groups in pairs to obtain several initial font image groups.
[0054] In this embodiment, if the initial image similarity corresponding to each font image subset is less than or equal to the similarity threshold, that is, there is no font image subset whose initial image similarity is greater than the similarity threshold, it means that the styles of the several font images input by the user may be different or even wrong, then it is necessary to obtain several new font images re-entered by the user, use the several new font images to update the several font images, and return to execute step S111, so as to filter out the target font image set based on the several new font images.
[0055] S2. Add the preset LoRA branch to the deep layer of the Unet network in the pre-trained diffusion model to obtain the initial style font generation model.
[0056] In this embodiment, a LoRA branch is added to the deep layer of the Unet network in the pre-trained diffusion model, thereby changing the network structure of the diffusion model and obtaining an initial style font generation model. The deep layer of the Unet network is any of the last two layers of the upsampling module and the downsampling module in the Unet network. In addition, because the deep layer of the Unet network has a greater impact on the style generated by the diffusion model, the LoRA branch is only added to the deep layer of the Unet network, and only the diffusion model is fine-tuned in a targeted manner. There is no need to add LoRA branches to all attention layers of the Unet network. This can more quickly fine-tune the initial style font generation model to the target style font generation model.
[0057] S3. Fine-tune the initial style font generation model based on the fine-tuning dataset to obtain a target style font generation model.
[0058] In this embodiment, the initial style font generation model is fine-tuned based on the fine-tuning dataset, so that the style generated by the initial style font generation model is closer to the style of the fine-tuning dataset, thereby fine-tuning the initial style font generation model to the target style font generation model. The fine-tuning dataset is constructed based on a number of font images, and the font images are a small number of font images input by the user (e.g., 3 to 5 font images). Therefore, the fine-tuning dataset contains a small number of target font images with the same style, and the number of parameters is small. This results in faster convergence and less data dependence when training the model, and enables the rapid fine-tuning of a target style font generation model capable of generating the target style using only a small number of font images with the same style.
[0059] In one embodiment, see Figure 4 , step S3 includes:
[0060] S31, randomly selecting a set of fine-tuning data from the fine-tuning data set and inputting it into the initial style font generation model to generate an image, thereby obtaining an initial style font image;
[0061] S32. Obtaining an image loss between the target font image in the fine-tuning data and the initial style font image based on a loss function of the diffusion model;
[0062] S33, obtaining a style loss between the target font image in the fine-tuning data and the initial style font image based on a preset style loss function;
[0063] S34. Generate a final loss of the initial style font generation model based on the image loss and the style loss;
[0064] S35. Using the final loss, adjust the parameters of the LoRA branch in the initial style font generation model to obtain an adjusted initial style font generation model;
[0065] S36: Update the initial style font generation model using the adjusted initial style font generation model, and return to the step of randomly selecting a set of fine-tuning data from the fine-tuning dataset and inputting it into the initial style font generation model to generate an image to obtain an initial style font image, until the final loss of the initial style font generation model meets a preset stopping condition, then stop adjusting and output the target style font generation model.
[0066] In this embodiment, when fine-tuning the initial style font generation model, since the fine-tuning data set includes several groups of fine-tuning data groups, and each group of fine-tuning data groups includes a target font picture and a target style description word set, it is first necessary to randomly select a group of fine-tuning data from the fine-tuning data set and input it into the initial style font generation model for picture generation. The initial style font generation model outputs the initial style font picture.
[0067] Next, the image loss between the target font image and the initial style font image in the set of fine-tuning data is obtained through the loss function of the diffusion model, that is, the image loss between the input image and the output image of the initial style font generation model is calculated. Preferably, the loss function of the diffusion model can adopt the MSE (mean-square error) loss function. At the same time, the style loss between the target font image and the initial style font image is obtained through a preset style loss function, and the final loss of the initial style font generation model is generated based on the image loss and the style loss, wherein the final loss can be the sum of the image loss and the style loss. Then, the parameters of the LoRA branch in the initial style font generation model are adjusted using the final loss. The LoRA branch can be represented as the product of two low-rank matrices. Only the parameters of the two low-rank matrices in the LoRA branch are adjusted, without adjusting all the parameters in the initial style font generation model. This can improve the fine-tuning efficiency, thereby quickly completing this round of fine-tuning training of the initial style font generation model and obtaining the adjusted initial style font generation model, so that the style generated by the adjusted initial style font generation model is closer to the style corresponding to the target style description word set.
[0068] Finally, the adjusted initial style font generation model is used to update the initial style font generation model, and the process returns to step S31 to begin the next round of fine-tuning training on the initial style font generation model. The training process is repeated until the final loss of the initial style font generation model meets a preset stopping condition. The adjustment is then stopped, and the initial style font generation model at this point is output as the final target style font generation model. The preset stopping condition may be that the final loss of the initial style font generation model meets a convergence condition or that the number of training rounds reaches a preset upper limit, which is not specifically limited here.
[0069] In one embodiment, step S33 includes:
[0070] Calculating the similarity between the target font image in the fine-tuning data and the initial style font image to obtain style image similarity;
[0071] The style loss is obtained according to the style loss function and the style picture similarity; wherein the style loss function is f(s)=1-s, f(s) is the style loss, and s is the style picture similarity.
[0072] In this embodiment, the similarity between the target font image in the fine-tuning data input to the initial style font generation model and the initial style font image corresponding to the output of the initial style font generation model is calculated to obtain the style image similarity. Among them, the font image features corresponding to the target font image and the initial style font image can be extracted respectively through a pre-trained style transfer model, and the similarity between the font image features corresponding to the target font image and the initial style font image can be calculated to obtain the style image similarity. Preferably, the style transfer model can be a StyTr2 model. The style image similarity is substituted into the style loss function f(s) = 1-s, where s is the style image similarity, and the style loss f(s) is calculated to maximize the style similarity between the input image and the output image of the initial style font generation model based on the style loss, so as to achieve the purpose of guiding the constraint to generate the target style.
[0073] S4. If text information is received, the text information is input into the target style font generation model to generate an image, thereby obtaining a target style font image.
[0074] In this embodiment, when using the target style font generation model, upon receiving text information input by a user on a display terminal, the target style font generation model is used to generate an image, thereby obtaining a target style font image. This allows the text information to be created as a font with a specific style, without the need for frequent iterations of the model's prompt input and without relying on high-quality prompts, thereby improving the efficiency of style font generation. For example, in the production of BU event posters, it is often necessary to embed a specific style of font into the copy, which requires the creation of a font in the desired style. By simply inputting the text information to be created into the fine-tuned and trained target style font generation model for image generation, a target style font image in the specific desired style can be quickly obtained.
[0075] For ease of description, take the application scenario of property insurance business marketing copy as an example: the text information is "Ping An Insurance builds a safety barrier for you", and the font in the text information needs to be created as a font with "glossy", "blue" and "simple" styles, and the fine-tuned and trained target style font generation model can generate font images with "glossy", "blue" and "simple" styles, then "Ping An Insurance builds a safety barrier for you" is input into the target style font generation model, and the target style font image output by the target style font generation model shows that the font of "Ping An Insurance builds a safety barrier for you" has "glossy", "blue" and "simple" styles. The font of "Ping An Insurance builds a safety barrier for you" with "glossy", "blue" and "simple" styles is embedded in the copy, so as to quickly create the font of the required style and improve the efficiency of marketing copy production.
[0076] It can be seen that in the above scheme, the present invention receives several font images and constructs a fine-tuning dataset based on the several font images; adds the preset LoRA branch to the deep layer of the Unet network in the pre-trained diffusion model to obtain an initial style font generation model; fine-tunes the initial style font generation model based on the fine-tuning dataset to obtain a target style font generation model; if text information is received, the text information is input into the target style font generation model for image generation to obtain the target style font image, and the required style font can be generated without frequently iterating the prompt input into the model, thereby improving the efficiency of style font generation.
[0077] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0078] See also Figure 5 , Figure 5This is a schematic diagram of the structure of a style font generation device provided by an embodiment of the present invention. The present invention also provides a style font generation device, which corresponds to the style font generation method in the above embodiment. Figure 5 The style font generation device includes a fine-tuning data construction unit 100, an initial model generation unit 200, a model fine-tuning training unit 300, and a target style font generation unit 400. The functional units are described in detail as follows:
[0079] A fine-tuning data construction unit 100 is configured to receive a plurality of font images and construct a fine-tuning data set based on the plurality of font images;
[0080] The initial model generation unit 200 is used to add a preset LoRA branch to the deep layer of the Unet network in the pre-trained diffusion model to obtain an initial style font generation model;
[0081] A model fine-tuning training unit 300 is configured to perform fine-tuning training on the initial style font generation model based on the fine-tuning dataset to obtain a target style font generation model;
[0082] The target style font generation unit 400 is configured to, upon receiving text information, input the text information into the target style font generation model to generate an image, thereby obtaining a target style font image.
[0083] In one embodiment, when executing the step of constructing a fine-tuning dataset based on the plurality of font images, the fine-tuning data construction unit 100 is specifically configured to:
[0084] Obtaining a target font image set based on the plurality of font images and a preset image screening strategy;
[0085] Obtaining prompt words corresponding to each target font image in the target font image set; wherein the prompt words include a plurality of style description words;
[0086] Obtaining style description words common to all target font images in the target font image set as target style description words, and forming a target style description word set;
[0087] The fine-tuning dataset is constructed based on the target font image set and the target style description word set.
[0088] In one embodiment, when executing the step of obtaining a target font image set based on the plurality of font images and a preset image screening strategy, the fine-tuning data construction unit 100 is specifically configured to:
[0089] Combining two or more of the font pictures in the plurality of font pictures into font picture groups to obtain a plurality of initial font picture groups;
[0090] Calculate the similarity between the two font images included in each initial font image group to obtain an initial image similarity set;
[0091] Clustering the plurality of font images according to the initial image similarity set to obtain at least one font image subset;
[0092] If there is a unique font picture subset whose corresponding initial picture similarity is greater than a preset similarity threshold, the font picture subset greater than the similarity threshold is used as the target font picture set;
[0093] If there are multiple font image subsets whose corresponding initial image similarities are greater than the similarity threshold, the font image subsets greater than the similarity threshold are used as candidate font image subsets, and the candidate font image subset containing the largest number of font images is selected as the target font image set.
[0094] In one embodiment, when executing the step of calculating the similarity between two font images included in each initial font image group to obtain an initial image similarity set, the fine-tuning data construction unit 100 is specifically configured to:
[0095] Input each initial font image group into the pre-trained style transfer model in sequence to obtain two font image features corresponding to each initial font image group;
[0096] For each initial font picture group, the similarity between two font picture features corresponding to the initial font picture group is calculated to obtain the initial picture similarity set.
[0097] In one embodiment, after performing the step of clustering the plurality of font images according to the initial image similarity set to obtain at least one font image subset, the fine-tuning data construction unit 100 is further configured to:
[0098] If the initial image similarities corresponding to each font image subset are less than or equal to the similarity threshold, then obtain several new font images, use the several new font images to update the several font images, and return to execute the step of combining the font images in the several font images into font image groups in pairs to obtain several initial font image groups.
[0099] In one embodiment, the model fine-tuning training unit 300 is specifically used to:
[0100] Randomly selecting a set of fine-tuning data from the fine-tuning dataset and inputting it into the initial style font generation model to generate an image, thereby obtaining an initial style font image;
[0101] Obtaining an image loss between a target font image in the fine-tuning data and the initial style font image based on a loss function of the diffusion model;
[0102] Obtaining a style loss between a target font image in the fine-tuning data and the initial style font image based on a preset style loss function;
[0103] generating a final loss of the initial style font generation model based on the image loss and the style loss;
[0104] Using the final loss, adjusting parameters of the LoRA branch in the initial style font generation model to obtain an adjusted initial style font generation model;
[0105] The adjusted initial style font generation model is used to update the initial style font generation model, and the step of randomly selecting a set of fine-tuning data from the fine-tuning dataset and inputting it into the initial style font generation model for image generation to obtain an initial style font image is returned to the step of obtaining an initial style font image until the final loss of the initial style font generation model meets a preset stopping condition, then the adjustment is stopped and the target style font generation model is output.
[0106] In one embodiment, when the model fine-tuning training unit 300 executes the step of obtaining the style loss between the target font image in the fine-tuning data and the initial style font image based on a preset style loss function, it is specifically configured to:
[0107] Calculating the similarity between the target font image in the fine-tuning data and the initial style font image to obtain style image similarity;
[0108] The style loss is obtained according to the style loss function and the style picture similarity; wherein the style loss function is f(s)=1-s, f(s) is the style loss, and s is the style picture similarity.
[0109] The present invention provides a style font generation device, which receives a number of font images and constructs a fine-tuning dataset based on the font images; adds a preset LoRA branch to the deep layer of a Unet network in a pre-trained diffusion model to obtain an initial style font generation model; fine-tunes the initial style font generation model based on the fine-tuning dataset to obtain a target style font generation model; if text information is received, the text information is input into the target style font generation model to generate an image to obtain the target style font image. The required style font can be generated without frequently iterating the prompt input into the model, thereby improving the efficiency of style font generation.
[0110] The specific definition of the stylized font generation device can be found in the definition of the stylized font generation method above and will not be repeated here. Each unit in the stylized font generation device described above can be implemented in whole or in part via software, hardware, or a combination thereof. Each of the above units can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above units.
[0111] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of a style font generation method.
[0112] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of the client side of a style font generation method.
[0113] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0114] receiving a plurality of font images, and constructing a fine-tuning dataset based on the plurality of font images;
[0115] Add the preset LoRA branch to the deep layer of the Unet network in the pre-trained diffusion model to obtain the initial style font generation model;
[0116] Fine-tune the initial style font generation model based on the fine-tuning dataset to obtain a target style font generation model;
[0117] If text information is received, the text information is input into the target style font generation model to generate an image, thereby obtaining a target style font image.
[0118] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0119] receiving a plurality of font images, and constructing a fine-tuning dataset based on the plurality of font images;
[0120] Add the preset LoRA branch to the deep layer of the Unet network in the pre-trained diffusion model to obtain the initial style font generation model;
[0121] Fine-tune the initial style font generation model based on the fine-tuning dataset to obtain a target style font generation model;
[0122] If text information is received, the text information is input into the target style font generation model to generate an image, thereby obtaining a target style font image.
[0123] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0124] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0125] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0126] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A method for generating a style font, characterized in that: include: receiving a plurality of font images, and constructing a fine-tuning dataset based on the plurality of font images; Add the preset LoRA branch to the deep layer of the Unet network in the pre-trained diffusion model to obtain the initial style font generation model; Fine-tune the initial style font generation model based on the fine-tuning dataset to obtain a target style font generation model; If text information is received, the text information is input into the target style font generation model to generate an image, thereby obtaining a target style font image.
2. The method for generating a style font according to claim 1, wherein: The constructing of a fine-tuning dataset based on the plurality of font images includes: Obtaining a target font image set based on the plurality of font images and a preset image screening strategy; Obtaining prompt words corresponding to each target font image in the target font image set; wherein the prompt words include a plurality of style description words; Obtaining style description words common to all target font images in the target font image set as target style description words, and forming a target style description word set; The fine-tuning dataset is constructed based on the target font image set and the target style description word set.
3. The method for generating a style font according to claim 2, wherein: The step of obtaining a target font image set based on the plurality of font images and a preset image screening strategy includes: Combining two or more of the font pictures in the plurality of font pictures into font picture groups to obtain a plurality of initial font picture groups; Calculate the similarity between the two font images included in each initial font image group to obtain an initial image similarity set; Clustering the plurality of font images according to the initial image similarity set to obtain at least one font image subset; If there is a unique font picture subset whose corresponding initial picture similarity is greater than a preset similarity threshold, the font picture subset greater than the similarity threshold is used as the target font picture set; If there are multiple font image subsets whose corresponding initial image similarities are greater than the similarity threshold, the font image subsets greater than the similarity threshold are used as candidate font image subsets, and the candidate font image subset containing the largest number of font images is selected as the target font image set.
4. The method for generating a style font according to claim 3, wherein: The calculating the similarity between the two font pictures included in each initial font picture group to obtain an initial picture similarity set includes: Input each initial font image group into the pre-trained style transfer model in sequence to obtain two font image features corresponding to each initial font image group; For each initial font picture group, the similarity between two font picture features corresponding to the initial font picture group is calculated to obtain the initial picture similarity set.
5. The method for generating a style font according to claim 3, wherein: After clustering the plurality of font images according to the initial image similarity set to obtain at least one font image subset, the method further includes: If the initial image similarities corresponding to each font image subset are less than or equal to the similarity threshold, then obtain several new font images, use the several new font images to update the several font images, and return to execute the step of combining the font images in the several font images into font image groups in pairs to obtain several initial font image groups.
6. The method for generating a style font according to claim 2, wherein: The fine-tuning training of the initial style font generation model based on the fine-tuning dataset to obtain a target style font generation model includes: Randomly selecting a set of fine-tuning data from the fine-tuning dataset and inputting it into the initial style font generation model to generate an image, thereby obtaining an initial style font image; Obtaining an image loss between a target font image in the fine-tuning data and the initial style font image based on a loss function of the diffusion model; Obtaining a style loss between a target font image in the fine-tuning data and the initial style font image based on a preset style loss function; generating a final loss of the initial style font generation model based on the image loss and the style loss; Using the final loss, adjusting parameters of the LoRA branch in the initial style font generation model to obtain an adjusted initial style font generation model; The adjusted initial style font generation model is used to update the initial style font generation model, and the step of randomly selecting a set of fine-tuning data from the fine-tuning dataset and inputting it into the initial style font generation model for image generation to obtain an initial style font image is returned to the step of obtaining an initial style font image until the final loss of the initial style font generation model meets a preset stopping condition, then the adjustment is stopped and the target style font generation model is output.
7. The method for generating a style font according to claim 6, wherein: The obtaining, based on a preset style loss function, a style loss between a target font image in the fine-tuning data and the initial style font image includes: Calculating the similarity between the target font image in the fine-tuning data and the initial style font image to obtain style image similarity; The style loss is obtained according to the style loss function and the style picture similarity; wherein the style loss function is f(s)=1-s, f(s) is the style loss, and s is the style picture similarity.
8. A style font generation device, characterized in that: include: a fine-tuning data construction unit, configured to receive a plurality of font images and construct a fine-tuning data set based on the plurality of font images; The initial model generation unit is used to add the preset LoRA branch to the deep layer of the Unet network in the pre-trained diffusion model to obtain the initial style font generation model; A model fine-tuning training unit, configured to perform fine-tuning training on the initial style font generation model based on the fine-tuning dataset to obtain a target style font generation model; The target style font generation unit is configured to, upon receiving text information, input the text information into the target style font generation model to generate an image, thereby obtaining a target style font image.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method for generating a style font according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for generating a style font according to any one of claims 1 to 7 are implemented.