Font generation method and device and font generation model training method and device
Patent Information
- Application Number
- CN202380012602.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-08-29
AI Technical Summary
The existing font generation model can only generate fonts of specific styles, which is difficult to meet users' diverse needs for font styles, and cannot adjust the style of generated fonts by themselves.
By introducing vector weighting concepts into the font generation model, we use the basic font style vector weighting sum of various basic font styles to generate synthetic style font pictures, and combine the feature extraction capabilities of content branches and style branches to generate synthetic style fonts that conform to the user's handwriting style.
It realizes the diversified generation of font styles, and users can adjust the style of generated fonts by themselves, improving the flexibility and diversity of font beautification effects.
Smart Images

Figure CN120569726A_ABST
Abstract
Description
Font generation method and device, font generation model training method and device Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to a font generation method and device, and a font generation model training method and device. Background Art
[0002] When displaying text, electronic devices can generate and display a font that conforms to a specific style based on the handwritten font input by the user in order to improve the display effect and aesthetics of the font.
[0003] At present, supervised training is usually used to train deep learning networks such as GAN (Generative Adversarial Networks) in order to obtain a font generation model for generating fonts in a specific style.
[0004] However, these font generation models are often limited to generating fonts in specific styles. For example, if the reference style is Kaiti, the model will only generate Kaiti-style fonts; if the reference style is Xingshu, the model will only generate Xingshu-style fonts. Therefore, current font generation models often struggle to meet users' diverse needs for font styles, for example, users cannot adjust the style of the generated fonts themselves, and improvements are urgently needed.
[0005] Summary of the Invention
[0006] In view of this, embodiments of the present invention provide a font generation model training method and device, and a font generation method and device to address the deficiencies in the related art.
[0007] According to a first aspect of an embodiment of the present invention, a font generation method is provided, comprising:
[0008] Obtain the target text content vector of the target handwritten font image based on the font generation model;
[0009] Inputting the target handwritten font image into the style branch of the font generation model to extract a target text style vector, and determining a handwriting weighted style vector based on the target text style vector, wherein the handwriting weighted style vector is obtained by weighted summation of basic font style vectors of multiple basic font styles;
[0010] Obtain a synthetic style font image generated by the decoder of the font generation model according to the target text content vector and the handwriting weight style vector.
[0011] According to a second aspect of an embodiment of the present invention, a method for training a font generation model is provided, comprising:
[0012] Using a first training sample to train an original font generation model to obtain a first font generation model, the original font generation model includes a content branch, a style branch, and a decoder, wherein a first sample content font image and a first sample style font image in the first training sample are input into the content branch and the style branch, respectively, and the decoder is configured to generate and output a first synthesized style font image corresponding to the first training sample based on a first sample text content vector extracted from the content branch and a first sample text style vector extracted from the style branch;
[0013] The first font generation model is trained using a second training sample to obtain a second font generation model, wherein the second training sample includes a second sample content font image and a second sample style font image, the second sample content font image is input into the content branch, and the second sample text content vector extracted from the content branch is input into the decoder; the sample weight style vector of the second sample style font image is input into the decoder, and the sample weight style vector is obtained by weighted summation of the sample basic font style vectors of multiple sample basic font styles; the decoder is used to generate and output a second synthetic style font image corresponding to the second training sample based on the second sample text content vector and the sample weight style vector.
[0014] According to a third aspect of an embodiment of the present invention, a font generation model is proposed, wherein the model is trained by any one of the methods described in the second aspect.
[0015] According to a fourth aspect of an embodiment of the present invention, a font generation device is provided, the device including one or more processors, wherein the processors are configured to:
[0016] Obtain the target text content vector of the target handwritten font image based on the font generation model;
[0017] Inputting the target handwritten font image into the style branch of the font generation model to extract a target text style vector, and determining a handwriting weighted style vector based on the target text style vector, wherein the handwriting weighted style vector is obtained by weighted summation of basic font style vectors of multiple basic font styles;
[0018] Obtain a synthetic style font image generated by the decoder of the font generation model according to the target text content vector and the handwriting weight style vector.
[0019] According to a fifth aspect of an embodiment of the present invention, a device for training a font generation model is provided, the device comprising one or more processors, wherein the processors are configured to:
[0020] Using a first training sample to train an original font generation model to obtain a first font generation model, the original font generation model includes a content branch, a style branch, and a decoder, wherein a first sample content font image and a first sample style font image in the first training sample are input into the content branch and the style branch, respectively, and the decoder is configured to generate and output a first synthesized style font image corresponding to the first training sample based on a first sample text content vector extracted from the content branch and a first sample text style vector extracted from the style branch;
[0021] The first font generation model is trained using a second training sample to obtain a second font generation model, wherein the second training sample includes a second sample content font image and a second sample style font image, the second sample content font image is input into the content branch, and the second sample text content vector extracted from the content branch is input into the decoder; the sample weight style vector of the second sample style font image is input into the decoder, and the sample weight style vector is obtained by weighted summation of the sample basic font style vectors of multiple sample basic font styles; the decoder is used to generate and output a second synthetic style font image corresponding to the second training sample based on the second sample text content vector and the sample weight style vector.
[0022] According to a sixth aspect of an embodiment of the present invention, an electronic device is proposed, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the method described in any one of the first or second aspects above.
[0023] According to a seventh aspect of an embodiment of the present invention, a non-volatile computer-readable storage medium is proposed, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in any one of the first aspect or the second aspect are implemented.
[0024] According to an embodiment of the present invention, the training process of the font generation model includes two stages: the first stage is used to train the original font generation model to obtain a first font generation model, so that the model has the ability to extract features from font images; the second stage is used to train the first font generation model to obtain a second font generation model, so that the model has the ability to reason about handwritten fonts. It can be seen that the first stage is a pre-training process, and the obtained first font generation model is an intermediate state in the complete training process, while the second stage is a formal training process, and the obtained second font generation model is the final training result of the complete training process.
[0025] Compared with the font generation model and its training method in the related art, when this solution uses the second training sample to train the second font generation model, the sample text style vector of the second sample style font image obtained by extracting the style branch is not directly input into the decoder, but the sample weighted style vector of the second sample style font image obtained by weighted summation of the sample basic font style vectors of multiple sample basic font styles is input into the decoder, so that the decoder in the trained second font generation model has the ability to infer handwritten fonts of multiple styles.
[0026] Furthermore, when the font generation model obtained by the above training process (i.e., the second font generation model) is used to generate a corresponding synthetic style font image for the user's target handwriting font image, the target text style vector of the image extracted by the style branch is not directly input into the decoder, but the handwriting weight style vector obtained by weighted summation of the basic font style vectors of multiple basic font styles is first determined according to the target text style vector, and then the vector is input into the decoder, thereby obtaining the synthetic style font image generated by the decoder according to the handwriting weight style vector and the target text content vector extracted by the content branch. The actual style of the synthetic style font image generated in this way is obtained by fusion of multiple basic font styles and the handwriting style of the target handwriting font image. It can be seen that the present invention introduces the concept of vector weight in the font generation method based on the font generation model, so that the actual style of the synthetic font style image finally generated can meet the user's diverse needs for font style, such as the user can adjust the style of the generated font by himself, etc., which significantly improves the flexibility of the font beautification effect.
[0027] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0029] FIG1 is a schematic diagram of the structure of a font generation model implemented based on a DG-Font network and its corresponding discriminator in the related art.
[0030] FIG. 2 is a schematic diagram showing a comparison between a defective font and a normal font according to an embodiment of the present invention.
[0031] FIG3 is a flow chart of a method for training a font generation model according to an embodiment of the present invention.
[0032] FIG4 is a schematic diagram showing candidate font images belonging to a certain font style in a font image library according to an embodiment of the present invention.
[0033] FIG. 5 is a schematic diagram illustrating a plurality of candidate font images in a font image library that belong to different font styles and have the same text content according to an embodiment of the present invention.
[0034] FIG6 is a schematic diagram showing the structure of a CNN and FPN according to an embodiment of the present invention.
[0035] FIG. 7 is a schematic structural diagram of an improved font generation model according to an embodiment of the present invention.
[0036] FIG8 is a schematic structural diagram of a Transformer encoder according to an embodiment of the present invention.
[0037] FIG9 is a schematic diagram showing a font image and a Chinese character outline thereof according to an embodiment of the present invention.
[0038] FIG10 is a schematic diagram showing a weight correspondence relationship between a font image and a plurality of basic font style vectors according to an embodiment of the present invention.
[0039] FIG. 11 is a flow chart showing a font generation method according to an embodiment of the present invention.
[0040] 12 to 14 are comparative diagrams showing effects of synthesized style font images corresponding to the style proportions of one or more target font styles and handwriting styles according to embodiments of the present invention.
[0041] FIG15 is a schematic structural diagram of a node device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0043] At present, deep learning networks such as GAN are usually trained in a supervised training manner in order to obtain a font generation model for generating specific style fonts. Taking the commonly used DG-Font network in the industry as an example, the working method and its disadvantages of this type of font generation model will be introduced in detail below. As shown in (a) of Figure 1, the internal structure of the DG-Font network includes a content branch, a style branch, and a decoder. Among them, the content branch and the style branch are essentially encoders for images, so they are also called content encoder and style encoder, respectively, for special note.
[0044] The content branch consists of three layers of deformable convolutional layers, two layers of residual blocks, and a normalization layer. Usually, the convolution kernel stride used in the latter two of the three layers of deformable convolutional layers is 2. At this time, this branch will perform two downsamplings on the input image (such as the "Chen" character input on the left). Taking the size of the input image as 80*80 as an example, the size of the feature image (i.e., Zc) after two downsampling encodings is 20*20*256. In addition, the dimension of its output channel is set to 128.
[0045] The style branch consists of a VGG1 network (including a 3*3 ordinary convolutional network and five layers of max pooling layers), one layer of average pooling layer, and one layer of fully connected layer. The dimension of the output channel of the last fully connected layer is set to 128. Still taking the input image with a size of 80*80 (such as the "Dong" character input on the left) as an example, after five downsampling poolings of the VGG1 network, it becomes an output image with a size of 2*2, and then becomes an output image with a size of 1*1 (the dimension of its output channel is 512) after average pooling, and then the channel dimension is reduced to 128 dimensions through a fully connected layer, and finally a feature vector with a size of 1*1*128 (i.e., Zs) is output. It can be seen that the style branch will encode the input image into a 128-dimensional feature vector.
[0046] The decoder consists of residual convolutional blocks and upsampling layers. Since the content branch changes the input image with a size of 80*80 to 20*20, the decoder restores the size of the 20*20 feature image to 80*80 (the same as the input image size) through two upsamplings.
[0047] There are also FDSC-1 and FDSC-2 connected between the content branch and the decoder. The specific structure of FDSC can be seen in (b) of Figure 1, which will not be elaborated here.
[0048] In addition, the style difference between the synthesized font image output by the DG-Font network shown in FIG. 1 (such as the character "Chen" output on the right) and the original stylized font image (i.e., the character "Dong" input on the left) can be judged by a discriminator. For example, the synthesized font image can be input into the discriminator shown in (c) of FIG. 1. The discriminator is equivalent to a classifier for font styles, and the output dimension is the number of possible values of the font style of the synthesized font image. When inputting a synthesized font image, according to its true style category, the score value corresponding to this category can be taken from the output dimension as the score for the authenticity of the synthesis of this image. Among them, the higher the score of the synthesized font image for any style category, the more likely it is that the font belongs to this style category, that is, the closer the style of the font is to this style category.
[0049] During the working process, the content branch of the DG-Font network is input with the user's handwritten font image, and the style branch is input with a font image of a reference style (with the same size as the handwritten font image). Thus, the content branch and the style branch respectively extract the content features and style features of the corresponding images and input them into the decoder. The size of the synthesized font image output by the decoder is the same as the size of the handwritten font image. The text content of the synthesized font image is the same as that of the handwritten font image (that is, the two represent the same character), while the font style of the synthesized font image is the same as the reference style. As shown in (a) of FIG. 1, the input content is the character "Figure", and the reference style is a regular script. The output content of the network decoder has not changed, but the style has become very similar to the regular script, that is, the actual style of the "Chen" output on the right is close to the regular script.
[0050] As can be seen from the foregoing introduction, font generation models such as the DG-Font network disclosed in the related art can often only generate fonts of specific styles. For example, if the provided reference style is regular script, the model can only generate fonts in the regular script style. If the provided reference style is running script, the model can only generate fonts in the running script style. Therefore, at present, the above font generation models often难以满足用户针对字体风格的多样化需求,如用户无法自行调整所生成字体的风格等,亟待改进。
[0051] If the single input of the style branch of the pre-trained DG-Font network is directly changed to multiple inputs without adjusting the model parameters, the quality of the generated fonts is generally poor and there are many defects. As shown in FIG. 2, different-style synthesized font images corresponding to the regular script character "Shao" are generated by the multi-input method, where the handwritten style and the regular script style are used as styles with different weights to generate multiple synthesized font images. It can be seen that most of the generated fonts have defects (such as incorrect number of strokes and / or relative positions), and only a very small part are normal fonts. Obviously, this method is not advisable.
[0052] To address the above-mentioned technical problems in the related art, the embodiments of the present invention propose a new method for training a font generation model, and, for a second font generation model trained by the above-mentioned method, a font generation method based on the model is also proposed. The "vector" described in the present invention is a "feature vector." Given that the embodiments of the font generation model training method and the font generation method described in the present invention involve many concepts, to avoid confusion, some of the concepts and their meanings that appear below are first recorded in Tables 1 and 2 below to facilitate reading below.
[0053] Table 1 Concepts and meanings involved in this invention
[0054] The training method of the aforementioned font generation model and the font generation method are described in detail below with reference to the accompanying drawings.
[0055] Figure 3 is a flow chart illustrating a font generation model training method according to an exemplary embodiment of the present invention. This training method can be applied to a model training system, which can include a client used by developers and a corresponding server. The server can be used to centrally manage training resources such as GPUs to efficiently complete the model training process. As shown in Figure 3, the method includes the following steps 302-304.
[0056] Step 302: Use the first training sample to train the original font generation model to obtain a first font generation model, wherein the original font generation model includes a content branch, a style branch, and a decoder. The first sample content font image and the first sample style font image in the first training sample are input into the content branch and the style branch, respectively. The decoder is used to generate and output a first synthetic style font image corresponding to the first training sample based on the first sample text content vector extracted from the content branch and the first sample text style vector extracted from the style branch.
[0057] First of all, it should be noted that, whether it is the pre-training process of the original font generation model or the formal training process of the first font generation model, the sample content font image input to the content branch and the sample style font image input to the style branch can be regarded as the reference images of the two branches respectively. Based on the above reference images, reverse learning can be performed during the training process of the model. For example, the features of the synthetic style font image finally output by the decoder can be compared with the above reference image to calculate the corresponding loss value loss according to the loss function of the model, and the loss value is used for the next training. It can be seen that the model training process described in the present invention (including the aforementioned pre-training and formal training process) is essentially supervised training.
[0058] In addition, the text in the font image described in the present invention can be text in any language, such as Chinese characters, English, Japanese, Korean, etc., and the present invention is not limited to this. In order to avoid the training process from failing to converge, the training samples need to have good consistency. Therefore, the texts corresponding to all the training samples used in a complete training process (that is, the complete process of obtaining the final second font generation model through pre-training and formal training from the original font generation model) can belong to the same language. For example, if all Chinese characters are used as training samples, the second font generation model finally trained can be used to generate a synthetic style Chinese character image; if all English is used as training samples, the second font generation model finally trained can be used to generate a synthetic style English image, etc., which will not be repeated. Of course, according to actual needs, the texts corresponding to all the training samples used in a complete training process can also belong to different languages, and the present invention is not limited to this. The following embodiments are explained using the example of all training samples being Chinese characters.
[0059] In one embodiment, the original font generation model described in the present invention can be a DG-Font network. In other words, the present invention can use the original DG-Font network as the initial training object for training. Of course, the original font generation model can also be constructed using other GAN networks other than the DG-Font network, or using any suitable type of deep learning network, and the present invention is not limited to this. The following embodiments use the DG-Font network for exemplary description.
[0060] In the case where the original font generation model is a DG-Font network, in order to improve the performance of the second font generation model finally obtained by training, the internal structure of the DG-Font network can be appropriately modified. For example, in view of the fact that this solution only focuses on the glyph (corresponding to the style) and content of the text in the font image and does not pay attention to its color, and the glyph of the text can be represented in black and white, in order to reduce the amount of calculation of each network layer within the model to save computing resources, the three-channel input of the DG-Font network can be modified to a single-channel input, so that the grayscale values 0 (black) and 255 (white) are used to represent the font handwriting. For example, in a font image with black text on a white background, the grayscale value of the white background is 255, and the grayscale value of the font area is 0. Of course, if you need to output a color font, you can also set the number of channels corresponding to the corresponding color (such as RGB three channels, CMYK four channels, etc.), which will not be repeated here.
[0061] For another example, in order to avoid the problem that the stroke spacing of the generated synthetic style font picture is too small due to the picture size being too small, which in turn causes the font to be blurred, the size of the input picture of the DG-Font network (i.e., the sample content font picture and the sample style font picture) can be uniformly adjusted from the default value of 80*80 to 160*160, that is, the size of the input picture is doubled (the area is increased to four times the original). In addition, the size of the font in the picture can be set to be slightly smaller than the size of the picture, such as 140*140, so as to present a better blank spacing area around the font in the synthetic style font picture, thereby improving the beauty of the font. For another example, in order to make the picture feature vector used in the font generation model more comprehensive and rich in reflecting the style characteristics of the font, the size of the style feature vector involved in the DG-Font network can also be uniformly adjusted from 128 dimensions to 256 dimensions, which will not be repeated.
[0062] In one embodiment, before training the font generation model, a font image library containing multiple candidate font images can be constructed first. For example, at least one standard style ttf (True Type Font) file can be collected from authorized public or private channels, and the collected ttf file can be used to generate candidate font images of the corresponding standard style. For example, ttf files of multiple font styles such as Songti, Kaiti, New Roman, Bold, and Lishu can be used to generate candidate font images of corresponding font styles. For another example, it is also possible to obtain candidate font images of at least one handwriting style generated by handwriting, which can be specifically generated by natural person handwriting or generated by a preset handwriting font generation model. In the above manner, a font image library containing candidate font images of multiple font styles can be constructed, and the number of candidate font images belonging to any font style in the library can be multiple. For example, when the font is Chinese characters, candidate font images corresponding to the 3755 Chinese characters in the first-level font library can be generated for each font style. For example, the font image library may include Chinese character images in 1,000 font styles, with 3,755 Chinese character images for each font style corresponding to 3,755 common Chinese characters. The font image library contains approximately 3 million candidate font images in total. Of course, the number of candidate font images included in any two font styles can be the same or different, and can be adjusted based on actual circumstances. This is not a limitation of the present invention.
[0063] As shown in FIG. 4, it is a schematic diagram of each candidate font image belonging to a certain font style in the font image library. It can be seen that each candidate font image has the same font style (i.e., the certain font style), while the font content of each candidate font image is different from each other, that is, each candidate font image represents a different Chinese character (such as the first character in the first row is the character "图", the second character in the first row is the character "括", etc.). As shown in FIG. 5, it is a schematic diagram of multiple candidate font images belonging to different font styles and having the same font content in the font image library. It can be seen that each of the above-mentioned multiple candidates represents the character "图", and any two of them belong to different font styles respectively. It can be understood that these different "图" characters are generated using different ttf files, so their styles are different. Continuing from the foregoing embodiments, the size of each candidate font image in the font image library is 160*160, the size of the Chinese character in any image is 140*140, and the image is a single-channel image with a white background and black characters.
[0064] Before training the original font generation model, it is necessary to first obtain the first training sample, that is, to obtain the first sample content font image and the first sample style font image included in the first training sample. Similarly, before training the first font generation model in the subsequent step 304, it is also necessary to first obtain the second training sample, that is, to include obtaining the second sample content font image and the second sample style font image included in the second training sample. Of course, considering that the complete training process includes the above two stages, the second training sample can be obtained before pre-training (that is, after obtaining the first training sample and the second training sample and then starting pre-training), or the second training sample can be obtained after pre-training and before the formal training starts. The present invention does not limit this.
[0065] In one embodiment, for any one of the first training sample and the second training sample, a sample content font image and a sample style font image in the training sample can be obtained from the aforementioned font image library. The first training sample and the second training sample can be obtained from the same font image library (the following embodiment is used as an example for explanation), or the first training sample and the second training sample can be obtained from different font image libraries. As mentioned above, all candidate font images in any font image library belong to multiple font styles, and each candidate font image belongs to one font style. Any two candidate font images may belong to the same or different font styles. Taking the acquisition of any training sample from any font image library as an example, any candidate font image belonging to the first font style can be selected from all candidate font images (contained in the font image library) as the sample style font image in the training sample, and any candidate font image (the font content of the image can be the same or different from the font content of any candidate font image of the aforementioned font style) can be selected from the candidate font images belonging to the second font style as the sample content font image in the training sample, wherein the first font style and the second font style are different font styles (i.e., the two font styles are different from each other). In this way, two candidate font images can be selected from the font image library as sample content font images and sample style font images, respectively, to constitute any of the training samples. At this time, any of the training samples actually contains the two selected images, and these two images are respectively input into the content branch and the style branch during the training phase.
[0066] Alternatively, any candidate font picture can be selected from all the candidate font pictures (contained in the font picture library), and the picture can be used as the sample content font picture and the sample style font picture in any of the training samples. In this way, you only need to select any candidate font picture in the font picture library, and you can use the picture as both the sample content font picture and the sample style font picture, thereby constituting any of the training samples - at this time, any of the training samples actually contains only one picture, that is, any of the selected candidate font pictures, which are input into the content branch and the style branch respectively during the training phase. It can be seen that this method requires fewer candidate font pictures to be selected, thereby simplifying the construction speed of the training samples and reducing their data volume, which helps to improve the overall training efficiency of the font generation model.
[0067] Among them, when selecting any candidate font image from all the candidate font images contained in the font image library, it can be selected randomly. Alternatively, in view of the fact that the number of the first training sample or the second training sample is multiple, multiple candidate font images can be selected from all the candidate font images contained in the font image library to form the corresponding training samples. In the process of selecting multiple candidate font images for constructing any training sample, each candidate font image can be randomly selected, or as many candidate font images as possible that are marked as high-frequency fonts (such as fonts corresponding to multiple Chinese characters that are relatively more frequently used among the aforementioned 3755 Chinese characters) can be selected. It is also possible to uniformly select multiple candidate font images in the image library at equal intervals, etc. The embodiment of the present invention does not limit this.
[0068] It should also be noted that the number of first training samples and second training samples can be multiple. Taking the first training sample as an example, a first sample style font image obtained in the above manner and its corresponding first sample content font image together constitute a first training sample. The pre-training process requires the use of multiple first training samples. The second training sample is similar and will not be described in detail. The pre-training process in the following embodiments takes any first training sample as an example, and the formal training process takes any second training sample as an example.
[0069] After obtaining the first training sample in the aforementioned manner, the original font generation model can be trained using the sample. For the content branch, style branch and decoder contained in the original font generation model, when using any first training sample for pre-training, the first sample content font image in the training sample can be input into the content branch, and the first sample style font image therein can be input into the style branch. Correspondingly, the content branch is used to extract the first sample text content vector of the first sample content font image (i.e., perform encoding processing) and pass it into the decoder; the style branch is used to extract the first sample text style vector of the first sample style font image (i.e., perform encoding processing) and pass it into the decoder. Further, the decoder is used to decode the received first sample text content vector and the second sample font style feature vector to generate a first synthetic style font image, and output the image as the result corresponding to any one of the first training samples. Exemplarily, the first sample text content vector extracted by the content branch can be 20*20*256 dimensional, and the first sample text style vector extracted by the style branch can be 1*256 dimensional.
[0070] The above explanation is only explained by taking the completion process of the input and output of any first training sample as an example. In the actual training process, multiple first training samples can be grouped and iterative training can be performed in groups. For example, the number of training epochs (an epoch means sending a group of samples into the network, and the network completes a forward calculation and back propagation process) can be set to 240, and the number of iterations iter for each epoch is set to 1500 times, that is, a total of 240 rounds of training are performed, and each round of training is iterated 1500 times (each time using a group of first training samples). For example, when the number of GPU cards in the training resources is 4, the batchsize can be set to 4*32 (that is, the 128 samples contained in each group are evenly distributed to 4 GPUs for calculation). The embodiment of the present invention does not limit the type of training resources (such as GPU models, etc.), and can be reasonably set according to actual needs.
[0071] In one embodiment, multiple first training samples can be used to perform multiple rounds of supervised training on the original font generation model. The number of the first training samples is multiple, and the multiple first training samples include at least one standard training sample and at least one handwritten training sample. The first sample style font image in the standard training sample is generated by a standard style ttf file, and the first sample style font image in the handwritten training sample is generated by handwriting. Based on this, when the first training samples are used to pre-train the original font generation model, if the number of training rounds of the original font generation model does not reach a second round number threshold, the standard training samples can be used to train the original font generation model; and if the number of training rounds of the original font generation model reaches the second round number threshold, at least the handwritten training samples can be used to train the original font generation model. The at least using the handwritten training samples for training includes using only the handwritten training samples for training in the remaining training rounds, or using both the handwritten training samples and the standard training samples for training in the remaining training rounds. It can be seen that when the number of supervised training rounds has not reached the round number threshold (such as the first 200 epochs out of 240 epochs), standard-style font images generated by ttf files are used for training; until the number of supervised training rounds reaches the round number threshold (such as the last 40 epochs out of 240 epochs), handwriting-style font images generated by collecting writing actions are used for training.
[0072] This approach introduces the handwriting style of a handwritten font during training, enabling the trained first font generation model to distinguish font images with handwritten styles. This improves the model's accuracy in extracting handwriting style features and generating synthetic font images that match the handwriting style. Furthermore, by introducing handwritten training samples after the aforementioned number of training rounds reaches a threshold, the pre-training process for the original font generation model becomes more stable, avoiding excessive fluctuations that could hinder convergence of the output style features, leading to training failure or excessive training time. This helps improve the efficiency of model training.
[0073] As previously mentioned, the original font generation model can be a DG-Font network, in which the style branch consists of a VGG1 network, an average pooling layer, and a fully connected layer. In fact, some embodiments of the present invention further improve the internal structure of the style branch to further enhance its ability to extract features from fonts of different styles.
[0074] In one embodiment, the style branch may include a CNN (Convolutional Neural Networks) and a connected FPN (Feature Pyramid Network). The CNN and the FPN each include S (S is an integer greater than 1) network layers, and there is a one-to-one lateral connection between the S network layers in the CNN and the S network layers in the FPN. Based on this structure, for any font image input into the style branch, the S network layers in the FPN can cooperate with the S network layers in the CNN to extract and output a feature vector of S dimensions as the text style vector of the any font image output by the style branch.
[0075] Figure 6 shows the internal structure of the FPN. The CNN on the left (i.e., the aforementioned VGG1 network) is used for bottom-up forward feature extraction, while the FPN on the right is used for top-down upsampling. In the FPN, each stage corresponds to a level of the feature pyramid, and the final layer of features in each stage is selected as the features corresponding to that level in the FPN. Lateral connections combine the upsampled features of the previous layer with the same resolution as the current layer through summation, thereby gradually propagating semantic information from higher layers to lower layers. The complete structure of the improved style branch is shown in Figure 7. Through these improvements, the style branch leverages both the strong semantic features of the top layer (which facilitates classification) and the high-resolution information of the bottom layer (which facilitates localization). This allows the model to simultaneously focus on local and global style information, improving its ability to perceive style features at different levels. The text style vectors extracted by the improved style branch achieve small intra-class distances and large inter-class distances. Specifically, the text style vectors of different font images belonging to the same font style are close in distance, while the text style vectors of different font images belonging to different font styles are far apart.
[0076] In another embodiment, a Transformer encoder can be added to the original font generation model, such as the original font generation model can also include a content Transformer encoder and / or a style Transformer encoder. The content Transformer encoder is used to connect the output end of the content branch and the input end of the encoder, and the style Transformer encoder is used to connect the output end of the style branch and the input end of the encoder. Taking the original font generation model including the content Transformer encoder and the style Transformer encoder as an example, the positions of the two encoders (i.e., Style Transformer Encoder and Content Transformer Encoder) in the original font generation model are shown in Figure 7. It should be noted that the improved font generation model also includes FDSC-1 and FDSC-2, and the specific structure of FDSC can still be seen in (b) of Figure 1.
[0077] The internal structure of any Transformer encoder is shown in Figure 8. Taking the style Transformer encoder as an example, the output dimension of the style feature after passing through the encoder remains consistent with the input dimension (i.e., the vector dimension remains unchanged). The introduction of the Transformer encoder improves the model's ability to perceive the correlation between different regions, thereby improving the model's feature expression capabilities.
[0078] In addition, when the number of supervised training rounds reaches a round threshold, at least during the training process using handwritten training samples, a contrast loss function can be calculated between different handwritten training samples and their corresponding first synthetic style font images based on the feature deviation between the two, and the function can be used as part of the loss function of the font generation model. Given that different people often have different handwriting styles, while different fonts written by the same person usually belong to the same handwriting style, the introduction of the contrast loss function can make the style vectors extracted by the style branch for different handwriting font images written by the same person as similar as possible, while the style vectors extracted for different handwriting font images written by different people are as different as possible, thereby enhancing the style branch's ability to distinguish different writing styles / handwriting.
[0079] In one embodiment, the loss function of the first font generation model may include a style loss function. For example, the first sample text style vector of the first sample style font image and the first synthetic text style vector of the first synthetic style font image may be obtained, and the loss value of the style loss function may be determined based on the deviation between the first sample text style vector and the first synthetic text style vector. The deviation between the first sample text style vector and the first synthetic text style vector may be represented by an indicator such as the Euclidean distance between the two vectors, the cosine similarity, or the SSIM (Structure Similarity Index Measure) of the target handwritten font image, which will not be described in detail.
[0080] It is understandable that the above deviation can be used to reflect the degree of style difference between the synthesized style font image output by the model during training and the first sample style font image input to the model (for example, the larger the Euclidean distance between the two vectors, the greater the style difference between the two font images). Therefore, the loss value of the style loss function can be determined based on the deviation and used as the loss function of the model, or as part of the loss function of the model (in this case, the loss function of the model may also include cycle consistency loss, standard MSE (Mean Square Error) loss, noise loss, etc.). In this way, the style of the synthesized style font image can be better constrained to remain consistent with the style of the first sample style font image (i.e., the reference style at this time), that is, the actual style of the synthesized style font image is as close as possible to the font style of the first sample style font image.
[0081] When obtaining the first sample text style vector of the first sample style font image and the first synthesized text style vector of the first synthesized style font image, the first sample style font image can be input into a feature extraction network including multiple network layers, and the feature vector extracted by at least one network layer (of the multiple network layers) can be obtained as the first sample text style vector (of the first sample style font image). Furthermore, the first synthesized style font image can be input into the feature extraction network, and the feature vector extracted by the at least one network layer can be obtained as the first synthesized text style vector (of the first synthesized style font image).
[0082] Among them, the process of extracting the first sample text style vector and the first synthetic text style vector using the feature extraction network is similar, and the first sample text style vector is used as an example for explanation. If only one feature vector is extracted from one network layer, the feature vector can be used as the first sample text style vector. If a feature vector is extracted from multiple network layers respectively (that is, multiple feature vectors are extracted), the mean of these feature vectors can be used as the first sample text style vector. Among them, the mean can also be an arithmetic mean or a weighted mean. In the case where the mean is a weighted average, the weight value of the feature vector corresponding to any network layer can be related to the position of the network layer in the multiple network layers. For example, the later the network layer is, the greater the weight value of the feature vector corresponding to it, so as to minimize the accumulation effect of style deviation and improve the accuracy of the extracted first sample text style vector.
[0083] Exemplarily, the feature extraction network can be the discriminator shown in (c) of Figure 1, which includes multiple network layers. Taking the last two layers as an example: the first sample style font image can be input into the discriminator, and the feature vectors (hereinafter referred to as the two first vectors) output by the last two layers are obtained; and the first synthetic style font image is also input into the discriminator, and the feature vectors (hereinafter referred to as the two second vectors) output by the last two layers are obtained. To this end, the mean of the two first vectors can be calculated to obtain the first mean vector, and the mean of the two second vectors can be calculated to obtain the second mean vector, and then the deviation (such as the Euclidean distance) between the first mean vector and the second mean vector is calculated as the deviation between the first sample text style vector and the first synthetic text style vector. Alternatively, the deviation of the corresponding vectors in the two first vectors and the two second vectors (the vectors corresponding to the same network layer are a pair of corresponding vectors) can be calculated, and then the mean of the two deviations is calculated as the deviation between the first sample text style vector and the first synthetic text style vector. Of course, the above deviation calculation method can be flexibly selected or adjusted according to actual conditions, and the present invention is not limited to this.
[0084] Alternatively, the feature extraction network can be a pre-trained CNN feature extraction network, such as a VGG network, having multiple network layers. In this case, the first sample font style image and the first synthetic font style image can be input into the network, respectively, and feature vectors output by at least one of the network layers can be obtained. The deviation between the first sample text style vector and the first synthetic text style vector can then be calculated based on these feature vectors. The specific calculation method can be found in the previous embodiment and will not be repeated here.
[0085] In one embodiment, the loss function of the first font generation model may include a connectivity loss function, the value of which is used to characterize the connectivity difference between the first synthetic style font image and the first synthetic style font image. This difference can be used to evaluate the severity of the continuity, breakage, and connectivity of the first synthetic style font. The connectivity loss function may include at least one of the following loss functions: for example, a connected component loss function. The loss value of the function may be determined by first determining the number and size of connected components in the first synthetic style font image, and then determining the number and size of connected components in the first sample style font image. The number and size of connected components in the two images are then compared to determine their deviation, thereby measuring the connectivity difference between the first sample style font image and the first synthetic style font image using the deviation. For example, if the deviation is the mean square error (MSE), a larger MSE indicates a greater connectivity difference between the two images, and the effect of the first synthetic style font image may be considered to be worse. Another example is the stroke path loss. The loss value of this function can be calculated based on the glyph features such as intersection, overlap, and spacing in the stroke paths of the first sample style font image and the first synthetic style font image. It is used to measure the consistency of the stroke paths of the two images. The specific calculation process can be found in the records in the relevant technology and will not be repeated here. Another example is the skeletonization loss function. To calculate the loss value of this function, it is necessary to use the skeletonization algorithm to convert the first sample style font image and the first synthetic style font image into skeleton shapes respectively, and then calculate the connectivity and detail information of the two skeleton shapes. Finally, the loss value is determined by the difference between the above information.
[0086] For another example, the path continuity loss function, the loss value of this function can be calculated by the continuity, breakage and connection of the stroke paths of the first sample style font image and the first synthetic style font image, such as a continuous path and fewer breaks correspond to a smaller loss value. The loss value of this function can be represented by the Hausdorff distance. Specifically, the sample position coordinates of each sample boundary point on the font outline of the first sample style font image and the synthetic position coordinates of each synthetic boundary point on the font outline of the first synthetic style font image can be determined; then the Hausdorff distance between the sample boundary point and the synthetic boundary point is calculated using the sample position coordinates and the synthetic position coordinates as the loss value of the connectivity loss function.
[0087] Before determining the font outlines of the first sample style font image and the first synthesized style font image, the two font images can be pre-processed by binarization, denoising, or resizing to facilitate subsequent feature extraction. Appropriate path extraction algorithms can then be used to extract the stroke paths in the two font images. Common path extraction algorithms include edge detection algorithms, projection algorithms, and connectivity-based algorithms. Edge detection algorithms can extract stroke boundaries in font images, and commonly used edge detection algorithms include Canny edge detection and Sobel operator edge detection. Projection algorithms can project the stroke paths in the font image horizontally or vertically to obtain stroke position and length information. Connectivity-based algorithms can connect pixels in the font image into paths and group adjacent pixels into stroke paths based on the connectivity relationship between pixels, such as using a connected component analysis algorithm to identify and extract connected pixel clusters as stroke paths. The specific use of the above-mentioned algorithms can be found in the relevant technical records and will not be repeated here.
[0088] For example, in order to extract the font outline of black characters on a white background in a color font image, the color font image is first converted into a grayscale image, and then the grayscale image is binarized to obtain a black and white image, in which the font area is pure black (pixel value is 0) and the background area is white (pixel value is 255). Then use the image processing library (such as opencv) or the outline extraction function in the algorithm to detect the font outline in the black and white image. Of course, specific standards or rules can also be used to filter out the final font outline from the above-mentioned font outlines extracted preliminarily. As shown in Figure 9, the Chinese characters "hungry", "E", "e", and "peak" in the upper row are the original font images (such as the aforementioned first sample style font image or the first synthetic style font image), and the lower row is the font outline corresponding to each extracted font.
[0089] After extracting the font outlines of the first sample style font image and the first synthetic style font image, the sample position coordinates of each sample boundary point on the font outline of the first sample style font image (i.e., the coordinates of each pixel point on the font outline in the image) and the synthetic position coordinates of each synthetic boundary point on the font outline of the first synthetic style font image (i.e., the coordinates of each pixel point on the font outline in the image) can be determined respectively; then the Hausdorff distance between the sample boundary point and the synthetic boundary point is calculated using the sample position coordinates and the synthetic position coordinates as the loss value of the connectivity loss function. Among them, the Hausdorff distance is the maximum distance from a point in a set (i.e., the sample boundary points) to the nearest point in another set (i.e., the corresponding nearest point in the synthetic boundary points). The specific calculation method of the Hausdorff distance can be found in the records in the relevant technology and will not be repeated here. The calculated Hausdorff distance is used to represent the degree of similarity in path connectivity between the font outline of the first sample style font image and the font outline of the first synthesized style font image. The smaller the distance, the greater the degree of similarity in path connectivity between the first sample style font image and the first synthesized style font image, and vice versa.
[0090] Through the aforementioned training and loss compensation process, the original font generation model can be trained to obtain a first font generation model. Compared to the original font generation model, the first font generation model has the ability to extract features from font images, that is, the content branch and style branch have the ability to extract font content features and font style features, respectively.
[0091] Step 304: Use a second training sample to train the first font generation model to obtain a second font generation model. The second training sample includes a second sample content font image and a second sample style font image. The second sample content font image is input into the content branch, and the second sample text content vector extracted from the content branch is input into the decoder. The sample weight style vector of the second sample style font image is input into the decoder. The sample weight style vector is obtained by weighted summation of the sample basic font style vectors of multiple sample basic font styles. The decoder is used to generate and output a second synthetic style font image corresponding to the second training sample based on the second sample text content vector and the sample weight style vector.
[0092] It is understandable that because the original font generation model to be trained includes a content branch, a style branch, and a decoder, and the training process does not change the internal structure of the model, the pre-trained first font generation model and the formally trained second font generation model also include the content branch, style branch, and decoder, respectively. Hereinafter, the content branch, style branch, and decoder described in the various embodiments corresponding to step 304 for the formal training process should be understood as the content branch, style branch, and decoder in the first font generation model. This is for clarification.
[0093] Similar to the aforementioned pre-training process, when the first font generation model is formally trained using the second training sample, the second sample content font image and the second sample style font image in the second training sample can be processed separately. The second sample content font image can be input into the content branch for feature extraction, and the second sample text content vector extracted by the content branch is input into the decoder. As shown in Figure 1, the second sample text content vector output by the content branch is directly input into the decoder; as shown in Figure 7, the second sample text content vector output by the content branch is processed by the content Transformer encoder before being input into the decoder.
[0094] For the second sample style font image, a sample weight style vector of the font image may be obtained and input into the decoder, wherein the sample weight style vector is obtained by weighted summation of sample basic font style vectors of multiple sample basic font styles.
[0095] Before obtaining the sample weight style vector of the second sample style font picture, it is necessary to first determine the sample basic font style vectors of each of the multiple sample basic font styles. For example, multiple basic font styles can be randomly selected from a plurality of font styles, or, in order to ensure that the selected multiple basic font styles can reflect the various font styles as accurately as possible, multiple basic font styles with as large a style difference as possible can also be selected from a plurality of font styles. Among them, the multiple font styles (i.e., candidate font styles) used to select the basic font style can be all or part of the font styles corresponding to the font picture library (i.e., all or part of the font styles to which the candidate font pictures contained in the font picture library belong), and the candidate font pictures in the font picture library are used to select the second sample content font picture and the second sample style font picture to constitute the second training sample.
[0096] In one embodiment, a font picture library for selecting sample basic font style vectors can be determined first. It is assumed that all candidate font pictures in the font picture library belong to M2 font styles, where each candidate font picture belongs to one font style, and the second sample style font picture is selected from the font picture library (that is, the font picture library is a font picture library for selecting the second sample style font picture), m2≤M2, and m2 and M2 are both positive integers. The font picture library can be pre-created in the manner described in the previous embodiment and will not be repeated here. Based on this, the font style vectors of each of the M2 font styles can be obtained (that is, M2 font style vectors corresponding to the M2 font styles are obtained), and the sample basic font style vectors of each of the m2 sample basic font styles can be determined from the obtained M2 font style vectors. Among them, the style branch of the aforementioned first font generation model can be used to extract the text style vectors of each candidate font picture in the font picture library, and the font style vectors of each of the M2 font styles can be calculated based on these extracted text style vectors. For example, all candidate font images are input into the style branch in sequence, and the text style vectors of each candidate font image output by the branch are obtained; for the text style vectors of S2 candidate font images belonging to the same font style, the average value (such as the arithmetic mean) of these S2 text style vectors can be calculated, and the calculation result (the average value of multiple vectors is still a vector) is used as the font style vector of the font style.
[0097] In addition, a variety of methods can be used to determine the sample basic font style vectors of each of the m2 sample basic font styles from the M2 font style vectors obtained. For example, the PCA (Principal Component Analysis) algorithm can be used to analyze the M2 font style vectors obtained, and the m2 main feature vectors obtained by analysis can be determined as the sample basic font style vectors of each of the m2 sample basic font styles. For another example, a clustering algorithm can be used to cluster the calculated M2 font style vectors, and the m2 cluster centers obtained by processing can be determined as the sample basic font style vectors of each of the m2 sample basic font styles (these m2 cluster centers are the m2 sample basic font style vectors, and the corresponding font styles are the corresponding sample basic font styles). The clustering algorithm can use Kmeans, Mini-Batch K-means, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), Expectation Maximization (EM) algorithm, etc. It is understandable that, because the m2 cluster centers are relatively evenly distributed in the M2 style feature vectors, the m2 sample basic font styles finally obtained can fully and comprehensively represent the M2 font styles as much as possible. Of course, m2 font styles can also be randomly selected from the M2 font styles as sample basic font styles, which will not be repeated here. For example, the M2 font styles can be Songti, Kaiti, New Roman, Heiti, Lishu, Fangzheng, Xingkai, etc., and the m2 basic font styles selected in the above manner can be Kaiti, New Roman, Lishu, etc.
[0098] For example, if the font image library corresponds to 1,000 font styles, then 10 font styles can be selected from these 1,000 font styles as sample base font styles in the above manner. The font style vectors of each of the 10 selected font styles become the sample base font style vectors of these 10 sample base font styles. For example, assuming that the 1,000 font styles include Songti and the style vector of Songti is T, then if Songti is selected as the sample base font style, the vector T becomes the sample base font style vector of Songti, the sample base font style.
[0099] In one embodiment, the second sample style font image can be processed in a variety of ways during the formal training process. For example, a sample font style vector of the sample font style to which the second sample style font image belongs can be first obtained, and sample basic font style vectors corresponding to the multiple sample basic font styles can be obtained; sample weight values corresponding to the multiple sample basic font styles can then be calculated based on the vector distances between the font style vector and each sample basic font style vector; and sample basic font style vectors of the multiple sample basic font styles can be weighted and summed according to the sample weight values to obtain the sample weight style vector. Finally, the sample weight style vector can be input into the decoder.
[0100] For example, following the above-mentioned embodiment of "m2 basic font styles are Kaiti, New Roman and Lishu", if the sample font style of the second sample style font image is "Songti", then the sample font style vector T of Songti is first obtained. 宋体 And the sample basic font style vector T0 of Kaiti, New Roman and Lishu respectively 楷体 、T0 新罗马 and T0 隶书 , and then calculate T Songti and T0 respectively 楷体 、T0 新罗马 and T0 隶书 The vector distances L1, L2 and L3 of each font are respectively determined, and then the sample weight values k1, k2 and k3 corresponding to Kaiti, New Roman and Lishu are determined according to the vector distances L1, L2 and L3 (such as those determined by normalization and softmax function and other related algorithms, for details, please refer to the relevant technical records). Finally, the weighted summation of the basic font style vectors of each sample is performed according to the weight value to obtain the sample weight style vector Tw of Songti. 宋体 =k1*T0 楷体 +k2*T0 新罗马 +k3*T0 隶书 .
[0101] It can be understood that the sample weight style vector of the second sample style font image determined by this method is used to reflect the weight relationship between the font style of the font image and multiple sample basic font styles (the weight relationship between styles), and each weight value in the sample weight style vector is positively correlated with the proportion of the multiple basic font styles in the font style of the font image.
[0102] In addition, this method requires pre-determining and storing the following font style vectors: the basic font style vectors of the plurality of basic font styles (such as the aforementioned T0 楷体 、T0 新罗马 and T0 隶书), and the font style vectors of at least one font style to which the candidate font images that may be selected as the second sample style font images belong (such as the aforementioned T 宋体 ). Based on this, during the formal training process, for any second sample style font image actually selected, the font style vector of the font style to which the font image belongs, which is pre-stored, can be directly read, without inputting the second sample style font image into the style branch, which helps to speed up the training speed of the formal training. It can be understood that at this time, for the same font style, no matter which font image belonging to this style is selected as the second sample style font image, the sample weight style vector of the second sample style font image is the same. Continuing from the foregoing embodiment, whether the second sample style font image is the character 'good' in Song typeface or the character 'bad' in Song typeface (or any other character in Song typeface), as long as the font image belongs to Song typeface, its sample weight style vector is the aforementioned Tw 宋体 . In this regard, it should be noted when creating the second training sample: If it is determined to calculate the sample weight style vector in the above manner during the subsequent formal training process, it should be ensured that the second sample content font images included in each second training sample are different from each other, so as to substantially form the second training sample.
[0103] For another example, the second sample style font image can also be input into the style branch to obtain the sample text style vector of the second sample style font image extracted by the style branch, and obtain the sample basic font style vectors of each of the multiple sample basic font styles; and, calculate the sample weight values corresponding to each of the multiple sample basic font styles according to the vector distances between the sample text style vector and each of the sample basic font style vectors respectively; furthermore, perform weighted summation on the sample basic font style vectors of each of the multiple sample basic font styles according to the sample weight values to obtain the sample weight style vector, and input the sample weight style vector into the decoder.
[0104] Exemplarily, still continuing from the foregoing embodiment of'm2 basic font styles are regular script, Times New Roman, and clerical script', if the second sample style font image is the character 'good' in Song typeface, then during the training process, first input this image into the style branch to obtain the sample text style vector T_Song(good) extracted by this branch, and obtain the sample basic font style vectors T0 楷体 、T0 新罗马 and T0 隶书 of regular script, Times New Roman, and clerical script respectively, and then calculate T 宋体(佳) and T0 楷体 、T0 新罗马 and T0 隶书Their respective vector distances L4, L5, and L6, and then determine the sample weight values k4, k5, and k6 corresponding to regular script, Times New Roman, and clerical script respectively according to the vector distances L4, L5, and L6 (the specific method is the same as above). Finally, perform a weighted sum of each sample base font style vector according to the weight values to obtain the sample weight style vector Tw of the Song typeface "Jia" character 宋体(佳) = k4 * T0 楷体 + k5 * T0 新罗马 + k6 * T0 隶书 .
[0105] As shown in Figure 10, the character "图" a on the left (actually a picture, the same below) is a Chinese character in Song typeface, and the ten different style characters "图" b0~b9 on the right are used to represent ten base font styles, and the weight values k0~k9 are the weight values corresponding to these ten base font styles respectively. Let's assume that the sample text style vector T of the character "图" a 宋体 (图) = [a0, a1,..., a255], the base font style vector T0 of the character "图" b0 belongs to = [b00, b01,..., b0255], the base font style vector T1 of the character "图" b1 belongs to = [b10, b11,..., b1255],... the base font style vector T9 of the character "图" b9 belongs to = [b90, b91,..., b9255], then the sample text style vector T 宋体(图) = k0 * T0 + k1 * T1 +... + k9 * T9.
[0106] Actually, similar to the relationship between the vectors shown in Figure 6, for the sample font style vector T of the Song typeface to which the character "图" a belongs 宋体 (a 1*256-dimensional vector), the sample base font style vectors T0~T9 of the base font styles to which the characters "图" b0~b9 belong respectively (1*256-dimensional vectors respectively), and the weight values of the sample font style vector T of the Song typeface relative to the sample base font style vectors T0~T9 are K0, K1, K2,..., K9 respectively, then the vectors satisfy the relationship: T 宋体 = K0 * T0 + K1 * T1 +... + K9 * T9.
[0107] It can be understood that the sample weight style vector of the second sample style font image determined by this method is used to reflect the weight relationship between the font image itself and multiple sample basic font styles (the weight relationship between text and style), and each weight value in the sample weight style vector is positively correlated with the proportion of the multiple basic font styles in the text style of the font image itself. It should be noted that the text styles of each font under the same font style are the same as the font style, but the text style vector of each font is not necessarily equal to the font style vector of the font style - this is because the font style vector of a certain font style is the average of the text style vectors of each font image belonging to the style.
[0108] In addition, this method requires pre-determining and storing the basic font style vectors of the multiple basic font styles (such as the aforementioned T0 楷体 、T0 新罗马 and T0 隶书 ), and the sample text style vector of the second sample style font picture (as mentioned above T 宋体(佳) ) is obtained by inputting the style branch during the training process.
[0109] Of course, for each candidate font image that may be selected as the second sample font style image, its respective weighted style vector can also be pre-calculated and stored after pre-training is completed and before formal training begins. Therefore, when any candidate font image is selected as the second sample font style image, the stored weighted style vector of the font image can be directly read to serve as the sample weighted style vector of the second sample font style image. No further details will be given.
[0110] In one embodiment, for any of the aforementioned sample font styles and multiple sample basic font styles, the font style vector of the font style can be determined in the following manner: a plurality of candidate font images belonging to any of the aforementioned font styles are respectively input into the style branch of the first font generation model to obtain the text style vectors of the plurality of candidate font images extracted by the style branch, and the average value of the plurality of obtained text style vectors is used as the font style vector of any of the font styles. Among them, the plurality of candidate font images can be all the candidate font images belonging to the font style, or can be part of the candidate font images belonging to the font style, such as high-frequency font images, etc., which will not be repeated. It can be seen from the aforementioned embodiment that the process of obtaining the font style vector can be performed in advance before formal training, or can be performed separately for each second sample style font image during the formal training process, which will not be repeated.
[0111] Similar to the pre-training process, reverse learning can also be performed during the formal training process. In addition to the aforementioned cycle consistency loss, standard MSE loss, and noise loss, the corresponding loss functions can also include one or more of the contrast loss function, style loss function, and connectivity loss function. The specific loss value calculation method is similar to the calculation method in the pre-training process and will not be repeated here.
[0112] It should also be noted that, given that the aforementioned pre-training process has already enabled the content branch and the style branch to have feature extraction capabilities, the reverse learning during the formal training process can shield these two branches, that is, in this reverse learning process, only the internal parameters of the decoder are adjusted without adjusting the internal parameters of the aforementioned two branches, so as to avoid the formal training process from adversely affecting the feature extraction capabilities that the two branches already have. Of course, if there is a need (such as the accuracy of the vector extracted by the content branch and / or style branch does not meet the requirements), the internal parameters of the content branch and / or style branch can also be adjusted during the formal training process. The above adjustment process can be flexibly adjusted according to actual conditions, and the present invention is not limited to this.
[0113] As can be seen from the aforementioned embodiments, the training process of the font generation model includes two stages: the first stage is used to train the original font generation model to obtain a first font generation model, so that the model has the ability to extract features from font images; the second stage is used to train the first font generation model to obtain a second font generation model, so that the model has the ability to reason about handwritten fonts. It can be seen that the first stage is the pre-training process, and the obtained first font generation model is an intermediate state in the complete training process, while the second stage is the formal training process, and the obtained second font generation model is the final training result of the complete training process.
[0114] Compared with the font generation model and its training method in the related art, when this solution uses the second training sample to train the second font generation model, the sample text style vector of the second sample style font image obtained by extracting the style branch is not directly input into the decoder, but the sample weighted style vector of the second sample style font image obtained by weighted summation of the sample basic font style vectors of multiple sample basic font styles is input into the decoder, so that the decoder in the trained second font generation model has the ability to infer handwritten fonts of multiple styles.
[0115] The embodiment of the present invention further provides a font generation model, which is trained by the font generation model training method described in any of the above embodiments. This font generation model is a second font generation model obtained through the above training.
[0116] The following describes in detail a font generation method based on the second font generation model (hereinafter referred to as the font generation model) with reference to the accompanying drawings. FIG7 is a flowchart of a font generation method according to an exemplary embodiment of the present invention. As shown in FIG11 , the method includes the following steps 1102-1106.
[0117] Step 1102: Obtain a target text content vector of a target handwritten font image according to a font generation model.
[0118] In one embodiment, the font generation model can be obtained based on DG-Font network training. The specific training method can be found in the above embodiment.
[0119] In one embodiment, the font generation model includes a content branch, a style branch, and a decoder. The style branch may include a CNN and its connected FPN. There is a one-to-one lateral connection between the S network layers in the CNN and the S network layers in the FPN. For any font image input into the style branch, the S-dimensional feature vectors output by the S network layers in the FPN are used as the text style vector of the font image output by the style branch, where S is an integer greater than 1.
[0120] In one embodiment, the font generation model may further include a content Transformer encoder and / or a style Transformer encoder, wherein the content Transformer encoder is used to connect the output end of the content branch and the input end of the encoder, and the style Transformer encoder is used to connect the output end of the style branch and the input end of the encoder.
[0121] In the actual application stage of the font generation model, the model can be used to provide font generation services. The process of providing font generation services is the process of inferring the target handwritten font image to generate a synthetic style font image. Among them, the font generation model can run in the server (such as as an internal functional component of the server itself, or as an external functional component callable by the server, and the server can run in the server or other servers). At this time, after receiving the font generation request initiated by the client for the target handwritten font image, the server can call the model to generate a synthetic style font image corresponding to the target handwritten font image, and further feed the image back to the client for the user to view and / or use. Alternatively, the font generation model can also run in a terminal (e.g., as an internal functional component of the client itself, or as an external functional component callable by the client, the client can run in the terminal. For example, the terminal can be a mobile phone, and the client can be an APP or mini-program running on the mobile phone). In this case, after the user initiates a font generation request for the target handwritten font image, the client can call the model to generate a corresponding synthetic style font image and output the image to the user for viewing and / or use, etc., which will not be described in detail. The following description takes the font generation model running on the client as an example.
[0122] Users can interact with the client through an interactive page provided by the client. For example, a user can perform a writing action on the handwriting page to input a target handwriting font image. For another example, a user can specify a desired font style on the style indication page to instruct the font generation model to generate a synthetic font image that is close to the desired font style. Of course, the client can also provide a batch input interface for handwriting fonts for users to use, so that users can batch input pre-written or pre-generated handwriting fonts as target handwriting font images.
[0123] It should be noted that a user may write or input multiple handwritten font images corresponding to different characters at once. In this case, each of these handwritten font images can be identified as a target handwritten font image, so that the font generation model can generate a corresponding synthetic style font image for each target handwritten font image. Specifically, the number of target handwritten font images can be one or more: if there is only one target handwritten font image, the font generation model can generate a synthetic style font image corresponding to that image; if there are multiple target handwritten font images, the font generation model can generate a synthetic style font image corresponding to each image. The following embodiments use a synthetic style font image as an example for explanation.
[0124] As shown in FIG. 12, the user continuously writes the four characters "science", "technology", "technique", and "art" in the handwriting area. The target handwritten font image described in this solution can be the font image corresponding to each of these characters respectively (that is, the font image generated by recognizing the handwriting of the above fonts).
[0125] Since the font generation model includes a content branch and a style branch, it is necessary to process the content and style of the target handwritten font image separately. Among them, step 1102 and its related embodiments are the processing processes for content, and step 1104 and its related embodiments are the processing processes for style. In the processing process for content, it is necessary to obtain the target text content vector of the target handwritten font image according to the font generation model.
[0126] In one embodiment, the target content font image of the target handwritten font image can be determined first, and then it is input into the content branch of the font generation model to extract the target text content vector. Among them, the target content font image of the target handwritten font image can be determined in various ways.
[0127] For example, since the above training process uses font images in handwritten style as training samples, the style branch in the trained font generation model has the ability to extract features of font images in handwritten style (that is, it can directly extract the style features of handwritten fonts). Based on this, the target handwritten font image can be used as the target content font image, that is, the target handwritten font image itself (as its target content font image) is input into the content branch of the font generation model, so that the style branch extracts its style vector as the target text content vector.
[0128] For another example, the target reference font image of the target handwritten font image may be obtained as the corresponding target content font image; in other words, the target reference font image of the target handwritten font image is obtained and input into the content branch of the font generation model to extract the target text content vector. The target reference font image belongs to a preset reference font style and has the same text content as the target handwritten font image. The reference font style can actually be any font style, such as regular script, Song typeface, etc., and the content branch has a high recognition accuracy for font images belonging to the reference font style. Exemplarily, the target text content of the target handwritten font image can be recognized first (that is, to recognize which specific character is in the target handwritten font image), and then a font image with the text content being the target text content is queried from the reference font image library corresponding to the reference font style (each font image in this library belongs to the reference font style) as the target reference font image. Among them, the target text content of the target handwritten font image can be recognized by a preset content recognition algorithm; or, the target text content of the target handwritten font image can also be recognized by a trained content recognition model, and this content recognition model can be trained by a supervised learning method and is used to recognize the text content of the characters included in the input handwritten font image.
[0129] Exemplarily, let's assume that the reference font style is regular script, and the reference font image library contains 3,755 Chinese character images in regular script. If the Chinese character in the target handwritten font image is recognized as the handwritten style "图" character, then a regular script "图" character image can be queried from each Chinese character image included in the reference font image library as the target reference font image (of the target handwritten font image). Through this method, the target text content of the target handwritten font image can be accurately recognized, thereby improving the content correctness of the finally output synthetic style font image.
[0130] If the content branch in the font generation model is directly connected to the decoder (as shown in Figure 1), then the target text content vector extracted by the content branch will be directly input into the decoder. If there is also a content Transformer encoder connected between the content branch and the decoder in the font generation model (as shown in Figure 7), then the target text content vector extracted by the content branch will be further processed by this encoder, and the corresponding processing result is input into the decoder.
[0131] Step 1104, input the target handwritten font image into the style branch of the font generation model to extract the target text style vector, and determine the handwritten weight style vector according to the target text style vector. The handwritten weight style vector is obtained by weighted summation of the respective basic font style vectors of multiple basic font styles.
[0132] During style processing, the target handwritten font image is first input into the style branch to extract the target text style vector. Based on this, a weighted sum of the basic font style vectors of multiple basic font styles is then determined, resulting in a handwritten weighted style vector. This vector is then input into the decoder. The basic font style vector for any basic font style is calculated using the basic text style vector of the basic font image belonging to that style, extracted from the style branch.
[0133] For example, before actual use, S basic font images belonging to any basic font style can be respectively input into the style branch of the font generation model to extract the corresponding basic text style vector, and then the average value (such as the arithmetic mean) of the extracted S basic text style vectors is calculated, and the calculation result (the average value of multiple vectors is still a vector) is used as the basic font style vector of any basic font style. For another example, given that the sample basic font style vectors of multiple sample basic font styles are used in the training process of the font generation model, when the multiple basic font styles are the multiple sample basic font styles, the sample basic font style vectors of the multiple sample basic font styles can be directly used as the basic font style vectors of the multiple basic font styles. This not only eliminates the need to calculate the basic font style vector again, but also improves the accuracy of the reasoning result (because the same basic font style vector is used in the reasoning stage and the training stage). Of course, the multiple basic font styles involved in the reasoning stage may not be exactly the same as the multiple sample basic font styles involved in the training stage. In this case, the basic font style vectors of the multiple basic font styles need to be determined before reasoning, which will not be repeated.
[0134] As mentioned above, the number of target handwriting font images can be multiple (such as the user writes multiple Chinese characters continuously). In this scenario, each target handwriting font image can be input into the style branch respectively to obtain the text style vector of each target handwriting font image extracted by the style branch (each target handwriting font image is extracted into a text style vector). In this regard, the average value of each extracted text style vector can be calculated as the target text style vector (corresponding to the multiple target handwriting font images). In this way, the final output synthetic style font image can be made to conform to the overall style of the multiple handwriting fonts continuously input by the user, and the font style is more unified and more beautiful.
[0135] The average value of each of the text style vectors can be a weighted average value. The weight value of any one of the text style vectors can be a preset value. This method helps to simplify the determination logic of the target text style vector, thereby improving the inference speed. The preset value (i.e., the weight value preset for each of the text style vectors) can be the same or different. In the case of being the same, the influence degree of each target handwritten font image on the target text style vector is the same. At this time, the weighted average value is the arithmetic average value. Or, the magnitude of the weight value of any one of the text style vectors can be positively correlated with the stroke complexity of the target handwritten font image corresponding to this text style vector. For example, the weight values of each target handwritten font image can be determined according to the stroke complexity of the Chinese characters in the target handwritten font image. Among them, the more complex the strokes of a Chinese character are, the more structures and style information it contains, and the greater the weight value of the target handwritten font image corresponding to it should be; on the contrary, the simpler the strokes of a Chinese character are, the smaller the weight value of the target handwritten font image corresponding to it should be. Taking Chinese characters as an example: The stroke complexity of the character "聚" is obviously greater than that of the character "中", so the weight value of the text style vector of the character "聚" should be greater than the weight value of the text style vector of the character "中". Taking English as an example again: The stroke complexity of the word "language" (which can be determined by the number of letters and the number of strokes of the letters in the word) is obviously greater than that of the word "at", so the weight value of the text style vector of the word "language" should be greater than the weight value of the text style vector of the word "at". Or, the magnitude of the weight value of any one of the text style vectors can also be positively correlated with the degree of being behind among the target handwritten font images corresponding to this text style vector. For example, the weight values of the text style vectors of each Chinese character are determined according to the writing order of multiple continuously written Chinese characters. Among them, the later the Chinese character is, the greater the weight value of its text style vector is, and the earlier the Chinese character is, the smaller the weight value of its text style vector is. Still taking Chinese characters as an example, for the characters "我", "幸", and "加" in the sentence "我们很荣幸地邀请您参加", the weight values of the text style vectors of the third character increase in turn.
[0136] In one embodiment, before reasoning, it is necessary to first determine a font image library for selecting the basic font style vectors of each of the multiple basic font styles. It may be assumed that all candidate font images in the font image library belong to M1 font styles, and each candidate font image belongs to one font style, m1≤M1, and m1 and M1 are both positive integers. The font image library can be pre-created in the manner described in the previous embodiment and will not be repeated here. Based on this, the font style vectors of each of the M1 font styles can be obtained (that is, M1 font style vectors corresponding one-to-one to the M1 font styles are obtained), and the basic font style vectors of each of the m1 basic font styles can be determined from the obtained M1 font style vectors. Among them, the style branch of the aforementioned font generation model can be used to extract the text style vectors of each candidate font image in the font image library, and the font style vectors of each of the M1 font styles can be calculated based on these text style vectors. For example, all candidate font images can be input into the style branch in sequence, and the text style vectors of each candidate font image output by the branch can be obtained; for the text style vectors of S1 candidate font images belonging to the same font style, the average value (such as the arithmetic mean) of these S1 text style vectors can be calculated, and the calculation result (the average value of multiple vectors is still a vector) can be used as the font style vector of the font style.
[0137] In addition, a variety of methods can be used to determine the basic font style vectors of each of the m1 basic font styles from the M1 font style vectors obtained. For example, the aforementioned PCA algorithm can be used to analyze the M1 font style vectors obtained, and the m1 main eigenvectors obtained by analysis can be determined as the basic font style vectors of each of the m1 basic font styles. For another example, a clustering algorithm can also be used to cluster the calculated M1 font style vectors, and the processed m1 cluster centers can be determined as the basic font style vectors of each of the m1 basic font styles (these m1 cluster centers are the m1 basic font style vectors, and the corresponding font styles are the corresponding basic font styles). The clustering algorithm can use Kmeans, Mini-Batch K-means, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), Expectation Maximization (EM), etc., which will not be repeated here. It is understandable that, because the m1 cluster centers are relatively evenly distributed in the M1 style feature vectors, the m1 basic font styles finally obtained can fully and comprehensively represent the M1 font styles as much as possible. Of course, m1 font styles can also be randomly selected from the M1 font styles as the basic font styles, which will not be repeated here. For example, the M1 font styles can be Songti, Kaiti, New Roman, Heiti, Lishu, Fangzheng, Xingkai, etc., and the m1 basic font styles selected in the above manner can be Kaiti, New Roman, Lishu, etc.
[0138] It can be understood that the process of determining m1 basic font style vectors from M1 font style vectors is similar to the process of determining m2 sample basic font style vectors from M2 sample font style vectors described in the aforementioned embodiment. In fact, the M1 can be equal to M2, and the m1 can also be equal to m2. Alternatively, instead of determining the m1 basic font style vectors from the M1 font style vectors in the above manner, the m2 sample basic font styles can be directly used as the basic font styles (in this case, m1=m2). Accordingly, the m2 sample basic font style vectors are respectively used as the basic font style vectors of the m2 basic font styles, which will not be repeated here.
[0139] In one embodiment, when determining the handwritten weight style vector according to the target text style vector, the handwritten weight values corresponding to the multiple basic font styles can be determined first according to the target text style vector and the basic font style vectors of the multiple basic font styles, and then the basic font style vectors of the multiple basic font styles can be weighted and summed according to the handwritten weight values to obtain the handwritten weight style vector. It can be understood that the magnitude of the handwritten weight value corresponding to the basic font style vector of any basic font style determines the proportion of the basic font style vector in the handwritten weight style vector, and also reflects the proportion of the basic font style in the target handwritten style of the target handwritten font image.
[0140] Among them, when determining the handwritten weight values corresponding to the multiple basic font styles according to the target text style vector and the basic font style vectors of the multiple basic font styles, the vector distances between the target text style vector and the basic font style vectors of the multiple basic font styles can be calculated, and the handwritten weight values corresponding to the multiple basic font styles can be calculated according to the respective vector distances. Among them, the magnitude of the handwritten weight value corresponding to any basic font style is negatively correlated with the magnitude of the vector distance corresponding to the font style.
[0141] Exemplarily, assume that the target handwritten font image is the handwritten character "科" shown in Figure 12, and its target text style is T 手写(科) , if the multiple basic font styles are regular script, running script and clerical script respectively, and the corresponding basic font style vectors are T0 楷体 , T0 行书 and T0 隶书 respectively, then the vector distances L7, L8 and L9 between T 手写(科) and T0 楷体 , T0 行书 and T0 隶书 can be calculated respectively. Furthermore, the handwritten weight values k7, k8 and k9 corresponding to regular script, running script and clerical script can be calculated according to the vector distances L7, L8 and L9, and specific calculations can be performed through relevant algorithms such as normalization and softmax functions, which will not be elaborated here. Further, the basic font style vectors can be weighted and summed according to the above handwritten weight values to obtain the handwritten weight style vector Tw 手写(科) = k7 * T0 楷体 + k8 * T0 行书 + k9 * T0 隶书 .
[0142] In another embodiment, when determining the handwritten weight style vector according to the target text style vector, the basic text style vectors of the basic font pictures that belong to multiple basic font styles and have the same text content as the target handwritten font picture can also be determined first, and then the handwritten weight values corresponding to each basic font picture can be determined according to the target text style vector and these basic text style vectors. Then, the basic text style vectors of each basic font picture are weighted and summed according to the handwritten weight values to obtain the handwritten weight style vector.
[0143] Exemplarily, assume that the target handwritten font picture is the handwritten character '学' shown in Figure 12, and its target text style is T 手写(学) , if the multiple basic font styles are regular script, running script, and official script respectively, the corresponding basic font pictures can be determined first, that is, the character '学' in regular script, the character '学' in running script, and the character '学' in official script. Then, these three font pictures are respectively input into the style branch to extract the basic text style vectors T0 楷体(学) , T0 行书(学) and T0 隶书(学) of them. Furthermore, the vector distances L10, L11, and L12 between T 手写(学) and T0 楷体(学) , T0 行书(学) and T0 隶书(学) can be calculated respectively. Furthermore, the handwritten weight values k10, k11, and k12 corresponding to each basic font picture can be calculated according to the vector distance (the specific algorithm is the same as above). Further, the basic text style vectors can be weighted and summed according to the handwritten weight values to obtain the handwritten weight style vector Tw 手写(学) = k10 * T0 楷体(学) + k11 * T0 行书(学) + k12 * T0 隶书(学) .
[0144] The handwriting weight style vector of the target handwritten font image can be determined by the aforementioned method. Similar to the aforementioned content branch, if the style branch in the font generation model is directly connected to the decoder (as shown in Figure 1), the handwriting weight style vector can be directly input into the decoder. If a content Transformer encoder is also connected between the style branch and the decoder in the font generation model (as shown in Figure 7), the handwriting weight style vector is input into the encoder for further processing, and the corresponding processing result will be input into the decoder. It should be noted that these two methods use the connection between the style branch and the decoder inside the font generation model to input the corresponding style vector into the decoder. In fact, given that the process of determining the handwriting weight style vector in the aforementioned embodiment may not require the style branch to temporarily participate in the calculation, a vector processing component can also be implemented by programming. The component is used to calculate the handwriting weight style vector and input it into the decoder in the manner described in the aforementioned embodiment - this method does not use the connection between the style branch and the decoder to achieve input, and will not be repeated.
[0145] Step 1106: Obtain a synthetic style font image generated by the decoder of the font generation model according to the target text content vector and the handwriting weight style vector.
[0146] After the aforementioned target text content vector and handwriting weight style vector (or the corresponding Transformer encoder's processing results of these two vectors) are respectively input into the decoder, the decoder can perform decoding processing based on the two vectors to generate a synthetic style font image and output it.
[0147] It can be seen from the embodiments of the aforementioned steps 1101-1106 that when the font generation model obtained by the above training process is used to generate a corresponding synthetic style font image for the user's target handwriting font image (i.e., the aforementioned reasoning), the target text style vector of the image extracted by the style branch is not directly input into the decoder, but the handwriting weight style vector obtained by weighted summation of the basic font style vectors of multiple basic font styles is first determined based on the target text style vector, and then the vector is input into the decoder, and then the decoder obtains the synthetic style font image generated by the handwriting weight style vector and the target text content vector extracted by the content branch. The actual style of the synthetic style font image generated in this way is obtained by fusion of multiple basic font styles and the target handwriting style of the target handwriting font image. It can be seen that the present invention introduces the concept of vector weight in the font generation method based on the font generation model, so that the actual style of the synthetic font style image finally generated can meet the user's diverse needs for font style, such as the user can adjust the style of the generated font by himself, etc., which significantly improves the flexibility of the font beautification effect.
[0148] In one embodiment, for the target handwritten font image, an expected weighted style vector of the desired font style can also be obtained, and a composite weighted style vector is determined based on the handwritten weighted style vector and the expected weighted style vector and input into the decoder. Thus, the decoder can generate a composite style font image based on the target text content vector and the composite weighted style vector. In this way, the font style of the resulting composite style font image can be made close to the desired font style, thereby achieving a synthesis (or fusion) of the desired font style with the actual style of the target handwritten font image.
[0149] Among them, the expected font style can be a default (or preset) font style. For example, the expected font style can be pre-set by the font generation model or the developer of the client and recorded in the local configuration file of the client. At this time, the default expected font style can be strongly related to the usage scenario of the synthetic style font image, that is, different expected font styles can be preset for different font generation scenarios. For example, when the user specifies that the synthetic style font image is used to generate an academic conference PPT, the expected font style can be defaulted to Microsoft Yahei; when the user specifies that the synthetic style font image is used to generate a wedding invitation, the expected font style can be defaulted to a national style font, etc., which will not be repeated. This method can improve the fit between the font style of the generated synthetic style font image and its usage scenario, thereby providing users with a richer and more diverse font generation experience. Alternatively, the expected font style can also be a user-specified font style. For example, the user can pre-specify a font style as the expected font style on the client's settings page before this reasoning; or during this reasoning process (before and after writing the target handwritten font image), specify a font style as the expected font style in the style execution page provided by the client. In this way, users are allowed to flexibly set the desired font style that they want to present in the synthesized style font image, thereby helping to meet the user's personalized style needs.
[0150] The expected weighted style vector can be obtained by weighted summing the basic font style vectors of each of the multiple basic font styles. For example, the style branch can be used to predetermine the basic font style vectors of each of the multiple basic font styles; and, for each candidate font style that may be defaulted or selected as the expected font style, the style branch can be used to predetermine the font style vectors of each of these candidate font styles, and then the weighted style vectors of each candidate font style are calculated based on the font style vector and each basic font style vector, and these weighted style vectors are stored. For example, for any candidate font style, the vector distance between the font style vector of the font style and the multiple basic font style vectors can be calculated, and the weight values corresponding to the multiple basic font styles are calculated based on each vector distance (wherein the size of the weight value corresponding to any basic font style is negatively correlated with the size of the vector distance corresponding to the font style), and then the multiple basic font style vectors are weighted summed according to the weight value to obtain the weighted style vector of the candidate font style. Based on this, when a candidate font style is selected as the expected font style, the pre-stored weighted style vector of the font style can be directly obtained as the corresponding expected weighted style vector. For another example, the basic font style vectors of each of the multiple basic font styles can be pre-calculated and stored, so that during the reasoning process, for the aforementioned expected font style selected by default or by the user, the font style vector of the style can be temporarily calculated using the style branch, and the expected weighted style vector of the expected font style can be further calculated using the aforementioned vector distance and weight value.
[0151] In addition, when determining the synthesized weight style vector based on the handwritten weight style vector and the expected weight style vector and inputting it into the decoder, the style proportion of the expected font style and the target handwritten style to which the target handwritten font picture belongs can be determined, and the handwritten weight style vector and the expected weight style vector are weighted and summed according to this style proportion to obtain the synthesized weight style vector. Among them, the sum of the respective ratios of the expected font style and the target handwritten style is 1. In addition, the expected font style may be one or more. In the case where there is only one expected weight style (such as regular script), the sum of the respective ratios of this expected font style and the target handwritten style is 1; while in the case where there are multiple expected weight styles (such as regular script and Song typeface), the sum of the respective ratios of each expected font style and the target handwritten style is 1. It can be understood that the higher the proportion of any style (that is, the larger the aforementioned ratio), the greater the proportion of this style in the actual style of the synthesized style font picture, which is intuitively reflected as the actual style of the synthesized style font picture is more similar to any style. Of course, the above style proportion can also be the default proportion or specified by the user according to their own意愿. In the case where the user is allowed to specify, the user can flexibly set the proportion of the expected font style and the target handwritten style in the actual style of the synthesized style font picture (that is, whether the style of the synthesized style font picture is closer to the expected font style or closer to the target handwritten style), so as to meet the diverse font style needs of the user and achieve a more optimized font aesthetic effect.
[0152] As shown in FIGS. 12 to 14, the user continuously wrote the four characters "science", "technology", "technique", and "art" in the handwriting area, and the client can recognize the font content of these four fonts. Further, if the user sets the ratios of the handwritten style (that is, the target handwritten style of the handwritten font picture) to regular script (that is, the default or user-specified expected font style) to be 0.99:0.01, 0.64:0.36, and 0.21:0.79 respectively, the display effect of the finally output synthesized style font picture can be seen in the synthesized area shown in FIGS. 12 to 14. By comparing FIGS. 12 to 14, it can be seen that the larger the proportion of the handwritten style, the closer the style of the synthesized style font picture is to the style of each handwritten font in the handwriting area (that is, the more similar it is to the handwritten style); the larger the proportion of regular script, the closer the style of the synthesized style font picture is to regular script (that is, the more similar it is to regular script). Of course, the user can also specify multiple expected font styles and their respective style proportions, which will not be elaborated here.
[0153] So far, the training method of the font generation model and the font generation method implemented using the font generation model have been introduced. Corresponding to the foregoing embodiments, the present invention also provides embodiments of a training device for the font generation model and a font generation device.
[0154] An embodiment of the present invention provides a device for training a font generation model, the device comprising one or more processors, wherein the processors are configured to:
[0155] Using a first training sample to train an original font generation model to obtain a first font generation model, the original font generation model includes a content branch, a style branch, and a decoder, wherein a first sample content font image and a first sample style font image in the first training sample are input into the content branch and the style branch, respectively, and the decoder is configured to generate and output a first synthesized style font image corresponding to the first training sample based on a first sample text content vector extracted from the content branch and a first sample text style vector extracted from the style branch;
[0156] The first font generation model is trained using a second training sample to obtain a second font generation model, wherein the second training sample includes a second sample content font image and a second sample style font image, the second sample content font image is input into the content branch, and the second sample text content vector extracted from the content branch is input into the decoder; the sample weight style vector of the second sample style font image is input into the decoder, and the sample weight style vector is obtained by weighted summation of the sample basic font style vectors of multiple sample basic font styles; the decoder is used to generate and output a second synthetic style font image corresponding to the second training sample based on the second sample text content vector and the sample weight style vector.
[0157] In one embodiment, the processor is specifically configured to:
[0158] Select any candidate font image from all candidate font images included in the font image library as the sample content font image and the sample style font image in any training sample; or
[0159] Selecting any candidate font image belonging to a first font style from all candidate font images included in the font image library as a sample style font image in any one of the training samples, and selecting any candidate font image from candidate font images belonging to a second font style as a sample content font image in any one of the training samples, wherein the first font style and the second font style are different font styles;
[0160] Wherein, all candidate font images in the font image library belong to multiple font styles, and each candidate font image belongs to one font style.
[0161] In one embodiment, the processor is specifically configured to:
[0162] Inputting the second sample style font image into the style branch to obtain a sample text style vector of the second sample style font image extracted by the style branch, and obtaining sample basic font style vectors of each of the multiple sample basic font styles; and calculating sample weight values corresponding to the multiple sample basic font styles respectively based on vector distances between the sample text style vector and each sample basic font style vector; or,
[0163] Obtaining a sample font style vector of the sample font style to which the second sample style font image belongs, and obtaining sample basic font style vectors corresponding to the plurality of sample basic font styles respectively; and calculating sample weight values corresponding to the plurality of sample basic font styles respectively based on vector distances between the font style vector and each sample basic font style vector;
[0164] The sample basic font style vectors of the plurality of sample basic font styles are weighted and summed according to the sample weight value to obtain the sample weight style vector, and the sample weight style vector is input into the decoder.
[0165] In one embodiment, the processor is specifically configured to:
[0166] Multiple candidate font images belonging to any one of the font styles are respectively input into the style branch of the first font generation model to obtain text style vectors of the multiple candidate font images extracted by the style branch, and the average value of the obtained multiple text style vectors is used as the font style vector of any one of the font styles.
[0167] In one embodiment, the number of types of the sample basic font styles is m2, and the processor is specifically configured to:
[0168] Determine a font image library, wherein all candidate font images in the font image library belong to M2 font styles, each candidate font image belongs to one font style, and the second sample style font image is selected from the font image library, m2≤M2, and both m2 and M2 are positive integers;
[0169] The font style vectors of the M2 font styles are obtained, and the sample basic font style vectors of the m2 sample basic font styles are determined from the obtained M2 font style vectors.
[0170] In one embodiment, the processor is specifically configured to:
[0171] Analyze the M2 obtained font style vectors using the PCA algorithm, and determine the m2 main feature vectors obtained by the analysis as the sample basic font style vectors of the m2 sample basic font styles; or,
[0172] The obtained M2 font style vectors are clustered using a clustering algorithm, and the m2 cluster centers obtained by the processing are determined as the sample basic font style vectors of the respective m2 sample basic font styles.
[0173] In one embodiment, the loss function of the first font generation model includes a style loss function, and the processor is specifically configured to:
[0174] A first sample text style vector of the first sample style font image and a first synthetic text style vector of the first synthetic style font image are obtained, and a loss value of the style loss function is determined based on a deviation between the first sample text style vector and the first synthetic text style vector.
[0175] In one embodiment, the processor is specifically configured to:
[0176] Inputting the first sample style font image into a feature extraction network comprising multiple network layers, and obtaining a feature vector extracted by at least one network layer as the first sample text style vector; and
[0177] The first synthetic style font image is input into the feature extraction network, and the feature vector extracted by the at least one network layer is obtained as the first synthetic text style vector.
[0178] In one embodiment, the processor is specifically configured to:
[0179] Performing multiple rounds of training on the original font generation model using multiple first training samples;
[0180] When the number of training rounds of the original font generation model reaches a first round number threshold, a first sample text style vector of the first sample style font image and a first synthetic text style vector of the first synthetic style font image are obtained.
[0181] In one embodiment, the loss function of the first font generation model includes a connectivity loss function, and the processor is specifically configured to:
[0182] Determining the sample position coordinates of each sample boundary point on the font outline of the first sample style font image and the composite position coordinates of each composite boundary point on the font outline of the first composite style font image;
[0183] The Hausdorff distance between the sample boundary point and the synthetic boundary point is calculated using the sample position coordinates and the synthetic position coordinates to serve as a loss value of the connectivity loss function.
[0184] In one embodiment, there are multiple first training samples, and the multiple first training samples include at least one standard training sample and at least one handwritten training sample. The first sample style font image in the standard training sample is generated by a standard style TTF file, and the first sample style font image in the handwritten training sample is generated by handwriting. The processor is specifically configured to:
[0185] When the number of training rounds of the original font generation model does not reach the second round number threshold, training the original font generation model using the standard training sample;
[0186] When the number of training rounds of the original font generation model reaches a second round number threshold, the original font generation model is trained using at least the handwritten training sample.
[0187] In one embodiment, the original font generation model is a DG-Font network.
[0188] In one embodiment, the style branch includes a convolutional neural network (CNN) and a connected feature pyramid network (FPN), wherein:
[0189] There is a one-to-one lateral connection between the S network layers in the CNN and the S network layers in the FPN. For any font image input into the style branch, the S-dimensional feature vectors output by the S network layers in the FPN are used as the text style vector of the any font image output by the style branch, where S is an integer greater than 1.
[0190] In one embodiment, the original font generation model further includes a content Transformer encoder and / or a style Transformer encoder, wherein:
[0191] The content Transformer encoder is used to connect the output end of the content branch and the input end of the encoder, and the style Transformer encoder is used to connect the output end of the style branch and the input end of the encoder.
[0192] An embodiment of the present invention further provides a font generation device, the device comprising one or more processors, the processors being configured to:
[0193] Obtain the target text content vector of the target handwritten font image based on the font generation model;
[0194] Inputting the target handwritten font image into the style branch of the font generation model to extract a target text style vector, and determining a handwriting weighted style vector based on the target text style vector, wherein the handwriting weighted style vector is obtained by weighted summation of basic font style vectors of multiple basic font styles;
[0195] Obtain a synthetic style font image generated by the decoder of the font generation model according to the target text content vector and the handwriting weight style vector.
[0196] In one embodiment, the processor is specifically configured to:
[0197] Input the target handwritten font image into the content branch of the font generation model to extract the target text content vector; or,
[0198] A target reference font image of a target handwritten font image is obtained, and the target reference font image is input into a content branch of a font generation model to extract a target text content vector, wherein the target reference font image belongs to a preset reference font style and has the same text content as the target handwritten font image.
[0199] In one embodiment, there are multiple target handwritten font images, and the processor is specifically configured to:
[0200] Input each target handwritten font image into the style branch to obtain the text style vector of each target handwritten font image extracted by the style branch;
[0201] An average value of the obtained text style vectors is calculated as the target text style vector.
[0202] In one embodiment, the processor is specifically configured to:
[0203] Calculate the weighted average of each text style vector obtained; where,
[0204] The weight value of any text style vector is a preset value; or,
[0205] The weight value of any text style vector is positively correlated with the stroke complexity of the target handwritten font image corresponding to the text style vector; or,
[0206] The weight value of any character style vector is positively correlated with the degree to which the target handwriting font image corresponding to the character style vector is positioned at the back among all target handwriting font images.
[0207] In one embodiment, the number of basic font styles is m1, and the processor is specifically configured to:
[0208] Determine a font image library, wherein all candidate font images in the font image library belong to M1 font styles, each candidate font image belongs to one font style, m1≤M1, and m1 and M1 are both positive integers;
[0209] The font style vectors of the M1 font styles are obtained, and the basic font style vectors of the m1 basic font styles are determined from the obtained M1 font style vectors.
[0210] In one embodiment, the processor is specifically configured to:
[0211] Analyze the obtained M1 font style vectors using the principal component analysis (PCA) algorithm, and determine the obtained m1 main feature vectors as the basic font style vectors of the m1 basic font styles; or,
[0212] The obtained M1 font style vectors are clustered using a clustering algorithm, and the obtained m1 cluster centers are determined as the basic font style vectors of the respective m1 basic font styles.
[0213] In one embodiment, the processor is specifically configured to:
[0214] Determining handwriting weight values corresponding to the multiple basic font styles respectively according to the target text style vector and the respective basic font style vectors of the multiple basic font styles;
[0215] A handwriting weight style vector is obtained by weighted summing of the basic font style vectors of the multiple basic font styles according to the handwriting weight value.
[0216] In one embodiment, the processor is specifically configured to:
[0217] Calculate the vector distances between the target text style vector and the basic font style vectors of multiple basic font styles, and calculate the handwriting weight values corresponding to the multiple basic font styles based on each vector distance, wherein the size of the handwriting weight value corresponding to any basic font style is negatively correlated with the size of the vector distance corresponding to the font style.
[0218] In one embodiment, the processor is specifically configured to: obtain an expected weight style vector of an expected font style, and determine a composite weight style vector according to the handwriting weight style vector and the expected weight style vector and input the composite weight style vector into the decoder;
[0219] The processor is further configured to obtain a synthetic style font image generated by a decoder of the font generation model according to the target text content vector and the synthetic weight style vector.
[0220] In one embodiment, the desired font style is a default font style or a font style specified by the user.
[0221] In one embodiment, the expected weighted style vector is obtained by weighted summing of the basic font style vectors of the multiple basic font styles.
[0222] In one embodiment, the processor is specifically configured to:
[0223] The style proportions of the expected font style and the target handwriting style to which the target handwriting font image belongs are determined, and a composite weighted style vector is obtained by weighted summing the handwriting weighted style vector and the expected weighted style vector according to the style proportions.
[0224] In one embodiment, the font generation model is obtained based on DG-Font network training.
[0225] In one embodiment, the style branch includes a convolutional neural network (CNN) and a connected feature pyramid network (FPN), wherein:
[0226] There is a one-to-one lateral connection between the S network layers in the CNN and the S network layers in the FPN. For any font image input into the style branch, the S-dimensional feature vectors output by the S network layers in the FPN are used as the text style vector of the font image output by the style branch, where S is an integer greater than 1.
[0227] In one embodiment, the font generation model includes a content Transformer encoder and / or a style Transformer encoder, wherein:
[0228] The content Transformer encoder is used to connect the output end of the content branch and the input end of the encoder, and the style Transformer encoder is used to connect the output end of the style branch and the input end of the encoder.
[0229] An embodiment of the present invention further proposes an electronic device, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the font generation model training method or font generation method described in any of the above embodiments.
[0230] An embodiment of the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the training method of the font generation model or the steps in the font generation method described in any of the above embodiments are implemented.
[0231] Regarding the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments of the relevant methods and will not be elaborated on here.
[0232] Figure 15 is a schematic block diagram of an apparatus 1500 for data storage or driving mode determination according to an embodiment of the present invention. For example, apparatus 1500 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0233] 15 , device 1500 may include one or more of the following components: a processing component 1502 , a memory 1504 , a power component 1506 , a multimedia component 1508 , an audio component 1510 , an input / output (I / O) interface 1512 , a sensor component 1514 , and a communication component 1516 .
[0234] Processing component 1502 generally controls the overall operation of device 1500, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. Processing component 1502 may include one or more processors 1520 to execute instructions to perform all or part of the steps of the above-described methods. In addition, processing component 1502 may include one or more modules to facilitate interaction between processing component 1502 and other components. For example, processing component 1502 may include a multimedia module to facilitate interaction between multimedia component 1508 and processing component 1502.
[0235] The memory 1504 is configured to store various types of data to support the operation of the device 1500. Examples of such data include instructions for any application or method operating on the device 1500, contact data, phone book data, messages, pictures, videos, etc. The memory 1504 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0236] The power supply component 1506 provides power to the various components of the device 1500. The power supply component 1506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device 1500.
[0237] The multimedia component 1508 includes a screen that provides an output interface between the device 1500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 1508 includes a front camera and / or a rear camera. When the device 1500 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0238] The audio component 1510 is configured to output and / or input audio signals. For example, the audio component 1510 includes a microphone (MIC) that is configured to receive external audio signals when the device 1500 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals may be further stored in the memory 1504 or transmitted via the communication component 1516. In some embodiments, the audio component 1510 further includes a speaker for outputting audio signals.
[0239] I / O interface 1512 provides an interface between processing component 1502 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include, but are not limited to, a home button, volume buttons, a start button, and a lock button.
[0240] Sensor assembly 1514 includes one or more sensors for providing various aspects of the status assessment of device 1500. For example, sensor assembly 1514 can detect the open / closed state of device 1500, the relative positioning of components, such as the display and keypad of device 1500. Sensor assembly 1514 can also detect changes in the position of device 1500 or a component of device 1500, the presence or absence of user contact with device 1500, the orientation or acceleration / deceleration of device 1500, and changes in the temperature of device 1500. Sensor assembly 1514 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 1514 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 1514 can also include an accelerometer, a gyroscope, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0241] The communication component 1516 is configured to facilitate wired or wireless communication between the device 1500 and other devices. The device 1500 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, 4G LTE, 6G NR or a combination thereof. In an exemplary embodiment, the communication component 1516 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1516 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0242] In an exemplary embodiment, the apparatus 1500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described methods.
[0243] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1504 including instructions, which can be executed by the processor 1520 of the apparatus 1500 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0244] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the embodiments disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow from the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered as exemplary only, with the true scope and spirit of the invention being indicated by the following claims.
[0245] It should be understood that the present invention is not limited to the exact construction described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
[0246] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a..." does not exclude the presence of other identical elements in the process, method, article or device that includes the element.
[0247] The above is a detailed introduction to the methods and devices provided in the embodiments of the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the methods and core ideas of the present invention. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the content of the present invention should not be understood as limiting the present invention.
Claims
1. A font generation method, comprising: Obtaining a target text content vector of a target handwritten font picture according to a font generation model; Inputting the target handwritten font picture into a style branch of the font generation model to extract a target text style vector, and determining a handwritten weight style vector according to the target text style vector, where the handwritten weight style vector is obtained by weighted summation of basic font style vectors of various basic font styles; Obtaining a synthesized style font picture generated by a decoder of the font generation model according to the target text content vector and the handwritten weight style vector.
2. The method according to claim 1, wherein obtaining a target text content vector of a target handwritten font picture according to a font generation model comprises: Inputting the target handwritten font picture into a content branch of the font generation model to extract a target text content vector; Or, Obtaining a target reference font picture of the target handwritten font picture, and inputting the target reference font picture into a content branch of the font generation model to extract a target text content vector, where the target reference font picture belongs to a preset reference font style and has the same text content as the target handwritten font picture.
3. The method according to claim 1, wherein the number of the target handwritten font pictures is multiple, and inputting the target handwritten font pictures into the style branch to extract a target text style vector comprises: Inputting each of the target handwritten font pictures into the style branch respectively to obtain text style vectors of each of the target handwritten font pictures extracted by the style branch; Calculating an average value of the obtained text style vectors as the target text style vector.
4. The method according to claim 3, wherein calculating the average value of the obtained text style vectors includes: Calculating a weighted average value of the obtained text style vectors; wherein: The weight value of any text style vector is a preset value; or, The magnitude of the weight value of any text style vector is positively correlated with the stroke complexity of the target handwritten font picture corresponding to the text style vector; or, The magnitude of the weight value of any text style vector is positively correlated with the degree of the target handwritten font picture corresponding to the text style vector being behind among all the target handwritten font pictures.
5. The method according to claim 1, wherein the number of types of the basic font styles is m1, and determining basic font style vectors of the various basic font styles comprises: Determining a font picture library, where all candidate font pictures in the font picture library belong to M1 font styles in total, each candidate font picture belongs to one font style, m1 ≤ M1, and both m1 and M1 are positive integers; Obtaining font style vectors of the M1 font styles respectively, and determining basic font style vectors of the m1 basic font styles from the obtained M1 font style vectors.
6. The method according to claim 5, wherein determining basic font style vectors of the m1 basic font styles from the obtained M1 font style vectors comprises: Analyzing the obtained M1 font style vectors by using a principal component analysis (PCA) algorithm, and determining the obtained m1 principal feature vectors as basic font style vectors of the m1 basic font styles respectively; or, Cluster the obtained M1 font style vectors using a clustering algorithm, and determine the m1 cluster centers obtained from the processing as the respective basic font style vectors of the m1 basic font styles.
7. The method according to claim 1, wherein determining the handwritten weight style vector according to the target text style vector comprises: Determining the respective handwritten weight values corresponding to the multiple basic font styles according to the target text style vector and the respective basic font style vectors of the multiple basic font styles; Performing weighted summation on the respective basic font style vectors of the multiple basic font styles according to the handwritten weight values to obtain a handwritten weight style vector.
8. The method according to claim 7, wherein determining the respective handwritten weight values corresponding to the multiple basic font styles according to the target text style vector and the respective basic font style vectors of the multiple basic font styles comprises: Calculating the vector distances between the target text style vector and the respective basic font style vectors of the multiple basic font styles, and calculating the respective handwritten weight values corresponding to the multiple basic font styles according to each vector distance, wherein the magnitude of the handwritten weight value corresponding to any basic font style is negatively correlated with the magnitude of the vector distance corresponding to this font style.
9. The method according to claim 1, Also included are: Obtaining an expected weight style vector of an expected font style, and determining a synthesized weight style vector according to the handwritten weight style vector and the expected weight style vector and inputting the synthesized weight style vector into the decoder; The obtaining of the synthesized style font image generated by the decoder of the font generation model according to the target text content vector and the handwritten weight style vector comprises: obtaining the synthesized style font image generated by the decoder of the font generation model according to the target text content vector and the synthesized weight style vector.
10. The method according to claim 9, The expected font style is a default font style or a font style specified by a user.
11. The method according to claim 9, wherein the expected weight style vector is obtained by performing weighted summation on the respective basic font style vectors of the multiple basic font styles.
12. The method according to claim 9, wherein determining the synthesized weight style vector according to the handwritten weight style vector and the expected weight style vector and inputting the synthesized weight style vector into the decoder comprises: Determining the style proportion of the expected font style and the target handwritten style to which the target handwritten font image belongs, and performing weighted summation on the handwritten weight style vector and the expected weight style vector according to the style proportion to obtain a synthesized weight style vector.
13. The method according to any one of claims 1-12, wherein the font generation model is trained based on the DG-Font network.
14. The method according to any one of claims 1-12, wherein the style branch comprises a convolutional neural network CNN and a feature pyramid network FPN connected thereto, wherein, There is a one-to-one lateral connection between the S network layers in the CNN and the S network layers in the FPN. For any font image input to the style branch, the S-dimensional feature vectors output by the S network layers in the FPN are used as the text style vector of the font image output by the style branch, where S is an integer greater than 1.
15. The method according to any one of claims 1-12, wherein the font generation model includes a content Transformer encoder and / or a style Transformer encoder, wherein, The content Transformer encoder is used to connect the output end of the content branch and the input end of the encoder, and the style Transformer encoder is used to connect the output end of the style branch and the input end of the encoder.
16. A training method for a font generation model, comprising: Training an original font generation model using a first training sample to obtain a first font generation model. The original font generation model includes a content branch, a style branch, and a decoder. The first sample content font image and the first sample style font image in the first training sample are respectively input to the content branch and the style branch. The decoder is used to generate and output a first synthesized style font image corresponding to the first training sample according to the first sample text content vector extracted by the content branch and the first sample text style vector extracted by the style branch; Training the first font generation model using a second training sample to obtain a second font generation model. The second training sample includes a second sample content font image and a second sample style font image. The second sample content font image is input to the content branch, and the second sample text content vector extracted by the content branch is input to the decoder; The sample weight style vector of the second sample style font image is input to the decoder. The sample weight style vector is obtained by weighted summation of the sample base font style vectors of various sample base font styles; The decoder is used to generate and output a second synthesized style font image corresponding to the second training sample according to the second sample text content vector and the sample weight style vector.
17. The method according to claim 16, for any one of the first training sample and the second training sample, obtaining the sample content font image and the sample style font image in the any one of the training samples, including: Selecting any candidate font image from all candidate font images included in the font image library as the sample content font image and the sample style font image in the any one of the training samples; or, Selecting any candidate font image belonging to a first font style from all candidate font images included in the font image library as the sample style font image in the any one of the training samples, and selecting any candidate font image belonging to a second font style from the candidate font images belonging to the second font style as the sample content font image in the any one of the training samples, where the first font style and the second font style are different font styles; Among them, all candidate font images in the font image library belong to multiple font styles, and each candidate font image belongs to one font style.
18. The method according to claim 16, wherein training the first font generation model with the second training sample to obtain a second font generation model includes: Inputting the second sample style font image into the style branch to obtain a sample text style vector of the second sample style font image extracted by the style branch, and obtaining sample base font style vectors of the various sample base font styles; and calculating sample weight values corresponding to the various sample base font styles according to vector distances between the sample text style vector and each of the sample base font style vectors respectively; or, Obtaining a sample font style vector of the sample font style to which the second sample style font image belongs, and obtaining sample base font style vectors corresponding to the various sample base font styles respectively; and calculating sample weight values corresponding to the various sample base font styles according to vector distances between the font style vector and each of the sample base font style vectors respectively; Performing weighted summation on the sample base font style vectors of the various sample base font styles according to the sample weight values to obtain a sample weighted style vector, and inputting the sample weighted style vector into the decoder.
19. The method according to claim 18, for any one of the sample font style and the various sample base font styles, determining a font style vector of the any one font style includes: Inputting multiple candidate font images belonging to the any one font style into the style branch of the first font generation model respectively to obtain text style vectors of the multiple candidate font images extracted by the style branch, and taking an average value of the obtained multiple text style vectors as the font style vector of the any one font style.
20. The method according to claim 16, wherein the number of types of the sample base font styles is m2, and determining sample base font style vectors of the various sample base font styles respectively includes: Determining a font image library, all candidate font images in the font image library belong to M2 types of font styles in total, each candidate font image belongs to one font style, the second sample style font image is selected from the font image library, m2 ≤ M2, and both m2 and M2 are positive integers; Obtaining font style vectors of the M2 types of font styles respectively, and determining sample base font style vectors of m2 sample base font styles respectively from the obtained M2 font style vectors.
21. The method according to claim 20, wherein determining sample base font style vectors of m2 sample base font styles respectively from the obtained M2 font style vectors includes: Analyzing the obtained M2 font style vectors by using a PCA algorithm, and determining m2 principal eigenvectors obtained by the analysis as sample base font style vectors of the m2 sample base font styles respectively; or, Cluster the obtained M2 font style vectors using a clustering algorithm, and determine the m2 cluster centers obtained from the processing as the sample base font style vectors of the respective m2 sample base font styles.
22. The method according to claim 16, wherein the loss function of the first font generation model includes a style loss function, and the method further includes: Obtain a first sample text style vector of the first sample style font image and a first synthesized text style vector of the first synthesized style font image, and determine a loss value of the style loss function based on a deviation between the first sample text style vector and the first synthesized text style vector.
23. The method according to claim 22, wherein the obtaining the first sample text style vector of the first sample style font image and the first synthesized text style vector of the first synthesized style font image includes: Input the first sample style font image into a feature extraction network including multiple network layers, and obtain at least one feature vector extracted by the network layer as the first sample text style vector; and Input the first synthesized style font image into the feature extraction network, and obtain at least one feature vector extracted by the network layer as the first synthesized text style vector.
24. The method according to claim 22, Training the original font generation model with the first training sample to obtain the first font generation model includes: Use a plurality of first training samples to perform multiple rounds of training on the original font generation model; The obtaining the first sample text style vector of the first sample style font image and the first synthesized text style vector of the first synthesized style font image includes: when the number of training rounds of the original font generation model reaches a first round threshold, obtain the first sample text style vector of the first sample style font image and the first synthesized text style vector of the first synthesized style font image.
25. The method according to claim 16, wherein the loss function of the first font generation model includes a connectivity loss function, and the method further includes: Determine sample position coordinates of each sample boundary point on the font outline of the first sample style font image and synthetic position coordinates of each synthetic boundary point on the font outline of the first synthesized style font image; Calculate the Hausdorff distance between the sample boundary point and the synthetic boundary point using the sample position coordinates and the synthetic position coordinates as a loss value of the connectivity loss function.
26. The method according to claim 16, wherein the number of the first training samples is multiple, and the multiple first training samples include at least one standard training sample and at least one handwritten training sample. The first sample style font image in the standard training sample is generated by a ttf file of a standard style, and the first sample style font image in the handwritten training sample is generated by handwriting. The training of the original font generation model using the first training sample includes: When the number of training rounds of the original font generation model has not reached a second round threshold, use the standard training sample to train the original font generation model; When the number of training rounds of the original font generation model reaches the second round threshold, at least use the handwritten training samples to train the original font generation model.
27. The method according to any one of claims 16-26, wherein the original font generation model is a DG-Font network.
28. The method according to any one of claims 16-24, wherein the style branch includes a convolutional neural network CNN and a feature pyramid network FPN connected thereto, wherein there is a one-to-one lateral connection between S network layers in the CNN and S network layers in the FPN. For any font image input to the style branch, the S-dimensional feature vectors output by the S network layers in the FPN are used as the text style vectors of the any font image output by the style branch, and S is an integer greater than 1.
29. The method according to any one of claims 16-24, wherein the original font generation model further includes a content Transformer encoder and / or a style Transformer encoder, wherein: the content Transformer encoder is used to connect the output end of the content branch and the input end of the encoder, and the style Transformer encoder is used to connect the output end of the style branch and the input end of the encoder.
30. A font generation model, which is trained by the method according to any one of claims 16-29.
31. A font generation device, the device includes one or more processors, and the processors are configured to: Obtain the target text content vector of the target handwritten font image according to the font generation model; Input the target handwritten font image into the style branch of the font generation model to extract the target text style vector, and determine the handwritten weight style vector according to the target text style vector. The handwritten weight style vector is obtained by weighted summation of the respective basic font style vectors of multiple basic font styles; Obtain the synthetic style font image generated by the decoder of the font generation model according to the target text content vector and the handwritten weight style vector.
32. A training device for a font generation model, the device includes one or more processors, and the processors are configured to: Train the original font generation model using the first training sample to obtain the first font generation model. The original font generation model includes a content branch, a style branch and a decoder. The first sample content font image and the first sample style font image in the first training sample are respectively input into the content branch and the style branch. The decoder is used to generate and output the first synthetic style font image corresponding to the first training sample according to the first sample text content vector extracted by the content branch and the first sample text style vector extracted by the style branch; Training the first font generation model using a second training sample to obtain a second font generation model, where the second training sample includes second sample content font images and second sample style font images, the second sample content font images are input into the content branch, and the second sample text content vectors extracted by the content branch are input into the decoder; The sample weight style vectors of the second sample style font images are input into the decoder, and the sample weight style vectors are obtained by weighted summation of the sample base font style vectors of various sample base font styles; The decoder is configured to generate and output second synthesized style font images corresponding to the second training sample according to the second sample text content vectors and the sample weight style vectors.
33. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to implement the font generation method described in any one of claims 1-15 or the training method of the font generation model described in any one of claims 16-29.
34. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the font generation method described in any one of claims 1-15 or the training method of the font generation model described in any one of claims 16-29.