An artistic word generation system for complex texture structures

By designing an art word generation system for complex texture structures, using the generative adversarial network model and detail refinement module, the problem that existing systems cannot generate complex style art words is solved, and the generation of art words with complex style effects is achieved, meeting the diverse publicity and advertising needs.

CN114943783BActive Publication Date: 2025-06-13NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210651537.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-09
Publication Date
2025-06-13
Estimated Expiration
2042-06-09

AI Technical Summary

Technical Problem

The existing art word generation system can only generate simple art word based on simple styles, and cannot generate art word based on complex texture structures with complex style effects.

Method used

A word art generation system for complex texture structures is designed, including input processing module, generative adversarial network model and detail refinement module. The input processing module generates black and white text masks and style blocks, generates an adversarial network model to process these inputs through the first and second generators, and the detail refinement module further processes through the structural refinement network and the texture refinement network to generate the final complex style art words.

Benefits of technology

It realizes the generation of complex style effects based on complex texture structures, breaking through the limitation that the existing system can only generate simple style art characters, and meeting the needs of diversified publicity and advertising.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114943783B_ABST
    Figure CN114943783B_ABST
Patent Text Reader

Abstract

The present application provides an artistic word generation system for complex texture structures, including an input processing module that processes the input source text to generate a black-and-white text mask, and uses the black-and-white text mask to process the input style picture to generate style patches; the first generator of the generative adversarial network model processes the black-and-white text mask and the style patches to generate style blocks with real edges expanded by a preset multiple; the second generator of the generative adversarial network model processes the style blocks to generate a black-and-white style mask of the style blocks; the detail refinement module includes a structure refinement network and a texture refinement network. The structure refinement network structurally refines the style blocks to generate intermediate artistic words; the texture refinement network texture-refines the intermediate artistic words according to the black-and-white style mask to generate final artistic words. In this way, by generating a prototype of the artistic word and then refining the structure and details of the artistic word prototype, artistic words with complex style effects based on complex texture structures are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of art word generation, and particularly to an art word generation system for complex texture structures. Background Art

[0002] With the development of visual art, the demand for using art words in various promotional advertisements is increasing. Therefore, it is necessary to quickly generate art words in corresponding styles according to different requirements. Currently, art words can be generated through neural network models.

[0003] For art word generation, the commonly used neural network model is the Generative Adversarial Network (GAN) model. GAN is an unsupervised deep learning model. There are usually two methods for generating art words using GAN: One method is to generate the target art word corresponding to the target text based on the source text and the corresponding source art word, and the style of the target art word is the same as that of the source art word. This method requires a special art word style dataset for GAN training, and after training, it can only generate art words based on the existing art word styles, which cannot meet the diverse usage requirements of various promotional advertisements. If complex art words are to be generated based on this method, a large-scale art font dataset with high-resolution images and diversification needs to be developed, and the cost is too high; Another method is to set simple stylized text, animate the static text image by referring to a style picture or a style video, and generate art words with controllable glyphs corresponding to the static text image by perceiving the shapes in the style video. Although this method breaks through the limitations of the existing art word styles to a certain extent, the control of the glyphs still remains at the level of simple style control without obvious structural features. If complex art words are to be generated based on this method, complex structures and textures need to be referred to, but this method is prone to distorting these complex art features.

[0004] In summary, the existing art word generation systems can only generate simple art words based on simple styles and cannot generate art words with complex style effects based on complex texture structures. Summary of the Invention

[0005] This application provides an art word generation system for complex texture structures, which can be used to solve the technical problem that the existing art word generation systems can only generate simple art words based on simple styles and cannot generate art words with complex style effects based on complex texture structures.

[0006] To solve the above technical problem, this application discloses the following technical solutions:

[0007] An art word generation system for complex texture structures, the generation system comprising: an input processing module, a generative adversarial network model, and a detail refinement module connected in sequence;

[0008] The input processing module is configured to process the input source text to generate a black-and-white text mask with smoothed edges, and process the input style image using the black-and-white text mask to generate style patches, where the style image is an image with complex textures;

[0009] The generative adversarial network model includes a first generator and a second generator. The first generator is configured to process the black-and-white text mask and the style patches to generate style blocks with real edges expanded by a preset multiple; the second generator is configured to process the style blocks to generate a black-and-white style mask of the style blocks;

[0010] The generative adversarial network model is a Pro-gen GAN network model trained through generative training. The generative training uses preset black-and-white mask patches and preset cropped style patches in a pre-created generative training set as inputs, and preset style blocks and preset black-and-white style mask blocks in the generative training set as outputs, and trains the generative adversarial network model until convergence;

[0011] The detail refinement module includes a structure refinement network and a texture refinement network. The structure refinement network is configured to perform structure refinement processing on the style blocks to generate intermediate art words; the texture refinement network is configured to perform texture refinement processing on the intermediate art words according to the black-and-white style mask to generate final art words;

[0012] The structure refinement network is a Structure Net network trained through style training. The style training uses images in a preset conventional image dataset as input content images, a preset source style image as a reference style image, and a stylized content image as an output, and trains the structure refinement network until convergence.

[0013] In an implementable manner, the generative training set is pre-created through the following method:

[0014] Select the original style image Y g , where the original style image Y g is an image with complex textures;

[0015] Obtain the original black-and-white mask M g of the original style image Y g , where the style part in the original style image Y g is a black area in the original black-and-white mask M g , and the original style image Y gThe background part in is the white area in the original black-and-white mask M g is the white area;

[0016] Select the local black-and-white mask M with the largest black area in the original black-and-white mask M according to the preset first size L×L g and the original style image Y l and the local style image Y corresponding to the local black-and-white mask M in g ; l ; l ;

[0017] Perform edge simplification processing on the original black-and-white mask M g to generate a first black-and-white mask with smooth edges

[0018] Perform edge simplification processing on the local black-and-white mask M l to generate a second black-and-white mask with smooth edges

[0019] According to the preset large-piece cutting method, cut multiple preset style large pieces y from the original style image Y g and the local style image Y l and obtain the corresponding preset black-and-white mask large pieces of each preset style large piece y in the original black-and-white mask M g and the local black-and-white mask M l and obtain the corresponding preset black-and-white mask large pieces m of each preset black-and-white mask large piece m in the first black-and-white mask and the second black-and-white mask and obtain the corresponding preset black-and-white mask large pieces m of each preset black-and-white mask large piece m in the second black-and-white mask s ;

[0020] Randomly cut preset style small pieces from each preset style large piece y The size of the preset style large piece y is a preset multiple of the size of the preset style small piece ;

[0021] Downsample each preset black-and-white mask large piece m s according to the preset multiple to obtain a preset black-and-white mask small piece The size of the black-and-white mask large piece m s is a preset multiple of the size of the preset black-and-white mask small piece ;

[0022] Cut the preset style small piece through the preset black-and-white mask small piece to obtain a preset cut style small piece ;

[0023] All preset cropping style patches and all preset black and white mask patches are determined as the generated training set.

[0024] In one implementable manner, the preset large patch cropping method includes:

[0025] Set the second size of the large patch to xN×xN, where xN < L, and x is the preset multiple;

[0026] Crop multiple large patches from the first reference image according to the first probability a. The first reference image includes the original style image Y g , the original black and white mask M g and the first black and white mask

[0027] Crop multiple large patches from the second reference image according to the second probability 1 - a. The second reference image includes the local style image Y l , the local black and white mask M l and the second black and white mask

[0028] In one implementable manner, the size of the preset small patch is N×N. The preset small patch includes the preset style patch the preset black and white mask patch and the preset cropping style patch

[0029] In one implementable manner, the first generator includes a convolutional layer, a residual module, a splicing layer, and a transposed convolutional layer set according to the generation requirements; the second generator includes a convolutional layer, a residual module, and a splicing layer set according to the generation requirements.

[0030] In one implementable manner, the generation training includes first generation training and second generation training, where:

[0031] In the first generation training, by setting the stride and dilation rate of the convolutional layer inside the first generator, the stride and dilation rate of the transposed convolutional layer, and the size of the convolutional kernel, the preset black and white mask patch and the preset cropping style patch are used as inputs, and the preset style large patch y is used as the output to process the black and white text mask and the style patch, and generate training for the style large patch with real edges expanded by the preset multiple;

[0032] The second generation training processes the style patch by setting the stride and dilation rate of the convolutional layers inside the second generator, taking the preset style patch y as the input and the preset black-and-white mask patch m as the output, to generate the black-and-white style mask of the style patch for training.

[0033] In an implementable manner, the generative adversarial network model further includes a discriminator, which is used to cooperate with the first generator to complete the first generation training.

[0034] In an implementable manner, the generation system further includes a deformable module disposed between the input processing module and the generative adversarial network model. The deformable module is used to control the deformation degree of the black-and-white text mask by adding noise and edge erosion to the black-and-white text mask.

[0035] In an implementable manner, the deformable module controls the deformation degree of the black-and-white text mask by adding noise and edge erosion to the black-and-white text mask through the following steps:

[0036] Erode the edge of the black-and-white text mask and add noise to the eroded edge to obtain a noisy black-and-white mask;

[0037] Set a vector f to dilate the noisy black-and-white mask to control the deformation degree of the black-and-white text mask; the vector f includes f 0 、f 1 and f 2 , where f 0 is the size of the core of erosion and dilation, f 1 is the degree of noise addition in the edge, and f 2 is to control the degree of internal noise addition.

[0038] In an implementable manner, the structure refinement network further includes an attention mechanism module and an image conversion network disposed before the structure refinement network, where:

[0039] The attention mechanism module is used to output element attention parameters according to the input tensor of the structure refinement network. The input tensor is the input content map, and the element attention parameters are used to enable the image conversion network to focus on the element part of the input content map;

[0040] The image conversion network is used to generate an output tensor according to the result of the element-wise multiplication of the input tensor and the element attention parameters. The output tensor is the stylized content map, and the output tensor is used to calculate the loss function value in combination with the reference style map. The loss function value is used for the training of the structure refinement network.

[0041] This application provides an artistic word generation system for complex texture structures, including an input processing module that processes the input source text to generate a black-and-white text mask, and uses the black-and-white text mask to process the input style image to generate style patches; the first generator of the generative adversarial network model processes the black-and-white text mask and the style patches to generate style blocks with real edges expanded by a preset multiple; the second generator of the generative adversarial network model processes the style blocks to generate a black-and-white style mask for the style blocks; the detail refinement module includes a structure refinement network and a texture refinement network. The structure refinement network structurally refines the style blocks to generate intermediate artistic words; the texture refinement network texture-refines the intermediate artistic words according to the black-and-white style mask to generate the final artistic words. In this way, by generating a prototype of the artistic word and then refining the structure and details of the artistic word prototype, artistic words with complex style effects based on complex texture structures are realized. Description of the Drawings

[0042] Figure 1 It is a schematic structural diagram of an artistic word generation system for complex texture structures provided by this application;

[0043] Figure 2 It is a schematic diagram for creating a generation training set of an artistic word generation system for complex texture structures provided by this application;

[0044] Figure 3 It is a schematic structural diagram of the generative adversarial network model of an artistic word generation system for complex texture structures provided by this application;

[0045] Figure 4 It is a schematic diagram of the first generation training of an artistic word generation system for complex texture structures provided by this application;

[0046] Figure 5 It is a schematic diagram of the style training of an artistic word generation system for complex texture structures provided by this application;

[0047] Figure 6 It is an architecture diagram of the structure refinement network of an artistic word generation system for complex texture structures provided by this application;

[0048] Figure 7 It is a schematic structural diagram of the deformable module of an artistic word generation system for complex texture structures provided by this application;

[0049] Figure 8 It is a schematic diagram of the test phase of an artistic word generation system for complex texture structures provided by this application;

[0050] Figure 9 , It is a schematic diagram of the final artistic word of an artistic word generation system for complex texture structures provided by this application. Detailed implementation manners

[0051] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe in detail the implementation manners of the present application with reference to the accompanying drawings.

[0052] The terms used in the following embodiments are only for the purpose of describing specific embodiments, and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include expressions such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of the present application, "at least one" and "one or more" mean one, two or more, and "a plurality" means two or more. The term " / and" is used to describe the association relationship of associated objects, indicating that three relationships can exist; for example, A / and B can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship.

[0053] Reference to "one embodiment" or "some embodiments" etc. described in this specification means that specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0054] To solve the technical problem that the existing artistic word generation system can only generate simple artistic words based on simple styles and cannot generate artistic words with complex style effects based on complex texture structures, the present application provides an artistic word generation system for complex texture structures. The following specifically describes the embodiments of the present application with reference to the accompanying drawings.

[0055] See Figure 1 , which is a schematic structural diagram of an artistic word generation system for complex texture structures provided by the present application;

[0056] As can be seen from Figure 1 the generation system includes: an input processing module, a generative adversarial network model, and a detail refinement module that are connected in sequence;

[0057] The input processing module is used to process the input source text to generate a black-and-white text mask with smoothed edges, and use the black-and-white text mask to process the input style image to generate style patches, where the style image is an image with complex textures;

[0058] The generative adversarial network model includes a first generator and a second generator. The first generator is used to process the black-and-white text mask and the style patches to generate style blocks with real edges expanded by a preset multiple; the second generator is used to process the style blocks to generate a black-and-white style mask of the style blocks;

[0059] The generative adversarial network model is a Pro-gen GAN (Gp) network model trained through generative training. The generative training uses preset black-and-white mask patches and preset cropped style patches in a pre-created generative training set as inputs, and preset style blocks and preset black-and-white style mask blocks in the generative training set as outputs, and trains the generative adversarial network model until convergence;

[0060] The detail refinement module includes a structure refinement network (Structure Net, Ns) and a texture refinement network (Texture Net, Nt). The structure refinement network is used to perform structure refinement processing on the style blocks to generate intermediate artistic characters; the texture refinement network is used to perform texture refinement processing on the intermediate artistic characters according to the black-and-white style mask to generate final artistic characters;

[0061] The structure refinement network is a Structure Net network trained through style training. The style training uses images in a preset conventional image dataset as input content images, a preset source style image as a reference style image, and a stylized content image as output, and trains the structure refinement network until convergence.

[0062] See Figure 2 , which is a schematic diagram for creating a generative training set of an artistic character generation system for complex texture structures provided by this application;

[0063] From Figure 2 it can be seen that the generative training set in this application is pre-created in the following manner:

[0064] Step 101, select the original style image Y g , where the original style image Y g is an image with complex textures;

[0065] Step 102, obtain the original black-and-white mask M g of the original style image Y g , where the original style image Y gThe style part in the original black-and-white mask M g is a black area, and the background part in the original style image Y g in the original black-and-white mask M g is a white area;

[0066] Step 103: Select the local black-and-white mask M g with the largest black area in the original black-and-white mask M l , and the original style image Y g and the local style image Y l corresponding to the local black-and-white mask M l ;

[0067] Specifically, through Steps 101 to 103, the contour features of the style elements of the style image are obtained.

[0068] Step 104: Perform edge simplification on the original black-and-white mask M g to generate a first black-and-white mask with smooth edges

[0069] Step 105: Perform edge simplification on the local black-and-white mask M l to generate a second black-and-white mask with smooth edges

[0070] Specifically, through Steps 104 to 105, the smoothness of the edges of the black-and-white mask of the original text is simulated, and the edge simplification is completed by Gaussian blur and the sigmoid(.) function.

[0071] Step 106: According to the preset large-block cropping method, crop multiple preset style large blocks y from the original style image Y g and the local style image Y l , and obtain the corresponding preset black-and-white mask large blocks of each preset style large block y in the original black-and-white mask M g and the local black-and-white mask M l , and obtain the corresponding preset black-and-white mask large blocks m of each preset black-and-white mask large block m in the first black-and-white mask and the second black-and-white mask at the corresponding positions, and the preset black-and-white mask large blocks m s ;

[0072] In this way, six-channel real training samples of the generated training set can be created from a single style picture.

[0073] Specifically, the preset large-block cropping method is completed through the following steps:

[0074] Step 601: Set the second size of the large block to xN×xN, where xN < L and x is the preset multiple.

[0075] Step 602: Crop multiple large blocks from the first reference image according to the first probability a. The first reference image includes the original style image Y g , the original black-and-white mask M g and the first black-and-white mask

[0076] Step 603: Crop multiple large blocks from the second reference image according to the second probability 1 - a. The second reference image includes the local style image Y l , the local black-and-white mask M l and the second black-and-white mask

[0077] In this way, the preset style large block y has actual style features, and the preset black-and-white mask large block m is used to describe the true contour of the style elements. Combine the preset style large block y and the preset black-and-white mask large block m to obtain the six-channel true training sample [y; m] of the generated training set.

[0078] Step 107: Randomly crop a preset style small block from each preset style large block y The size of the preset style large block y is a preset multiple of the size of the preset style small block multiple.

[0079] Step 108: Downsample each preset black-and-white mask large block m s according to the preset multiple to obtain a preset black-and-white mask small block The size of the black-and-white mask large block m s is a preset multiple of the size of the preset black-and-white mask small block multiple.

[0080] Step 109: Crop the preset style small block through the preset black-and-white mask small block to obtain a preset cropped style small block

[0081] Step 110: Determine all the preset cropped style small blocks and all the preset black-and-white mask small blocks as the generated training set.

[0082] Specifically, the size of the preset small block is N×N, and the preset small block includes the preset style small block the preset black-and-white mask small block and the preset cropped style small block

[0083] Thus, due to the simplification of the black-and-white mask edge, the preset cropping style patch and the preset black-and-white mask patch have smooth contours, and their characteristics are similar to those of the black-and-white mask of the original text and its cropping style patch where the superscript ’ represents the input or output data in the test phase. The preset cropping style patch and the preset black-and-white mask patch are combined to generate the six-channel input data of the training set

[0084] See Figure 3 , which is a schematic structural diagram of the generative adversarial network model of an artistic word generation system for complex texture structures provided by this application;

[0085] It can be seen from Figure 3 that the first generator G p1 includes a convolutional layer, a residual module, a splicing layer, and a transposed convolutional layer set according to the generation requirements; the second generator G p2 includes a convolutional layer, a residual module, and a splicing layer set according to the generation requirements.

[0086] Specifically, the internal structure of the first generator is, in sequence, a first convolutional layer, a second convolutional layer, a third convolutional layer, multiple residual modules, a first splicing layer, a first transposed convolutional layer, a second transposed convolutional layer, and a fourth convolutional layer;

[0087] The second generator includes a fifth convolutional layer, a residual module, a sixth convolutional layer, and a second splicing layer connected in sequence;

[0088] The generation training includes first generation training and second generation training, where:

[0089] In the first generation training, by setting the stride and dilation rate of the convolutional layer inside the first generator, the stride and dilation rate of the transposed convolutional layer, and the size of the convolutional kernel, the preset black-and-white mask patch and the preset cropping style patch are used as inputs, and the preset style patch y is used as the output to process the black-and-white text mask and the style patch, and generate training for the style patch with the real edge expanded by a preset multiple;

[0090] Specifically, Sidj indicates that the stride of this layer is i and the dilation rate is j. Among them, a convolutional layer with a stride of 2 can downsample features, while a transposed convolutional layer with a stride of 2 can upsample features. kx represents that the convolutional kernel size of the convolutional layer is x×x, and ty represents that the convolutional kernel size of the transposed convolutional layer is y×y.

[0091] See Figure 4 , which is the first generation training schematic diagram of an artistic word generation system for complex texture structures provided by this application;

[0092] It can be seen from Figure 4 that during the generation training stage, the input uses the preset cropped style patches and the preset black and white mask patches in the generation training set. The output is the preset style large patch y and the preset black and white mask large patch m that are a preset multiple larger in the generation training set. Specifically, the first generator G p1 corresponds to the output of the preset style large patch y, and the second generator G p2 corresponds to the output of the preset black and white mask large patch m. The function of the first generator G p1 is to obtain a large patch image with the content expanded by a preset multiple and having real edges based on the edge-smoothing mask and the small style image. The function of the second generator G p2 is to extract the real black and white mask of this large patch image.

[0093] The second generation training realizes the training of generating the black and white style mask of the style large patch by setting the stride and dilation rate of the convolutional layer inside the second generator, using the preset style large patch y as the input and the preset black and white mask large patch m as the output to process the style large patch.

[0094] Specifically, the generative adversarial network model further includes a discriminator Dp, and the discriminator Dp is used to cooperate with the first generator to complete the first generation training.

[0095] See Figure 5 , which is the style training schematic diagram of an artistic word generation system for complex texture structures provided by this application;

[0096] It can be seen from Figure 5 that specifically, the structure refinement network is a Structure Net network trained through style training. The style training uses the pictures in the preset conventional picture dataset as the input content pictures, such as coco dataset, uses the preset source style pictures as the reference style pictures, and uses the stylized content pictures as the output, and trains the structure refinement network until convergence.

[0097] See Figure 6, which is the structural refinement network architecture diagram of an artistic word generation system for complex texture structures provided by this application;

[0098] It can be seen from Figure 6 that specifically, the structural refinement network further includes an attention mechanism module and an image transform net (ITN) arranged before the structural refinement network, where:

[0099] The attention mechanism module is used to output element attention parameters according to the input tensor of the structural refinement network. The input tensor is the input content map, and the element attention parameters are used to enable the image transform network to focus on the element part of the input content map;

[0100] Specifically, the attention mechanism module includes an average pooling layer, a seventh convolutional layer, a relu activation function layer, an eighth convolutional layer, and a sigmoid layer connected in sequence. In the figure, k1 represents that the convolution kernel of the convolutional layer is 1, Och1 represents that the output channel of the seventh convolutional layer is 1, Och3 represents that the output channel of the eighth convolutional layer is 3, and the relu activation function layer and the sigmoid layer represent the corresponding non-linear layers.

[0101] The image transform network is used to generate an output tensor according to the result of the element-wise multiplication of the input tensor and the element attention parameters. The output tensor is the stylized content map, and the output tensor is used to calculate the loss function value in combination with the reference style map. The loss function value is used for the training of the structural refinement network.

[0102] See Figure 7 , which is the structural schematic diagram of the deformable module of an artistic word generation system for complex texture structures provided by this application;

[0103] It can be seen from Figure 7 that the generation system further includes a deformable module arranged between the input processing module and the generative adversarial network model. The deformable module is used to control the deformation degree of the black and white text mask by adding noise and edge erosion to the black and white text mask.

[0104] Specifically, the deformable module realizes the control of the deformation degree of the black and white text mask by adding noise and edge erosion to the black and white text mask through the following steps:

[0105] Erode the edge of the black and white text mask and add noise to the eroded edge to obtain a noisy black and white mask;

[0106] Set the vector f to expand the noise black and white mask, achieving control over the deformation degree of the black and white text mask; the vector f includes f 0 , f 1 and f 2 , where f 0 is the size of the core for erosion and dilation, f 1 is the degree of noise addition in the edge, and f 2 is to control the degree of internal noise addition.

[0107] Specifically, the three factors in the deformable module are related to style control. Two of them are related to edge erosion, and the other is related to the degree of internal deformation. Specifically, as Figure 7 (a) shows, first, control the edge of the eroded black and white mask through f 0 , and add noise to the eroded edge through f 1 . Secondly, expand the black and white mask with noise. The scale of edge deformation is controlled by (f 0 , f 1 ). For internal deformation, noise will be added inside the text, where f 2 controls the degree of internal noise addition. The combination of (f 0 , f 1 , f 2 ) is named vector f, which determines the degree of deformation. Therefore, as Figure 7 (b) shows, various scales of text black and white masks can be obtained by changing vector f. It should be noted that the deformable module is only used to process the text black and white mask in the test stage, and then the image is cropped with the style-variable black and white mask to obtain a rough style text. As Figure 7 (c) shows, multi-scale artistic texts are obtained through the style-variable black and white mask and the cropped texture without retraining the network.

[0108] See Figure 8 , which is a schematic diagram of the test stage of an artistic word generation system for complex texture structures provided by this application;

[0109] See Figure 9 , which is a schematic diagram of the final artistic word of an artistic word generation system for complex texture structures provided by this application.

[0110] As Figure 8 , Figure 9 shows, this application sets a test stage to verify the results of the generation training and style training. In the test stage, under the control of vector f, the binary text mask and the style image enter the deformable module to obtain a rough style text prototype, which is successively fed into Gp, Ns, and Nt for forward inference operations. Finally, the output of Nt is a stylized final artistic word.

[0111] The present application provides an artistic word generation system for complex texture structures, including an input processing module that processes the input source text to generate a black-and-white text mask, and uses the black-and-white text mask to process the input style picture to generate style patches; the first generator of the generative adversarial network model processes the black-and-white text mask and the style patches to generate style blocks with real edges expanded by a preset multiple; the second generator of the generative adversarial network model processes the style blocks to generate a black-and-white style mask of the style blocks; the detail refinement module includes a structure refinement network and a texture refinement network. The structure refinement network structurally refines the style blocks to generate intermediate artistic words; the texture refinement network texture-refines the intermediate artistic words according to the black-and-white style mask to generate the final artistic words. In this way, by generating a prototype of the artistic word and then refining the structure and details of the artistic word prototype, artistic words with complex style effects based on complex texture structures are realized.

[0112] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed by the present invention; the specification and examples are only regarded as exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.

[0113] It should be understood that the present invention is not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope; the scope of the present invention is only limited by the appended claims.

Claims

1. An artistic word generation system for complex texture structures, characterized in that, the generation system includes: an input processing module, a generative adversarial network model, and a detail refinement module connected in sequence; the input processing module is used to process the input source text to generate a black-and-white text mask with smoothed edges, and use the black-and-white text mask to process the input style image to generate style patches, where the style image is an image with complex textures; the generative adversarial network model includes a first generator and a second generator. The first generator is used to process the black-and-white text mask and the style patches to generate style blocks with real edges expanded by a preset multiple; the second generator is used to process the style blocks to generate a black-and-white style mask of the style blocks; the generative adversarial network model is a Pro-gen GAN network model trained through generation training. The generation training uses preset black-and-white mask patches and preset cropped style patches in a pre-created generation training set as inputs, and preset style blocks and preset black-and-white style mask blocks in the generation training set as outputs, and trains the generative adversarial network model until convergence; the detail refinement module includes a structure refinement network and a texture refinement network. The structure refinement network is used to perform structure refinement processing on the style blocks to generate intermediate artistic words; the texture refinement network is used to perform texture refinement processing on the intermediate artistic words according to the black-and-white style mask to generate final artistic words; the structure refinement network is a Structure Net network trained through style training. The style training uses the images in a preset conventional image dataset as input content images, a preset source style image as a reference style image, and a stylized content image as the output, and trains the structure refinement network until convergence; the generation system further includes a deformable module disposed between the input processing module and the generative adversarial network model. The deformable module is used to control the deformation degree of the black-and-white text mask by adding noise and edge erosion to the black-and-white text mask; the deformable module realizes the control of the deformation degree of the black-and-white text mask by adding noise and edge erosion to the black-and-white text mask through the following steps: Erode the edges of the black-and-white text mask and add noise to the eroded edges to obtain a noisy black-and-white mask; Set the vector f to expand the noise black-and-white mask, achieving control over the deformation degree of the black-and-white text mask; the vector f includes f 0 , f 1 and f 2 , where f 0 is the size of the core for erosion and expansion, f 1 is the degree of noise addition in the edge, and f 2 is for controlling the degree of internal noise addition.

2. The artistic word generation system for complex texture structures according to claim 1, characterized in that, the generation training set is pre-created in the following manner: Select the original-style image Y g , where the original-style image Y g is an image with complex textures; Obtain the original style image Y g of the original black and white mask M g , and the style part in the original style image Y g is a black area in the original black and white mask M g , and the background part in the original style image Y g is a white area in the original black and white mask M g ; Select the original black-and-white mask M according to a preset first size L×L g The partial black-and-white mask M with the largest black area in l , and the original style image Y g In the corresponding partial style image Y of the partial black-and-white mask M l ; l ; Perform edge simplification processing on the original black and white mask M g to generate a first black and white mask with smooth edges Perform edge simplification processing on the local black-and-white mask M l to generate a second black-and-white mask with smooth edges According to the preset large-piece cutting method, from the original style image Y g and the local style image Y l multiple preset style large pieces y are cut out, and for each preset style large piece y, the corresponding preset black-and-white mask large piece in the original black-and-white mask M g and the local black-and-white mask M l is obtained, and for each preset black-and-white mask large piece m, the corresponding preset black-and-white mask large piece m in the first black-and-white mask and the second black-and-white mask is obtained s ; Randomly cut out a preset style small piece from each preset style large piece y The size of the preset style large piece y is a preset multiple of the size of the preset style small piece times; For each preset black-and-white mask large block m s Downsample according to the preset multiple to obtain a preset black-and-white mask small block The black-and-white mask large block m s has a size that is a preset multiple of the size of the preset black-and-white mask small block ; Through the preset black-and-white mask patch Crop the preset style patch Obtain the preset cropped style patch All preset clipping style chunks and all preset black and white mask chunks are determined as the generated training set.

3. The artistic word generation system for complex texture structures according to claim 2, characterized in that, the preset large block cropping method includes: Set the second size of the large block to xN×xN, where xN < L, where x is the preset multiple; Crop a plurality of large blocks from the first reference image according to the first probability a, where the first reference image includes the original style image Y g , the original black and white mask M g and the first black and white mask Crop out a plurality of large blocks from the second reference image according to the second probability 1 - a, where the second reference image includes the local style image Y l , the local black and white mask M l and the second black and white mask 4. The artistic word generation system for complex texture structures according to claim 3, characterized in that, The size of the preset small block is N×N, and the preset small block includes the preset style small block the preset black-and-white mask small block and the preset cutting style small block 5. The artistic word generation system for complex texture structures according to claim 1, characterized in that, The first generator includes a convolutional layer, a residual module, a splicing layer, and a transposed convolutional layer set according to generation requirements; the second generator includes a convolutional layer, a residual module, and a splicing layer set according to generation requirements.

6. An artistic word generation system for complex texture structures according to claim 5, wherein, the generation training includes first generation training and second generation training, wherein: The first generation training sets the stride and dilation rate of the internal convolutional layer of the first generator, the stride and dilation rate of the transposed convolutional layer, and the size of the convolutional kernel, and uses the preset black and white mask patch and the preset cropping style patch as inputs and the preset style large patch y as the output, to process the black and white text mask and the style patch, and generate the training for the style large patch that expands the real edges by a preset multiple. In the second generation training, by setting the stride and dilation rate of the convolutional layer inside the second generator, using the preset style block y as the input and the preset black-and-white mask block m as the output, the style block is processed to generate the training of the black-and-white style mask of the style block.

7. An artistic word generation system for complex texture structures according to claim 6, wherein, the generative adversarial network model further includes a discriminator, and the discriminator is used to cooperate with the first generator to complete the first generation training.

8. An artistic word generation system for complex texture structures according to claim 1, wherein, the structure refinement network further includes an attention mechanism module and an image conversion network arranged before the structure refinement network, wherein: The attention mechanism module is used to output element attention parameters according to the input tensor of the structure refinement network. The input tensor is the input content map, and the element attention parameters are used to enable the image conversion network to focus on the element part of the input content map; The image conversion network is used to generate an output tensor according to the result of the element-wise multiplication of the input tensor and the element attention parameters. The output tensor is the stylized content map, and the output tensor is used to calculate the loss function value in combination with the reference style map, and the loss function value is used for the training of the structure refinement network.

Citation Information

Patent Citations

  • Artistic text image generation method based on neural style migration

    CN111553837A

  • Wordart image synthesis system and method based on generative adversarial network

    CN114037644A