A method for generating editable conditional printing images
By designing a print image generation model containing generator and discriminator, using the multi-head attention mechanism and the YT attention mechanism to fusion and parse elements and style characteristics, the problem that the existing technology cannot decouple fabric semantic information and generate specific style printing is solved, and personalized design and support for diversified needs is achieved.
Patent Information
- Application Number
- CN202411724053.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-11-28
AI Technical Summary
The prior art cannot decouple fabric semantic information, cannot generate a printing style with a specific style that meets customer needs, and cannot achieve personalized design and diversified needs.
An editable conditional printed image generation method is proposed. By designing a printed image generation model containing a generator and a discriminator, the multi-head attention mechanism and the YT attention mechanism are used to fusion and parse elements and style features, and combined with a unique data set and loss function, the generation of printed images is realized.
It realizes the decoupling of fabric semantic information, generates a printing style with a specific style that meets customer needs, supports personalized design and diversified needs, and improves the refinement and design sense of printed images.
Smart Images

Figure CN119227549B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method for generating an editable conditional printing image, and belongs to the field of image generation. Background Art
[0002] Pattern design and printing design are important links in fabric design. For the textile and apparel industry, the high fashion and timeliness determine that pattern and printing design needs to be as fast as possible. The traditional design method is that the designer designs the pattern and embodies it on the fabric by designing the weaving method, which has a long cycle and high manpower and time costs. Using artificial intelligence to assist designers in designing patterns can save a lot of time, quickly integrate design elements and styles according to specific task requirements, and generate a variety of design patterns for designers to choose and modify. Usually, this image generation method is implemented by generative adversarial networks (GAN) or diffusion models. Goodfellow et al. proposed the generative adversarial network theory (GAN) in "Generative adversarial networks", which realizes the generation of images by promoting each other between the generator and the discriminator, providing new ideas for image style transfer. After Radford et al. proposed the deep convolutional generative adversarial network (DCGAN) in "Unsupervised representation learning with deepconvolutional generative adversarial networks", the framework of end-to-end image translation was proposed and continuously developed. Mirza et al.'s "Conditional generative adversarial nets" proposed a CGAN network with conditional constraints based on GAN, which can guide model generation through conditional information.
[0003] If the input is a design image for constraint, a specified style image needs to be generated. This method is usually called image translation, which is mainly divided into three types: supervised, unsupervised, and multi-domain generation. Supervised image translation is usually implemented by a dataset composed of paired images; unsupervised image translation does not need to rely on paired training data, and is implemented based on deep learning models such as generative adversarial networks and autoencoders; multi-domain image translation can usually meet the task requirements of generating a variety of different conditions. In order to better realize the style transfer of images; Isola et al. proposed the Pix2pix algorithm model in "Image-to-image translation with conditional adversarial networks", which improved the controllability of generated images by establishing a database of paired images with different styles. Pix2pix is essentially a conditional generative adversarial network (CGAN), which uses a supervised method and has a relatively good performance in image translation tasks. Wang et al. proposed a perceptual adversarial network (PAN) in "Perceptual adversarial networks for image-to-image transformation", adding perceptual loss to Pix2pix to achieve universal image conversion. Zhu et al. proposed a matching multimodal image translation method (BicycleGAN) in "Toward multimodal image-to-image translation". This method uses a combination of two models to force the generator not to ignore noise and use noise to generate diverse results. However, due to the difficulty and high cost of establishing matching datasets for some tasks, Zhu et al. proposed an unsupervised training method for unmatched image datasets in "Unpaired image-to-image translation using cycle-consistent adversarial networks". Based on Pix2pix, the network was further improved, and a method of simultaneously training two different sets of generators and discriminators to exchange information was proposed. Cycle consistency loss was introduced to solve the problem that the generation network needs to pair images, and image style conversion was achieved using CycleGAN.However, the constraint effect of using only cycle consistency loss is weak; StyleGAN et al. proposed a novel generator architecture based on CycleGAN in "Analyzing and improving the image quality of styleGAN", which affects the random generation of details through noise during upsampling, and achieves style diversity and detail diversity without changing the main content of the image; Sem-GAN et al. built a semantically consistent framework in "semantically-consistent image-to-image translation", and used semantic information to constrain the generation of image conditions; at the same time, Tang et al. proposed a generation method (AttentionGAN) combined with the attention mechanism in "AttentionGAN: unpaired image-to-image translation using attention-guided generative ad versarial networks", which solved the shortcoming of previous algorithms that failed to translate high-level semantic information of images, and adopted a combination of output and attention mask to minimize background changes while translating images; Emami et al. introduced a spatial attention mechanism in the discriminator in "SPA-GAN: spatial attention GAN for image-to-image translation", which strengthened the ability of the discriminator.
[0004] The above-mentioned supervised image translation and multi-domain image translation can only solve the one-to-one mapping problem, but the multi-domain translation problem cannot be solved. Based on this problem, Choi et al. proposed StarGAN in "Stargan: Unified generative adversarial networks for multi-domain image-to-image translation", which realized the simultaneous translation of datasets in different domains in a single network, and trained the network's multi-domain translation ability by letting the discriminator output image categories. Although the multi-domain image translation framework can preserve the structural information of the source domain image, it cannot transfer the style of the translated image well; therefore, Sun et al. proposed multimodal unsupervised image-to-image translation without independent style encoder (MNISE-GAN) in "Multimodal unsupervised image-to-image translation without independent style encoder", which enhanced the style generation ability. Huang et al. used the MUNIT network in "Multimodal unsupervised image-to-image translation" to try to decouple the image translation process; Hu et al. believed in "Latent Style: multi-style image transfer via latentstyle coding and skip connection" that the code after style information is decoupled is random noise, and added self-attention and skip connection structures to the original MUNIT, so that the network pays more attention to global and detail information.
[0005] Fang et al. established three networks, namely generator, discriminator and classifier, in "Triple-GAN: progressive face aging with triple translation loss". The classifier predicts the label of the generated image and can generate a variety of real or non-real images, which can be applied to the generation of prints. The Chinese patent with publication number CN118628336A discloses an image style transfer method based on residual network, which uses a multi-scale transformation MST consistent fusion method to fuse embroidery and flat prints in latent space and realize image reconstruction through a decoder; however, it has the disadvantages of simple structure, poor constraints, and inability to achieve element editability.
[0006] Although Zhang Jiawei et al. achieved the generation of print images through fine-tuning in "Design of Print Pattern Generation Method Based on Diffusion Model", the fine-tuning model cannot understand the semantics of prints and fabrics well without element controllability.
[0007] A Chinese patent with publication number CN102360399A discloses a method for generating patterns of printed fabrics based on a generalized Mandelbrot set. The method constructs a generalized Mandelbrot set fractal image and its local detail image as basic elements, designs textile patterns by changing parameters in the iterative formula, configuration of color values, image magnification area and magnification, and realizes rapid design of printing. However, this method has strong image limitations and a single style, and cannot achieve personalized design.
[0008] A Chinese patent with publication number CN106709171A discloses a method for generating a printed pattern based on repeated pattern discovery. The method constructs a layout template for the repeated printed object, performs multi-granular quadrilateral meshing on the outline, solves the optimal layout, and then calculates the affine transformation of the object instance. The hierarchical relationship between the instances is used during splicing to perform instance drawing to achieve synthesis of the printed pattern. However, this method constructs a layout template for the repeated template, cannot achieve personalized generation, and the style is limited to the image set, with poor scalability.
[0009] In summary, the existing technology cannot achieve the decoupling of fabric semantic information and generate a printing style with a specific style that meets customer needs. Summary of the invention
[0010] In order to solve one or more of the above problems, the present invention proposes an editable conditional print image generation method, which can assist print designers in designing print images with different styles and patterns. Compared with existing models and methods, it can achieve the decoupling of fabric semantic information and generate a print style with a specific style that meets customer needs, so as to meet personalized design and diversified needs.
[0011] The first object of the present invention is to provide a print image generation model, which is implemented based on the method proposed by the present invention. The generation model includes a generator and a discriminator, wherein the generator structure is as follows:
[0012] like Figure 3 As shown, the generator includes a first encoder, a feature fusion module, a multi-head attention mechanism and a decoder; the first encoder accepts noise input and generates a noise latent space vector in the latent space; Figure 4As shown, the feature fusion module includes a second encoder and a YT attention mechanism; different elements and specified style feature maps are selected in the data set and merged on the first path channel after preprocessing, and multiple merged feature maps are input into the YT attention mechanism; at the same time, the second decoder is used in the second path to realize the extraction and information exchange of multi-dimensional information, and after merging with the feature map output by the YT attention mechanism of the first path, an element combined style feature map is output, and it is merged with the pre-encoded color information vector to form a latent space vector; the latent space vector and the noise latent space vector are input into the multi-head attention mechanism together, and the encoded conditional information vector is added at this time, and the latent space vector of the target image is output, and the latent space vector of the target function is decoded to obtain the target image.
[0013] In one embodiment, the first encoder structure is a three-layer convolution layer, the first two layers use a 3x3 convolution kernel, the first layer has a step size of 1, the second layer has a step size of 2, the third layer uses a 5x5 convolution kernel, and a residual connection is used after each convolution layer to ensure image clarity.
[0014] In one embodiment, the second encoder structure is three paths:
[0015] Path 1 (low-scale feature extraction):
[0016] The first convolution layer uses a 3x3 convolution kernel with a stride of 1 and outputs 64 channels.
[0017] The second convolution layer uses a 3x3 convolution kernel with a stride of 2 and outputs 128 channels.
[0018] Interaction layer: The output of the first layer is added to the output of the second layer through residual connection to enhance the low-level features.
[0019] Path 2 (mesoscale feature extraction):
[0020] The first convolution layer uses a 5x5 convolution kernel with a stride of 1 and outputs 128 channels.
[0021] The second convolution layer uses a 5x5 convolution kernel with a stride of 2 and outputs 256 channels.
[0022] Interaction layer: Connect the output of path one with the first layer output of path two to promote the communication and fusion of features.
[0023] Path three (high-scale feature extraction):
[0024] The first convolution layer uses a 7x7 convolution kernel with a stride of 1 and outputs 256 channels.
[0025] The second convolution layer uses a 7x7 convolution kernel with a stride of 2 and outputs 512 channels.
[0026] Interaction layer: The output of path 2 is added to the output of the first layer of path 3 through residual connection to enhance the interaction of mid- and high-level features. The encoder realizes multi-scale feature extraction: features are extracted in parallel through convolution kernels of different sizes; interactive residual connection: different paths share information through the interaction layer to improve the richness of feature expression.
[0027] In one embodiment, the multi-head attention mechanism structure contains three parts: key, query and value, which respectively represent the features of different dimensions in the input data. Through linear mapping, self-attention is calculated and concatenated, and finally the image latent space vector is output through linear transformation.
[0028] In one embodiment, the color information is encoded by encoding the information of the three RGB channels to form a latent space vector with a channel and a width of 1, which is similar to the size of the word vector.
[0029] In one embodiment, the conditional information is encoded using a text vector encoding method, and the conditional information is encoded using the text encoding function provided by pytorch.
[0030] In one embodiment, the decoder structure is a six-layer deconvolution upsampling structure with a skip connection structure to enhance feature expression capability and information flow.
[0031] In one embodiment, the YT attention mechanism is as follows Figure 3 As shown in the figure, it consists of two interconnected paths: one is the local path, which performs spatial attention feature extraction and maximum pooling channel feature extraction on a part of the input feature channel; the other is the global path, which performs Fourier transform and average pooling channel feature extraction on another part of the input feature channel to obtain element global information, so as to further extract deep features. Each path can capture complementary information with different receptive fields, and the information exchange between these paths is carried out internally, so as to achieve non-local receptive fields and cross-region and cross-scale fusion.
[0032] The second object of the present invention is to provide a method for training a print image generation model, which is implemented based on the technical solution of the present invention and establishes a unique data set and loss function;
[0033] In one embodiment, the data set content consists of three different data sets:
[0034] The first data set: element data set, uses CNN to extract features from elements in different print images, and saves the feature maps of their high-level output as data sets. The same element labels are placed in the same folder. More than 50 different element feature maps of different styles are saved in the data set.
[0035] The second data set: style images, by collecting and annotating print images of different styles, the data set size is 20 different styles of prints, each style contains more than 10 print images containing different elements.
[0036] The third dataset: real constrained images, each image contains a label and is uniquely representative. The dataset size is 400.
[0037] In one embodiment, the loss function is composed of four loss functions with different effects:
[0038] The total loss function L is expressed as follows:
[0039] (1.1)
[0040] Where L1 represents the first loss function, which is used to constrain the authenticity of the printed image, L2 represents the second loss function, L3 represents the third loss function, and the two are used to jointly constrain the circularity of the printed image, and L4 represents the fourth loss function. represents the weight of the second loss function L2 in the total loss function L, represents the weight of the third loss function L3 in the total loss function L, Represents the weight of the fourth loss function L4 in the total loss function L.
[0041] The first loss function L1 is introduced as follows:
[0042] (1.2)
[0043] Among them, G represents the generator, D represents the discriminator, x represents the picture sampled from the real data set, y represents the conditional information, z represents the noise, E x~Pdata(x) represents the expected value of x with respect to the data distribution Pdata(x), Pdata(x) represents the probability distribution of x, E z~Pz represents the expected value of z with respect to the data distribution Pz, and Pz represents the probability distribution of z.
[0044] The second loss function L2 is introduced to further constrain the cyclicity of the generated stamp image. When calculating the second loss function L2, multiple fixed-size cyclic stamp areas are intercepted on the generated stamp image and passed through a network to represent them as low-dimensional vectors. The cosine distance between the vectors is calculated to characterize the similarity between multiple images. The loss function L2 is expressed as:
[0045] (1.3)
[0046] Among them, p and q represent the low-dimensional vector representations of two cyclic printing area images, p i represents the i-th pixel of the low-dimensional feature p, q i represents the i-th pixel of the low-dimensional feature q, and n represents the number of pixels of the low-dimensional image features q and p, which are assumed to be the same and equal to n.
[0047] The third loss function L3 is introduced to perform a CNN feature extraction (feature extraction of different receptive fields) on the generated image, and Fourier transform is performed on each downsampled feature map, and their similar repetition frequencies are detected and compared to constrain them to specific cyclic images and image sizes;
[0048] (1.4)
[0049] Among them, f(k i ) represents the frequency characteristics of the Fourier transform acquisition of the feature map obtained by the i-th downsampling of image k. Similarly, f(k i-1 ) is the frequency characteristic of the feature map obtained by Fourier transform after the i-1th downsampling, and n is the set number of sampling times.
[0050] The fourth loss function L4 is introduced to constrain the semantic information of the printed image to enhance its information expression ability;
[0051] (1.5)
[0052] Among them, ||·||1 represents the L1 norm, C l Indicates the number of channels of the image, H l Represents the length of the image, W l represents the width of the image, G(z) represents the output content of the generator with noise as input, x represents the real image in the data set, and l represents the feature map output on different VGG layers. The goal of this function is to pass the generated image and the real image through the pre-trained VGG network, and perform the first loss function L1 regularization on the feature map output by a specific layer. It can be used to effectively compare whether the semantic information of the generated image is complete or not. The principle is that the pre-trained VGG network will extract low-level features such as edge information in the first few layers, while higher layers will output features with high-level semantic information. This loss function is improved from the perceptual loss and can better constrain this model.
[0053] Beneficial effects of the present invention:
[0054] 1. Through downsampling, different high-level features of each element of the printed plane image are extracted, and the elements and style information in the printed plane image are captured through multi-scale feature extraction and information exchange. Then, the latent space vectors corresponding to different printed plane image features are analyzed, and the fusion of specific styles is realized in the latent space. The decoder is used to decode the printed design image, realizing the decoupling of element information, which can meet personalized design and diversified needs and is scalable.
[0055] 2. The multi-input design of the model can help generate more detailed and more designed print images, providing designers with more diverse design ideas. During the decoding process, the model needs to gradually upsample the feature map so that the low-dimensional potential representation becomes a high-resolution image. In order to ensure the consistency of multi-scale features, this model introduces a multi-head attention mechanism to select the most useful features at multiple scales. Combine the output of the encoder with the features being processed by the decoder. This can help the model combine the global information in the encoder and improve the feature representation in the decoding stage.
[0056] 3. The dual-path approach further realizes the feature fusion and analysis of elements and styles, improving the expressiveness of the model. Each path of the YT attention mechanism can capture complementary information with different receptive fields, and the information exchange between these paths is carried out internally, thereby achieving non-local receptive fields and cross-region and cross-scale fusion.
[0057] 4. Combine elements, style vectors and noise latent space vectors through multi-head attention to fuse multi-scale features. When generating an image containing multiple printed elements, the multi-head attention mechanism can help the model ensure the style consistency and coordination between different printed elements.
[0058] 5. Through a uniquely designed loss function, multiple loss functions jointly constrain the authenticity, recyclability and semantic information integrity of the printed image generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0060] Figure 1 It is a diagram of the generator architecture proposed by the present invention;
[0061] Figure 2 It is a structural diagram of the feature fusion module proposed in the present invention;
[0062] Figure 3This is the YT attention mechanism architecture diagram proposed by the present invention;
[0063] Figure 4 This is the architecture diagram of the generation model proposed in the present invention. DETAILED DESCRIPTION
[0064] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0065] Embodiment 1
[0066] This embodiment provides a print image generation model, which is implemented based on the method proposed in the present invention. The generation model includes two parts: a generator and a discriminator.
[0067] like Figure 1 As shown, the generator includes a first encoder, a feature fusion module, a multi-head attention mechanism and a decoder; Figure 2 As shown, the feature fusion module includes a second encoder and a YT attention mechanism.
[0068] The first encoder structure is a three-layer convolutional layer with residual connections, which encodes the randomly distributed noise input to form a noise latent vector;
[0069] The second encoder structure consists of three paths, and the interaction of the three paths further improves the calculation of the YT attention mechanism on the element and style feature map. The encoder implements multi-scale feature extraction: extracting features in parallel through convolution kernels of different sizes; interactive residual connection: different paths share information through interactive layers to improve the richness of feature expression.
[0070] The feature fusion module uses a 1x1 convolutional layer to map the concatenated feature maps to a unified feature dimension to reduce redundant information.
[0071] like Figure 3 As shown in the figure, the YT attention mechanism consists of two interconnected paths: one is the local path, which performs spatial attention feature extraction and maximum pooling channel feature extraction on a part of the input feature channel; the other is the global path, which performs Fourier transform and average pooling channel feature extraction on another part of the input feature channel to obtain element global information.
[0072] The multi-head attention mechanism consists of three parts: key, query and value, which represent the features of different dimensions in the input data respectively. Through linear mapping, self-attention is calculated and concatenated, and finally the image latent space vector is output through linear transformation.
[0073] The decoder can map the high-dimensional feature vector in the latent space back to the data space to reconstruct the print design image.
[0074] First encoder:
[0075] The first convolution layer uses multiple 3x3 convolution kernels with a stride of 1, a padding of 1, and outputs 128 channels to extract local features.
[0076] The second convolution layer uses multiple 3x3 convolution kernels with a step size of 2 and outputs 256 channels, reducing the spatial dimension while increasing the feature abstraction capability.
[0077] The third convolution layer uses multiple 5x5 convolution kernels with a step size of 2 and outputs 512 channels to further extract global features.
[0078] Residual layer: Add residual connections between convolutional layers to ensure effective information transfer and avoid gradient vanishing. After each convolutional block, connect the input and output for addition operation.
[0079] Second encoder:
[0080] Path 1 (low-scale feature extraction):
[0081] The first convolution layer uses a 3x3 convolution kernel with a stride of 1 and outputs 64 channels.
[0082] The second convolution layer uses a 3x3 convolution kernel with a stride of 2, outputs 128 channels, and reduces the spatial dimension.
[0083] Path 2 (mesoscale feature extraction):
[0084] The first convolution layer uses a 5x5 convolution kernel with a stride of 1 and outputs 128 channels.
[0085] The second convolution layer uses a 5x5 convolution kernel with a stride of 2 and outputs 256 channels to further reduce the spatial dimension.
[0086] Path three (high-scale feature extraction):
[0087] The first convolution layer uses a 7x7 convolution kernel with a stride of 1 and outputs 256 channels.
[0088] The second convolution layer uses a 7x7 convolution kernel with a stride of 2 and outputs 512 channels.
[0089] Feature fusion module:
[0090] The feature maps from the three convolution paths are concatenated to form a large feature representation. A 1x1 convolution layer is used to map the concatenated feature maps to a unified feature dimension to reduce redundant information.
[0091] The feature map processed by the attention mechanism passes through the fully connected layer, and the output shape is a high-dimensional feature vector of the same size as the noise latent vector, which is used for subsequent generation or classification tasks.
[0092] Decoder:
[0093] All six layers have 4x4 deconvolution kernels with a step size of 2; the input feature map has a shape of [batch_size, 512, 4, 4], which is the initial low-resolution feature map after full connection mapping.
[0094] First deconvolution layer: 4x4 deconvolution kernel, stride 2, output 256 channels, output shape [batch_size, 256, 8, 8];
[0095] Second layer of deconvolution: 4x4 deconvolution kernel, stride 2, output 128 channels, output shape [batch_size, 128, 16, 16];
[0096] The third deconvolution layer: 4x4 deconvolution kernel, stride 2, output 64 channels, output shape [batch_size, 64, 32, 32];
[0097] Fourth layer of deconvolution: 4x4 deconvolution kernel, stride 2, output 32 channels, output shape [batch_size, 32, 64, 64];
[0098] Fifth deconvolution layer: 4x4 deconvolution kernel, stride 2, output 16 channels, output shape [batch_size, 16, 128, 128];
[0099] The sixth deconvolution layer (final layer): 4x4 deconvolution kernel, stride 2, output 3 channels (RGB image), output shape is [batch_size, 3, 224, 224].
[0100] Discriminator:
[0101] Input layer: The input data accepted by the discriminator is a three-dimensional tensor with the required feature map size.
[0102] Feature extraction layer:
[0103] The first convolutional layer: The size of the convolution kernel is set to 3×3 and the stride is 2. This layer extracts the preliminary features of the input image through the convolution operation, and the ReLU activation function is applied after the convolution layer to introduce nonlinear features.
[0104] First pooling layer: The maximum pooling operation is used, and the pooling window is 2×2 to halve the size of the feature map.
[0105] Second convolutional layer: Set the convolution kernel size to 3×3 and the stride to 2 to further extract higher-level features, and the activation function applies the ReLU activation function again.
[0106] The second pooling layer: uses the maximum pooling operation with a pooling window of 2×2.
[0107] Flatten layer: Flattens the output of the last pooling layer into a one-dimensional vector of length 1024.
[0108] Fully connected layer: It is set with 128 neurons and the ReLU activation function is applied to further learn feature representation.
[0109] Output layer: The final output layer of the discriminator of the present invention includes 1 neuron, and uses the Sigmoid activation function to compress the output value to between [0, 1], indicating the probability that the input sample is a true sample.
[0110] like Figure 4 As shown in FIG, the specific process of the generative model is as follows: the noise signal and the conditional information signal are input into the generator, the generator generates an image which is input into the discriminator, and the discriminator discriminates the generated image based on the real image.
[0111] Embodiment 2
[0112] This embodiment provides a method for training a stamp image generation model; the generation model training method is implemented based on the method proposed in the present invention, and the training method includes:
[0113] A unique dataset and loss function were established. The dataset content consists of three different datasets:
[0114] The first data set: element data set, uses CNN to extract features from elements in different print images, and saves the feature maps of their high-level output as data sets. The same element labels are placed in the same folder. More than 50 different element feature maps of different styles are saved in the data set.
[0115] The second data set: style images, by collecting and annotating print images of different styles, the data set size is 20 different styles of prints, each style contains more than 10 print images containing different elements.
[0116] The third dataset: real constrained images, each image contains a label and is uniquely representative. The dataset size is 400.
[0117] Loss function:
[0118] The total loss function L is expressed as follows:
[0119] (1.1)
[0120] Where L1 represents the first loss function, which is used to constrain the authenticity of the printed image, L2 represents the second loss function, L3 represents the third loss function, and the two are used to jointly constrain the circularity of the printed image, and L4 represents the fourth loss function. represents the weight of the second loss function L2 in the total loss function L, represents the weight of the third loss function L3 in the total loss function L, Represents the weight of the fourth loss function L4 in the total loss function L.
[0121] The expression of the first loss function L1 is introduced as:
[0122] (1.2)
[0123] Among them, G represents the generator, D represents the discriminator, x represents the picture sampled from the real data set, y represents the conditional information, z represents the noise, E x~Pdata(x) represents the expected value of x with respect to the data distribution Pdata(x), Pdata(x) represents the probability distribution of x, E z~Pz represents the expected value of z with respect to the data distribution Pz, and Pz represents the probability distribution of z.
[0124] The second loss function L2 is introduced to further constrain the cyclicity of the generated stamp image. When calculating the second loss function L2, multiple fixed-size cyclic stamp areas are intercepted on the generated stamp image and passed through a network to represent them as low-dimensional vectors. The similarity between multiple images is characterized by calculating the cosine distance between the vectors. The expression of the second loss function L2 is:
[0125] (1.3)
[0126] Among them, p and q represent the low-dimensional vector representations of two cyclic printing area images, p i represents the i-th pixel of the low-dimensional feature p, q i represents the i-th pixel of the low-dimensional feature q, and n represents the number of pixels of the low-dimensional image features q and p, which are assumed to be the same and equal to n.
[0127] The third loss function L3 is introduced to perform a CNN feature extraction (feature extraction of different receptive fields) on the generated image, and perform Fourier transform on each downsampled feature map, and detect their similar repetition frequencies for comparison to constrain them to specific cyclic images and image sizes. The expression of the third loss function L3 is:
[0128] (1.4)
[0129] Among them, f(ki ) represents the frequency characteristics of the Fourier transform acquisition of the feature map obtained by the i-th downsampling of image k. Similarly, f(k i-1 ) is the frequency characteristic of the feature map obtained by Fourier transform after the i-1th downsampling, and n is the set number of sampling times.
[0130] The fourth loss function L4 is introduced to constrain the semantic information of the stamp image to enhance its information expression ability. The expression of the fourth loss function L4 is:
[0131] (1.5)
[0132] Among them, ||·||1 represents the L1 norm, C l Indicates the number of channels of the image, H l Represents the length of the image, W l represents the width of the image, G(z) represents the output content of the generator with noise as input, x represents the real image in the data set, and l represents the feature map output on different VGG layers. The goal of this function is to pass the generated image and the real image through the pre-trained VGG network, and perform the first loss function L1 regularization on the feature map output by a specific layer. It can be used to effectively compare whether the semantic information of the generated image is complete or not. The principle is that the pre-trained VGG network will extract low-level features such as edge information in the first few layers, while higher layers will output features with high-level semantic information. This loss function is improved from the perceptual loss and can better constrain this model.
[0133] Some steps in the embodiments of the present invention may be implemented using software, and the corresponding software program may be stored in a readable storage medium, such as a CD or a hard disk.
[0134] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for generating an editable conditional stamp image, characterized in that: The method comprises: Step 1: construct the first data set, the second data set and the third data set; The element features in different stamp images are extracted through CNN, and the high-level output feature maps are collected as the first data set; Collect and annotate print images of different styles as the second data set; Collecting true constraint images containing labels as the third data set; A print image generation model is constructed by using the first data set, the second data set and the third data set, and a loss function is introduced to train the print image generation model; Step 2: Encode the printing image and the digital printing conditions determined according to the specific requirements of the task into embedding vectors and input them into the trained printing image generation model as guidance information; Step 3: The printing image generation model outputs a printing design image generated according to the digital printing conditions; The generation model in step 2 includes a generator and a discriminator, the generator includes: a first encoder, a feature fusion module, a multi-head attention mechanism and a decoder; the feature fusion module includes a second encoder and a YT attention mechanism; The first encoder structure is a three-layer convolutional layer with residual connections, which encodes randomly distributed noise input to form a noise latent vector; The second encoder structure includes three paths: a low-scale feature extraction path, a mid-scale feature extraction path, and a high-scale feature extraction path, and different paths of the second encoder share information through an interaction layer; The interactive layer in the low-scale feature extraction path: adding the output of the first convolution layer and the output of the second convolution layer in the low-scale feature extraction path through a residual connection; The interactive layer in the mesoscale feature extraction path: connecting the output of the low-scale feature extraction path with the first layer output of the mesoscale feature extraction path; The interactive layer in the high-scale feature extraction path: adding the output of the medium-scale feature extraction path to the output of the first layer in the high-scale feature extraction path through a residual connection; The feature fusion module concatenates the feature maps from the three convolution paths; The YT attention mechanism includes two interconnected paths: one is a local path, which performs spatial attention feature extraction and maximum pooling channel feature extraction on a part of the input feature channel; the other is a global path, which performs Fourier transform and average pooling channel feature extraction on another part of the input feature channel; The multi-head attention mechanism includes keys, queries and values, which represent the features of different dimensions in the input data. Through linear mapping, self-attention is calculated and concatenated, and the linear transformation outputs the image latent space vector; The decoder is a six-layer deconvolution upsampling with a skip connection structure.
2. The method for generating an editable conditional stamp image according to claim 1, characterized in that: The step 2 is specifically as follows: S1: The first encoder accepts the noise input and generates a noise latent space vector in the latent space; S2: dual-path processing of element and style feature maps: First path: Select different elements in the data set and specify the style feature map, merge them on the first path implementation channel after preprocessing, and input multiple merged feature maps into the YT attention mechanism; Second path: The elements and style feature maps are used in the second path to extract multi-dimensional information and exchange information through the second encoder; The outputs of the two paths are fused into an element-combined style feature map, which is then combined with the color information vector to form a latent space vector; S3: The latent space vector and the noise latent space vector are input into the multi-head attention mechanism, and then the conditional information vector is added to output the latent space vector of the target image; S4: Decode the latent space vector of the target function to obtain the target image.
3. The method for generating an editable conditional stamp image according to claim 2, characterized in that: The total loss function L is composed of four loss functions with different effects: Where L1 represents the first loss function, L2 represents the second loss function, L3 represents the third loss function, and L4 represents the fourth loss function. represents the weight of the second loss function L2 in the total loss function L, represents the weight of the third loss function L3 in the total loss function L, Represents the weight of the fourth loss function L4 in the total loss function L.
4. The method for generating an editable conditional stamp image according to claim 3, characterized in that: The expression of the first loss function L1 is: Among them, G represents the generator, D represents the discriminator, x represents the picture sampled from the real data set, y represents the conditional information, z represents the noise, E x~Pdata(x) represents the expected value of x with respect to the data distribution Pdata(x), Pdata(x) represents the probability distribution of x, E z~Pz represents the expected value of z with respect to the data distribution Pz, and Pz represents the probability distribution of z.
5. The method for generating an editable conditional stamp image according to claim 4, characterized in that: The expression of the second loss function L2 is: Among them, p and q represent the low-dimensional vector representations of two cyclic printing area images, p i represents the i-th pixel of the low-dimensional feature p, q i represents the i-th pixel of the low-dimensional feature q, and n represents the number of pixels of the low-dimensional image features q and p, which are assumed to be the same and equal to n.
6. The method for generating an editable conditional stamp image according to claim 5, characterized in that: The expression of the third loss function L3 is: Among them, f(k i ) represents the frequency characteristics of the Fourier transform acquisition of the feature map obtained by the i-th downsampling of image k, f(k i-1 ) is the frequency characteristic of the feature map obtained by Fourier transform after the i-1th downsampling, and n is the set number of sampling times.
7. The method for generating an editable conditional stamp image according to claim 6, characterized in that: The expression of the fourth loss function L4 is: Among them, ||·||1 represents the L1 norm, C l Indicates the number of channels of the image, H l Represents the length of the image, W l Represents the width of the image, G(z) represents the output content of the generator with noise as input, x represents the real image in the dataset, and l represents the feature map output on different VGG layers.
8. The method for generating an editable conditional stamp image according to claim 7, characterized in that: By determining the conditional input according to different application scenarios, the high-level features of each element mapping are calculated and adjusted in the latent space, and the decoder is used to reconstruct the print design image containing different numbers and styles of elements. The randomly distributed noise input is used to generate different print design images that meet the conditional requirements.
9. The method for generating an editable conditional stamp image according to claim 8, characterized in that: Different high-level features of each element of the printed plane image are extracted through downsampling, and different scale information in the printed plane image is captured through multi-scale feature extraction and information exchange, and the latent space vectors corresponding to different printed plane image features are analyzed.
Citation Information
Patent Citations
Generation method of printed fabric patterns based on generalized Mandelbrot set
CN102360399A
Repeat mode discovery-based printed pattern generation method
CN106709171A
Image style migration method based on residual network
CN118628336A
Unsupervised segmentation method for surface embossed character image
CN113627436A
Printing image retrieval method based on depth feature fusion
CN115544286A