Embroidery-imitating style generation method of printed pattern
Through the embroidery style printed image generation model using a multi-encoder structure and a multi-channel information fusion attention mechanism, the problem of high-quality, controllable element embroidery image generation in the existing technology is solved, and clear and delicate embroidery style image generation is achieved, and flexibility and diversity are provided.
Patent Information
- Application Number
- CN202510283862.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to achieve high-quality and controllable embroidery image generation, and the style diversity and universality are poor.
An embroidered printed image generation model with a multi-encoder structure and a multi-channel information fusion attention mechanism is adopted. This model uses dual stages of information decoupling and feature fusion to capture and decouple complex relationships between different features.
The generated embroidery style images have clear and delicate textures, which can truly reflect the characteristics of embroidery art, and achieve higher precision embroidery image generation. At the same time, they have flexibility and diversity, and are suitable for pattern generation in different styles.
Smart Images

Figure CN120219147A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for generating an embroidered style of a printed pattern, belonging to the field of computer image generation. Background Art
[0002] The printing process is a process of applying printed patterns on textiles with dyes or pigments. With the development of technology and the times, digital printing has gradually replaced traditional printing methods, featuring high efficiency and low cost. Digital printing requires designers to design printed images and perform printing based on computer and numerical control technologies. Embroidery is a general term for various decorative patterns embroidered on fabrics with needles and threads, usually completed manually, which requires a large amount of time and labor costs. Therefore, designing printed images with an embroidery effect can weave textiles with an embroidery effect through digital printing, which has great economic benefits and saves a large amount of time and labor costs.
[0003] Designing printed images with an embroidery effect requires experienced embroidery designers and time costs. By leveraging artificial intelligence to convert ordinary flat printed images into printed images with an embroidered style, it can effectively save the time and labor costs of embroidery designers and efficiently assist embroidery designers in designing embroidered printed images.
[0004] In recent years, the further development of artificial intelligence has brought new methods for the generation of embroidered style printed images. Goodfellow et al. proposed the generative adversarial network theory (GAN) in "Generative adversarial network", which realizes the generation of pictures by promoting each other between the generator and the discriminator, providing a new idea for picture style conversion. It is a possibility to achieve the transformation of ordinary printing to an embroidered style through the GAN network.
[0005] Chinese Patent with Publication No. CN118628336A discloses an image style transfer method based on a residual network, which fuses embroidery and flat printing in the latent space by a consistent fusion method of multi-scale transformation MST and realizes image reconstruction through a decoder; although the structure is simple, the constraint is poor and the element editability cannot be achieved.
[0006] The generation of embroidery-style prints depends on the constraints of the original image, and this process is similar to the method of image translation. The goal of image translation is to achieve domain transfer. This task is to transform a flat print pattern into a more three-dimensional print pattern with an embroidery effect. After the deep convolutional generative adversarial network (DCGAN) proposed by Radford et al. in "Unsupervised representation learning with deep convolutional generative adversarial networks", the framework of end-to-end image translation was proposed and has been continuously developed.
[0007] Chinese Patent No. CN118015127A discloses a method and device for synthesizing a design pattern with embroidery texture and a texture, and an electronic device. The generator network includes a content encoder and a style encoder, which are used to downsample the content image and the style image to obtain features, and combine the features through a converter to finally generate an embroidery-style image, as Figure 1 shown. However, due to the singularity of the loss function, it is impossible to control the generation of element content, and the style scalability is poor.
[0008] Chinese Patent No. CN109308380A discloses a method for simulating an embroidery art style based on non-photorealism. By superimposing an edge image on an embroidery line texture image and then performing concavo-convex processing, an embroidery texture image with a 2.5D stereoscopic visual effect is obtained, and then different background images such as fabrics are input, and the embroidery texture image and the background image are fused to obtain a final embroidery art effect image; as Figure 2 shown, Figure 2 in (a) is the original image, Figure 2 in (b) and (c) are the generated images with different fabrics as the background. Since there is a relatively complex non-linear relationship between different ordinary print images and embroidery-style images, the method of superimposing edge images and concavo-convex processing cannot well fit the embroidery style, resulting in a weak embroidery style and lack of universality, and the style diversity cannot be achieved.
[0009] Chinese Patent No. CN117094882A discloses a method, system, device and medium for lossless digital embroidery image style transfer. By constructing a style transfer network model including a reversible residual module and a style conversion module based on an attention mechanism, the reversible residual module can map the first feature map of the content image and the second feature map of the embroidery image, and after style transfer, the target image is regressed; however, the style is limited and the scalability is poor.
[0010] In summary, the above approach has not yet achieved the high-quality generation of embroidered images of controllable elements in a mature manner, and a method for generating embroidered images that can achieve stylized design is urgently needed to be proposed. Summary of the Invention
[0011] To solve the problems existing in the above-mentioned prior art, the present invention provides a method for generating an embroidered style of a printed pattern, which mainly includes an embroidered style printed image generation model. The generation model includes a first generator, a second generator, a first discriminator, and a second discriminator; the first generator and the second generator have a two-stage information decoupling and feature fusion, and are of a multi-encoder structure, including a first encoder, a second encoder, a third encoder, a multi-channel information fusion attention mechanism, a cross-scale feature communication operator, and a decoder.
[0012] The first object of the present invention is to provide an embroidered style printed image generation model, which is implemented based on the method proposed by the present invention and includes the following steps: S1. Respectively collect embroidered images and flat printed images to construct an embroidered image data set and a flat printed image data set; S2. Use the embroidered image data set and the flat printed image data set obtained in S1 for training to obtain an embroidered style printed image generation model; S3. Encode and input the printed image into the generation model; S4. The embroidered style printed image generation model outputs a printed image with an embroidered style; In one implementation, the embroidered style printed image generation model includes a first generator, a second generator, a first discriminator, and a second discriminator; the inputs of the first generator and the second discriminator are respectively the original printed image and the ordinary printed image , and the inputs of the second generator and the first discriminator are respectively the original embroidered image and the embroidered effect printed image .
[0013] In one implementation, the structures of the first encoder and the second encoder in the first generator and the second generator are both three-layer convolutional layers with residual connections, which are respectively a low-scale feature extraction path with a 3x3 convolutional kernel, a medium-scale feature extraction path with a 5x5 convolutional kernel, and a high-scale feature extraction path with a 7x7 convolutional kernel; The input of the first encoder is the original printed image , and the output is the original printed image latent space vector ; the input of the second encoder is the Canny contour feature map , and the output is the Canny contour feature map latent space vector ; The third encoder structure is a multi-dimensional structure with multi-scale information exchange constructed by combining a convolutional layer, a fully connected layer, and a channel fusion layer. The input is an element feature map , and the output is the hidden space vector of the element feature map . The calculation method is as follows:
[0014]
[0015]
[0016] Among them, is the first encoder, is the second encoder. The calculation method is as follows:
[0017] Among them, is the low-scale feature extraction calculation of the first encoder, is the medium-scale feature extraction calculation of the first encoder, is the high-scale feature extraction calculation of the first encoder, is the weight matrix of the first encoder, is the dimension merging operation, is the non-linear mapping calculation;
[0018] Among them, is the low-scale feature extraction calculation of the second encoder, is the medium-scale feature extraction calculation of the second encoder, is the high-scale feature extraction calculation of the second encoder, is the weight matrix of the second encoder, is the dimension merging operation, is the non-linear mapping calculation; is the third encoder. The calculation method is as follows:
[0019] Among them, is the low-scale feature extraction calculation of the third encoder, is the element feature map, is the weight matrix of the third encoder, is the bias of the third encoder, is the dimension flattening operation, is the fully connected calculation.
[0020] In one implementation, the multi-channel information fusion attention mechanism receives the hidden space vector of the original printed image 、The latent space vector of the Canny contour feature map and the latent space vector of the element feature map are three inputs, and one feature vector is output , and its calculation formula is as follows:
[0021]
[0022]
[0023] Among them, represents the dimensionality transformation convolution calculation, represents the dimensionality transformation convolution calculation, represents the dimensionality transformation convolution calculation; represents the intermediate vector of the original printed image, represents the intermediate vector of the Canny contour feature map, represents the intermediate vector of the element feature map; the purpose of the dimensionality transformation convolution calculation is to make , , have the same dimensionality size; For , , perform calculations to obtain the output of the th attention branch , and its calculation method is:
[0024] where is the relative position embedding matrix; is the dimensionality of the input, is the non-linear transformation method, is the output of the attention branch, is the serial number of the branch, ; Concatenate and project the outputs of all attention branches to obtain the feature vector , and the calculation method is as follows:
[0025] where is the weight matrix, is the dynamically learned branch weight, is the fusion weight, which determines which branch outputs are more important through the learned attention layer, is the corresponding The output of the attention branch, corresponds to the output of the attention branch, corresponds to the attention branch, is a merging operation, is the serial number of the branch, , is the output of the multi-channel information fusion attention mechanism.
[0026] In one implementation, the cross-scale feature communication operator has three modules: a Fourier transform module, a spatial attention feature extraction module, and a channel attention segmentation module; the input of the cross-scale feature communication operator is a feature vector , which is split along the channels into a frequency feature vector and a spatial feature vector ; among them, the frequency feature vector is first input into the channel attention module, and after calculation, it is input into the Fourier transform module, and finally the frequency feature map is output. The spatial feature vector is input into the spatial attention feature extraction module after being calculated by the channel attention module, and finally the spatial feature map is output; The calculation method is as follows:
[0027]
[0028] Among them, M fu represents the Fourier transform module, M s represents the spatial attention feature extraction module, M c represents the channel attention segmentation module; The obtained frequency feature map and the spatial feature map are added to obtain the output of the cross-scale feature communication operator;
[0029] Decoding gives the embroidered effect printed image , and the calculation method of the decoder is as follows:
[0030] Among them, is a convolution operation with a 4x4 convolution kernel, is a convolution operation with a 3x3 convolution kernel, and are different weight matrices, and is the embroidered effect printed image.
[0031] In one embodiment, the three inputs corresponding to the first encoder, the second encoder, and the third encoder in the first generator and the second generator are the original printed image , the Canny contour feature map and the element feature map ; among them, the Canny contour feature map is obtained by performing a series of image processing means such as Otsu threshold method and Canny edge feature extraction on the printed image, and the element feature map is obtained by extracting features of elements in different printed images through CNN, and the high-level feature map is taken as the element feature map. The CNN is trained in a supervised manner in advance.
[0032] The second object of the present invention is to provide a training method for an embroidered style printed image generation model, and the training method of this model is implemented based on the embroidered style printed image generation model proposed by the present invention.
[0033] In one embodiment, the data set used includes a first data set and a second data set: First data set: a flat printed image data set, collecting ordinary printed images as data set samples, denoted as B ; Second data set: an embroidered image data set, collecting image samples with typical embroidered image characteristics, denoted as A ; In one embodiment, the loss function used is as follows: The total loss function is composed of 4 different loss functions: (1.13) Among them, G represents the first generator, F represents the second generator, D 1 represents the first discriminator, D 2 represents the second discriminator, C represents edge information image extraction for a specific image, represents the third loss function L 3 is the weight of the third loss function in the total loss function, represents the fourth loss function L 4 is the weight of the fourth loss function in the total loss function.
[0034] Introduce the first loss function L 1 and the second loss function L 2, so as to constrain the authenticity of the generated embroidered image. The first loss function L 1 and the second loss functionL After the discriminator needs to be trained to be able to effectively distinguish between real and fake embroidery images, that is, the generator is trained to generate pictures and the real images in the original embroidery dataset are jointly input into the discriminator. When the discriminator cannot make an effective distinction, at this time, the loss function is the smallest, and the difference between the image generated by the generator and the target embroidery real image is the smallest; The first loss function L 1 and the second loss function L 2 are expressed as: (1.14) (1.15) Among them, x represents the original printed image, y represents the original embroidery image, c represents the edge information image, and take x ∈ B, y ∈ A, E x~Pdata(x) represents x the expected value with respect to the data distribution P data (x) of, P data (x) represents x the probability distribution of, E y~Pdata(y) represents y the expected value with respect to the data distribution P data (y) of, P data (y) represents y the probability distribution of; The third loss function L 3 and the fourth loss function L 4 are introduced to constrain the element integrity of the generated embroidery image; The third loss function L 3 is expressed as: (1.16) Among them, ||·||1 represents the L1 norm; the first loss function L 1 and the second loss function L 2 can ensure the "realistic sense" of the generated image, but cannot guarantee the "corresponding relationship" between the generated image and the input image. The third loss function L 3 can input the generated embroidery printed image into the second generator and compare its output with the floor plan in the first dataset, so as to ensure the element integrity of the generated image. The third loss functionL 3 represents the input x and y the difference from the original image after conversion in two directions. The third loss function L The smaller the value of 3, the smaller the difference from the original image. At this time, the two generators can be effectively constrained; The fourth loss function L The expression of 4 is: (1.17) The loss function L The role of 4 is to further constrain the content consistency of the generated embroidered image. Its core idea is to enable an image to be reconstructed back to the original image after conversion in two directions, thereby maintaining the structural information of the image. At the same time, we found that the most important feature difference between ordinary printed images and embroidered images lies in the edge information. The embroidered style will have more complex and deeper edge information. Therefore, comparing the edge features of the secondarily reconstructed image with the original image in advance can effectively constrain the generator; Compare the edge information image of the flat print and the edge information image of the generated embroidered-style print to generate a flat print image. The boundary between elements and the background is strong. By comparing the edge information images, the integrity of the content and elements can be effectively guaranteed. When generating embroidered fabric images with different numbers and styles of elements, the weight coefficient before the fourth loss function L 4 can be reduced to 0, thus ensuring the diversity of generation.
[0035] Advantages of the present invention: High-quality generation: Through the multi-scale residual and attention mechanisms, the generated embroidered-style images have clear and delicate textures, and can truly reflect the characteristics of embroidery art; compared with the existing ones, higher-precision embroidered images can be achieved. The convolutional self-attention mechanism improves the judgment ability for complex details, enhances the robustness of the model, and improves the model stability.
[0036] Flexibility and diversity: The system can generate diverse embroidered styles, meet personalized design needs, be applicable to the generation of patterns in different styles, and has strong scalability.
[0037] No need to match the dataset: The cyclic supervision structure of the generation model can use non-matching datasets, thereby reducing costs.
[0038] Controllable element content: The generator has two stages of information decoupling and feature fusion, and the structure of multiple encoders enables the model to capture the complex relationships between different features, thereby achieving deeper feature decoupling. The two-stage feature fusion can enable the model to further understand the dependence relationship between features, and the shallow feature information and deep semantic information are further communicated, thereby realizing the controllability of the generated elements. Brief Description of the Drawings
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0040] Figure 1 is the effect diagram of generating an image by using the design pattern synthesis method with an embroidered texture in the prior art; Figure 2 is the effect diagram of generating an image and the original image by using the non-photorealistic embroidered art style simulation method in the prior art; Figure 3 is the schematic structural diagram of the embroidered style printing image generation model provided in the first embodiment of the present invention; Figure 4 is the schematic structural diagram of the generator in the embroidered style printing image generation model provided in the first embodiment of the present invention; Figure 5 is the comparison diagram of the effects of the embroidered printed fabric and the ordinary printed fabric generated by the embroidered style printing image generation model provided in the first embodiment of the present invention. Detailed Embodiments
[0041] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the drawings.
[0042] First Embodiment This embodiment provides an embroidered style printing image generation model. As Figure 3 shown, the generation model includes a first generator, a second generator, a first discriminator, and a second discriminator; the inputs of the first generator and the second discriminator are the original printed image and the ordinary printed image , and the inputs of the second generator and the first discriminator are the original embroidered image and the embroidered effect printed image .
[0043] As Figure 4 shown, both the first generator and the second generator include a first encoder, a second encoder, a third encoder, a multi-channel information fusion attention mechanism, a cross-scale feature communication operator, and a decoder.
[0044] Among them, the first and second encoders in the first and second generators are both three-layer convolutional layers with residual connections, namely the low-scale feature extraction path with a 3x3 convolutional kernel, the medium-scale feature extraction path with a 5x5 convolutional kernel, and the high-scale feature extraction path with a 7x7 convolutional kernel; the input of the first encoder is the original printed image , and the output is the latent space vector of the original printed image ; the input of the second encoder is the Canny contour feature map , and the output is the latent space vector of the Canny contour feature map ; The structure of the third encoder is a multi-dimensional structure with multi-scale information exchange constructed by combining a convolutional layer, a fully connected layer, and a channel fusion layer. The input is the element feature map , and the output is the latent space vector of the element feature map , and the calculation method is as follows:
[0045]
[0046]
[0047] Among them, is the first encoder, is the second encoder, and the calculation method is as follows:
[0048] Among them, is the low-scale feature extraction calculation of the first encoder, is the medium-scale feature extraction calculation of the first encoder, is the high-scale feature extraction calculation of the first encoder, is the weight matrix of the first encoder, is the dimension merging operation, is the non-linear mapping calculation;
[0049] Among them, is the low-scale feature extraction calculation of the second encoder, is the medium-scale feature extraction calculation of the second encoder, is the high-scale feature extraction calculation of the second encoder, is the weight matrix of the second encoder, is the dimension merging operation, is the non-linear mapping calculation, is the third encoder, and the calculation method is as follows:
[0050] Among them, is the low-scale feature extraction calculation of the third encoder, is the element feature map, is the weight matrix of the third encoder, is the bias of the third encoder, is the dimension flattening operation, is the fully connected calculation.
[0051] The multi-channel information fusion attention mechanism receives the original printed image latent space vector , the Canny contour feature map latent space vector , and the element feature map latent space vector as three inputs, and outputs a feature vector , and its calculation formula is as follows:
[0052]
[0053]
[0054] Among them, represents the dimension transformation convolution calculation, represents the dimension transformation convolution calculation, represents the dimension transformation convolution calculation; represents the intermediate vector of the original printed image, represents the intermediate vector of the Canny contour feature map, represents the intermediate vector of the element feature map; the purpose of the dimension transformation convolution calculation is to make , , have the same dimension size; For , , , calculate to obtain the output of the th attention branch, and its calculation method is:
[0055] Among them is the relative position embedding matrix; is the dimension of the input, is the non-linear transformation method, is the output of the attention branch, is the serial number of the branch, ; Concatenate and project the outputs of all attention branches to obtain the feature vector , and the calculation method is as follows:
[0056] where is the weight matrix, is the dynamically learned branch weight, is the fusion weight, which determines which branch outputs are more important through the learned attention layer, corresponds to the output of the attention branch, corresponds to the output of the attention branch, corresponds to the attention branch, is the merge operation, is the serial number of the branch, , is the output of the multi-channel information fusion attention mechanism.
[0057] The cross-scale feature communication operator has three modules: the Fourier transform module, the spatial attention feature extraction module, and the channel attention segmentation module; the input of the cross-scale feature communication operator is the feature vector , which is split along the channels into the frequency feature vector and the spatial feature vector ; among them, the frequency feature vector is first input into the channel attention module, and after calculation, it is input into the Fourier transform module, and finally the frequency feature map is output. The spatial feature vector is input into the spatial attention feature extraction module after calculation by the channel attention module, and finally the spatial feature map is output; The calculation method is as follows:
[0058]
[0059] where M fu represents the Fourier transform module, M s represents the spatial attention feature extraction module, M c represents the channel attention segmentation module; Add the obtained frequency feature map and the spatial feature map to obtain the output of the cross-scale feature communication operator;
[0060] Decode to obtain the embroidered effect printed image , and the calculation method of the decoder is as follows:
[0061] Among them, is a convolution operation with a 4x4 convolution kernel, is a convolution operation with a 3x3 convolution kernel, and are different weight matrices, is the embroidered effect printed image.
[0062] The three inputs corresponding to the first encoder, the second encoder, and the third encoder in the first generator and the second generator are the original printed image , the Canny contour feature map and the element feature map ; among them, the Canny contour feature map is obtained by performing a series of image processing means such as Otsu threshold method and Canny edge feature extraction on the printed image, and the element feature map is obtained by extracting features of elements in different printed images through CNN, and the high-level feature map is taken as the element feature map. The CNN is trained in a supervised manner in advance.
[0063] The decoder can map the high-dimensional feature vector in the latent space back to the data space to achieve the embroidery style transfer of the printed design image.
[0064] The generation steps of the embroidered style printed image generation model are as follows: The first generator receives the input of the original printed image and inputs it together with the edge information image of the flat printed image, and uses noise and other guiding information as an aid to generate a printed image with an embroidered style, and inputs the generated embroidered style into the first discriminator; The second generator receives the original embroidered printed image in the second dataset and outputs a normal flat printed image, and is assisted by guiding information to learn to generate a normal flat printed image; The task of the first discriminator is to identify the difference between the image generated by the first generator and the images in the embroidered image dataset based on the embroidered image dataset, and to improve the prediction accuracy as much as possible; The task of the second discriminator is to identify the difference between the image generated by the second generator and the images in the flat printed image dataset based on the flat printed image dataset, and to improve the prediction accuracy as much as possible; The structures of the first discriminator and the second discriminator are: Input layer: Receive two images, the embroidered effect printed image and the original embroidery image , connect them together to form a combined input, concatenate in the channel dimension to form a new input graph.
[0065] Convolutional layer: including the first convolutional layer and the second convolutional layer; First convolutional layer: Use multiple 3x3 convolutional kernels to extract features, combined with Batch Normalization and ReLU activation functions to improve stability; Second convolutional layer: Continue to use 3x3 convolutional kernels, with the number of output channels being 128. Also apply Batch Normalization and ReLU activation.
[0066] Pooling layer: Use the max pooling layer for spatial dimensionality reduction, reduce the size of the feature map, and improve computational efficiency.
[0067] Fully connected layer: including the flattening layer, the first fully connected layer and the second fully connected layer; Flattening layer: Flatten the final feature map into a one-dimensional vector.
[0068] First fully connected layer: Map the flattened vector to a smaller dimension, such as 256, using the ReLU activation function.
[0069] Second fully connected layer: Output layer, use the Sigmoid activation function to generate a probability value between 0 and 1, representing the input embroidered effect printed image and the original printed image similarity.
[0070] Accept the input of the embroidered effect printed image and the original embroidery image as guiding information, aiming to distinguish the difference between the real embroidery image and the generated embroidery image.
[0071] The first encoder and the second encoder are structures with feature extraction paths at low, medium, and high scales; Path 1 (low-scale feature extraction): First layer of convolution: Use a 3x3 convolutional kernel, with a stride of 1, and output 64 channels.
[0072] Second layer of convolution: Use a 3x3 convolutional kernel, with a stride of 2, and output 128 channels, reducing the spatial dimension.
[0073] Path 2 (medium-scale feature extraction): First layer of convolution: Use a 5x5 convolutional kernel, with a stride of 1, and output 128 channels.
[0074] Second - layer Convolution: Use a 5x5 convolutional kernel with a stride of 2, output 256 channels, and further reduce the spatial dimension.
[0075] Path Three (High - scale Feature Extraction): First - layer Convolution: Use a 7x7 convolutional kernel with a stride of 1, output 256 channels.
[0076] Second - layer Convolution: Use a 7x7 convolutional kernel with a stride of 2, output 512 channels.
[0077] Feature Fusion Module: Concatenate the feature maps from the three convolutional paths to form a large feature representation.
[0078] Use a 1x1 convolutional layer to map the concatenated feature maps to a unified feature dimension and reduce redundant information.
[0079] The feature maps processed by the multi - channel information fusion attention mechanism pass through a fully - connected layer, and a high - dimensional feature vector with the same size as the noise latent vector is output for subsequent generation or classification tasks.
[0080] The structure of the third encoder is as follows: Convolutional Layers: Include the first convolutional layer and the second convolutional layer; First Convolutional Layer: Use a 3x3 convolutional kernel, and increase the number of channels to num_features. The role of this layer is to extract local features in the feature map; Second Convolutional Layer: Also use a 3x3 convolutional kernel to further process the feature map and keep the number of channels as num_features. Through this layer, more complex features can be captured.
[0081] Channel Fusion Layer: Use a 1x1 convolutional kernel to compress the number of channels to num_features / / 2. The purpose of this layer is to fuse the information of different feature maps in the channel dimension and reduce redundancy.
[0082] Flatten Layer: Flatten the fused feature map into a one - dimensional vector for input to the fully - connected layer. The dimension at this time is (8, num_features / / 2 * height * width), where height and width are the spatial dimensions of the feature map.
[0083] Fully - connected Layers: Include the first fully - connected layer and the second fully - connected layer; First Fully - connected Layer: Input the flattened vector into the fully - connected layer, and output a 128 - dimensional intermediate vector. Use the ReLU activation function to increase non - linearity; Second Fully - connected Layer: Map the 128 - dimensional vector to the latent space vector.
[0084] The input of the decoder is a feature map with a shape of [8, 512, 4, 4], which is the initial low-resolution feature map after fully connected mapping.
[0085] Transposed convolution layers: including the first transposed convolution layer, the second transposed convolution layer, the third transposed convolution layer, the fourth transposed convolution layer, the fifth transposed convolution layer, and the sixth transposed convolution layer; The first transposed convolution: Transposed convolution layer: a 4x4 transposed convolution kernel, a stride of 2, 256 output channels, and an output shape of [8, 256, 8, 8].
[0086] The second transposed convolution: Transposed convolution layer: a 4x4 transposed convolution kernel, a stride of 2, 128 output channels, and an output shape of [8, 128, 16, 16].
[0087] The third transposed convolution: Transposed convolution layer: a 4x4 transposed convolution kernel, a stride of 2, 64 output channels, and an output shape of [8, 64, 32, 32].
[0088] The fourth transposed convolution: Transposed convolution layer: a 3x3 transposed convolution kernel, a stride of 2, 32 output channels, and an output shape of [8, 32, 64, 64].
[0089] The fifth transposed convolution: Transposed convolution layer: a 3x3 transposed convolution kernel, a stride of 2, 16 output channels, and an output shape of [8, 16, 128, 128].
[0090] The sixth transposed convolution (final layer): Transposed convolution layer: a 3x3 transposed convolution kernel, a stride of 2, 3 output channels (RGB image), and an output shape of [3, 224, 224].
[0091] Embodiment 2 This embodiment provides a training method for an embroidery-style printed image generation model, and the training method of this model is implemented based on the embroidery-style printed image generation model proposed in Embodiment 1. The training method includes: (1) Construct an embroidery image dataset and a flat printed image dataset; The first dataset: a flat printed image dataset, collecting ordinary printed images as dataset samples, a total of 900 images, denoted as B; The second dataset: an embroidery image dataset, collecting 500 image samples with typical embroidery image characteristics, denoted as A; (2) Introduce a loss function to train the model; The total loss function is composed of 4 different loss functions: (1.13) Among them G represents the first generator F represents the second generator D 1 represents the first discriminator D 2 represents the second discriminator, and C represents edge information image extraction for a specific image represents the third loss function L 3 is the weight of the third loss function in the total loss function represents the fourth loss function L 4 is the weight of the fourth loss function in the total loss function
[0092] Introduce the first loss function L 1 and the second loss function L 2, so as to constrain the authenticity of the generated embroidery image The first loss function L 1 and the second loss function L 2 are expressed as: (1.14) (1.15) Among them x represents the original image y represents the original embroidery image, and take x ∈ B , y ∈ A , c represents the edge information image E x~Pdata(x) represents x the expected value with respect to the data distribution P data (x) of P data (x) represents x the probability distribution of E y~Pdata(y) represents y the expected value with respect to the data distribution P data (y) of P data (y) represents y the probability distribution of; Introduce the third loss function L 3 and the fourth loss function L 4, so as to constrain the element integrity of the generated embroidery image The third loss function L 3 is expressed as: (1.16) Among them, ||·||1 represents the L1 norm; the third loss function L The function of 3 is to input the generated embroidered printed image into the second generator and compare its output with the floor plan in the first dataset, so as to ensure the element integrity of the generated image; The fourth loss function L The expression of 4 is: (1.17) The fourth loss function L The function of 4 is to further constrain the content consistency of the generated embroidered image.
[0093] Compare the edge information image of the flat print and the edge information image of the generated embroidered-style printed image to generate a flat print image. The boundary between elements and the background is strong. By comparing the edge information images, the content and element integrity can be effectively ensured. When generating embroidered fabric images with different numbers and styles of elements, the weight coefficient before the fourth loss function L 4 can be reduced to 0, so as to ensure the diversity of generation.
[0094] Such as Figure 5 shown, Figure 5 where a is the effect diagram of an ordinary flat printed fabric, Figure 5 where b is the effect diagram of the embroidered-style print obtained by the method proposed in the present invention on the fabric. It can be seen that compared with the ordinary printed fabric, the embroidered-style image of the present invention has clearer and finer textures and can better reflect the characteristics of embroidery art.
[0095] Some steps in the embodiments of the present invention can be implemented by software, and the corresponding software program can be stored in a readable storage medium, such as a CD or a hard disk, etc.
[0096] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for generating an embroidery-like style of a printed pattern, characterized in that: The method comprises: Step 1: Collect original embroidery images to construct an original embroidery image dataset, collect original print images to construct an original print image dataset, and construct an embroidery style print image generation model by training the original embroidery image dataset and the original print image dataset; Step 2: Introduce a loss function to constrain the embroidery style print image generation model obtained in step 1; Step 3: Encode the original print image and input it into the embroidery style print image generation model trained in step 2 to obtain a print image with embroidery style.
2. The method according to claim 1, characterized in that The embroidery style print image generation model in step 1 includes a first generator, a second generator, a first discriminator and a second discriminator; The input of the first generator is the original stamp image and normal printed images , the output is an embroidery effect printed image , the input of the second generator is the original embroidery image and embroidery effect printed images , output as a normal printed image ; The first generator and the second generator have two stages of information decoupling and feature fusion, and are a multi-encoder structure, including a first encoder, a second encoder, a third encoder, a multi-channel information fusion attention mechanism, a cross-scale feature exchange operator and a decoder; The first discriminator and the second discriminator each include an input layer, a convolutional layer, a downsampling layer and a fully connected layer.
3. The method according to claim 2, characterized in that The structures of the first encoder and the second encoder are both three convolutional layers with residual connections, which are respectively a low-scale feature extraction path of a 3x3 convolution kernel, a medium-scale feature extraction path of a 5x5 convolution kernel, and a high-scale feature extraction path of a 7x7 convolution kernel; The input of the first encoder is the original printed image , the output is the original printed image latent space vector ; The input of the second encoder is the Canny contour feature map , the output is the Canny contour feature map latent space vector .
4. The method according to claim 2, characterized in that: The third encoder structure is a multi-dimensional structure with multi-scale information exchange constructed by combining convolutional layers and fully connected layers with channel fusion layers, and the input is the element feature map , the output is the element feature map latent space vector , calculated as follows: in, is the first encoder, For the second encoder, the calculation method is as follows: in, Calculate the low-scale feature extraction for the first encoder, Calculate the scale feature extraction for the first encoder, Calculate the high-scale feature extraction for the first encoder, is the weight matrix of the first encoder, is a dimension merge operation, It is a nonlinear mapping calculation; in, Calculate the low-scale feature extraction for the second encoder, Calculate the scale feature extraction for the second encoder, Calculate the high-scale feature extraction for the second encoder, is the weight matrix of the second encoder, is a dimension merge operation, It is a nonlinear mapping calculation; For the third encoder, the calculation method is as follows: in, Calculate the low-scale feature extraction for the third encoder, is the element characteristic diagram, is the weight matrix of the third encoder, is the bias of the third encoder, is the dimension flattening operation, It is a fully connected calculation.
5. The method according to claim 2, characterized in that: The multi-channel information fusion attention mechanism accepts the original stamp image latent space vector , Canny contour feature map latent space vector And the element feature map latent space vector Three inputs, one feature vector output , which is calculated as follows: in, represent Dimension transformation convolution calculation, represent Dimension transformation convolution calculation, represent Dimension transformation convolution calculation; Represents the middle vector of the original printed image, Represents the intermediate vector of the Canny contour feature map, Represents the intermediate vector of the element feature map; the purpose of the dimension transformation convolution is to make , , have the same dimensions; right , , Calculate and get The output of the attention stream , which is calculated as: in is the relative position embedding matrix; is the dimension of the input, It is a nonlinear transformation method. is the output of the attention stream, is the sequence number of the tributary, ; The outputs of all attention branches are concatenated and projected to obtain the feature vector , calculated as follows: in is the weight matrix, is the dynamically learned tributary weight, is the fusion weight, which determines which tributary outputs are more important through the learned attention layer. For the corresponding The attention branch output, For the corresponding The attention branch output, For the corresponding The attention stream, is a merge operation, is the sequence number of the tributary, , It is the output of the multi-channel information fusion attention mechanism.
6. The method according to claim 2, characterized in that The cross-scale feature exchange operator includes three modules: a Fourier transform module, a spatial attention feature extraction module, and a channel attention module; the input of the cross-scale feature exchange operator is a feature vector , is split into frequency feature vectors along the channel and spatial eigenvectors ; The frequency eigenvector First, the channel attention module is input, and after calculation, it is input into the Fourier transform module, and finally the frequency feature map is output. , the spatial eigenvector After calculation by the channel attention module, it is input into the spatial attention feature extraction module, and finally the spatial feature map is output. ; The calculation is as follows: in, M fu represents the Fourier transform module, M s represents the spatial attention feature extraction module, M c represents the channel attention segmentation module; The obtained frequency characteristic diagram and spatial feature map Add together to get the output of the cross-scale feature exchange operator ; 。 7. The method according to claim 2, characterized in that The decoder converts the output of the cross-scale feature communication operator into Decode to get the embroidery effect printed image , the decoder is calculated as follows: in, is a convolution operation with a convolution kernel of 4x4. is a convolution operation with a convolution kernel of 3x3. is the first decoder weight matrix, is the weight matrix of the second decoder, Printed image for embroidery effect.
8. The method according to claim 1, characterized in that The loss function is introduced to constrain the training of the embroidery style print image generation model. The total loss function L Including the first loss function L 1. Second loss function L 2. The third loss function L 3 and the fourth loss function L 4. Total loss function L The expression is: in, G represents the first generator, F represents the second generator, D 1 represents the first discriminator, D 2 represents the second discriminator, A represents a flat printed image dataset, B represents the embroidery image dataset, C represents edge information image extraction for a specific image, Represents the third loss function L 3 is the weight in the total loss function, Represents the fourth loss function L 4 is the weight in the overall loss function.
9. The method according to claim 8, characterized in that The first loss function L 1. Second loss function L 2. The third loss function L 3 and the fourth loss function L The expression for 4 is: in, x Represents the original printed image, y Represents the original embroidery image, taking x∈ B , y∈ A , c Represents the edge information of the image, E x~Pdata(x) represent x About data distribution P data (x) The expected value of P data (x) represent x The probability distribution of E y~Pdata(y) represent y About data distribution P data (y) The expected value of P data (y) represent y The probability distribution of ||·||1 represents the L1 norm; the first loss function L 1 and the second loss function L 2 Constraining the authenticity of the generated embroidery image, the third loss function L 3. Input the generated embroidery print image into the second generator and compare its output with the planar image in the dataset to ensure the element integrity of the generated image. The fourth loss function L 4. The edge information image of the flat print is compared with the edge information image of the flat print image generated by the embroidery style print image to ensure the integrity of the content and elements.
Citation Information
Patent Citations
Simulation method of embroidery art style based on non-realistic feeling
CN109308380A
Lossless digital embroidery image style migration method, system, equipment and medium
CN117094882A
Method and device for synthesizing design pattern with embroidery texture and electronic equipment
CN118015127A
Image style migration method based on residual network
CN118628336A