A style transfer method and device for digital Xiang embroidery images based on regional semantics

Through the Hunan embroidery image style transfer method based on regional semantics, semantic segmentation and lightweight reversible network extraction features are used, and combined with the AdaIN migration algorithm, images with rich Hunan embroidery styles are generated, solving the problem of insufficient local style characteristics in Hunan embroidery image style transfer and improving the visual effect.

CN119741188BActive Publication Date: 2025-08-19CENT SOUTH UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411700363.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-08-19
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

In the transfer of Hunan embroidery image style, it is difficult to effectively capture local style characteristics, especially the transfer effect of needle method is limited, and the semantic relationship of objects in the image is not fully considered, resulting in poor visual effects.

Method used

The style transfer method based on regional semantics is adopted, and the semantic region of the Hunan embroidery image is matched through the semantic segmentation network and the CLIP network, features are extracted in combination with lightweight reversible networks, and global and local style transfers are carried out through the AdaIN migration algorithm, and the overall stylized features are integrated to generate Hunan embroidery image.

Benefits of technology

The generated Hunan embroidery style images have more unique style characteristics and texture details of Hunan embroidery, which solves the problem of insufficient transfer of local style characteristics and improves the visual effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741188B_ABST
    Figure CN119741188B_ABST
Patent Text Reader

Abstract

The present application relates to a method and device for style transfer of digital Xiang embroidery images based on regional semantics. The method finds the semantic information of the closest semantic segmentation region from the digital Xiang embroidery image based on the semantic information of the semantic segmentation region in the content image, implements regional style transfer for different semantic segmentation regions, and obtains local stylized features; performs global style transfer based on the features from the content image and the features of the digital Xiang embroidery image to obtain global stylized features, then obtains overall stylized features by fusing the local stylized features with the global stylized features, and finally generates a Xiang embroidery style image with more style characteristics and texture details unique to Xiang embroidery based on the overall stylized features, thereby improving the visual effect of the final generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of image processing technology, and in particular to a method and device for style transfer of digital Hunan embroidery images based on regional semantics. Background Art

[0002] Hunan embroidery, as one of the four famous embroidery styles in China, is famous for its exquisite craftsmanship and unique artistic style.

[0003] Style transfer is a popular image generation technology. It transfers the artistic style of an artistic image to another content image, aiming to use the artistic style of the artistic image to re-present the objects in the content image, thereby generating a new artistic image.

[0004] Currently, deep convolutional networks pre-trained on the ImageNet dataset are primarily used to extract features. Style is represented by calculating high-order statistics such as the Gram matrix and mean-variance on these features, laying the foundation for style transfer methods based on deep convolutional networks. However, these methods are limited in their ability to capture local stylistic characteristics, particularly those related to stitching techniques in Xiang embroidery. Furthermore, the texture, color, and stitching techniques of different semantic objects in artistic images vary widely. If style transfer is performed solely from the perspective of overall stylistic characteristics, without considering the semantic relationship between the Xiang embroidery style image and the objects in the content image, the resulting visual quality will be significantly affected. Summary of the Invention

[0005] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0006] The present application aims to at least solve the technical problems existing in the prior art. To this end, the present application proposes a method and apparatus for style transfer of digital Xiang embroidery images based on regional semantics. The method finds the semantic information of the closest semantic segmentation region in the digital Xiang embroidery image based on the semantic information of the semantic segmentation region in the content image, implements regional style transfer for different semantic segmentation regions, obtains local stylized features, and performs global style transfer based on the features from the content image and the features of the digital Xiang embroidery image to obtain global stylized features. The style transfer model generates a Xiang embroidery-style image with more unique stylistic features and texture details of Xiang embroidery by integrating the overall stylized features of the local and global stylized features.

[0007] A first aspect of an embodiment of the present application provides a method for style transfer of a digital Hunan embroidery image based on regional semantics, the method comprising:

[0008] Acquire a target content image and a target digital Xiang embroidery image;

[0009] The target content image and the target digital Xiang embroidery image are input into a preset style transfer model, and the style transfer model outputs a Xiang embroidery style image corresponding to the target content image; the training process of the style transfer model includes:

[0010] Extracting a plurality of first semantic segmentation regions from a first image and extracting a plurality of second semantic segmentation regions from a second image; the first image is a content image for training, and the second image is a digital Xiang embroidery image for training;

[0011] Perform semantic matching on the plurality of first semantic segmentation regions and the plurality of second semantic segmentation regions to obtain a plurality of semantic matching pairs; any one of the semantic matching pairs includes: one first semantic segmentation region and the second semantic segmentation region closest thereto;

[0012] extracting a first image feature of the first image and extracting a second image feature of the second image;

[0013] Performing global style transfer based on the first image feature and the second image feature to obtain a first stylized feature;

[0014] Calculating a first pixel dot product result between a first semantic segmentation region mask corresponding to a first semantic segmentation region in each of the semantic matching pairs and the first image feature, and a second pixel dot product result between a second semantic segmentation region mask corresponding to a second semantic segmentation region and the second image feature, and performing regional style transfer based on the first pixel dot product result and the second pixel dot product result to obtain a second stylized feature corresponding to each of the semantic matching pairs;

[0015] fusing the first stylized feature and all the second stylized features to obtain an overall stylized feature;

[0016] A Hunan embroidery style image corresponding to the first image is generated according to the overall stylized features.

[0017] This embodiment has at least the following beneficial effects:

[0018] This method provides a style transfer model. Through this style transfer model, the content image and the digital Xiang embroidery image can be semantically segmented first, and the semantic segmentation areas of the content image and the semantic segmentation areas of the digital Xiang embroidery image can be extracted respectively; then the semantic segmentation areas of the content image and the semantic segmentation areas of the digital Xiang embroidery image are semantically matched to find the semantic segmentation areas of the digital Xiang embroidery image that are most similar to the semantic segmentation areas of the content image; secondly, on the one hand, global style transfer can be performed based on the respective image features of the content image and the digital Xiang embroidery image, and on the other hand, local regional style transfer can be performed based on the semantic matching pairs, the content image, and the digital Xiang embroidery image; finally, the stylized features obtained by global style transfer and the stylized features obtained by regional style transfer are fused, and a Xiang embroidery style image corresponding to the content image is generated according to the fusion result. This method finds the semantic information of the closest semantic segmentation area (such as texture, stitch and color) from the digital Xiang embroidery image based on the semantic information of the semantic segmentation area in the content image, realizes regional style transfer for different semantic segmentation areas, and obtains local stylized features; performs global style transfer based on the features from the content image and the features of the digital Xiang embroidery image, and obtains global stylized features, so that the style transfer model generates Xiang embroidery style images with more unique style characteristics and texture details of Xiang embroidery by integrating the overall stylized features of local stylized features and global stylized features.

[0019] In some embodiments, extracting a first image feature of the first image and extracting a second image feature of the second image include:

[0020] Inputting the first image into a lightweight reversible network to extract first image features of the first image through a forward process of the lightweight reversible network;

[0021] Inputting the second image into the lightweight reversible network to extract second image features of the second image through a forward process of the lightweight reversible network;

[0022] The lightweight reversible network includes a plurality of stacked reversible units, and the forward process of the lightweight reversible network includes:

[0023] Y k+1 =Concat(Y k+1 [c+1:C],Y k+1 [1:c]);

[0024] Y k+1 [c+1:C]=Y k [c+1:C]+H(Y k [1:c]);

[0025] Y k+1 [1:c]=Y k[1:c];

[0026] Among them, Y k+1 is the output of the k+1th reversible unit, Concat is the channel concatenation function, [1:c] is the 1st to cth layer channel of any reversible unit, [c+1:C] is the cth to Cth layer channel of any reversible unit, Y k+1 [c+1:C] is the output of the k+1th reversible unit at [1:c], Y k+1 [1:c] is the output of the k+1th reversible unit at [c+1:C], and H is the mapping function.

[0027] In some embodiments, generating a Hunan embroidery style image corresponding to the first image based on the overall stylized features includes:

[0028] Inputting the overall stylized features into the lightweight reversible network to generate a Hunan embroidery style image corresponding to the first image through a reverse process of the lightweight reversible network;

[0029] The reverse process of the lightweight reversible network includes:

[0030] Y k =Concat(Y k [c+1:C],Y k [1:c]);

[0031] Y k [c+1:C]=Y k+1 [1:c]-H(Y k+1 [1:c]);

[0032] Y k [1:c]=Y k+1 [1:c];

[0033] The last reversible unit in the reverse process outputs a Hunan embroidery style image corresponding to the first image.

[0034] In some embodiments, the loss function of the style transfer model is a content loss function Style loss function First consistency loss function and the second consistency loss function The weighted sum of

[0035]

[0036] Where N is the total number of lightweight reversible networks, For a lightweight reversible network, from the first image I c Corresponding Xiang embroidery style image I csThe i-th layer features extracted from For a lightweight reversible network, from the first image I c The i-th layer features extracted from For a lightweight reversible network, from the second image I s The i-th layer features extracted from Characterized by The mean of Features The mean of Features The variance of Characterized by The variance, ||·|| is the norm, ||·||2 is the 2-norm, I cc The second image I s and the first image I c The first image I is used c As the input generated Xiang embroidery style image, I ss The second image I s and the first image I c The second image I is used s As input, the generated Xiang embroidery style image is For lightweight reversible networks from I cc The i-th layer features extracted from For lightweight reversible networks from I ss The i-th layer features extracted from .

[0037] In some embodiments, extracting a plurality of first semantic segmentation regions of the first image and a plurality of second semantic segmentation regions of the second image includes:

[0038] Extracting a plurality of first semantic segmentation region masks of the first image according to the first semantic segmentation network;

[0039] Extracting a plurality of second semantic segmentation region masks of the second image through a second semantic segmentation network; wherein the second semantic segmentation network is a deep learning SAM network;

[0040] Performing pixel-by-pixel multiplication of the first semantic segmentation region masks with the first image to obtain a plurality of first semantic segmentation regions;

[0041] Perform pixel-by-pixel multiplication on the second image and the plurality of second semantic segmentation region masks to obtain a plurality of second semantic segmentation regions.

[0042] In some embodiments, the process of calculating the first stylized feature includes:

[0043]

[0044] Among them, AdaIN is the AdaIN migration function, F c is the first image feature, F s is the second image feature, It is the first stylized feature.

[0045] In some embodiments, the process of calculating the second stylized feature includes:

[0046]

[0047] Among them, AdaIN is the AdaIN migration function, Dawn(m cmatch(j) )⊙F c Dawn(m cmatch(j) ) and F c The first pixel multiplication result, Dawn(m cmatch(j) ) is m cmatch(j) Downsampling, m cmatch(j) is the first semantic segmentation region mask j, F c is the first image feature, Dawn(m sj )⊙F s Dawn(m sj ) and F s The second pixel multiplication result, Dawn(m sj ) is m sj Downsampling, m sj is the second semantic segmentation region mask j, F s is the second image feature, m cmatch(j ) and the second semantic segmentation region corresponding to msj form a pair of semantic matching pairs j; is the second stylized feature corresponding to the semantic matching pair j.

[0048] In some embodiments, performing semantic matching on the plurality of first semantic segmentation regions and the plurality of second semantic segmentation regions to obtain a plurality of semantic matching pairs includes:

[0049] Inputting the first semantic labels corresponding to the plurality of second semantic segmentation regions and the plurality of first semantic segmentation regions into a multimodal CLIP network to obtain a plurality of matching results output by the multimodal CLIP network, wherein the process of generating any one matching result by the multimodal CLIP network includes:

[0050] C match(j)L =argmax CLIP(C iL , S j );

[0051] Among them, C iL is the first semantic label corresponding to the first semantic segmentation region i, Sj is the second semantic segmentation region j, C match(j)L C iL and S j The matching result, argmax is the maximum value function, CLIP() is the multimodal CLIP network function;

[0052] The first semantic label in each matching result is replaced with the corresponding first semantic segmentation region to obtain a plurality of semantic matching pairs.

[0053] In some embodiments, the fusing of the first stylized feature and all the second stylized features to obtain an overall stylized feature includes:

[0054]

[0055] in, is the first stylized feature, To fuse all the features after the second stylized features, Concat is the channel splicing function, avg is the average pooling function, and CONVS is the convolution function.

[0056] A second aspect of the embodiments of the present application provides a style transfer device for digital Hunan embroidery images based on regional semantics, the device comprising:

[0057] A data acquisition module is used to acquire a target content image and a target digital Xiang embroidery image;

[0058] The style transfer module is used to input the target content image and the target digital Hunan embroidery image into a preset style transfer model, so that the style transfer model outputs a Hunan embroidery style image corresponding to the target content image; the training process of the style transfer model includes:

[0059] Extracting a plurality of first semantic segmentation regions from a first image and extracting a plurality of second semantic segmentation regions from a second image; the first image is a content image for training, and the second image is a digital Xiang embroidery image for training;

[0060] Perform semantic matching on the plurality of first semantic segmentation regions and the plurality of second semantic segmentation regions to obtain a plurality of semantic matching pairs; any one of the semantic matching pairs includes: one first semantic segmentation region and the second semantic segmentation region closest thereto;

[0061] extracting a first image feature of the first image and extracting a second image feature of the second image;

[0062] Performing global style transfer based on the first image feature and the second image feature to obtain a first stylized feature;

[0063] Calculating a first pixel dot product result between a first semantic segmentation region mask corresponding to a first semantic segmentation region in each of the semantic matching pairs and the first image feature, and a second pixel dot product result between a second semantic segmentation region mask corresponding to a second semantic segmentation region and the second image feature, and performing regional style transfer based on the first pixel dot product result and the second pixel dot product result to obtain a second stylized feature corresponding to each of the semantic matching pairs;

[0064] fusing the first stylized feature and all the second stylized features to obtain an overall stylized feature;

[0065] A Hunan embroidery style image corresponding to the first image is generated according to the overall stylized features. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0067] Figure 1 This is a flowchart of the style transfer method of digital Xiang embroidery images based on regional semantics provided by this application;

[0068] Figure 2 This is a flowchart of the training style transfer model provided by this application;

[0069] Figure 3 This is a schematic diagram of the VGG encoded content leakage provided by this application;

[0070] Figure 4 is a schematic diagram of the deviation of style conversion provided by this application;

[0071] Figure 5 This is a flowchart of another method for style transfer of digital Xiang embroidery images based on regional semantics provided by this application;

[0072] Figure 6 This is a structural diagram of a style transfer device for digital Xiang embroidery images based on regional semantics provided by this application;

[0073] Figure 7 It is a structural diagram of the electronic device provided in this application. DETAILED DESCRIPTION

[0074] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0075] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0076] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0077] like Figure 1 and Figure 2 One embodiment of the present application provides a method for style transfer of digital Xiang embroidery images based on regional semantics, the method comprising:

[0078] Step S110, acquiring a target content image and a target digital Xiang embroidery image;

[0079] Step S120 , inputting the target content image and the target digital Hunan embroidery image into a preset style transfer model, and obtaining a Hunan embroidery style image corresponding to the target content image output by the style transfer model.

[0080] It should be noted that the above steps S110 and S120 are the process of using the style transfer model.

[0081] The training process of the style transfer model includes:

[0082] Step S1210, extracting a plurality of first semantic segmentation regions from a first image and extracting a plurality of second semantic segmentation regions from a second image; the first image is a content image used for training, and the second image is a digital Xiang embroidery image used for training;

[0083] Step S1220: semantically matching the plurality of first semantic segmentation regions with the plurality of second semantic segmentation regions to obtain a plurality of semantic matching pairs; any semantic matching pair includes: a first semantic segmentation region and the second semantic segmentation region closest thereto;

[0084] Step S1230, extracting a first image feature of the first image and extracting a second image feature of the second image;

[0085] Step S1240, performing global style transfer based on the first image feature and the second image feature to obtain a first stylized feature;

[0086] Step S1250: Calculate a first pixel dot product result between the first semantic segmentation region mask and the first image feature corresponding to the first semantic segmentation region in each semantic matching pair, and a second pixel dot product result between the second semantic segmentation region mask and the second image feature corresponding to the second semantic segmentation region, and perform regional style transfer based on the first pixel dot product result and the second pixel dot product result to obtain a second stylized feature corresponding to each semantic matching pair;

[0087] Step S1260: fusing the first stylized feature and all the second stylized features to obtain an overall stylized feature;

[0088] Step S1270: Generate a Hunan embroidery style image corresponding to the first image based on the overall stylized features.

[0089] In step S1210, a plurality of first semantic segmentation regions of the first image and a plurality of second semantic segmentation regions of the second image may be extracted through a conventional semantic segmentation network (such as FCN, UNet, SegNet).

[0090] In some embodiments, the step of extracting a plurality of first semantically segmented regions of the first image and a plurality of second semantically segmented regions of the second image comprises:

[0091] Step S1211: extract multiple first semantic segmentation region masks of the first image according to the first semantic segmentation network.

[0092] Step S1212: extract multiple second semantic segmentation region masks of the second image through a second semantic segmentation network; the second semantic segmentation network is a SAM network.

[0093] Step S1213 : performing pixel-by-pixel multiplication of the plurality of first semantic segmentation region masks with the first image to obtain a plurality of first semantic segmentation regions.

[0094] Step S1214 : performing pixel-by-pixel multiplication of the plurality of second semantic segmentation region masks with the second image to obtain a plurality of second semantic segmentation regions.

[0095] Because the first image is a content image, in this embodiment, a conventional semantic segmentation network can be used as the first semantic segmentation network. When the first semantic segmentation network outputs a semantic segmentation region mask, it also outputs a semantic label corresponding to the region.

[0096] Since conventional semantic segmentation networks are only trained on natural images, they rarely use digital Xiang embroidery images during training. If the semantic segmentation labels for Xiang embroidery images are remade, it will not only be time-consuming and labor-intensive, but the digital Xiang embroidery images that can be collected are also very limited compared to the hundreds of thousands of images required for semantic segmentation model training. Moreover, the training labels of the semantic segmentation network are fixed, and only these labels can be used for prediction. This will greatly affect the accuracy of the style area, which is not only reflected in whether the segmented area is complete and accurate, but also in the judgment of the semantic information in the area. Therefore, this embodiment uses a large visual model (SegmentAnything, SAM) that has been trained with massive data and has a strong zero-sample segmentation capability to automatically segment the digital Xiang embroidery image (i.e., the second image) to obtain different semantic areas.

[0097] Through this step, the first semantic segmentation region of the first image and its semantic label, and the second semantic segmentation region of the second image can be obtained.

[0098] In step S1220, the purpose of semantic matching is to find each first semantic segmentation region and its closest second semantic segmentation region. In Hunan embroidery images, different semantic regions have diverse and unique stylistic features such as stitches, colors, and textures, and are often not identical. Therefore, during the style transfer process, this embodiment first finds the content semantic region that is most semantically similar to the second semantic region of the second image based on the degree of semantic similarity, and then performs regionalized style transfer.

[0099] In some embodiments, semantic matching is performed on the plurality of first semantic segmentation regions and the plurality of second semantic segmentation regions to obtain a plurality of semantic matching pairs, including:

[0100] Step S1221: Input the first semantic labels corresponding to the plurality of second semantic segmentation regions and the plurality of first semantic segmentation regions into the CLIP network to obtain a plurality of matching results output by the CLIP network. The process of generating any matching result by the CLIP network includes:

[0101] C match(j)L =argmax CLIP(C iL , S j );

[0102] Among them, C iL is the first semantic label corresponding to the first semantic segmentation region i, S j is the second semantic segmentation region j, C match(j)L C iL and S j The matching result of CLIP is the function corresponding to the CLIP network;

[0103] Step S1222 : replacing the first semantic label in each matching result with the corresponding first semantic segmentation region to obtain a plurality of semantic matching pairs.

[0104] In order to match semantic regions, this embodiment uses the CLIP network with cross-modal understanding capabilities. CLIP (Contrastive Language-Image Pre-training) is an advanced multimodal pre-training large model proposed by OpenAI. It is trained by a large number of image-text pairs collected from the Internet, and learns rich visual and language feature representations, so that it can understand the content of the image and the text description related to it. The core idea of the CLIP model is to use contrastive learning to train a network that can map images and texts to a common feature space. In this feature space, the embedding vectors of semantically related images and texts will be close to each other, while unrelated ones will be far away from each other. Because CLIP learns rich visual and language representations, it can achieve zero-sample learning on different image classification tasks without additional training. This means that CLIP can directly classify unseen categories. Therefore, the matching process can be converted to using CLIP to classify each region.

[0105] In step S1230 , first image features of the first image are extracted and second image features of the second image are extracted.

[0106] In some embodiments, a common encoder-decoder structure may be used to extract image features or decode an image based on image features, which will not be described in detail here.

[0107] In some embodiments, extracting a first image feature of the first image and extracting a second image feature of the second image include:

[0108] Step S1231: Input the first image into a lightweight reversible network to extract first image features of the first image through a forward process of the lightweight reversible network.

[0109] Step S1232: inputting the second image into the lightweight reversible network to extract second image features of the second image through a forward process of the lightweight reversible network;

[0110] The lightweight reversible network consists of multiple stacked reversible units. The forward process of the lightweight reversible network includes:

[0111] Y k+1 =Concat(Y k+1 [c+1:C],Y k+1 [1:c]);

[0112] Y k+1 [c+1:C]=Yk [c+1:C]+H(Y k [1:c]);

[0113] Y k+1 [1:c]=Y k [1:c];

[0114] Among them, Y k+1 is the output of the k+1th reversible unit, Concat is the channel concatenation function, [1:c] is the 1st to cth layer channel of any reversible unit, [c+1:C] is the cth to Cth layer channel of any reversible unit, Y k+1 [c+1:C] is the output of the k+1th reversible unit at [1:c], Y k+1 [1:c] is the output of the k+1th reversible unit at [c+1:C], and H is the mapping function.

[0115] Style transfer methods based on common encoder-decoder structures may suffer from content leakage issues. Causes include:

[0116] The encoder and decoder architecture can use the encoder to encode the image into features, and then use the decoder to decode the features into an image. The image obtained after decoding is very similar to the original image. However, if the decoded image is input into the encoder and decoder again, after multiple rounds of iteration, the final image will be disorganized. The effect will be more obvious as the number of rounds increases. At the same time, there are deviations in the decoding process. The two work together to cause content leakage problems. Even if a first image (i.e., content image) and a second image (i.e., style image) are used to generate a Xiang embroidery style image through a style transfer model, if the Xiang embroidery style image is used as the first image again, and the second image continues to be used as the second image for multiple rounds of style transfer operations, it is found that the original content of the first image is missing in the final Xiang embroidery style image, and the content of the second image is introduced instead. For details, see Figure 3 and Figure 4 . Figure 3 The top and bottom rows show the process of content leakage of two images from left to right.

[0117] In order to solve the existing content leakage problem while taking into account the feature extraction capability and computational efficiency, this embodiment provides a lightweight reversible network, which extracts the content features in the first image and the Hunan embroidery style features in the second image through the forward process of the lightweight reversible network, thereby extracting pure content features (i.e., first image features) and style features (i.e., second image features), maintaining the integrity of the features during the forward propagation process, and preventing content leakage problems, that is, performing lossless style transfer.

[0118] In step S1240 , global style transfer is performed based on the first image feature and the second image feature to obtain a first stylized feature.

[0119] In some embodiments, a biased transfer network may be used for global or local style transfer.

[0120] In some embodiments, the process of calculating the first stylized feature includes:

[0121]

[0122] Among them, AdaIN is the AdaIN migration function, F c is the first image feature, F s is the second image feature, It is the first stylized feature.

[0123] Although a biased transfer network can be used to perform global or local style transfer, it may cause content leakage. Therefore, the unbiased AdaIN transfer algorithm in this embodiment performs global style transfer to achieve lossless style transfer.

[0124] In step S1250, a second stylized feature corresponding to each semantic matching pair is calculated. In some embodiments, the process of calculating the second stylized feature includes:

[0125]

[0126] Among them, AdaIN is the AdaIN migration function, Dawn(m cmatch(j) )⊙F c Dawn(m cmatch(j) ) and F c The first pixel multiplication result, Dawn(m cmatch(j) ) is m cmatch(j) Downsampling, m cmatch(j) is the first semantic segmentation region mask j, F c is the first image feature, Dawn(m sj )⊙F s Dawn(m sj ) and F s The second pixel multiplication result, Dawn(m sj ) is m sj Downsampling, m sj is the second semantic segmentation region mask j, F s is the second image feature, m cmatch(j) The corresponding first semantic segmentation area and m sj The corresponding second semantic segmentation regions form a pair of semantic matching pairs j; is the second stylized feature corresponding to the semantic matching pair j.

[0127] The unbiased AdaIN migration algorithm of this embodiment performs local style migration and can achieve lossless style migration.

[0128] Because the textures, colors, stitches, and other stylistic characteristics of different semantic objects in the second image (i.e., the digital Xiang embroidery image) vary, if style transfer is performed solely from the perspective of global stylistic features without considering the semantic relationship between the objects in the second image and the first image, the resulting visual effect will be significantly affected.

[0129] In the above steps, global first stylized features were extracted. This embodiment builds on this by extracting local second stylized features corresponding to each semantically matched pair. By combining these global first and local second stylized features in the image generation process, style transfer is not only achieved from a holistic perspective, but also limited to semantically matched regions. In particular, stylistic features such as stitching, color, and texture within different semantic regions of the second image are transferred, resulting in a Xiang embroidery-style image with more unique stylistic characteristics and textural details.

[0130] In step S1260 , the first stylized feature and all the second stylized features are fused to obtain an overall stylized feature.

[0131] In some embodiments, the first stylized feature and all the second stylized features are fused to obtain an overall stylized feature, including:

[0132]

[0133] in, is the first stylized feature, To fuse all the features after the second stylized features, Concat is the channel splicing function, avg is the average pooling function, and CONVS is the convolution function.

[0134] In this embodiment, an adaptive fusion process is designed to adaptively fuse the local second stylized feature and the global first stylized feature. The adaptive fusion process includes: firstly, the global first stylized feature and the second stylized feature of localized regions The concatenation is performed in the channel dimension, followed by an average pooling operation avg(.) and three connected convolution operations CONVS(·) for coordination and fusion.

[0135] In step S1270, a Hunan embroidery style image corresponding to the first image is generated based on the overall stylized features.

[0136] In some embodiments, this can be implemented through a common encoder-decoder structure.

[0137] In some embodiments, generating a Hunan embroidery style image corresponding to the first image based on the overall stylized features includes:

[0138] Inputting the overall stylized features into a lightweight reversible network to generate a Hunan embroidery style image corresponding to the first image through a reverse process of the lightweight reversible network;

[0139] The reverse process of the lightweight reversible network includes:

[0140] Y k =Concat(Y k [c+1:C],Y k [1:c]);

[0141] Y k [c+1:C]=Y k+1 [1:c]-H(Y k+1 [1:c]);

[0142] Y k [1:c]=Y k+1 [1:c];

[0143] The last reversible unit of the reverse process outputs a Hunan embroidery style image corresponding to the first image.

[0144] This embodiment not only uses the forward process of a lightweight reversible network for feature extraction, but also uses the reverse process of a lightweight reversible network for image reconstruction. The fully reversible design reduces the content and style leakage problems in the entire style transfer process.

[0145] After local style transfer, the style characteristics of each region will be more obvious. When there are large differences in the semantics of two adjacent regions, the stitching, color, etc. of the adjacent regions will not be very coherent. Therefore, the local style transfer result (i.e., the second stylized feature) will produce visually incoherent edges between adjacent regions containing different object categories, affecting the overall visual effect. In order to coordinate the consistency of the styles of each region, in some embodiments, the loss function of the style transfer model is the content loss function Style loss function First consistency loss function and the second consistency loss function The weighted sum of

[0146]

[0147]

[0148] Where N is the total number of lightweight reversible networks, For a lightweight reversible network, from the first image I c Corresponding Xiang embroidery style image I cs The i-th layer features extracted from For a lightweight reversible network, from the first image I c The i-th layer features extracted from For a lightweight reversible network, from the second image I s The i-th layer features extracted from Characterized by The mean of Characterized by The mean of Characterized by The variance of Characterized by The variance, ||·|| is the norm, ||·||2 is the 2-norm, I cc The second image I s and the first image I c The first image I is used c As the input generated Xiang embroidery style image, I ss The second image I s and the first image I c The second image I is used s As input, the generated Xiang embroidery style image is For lightweight reversible networks from I cc The i-th layer features extracted from For lightweight reversible networks from I ss The i-th layer features extracted from .

[0149] Setting the content loss function The purpose is to preserve the content of the first image in the generated Xiang embroidery style image and use content loss to shorten the distance between the generated Xiang embroidery style image and the second image features.

[0150] Setting the style loss function The purpose is to enable the style transfer model to learn the ability to transfer style.

[0151] Set the first consistency loss function and the second consistency loss function The purpose is to enhance the model's ability to extract features, so that the content and style features extracted by the model from the same image can be used to reconstruct the original image. At this time, the generated Xiang embroidery style image should be as consistent as possible with the original image (the first image).

[0152] like Figure 5One embodiment of the present application provides a method for transferring the style of a digital Xiang embroidery image based on regional semantics, the method comprising the following steps:

[0153] Step S910: Construct a style transfer model. The style transfer model includes: a first semantic segmentation network, a second semantic segmentation network, a CLIP network, a transfer network, an adaptive fusion network, and a lightweight reversible network.

[0154] Step S920: extract the first image I c The first semantic segmentation region and the second image I s The first image is the content image used for training, and the second image is the digital Xiang embroidery image used for training.

[0155] The first image I c Input into the first semantic segmentation network to obtain the semantic region mask and the label corresponding to each region:

[0156] m c ={m c1 , m c2 ,...,m ci ,...,m cn};

[0157] CL={C 1L , C 2L ,...,C iL ,...,C nL};

[0158] Among them, cn is the total segmentation area, m ci represents the mask of the i-th first semantic segmentation region, and iL represents the semantic category corresponding to the i-th first semantic segmentation region.

[0159] The second image I s Input into the second semantic segmentation network (ie SAM) to obtain the second semantic segmentation region mask:

[0160] m s ={m s1 , m s2 ,...,m sj ,...,m so};

[0161] so is the total number of partitioned regions of SAM.

[0162] After obtaining the mask of each region, the mask and the image are multiplied pixel by pixel to obtain the segmented region corresponding to the image.

[0163] The first semantic segmentation region set can be expressed as:

[0164] C={C1,C2,...,C i ,...,C n};

[0165] Where n is the total number of segmented regions. i The formula is:

[0166] C i =m ci ⊙I c ;

[0167] The second semantic segmentation region set can be expressed as:

[0168] S={S1,S2,…,S j ,…,S m};

[0169] Where m is the total number of segmented areas. The formula is:

[0170] S j =m sj ⊙I s ;

[0171] Step S930: Input the second semantic segmentation region set S obtained by the second semantic segmentation network and all the semantic labels CL of the first semantic segmentation region set C into the CLIP network. The semantic label with the highest probability output by the CLIP network is the final matching result.

[0172] The specific formula is:

[0173]

[0174] So far, we have found the second semantic segmentation area S j The label C corresponding to the semantically closest first semantic segmentation region match(j)L , so we can get C match(j)L The corresponding first semantic segmentation area C match(j) , and finally we can get the semantic matching pair P j =(C match(j) , S j ).

[0175] For each second semantic segmentation region S in the second semantic segmentation region set S j Repeat the above process to obtain the set P of all semantic matching pairs.

[0176] Step S940: The first image I c and the second image I s Input into a lightweight reversible network.

[0177] The lightweight reversible network consists of designed reversible units, which can losslessly extract image features through the radial changes of the forward process of the reversible units.

[0178] Specific operations include:

[0179] Y k+1 [c+1:C]=Y k [c+1:C]+H(Y k [1:c]);

[0180] Y k+1 [1:c]=Y k [1:c];

[0181] Y k+1 =Concat(Y k+1 [c+1:C],Y k+1 [1:c]);

[0182] The explanation of the relevant characters of the formula has been introduced in the above embodiment and will not be described in detail here;

[0183] The extraction capability of reversible units is largely related to the extraction capability of H. If the H function is too simple, it will affect the feature extraction capability of the model, and a complex mapping function will also increase the computational complexity of the model. Therefore, while considering both the extraction capability and the computational complexity, the component module of the lightweight network SqueezeNet is used as H.

[0184] During each forward propagation, the features of the [1:c] channel are not extracted, which will seriously affect the model's extraction capabilities. Therefore, before each feature is input into the lightweight reversible network, all channels are randomly shuffled. At the same time, the value of c is set to a different number for each module, further improving the lightweight reversible network's ability to learn information from different channels, thereby improving the network's extraction capabilities. Subsequently, all features of [1:C] are spliced together and fed into a reversible unit. The extraction process of the lightweight reversible network can be expressed as:

[0185] Repeat the above process until the last reversible unit outputs the required first image feature F c and the second image feature F of the second image s .

[0186] In step S930 , the transfer network is used to perform style transfer between semantic regions and global style transfer.

[0187] Biased transfer networks can lead to content leakage problems, because the unbiased AdaIN transfer network is used for style transfer during the transfer process. Therefore, let the first stylized feature be The calculation formula can be expressed as:

[0188]

[0189] For the semantic matching pair P j =(C match(j) , S j ), the second stylized feature of its region It can be expressed as:

[0190]

[0191] Repeat the above operation to get the second stylized feature

[0192] Step S960: Adaptively fuse the first stylized features and all second stylized features Right now

[0193]

[0194] Step S970: transform the overall stylized feature F cs In the reverse process of the lightweight reversible network fed into it.

[0195] This process is the inverse transformation of the affine transformation in the forward propagation process, and the entire process is completely reversible.

[0196] Its operation can be expressed as:

[0197] Y k [c+1:C]=Y k+1 [1:c]-H(Y k+1 [1:c]);

[0198] Y k [1:c]=Y k+1 [1:c];

[0199] Y k =Concat(Y k [c+1:C],Y k [1:c]);

[0200] In the lightweight reversible units of the model stack, the above affine transformation inverse process is repeated until the last reversible unit can output the final Xiang embroidery style image I cs .

[0201] During the training process, it is necessary to calculate the gradient of the loss function and use the backpropagation algorithm to update the parameters in the network.

[0202]

[0203] To enable the model to learn the ability to transfer style, style loss is used:

[0204]

[0205] In order to enhance the model's ability to extract features, consistency loss is applied to the first and second images respectively, as follows:

[0206]

[0207] Step S980: Model application.

[0208] For example, the target content image and the target digital Xiang embroidery image are input into a preset style transfer model, and the style transfer model outputs a Xiang embroidery style image corresponding to the target content image.

[0209] This method first obtains semantically segmented regions of the content image with semantic labels. It then uses the SAM algorithm, which has strong zero-shot segmentation capabilities, to segment the digital embroidery image, obtaining semantically segmented regions of the unlabeled digital embroidery image. Subsequently, the CLIP model is used to match the semantically segmented regions of the content image with those of the digital embroidery image based on the cosine distance between features, resulting in region-based semantic matching pairs. Simultaneously, a lightweight reversible network is used to extract features of the content image and the digital embroidery image using the forward pass of the reversible network. Subsequently, guided by the mask image, style transfer and fusion are performed between the matching semantically segmented regions of the content image and the digital embroidery image in the semantic matching pairs, obtaining local stylized features. In this way, based on the semantic information of the semantically segmented regions of the content image, the most appropriate stitches, colors, and other features are found in the digital embroidery image to achieve style transfer for different regions, which is highly consistent with the creative process of embroidery. Next, the extracted content image features are combined with the features of the digital Xiang embroidery image to perform global style transfer. The resulting global stylized features are then adaptively fused with the local stylized features to achieve consistency in the styles of each region. Finally, the fused overall stylized features are transformed into the final Xiang embroidery-style image through a lightweight reversible network's inverse process, which is completely reversible from the forward process. This enables lossless style transfer in conjunction with the reversible network's forward process, addressing the content leakage issue.

[0210] like Figure 6 As shown, one embodiment of the present application provides a style transfer device for digital Hunan embroidery images based on regional semantics, the device comprising:

[0211] The data acquisition module 1100 is used to acquire the target content image and the target digital Xiang embroidery image;

[0212] The style transfer module 1200 is used to input the target content image and the target digital Xiang embroidery image into a preset style transfer model, and obtain the Xiang embroidery style image corresponding to the target content image output by the style transfer model. The training process of the style transfer model includes:

[0213] Extracting a plurality of first semantic segmentation regions from a first image and extracting a plurality of second semantic segmentation regions from a second image; the first image is a content image for training, and the second image is a digital Xiang embroidery image for training;

[0214] Perform semantic matching on the plurality of first semantic segmentation regions and the plurality of second semantic segmentation regions to obtain a plurality of semantic matching pairs; any semantic matching pair includes: a first semantic segmentation region and the second semantic segmentation region closest thereto;

[0215] extracting a first image feature of the first image and extracting a second image feature of the second image;

[0216] Performing global style transfer based on the first image feature and the second image feature to obtain a first stylized feature;

[0217] Calculate the first pixel dot product result of the first semantic segmentation region mask and the first image feature corresponding to the first semantic segmentation region in each semantic matching pair, and the second pixel dot product result of the second semantic segmentation region mask and the second image feature corresponding to the second semantic segmentation region, and perform regional style transfer based on the first pixel dot product result and the second pixel dot product result to obtain the second stylized feature corresponding to each semantic matching pair;

[0218] Fusing the first stylized feature and all the second stylized features to obtain an overall stylized feature;

[0219] A Hunan embroidery style image corresponding to the first image is generated according to the overall stylized features.

[0220] It should be noted that the present device embodiment and the above method embodiment are based on the same inventive concept, so the content of the above method embodiment is also applicable to the present device embodiment and will not be repeated here.

[0221] like Figure 7 , an embodiment of the present application further provides an electronic device, the electronic device comprising:

[0222] at least one memory;

[0223] at least one processor;

[0224] at least one program;

[0225] The programs are stored in the memory, and the processor executes at least one program to implement the above-mentioned regional semantics-based style transfer method for digital Hunan embroidery images in the present disclosure.

[0226] The electronic device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a car computer, etc.

[0227] The electronic device according to the embodiment of the present application is described in detail below.

[0228] The processor 1600 may be implemented as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention.

[0229] Memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). Memory 1700 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in memory 1700 and is called by processor 1600 to execute the regional semantics-based digital Xiang embroidery image style transfer method of the embodiment of the present invention.

[0230] Input / output interface 1800, used for information input and output;

[0231] Communication interface 1900, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0232] bus 2000 , which transmits information between various components of the device (e.g., processor 1600 , memory 1700 , input / output interface 1800 , and communication interface 1900 );

[0233] The processor 1600 , the memory 1700 , the input / output interface 1800 , and the communication interface 1900 are connected to each other in communication within the device via the bus 2000 .

[0234] An embodiment of the present invention further provides a storage medium, which is a computer-readable storage medium and stores computer-executable instructions. The computer-executable instructions are used to enable a computer to execute the above-mentioned style transfer method of digital Hunan embroidery images based on regional semantics.

[0235] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0236] The embodiments described in the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.

[0237] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0238] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0239] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0240] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0241] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0242] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0243] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0244] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling an electronic device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0245] The above is a specific description of the preferred implementation of the embodiments of the present application, but the embodiments of the present application are not limited to the above-mentioned implementation methods. Technical personnel familiar with the art can also make various equivalent modifications or substitutions without violating the spirit of the embodiments of the present application. These equivalent modifications or substitutions are all included in the scope defined by the claims of the embodiments of the present application.

Claims

1. A style transfer method for digital Xiang embroidery images based on regional semantics, characterized by: The method comprises: Acquire a target content image and a target digital Xiang embroidery image; The target content image and the target digital Xiang embroidery image are input into a preset style transfer model, and the style transfer model outputs a Xiang embroidery style image corresponding to the target content image; the training process of the style transfer model includes: Extracting a plurality of first semantic segmentation regions from a first image and extracting a plurality of second semantic segmentation regions from a second image; the first image is a content image for training, and the second image is a digital Xiang embroidery image for training; Perform semantic matching on the plurality of first semantic segmentation regions and the plurality of second semantic segmentation regions to obtain a plurality of semantic matching pairs; any one of the semantic matching pairs includes: one first semantic segmentation region and the second semantic segmentation region closest thereto; extracting a first image feature of the first image and extracting a second image feature of the second image; Performing global style transfer based on the first image feature and the second image feature to obtain a first stylized feature; Calculating a first pixel dot product result between a first semantic segmentation region mask corresponding to a first semantic segmentation region in each of the semantic matching pairs and the first image feature, and a second pixel dot product result between a second semantic segmentation region mask corresponding to a second semantic segmentation region and the second image feature, and performing regional style transfer based on the first pixel dot product result and the second pixel dot product result to obtain a second stylized feature corresponding to each of the semantic matching pairs; fusing the first stylized feature and all the second stylized features to obtain an overall stylized feature; A Hunan embroidery style image corresponding to the first image is generated according to the overall stylized features.

2. The style transfer method of digital Hunan embroidery images based on regional semantics according to claim 1 is characterized in that: The extracting a first image feature of the first image and extracting a second image feature of the second image includes: Inputting the first image into a lightweight reversible network to extract first image features of the first image through a forward process of the lightweight reversible network; Inputting the second image into the lightweight reversible network to extract second image features of the second image through a forward process of the lightweight reversible network; The lightweight reversible network includes a plurality of stacked reversible units, and the forward process of the lightweight reversible network includes: ; ; ; in, For the The output of the reversible unit, is the channel splicing function, is the 1st to cth layer channel of any reversible unit, is the cth to Cth layer channel of any reversible unit, For the Reversible units in The output, For the Reversible units in The output, is the mapping function.

3. The style transfer method of digital Xiang embroidery images based on regional semantics according to claim 2 is characterized in that: Generating a Hunan embroidery style image corresponding to the first image according to the overall stylized features includes: Inputting the overall stylized features into the lightweight reversible network to generate a Hunan embroidery style image corresponding to the first image through a reverse process of the lightweight reversible network; The reverse process of the lightweight reversible network includes: ; ; ; The last reversible unit in the reverse process outputs a Hunan embroidery style image corresponding to the first image.

4. The method for style transfer of digital Hunan embroidery images based on regional semantics according to claim 3 is characterized in that: The loss function of the style transfer model is the content loss function , style loss function , the first consistency loss function and the second consistency loss function The weighted sum of ; ; ; ; in, is the total number of lightweight reversible networks, For a lightweight reversible network from the first image Corresponding Xiang embroidery style images The extracted Layer features, For a lightweight reversible network from the first image The extracted Layer features, For a lightweight reversible network from the second image The extracted Layer features, Characterized by The mean of Characterized by The mean of Characterized by The variance of Characterized by The variance of is the norm, is the 2-norm, For the second image and the first image Use the first image As input, the generated Xiang embroidery style image is For the second image and the first image Use the second image As input, the generated Xiang embroidery style image is For lightweight reversible network The extracted Layer features, For lightweight reversible network The extracted Layer features.

5. The method for style transfer of digital Hunan embroidery images based on regional semantics according to claim 1 is characterized in that: The extracting of a plurality of first semantic segmentation regions of the first image and a plurality of second semantic segmentation regions of the second image comprises: Extracting a plurality of first semantic segmentation region masks of the first image according to the first semantic segmentation network; Extracting a plurality of second semantic segmentation region masks of the second image through a second semantic segmentation network; wherein the second semantic segmentation network is a deep learning SAM network; Performing pixel-by-pixel multiplication of the first semantic segmentation region masks with the first image to obtain a plurality of first semantic segmentation regions; Perform pixel-by-pixel multiplication on the second image and the plurality of second semantic segmentation region masks to obtain a plurality of second semantic segmentation regions.

6. The method for style transfer of digital Hunan embroidery images based on regional semantics according to claim 5 is characterized in that: The process of calculating the first stylized feature includes: ; in, is the AdaIN migration function, is the first image feature, is the second image feature, It is the first stylized feature.

7. The method for style transfer of digital Hunan embroidery images based on regional semantics according to claim 5 is characterized in that: The process of calculating the second stylized feature includes: ; in, is the AdaIN migration function, for and The first pixel multiplication result is for Downsampling, is the first semantic segmentation region mask , is the first image feature, for and The second pixel multiplication result is for Downsampling, is the second semantic segmentation region mask , is the second image feature, The corresponding first semantic segmentation area and The corresponding second semantic segmentation area forms a pair of semantic matching pairs ; Semantic matching pair The corresponding second stylized feature.

8. The method for style transfer of digital Hunan embroidery images based on regional semantics according to claim 5 is characterized in that: The semantically matching the plurality of first semantic segmentation regions with the plurality of second semantic segmentation regions to obtain a plurality of semantic matching pairs includes: Inputting the first semantic labels corresponding to the plurality of second semantic segmentation regions and the plurality of first semantic segmentation regions into a multimodal CLIP network to obtain a plurality of matching results output by the multimodal CLIP network, wherein the process of generating any one matching result by the multimodal CLIP network includes: ; in, is the first semantic segmentation area The corresponding first semantic label, is the second semantic segmentation area , for and The matching results, To obtain the maximum value function, It is a multimodal CLIP network function; The first semantic label in each matching result is replaced with the corresponding first semantic segmentation region to obtain a plurality of semantic matching pairs.

9. The method for style transfer of digital Hunan embroidery images based on regional semantics according to claim 5, characterized in that: The fusing of the first stylized feature and all the second stylized features to obtain an overall stylized feature includes: ; in, is the first stylized feature, To fuse all the features after the second stylized features, is the channel splicing function, is the average pooling function, is the convolution function.

10. A style transfer device for digital Hunan embroidery images based on regional semantics, characterized in that: The device comprises: A data acquisition module is used to acquire a target content image and a target digital Xiang embroidery image; The style transfer module is used to input the target content image and the target digital Hunan embroidery image into a preset style transfer model, so that the style transfer model outputs a Hunan embroidery style image corresponding to the target content image; the training process of the style transfer model includes: Extracting a plurality of first semantic segmentation regions from a first image and extracting a plurality of second semantic segmentation regions from a second image; the first image is a content image for training, and the second image is a digital Xiang embroidery image for training; Perform semantic matching on the plurality of first semantic segmentation regions and the plurality of second semantic segmentation regions to obtain a plurality of semantic matching pairs; any one of the semantic matching pairs includes: one first semantic segmentation region and the second semantic segmentation region closest thereto; extracting a first image feature of the first image and extracting a second image feature of the second image; Performing global style transfer based on the first image feature and the second image feature to obtain a first stylized feature; Calculating a first pixel dot product result between a first semantic segmentation region mask corresponding to a first semantic segmentation region in each of the semantic matching pairs and the first image feature, and a second pixel dot product result between a second semantic segmentation region mask corresponding to a second semantic segmentation region and the second image feature, and performing regional style transfer based on the first pixel dot product result and the second pixel dot product result to obtain a second stylized feature corresponding to each of the semantic matching pairs; fusing the first stylized feature and all the second stylized features to obtain an overall stylized feature; A Hunan embroidery style image corresponding to the first image is generated according to the overall stylized features.

Citation Information

Patent Citations

  • Image style migration model training method, image style migration method and device

    CN112734627A

  • Image style migration model training method, system and device and storage medium

    CN114494789A