A method and device for replacing clothing fabric
By adopting the C-TransUNet and Shared Attention Stable Diffusion model in the clothing fabric replacement technology, combined with the Attention-Enhanced Thin Plate Spline algorithm, the problems of poor fabric customization and difficult processing of complex spatial relationships in the existing technology are solved, and high-quality clothing fabric replacement and personalized customization are achieved.
Patent Information
- Application Number
- CN202410941842.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-15
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-07-15
AI Technical Summary
The existing clothing fabric replacement technology is difficult to achieve personalized customization while retaining the overall style features and local texture details. The traditional method is not effective when dealing with complex spatial relationships, resulting in unnatural occlusion and blurred boundaries in the replacement effect.
The semantic segmentation network based on C-TransUNet and the Shared Attention Stable Diffusion model are adopted, combined with the Attention-Enhanced Thin Plate Spline algorithm, the precise segmentation, personalized generation and style transfer of clothing fabrics are realized, while edge repair is carried out to improve the authenticity of the replacement effect and the quality of personalized customization.
It significantly improves the authenticity of clothing images and the quality of personalized customization, optimizes the user experience, reduces inventory and logistics costs, and improves the robustness of the system and the clarity of the image.
Smart Images

Figure CN118941799B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of computer and network technology, and in particular to a method and device for replacing clothing fabrics. Background Art
[0002] With the booming development of e-commerce and the growing demand for personalization, clothing fabric replacement technology has become an important innovation direction in the clothing industry. It not only provides a more convenient and personalized shopping experience, but also significantly reduces inventory costs and logistics pressure. Current clothing fabric replacement technology can be divided into two research directions: 2D clothing image synthesis and 3D clothing reconstruction.
[0003] In the process of replacing clothing fabrics, the occlusion relationship and boundary processing between clothing and the human body are key factors affecting the authenticity of the replacement effect. When some parts of clothing overlap with the human body, traditional two-dimensional clothing image synthesis algorithms often find it difficult to accurately handle this complex spatial relationship, resulting in unnatural occlusion and blurred boundaries in the replacement effect. Three-dimensional clothing reconstruction methods focus on the reconstruction of the overall style of clothing and lack the migration of fabric texture and shadows. It is often difficult to achieve an effect comparable to the real world, resulting in a lack of realism in the generated clothing, which limits the application scope and user experience of clothing fabric replacement technology.
[0004] Personalized customization is one of the important trends in the clothing industry. Users hope to be able to customize unique clothing according to their preferences and needs. However, the existing clothing fabric replacement technology mainly uses TPS and optical flow-based deformation algorithms in fabric customization methods, but its generation model lacks self-correction capabilities. For example, it has poor results when processing details such as flannel and threads, and the mapping traces are more obvious. It cannot meet users' high-quality requirements for fabric color, pattern or texture, which limits the personalized customization capabilities of clothing fabric replacement technology.
[0005] In summary, existing clothing fabric replacement technologies find it difficult to personalize clothing images based on customized fabrics while retaining overall style features and local texture details. Summary of the invention
[0006] In order to solve the problems existing in the prior art, the present invention proposes a method and device for replacing clothing fabrics.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] In one aspect, the present invention provides a method for replacing clothing fabrics, comprising the following steps:
[0009] S1: Perform grayscale calculation on the clothing image data in the database to obtain a grayscale image, and use the Mosaic method to perform data enhancement on the RGB image data and the grayscale image data respectively. After image preprocessing, they are input into the C-TransUNet network together, and the model clothing image and fabric image input by the user are obtained at the same time;
[0010] S2: Accurate segmentation of model clothing images is achieved by building a semantic segmentation network based on C-TransUNet. The C-TransUNet network model consists of convolutional layers, shallow feature fusion (SFF) modules, fused visual Transformer networks (FVit) and decoder modules. Convolutional layers and shallow feature fusion (SFF) modules are used to extract and fuse shallow features, and the FVit network is used to learn cross-channel deep representations with high inter-class separability and low intra-class diversity. The decoder module restores the spatial information of the input image with higher accuracy.
[0011] S3: Use the binary cross entropy loss function to optimize the parameters of the network in S2 to further improve the segmentation performance of clothes in clothing images and save the trained model weights;
[0012] S4: Load the model weights and network structure trained in S3, and input the model clothing image input by the user into the trained C-TransUNet to obtain a mask image of the clothing in the clothing image;
[0013] S5: Design a Shared Attention Stable Diffusion model to perform a diffusion operation on the input image using a diffusion operator q, and then apply a single reverse operator p θ To complete a diffusion operation, the four-way continuous generation of the user input fabric graph is realized through multiple continuous diffusion operations. The Shared Attention Stable Diffusion model is:
[0014] Project the features to the query Q∈m×d through a linear layer k , key K∈m×d k The sum V∈m×d k Then, the attention is calculated using the following formula:
[0015]
[0016] Where, d k is the dimension of Q and K. Intuitively, each image feature is updated by V weighted and updated, and its weight depends on the correlation between the projected query Q and the key K. Each self-attention layer contains multiple attention heads. By concatenating the outputs of multiple attention heads and projecting them back to the image feature space dh To calculate the residual:
[0017]
[0018] S6: Using the Attention-Enhanced Thin Plate Spline algorithm, fabric style transfer is achieved while maintaining fabric details. Based on the mask image of the clothing image, the personalized fabric input by the user is transferred to the model clothing image to achieve more realistic fabric and shadow effects. The Attention-Enhanced Thin Plate Spline algorithm is:
[0019] Approximate image style transfer by minimizing transformation energy through a unified model:
[0020] ε=ε T +λε d
[0021]
[0022] In the formula, ε represents the total energy of the expected transformation, (x rec ,y rec )and Represent the points in the source domain S and the target domain T respectively, ε T represents the data penalty energy, ε d Denotes the distortion energy, the above formula with the lowest total energy is the desired transformation, where the hyperparameter λ is designed to balance the energy between data penalty and distortion;
[0023] S7: The blurred edges of the model clothing image after migration are repaired using the SA-SD model according to the mask image of the clothing image, and the clothing model image after fabric style migration and edge repair is output.
[0024] Optionally, the specific operation steps of performing Mosaic data enhancement on the clothing image in step S1 are:
[0025] S11: Randomly read four clothing images from the clothing image dataset;
[0026] S12: performing inversion (flipping the original image left and right), scaling (scaling the size of the original image), and color gamut change (changing the brightness, saturation, and hue of the original image) operations on the four images respectively;
[0027] S13: stitching together the four images transformed in S12, with the first image placed at the upper left, the second image placed at the lower left, the third image placed at the lower right, and the fourth image placed at the upper right;
[0028] S14: After the four images are arranged, a fixed area is intercepted in a matrix manner and stitched into a new image, which retains the characteristics and distribution of the original data. During the stitching process, sometimes there will be overlapping images, that is, the image exceeds the edge between the two images (the artificially set dividing line). This part of the image will be deleted after data enhancement. The input clothing image is preprocessed through Mosaic data enhancement. The processed image will contain richer clothing background and fabric details, while making the easily detected clothing targets relatively smaller. When performing normalization calculations in the batch normalization layer (BN), the data of the four images can be calculated simultaneously, thereby improving the robustness of the algorithm.
[0029] Optionally, the convolutional layer and shallow feature fusion (SFF) module in the semantic segmentation network based on C-TransUNet described in step S2 is:
[0030] Use X∈R H×W×3 and Y∈R H×W×1 Represents an RGB image and its corresponding grayscale image data (DSM image), where H and W are the height and width of the input, the RGB image is three channels, and the DSM image data is one channel. The proposed C-TransUNet adopts a dual-branch architecture and designs a dual-branch encoder. First, one branch is used to extract multi-scale features from the DSM modality. Each branch of the encoder consists of four convolutional layers. Among them, the size is The downsampled feature map is generated by the i-th encoder layer, where i is the layer index of the CNN encoder. Next, the shallow features of the grayscale modality extracted by the convolution operation are fused into the features of the main modality (i.e., RGB) using the SFF module, and the fused features are input before the next RGB image encoder branch. Specifically, for the RGB and DSM shallow features input to the i-th SFF module, the input channel size is C i , feature compression is performed through two global average pooling operations with a kernel size of 1×1 to aggregate global information, and ReLU and Sigmoid functions are used for activation, and then the RGB and DSM features are weighted and element-wise added to generate the final fused shallow features. Finally, the output of the SFF module is directly fed into the corresponding decoder layer by utilizing skip connections, which is designed to recover detailed local and contextual information.
[0031] Optionally, the FVit network in the C-TransUNet-based semantic segmentation network described in step S2 is:
[0032] For a given x I and I Respectively indicate the size RGB and DSM feature maps, where I and C I Represents the layer index and output channel size of the last layer in the CNN backbone network. First, two linear layers and a reshape operation are used to transform x I and I Tokenized. Specifically, the linear layer converts the input channel size from C I Change to C hid , the reshape operation flattens the output of the linear layer into two two-dimensional sequences, denoted as and Size C hid ×L, where is the sequence length; specific position codes are then added to and to retain the position information and input it into FVit.
[0033] The input of the FVit encoder passes through three stages in sequence, including the first stage dynamic attention layer (D-SA) for deep feature enhancement, the second stage adaptive cross fusion attention layer (Ada-CFA) for deep feature fusion, and the third stage D-SA layer for fusion feature enhancement, with 3, 6, and 3 layers respectively. and represents the hidden features of the nth layer in the RGB branch and the DSM branch, where n∈(1,2,…,12). It is worth noting that this process preserves the dimension of the feature map as C throughout FVit. hid ×L. Specifically, the D-SA layer consists of two dynamic attention modules (D-SA), two multi-layer perceptron modules (MLP), and a layer normalization (LN) layer. and Given the multi-modal feature input represented by , the D-SA layer is designed to derive the global relationship of each modality using a multi-head dynamic attention mechanism.
[0034] After performing deep feature enhancement in the first-stage D-SA layer, FVit further uses the second-stage Ada-CFA layer to fuse multimodal features in an abstract semantic space with rich contextual information. In this deep feature fusion stage, cross-attention (CA) and self-attention (SA) are simultaneously calculated in the Ada-CFA module to learn the correlation between the primary modality RGB and the auxiliary modality DSM.
[0035] Finally, the fused feature map is enhanced by the D-SA layer in the third stage. The specific calculation process is the same as in the first stage to enhance the fused feature maps of the RGB branch and the DSM branch respectively. The final output of FVit is expressed as is the feature map derived from the last D-SA layer. Based on the proposed FVit, rich contextual information extracted from multimodal data is deeply fused before being fed into the cascaded decoder.
[0036] Optionally, the decoder in the semantic segmentation network based on C-TransUNet in step S2 is:
[0037] The decoder recovers the hidden fusion features of the final segmentation process by utilizing multiple upsampling modules. The decoder first uses a reconstruction module to transform the 2-D input sequence z N Reshape to size A 3-D tensor, where C dec is the number of channels of the first block in the input decoder. Afterwards, multiple cascaded decoder blocks restore the spatial resolution to H×W by connecting the skip connections from the corresponding CNN backbone layers. Each decoder block consists of an upsampling operator, a convolutional (Conv) layer, and a ReLU layer. Finally, the segmentation head performs the final semantic prediction.
[0038] Optionally, the binary cross entropy loss function calculation formula in step S3 is:
[0039]
[0040] In the formula, M represents the number of classification categories, y ic is a sign function. When the true category of sample i is equal to category c, it takes 1, otherwise it takes 0. ic is the predicted probability value that the observed sample i belongs to category c; the total loss of the C-TransUNet network model is calculated by using the binary cross entropy loss function, and the parameters of the network model are updated and optimized using the back propagation algorithm and gradient descent algorithm. The purpose of network training is to minimize the total network loss function.
[0041] Optionally, in step S5, the Shared Attention Stable Diffusion (SA-SD) model is:
[0042] The existing Wensheng graph diffusion model mainly adopts the U-Net architecture, which consists of convolutional layers and transformer attention blocks; deep image features These self-attention layers attend to each other and the cross-attention layers attend to the contextual text embeddings; the self-attention layers are modified so that the deep features are updated through mutual self-attention. First, the features are projected to the query Q∈m×d through a linear layer k , key K∈m×d k The sum V∈m×d k Then, the attention is calculated using the following formula:
[0043]
[0044] Where, d k is the dimension of Q and K. Intuitively, each image feature is updated by the weighted sum of V, where the weight depends on the correlation between the projected query Q and the key K. In practice, each self-attention layer contains multiple attention heads, and then the outputs of multiple attention heads are concatenated and projected back to the image feature space d h To calculate the residual:
[0045]
[0046] The goal of the method of the invention is to generate a set of images which is aligned with a set of input text hints y 1 ,y 2 …,y n And have a consistent style interpretation with each other, that is, they are consistent with each other and with the style of the input text on top. The traditional method of generating style-aligned image sets with different contents is to use the same style description in the text prompt. The method of the present invention is achieved by sharing attention layers between generated images and allowing them to communicate with each other. However, the present invention notes that by enabling full attention sharing, the quality of the generated set may be impaired, resulting in content leakage between images. In order to limit content leakage and allow sharing of different sets, the deep image features are re-defined to only focus on one image in the generated set (usually the first one in the batch). That is, the target image features Only focuses on its own features and the features of one reference image in the set. However, since focusing on only one image in the set will cause less attention flow from the reference image to the target image, the styles of different images are less consistent.
[0047] In order to achieve a balance of attention reference, the present invention designs an adaptive normalization operation (ADAIN) using the query Q of the reference image r and keyword K r To normalize the query Q of the target image t and keyword K t , thereby calculating the corresponding shared attention to address the above challenges.
[0048] Optionally, the Attention-Enhanced Thin Plate Spline algorithm in step S6 is:
[0049] Approximate image style transfer by minimizing transformation energy through a unified model:
[0050] ε=ε T +λε d
[0051]
[0052] Where ε represents the total energy of the expected transformation. rec ,y rec )and Represent the points in the source domain S and the target domain T respectively, ε T represents the data penalty energy, ε d Denotes the distortion energy. The above formula with the lowest total energy is the desired transformation, where the hyperparameter λ is designed to balance the energy between data penalty and distortion.
[0053] Since the TPS transformation allows one image to be transferred to another with minimal distortion given two sets of control points in the corresponding images, and provides a more complex, flexible and nonlinear representation, it has the problem of easily ignoring image details. To this end, the present invention proposes to use the attention-enhanced TPS transformation (A-TPS) to achieve end-to-end unsupervised high-quality fabric style transfer.
[0054] A-TPS brings additional flexibility to TPS by incorporating attention scores, producing more natural corrections by content-aware and adaptively weighting control points when performing the transformation. Since correction and recognition are jointly optimized, this flexibility can better guide parameter updates to achieve high-quality style transfer.
[0055] Optionally, the method for edge repairing using Shared Attention Stable Diffusion in step S7 is:
[0056] S71: inputting a clothing model image and a clothing mask image for which fabric migration has been completed but needs to be repaired;
[0057] S72: extracting edge regions in the image by transparency superposition and generating a preliminary edge mask;
[0058] S73: Generate area masks that need to be repaired according to the edge mask, ensuring that these areas include edges and surrounding parts that need to be smoothed and repaired;
[0059] S74: using the clothing model image after fabric migration and the repaired region mask as input data of the SA-SD model;
[0060] S75: inputting the input data into the SA-SD model to perform edge repair to generate a repaired image, thereby filling and smoothing the edge area so that it transitions naturally with the surrounding area;
[0061] S76: extracting the repaired image, combining the repaired edge area with the original image by using transparency overlay, and forming a final repaired image.
[0062] On the other hand, the present invention also provides a clothing fabric replacement device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor runs the computer program, the steps of the above-mentioned clothing fabric method based on Shared Attention Stable Diffusion and C-TransUNet network are implemented.
[0063] Compared with the prior art, the present invention has the following beneficial effects:
[0064] The grayscale image of the clothing image data in the database is calculated and Mosaic data enhancement is performed; then the enhanced RGB image data and the grayscale image data are input into a U-shaped network (C-TransUNet) based on convolution and Transformer to learn and fuse multi-scale shallow and deep features; the binary cross entropy loss function is used to update and optimize the model parameters in real time, so that the model can achieve accurate segmentation of local texture details and global semantics; again, the model clothing image input by the user is input into the trained C-TransUNet to obtain the mask of the clothing image; next, the personalized fabric image input by the user is generated in four directions continuously through the redesigned Shared Attention Stable Diffusion (SA-SD) to maintain consistency with the model clothing input by the user. Figure 1 The size of the model clothing image is consistent; then, the Attention-Enhanced Thin Plate Spline (A-TPS) algorithm is used to achieve fabric style transfer while maintaining fabric details, so as to achieve more realistic fabric and shadow effects on the model clothing image; finally, SA-SD is used to repair the edges of the migrated model clothing image according to the mask image of the clothing image to reduce the problems of edge blur and occlusion.
[0065] The present invention significantly improves the realism of clothing images and the quality of personalized customization through an innovative clothing fabric replacement method, while optimizing the user experience and reducing inventory and logistics costs. It also improves the robustness of the system and the clarity of the image through efficient image segmentation and edge repair technology, thereby bringing cost-effectiveness and market competitiveness to the clothing industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is a logic flow chart of the present invention;
[0067] Figure 2It is the network structure diagram of C-TransUNet of the present invention;
[0068] Figure 3 It is a structural diagram of the SSF module in the C-TransUNet network of the present invention;
[0069] Figure 4 It is a structural diagram of the D-SA module in the C-TransUNet network of the present invention;
[0070] Figure 5 It is a structural diagram of the Ada-CFA module in the C-TransUNet network of the present invention;
[0071] Figure 6 It is the dynamic self-attention result graph of the present invention;
[0072] Figure 7 is the adaptive cross-fusion attention structure diagram of the present invention;
[0073] Figure 8 It is the overall framework flow chart of the present invention. DETAILED DESCRIPTION
[0074] The technical solutions in the embodiments of the present invention will be described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0075] like Figure 1-8 As shown, a method for replacing clothing fabrics provided in an embodiment includes the following steps:
[0076] (1) First, obtain the clothing image data and its corresponding true labels from the database and divide them into training set, validation set and test set according to 8:1:1; then calculate the grayscale value of the clothing image data and convert it from RGB image to grayscale image. The calculation formula is:
[0077]
[0078] Where R, G, and B represent the pixel values of the three color channels of the image respectively.
[0079] Next, the grayscale image and RGB image are enhanced using the Mosaic method. Four images are randomly selected from the dataset and are denoted as I 1 ,I 2 ,I 3 ,I 4 , select a center point (cx, cy) in the new image, which divides the four images into four quadrants, scale each image to fit the size of the new image, and stitch the four images into a new image I according to the center point (cx, cy) and the quadrant positionmosaic , thereby achieving data enhancement, and increasing the diversity of training data by mixing different images using the Mosaic method to improve the generalization ability of the model.
[0080] (2) The user manually uploads the model clothing image and the corresponding fabric image that need to be replaced in the system. After entering the system, the grayscale image of the model clothing image will be automatically calculated, and the model clothing image, the grayscale image of the model clothing image and the fabric image will be saved in the system for subsequent operations.
[0081] (3) The data-enhanced clothing images and grayscale images are input into the carefully designed C-TransUNet network, in which the convolutional layer, shallow feature fusion (SFF) module, FVit network and decoder module are introduced to locate and learn the target features.
[0082] First, define X∈R H×W×3 and Y∈R H×W×1 Respectively represent the RGB image and its corresponding grayscale image data (DSM image), where H and W are the height and width of the input, the RGB image is a color image with three channels, and the DSM image data is a grayscale image with one channel. The C-TransUNet proposed in the present invention adopts a dual-branch architecture and carefully designs a dual-branch encoder for image inputs of different modalities. First, one branch is used to extract multi-scale features from each modality. Specifically, each branch of the encoder consists of four convolutional layers for multi-scale feature extraction, where the size of the downsampled feature map generated by the i-th encoder layer is i is the layer index of the CNN encoder. The shallow features extracted by the convolution operation are then input to the SFF module for feature fusion. Notably, the features derived from the grayscale image modality are fused into the features from the main modality (i.e., RGB) before being input to the next RGB image encoder branch. In addition, the output of the SFF module is directly fed into the corresponding decoder layer by utilizing skip connections, which aims to recover detailed local and contextual information.
[0083] The SFF module first performs global average pooling (Global AvgPool) on the RGB and DSM branches to aggregate global information. Specifically, for a given i-th SFF module, its input channel size is C i , feature compression is performed through two global average pooling operations with a kernel size of 1×1, and then ReLU and Sigmoid functions are used for activation. Finally, the RGB and DSM features are weighted and added element by element to generate the final fused shallow features.
[0084] For a given x I andI Respectively indicate the size RGB and DSM feature maps, where I and C I Represents the layer index and output channel size of the last layer in the CNN backbone network. First, two linear layers and a reshape operation are used to transform x I and I The linear layer transforms the input channel size from C I Change to C hid , and then a reshape operation is performed to flatten the output of the linear layer into two two-dimensional sequences, respectively. and Indicates that the size is C hid ×L, where is the sequence length. Add the specific position code to and to retain the position information. After that, the sequence with the position information added and Enter into FVit.
[0085] The input of the FVit encoder goes through three stages in sequence, including the first stage dynamic attention layer (D-SA) for deep feature enhancement, the second stage adaptive cross fusion attention layer (Ada-CFA) for deep feature fusion, and the third stage D-SA layer for fusion feature enhancement, with 3, 6, and 3 layers respectively. and represents the hidden features of the nth layer in the RGB branch and the DSM branch, where n∈(1,2,…,12). It is worth noting that this process preserves the dimension of the feature map as C throughout FVit. hid ×L. Specifically, the D-SA layer consists of two dynamic attention modules (D-SA), two multi-layer perceptron modules (MLP), and a layer normalization (LN) layer. and The D-SA layer is designed to use a multi-head dynamic attention mechanism to derive the global relationship of each modality for multi-modal feature input. Mathematically, the output of the nth D-SA layer when n=1,2,3 can be written as follows:
[0086]
[0087] In the formula, D-SA() represents dynamic attention calculation. Specifically, taking the calculation of RGB features on a single attention head as an example, first, the input feature Projection is the initial tensor query Q dsa , key K dsa Sum value V dsa :
[0088]
[0089] In the formula, is the linear mapping weight matrix. Then, the attention maps of multiple attention heads are summed to obtain the global attention A:
[0090]
[0091] In the formula, each value in the global attention A is A[i,j], which represents the total attention of query token i to key token j. Next, the matrix of global attention A is derived as the adjacency matrix A of the region-to-region affinity graph r . Adjacency matrix A r The value in represents the degree to which two regions are semantically related. Then, the adjacency matrix is pruned by using the row-by-row topk operator to retain only the top k connections of each region as the affinity graph to derive the routing index matrix I r :
[0092] I r =topkindex(A r )
[0093] In the formula, I r The i-th row of contains the k indices of the most relevant regions of the i-th region, and the k value will dynamically float with the size of the feature map to ensure that the routing index object contains only the most critical dynamic information. Next, by collecting and concentrating all key and value tensors, the attention of each query token in region i resides on the i-th region. r (i,1),I r (i,2),…,I r In the k routing areas indexed by (i,k):
[0094]
[0095] In the formula is the aggregated key-value tensor, so we can focus our attention on the collected key-value pairs. Finally, the aggregated key-value tensor is used to calculate the dynamic attention map feature F sad :
[0096]
[0097] Where, d sad Represents the normalization parameter.
[0098] After the first-stage D-SA layer performs deep feature enhancement, FVit further uses the second-stage Ada-CFA layer to fuse multimodal features in the abstract semantic space with rich contextual information. In this deep feature fusion stage, cross attention (CA) and self-attention (SA) are simultaneously calculated in the Ada-CFA module to learn the correlation between the main modality RGB and the auxiliary modality DSM. The output of the Ada-CFA layer can be written as:
[0099]
[0100] set up and Then the fusion output of the n∈(4,5,…,9)th layer can be written as:
[0101]
[0102] With the support of multi-head design, the proposed Ada-CFA module adopts multi-head design to input its multimodal features. and Divide into H equal segments, using and Denote h = 1, 2, ..., H, where H is the number of heads. Since the operation of each head is the same, the index h will be omitted when discussing the fusion attention module in the multi-head mechanism below. Next, use linear projection to and To calculate the two sets of matrices {Q x ,K x ,V x} and {Q y ,K y ,V y}. Then, derive D-SA information for both modalities (e.g. dsa x ,dsa y ) and CA information (e.g. ca x ,ca y ). More specifically, D-SA uses {Q x ,K x ,V x} and {Q y ,K y ,V y} calculates the information within the pattern, while CA uses {Q x ,K x ,V x} and {Q y ,K y ,V y}Calculate the information between patterns. This process realizes the extraction and fusion of deep features through a module. The process in the Ada-CFA module can be expressed as:
[0103]
[0104] Where d is the normalization parameter, and and(·) T They are the Softmax function and the matrix transposition operator. From the above formula, we can find that the interrelated guidance mechanism uses the method of exchanging some matrices to achieve the guidance fusion of the two modes. Next, the following adaptive mechanism is proposed to fuse D-SA and CA:
[0105]
[0106] In the formula, and are learnable weighting coefficients used to balance the contributions of D-SA and CA, respectively.
[0107] Subsequently, the fused feature map is enhanced by the D-SA layer in the third stage. The specific calculation process is the same as in the first stage to enhance the fused feature maps of the RGB branch and the DSM branch respectively. The final output of FVit is expressed as is the feature map derived from the last D-SA layer. Based on the proposed FVit, rich contextual information extracted from multimodal data is deeply fused before being fed into the cascaded decoder.
[0108] Finally, the decoder recovers the hidden fusion features of the final segmentation process by utilizing multiple upsampling modules. More specifically, the decoder first uses a reconstruction module to transform the 2-D input sequence z N Reshape to size A 3-D tensor, where C dec is the first block in the decoder of the input channel number. After that, multiple cascaded decoder blocks restore the spatial resolution to H×W by connecting the skip connections from the corresponding CNN backbone layers. Each decoder block consists of an upsampling operator, a convolutional (Conv) layer, and a ReLU layer. Finally, the segmentation head performs the final semantic prediction.
[0109] (4) The C-TransUNet network is trained using the binary cross entropy loss function, which is as follows:
[0110]
[0111] In the formula, M represents the number of classification categories, y ic is a sign function. When the true category of sample i is equal to category c, it takes 1, otherwise it takes 0. icis the predicted probability value of the observed sample i belonging to category c; the total loss of the C-TransUNet network model is calculated by using the binary cross entropy loss function, and the parameters of the network model are updated and optimized using the back propagation algorithm and gradient descent algorithm;
[0112] The goal of network training is to minimize the total loss function of the network.
[0113] (5) The shared self-attention layer in the Shared Attention Stable Diffusion (SA-SD) model is:
[0114] Define the generated set of images as a set The set of input text prompt words corresponding to a group is {y 1 ,y 2 …,y n}, from the collection Deep features of the self-attention layer The query Q is obtained by projection i , key K i Sum value V i ,Then, The attention update is given by:
[0115] Attention(Q i ,K 1…n ,V 1...n )
[0116] In the formula, K 1…n =[K 1 ,K 2 ,…K n ] T ,V 1…n =[V 1 ,V 2 ,…V n ] T However, sharing weights may cause content leakage and thus damage the quality of the generated set. In order to limit content leakage and allow sharing of different sets, the target image features are set to It only focuses on the features of itself and one reference image in the set. However, since it only focuses on one image in the set, the attention flow from the reference image to the target image is small, resulting in inconsistency in the styles of different images. In order to achieve a balance of attention reference, the present invention designs an adaptive normalization operation (ADAIN) using the query Q of the reference image. r and keyword K r To normalize the query Q of the target image t and keyword K t :
[0117]
[0118] In the formula, the AdaIN operation is defined as:
[0119]
[0120] In the formula, are the mean and standard deviation of the query and key across different pixels. Finally, the shared attention is defined as:
[0121]
[0122] In the formula, V rt =[V r ,V t T . By using the shared attention calculation to replace the calculation of the attention layer in the existing StableDiffusion model, a set of images with a unified style can be better generated, which is beneficial to generating a fabric map with higher fine-grainedness
[0123] The method for realizing the transfer of clothing fabric style and detail features based on the attention-enhanced TPS transformation in (6) is as follows:
[0124] The transfer of image style is a non-linear and non-rigid mapping process, which cannot be represented by a simple linear transformation. In addition, it is challenging to establish an accurate parametric model to supervise and estimate the change process of image style. From the perspective of energy minimization, the transfer of image style can be approximated by minimizing the transformation energy through a unified model:
[0125] ε = ε T + λε d
[0126]
[0127] In the formula, ε represents the total energy of the expected transformation. (x rec ,y rec ) and represent points in the source domain S and the target domain T respectively. ε T represents the data penalty energy, and ε d represents the distortion energy. The formula with the lowest total energy above is the required transformation, where the hyperparameter λ is designed to balance the energy between data penalty and distortion
[0128] Since the TPS transformation allows one image to be transferred to another image with minimal distortion given two sets of control points in the corresponding images, and provides a more complex, flexible and nonlinear representation, the present invention proposes to use the attention-enhanced TPS transformation (A-TPS) to achieve end-to-end unsupervised fabric style transfer.
[0129] Assume Q=[Q 1 ,Q 2 ,…Q N ] and P = [P 1 ,P 2 ,…P N ] are the source and target points in the migration process, then the data item ε T It can be calculated by After aligning the control points, A-TPS finds the minimum distortion term ε by the following method. d The optimal interpolation transformation of :
[0130]
[0131] In the above formula, the second-order derivative is used to express the distortion deviation of each target point and constrain the cumulative global minimum. Then we can get the spatial deformation function parameterized by the control point as follows:
[0132]
[0133] Where u is a point on the corrected wide-angle image. and is the transformation parameter, which can be derived by minimizing equation T. In addition, U(r) is the radial basis function, which represents the influence of the control point on u: U(r) = r 2 logr 2 .
[0134] To this end, the present invention defines N = (U + 1) × (V + 1) control points that can be connected to form a grid, and the network learns to predict the image grid after correction migration. Then, A-TPS transforms the predicted grid into a regular grid on the target domain image, where the control points are set to be evenly distributed on the image.
[0135] For the network, the present invention first selects the pre-trained ResNet50 as the backbone network to extract high-level semantic features. Assuming that the size of the input clothing model image is W×H, the size of the backbone network output feature map is These feature maps are then fed to the motion head to predict the (U+1)×(V+1) control points on the rectified image. Four convolutional layers with a kernel size of 3 and queue normalization are stacked to aggregate high-level features, where all channel dimensions are set to 512. Subsequently, an attention layer and three fully connected layers with the following number of units: 4096, 2048, (U+1)×(V+1)×2 are used to predict the coordinates of all control points. With the estimated control points and attention scores, the present invention develops a more flexible attention-enhanced TPS conversion method (A-TPS). For each position P in the rectified feature map i ′, its original position P i Determined by the geometric transformation given by the following equation:
[0136] P i =G·F(P i ′)
[0137] Where P i Calculated by G, F() contains the attention scores as follows:
[0138]
[0139] Where N′ is a set of predicted control points after full connection, K=WH is the number of control points. S is a square matrix whose elements s ij =Eu(||n i -n j ||) is defined as the number of control points N that are used to calculate n i With n j The radial basis kernel is the Euclidean distance between .
[0140]
[0141] In the formula, A i,k is the attention score between the i-th position and the k-th control point. λ and β are hyperparameters, which are empirically set to 0.5 and 1, respectively. When λ = 0, the equation changes to the traditional TPS.
[0142] A-TPS brings additional flexibility to TPS by incorporating attention scores. This produces more natural corrections by content-aware and adaptively weighting the control points when performing the transformation. Since correction and recognition are jointly optimized, this flexibility can better guide parameter updates for high-quality style transfer.
[0143] (7) The method for edge repair using Shared Attention Stable Diffusion (SA-SD) is:
[0144] First, input the image of the mannequin after fabric migration and the clothing mask map that needs to be repaired. Then, the edge area in the image is extracted by transparency superposition, and a preliminary edge mask is generated. Next, the area mask that needs to be repaired is generated based on the edge mask to ensure that these areas contain the edges and the parts around them that need to be smoothed and repaired. The image of the mannequin after fabric migration and the repair area mask are used as input data for the SA-SD model. Next, the input data is input into the SA-SD model for edge repair to generate a repaired image, fill and smooth the edge area, and make it transition naturally with the surrounding area. Finally, extract the repaired image, combine the repaired edge area with the original image by transparency superposition, and output the final repaired image.
[0145] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A method for replacing clothing fabrics, characterized in that: The following steps are involved: S1: Perform grayscale calculation on the clothing image data in the database to obtain a grayscale image, and use the Mosaic method to perform data enhancement on the RGB image data and the grayscale image data respectively. After image preprocessing, they are input into the C-TransUNet network together, and the model clothing image and fabric image input by the user are obtained at the same time; S2: Accurate segmentation of model clothing images is achieved by building a semantic segmentation network based on C-TransUNet. The C-TransUNet network model consists of a convolutional layer, a shallow feature fusion module, a fused visual Transformer network, and a decoder module. The convolutional layer and the shallow feature fusion module are used to extract and fuse shallow features, the fused visual Transformer network is used to learn cross-channel deep representations with high inter-class separability and low intra-class diversity, and the decoder module restores the spatial information of the input image with higher accuracy. S3: Use the binary cross entropy loss function to optimize the parameters of the network in S2 to further improve the segmentation performance of clothes in clothing images and save the trained model weights; S4: Load the model weights and network structure trained in S3, and input the model clothing image input by the user into the trained C-TransUNet to obtain a mask image of the clothing in the clothing image; S5: Design a Shared Attention Stable Diffusion model to perform a diffusion operation on the input image using a diffusion operator q, and then apply a single reverse operator p θ To complete a diffusion operation, the four-way continuous generation of the user input fabric graph is realized through multiple continuous diffusion operations. The Shared Attention Stable Diffusion model is: Project the features to the query Q∈m×d through a linear layer k , key K∈m×d k The sum V∈m×d k Then, the attention is calculated using the following formula: Where, d k is the dimension of Q and K. Intuitively, each image feature is updated by V weighted and updated, and its weight depends on the correlation between the projected query Q and the key K. Each self-attention layer contains multiple attention heads. By concatenating the outputs of multiple attention heads and projecting them back to the image feature space d h To calculate the residual: S6: Using the Attention-Enhanced Thin Plate Spline algorithm, fabric style transfer is achieved while maintaining fabric details. Based on the mask image of the clothing image, the personalized fabric input by the user is transferred to the model clothing image to achieve more realistic fabric and shadow effects. The Attention-Enhanced Thin Plate Spline algorithm is: Approximate image style transfer by minimizing transformation energy through a unified model: e=e T +le d T:(x rec ,y rec )→(x rec 2 ,y rec 2 ) In the formula, ε represents the total energy of the expected transformation, (x rec ,y rec ) and (x rec 2 ,y rec 2 ) represent the points in the source domain S and the target domain T respectively, ε T represents the data penalty energy, ε d Denotes the distortion energy, the above formula with the lowest total energy is the desired transformation, where the hyperparameter λ is designed to balance the energy between data penalty and distortion; S7: The blurred edges of the model clothing image after migration are repaired using the SA-SD model according to the mask image of the clothing image, and the clothing model image after fabric style migration and edge repair is output.
2. A method for replacing clothing fabric according to claim 1, characterized in that: The specific operation steps of performing Mosaic data enhancement on the clothing image in step S1 are: S11: Randomly read four clothing images from the clothing image dataset; S12: respectively flipping the four images left and right, scaling the original images, and changing the brightness, saturation, and hue of the original images; S13: stitching together the four images transformed in S12, with the first image placed at the upper left, the second image placed at the lower left, the third image placed at the lower right, and the fourth image placed at the upper right; S14: After the four images are arranged, fixed areas of the four images are cut out using a matrix method, and then they are spliced together to form a new image, which contains the original data features and data distribution.
3. A method for replacing clothing fabric according to claim 1, characterized in that: The convolutional layer and shallow feature fusion module in the semantic segmentation network based on C-TransUNet in step S2 is: Use X∈R H×W×3 and Y∈R H×W×1 Represents an RGB image and its corresponding grayscale image data, where H and W are the height and width of the input, the RGB image has three channels, and the grayscale image data has one channel; first, a dual-branch encoder is designed for C-TransUNet based on a dual-branch architecture. Each branch of the encoder consists of four convolutional layers, which is conducive to extracting multi-scale features of the two modalities; then, the shallow features of the grayscale image and the RGB image are extracted through convolution operations and input into the shallow feature fusion module to generate fused shallow features, which are then input before the next RGB image encoder branch; finally, the output of the shallow feature fusion module is directly fed into the corresponding decoder layer using skip connections to recover detailed local and contextual information.
4. A method for replacing clothing fabric according to claim 1, characterized in that: The fused visual Transformer network in step S2 is: For a given x I and I Respectively indicate the size RGB and DSM feature maps, where I and C I Represent the layer index and output channel size of the last layer in the CNN backbone network respectively; First, two linear layers and a reshape operation are used to transform x I and I Vectorize it into and Size C hid ×L, where is the sequence length; specific position codes are then added to and The position information is retained and input into the fusion visual Transformer network, aiming to extract rich contextual information from multimodal data and deeply fuse it; Specifically, the input of the fusion vision Transformer network encoder goes through three stages in sequence, including the first stage dynamic attention layer for deep feature enhancement, the second stage adaptive cross fusion attention layer for deep feature fusion, and the third stage dynamic attention layer for fusion feature enhancement, with 3, 6 and 3 layers respectively; and represents the hidden features of the nth layer in the RGB branch and the DSM branch, where n∈(1,2,…,12). It is worth noting that this process keeps the dimension of the feature map as C throughout the FVit. hid ×L; the first stage dynamic attention layer consists of two dynamic attention modules, two multi-layer perceptron modules and a layer normalization layer; for a given and Given the multi-modal feature input represented by , the first stage dynamic attention layer is designed to derive the global relationship of each modality using a multi-head dynamic attention mechanism; After the first-stage dynamic attention layer performs deep feature enhancement, the fused visual Transformer network further uses the second-stage adaptive cross-fusion attention layer to fuse multimodal features in the abstract semantic space with rich contextual information by simultaneously calculating cross-attention and self-attention to learn the correlation between the primary modality RGB and the auxiliary modality DSM; Finally, the fused feature map is enhanced by the third stage dynamic attention layer. The specific calculation process is the same as in the first stage to enhance the fused feature maps of the RGB branch and the DSM branch respectively. The final output of the fused visual Transformer network is expressed as 5. A method for replacing clothing fabric according to claim 1, characterized in that: The decoder in the semantic segmentation network based on C-TransUNet in step S2 is: The decoder recovers the hidden fusion features of the final segmentation process by utilizing multiple upsampling modules. The decoder first uses a reconstruction module to transform the 2-D input sequence z N Reshape to size A 3-D tensor, where C dec is the number of channels of the first block in the input decoder; after that, multiple cascaded decoder blocks restore the spatial resolution to H×W by connecting the skip connections from the corresponding CNN backbone layers. Each decoder block consists of an upsampling operator, a convolutional layer, and a ReLU layer. Finally, the segmentation head performs the final semantic prediction.
6. A method for replacing clothing fabric according to claim 1, characterized in that: The binary cross entropy loss function calculation formula in step S3 is: In the formula, M represents the number of classification categories, y ic is a sign function. When the true category of sample i is equal to category c, it takes 1, otherwise it takes 0. ic is the predicted probability value that the observed sample i belongs to category c; the total loss of the C-TransUNet network model is calculated by using the binary cross entropy loss function, and the parameters of the network model are updated and optimized using the back propagation algorithm and gradient descent algorithm. The purpose of network training is to minimize the total network loss function.
7. A method for replacing clothing fabric according to claim 1, characterized in that: The method for edge repairing using Shared Attention Stable Diffusion in step S7 is: S71: inputting a clothing model image and a clothing mask image for which fabric migration has been completed but needs to be repaired; S72: extracting edge regions in the image by transparency superposition and generating a preliminary edge mask; S73: Generate area masks that need to be repaired according to the edge mask, ensuring that these areas include edges and surrounding parts that need to be smoothed and repaired; S74: using the clothing model image after fabric migration and the repaired region mask as input data of the SA-SD model; S75: inputting the input data into the SA-SD model to perform edge repair to generate a repaired image, thereby filling and smoothing the edge area so that it transitions naturally with the surrounding area; S76: extracting the repaired image, combining the repaired edge area with the original image by using transparency overlay, and forming a final repaired image.
8. A device using the method for replacing clothing fabrics as described in any one of claims 1 to 7, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor runs the computer program, the steps of the clothing fabric method based on Shared Attention Stable Diffusion and C-TransUNet network are implemented.
Citation Information
Patent Citations
Semantic image segmentation method and system based on edge enhancement
CN111462126A
Repeated synthesis of image using direct deformation of image, pass discriminator and coordinate-based remodelling
RU2726160C1