Cross-attention-based multi-size window transform network method and system for cloth image registration
By using a multi-size window Transformer network based on cross-attention, the problem of imprecise feature extraction in image registration is solved, achieving efficient image registration and defect recognition, and improving fabric production efficiency.
Patent Information
- Application Number
- CN202310933471.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-27
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-07-27
AI Technical Summary
Existing image registration methods based on the Transformer structure only focus on the correlation of a single image and ignore the mapping relationship between image pairs. This results in imprecise feature extraction and ineffective image registration, especially in the case of identifying defective fabrics in fabric production, where the methods are inefficient.
A multi-size window Transformer network based on cross-attention is adopted. Image correspondence is learned through cross-attention, local transformation is performed using multi-size windows to obtain detailed information, and feature matching is performed by sharing parameters through a feature fusion module. Finally, image registration and defect recognition are performed using deformation field.
It improves the accuracy and efficiency of image registration, effectively identifies defective fabrics, and increases production efficiency.
Smart Images

Figure CN116934820B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of fabric image registration technology, specifically relating to an attention mechanism Transformer structure and a deformable image registration method and system with multi-size windows. Background Technology
[0002] With the rapid development of digital image acquisition technology, it is now possible to easily obtain image data from different perspectives and time points. This image data plays a crucial role in many computer vision fields, such as marine resource exploration, medical image diagnosis, remote sensing image processing, and target anomaly detection. For example, sonar images are already used in many underwater tasks such as seabed target detection, target tracking, and path planning; medical images play an important role in image applications and pathological analysis; and remote sensing images are widely used in mapping, environmental monitoring, and weather forecasting. However, due to varying image acquisition conditions, images may undergo transformations such as rotation, translation, scaling, and distortion, and complex nonlinear relationships may even exist between image pairs, leading to incomplete image matching and making subsequent image analysis and processing difficult. Therefore, image registration is necessary before image analysis and processing.
[0003] Image registration has become a crucial problem in computer vision, with wide applications in video analysis, pattern recognition, and moving target detection. However, during image acquisition, environmental complexity and equipment limitations can lead to noise contamination and various forms of distortion. These factors result in images with low signal-to-noise ratios, low resolution, and indistinct texture features, and complex nonlinear relationships exist between images from different viewpoints. Furthermore, the rapid increase in global image data volume and the expanding applications of image data place higher demands on the speed and accuracy of image registration methods, posing significant challenges and necessitating continuous improvement of image registration techniques.
[0004] Traditional image registration is based on point features such as SIFT, SURF, and ORB. Point features effectively reduce the number of false matches under the specific characteristics of image registration, achieving the goal of image registration by establishing an image transformation model. With the development of deep learning, neural networks have been used for image registration. Image registration extracts image features through neural networks. Therefore, it is superior to traditional image registration methods. Supervised image registration methods obtain deformation model parameters between input images through neural networks to achieve image registration. Unsupervised image registration methods do not require manually constructing image deformation models and evaluating image matching through similarity. Due to the complex nonlinear transformations between images in image registration, constructing parametric deformation models is often challenging. In recent years, unsupervised image registration based on deformation fields has received increasing attention. Deformation fields achieve image matching by constructing the vector displacement of each pixel in the image to be registered.
[0005] While existing attention-based Transformer architectures can perform image matching, traditional Transformers still employ the same attention mechanism as single-image tasks, focusing only on the relevance of a single image while ignoring the mapping relationships between image pairs. This limits the Transformer's ability to find effective registration features for fine-grained registration. Furthermore, in the process of extracting image features, the global correspondence approach cannot extract features precisely, limiting the correspondence between different pieces of information between images and potentially leading to problems such as the loss of key structures and details. Summary of the Invention
[0006] To address the issues of focusing only on the relevance of a single image and the loss of some features, this invention proposes an image registration method and system based on a multi-size window Transformer network using cross-attention. First, this invention learns the correspondence between images through cross-attention, utilizing its attention mechanism to calculate the relevance of image pairs, prompting automatic feature matching within the network. Second, a feature fusion module based on cross-attention continuously matches and fuses features, integrating two input features into a single attention information and sharing parameters for feature matching. Finally, multi-size windows are used to focus on the local transformations of deformable registration, acquiring detailed information while constraining the attention calculation between the basic window and search windows of different sizes. This invention improves the accuracy of image registration and is beneficial for identifying defective fabrics in the fabric production process, thereby increasing production efficiency.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] A method for fabric image registration using a multi-size window Transformer network based on cross-attention includes the following steps:
[0009] S1. Process real fabric images and divide them into training and test sets;
[0010] S2. Create a dual-channel Transformer network to divide the input image pair into image blocks of the same size and encode them linearly, and then extract features from the fixed image and the moving image respectively;
[0011] S3. Feature blocks from the dual-channel network exchange their input order and obtain cross attention through the multi-size window method in two Cross AttentionTransformers (CAT), fusing the two input features into one attention information;
[0012] S4. The cross-fused feature blocks are aggregated using a skip connection method to finally obtain the output deformation field;
[0013] S5. Use the obtained deformation field and spatial transformation network to deform the fabric image to obtain the registered image, and calculate the similarity between the fixed image and the registered image;
[0014] S6. Perform a difference operation between the registered fabric image and the fixed image, and identify defective fabrics based on the pixels of the differenced image.
[0015] Furthermore, in step S1, data processing includes image cropping to obtain training and test sets, data augmentation of the training set, and inputting the augmented training dataset into the network.
[0016] Furthermore, in step S2, features of the moving and stationary images are extracted using dual parallel networks. The two networks communicate through feature fusion, and their operational mechanisms are identical. These two parallel networks follow the encoding and decoding parts of the UNet structure, but replace convolutions with Cross Attention Transformer blocks. These blocks play a crucial role in the attention feature fusion module between the two networks, promoting automatic feature matching within the network. The network of this invention not only vertically exchanges cross-image information but also maintains horizontal refinement functionality. Because the mechanisms of the upper and lower parallel networks are identical, one of the networks will be described below, referred to as the single-channel network.
[0017] The single-channel Transformer architecture proceeds as follows: First, the input color image is cropped into image patches without repeating regions using an image patch segmentation module. Each image patch can be considered a marker, its function being to connect the RGB values of the input image pixels. In the single-channel network, the image patch size is set to 4×4, so the feature dimension of a single image patch is 48. A linear embedding layer is applied to the segmented image patches, its function being to map the 48-dimensional image patches to an arbitrary dimension (C). Features are extracted from these image patch markers using several improved Transformer blocks, without changing the number of Transformer block markers (H / 4×W / 4), and this process, along with the linear embedding module, is referred to as "Step 1".
[0018] As the network deepens, an image patch merging module is used to reduce the number of labels in order to obtain multi-level features. Assuming the input to the image patch merging module is a 4×4 feature map, the module first concatenates patches of the same color together to form four 2×2 image patches; then it performs a normalization operation on the features of these four image patches; finally, it performs a linear transformation through a linear layer. At this point, the number of labels is reduced to four times the previous number (2× resolution downsampling), and the output dimension becomes 2C.
[0019] Next, a Transformer block is applied for feature swapping, maintaining a resolution of H / 8 × W / 8. The image block merging module and the feature transformation Transformer block are denoted as "Step 2". The above "Step 2" process is repeated twice, denoted as "Step 3" and "Step 4", with output resolutions of H / 16 × W / 16 and H / 32 × W / 32 respectively.
[0020] Furthermore, in step S3, two CAT blocks are used to fuse the two input features into a single attention information, sharing parameters for feature matching; and the multi-size window method in the CAT block is used to achieve accurate local correspondence, ultimately generating a refined deformation flow field. The moving image features T from the parallel sub-network... m and fixed image features T f By swapping the input order, mutual attention is obtained through two CAT blocks. Then, the outputs of the other two attention blocks are returned to the original channels to obtain the fused feature T. mf and T fm This prepares for further, more in-depth information exchange. In a feature fusion module, there are a total of k communications to obtain sufficient mutual information. Through the attention-based feature fusion module between the two networks, features from different networks with different semantic information frequently exchange information. Therefore, the network of this invention can continuously learn multi-level semantic features for final, fine-grained registration.
[0021] This novel attention mechanism, CAT, enables the full exchange of information between image pairs, balancing the representativeness and multi-scale nature of matching features. Assume that b and s are divided into two sets of windows in different ways, with a basic window set S. ba and search window set S se This is used for the next window-based attention computation. The purpose of the CAT block is to compute new feature labels that have corresponding relevance from input feature b to feature s through an attention mechanism. ba and S se They have the same number, but different window sizes. (The last part, "S," appears to be a typo and can be left as is.) ba Each base window is projected onto the query set query, and each search window is projected onto the knowledge set keys and values through a linear layer. Then, window-based multi-head cross-attention (W-MCA) is used to compute cross-attention between two windows and adds the attention to the base windows, allowing each base window to obtain the corresponding weighted information from the search windows. Finally, the new output set is sent to a multilayer perceptron with GELU nonlinearity to improve its learning ability. A LayerNorm (LN) layer is used before each W-MCA and each MLP module to ensure that each layer is effective.
[0022] Multi-size window partitioning includes two different methods, window partitioning (WP) and window region partitioning (WAP), to divide the input feature labels b and s into windows of different sizes. WP partitions the feature labels directly into a base window set S of size n×h×w. ba In the WAP, the window size increases with the magnification factors α and β. Therefore, the sizes of the base and search windows are calculated as follows:
[0023] h ba ,w ba =h,w
[0024] h se ,w se =α·h,β·w
[0025] Among them, h ba w ba The size of the base window, and h se w se The size of the search window; to obtain the same number of two window sets, WAP utilizes a sliding window and sets the stride to the base window size, therefore S se The size is n×α·h×β·w. By using corresponding windows of different sizes, the CAT block effectively calculates the cross-attention between two feature labels, achieving accurate information exchange without large-span search.
[0026] Attention is a function that maps a query and a set of key-value pairs to an output, where the query, key, value, and output are all vectors. The W-MCA proposed in this invention calculates the cross-attention between a base window and a search window to obtain an accurate correspondence. K, Q, and V represent the features mapped from the image patch, where K represents the features mapped from the base window, and Q and V come from the search window. The calculated result is a weighted sum, where the weight assigned to each value is calculated by a compatibility function between the query and the corresponding key.
[0027] W-MCA employs multi-head attention to fully represent the subspace, performing a dot product operation between the query and the key. First, each key is divided by... Next, a softmax function is used to obtain the weights of these values. Therefore, the cross-attention calculation is expressed as:
[0028]
[0029] Among them, Q ba K se V se It consists of a query matrix, a key matrix, and a value matrix. Q ba ∈R n×s×c It is S ba and K se linear projection, V se ∈R n×μ·s×c It is S se The linear projections of s = h × w and μ = α · β, where c is the dimension of each feature label.
[0030] Furthermore, in step S4, the cross-fused feature blocks are aggregated using a skip connection method to finally obtain the output deformation field;
[0031] Furthermore, in step S5, the network's loss function consists of two parts: first, the similarity loss, denoted by Mean Squared Error (MSE), which measures the similarity between moving and stationary images and penalizes the differences between them; and second, the regularization loss, which consists of a hyperparameter and a regularization term. The regularization term adds a smoothness constraint to the estimated deformation field to prevent the deformation field from folding too much.
[0032] MSE represents the expected value of the squared difference between the true value and the estimated value. A smaller MSE value indicates better prediction accuracy. The mean squared error between the moving image and the predicted image is expressed as:
[0033]
[0034] Where P represents a pixel in the moving and stationary images, and Ω represents the entire image region.
[0035] Regularization penalizes folds in a deformation field, and is expressed as:
[0036]
[0037] Where R(θ) is a regularization term, This represents the gradient at point P in the X and Y directions. If we use... Let the coefficients of the loss regularization term represent the loss function, then the loss function can be expressed as:
[0038]
[0039] Furthermore, in step S6, the registered fabric image is differentially analyzed with the fixed image, and the defective fabric image is identified based on the pixels of the differential image. This involves setting a threshold; the threshold and window size are pre-defined, and the average pixel value within the window is checked sequentially by sliding the window. If the average value exceeds the threshold, the image is considered defective; otherwise, it is considered defective.
[0040] This invention also discloses a multi-size window Transformer network fabric image registration system based on cross-attention, used to perform the above method, which includes the following modules:
[0041] Dataset creation module: Crops the fabric images and further divides them into training and test sets;
[0042] Dual-channel Transformer structure module: Creates a dual-channel Transformer structure network, which divides the input image pair into image blocks of the same size and encodes them linearly, and then extracts features from the stationary image and the moving image respectively;
[0043] Feature fusion module: Feature blocks from the dual-channel network exchange their input order and obtain cross attention through the multi-size window method in two CrossAttention Transformers (CAT), fusing the two input features into one attention information;
[0044] Feature aggregation module: Aggregates features between cross-fused feature blocks using a skip connection method to finally obtain the output deformation field;
[0045] Training module: The model is trained using mean squared error loss and regularization loss;
[0046] Defect detection module: Performs a difference operation between the registered fabric image and a fixed image, and identifies defective fabric based on the pixels of the differenced image.
[0047] Compared with existing technologies, the fabric image registration method and system based on a multi-size window Transformer network using cross-attention in this invention firstly utilizes a cross-attention-based Transformer block to fuse information at different scales, effectively addressing the mapping problem between image pairs; secondly, it employs a multi-size window to focus on deformable local transformations, acquiring detailed features to improve registration results. This invention, based on a cross-attention Transformer architecture, continuously matches and fuses image pair features, integrating two input features into a single attention information and sharing parameters for feature matching, thus better addressing image registration problems and completing the task of identifying defective fabrics. Attached Figure Description
[0048] Figure 1 This is a flowchart of the fabric image registration method based on cross-attention multi-size window Transformer network provided in Embodiment 1 of the present invention.
[0049] Figure 2 This is a schematic diagram of the single-channel Transformer architecture in step S12 provided in Embodiment 1 of the present invention.
[0050] Figure 3 This is a schematic diagram of the feature fusion module in step S13 provided in Embodiment 1 of the present invention.
[0051] Figure 4 This is a schematic diagram of the multi-size window method in step S13 provided in Embodiment 1 of the present invention.
[0052] Figure 5 This is a schematic diagram of the multi-size window operation for cross-attention in step S13 provided in Embodiment 1 of the present invention.
[0053] Figure 6 This is a block diagram of a multi-size window Transformer network fabric image registration system based on cross-attention, provided in Embodiment 1 of the present invention. Detailed Implementation
[0054] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.
[0055] The purpose of this invention is to address the shortcomings of existing technologies by providing a method and system for fabric image registration using a multi-size window Transformer network based on cross-attention.
[0056] Example 1
[0057] like Figure 1 As shown, this embodiment provides a method for fabric image registration using a multi-size window Transformer network based on cross-attention. The specific implementation process includes the following steps:
[0058] S11. The pre-collected fabric images are cropped to the specified size and further divided into training and test sets, and the training set is augmented with data.
[0059] S12. Create a dual-channel Transformer network to divide the input image pair into image blocks of the same size and encode them linearly, and then extract features from the fixed image and the moving image respectively;
[0060] S13. The feature blocks from the dual-channel network exchange the order of their inputs and use the multi-size window method in two Cross Attention Transformers (CAT) to obtain cross attention, fusing the two input features into one attention information;
[0061] S14. Use skip connections to aggregate the cross-fused feature blocks separately to obtain the output deformation field;
[0062] S15. Constrain the deformation field using smoothness loss and train it using similarity loss;
[0063] S16. Use the pixels obtained by performing a difference operation between the registered fabric image and the fixed image to identify defective fabric.
[0064] The specific steps of this embodiment are described below:
[0065] In step S11, the acquired fabric images are cropped to a size of 512x512 and divided into training and test sets at a 5:1 ratio. Simultaneously, data augmentation is applied to the training dataset by performing a radial transformation on the training images and then adding elastic deformation. The scaling factor α and elastic coefficient σ can be adjusted or changed according to different acquired fabric images.
[0066] In step S12, features of the moving and stationary images are extracted using dual parallel networks. The two networks communicate through a feature fusion module, and their operational mechanisms are identical. These two parallel networks follow the encoding and decoding parts of the UNet structure, but replace convolutions with CAT (Cross Attention Transformer) blocks. These blocks play a crucial role in the attention feature fusion module between the two networks, facilitating automatic feature matching within the network. The network not only vertically exchanges cross-image information but also maintains horizontal refinement functionality. Because the mechanisms of the upper and lower parallel networks are identical, one of the networks will be described below, referred to as the single-channel network.
[0067] Figure 2 The single-channel Transformer architecture is shown, and its process is as follows: First, the input color image is cropped into image patches without repeating regions using an image patch segmentation module. Each image patch can be considered a marker, its function being to connect the RGB values of the input image pixels. In the single-channel network, the image patch size is set to 4×4, so the feature dimension of a single image patch is 48. Then, a linear embedding layer is applied to the segmented image patches, its function being to map the 48-dimensional image patches to an arbitrary dimension (C). Features are extracted from these image patch markers through several improved Transformer blocks, without changing the number of Transformer block markers (H / 4×W / 4), and this, along with the linear embedding module, is referred to as "Step 1".
[0068] As the network becomes more sophisticated, an image patch merging module is used to reduce the number of labels in order to obtain multi-level features. Assuming the input to the image patch merging module is a 4×4 feature map, the module first concatenates patches of the same color together to form four 2×2 image patches; then it normalizes the features of these four patches; finally, it performs a linear transformation through a linear layer. At this point, the number of labels is reduced to four times the previous number (2× resolution downsampling), and the output dimension becomes 2C.
[0069] Next, a Transformer block is applied for feature swapping, maintaining a resolution of H / 8 × W / 8. The image block merging module and the feature transformation Transformer block are denoted as "Step 2". The above "Step 2" process is repeated twice, denoted as "Step 3" and "Step 4", with output resolutions of H / 16 × W / 16 and H / 32 × W / 32 respectively.
[0070] In step S13, two CAT blocks are used to fuse the two input features into a single attention information, sharing parameters for feature matching; and the multi-size window method in the CAT block is used to achieve accurate local correspondence, ultimately generating a refined deformation flow field. Figure 3 As shown in (a), the moving image features T from the parallel subnetwork m and fixed image features T f By swapping the input order, mutual attention is obtained through two CAT blocks. The outputs of the other two attention blocks are returned to the original channels to obtain the fused feature T. mf and T fm This prepares for further, more in-depth information exchange. In a feature fusion module, there are a total of k communications to obtain sufficient mutual information. Through the attention-based feature fusion module between the two networks, features from different networks with different semantic information frequently exchange information. Therefore, the network of this invention can continuously learn multi-level semantic features for final, fine-grained registration.
[0071] This novel attention mechanism, CAT, enables the full exchange of information between image pairs, balancing the representativeness and multi-scale nature of matching features. For example... Figure 3 As shown in (b), b and s are divided into two groups of windows in different ways, and the basic window set S ba and search window set S se This is used for the next window-based attention computation. The purpose of the CAT block is to compute new feature labels that have corresponding relevance from input feature b to feature s through an attention mechanism. ba and S se They have the same number, but different window sizes. (The last part, "S," appears to be a typo and can be left as is.) ba Each base window is projected onto the query set query, and each search window is projected onto the knowledge set keys and values through a linear layer. Then, window-based multi-head cross-attention (W-MCA) is used to compute cross-attention between two windows and adds the attention to the base windows, allowing each base window to obtain the corresponding weighted information from the search windows. Finally, the new output set is sent to a multilayer perceptron with GELU nonlinearity to improve its learning ability. A LayerNorm (LN) layer is used before each W-MCA and each MLP module to ensure that each layer is effective.
[0072] Multi-size window partitioning includes two different methods, window partitioning (WP) and window region partitioning (WAP), to divide the input feature labels b and s into windows of different sizes. For example... Figure 4 As shown, the WP partition feature labels are directly entered into the base window set S of size n×h×w. ba In the WAP, the window size increases with the magnification factors α and β. Therefore, the sizes of the base and search windows are calculated as follows:
[0073] h ba ,w ba =h,w
[0074] hse ,w se =α·h,β·w
[0075] Among them, h ba w ba h represents the size of the basic window. se w se S represents the size of the search window, where α and β are magnification factors. To obtain the same number of two window sets, WAP utilizes a sliding window and sets the stride to the base window size, therefore S se The size is n×α·h×β·w. By using corresponding windows of different sizes, the CAT block effectively calculates the cross-attention between two feature labels, achieving accurate information exchange without large-span search.
[0076] Attention refers to a function that maps a query and a set of key-value pairs to an output, where the query, key, value, and output are all in vector form. For example... Figure 5 As shown, the proposed W-MCA computes the cross-attention between the base window and the search window to obtain an accurate correspondence. K, Q, and V represent the features mapped from the image patch, where K represents the features mapped from the base window, and Q and V come from the search window. The computed value is a weighted sum, where the weight assigned to each value is calculated by a compatibility function between the query and the corresponding key.
[0077] W-MCA employs multi-head attention to fully represent the subspace, performing a dot product operation between the query and the key. First, each key is divided by... Next, a softmax function is used to obtain the weights of these values. Therefore, the cross-attention calculation is expressed as:
[0078]
[0079] Among them, Q ba K se V se It consists of a query matrix, a key matrix, and a value matrix. Q ba ∈R n×s×c It is S ba and K se linear projection, V se ∈R n×μs×c It is S se The linear projections of s = h × w and μ = α · β, where c is the dimension of each feature label.
[0080] In step S14, the cross-fused feature blocks are aggregated separately using skip connections, and finally the deformation field is obtained through convolution.
[0081] In step S15, the network's loss function consists of two parts: first, the similarity loss, denoted by Mean Squared Error (MSE), which measures the similarity between moving and stationary images and penalizes the differences between them; and second, the regularization loss, which consists of a hyperparameter and a regularization term. The regularization term adds a smoothness constraint to the estimated deformation field to prevent the deformation field from folding too much.
[0082] MSE represents the expected value of the squared difference between the true value and the estimated value. A smaller MSE value indicates better prediction accuracy. The mean squared error between the moving image and the predicted image is expressed as:
[0083]
[0084] Where P represents a pixel in the moving and stationary images, and Ω represents the entire image region.
[0085] Regularization penalizes folds in a deformation field, and is expressed as:
[0086]
[0087] Where R(θ) is a regularization term, This represents the gradient at point P in the X and Y directions. If we use... Let the coefficients of the loss regularization term represent the loss function, then the loss function can be expressed as:
[0088]
[0089] In step S16, the registered fabric image is differentially analyzed with the fixed image, and the defective fabric image is identified based on the pixels of the differential image. This embodiment adopts the idea of setting a threshold. The threshold size and window size can be set empirically. By sliding the window, it is determined whether the average value of the pixels within the window exceeds the threshold. If it exceeds the threshold, the image has defects; otherwise, it does not.
[0090] This embodiment proposes a fabric image registration method based on a multi-size window Transformer network with cross-attention. Specifically, it utilizes a Transformer block based on cross-attention to fuse information at different scales, effectively addressing the mapping problem between image pairs. Furthermore, it employs a multi-size window to focus on deformable local transformations, acquiring detailed features to improve the registration performance of complex fabric images. In the subsequent task of identifying defective fabrics, the defective fabric image can be determined simply by subtracting the registered image from a fixed image.
[0091] Example 2
[0092] like Figure 6As shown, this embodiment provides a multi-size window Transformer network fabric image registration system based on cross-attention, which is used to execute the method of Embodiment 1, and specifically includes the following modules:
[0093] Dataset creation module: Crops the fabric images and further divides them into training and test sets;
[0094] Dual-channel Transformer structure module: Creates a dual-channel Transformer structure network, which divides the input image pair into image blocks of the same size and encodes them linearly, and then extracts features from the stationary image and the moving image respectively;
[0095] Feature fusion module: Feature blocks from the dual-channel network exchange their input order and obtain cross attention through the multi-size window method in two CrossAttention Transformers (CAT), fusing the two input features into one attention information;
[0096] Feature aggregation module: Aggregates features between cross-fused feature blocks using a skip connection method to finally obtain the output deformation field;
[0097] Training module: The model is trained using mean squared error loss and regularization loss;
[0098] Defect detection module: Performs a difference operation between the registered fabric image and a fixed image, and identifies defective fabric images based on the pixels of the differenced image.
[0099] The modules in this embodiment are described in detail below.
[0100] In the dataset creation module, the acquired fabric images are cropped to a size of 512x512 and divided into training and test sets at a 5:1 ratio. Data augmentation is applied to the training dataset by performing a radial transformation on the images and then adding elastic deformation. The scaling factor α and elastic coefficient σ can be adjusted or changed based on the different fabric images acquired.
[0101] In the dual-channel Transformer architecture, two parallel networks are used to extract features from moving and stationary images respectively. The two networks communicate through a feature fusion module, and their operational mechanisms are identical. These two parallel networks follow the encoding and decoding parts of the UNet architecture, but replace convolutions with CAT (Cross Attention Transformer) blocks. These blocks play a crucial role in the attention feature fusion module between the two networks, facilitating automatic feature matching within the network. The network not only vertically exchanges cross-image information but also maintains horizontal refinement functionality. Because the mechanisms of the parallel networks are identical, we will now introduce one of the networks, referred to below as the single-channel network.
[0102] Figure 2 The single-channel Transformer architecture is shown, and its process is as follows: First, the input color image is cropped into image patches without repeating regions using an image patch segmentation module. Each image patch can be considered a marker, its function being to connect the RGB values of the input image pixels. In the single-channel network, the image patch size is set to 4×4, so the feature dimension of a single image patch is 48. Then, a linear embedding layer is applied to the segmented image patches, its function being to map the 48-dimensional image patches to an arbitrary dimension (C). Features are extracted from these image patch markers through several improved Transformer blocks, without changing the number of Transformer block markers (H / 4×W / 4), and this, along with the linear embedding module, is referred to as "Step 1".
[0103] As the network deepens, an image patch merging module is used to reduce the number of labels in order to obtain multi-level features. Assuming the input to the image patch merging module is a 4×4 feature map, the module first concatenates patches of the same color together to form four 2×2 image patches; then it performs a normalization operation on the features of these four image patches; finally, it performs a linear transformation through a linear layer. At this point, the number of labels is reduced to four times the previous number (2× resolution downsampling), and the output dimension becomes 2C.
[0104] Next, a Transformer block is applied for feature swapping, maintaining a resolution of H / 8 × W / 8. The image block merging module and the feature transformation Transformer block are denoted as "Step 2". The above "Step 2" process is repeated twice, denoted as "Step 3" and "Step 4", with output resolutions of H / 16 × W / 16 and H / 32 × W / 32 respectively.
[0105] In the feature fusion module, two CAT blocks are used to fuse two input features into a single attention information, sharing parameters for feature matching; and the multi-size window method in the CAT block is used to achieve accurate local correspondence, ultimately generating a refined deformation flow field. For example... Figure 3 As shown in (a), the moving image features T from the parallel subnetwork m and fixed image features T f By swapping the input order, mutual attention is obtained through two CAT blocks. The outputs of the other two attention blocks are returned to the original channels to obtain the fused feature T. mf and T fm This prepares for further, more in-depth information exchange in the next step. Within a feature fusion module, there are a total of k communications to obtain sufficient mutual information. Through the attention-based feature fusion module between the two networks, features from different networks with varying semantic information frequently exchange information, allowing the network to continuously learn multi-level semantic features for final, fine-grained registration.
[0106] This novel attention mechanism, CAT, enables the full exchange of information between image pairs, balancing the representativeness and multi-scale nature of matching features. For example... Figure 3 As shown in (b), b and s are divided into two groups of windows in different ways, and the basic window set S ba and search window set S se This is used for the next window-based attention computation. The purpose of the CAT block is to compute new feature labels that have corresponding relevance from input feature b to feature s through an attention mechanism. ba and S se They have the same number, but different window sizes. (The last part, "S," appears to be a typo and can be left as is.) ba Each base window is projected onto the query set query, and each search window is projected onto the knowledge set keys and values through a linear layer. Then, window-based multi-head cross-attention (W-MCA) is used to compute cross-attention between two windows and adds the attention to the base windows, allowing each base window to obtain corresponding weighted information from the search windows. Finally, the new output set is sent to a multilayer perceptron with GELU nonlinearity to improve its learning ability. A LayerNorm (LN) layer is used before each W-MCA and each MLP module to ensure that each layer is effective.
[0107] Multi-size window partitioning includes two different methods, window partitioning (WP) and window region partitioning (WAP), to divide the input feature labels b and s into windows of different sizes. For example... Figure 4 As shown, the WP partition feature labels are directly entered into the base window set S of size n×h×w. ba In the WAP, the window size increases with the magnification factors α and β. Therefore, the sizes of the base and search windows are calculated as follows:
[0108] h ba ,w ba =h,w
[0109] h se ,w se =α·h,β·w
[0110] Among them, h ba w ba The size of the base window, and h se w se The size of the search window. To obtain the same number of two window sets, WAP utilizes a sliding window and sets the stride to the base window size, therefore S se The size is n×α·h×β·w. By using corresponding windows of different sizes, the CAT block effectively calculates the cross-attention between two feature labels, achieving accurate information exchange without large-span search.
[0111] Attention refers to a function that maps a query and a set of key-value pairs to an output, where the query, key, value, and output are all in vector form. For example... Figure 5 As shown, the proposed W-MCA computes the cross-attention between the base window and the search window to obtain an accurate correspondence. K, Q, and V represent the features mapped from the image patch, where K represents the features mapped from the base window, and Q and V come from the search window. The computed value is a weighted sum, where the weight assigned to each value is calculated by a compatibility function between the query and the corresponding key.
[0112] W-MCA employs multi-head attention to fully represent the subspace, performing a dot product operation between the query and the key. First, each key is divided by... Next, a softmax function is used to obtain the weights of these values. Therefore, the cross-attention calculation is expressed as:
[0113]
[0114] Among them, Q ba K se V se It consists of a query matrix, a key matrix, and a value matrix. Q ba ∈R n×s×c It is S ba and K se linear projection, V se ∈R n×μ·s×c It is S se The linear projections of s = h × w and μ = α · β, where c is the dimension of each feature label.
[0115] In the feature aggregation module, skip connections are used to aggregate the cross-fused feature blocks separately, and finally the deformation field is obtained through convolution.
[0116] In the training module, the network's loss function consists of two parts: first, the similarity loss, denoted by Mean Squared Error (MSE), which measures the similarity between moving and stationary images and penalizes the differences between them; and second, the regularization loss, which consists of a hyperparameter and a regularization term. The regularization term adds a smoothness constraint to the estimated deformation field to prevent the deformation field from folding too much.
[0117] MSE represents the expected value of the squared difference between the true value and the estimated value. A smaller MSE value indicates better prediction accuracy. The mean squared error between the moving image and the predicted image is expressed as:
[0118]
[0119] Where P represents a pixel in the moving and stationary images, and Ω represents the entire image region.
[0120] Regularization penalizes folds in a deformation field, and is expressed as:
[0121]
[0122] Where R(θ) is a regularization term, This represents the gradient at point P in the X and Y directions. If we use... Let the coefficients of the loss regularization term represent the loss function, then the loss function can be expressed as:
[0123]
[0124] In the defect detection module, the registered fabric image is differentially analyzed with a fixed image. The defective fabric image is then identified based on the pixel count of the differential image. A threshold-based approach is used, with the threshold and window size set empirically. The average pixel count within the sliding window is checked sequentially to see if it exceeds the threshold. If it does, the image is considered defective; otherwise, it is not.
[0125] This embodiment proposes a fabric image registration system based on a multi-size window Transformer network using cross-attention. Specifically, it utilizes a cross-attention-based Transformer block to fuse information at different scales, effectively addressing the mapping problem between image pairs. Furthermore, it employs a multi-size window to focus on deformable local transformations, acquiring detailed features to improve the registration performance of complex fabric images. In the subsequent task of identifying defective fabrics, the defective fabric image can be determined simply by subtracting the registered image from a fixed image.
[0126] In summary, compared with existing technologies, the present invention, a fabric image registration method and system based on a multi-size window Transformer network with cross-attention, enhances a small number of fabric images without requiring cumbersome large-scale data collection, and uses a Transformer network for accurate image registration. Specifically, it utilizes a cross-attention-based Transformer block to fuse information at different scales, effectively addressing the mapping problem between image pairs; furthermore, it employs a multi-size window to focus on deformable local transformations, acquiring detailed features to improve the registration performance of complex fabric images. These two points enhance the accuracy of fabric image registration. The invention also maximizes the ease of use and flexibility of the model through modular design.
[0127] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A fabric image registration method based on a multi-size window Transformer network with cross-attention, characterized in that, Including the following steps: S1. Process the fabric image pairs and divide them into training and test sets; S2. Create a dual-channel Transformer network to divide the input image pair into image blocks of the same size and encode them linearly, and then extract features from the fixed image and the moving image respectively; S3. Feature blocks from the dual-channel network exchange their input order and obtain cross-attention through the multi-size window method in both CATs. The two CAT blocks then fuse the two input features into a single attention information, sharing parameters for feature matching. The multi-size window method includes window partitioning (WP) and window region partitioning (WAP) to divide the input feature labels b and s into windows of different sizes. The WP-partitioned feature labels directly enter the base window set S of size n×h×w. ba In the WAP window, the size increases with the magnification factors α and β; therefore, the sizes of the base and search windows are calculated as follows: Among them, h ba w ba h is the size of the base window. se w se The size of the search window; to obtain the same number of two window sets, WAP utilizes a sliding window and sets the stride to the base window size, therefore S se The size is n×α·h×β·w; S4. Aggregate the cross-fused feature blocks using a skip connection method to obtain the output deformation field; S5. Train the model using mean squared error loss and regularization loss; S6. Perform a difference operation between the registered fabric image and the fixed image, and identify defective fabrics based on the pixels of the differenced image.
2. The fabric image registration method based on a multi-size window Transformer network with cross-attention according to claim 1, characterized in that, In step S1, the fabric images are cropped and divided into training and testing sets, and data augmentation is performed on the resulting training set.
3. The fabric image registration method based on cross-attention multi-size window Transformer network according to claim 2, characterized in that, In step S2, features of moving and fixed images are extracted using a dual parallel network. The two networks communicate through a feature fusion module, and the two networks operate with the same mechanism.
4. The fabric image registration method based on cross-attention multi-size window Transformer network according to claim 3, characterized in that, The mechanisms of parallel networks are the same, with one of the networks referred to as a single-channel network, as follows: The input color image is cropped into image patches without repeating regions by the image patch segmentation module. Each image patch is treated as a marker, which connects the RGB values of the input image pixels. In the single-channel network, the image patch size is set to 4×4, so the feature dimension of a single image patch is 48. A linear embedding layer is applied to the segmented image patches, which maps the 48-dimensional image patches to an arbitrary dimension C. Features are extracted from these image patch markers through several improved Transformer blocks without changing the number of Transformer block markers H / 4×W / 4, and this is referred to as "Step 1" along with the linear embedding module. Assuming the input image patch merging module is a 4×4 feature map, the image patch merging module first stitches together patches of the same color to form four 2×2 image patches; then it connects the features of these four image patches and performs a normalization operation; finally, it performs a linear transformation through a linear layer; at this time, the number of labels will be reduced to 4 times the previous number, and the output dimension will become 2C. Apply a Transformer block to perform feature exchange while maintaining a resolution of H / 8×W / 8; denote the image block merging module and the Transformer block for feature transformation as "Step 2"; repeat the above "Step 2" process twice, denoted as "Step 3" and "Step 4", at which point the output resolutions are H / 16×W / 16 and H / 32×W / 32, respectively.
5. The fabric image registration method based on a multi-size window Transformer network with cross-attention according to any one of claims 1-4, characterized in that, In step S4, the cross-fused feature blocks are aggregated separately using skip connections, and finally the deformation field is obtained through convolution.
6. The fabric image registration method based on a multi-size window Transformer network with cross-attention according to any one of claims 1-4, characterized in that, In step S5, the network's loss function consists of two parts: one is the similarity loss, denoted by MSE, which measures the similarity between moving and stationary images and penalizes the difference between them; the other is the regularization loss, which consists of a hyperparameter and a regularization term. The regularization term adds a smoothness constraint to the estimated deformation field to prevent the deformation field from folding too much. MSE represents the expected value of the squared difference between the true value and the estimated value. The mean squared error between the moving image and the predicted image is expressed as: Where P represents the pixel in the moving and stationary images, Represents the entire image region; Regularization penalizes folds in a deformation field, and is expressed as: in, It is a regular expression term. This represents the gradient at point P in the X and Y directions; if we use Let the coefficients of the loss regularization term represent the loss function, then the loss function can be expressed as: Among them, I f Indicates a fixed image, I w This represents the image after deformation.
7. The fabric image registration method based on a multi-size window Transformer network with cross-attention according to any one of claims 1-4, characterized in that, In step S6, the threshold size and window size are set, and the average value of the pixels in the window is judged sequentially by sliding the window to see if it exceeds the threshold. If it exceeds the threshold, the image has defects; otherwise, it does not have defects.
8. A fabric image registration system based on a multi-size window Transformer network with cross-attention, used to perform the method of claim 1, characterized in that... Includes the following modules: Dataset creation module: Processes fabric image pairs and divides them into training and test sets; Dual-channel Transformer structure module: Creates a dual-channel Transformer structure network, which divides the input image pair into image blocks of the same size and encodes them linearly, extracting features from the stationary and moving images respectively; Feature fusion module: Feature blocks from the dual-channel network exchange their input order and obtain cross attention through the multi-size window method in the two CATs. The two CAT blocks are used to fuse the two input features into one attention information and share parameters for feature matching. Feature aggregation module: Aggregates features between cross-fused feature blocks using a skip connection method to obtain the output deformation field; Training module: The model is trained using mean squared error loss and regularization loss; Defect detection module: Performs a difference operation between the registered fabric image and a fixed image, and identifies defective fabric based on the pixels of the differenced image.
Citation Information
Patent Citations
Induced iteration forward-looking sonar image registration method and system based on multi-scale slender network
CN116228968A
ViT and sliding window attention fusion-based visual pointer understanding method and system
CN116258931A