A Transformer Method for the Tracking Structure in Image Inpainting
By introducing a Transformer method that tracks structures in image repair, combining structure enhancement modules and synchronous tracking dual-axis Transformer, the problem of difficulty in effectively recovering large-area missing images in existing technology is solved, and efficient image repair in complex scenarios is achieved. The generated images have consistency in structure and texture.
Patent Information
- Application Number
- CN202211394375.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-08
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-11-08
AI Technical Summary
In the prior art, it is difficult to effectively restore large-area missing images in image repair, especially in complex scenarios, it is difficult to generate semantic images.
A Transformer method for tracking structures for image repair is proposed. The image edge and HOG features are restored through the SEM module (SEM), and the image completion is completed using a synchronous tracking biaxial Transformer (STT), combining the structure-texture cross-attention module and the channel space biaxial attention module to realize the synchronous extraction and interaction of structure and texture.
It realizes efficient recovery of large-area missing images in complex scenarios. The generated images have consistency in structure and texture, avoids non-overlapping artifacts at the hole boundaries, and significantly improves the quality and efficiency of image repair.
Smart Images

Figure CN115619685B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image inpainting methods based on deep learning, and specifically to a Transformer method with a tracking structure for image inpainting. Background Art
[0002] Image inpainting is a fundamental low-level vision task whose main goal is to fill in the missing regions of an image while making the restored image semantically appropriate and visually pleasing. It is widely applied in many practical scenarios, such as object removal, photo editing, and image restoration. Traditional methods solve this challenging task by searching for similar patches from known regions to construct the image, but simply in this way, it is difficult to repair images with large missing areas, and when faced with a more complex image scene, it is also difficult to generate semantically reasonable images.
[0003] In recent years, convolutional neural networks (CNNs) have shown their advantages in understanding rich high-level features of images through training on large-scale datasets. However, the performance of CNN models still has bottlenecks: 1) The local inductive prior and spatially invariant kernels of convolutional operations make it difficult to restore the overall structure of the image. 2) Previous methods that utilize structural information view the fusion between structural features and subsequent feature extraction from isolated perspectives, making it difficult to convey globally consistent complementary information to help each other. 3) Some pioneering works utilize attention mechanisms to simulate long-term dependencies to solve these problems. However, the attention mechanism is only applicable to relatively small latent feature maps, where the long-range modeling ability of the model is not fully considered.
[0004] Comparing the application of the attention mechanism in CNNs, Transformer is a natural architecture for solving the long-range modeling problem, and recent progress has utilized the Transformer architecture for image inpainting tasks. Nevertheless, considering that Transformer requires a huge memory footprint, existing work still relies on CNNs for general feature extraction and only uses Transformer for high-dimensional spatial expression. Therefore, the restored image structure and texture are rough, and a complete long-range interaction has not been established yet.
[0005] Based on the above problems, the present invention proposes a Transformer method with a tracking structure for image inpainting. Summary of the Invention
[0006] (1) Technical Problems to be Solved
[0007] Aiming at the deficiencies of the prior art, the present invention provides a Transformer method with a tracking structure for image inpainting, which solves the problems described in the above background art.
[0008] (2) Technical solution
[0009] To achieve the purpose of solving the problems described in the above background technology, the present invention provides the following technical solution: A Transformer method for a tracking structure in image restoration, and the Transformer of this tracking structure for the image restoration method includes the following steps:
[0010] S1: Let be the real image, M ∈ {0, 1} H×W×1 be the mask (the missing area is 0, otherwise it is 1), I in = I gt ⊙M represents the damaged image, Y m = Y gt ⊙M, H m = H gt ⊙M and E m = E gt ⊙M respectively represent the missing gray, HOG, and Canny Edge images;
[0011] S2: After splicing the above three images and inputting them into the SEM, the restored edges E out and H out features are used as the sketch space vectors, and the formula is [E out , H out = SANet(E m , H m , Y m );
[0012] S3: STT connects the damaged image I in , the restored structural image H out and E out to finally generate the output image I out , and the formula is I out = STT(I in , H out , E out ), and the number of channels C = 24.
[0013] Preferably, in the above S2, the structure enhancement module (SEM) restores the image edges and HOG as the auxiliary structural features of the core STT. The input missing grayscale image Y m , HOG image H m and Canny edge E m are applied with a convolutional head to generate a feature map with a size of 1 / 8, reducing the computational amount of the standard self-attention. The channel-based self-attention captures the global structural information in the low-resolution feature space, and the convolutional tail uses transposed convolution to upsample these features to the output structures E out and H out , to optimize the predicted sketch structure:
[0014]
[0015] where E gt and H gt are the complete edge and HOG images respectively. The binary cross-entropy (BCE) and l1 loss are used to reconstruct the complete edge and HOG features respectively. In the experiment, λ h = 0.1,
[0016] HOG carves the distribution of the gradient direction and the edge direction within the sub-region, which is achieved by subtracting adjacent pixels (gradient filtering). Its main feature is to capture the local shape and appearance and maintain good robustness to geometric changes. Even without precise knowledge of the corresponding gradient and edge positions, HOG can well represent the appearance and shape of local objects.
[0017] Preferably, in the S3, the proposed Synchronous Tracking Twin-Axis Transformer (STT) is a U-Net architecture following the encoder-decoder style. The structural information helps the preliminary contour recovery in the early stage of image inpainting. An encoder with 24 basic Transformer blocks is designed, and each block consists of a Structure-Texture Cross-Attention Module (STCM). Its image completion flow includes a Channel-Spatial Twin-Axis Attention Module (CSPC). A decoder with 20 basic Transformer blocks is designed, and each block only contains CSPC.
[0018] Preferably, the description of the STCM is as follows: The recovered structural features contain the complete gradient distribution and edge direction. The STCM (the key component of STT) is designed to synchronously capture the long-range dependencies on both structure and texture respectively. In addition to self-attention, STCM introduces the cross-attention method to guide texture extraction by tracking the structure. I in 、E out and H out represent the input of STCM. Different from the original multi-head attention module, STCM performs dual-path attention operations on two separate streams: the image completion stream and the structure target stream. For the image completion stream, a Channel-Spatial Twin-Axis Attention Module is designed to capture the correlation between channels and spaces. STCM can perform self-attention on each stream to capture texture and target-specific structure, and STCM performs cross-attention on the two streams to fuse their interaction information.
[0019] Encode I in as the texture token of the image completion stream, and encode E out and H outThe structural tokens encoded as the structural target stream perform lightweight depth convolution projection on each feature map. Different from the patch-based MLP embedding method, this lightweight convolution can provide useful local perception bias for the transformer. Apply 3×3 depth convolution to the query, key, and value embeddings respectively, and transfer the Q t , K t , and V t are represented as textures to be completed. Q s , K s , and V s are represented as target structures. Transfer the structural information from the structural target stream to the image completion stream, and propose a residual addition method to achieve cross-attention, which is defined as:
[0020] K c = αK s + K t (2)
[0021] V c = βV s + V t (3)
[0022] where α and β are learnable scaling parameters used to control the fusion rate.
[0023] Use the structural target stream to improve the performance of the image completion stream. The cross-attention formula is as follows:
[0024] Attention t (Q t , K c , V c ) = V c · Softmax(K c · Q t / μ t ) (4)
[0025] Attention s (Q s , K s , V s ) = V s · Softmax(K s · Q s / μ s ) (5)
[0026] where μ t and μ s are learnable scaling parameters, and Attention t and Attention s are the attention maps of the structural target stream and the image completion stream respectively;
[0027] Connect the texture marker and the structure marker and input them into the feed-forward network. For the input of the next round, the obtained features are divided into two parts, namely structure features and texture features, according to the channels.
[0028] Preferably, the description of the CSPC is as follows: Channel-Spatial Two-Axis Attention Module (CSPC): Effectively fuse information from channels and space, and design a channel-spatial two-axis attention module; combine channel-wise attention and spatial window attention to form a biaxial self-attention mechanism. Given an input feature, divide it into two parts according to channels. On the channel axis, perform self-attention across channels. The channel-wise self-attention can be defined as:
[0029]
[0030] where represent the query, key, and value respectively, μ is a learnable scaling parameter, and the computational complexity of the channel-wise self-attention is O(C 2 WH), C 2 is a constant;
[0031] On the spatial axis, use spatial window attention to capture spatial dependencies. The window is obtained by equally dividing the image in a non-overlapping manner. Suppose there are N w different windows, and each window contains P w patches. The total number of patches obtained is: P = P w *N w The spatial window attention can be defined as:
[0032]
[0033] where Q i K i represent the query, key, and value of the spatial window attention respectively, C h represents the number of channels per head, and the computational complexity of the spatial window self-attention is O(2CPP w +4C 2 P), which is a linear complexity;
[0034] Use a feed-forward network to fuse the above features, and apply a 3×3 depth convolution to encode the input features, which helps to learn information about the local spatial context. Given the structure features generated by the structure target stream the channel features generated by the image completion stream and the spatial features This feed-forward network is represented as:
[0035] X = Concat(X st , X ch , Xsp ) (8)
[0036]
[0037]
[0038] where W p (·) represents a 1×1 pointwise convolution, and W d (·) represents a 3×3 depthwise convolution, ⊙ is element-wise multiplication, LN is layer normalization, ⊙ is the element-wise product of two parallel paths of the convolutional layer, and the feed-forward network can mix different features and control the information flow at each level, allowing each level to focus on complementing the details of other levels.
[0039] Preferably, the core objective of the Transformer method for the tracking structure in image inpainting is to design a Tracking Structure Transformer (TSFormer), which allows synchronous extraction of structural and texture features. Among them, the texture is extracted by tracking the structure, and the inpainted image is consistent in structure and texture, avoiding non-overlapping artifacts at the hole boundaries. A novel synchronous self-attention method is proposed to extract texture and structure in parallel, and a cross-attention method is proposed to allow them to interact. The overall framework of the proposed TSFormer consists of two networks: a Structure Enhancement Module (SEM) and a Synchronous Tracking Biaxial Transformer (STT). The SEM aims to restore the image structure, including the histograms of edge and oriented gradient (HOG) features. The proposed core network (STT) includes a Structure-Texture Synchronous Attention Module and a Channel-Spatial Biaxial Attention Module.
[0040] Preferably, the Tracking Structure Transformer (TSFormer) includes three core designs. Considering that the Histogram of Oriented Gradients (HOG) can depict the gradient direction distribution and edge direction of local sub-regions, HOG is first introduced in image inpainting, and a Structure Enhancement Module (SEM) is constructed to restore the overall image edges and HOG in the sketch space. Secondly, a Structure-Texture Cross-Attention Module (STCM) is proposed, which aims to track the image structure and perform intrinsic communication, allowing feature extraction to be more specific to structural targets. A gating mechanism is proposed to dynamically transmit structural information. In the synchronous module, a novel Channel-Spatial Biaxial Attention Module (CSPC) is proposed to allow effective co-learning of channel and spatial visual cues.
[0041] Preferably, the one tracking structure Transformer (TSFormer) includes a structure enhancement module (SEM) and a synchronous tracking biaxial Transformer (STT). In the SEM, Edge and Histogram of Oriented Gradients (HOG) are used as structural features to assist the STT network. In the STT network, a structure texture cross-attention module (STCM) is proposed to track the image structure and perform inherent communication, allowing feature extraction to be more specific to the structural target. And in the synchronous module, a novel channel space biaxial attention module (CSPC) is proposed to allow effective co-learning of channel and spatial visual cues.
[0042] Another technical problem to be solved by the present invention is to provide a Transformer for tracking structures for an image inpainting method, including the following steps:
[0043] S1: Let be the real image, M ∈ {0, 1} H×W×1 be the mask (the missing area is 0, otherwise 1), I in = I gt ⊙M represents the damaged image, Y m = Y gt ⊙M, H m = H gt ⊙M and E m = E gt ⊙M represent the missing gray, HOG, and Canny Edge images respectively;
[0044] S2: After splicing the above three images and inputting them into the SEM, the restored edge E out and H out features are used as sketch space vectors, and the formula is [E out , H out = SEM(E m , H m , Y m );
[0045] The structure enhancement module (SEM) restores the image edge and HOG as auxiliary structural features of the core STT. The input missing grayscale image Y m , HOG image H m and Canny edge E m are applied with a convolutional head to generate a feature map of 1 / 8 size, reducing the computational amount of standard self-attention. The channel-based self-attention captures global structural information in the low-resolution feature space, and the convolutional tail uses transposed convolution to upsample these features to the output structures E out and H out to optimize the predicted sketch structure: to optimize the predicted sketch structure:
[0046]
[0047] where E gt and H gt are the complete Edge and HOG images respectively. The binary cross-entropy (BCE) and l1 loss are used to reconstruct the sharp Edge and HOG features respectively. In the experiment, λ h = 0.1,
[0048] HOG carves the distribution of the gradient direction and the edge direction within the sub-region, which is achieved by subtracting adjacent pixels (gradient filtering). Its main feature is to capture the local shape and appearance and maintain good robustness to geometric changes. Even without precise knowledge of the corresponding gradient and edge positions, HOG can well represent the appearance and shape of local objects;
[0049] S3: STT connects the damaged image I in , the restored structure image H out and E out to finally generate the output image I out , and the formula is I out = STT(I in ,H out ,E out ), and the number of channels C = 24. The proposed Synchronous Tracking Transformer (STT) is a U-Net architecture following the encoder-decoder style. The structural information helps the initial contour recovery in the early stage of image inpainting. An encoder with 24 basic Transformer blocks is designed, and each block consists of a Structure-Texture Cross-Attention Module (STCM). Its image completion flow includes a Channel-Spatial Cross-Attention Module (CSPC). A decoder with 20 basic Transformer blocks is designed, and each block only contains CSPC. The description of the STCM is as follows: The restored structural features contain the complete gradient distribution and edge direction. The STCM (a key component of STT) is designed to synchronously capture the long-range dependencies on both structure and texture respectively. In addition to self-attention, the STCM introduces the cross-attention method to guide texture extraction by tracking the structure. I in 、E out and H outRepresents the input of STCM. Different from the original multi - head attention module, STCM performs dual - path attention operations on two separate streams: the image completion stream and the structure target stream. For the image completion stream, a channel - spatial biaxial attention module is designed to capture the correlation between channels and space. STCM can perform self - attention on each stream to capture textures and target - specific structures. STCM performs cross - attention on the two streams to fuse their interaction information.
[0050] Encode I in as texture tokens for the image completion stream, and encode E out and H out as structure tokens for the structure target stream. Perform lightweight depth - wise convolution projection on each feature map. Different from the patch - based MLP embedding method, this lightweight convolution can provide useful local perception bias for the transformer. Apply 3×3 depth - wise convolution to the query, key, and value embeddings respectively, and represent Q t , K t and V t as textures to be completed, and represent Q s , K s and V s as target structures. Transmit the structure information from the structure target stream to the image completion stream, and propose a residual addition method to achieve cross - attention, which is defined as:
[0051] K c =αK s +K t (2)
[0052] V c =βV s +V t (3)
[0053] where α and β are learnable scaling parameters used to control the fusion rate.
[0054] Utilize the structure target stream to improve the performance of the image completion stream. The cross - attention formula is as follows:
[0055] Attention t (Q t , K c , V c ) = V c ·Softmax(K c ·Q t / μ t ) (4)
[0056] Attention s (Q s , K s , V s) = V s · Softmax(K s · Q s / μ s ) (5)
[0057] where μ t and μ s are learnable scaling parameters, Attention t and Attention s are the attention maps of the structural target stream and the image completion stream respectively;
[0058] Connect the texture label and the structure label and input them into the feed-forward network. For the input of the next round, the obtained features are divided into two parts: structural features and texture features according to channels; The description of the CSPC is: Channel-Spatial Two-Axis Attention Module (CSPC): Effectively fuse information from channels and space, and design a channel-spatial two-axis attention module; Combine channel-wise attention and spatial window attention to form a biaxial self-attention mechanism. Given an input feature, divide it into two parts according to channels. On the axis of channels, perform self-attention across channels. The channel-wise self-attention can be defined as:
[0059]
[0060] where represent query, key, and value respectively, μ is a learnable scaling parameter, and the computational complexity of channel-wise self-attention is O(C 2 WH), C 2 is a constant;
[0061] On the spatial axis, use spatial window attention to capture spatial dependencies. The window is obtained by equally dividing the image in a non-overlapping manner. Assume there are N w different windows, and each window contains P w patches. The total number of patches obtained is: P = P w * N w The spatial window attention can be defined as:
[0062]
[0063] where Q i K i represent the query, key, and value of the spatial window attention respectively, C h represents the number of channels per head, and the computational complexity of spatial window self-attention is O(2CPP w + 4C 2 P), which is a linear complexity;
[0064] Fuse the above features using a feed - forward network and apply 3×3 depth convolution to encode the input features, which helps to learn information about the local spatial context, and the structural features generated by the given structural target stream The channel features generated by the image completion stream and the spatial features This feed - forward network is represented as:
[0065] X = Concat(X st , X ch , X sp ) (8)
[0066]
[0067]
[0068] where W p (·) represents a 1×1 point - wise convolution, W d(·) represents a 3×3 depthwise convolution, ⊙ is the element-wise multiplication, LN is the layer normalization, ⊙ is the element product of two parallel paths of the convolutional layer. The feed-forward network can mix different features and control the information flow at each level, allowing each level to focus on complementing the details of other levels; the core objective of the Transformer with the tracking structure for the image inpainting method is to design a Tracking Structure Transformer (TSFormer), which allows synchronous extraction of structural and texture features, where the texture is extracted by tracking the structure, and the inpainted image is consistent in structure and texture, avoiding non-overlapping artifacts at the hole boundaries. A novel synchronous self-attention method is proposed to extract texture and structure in parallel, and a cross-attention method is proposed to allow them to interact. The overall framework of the proposed TSFormer consists of two networks: the Structure Enhancement Module (SEM) and the Synchronous Tracking Biaxial Transformer (STT). The SEM aims to restore the image structure, including the histograms of edge and oriented gradient (HOG) features. The proposed core network (STT) includes a Structure Texture Synchronous Attention Module and a Channel Space Biaxial Attention Module; the Tracking Structure Transformer (TSFormer) includes three core designs. Considering that the Histogram of Oriented Gradients (HOG) can carve the gradient direction distribution and edge direction of local sub-regions, HOG is first introduced in image inpainting, and a Structure Enhancement Module (SEM) is constructed to restore the overall image edges and HOG in the sketch space. Secondly, a Structure Texture Cross-Attention Module (STCM) is proposed, which aims to track the image structure and perform intrinsic communication, allowing feature extraction to be more specific to the structural target. A gating mechanism is proposed to dynamically transmit structural information. In the synchronous module, a novel Channel Space Biaxial Attention Module (CSPC) is proposed to allow effective co-learning of channel and spatial visual cues; the Tracking Structure Transformer (TSFormer) includes the Structure Enhancement Module (SEM) and the Synchronous Tracking Biaxial Transformer (STT). In the SEM, Edge and the Histogram of Oriented Gradients (HOG) are used as structural features to assist the STT network. In the STT network, a Structure Texture Cross-Attention Module (STCM) is proposed, which aims to track the image structure and perform intrinsic communication, allowing feature extraction to be more specific to the structural target. And in the synchronous module, a novel Channel Space Biaxial Attention Module (CSPC) is proposed to allow effective co-learning of channel and spatial visual cues. Description of the Drawings
[0069] Figure 1 : Overview of the backbone network (TSFormer);
[0070] Figure 2 : Diagram of the Structure Texture Cross-Attention (STCM) module;
[0071] Figure 3 : Channel Space Pyramid Attention (CSPC) Module Diagram
[0072] Figure 4 : Comparison of the repair effect of the present invention on irregular holes with existing deep learning-based image repair techniques
[0073] Figure 5 : Comparison of the present invention with existing deep learning-based image repair techniques in face repair
[0074] Figure 6 : Comparison of the present invention with existing deep learning-based image repair techniques in building repair
[0075] (III) Beneficial Effects
[0076] Compared with the prior art, the present invention provides a Transformer with a tracking structure for an image repair method, having the following beneficial effects:
[0077] 1. For the image repair method of the Transformer with the tracking structure, the present invention designs an end-to-end tracking structure Transformer (TSFormer) for image repair, which includes a Structure Enhancement Module (SEM) and a Synchronous Tracking Pyramid Transformer (STT). Specifically, in the SEM, this patent uses Edge and Histogram of Oriented Gradients (HOG) as structural features to assist the STT network. In the STT network, this patent proposes a Structure Texture Cross-Attention Module (STCM) aimed at tracking image structures and performing intrinsic communication. This synchronization allows feature extraction to be more specific to structural targets. And in the synchronization module, this patent proposes a novel Channel Space Pyramid Attention Module (CSPC) to allow effective co-learning of channel and spatial visual cues.
[0078] 2. For the image repair method of the Transformer with the tracking structure, by using the histogram of the generated edges and oriented gradients (HOG) features in the missing area by this network as the sketch tensor space and utilizing HOG features in the image repair task, it can provide gradient direction or edge direction distribution for local sub-regions. A Synchronous Tracking Pyramid Transformer (STT) is designed for unified feature extraction and structural feature fusion.
[0079] 3. The Transformer of this tracking structure performs feature extraction and interaction of structural features for the image inpainting method. Self-attention is responsible for extracting features of the image texture or image structure area, and cross-attention enables them to transfer feature information to each other, making the feature extraction target the specified structural target. An incremental training strategy is adopted to dynamically transfer effective structural information to the inpainting model, and a low-complexity channel-space biaxial attention module is designed to capture channel and spatial interactions in parallel. Our design intention is to establish long-range relationships, which can be applied to the entire backbone network with linear complexity. Detailed implementation manners
[0080] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0081] Embodiment: S1: Let be the real image, M ∈ {0, 1} H×W×1 be the mask where the missing area is 0, otherwise 1, I in = I gt ⊙M represents the damaged image, Y m = Y gt ⊙M, H m = H gt ⊙M and E m = E gt ⊙M respectively represent the missing gray, HOG, and Canny Edge images;
[0082] S2: After splicing the above three images and inputting them into the SEM, the restored edge E out and H out features are used as the sketch space vectors, and the formula is [E out , H out = SANet(E m , H m , Y m );
[0083] The structure enhancement module (SEM) restores the image edge and HOG as the auxiliary structural features of the core STT. The input missing grayscale image Y m , HOG image H m and Canny edge E m, a convolutional head is applied to generate feature maps of 1 / 8 size, reducing the computational complexity of the standard self-attention. Channel-based self-attention captures global structural information in the low-resolution feature space. The convolutional tail uses transposed convolution to upsample these features to the output structure E out and H out , to optimize the predicted sketch structure:
[0084]
[0085] where E gt and H gt are the complete edge Edge and HOG images respectively. Binary cross-entropy (BCE) and l1 loss are used to reconstruct the sharp edge Edge and HOG features respectively. In the experiment, λ h = 0.1,
[0086] HOG carves the distribution of gradient directions and edge directions within sub-regions, achieved by subtracting adjacent pixel gradient filtering. Its main feature is to capture local shape and appearance, maintaining good robustness to geometric changes. Even without precise knowledge of the corresponding gradient and edge positions, HOG can well represent the appearance and shape of local objects;
[0087] S3: STT connects the damaged image I in , the restored structure image H out and E out to finally generate the output image I out , with the formula I out = STT(I in ,H out ,E out ), and the number of channels C = 24; The proposed Synchronous Tracking Biaxial Transformer (STT) follows the encoder-decoder architecture in a U-Net style. Structural information helps with the initial contour recovery in the early stage of image inpainting. An encoder with 24 basic Transformer blocks is designed, and each block consists of a Structure-Texture Cross-Attention Module (STCM). Its image completion sequence includes a Channel-Spatial Biaxial Attention Module (CSPC). A decoder with 20 basic Transformer blocks is designed, and each block only contains CSPC; The description of STCM is: The restored structural features contain the complete gradient distribution and edge directions. STCM is designed and is a key component of STT, which can synchronously capture the long-range dependencies on structure and texture respectively. In addition to self-attention, STCM introduces the cross-attention method to guide texture extraction by tracking the structure. I in 、E out and H outRepresents the input of STCM. Different from the original multi - head attention module, STCM performs dual - path attention operations on two separate streams: the image completion stream and the structure target stream. For the image completion stream, a channel - spatial biaxial attention module is designed to capture the correlation between channels and space. STCM can perform self - attention on each stream to capture texture and target - specific structures. STCM performs cross - attention on the two streams to fuse their interaction information.
[0088] Encode I in as texture tokens for the image completion stream, and encode E out and H out as structure tokens for the structure target stream. Perform lightweight depth - wise convolution projection on each feature map. Different from the patch - based MLP embedding method, this lightweight convolution can provide useful local perception bias for the transformer. Apply 3×3 depth - wise convolution to the query, key, and value embeddings respectively. Represent Q t , K t and V t as textures to be completed, and represent Q s , K s and V s as target structures. Transmit the structure information from the structure target stream to the image completion stream. A residual addition method is proposed to achieve cross - attention, which is defined as:
[0089] K c =αK s +K t (2)
[0090] V c =βV s +V t (3)
[0091] where α and β are learnable scaling parameters used to control the fusion rate.
[0092] Utilize the structure target stream to improve the performance of the image completion stream. The cross - attention formula is as follows:
[0093] Attention t (Q t ,K c ,V c )=V c ·Softmax(K c ·Q t / μ t ) (4)
[0094] Attention s (Q s ,K s ,V s) = V s ·Softmax(K s ·Q s / μ s ) (5)
[0095] where μ t and μ s are learnable scaling parameters, Attention t and Attention s are the attention maps of the structural target stream and the image completion stream respectively;
[0096] Connect the texture label and the structure label and input them into the feed-forward network. For the input of the next round, the obtained features are divided into two parts: structural features and texture features according to the channels; The description of CSPC is: Channel-Spatial Two-Axis Attention Module (CSPC): Effectively fuse information from channels and space, and design a channel-spatial two-axis attention module; Combine channel-wise attention and spatial window attention to form a biaxial self-attention mechanism. Given an input feature, divide it into two parts according to the channels. On the channel axis, perform self-attention across channels. The channel-wise self-attention can be defined as:
[0097]
[0098] where represent the query, key, and value respectively, μ is a learnable scaling parameter, and the computational complexity of the channel-wise self-attention is O(C 2 WH), C 2 is a constant;
[0099] On the spatial axis, use spatial window attention to capture spatial dependencies. The window is obtained by equally dividing the image in a non-overlapping manner. Suppose there are N w different windows, and each window contains P w patches. The total number of patches obtained is: P = P w *N w The spatial window attention can be defined as:
[0100]
[0101] where Q i K i represent the query, key, and value of the spatial window attention respectively, C h represents the number of channels per head, and the computational complexity of the spatial window self-attention is O(2CPP w +4C 2 P), which is a linear complexity;
[0102] Fuse the above features using a feed - forward network and apply 3×3 depth convolution to encode the input features, which helps to learn information about the local spatial context and the structural features generated by the given structural target stream Channel features generated by image completion and spatial features This feed - forward network is represented as:
[0103] X = Concat(X st , X ch , X sp ) (8)
[0104]
[0105]
[0106] where W p (·) represents 1×1 point - wise convolution, W d(·) represents a 3×3 depthwise convolution, ⊙ is element-wise multiplication, LN is layer normalization, ⊙ is the element-wise product of two parallel paths of the convolutional layer. The feed-forward network can mix different features and control the information flow at each level, allowing each level to focus on complementing the details of other levels; for the image inpainting method, the core goal of the Transformer with a tracking structure is to design a Tracking Structure Transformer (TSFormer), which allows for the synchronous extraction of structural and texture features, where the texture is extracted by tracking the structure, and the inpainted image is consistent in terms of structure and texture, avoiding non-overlapping artifacts at the hole boundaries. A novel synchronous self-attention method is proposed to extract texture and structure in parallel, and a cross-attention method is proposed to allow them to interact. The overall framework of the proposed TSFormer consists of two networks: a Structure Enhancement Module (SEM) and a Synchronous Tracking Biaxial Transformer (STT). The SEM aims to restore the image structure, including the histogram of Edge and oriented gradient (HOG) features. The proposed core network STT includes a Structure-Texture Synchronous Attention Module and a Channel-Spatial Biaxial Attention Module; a Tracking Structure Transformer (TSFormer), which includes three core designs. Considering that the Histogram of Oriented Gradients (HOG) can carve the gradient direction distribution and edge direction of local sub-regions, HOG is first introduced in image inpainting, and a Structure Enhancement Module (SEM) is constructed to restore the overall image edges and HOG in the sketch space. Secondly, a Structure-Texture Cross-Attention Module (STCM) is proposed, which aims to track the image structure and perform intrinsic communication, allowing feature extraction to be more specific to the structural target. A gating mechanism is proposed to dynamically transmit structural information. In the synchronous module, a novel Channel-Spatial Biaxial Attention Module (CSPC) is proposed to allow for the effective co-learning of channel and spatial visual cues; a Tracking Structure Transformer (TSFormer), which includes a Structure Enhancement Module (SEM) and a Synchronous Tracking Biaxial Transformer (STT). In the SEM, Edge and the histogram of oriented gradient HOG are used as structural features to assist the STT network. In the STT network, a Structure-Texture Cross-Attention Module (STCM) is proposed, which aims to track the image structure and perform intrinsic communication, allowing feature extraction to be more specific to the structural target, and in the synchronous module, a novel Channel-Spatial Biaxial Attention Module (CSPC) is proposed to allow for the effective co-learning of channel and spatial visual cues.
[0107] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made therein without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A Transformer method for a tracking structure in image inpainting, characterized in that, Including a tracking structure Transformer, the tracking structure Transformer includes a structure enhancement module (SEM) and a synchronous tracking biaxial Transformer (STT), and comprises the following steps: S1: Let be the real image, M ∈ {0, 1} H×W×1 be the mask, I in = I gt ⊙ M represents the damaged image, Y m = Y gt ⊙ M, H m = H gt ⊙ M and E m = E gt ⊙ M respectively represent the missing grayscale, HOG, and Canny Edge images; where E gt and H gt are the complete Edge and HOG images respectively, and ⊙ is the element-wise product of two parallel paths of the convolutional layer; S2: Input the three stitched images into the SEM to obtain the restored edge E out and the restored HOGH out features as the sketch space vectors, with the formula [E out , H out = SEM(E m , H m , Y m ); S3: The STT takes the damaged image I in , restores the structural image H out and E out to connect them, and finally generates the output image I out , with the formula I out = STT(I in , H out , E out ), and the number of channels C = 24; In S2, the SEM restored image edge and HOG are used as auxiliary structural features of the STT, and the missing grayscale image Y of the input m , the HOG image H m and the Canny edge E m , apply a convolutional head to generate a feature map of 1 / 8 size. The channel-based self-attention captures global structural information in the low-resolution feature space, and the convolutional tail uses transposed convolution to upsample these features to the output structures E out and H out , to optimize the predicted sketch structure: The binary cross-entropy (BCE) and l1 loss are used respectively to reconstruct the complete Edge and HOG features, and λ is taken as h 0.1 in the experiment. HOG carves the distribution of the gradient direction and the edge direction within the sub-region, and captures the local shape and appearance by subtracting adjacent pixels.
2. The Transformer method for a tracking structure used in image restoration according to claim 1, wherein In S3, STT is a U-Net architecture following the encoder-decoder style. An encoder with 24 basic Transformer blocks is designed, and each block consists of a structure texture cross-attention module (STCM). Its image completion stream includes a channel space biaxial attention module (CSPC). A decoder with 20 basic Transformer blocks is designed, and each block only contains CSPC.
3. The Transformer method for a tracking structure for image restoration according to claim 2, wherein, To restore the structural features with a complete gradient distribution and edge direction, STCM is designed to synchronously capture the long-range dependencies on both structure and texture respectively, including self-attention, and cross-attention is introduced to guide texture extraction by tracing the structure. I in 、E out and H out represent the inputs of STCM. STCM performs dual-path attention operations on two separate streams: the image completion stream and the structure target stream. For the image completion stream, a channel-spatial biaxial attention module is designed to capture the correlation between channels and space. STCM can perform self-attention on each stream to capture texture and target-specific structures, and STCM performs cross-attention on the two streams to fuse their interaction information. Encode I in as a texture marker for the image completion stream, and encode E out and H out as structure markers for the structure target stream. Perform lightweight depth convolution projection on each feature map, apply 3×3 depth convolution to the query, key, and value embeddings respectively, and represent Q t , K t and V t as textures to be completed. Represent Q s , K s and V s as target structures. Transmit the structure information from the structure target stream to the image completion stream, and propose a residual addition method to achieve cross-attention, which is defined as: K c = αK s + K t (2) V c = βV s + V t (3) Where α and β are learnable scaling parameters used to control the fusion rate. Using the structure target stream to improve the performance of the image completion stream, the cross-attention formula is as follows: Attention t (Q t ,K c ,V c ) = V c ·Softmax(K c ·Q t / μ t ) (4) Attention s (Q s ,K s ,V s ) = V s ·Softmax(K s ·Q s / μ s ) (5) where μ t and μ s are learnable scaling parameters, Attention t and Attention s are the attention maps of the structural target flow and the image completion flow respectively; Connect the texture token and the structure token and input them into the feed-forward network as the input for the next round. The obtained features are divided into two parts, namely structure features and texture features, according to the channels.
4. The Transformer method for a tracking structure for image restoration according to claim 3, characterized in that, The CSPC combines per-channel attention and spatial window attention to form a biaxial self-attention mechanism. Given an input feature, it is divided into two parts according to the channels. On the axis of the channels, self-attention is performed across channels. The per-channel self-attention can be defined as: where represent query, key, and value respectively, μ is a learnable scaling parameter, and the computational complexity of channel-wise self-attention is O(C 2 WH), C 2 is a constant; On the spatial axis, spatial window attention is used to capture spatial dependencies. The windows are obtained by equally dividing the image in a non-overlapping manner. There are N w different windows, and each window contains P w patches, resulting in the total number of patches: P = P w * N w Spatial window attention can be defined as: Among them respectively represent the query, key, and value of the spatial window attention, and C h represents the number of channels for each head. The computational complexity of the spatial window self-attention is O(2CPP w +4C 2 P), which is a linear complexity; Fuse the above features using a feed-forward network, and apply a 3×3 depth convolution to encode the input features, which helps to learn information about the local spatial context, and the structural features generated by the given structural target stream The channel features generated by the image completion stream and the spatial features This feed-forward network is represented as: X = Concat(X st , X ch , X sp ) (8) Among which W p (·) represents 1×1 pointwise convolution, W d (·) represents 3×3 depthwise convolution, LN is layer normalization, ⊙ is the element-wise product of two parallel paths of the convolutional layer, and the feed-forward network can mix different features and control the information flow at each level, allowing each level to focus on complementing the details of other levels.
Citation Information
Patent Citations
Incremental image restoration method based on wireframe and edge structure
CN114399436A
Image blind restoration method based on semantic inconsistency detection
CN114897738A