A Four-Branch Multi-Source Image Patch Matching Method Based on Scale Separation Metric Learning
The SSML-QNet model addresses the challenge of scale-specific feature encoding in multiple-source image matching by using a four-branch network with scale-separated metric learning, enhancing precision in VIS-NIR image matching.
Patent Information
- Application Number
- CN202310216819.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-08
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-03-08
AI Technical Summary
Existing multi-source image matching methods are difficult to effectively learn common and private features between image pairs, and cannot accurately measure the similarity of features of different scales, resulting in insufficient matching accuracy.
A four-branch multi-source image block matching method based on scale separation metric learning is adopted, and common and private features are extracted using twin and pseudo-twin networks, and different scale features are encoded and fused through scale separation metric learning module and multi-scale feature fusion module to predict the matching results.
The accuracy of multi-source image matching is improved, especially in visible-near-infrared image matching, FPR95 is significantly better than existing algorithms, reducing false positive rates and improving matching performance.
Smart Images

Figure CN116310438B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and specifically relates to a four-branch multi-source image patch matching method based on scale-separated metric learning. Background Art
[0002] Image matching aims to correspond pixel points representing the same physical position in the image to be matched and the reference image one by one. It is the premise of multi-source information fusion and collaborative analysis, and also the basis of multi-modal three-dimensional reconstruction and change detection, and has wide applications in military exploration, security and other fields. Due to the different imaging mechanisms between different sensors, multi-source images have strong appearance differences. For example, visible light images can provide more detailed information about the scene, but are greatly affected by illumination. Infrared images can reflect the thermal characteristics of the target, are insensitive to the brightness change characteristics of the scene, and can penetrate fog, haze, smoke, etc., but have low resolution. These significant differences make multi-source image matching extremely challenging.
[0003] Traditional methods based on manually designed descriptors rely heavily on human prior knowledge and easily ignore the pattern information hidden in the data. In recent years, image registration methods based on deep learning have made great progress due to the powerful feature representation and learning ability of deep neural networks.
[0004] According to the network output, existing methods can be divided into two categories: descriptor learning methods and metric learning methods. Descriptor learning methods use convolutional neural networks to extract high-level features from images, match images based on feature distances, and determine correct matching pairs when the matching distance is less than a pre-selected threshold. This method can automatically learn feature extractors according to image characteristics, without manual feature design, more flexible and adaptable. For example, Tian et al. proposed a descriptor learning model in the literature "L2-net: Deep learning of discriminative patch descriptor in euclidean space, IEEE conference on computer vision and pattern recognition. (2017), 661-669." This model adopts a progressive sampling strategy, enabling the network to access as many training samples as possible within a limited number of iterations. At the same time, L2 distance is used to constrain feature descriptors to reduce the risk of overfitting. Metric learning mainly uses original image pairs or generated feature descriptors as inputs, learns the similarity measurement relationship between image patches, converts the matching task into a binary classification task, and outputs the matching labels of image patch pairs. This method does not require the design of a measurement criterion and can directly learn the mapping from image patch pairs to matching labels through an end-to-end network structure. For example, Ko et al. proposed a domain-invariant image matching model in the literature "Spectral-invariant matching network. Information Fusion, (2023), 91: 623-632.", designed a domain transformation model to transform the input image from one domain to another, and used a dual Siamese network to extract discriminative features from the original image and the transformed image, achieving good performance in visible light and near-infrared image matching.
[0005] However, existing multi-source image matching research methods still have the following deficiencies: First, due to different imaging mechanisms between different sensors, multi-source images have large appearance differences, and existing methods are difficult to learn the common and private features between image pairs. Second, when existing methods measure similarity, they often consider features as a whole, and different scale features of the same image pair may have different similarities, and general measurement methods cannot accurately measure image similarity. Summary of the Invention
[0006] Technical Problems to be Solved
[0007] Aiming at the accuracy problem of multi-source image matching, the present invention provides a four-branch multi-source image patch matching method based on scale-separated metric learning for multi-source image matching.
[0008] Technical solution
[0009] A four-branch multi-source image patch matching method based on scale-separated metric learning, which is characterized by including a multi-modal feature extraction module, a scale-separated metric learning module, and a multi-scale feature fusion and prediction module; the steps are as follows:
[0010] Step 1: Input different modal image patches T1 and T2 with the number of channels and resolution size of C×H×W into the twin sub-network and pseudo-twin sub-network in the multi-modal feature extraction module to learn the common features and private features of the images; where the twin sub-network represents two branches with shared parameters and the same structure, and the pseudo-twin sub-network represents two branches with non-shared parameters and the same structure; then stack the four feature maps with high-level semantic information extracted along the channel dimension to obtain a feature vector F, with the number of channels and resolution size of 512×H / 4×W / 4.
[0011] Step 2: Input the feature vector F into the scale-separated metric learning module and split it through convolutional kernels of 3×3, 5×5, 7×7, and 9×9 to obtain four groups of different-scale feature maps (F0, F1, F2, F3), and the number of channels of each group of feature maps becomes 1 / 4 of the original, and the resolution remains unchanged; then each scale feature sequentially performs coordinate attention CA and channel attention SE, encodes the correlation at each group of coordinates and channels, and stacks them along the channel dimension to obtain a feature vector F ssml As the output, with a size of 512×H / 4×W / 4;
[0012] Step 3: Input the feature vector F learned by the scale-separated metric learning module ssml into the multi-scale feature fusion and prediction module, sequentially pass through coordinate attention and 1×1 convolution, and the obtained feature map is added to itself in a residual manner to generate the final fused feature descriptor, with a size of 8×H / 4×W / 4; finally, use three fully connected layers with sizes of 512, 128, and 2 respectively to predict the final matching result of the network.
[0013] A further technical solution of the present invention: In step 1, each branch includes a convolutional layer 1, an instance normalization layer 1, and a batch normalization layer 1; a convolutional layer 2, an instance normalization layer 2, a batch normalization layer 2, and a max pooling layer 2; a convolutional layer 3, an instance normalization layer 3, and a batch normalization layer 3; a convolutional layer 4, a batch normalization layer 4; a convolutional layer 5, a batch normalization layer 5, and a max pooling layer 5; a convolutional layer 6, a batch normalization layer 6; a convolutional layer 7, a batch normalization layer 7. Among them, all convolutional layers use 3×3 convolutional kernels, the stride of convolution is 1, and the stride of the pooling layer is 2.
[0014] Further technical solution of the present invention: The specific process of sequentially executing coordinate attention CA at each scale in step 2 is as follows: First, perform spatial global average pooling on four groups of feature maps with different scales (F0, F1, F2, F3) along the X coordinate direction and the Y coordinate direction respectively to generate two feature maps with sizes of C×H×1 and C×1×W respectively, where C represents the number of channels, and H and W represent the height and width of the image respectively; Second, stack along the channel dimension and then pass through a 1×1 convolution, batch normalization, and Sigmoid activation function to learn the weights of objects at different scales, generating a feature map with a size of C×1×(H×W); Then, perform 1×1 convolution and Sigmoid activation function on it again along the X coordinate direction and the Y coordinate direction respectively to obtain the attention weights in the X coordinate direction and the attention weights in the Y coordinate direction, with sizes of C×H×1 and C×1×W respectively; Finally, multiply the weight vector by the input feature map to obtain the final feature map with spatial information.
[0015] A computer system, characterized by comprising: one or more processors, a computer-readable storage medium for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the above method.
[0016] A computer-readable storage medium, characterized by storing computer-executable instructions, which are used to implement the above method when executed.
[0017] Beneficial effects
[0018] A four-branch multi-source image patch matching method based on scale separation metric learning provided by the present invention designs a high-precision multi-source image patch matching network model SSML-QNet. A four-branch network composed of twins and pseudo-twins is used to extract common features and private features between cross-modal images. A scale separation metric learning module is designed to encode different-scale features respectively, and cross-modal consistency features are extracted through coordinate and channel attention at each scale. Then, a multi-scale feature fusion module composed of coordinate attention and convolution is used to fuse different-scale features and predict the matching result. In terms of accuracy, the SSML-QNet of the present invention is significantly better than existing algorithms in the VIS-NIR multi-modal dataset, and the FPR95 reaches 0.73.
[0019] The present invention can extract common features and private features in multi-source images, and eliminate the influence caused by interference factors such as sensor imaging mechanism, illumination, and shadow.
[0020] The model proposed by the present invention can encode the similarity of different-scale features between the same group of image patch pairs, and solve the influence caused by factors such as scale. Brief description of the drawings
[0021] The accompanying drawings are only for the purpose of showing specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference numerals represent the same components.
[0022] Figure 1 is the network structure diagram of an embodiment of the present invention;
[0023] Figure 2 is the structure diagram of the scale separation metric learning module of the present invention;
[0024] Figure 3 is the structure diagram of the multi-scale feature fusion and prediction module of the present invention;
[0025] Figure 4 is the structure diagram of the coordinate attention of the present invention;
[0026] Figure 5 is the structure diagram of the channel attention of the present invention. Specific Embodiments
[0027] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0028] The present invention proposes a four-branch multi-source image patch matching method based on scale separation metric learning, and designs a high-precision multi-source image patch matching network model SSML-QNet, which includes a four-branch multi-modal feature extraction module, a scale separation metric learning module, and a multi-scale feature fusion and prediction module. The four-branch multi-modal feature extraction module uses twin sub-networks and pseudo-twin sub-networks to extract the common features and private features between the input multi-modal image pairs T1 and T2, where the twin sub-networks represent two branches with shared parameters and the same structure, and the pseudo-twin sub-networks represent two branches with non-shared parameters and the same structure. Then, the four feature maps with high-level semantic information extracted are stacked along the channel dimension to obtain the feature vector F. The scale separation metric learning module is used to split the feature vector F through convolution kernels of 3×3, 5×5, 7×7, and 9×9 to obtain four groups of different-scale feature maps (F0, F1, F2, F3), and then coordinate attention (CA) and channel attention (SE) are sequentially performed on each scale to encode the correlation at each coordinate and channel level and stack them along the feature channels to obtain the feature vector F ssml as the output. The multi-scale feature fusion and prediction module is used to process the feature vector F output by the scale separation metric learning module ssmlAfter fusing through coordinate attention and 1×1 convolution, it is then added to itself in a residual manner to generate the final fused feature descriptor. Finally, three fully connected layers with sizes of 512, 128, and 2 are used to predict the final matching result of the network.
[0029] It includes the following steps:
[0030] Step 1: Input the image patches T1 and T2 of different modalities into the twin sub-networks and pseudo-twin sub-networks of the four-branch multi-modal feature extraction module to learn the common features and private features of the images. Among them, the twin sub-networks represent two branches with shared parameters and the same structure, and the pseudo-twin sub-networks represent two branches with non-shared parameters and the same structure. Then, the four feature maps with high-level semantic information extracted are stacked along the channel dimension to obtain the feature vector F.
[0031] Step 1-1: The four-branch multi-modal feature extraction module includes four identical sub-branches. The structure of each branch is: convolutional layer 1, instance normalization layer 1, batch normalization layer 1; convolutional layer 2, instance normalization layer 2, batch normalization layer 2, max pooling layer 2; convolutional layer 3, instance normalization layer 3, batch normalization layer 3; convolutional layer 4, batch normalization layer 4; convolutional layer 5, batch normalization layer 5, max pooling layer 5; convolutional layer 6, batch normalization layer 6; convolutional layer 7, batch normalization layer 7. Among them, all convolutional layers use 3×3 convolutional kernels, and the stride of convolution is 1.
[0032] Step 2: Construct a scale separation metric learning module: The module passes the feature vector F output by the four-branch multi-modal feature extraction module through the scale separation metric learning module to output a discriminative feature vector F ssml .
[0033] Step 2-1: The scale separation metric learning module splits the feature vector F successively through convolutional kernels of 3×3, 5×5, 7×7, and 9×9 to obtain four groups of feature maps with different scales (F0, F1, F2, F3), and the number of channels of each group of feature maps becomes 1 / 4 of the original.
[0034] F0,F1,F2,F3 = f 3×3 (F),f 5×5 (F),f 7×7 (F),f 9×9 (F)(2)
[0035] Among them, f 3×3 , f 5×5 , f 7×7 , f 9×9 represent convolutional layers of 3×3, 5×5, 7×7, and 9×9 respectively.
[0036] Step 2-2: The four groups of feature maps (F0, F1, F2, F3) respectively pass through coordinate attention (CA) to learn the spatial layout information of the image. Coordinate attention first performs average pooling along the X coordinate direction and the Y coordinate direction respectively to generate two feature maps with sizes of C×H×1 and C×1×W respectively, where C represents the number of channels, and H and W represent the height and width of the image respectively. Secondly, stack along the channel dimension and then pass through a 1×1 convolution, batch normalization, and Sigmoid activation function to generate a feature map with a size of C×1×(H+W); then, perform 1×1 convolution and Sigmoid activation function along the X coordinate direction and the Y coordinate direction respectively again to obtain the attention weights in the X coordinate direction and the attention weights in the Y coordinate direction, with sizes of C×H×1 and C×1×W respectively. Finally, multiply the weight vector by the input feature map to obtain the feature map with spatial information.
[0037]
[0038]
[0039] where F(i,j) represents the feature vector output by the four-branch multi-modal feature extraction module; H, W represent the width and height of the image; f 1×1 and respectively represent a 1×1 convolution layer and element-wise multiplication; δ represents batch normalization; σ represents the Sigmoid activation function; T h , T w represent encoding each channel along the X coordinate direction and the Y coordinate direction respectively.
[0040] Step 2-3: The four groups of feature maps with spatial information respectively pass through channel attention (SE) to encode and output the channel-level correlation of each group, that is, global average pooling, fully connected layer, ReLU activation function, fully connected layer, and Sigmoid activation function. Further obtain the channel attention vector, and the attention vector represents the correlation between feature maps. The dimension sizes of the network input and output remain unchanged.
[0041]
[0042] where CA(i,j) represents the feature map after coordinate attention encoding; H, W represent the width and height of the image; FC represents the fully connected layer; γ represents the ReLU activation function; σ represents the Sigmoid activation function
[0043] Step 3: The feature vector F extracted by the scale separation metric learning module ssmlAfter passing through coordinate attention and 1×1 convolution in sequence, the obtained feature map is added to itself in a residual manner to generate the final fused feature descriptor. Finally, three fully connected layers with sizes of 512, 128, and 2 are used to predict the final matching result of the network.
[0044] When training the network, the cross-entropy loss L is calculated using the prediction result and the true matching label. bc , and then backpropagation is performed, repeating the iteration until the iteration times reach the set initial value to determine that the training is completed.
[0045]
[0046] where represents the prediction result, and y represents the true label.
[0047] To enable those skilled in the art to better understand the present invention, the present invention will be described in detail below with reference to specific embodiments.
[0048] Embodiment:
[0049] As Figure 1 shown, in view of the problem of insufficient accuracy of the current multi-source image matching result, the present invention designs a model of a four-branch multi-source image block matching method based on scale-separated metric learning. It includes three parts: a four-branch multi-modal feature extraction module, a scale-separated metric learning module, and a multi-scale feature fusion and prediction module, as shown in Figure 1 , 2 , and Figure 3 respectively. The specific method includes the following steps:
[0050] Step 1, input the image blocks T1 and T2 of different modalities into the twin sub-network and pseudo-twin sub-network of the four-branch multi-modal feature extraction module to extract the common features and private features of the images. The twin sub-network represents two branches with shared parameters and the same structure, and the pseudo-twin sub-network represents two branches with non-shared parameters and the same structure. Then, the four feature maps with high-level semantic information extracted are stacked along the channel dimension to obtain the feature vector F.
[0051] Step 2, split the feature vector F through convolution kernels of 3×3, 5×5, 7×7, and 9×9 to obtain four groups of feature maps with different scales (F0, F1, F2, F3). Then, coordinate attention (CA) and channel attention (SE) are sequentially performed on each scale to encode the coordinate and channel-level correlations of each group and stack them along the channel to obtain the feature vector F ssml as the output.
[0052] Step 3, the feature vector F extracted by the scale-separated metric learning module ssmlAfter passing through coordinate attention and 1×1 convolution in sequence, the obtained feature map is added to itself in a residual manner to generate the finally fused feature descriptor. Finally, three fully connected layers with sizes of 512, 128, and 2 are used to predict the final matching result of the network.
[0053] In this embodiment, the execution network of steps 1-3 is abbreviated as SSML-QNet. The execution processes of steps 1-step 3 will be further described in detail below in combination with the structure of SSML-QNet.
[0054] In this embodiment, in step 1, the multi-source image blocks T1 and T2 with a size of H×W×3 are respectively input into the twin sub-network and the pseudo-twin sub-network to learn the common features and private features of the images. The twin sub-network refers to a two-branch structure with shared parameters, and the pseudo-twin sub-network refers to a two-branch structure without shared parameters. The four branch structures are the same. The input images of each branch sequentially pass through convolutional layer 1, instance normalization layer 1, batch normalization layer 1; convolutional layer 2, instance normalization layer 2, batch normalization layer 2, max pooling layer 2; convolutional layer 3, instance normalization layer 3, batch normalization layer 3; convolutional layer 4, batch normalization layer 4; convolutional layer 5, batch normalization layer 5, max pooling layer 5; convolutional layer 6, batch normalization layer 6; convolutional layer 7, batch normalization layer 7. Among them, all convolutional layers use 3×3 convolutional kernels, and the stride of convolution is 1. The four feature maps with high-level semantic information extracted are stacked along the channels to obtain the feature vector F, whose feature resolution and number of channels are H×W×512.
[0055] See Figure 1 , in step 2 of this embodiment, the feature vector F output by the four-branch multi-modal feature extraction module is sent into the scale separation metric learning module, and the scale separation metric learning module is as Figure 2 shown. First, the feature vector F is split by convolutional kernels of 3×3, 5×5, 7×7, and 9×9 in sequence to obtain four groups of feature maps with different scales (F0, F1, F2, F3), and the number of channels of each group of feature maps becomes 1 / 4 of the original. Then, each scale feature sequentially performs coordinate attention (CA) and channel attention (SE) to encode the correlation at each coordinate and channel level and stack them along the channels to obtain the feature vector F ssml as the output.
[0056] In this embodiment, see Figure 4, Coordinate attention can capture long-range dependencies with precise position information. First, spatial global average pooling is performed separately in the X-coordinate direction and the Y-coordinate direction to generate two feature maps with sizes of C×H×1 and C×1×W respectively. Secondly, they are stacked along the channel dimension and passed through a 1×1 convolution, batch normalization, and the Sigmoid activation function to generate a feature vector of size C×1×(H + W). Then, through 1×1 convolution and the Sigmoid activation function in the X-coordinate direction and the Y-coordinate direction respectively, the attention weights in the X-coordinate direction and the Y-coordinate direction are obtained, with sizes of C×H×1 and C×1×W respectively. Finally, the weight vector is multiplied by the input feature map to obtain a feature vector with spatial information.
[0057] In this embodiment, referring to Figure 5 , Channel attention (SE) can encode and output the channel-level correlations of each group to obtain the correlations between feature maps. This module includes global average pooling, fully connected layers, the ReLU activation function, fully connected layers, and the Sigmoid activation function. The dimension sizes of the network input and output remain unchanged.
[0058] In this embodiment, referring to Figure 3 , Step 3 fuses the feature vector F ssml extracted by the scale separation metric learning module through coordinate attention and 1×1 convolution. The obtained feature map is added to itself in a residual manner to generate the final fused feature descriptor. Finally, three fully connected layers with sizes of 512, 128, and 2 are used to predict the final matching result of the network.
[0059] To verify the effectiveness of the SSML-QNet network, this embodiment uses the publicly available dataset VIS-NIR (visible light - near infrared) for the training and testing of the network framework and compares it with other methods. The VIS-NIR dataset contains 9 categories: countryside, field, forest, indoor, mountain, old building, street, city, and lake, with more than 1 million pairs of image patches. This embodiment uses the countryside category for training and the other eight categories for testing, and the size of all image patches is 64×64 pixels.
[0060] The algorithm proposed in this embodiment is compared with 9 mainstream multi-source image registration methods. The specific results are shown in Table 1. The evaluation index is FPR95, which specifically refers to the false positive rate (FPR95) when the true positive rate (positive recall rate) is equal to 95%, and is used to quantitatively evaluate the matching performance of the network. The lower the index, the better the performance.
[0061] Table 1 Comparison of the method proposed in this embodiment with other mainstream multi-source image patch matching methods
[0062]
[0063] As can be seen from Table 1, the algorithm proposed in this embodiment is significantly better than other comparison algorithms. Compared with the existing best publicly available method MFD-Net, the average value of FPR95 of this method is reduced by 0.15. The results show that by using the proposed scale separation metric learning module and fusion strategy, this network can effectively extract and measure the similarity of multimodal image patches.
[0064] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or replacements, and these modifications or replacements should be covered within the protection scope of the present invention.
Claims
1. A four-branch multi-source image patch matching method based on scale separation metric learning, characterized in that It includes a multi-modal feature extraction module, a scale separation metric learning module, and a multi-scale feature fusion and prediction module; the steps are as follows: Step 1: Input different modal image patches T1 and T2 with the number of channels and resolution size of C×H×W into the twin sub-networks and pseudo-twin sub-networks in the multi-modal feature extraction module respectively to learn the common features and private features of the images; Among them, the twin sub-networks represent two branches with shared parameters and the same structure, and the pseudo-twin sub-networks represent two branches with non-shared parameters and the same structure; then stack the four feature maps with high-level semantic information extracted along the channel dimension to obtain a feature vector F, with the number of channels and resolution size of 512×H / 4×W / 4; Step 2: Input the feature vector F into the scale separation metric learning module and split it through convolutional kernels of 3×3, 5×5, 7×7, and 9×9 to obtain four groups of feature maps with different scales (F0, F1, F2, F3). The number of channels of each group of feature maps becomes 1 / 4 of the original, and the resolution remains unchanged. Then, coordinate attention CA and channel attention SE are sequentially performed on each scale feature to encode the coordinate and channel-level correlations of each group and stack them along the channel dimension to obtain the feature vector F ssml As the output, the size is 512×H / 4×W / 4; Step 3: Input the feature vector F learned by the scale separation metric learning module ssml into the multi-scale feature fusion and prediction module. It passes through coordinate attention and 1×1 convolution in sequence. The resulting feature map is added to itself in a residual manner to generate the final fused feature descriptor, with a size of 8×H / 4×W / 4. Finally, three fully connected layers with sizes of 512, 128, and 2 are used to predict the final matching result of the network.
2. The four-branch multi-source image patch matching method based on scale separation metric learning according to claim 1, characterized in that: Each branch in Step 1 includes a convolutional layer 1, an instance normalization layer 1, and a batch normalization layer 1; a convolutional layer 2, an instance normalization layer 2, a batch normalization layer 2, and a max pooling layer 2; a convolutional layer 3, an instance normalization layer 3, and a batch normalization layer 3; a convolutional layer 4, and a batch normalization layer 4; a convolutional layer 5, a batch normalization layer 5, and a max pooling layer 5; a convolutional layer 6, and a batch normalization layer 6; a convolutional layer 7, and a batch normalization layer 7; where the convolutional layers all use 3×3 convolutional kernels, the stride of the convolution is 1, and the stride of the pooling layer is 2.
3. The four-branch multi-source image patch matching method based on scale separation metric learning according to claim 1, characterized in that: The specific process of coordinate attention CA is executed sequentially for each scale in Step 2 as follows: First, perform spatial global average pooling on the four groups of feature maps with different scales (F0, F1, F2, F3) along the X coordinate direction and the Y coordinate direction respectively to generate two feature maps with sizes of C×H×1 and C×1×W respectively, where C represents the number of channels, and H and W represent the height and width of the image respectively; Second, stack along the channel dimension and then pass through a 1×1 convolution, batch normalization, and Sigmoid activation function to learn the weights of objects at different scales, generating a feature map with a size of C×1×(H+W); Then, perform 1×1 convolution and Sigmoid activation function on it again along the X coordinate direction and the Y coordinate direction respectively to obtain the attention weights in the X coordinate direction and the attention weights in the Y coordinate direction, with sizes of C×H×1 and C×1×W respectively; Finally, multiply the weight vectors by the input feature maps to obtain the final feature maps with spatial information.
4. A computer system, characterized in that It includes: One or more processors, a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in claim 1.
5. A computer-readable storage medium, characterized in that Stored with computer-executable instructions, the instructions are used to implement the method described in claim 1 when executed.
Citation Information
Patent Citations
Feature decomposition method and system for multi-modal image block matching
CN113221923A
Remote sensing image change detection method based on twinborn multi-scale difference feature fusion
CN113420662A