Feature matching method based on differential wavelet transform and dynamic multi-scale feature fusion
Through the feature matching method of differential wavelet transformation and dynamic multi-scale features, the accuracy problem of image matching in complex backgrounds is solved, and the feature matching effect with high accuracy and robustness is achieved.
Patent Information
- Application Number
- CN202510767462.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The prior art has noise and blurring of key features in image matching under complex backgrounds, resulting in a decrease in matching accuracy. Especially under the influence of factors such as lighting changes and viewing angle changes, it is difficult to accurately match image features.
A feature matching method based on the fusion of differential wavelet transform and dynamic multi-scale feature is adopted. By constructing a feature matching model, including backbone network, key point detection network and descriptor generation network, differential wavelet transform is used to separate noise and key features, combining dynamic multi-scale feature fusion and adaptive multi-scale descriptor enhancement module, the feature matching accuracy and robustness are improved.
Effectively remove redundant information, enhance key features, improve feature matching accuracy and stability, adapt to scale changes under different perspectives, improve the quality and expression ability of descriptors, and achieve high-precision feature matching.
Smart Images

Figure CN120279288A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of feature matching, and specifically to a feature matching method based on differential wavelet transform and dynamic multi-scale feature fusion. Background Art
[0002] In the field of computer vision, image matching is a crucial task and often serves as the primary step for many complex vision tasks, being widely applied in fields such as 3D reconstruction, motion estimation, image stitching, and object detection. The purpose of image feature matching is to accurately match the feature points in different images. Usually, these feature points include corner points, edge points, etc., to find the corresponding relationships between images. According to images under different imaging conditions, image matching can be classified into the following categories: image matching based on different imaging times, image matching based on different perspectives, and template-based image matching.
[0003] Feature matching is an image matching method that can extract features in pictures (such as point features, line features, etc.). By matching the features in the images, the spatial and structural relationships between the images can be inferred. In many computer vision tasks with high-precision requirements, the feature matching accuracy is crucial and will affect the effects of these computer vision tasks.
[0004] The development of deep learning has brought a huge breakthrough to feature matching technology. The feature matching method based on deep learning can learn the deep features of the target object by constructing a deep neural network architecture and optimize them by fusing high-level semantic features, breaking through the limitations of traditional methods in feature extraction and performing better in terms of matching accuracy and robustness. However, due to differences in the source time, shooting perspective, acquisition sensor, etc. of the images to be matched, the difficulty of image matching has increased. In these environments with complex backgrounds and application scenarios, the existing technologies have deficiencies in distinguishing noise and key features, resulting in the key features being blurred and reducing the matching accuracy. Therefore, how to reduce the incorrect matching caused by problems such as light changes and perspective changes is an urgent problem to be solved in feature matching. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the present invention provides a feature matching method based on differential wavelet transform and dynamic multi-scale feature fusion, aiming to solve the problems in the background art.
[0006] To achieve the above object, the present invention provides the following technical solutions: A feature matching method based on differential wavelet transform and dynamic multi-scale feature fusion, comprising the following steps: reading an image to be matched; constructing a feature matching model; the feature matching model includes a backbone network, a key point detection network, and a descriptor generation network; the backbone network consists of a feature extraction module and a sixth convolutional block; inputting the image to be matched into the feature extraction module in the backbone network to generate a feature map, passing the feature map through the sixth convolutional block to generate a reliability map; inputting the image to be matched into the key point detection network to generate a key point heat map, inputting the key point heat map and the reliability map into the descriptor generation network to screen out key points, and finally extracting the feature descriptors of the image to be matched corresponding to the screened key points from the feature map, processing the feature descriptors of the image to be matched through an adaptive multi-scale descriptor enhancement module to obtain enhanced descriptors of the image to be matched, and finally performing brute-force matching on the enhanced descriptors of the image to be matched; setting the number of images to be matched read as two, and finally performing brute-force matching on the enhanced descriptors of the two images to be matched obtained.
[0007] Further, the feature extraction module includes a first convolutional block, a second convolutional block, a third convolutional block, a fourth convolutional block, and a fifth convolutional block; passing the image to be matched through the first convolutional block and the second convolutional block respectively, adding the outputs of the first convolutional block and the second convolutional block to obtain an added feature, and passing the added feature through the third convolutional block, the fourth convolutional block, and the fifth convolutional block in sequence to obtain a feature map F; Among them, the first convolutional block consists of a first convolutional layer, a first differential wavelet transform downsampling module, a first dynamic multi-scale feature fusion module, a second dynamic multi-scale feature fusion module, and a second differential wavelet transform downsampling module connected in sequence; passing the image to be matched through the first convolutional layer, the first differential wavelet transform downsampling module, the first dynamic multi-scale feature fusion module, the second dynamic multi-scale feature fusion module, and the second differential wavelet transform downsampling module in sequence to obtain the output of the first convolutional block; Among them, the second convolutional block consists of a first average pooling layer and a second convolutional layer connected in sequence; passing the image to be matched through the first average pooling layer and the second convolutional layer in sequence to obtain the output of the second convolutional block; Among them, the third convolutional block consists of a second dynamic multi-scale feature fusion module and a third dynamic multi-scale feature fusion module connected in sequence; passing the added feature through the second dynamic multi-scale feature fusion module and the third dynamic multi-scale feature fusion module in sequence to obtain the output of the third convolutional block; Among them, the fourth convolutional block is composed of a third difference wavelet transform downsampling module, a fourth dynamic multi-scale feature fusion module, and a fifth dynamic multi-scale feature fusion module connected in sequence; the output of the third convolutional block is successively passed through the third difference wavelet transform downsampling module, the fourth dynamic multi-scale feature fusion module, and the fifth dynamic multi-scale feature fusion module to obtain the output of the fourth convolutional block. Among them, the fifth convolutional block is composed of a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer connected in sequence; the output of the fourth convolutional block is successively passed through the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer to obtain the output of the fifth convolutional block. Among them, the sixth convolutional block is composed of a sixth convolutional layer, a seventh convolutional layer, and an eighth convolutional layer connected in sequence; the feature map F is successively passed through the sixth convolutional layer, the seventh convolutional layer, and the eighth convolutional layer and then passed through the sigmoid function to obtain the reliability map R.
[0008] Furthermore, the first difference wavelet transform downsampling module, the second difference wavelet transform downsampling module, the third difference wavelet transform downsampling module, and the fourth difference wavelet transform downsampling module have the same structure; the first difference wavelet transform downsampling module includes a ninth convolutional layer, a discrete wavelet transform module, a tenth convolutional layer, an edge information enhancement module, and an eleventh convolutional layer. The processing flow of the first difference wavelet transform downsampling module is as follows: The output of the first convolutional layer is passed through the ninth convolutional layer, and the output of the ninth convolutional layer is input into the discrete wavelet transform module for frequency domain decomposition to obtain low-frequency information and high-frequency information. The high-frequency information is further decomposed into horizontal high-frequency information, vertical high-frequency information, and diagonal high-frequency information. The horizontal high-frequency information, vertical high-frequency information, and low-frequency information are concatenated along the channel dimension to obtain a high-frequency information feature map. The high-frequency information feature map is successively passed through the tenth convolutional layer and the edge information enhancement module. The output of the edge information enhancement module and the output of the tenth convolutional layer are multiplied. The multiplied output is then passed through the eleventh convolutional layer to obtain the output of the first difference wavelet transform downsampling module.
[0009] Furthermore, the edge information enhancement module includes a second average pooling layer and a twelfth convolutional layer; the processing flow of the edge information enhancement module is as follows: the output of the tenth convolutional layer is passed through the second average pooling layer, the output of the second average pooling layer is subtracted by the output of the tenth convolutional layer. The subtracted output is first adjusted by the twelfth convolutional layer and then activated by the Sigmoid function to obtain a weight map. The weight map is added by 1 to obtain the output of the edge information enhancement module.
[0010] Furthermore, the first dynamic multi-scale feature fusion module, the second dynamic multi-scale feature fusion module, the third dynamic multi-scale feature fusion module, the fourth dynamic multi-scale feature fusion module, and the fifth dynamic multi-scale feature fusion module have the same structure; the first dynamic multi-scale feature fusion module includes a first depthwise separable convolution, a Split function, a second depthwise separable convolution, a third depthwise separable convolution, a thirteenth convolutional layer, a fourteenth convolutional layer, a fifteenth convolutional layer, and a SiLU activation function.
[0011] Furthermore, the processing flow of the first dynamic multi-scale feature fusion module is as follows: the outputs of the first differential wavelet transform downsampling module are respectively passed through the first depthwise separable convolution and the fourteenth convolutional layer. The output of the first depthwise separable convolution is divided into a first group of features, a second group of features, and a third group of features in proportion in the channel dimension through the Split function. The second group of features and the third group of features are respectively passed through the second depthwise separable convolution and the third depthwise separable convolution. The outputs of the second depthwise separable convolution and the third depthwise separable convolution are concatenated with the first group of features in the channel dimension. The concatenated result is input into the thirteenth convolutional layer. The outputs of the thirteenth convolutional layer and the fourteenth convolutional layer are both multiplied after passing through a SiLU activation function. The multiplied result passes through the fifteenth convolutional layer to obtain the output of the first dynamic multi-scale feature fusion module.
[0012] Furthermore, the specific process of inputting the key point heat map K and the reliability map R into the descriptor generation network to screen out key points, and finally extracting the feature descriptors of the corresponding images to be matched from the feature map F according to the screened key points is as follows: non-maximum suppression is performed on the key point heat map K, and the local response peak points are selected as candidate key points; then the response values of the candidate key points are selected from the key point heat map K using nearest neighbor interpolation, and the feature responses of the candidate key points are selected from the reliability map R using bilinear interpolation. The response values of the candidate key points selected from the key point heat map K and the feature responses of the candidate key points selected from the reliability map R are multiplied to obtain a comprehensive score. According to the sorting and screening of the comprehensive scores, key points are obtained, and then the corresponding feature descriptors of the images to be matched are extracted from the feature map F using bilinear interpolation according to the screened key points.
[0013] Furthermore, the adaptive multi-scale descriptor enhancement module includes a dynamic hyperbolic tangent function, a first Mamba layer, a second Mamba layer, a third Mamba layer, a linear layer, and a GeLU activation function; the processing flow of the adaptive multi-scale descriptor enhancement module is as follows: the feature descriptor of the image to be matched is processed by the dynamic hyperbolic tangent function, and the output after passing through the dynamic hyperbolic tangent function is added to the feature descriptor of the image to be matched to obtain a first added descriptor. The first added descriptor is respectively input into the first Mamba layer, the second Mamba layer, and the third Mamba layer. The outputs of the first Mamba layer, the second Mamba layer, and the third Mamba layer are averaged and fused. The output after average fusion is added to the first added descriptor to obtain a second added descriptor. The second added descriptor is input into the linear layer for processing. The output of the linear layer is concatenated with the second added descriptor to obtain a third added descriptor. The output after passing the third added descriptor through the GeLU activation function is concatenated with the third added descriptor again to obtain the output of the adaptive multi-scale descriptor enhancement module, that is, the enhanced descriptor of the image to be matched.
[0014] Furthermore, the specific process of inputting the image to be matched into the key point detection network to generate the key point heat map K is as follows: the image to be matched is divided into multiple grid cells, and each grid cell is reshaped into a feature vector , and then the feature vector passes through four sixteenth convolutional layers to obtain convolutional features , and then the convolutional features are processed by the softmax function and expanded to the resolution of the image to be matched to obtain the key point heat map K.
[0015] Furthermore, the Euclidean distance is used to complete the brute-force matching of the enhanced descriptors of two images to be matched; for the enhanced descriptors of two given images to be matched, the similarity is measured by calculating the Euclidean distance between the two for matching.
[0016] Compared with the existing technologies, the present invention has the following beneficial effects:
[0017] (1) The present invention constructs a new dual-branch collaborative architecture. From the perspective of dual-branch collaboration between feature detection and feature description, the key point detection and feature description are processed separately. The key point detection network focuses on low-level features, avoiding the interference of high-level features and improving the accuracy of key point detection. At the same time, through the feature extraction backbone network composed of a differential wavelet transform downsampling module and a dynamic multi-scale feature fusion module, the discrete wavelet transform is used to separate noise and key features, combined with differential operations to initially remove redundant information and enhance key features. Then, a multi-branch structure is used to extract multi-scale features, and the contributions of each branch are adjusted by dynamic weighting to further remove redundant information and achieve efficient multi-scale feature fusion to adapt to scale changes under different perspectives. In the descriptor generation network, by designing an adaptive multi-scale descriptor enhancement module, using various mechanisms such as dynamic feature adjustment, multi-scale feature fusion, non-linear activation, and adaptive feature fusion, the quality and expression ability of the descriptor are improved, and the accuracy of feature matching is increased.
[0018] (2) The differential wavelet transform downsampling module proposed by the present invention uses the discrete wavelet transform to decompose the features in the frequency domain, separates noise and key features, and combines the differential operation of calculating the difference between the feature and the neighborhood average pooling, effectively removing redundant information and enhancing key features, enabling the descriptor to more accurately capture the key features of the image, suppressing the contrast distortion caused by light changes, improving the stability of feature response in uneven illumination scenarios, and increasing the accuracy of subsequent feature matching.
[0019] (3) The dynamic multi-scale feature fusion module proposed by the present invention divides the input channels into three parts, and uses depthwise separable convolutions of different sizes to extract multi-scale features respectively, avoiding the high computational cost of full-channel convolutions and reducing the computational amount through grouped convolutions. At the same time, the contributions of each branch are adjusted by dynamic weighting, and the activation function is used to smoothly amplify key features and suppress irrelevant information, achieving efficient multi-scale feature fusion, enhancing the scale invariance of the descriptor, and improving the accuracy of feature matching under perspective changes.
[0020] (4) The adaptive multi-scale descriptor enhancement module proposed by the present invention uses a dynamic hyperbolic tangent function to perform non-linear dynamic scaling on the input descriptor, adaptively adjusts the feature distribution, effectively suppresses the feature ambiguity caused by complex illumination and perspective changes, and improves the robustness of the descriptor. At the same time, the mamba layer with local convolutions of different sizes is used to capture feature information in different receptive fields, and the multi-scale information is integrated by average fusion to achieve the integration of multi-scale information of the descriptor, enhancing the expression ability of the descriptor for local geometric structures and global semantic information, and achieving high-precision feature matching in complex scenarios with combined perspective and illumination changes. Brief Description of the Drawings
[0021] Figure 1 Schematic diagram of the feature matching model structure of the present invention.
[0022] Figure 2 Schematic diagram of the feature extraction module structure of the present invention.
[0023] Figure 3 Schematic diagram of the differential wavelet transform downsampling module structure of the present invention.
[0024] Figure 4 Schematic diagram of the edge information enhancement module structure of the present invention.
[0025] Figure 5 Schematic diagram of the dynamic multi-scale feature fusion module structure of the present invention.
[0026] Figure 6 Schematic diagram of the adaptive multi-scale descriptor enhancement module structure of the present invention. Detailed implementation manners
[0027] The present invention provides a technical solution: a feature matching method based on differential wavelet transform and dynamic multi-scale feature fusion, including the following steps:
[0028] Read the images to be matched (set to two); construct a feature matching model, as Figure 1 shown; the feature matching model includes a backbone network, a key point detection network, and a descriptor generation network; the backbone network consists of a feature extraction module and a sixth convolutional block; input the images to be matched into the feature extraction module in the backbone network to generate a feature map F, pass the feature map F through the sixth convolutional block to generate a reliability map R; input the images to be matched into the key point detection network to generate a key point heat map K, input the key point heat map K and the reliability map R into the descriptor generation network to screen out key points, and finally extract the feature descriptors of the corresponding images to be matched from the feature map F according to the screened key points, process the feature descriptors of the images to be matched through an adaptive multi-scale descriptor enhancement module to obtain the enhanced descriptors of the images to be matched, and finally perform brute-force matching on the enhanced descriptors of the two images to be matched; through this process, a set of key points and feature descriptors with high accuracy and high response can be generated, thereby realizing efficient and accurate feature matching.
[0029] Among them, the brute-force matching of the enhanced descriptors of the two images to be matched is completed using the Euclidean distance; for the given enhanced descriptors of the two images to be matched, the similarity is measured by calculating the Euclidean distance between the two for matching.
[0030] Among them, the specific process of inputting the images to be matched into the key point detection network to generate a key point heat map K is: divide the images to be matched into multiple 8×8 grid cells, and reshape each grid cell into a 64-dimensional feature vector , which is used to model the distribution of key points in each grid, and then the feature vector Pass through four 1×1 sixteenth convolutional layers to obtain convolutional features , and then the convolutional features After being processed by the softmax function, it is expanded to the resolution of the original input image (image to be matched) to obtain the key point heat map K.
[0031] As Figure 2 shown, the feature extraction module includes a first convolutional block, a second convolutional block, a third convolutional block, a fourth convolutional block, and a fifth convolutional block; the input (image to be matched) passes through the first convolutional block and the second convolutional block respectively, and the outputs of the first convolutional block and the second convolutional block are added together to obtain an added feature, and the added feature passes through the third convolutional block, the fourth convolutional block, and the fifth convolutional block in sequence to obtain the feature map F.
[0032] Among them, the first convolutional block is composed of a first convolutional layer (3×3), a first discrete wavelet transform downsampling module, a first dynamic multi-scale feature fusion module, a second dynamic multi-scale feature fusion module, and a second discrete wavelet transform downsampling module connected in sequence; the input (image to be matched) passes through the first convolutional layer (3×3), the first discrete wavelet transform downsampling module, the first dynamic multi-scale feature fusion module, the second dynamic multi-scale feature fusion module, and the second discrete wavelet transform downsampling module in sequence to obtain the output of the first convolutional block.
[0033] Among them, the second convolutional block is composed of a first average pooling layer (4×4) and a second convolutional layer (1×1) connected in sequence; the input (image to be matched) passes through the first average pooling layer and the second convolutional layer (1×1) in sequence to obtain the output of the second convolutional block.
[0034] Among them, the third convolutional block is composed of a second dynamic multi-scale feature fusion module and a third dynamic multi-scale feature fusion module connected in sequence; the input (added feature) passes through the second dynamic multi-scale feature fusion module and the third dynamic multi-scale feature fusion module in sequence to obtain the output of the third convolutional block.
[0035] Among them, the fourth convolutional block is composed of a third discrete wavelet transform downsampling module, a fourth dynamic multi-scale feature fusion module, and a fifth dynamic multi-scale feature fusion module connected in sequence; the input (output of the third convolutional block) passes through the third discrete wavelet transform downsampling module, the fourth dynamic multi-scale feature fusion module, and the fifth dynamic multi-scale feature fusion module in sequence to obtain the output of the fourth convolutional block.
[0036] Among them, the fifth convolutional block consists of a third convolutional layer (3×3), a fourth convolutional layer (3×3), and a fifth convolutional layer (1×1) connected in sequence; the input (the output of the fourth convolutional block) passes through the third convolutional layer (3×3), the fourth convolutional layer (3×3), and the fifth convolutional layer (1×1) in sequence to obtain the output of the fifth convolutional block.
[0037] Among them, the sixth convolutional block consists of a sixth convolutional layer (1×1), a seventh convolutional layer (1×1), and an eighth convolutional layer (1×1) connected in sequence; the input (feature map F) passes through the sixth convolutional layer (1×1), the seventh convolutional layer (1×1), and the eighth convolutional layer (1×1) in sequence and then passes through the sigmoid function to obtain the reliability map R.
[0038] Among them, after the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, the fifth convolutional layer, the sixth convolutional layer, the seventh convolutional layer, and the eighth convolutional layer, a batch normalization layer (BN) and a ReLU activation function are connected in sequence.
[0039] As Figure 3 shown, the first differential wavelet transform downsampling module (DWTD), the second differential wavelet transform downsampling module, the third differential wavelet transform downsampling module, and the fourth differential wavelet transform downsampling module have the same structure; the first differential wavelet transform downsampling module includes a ninth convolutional layer (1×1), a discrete wavelet transform module, a tenth convolutional layer (1×1), an edge information enhancement module, and an eleventh convolutional layer (1×1); the processing flow of the first differential wavelet transform downsampling module is: the input (the output of the first convolutional layer ) passes through the ninth convolutional layer, and the output of the ninth convolutional layer is input into the discrete wavelet transform module for frequency domain decomposition to obtain low-frequency information and high-frequency information , which can be expressed as: ; In the formula, represents the discrete wavelet forward transform function.
[0040] The high-frequency information is further decomposed into horizontal high-frequency information , vertical high-frequency information and diagonal high-frequency information .
[0041] Since the diagonal high-frequency information mainly contains large-scale smoothing information and its contribution to the extraction of fine-grained edge features is limited, it is directly discarded to reduce redundancy, and only the other three pieces of information are retained, focusing on details and edge features.
[0042] The horizontal high-frequency information , the vertical high-frequency information and the low-frequency information Concatenate along the channel dimension to obtain a high-frequency information feature map. Pass the high-frequency information feature map through the tenth convolutional layer and the edge information enhancement module in sequence. Multiply the output of the edge information enhancement module and the output of the tenth convolutional layer. Pass the multiplied output through the eleventh convolutional layer to obtain the output of the first differential wavelet transform downsampling module.
[0043] The second differential wavelet transform downsampling module, the third differential wavelet transform downsampling module, the fourth differential wavelet transform downsampling module, and the first differential wavelet transform downsampling module have the same structure, and the processing flow will not be elaborated here.
[0044] As Figure 4 shown, the edge information enhancement module includes a second average pooling layer and a twelfth convolutional layer. The processing flow of the edge information enhancement module is as follows: Pass the input (the output of the tenth convolutional layer) through the second average pooling layer, subtract the output of the second average pooling layer using the input (the output of the tenth convolutional layer), first adjust the subtracted output through the twelfth convolutional layer, and then activate it through the Sigmoid function to obtain a weight map , add 1 to the weight map (since the output of the sigmoid function is between 0 and 1, adding 1 changes it to between 1 and 2, which can prevent errors such as division by zero in subsequent calculations) to obtain the output of the edge information enhancement module ; This design of the edge information enhancement module can enhance the edge information while maintaining the integrity of the feature distribution, providing a more balanced expression ability for the descriptor.
[0045] The differential wavelet transform downsampling module not only increases the feature representation form by transforming in the feature space domain to increase the frequency domain information, but also effectively removes noise and redundant information.
[0046] As Figure 5As shown, the first dynamic multi-scale feature fusion module (DMFF), the second dynamic multi-scale feature fusion module, the third dynamic multi-scale feature fusion module, the fourth dynamic multi-scale feature fusion module, and the fifth dynamic multi-scale feature fusion module have the same structure; the first dynamic multi-scale feature fusion module includes a first depthwise separable convolution (3×3), a Split function, a second depthwise separable convolution (5×5), a third depthwise separable convolution (7×7), a thirteenth convolutional layer (1×1), a fourteenth convolutional layer (1×1), a fifteenth convolutional layer (1×1), and a SiLU activation function; the processing flow of the first dynamic multi-scale feature fusion module is as follows: the input (the output of the first differential wavelet transform downsampling module) is respectively passed through the first depthwise separable convolution and the fourteenth convolutional layer, the output of the first depthwise separable convolution is divided into a first group of features, a second group of features, and a third group of features in the channel dimension according to (1:1:2) through the Split function, the second group of features and the third group of features are respectively passed through the second depthwise separable convolution and the third depthwise separable convolution, the outputs of the second depthwise separable convolution and the third depthwise separable convolution are concatenated with the first group of features in the channel dimension, the concatenated result is input into the thirteenth convolutional layer, the outputs of the thirteenth convolutional layer and the fourteenth convolutional layer are both multiplied after passing through a SiLU activation function, and the multiplied result is passed through the fifteenth convolutional layer to obtain the output of the first dynamic multi-scale feature fusion module.
[0047] Among them, the specific process of inputting the keypoint heatmap K and the reliability map R into the descriptor generation network to screen out keypoints, and finally extracting the feature descriptors of the corresponding images to be matched from the feature map F according to the screened keypoints is as follows: perform non-maximum suppression on the keypoint heatmap K, and select the local response peak points as candidate keypoints; then use nearest neighbor interpolation to select the response values of the candidate keypoints from the keypoint heatmap K, use bilinear interpolation to select the feature responses of the candidate keypoints from the reliability map R, multiply the response values of the candidate keypoints selected from the keypoint heatmap K and the feature responses of the candidate keypoints selected from the reliability map R to obtain a comprehensive score, screen according to the ranking of the comprehensive score to obtain a group of keypoints with high accuracy and high response, and then use bilinear interpolation to extract the feature descriptors of the corresponding images to be matched from the feature map F according to the screened keypoints.
[0048] Such as Figure 6As shown, the Adaptive Multi-Scale Descriptor Enhancement Module (AMDE) includes a dynamic hyperbolic tangent function, a first Mamba layer (with a convolutional kernel size of 2), a second Mamba layer (with a convolutional kernel size of 3), a third Mamba layer (with a convolutional kernel size of 4), a linear layer, and a GeLU activation function; the processing flow of the Adaptive Multi-Scale Descriptor Enhancement Module is as follows: The input (the feature descriptor of the image to be matched) is processed by the dynamic hyperbolic tangent function (Dy-Tanh), which dynamically adjusts the activation intensity and slope according to the input features, thereby introducing non-linearity and enhancing the feature expression ability of the model. The output of the dynamic hyperbolic tangent function is added to the input (the feature descriptor of the image to be matched) to obtain the first added descriptor. The first added descriptor is respectively input into the first Mamba layer, the second Mamba layer, and the third Mamba layer. The outputs of the first Mamba layer, the second Mamba layer, and the third Mamba layer are averaged and fused. The output after average fusion is added to the first added descriptor to obtain the second added descriptor. The second added descriptor is input into the linear layer for processing. The output of the linear layer is concatenated with the second added descriptor to obtain the third added descriptor. The output after passing the third added descriptor through the GeLU activation function is concatenated with the third added descriptor again to obtain the output of the Adaptive Multi-Scale Descriptor Enhancement Module, that is, the enhanced descriptor of the image to be matched.
[0049] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. Feature matching method based on differential wavelet transform and dynamic multi-scale feature fusion, characterized in that, The steps are as follows: reading the image to be matched; constructing a feature matching model; the feature matching model includes a backbone network, a key point detection network, and a descriptor generation network; the backbone network consists of a feature extraction module and a sixth convolutional block; inputting the image to be matched into the feature extraction module in the backbone network to generate a feature map, passing the feature map through the sixth convolutional block to generate a reliability map; inputting the image to be matched into the key point detection network to generate a key point heat map, inputting the key point heat map and the reliability map into the descriptor generation network to screen out key points, and finally extracting the feature descriptors of the image to be matched corresponding to the key points screened out from the feature map, processing the feature descriptors of the image to be matched through an adaptive multi-scale descriptor enhancement module to obtain enhanced descriptors of the image to be matched, and finally performing brute-force matching on the enhanced descriptors of the image to be matched; Set the number of images to be matched read as two, and finally perform brute-force matching on the enhanced descriptors of the two images to be matched obtained.
2. The feature matching method based on differential wavelet transform and dynamic multi-scale feature fusion according to claim 1, wherein: The feature extraction module includes a first convolutional block, a second convolutional block, a third convolutional block, a fourth convolutional block, and a fifth convolutional block; passing the image to be matched through the first convolutional block and the second convolutional block respectively, adding the outputs of the first convolutional block and the second convolutional block to obtain an added feature, and passing the added feature through the third convolutional block, the fourth convolutional block, and the fifth convolutional block in sequence to obtain a feature map F; Among them, the first convolutional block consists of a first convolutional layer, a first differential wavelet transform downsampling module, a first dynamic multi-scale feature fusion module, a second dynamic multi-scale feature fusion module, and a second differential wavelet transform downsampling module connected in sequence; passing the image to be matched through the first convolutional layer, the first differential wavelet transform downsampling module, the first dynamic multi-scale feature fusion module, the second dynamic multi-scale feature fusion module, and the second differential wavelet transform downsampling module in sequence to obtain the output of the first convolutional block; Among them, the second convolutional block consists of a first average pooling layer and a second convolutional layer connected in sequence; passing the image to be matched through the first average pooling layer and the second convolutional layer in sequence to obtain the output of the second convolutional block; Among them, the third convolutional block consists of a second dynamic multi-scale feature fusion module and a third dynamic multi-scale feature fusion module connected in sequence; passing the added feature through the second dynamic multi-scale feature fusion module and the third dynamic multi-scale feature fusion module in sequence to obtain the output of the third convolutional block; Among them, the fourth convolutional block consists of a third differential wavelet transform downsampling module, a fourth dynamic multi-scale feature fusion module, and a fifth dynamic multi-scale feature fusion module connected in sequence; passing the output of the third convolutional block through the third differential wavelet transform downsampling module, the fourth dynamic multi-scale feature fusion module, and the fifth dynamic multi-scale feature fusion module in sequence to obtain the output of the fourth convolutional block; Among them, the fifth convolutional block consists of a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer connected in sequence; passing the output of the fourth convolutional block through the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer in sequence to obtain the output of the fifth convolutional block; Among them, the sixth convolutional block consists of a sixth convolutional layer, a seventh convolutional layer, and an eighth convolutional layer connected in sequence; the feature map F is successively passed through the sixth convolutional layer, the seventh convolutional layer, and the eighth convolutional layer and then passed through the sigmoid function to obtain a reliability map.
3. The feature matching method based on differential wavelet transform and dynamic multi-scale feature fusion according to claim 2, wherein: The first differential wavelet transform downsampling module, the second differential wavelet transform downsampling module, the third differential wavelet transform downsampling module, and the fourth differential wavelet transform downsampling module have the same structure; the first differential wavelet transform downsampling module includes a ninth convolutional layer, a discrete wavelet transform module, a tenth convolutional layer, an edge information enhancement module, and an eleventh convolutional layer; The processing flow of the first differential wavelet transform downsampling module is as follows: The output of the first convolutional layer is passed through the ninth convolutional layer, and the output of the ninth convolutional layer is input into the discrete wavelet transform module for frequency domain decomposition to obtain low-frequency information and high-frequency information; The high-frequency information is further decomposed into horizontal high-frequency information, vertical high-frequency information, and diagonal high-frequency information; The horizontal high-frequency information, vertical high-frequency information, and low-frequency information are concatenated along the channel dimension to obtain a high-frequency information feature map. The high-frequency information feature map is successively passed through the tenth convolutional layer and the edge information enhancement module. The output of the edge information enhancement module and the output of the tenth convolutional layer are multiplied, and the multiplied output is then passed through the eleventh convolutional layer to obtain the output of the first differential wavelet transform downsampling module.
4. The feature matching method based on differential wavelet transform and dynamic multi-scale feature fusion according to claim 3, wherein: The edge information enhancement module includes a second average pooling layer and a twelfth convolutional layer; the processing flow of the edge information enhancement module is as follows: the output of the tenth convolutional layer is passed through the second average pooling layer, the output of the second average pooling layer is subtracted by the output of the tenth convolutional layer, the subtracted output is first adjusted by the twelfth convolutional layer, and then activated by the Sigmoid function to obtain a weight map, and 1 is added to the weight map to obtain the output of the edge information enhancement module.
5. The feature matching method based on differential wavelet transform and dynamic multi-scale feature fusion according to claim 4, wherein: The first dynamic multi-scale feature fusion module, the second dynamic multi-scale feature fusion module, the third dynamic multi-scale feature fusion module, the fourth dynamic multi-scale feature fusion module, and the fifth dynamic multi-scale feature fusion module have the same structure; the first dynamic multi-scale feature fusion module includes a first depthwise separable convolution, a Split function, a second depthwise separable convolution, a third depthwise separable convolution, a thirteenth convolutional layer, a fourteenth convolutional layer, a fifteenth convolutional layer, and a SiLU activation function.
6. The feature matching method based on differential wavelet transform and dynamic multi-scale feature fusion according to claim 5, wherein: The processing flow of the first dynamic multi-scale feature fusion module is as follows: The outputs of the first difference wavelet transform downsampling module are respectively passed through the first depthwise separable convolution and the fourteenth convolutional layer. The output of the first depthwise separable convolution is divided into a first group of features, a second group of features, and a third group of features in proportion in the channel dimension through the Split function. The second group of features and the third group of features are respectively passed through the second depthwise separable convolution and the third depthwise separable convolution. The outputs of the second depthwise separable convolution and the third depthwise separable convolution are concatenated with the first group of features in the channel dimension. The concatenated result is input into the thirteenth convolutional layer. The outputs of the thirteenth convolutional layer and the fourteenth convolutional layer are both multiplied after passing through a SiLU activation function. The result after multiplication is passed through the fifteenth convolutional layer to obtain the output of the first dynamic multi-scale feature fusion module.
7. The feature matching method based on differential wavelet transform and dynamic multi-scale feature fusion according to claim 6, characterized in that: The specific process of inputting the key point heat map and the reliability map into the descriptor generation network to screen out key points and finally extracting the feature descriptors of the corresponding images to be matched from the feature map F according to the screened key points is as follows: Perform non-maximum suppression on the key point heat map and select the local response peak points as candidate key points. Then, use nearest neighbor interpolation to select the response values of the candidate key points from the key point heat map, and use bilinear interpolation to select the feature responses of the candidate key points from the reliability map. Multiply the response values of the candidate key points selected from the key point heat map and the feature responses of the candidate key points selected from the reliability map to obtain a comprehensive score. According to the sorting and screening of the comprehensive score, key points are obtained. Then, according to the screened key points, bilinear interpolation is used to extract the corresponding feature descriptors of the images to be matched from the feature map F.
8. The feature matching method based on differential wavelet transform and dynamic multi-scale feature fusion according to claim 7, wherein: The adaptive multi-scale descriptor enhancement module includes a dynamic hyperbolic tangent function, a first Mamba layer, a second Mamba layer, a third Mamba layer, a linear layer, and a GeLU activation function. The processing flow of the adaptive multi-scale descriptor enhancement module is as follows: The feature descriptors of the images to be matched are processed by the dynamic hyperbolic tangent function. The output after passing through the dynamic hyperbolic tangent function is added to the feature descriptors of the images to be matched to obtain a first added descriptor. The first added descriptor is respectively input into the first Mamba layer, the second Mamba layer, and the third Mamba layer. The outputs of the first Mamba layer, the second Mamba layer, and the third Mamba layer are averaged and fused. The output after average fusion is added to the first added descriptor to obtain a second added descriptor. The second added descriptor is input into the linear layer for processing. The output of the linear layer is concatenated with the second added descriptor to obtain a third added descriptor. The output after passing the third added descriptor through the GeLU activation function is concatenated with the third added descriptor again to obtain the output of the adaptive multi-scale descriptor enhancement module, that is, the enhanced descriptor of the images to be matched.
9. The feature matching method based on differential wavelet transform and dynamic multi-scale feature fusion according to claim 8, wherein: The specific process of inputting the image to be matched into the key point detection network to generate the key point heat map is as follows: The image to be matched is divided into multiple grid units, and each grid unit is reshaped into a feature vector , and then the feature vector passes through four sixteenth convolutional layers to obtain convolutional features . Then, the convolutional features are processed by the softmax function and expanded to the resolution of the image to be matched to obtain the key point heat map.
10. The feature matching method based on differential wavelet transform and dynamic multi-scale feature fusion according to claim 9, wherein: Use the Euclidean distance to complete the brute-force matching of the enhanced descriptors of two images to be matched; for the given enhanced descriptors of two images to be matched, calculate the Euclidean distance between the two to measure the similarity for matching.
Citation Information
Patent Citations
Dynamic illumination face image quality enhancement method based on multi-scale attention mechanism
CN115880225A
Underwater image enhancement method of Mama hybrid architecture based on space-frequency fusion
CN118710507A
Vehicle positioning method based on multi-view and multi-scale feature fusion
CN119027776A
Multi-scale fusion image enhancement method based on discrete wavelet transform and deep network
CN119048380A
Remote sensing landform enhancement algorithm
CN119991458A
Cited By
Multi-branch fusion feature matching method based on wavelet pooling downsampling
CN120526276A
PRNU anonymity method based on multi-scale and hierarchical feature fusion
CN120599058A
Complex underwater side-scan sonar exploration detection method and device based on multi-dimensional attention collaborative lightweight anti-noise detection framework
CN120871150A
A complex underwater side-scan sonar exploration detection method and device based on a multi-dimensional attention collaborative lightweight anti-noise detection framework
CN120871150B
Feature matching method based on double-branch feature extraction and channel feature enhancement
CN121616861A