A method and apparatus for large-format image registration and fusion
By establishing a convolutional neural network and sparse sampling technology, the problem of low efficiency in large-format image registration and fusion was solved, achieving efficient image fusion results and reducing memory and time requirements.
Patent Information
- Application Number
- CN202310389302.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-13
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2043-04-13
AI Technical Summary
Existing technologies are inefficient in large-format image registration and fusion, and require a lot of memory and time. They also suffer from errors and insufficient processing efficiency, especially in the digital preservation of cultural relics.
An image fusion method based on sparse sampling is adopted. A convolutional neural network is established to extract the location information of feature points in the source image, and the location of the target image is calculated by using a correlation filter. The image blocks are decomposed to generate a low-resolution image, and Poisson fusion is used to generate a sparse edge image. Finally, color pixel fusion is performed at a high resolution.
It improves the efficiency of large-format image registration and fusion, reduces memory and time requirements, and maintains visual consistency, thus enabling rapid fusion of large-format images.
Smart Images

Figure CN116612160B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to a large-format image registration and fusion method and device. BACKGROUND
[0002] China is one of the ancient civilizations, among the precious wealth left by our predecessors, cultural relics are a typical material cultural heritage, which plays an important role in traditional culture research, inheritance, appreciation and education, etc. In contemporary China, cultural relics are an integral part of cultural self-confidence. The rapid development of information and computer technology makes the digitization of cultural relics an important part of the entire digital cultural relic collection, digital restoration and high-fidelity copying work. With the advent of artificial intelligence, using advanced digital and intelligent technology to digitize and effectively protect China's ancient cultural heritage not only overcomes the difficulties and challenges of traditional methods, but also solves problems that cannot be solved by manpower in a faster and more efficient way, which has important practical significance.
[0003] At present, there are still many bottleneck problems in the protection of large-format cultural relics from theoretical methods to equipment development. Among them, for traditional methods, due to the sensitivity to spectral information changes, the extraction of feature points under the load environment often introduces errors, etc. When performing image registration, the registration result is good for small areas, but the efficiency is low for large-scale large-format image registration. After large-format image registration, a large amount of memory and time is required for fusion, and the processing efficiency is low. SUMMARY
[0004] The purpose of the present application is to overcome the shortcomings of the prior art, and the present application provides a large-format image registration and fusion method and device. By extracting the position information of the feature points on the source image on the target image, and using an image fusion method based on sparse sampling, the fast registration and fusion of large-format target images are realized, and the efficiency of large-format high-resolution image registration and fusion in the field of cultural relic restoration is improved.
[0005] The present application provides a large-format image registration and fusion method, which comprises:
[0006] A convolutional neural network is established, and a source image is input into the convolutional neural network to extract convolutional layer feature information at different depths in the source image;
[0007] The convolutional layer feature information at different depths is filtered based on a correlation filter, and the position information of the feature point pair on the source image on the target image is calculated;
[0008] The target image is registered based on the position information;
[0009] The target image is decomposed into a plurality of image blocks, two corresponding low-resolution images are generated, and the two low-resolution images are poisson fused into a sparse edge image;
[0010] The internal pixel point information of the plurality of quadrilateral image blocks is calculated, the sparse edge image is combined for fusion, and a final fusion image is generated.
[0011] Further, the convolutional neural network is established, the source image is input into the convolutional neural network, and the convolutional layer feature information of different depths in the source image is extracted, including:
[0012] A plurality of small area images centered on feature points are collected as a training set;
[0013] A VGG-Net deep convolutional neural network model is established, and the deep convolutional neural network model is trained based on the training set;
[0014] The convolutional layer feature information of different depths in the source image is extracted based on the third, fourth and fifth layers of the trained deep convolutional neural network model.
[0015] Further, the convolutional layer feature information of different depths in the source image is extracted based on the third, fourth and fifth layers of the trained deep convolutional neural network model, including:
[0016] The third, fourth and fifth layers of the trained deep convolutional neural network model are combined to generate a plurality of combined convolutional layers with a convolutional kernel size of 7x7, and the convolutional layer feature information of different depths in the source image is extracted based on the combined convolutional layers.
[0017] Further, the convolutional layer feature information of different depths is filtered based on the correlation filter, and the position information of the feature points on the source image in the target image is calculated, including:
[0018] A correlation filter model is established based on the convolutional layer feature information of different depths in the source image, a correlation filtering algorithm is used to calculate the similarity of the feature point pairs in each depth of the source image, and the position of the maximum similarity response value is determined;
[0019] Based on the position information of the maximum similarity response value, the convolutional layer feature information of different depths in the source image is combined to comprehensively calculate the position information of the feature points on the source image in the target image.
[0020] Further, the related filtering model is established based on the convolution layer feature information of different depths in the source image, a correlation filtering algorithm is used to calculate the similarity of the to-be-matched points in each depth of the source image, and the position where the maximum similarity response value appears is determined.
[0021] The feature point pairs in the source image are matched by using the kernel correlation filtering algorithm, and the similarity response values of each feature point pair are calculated.
[0022] It is judged whether the similarity response values are reasonable, and the incorrect feature point pairs are deleted.
[0023] The similarity response values of the reasonable feature point pairs in the source image are sorted, the feature point pair with the maximum similarity response value is obtained, and the position information of the feature point pair with the maximum similarity response value is obtained.
[0024] Further, the target image is registered based on the position information, which includes:
[0025] The feature point pairs with the same position information of the feature point pair with the maximum similarity response value in the source image are matched, and the target image where the feature point pairs with the same position information are located is registered.
[0026] Further, the target image is decomposed into a plurality of image blocks, two corresponding low-resolution images are generated, and the two low-resolution images are Poisson fused into a sparse edge image, which includes:
[0027] The target image is decomposed into a plurality of quadrilateral image blocks, the pixel points inside each of the plurality of quadrilateral image blocks are discarded, and a plurality of quadrilateral image blocks retaining boundary pixel points are generated;
[0028] The plurality of quadrilateral image blocks retaining boundary pixel points are constructed into a low-resolution image in the horizontal direction and a low-resolution image in the vertical direction;
[0029] The two low-resolution images are Poisson fused into a sparse edge image.
[0030] Further, the plurality of quadrilateral image blocks retaining boundary pixel points are constructed into a low-resolution image in the horizontal direction and a low-resolution image in the vertical direction, which includes:
[0031] A complete binary tree structure model is constructed, the plurality of quadrilateral image blocks retaining boundary pixel points are input into the complete binary tree structure model, and a low-resolution image in the horizontal direction and a low-resolution image in the vertical direction are generated.
[0032] Further, the internal pixel point information of the plurality of quadrilateral image blocks is calculated, and the sparse edge image is fused to generate a final fused image, which includes:
[0033] The internal pixel point information of each of the plurality of quadrilateral image blocks is interpolated by using a GPU parallel method, and the basic information value of the sparse edge image is combined to calculate the basic information value of the final fusion image.
[0034] The application further provides a device for large-format image registration and fusion.
[0035] The feature extraction module is configured to establish a convolutional neural network, input a source image into the convolutional neural network, and extract convolutional layer feature information at different depths in the source image.
[0036] The correlation filtering module is configured to filter the convolutional layer feature information at different depths to calculate position information of a feature point pair on the source image on the target image.
[0037] The registration module is configured to register the target image based on the position information.
[0038] The sparse edge image generation module is configured to decompose the target image into a plurality of image blocks, generate two corresponding low-resolution images, and Poisson fuse the two low-resolution images into a sparse edge image.
[0039] The fusion module is configured to calculate internal pixel point information of the plurality of quadrilateral image blocks, fuse the sparse edge image, and generate a final fusion image.
[0040] The application establishes a VGG-Net deep convolutional neural network model, extracts convolutional layer feature information at different depths in a source image through a combination of third, fourth and fifth convolutional layers, and comprehensively calculates position information of a feature point on the source image on a target image based on a correlation filtering algorithm, so that the extraction accuracy is high and the efficiency of large-format image registration is improved. Based on a sparse sampling method, the internal pixel points of the target image are discarded, the target image is synthesized at a low resolution, and the target image is fused into a sparse edge image in a Poisson fusion manner. The color pixel points are fused at a high resolution on the original target image to generate a final fusion image, which greatly improves the efficiency of gradient domain synthesis of large-format images, reduces the memory and time required for fusion operation, achieves fast fusion of large-format images while achieving the same visual effect, and improves the efficiency of large-format image fusion. BRIEF DESCRIPTION OF DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings described below only illustrate some of the embodiments of the present application, and all other drawings obtained by those of ordinary skill in the art without creative effort based on these drawings are within the scope of the present application.
[0042] Figure 1 is a method flowchart of large-format image registration and fusion in the embodiments of the present application;
[0043] Figure 2 is a flowchart of extracting convolutional layer feature information of different depths in a source image in the embodiments of the present application;
[0044] Figure 3 is a flowchart of calculating position information of feature points on a source image in the embodiments of the present application;
[0045] Figure 4 is a flowchart of obtaining position information of a feature point pair with the largest similarity response value in the embodiments of the present application;
[0046] Figure 5 is a tracking flowchart of a kernel correlation filtering algorithm in the embodiments of the present application;
[0047] Figure 6 is a flowchart of obtaining a sparse edge image in the embodiments of the present application;
[0048] Figure 7 is a device structure schematic diagram of large-format image registration and fusion in the embodiments of the present application. DETAILED DESCRIPTION
[0049] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of the present application.
[0050] In the present application, it should be understood that terms such as "include" or "have" are intended to indicate that there exist the features, numbers, steps, actions, components, parts or combinations thereof disclosed in the specification, and do not exclude the possibility that one or more other features, numbers, steps, actions, components, parts or combinations thereof exist or are added.
[0051] In addition, it should be further understood that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.
[0052] Example 1
[0053] This invention relates to a method for large-format image registration and fusion, comprising: establishing a convolutional neural network; inputting a source image into the convolutional neural network; extracting feature information of convolutional layers at different depths in the source image; filtering the feature information of the convolutional layers at different depths based on a correlation filter; calculating the position information of feature point pairs on the source image in the target image; registering the target image based on the position information; decomposing the target image into several image blocks to generate two corresponding low-resolution images; and merging the two low-resolution images into a sparse-edge image using Poisson fusion; calculating the internal pixel information of the several quadrilateral image blocks; and fusing them with the sparse-edge image to generate the final fused image.
[0054] In one optional implementation of this embodiment, such as Figure 1 As shown, Figure 1 A flowchart of a large-format image registration and fusion method according to an embodiment of the present invention is shown, including the following steps:
[0055] S101. Establish a convolutional neural network, input the source image into the convolutional neural network, and extract feature information of convolutional layers of different depths in the source image;
[0056] In one optional implementation of this embodiment, such as Figure 2 As shown, Figure 2 The flowchart illustrating the extraction of feature information from convolutional layers of different depths in a source image according to an embodiment of the present invention includes the following steps:
[0057] S201. Collect several small region images centered on feature points as a training set;
[0058] Specifically, several processed, small-area images centered on feature points are collected through multiple channels, including the point set domain of the source image and the point set domain of the quasi-target image.
[0059] S202. Establish a VGG-Net deep convolutional neural network model, and train the deep convolutional neural network model based on the training set;
[0060] In an optional implementation of this embodiment, a VGG-Net (Visual Geometry Group Network) deep convolutional neural network model based on the Convolutional Neural Network (CNN) algorithm is established, and the VGG-Net deep convolutional neural network model is trained based on the training set collected in step S201.
[0061] Specifically, the VGG-Net deep convolutional neural network model is a super-deep convolutional neural network model for large-scale recognition, mainly including VGG-16 and VGG-19, and in the embodiment, the VGG-16 is adopted, and the VGG-16 contains 13 convolutional layers and 3 fully connected layers.
[0062] S203, the convolutional layer feature information of different depths in the source image is extracted based on the third, fourth and fifth layers in the trained deep convolutional neural network model.
[0063] In an optional implementation of the embodiment, the third, fourth and fifth layers in the trained deep convolutional neural network model are combined to generate a combined convolutional layer with several convolutional kernels with a size of 7*7, and the convolutional layer feature information of different depths in the source image is extracted based on the combined convolutional layer.
[0064] Specifically, the third, fourth and fifth convolutional layers of the VGG-16 are each provided with 512 convolutional kernels with a size of 3*3, in this step, a combined convolutional layer with 512 convolutional kernels with a size of 7*7 is generated, and the convolutional layer feature information of different depths in the source image is extracted based on the combined convolutional layer.
[0065] It should be noted that, as the depth of the convolutional layer of the CNN algorithm increases, the expression ability of the spatial information gradually decreases, and the expression ability of the semantic information gradually increases, therefore, for the registration of a large-format image, the third, fourth and fifth convolutional layers of the VGG-16 based on the convolutional neural network are combined, and the extraction accuracy of the position information is high.
[0066] S102, the convolutional layer feature information of different depths is filtered based on a correlation filter, and the position information of the feature point pair on the source image on the target image is calculated;
[0067] In an optional implementation of the embodiment, as shown in Figure 3 Figure 3 a flow chart for calculating the position information of the feature point on the source image on the target image in the embodiment is shown, including the following steps:
[0068] S301, a correlation filtering model is established based on the convolutional layer feature information of different depths in the source image, a correlation filtering algorithm is used to calculate the similarity of the feature point pair in each depth of the source image, and the position of the maximum correlation degree response value is determined;
[0069] Specifically, the outputs of multiple convolution kernels in a single convolution layer in the VGG-16 are used for multi-channel features to establish a correlation filter model to calculate the similarity of feature point pairs between different depth convolution layers in the source image, and the position of the feature point pair with the maximum similarity response value is determined by finding the position of the maximum similarity response value.
[0070] In an optional implementation of the embodiment, as shown in Figure 4 Figure 4 A flowchart of a process of obtaining the position information of the feature point pair with the maximum similarity response value in the embodiment of the application is shown, including the following steps:
[0071] S401, matching the feature point pairs in the source image by using the kernel correlation filter algorithm, and calculating the similarity response value of each feature point pair;
[0072] In an optional implementation of the embodiment, the kernel correlation filter algorithm (KCF, Kernel Correlation Filter) is used to match the feature point pairs in the source image. The algorithm combines the kernel function with the correlation filter, uses a cyclic matrix to obtain training samples, increases the number of negative samples, trains a discriminant classifier, judges whether the tracking is the target or the background information of the target by the classifier, uses the fast Fourier transform (FFT) to improve the calculation speed of the algorithm, converts the ridge regression to a nonlinear space, reduces the network complexity, and has very outstanding performance in tracking effect and tracking speed.
[0073] Specifically, Figure 5 A tracking flowchart of the kernel correlation filter algorithm is shown. First, the feature information of the source image is input, the corresponding image information matched with the feature point pairs in the source image is obtained by the fast Fourier transform (FFT) and the inverse fast Fourier transform (IFFT), the image information is input into the correlation filter to perform the fast Fourier transform again, the position information of each feature point pair in the source image is obtained after multiple cycles, and the similarity response value of each feature point pair in the source image is calculated.
[0074] S402, judging whether the similarity response value is reasonable, and deleting the incorrect feature point pairs;
[0075] In an optional implementation of the embodiment, the similarity response values calculated are reasonably judged according to the past calculation data, the feature point pairs corresponding to the unreasonable similarity response values are screened out and deleted.
[0076] S403, sorting the similarity response values of the reasonable feature point pairs in the source image to obtain the feature point pair with the maximum similarity response value, and obtaining the position information of the feature point pair with the maximum similarity response value.
[0077] In an optional implementation of this embodiment, the similarity response values of the reasonable feature point pairs determined in step S402 are sorted to obtain the feature point pair with the largest similarity response value, and the position information of the feature point pair with the largest similarity response value is extracted accordingly.
[0078] S302. Based on the location information of the maximum similarity response value, and combined with the feature information of convolutional layers of different depths in the source image, the location information of the feature points on the source image on the target image is calculated.
[0079] In an optional implementation of this embodiment, based on the location information of the feature point pair with the largest similarity response value obtained in step S403, and combined with the feature information of convolutional layers of different depths in the source image, the final location information of the feature points on the source image on the target image is obtained by comprehensively considering the judgment results of multiple layers.
[0080] S103. Register the target image based on the location information;
[0081] In an optional implementation of this embodiment, feature point pairs with the same location information as the feature point pairs with the largest similarity response values on the source image are matched, and the target image containing the feature point pairs with the same location information is registered.
[0082] S104. Decompose the target image into several image blocks to generate two corresponding low-resolution images, and then merge the two low-resolution images into a sparse edge image using Poisson fusion.
[0083] In an optional implementation of this embodiment, the target image is synthesized at a low resolution to obtain a sparse edge image.
[0084] Specifically, a sparse sampling method is used here to synthesize the target image at low resolution. By exploiting the sparsity of the target image at low resolution, discrete image samples of the target image are obtained by random sampling under the condition of a much smaller Nyquist sampling rate. These samples are then reconstructed using a nonlinear reconstruction algorithm to synthesize a sparse edge image.
[0085] Specifically, such as Figure 6 As shown, Figure 6 A flowchart illustrating the process of obtaining a sparse edge image in an embodiment of the present invention is shown, including the following steps:
[0086] S601. Decompose the target image into several quadrilateral image blocks, discard the pixels inside each quadrilateral image block, and generate several quadrilateral image blocks that retain boundary pixels.
[0087] In one optional implementation of this embodiment, the target image is decomposed into several quadrilateral image blocks of the same size according to the proportion. The pixels inside these quadrilateral image blocks are discarded, and only the pixels at the boundaries are retained, thus generating several quadrilateral image blocks with retained boundary pixels.
[0088] S602. The quadrilateral image blocks with reserved boundary pixels are used to form a low-resolution image with one in the horizontal direction and one in the vertical direction.
[0089] In one optional implementation of this embodiment, a complete binary tree structure model is constructed, and the quadrilateral image blocks with reserved boundary pixels are input into the complete binary tree structure model to generate low-resolution images in the horizontal direction and low-resolution images in the vertical direction.
[0090] Specifically, to better organize the boundary images, a complete binary tree structure model based on sparse edges is constructed. The main node of the tree contains basic information about the boundary images, including size, image format, and position in the disk. The child nodes contain the corresponding sparse edge images. In this embodiment, the quadrilateral image blocks that retain the boundary pixels are used as the main nodes, and then the child nodes generate low-resolution images in the horizontal direction and low-resolution images in the vertical direction.
[0091] S603. The two low-resolution images are fused into a sparse-edge image using a Poisson fusion method.
[0092] In one optional implementation of this embodiment, Poisson fusion is an image fusion algorithm in the field of image processing. Its main advantage is that it can obtain naturally generated results without precise image matting. The Poisson fusion equation can be expressed as:
[0093]
[0094] In the formula, Represents the Laplace operator. Given a function, when When the time is equal to 1, it is called the Laplace equation, and the formula for calculating the two-dimensional Laplace equation is as follows:
[0095]
[0096] It should be noted that the Laplace equation is a special case of the Poisson equation.
[0097] S105. Calculate the internal pixel information of the plurality of quadrilateral image blocks, and fuse them with the sparse edge image to generate the final fused image.
[0098] In an optional implementation of the embodiment, GPU parallel method is adopted to interpolate the internal pixel point information of each of the plurality of quadrilateral image blocks, and the basic information value of the sparse edge image is combined to calculate the basic information value of the final fusion image.
[0099] Specifically, since the calculation of the internal pixel points of the quadrilateral image block is independent, parallel algorithm is used, and by solving the Poisson equation of the low resolution image and interpolating the quadrilateral image block, the time and memory consumption can be greatly reduced while achieving the same visual effect.
[0100] In summary, the embodiment one of the present application proposes a large-format image registration and fusion method, establishes a VGG-Net deep convolutional neural network model, extracts the convolutional layer feature information of different depths in the source image through the combination of the third, fourth and fifth convolutional layers, and based on the correlation filtering algorithm, the position information of the feature points on the target image is calculated, the extraction accuracy is high, and the efficiency of large-format image registration is improved; based on the sparse sampling method, the pixel points inside the target image are discarded, and the synthesis is performed at a low resolution, and the Poisson fusion method is used to fuse into a sparse edge image, and the color pixel points are fused on the original target image at a high resolution to generate the final fusion image, which greatly improves the efficiency of large-format image synthesis in the gradient domain, reduces the memory and time required for fusion operation, achieves the same visual effect while realizing the rapid fusion of large-format images, and improves the efficiency of large-format image fusion.
[0101] Embodiment two
[0102] The embodiment of the present application also relates to a large-format image registration and fusion device, as shown in Figure 7 , the device structure schematic diagram of the large-format image registration and fusion device in the embodiment of the present application is shown, and the device comprises: Figure 7 The device structure schematic diagram of the large-format image registration and fusion device in the embodiment of the present application is shown, and the device comprises:
[0103] The feature extraction module 10 is used to establish a convolutional neural network, input the source image into the convolutional neural network, and extract the convolutional layer feature information of different depths in the source image.
[0104] The correlation filtering module 20 is used to filter the convolutional layer feature information of different depths, and calculate the position information of the feature point pair on the target image.
[0105] The registration module 30 is used to register the target image based on the position information.
[0106] The sparse edge image generation module 40 is configured to decompose the target image into a plurality of image blocks, generate corresponding two low-resolution images, and poisson fuse the two low-resolution images into a sparse edge image.
[0107] The fusion module 50 is configured to calculate internal pixel point information of the plurality of quadrilateral image blocks, fuse the internal pixel point information in combination with the sparse edge image, and generate a final fused image.
[0108] In summary, the second embodiment of the present application proposes a large-format image registration and fusion device for performing the large-format image registration and fusion method described above, establishing a VGG-Net deep convolutional neural network model, extracting convolutional layer feature information at different depths in the source image through the third, fourth and fifth convolutional layer combinations, and calculating the position information of the feature points on the source image on the target image based on a correlation filtering algorithm, which has high extraction accuracy and improves the efficiency of large-format image registration; based on the sparse sampling method, the pixel points inside the target image are discarded, and the target image is synthesized at a low resolution, and the poisson fusion method is used to fuse the target image into a sparse edge image, and the color pixel points are fused at a high resolution on the original target image to generate a final fused image, which greatly improves the efficiency of large-format image registration in the gradient domain, reduces the memory and time required for fusion operation, achieves the same visual effect while realizing fast fusion of large-format images, and improves the efficiency of large-format image fusion.
[0109] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiments can be completed by programs instructing related hardware, and the programs can be stored in a computer readable storage medium, which can include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0110] In addition, the embodiments of the present application are described in detail above, and specific examples are used in this paper to describe the principles and implementation modes of the present application; the above description of the embodiments is only used to help understand the method and core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed; in summary, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A method of large format image registration and fusion, characterized by, The method comprises: establishing a convolutional neural network, inputting a source image into the convolutional neural network, and extracting convolutional layer feature information at different depths in the source image; filtering the convolutional layer feature information at different depths based on a correlation filter to calculate position information of a feature point pair on the source image on a target image; registering the target image based on the position information; decomposing the target image into a plurality of quadrilateral image blocks, generating corresponding two low-resolution images, and Poisson fusing the two low-resolution images into a sparse edge image; calculating internal pixel point information of the plurality of quadrilateral image blocks, combining the sparse edge image to generate a final fused image.
2. The method of large format image registration and fusion of claim 1, wherein, The establishment of the convolutional neural network, the input of the source image into the convolutional neural network, and the extraction of the convolutional layer feature information at different depths in the source image comprise: collecting a plurality of small area images centered on feature points as a training set; establishing a VGG-Net deep convolutional neural network model, training the deep convolutional neural network model based on the training set; extracting the convolutional layer feature information at different depths in the source image based on the third, fourth and fifth layers of the trained deep convolutional neural network model.
3. The method of large format image registration and fusion of claim 2, wherein, The extraction of the convolutional layer feature information at different depths in the source image based on the third, fourth and fifth layers of the trained deep convolutional neural network model comprises: combining a plurality of convolution kernels with a size of 3x3 in the third, fourth and fifth layers of the trained deep convolutional neural network model to generate a combined convolution layer with a plurality of convolution kernels with a size of 7x7, and extracting the convolutional layer feature information at different depths in the source image based on the combined convolution layer.
4. The method of large format image registration and fusion of claim 1, wherein, The filtering of the convolutional layer feature information at different depths based on the correlation filter to calculate the position information of the feature point on the source image on the target image comprises: establishing a correlation filtering model based on the convolutional layer feature information at different depths in the source image, calculating the similarity of the feature point pair in each depth of the source image using a correlation filtering algorithm, and determining the position of the maximum similarity response value; based on the position information of the maximum similarity response value, combining the convolutional layer feature information at different depths in the source image, and comprehensively calculating the position information of the feature point on the source image on the target image.
5. The method of large format image registration and fusion of claim 4, wherein, The establishment of the correlation filtering model based on the convolutional layer feature information at different depths in the source image, the calculation of the similarity of the matching point in each depth of the source image using the correlation filtering algorithm, and the determination of the position of the maximum similarity response value comprise: matching the feature point pair in the source image using a kernel correlation filtering algorithm and calculating the similarity response value of each feature point pair; determining whether the similarity response value is reasonable and deleting the incorrect feature point pair; sorting the similarity response values of the reasonable feature point pairs in the source image to obtain the feature point pair with the maximum similarity response value, and correspondingly obtaining the position information of the feature point pair with the maximum similarity response value.
6. The method of large format image registration and fusion of claim 5, wherein, The registration of the target image based on the position information comprises: Matching feature point pairs with same position information of feature point pairs with maximum similarity response value on the source image, and registering target images where the feature point pairs with same position information are located.
7. The method of large format image registration and fusion of claim 1, wherein, The step of decomposing the target image into a plurality of image blocks, generating corresponding two low-resolution images, and Poisson fusing the two low-resolution images into a sparse edge image includes: decomposing the target image into a plurality of quadrilateral image blocks, discarding pixel points inside each of the plurality of quadrilateral image blocks, and generating a plurality of quadrilateral image blocks with retained boundary pixel points; constructing a complete binary tree structure model, inputting the plurality of quadrilateral image blocks with retained boundary pixel points into the complete binary tree structure model, and generating a low-resolution image in the horizontal direction and a low-resolution image in the vertical direction. The step of decomposing the target image into a plurality of image blocks, generating corresponding two low-resolution images, and Poisson fusing the two low-resolution images into a sparse edge image includes:
8. The method of large format image registration and fusion of claim 7, wherein, constructing a complete binary tree structure model, inputting the plurality of quadrilateral image blocks with retained boundary pixel points into the complete binary tree structure model, and generating a low-resolution image in the horizontal direction and a low-resolution image in the vertical direction. The step of calculating internal pixel point information of the plurality of quadrilateral image blocks, combining the sparse edge image for fusion, and generating a final fused image includes:
9. The method of large format image registration and fusion of claim 1, wherein, using a GPU parallel method to interpolate internal pixel point information of each of the plurality of quadrilateral image blocks, combining basic information values of the sparse edge image, and calculating a basic information value of a final fused image. The device includes:
10. An apparatus for large format image registration and fusion, comprising: a feature extraction module configured to establish a convolutional neural network, input a source image into the convolutional neural network, and extract convolutional layer feature information at different depths in the source image; a correlation filtering module configured to filter process the convolutional layer feature information at different depths, and calculate position information of feature point pairs on a target image; a registration module configured to register the target image based on the position information; a sparse edge image generation module configured to decompose the target image into a plurality of quadrilateral image blocks, generate corresponding two low-resolution images, and Poisson fuse the two low-resolution images into a sparse edge image; a fusion module configured to calculate internal pixel point information of the plurality of quadrilateral image blocks, combine the sparse edge image for fusion, and generate a final fused image.
Citation Information
Patent Citations
A hierarchical remote sensing image fusion method using layer-by-layer iterative super-resolution is presented
CN109509160A
Face super-resolution method and device based on combined learning
CN110580680A