Remote sensing image matching method based on comparative learning
Through the method based on contrast learning, the joint training of feature descriptor extraction network and twin network is solved, and the existing remote sensing image matching method is difficult to maintain rotation invariance, achieving high-precision and high-root multi-source remote sensing image matching.
Patent Information
- Application Number
- CN202510460152.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing remote sensing image matching methods based on supervised deep learning are difficult to maintain rotational invariance and other problems.
Using a method based on contrast learning, by establishing two sets of one-to-one corresponding and center-aligned training image data sets, the network is extracted using feature descriptors to calculate the loss function, and the twin network is jointly trained to learn common features between multi-source images.
It realizes high-precision and high-root multi-source remote sensing image matching, which can effectively deal with problems such as rotation and geometric transformation.
Smart Images

Figure CN119992136A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of remote sensing image matching, and in particular to a remote sensing image matching method based on contrast learning. Background Art
[0002] With the rapid development of aerospace and remote sensing technology, the means of acquiring remote sensing images are increasing and the types are becoming more diverse. Due to the differences in equipment technology and imaging mechanisms of various sensors, remote sensing images from a single data source are difficult to fully reflect the characteristics of ground objects. In order to make full use of multi-source remote sensing data obtained by different types of sensors and achieve integration and information complementarity, multi-source remote sensing images need to be matched.
[0003] Multi-source remote sensing image matching refers to the process of spatial alignment and feature correspondence of multi-sensor remote sensing images of the same area acquired at different times, different perspectives or different sensor conditions. In the prior art, the methods of multi-source remote sensing image matching include traditional methods based on manual feature technology and methods based on deep learning. Traditional methods are based on features or regional templates, which rely on manually designed features. For matching remote sensing images of different sensors and different modalities, these manual features usually need to be redesigned. The method based on deep learning extracts deep features from multi-source remote sensing images, which has better versatility than manual features. Compared with traditional manual features, the use of deep neural networks can extract more abstract and robust deep features of images, which is conducive to improving the performance of image matching. Therefore, exploring the precise matching technology of multi-modal remote sensing images expressed by deep features has important research significance and application value. Summary of the invention
[0004] The embodiments of the present application provide a remote sensing image matching method based on contrastive learning to solve the problem that the existing remote sensing image matching method based on supervised deep learning is difficult to maintain rotation invariance.
[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.
[0006] According to a first aspect of an embodiment of the present application, a remote sensing image matching method based on contrastive learning is provided, comprising: Establishing a first training image data set and a second training image data set, wherein the image data in the two training image data sets correspond to each other and are centrally aligned; Extracting image pairs for training from the first training image data set and the second training image data set; Based on the image pair, calculating a loss function using a feature descriptor extraction network; Training the feature descriptor extraction network based on a loss function; Two sets of multi-source remote sensing images are obtained, and remote sensing image matching is performed on the two sets of multi-source remote sensing images using the trained feature descriptor extraction network.
[0007] In some embodiments of the present application, based on the above solution, extracting image pairs for training from the first training image dataset and the second training image dataset includes: Select a first image from the first training image dataset , and select from the second training image dataset the image corresponding to the first image The corresponding second image ; For the first image And the second image Perform the same random geometric transformation to generate the distorted first image and the distorted second image ; The first image The second image , the distorted first image and the distorted second image as image pairs for training.
[0008] In some embodiments of the present application, based on the aforementioned solution, the step of calculating the loss function based on the image pair using a feature descriptor extraction network includes: The first image and the distorted first image Input into the first twin network to extract the first image Feature descriptor and distorted first image Feature descriptor ; Feature descriptor based and feature descriptors Calculate the first image and the distorted first image Loss function between two homologous images ; The second image and the distorted second image Input into the second twin network to extract the second image Feature descriptor and the distorted second image Feature descriptor ; Feature descriptor based and feature descriptors Calculate the second image and the distorted second image Loss function between two homologous images ; Feature descriptor based and feature descriptors Calculate the first image And the second image Loss function between two heterogeneous images , and feature descriptor-based and feature descriptors Calculate the distorted first image and the distorted second image Loss function between two heterogeneous images ; Based on the loss function , loss function , loss function And the loss function Calculate the joint loss function ; Among them, the first twin network is composed of two feature descriptor extraction networks that share a first weight in parallel; the second twin network is composed of two feature descriptor extraction networks that share a second weight in parallel.
[0009] In some embodiments of the present application, based on the above scheme, the first image and the distorted first image Input into the first twin network to extract the first image Feature descriptor and distorted first image Feature descriptor ,include: The first image and the distorted first image Randomly crop the image into 128×128 pixel patches with aligned center points, and then apply a circular mask to the image patches; The image block is input into the first twin network for feature extraction to obtain the first image with a length of 256 dimensions. Feature descriptor and distorted first image Feature descriptor .
[0010] In some embodiments of the present application, based on the above solution, the second image and the distorted second image Input into the second twin network to extract the second image Feature descriptor and the distorted second image Feature descriptor ,include: The second image and the distorted second image Randomly crop the image into 128×128 pixel patches with aligned center points, and then apply a circular mask to the image patches; The image block is input into the second twin network for feature extraction to obtain a second image with a length of 256 dimensions. Feature descriptor and the distorted second image Feature descriptor .
[0011] In some embodiments of the present application, based on the aforementioned solution, the training of the feature descriptor extraction network based on the loss function includes: Using the loss function Training the first twin network; Using the loss function Training the second twin network; The first twin network and the second twin network are jointly trained using the joint loss function.
[0012] In some embodiments of the present application, based on the aforementioned scheme, during the training of the first twin network, the two feature descriptor extraction networks constituting the first twin network share a first weight, and during the training of the second twin network, the two feature descriptor extraction networks constituting the second twin network share a second weight, and during the joint training process, weights are not shared between the first twin network and the second twin network. In some embodiments of the present application, based on the above-mentioned solution, the method of performing remote sensing image matching on two sets of multi-source remote sensing images using the trained feature descriptor extraction network includes: Based on the trained feature descriptor extraction network, feature descriptors of two groups of multi-source remote sensing images are extracted respectively to obtain two feature descriptor sets; The feature descriptors in two feature descriptor sets are paired to complete remote sensing image matching.
[0013] In some embodiments of the present application, based on the aforementioned solution, the feature descriptor extraction network based on the training respectively extracts feature descriptors of two groups of multi-source remote sensing images to obtain two feature descriptor sets, including: The pixel coordinates of multiple feature points in two sets of multi-source remote sensing images are extracted using feature point detection algorithms. Based on the pixel coordinates of each feature point, multiple image blocks are cropped from two sets of multi-source remote sensing images with each feature point as the center; Multiple image blocks are input into the trained feature descriptor extraction network to extract the radiation and geometric invariant feature vectors of each feature point as the feature descriptor of each feature point.
[0014] In some embodiments of the present application, based on the above solution, pairing the feature descriptors in the two feature descriptor sets to complete remote sensing image matching includes: Calculate the Euclidean distance between multiple feature descriptors on two sets of multi-source images; Assign the feature point with the closest Euclidean distance to each feature point in the two sets of multi-source images as the initial matching point pair; Based on the initial matching point pair, an erroneous matching elimination method is adopted to eliminate the error points in the initial matching point pair, and the final retained initial matching point pair is used as output to complete the matching.
[0015] According to a second aspect of an embodiment of the present application, a remote sensing image matching device based on contrast learning is provided, comprising: An establishing unit, used for establishing a first training image data set and a second training image data set, wherein the image data in the two sets of training image data sets correspond to each other one by one and are centrally aligned; A first extraction unit, configured to extract image pairs for training from the first training image data set and the second training image data set; A second extraction unit, configured to calculate a loss function based on the image pair using a feature descriptor extraction network; A training unit, used for training the feature descriptor extraction network based on a loss function; The matching unit is used to obtain two sets of multi-source remote sensing images and use the trained feature descriptor extraction network to perform remote sensing image matching on the two sets of multi-source remote sensing images.
[0016] According to a third aspect of an embodiment of the present application, there is provided an electronic device, including: a memory and a processor; The memory is used to store computer instructions; The processor is used to call the computer instructions stored in the memory so that the electronic device executes the method as described in the first aspect.
[0017] The technical solution of this application has the following beneficial effects: This application adopts the form of combining two twin networks for joint training, effectively learns the common features between multi-source images, updates the parameters of each model network by back propagation, and achieves high-precision and high-robust multi-source remote sensing image matching.
[0018] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings: Figure 1 A schematic diagram of a process of a remote sensing image matching method based on contrast learning according to an embodiment of the present application is shown; Figure 2 A schematic diagram of a process for calculating a loss function using a feature descriptor extraction network according to an embodiment of the present application is shown; Figure 3 A schematic diagram of the structure of two twin networks according to an embodiment of the present application is shown; Figure 4 A schematic diagram of a construction process of a descriptor matrix according to an embodiment of the present application is shown; Figure 5 A block diagram of a remote sensing image matching device based on contrast learning according to an embodiment of the present application is shown; Figure 6 A block diagram of an electronic device according to an embodiment of the present application is shown; Figure 7 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown. DETAILED DESCRIPTION
[0020] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more comprehensive and complete and fully convey the concept of the example embodiments to those skilled in the art.
[0021] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0022] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0023] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the objects used in this way can be interchanged where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those shown or described.
[0025] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all the embodiments. Based on the embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0026] Some embodiments of the present application will be described in detail below in conjunction with the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.
[0027] See also Figure 1 , shows a flow chart of a remote sensing image matching method based on contrast learning according to an embodiment of the present application.
[0028] like Figure 1As shown, a remote sensing image matching method based on contrast learning is demonstrated, including steps S100 to S500.
[0029] refer to Figure 1 , step S100, establishing a first training image data set and a second training image data set, the image data in the two sets of training image data sets correspond one to one and are aligned in center.
[0030] It should be noted that the image data in the first training image data set and the second training image data set in this embodiment include image data acquired by heterogeneous or non-heterogeneous remote sensing platforms such as visible light, infrared, SAR, hyperspectral, and multispectral.
[0031] Continue to refer Figure 1 , step S200, extracting image pairs for training from the first training image data set and the second training image data set.
[0032] In some feasible embodiments, based on the above solution, extracting image pairs for training from the first training image dataset and the second training image dataset includes: Select a first image from the first training image dataset , and select from the second training image dataset the image corresponding to the first image The corresponding second image ; For the first image And the second image Perform the same random geometric transformation to generate the distorted first image and the distorted second image ; The first image The second image , the distorted first image and the distorted second image as image pairs for training.
[0033] It is understandable that the first image With the second image Center alignment.
[0034] Exemplarily, taking the matching of optical images and synthetic aperture radar (SAR) images as an example, this embodiment uses fixed-resolution images a and b as original images, and the pixel-level correspondence between them is known. Random affine transformations including large rotations and scaling are added to images a and b, and corresponding pixel correspondences are generated at the same time to obtain corresponding distorted images a', b'. After being processed by the method provided in this embodiment, a descriptor with geometric consistency and radiation consistency corresponding to the original image is obtained. The multi-source remote sensing image dataset contains multiple pairs of regional images similar to the above-mentioned images a and b.
[0035] It should be understood that other embodiments of the present application, including but not limited to the matching of multi-source optical images, the matching of optical images and infrared images, and the matching method provided by the present application should all be within the protection scope of the present invention.
[0036] Continue to refer Figure 1 , step S300, based on the image pair, using a feature descriptor extraction network to calculate a loss function.
[0037] For example, see Figure 2 The specific process of using feature descriptors to extract the network to calculate the loss function is as follows Figure 2 shown.
[0038] In some feasible embodiments, based on the above scheme, the step of calculating the loss function based on the image pair using a feature descriptor extraction network includes: The first image and the distorted first image Input into the first twin network to extract the first image Feature descriptor and distorted first image Feature descriptor ; Feature descriptor based and feature descriptors Calculate the first image and the distorted first image Loss function between two homologous images ; The second image and the distorted second image Input into the second twin network to extract the second image Feature descriptor and the distorted second image Feature descriptor ; Feature descriptor based and feature descriptors Calculate the second image and the distorted second image Loss function between two homologous images ; Feature descriptor based and feature descriptors Calculate the first image And the second image Loss function between two heterogeneous images , and feature descriptor-based and feature descriptors Calculate the distorted first image and the distorted second image Loss function between two heterogeneous images ; Based on the loss function , loss function , loss function And the loss function Calculate the joint loss function ; Among them, the first twin network is composed of two feature descriptor extraction networks that share a first weight in parallel; the second twin network is composed of two feature descriptor extraction networks that share a second weight in parallel.
[0039] In some feasible embodiments, based on the above scheme, the first image and the distorted first image Input into the first twin network to extract the first image Feature descriptor and distorted first image Feature descriptor ,include: The first image and the distorted first image Randomly crop the image into 128×128 pixel patches with aligned center points, and then apply a circular mask to the image patches; The image block is input into the first twin network for feature extraction to obtain the first image with a length of 256 dimensions. Feature descriptor and distorted first image Feature descriptor .
[0040] In some feasible embodiments, based on the above solution, the second image and the distorted second image Input into the second twin network to extract the second image Feature descriptor and the distorted second image Feature descriptor ,include: The second image and the distorted second image Randomly crop the image into 128×128 pixel patches with aligned center points, and then apply a circular mask to the image patches; The image block is input into the second twin network for feature extraction to obtain a second image with a length of 256 dimensions. Feature descriptor and the distorted second image Feature descriptor .
[0041] Exemplarily, a first image a and a distorted first image are provided. The specific extraction process of the feature descriptor: The first image and distorted first image The image patches are randomly cropped to 128×128 pixel patches with their center points aligned, and then circular masks are applied to the image patches to make the model adaptable to changes in the input image angle.
[0042] The first image and distorted first image Input to the first twin network to generate deep features and output a feature descriptor with a length of 256 dimensions and .
[0043] like Figure 3As shown in the figure, the feature extraction part of the first twin network consists of 9 groups of interconnected convolution blocks. The first convolution block includes a normal convolution, a batch normalization layer, and a SiLU activation function. The second to eighth convolution blocks consist of 1×1 convolution, depth convolution, Squeeze-and-Excitation (SE) blocks of channel attention mechanism, and Dropout. The SE block consists of a global average pooling and two fully connected layers. The last convolution block includes a normal convolution layer and an L2 normalization layer. In order to adapt the network to output a one-dimensional vector suitable for feature descriptor extraction, the last layer operation is changed from the original "1×1 convolution + pooling + full connection" mode to "4×4 global convolution + L2 normalization" of Conv+L2Norm operation. In this example, an image block with a size of 128×128×1 pixels centered on the feature point is input, and the first twin network outputs a feature descriptor of 1×1×256 length used to describe the geometric and radiation invariant features of the feature point.
[0044] It should be noted that Figure 3 In the figure, ConvBS indicates that the convolution block has a batch normalization layer after the convolution operation, MBConv indicates an inverted linear bottleneck layer with depthwise separable convolution, and ConvL indicates that the convolution block has an L2 normalization layer after the convolution operation.
[0045] It should be noted that the second image and the distorted second image The specific extraction process of the feature descriptor is similar to that of the first image and distorted first image The specific extraction process of the feature descriptor is the same as that of the second image. and the distorted second image The feature descriptor is extracted by inputting into the second twin network.
[0046] For example, a loss function is provided The calculation process: Loss Function The calculation formula is as follows: ; (1) in, represents Triplet Margin Loss (triplet loss), represents the regularization term loss.
[0047] The calculation process is as follows: To calculate the feature descriptor and Taking the radiation consistency of as an example, construct the descriptor matrix based on the training batch, such as Figure 3 As shown: Assume the training batch size is , The image block is extracted through the feature descriptor network, and the output is For the feature descriptor, the formula (2) is used to calculate Each descriptor in the sequence and Each descriptor in the sequence The distance is Descriptor distance , forming a descriptor matrix based on the training batch .
[0048] . (2) Construct positive and negative sample pairs. Figure 4 As shown, all elements on the main diagonal of the matrix represent quantities The distance between the positive sample pairs (matching pairs) is marked in green font in the figure, and all elements on the non-main diagonal line represent the number The distance of negative sample pairs (non-matching pairs) is marked in black font in the figure. In the calculation of Triplet Margin Loss, each positive sample corresponds to only one negative sample, and it is not necessary to use all negative sample pairs to calculate the distance. Therefore, it is necessary to find the most difficult negative sample pair for each positive sample pair. For example, its row index is 1 and column index is 1, so we traverse all the elements in the row of the distance matrix respectively. and all elements in the column , this The elements on the non-diagonal lines are sorted by distance value, and the element with the smallest distance value is selected as the negative sample.
[0049] Finally, the Triplet Margin Loss based on the training batch is calculated. In a training batch, the positive and negative sample pairs are repeatedly constructed. For each positive sample pair element on the main diagonal, a corresponding negative sample pair is found. To facilitate the subsequent formula expression, the distance between the positive sample pairs is recorded as , the distance between negative samples is ,in Represents two descriptors (feature vectors) and The angle between them is calculated as: ; (3) ; (4) in, A hyperparameter representing the target loss Margin. It can reflect the difficulty of obtaining samples.
[0050] L2 normalization is beneficial to enhance the consistency of matching sample pairs. The performance of L2 normalized descriptors is better than that of early non-normalized descriptors. However, L2 normalization also causes information loss parallel to the feature direction. Therefore, an improved Triplet Margin Loss proposed in HyNet is adopted. This loss proposes a hybrid similarity metric based on Triplet Margin Loss. , instead of the L2 distance metric between descriptors or related metrics , the formula is as follows: ; (5) in, is a hyperparameter that balances the distance metric and the correlation metric. The Triplet Margin Loss formula improved by the hybrid similarity metric is as follows: ; (6) At the same time, the L2 regularization term that calculates the difference between non-normalized features is added to provide appropriate constraints for L2 normalization, which reduces the adverse effects of L2 normalization on descriptor learning. L2 regularization is conducive to adding constraints to descriptors, inhibiting their changes with the scaling of image intensity, and promoting the robustness of the network to changes in image intensity, thereby improving the learning of descriptors. The calculation formula of the L2 regularization term is as follows: ; (7) in and is a pair of positive sample descriptors before L2 normalization in the network structure.
[0051] It is understandable that the loss function With loss function The calculation process is the same, the difference is only in the parameters.
[0052] For example, the loss function The calculation formula is as follows: ; (8) in, represents Triplet Margin Loss (triplet loss), represents the regularization term loss.
[0053] Joint loss function The calculation formula is as follows: ; (9) Continue to refer Figure 1 , step S400, training the feature descriptor extraction network based on the loss function.
[0054] In some feasible embodiments, based on the above solution, the training of the feature descriptor extraction network based on the loss function includes: Using the loss function Training the first twin network; Using the loss function Training the second twin network; The first twin network and the second twin network are jointly trained using the joint loss function.
[0055] In some feasible embodiments, based on the aforementioned scheme, during the training of the first twin network, the two feature descriptor extraction networks constituting the first twin network share a first weight, and during the training of the second twin network, the two feature descriptor extraction networks constituting the second twin network share a second weight, and during the joint training process, weights are not shared between the first twin network and the second twin network.
[0056] Continue to refer Figure 1 , step S500, obtaining two groups of multi-source remote sensing images, and using the trained feature descriptor extraction network to perform remote sensing image matching on the two groups of multi-source remote sensing images.
[0057] In some feasible embodiments, based on the above scheme, the method of performing remote sensing image matching on two sets of multi-source remote sensing images using the trained feature descriptor extraction network includes: Based on the trained feature descriptor extraction network, feature descriptors of two groups of multi-source remote sensing images are extracted respectively to obtain two feature descriptor sets; The feature descriptors in two feature descriptor sets are paired to complete remote sensing image matching.
[0058] In some feasible embodiments, based on the above scheme, the feature descriptor extraction network based on the training respectively extracts feature descriptors of two groups of multi-source remote sensing images to obtain two feature descriptor sets, including: The pixel coordinates of multiple feature points in two sets of multi-source remote sensing images are extracted using feature point detection algorithms. Based on the pixel coordinates of each feature point, multiple image blocks are cropped from two sets of multi-source remote sensing images with each feature point as the center; Multiple image blocks are input into the trained feature descriptor extraction network to extract the radiation and geometric invariant feature vectors of each feature point as the feature descriptor of each feature point.
[0059] In some feasible embodiments, based on the above solution, pairing the feature descriptors in the two feature descriptor sets to complete remote sensing image matching includes: Calculate the Euclidean distance between multiple feature descriptors on two sets of multi-source images; Assign the feature point with the closest Euclidean distance to each feature point in the two sets of multi-source images as the initial matching point pair; Based on the initial matching point pair, an erroneous matching elimination method is adopted to eliminate the error points in the initial matching point pair, and the final retained initial matching point pair is used as output to complete the matching.
[0060] It should be noted that the false matching elimination method may be RANSAC, GENSAC, least squares method, etc.
[0061] The following describes an embodiment of the device of the present application, which can be used to execute a remote sensing image matching method based on contrast learning in the above embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the above method of the present application.
[0062] Reference Figure 5 As shown, according to an embodiment of the present application, a remote sensing image matching device 500 based on contrast learning includes: An establishing unit 501 is used to establish a first training image data set and a second training image data set, wherein the image data in the two training image data sets correspond to each other and are aligned in center; A first extraction unit 502, configured to extract image pairs for training from the first training image data set and the second training image data set; A second extraction unit 503 is used to calculate a loss function based on the image pair using a feature descriptor extraction network; A training unit 504, configured to train the feature descriptor extraction network based on a loss function; The matching unit 505 is used to obtain two groups of multi-source remote sensing images and perform remote sensing image matching on the two groups of multi-source remote sensing images using the trained feature descriptor extraction network.
[0063] like Figure 6 As shown, an embodiment of the present application also provides an electronic device 600, including a memory 610, a processor 620, and a computer program 611 stored in the memory 610 and executable on the processor. When the processor 620 executes the computer program 611, the steps of the above-mentioned remote sensing image matching method based on contrast learning are implemented.
[0064] Since the electronic device introduced in this embodiment is a device used to implement a remote sensing image matching device based on contrast learning in the embodiment of the present application, based on the method introduced in the embodiment of the present application, the technical personnel in this field can understand the specific implementation mode of the electronic device of this embodiment and its various variations. Therefore, how the electronic device implements the method in the embodiment of the present application is not introduced in detail here. As long as the equipment used by the technical personnel in this field to implement the method in the embodiment of the present application is within the scope of protection of this application.
[0065] During the specific implementation process, when the computer program 611 is executed by a processor, any implementation method in the embodiments corresponding to the first aspect can be implemented.
[0066] Figure 7 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown.
[0067] It should be noted that Figure 7 The computer system 700 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0068] like Figure 7 As shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage part 708 to the random access memory (RAM) 703, such as executing the method described in the above embodiment. In the RAM 703, various programs and data required for system operation are also stored. The CPU 701, the ROM 702 and the RAM 703 are connected to each other through the bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.
[0069] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read therefrom is installed into the storage section 708 as needed.
[0070] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 709, and / or installed from a removable medium 711. When the computer program is executed by the central processing unit (CPU) 701, various functions defined in the system of the present application are executed.
[0071] It should be noted that the computer-readable medium shown in the embodiment of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0072] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. Wherein, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0073] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.
[0074] As another aspect, the present application also provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes a remote sensing image matching method based on contrast learning described in the above embodiment.
[0075] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements a remote sensing image matching method based on contrast learning described in the above embodiment.
[0076] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.
[0077] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software, or by combining software with necessary hardware. Therefore, the technical solution according to the implementation method of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the implementation method of the present application.
[0078] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary techniques in the art that are not disclosed in the present application. It should be understood that the present application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A remote sensing image matching method based on contrastive learning, characterized in that: include: Establishing a first training image data set and a second training image data set, wherein the image data in the two training image data sets correspond to each other and are centrally aligned; Extracting image pairs for training from the first training image data set and the second training image data set; Based on the image pair, calculating a loss function using a feature descriptor extraction network; Training the feature descriptor extraction network based on a loss function; Two sets of multi-source remote sensing images are obtained, and remote sensing image matching is performed on the two sets of multi-source remote sensing images using the trained feature descriptor extraction network.
2. The method according to claim 1, characterized in that The step of extracting image pairs for training from the first training image data set and the second training image data set includes: Select a first image from the first training image dataset , and select from the second training image dataset the image corresponding to the first image The corresponding second image ; For the first image And the second image Perform the same random geometric transformation to generate the distorted first image and the distorted second image ; The first image The second image , the distorted first image and the distorted second image as image pairs for training.
3. The method according to claim 2, characterized in that The step of calculating the loss function based on the image pair using a feature descriptor extraction network includes: The first image and the distorted first image Input into the first twin network to extract the first image Feature descriptor and distorted first image Feature descriptor ; Feature descriptor based and feature descriptors Calculate the first image and the distorted first image Loss function between two homologous images ; The second image and the distorted second image Input into the second twin network to extract the second image Feature descriptor and the distorted second image Feature descriptor ; Feature descriptor based and feature descriptors Calculate the second image and the distorted second image Loss function between two homologous images ; Feature descriptor based and feature descriptors Calculate the first image And the second image Loss function between two heterogeneous images , and feature descriptor-based and feature descriptors Calculate the distorted first image and the distorted second image Loss function between two heterogeneous images ; Based on the loss function , loss function , loss function And the loss function Calculate the joint loss function ; Among them, the first twin network is composed of two feature descriptor extraction networks that share a first weight in parallel; the second twin network is composed of two feature descriptor extraction networks that share a second weight in parallel.
4. The method according to claim 3, characterized in that The first image and the distorted first image Input into the first twin network to extract the first image Feature descriptor and distorted first image Feature descriptor ,include: The first image and the distorted first image Randomly crop the image into 128×128 pixel patches with aligned center points, and then apply a circular mask to the image patches; The image block is input into the first twin network for feature extraction to obtain the first image with a length of 256 dimensions. Feature descriptor and distorted first image Feature descriptor .
5. The method according to claim 3, characterized in that: The second image and the distorted second image Input into the second twin network to extract the second image Feature descriptor and the distorted second image Feature descriptor ,include: The second image and the distorted second image Randomly crop the image into 128×128 pixel patches with aligned center points, and then apply a circular mask to the image patches; The image block is input into the second twin network for feature extraction to obtain a second image with a length of 256 dimensions. Feature descriptor and the distorted second image Feature descriptor .
6. The method according to claim 3, characterized in that The training of the feature descriptor extraction network based on the loss function includes: Using the loss function Training the first twin network; Using the loss function Training the second twin network; The first twin network and the second twin network are jointly trained using the joint loss function.
7. The method according to claim 6, characterized in that During the training of the first twin network, the two feature descriptor extraction networks constituting the first twin network share a first weight. During the training of the second twin network, the two feature descriptor extraction networks constituting the second twin network share a second weight. During the joint training, no weights are shared between the first twin network and the second twin network.
8. The method according to claim 1, characterized in that The method of performing remote sensing image matching on two sets of multi-source remote sensing images using the trained feature descriptor extraction network includes: Based on the trained feature descriptor extraction network, feature descriptors of two groups of multi-source remote sensing images are extracted respectively to obtain two feature descriptor sets; The feature descriptors in two feature descriptor sets are paired to complete remote sensing image matching.
9. The method according to claim 8, characterized in that The feature descriptor extraction network based on the training is used to extract feature descriptors of two groups of multi-source remote sensing images respectively, and obtain two feature descriptor sets, including: The pixel coordinates of multiple feature points in two sets of multi-source remote sensing images are extracted using feature point detection algorithms. Based on the pixel coordinates of each feature point, multiple image blocks are cropped from two sets of multi-source remote sensing images with each feature point as the center; Multiple image blocks are input into the trained feature descriptor extraction network to extract the radiation and geometric invariant feature vectors of each feature point as the feature descriptor of each feature point.
10. The method according to claim 8, characterized in that The step of pairing the feature descriptors in the two feature descriptor sets to complete remote sensing image matching includes: Calculate the Euclidean distance between multiple feature descriptors on two sets of multi-source images; Assign the feature point with the closest Euclidean distance to each feature point in the two sets of multi-source images as the initial matching point pair; Based on the initial matching point pair, an erroneous matching elimination method is adopted to eliminate the error points in the initial matching point pair, and the final retained initial matching point pair is used as output to complete the matching.
Citation Information
Patent Citations
Optical-SAR remote sensing image cross-modal retrieval method based on modal common characteristics
CN115129917A
Multi-source image block matching method based on double-branch parallel depth interaction cooperation
CN116597177A
Unsupervised learning method based on multi-modal spectral image registration
CN117994302A
Contrast learning-based different-source image matching method and system
CN118334362A