Image registration method and electronic device

By training a lightweight student model through knowledge distillation and combining it with a geometric consistency algorithm, the problem of low registration accuracy of multimodal remote sensing images was solved, achieving efficient image registration on resource-constrained platforms and improving the accuracy and efficiency of remote sensing image fusion.

CN121414807BActive Publication Date: 2026-04-10INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Due to the differences in imaging mechanisms of different remote sensing devices, remote sensing images have different characteristics, which makes multimodal image fusion and interpretation difficult. Existing technologies struggle to achieve high-precision image registration, especially on resource-constrained embedded platforms, where computational complexity is high and registration accuracy is low.

Method used

A lightweight student model is trained using a knowledge distillation strategy. The student model is optimized by aligning the loss terms of the output layer and intermediate layers. Combined with the geometric consistency algorithm, multimodal images are processed in blocks. The target student model is used to process candidate matching pairs to obtain feature vectors, thereby achieving accurate registration of image transformation information.

Benefits of technology

It reduces computational complexity, improves image registration accuracy, meets the computational and storage resource limitations of embedded platforms, and achieves efficient multimodal image registration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121414807B_ABST
    Figure CN121414807B_ABST
Patent Text Reader

Abstract

The application provides an image registration method and electronic equipment, which can be applied to the fields of remote sensing, computer vision, digital image processing, artificial intelligence and embedding. The method comprises the following steps: determining matched candidate first modality image blocks and candidate second modality image blocks from a plurality of first modality image blocks and a plurality of second modality image blocks according to the first positions of the plurality of first modality image blocks and the second positions of the plurality of second modality image blocks, to obtain a plurality of candidate matching pairs; processing the plurality of candidate matching pairs by using a target student model to obtain a plurality of first modality feature vectors and a plurality of second modality feature vectors; and obtaining an image registration result indicating image transformation information between a first modality image and a second modality image according to the plurality of first modality feature vectors and the plurality of second modality feature vectors.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the fields of remote sensing, computer vision, digital image processing, artificial intelligence and embedded, and in particular to an image registration method and electronic equipment. BACKGROUND

[0002] Due to the difference of imaging mechanism of different remote sensing devices, the remote sensing images have different characteristics, so it is more and more important to fuse and interpret multi-modal remote sensing images. Multi-modal image registration is the premise of multi-modal image fusion and interpretation, and is of great significance to multi-modal image fusion.

[0003] Since the image registration between different modal remote sensing images is a pre-processing operation of image fusion, the image registration accuracy will affect the effectiveness and reliability of subsequent feature association, feature fusion and decision reasoning. Therefore, there is an urgent need for a method that can improve the image registration accuracy of multi-modal images. SUMMARY

[0004] In view of the above problems, the embodiments of the present application provide an image registration method and electronic equipment.

[0005] According to a first aspect of the embodiments of the present application, an image registration method is provided, comprising: determining matched candidate first modal image blocks and candidate second modal image blocks from a plurality of first modal image blocks and a plurality of second modal image blocks according to the first positions of the plurality of first modal image blocks and the second positions of the plurality of second modal image blocks, obtaining a plurality of candidate matching pairs; processing the plurality of candidate matching pairs by using a target student model to obtain a plurality of first modal feature vectors and a plurality of second modal feature vectors, the target student model being obtained by training a student model using a loss function comprising an output layer alignment loss term and an intermediate layer alignment loss term, the output layer alignment loss term being used to align a teacher feature quantity and a student feature vector obtained by processing a sample image block using a teacher model and a student model, the intermediate layer alignment loss term being used to align an intermediate teacher feature vector and an intermediate student feature vector obtained by processing the sample image block using the teacher model and the student model; obtaining an image registration result indicating image transformation information between the first modal image and the second modal image according to the plurality of first modal feature vectors and the plurality of second modal feature vectors.

[0006] According to the embodiments of the present application, the output layer alignment loss term is determined according to a similarity subterm and a constraint subterm, the similarity subterm is used to determine the similarity between the teacher feature vector and the student feature vector, and the constraint subterm is determined according to the model parameters of the student model; the intermediate layer alignment loss term is used to determine the similarity between the intermediate teacher feature vector and the intermediate student feature vector.

[0007] According to an embodiment of the present application, the sample image block comprises a sample first modality image block and a sample second modality image block; the student feature vector comprises a first student first modality feature vector of the sample first modality image block and a first student second modality feature vector of the sample second modality image block, the intermediate student feature vector comprises a first intermediate student first modality feature vector of the sample first modality image block and a first intermediate student second modality feature vector of the sample second modality image block; the student model comprises a student first modality branch, a student second modality branch and an interactive spatial attention module, the student first modality branch comprises a student first modality channel alignment unit and a student first modality adaptation unit, the student second modality branch comprises a student second modality channel alignment unit and a student second modality adaptation unit; the first student first modality feature vector is obtained by adjusting the number of feature channels of a second student first modality feature vector obtained by processing the sample first modality image block through other parts of the student first modality branch and the interactive spatial attention module, using the student first modality adaptation unit, the first student second modality feature vector is obtained by adjusting the number of feature channels of a second student second modality feature vector obtained by processing the sample second modality image block through other parts of the student second modality branch and the interactive spatial attention module, using the student second modality adaptation unit; the first intermediate student first modality feature vector is obtained by adjusting the number of feature channels of a second intermediate student first modality feature vector obtained by processing the sample first modality image block through other parts of the student first modality branch and the interactive spatial attention module, using the student first modality channel alignment unit, the first intermediate student second modality feature vector is obtained by adjusting the number of feature channels of a second intermediate student second modality feature vector obtained by processing the sample second modality image block through other parts of the student second modality branch and the interactive spatial attention module, using the student second modality channel alignment unit.

[0008] According to an embodiment of the present application, the sample image block comprises a sample first modality image block and a sample second modality image block; the teacher feature vector comprises a first teacher first modality feature vector of the sample first modality image block and a first teacher second modality feature vector of the sample second modality image block, and the intermediate teacher feature vector comprises a first intermediate teacher first modality feature vector of the sample first modality image block and a first intermediate teacher second modality feature vector of the sample second modality image block; the teacher model comprises a teacher first modality branch, a teacher second modality branch and an interactive feature fusion module; the first teacher first modality feature vector is obtained by processing a second teacher first modality feature vector corresponding to the sample first modality image block by using the teacher first modality branch and the interactive feature fusion module, and the first teacher second modality feature vector is obtained by processing a second teacher second modality feature vector corresponding to the sample second modality image block by using the teacher second modality branch and the interactive feature fusion module; the first intermediate teacher first modality feature vector is obtained by processing a second intermediate teacher first modality feature vector output by the target intermediate layer by using the teacher first modality branch and the interactive feature fusion module, and the first intermediate teacher second modality feature vector is obtained by processing a second intermediate teacher second modality feature vector output by the target intermediate layer by using the teacher second modality branch and the interactive feature fusion module.

[0009] According to an embodiment of the present application, the image registration method further comprises: performing scale normalization processing on the first modality image and the second modality image to obtain a first modality down-sampling image and a second modality down-sampling image; and performing block processing on the first modality down-sampling image and the second modality down-sampling image to obtain a plurality of first modality image blocks and a plurality of second modality image blocks.

[0010] According to an embodiment of the present application, the image registration result indicating the image transformation information between the first modality image and the second modality image is obtained according to the plurality of first modality feature vectors and the plurality of second modality feature vectors, comprising: determining a cosine similarity matrix according to the plurality of first modality feature vectors and the plurality of second modality feature vectors; in the case that the maximum value feature vector determined in the i th row of the cosine similarity matrix is the maximum value feature vector in the j th column, determining a pixel point corresponding to the i th first modality feature vector and a pixel point corresponding to the j th second modality feature vector as an initial matching point pair; and processing the initial matching point pair by using a geometric consistency algorithm to obtain the image registration result indicating the image transformation information between the first modality image and the second modality image.

[0011] According to an embodiment of the present application, the initial matching point pairs are processed by using a geometric consistency algorithm to obtain an image registration result indicating image transformation information between the first modality image and the second modality image, comprising: in the case of i>1, determining the i-th round of matching point pairs according to the (i-1)-th round of second modality projection pixel point set and the second modality pixel point set, the second modality pixel point set being obtained according to the initial matching point pairs, the first round of second modality projection pixel point set being obtained according to the first round of affine transformation matrix processed on the initial matching point pairs, the first round of affine transformation matrix being obtained according to the least square method processed on the first round of matching point pairs determined from the initial matching point pairs; processing the i-th round of matching point pairs by using the least square method to obtain the i-th round of affine transformation matrix; processing the initial matching point pairs according to the i-th round of affine transformation matrix to obtain the i-th round of second modality projection pixel point set; in the case that the pixel difference between the projection pixel point in the i-th round of second modality projection pixel point set and the pixel point in the second modality pixel point set satisfies a preset pixel threshold, determining the i-th round of affine transformation matrix as the image registration result.

[0012] According to an embodiment of the present application, the first position is the longitude and latitude coordinates of the first modality center point, and the second position is the longitude and latitude coordinates of the second modality center point; the matched candidate first modality image block and the candidate second modality image block are determined from the plurality of first modality image blocks and the plurality of second modality image blocks according to the first positions of the plurality of first modality image blocks and the second positions of the plurality of second modality image blocks, to obtain a plurality of candidate matching pairs, comprising: determining the distance between each first modality image block and any second modality image block in the plurality of second modality image blocks according to the longitude and latitude coordinates of the first modality center point of each first modality image block and the longitude and latitude coordinates of the second modality center point of each second modality image block in the second modality image, to obtain a plurality of spatial distances; in the case that the spatial distance is less than or equal to a distance threshold, the first modality image block and the second modality image block corresponding to the spatial distance are respectively taken as the candidate first modality image block and the candidate second modality image block, to obtain the candidate matching pair.

[0013] The second aspect of the embodiment of the present application provides an image registration device, comprising: a candidate matching pair module, configured to determine matched candidate first modality image blocks and second modality image blocks from a plurality of first modality image blocks and a plurality of second modality image blocks according to respective first positions of the plurality of first modality image blocks and respective second positions of the plurality of second modality image blocks of a first modality image and a second modality image, to obtain a plurality of candidate matching pairs; a feature vector module, configured to process the plurality of candidate matching pairs by using a target student model to obtain a plurality of first modality feature vectors and a plurality of second modality feature vectors, the target student model being obtained by training a student model by using a loss function comprising an output layer alignment loss term and an intermediate layer alignment loss term, the output layer alignment loss term being used to align a teacher feature vector and a student feature vector obtained by processing a sample image block by using a teacher model and the student model, and the intermediate layer alignment loss term being used to align an intermediate teacher feature vector and an intermediate student feature vector obtained by processing the sample image block by using the teacher model and the student model; and a registration module, configured to obtain an image registration result indicating image transformation information between the first modality image and the second modality image according to the plurality of first modality feature vectors and the plurality of second modality feature vectors.

[0014] The third aspect of the embodiment of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method.

[0015] The fourth aspect of the embodiment of the present application further provides a computer program product, comprising computer programs or instructions, wherein the computer programs or instructions are executed by a processor to implement the steps of the method.

[0016] According to the embodiment of the present application, by dividing the first modality image and the second modality image into a plurality of image blocks respectively, the huge computing load caused by the large size characteristics of the image can be reduced, so that the large size image does not occupy a large amount of memory to cause high computing delay, according to the first position of the first modality image block and the second position of the second modality image block, the matched candidate first modality image block and the candidate second modality image block can be determined from the plurality of first modality image blocks and the plurality of second modality image blocks, a plurality of candidate matching pairs are obtained, through rough matching, invalid matching can be further reduced, and the subsequent processing efficiency is improved. Using the target student model to process a plurality of candidate matching pairs, a plurality of first modality feature vectors and a plurality of second modality feature vectors can be obtained, the target student model is obtained by training the student model using the loss function including the output layer alignment loss term and the intermediate layer alignment loss term, the output layer alignment loss term is used to align the teacher feature vector and the student feature vector obtained by processing the sample image block using the teacher model and the student model, and the intermediate layer alignment loss term is used to align the intermediate teacher feature vector and the intermediate student feature vector obtained by processing the sample image block using the teacher model and the student model. By using the above multi-level distillation strategy, the semantic gap between the complex teacher model and the lightweight student model can be bridged, so that the target student model can still inherit the robust understanding ability of the teacher model to the cross-modality geometric deformation under the condition of reducing the number of parameters. Finally, according to the plurality of first modality feature vectors and the plurality of second modality feature vectors, the image registration result indicating the image transformation information between the first modality image and the second modality image can be obtained, the whole process reduces the computing complexity and improves the image registration accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0017] The above and other objects, features and advantages of the present application will become more apparent from the following description when taken in conjunction with the accompanying drawings, in which:

[0018] Figure 1 An application scenario diagram of the image registration method and the image registration device according to the embodiment of the present application is shown.

[0019] Figure 2 A flowchart of the image registration method according to the embodiment of the present application is shown.

[0020] Figure 3 A structure diagram of the teacher model according to the embodiment of the present application is shown.

[0021] Figure 4 A structure diagram of the target student model according to the embodiment of the present application is shown.

[0022] Figure 5 A structure diagram of the knowledge distillation according to the embodiment of the present application is shown.

[0023] Figure 6 A structural block diagram of an image registration apparatus according to an embodiment of the present application is shown.

[0024] Figure 7 A block diagram of an electronic device suitable for implementing an image registration method according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0025] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. It is to be understood, however, that these descriptions are merely exemplary and are intended to illustrate the scope of the present application, not to limit it. In the following detailed description of the embodiments of the present application, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to one skilled in the art that the embodiments of the present application can be practiced without these specific details. In other instances, well-known structures and functions have been omitted or simplified in order not to obscure the concepts of the present application.

[0026] The terms used herein are merely used to describe specific embodiments, and are not intended to limit the present application. The terms "include" and "have" and the like used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0027] All terms used herein, including technical and scientific terms, have the same meanings as those generally understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having meanings consistent with the context of the present specification, and should not be interpreted in an idealized or excessively formal manner.

[0028] In the case of using expressions similar to "at least one of A, B, and C, etc.", it should generally be interpreted to include at least one of the items enumerated, but not limited to the items enumerated (e.g., "a system having at least one of A, B, and C" should include a system having A alone, a system having B alone, a system having C alone, a system having A and B together, a system having A and C together, a system having B and C together, and / or a system having A, B, and C together, etc.).

[0029] Image registration is a technique for spatially aligning images of the same scene acquired at different viewing angles, times, or sensor conditions, aiming to establish reliable feature correspondence. In the field of remote sensing, with the enrichment of earth observation means and the increase of revisit frequency, change detection, image fusion, or image stitching, which rely on high-precision collaborative analysis of different spatio-temporal and multi-modal remote sensing images, make remote sensing image registration an important and fundamental link.

[0030] Multi-modal remote sensing images can include not only visible light images, but also SAR (Synthetic Aperture Radar) images and infrared images. Due to the differences in imaging mechanisms, different modal images usually have complementary information characteristics. For example, visible light images have rich spectral information and are easy to interpret, and are suitable for fine object recognition, but are susceptible to light and weather. SAR images have all-weather and all-day imaging capabilities, are sensitive to surface geometric structure and dielectric properties, but have coherent speckle noise and are relatively difficult to interpret. Infrared images capture thermal radiation information on the ground, are extremely sensitive to temperature differences, can image at night or in low light conditions, and are suitable for geothermal anomaly detection; but their spatial resolution is usually low, and they are susceptible to atmospheric attenuation, changes in ground emissivity, and other factors.

[0031] It is increasingly important to promote the deployment of deep learning-based multi-modal image registration technology to resource-constrained edge side. Taking a spaceborne platform as an example, high-time-efficiency remote sensing monitoring tasks require the joint observation of optical and SAR satellites and the real-time completion of key data processing on board before being downloaded. In this scenario, it is necessary to achieve high-precision and high-efficiency registration of large-size visible light images and SAR images while meeting the strict computing, storage, and power consumption constraints of the embedded platform on board. Therefore, it is crucial to develop a lightweight registration network with high precision and low resource consumption and its embedded deployment method, which not only deepens the theory and compression technology of lightweight models, but also provides important support for real-time applications such as on-board fusion, on-board geographic correction, and cross-modal navigation. In addition, remote sensing images usually have large size characteristics. Directly processing the entire large-size image on the embedded platform on board will consume a lot of memory and have high computational delay, making it difficult to meet the requirements of high-time-efficiency tasks. If the network input is simply adapted by downsampling, a large amount of key details will be lost, resulting in a decrease in registration accuracy. Low registration accuracy can cause spatial displacement in the fused image, and in severe cases, it can cause misjudgment of ground objects or failure of change detection. Therefore, it is necessary to design a registration process suitable for large-size images, which decomposes the global registration problem of large-size images into local problems that can be computed in parallel and can be carried by the embedded platform, while ensuring accuracy.

[0032] In view of the huge computing load brought by the large size characteristics of remote sensing images, it is difficult to directly deploy a complex registration network on an embedded platform. Model compression is a potential way to adapt to embedded platforms by using a knowledge distillation strategy, that is, by designing a lightweight student model to learn the knowledge of a complex teacher model to reduce the number of parameters and the amount of calculation. However, this method faces the serious challenge of "knowledge gap": the teacher model usually has a deep structure, and its deep feature map contains high-order geometric semantic information about complex scenes (such as building deformation). In order to meet the resource constraints of embedded platforms, the student model is often designed to be more shallow and compact, resulting in a limited receptive field and insufficient representation ability, making it difficult to effectively capture and transfer the rich semantic knowledge of the deep layers of the teacher model. The significant dimensional gap between the deep high-order semantics of the teacher model and the shallow low-order features of the student model and the architecture mismatch make the knowledge transfer inefficient, forming a "knowledge gap", making it difficult for the student model to fully inherit the excellent performance of the teacher model, resulting in poor learning effect, and the generalization ability and final registration accuracy of the model are easily damaged, and the expected goal of knowledge distillation is difficult to achieve.

[0033] In order to at least partially solve the technical problems existing in the related art, the present application provides an image registration method. The method comprises: determining matched candidate first modality image blocks and candidate second modality image blocks from a plurality of first modality image blocks and a plurality of second modality image blocks according to the respective first positions of the plurality of first modality image blocks and the respective second positions of the plurality of second modality image blocks, obtaining a plurality of candidate matching pairs; processing the plurality of candidate matching pairs using a target student model to obtain a plurality of first modality feature vectors and a plurality of second modality feature vectors, the target student model being obtained by training a student model using a loss function comprising an output layer alignment loss term and an intermediate layer alignment loss term, the output layer alignment loss term being used to align a teacher feature vector and a student feature vector obtained by processing a sample image block using a teacher model and a student model, and the intermediate layer alignment loss term being used to align an intermediate teacher feature vector and an intermediate student feature vector obtained by processing the sample image block using the teacher model and the student model; and obtaining an image registration result indicating image transformation information between the first modality image and the second modality image according to the plurality of first modality feature vectors and the plurality of second modality feature vectors.

[0034] In the technical solutions of the embodiments of the present application, the user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved are all information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards, necessary security measures are taken, do not violate public order and good customs, and corresponding operation portals are provided for the user to choose authorization or refusal.

[0035] In the scenario of making automated decisions using personal information, the method, device and system provided by the embodiments of the present application all provide corresponding operation portals for the user to choose to agree or refuse the automated decision result; if the user chooses to refuse, the expert decision process is entered. The expression "automated decision" here refers to the activity of automatically analyzing, evaluating the behavior habits, interests and hobbies or economic, health, credit status of a person through a computer program and making decisions. The expression "expert decision" here refers to the activity of making decisions by personnel who are engaged in a certain field of work, have special experience, knowledge and skills and reach a certain professional level.

[0036] Figure 1 An application scenario diagram of an application scenario of an image registration method and an image registration apparatus according to an embodiment of the present application is shown.

[0037] It should be noted that, Figure 1 The shown are only examples of application scenarios to which the embodiments of the present application can be applied, to help understand the technical content disclosed in the present application, but do not mean that the embodiments of the present application cannot be applied to other devices, systems, environments or application scenarios. For example, in another embodiment, the application scenario of the image registration method and the image registration apparatus can include a terminal device, but the terminal device can not need to interact with the server, that is, the image registration method and the apparatus provided by the embodiments of the present application can be implemented.

[0038] As Figure 1 shown, the application scenario 100 according to the embodiment can include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104 and a server 105. The network 104 is a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the server 105. The network 104 can include various connection types. For example, wired, wireless communication link or optical fiber cable, etc.

[0039] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103. For example, a shopping application, a web browser application, a search application, an instant messaging tool, an email client, or a social platform software, etc.

[0040] The first terminal device 101, the second terminal device 102, or the third terminal device 103 can be various electronic devices with a display screen and supporting web browsing. For example, a smartphone, a tablet computer, a laptop computer, or a desktop computer, etc.

[0041] The server 105 can be a server providing various services, for example, a background management server supporting a website browsed by the user using the first terminal device 101, the second terminal device 102, or the third terminal device 103 (only as an example). The background management server can analyze and process the received user request and other data, and feed back the processing result (for example, a webpage, information, or data, etc. obtained or generated according to the user request) to the terminal device.

[0042] Optionally, the image registration method provided by the embodiments of the present application can be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the image registration apparatus provided by the embodiments of the present application can also be arranged in the first terminal device 101, the second terminal device 102, or the third terminal device 103.

[0043] Optionally, the image registration method provided by the embodiments of the present application can also be executed by the server 105. Correspondingly, the image registration apparatus provided by the embodiments of the present application can also be arranged in the server 105. In addition, the image registration method provided by the embodiments of the present application can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the image registration apparatus provided by the embodiments of the present application can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0044] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the above-mentioned system is only illustrative. According to the implementation needs, there can be any number of terminal devices, networks, and servers.

[0045] It should be noted that the sequence numbers of the various operations in the following method are only used to represent the operations for description, and should not be regarded as the execution sequence of the various operations. The method does not need to be executed in the order shown unless explicitly indicated.

[0046] Figure 2 A flow chart of an image registration method according to an embodiment of the present application is shown.

[0047] As shown in Figure 2 , the image registration method of this embodiment includes operations S210-S230, which can be executed by an electronic device. For example, the electronic device can be the first terminal device 101, the second terminal device 102, or the third terminal device 103 in Figure 1 . Alternatively, the electronic device can also be the server 105 in Figure 1 .

[0048] In operation S210, candidate matched first modality image blocks and candidate matched second modality image blocks are determined from a plurality of first modality image blocks and a plurality of second modality image blocks according to the first positions of the plurality of first modality image blocks and the second positions of the plurality of second modality image blocks, to obtain a plurality of candidate matching pairs.

[0049] In operation S220, the plurality of candidate matching pairs are processed using a target student model to obtain a plurality of first modality feature vectors and a plurality of second modality feature vectors.

[0050] In operation S230, an image registration result indicating image transformation information between the first modality image and the second modality image is obtained according to the plurality of first modality feature vectors and the plurality of second modality feature vectors.

[0051] The first modality image blocks can be represented as the first modality image being divided into a plurality of image blocks of the same size according to a preset rule. The second modality image blocks can be represented as the second modality image being divided into a plurality of image blocks of the same size according to a preset rule.

[0052] The first positions can represent the latitude and longitude coordinates corresponding to the first modality image blocks, for example, can be the latitude and longitude coordinates of the center point, the latitude and longitude coordinates of the upper left corner, or the latitude and longitude coordinates of the upper right corner of the first modality image blocks, and the like. The second positions can represent the latitude and longitude coordinates corresponding to the second modality image blocks, for example, can be the latitude and longitude coordinates of the center point, the latitude and longitude coordinates of the upper left corner, or the latitude and longitude coordinates of the upper right corner of the second modality image blocks, and the like. It should be noted that the sizes of the plurality of image blocks corresponding to the first modality image and the second modality image are uniform. The image block size can be adjusted according to system resources. In the case of sufficient system resources, large-size image blocks can be used to improve feature integrity. In the case of insufficient system resources, the image can be divided into a plurality of small-size image blocks.

[0053] The matched candidate first modality image block and the candidate second modality image block can be determined from the plurality of first modality image blocks and the plurality of second modality image blocks according to the first positions of the plurality of first modality image blocks and the second positions of the plurality of second modality image blocks, respectively, to obtain a plurality of candidate matching pairs. For example, the first modality image is divided into four first modality image blocks (A, B, C and D), and the second modality image is divided into four second modality image blocks (a, b, c and d). According to the comparison of the plurality of first positions and the plurality of second positions, A and b are determined as a candidate matching pair, and C and d are determined as a candidate matching pair.

[0054] The target student model is obtained by training the student model using a loss function including an output layer alignment loss term and an intermediate layer alignment loss term. The output layer alignment loss term is used to align a teacher feature vector and a student feature vector obtained by processing a sample image block using a teacher model and a student model. The intermediate layer alignment loss term is used to align an intermediate teacher feature vector and an intermediate student feature vector obtained by processing the sample image block using the teacher model and the student model. The target student model is used to process the plurality of candidate matching pairs to obtain a plurality of first modality feature vectors and a plurality of second modality feature vectors. According to the plurality of first modality feature vectors and the plurality of second modality feature vectors, an image registration result indicating image transformation information between the first modality image and the second modality image can be obtained. The image transformation information can include a first modality pixel point coordinate in the first modality image and a second modality pixel point coordinate registered with the first modality pixel point coordinate, and an affine transformation matrix. The affine transformation matrix can be used to determine the pixel point position in the first modality image registered with the second modality image.

[0055] According to the embodiment of the present application, by dividing the first modality image and the second modality image into a plurality of image blocks respectively, the huge computational load caused by the large size characteristics of the image can be reduced, so that the large size image does not monopolize a large amount of memory to cause high computational delay, according to the first position of the first modality image block and the second position of the second modality image block, the matched candidate first modality image block and the candidate second modality image block can be determined from the plurality of first modality image blocks and the plurality of second modality image blocks, to obtain a plurality of candidate matching pairs, through rough matching, invalid matching can be further reduced, and the subsequent processing efficiency can be improved. By processing a plurality of candidate matching pairs using a target student model, a plurality of first modality feature vectors and a plurality of second modality feature vectors can be obtained, the target student model is obtained by training a student model using a loss function including an output layer alignment loss term and an intermediate layer alignment loss term, the output layer alignment loss term is used to align a teacher feature vector and a student feature vector obtained by processing a sample image block using a teacher model and a student model, and the intermediate layer alignment loss term is used to align an intermediate teacher feature vector and an intermediate student feature vector obtained by processing a sample image block using a teacher model and a student model, by using the above multi-level distillation strategy, the semantic gap between a complex teacher model and a lightweight student model can be bridged, so that the target student model can still inherit the robust understanding ability of the teacher model to cross-modality geometric deformation under the condition of reducing the number of parameters. Finally, according to the plurality of first modality feature vectors and the plurality of second modality feature vectors, an image registration result indicating the image transformation information between the first modality image and the second modality image can be obtained, the whole process reduces the computational complexity and improves the image registration accuracy.

[0056] The student model can include a student first modality branch, a student second modality branch and an interactive spatial attention module. The student first modality branch can include a student first modality channel alignment unit and a student first modality adaptation unit, and the student second modality branch can include a student second modality channel alignment unit and a student second modality adaptation unit.

[0057] The sample image block can include a sample first modality image block and a sample second modality image block; the student feature vector can include a first student first modality feature vector of the sample first modality image block and a first student second modality feature vector of the sample second modality image block, and the intermediate student feature vector includes a first intermediate student first modality feature vector of the sample first modality image block and a first intermediate student second modality feature vector of the sample second modality image block.

[0058] The first student first modality feature vector is obtained by adjusting the feature channel number of the second student first modality feature vector obtained by processing the sample first modality image block through the other part of the student first modality branch and the interactive spatial attention module, using the student first modality adaptation unit. The first student second modality feature vector is obtained by adjusting the feature channel number of the second student second modality feature vector obtained by processing the sample second modality image block through the other part of the student second modality branch and the interactive spatial attention module, using the student second modality adaptation unit.

[0059] The first intermediate student first modality feature vector is obtained by adjusting the feature channel number of the second intermediate student first modality feature vector obtained by processing the sample first modality image block through the other part of the student first modality branch and the interactive spatial attention module, using the student first modality channel alignment unit. The first intermediate student second modality feature vector is obtained by adjusting the feature channel number of the second intermediate student second modality feature vector obtained by processing the sample second modality image block through the other part of the student second modality branch and the interactive spatial attention module, using the student second modality channel alignment unit.

[0060] According to the embodiments of the present application, for the first modality image, the first intermediate student first modality feature vector can be aligned with the channel number of the first intermediate teacher first modality feature vector by using the student first modality channel alignment unit, and the second student first modality feature vector obtained by the first modality image block can be aligned with the channel number of the first intermediate teacher first modality feature vector by using the student first modality adaptation unit. For the second modality image, the first intermediate student second modality feature vector can be aligned with the channel number of the first intermediate teacher second modality feature vector by using the student second modality channel alignment unit, and the second student second modality feature vector obtained by the second modality image block can be aligned with the channel number of the first intermediate teacher second modality feature vector by using the student second modality adaptation unit, so as to realize the consistency of the feature space dimension of the target intermediate layer and the output layer, and take the output features of the teacher model target intermediate layer and the output layer as the soft target (i.e., the result obtained by processing the first modality image or the second modality image by using the lightweight student model is as same as the result obtained by processing the first modality image or the second modality image by using the teacher model as possible), so that the student model can still inherit the robust understanding ability of the teacher model to the cross-modality geometric deformation in the case of reducing the parameter amount.

[0061] According to an embodiment of the present application, the sample image block comprises a sample first modality image block and a sample second modality image block; the teacher feature vector comprises a first teacher first modality feature vector of the sample first modality image block and a first teacher second modality feature vector of the sample second modality image block, the intermediate teacher feature vector comprises a first intermediate teacher first modality feature vector of the sample first modality image block and a first intermediate teacher second modality feature vector of the sample second modality image block; the teacher model comprises a teacher first modality branch, a teacher second modality branch and an interactive feature fusion module; the first teacher first modality feature vector is obtained by processing a second teacher first modality feature vector corresponding to the sample first modality image block by using the teacher first modality branch and the interactive feature fusion module, and the first teacher second modality feature vector is obtained by processing a second teacher second modality feature vector corresponding to the sample second modality image block by using the teacher second modality branch and the interactive feature fusion module; the first intermediate teacher first modality feature vector is obtained by processing a second intermediate teacher first modality feature vector output by the target intermediate layer by using the teacher first modality branch and the interactive feature fusion module, and the first intermediate teacher second modality feature vector is obtained by processing a second intermediate teacher second modality feature vector output by the target intermediate layer by using the teacher second modality branch and the interactive feature fusion module.

[0062] Figure 3 A structure diagram of the teacher model according to an embodiment of the present application is shown.

[0063] As shown in Figure 3 , the teacher model adopts a pseudo-twin architecture with cross-modality feature fusion capability, and realizes deep semantic association modeling of the sample first modality image and the sample second modality image through an interactive feature fusion module. The teacher model is fully trained on an image registration dataset, forming reliable feature expression capability. The encoder of the teacher model is based on a double-branch ResNet-34 backbone network, and an interactive feature fusion module is introduced. L1, L2 and L3 in the encoder are network architectures and have the same architecture, and the network structure of each of them includes a first convolution unit, a second convolution unit and a maximum pooling unit connected in series. The decoder includes a plurality of L4 and L5, and the specific structure of each of them is the same, and the network structure is composed of three 3x3 convolutions and one 1x1 convolution.

[0064] According to an embodiment of the present application, the teacher model adopts a pseudo-twin architecture with cross-modality feature fusion capability, and through the interactive feature fusion module, deep semantic association modeling of the first modality and the second modality image can be realized.

[0065] According to an embodiment of the present application, the output layer alignment loss term is determined according to a similarity subterm and a constraint subterm, the similarity subterm is used to determine the similarity between the teacher feature vector and the student feature vector, and the constraint subterm is determined according to the model parameter of the student model; and the intermediate layer alignment loss term is used to determine the similarity between the intermediate teacher feature vector and the intermediate student feature vector.

[0066] The output layer alignment loss term is determined according to a similarity subterm and a constraint subterm, as shown in formula (1).

[0067] (1);

[0068] wherein, represents the output layer alignment loss term, represents the similarity subterm, represents the constraint subterm, represents the student feature vector, represents the teacher feature vector, represents the model parameter of the student model.

[0069] According to the above formula, the similarity subterm can be used to determine the similarity between the teacher feature vector and the student feature vector, and the constraint subterm is determined according to the model parameter of the student model.

[0070] The intermediate layer alignment loss term is used to determine the similarity between the intermediate teacher feature vector and the intermediate student feature vector, as shown in formula (2).

[0071] (2);

[0072] wherein, represents the intermediate layer alignment loss, is the intermediate student feature vector, is the intermediate teacher feature vector.

[0073] According to an embodiment of the present application, by training the student model by using the output layer alignment loss term and the intermediate layer alignment loss term, the semantic gap between the complex teacher model and the lightweight student model can be bridged, so that the student model can still inherit the robust understanding ability of the teacher model for cross-modal geometric deformation while reducing the number of parameters.

[0074] Figure 4 FIG. 1 shows a structure diagram of a target student model according to an embodiment of the present application.

[0075] Figure 5 FIG. 2 shows a structure diagram of knowledge distillation according to an embodiment of the present application.

[0076] As Figure 4As shown, the student model adopts a lightweight pseudo-twin architecture, the double-branch ResNet-34 backbone network in the teacher model is replaced by MobileNetV3-Small, and the interactive feature fusion module is simplified into an interactive spatial attention mechanism and a lightweight feature alignment module (including a student first modality channel alignment unit, a student first modality adaptation unit, a student second modality channel alignment unit and a student second modality adaptation unit). M1 and M2 in the encoder are network architectures and have the same architecture, and both contain a first convolution unit, a second convolution unit and a max pooling unit in series. The decoder includes multiple M4 and M3, and the specific structures of M4 and M3 are the same, and the network structure is composed of a 3x3 convolution and a 1x1 convolution.

[0077] As shown in Figure 5 A learnable 1x1 convolution adaptation layer is added at the end of the student model, i.e., a student first modality adaptation unit or a student second modality adaptation unit. The function of this layer is to dynamically map the channel number of the student model output feature to the same dimension as the teacher model output feature, realizing the basic matching of the feature space dimension.

[0078] A channel alignment unit (CAU) is inserted between the target intermediate layer (i.e., L1 in Figure 5 The CAU is composed of a 1x1 convolution layer followed by a batch normalization layer. Its function is to map the channel number of the student model output feature at the target intermediate layer to the same dimension as the intermediate feature provided by the teacher model at this layer through a 1x1 convolution.

[0079] The first modality image can be input into the student model and the teacher model. The output feature of the target intermediate layer of the student model is aligned with the output feature of the target intermediate layer of the teacher model using the student first modality channel alignment unit. According to the intermediate layer alignment loss term, the similarity between the intermediate teacher feature vector and the intermediate student feature vector is determined, and then the intermediate layer alignment loss value is determined. The output feature of the output layer of the student model is aligned with the output feature of the output layer of the teacher model using the student first modality adaptation unit, and the output layer alignment loss value is calculated according to the output layer alignment loss term. Finally, the student model can be trained according to the sum of the intermediate layer alignment loss value and the output layer alignment loss value corresponding to the first modality image to bridge the semantic gap between the complex teacher model and the lightweight student model. The training steps of the second modality image and the first modality image are the same, and here, they are not described again.

[0080] According to an embodiment of the present application, the image registration method further comprises: performing scale normalization processing on the first modality image and the second modality image to obtain a first modality down-sampled image and a second modality down-sampled image; and performing block division on the first modality down-sampled image and the second modality down-sampled image respectively to obtain a plurality of first modality image blocks and a plurality of second modality image blocks.

[0081] The first modality image and the second modality image under different viewing angles, time points or sensor conditions of the same target scene can be acquired. The bilinear interpolation method is used for scale normalization processing on the first modality image and the second modality image. The down-sampling rate can be 0.5, and the down-sampling rate can be dynamically adjusted in the range of 0.3 to 0.7 according to the memory capacity of the embedded platform. For example, the first modality image (3 channels) and the second modality image (single channel) with an input resolution of 1024x1024, the down-sampling rate is 0.5, and the image size can be compressed to 512x512, thereby reducing the subsequent calculation load.

[0082] Meanwhile, the first modality down-sampled image and the second modality down-sampled image can also be subjected to structured block division and padding, and the structured block division and padding of the first modality down-sampled image and the second modality down-sampled image are the same. For example, the first modality down-sampled image can be divided into non-overlapping grids, for example, the block size is 64x64 pixels. Zero padding technology is used for the edge region to ensure that all block sizes are uniform. It should be noted that the block size can be adjusted to 128x128 or 256x256 according to the hardware performance, for example, when the GPU resource is sufficient, large size block division is used to improve the feature extraction integrity.

[0083] According to an embodiment of the present application, by performing scale normalization processing on the first modality image and the second modality image, the subsequent calculation load can be reduced, and by performing block division and padding on the first modality down-sampled image and the second modality down-sampled image, a plurality of first modality image blocks and a plurality of second modality image blocks can be obtained, so that more edge structural features are retained under the premise of resource limitation.

[0084] According to an embodiment of the present application, the first position is a first modality center point latitude and longitude coordinate, and the second position is a second modality center point latitude and longitude coordinate; and the determining of the matched candidate first modality image block and the candidate second modality image block from the plurality of first modality image blocks and the plurality of second modality image blocks according to the first position of each first modality image block and the second position of each second modality image block comprises: determining the distance between each first modality image block and each second modality image block according to the first modality center point latitude and longitude coordinate of each first modality image block and the second modality center point latitude and longitude coordinate of each second modality image block, to obtain a plurality of spatial distances; and in the case that the spatial distance is less than or equal to a distance threshold, taking each of the first modality image block and the second modality image block corresponding to the spatial distance as the candidate first modality image block and the candidate second modality image block, to obtain the candidate matching pair.

[0085] The first position is a first modality center point latitude and longitude coordinate corresponding to the center point of the first modality image block, and the second position is a second modality center point latitude and longitude coordinate corresponding to the center point of the second modality image block. The first position can not only represent the latitude and longitude coordinate corresponding to the center point of the first modality image block, but also represent the latitude and longitude coordinate corresponding to other positions of the first modality image block, such as the upper left corner or the lower right corner, and the second position is the same.

[0086] The first modality center point latitude and longitude coordinate can be bound for each first modality image block, and the second modality center point latitude and longitude coordinate can be bound for each second modality image block. According to the first modality center point latitude and longitude coordinate of each first modality image block and the second modality center point latitude and longitude coordinate of each second modality image block, the spatial distance between each first modality image block and each second modality image block can be determined, to obtain a plurality of spatial distances, as shown in formula (3).

[0087] 2(3);

[0088] wherein, represents the first modality center point latitude and longitude coordinate corresponding to the first modality image block, represents the second modality center point latitude and longitude coordinate corresponding to the second modality image block, represents the spatial distance between the first modality image block and the second modality image block.

[0089] In the case that the spatial distance is less than or equal to a distance threshold, each of the first modality image block and the second modality image block corresponding to the spatial distance can be taken as the candidate first modality image block and the candidate second modality image block, to obtain the candidate matching pair. The distance threshold The calculation is shown as formula (4).

[0090] (4);

[0091] wherein, is a geo-location error of the first modality image, is a geo-location error of the second modality image, determined according to a nominal positioning accuracy index of the satellite platform. S is a physical side length of the image block, , N is an image size, and R is a resolution. k is a safety coefficient, and is taken as 1.5 by default, represents the first modality image, represents the second modality image.

[0092] According to the embodiment of the application, by using the spatial distance matching of the respective first modality center point longitude and latitude coordinates of the first modality image blocks and the respective second modality center point longitude and latitude coordinates of the second modality image blocks, the candidate matching pairs are obtained, a plurality of invalid matching pairs can be filtered, and the efficiency and reliability of the registration process are improved.

[0093] According to the embodiment of the application, the image registration result indicating the image transformation information between the first modality image and the second modality image is obtained according to the plurality of first modality feature vectors and the plurality of second modality feature vectors, including: determining a cosine similarity matrix according to the plurality of first modality feature vectors and the plurality of second modality feature vectors; in the case that the maximum value feature vector determined in the ith row of the cosine similarity matrix is the maximum value feature vector of the jth column, determining the pixel point corresponding to the ith first modality feature vector and the pixel point corresponding to the jth second modality feature vector as an initial matching point pair; processing the initial matching point pair by using a geometric consistency algorithm to obtain the image registration result indicating the image transformation information between the first modality image and the second modality image.

[0094] The cosine similarity matrix can be determined according to the plurality of first modality feature vectors and the plurality of second modality feature vectors, as shown in formula (5).

[0095] (5);

[0096] wherein, X represents the cosine similarity matrix, represents the ith first modality feature vector, represents the jth second modality feature vector.

[0097] In the case that the maximum eigenvector determined in the i-th row of the cosine similarity matrix is also the maximum eigenvector of the j-th column, the pixel point corresponding to the i-th first modality eigenvector and the pixel point corresponding to the j-th second modality eigenvector are determined as an initial matching point pair, as shown in formula (6).

[0098] (6);

[0099] wherein, represents the initial matching point pair, represents the pixel point corresponding to the k-th column maximum eigenvector, represents the pixel point corresponding to the k-th row maximum eigenvector.

[0100] The initial matching point pair is processed by using a geometric consistency algorithm, and an image registration result indicating image transformation information between the first modality image and the second modality image can be obtained. This process synchronously fuses the latitude and longitude coordinates and performs secondary verification on multiple candidate matching block pairs, and eliminates the interference of abnormal matching. The image transformation information can include a pixel point set after registration of the first modality image and the second modality image and an affine transformation matrix.

[0101] According to the embodiments of the present application, in the case that the maximum eigenvector determined in the i-th row of the cosine similarity matrix is also the maximum eigenvector of the j-th column, the pixel point corresponding to the i-th first modality eigenvector and the pixel point corresponding to the j-th second modality eigenvector are determined as an initial matching point pair, bidirectional constraints can be performed in the row direction and the column direction, the high precision and reliability of the registration result are improved, and multiple initial matching point pairs can be checked again based on a geometric consistency algorithm, and the precision of the image registration result is improved.

[0102] According to an embodiment of the present application, the initial matching point pairs are processed by using a geometric consistency algorithm to obtain an image registration result indicating image transformation information between the first modality image and the second modality image, comprising: in the case of i>1, determining the i-th round of matching point pairs according to the (i-1)-th round of second modality projection pixel point set and the second modality pixel point set, the second modality pixel point set being obtained according to the initial matching point pairs, the 1st round of second modality projection pixel point set being obtained according to the 1st round of affine transformation matrix processed from the initial matching point pairs, the 1st round of affine transformation matrix being obtained according to the least square method processed from the 1st round of matching point pairs determined from the initial matching point pairs; processing the i-th round of matching point pairs by using the least square method to obtain the i-th round of affine transformation matrix; processing the initial matching point pairs according to the i-th round of affine transformation matrix to obtain the i-th round of second modality projection pixel point set; in the case that the pixel difference between the projection pixel point in the i-th round of second modality projection pixel point set and the pixel point in the second modality pixel point set satisfies a preset pixel threshold, determining the i-th round of affine transformation matrix as the image registration result.

[0103] The plurality of matching point pairs can be randomly determined from the initial matching point pairs and taken as the first round of matching point pairs. The first round of affine transformation matrices can be obtained by processing the first round of matching point pairs according to the least square method. The first round of second modality projection pixel point sets can be obtained by processing the initial matching point pairs according to the first round of affine transformation matrices. In the case of i>1, the i-th round of matching point pairs can be determined according to the (i-1)-th round of second modality projection pixel point sets and the second modality pixel point set, and the second modality pixel point set is obtained according to the initial matching point pairs and stores the initial matching pixel point pairs. For example, there are N initial matching point pairs in total, and M matching point pairs can be randomly determined. Each matching point pair can include a first modality pixel point coordinate in the first modality image and a second modality pixel point coordinate in the second modality image preliminarily registered with the first modality pixel point coordinate. The M matching point pairs can be determined as the first round of matching point pairs ([a1, b1]...[am, bm]), wherein a1 to am are first modality pixel points, and b1 to bm are second modality pixel points. The corresponding first round of affine transformation matrices can be obtained by processing the first round of matching point pairs based on the least square method. The first round of second modality projection pixel point sets [B1, B2...Bn] can be obtained by processing the N initial matching point pairs according to the first round of affine transformation matrices, and B1 to Bn are the first round of pixel point sets obtained by processing the initial matching point pairs using the first round of affine transformation matrices. The second modality pixel point set is [b1, b2...bn]. The pixel differences between B1 and b1, B2 and b2,..., and Bn and bn can be calculated respectively. The matching point pairs with pixel differences less than a preset pixel threshold can be determined as the second round of matching point pairs from the plurality of pixel differences. The i-th round of affine transformation matrices can be obtained by processing the i-th round of matching point pairs according to the least square method according to the following steps.

[0104] The i-th round of second modality projection pixel point sets can be obtained by processing the initial matching point pairs according to the i-th round of affine transformation matrices, and the iteration is continued until the preset number of pixel differences are all less than the preset pixel threshold or the iteration number meets the condition, and the corresponding i-th round of affine transformation matrices are determined as the image registration result.

[0105] According to the embodiments of the present application, through multiple rounds of iteration, the pixel difference between the projection pixel point in the i-th round of second modality projection pixel point set and the pixel point in the second modality pixel point set meets the preset pixel threshold, which can filter the wrong matching point pairs and solve the problem of spatial deviation caused by low image registration accuracy, thereby meeting the demand of high-time-efficiency remote sensing monitoring tasks.

[0106] According to an embodiment of the present invention, the first modal image is a visible light image, and the second modal image can be a SAR image or an infrared image. There is no limitation on this, and it can be set according to the specific situation.

[0107] Figure 6 A structural block diagram of an image registration apparatus according to an embodiment of the present invention is shown.

[0108] like Figure 6 As shown, the image registration device of this embodiment includes a candidate matching pair module 610, a feature vector module 620, and a registration module 630.

[0109] The candidate matching pair module 610 is used to determine matching candidate first modal image blocks and second modal image blocks from multiple first modal image blocks and multiple second modal image blocks according to the first positions of each of the multiple first modal image blocks of the first modal image and the second positions of each of the multiple second modal image blocks of the second modal image, so as to obtain multiple candidate matching pairs.

[0110] The feature vector module 620 is used to process multiple candidate matching pairs using the target student model to obtain multiple first-mode feature vectors and multiple second-mode feature vectors. The target student model is obtained by training the student model using a loss function that includes an output layer alignment loss term and an intermediate layer alignment loss term. The output layer alignment loss term is used to align the teacher feature vector and student feature vector obtained by processing sample image patches using the teacher model and student model. The intermediate layer alignment loss term is used to align the intermediate teacher feature vector and intermediate student feature vector output by the target intermediate layer obtained by processing sample image patches using the teacher model and student model.

[0111] The registration module 630 is used to obtain an image registration result that indicates the image transformation information between the first modality image and the second modality image based on multiple first modality feature vectors and multiple second modality feature vectors.

[0112] According to the embodiment of the present application, by dividing the first and second modal images into a plurality of image blocks, the large size characteristics of the images can be reduced, so that the large size images do not monopolize a large amount of memory and cause high computational delay. According to the first position of the first modal image block and the second position of the second modal image block, the matching candidate first modal image block and the candidate second modal image block can be determined from the plurality of first modal image blocks and the plurality of second modal image blocks, and a plurality of candidate matching pairs are obtained. Through rough matching, invalid matching can be further reduced, and the efficiency of subsequent processing can be improved. By processing the plurality of candidate matching pairs using the target student model, a plurality of first modal feature vectors and a plurality of second modal feature vectors can be obtained. The target student model is obtained by training the student model using a loss function including an output layer alignment loss term and an intermediate layer alignment loss term. The output layer alignment loss term is used to align the teacher feature vector and the student feature vector obtained by processing the sample image block using the teacher model and the student model. The intermediate layer alignment loss term is used to align the intermediate teacher feature vector and the intermediate student feature vector obtained by processing the sample image block using the teacher model and the student model. By using the above multi-level distillation strategy, the semantic gap between the complex teacher model and the lightweight student model can be bridged, so that the target student model can still inherit the robust understanding ability of the teacher model for cross-modal geometric deformation under the condition of reducing the number of parameters. Finally, according to the plurality of first modal feature vectors and the plurality of second modal feature vectors, the image registration result indicating the image transformation information between the first modal image and the second modal image can be obtained. The whole process reduces the computational complexity and improves the image registration accuracy.

[0113] The feature vector module 620 includes: the output layer alignment loss term is determined according to the similarity subterm and the constraint subterm, the similarity subterm is used to determine the similarity between the teacher feature vector and the student feature vector, and the constraint subterm is determined according to the model parameters of the student model. The intermediate layer alignment loss term is used to determine the similarity between the intermediate teacher feature vector and the intermediate student feature vector.

[0114] The student feature vector includes the first student first modal feature vector of the sample first modal image block and the first student second modal feature vector of the sample second modal image block, and the intermediate student feature vector includes the first intermediate student first modal feature vector of the sample first modal image block and the first intermediate student second modal feature vector of the sample second modal image block.

[0115] The student model includes a student first modal branch, a student second modal branch, and an interactive spatial attention module. The student first modal branch includes a student first modal channel alignment unit and a student first modal adaptation unit. The student second modal branch includes a student second modal channel alignment unit and a student second modal adaptation unit.

[0116] The first student first modality feature vector is obtained by adjusting the number of feature channels of the second student first modality feature vector obtained by processing the sample first modality image block via the other part of the student first modality branch and the interactive spatial attention module, and the first student second modality feature vector is obtained by adjusting the number of feature channels of the second student second modality feature vector obtained by processing the sample second modality image block via the other part of the student second modality branch and the interactive spatial attention module.

[0117] The first intermediate student first modality feature vector is obtained by adjusting the number of feature channels of the second intermediate student first modality feature vector obtained by processing the sample first modality image block via the other part of the student first modality branch and the interactive spatial attention module, and the first intermediate student second modality feature vector is obtained by adjusting the number of feature channels of the second intermediate student second modality feature vector obtained by processing the sample second modality image block via the other part of the student second modality branch and the interactive spatial attention module.

[0118] The feature vector module 620 includes: the sample image block includes the sample first modality image block and the sample second modality image block.

[0119] The teacher feature vector includes the first teacher first modality feature vector of the sample first modality image block and the first teacher second modality feature vector of the sample second modality image block, and the intermediate teacher feature vector includes the first intermediate teacher first modality feature vector of the sample first modality image block and the first intermediate teacher second modality feature vector of the sample second modality image block.

[0120] The teacher model includes a teacher first modality branch, a teacher second modality branch, and an interactive feature fusion module.

[0121] The first teacher first modality feature vector is obtained by processing the second teacher first modality feature vector corresponding to the sample first modality image block by using the teacher first modality branch and the interactive feature fusion module, and the first teacher second modality feature vector is obtained by processing the second teacher second modality feature vector corresponding to the sample second modality image block by using the teacher second modality branch and the interactive feature fusion module.

[0122] The first intermediate teacher first modality feature vector is obtained by processing the second intermediate teacher first modality feature vector of the target intermediate layer output by using the teacher first modality branch and the interactive feature fusion module, and the first intermediate teacher second modality feature vector is obtained by processing the second intermediate teacher second modality feature vector of the target intermediate layer output by using the teacher second modality branch and the interactive feature fusion module.

[0123] The image registration apparatus of this embodiment comprises a normalization module and a block module.

[0124] The normalization module is configured to perform scale normalization on the first modality image and the second modality image to obtain a first modality down-sampled image and a second modality down-sampled image.

[0125] The block module is configured to block the first modality down-sampled image and the second modality down-sampled image respectively to obtain a plurality of first modality image blocks and a plurality of second modality image blocks.

[0126] The registration module 630 comprises a cosine similarity matrix sub-module, an initial matching point pair determination sub-module, and a registration sub-module.

[0127] The cosine similarity matrix sub-module is configured to determine a cosine similarity matrix according to the plurality of first modality feature vectors and the plurality of second modality feature vectors.

[0128] The initial matching point pair determination sub-module is configured to, in a case where the maximum value feature vector in the i-th row of the cosine similarity matrix is the maximum value feature vector in the j-th column, determine a pixel point corresponding to the i-th first modality feature vector and a pixel point corresponding to the j-th second modality feature vector as an initial matching point pair.

[0129] The registration sub-module is configured to process the initial matching point pair by using a geometric consistency algorithm to obtain an image registration result indicating image transformation information between the first modality image and the second modality image.

[0130] The registration sub-module comprises an i-th round matching point pair determination unit, an i-th round affine transformation matrix obtaining unit, an i-th round second modality obtaining unit, and a registration result determination unit.

[0131] The i-th round matching point pair determination unit is configured to, in a case where i>1, determine an i-th round matching point pair according to an (i-1)-th round second modality projected pixel point set and a second modality pixel point set, the second modality pixel point set being obtained according to the initial matching point pair, the 1st round second modality projected pixel point set being obtained according to the 1st round affine transformation matrix processed on the initial matching point pair, and the 1st round affine transformation matrix being obtained according to a least square method processed on a 1st round matching point pair determined from the initial matching point pair.

[0132] The i-th round affine transformation matrix obtaining unit is configured to obtain an i-th round affine transformation matrix according to a least square method processed on the i-th round matching point pair.

[0133] The i-th round second modality obtaining unit is configured to obtain an i-th round second modality projected pixel point set according to the i-th round affine transformation matrix processed on the initial matching point pair.

[0134] The registration result determination unit is configured to determine the affine transformation matrix in the i-th round as the image registration result when the pixel difference between the projection pixel point in the i-th round of the second modality projection pixel point set and the pixel point in the second modality pixel point set satisfies the preset pixel threshold.

[0135] The first position is the longitude and latitude coordinates of the first modality center point, and the second position is the longitude and latitude coordinates of the second modality center point.

[0136] The candidate matching pair module 610 includes a spatial distance determination sub-module and a candidate matching pair obtaining sub-module.

[0137] The spatial distance determination sub-module is configured to determine the distance between each first modality image block and each second modality image block according to the longitude and latitude coordinates of the first modality center point of each first modality image block of the first modality image and the longitude and latitude coordinates of the second modality center point of each second modality image block of the second modality image, to obtain a plurality of spatial distances.

[0138] The candidate matching pair obtaining sub-module is configured to, when the spatial distance is less than or equal to the distance threshold, take each of the first modality image block and the second modality image block corresponding to the spatial distance as a candidate first modality image block and a candidate second modality image block, to obtain a candidate matching pair.

[0139] According to embodiments of the present application, any of the candidate matching pair module 610, the feature vector module 620 and the registration module 630 can be combined in one module, or any of the modules can be split into multiple modules. Alternatively, at least part of the function of one or more of the modules can be combined with at least part of the function of other modules, and implemented in one module. According to embodiments of the present application, at least one of the candidate matching pair module 610, the feature vector module 620 and the registration module 630 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on board, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging a circuit, etc. hardware or firmware, or in any one of software, hardware and firmware or in any appropriate combination of any of them. Alternatively, at least one of the candidate matching pair module 610, the feature vector module 620 and the registration module 630 can be at least partially implemented as a computer program module which, when executed, can perform the corresponding functions.

[0140] Figure 7 A block diagram of an electronic device suitable for implementing the gyroscope drift error compensation method according to an embodiment of the present application is shown.

[0141] like Figure 7 As shown, an electronic device 700 according to an embodiment of the present invention includes a processor 701, which can perform various appropriate actions and processes according to a program stored in ROM 702 (i.e., read-only memory) or a program loaded from storage portion 708 into RAM 703 (i.e., random access memory). The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0142] RAM 703 stores various programs and data required for the operation of electronic device 700. Processor 701, ROM 702, and RAM 703 are interconnected via bus 704. Processor 701 executes various operations of the method flow according to embodiments of the present invention by executing programs in ROM 702 and / or RAM 703. It should be noted that programs may also be stored in one or more memories other than ROM 702 and RAM 703. Processor 701 may also execute various operations of the method flow according to embodiments of the present invention by executing programs stored in one or more memories.

[0143] According to an embodiment of the present invention, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to a bus 704. The electronic device 700 may also include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 710 as needed so that computer programs read from it can be installed into the storage section 708 as needed.

[0144] The application further provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments, or can exist independently without being assembled into the device / apparatus / system. The computer readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the application.

[0145] According to the embodiments of the application, the computer readable storage medium can be a non-volatile computer readable storage medium, which can include, but is not limited to, a portable computer diskette, a hard disk, a random access memory (RAM 703), a read-only memory (ROM 702), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination thereof. In this application, a computer readable storage medium can be any tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. For example, in the embodiments of the application, a computer readable storage medium can include the ROM 702 and / or the RAM 703 described above, and / or one or more memory devices other than the ROM 702 and the RAM 703.

[0146] The embodiments of the application also include a computer program product, which includes a computer program containing program codes for executing the methods shown in the flowcharts. When the computer program product is run in a computer system, the program codes are used to make the computer system implement the methods provided by the embodiments of the application.

[0147] The above functions defined in the system / apparatus of the embodiments of the application are performed when the computer program is executed by the processor 701. According to the embodiments of the application, the system, apparatus, module, unit, etc. described above can be implemented by computer program modules.

[0148] In one embodiment, the computer program can rely on a tangible storage medium such as an optical storage device, a magnetic storage device, etc. In another embodiment, the computer program can also be transmitted, distributed, and downloaded in the form of a signal on a network medium, and be downloaded and installed through the communication part 709, and / or installed from the detachable medium 711. The program codes contained in the computer program can be transmitted by any appropriate network medium, including but not limited to wireless, wired, etc., or any appropriate combination thereof.

[0149] In such embodiments, the computer program can be downloaded and installed from the network via the communication section 709, and / or installed from the removable media 711. When the computer program is executed by the processor 701, the above-described functions defined in the system of the embodiments of the present application are executed. The system, device, apparatus, module, unit, and the like described above can be realized by the computer program modules according to the embodiments of the present application.

[0150] According to the embodiments of the present application, the program code for executing the computer program provided by the embodiments of the present application can be written in any combination of one or more programming languages, and specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming language, and / or assembly / machine language. The programming language includes, but is not limited to, such as Java, C++, python, "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case involving a remote computing device, the remote computing device can be connected to the user computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, connected to the Internet through an Internet service provider).

[0151] The flowcharts and block diagrams in the drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that shown in the figures. For example, two blocks noted in succession can actually be executed substantially concurrently, or they can sometimes be executed in reverse order, depending on the functionality involved. It should also be noted that each block in the flowcharts or block diagrams, and combinations of blocks in the flowcharts or block diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0152] Those skilled in the art can understand that the features described in various embodiments of the present application can be combined and / or integrated in various combinations and / or integrations, even if such combinations or integrations are not explicitly described in the present application. In particular, the features described in various embodiments of the present application can be combined and / or integrated in various combinations and / or integrations without departing from the spirit and teachings of the present application. All such combinations and / or integrations fall within the scope of the present application.

[0153] The embodiments of the application have been described. However, these embodiments are merely for illustration and are not intended to limit the scope of the application. Although each embodiment is described above separately, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Various alternatives and modifications to the embodiments described herein will be apparent to those skilled in the art in view of the foregoing without departing from the scope of the application.

Claims

1. An image registration method, characterized in that, The method includes: Based on the first positions of each of the plurality of first modal image blocks in the first modal image and the second positions of each of the plurality of second modal image blocks in the second modal image, matching candidate first modal image blocks and candidate second modal image blocks are determined from the plurality of first modal image blocks and the plurality of second modal image blocks to obtain a plurality of candidate matching pairs; The target student model is used to process the multiple candidate matching pairs to obtain multiple first modality feature vectors and multiple second modality feature vectors. The target student model is obtained by training a student model using a loss function that includes an output layer alignment loss term and an intermediate layer alignment loss term. The output layer alignment loss term is used to align the teacher feature vector and student feature vector obtained by processing sample image patches using the teacher model and the student model. The intermediate layer alignment loss term is used to align the intermediate teacher feature vector and intermediate student feature vector output by the target intermediate layer obtained by processing the sample image patches using the teacher model and the student model. Based on the plurality of first modality feature vectors and the plurality of second modality feature vectors, an image registration result indicating the image transformation information between the first modality image and the second modality image is obtained.

2. The method according to claim 1, characterized in that, The output layer alignment loss term is determined based on a similarity sub-term and a constraint sub-term. The similarity sub-term is used to determine the similarity between the teacher feature vector and the student feature vector, and the constraint sub-term is determined based on the model parameters of the student model. The intermediate layer alignment loss term is used to determine the similarity between the intermediate teacher feature vector and the intermediate student feature vector.

3. The method according to claim 1 or 2, characterized in that, The sample image block includes a first modality image block and a second modality image block; The student feature vector includes the first student first modality feature vector of the first modality image block of the sample and the first student second modality feature vector of the second modality image block of the sample; the intermediate student feature vector includes the first intermediate student first modality feature vector of the first modality image block of the sample and the first intermediate student second modality feature vector of the second modality image block of the sample. The student model includes a student first modality branch, a student second modality branch, and an interactive spatial attention module. The student first modality branch includes a student first modality channel alignment unit and a student first modality adaptation unit. The student second modality branch includes a student second modality channel alignment unit and a student second modality adaptation unit. The first student first modality feature vector is obtained by adjusting the number of feature channels of the second student first modality feature vector obtained by processing the sample first modality image block through the other parts of the student first modality branch and the interactive spatial attention module using the student first modality adaptation unit. The first student second modality feature vector is obtained by adjusting the number of feature channels of the second student second modality feature vector obtained by processing the sample second modality image block through the other parts of the student second modality branch and the interactive spatial attention module using the student second modality adaptation unit. The first intermediate student first modality feature vector is obtained by adjusting the number of feature channels of the second intermediate student first modality feature vector output by the target intermediate layer, which is obtained by processing the sample first modality image block through the other parts of the student first modality branch and the interactive spatial attention module. The first intermediate student second modality feature vector is obtained by adjusting the number of feature channels of the second intermediate student second modality feature vector output by the target intermediate layer, which is obtained by processing the sample second modality image block through the other parts of the student second modality branch and the interactive spatial attention module.

4. The method according to claim 1 or 2, characterized in that, The sample image block includes a first modality image block and a second modality image block; The teacher feature vector includes the first teacher first modality feature vector of the first modality image block of the sample and the first teacher second modality feature vector of the second modality image block of the sample; the intermediate teacher feature vector includes the first intermediate teacher first modality feature vector of the first modality image block of the sample and the first intermediate teacher second modality feature vector of the second modality image block of the sample. The teacher model includes a first modality branch, a second modality branch, and an interactive feature fusion module. The first teacher first modality feature vector is obtained by processing the second teacher first modality feature vector corresponding to the sample first modality image block using the teacher first modality branch and the interactive feature fusion module; the first teacher second modality feature vector is obtained by processing the second teacher second modality feature vector corresponding to the sample second modality image block using the teacher second modality branch and the interactive feature fusion module. The first intermediate teacher first modality feature vector is obtained by processing the second intermediate teacher first modality feature vector output by the target intermediate layer using the teacher first modality branch and the interactive feature fusion module. The first intermediate teacher second modality feature vector is obtained by processing the second intermediate teacher second modality feature vector output by the target intermediate layer using the teacher second modality branch and the interactive feature fusion module.

5. The method according to claim 1 or 2, characterized in that, The method further includes: The first modality image and the second modality image are scale-normalized to obtain the first modality downsampled image and the second modality downsampled image; The first modal downsampled image and the second modal downsampled image are divided into blocks to obtain multiple first modal image blocks and multiple second modal image blocks.

6. The method according to claim 1 or 2, characterized in that, The step of obtaining an image registration result indicating image transformation information between the first modality image and the second modality image based on the plurality of first modality feature vectors and the plurality of second modality feature vectors includes: A cosine similarity matrix is ​​determined based on multiple first modality feature vectors and multiple second modality feature vectors; If the maximum value eigenvector determined in the i-th row of the cosine similarity matrix is ​​also the maximum value eigenvector in the j-th column, the pixel corresponding to the i-th first modality eigenvector and the pixel corresponding to the j-th second modality eigenvector are determined as the initial matching point pair. The initial matching point pairs are processed using a geometric consistency algorithm to obtain an image registration result that indicates the image transformation information between the first modal image and the second modal image, where i and j are positive integers.

7. The method according to claim 6, characterized in that, The process of using a geometric consistency algorithm to process the initial matching point pairs yields an image registration result that indicates image transformation information between the first modality image and the second modality image, including: When i>1, the matching point pair for the i-th round is determined based on the second modal pixel set and the second modal projection pixel set for the (i-1)th round, wherein the second modal pixel set is obtained based on the initial matching point pair; The i-th round of matching point pairs is processed using the least squares method to obtain the i-th round of affine transformation matrix; The initial matching point pair is processed according to the i-th round affine transformation matrix to obtain the i-th round second modal projection pixel set; If the pixel difference between the projected pixels in the second modal projection pixel set of the i-th round and the pixels in the second modal pixel set satisfies a preset pixel threshold, the affine transformation matrix of the i-th round is determined as the image registration result. The first round of second modal projection pixel set is obtained by processing the initial matching point pair according to the first round of affine transformation matrix. The first round of affine transformation matrix is ​​obtained by processing the first round of matching point pair determined from the initial matching point pair according to the least squares method.

8. The method according to claim 1 or 2, characterized in that, The first position is the latitude and longitude coordinates of the center point of the first mode, and the second position is the latitude and longitude coordinates of the center point of the second mode; The step of determining matching candidate first modality image blocks and candidate second modality image blocks from the plurality of first modality image blocks and the plurality of second modality image blocks based on the first positions of the plurality of first modality image blocks of the first modality image and the second positions of the plurality of second modality image blocks of the second modality image to obtain a plurality of candidate matching pairs includes: Based on the latitude and longitude coordinates of the center points of the first modal of each of the multiple first modal image blocks of the first modal image and the latitude and longitude coordinates of the center points of the second modal of each of the multiple second modal image blocks of the second modal image, the distance between any first modal image block among the multiple first modal image blocks and any second modal image block among the multiple second modal image blocks is determined, thereby obtaining multiple spatial distances; When the spatial distance is less than or equal to the distance threshold, the first modal image block and the second modal image block corresponding to the spatial distance are respectively regarded as candidate first modal image block and candidate second modal image block to obtain the candidate matching pair.

9. The method according to claim 1, characterized in that, The first modal image is a visible light image.

10. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Image processing model training method and device, electronic equipment and storage medium

    CN117036181A

  • Image registration method, device and equipment applied to multiple modes

    CN117557604A