Universal cross-modal image matching method and system for any source

Through the HOMO algorithm, the MOM and GPolar descriptors are used to solve the universality problem of cross-modal image matching, and high-precision, robustness and stability matching between different modal images is achieved. It is suitable for remote sensing, medicine and computer vision and other fields.

CN120375019APending Publication Date: 2025-07-25BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510548167.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing cross-modal image matching methods lack versatility and are difficult to effectively handle rotation, scale and texture differences under different sensors and imaging conditions. The deep learning methods rely on training data and are insufficiently adaptable.

Method used

The HOMO algorithm is used to achieve stable matching of cross-modal images through image preprocessing, feature extraction, key point detection and matching based on the subject direction feature map (MOM), generalized polar coordinate descriptor (GPolar) and multi-scale strategy (MsS).

Benefits of technology

It achieves high precision, robustness and stability matching between multiple modal images, can adapt to rotation, scale and viewing angle differences, and is widely used in fields such as remote sensing, medicine and computer vision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375019A_ABST
    Figure CN120375019A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal image matching method and system universal for any sources, and the method comprises the steps: carrying out the normalization and denoising processing of an input image, converting the input image into a single-channel gray image, and guaranteeing the stability and consistency of data; thirdly, generating a main body direction feature MOM of the image by using a multi-scale Gaussian pyramid, and extracting an invariant image feature; a key point in a cross-modal image is accurately positioned by calculating an MOM differential pyramid and combining a PC-ShiTomasi feature point detector. Further, a GPlar descriptor is used to extract feature vectors of rotation invariance and scale robustness of each key point. Through a multi-scale matching strategy and an RANSAC algorithm, accurate key point matching and estimation of a spatial transformation model are finally realized. The method has strong cross-modal adaptive capacity, can effectively deal with the problems of modal difference, scale change and rotation between images, and has wide application prospects in registration of remote sensing images, medical images and multi-source images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - field of image processing and machine vision, and particularly relates to a cross - modal image matching method and system that are universal for any source. Background Art

[0002] Image invariant feature extraction and matching is a basic technology in the fields of image processing and machine vision, and is crucial for direct application tasks such as image registration, homography estimation, camera localization, structure from motion, etc., as well as indirect application tasks such as image fusion and joint analysis of multi - source data. Represented by SIFT (Scale - Invariant Feature Transform), the algorithm extracts local features from the neighborhood of specific positions in the image, and these features remain invariant when the scene changes (such as moving, rotating, scaling, etc.), so as to match and correspond the features between different - source images. In the early stage, the algorithm mainly processed single - modal images, mainly optical images, and also considered some multi - source problems such as illumination, color or phase differences. In recent years, algorithms based on deep neural networks have attracted attention, especially the use of Transformer (attention mechanism) has shown advantages in multi - view scenarios, which globally perceives potential 3D patterns and handles perspective differences. However, the complexity of cross - modal problems is higher because it is no longer limited to the imaging environment but extends to different sensors.

[0003] Cross-modal image matching, that is, images captured by different sensors, faces greater challenges. With the rapid development of sensor technology, a variety of image modalities have emerged. Integrating general cross-modal images by automatically and effectively establishing spatial relationships has become a key task in image feature matching. The demand for cross-modal matching initially focused on medical imaging, such as matching computed tomography (CT) and magnetic resonance imaging (MRI) scans of patients. Subsequently, cross-modal image research in various fields has gradually received attention. The matching of cross-modal images such as visible light and infrared, or optical and synthetic aperture radar (SAR), lidar (LiDAR) images, etc., is one of the most common requirements. When the modality changes, the extracted features should be able to withstand the complex distortions caused by different sensor imaging mechanisms while maintaining uniqueness to ensure successful matching in the feature set. Achieving a balance between invariance and uniqueness is the key to the effectiveness of cross-modal algorithms. In recent years, cross-modal image feature matching using deep learning techniques has attracted attention. A state-of-the-art matching network model XoFTR (Cross-modal Local Feature Transformer) based on the improvement of the LoFTR (Local Feature Transformer) model has achieved the matching of multi-view visible light and thermal infrared images. However, these methods have limitations in the input of specific image modalities, which raises a question: whether there is a general feature or model that can handle a wide range of cross-modal data without discrimination. The ability to accurately match image key points under different conditions (such as lighting, viewing angle, or sensor type) is the core requirement of the algorithm. Currently, the research on general cross-modal feature matching is still immature, and there is still a lack of a basic general paradigm.

[0004] The proposal of general cross-modal image matching faces the most complex difficulties. It is still a challenge to implement a method that is applicable and effective for various cross-modal images and can simultaneously and effectively handle spatial distortions such as rotation and scale differences. In the field of deep learning, the establishment of general cross-modal image matching models is restricted by the lack of cross-modal datasets and the difficulty of creating them; for non-data-driven methods, there are few studies that can clearly determine whether there is a basic feature that is highly invariant in any modal image data. If there is such a technology that is not restricted by the modality type, then any images containing the same scene can automatically establish spatial relationships, providing strong support for a wide range of multi-source, multi-modal image analysis task scenarios in various application fields. Summary of the Invention

[0005] To solve the technical problems existing in the background art, the present invention aims to provide a cross-modal image matching method and system that is universal for any source, capable of effectively processing non-specific sources and adapting to a general image feature extraction and matching framework for rotation, scale, and texture differences. It does not rely on black-box network model learning but rather a fixed-computation feature that is clearly defined. An arbitrary-source universal cross-modal image matching algorithm based on the Homomorphism of Organized Major Orientation (HOMO) provides a general solution for image matching in various fields such as remote sensing, medicine, and computer vision to handle various modal differences and spatial distortions. Based on the local direction information of the image, HOMO designs a global non-data-driven algorithm process that takes a Major Orientation Map (MOM) as the core, combines a novel Generalized-Polar (GPolar) descriptor, and a Multi-scale Strategy (MsS) to effectively address key issues such as modal differences, rotation, scale differences, geometric distortions, and image noise. This algorithm demonstrates excellent cross-modal scene robustness, stability, generalization, and the ability to handle complex image differences. HOMO performs well in image matching in multiple fields such as remote sensing, medicine, and computer vision, as well as under different imaging conditions, thus providing stable and reliable technical support for a wide range of multi-source image application scenarios.

[0006] To solve the technical problems, the technical solution of the present invention is:

[0007] An arbitrary-source universal cross-modal image matching method, the method comprising:

[0008] S1: Based on the preprocessed single-channel grayscale image, construct a Gaussian scale pyramid of the input image, generate multiple scale images through downsampling and Gaussian blur, calculate the major orientation feature MOM of each layer of the image in the scale space, extract the core invariant features of the image, and convert them into a pixel-level feature domain to form the MOM pyramid of the image;

[0009] S2: Based on the MOM pyramid of the image, calculate the difference of the multi-scale MOM to generate a MOM difference pyramid (DoM), normalize and weight each layer of the DoM, calculate the MOM change weight for enhancing the localization of feature points; use the PC-ShiTomasi feature point detector, combine the MOM change weight to enhance the key point response of the PC-ShiTomasi, and accurately locate stable and significant key points through local non-maximum suppression LNMS;

[0010] S3: Take the coordinates and response values of the key points detected in step S2 as inputs. At each scale level, extract GPolar descriptors for the detected key points. This descriptor includes a basic structure and a deep structure, ensuring rotational invariance and scale robustness. Through the Oriented Consistency Discriminant (OCD), output the GPolar descriptor for each key point, including multi-scale and multi-level feature vectors.

[0011] S4: Adopt a multi-scale strategy (MsS) to iteratively strengthen key point matching layer by layer, improving the matching accuracy and stability. Use the Nearest Neighbor (NN) algorithm for feature matching based on the Euclidean distance, and remove the incorrect matching points through the RANSAC algorithm to obtain the matched key point pairs, i.e., the corresponding point pairs in the cross-modal images.

[0012] S5: According to the spatial coordinates of the matched key point pairs, select a spatial transformation model. Estimate the parameters of the transformation model through the spatial coordinate relationship between the matching points, calculate the spatial matching relationship between the images, and finally output the parameters of the spatial transformation model for the cross-modal images for image registration.

[0013] A cross-modal image matching system applicable to any source, which is applied to any of the above methods. The system includes:

[0014] Image preprocessing module: Perform normalization and summation processing on the multi-source cross-modal images to be matched, convert them into single-channel grayscale images, and use Gaussian filtering for denoising to obtain the preprocessed single-channel grayscale images.

[0015] Gaussian scale pyramid construction module: Based on the preprocessed single-channel grayscale images, construct a Gaussian scale pyramid of the input images. Generate multiple scale images through downsampling and Gaussian blur, calculate the Main Orientation Moments (MOM) of the main direction features of each layer of images in the scale space, extract the core invariant features of the images, and convert them into a pixel-level feature domain to form the MOM pyramid of the images.

[0016] MOM difference pyramid generation and feature point detection module: Based on the MOM pyramid of the images, calculate the differences of the multi-scale MOM to generate a MOM difference pyramid (DoM). Normalize and weight each layer of the DoM, calculate the MOM change weight for enhancing the localization of feature points. Adopt a PC-ShiTomasi feature point detector, combine the MOM change weight to enhance the key point response of the PC-ShiTomasi, and accurately locate the stable and significant key points through Local Non-Maximum Suppression (LNMS).

[0017] Key point descriptor extraction module: Taking the coordinates and response values of the above-detected key points as input, at each scale level, GPolar descriptors are extracted for the detected key points. This descriptor includes a basic structure and a deep structure, ensuring rotational invariance and scale robustness. Through the orientation consistency discriminant OCD, the GPolar descriptor of each key point is output, including a multi-scale and multi-level feature vector;

[0018] Key point matching module: Adopting the multi-scale strategy MsS, iteratively strengthening key point matching layer by layer to improve the matching accuracy and stability. Using the nearest neighbor NN algorithm for feature matching based on the Euclidean distance, and removing incorrect matching points through the RANSAC algorithm to obtain the matched key point pairs, that is, the corresponding point pairs in the cross-modal images;

[0019] Spatial transformation model estimation module: According to the spatial coordinates of the matched key point pairs, a spatial transformation model is selected. By the spatial coordinate relationship between the matching points, the parameters of the transformation model are estimated, and the spatial matching relationship between the images is calculated. Finally, the spatial transformation model parameters of the cross-modal images are output for image registration.

[0020] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements an arbitrary-source general cross-modal image matching method described in any one of the above.

[0021] A computer-readable storage medium stores a computer program. When the program is executed by a processor, it implements an arbitrary-source general cross-modal image matching method described in any one of the above.

[0022] Compared with the prior art, the advantages of the present invention are as follows:

[0023] The proposed HOMO algorithm for the group structure main direction homomorphy feature in the present invention shows remarkable effects and advantages in a wide range of cross-modal image matching tasks in multiple application fields including remote sensing, medicine, computer vision, etc. Compared with other traditional methods and the latest deep learning algorithms, HOMO achieves higher accuracy, robustness, stability, and generalization, effectively coping with common challenges such as modal differences, rotation, scale differences, spatial distortion, and image noise. The following is a detailed elaboration of the specific effects and experimental data:

[0024] 1. Advantage of invariant feature extraction

[0025] Inspired by the human visual system, a new invariant feature extraction method with the direction feature as the core is proposed, namely the main direction feature map (MOM), which has high invariance and interpretability in the low-level feature domain of cross-modal images, providing a solid guidance and a new perspective for the method design of cross-modal image invariant feature extraction.

[0026] 2. General Image Matching Framework

[0027] Based on MOM, an end-to-end full-chain general cross-modal image matching algorithm framework HOMO was developed. By using a series of sub-module algorithms such as DoM, GPolar, and MsS, it effectively addressed key issues such as modal differences, arbitrary rotations, significant scale differences, and local spatial distortions in cross-modal images. Comprehensive experiments show that the proposed HOMO has extremely high effectiveness, stability, and generalization on a wide range of cross-modal data, and can accommodate dozens of image modalities.

[0028] 3. Cross-modal Feature Invariance Performance

[0029] The feature extraction of HOMO has excellent cross-modal general invariance. Compared with other existing representative traditional methods such as SIFT, PIIFD, PSO-SIFT, HOPC, CFOG, RIFT, LNIFT, CoFSM, SRIF, etc., HOMO significantly improves the matching performance and can obtain the most effective matching points in all cross-modal categories, providing reliable reference information for the estimation of spatial model parameters in image matching. When the multi-scale strategy is removed, it is still superior to other methods in most cases and can outperform other multi-scale methods such as MS-PIIFD, MS-HLMO, etc. As a non-data-driven hand-designed framework, HOMO also shows significant advantages compared with the current latest deep learning-based neural network matching methods such as SuperPoint, SuperGlue, LightGlue, LoFTR, ELoFTR, ASpanFormer, TopicFM+, XoFTR, etc. Data-driven models highly rely on the input training data and are restricted by cross-modal datasets, all showing extremely unbalanced modal adaptability. In contrast, HOMO has excellent generalization ability among all modalities, with stable matching effects, strong versatility, high degrees of freedom, and extremely high application potential.

[0030] 4. Algorithm Spatial Distortion Robustness Performance

[0031] HOMO has strong rotational invariance and can be effectively matched between cross-modal images with any rotational angle difference within the complete 0 - 360° (±180°) interval, being minimally affected by image rotation. In contrast, existing methods, especially neural network algorithms, exhibit poor rotational invariance. When the rotation exceeds ±45°, the performance drops sharply, leading to matching failure. HOMO can handle large scale differences and can be effectively matched between cross-modal images with a scale (resolution) ratio of up to 4 times, showing good scale robustness. HOMO can handle significant viewing angle differences and can be effectively matched between cross-modal images with a viewing angle difference of up to ±45°, and maintain a high level of performance within a 30° viewing angle difference, showing good viewing angle robustness.

[0032] Considering the overall effect, HOMO, as a traditional framework, shows better generalization and stability when differences occur, which is mainly attributed to its non-data-driven nature and rigid descriptors. In contrast, network methods only show applicability under specific conditions, and their performance depends on the input training samples, while HOMO is not restricted. Overall, HOMO demonstrates the best comprehensive performance under the vast majority of spatial distortion conditions.

[0033] 5. Matching performance on real image data

[0034] HOMO shows the best comprehensive performance on current publicly available datasets such as the MRSI dataset, SRIF dataset, MIMD dataset, etc. The modalities it can effectively adapt to include visible light RGB, panchromatic, thermal infrared, near infrared, shortwave infrared, multispectral, hyperspectral, SAR, LiDAR-DEM, LiDAR-DSM, digital maps, day-night, clear-fog, cross-seasonal, real scene-painting, scene-label, medical images such as CT, Single Photon Emission Computed Tomography (SPECT), MRI, fluorescence, microscopic tissue staining, etc., dozens of cross-modal image data, showing strong cross-modal generalization ability, which highlights the potential and prospects of HOMO in multi-source imaging applications. Comprehensive experiments fully demonstrate that the proposed HOMO exhibits superior performance in cross-modal invariance, algorithm robustness, stability, and generalization.

[0035] In summary, the high precision, strong cross-modal generalization ability, and good complex spatial distortion adaptation ability of HOMO give it significant technical advantages in cross-modal image matching of any source, providing solid technical support for subsequent related tasks of multi-source image joint applications. Description of the Drawings

[0036] Figure 1, The main technical roadmap of a cross-modal image matching method that is universal for any source in the present invention;

[0037] Figure 2 , The basic structure diagram of the GPolar descriptor;

[0038] Figure 3 , The deep structure diagram of the GPolar descriptor;

[0039] Figure 4 , The schematic diagram of the MsS process. Detailed implementation manners

[0040] The following describes the specific implementation manners of the present invention in conjunction with embodiments:

[0041] It should be noted that the structures, ratios, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those skilled in this technology to understand and read, and are not used to limit the limited conditions under which the present invention can be implemented. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.

[0042] At the same time, the terms such as "upper", "lower", "left", "right", "middle", and "one" cited in this specification are only for the convenience of clear narration, and are not used to limit the scope under which the present invention can be implemented. The change or adjustment of their relative relationships, without substantial change in the technical content, should also be regarded as the scope under which the present invention can be implemented.

[0043] Embodiment 1:

[0044] The present invention proposes a cross-modal image matching algorithm that is universal for any source based on the homomorphic feature HOMO of the fabric main body direction, and provides a general solution for image matching in the fields of remote sensing, medicine, computer vision, etc. to cope with various modal differences and spatial distortions. The overall process is as Figure 1 shown. The core technologies of this technical solution mainly include four parts: basic feature map, key point detection, descriptor extraction, and key point matching. The algorithm process is mainly divided into six steps:

[0045] Step 1: Data input and preprocessing

[0046] Sum and normalize all channel (band) dimensions of the input cross-modal image data uniformly, and generate single-channel grayscale images as the input samples of the algorithm, and perform simple image denoising using Gaussian filtering.

[0047] Step 2: Calculation of the HOMO-basic feature map

[0048] Build the scale space of the input image, that is, generate the Gaussian scale pyramid of the image through downsampling and Gaussian blurring. Calculate the main direction feature MOM of each layer of the scale image in the scale space, extract the core invariant features, transform the input image into the pixel-level feature domain, and generate the MOM pyramid of the image.

[0049] Step 3: HOMO - Key point detection

[0050] Perform interlayer differences on the multi-scale MOM to generate the Difference of MOM (DoM) pyramid. Sum and normalize all layers of the DoM, and calculate the MOM change weight. This weight can be applied to any pixel-level feature detector to enhance key point localization. Use a Phase Congruency (PC)-based ShiTomasi feature point detector (PC-ShiTomasi) to calculate the key point response, multiply it by the MOM change weight to enhance the key point response, and finally detect and locate the key points in the input image through Local Non-Maximum Suppression (LNMS).

[0051] Step 4: HOMO - Descriptor extraction

[0052] Extract the Generalized Polar (GPolar) descriptor for the detected key points in each layer of the scale in the MOM pyramid. This descriptor assigns a descriptor vector to each key point through a basic structure and a deep structure, endowing rotation invariance and scale robustness, and uses an Orientation Consistency Discrimination (OCD) module to resist the descriptor flipping problem.

[0053] Step 5: HOMO - Key point matching

[0054] Use the key point matcher with the multi-scale strategy MsS to enable the key points to be iteratively enhanced layer by layer for matching, and also provide multi-scale accommodation ability for HOMO. Among them, use the Nearest Neighbor (NN) algorithm as the feature matcher, perform one-by-one matching between feature vectors through Euclidean distance measurement, and use the Random Sample Consensus (RANSAC) algorithm to eliminate mismatches and remove incorrect matching points.

[0055] Step 6: Spatial transformation model estimation

[0056] After obtaining a set of matching points between cross-modal images, a spatial transformation (mapping) model is selected, such as similarity transformation, affine transformation (first-order polynomial), projective transformation (projective transformation, homographic transformation), polynomial transformation, etc. The parameters of the transformation model are estimated through the spatial coordinate correspondence between the matching points to obtain the spatial matching relationship of the cross-modal images.

[0057] Embodiment 2:

[0058] This Embodiment 2 is a specific expansion of Embodiment 1 and specifically includes the following steps:

[0059] Step 1: Data input and preprocessing

[0060] Sum and normalize all channel (band) dimensions of the input cross-modal image data uniformly, and generate single-channel grayscale images as the input samples of the algorithm. Gaussian filtering is used for simple image denoising.

[0061] Step 2: HOMO - basic feature map calculation

[0062] Build the scale space of the input image, that is, generate the Gaussian scale pyramid of the image through downsampling and Gaussian blur. Calculate the MOM of each layer of scale image in the scale space, extract the core invariant features, and transform the input image into the pixel-level feature domain to generate the MOM pyramid of the image. The extraction process of MOM is mainly divided into three steps: alienated LogGabor weighting, mean square weighting, and odd-even feature coupling.

[0063] 1) Alienated LogGabor weighting

[0064] According to the principle of LogGabor (LG), the 2D - LG kernel function is:

[0065]

[0066] where (ρ, θ) are polar coordinates; s is the scale index and o is the orientation index; ρ s and θ o are the central frequency and the preferred direction at scale s and orientation o respectively; σ p and σ θ are the bandwidths in ρ and θ respectively. The LG in Equation (1) is defined in polar coordinates, and its Cartesian coordinate representation is:

[0067]

[0068] where for each pixel (x, y), the real part formed by even - LG and the imaginary part formed by odd - LG are composed. i represents the imaginary unit.

[0069] First, the single-band input image I(x, y) is filtered by a predefined LG with scale s and orientation o:

[0070]

[0071] where the real part can be physically interpreted as the "edge" response at a specific scale s and orientation o, and the imaginary part can be interpreted as the "contour" response, with * denoting convolution. Then, the LG response is projected onto the orthogonal x and y directions according to the orientation o. First, the odd-LG response is directionally weighted:

[0072]

[0073] where the responses at different scales s for each orientation o are added together.

[0074] However, due to the non-directional nature of the even-LG response, the directly weighted even-LG response does not have rotational invariance. To solve this problem, a modified even-LG weighting method is defined. For ease of expression, a sign function is defined:

[0075]

[0076] The modified even-LG component is defined as follows:

[0077]

[0078] where ° denotes the Hadamard product. In the above definition, the sign of the even-LG response is flipped according to the sign of the odd-LG response at the corresponding pixel position. For the same position in the image, due to the property of the even function, the even-LG responses with an angular difference exceeding 180° must have the same sign, while the odd-LG response at that position must have the opposite sign in the corresponding direction (unless the value is 0). Therefore, using the sign of the odd-LG response as a reference ensures that the opposite directions of the even-LG response are given opposite signs accordingly.

[0079] This method uses the odd-LG response to guide the weighting of the even-LG response, thereby generating a modified even-LG response component that has the property of having opposite signs in the opposite direction projection components. The modified even-LG response is further used for directional weighting:

[0080]

[0081] Then, the magnitude and orientation maps of the odd imaginary and even real parts of the LG response are obtained:

[0082]

[0083] Two "gradient" features can be considered to be obtained, and their angular values are in the range of [0, π) or correspond to the local structure direction of the image. Both the odd part and the even part exhibit rotational invariance, and multi-scale weighting enhances stability and scale robustness.

[0084] 2) Mean square weighting

[0085] The feature maps obtained in the previous stage, whether odd or even, can be used for descriptor statistics to achieve cross-modal matching. However, the stability and generalization ability of the features are still relatively poor, and their roles have not been fully exerted. To further enhance the features, a gradient weighting method, namely Averaging Squared Gradient (ASG), is adopted.

[0086] ASG is a local gradient weighting method. The elementary gradients of the image along the x and y directions, namely G x and G y , are calculated as follows:

[0087]

[0088] where I(x, y) represents the input single-layer grayscale image. The magnitude and direction of its gradient, G ρ and are expressed as:

[0089]

[0090] Assume that a new gradient is calculated by squaring the original magnitude and doubling the angle, which is expressed as:

[0091]

[0092] In this transformation, all vectors that originally had opposite angles (differing by π) in the range [-π, π] are now aligned to the same angle (differing by 2π). Then, the new gradients G s,x and G s,y along the x and y directions are expressed as:

[0093]

[0094] Substituting Equation (12) into Equation (13), we get:

[0095]

[0096] Then, by substituting Equation (11) into Equation (14), the relationship between the new gradient and the original gradient in Cartesian coordinates is obtained:

[0097]

[0098] The key purpose is to weight the local gradient information. In ASG, the local weighted squared gradients G σ,s,x and G σ,s,y along the x and y directions are expressed as:

[0099]

[0100] where W σ is a Gaussian window with variance σ, denotes the neighborhood weighting of applying W σ to each position in the input matrix, which is also understood as Gaussian filtering. Therefore, the direction of this gradient is:

[0101]

[0102] where ∠(X,Y) is defined as:

[0103]

[0104] such that is in the range of (-π,π). According to Equation (12), the weighted original gradient is now calculated as:

[0105]

[0106] Based on ASG, the present invention defines a multi-scale weighting scheme, in which G σ,ρ is jointly weighted at multiple scales. However, since direct summation is not feasible, the components G σ,s,x and G σ,s,y for each scale (Gaussian window variance) σ are weighted:

[0107]

[0108] where λ σ is the weight of scale σ. Then, according to Equations (17), (19), and (20), the Averaging Squared Weighting (ASW) feature map is obtained:

[0109]

[0110] 3) Odd-even feature coupling

[0111] The ASW of the previous stage, as shown in Equation (21), is in functional form:

[0112]

[0113] where \(x\) and \(Y\) represent two two-dimensional matrices of equal size as inputs. Then, ASW is applied to the odd-LG and even-LG responses obtained in the first stage, substituting equations (4) and (7) into equation (22), thus obtaining a method for calculating the average square LogGabor features:

[0114]

[0115] In this way, the odd and even features respectively provide a stable representation of the directions of significant local edges and contour structures in the image. Finally, a coupling operation is applied to these two feature maps to obtain the MOM:

[0116]

[0117] where the dominant feature (odd or even) at each position is determined and its direction is assigned as the local direction of the image. The calculation of MOM can use the sign function sgn * Expressed as:

[0118]

[0119] At this stage, the final feature map \(M\) MOM is obtained, which is the core of the HOMO algorithm and is a low-level pixel-level invariant feature.

[0120] Step 3: HOMO - Key Point Detection

[0121] Extract the MOM in the Gaussian scale space, providing a MOM pyramid that reflects the local main directions at different scales. By calculating the differences (absolute values) of each layer of the MOM pyramid and summing them, a change response \(DoM\) of the MOM at different scales is obtained. Sum and normalize all layers of the \(DoM\) to calculate the MOM change weight. The reciprocal of the MOM change weight is used as a weight, and then multiplied by the key point response to identify the feature positions with prominent and stable structures, enhancing the key point response. This weight can be applied to any pixel-level feature detector to enhance key point localization. In this process, a PC-ShiTomasi feature point detection algorithm is used to calculate the key point response, which is multiplied by the MOM change weight, and finally the key points are detected and located in the input image through LNMS.

[0122] Step 4: HOMO: - Descriptor Extraction

[0123] Extract the GPolar descriptor for the detected key points at each scale level in the MOM pyramid. This descriptor assigns a descriptor vector to each key point through a basic structure and a deep structure, endowing rotation invariance and scale robustness, and uses an Orientation Consistency Discrimination (OCD) module to resist the descriptor flipping problem. The extraction process of GPolar is mainly divided into three modules: basic structure extraction, deep structure extraction, and cross-modal rotation adaptation.

[0124] 1) Basic structure: The basic structure of GPolar has a flexible circular polar coordinate design, as Figure 2 shown. It consists of a central circular unit A 0 and an outer ring evenly divided into fan-shaped units, denoted as where i and j are the unit indices along the radial and angular directions respectively. N A is the number of fan-shaped units in each outer ring. It should be noted that this structure is not fixed, and both the radial and angular subdivisions can be extended. The radii of the central and outer regions are denoted as R0, R1, and R2 respectively. The area of each unit is set to be equal, which fixes the relationship between R0, R1, R2, and N A :

[0125]

[0126] so that the number of pixels for counting is roughly the same, mainly for the fair weights of the units.

[0127] For each detected key point, GPolar extracts local features in the neighborhood with a radius of R2 centered on it, thus generating a descriptor vector. The MOM value within each unit, regarded as or the orientation angle within [0, π), is uniformly quantized into φ k , k = 1, 2,..., N O . A histogram vector with N directions is constructed for each unit O . Then, these vectors are weighted with a Gaussian kernel along the θ direction. Therefore, the basic structure of GPolar generates a descriptor vector of length (1 + 2N ) × N A for each key point. O

[0128] 2) Deep structure: Based on the basic structure of GPolar, a shallow-deep-global multi-level feature acquisition strategy is implemented, as Figure 3As shown. Generally speaking, deep feature extraction in GPolar involves gradually increasing the receptive field of dense units in the basic structure. The polar coordinate arrangement of GPolar expands along the radial direction ρ and the angular direction θ of the polar coordinate system. Specifically, in the outer ring, the descriptor vectors of adjacent units along the ρ direction by 2 units and the θ direction by N d are weighted and averaged together with the central unit vector H : 0 together with the central unit vector H

[0129]

[0130] where the normalization factor N d ensures equal weights for all units, and the weighting factor λ d controls the weight ratio of the extracted deep features relative to H 0 , and to prevent the imbalance of unit weights from reducing the invariance. Compared with the shallow perception performed by the basic structure, it realizes the deep perception of the image.

[0131] In addition, all deep descriptors are finally weighted and averaged:

[0132]

[0133] The final H 4 represents the global perception of MOM within the neighborhood of the key point. Finally, the multi-level descriptors are concatenated to generate a complete GPolar descriptor for each key point, with a total length of (2 + 3N A ) × N O :

[0134]

[0135] In most hand-designed methods, the statistical features of sub-regions are isolated. Although this usually has the least impact on high-quality images, local spatial distortions or deformations may cause significant changes in local features within each sub-region. Since the descriptor structure is rigid, this reduces its invariance. The design of the deep structure alleviates this problem because the gradually increasing receptive fields at different levels provide adaptability to local distortions, thereby reducing the overall impact on the descriptor.

[0136] 3) Cross-modal rotation adaptation: Assume that the image may be rotated arbitrarily within the range of [0, 2π). To achieve rotational invariance, each key point is assigned a reference direction θ0 ∈ [0, π). Taking advantage of the fact that MOM inherently encodes direction information, θ0 is obtained by obtaining N within the entire GPolar window (with a radius of R2) ODetermined by the statistical histogram of each unit. The direction with the highest proportion is selected as the main reference direction, and the additional directions with a proportion exceeding 80% of the main direction are regarded as auxiliary reference directions. Then, based on θ0, the entire spatial structure and angle division of GPolar are rotated to achieve effective rotational invariance.

[0137] For cross-modal data, restricting the angle to will lead to a serious descriptor inversion problem. This phenomenon occurs in two cases: (i) Rotation of the image causes θ0 to exceed (ii) Modal differences cause θ0 to be close to crossing the boundary or Both of these cases will cause descriptor inversion due to the mechanisms of MOM and GPolar. In the present invention, an Orientation Consistency Discrimination (OCD) module is proposed. The extracted shallow descriptors and are divided into two halves along the direction defined by the line passing through θ0.

[0138]

[0139] where σ 2 (X) is the variance calculated from all elements in matrix X. Then, deep descriptors are extracted through equations (27) and (28). This variance reflects the degree of dispersion of the main direction within each unit, which also indicates the complexity of the texture. For example, if there are significant edge features, the MOM value in this region tends to be consistent with the direction of the edge, resulting in a smaller variance.

[0140] This method assumes that the texture complexity within the regions on both sides of the key point reference direction is likely to remain consistent between cross-modal images. Therefore, through OCD, the variance value is used to verify whether the reference direction has flipped and to determine whether to swap the descriptors of the two regions.

[0141] Step 5: HOMO - Key Point Matching

[0142] The present invention adopts a multi-scale strategy MsS to enable feature points to iteratively strengthen the matching scale by scale layer, and also provides multi-scale accommodation ability for HOMO, thereby achieving scale difference robustness in cross-modal matching. In addition to key point detection, the overall process of HOMO is carried out in the scale space, where the Gaussian pyramid and MOM pyramid of the image are directly obtained in the construction of DoM. N oct downsampling groups and N lyr layers of Gaussian blurring are predefined, and the total number of matching operations is times.

[0143] As Figure 4 shown, the iterative mechanism uses the input key point i from the matching of the previous scalek , k = 1, …, N M information s of k , k = 1, …, N M -1 to guide the matching at the subsequent scale. Key points that are effectively matched at the previous scale do not require further descriptor extraction or participate in the matching at the next scale. In addition, the effective matches at each layer are regarded as the spatial reference from the current layer to the next layer to guide the matching at the next layer. Specifically, these matching pairs are input into the outlier removal at the next layer and processed continuously as samples participating in the calculation. The multi-scale matching is carried out simultaneously with the descriptor extraction, so the matched key points s k filtered from i k+1 reduce the burden of unnecessary descriptor extraction. This method accelerates the establishment of stable spatial relationships between key points and reduces the computational overhead by filtering layer by layer. This process can be regarded as continuously refining the effective matches from the initial set of key points.

[0144] Finally, the matches o at all scales k , k = 1, …, N M are unioned, and then outlier removal is performed to generate a set of matching point pairs. Among them, the Nearest Neighbor (NN) algorithm is used as the feature matcher, and the feature vectors are matched one by one through the Euclidean distance measurement, and the Random Sample Consensus (RANSAC) algorithm is used for mismatch elimination to remove the incorrect matching points.

[0145] Step 6: Spatial transformation model estimation

[0146] After obtaining a set of matching points between cross-modal images, a spatial transformation (mapping) model is selected, such as similarity transformation, affine transformation (first-order polynomial), projective transformation (projective transformation, homography transformation), polynomial transformation, etc. The parameters of the transformation model are estimated through the spatial coordinate correspondence relationship between the matching points, and the spatial matching relationship of the cross-modal images is obtained, providing a key data basis for subsequent related tasks of multi-source image joint application.

[0147] Embodiment 2:

[0148] The present invention provides a cross-modal image matching system that is universal for any source, and this system can be used to implement the above-mentioned cross-modal image matching method that is universal for any source. Specifically, it includes:

[0149] Image preprocessing module: Normalize and sum the multi-source cross-modal images to be matched, convert them into single-channel grayscale images, and use Gaussian filtering for denoising to obtain the preprocessed single-channel grayscale images.

[0150] Gaussian Scale Pyramid Construction Module: Based on the preprocessed single-channel grayscale image, construct the Gaussian scale pyramid of the input image. Generate multiple scale images through downsampling and Gaussian blur. Calculate the main direction feature MOM of each layer of the image in the scale space, extract the core invariant features of the image, and convert them into a pixel-level feature domain to form the MOM pyramid of the image;

[0151] MOM Difference Pyramid Generation and Feature Point Detection Module: Based on the MOM pyramid of the image, calculate the difference of multi-scale MOM to generate the MOM difference pyramid (DoM). Normalize and weight each layer of the DoM, calculate the MOM change weight for enhancing the localization of feature points. Use the PC-ShiTomasi feature point detector, combine the MOM change weight to enhance the key point response of PC-ShiTomasi, and accurately locate stable and significant key points through local non-maximum suppression LNMS;

[0152] Key Point Descriptor Extraction Module: Take the coordinates and response values of the detected key points as input. At each scale level, extract the GPolar descriptor for the detected key points. This descriptor includes the basic structure and the deep structure, ensuring rotation invariance and scale robustness. Through the orientation consistency discrimination OCD, output the GPolar descriptor of each key point, including multi-scale and multi-level feature vectors;

[0153] Key Point Matching Module: Adopt the multi-scale strategy MsS, iteratively strengthen key point matching layer by layer to improve the matching accuracy and stability. Use the nearest neighbor NN algorithm for feature matching based on the Euclidean distance, and remove the wrong matching points through the RANSAC algorithm to obtain the matching key point pairs, that is, the corresponding point pairs in the cross-modal images;

[0154] Spatial Transformation Model Estimation Module: According to the spatial coordinates of the matching key point pairs, select the spatial transformation model. Estimate the parameters of the transformation model through the spatial coordinate relationship between the matching points, calculate the spatial matching relationship between the images, and finally output the parameters of the spatial transformation model of the cross-modal images for image registration.

[0155] Example 3:

[0156] This embodiment provides a terminal device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function. The processor described in the embodiments of the present invention can be used for the operation of an arbitrary-source general cross-modal image matching method, including the following steps:

[0157] S1: Perform normalization and summation processing on the multi-source cross-modal images to be matched, convert them into single-channel grayscale images, and use Gaussian filtering for denoising to obtain the preprocessed single-channel grayscale images.

[0158] S2: Based on the preprocessed single-channel grayscale images, construct a Gaussian scale pyramid of the input images, generate multiple scale images through downsampling and Gaussian blur, calculate the main direction feature MOM of each layer of images in the scale space, extract the core invariant features of the images, and convert them into a pixel-level feature domain to form the MOM pyramid of the images;

[0159] S3: Based on the MOM pyramid of the images, calculate the difference of the multi-scale MOMs to generate a MOM difference pyramid (DoM), normalize and weight each layer of the DoM, calculate the MOM change weight for enhancing the localization of feature points; use the PC-ShiTomasi feature point detector, combine the MOM change weight to enhance the key point response of the PC-ShiTomasi, and accurately locate the stable and significant key points through local non-maximum suppression LNMS;

[0160] S4: Take the coordinates and response values of the key points detected in step S3 as inputs. At each layer of scale, extract the GPolar descriptors for the detected key points. The descriptor includes a basic structure and a deep structure to ensure rotation invariance and scale robustness. Through the orientation consistency discriminant OCD, output the GPolar descriptors of each key point, including multi-scale and multi-level feature vectors;

[0161] S5: Adopt the multi-scale strategy MsS, iteratively strengthen the key-point matching layer by layer to improve the matching accuracy and stability. Use the nearest neighbor (NN) algorithm to perform feature matching based on the Euclidean distance, and remove the mismatched points through the RANSAC algorithm to obtain the matched key-point pairs, that is, the corresponding point pairs in the cross-modal images.

[0162] S6: According to the spatial coordinates of the matched key-point pairs, select a spatial transformation model. Estimate the parameters of the transformation model through the spatial coordinate relationship between the matching points, calculate the spatial matching relationship between the images, and finally output the parameters of the spatial transformation model of the cross-modal images for image registration.

[0163] Example 4:

[0164] This example provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is the memory device in the terminal device, used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, and this storage space stores the operating system of the terminal. And in this storage space, there is also stored one or more instructions suitable for being loaded and executed by the processor. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory.

[0165] One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the cross-modal image matching method for any source as described in the above example; one or more instructions in the computer-readable storage medium are loaded and executed by the processor to perform the following steps:

[0166] S1: Perform normalization and summation processing on the multi-source cross-modal images to be matched, convert them into single-channel grayscale images, and use Gaussian filtering for denoising to obtain the preprocessed single-channel grayscale images.

[0167] S2: Based on the preprocessed single-channel grayscale images, construct a Gaussian scale pyramid of the input images, generate multiple scale images through downsampling and Gaussian blur, calculate the main orientation feature MOM of each layer of images in the scale space, extract the core invariant features of the images, and convert them into a pixel-level feature domain to form the MOM pyramid of the images.

[0168] S3: Based on the MOM pyramid of the image, calculate the difference of multi-scale MOM to generate a Difference of MOM pyramid (DoM). Normalize and weight each layer of the DoM, calculate the MOM change weight for enhancing the localization of feature points. Adopt the PC-ShiTomasi feature point detector, combine with the MOM change weight to enhance the key point response of PC-ShiTomasi, and accurately locate stable and significant key points through Local Non-Maximum Suppression (LNMS).

[0169] S4: Take the coordinates and response values of the key points detected in step S3 as input. At each scale level, extract GPolar descriptors for the detected key points. This descriptor includes a basic structure and a deep structure to ensure rotation invariance and scale robustness. Through Orientation Consistency Discrimination (OCD), output the GPolar descriptor of each key point, including multi-scale and multi-level feature vectors.

[0170] S5: Adopt a multi-scale strategy (MsS), iteratively strengthen key point matching layer by layer to improve the matching accuracy and stability. Use the Nearest Neighbor (NN) algorithm for feature matching based on the Euclidean distance, and remove incorrect matching points through the RANSAC algorithm to obtain the matched key point pairs, i.e., the corresponding point pairs in cross-modal images.

[0171] S6: According to the spatial coordinates of the matched key point pairs, select a spatial transformation model. Estimate the parameters of the transformation model through the spatial coordinate relationship between the matching points, calculate the spatial matching relationship between the images, and finally output the parameters of the spatial transformation model of the cross-modal images for image registration.

[0172] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can be in the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0173] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for realizing the process Figure 1One or more processes and / or blocks Figure 1 Apparatus for the functions specified in one or more blocks

[0174] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the processes Figure 1 One or more processes and / or blocks Figure 1 The functions specified in one or more blocks

[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the processes Figure 1 One or more processes and / or blocks Figure 1 The functions specified in one or more blocks

[0176] The preferred embodiments of the present invention have been described in detail above, but the present invention is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention

[0177] Many other changes and modifications can be made without departing from the concept and scope of the present invention. It should be understood that the present invention is not limited to specific embodiments, and the scope of the present invention is defined by the appended claims

Claims

1. A cross-modal image matching method that is universal for any source, characterized in that, The method includes: S1: Based on the preprocessed single-channel grayscale image, construct a Gaussian scale pyramid of the input image. Generate multiple scale images through downsampling and Gaussian blur. Calculate the main direction feature MOM of each layer of the image in the scale space, extract the core invariant features of the image, and convert them into a pixel-level feature domain to form the MOM pyramid of the image. S2: Based on the MOM pyramid of the image, calculate the difference of the multi-scale MOM to generate the MOM difference pyramid (DoM). Normalize and weight each layer of the DoM, calculate the MOM change weight for enhancing the localization of feature points. Use the PC-ShiTomasi feature point detector, combine with the MOM change weight to enhance the key point response of PC-ShiTomasi, and accurately locate stable and significant key points through local non-maximum suppression LNMS. S3: Take the coordinates and response values of the key points detected in step S2 as inputs. At each scale level, extract the GPolar descriptor for the detected key points. This descriptor includes a basic structure and a deep structure to ensure rotation invariance and scale robustness. Through the orientation consistency discriminant OCD, output the GPolar descriptor of each key point, including a multi-scale and multi-level feature vector. S4: Adopt the multi-scale strategy MsS, iteratively strengthen the key point matching layer by layer to improve the matching accuracy and stability. Use the nearest neighbor NN algorithm for feature matching based on the Euclidean distance, and remove the wrong matching points through the RANSAC algorithm to obtain the matching key point pairs, that is, the corresponding point pairs in the cross-modal images. S5: According to the spatial coordinates of the matching key point pairs, select a spatial transformation model. Estimate the parameters of the transformation model through the spatial coordinate relationship between the matching points, calculate the spatial matching relationship between the images, and finally output the parameters of the spatial transformation model of the cross-modal images for image registration.

2. The cross-modal image matching method that is universal for any source according to claim 1, wherein Perform normalization and summation processing on the multi-source cross-modal images to be matched, convert them into single-channel grayscale images, and use Gaussian filtering for denoising to obtain the preprocessed single-channel grayscale images.

3. An arbitrary-source general cross-modal image matching method according to claim 1, characterized in that The specific steps of step S1 include: S101: Anisotropic LogGabor weighting According to the principle of LogGabor, the 2D-LG kernel function is: where (ρ, θ) are polar coordinates; s is the scale index and o is the orientation index; ρ s and θ o are the center frequency and the preferred orientation at scale s and orientation o, respectively; σ ρ and σ θ are the bandwidths in ρ and θ, respectively; LG in Equation (1) is defined in polar coordinates and its Cartesian coordinate representation is: wherein each pixel (x, y) is composed of the real part formed by even-LG and the imaginary part formed by odd-LG, where i represents the imaginary unit; First, the single-band input image I(x, y) is filtered by a predefined with scale s and orientation o: where the real part has a physical interpretation as the edge response at a specific scale s and orientation o, while the imaginary part can be interpreted as the contour response, * denotes convolution, and then, the LG response is projected onto the orthogonal x and y directions according to the orientation o. First, the odd-LG response is directionally weighted: where the responses at different scales s in each direction o are added together; Define an anisotropic even-LG weighting method. For convenience of expression, a sign function is defined: The anisotropic even-LG component is defined as follows: wherein represents the Hadamard product. In the above definition, the sign of the even-LG response is flipped according to the sign of the odd-LG response at the corresponding pixel position. Using the sign of the odd-LG response as a reference, it is ensured that the opposite direction of the even-LG response is correspondingly assigned the opposite sign; This method uses the odd-LG response to guide the weighting of the even-LG response, thereby generating a differentiated even-LG response component. This component has the characteristic of having an opposite sign in the projection component in the opposite direction; further use the differentiated even-LG response for direction weighting: Then, obtain the amplitude and orientation maps of the odd imaginary part and even real part of the LG response: Finally, two gradient features are obtained, and their angular values are in [0, π) or correspond to the local structure direction of the image; both the odd part and the even part exhibit rotational invariance, and multi-scale weighting enhances stability and scale robustness; S102: Mean square weighting To further enhance the features, the Average Squared Gradient (ASG) is adopted; ASG is a local gradient weighting method, and the elementary gradients of the image in the x and y directions, namely G x and G y , are calculated as follows: where I(x, y) represents the input single-layer grayscale image, and the magnitude and direction of its gradient, G ρ and are expressed as: Assume that a new gradient is calculated by squaring the original amplitude and doubling the angle, denoted as: In this transformation, all vectors that originally had opposite angles (differing by π) within the range [-π, π] are now aligned to the same angle (differing by 2π); then, the new gradients G s,x and G s,y are expressed as: Substitute equation (12) into equation (13) to get: Then, substitute equation (11) into equation (14) to obtain the relationship between the new gradient and the original gradient in Cartesian coordinates: The key purpose is to weight the local gradient information; in ASG, the local weighted squared gradients G σ,s,x and G σ,s,y are expressed as: where W σ is a Gaussian window with variance σ, denotes applying W σ to the neighborhood weighting at each position in the input matrix, which is also understood as Gaussian filtering. Therefore, the direction of this gradient is: where ∠(X,Y) is defined as: such that In the range (-π, π), according to Equation (12), the weighted original gradient is now calculated as: Based on the ASG, a multi-scale weighting scheme is defined, where G σ,ρ is jointly weighted at multiple scales, and for each scale, the component G of the Gaussian window variance σ σ,s,x and G σ,s,y are weighted as follows: where λ σ is the weight of the scale σ. Then, according to Eqs. (17), (19), and (20), the average squared weighted ASW feature map is obtained: S103: Odd-even feature coupling The ASW in the previous stage, as shown in equation (21), is in the form of a function: Where X and Y represent two two-dimensional matrices of equal size as inputs:, then, ASW is applied to the odd-LG and even-LG responses obtained in the first stage, substituting equations (4) and (7) into equation (22), thus obtaining a method for calculating average square LogGabor features: In this way, the odd and even features respectively provide stable representations of the directions of significant local edges and contour structures in the image. Finally, a coupling operation is applied to these two feature maps to obtain the MOM: Among them, the dominant feature of each position is determined, and its direction is assigned as the local direction of the image; the calculation of MOM can use the sign function sgn * It is expressed as: At this stage, the final feature map M is obtained MOM , which, as the core of the HOMO algorithm, is a low-level pixel-level invariant feature.

4. An arbitrary-source general cross-modal image matching method according to claim 1, characterized in that The step S2 specifically includes: Extract the MOM in the Gaussian scale space, providing a MOM pyramid that reflects the local main directions at different scales; by calculating the differences (i.e., absolute values) of each layer of the MOM pyramid and summing them, a change response DoM of the MOM at different scales is obtained. Sum and normalize all layers of the DoM to calculate the MOM change weight; the reciprocal of the MOM change weight is used as the weight, and then multiplied by the key point response to identify the feature positions with prominent and stable structures, enhancing the key point response. This weight can be applied to any pixel-level feature detector to enhance key point localization; use the PC-ShiTomasi feature point detection algorithm to calculate the key point response, multiply it by the MOM change weight, and finally detect and locate the key points in the input image through LNMS.

5. An arbitrary-source general cross-modal image matching method according to claim 1, characterized in that, The step S3 specifically includes: S301: Basic structure The basic structure of GPolar has a flexible circular polar coordinate design, which consists of a central circular unit A 0 and an outer ring evenly divided into sector units, which are respectively denoted as where i and j are the unit indices along the radial and angular directions respectively, and N A is the number of sector units in each outer ring. This structure is not fixed, and both the radial and angular subdivisions can be extended. The radii of the central and outer regions are denoted as R0, R1, and R2 respectively; the area of each unit is set to be equal, which fixes the relationship between R0, R1, R2, and N A as follows: So that the number of pixels for counting is roughly the same, mainly used for the fair weights of the units; For each detected key point, GPolar extracts local features within the neighborhood of radius R2 centered on it, thereby generating a descriptor vector. The MOM value within each cell is regarded as or the direction angle within [0, π) is uniformly quantized to φ k , k = 1, 2, …, N O , for each cell constructs a histogram vector with N O directions Then, these vectors are weighted with a Gaussian kernel along the θ direction. Therefore, the basic structure of GPolar generates a descriptor vector of length (1 + 2N A ) × N O for each key point; S302: Deep structure A shallow-deep-global multi-level feature acquisition strategy is implemented based on the basic structure of GPolar. The deep feature extraction in GPolar involves gradually increasing the receptive field of the dense units in the basic structure. The polar coordinate arrangement of GPolar expands along the radial direction ρ and the angular direction θ of the polar coordinate system. Specifically, in the outer ring, the descriptor vectors of adjacent units that are 2 units along the ρ direction and n d units along the θ direction are weighted and averaged together with the central unit vector H 0 : Among them, the normalization factor N d ensures that the weights of all units are equal, and the weighting factor λ d controls the extracted deep features relative to H 0 , and the weight ratio of, to prevent the imbalance of unit weights from reducing the invariance. Compared with the shallow perception performed by the basic structure, it realizes the deep perception of the image; In addition, all deep descriptors are finally weighted and averaged: Final H 4 represents the global perception of MOM within the neighborhood of key points. Finally, the multi-level descriptors are concatenated to generate a complete GPolar descriptor for each key point, with a total length of (2 + 3N A ) × N O : S303: Cross-modal rotation adaptation: Assume that the image can be arbitrarily rotated within the range of [0, 2π). To achieve rotational invariance, each key point is assigned a reference direction θ0 ∈ [0, π). Taking advantage of the fact that MOM inherently encodes direction information, θ0 is determined by obtaining a statistical histogram with N O cells within the entire GPolar window. The direction with the highest proportion is selected as the main reference direction, and additional directions with a proportion exceeding 80% of the main direction are regarded as auxiliary reference directions. Then, based on θ0, the entire spatial structure and angular division of GPolar are rotated, thereby achieving effective rotational invariance; Propose an Orientation Consistency Discriminator (OCD) module and extract shallow descriptors and are divided into two halves along the direction defined by the line passing through θ0: where σ 2 (X) is the variance calculated from all elements in matrix X. Then, the deep layer extracts descriptors through equations (27) and (28). This variance reflects the degree of dispersion of the main direction within each cell, indicating the complexity of the texture. Through OCD, the variance value is used to verify whether the reference direction has flipped and to determine whether to exchange the descriptors of the two regions.

6. A cross-modal image matching method that is universal for any source according to claim 1, characterized in that In the step S4, except for key point detection, the overall process of HOMO is carried out in the scale space, where the Gaussian pyramid and MOM pyramid of the image are directly obtained in the construction of DoM; N oct downsampling groups and N lyr layers of Gaussian blur are predefined, and the total matching operation is carried out times; The iterative mechanism uses the input keypoint i from the previous scale match k ,k=1,…,N M Information k ,k=1,…,N M -1 is used to guide the matching of subsequent scales. The key points that are effectively matched at the previous scale do not require further descriptor extraction or participate in the matching of the next scale. In addition, the effective matches of each layer are regarded as spatial references from the current layer to the next layer to guide the matching of the next layer; specifically, these matching pairs are input into the outlier removal of the next layer and are continuously processed as samples involved in the calculation; multi-scale matching is performed simultaneously with descriptor extraction, so the matched key points s k from i k+1 The key points are filtered out, which speeds up the establishment of stable spatial relationships between key points and reduces the computational overhead by filtering layer by layer. Finally, unionize the matches at all scales k , k = 1, …, N M to generate a set of matching point pairs by performing outlier removal after unionization; among them, the nearest neighbor (NN) algorithm is used as the feature matcher, pairwise matching between feature vectors is performed by measuring the Euclidean distance, and the random sample consensus (RANSAC) algorithm is used to eliminate mismatches and remove incorrect matching points.

7. A cross-modal image matching system that is universal for any source, characterized in that, The system is applied to the method described in any one of claims 1-6. The system includes: Image preprocessing module: Perform normalization and summation processing on the multi-source cross-modal images to be matched, convert them into single-channel grayscale images, and use Gaussian filtering for denoising to obtain the preprocessed single-channel grayscale images. Gaussian scale pyramid construction module: Based on the preprocessed single-channel grayscale image, construct a Gaussian scale pyramid of the input image, generate multiple scale images through downsampling and Gaussian blurring, calculate the main direction feature MOM of each layer of the image in the scale space, extract the core invariant features of the image, and convert them into a pixel-level feature domain to form the MOM pyramid of the image; MOM difference pyramid generation and feature point detection module: Based on the MOM pyramid of the image, calculate the differences of the multi-scale MOM to generate the MOM difference pyramid (DoM), normalize and weight each layer of the DoM to calculate the MOM change weight for enhancing the localization of feature points; use the PC-ShiTomasi feature point detector, combine the MOM change weight to enhance the key point response of PC-ShiTomasi, and accurately locate the stable and significant key points through local non-maximum suppression LNMS; Key point descriptor extraction module: Taking the coordinates and response values of the above-detected key points as input, at each layer scale, GPolar descriptors are extracted for the detected key points. This descriptor includes a basic structure and a deep structure to ensure rotational invariance and scale robustness. Through the orientation consistency discriminant OCD, the GPolar descriptor of each key point is output, including a multi-scale and multi-level feature vector. Key point matching module: Adopting the multi-scale strategy MsS, iteratively strengthening key point matching layer by layer to improve the matching accuracy and stability. Using the nearest neighbor NN algorithm for feature matching based on the Euclidean distance, and removing the wrong matching points through the RANSAC algorithm to obtain the matching key point pairs, that is, the corresponding point pairs in the cross-modal images. Spatial transformation model estimation module: According to the spatial coordinates of the matching key point pairs, select a spatial transformation model, estimate the parameters of the transformation model through the spatial coordinate relationship between the matching points, calculate the spatial matching relationship between the images, and finally output the spatial transformation model parameters of the cross-modal images for image registration.

8. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a cross-modal image matching method for any source generality as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the program is executed by the processor, it implements a cross-modal image matching method for any source generality as described in any one of claims 1 to 6.