An image feature matching and fusion method and system based on a multi-modal large model

By using a multimodal large model image feature matching and fusion method, the problems of noise spillover and detail drift in multimodal matching and fusion technology are solved. The method achieves directional alignment and spatial consistency of multimodal data in a unified coordinate system, ensuring the stability and reliability of the fusion process.

CN120932053BActive Publication Date: 2025-12-23西安圣瞳科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511452778.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2025-12-23
Estimated Expiration
2045-10-13

AI Technical Summary

Technical Problem

Existing multimodal matching and fusion technologies face challenges in boundary fidelity and time-varying robustness. The propagation of cross-modal information in structural boundaries and time-varying regions lacks continuous self-inhibition and regional constraints, and the handling of scale and dimension correlation between modes is insufficient, leading to noise spillover and detail drift.

Method used

By employing an image feature matching and fusion method based on a multimodal large model, we acquire intrinsic and extrinsic parameters of different modal data and imaging devices, extract initial features, generate prior information, construct a geometric bias matrix and spatial scale in a unified coordinate system, calculate cross-modal correlation maps and multiplicative injection coefficients, update features and generate preliminary fusion features, construct an information hiding mask for decoupling, generate a cleaned matching set, and achieve final fusion through shared evidence maps and gating adjustment.

Benefits of technology

It achieves directional alignment and spatial consistency of multimodal data in a unified coordinate system, avoids the offset of different modal data in geometric structure, ensures the stability and reliability of the fusion process, and realizes conformal transmission and anomaly suppression of multimodal data in the geometric principal direction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932053B_ABST
    Figure CN120932053B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multimodal big model's image feature matching and fusion method and system, it is related to image processing technical field, including, obtain different modal data and the internal and external parameters of imaging device, extract the initial feature of each modal data, generate prior information;Geometric bias matrix and space scale are constructed in unified coordinate system, based on the initial feature of each modal data and prior information, calculate cross-modal correlation graph and multiplicative injection coefficient, update to obtain the optimized feature of each modal data, normalization generates preliminary fusion feature;Information hiding mask is constructed on preliminary fusion feature, decouples along the geometric main direction of geometric structure optimization feature, and adopts double criterion to select channel and obtain task selected description.The application realizes the direction alignment and space consistency of multimodal data in unified coordinate system by introducing geometric guide calculation and space constraint modeling, guarantees the stability and reliability of fusion process.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of multi-modal information processing, in particular to an image feature matching and fusion method and system based on a multi-modal large model. BACKGROUND

[0002] Multi-modal large models are increasingly widely used in the field of image processing. Visible light images, depth images and infrared images are introduced simultaneously to fuse multi-dimensional information such as texture features, geometric features and radiation features. Existing technologies establish correspondence between modalities through correlation measurement and geometric constraints, and realize scale unification and feature alignment by combining logarithmic domain modeling and robust statistical methods, thereby supporting application scenarios such as image retrieval, image sorting and geometric registration.

[0003] In engineering practice, the conventional process still faces two types of difficulties: firstly, the propagation of cross-modal information in structural boundaries and time-varying regions lacks continuous self-suppression and regional constraints, which easily causes noise overflow and detail drift; secondly, the processing of scale and dimension correlation between modalities is insufficient, which affects the consistency of measurement and the stability of registration. Existing multi-modal matching and fusion technologies face engineering challenges in boundary fidelity and time-varying robustness. SUMMARY

[0004] In view of the above existing problems, the present application is proposed.

[0005] Therefore, the present application provides an image feature matching and fusion method based on a multi-modal large model to solve the problems of cross-modal noise overflow, detail drift and insufficient scale and dimension correlation processing in the image processing process.

[0006] To solve the above technical problems, the present application provides the following technical solutions:

[0007] In a first aspect, the present application provides an image feature matching and fusion method based on a multi-modal large model, which includes obtaining different modal data and internal and external parameters of an imaging device, extracting initial features of each modal data, and generating prior information;

[0008] A geometric bias matrix and a spatial scale are constructed in a unified coordinate system. Based on the initial features of each modal data and the prior information, a cross-modal correlation graph and a multiplicative injection coefficient are calculated, and the features of each modal data are updated to obtain optimized features. The preliminary fusion features are normalized to generate initial fusion features;

[0009] An information hiding mask is constructed on the preliminary fusion features, and decoupling is performed along the geometric main direction of the geometric structure optimization features. A double-criterion screening channel is used to obtain task-selected descriptions;

[0010] An initial matching set is generated under geometric correlation constraints. Residual values are obtained using a closed-form geometric trial solution. After spatial median robustness, a cleaned matching set is generated.

[0011] The conformal transmission field is constructed by the shared evidence graph taking the median of the evidence rank percentile value, data sharing is completed, the features of each modality data are optimized again, and the final fusion features are generated by median aggregation in the logarithmic domain;

[0012] The final fusion features are spliced with the task selected description and unitized to generate the task features of each position, and the registration coordinate transformation matrix is generated in a closed method according to the cleaned matching set.

[0013] As a preferred scheme of the image feature matching and fusion method based on a multi-modal large model, the initial features of each modality data include receiving modality data of visible light images, depth images and infrared images, and using a public convolutional neural network to extract texture color initial features of the visible light images and thermal radiation initial features of the infrared images.

[0014] The geometric structure initial features of the depth images are obtained by calculating the horizontal difference, vertical difference, unit normal three components and curvature intensity of the depth images.

[0015] As a preferred scheme of the image feature matching and fusion method based on a multi-modal large model, the generation of prior information includes estimating the average pixel motion amount between the current frame and the adjacent frame, combining the inter-frame time difference and processing through a smoothing function mapping to obtain a time confidence.

[0016] The continuous effectiveness weight of each modality data is generated according to the respective observation, and the time confidence is used to form the prior information.

[0017] As a preferred scheme of the image feature matching and fusion method based on a multi-modal large model, the normalization generates preliminary fusion features, which includes reading the initial features and prior information of each modality data in a unified coordinate system, combining the geometric bias matrix, spatial scale and current temperature, calculating the cross-modality correlation graph, and establishing the evidence channel of the source transformation target.

[0018] The consistency of the source direction and the target own direction is compared to obtain a multiplicative injection coefficient, a small dose of multiplicative update is performed on each modality data to obtain the optimized features of each modality data, the features are aggregated by taking the median of each position in the logarithmic domain, and then restored to the original data domain to obtain the preliminary fusion features.

[0019] As a preferred scheme of the image feature matching and fusion method based on the multi-modal large model, wherein: the task selection description includes performing multi-scale analysis and normalization processing on the preliminary fusion feature, obtaining a weighted redundancy combined with prior information, and marking the spatial position as a shielding area and a fidelity area according to the relationship between the weighted redundancy and the original redundancy to generate an information hiding mask.

[0020] According to the geometric main direction of the geometric structure optimization feature, the preliminary fusion feature is split into a fine-grained feature channel and a high semantic feature channel through median filtering decomposition within a neighborhood range.

[0021] The conditional distribution is calculated on the cross-modal consistent samples and inconsistent samples, the information gain of each channel is calculated based on the difference of the conditional distribution, and the mutual information index of each channel and the task agent variable is calculated.

[0022] The median of the information gain and the median of the mutual information index are used as double criteria to splice the selected fine-grained feature channel and high semantic feature channel in channel order to generate a task selection description.

[0023] As a preferred scheme of the image feature matching and fusion method based on the multi-modal large model, wherein: the initial matching set is generated under the geometric correlation constraint, including: taking the optimized features of each modal data and the preliminary fusion feature as candidate matching points, adjusting based on the cross-modal correlation graph and the current temperature parameter, determining the main peak position in the pairing relationship between the target position and the source position, and screening the candidate matching points in combination with the geometric bias matrix and the spatial scale constraint.

[0024] For each target position, only the paired relationship with the source position as the main peak is retained, and the peak value of the cross-modal correlation graph, the displacement reading of the geometric bias matrix, the reading of the prior information and the reading of the information hiding mask are recorded for each pair of target position and source position pairing relationship to generate an initial matching set.

[0025] As a preferred scheme of the image feature matching and fusion method based on the multi-modal large model, wherein: the generation of the cleaned matching set includes selecting a closed-form geometric trial solution method under scene conditions, calculating the residual value of the candidate matching point, and mapping the residual value to a rank percentile value, and performing numerical combination calculation with the peak value of the cross-modal correlation graph, the displacement reading of the geometric bias matrix, the reading of the prior information and the reading of the information hiding mask to obtain updated soft weights.

[0026] The residual value is subjected to neighborhood median aggregation to form a regional median re-projection residual field, and the updated soft weights are adjusted according to the regional residual information to generate a cleaned matching set.

[0027] As a preferred scheme of the image feature matching and fusion method based on the multi-modal large model, the generating final fusion features by median aggregation in the logarithmic domain comprises: constructing a shared evidence graph by median statistics of an evidence rank percentile value, and writing shared strength generated by combining a cross-modal correlation graph, a multiplicative injection coefficient, a consistency index, and a reading of prior information into the shared evidence graph.

[0028] In the logarithmic domain, the effective components of the source position optimization features are guided and filtered along the geometric direction of the target position and the boundary to obtain a conformal transmission field;

[0029] The shared evidence graph is used as a gating strength to compress and suppress the conformal transmission field, and is linked with an information hiding mask to obtain a shared update to the target position;

[0030] The shared update is performed on each modality data respectively to obtain re-optimized features, and the features are restored to the original numerical value domain after taking the median at each spatial position in the logarithmic domain to generate final fusion features.

[0031] As a preferred scheme of the image feature matching and fusion method based on the multi-modal large model, the generating final fusion features by median aggregation in the logarithmic domain comprises: constructing a shared evidence graph by median statistics of an evidence rank percentile value, and writing shared strength generated by combining a cross-modal correlation graph, a multiplicative injection coefficient, a consistency index, and a reading of prior information into the shared evidence graph.

[0032] The multiplicative gating factor is used as an adjustment coefficient to proportionally scale and adjust the local rank of each channel to obtain a feature component after gating adjustment, and a task component is formed by applying an information hiding mask constraint to the feature component after gating adjustment.

[0033] In a second aspect, the present application provides an image feature matching and fusion system based on a multi-modal large model, comprising,

[0034] The acquisition processing module is configured to acquire different modality data and internal and external parameters of the imaging device, extract initial features of the modality data, and generate prior information.

[0035] The spatial alignment module is configured to construct a geometric bias matrix and a spatial scale, calculate a cross-modal correlation graph and a multiplicative injection coefficient based on the initial features of the modality data and the prior information, obtain optimized features of the modality data, and generate preliminary fusion features.

[0036] A hidden decoupling module is configured to construct an information hiding mask on the preliminary fused features, decouple along the geometric main direction of the geometric structure, and use a double criterion to select channels to obtain task selected descriptions.

[0037] A matching feedback module is configured to generate an initial matching set under geometric correlation constraints, obtain residual values using a closed-form geometric trial solution, and generate a cleaned matching set after spatial median robustness.

[0038] A shared fusion module is configured to construct a shared evidence graph and a conformal transmission field, complete data sharing, obtain features optimized again from each modality data, and generate final fused features in a logarithmic domain.

[0039] A task solving module is configured to concatenate and unitize the final fused features and the task selected descriptions to generate position-by-position task features, and generate a registration coordinate transformation matrix in a closed-form method according to the cleaned matching set.

[0040] The present application has the following advantages: by introducing geometric guidance calculation and space constraint modeling, the directional alignment and spatial consistency of multi-modal data in a unified coordinate system are realized, the offset of different modal data in the geometric structure is avoided, and the stability and reliability of the fusion process are ensured; by constructing a shared evidence graph and combining median statistics and gate regulation, conformal transmission and abnormal suppression of multi-modal data in the geometric main direction are realized. BRIEF DESCRIPTION OF DRAWINGS

[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0042] Fig. 1 The flowchart of the image feature matching and fusion method based on the multi-modal large model.

[0043] Fig. 2 The schematic diagram of the image feature matching and fusion system based on the multi-modal large model.

[0044] Fig. 3 The flowchart of generating the task selected description.

[0045] Fig. 4 The flowchart of generating the cleaned matching set. DETAILED DESCRIPTION

[0046] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification.

[0047] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without the specific details set forth in this description. In other instances, well-known methods have not been described in detail in order not to unnecessarily obscure aspects of the present application.

[0048] Secondly, the "one embodiment" or "an embodiment" referred to herein means a specific feature, structure, or characteristic under discussion. Thus, "one embodiment" does not mean a single embodiment or that a feature, structure, or characteristic is required in all or in a single embodiment. Also, since numerous specific details of implementation are set forth herein, it is understood that no component, or component of a system, is required to be made or work in accordance with a specific implementation unless explicitly so stated herein.

[0049] Reference Figs. 1-4 For one embodiment of the present application, the embodiment provides a multi-modal large model-based image feature matching and fusion method, comprising the following steps:

[0050] S1, acquire the internal and external parameters of different modal data and imaging devices, extract the initial features of each modal data, and generate prior information.

[0051] Further, receive three modal data of visible light image, depth image and infrared image, and save the internal and external parameter matrices of the three imaging devices respectively.

[0052] Based on the depth image, the pixel position and its corresponding depth value on the depth image are back projected into a three-dimensional point, and the three-dimensional point is converted into a unified coordinate system through the external parameter matrix of the depth imaging device; then the three-dimensional point in the unified coordinate system is projected to the imaging plane of the visible light image and the infrared image through the internal and external parameter matrices of the visible light imaging device and the infrared imaging device, to obtain the pixel position corresponding to the depth image position; after processing, the three modal data establish a position-by-position corresponding relationship in the geometric space.

[0053] Based on the position-by-position corresponding relationship, the aligned visible light image, infrared image and depth image, as well as their position-by-position validity masks and corresponding indexes, are divided into a training set and a validation set; the visible light image is subjected to distortion correction, color correction and gain correction; the infrared image is subjected to black level correction and gain correction, and is converted into physical radiation according to the calibration curve; samples are generated according to the unified resolution and cropping rules, and the position-by-position validity masks and corresponding indexes are output synchronously.

[0054] The initial features of each modal data include using a public convolutional neural network (such as ResNet-50 and MobileNetV2) to extract the texture color initial features of the visible light image and the thermal radiation initial features of the infrared image respectively.

[0055] The initial features of the geometric structure of the depth image are obtained by calculating the horizontal difference, vertical difference, unit normal three-component and curvature strength of the depth image.

[0056] Specifically, on the depth image, the horizontal difference and vertical difference of the depth image are calculated by using difference operators respectively, and the three-dimensional point set obtained by inverse projection of the pixel neighborhood is approximated as a tangent plane, and the unit normal three-component is obtained based on the geometric relationship of the tangent plane; at the same time, the curvature strength is approximated and calculated by using a discrete Laplace operator on the depth image, and the unit normal three-component, the curvature strength and the depth effectiveness weight are stacked in the order of the channel to obtain the initial features of the geometric structure of the depth image.

[0057] The visible light image feature extraction network (based on ResNet-50) and the infrared image feature extraction network (based on MobileNetV2) are set; in the visible light image feature extraction network, the publicly pre-trained ResNet-50 is loaded, the classification layer is deleted, and the middle and high layer convolution is retained; the middle and high layer convolution result of the visible light image feature extraction network is sequentially mapped to a dense representation with a fixed number of channels by a point convolution layer, a normalization layer and a nonlinear activation, which is defined as the initial texture color feature; in the infrared image feature extraction network, the publicly pre-trained MobileNetV2 is loaded, the first layer is changed to a single channel or adapted by channel replication; in the backbone of the infrared image feature extraction network, the last resolution reduction operation is cancelled, and the convolution at this place is set to not change the spatial resolution; the equal-interval holes are set in the subsequent continuous convolution, which maintains the feature map resolution while maintaining the receptive field, so that the convolution result obtained by the infrared image feature extraction network is the same as the spatial sampling interval of the visible light image feature extraction network; the backbone convolution result of the infrared image feature extraction network is sequentially mapped to a dense representation with a fixed number of channels by a point convolution layer, a normalization layer and a nonlinear activation, which is defined as the initial thermal radiation feature.

[0058] For the visible light image feature extraction network and the infrared image feature extraction network, combined with the position-by-position correspondence and the position-by-position effectiveness mask, the training target is defined as follows: the intra-modal consistency constraint is performed on the feature maps of the visible light image feature extraction network and the infrared image feature extraction network, two light-weight enhanced samples are generated for the same original image, the feature vectors of the feature maps of the two sets of feature extraction networks at the same spatial position are used as positive samples by using the position-by-position correspondence index, and the feature vectors at other positions are used as control pairs, and the similarity is used as a measure to reduce the distance between the positive sample vectors while suppressing the similarity of the control pairs; the cross-modal alignment constraint is performed on the feature maps of the visible light image feature extraction network and the infrared image feature extraction network, the feature vectors at the same spatial position of the feature maps of the two sets of feature extraction networks are paired in a unified coordinate system, and the distance between the paired vectors is minimized, thereby realizing cross-modal alignment; the boundary consistency regularization is performed on the feature maps of the visible light image feature extraction network and the infrared image feature extraction network, based on the unit normal three-component and the curvature intensity calculated from the depth image, at each spatial position, the feature maps of the two sets of feature extraction networks are respectively subjected to finite difference along the normal and tangent directions, and the directional consistency constraint is applied to the normal difference, and the smoothing constraint is applied to the tangent difference; the thermal radiation monotonicity constraint is performed on the feature maps of the infrared image feature extraction network, and the scalar reading corresponding to the position with higher radiation is required to be not lower than the position with lower radiation; the relative importance weights are set for the intra-modal consistency constraint, the cross-modal alignment constraint, the boundary consistency regularization, and the thermal radiation monotonicity constraint, respectively; the four losses are proportionally superimposed to form the overall loss according to the relative importance, and the contribution of the missing, occluded and mismatched positions is shielded by the position-by-position effectiveness mask; the overall loss is used to update the trainable parameters of the visible light image feature extraction network and the infrared image feature extraction network.

[0059] After the overall loss is determined, the parameter update is performed in the following order: first, only the point convolution layer, the normalization layer and the nonlinear activation layer are updated, and the remaining convolution layers of the two sets of feature extraction networks are kept unchanged; then the convolution layers located in the middle layer are unlocked and updated jointly with the aforementioned three types of layers; finally, the convolution layers close to the output end are unlocked, and the whole network fine-tuning is completed; the batch data is organized by the position-by-position correspondence, and two light-weight enhanced views are generated for each sample, the enhancement is limited to monotonic mapping of brightness and contrast, limited translation and rotation, and random cropping, and the transformation that destroys the position-by-position correspondence is prohibited.

[0060] Optimization adopts an adaptive method with momentum, the learning rate decreases stage by stage with the unlocking range, only updating three types of layers adopts a larger step size, reducing the step size after unlocking the middle layer, and further reducing after unlocking the high layer; during training, the cross-modal alignment error and the intra-modal consistency index are calculated on the validation set at fixed intervals, when the change amplitude of the two indicators in continuous evaluation continues to decrease and remains stable, and there is no longer a downward trend, it is determined to converge and stop updating, otherwise continue iteration; after training, save all parameters of the visible light image feature extraction network and the infrared image feature extraction network, as well as the configuration of the point convolution layer, the normalization layer and the nonlinear activation layer, for generating texture color initial features and thermal radiation initial features.

[0061] In the inference stage, the visible light image and the infrared image which are geometrically and radiometrically corrected and unified in size are input, and the texture color initial features and the thermal radiation initial features are output by the two sets of feature extraction networks respectively, together with the geometric structure initial features calculated from the depth image, as the initial features of each modal data.

[0062] Further, generating prior information includes estimating the average pixel motion between the current frame and the adjacent frame, combining the inter-frame time difference and processing through a smoothing function to obtain a time confidence.

[0063] Each modal data generates a visible light image continuity effectiveness weight, a depth image continuity effectiveness weight and an infrared image continuity effectiveness weight according to its own observation, and together with the time confidence, it constitutes the prior information.

[0064] In the unified coordinate system, the visible light image, the depth image and the infrared image are respectively executed radiometric consistency and adaptive denoising integration processing, specifically, the black level correction, gain correction and dark corner correction are performed on the visible light image and the infrared image; on the infrared image, the original count value is converted into physical radiation according to the calibration curve; the experience cumulative distribution of the current frame is monotonically mapped to the reference distribution through fixed quantile anchor points, realizing the consistency of brightness aperture and radiation aperture; for the depth image, no radiation calibration is performed, only the amplitude aperture is unified, and the hole filling and abnormal flying point correction are adopted; on the visible light image, the depth image and the infrared image, the unified processing aperture is adopted to perform edge preservation, distribution self-calibration and time-consistent light denoising, and the processing result presents a smooth effect in the flat area and maintains structural fidelity in the geometric boundary and thermal boundary.

[0065] Between the current frame and the adjacent frame, the motion amplitude of all valid pixel positions is calculated, and the motion amplitude is averaged to obtain the average pixel motion amount, combined with the inter-frame time difference, and processed through a smoothing function mapping to obtain the time confidence; the visible light image continuous validity weight is obtained by combining the gradient amplitude robust normalization result and the monotonic mapping of the brightness quantile coordinate; the depth image continuous validity weight is calculated by self-calibration smoothing mapping of local depth fluctuation degree, and a linear penalty is applied to obtain the missing ratio; the infrared image continuous validity weight is obtained by mapping the infrared pixel value to the upper tail coordinate of the empirical distribution, and applying a smoothing top suppression.

[0066] In the unified coordinate system, the initial features of the visible light image, the depth image and the infrared image are processed by a bicubic interpolation method for resolution consistency, and the initial features of the three modal data are adjusted to the same resolution; the continuous validity weights of the visible light image, the depth image and the infrared image are processed by a bilinear interpolation method for size consistency, and the continuous validity weights of the three modal data are synchronized to the same size.

[0067] S2, in the unified coordinate system, a geometric bias matrix and a spatial scale are constructed, and based on the initial features and prior information of each modal data, a cross-modal correlation map and a multiplicative injection coefficient are calculated, and the optimized features of each modal data are updated to generate a preliminary fusion feature.

[0068] Specifically, the normalized preliminary fusion feature includes reading the initial features and prior information of each modal data in the unified coordinate system, calculating the cross-modal correlation map combined with the geometric bias matrix, the spatial scale and the current temperature, and establishing the evidence channel of the source transformation target.

[0069] The consistency of the source direction and the target self-direction is compared to obtain the multiplicative injection coefficient, and a small dose of multiplicative update is performed on each modal data to obtain the optimized features of each modal data, which are aggregated in the logarithmic domain by taking the median value position by position, and then restored to the original data domain to obtain the preliminary fusion feature.

[0070] Further, in the unified coordinate system, a position-by-position correspondence table is defined first, which records the pixel coordinates and continuous validity weights of the visible light image, the depth image and the infrared image on their respective imaging planes for each spatial index; when any spatial index is selected and the target modal and the source modal are specified, the target position is taken from the target modal pixel coordinates in the correspondence table, and the source position is taken from the source modal pixel coordinates in the correspondence table; only when the target position and the source position are both in the effective imaging range and the continuous validity weights of the two positions are greater than zero, it is determined as a pair of valid target position and source position.

[0071] two types of guidance information are established for each pair of valid target position and source position; the first type is geometric guidance, a geometric bias matrix is formed by the spatial coordinate difference between the target position and the source position; the second type is reliability and time guidance, the continuous validity weight of the visible light image, the continuous validity weight of the depth image, the continuous validity weight of the infrared image and the time confidence are combined into a continuous reliability reading index at the source position as the spatial scale.

[0072] Based on the geometric bias matrix and the spatial scale, the search neighborhood of the source position in the coordinate system of the target position is determined; the amplitude normalization processing is performed on the initial feature of the target position and the initial feature of each source position in the search neighborhood, and the direction consistency reading is calculated as the content similarity; the content similarity and the weight of the geometric bias matrix at the corresponding source position are monotonically combined to obtain the original correlation reading; the normalized mapping controlled by the temperature adjustment coefficient is applied to the original correlation reading in the search neighborhood to form the cross-modal similarity distribution.

[0073] The temperature adjustment coefficient is jointly determined at each target position according to the spatial scale, the time confidence, the continuous validity weight of the visible light image, the continuous validity weight of the depth image, the continuous validity weight of the infrared image, the information hiding mask value and the geometric main direction intensity of the geometric structure optimization feature; when the spatial scale is large, the time confidence is low, or any of the continuous validity weights is low, the information hiding mask is marked as should be shielded, the texture is weak or there is missing measurement value, the temperature adjustment coefficient is increased; when the spatial scale is small, the time confidence is high, and the three continuous validity weights are reliable, the information hiding mask is marked as should be faithful, and it is located at the boundary or high gradient position, the temperature adjustment coefficient is reduced; the example value range of the temperature adjustment coefficient is set to 0.5 to 3.0, the conventional working interval is set to 0.6 to 2.0, and the baseline is 1.0; the reasons for selection are as follows: first, to ensure that the normalized mapping maintains numerical stability within the similarity reading range and avoids gradient abnormalities; second, when the temperature adjustment coefficient is less than about 0.5, the distribution is too sharp, and noise easily triggers peak jumping, which is not conducive to weak texture area candidate coverage; third, when the temperature adjustment coefficient is higher than about 3.0, the distribution is too flat, and it is difficult to distinguish the content difference and the geometric bias; fourth, taking 1.0 as the baseline facilitates alignment with the normalized aperture without temperature adjustment, and converges to 0.6 to 0.9 when reliability is improved or boundary is strengthened to highlight the main peak, and expands to 1.2 to 2.0 when reliability is reduced or texture is sparse to increase candidate coverage.

[0074] The cross-modal correlation graph is generated by superimposing a geometric bias matrix and a spatial scale to form a smooth offset control, and performing normalization processing; the cross-modal correlation graph is used to explain how the target position aggregates evidence on the source information, wherein the geometric bias matrix determines the sensitivity of the correlation to the spatial distance, the temperature parameter adjusts the concentration or dispersion of the similarity distribution, and the spatial scale sets the range of action of the geometric bias.

[0075] Further, in the coordinate system of the cross-modal correlation graph, the information of the source position is aggregated to obtain a candidate injection amount of the source pointing to the target; at each target position, the direction of the source position is compared with the direction of the target itself to obtain a continuous reading of direction coherence; after the continuous reading is compressed to a stable interval by a smoothing function, it is mapped to a multiplicative injection coefficient, and an example value range is between 0.9 and 1.1.

[0076] When the source direction is consistent with the target direction, the multiplicative injection coefficient is slightly greater than 1, indicating that the direction of the target position is slightly amplified; when there is a slight difference between the source and target directions, the multiplicative injection coefficient is close to 1, indicating that it is almost unchanged; when the source and target directions are opposite, the multiplicative injection coefficient is less than 1, indicating that the direction of the target position is slightly suppressed.

[0077] At each target position, the reference contribution of the target position itself and the multiplicative injection coefficients of the two source positions are combined into a three-element set, and the final multiplicative change amount of the target position is determined by median statistics; when the correction amounts of the two modalities are different, or one of the modalities is disturbed by noise, the median will automatically select the middle change amount, so that the characteristics of the target position only undergo a small multiplicative adjustment, and the multiplicative adjustment acts on the mutual update between the initial characteristics of the geometric structure, the initial characteristics of the thermal radiation, and the initial characteristics of the texture color in turn, to obtain the optimized characteristics of the modal data, including the optimized texture color characteristics, the optimized thermal radiation characteristics, and the optimized geometric structure characteristics.

[0078] The optimized characteristics of the modal data are normalized and converted to the logarithmic domain channel by channel, and the optimized characteristics of the modal data are aggregated by position in the logarithmic domain, and finally returned to the original numerical domain to generate preliminary fusion features.

[0079] S3, constructing an information hiding mask on the preliminary fusion features, decoupling along the geometric main direction of the optimized geometric structure, and using a double-criterion screening channel to obtain a task-selected description.

[0080] Specifically, the task-selected description includes performing multi-scale analysis and normalization processing on the preliminary fusion features, obtaining a weighted redundancy in combination with prior information, and according to the relationship between the weighted redundancy and the original redundancy, marking the spatial position as a shielding area and a fidelity area to generate an information hiding mask.

[0081] According to the geometric main direction of the geometric structure optimization feature, the median filtering decomposition is performed within the neighborhood range, the preliminary fusion feature is split into fine-grained feature channels and high semantic feature channels, and the information gain of each channel is calculated by combining the statistical results of the task agent variable on the cross-modal consistent samples and the cross-modal inconsistent samples.

[0082] The fine-grained feature channels and the high semantic feature channels screened out are spliced according to the channel order by adopting the double criteria of median and mutual information, and a task selected description is generated.

[0083] Further, in the preliminary fusion feature, the contents that do not participate in subsequent matching and recognition, and even cause interference are filtered out, for example, large-area flat areas, periodic repeated textures, fine high-frequency information caused by noise and obvious imaging artifacts; a pyramid structure is constructed by adopting Gaussian filtering and Laplace filtering to analyze the preliminary fusion feature on multiple scales, on each scale, the high-frequency intensity reading and the dispersion reading of the corresponding scale are calculated respectively, and the median absolute deviation method is used to normalize the high-frequency intensity reading and the dispersion reading respectively; through the multi-scale analysis and normalization processing, a unified measurement standard is established between different images, different devices and different scales; the area with weak high frequency is compressed after normalization, indicating information redundancy; the area with strong high frequency is highlighted after normalization, indicating that it contains more effective information.

[0084] In the unified coordinate system, a number of representative positions are selected by adopting the grid and texture layering method, the scale redundancy is calculated for each spatial position by scale, and the discrete redundancy reading arranged by scale is formed into a monotone non-decreasing continuous curve by interpolation; for any spatial position, the redundancy reading sequence arranged by scale is obtained and scale consistency correction is performed; after the correction is completed, the redundancy reading sequence of each spatial position and each channel is aggregated between scales in the logarithmic domain in a median manner; then the aggregation result is restored to the original numerical domain to obtain the multi-scale consistent redundancy.

[0085] After obtaining the multi-scale consistent redundancy, the prior information is used as a multiplicative factor, and the visible light image continuous effectiveness weight, the depth image continuous effectiveness weight, the infrared image continuous effectiveness weight and the time confidence are jointly applied to the multi-scale consistent redundancy to obtain a weighted redundancy; the weighted redundancy is used for region marking of the spatial position, when the three modal data are reliable at the same spatial position and the time confidence is consistent, the weighted redundancy is close to the original redundancy, and the spatial position is marked as a region that should be faithful; when any modal data is unreliable at the same spatial position or the time confidence is inconsistent, the weighted redundancy is lower than the original redundancy, and the spatial position is marked as a region that should be shielded.

[0086] In the spatial dimension, the weighted redundancy is locally calibrated by the median in the small window, and the geometric principal direction of the geometric optimization feature is constrained along the geometric structure to avoid crossing the real boundary, and finally the result is cropped to the normalized interval to generate the information hiding mask.

[0087] Further, in the geometric principal direction neighborhood of the geometric optimization feature, a median filtering operation is performed to obtain a filtering result; the filtering result is taken as a low-frequency component, and the difference between the preliminary fusion feature and the low-frequency component is taken as a high-frequency component; the low-frequency component is regarded as a high semantic component, and the high-frequency component is regarded as a fine-grained component.

[0088] In the preliminary fusion feature, the fine-grained component and the high semantic component are split into a fine-grained feature channel and a high semantic feature channel; positive normalization is performed on each channel, and the channel readings are converted into a logarithmic domain.

[0089] The geometric principal direction is calculated based on the horizontal difference, vertical difference, unit normal three components and curvature strength of the depth image; a local neighborhood is taken at each spatial position to construct a gradient vector set composed of the horizontal difference and the vertical difference; the weighted covariance matrix is calculated by taking the curvature strength as the weight; the eigenvector corresponding to the minimum eigenvalue of the weighted covariance matrix is determined as the geometric principal direction, and the eigenvector orthogonal to the geometric principal direction is taken as the normal direction.

[0090] In the spatial position where the information hiding mask is small, the fine-grained component is retained as the dominant of the fine-grained feature; in the spatial position where the information hiding mask is large, the high semantic component is retained as the dominant of the high semantic feature.

[0091] In each spatial position, a task proxy variable is marked, which is used to determine whether the spatial position belongs to the real structure consistent across modalities; specifically, when the principal direction of the visible light image and the principal direction of the depth image remain consistent, the principal direction of the visible light image and the principal direction of the infrared image remain consistent, and the continuous validity weights of the three modal data are all indicative of validity and consistent time confidence, it is determined that the current spatial position belongs to the real structure consistent across modalities, and the task proxy variable is marked as positive; if the conditions are not met, it is determined that the current spatial position does not belong to the real structure consistent across modalities, and the task proxy variable is marked as negative.

[0092] In the fine-grained feature channel and the high semantic feature channel, according to the marking result of the task proxy variable, the spatial position is divided into cross-modality consistent samples and cross-modality inconsistent samples; the conditional distribution of the channel feature value is counted for the two types of samples respectively, and according to the difference between the conditional distributions of the two types of samples, the contribution of each channel in reducing uncertainty is calculated, which is defined as the information gain of the channel, and the information gain represents the action strength of the channel in distinguishing the cross-modality consistent structure and the cross-modality inconsistent structure.

[0093] In the space division of the information hiding mask, the spatial positions marked as the fidelity region are further calculated for mutual information between each channel and the task agent variable as the mutual information indicator.

[0094] After obtaining the information gain and the mutual information indicator, a median and mutual information double criterion is adopted to perform double screening on all fine-grained feature channels and high semantic feature channels; first, the information gain of the channel must not be lower than the median of the information gain of all channels; second, the mutual information indicator of the channel must not be lower than the median of the mutual information indicator of all channels; the fine-grained feature channels and the high semantic feature channels satisfying the two criteria are retained and spliced according to the original channel order to generate a task selected description.

[0095] S4, generating an initial matching set under the constraint of geometric correlation, obtaining residual values by using a closed-form geometric trial solution, and generating a cleaned matching set through spatial median robustness.

[0096] Specifically, generating an initial matching set under the constraint of geometric correlation includes taking the optimized features of each modality data and the preliminary fusion features as candidate carriers, finding the main peak position of each target position in the source position based on the cross-modal correlation graph and combined with the current temperature parameter adjustment in the pairing relationship between the target position and the source position, and combining the geometric bias matrix and the spatial scale constraint to screen the candidate matching points.

[0097] For each target position, only the paired relationship with the source position as the main peak is retained, and evidence information is recorded for each paired relationship of the target position and the source position, the evidence information including the peak value of the cross-modal correlation graph, the displacement reading of the geometric bias matrix, the reading of the prior information, and the reading of the information hiding mask, to generate an initial matching set.

[0098] Further, in the unified coordinate system, the texture color optimization feature, the geometric structure optimization feature, the thermal radiation optimization feature, and the preliminary fusion feature are jointly used as the matchable candidate carriers; based on the cross-modal correlation graph, in the pairing relationship between the target position and the source position, combined with the current temperature parameter adjustment, the main peak position of each target position on the source is read, and at the same time, the reference displacement vector provided by the geometric bias matrix is used, and combined with the numerical value of the spatial scale, when the distance between the coordinate difference of the source position and the target position and the reference displacement vector does not exceed the numerical value of the spatial scale, the source position and the target position form a candidate matching point; for each target position, only the paired relationship with the source position as the main peak is retained, wherein the definition of mutual main peak is that the responses of the two positions in the corresponding row or column are both maximum values or local peak values.

[0099] All the cross-modal candidate matching point pairs of mutual main peaks are recorded, each candidate matching point pair including a peak value of the cross-modal correlation graph, a displacement reading of the geometric bias matrix, a reading of the prior information, and a reading of the information hiding mask; all the candidate matching point pairs are recorded to form an initial matching set, which not only contains candidate positions jointly defined by the main peaks and geometric neighbors, but also carries evidence information derived from the time and modal reliability determination, for matching reliability evaluation.

[0100] Further, on the basis of the initial matching set, a geometric constraint is used to construct a closed geometric alignment trial solution, and a residual value of the candidate matching point is calculated to measure the projection error size of the candidate matching point after geometric alignment; when the geometric structure optimization feature or external information can provide reliable three-dimensional spatial structure, a three-dimensional spatial point set alignment method is used for similarity calculation; when the scene is closer to planar imaging, an alignment method under planar constraint is used for calculation; the result of the geometric alignment calculation is a three-dimensional spatial similarity transformation matrix or a two-dimensional homography matrix, which is used to predict the projection position of the candidate matching point.

[0101] The external information includes internal and external parameters of the imaging device, a known geometric reference surface of the scene, or three-dimensional structure information obtained by an external sensor.

[0102] For each candidate matching point pair in the initial matching set, a soft weight is assigned; the calculation of the soft weight is as follows: the peak value of the cross-modal correlation graph, the displacement reading of the geometric bias matrix, the reading of the prior information, and the reading of the information hiding mask are sequentially subjected to normalization processing, and the normalized values are respectively subjected to natural logarithm operation to obtain values in the logarithmic domain; then, the logarithmic domain values are added item by item to obtain a logarithmic value of the comprehensive weight, and the logarithmic value of the comprehensive weight is subjected to exponential operation to restore the soft weight of the candidate matching point.

[0103] After obtaining the soft weight, residual calculation is performed on each candidate matching point pair in the initial matching set. Specifically, first, based on the geometric structure optimization feature, the candidate matching point is back-projected into a three-dimensional point set in a unified coordinate system, and combined with the internal parameter matrix and the external parameter matrix of the imaging device, the three-dimensional point is projected back to the image plane to obtain a predicted pixel position; second, the predicted pixel position is compared with the true pixel position of the candidate matching point in the corresponding modality to calculate the pixel coordinate difference; when the three-dimensional point set alignment method is used, the difference is expressed as a three-dimensional Euclidean distance, and when the planar constraint alignment method is used, the difference is expressed as a two-dimensional projection deviation; the difference is the geometric projection residual or the re-projection residual of the candidate matching point.

[0104] After obtaining the residuals of all candidate matching points, all residual values are collected into the same distribution domain, and the residual values in the distribution domain are sorted in ascending order; the cumulative proportion of each residual value in the sorted sequence is calculated to obtain the corresponding empirical cumulative distribution function value; the empirical cumulative distribution function value is taken as the rank percentile value of the residual value, representing the relative position of the residual in the overall error distribution.

[0105] After obtaining the rank percentile value of the residual value, the initial matching set is re-weighted and regionally robustly processed. Specifically, the rank percentile value of the residual value of each pair of candidate matching points is converted into a continuous coefficient through monotone mapping and is limited within the normalized interval range; the converted continuous coefficient is sequentially combined with the peak value of the cross-modal correlation graph, the displacement reading of the geometric bias matrix, the reading of the prior information, and the reading of the information hiding mask to obtain the updated soft weight.

[0106] In the image plane, the rank percentile value of the residual value is spatially aggregated in the neighborhood median manner to form a position-by-position region median re-projection residual field; the region median re-projection residual field is combined with the cross-modal correlation graph, the geometric bias matrix, and the spatial scale to obtain the correction effect on the region scale; the convergence effect is presented at positions with larger region scale, and the relaxation effect is presented at positions with smaller region scale, thereby realizing the spatial robust correction of the matching residual.

[0107] After completing the consistent re-weighting and regional robust processing, the residual value of the candidate matching point, the continuous coefficient formed by the rank percentile value of the residual value, and the updated soft weight are uniformly arranged to form a new matching set; for the matching points that meet the conditions in the mutual consistency determination and the region median residual is in the stable range, the residual value, the continuous coefficient formed by the rank percentile value of the residual value, and the updated soft weight are retained and marked as high-reliability matching points; for the matching points that do not meet the conditions in the mutual consistency determination or the region median residual fluctuates too much, the residual value, the continuous coefficient formed by the rank percentile value of the residual value, and the updated soft weight are uniformly marked as low-reliability and arranged to the end of the weight sequence in the matching point set; the final cleaned matching set contains high-reliability matching points and low-reliability matching points, and each pair of matching points corresponds to the residual value, the continuous coefficient formed by the rank percentile value of the residual value, and the updated soft weight.

[0108] S5, a conformal transmission field is constructed by taking the median of the evidence rank percentile value to share the evidence graph, data sharing is completed, and the features of the optimized modal data are obtained, and the final fusion features are generated by median aggregation in the logarithmic domain.

[0109] Specifically, the final fusion features generated by median aggregation in the logarithmic domain include constructing a shared evidence graph by median statistics of the evidence rank percentile value.

[0110] The cross-modal correlation map, the multiplicative injection coefficient, the consistency index, and the reading of the prior information are combined to generate a shared intensity.

[0111] In the logarithmic domain, the effective components of the source position optimization feature are guided and filtered along the geometric direction of the target position and the boundary to obtain a conformal transfer field.

[0112] The shared evidence map is used as a gating intensity to compress and suppress the conformal transfer field, and is combined with an information hiding mask to obtain a shared update of the target position.

[0113] The shared update is performed on each modality data respectively to obtain a re-optimized feature, and the spatial positions in the logarithmic domain are taken as a median to restore to the original numerical value domain to generate a final fusion feature.

[0114] Further, in a unified coordinate system, a shared evidence map is constructed based on the correspondence between the target position and the source position; the generation of the shared evidence map relies on two types of information: the first type is content consistency evidence, which is derived from the content coupling degree between the target position and the source position represented by the cross-modal correlation map peak value and the multiplicative injection coefficient; the second type is prior reliability evidence, which is derived from the prior information; the two types of evidence are sorted and counted in a unified scale to obtain an evidence rank percentile value; the evidence rank percentile value is median counted to form a shared intensity as a reading of the shared evidence map.

[0115] In the coordinate system of the target position, the geometric main direction of the target position is calculated based on the geometric structure optimization feature, and the geometric main direction represents the main direction of the texture or the boundary; the geometric bias matrix is used to provide a reference displacement vector, and a neighborhood range is defined by the spatial scale to constrain the geometric difference between the source position and the target position; under the constraint condition, the effective components of the source position optimization feature are gradually transferred in the neighborhood range according to the geometric main direction of the target position to form a propagation path consistent in direction; in the transfer process, the geometric bias matrix limits the displacement deviation between the source position and the target position, and the spatial scale limits the effective range of propagation to ensure that the transfer process maintains local geometric consistency; finally, a numerical field continuously distributed in the geometric main direction of the target position is obtained, and the numerical field is truncated by the spatial scale constraint in the boundary region to avoid diffusion interference of the source position optimization feature across the real boundary, and the numerical field is defined as a conformal transfer field.

[0116] The shared evidence map and the conformal transfer field are combined to complete a small dose of shared update of the target position, specifically, first, the change amount transferred by the source position is calculated based on the numerical baseline of the target position; second, the shared intensity of the shared evidence map is used as a gating intensity to adjust the amplitude of the change amount in the spatial distribution; and finally, the shared update of the source position to the target position is formed.

[0117] The information hiding mask is hidden in the shared update as a synchronization constraint. In a region with a large information hiding mask value, a high semantic component is preserved as the dominant component. In a region with a small information hiding mask value, a fine-grained component is preserved as the dominant component.

[0118] Based on the shared update result, cross updates between the texture color optimization feature, the geometric structure optimization feature, and the thermal radiation optimization feature are sequentially completed to obtain re-optimized features of each modality data. After re-optimization is completed, the re-optimized features of each modality data are normalized by channel and are converted into a logarithmic domain. In the logarithmic domain, median aggregation is performed by spatial position, and the result is restored to the original value domain to generate a final fusion feature.

[0119] Further, under the boundary condition, when the source position and the target position are obviously inconsistent in content consistency or prior reliability, the value of the shared evidence graph tends to zero, and the sharing process tends to stagnate. When the cross-modality time sequence is out of sync and causes a value mutation of the source position, the gating of the shared evidence graph performs automatic compression on the abnormal difference, and combines the shielding mechanism of the information hiding mask to avoid the spread of unstable information. When the cross-modality structure is consistent and the reliability reading is close to a stable value, the gating of the shared evidence graph promotes the shared update in a small dose to prevent excessive spread.

[0120] S6, splicing and unitizing the final fusion feature and the task selected description to generate a location-by-location task feature, and generating a registration coordinate transformation matrix in a closed form method according to the cleaned matching set.

[0121] Specifically, generating the location-by-location task feature includes calculating a scene pass factor based on the cleaned matching set, and using the scene pass factor as a multiplicative gating factor to normalize and inverse hyperbolic sine transform the final fusion feature and the decoupled task selected description by channel, and performing empirical rank statistics using a spatial scale to define a neighborhood to obtain local ranks of each channel.

[0122] Scaling and amplitude adjusting the local ranks of each channel by using the multiplicative gating factor as a regulation coefficient to obtain feature components after gating and regulation, and applying an information hiding mask constraint to the feature components after gating and regulation to form task components, and splicing and unitizing all the task components by channel to generate the location-by-location task feature.

[0123] Further, in a unified coordinate system, the final fusion feature and the task selected description are spliced by channel. When splicing, the final fusion feature is arranged first, and then the fine-grained feature channels and the high semantic feature channels in the task selected description are arranged in the order of modality categories, so that the splicing order of each spatial position is uniquely determined. After splicing is completed, the spliced result is unitized as the initial input for generating the task feature.

[0124] Based on the matched set after cleaning, the scene passing degree factor is calculated. Specifically, the residual values of all candidate matching points in the matched set after cleaning are normalized, converted into rank percentile values of the residual values, and the rank percentile values of all residual values are averaged to obtain the scene passing degree factor. The scene passing degree factor, together with the main peak value of the cross-modal correlation graph and the multiplicative injection coefficient, generates a position-by-position multiplicative gating factor.

[0125] On the initial input of the task feature generation, the positive value and the inverse hyperbolic sine transformation are performed channel by channel. Based on the spatial scale, the empirical rank statistics is performed on each channel in the local neighborhood to obtain the local rank of each channel. The multiplicative gating factor is used as an adjustment coefficient to adjust the numerical range of the local rank of each channel, and the local rank is proportionally scaled and amplitude adjusted to obtain the gated and adjusted feature component.

[0126] Information hiding mask constraint is used on the gated and adjusted feature component. In the area with a larger information hiding mask value, high semantic components are mainly retained. In the area with a smaller information hiding mask value, fine-grained components are mainly retained. After the information hiding mask constraint, all gated and adjusted feature components are reassembled and unitized to generate a position-by-position task feature.

[0127] Based on the position-by-position task feature, the matched set after cleaning is combined to calculate the registration coordinate transformation matrix using a closed-form method. When the matched set after cleaning contains three-dimensional spatial structure constraints, a three-dimensional point set alignment method is used to generate a three-dimensional similarity transformation matrix. When the matched set after cleaning meets the plane imaging condition, a direct linear transformation method is used to generate a two-dimensional homography matrix.

[0128] The registration coordinate transformation matrix is used to unify the coordinates of the visible light image, the depth image and the infrared image to the same geometric reference system, ensuring the spatial consistency of the multi-modal data.

[0129] The embodiment also provides an image feature matching and fusion system based on a multi-modal large model, comprising:

[0130] The acquisition processing module is configured to acquire different modal data and internal and external parameters of the imaging device, extract initial features of the modal data, and generate prior information.

[0131] The spatial alignment module is configured to construct a geometric bias matrix and a spatial scale, calculate a cross-modal correlation graph and a multiplicative injection coefficient based on the initial features of the modal data and the prior information, obtain optimized features of the modal data, and generate preliminary fusion features.

[0132] A hidden decoupling module is configured to construct an information hiding mask on the preliminary fused feature, decouple along the geometric main direction of the geometric structure, and use a double-criterion channel selection to obtain a task-selected description.

[0133] A matching feedback module is configured to generate an initial matching set under geometric correlation constraints, obtain a residual value using a closed-form geometric trial solution, and generate a cleaned matching set after spatial median robustness.

[0134] A shared fusion module is configured to construct a shared evidence graph and a conformal transmission field, complete data sharing, obtain features of each modality data optimized again, and generate a final fused feature in a logarithmic domain.

[0135] A task solving module is configured to concatenate and unitize the final fused feature and the task-selected description to generate a position-by-position task feature, and generate a registration coordinate transformation matrix in a closed-form method according to the cleaned matching set.

[0136] In summary, the present application achieves directional alignment and spatial consistency of multi-modal data in a unified coordinate system by introducing geometric guidance calculation and spatial constraint modeling, avoids offset of different modal data in geometric structure, and guarantees stability and reliability of the fusion process. The present application achieves conformal transmission and abnormal suppression of multi-modal data in the geometric main direction by constructing a shared evidence graph and combining median statistics and gate regulation.

[0137] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit it. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the present application, which should be covered in the scope of the claims of the present application.

Claims

1. An image feature matching and fusion method based on a multimodal large model, characterized in that: include, Acquire different modal data and the intrinsic and extrinsic parameters of imaging devices, extract the initial features of each modal data, and generate prior information; In a unified coordinate system, a geometric bias matrix and spatial scale are constructed. Based on the initial features and prior information of each modal data, cross-modal correlation maps and multiplicative injection coefficients are calculated. After updating, the optimized features of each modal data are obtained, and the preliminary fusion features are generated by normalization. An information hiding mask is constructed on the initial fusion features, and the geometric main direction of the optimized features is decoupled. A dual-criteria filtering channel is used to obtain the task's selected description. An initial matching set is generated under geometric correlation constraints. The residual value is obtained by closed geometric trial solution. After spatial median robustness, a cleaned matching set is generated. A conformal transmission field is constructed by taking the median of the shared evidence graph with the rank percentile values ​​of the evidence, data sharing is completed, features of each modality data are obtained and optimized again, and the final fusion features are generated by aggregation in the logarithmic domain. The final fused features are concatenated with the selected task descriptions and normalized to generate position-wise task features. The registration coordinate transformation matrix is ​​then generated using a closed-form method based on the cleaned matching set. The normalization generation of preliminary fusion features includes reading the initial features and prior information of each modal data in a unified coordinate system, combining the geometric bias matrix, spatial scale and current temperature to calculate the cross-modal correlation map and establish an evidence channel for the source conversion target. By comparing the consistency between the source direction and the target's own direction, the multiplicative injection coefficient is obtained. Small-dose multiplicative updates are performed on each modality data to obtain the optimized features of each modality data. The features are then aggregated in the logarithmic domain by taking the median position one by one, and then restored back to the original data domain to obtain the preliminary fusion features. The task selection description includes performing multi-scale analysis and normalization processing on the initial fusion features, obtaining weighted redundancy by combining prior information, and marking spatial locations as areas to be shielded and areas to be preserved based on the relationship between weighted redundancy and original redundancy, thereby generating an information hiding mask. Based on the geometric principal direction of the optimized features of the geometric structure, median filtering decomposition is performed in the neighborhood to split the preliminary fused features into fine-grained feature channels and high semantic feature channels. Statistical conditional distributions are calculated on cross-modal consistent and inconsistent samples. Based on the differences in conditional distributions, the information gain of each channel is calculated, and the mutual information index between each channel and the task proxy variable is calculated. Using the median of information gain and the median of mutual information index as dual criteria, the selected fine-grained feature channels and high semantic feature channels are concatenated in channel order to generate a carefully selected task description. The process of generating the cleaned matching set includes selecting a closed geometric trial solution method under the scenario conditions, calculating the residual values ​​of the candidate matching points, mapping the residual values ​​to rank percentile values, and performing numerical combination calculations with the peak value of the cross-modal correlation map, the displacement reading of the geometric bias matrix, the reading of the prior information, and the reading of the information hiding mask to obtain updated soft weights. The spatial median robustness includes performing neighborhood median aggregation on the residual values ​​to form a regional median reprojection residual field, and adjusting and updating the soft weights according to the regional residual information to generate a cleaned matching set. The process of generating the final fusion feature by median aggregation in the logarithmic domain includes constructing a shared evidence graph by median statistics of the rank percentile values ​​of evidence, and writing the shared strength generated by combining the cross-modal correlation graph, multiplicative injection coefficient, consistency index and prior information readings into the shared evidence graph. In the logarithmic domain, the effective components of the source location optimization features are guided and filtered along the geometric direction and boundary of the target location to obtain a conformal transmission field. Using the shared evidence map as the gating strength, the conformal transmission field is compressed and suppressed, and linked with the information hiding mask to obtain the shared update of the target position; Shared updates are performed on each modality of data to obtain further optimized features. The features are then restored to the original numerical domain by taking the median of each spatial position in the logarithmic domain, generating the final fused features.

2. The image feature matching and fusion method based on a multimodal large model as described in claim 1, characterized in that: The initial features for extracting each modal data include receiving modal data from visible light images, depth images, and infrared images, and using a publicly available convolutional neural network to extract the initial texture and color features of the visible light image and the initial thermal radiation features of the infrared image, respectively. The initial geometric features of the depth image are obtained by calculating the horizontal difference, vertical difference, unit normal three components, and curvature intensity of the depth image.

3. The image feature matching and fusion method based on a multimodal large model as described in claim 2, characterized in that: The generation of prior information includes estimating the average pixel motion between the current frame and adjacent frames, combining the inter-frame time difference and processing it through a smoothing function to obtain the time confidence. Each modal data generates a continuous validity weight based on its own observations, which, together with the time confidence level, constitutes prior information.

4. The image feature matching and fusion method based on a multimodal large model as described in claim 3, characterized in that: The process of generating an initial matching set under geometric correlation constraints includes using the optimized features and preliminary fusion features of each modal data as candidate carriers, determining the main peak position based on the cross-modal correlation map and the current temperature parameter adjustment in the pairing relationship between the target position and the source position, and screening candidate matching points in combination with the geometric bias matrix and spatial scale constraints. For each target location, only the pairing relationships with the source location that are mutually dominant are retained. For each pair of target locations and source locations, the peak value of the cross-modal correlation map, the displacement reading of the geometric bias matrix, the reading of the prior information, and the reading of the information hiding mask are recorded to generate an initial matching set.

5. The image feature matching and fusion method based on a multimodal large model as described in claim 4, characterized in that: The generated position-by-position task features include: The scene passability factor is calculated based on the cleaned matching set, and the scene passability factor is used as a multiplicative gating factor to select and describe the final fusion features and decoupling tasks. Positive value transformation and inverse hyperbolic sine transform are performed on each channel, and empirical rank statistics are performed on the neighborhood defined by spatial scale to obtain the local rank of each channel. Using a multiplicative gating factor as an adjustment coefficient, the local rank of each channel is scaled and the amplitude is adjusted to obtain the gated feature components. An information hiding mask constraint is applied to the gated feature components to form task components. All task components are spliced ​​together by channel and normalized to generate position-by-position task features.

6. An image feature matching and fusion system based on a multimodal large model, based on the image feature matching and fusion method based on a multimodal large model as described in any one of claims 1 to 5, characterized in that: include, The acquisition and processing module is used to acquire data from different modalities and the intrinsic and extrinsic parameters of the imaging device, extract the initial features of each modal data, and generate prior information. The spatial alignment module is used to construct the geometric bias matrix and spatial scale. Based on the initial features and prior information of each modality data, it calculates the cross-modal correlation map and multiplicative injection coefficients to obtain the optimized features of each modality data and generate preliminary fusion features. The hidden decoupling module is used to construct an information hiding mask on the initial fused features, decouple along the geometric main direction of the optimized features, and use dual-criteria filtering channels to obtain the task's selected description. The matching feedback module is used to generate an initial matching set under geometric correlation constraints, obtain residual values ​​using closed geometric trial solutions, and generate a cleaned matching set after spatial median robustness. The shared fusion module is used to construct a shared evidence map and a conformal transmission field, complete data sharing, obtain features of each modality data for further optimization, and aggregate them in the logarithmic domain to generate the final fusion features. The task solution module is used to concatenate the final fused features with the selected task description and normalize them to generate position-by-position task features, and generate the registration coordinate transformation matrix using a closed-form method based on the cleaned matching set.

Citation Information

Patent Citations

  • Self-adaptive alignment cross-modal vision-language ship intelligent man-machine interaction method

    CN119357897A

  • Multi-modal image registration method based on self-modal correlation and cross-modal estimation

    CN119762559A