Image feature matching and fusion method and system based on multi-modal large model
By employing a multimodal large-scale model image feature matching and fusion method, the problems of cross-modal noise spillover and detail drift are solved, achieving stable and reliable fusion of multimodal data and ensuring the accuracy and consistency of the image processing process.
Patent Information
- Application Number
- CN202511452778.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-10-13
AI Technical Summary
Existing multimodal image processing techniques lack continuous self-restraint and regional constraints in cross-modal information propagation, leading to noise spillover and detail drift. Insufficient handling of scale and dimensional correlation between modalities also affects metric consistency and registration stability.
By using an image feature matching and fusion method based on a multimodal large model, initial features of different modal data are obtained, prior information is generated, a geometric bias matrix and spatial scale are constructed in a unified coordinate system, cross-modal correlation maps are calculated and multiplicative injection intensity is updated, an information hiding mask is constructed for decoupling, a cleaned matching set is generated, and the final fused features are achieved through shared evidence maps and gating adjustment.
It achieves directional alignment and spatial consistency of multimodal data in a unified coordinate system, avoids the offset of modal data in geometric structure, ensures the stability and reliability of the fusion process, and realizes conformal transmission and anomaly suppression of multimodal data in the main geometric direction.
Smart Images

Figure CN120932053A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal information processing technology, and in particular to an image feature matching and fusion method and system based on a multimodal large model. Background Technology
[0002] Multimodal large models are increasingly used in image processing, simultaneously incorporating visible light, depth, and infrared images to fuse multi-dimensional information such as texture, geometric, and radiometric features. Existing technologies establish intermodal correspondences through correlation metrics and geometric constraints, and combine logarithmic domain modeling with robust statistical methods to achieve scale uniformity and feature alignment, thereby supporting applications such as image retrieval, image ranking, and geometric registration.
[0003] In engineering practice, conventional processes still face two main challenges: First, the propagation of cross-modal information across structural boundaries and time-varying regions lacks continuous self-restraint and regional constraints, easily leading to noise spillover and detail drift. Second, insufficient handling of scale and dimensional correlations between modes affects metric consistency and registration stability. Existing multimodal matching and fusion technologies face engineering challenges in terms of boundary fidelity and time-varying robustness. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides an image feature matching and fusion method based on a multimodal large model to solve the problems of cross-modal noise spillover, detail drift, and insufficient scale and dimension correlation processing in image processing.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides an image feature matching and fusion method based on a multimodal large model, which includes acquiring different modal data and the intrinsic and extrinsic parameters of the imaging device, extracting the initial features of each modal data, and generating prior information; A geometric bias matrix and spatial scale are constructed in a unified coordinate system. Based on the initial features and prior information of each modal data, the cross-modal correlation map and multiplicative injection intensity are calculated. After updating, the optimized features of each modal data are obtained, and the preliminary fusion features are generated by normalization. An information hiding mask is constructed on the initial fusion features, and the geometric main direction of the optimized features is decoupled. A dual-criteria filtering channel is used to obtain the task's selected description. An initial matching set is generated under geometric correlation constraints. The residual value is obtained by closed geometric trial solution. After spatial median robustness, a cleaned matching set is generated. A conformal transmission field is constructed by taking the median of the shared evidence graph with the rank percentile values of the evidence, data sharing is completed, features of each modality data are obtained and optimized again, and the final fusion features are generated by aggregation in the logarithmic domain. The final fused features are concatenated with the selected task descriptions and normalized to generate position-wise task features. The registration coordinate transformation matrix is then generated using a closed-form method based on the cleaned matching set.
[0007] As a preferred embodiment of the image feature matching and fusion method based on a multimodal large model described in this invention, the initial features of each modality data extraction include receiving modality data of visible light images, depth images and infrared images, and using a publicly available convolutional neural network to extract the initial texture and color features of the visible light image and the initial thermal radiation features of the infrared image, respectively. The initial geometric features of the depth image are obtained by calculating the horizontal difference, vertical difference, unit normal three components, and curvature intensity of the depth image.
[0008] As a preferred embodiment of the image feature matching and fusion method based on a multimodal large model described in this invention, the generation of prior information includes estimating the average pixel motion between the current frame and adjacent frames, combining the inter-frame time difference and processing it through a smoothing function to obtain the time confidence. Each modal data generates a continuous validity weight based on its own observations, which, together with the time confidence level, constitutes prior information.
[0009] As a preferred embodiment of the image feature matching and fusion method based on a multimodal large model described in this invention, the normalization generation of preliminary fusion features includes: reading the initial features and prior information of each modality data in a unified coordinate system, combining the geometric bias matrix, spatial scale and current temperature to calculate a cross-modal correlation map and establish an evidence channel for the source-transformation target; By comparing the consistency between the source direction and the target's own direction, the multiplicative injection coefficient is obtained. Small-dose multiplicative updates are performed on each modality data to obtain the optimized features of each modality data. The features are then aggregated in the logarithmic domain by taking the median position one by one, and then restored back to the original data domain to obtain the preliminary fusion features.
[0010] As a preferred embodiment of the image feature matching and fusion method based on a multimodal large model described in this invention, the task selection description includes: performing multi-scale analysis and normalization processing on the preliminary fused features, obtaining weighted redundancy by combining prior information, and marking the spatial location as the area to be shielded and the area to be preserved according to the relationship between the weighted redundancy and the original redundancy, thereby generating an information hiding mask; Based on the geometric principal direction of the optimized features of the geometric structure, median filtering decomposition is performed in the neighborhood to split the preliminary fused features into fine-grained feature channels and high semantic feature channels. Statistical conditional distributions are calculated on cross-modal consistent and inconsistent samples. Based on the differences in conditional distributions, the information gain of each channel is calculated, and the mutual information index between each channel and the task proxy variable is calculated. Using the median of information gain and the median of mutual information as dual criteria, the selected fine-grained feature channels and high semantic feature channels are concatenated in channel order to generate a carefully selected task description.
[0011] As a preferred embodiment of the image feature matching and fusion method based on multimodal large model described in this invention, wherein: the generation of the initial matching set under geometric correlation constraints includes taking the optimized features of each modality data and the preliminary fused features as candidate carriers, determining the main peak position in the pairing relationship between the target position and the source position based on the cross-modal correlation map and the current temperature parameter adjustment, and screening candidate matching points in combination with the geometric bias matrix and spatial scale constraints; For each target location, only the pairing relationships with the source location that are mutually dominant are retained. For each pair of target locations and source locations, the peak value of the cross-modal correlation map, the displacement reading of the geometric bias matrix, the reading of the prior information, and the reading of the information hiding mask are recorded to generate an initial matching set.
[0012] As a preferred embodiment of the image feature matching and fusion method based on multimodal large model described in this invention, the generation of the cleaned matching set includes: selecting a closed geometric trial solution method under scene conditions, calculating the residual value of the candidate matching point, mapping the residual value to the rank percentile value, and performing numerical combination calculation with the peak value of the cross-modal correlation map, the displacement reading of the geometric bias matrix, the reading of the prior information, and the reading of the information hiding mask to obtain the updated soft weight; The residual values are aggregated in the neighborhood to form a regional median reprojection residual field. The soft weights are then adjusted and updated based on the regional residual information to generate a cleaned matching set.
[0013] As a preferred embodiment of the image feature matching and fusion method based on a multimodal large model described in this invention, the step of generating the final fused feature by median aggregation in the logarithmic domain includes constructing a shared evidence map by median statistics of the rank percentile values of evidence, and writing the shared strength generated by combining the cross-modal correlation map, multiplicative injection coefficient, consistency index and prior information readings into the shared evidence map. In the logarithmic domain, the effective components of the source location optimization features are guided and filtered along the geometric direction and boundary of the target location to obtain a conformal transmission field. Using the shared evidence map as the gating strength, the conformal transmission field is compressed and suppressed, and linked with the information hiding mask to obtain the shared update of the target position; Shared updates are performed on each modality of data to obtain further optimized features. The features are then restored to the original numerical domain by taking the median of each spatial position in the logarithmic domain, generating the final fused features.
[0014] As a preferred embodiment of the image feature matching and fusion method based on a multimodal large model described in this invention, the generation of position-wise task features includes: calculating the scene passability factor based on the cleaned matching set, using the scene passability factor as a multiplicative gating factor, carefully describing the final fused features and decoupled tasks, performing positiveization and inverse hyperbolic sine transform on each channel, and performing empirical rank statistics on the neighborhood defined by the spatial scale to obtain the local rank of each channel; Using a multiplicative gating factor as an adjustment coefficient, the local rank of each channel is scaled and the amplitude is adjusted to obtain the gated feature components. An information hiding mask constraint is applied to the gated feature components to form task components. All task components are spliced together by channel and normalized to generate position-by-position task features.
[0015] Secondly, this invention provides an image feature matching and fusion system based on a multimodal large model, comprising: The acquisition and processing module is used to acquire data from different modalities and the intrinsic and extrinsic parameters of the imaging device, extract the initial features of each modal data, and generate prior information. The spatial alignment module is used to construct the geometric bias matrix and spatial scale. Based on the initial features and prior information of each modality data, it calculates the cross-modal correlation map and multiplicative injection intensity, obtains the optimized features of each modality data, and generates preliminary fusion features. The hidden decoupling module is used to construct an information hiding mask on the initial fused features, decouple along the geometric main direction of the optimized features, and use dual-criteria filtering channels to obtain the task's selected description. The matching feedback module is used to generate an initial matching set under geometric correlation constraints, obtain residual values using closed geometric trial solutions, and generate a cleaned matching set after spatial median robustness. The shared fusion module is used to construct a shared evidence map and a conformal transmission field, complete data sharing, obtain features of each modality data for further optimization, and aggregate them in the logarithmic domain to generate the final fusion features. The task solution module is used to concatenate the final fused features with the selected task description and normalize them to generate position-by-position task features, and generate the registration coordinate transformation matrix using a closed-form method based on the cleaned matching set.
[0016] The beneficial effects of this invention are as follows: by introducing geometrically guided computation and spatial constraint modeling, the orientation alignment and spatial consistency of multimodal data in a unified coordinate system are achieved, avoiding the offset of different modal data in geometric structure and ensuring the stability and reliability of the fusion process; by constructing a shared evidence graph and combining median statistics and gating adjustment, the conformal transmission and anomaly suppression of multimodal data in the geometric principal direction are achieved. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of an image feature matching and fusion method based on a multimodal large model.
[0019] Figure 2 This is a schematic diagram of an image feature matching and fusion system based on a multimodal large model.
[0020] Figure 3 A flowchart for selecting and describing the task to be generated.
[0021] Figure 4 The flowchart for generating the cleaned matching set. Detailed Implementation
[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0025] Reference Figures 1-4 This is one embodiment of the present invention, which provides an image feature matching and fusion method based on a multimodal large model, including the following steps: S1. Acquire different modal data and the intrinsic and extrinsic parameters of the imaging device, extract the initial features of each modal data, and generate prior information.
[0026] Furthermore, it receives three modal data streams: visible light image, depth image, and infrared image, and saves the intrinsic parameter matrix and extrinsic parameter matrix of each of the three imaging devices.
[0027] Based on the depth image, the pixel positions and their corresponding depth values on the depth image are back-projected into three-dimensional points. The three-dimensional points are then converted into a unified coordinate system using the extrinsic parameter matrix of the depth imaging device. Then, the three-dimensional points in the unified coordinate system are projected onto the imaging plane of the visible light image and the infrared image using the extrinsic parameter matrices of the visible light imaging device and the infrared imaging device, thus obtaining the pixel positions corresponding to the positions in the depth image. After processing, a position-by-position correspondence is established between the three modal data in geometric space.
[0028] Based on the positional correspondence, the aligned visible light image, infrared image, and depth image, along with their positional validity masks and corresponding indices, are divided into training and validation sets. Distortion correction, color correction, and gain correction are performed on the visible light image. Black level correction and gain correction are performed on the infrared image, and the results are converted into physical radiation quantities according to the calibration curve. Samples are generated according to a uniform resolution and cropping rules, and the positional validity masks and corresponding indices are output synchronously.
[0029] Extracting initial features for each modality of data includes using publicly available convolutional neural networks (such as ResNet-50 and MobileNetV2) to extract initial texture and color features for visible light images and initial thermal radiation features for infrared images, respectively.
[0030] The initial geometric features of the depth image are obtained by calculating the horizontal difference, vertical difference, unit normal three components, and curvature intensity of the depth image.
[0031] Specifically, on the depth image, the difference operator is used to calculate the horizontal and vertical differences of the depth image respectively, and the three-dimensional point set obtained by backprojection of the pixel neighborhood is approximated as the tangent plane. Based on the geometric relationship of the tangent plane, the three components of the unit normal are obtained. At the same time, the discrete Laplacian operator is used to approximate the curvature intensity on the depth image. The three components of the unit normal, the curvature intensity, and the depth validity weight are stacked in channel order to obtain the initial features of the geometric structure of the depth image.
[0032] A visible light image feature extraction network (based on ResNet-50) and an infrared image feature extraction network (based on MobileNetV2) are set up. In the visible light image feature extraction network, a publicly available pre-trained ResNet-50 is loaded, the classification layer is removed, and only the mid-to-high-level convolutions are retained. The mid-to-high-level convolution results of the visible light image feature extraction network are sequentially passed through point convolutional layers, normalization layers, and non-linear activation layers to map a dense representation with a fixed number of channels, which is defined as the initial texture color feature. In the infrared image feature extraction network, a publicly available pre-trained MobileNetV2 is loaded, and the first layer... The configuration is changed to single-channel or adapted using channel replication. In the backbone of the infrared image feature extraction network, the last resolution reduction operation is canceled, and the convolution here is set to not change the spatial resolution. In the subsequent continuous convolutions, equally spaced holes are set to maintain the feature map resolution while maintaining the receptive field, so that the convolution result obtained by the infrared image feature extraction network is the same as that of the visible light image feature extraction network in terms of spatial sampling interval. The convolution result of the backbone of the infrared image feature extraction network is sequentially passed through point convolutional layers, normalization layers, and nonlinear activation mapping to a dense representation with a fixed number of channels, which is defined as the initial feature of thermal radiation.
[0033] For visible light image feature extraction networks and infrared image feature extraction networks, combining positional correspondence and positional validity masking, the training objectives are defined as follows: Intramodal consistency constraints are applied to the feature maps of the visible light and infrared image feature extraction networks. Two lightweight augmented samples are generated for the same original image. Using positional correspondence indexing, feature vectors at the same spatial position in the feature maps of the two networks are used as positive samples, and feature vectors at other positions are used as control pairs. Similarity is used as a metric to reduce the distance between positive sample vectors while suppressing the similarity of control pairs. Cross-modal alignment constraints are applied to the feature maps of the visible light and infrared image feature extraction networks. Under a unified coordinate system, feature vectors at the same spatial position in the feature maps of the two networks are paired, and the distance between paired vectors is minimized to achieve cross-modal alignment. Boundary consistency regularization is applied to the feature maps of the visible light image feature extraction network and the infrared image feature extraction network. Based on the unit normal three components and curvature intensity calculated from the depth image, finite differences are performed on the feature maps of the two feature extraction networks along the normal and tangential directions at each spatial location. A direction consistency constraint is applied to the normal difference, and a smoothing constraint is applied to the tangential difference. A thermal radiation monotonic constraint is applied to the feature map of the infrared image feature extraction network, requiring that the scalar reading corresponding to the location with higher radiation is not lower than that of the location with lower radiation. Relative importance weights are set for the intramodal consistency constraint, cross-modal alignment constraint, boundary consistency regularization, and thermal radiation monotonic constraint. The four losses are summed proportionally according to their relative importance to form the overall loss, and the contribution of missing, occluded, and mismatched locations is masked by a positional validity mask. The overall loss is used to update the trainable parameters of the visible light image feature extraction network and the infrared image feature extraction network.
[0034] After the overall loss is determined, the parameters are updated in the following order: first, only the point convolutional layers, normalization layers, and non-linear activation layers are updated, while the remaining convolutional layers of the two feature extraction networks remain unchanged; then, the convolutional layers located in the middle layers are unlocked and updated jointly with the aforementioned three types of layers; finally, the convolutional layers near the output are unlocked to complete the fine-tuning of the entire network; the batch data is organized by a positional correspondence, and two lightweight enhanced views are generated for each sample. The enhancement is limited to a monotonic mapping of brightness and contrast, limited translation and rotation, and random cropping, and transformations that violate the positional correspondence are prohibited.
[0035] The optimization employs an adaptive method with momentum, where the learning rate decreases progressively with the unlocking range. A larger step size is used only for updating the three classes of layers, decreasing the step size after unlocking the intermediate layers and further decreasing it after unlocking the higher layers. During training, cross-modal alignment error and intra-modal consistency indices are calculated on the validation set at fixed intervals. When the magnitude of change in both indices continuously decreases and stabilizes in continuous evaluation, and no longer shows a downward trend, convergence is determined and updates are stopped; otherwise, iteration continues. After training, all parameters of the visible light image feature extraction network and the infrared image feature extraction network, as well as the configurations of point convolutional layers, normalization layers, and nonlinear activation layers, are saved to generate initial texture color features and initial thermal radiation features.
[0036] During the inference phase, the input visible light image and infrared image, which have been geometrically and radiometrically corrected and uniformly sized, are used as the initial features for each modality data. The two sets of feature extraction networks output the initial features of texture color and thermal radiation, respectively. Together with the initial features of geometric structure calculated from the depth image, they serve as the initial features for each modality data.
[0037] Furthermore, generating prior information includes estimating the average pixel motion between the current frame and adjacent frames, combining the inter-frame time difference and processing it through a smoothing function to obtain temporal confidence.
[0038] Each modal data generates continuous validity weights for visible light images, depth images, and infrared images based on its respective observations, and these weights, together with the time confidence level, constitute prior information.
[0039] Within a unified coordinate system, radiometric uniformity and adaptive denoising are integrated into three modal data streams: visible light, depth, and infrared images. Specifically, black level correction, gain correction, and vignetting correction are performed on visible light and infrared images. On the infrared image, the original count values are converted into physical radiance values according to the calibration curve. The empirical cumulative distribution of the current frame is monotonically mapped to the reference distribution through fixed quantile anchor points to achieve uniformity between luminance and radiance apertures. For the depth image, no radiometric calibration is performed; only amplitude aperture is unified, and hole filling and anomalous flying point correction are used. On the three modal data streams of visible light, depth, and infrared images, a unified processing approach is used to perform lightweight denoising with edge preservation, distribution self-calibration, and temporal consistency. The processing results show a smooth effect in flat areas and maintain structural fidelity at geometric and thermal boundaries.
[0040] Between the current frame and adjacent frames, the motion amplitude of all valid pixel positions is calculated, and the motion amplitude is averaged to obtain the average pixel motion. Combined with the inter-frame time difference, the time confidence is obtained through smoothing function mapping. The continuous validity weight of visible light image is obtained by combining the robust normalization result of gradient amplitude with the monotonic mapping of brightness quantile coordinates. The continuous validity weight of depth image is calculated by self-calibrated smoothing mapping of local depth fluctuation and obtained by applying a linear penalty to the missing measurement ratio. The continuous validity weight of infrared image is obtained by mapping infrared pixel values to the upper tail coordinates of empirical distribution and applying smoothing clipping suppression.
[0041] Within a unified coordinate system, the initial features of the visible light image, depth image, and infrared image are processed using bicubic interpolation to achieve resolution uniformity, adjusting the initial features of the three modal data to the same resolution. For the continuous validity weights of the visible light image, depth image, and infrared image, the size is uniformized using bilinear interpolation, synchronizing the continuous validity weights of the three modal data to the same size.
[0042] S2. Construct a geometric bias matrix and spatial scale in a unified coordinate system. Based on the initial features and prior information of each modal data, calculate the cross-modal correlation map and multiplicative injection intensity. After updating, obtain the optimized features of each modal data and normalize to generate preliminary fusion features.
[0043] Specifically, the normalization generation of preliminary fusion features includes reading the initial features and prior information of each modality data in a unified coordinate system, combining the geometric bias matrix, spatial scale and current temperature to calculate the cross-modal correlation map and establish an evidence channel for the source transformation target.
[0044] By comparing the consistency between the source direction and the target's own direction, the multiplicative injection coefficient is obtained. Small-dose multiplicative updates are performed on each modality data to obtain the optimized features of each modality data. The features are then aggregated in the logarithmic domain by taking the median position one by one, and then restored back to the original data domain to obtain the preliminary fusion features.
[0045] Furthermore, under a unified coordinate system, a position-by-position correspondence table is first defined to record the pixel coordinates and continuous validity weights of the visible light image, depth image, and infrared image on their respective imaging planes for each spatial index. When any spatial index is selected and the target mode and source mode are specified, the target position is taken from the target mode pixel coordinates in the correspondence table, and the source position is taken from the source mode pixel coordinates in the correspondence table. Only when both the target position and the source position are within the effective imaging range and the continuous validity weights of the two locations are greater than zero are they considered a pair of valid target and source positions.
[0046] Two types of guidance information are established for each pair of valid target and source locations. The first type is geometric guidance, which forms a geometric bias matrix through the spatial coordinate difference between the target and source locations. The second type is reliability and time guidance, which combines the continuous validity weights of the visible light image, the continuous validity weights of the depth image, the continuous validity weights of the infrared image, and the time confidence at the source location to form a continuous reliability reading index as a spatial scale.
[0047] Based on the geometric bias matrix and spatial scale, a search neighborhood of the source location is determined within the coordinate system of the target location. Amplitude normalization is performed on the initial features of the target location and the initial features of each source location within the search neighborhood, and directional consistency readings are calculated as content similarity. The content similarity and the weights of the geometric bias matrix at the corresponding source locations are monotonically combined to obtain the original correlation readings. A normalized mapping controlled by a temperature adjustment coefficient is applied to the original correlation readings within the search neighborhood to form a cross-modal similarity distribution.
[0048] The temperature adjustment coefficient at each target location is jointly determined based on spatial scale, temporal confidence, continuous validity weights of visible light images, continuous validity weights of depth images, continuous validity weights of infrared images, information hiding mask values, and the geometric principal direction intensity of geometric structure optimization features. The temperature adjustment coefficient is increased when the spatial scale is large, the temporal confidence is low, any continuous validity weight is low, the information hiding mask is marked as needing to be masked, the texture is weak, or there are missing measurements. Conversely, the temperature adjustment coefficient is decreased when the spatial scale is small, the temporal confidence is high, the three continuous validity weights are reliable, the information hiding mask is marked as needing to be preserved, or the location is at a boundary or high gradient position. An example range of values for the temperature adjustment coefficient is provided. The temperature adjustment factor is set to 0.5 to 3.0, the normal operating range is set to 0.6 to 2.0, and the baseline is 1.0. The reasons for this selection are as follows: First, it ensures that the normalized mapping remains numerically stable within the similarity reading range and avoids gradient anomalies. Second, when the temperature adjustment factor is below about 0.5, the distribution becomes too sharp, and noise can easily trigger peak abrupt changes, which is not conducive to candidate coverage of weak texture regions. Third, when the temperature adjustment factor is above about 3.0, the distribution becomes too flat, making it difficult to distinguish between content differences and geometric biases. Fourth, using 1.0 as the baseline facilitates alignment with the normalization caliber without temperature adjustment, converges to 0.6 to 0.9 to highlight the main peak when reliability is improved or the boundary is strengthened, and expands to 1.2 to 2.0 to increase candidate coverage when reliability decreases or the texture is sparse.
[0049] By superimposing the geometric bias matrix with the spatial scale, a smooth offset control is formed, and normalization is performed to generate a cross-modal correlation map. The cross-modal correlation map is used to explain how the target location aggregates evidence on the source information. The geometric bias matrix determines the sensitivity of the correlation to spatial distance, the temperature parameter adjusts the concentration or dispersion of similarity distribution, and the spatial scale sets the range of the geometric bias.
[0050] Furthermore, within the coordinate system of the cross-modal correlation map, information from the source location is aggregated to obtain candidate injection quantities from the source pointing to the target; at each target location, the direction of the source location is compared with the target's own direction to obtain directionally coherent continuous readings; after the continuous readings are compressed to a stable interval through a smoothing function, they are mapped to multiplicative injection coefficients, with an example value range of 0.9 to 1.1.
[0051] When the source direction is the same as the target direction, the multiplicative injection coefficient is slightly greater than 1, indicating a slight amplification of the target position's direction; when there is a slight difference between the source and target directions, the multiplicative injection coefficient is close to 1, indicating almost no movement; when the source and target directions are opposite, the multiplicative injection coefficient is less than 1, indicating a slight suppression of the target position's direction.
[0052] At each target location, the baseline contribution of the target location itself is combined with the multiplicative injection coefficients of the two source locations into a ternary set. The final multiplicative change of the target location is determined by median statistics. When there is a difference in the correction amount between the two modes, or when one of the modes is affected by noise, the median will automatically select the middle change amount so that the features of the target location undergo only a small multiplicative adjustment. The multiplicative adjustment is applied sequentially to the mutual update between the initial features of geometric structure, the initial features of thermal radiation, and the initial features of texture color, resulting in the optimized features of each modality, including optimized features of texture color, optimized features of thermal radiation, and optimized features of geometric structure.
[0053] The optimized features of each modality are positiveized channel by channel and transferred to the logarithmic domain. In the logarithmic domain, the optimized features of each modality are aggregated position by position according to the median. Finally, the original numerical domain is returned to generate preliminary fused features.
[0054] S3. Construct an information hiding mask based on the initial fusion features, decouple the features along the geometric main direction of the geometric structure optimization, and use dual-criteria filtering channels to obtain the task-selected description.
[0055] Specifically, the task description includes performing multi-scale analysis and normalization on the initial fusion features, obtaining weighted redundancy by combining prior information, and marking spatial locations as areas to be shielded and areas to be preserved based on the relationship between weighted redundancy and original redundancy, thereby generating information hiding masks.
[0056] Based on the geometric principal direction of the optimized features, median filtering decomposition is performed in the neighborhood to split the preliminary fused features into fine-grained feature channels and high semantic feature channels. The information gain of each channel is calculated by combining the statistical results of the task proxy variables on cross-modal consistent samples and cross-modal inconsistent samples.
[0057] Using the median and mutual information as dual criteria, the selected fine-grained feature channels and high semantic feature channels are concatenated in channel order to generate a carefully selected task description.
[0058] Furthermore, in the initial fusion feature analysis, content that does not participate in subsequent matching and recognition, or even causes interference, is filtered out. Examples include large flat areas, periodically repeating textures, fragmented high-frequency information caused by noise, and obvious imaging artifacts. A pyramid structure is constructed using Gaussian and Laplacian filtering to analyze the initial fusion features at multiple scales. At each scale, the high-frequency intensity readings and dispersion readings are calculated, and the median absolute deviation method is used to normalize both. Through multi-scale analysis and normalization, a unified metric is established across different images, devices, and scales. Regions with weaker high frequencies are compressed after normalization, indicating information redundancy; regions with stronger high frequencies are highlighted after normalization, indicating richer and more effective information.
[0059] Within a unified coordinate system, several representative locations are selected using a grid and texture layering approach. Scale redundancy is calculated for each spatial location on a scale-by-scale basis, and the discrete redundancy readings arranged by scale index are interpolated to form a monotonically non-decreasing continuous curve. For any spatial location, a redundancy reading sequence arranged by scale is obtained, and scale consistency correction is performed. After correction, the redundancy reading sequences for each spatial location and each channel are aggregated across scales in the logarithmic domain using the median. The aggregated results are then restored to the original numerical domain to obtain multi-scale consistent redundancy.
[0060] After obtaining the multi-scale consistent redundancy, prior information is used as a multiplicative factor to combine the continuous validity weights of the visible light image, depth image, and infrared image, along with the temporal confidence, onto the multi-scale consistent redundancy, resulting in a weighted redundancy. The weighted redundancy is used to label spatial locations. When all three modal data are reliable at the same spatial location and the temporal confidence is consistent, the weighted redundancy is close to the original redundancy, and the spatial location is marked as a region that should be preserved. When any modal data is unreliable at the same spatial location or the temporal confidence is inconsistent, the weighted redundancy is lower than the original redundancy, and the spatial location is marked as a region that should be shielded.
[0061] In the spatial dimension, the weighted redundancy is locally calibrated by the mean value of the small window, and the geometric principal direction of the optimized features along the geometric structure is constrained to avoid crossing the real boundary. Finally, the result is cropped to the normalized interval to generate an information hiding mask.
[0062] Furthermore, within the geometric principal direction neighborhood of the geometrically optimized features, median filtering is performed to obtain the filtering result. The filtering result is used as the low-frequency component, and the difference between the preliminary fused features and the low-frequency component is used as the high-frequency component. The low-frequency component is regarded as a high-semantic component, and the high-frequency component is regarded as a fine-grained component.
[0063] In the initial fusion feature, the fine-grained components and high semantic components are split into fine-grained feature channels and high semantic feature channels; positive value is performed on each channel, and the channel readings are transferred to the logarithmic domain.
[0064] The geometric principal direction is calculated based on the horizontal difference, vertical difference, unit normal components, and curvature intensity of the depth image. At each spatial location, a local neighborhood is taken to construct a gradient vector set composed of the horizontal and vertical differences. The weighted covariance matrix is calculated using curvature intensity as the weight. The eigenvector corresponding to the smallest eigenvalue of the weighted covariance matrix is determined as the geometric principal direction, and the eigenvector orthogonal to the geometric principal direction is taken as the normal direction.
[0065] In spatial locations with smaller information hiding masks, fine-grained components are retained as the dominant fine-grained features; in spatial locations with larger information hiding masks, high-semantic components are retained as the dominant high-semantic features.
[0066] At each spatial location, a task proxy variable is labeled to determine whether the spatial location belongs to a cross-modal consistent real structure. Specifically, when the main direction of the visible light image is consistent with the main direction of the depth image, the main direction of the visible light image is consistent with the main direction of the infrared image, and the continuous validity weights of the three modal data all indicate validity and the time confidence is consistent, the current spatial location is determined to belong to a cross-modal consistent real structure, and the task proxy variable is labeled as positive. If the conditions are not met, the current spatial location is determined not to belong to a cross-modal consistent real structure, and the task proxy variable is labeled as negative.
[0067] In the fine-grained feature channel and the high semantic feature channel, the spatial location is divided into cross-modal consistent samples and cross-modal inconsistent samples according to the labeling results of the task proxy variables. The conditional distribution of the channel feature values is statistically analyzed for the two types of samples. Based on the difference in the conditional distribution of the two types of samples, the contribution of each channel in reducing uncertainty is calculated. The contribution is defined as the information gain of the channel. The information gain represents the strength of the channel's role in determining cross-modal consistent structures and cross-modal inconsistent structures.
[0068] In the spatial partitioning of the information hiding mask, the spatial location marked as the area to be preserved is further calculated as the mutual information between each channel and the task proxy variable, which serves as the mutual information index.
[0069] After obtaining the information gain and mutual information indices, a dual criterion of median and mutual information is used to perform double screening on all fine-grained feature channels and high semantic feature channels. First, the information gain of a channel must not be lower than the median of the information gains of all channels. Second, the mutual information index of a channel must not be lower than the median of the mutual information indices of all channels. Fine-grained feature channels and high semantic feature channels that meet both criteria are retained and concatenated in the original channel order to generate a selected description of the task.
[0070] S4. Generate an initial matching set under geometric correlation constraints, obtain residual values using closed geometric trial solutions, and generate a cleaned matching set after spatial median robustness.
[0071] Specifically, generating the initial matching set under geometric correlation constraints includes using the optimized features and preliminary fused features of each modality data as candidate carriers, and based on the cross-modal correlation map and combined with the current temperature parameter adjustment, finding the main peak position in the source position for each target position based on the pairing relationship between the target position and the source position, and screening candidate matching points by combining the geometric bias matrix and spatial scale constraints.
[0072] For each target location, only the pairing relationships with the source location that are mutually dominant are retained. Evidence information is recorded for each pair of target and source locations. The evidence information includes the peak value of the cross-modal correlation map, the displacement reading of the geometric bias matrix, the reading of the prior information, and the reading of the information hiding mask, and an initial matching set is generated.
[0073] Furthermore, within a unified coordinate system, texture color optimization features, geometric structure optimization features, thermal radiation optimization features, and preliminary fusion features are collectively used as matching candidate carriers. Based on the cross-modal correlation map, in the pairing relationship between target and source locations, combined with the current temperature parameter adjustment, the main peak position of each target location on the source is read. At the same time, using the reference displacement vector provided by the geometric offset matrix and combined with the spatial scale value, when the distance between the coordinate difference between the source and target locations and the reference displacement vector does not exceed the spatial scale value, the source and target locations are combined into candidate matching points. For each target location, only the pairing relationship with the source location that is mutually dominant peak is retained, where mutually dominant peak is defined as the response of both locations in the corresponding row or column being the maximum value, or being a local peak.
[0074] All cross-modal candidate matching point pairs that are mutually dominant peaks are recorded. Each candidate matching point pair includes the peak value of the cross-modal correlation map, the displacement reading of the geometric bias matrix, the reading of the prior information, and the reading of the information hiding mask. All candidate matching point pairs are recorded to form an initial matching set. The initial matching set not only includes the candidate positions jointly defined by the dominant peak and the geometric nearest neighbor, but also carries evidence information from the time and modal reliability determination, which is used for matching reliability assessment.
[0075] Furthermore, based on the initial matching set, a closed geometric alignment trial solution is constructed using geometric constraints. The residual values of candidate matching points are calculated to measure the projection error of the candidate matching points after geometric alignment. When the geometric structure optimization features or external information can provide a reliable three-dimensional spatial structure, a three-dimensional spatial point set alignment method is used for similarity calculation. When the scene is closer to planar imaging, an alignment method under planar constraints is used for calculation. The result of the geometric alignment calculation is a three-dimensional spatial similarity transformation matrix or a two-dimensional homography matrix, which is used to predict the projection position of the candidate matching points.
[0076] External information includes the internal and external parameters of the imaging device, the known geometric reference surface of the scene, or the three-dimensional structural information obtained by external sensors.
[0077] For each pair of candidate matching points in the initial matching set, a soft weight is assigned. The soft weight is calculated as follows: the peak value of the cross-modal correlation map, the displacement reading of the geometric bias matrix, the reading of the prior information, and the reading of the information hiding mask are normalized sequentially. The normalized values are then subjected to the natural logarithm operation to obtain the values in the logarithmic domain. Subsequently, the logarithmic domain values are accumulated one by one to obtain the logarithmic value of the comprehensive weight. The logarithmic value of the comprehensive weight is then subjected to the exponential operation to restore the soft weight of the candidate matching point.
[0078] After obtaining the soft weights, residual calculation is performed on each pair of candidate matching points in the initial matching set. Specifically, firstly, based on the geometric structure optimization features, the candidate matching points are back-projected into a three-dimensional point set in a unified coordinate system. Combined with the intrinsic and extrinsic parameter matrices of the imaging device, the three-dimensional points are projected back onto the image plane to obtain the predicted pixel positions. Secondly, the predicted pixel positions are compared with the actual pixel positions of the candidate matching points in the corresponding modalities to calculate the pixel coordinate differences. When the three-dimensional point set alignment method is used, the difference is expressed as a three-dimensional Euclidean distance. When the planar constraint alignment method is used, the difference is expressed as a two-dimensional projection deviation. The difference is the geometric projection residual or reprojection residual of the candidate matching points.
[0079] After obtaining the residuals of all candidate matching points, all residual values are collected into the same distribution domain, and the residual values in the distribution domain are sorted in ascending order; the cumulative proportion of each residual value in the sorted sequence is calculated to obtain its corresponding empirical cumulative distribution function value; the empirical cumulative distribution function value is used as the rank percentile value of the residual value to represent the relative position of the residual in the overall error distribution.
[0080] After obtaining the rank percentile values of the residuals, the initial matching set is subjected to consistent reweighting and regional robustness processing. Specifically, the rank percentile values of the residuals of each pair of candidate matching points are converted into continuous coefficients through monotonic mapping and limited to the normalized interval. The converted continuous coefficients are then numerically combined with the peak value of the cross-modal correlation map, the displacement reading of the geometric bias matrix, the reading of the prior information, and the reading of the information hiding mask to obtain the updated soft weights.
[0081] On the image plane, the rank percentile values of the residuals are spatially aggregated according to the neighborhood median to form a position-by-position regional median reprojection residual field. The regional median reprojection residual field is combined with the cross-modal correlation map, geometric bias matrix and spatial scale to obtain the correction effect at the regional scale. It exhibits a convergence effect at the larger regional scale and a relaxation effect at the smaller regional scale, thereby achieving spatial robustness correction of the matching residuals.
[0082] After completing the consistency reweighting and regional robustness processing, the residual values of candidate matching points, the continuous coefficients formed by the rank percentile values of the residual values, and the updated soft weights are uniformly organized to form a new matching set. For matching points that meet the conditions in the mutual consistency judgment and whose regional median residuals are within a stable range, their residual values, the continuous coefficients formed by the rank percentile values of the residual values, and the updated soft weights are retained and marked as high-reliability matching points. For matching points that do not meet the conditions in the mutual consistency judgment or whose regional median residuals fluctuate too much, their residual values, the continuous coefficients formed by the rank percentile values of the residual values, and the updated soft weights are uniformly marked as low-reliability and arranged at the end of the weight sequence in the matching point set. The final cleaned matching set contains both high-reliability and low-reliability matching points, and each pair of matching points corresponds to a residual value, the continuous coefficients formed by the rank percentile values of the residual values, and the updated soft weights.
[0083] S5. Construct a conformal transmission field by taking the median of the shared evidence graph with the rank percentile values of the evidence, complete data sharing, obtain the features of each modality data for further optimization, and aggregate them in the logarithmic domain to generate the final fused features.
[0084] Specifically, in the logarithmic domain, the aggregation of median values to generate the final fusion features includes constructing a shared evidence graph through median statistics of the rank percentile values of evidence; The shared strength is generated by combining cross-modal correlation plots, multiplicative injection coefficients, consistency indices, and prior information readings.
[0085] In the logarithmic domain, the effective components of the source location optimization features are guided and filtered along the geometric direction and boundary of the target location to obtain a conformal transmission field.
[0086] Using the shared evidence map as the gating strength, the conformal transmission field is compressed and suppressed, and linked with the information hiding mask to obtain a shared update of the target location.
[0087] Shared updates are performed on each modality of data to obtain further optimized features. The features are then restored to the original numerical domain by taking the median of each spatial position in the logarithmic domain, generating the final fused features.
[0088] Furthermore, within a unified coordinate system, a shared evidence map is constructed based on the correspondence between the target location and the source location. The generation of the shared evidence map depends on two types of information: the first type is content consistency evidence, which comes from the content coupling degree between the target location and the source location represented by the peak value of the cross-modal correlation map and the multiplicative injection coefficient; the second type is prior reliability evidence, which comes from prior information. The two types of evidence are ranked and statistically analyzed on a unified scale to obtain the evidence rank percentile values. The evidence rank percentile values are then statistically analyzed using the median to form the shared strength, which serves as the reading of the shared evidence map.
[0089] In the coordinate system of the target location, the principal geometric direction of the target location is calculated based on the geometric structure optimization features. The principal geometric direction represents the main direction of the texture or boundary. Using the reference displacement vector provided by the geometric offset matrix and combined with the neighborhood range limited by the spatial scale, the geometric difference between the source location and the target location is constrained. Under the constraint conditions, the effective components of the optimization features of the source location are gradually transmitted in the neighborhood range according to the principal geometric direction of the target location, forming a propagation path with consistent direction. During the transmission process, the geometric offset matrix limits the displacement deviation between the source location and the target location, and the spatial scale limits the effective range of propagation, ensuring that the transmission process maintains local geometric consistency. Finally, a numerical field continuously distributed in the principal geometric direction of the target location is obtained. The numerical field is truncated in the boundary region by the spatial scale constraint to avoid the source location optimization features from crossing the real boundary and causing diffusion interference. The numerical field is defined as a conformal transmission field.
[0090] By combining the shared evidence map with the conformal transport field, a small-dose shared update of the target location is completed. Specifically, firstly, the change in the transport from the source location is calculated based on the numerical baseline of the target location; secondly, the shared intensity of the shared evidence map is used as the gate control intensity to adjust the amplitude of the change in spatial distribution; finally, a shared update from the source location to the target location is formed.
[0091] The information hiding mask is used as a synchronization constraint in the shared update. In regions with large information hiding mask values, high semantic components are mainly retained; in regions with small information hiding mask values, fine-grained components are mainly retained.
[0092] Based on the shared update results, the cross-update between texture color optimization features, geometric structure optimization features and thermal radiation optimization features is completed in sequence to obtain the re-optimized features of each modality data. After the re-optimization is completed, the re-optimized features of each modality data are positiveized channel by channel and transferred to the logarithmic domain. Median aggregation is performed spatially in the logarithmic domain, and the results are restored to the original numerical domain to generate the final fused features.
[0093] Furthermore, under boundary conditions, when the source and target locations are significantly inconsistent in terms of content consistency or prior reliability, the value of the shared evidence graph approaches zero, and the sharing process tends to stagnate. When cross-modal temporal asynchrony causes abrupt changes in the value of the source location, the gating of the shared evidence graph will automatically compress abnormal differences and, combined with the shielding mechanism of information hiding mask, prevent the spread of unstable information. When the cross-modal structure remains consistent and the reliability reading is close to a stable value, the gating of the shared evidence graph will promote sharing updates in a small dose to prevent excessive diffusion.
[0094] S6. The final fused features are concatenated with the selected task descriptions and normalized to generate position-wise task features. The registration coordinate transformation matrix is generated using a closed-form method based on the cleaned matching set.
[0095] Specifically, generating position-wise task features includes calculating the scene passability factor based on the cleaned matching set, using the scene passability factor as a multiplicative gating factor, selecting and describing the final fused features and decoupled tasks, performing positiveization and inverse hyperbolic sine transform on each channel, and defining the neighborhood empirical rank statistics using spatial scale to obtain the local rank of each channel.
[0096] Using a multiplicative gating factor as an adjustment coefficient, the local rank of each channel is scaled and the amplitude is adjusted to obtain the gated feature components. An information hiding mask constraint is applied to the gated feature components to form task components. All task components are spliced together by channel and normalized to generate position-by-position task features.
[0097] Furthermore, within a unified coordinate system, the final fused features and the task-selected descriptions are concatenated by channels. During concatenation, the final fused features are arranged first, followed by the fine-grained feature channels and high semantic feature channels in the task-selected descriptions, according to the order of modal categories, ensuring that the concatenation order at each spatial location is uniquely determined. After concatenation, the concatenation result is normalized and used as the initial input for task feature generation.
[0098] Based on the cleaned matching set, the scene passability factor is calculated. Specifically, the residual values of all candidate matching points in the cleaned matching set are normalized and converted into the rank percentile values of the residual values. The rank percentile values of all residual values are averaged to obtain the scene passability factor. The scene passability factor, together with the main peak value of the cross-modal correlation map and the multiplicative injection coefficient, generates a position-by-position multiplicative gating factor.
[0099] On the initial input for task feature generation, positive transformation and inverse hyperbolic sine transform are performed channel by channel. Based on the spatial scale, empirical rank statistics are performed on each channel in the local neighborhood to obtain the local rank of each channel. The multiplicative gating factor is used as the adjustment coefficient, which is applied to the numerical range of the local rank of each channel. The local rank is scaled and the amplitude is adjusted to obtain the gated feature components.
[0100] Information hiding mask constraints are applied to the gated feature components. In regions with large information hiding mask values, high semantic components are retained as the main component, while in regions with small information hiding mask values, fine-grained components are retained as the main component. After information hiding mask constraints, all gated feature components are reassembled and normalized to generate position-wise task features.
[0101] Based on the position-by-position task features, and combined with the cleaned matching set, a closed-form method is used to calculate the registration coordinate transformation matrix. When the cleaned matching set contains three-dimensional spatial structure constraints, a three-dimensional point set alignment method is used to generate a three-dimensional similarity transformation matrix. When the cleaned matching set meets the planar imaging conditions, a direct linear transformation method is used to generate a two-dimensional homography matrix.
[0102] The registration coordinate transformation matrix is used to unify the coordinates of visible light images, depth images, and infrared images to the same geometric reference system, ensuring the spatial consistency of multimodal data.
[0103] This embodiment also provides an image feature matching and fusion system based on a multimodal large model, including: The acquisition and processing module is used to acquire data from different modalities and the intrinsic and extrinsic parameters of the imaging device, extract the initial features of each modal data, and generate prior information. The spatial alignment module is used to construct the geometric bias matrix and spatial scale. Based on the initial features and prior information of each modality data, it calculates the cross-modal correlation map and multiplicative injection intensity, obtains the optimized features of each modality data, and generates preliminary fusion features. The hidden decoupling module is used to construct an information hiding mask on the initial fused features, decouple along the geometric main direction of the optimized features, and use dual-criteria filtering channels to obtain the task's selected description. The matching feedback module is used to generate an initial matching set under geometric correlation constraints, obtain residual values using closed geometric trial solutions, and generate a cleaned matching set after spatial median robustness. The shared fusion module is used to construct a shared evidence map and a conformal transmission field, complete data sharing, obtain features of each modality data for further optimization, and aggregate them in the logarithmic domain to generate the final fusion features. The task solution module is used to concatenate the final fused features with the selected task description and normalize them to generate position-by-position task features, and generate the registration coordinate transformation matrix using a closed-form method based on the cleaned matching set.
[0104] In summary, this invention achieves directional alignment and spatial consistency of multimodal data in a unified coordinate system by introducing geometrically guided computation and spatial constraint modeling, thus avoiding geometrical offsets between different modal data and ensuring the stability and reliability of the fusion process. Furthermore, by constructing a shared evidence graph and combining median statistics and gating adjustment, it achieves conformal transmission and anomaly suppression of multimodal data in the geometric principal direction.
[0105] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. An image feature matching and fusion method based on a multimodal large model, characterized in that: include, Acquire different modal data and the intrinsic and extrinsic parameters of imaging devices, extract the initial features of each modal data, and generate prior information; A geometric bias matrix and spatial scale are constructed in a unified coordinate system. Based on the initial features and prior information of each modal data, the cross-modal correlation map and multiplicative injection intensity are calculated. After updating, the optimized features of each modal data are obtained, and the preliminary fusion features are generated by normalization. An information hiding mask is constructed on the initial fusion features, and the geometric main direction of the optimized features is decoupled. A dual-criteria filtering channel is used to obtain the task's selected description. An initial matching set is generated under geometric correlation constraints. The residual value is obtained by closed geometric trial solution. After spatial median robustness, a cleaned matching set is generated. A conformal transmission field is constructed by taking the median of the shared evidence graph with the rank percentile values of the evidence, data sharing is completed, features of each modality data are obtained and optimized again, and the final fusion features are generated by aggregation in the logarithmic domain. The final fused features are concatenated with the selected task descriptions and normalized to generate position-wise task features. The registration coordinate transformation matrix is then generated using a closed-form method based on the cleaned matching set.
2. The image feature matching and fusion method based on a multimodal large model as described in claim 1, characterized in that: The initial features for extracting each modal data include receiving modal data from visible light images, depth images, and infrared images, and using a publicly available convolutional neural network to extract the initial texture and color features of the visible light image and the initial thermal radiation features of the infrared image, respectively. The initial geometric features of the depth image are obtained by calculating the horizontal difference, vertical difference, unit normal three components, and curvature intensity of the depth image.
3. The image feature matching and fusion method based on a multimodal large model as described in claim 2, characterized in that: The generation of prior information includes estimating the average pixel motion between the current frame and adjacent frames, combining the inter-frame time difference and processing it through a smoothing function to obtain the time confidence. Each modal data generates a continuous validity weight based on its own observations, which, together with the time confidence level, constitutes prior information.
4. The image feature matching and fusion method based on a multimodal large model as described in claim 3, characterized in that: The normalization generation of preliminary fusion features includes reading the initial features and prior information of each modal data in a unified coordinate system, combining the geometric bias matrix, spatial scale and current temperature to calculate the cross-modal correlation map and establish an evidence channel for the source conversion target. By comparing the consistency between the source direction and the target's own direction, the multiplicative injection coefficient is obtained. Small-dose multiplicative updates are performed on each modality data to obtain the optimized features of each modality data. The features are then aggregated in the logarithmic domain by taking the median position one by one, and then restored back to the original data domain to obtain the preliminary fusion features.
5. The image feature matching and fusion method based on a multimodal large model as described in claim 4, characterized in that: The task selection description includes performing multi-scale analysis and normalization processing on the initial fusion features, obtaining weighted redundancy by combining prior information, and marking spatial locations as areas to be shielded and areas to be preserved based on the relationship between weighted redundancy and original redundancy, thereby generating an information hiding mask. Based on the geometric principal direction of the optimized features of the geometric structure, median filtering decomposition is performed in the neighborhood to split the preliminary fused features into fine-grained feature channels and high semantic feature channels. Statistical conditional distributions are calculated on cross-modal consistent and inconsistent samples. Based on the differences in conditional distributions, the information gain of each channel is calculated, and the mutual information index between each channel and the task proxy variable is calculated. Using the median of information gain and the median of mutual information as dual criteria, the selected fine-grained feature channels and high semantic feature channels are concatenated in channel order to generate a carefully selected task description.
6. The image feature matching and fusion method based on a multimodal large model as described in claim 5, characterized in that: The process of generating an initial matching set under geometric correlation constraints includes using the optimized features and preliminary fusion features of each modal data as candidate carriers, determining the main peak position based on the cross-modal correlation map and the current temperature parameter adjustment in the pairing relationship between the target position and the source position, and screening candidate matching points in combination with the geometric bias matrix and spatial scale constraints. For each target location, only the pairing relationships with the source location that are mutually dominant are retained. For each pair of target locations and source locations, the peak value of the cross-modal correlation map, the displacement reading of the geometric bias matrix, the reading of the prior information, and the reading of the information hiding mask are recorded to generate an initial matching set.
7. The image feature matching and fusion method based on a multimodal large model as described in claim 6, characterized in that: The process of generating the cleaned matching set includes selecting a closed geometric trial solution method under the scenario conditions, calculating the residual values of the candidate matching points, mapping the residual values to rank percentile values, and performing numerical combination calculations with the peak value of the cross-modal correlation map, the displacement reading of the geometric bias matrix, the reading of the prior information, and the reading of the information hiding mask to obtain updated soft weights. The residual values are aggregated in the neighborhood to form a regional median reprojection residual field. The soft weights are then adjusted and updated based on the regional residual information to generate a cleaned matching set.
8. The image feature matching and fusion method based on a multimodal large model as described in claim 7, characterized in that: The process of generating the final fusion feature by median aggregation in the logarithmic domain includes constructing a shared evidence graph through median statistics of the rank percentile values of evidence, and writing the shared strength generated by combining the cross-modal correlation graph, multiplicative injection coefficient, consistency index and prior information readings into the shared evidence graph. In the logarithmic domain, the effective components of the source location optimization features are guided and filtered along the geometric direction and boundary of the target location to obtain a conformal transmission field. Using the shared evidence map as the gating strength, the conformal transmission field is compressed and suppressed, and linked with the information hiding mask to obtain the shared update of the target position; Shared updates are performed on each modality of data to obtain further optimized features. The features are then restored to the original numerical domain by taking the median of each spatial position in the logarithmic domain, generating the final fused features.
9. The image feature matching and fusion method based on a multimodal large model as described in claim 8, characterized in that: The generated position-by-position task features include: The scene passability factor is calculated based on the cleaned matching set, and the scene passability factor is used as a multiplicative gating factor to select and describe the final fusion features and decoupling tasks. Positive value transformation and inverse hyperbolic sine transform are performed on each channel, and empirical rank statistics are performed on the neighborhood defined by spatial scale to obtain the local rank of each channel. Using a multiplicative gating factor as an adjustment coefficient, the local rank of each channel is scaled and the amplitude is adjusted to obtain the gated feature components. An information hiding mask constraint is applied to the gated feature components to form task components. All task components are spliced together by channel and normalized to generate position-by-position task features.
10. An image feature matching and fusion system based on a multimodal large model, based on the image feature matching and fusion method based on a multimodal large model as described in any one of claims 1 to 9, characterized in that: include, The acquisition and processing module is used to acquire data from different modalities and the intrinsic and extrinsic parameters of the imaging device, extract the initial features of each modal data, and generate prior information. The spatial alignment module is used to construct the geometric bias matrix and spatial scale. Based on the initial features and prior information of each modality data, it calculates the cross-modal correlation map and multiplicative injection intensity, obtains the optimized features of each modality data, and generates preliminary fusion features. The hidden decoupling module is used to construct an information hiding mask on the initial fused features, decouple along the geometric main direction of the optimized features, and use dual-criteria filtering channels to obtain the task's selected description. The matching feedback module is used to generate an initial matching set under geometric correlation constraints, obtain residual values using closed geometric trial solutions, and generate a cleaned matching set after spatial median robustness. The shared fusion module is used to construct a shared evidence map and a conformal transmission field, complete data sharing, obtain features of each modality data for further optimization, and aggregate them in the logarithmic domain to generate the final fusion features. The task solution module is used to concatenate the final fused features with the selected task description and normalize them to generate position-by-position task features, and generate the registration coordinate transformation matrix using a closed-form method based on the cleaned matching set.
Citation Information
Patent Citations
Self-adaptive alignment cross-modal vision-language ship intelligent man-machine interaction method
CN119357897A
Multi-modal image registration method based on self-modal correlation and cross-modal estimation
CN119762559A
Multi-modal bill processing method based on dynamic knowledge enhancement
CN120470018A
Visible light and infrared image fusion method based on cross-modal dynamic collaboration
CN120525735A
Evidence obtaining method and system based on image processing
CN120766124A
Cited By
LiDAR point cloud and image registration method based on feature constraint
CN121883547A