Visible light infrared multi-source image target fusion detection method and system
By combining the PROSAC algorithm and a dual-channel decoupling model with the DS evidence fusion engine, the problems of feature loss and conflict handling in visible light and infrared image target detection are solved, achieving high-precision and efficient target detection.
Patent Information
- Application Number
- CN202511605524.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2045-11-05
AI Technical Summary
Existing target detection methods for visible light and infrared images suffer from feature extraction models based on single-channel or multi-channel coupling structures, leading to the loss of key features or an increase in redundant information. Furthermore, the lack of effective conflict handling mechanisms results in insufficient target detection accuracy and precision.
The PROSAC optimization algorithm is used for coordinate calculation and location association, a dual-channel decoupling model is used for feature extraction and fusion, the DS evidence fusion engine is combined to handle multi-source evidence conflicts, and target fusion detection results are generated through manifold mapping and morphological operations.
It significantly improves the accuracy and processing efficiency of target detection, avoids misjudgment between modalities, and enhances the accuracy of target detection.
Smart Images

Figure CN121330271A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image target detection, and particularly relates to a visible light and infrared multi-source image target fusion detection method and system. BACKGROUND
[0002] In recent years, with the rapid development of remote sensing technology, the acquisition of visible light and infrared images has become increasingly convenient. Visible light images mainly focus on the spatial distribution and texture features of objects, while infrared images emphasize thermal radiation information caused by temperature differences. Visible light and infrared images provide important information support for environmental monitoring, military reconnaissance, disaster assessment and other fields. In particular, in the aspect of target detection, by fusing image information of different wavebands, the accuracy and reliability of target recognition can be effectively improved.
[0003] At present, in the field of target detection of visible light and infrared images, there are still some problems to be solved. Firstly, most of the existing feature extraction models are based on single-channel or multi-channel coupling structure, which leads to the loss of key features or the increase of redundant information, limiting the precision and efficiency of target detection. Secondly, there is a lack of effective conflict processing mechanism, which leads to the inability to eliminate the conflict between different modal information, so that misjudgment phenomenon is prone to occur in the multi-source evidence fusion stage, reducing the accuracy of target detection. SUMMARY
[0004] In view of the above existing problems, the present application is proposed.
[0005] Therefore, the present application provides a visible light and infrared multi-source image target fusion detection method to solve the problems of low target detection precision and insufficient target detection accuracy.
[0006] To solve the above technical problems, the present application provides the following technical solutions: In a first aspect, the present application provides a visible light and infrared multi-source image target fusion detection method, which comprises performing radiation correction and spatial registration on a multi-source image set to form an aligned multi-image sequence, and using a PROSAC optimization algorithm for coordinate solution and position association to obtain preliminary screening target position information; The preliminary screening target position information is input into a dual-channel decoupling model, the visible light channel layer performs multi-scale context perception and spatial-spectral joint analysis, the infrared channel layer performs thermal contour extraction and thermal radiation intensity calibration, and a target decoupling feature vector is output; The target decoupling feature vector is subjected to manifold mapping to obtain a multi-source evidence body, the multi-source evidence body is input into a D-S evidence fusion engine for conflict factor calculation to generate a target fusion probability value; The target fusion probability value is subjected to binaryzation judgment to obtain a target binary mask, and morphological closing operation and connected domain analysis are performed on the target binary mask to form a target fusion detection result.
[0007] As a preferred scheme of the visible light infrared multi-source image target fusion detection method, the multi-source image set comprises a visible light image, an infrared image and a multi-spectral image.
[0008] As a preferred scheme of the visible light infrared multi-source image target fusion detection method, the method comprises the following steps: The visible light image, the infrared image and the multi-spectral image are subjected to homomorphic filtering, radiation correction and wavelet transformation, and a multi-scale corrected image is outputted; The multi-scale corrected image is subjected to bicubic interpolation resampling and spatial registration, and an aligned multi-image sequence is formed; The aligned multi-image sequence is subjected to feature point screening and spatial coordinate projection using a PROSAC optimization algorithm, and a preliminary screening target positioning vector is generated; Spatial buffer analysis is performed on the preliminary screening target positioning vector, position correlation data is acquired, and the position correlation data is subjected to confidence fusion, and preliminary screening target position information is generated.
[0009] As a preferred scheme of the visible light infrared multi-source image target fusion detection method, the method comprises the following steps: A visible light channel layer and an infrared channel layer are built, multi-scale feature decoupling and hierarchical stacking are performed, and a dual-channel decoupling model is constructed; The preliminary screening target position information is inputted into the dual-channel decoupling model, the visible light channel layer applies a dilated pyramid convolution to perform multi-scale context perception and spatial-spectral joint analysis, and a high-resolution texture feature map is acquired; The infrared channel layer performs thermal contour extraction and thermal radiation intensity calibration through deformable convolution, and generates a thermal response feature map; The high-resolution texture feature map and the thermal response feature map are subjected to bidirectional cross-modal fusion, and a spectral joint feature representation is formed; The spectral joint feature representation is subjected to multi-level spatial information aggregation using a global average pooling, and a target decoupling feature vector is outputted.
[0010] As a preferred scheme of the visible light infrared multi-source image target fusion detection method, the method comprises the following steps: The target decoupling feature vector is subjected to manifold embedding and nonlinear mapping using a t-SNE method, and a manifold coordinate representation is generated; The manifold coordinate representation is subjected to selective pseudo-labeling, and a multi-source evidence body is acquired.
[0011] As a preferred scheme of the visible light infrared multi-source image target fusion detection method, the target fusion probability value is generated, and the target fusion probability value specifically includes the following steps, The D-S evidence fusion engine performs basic probability assignment and weighted aggregation on the multi-source evidence body to generate a weighted evidence body. The conflict factor of the weighted evidence body is calculated, and a comprehensive BPA score is output. The comprehensive BPA score is converted into a Pignistic probability to generate a target fusion probability value.
[0012] As a preferred scheme of the visible light infrared multi-source image target fusion detection method, the target fusion detection result is formed, and the target fusion detection result specifically includes the following steps, The target fusion probability value is binarized and determined by the Poisson probability method to obtain a target binary mask. The target binary mask is subjected to a morphological closing operation to generate an optimized mask, the optimized mask is subjected to connected domain labeling and contour extraction to form a target fusion detection result.
[0013] In a second aspect, the present application provides a visible light infrared multi-source image target fusion detection system, which comprises a position preliminary screening module for performing radiation correction and spatial registration on a multi-source image set to form an aligned multi-image sequence, and using a PROSAC optimization algorithm to perform longitude and latitude coordinate calculation and position association to obtain preliminary screening target position information. A feature decoupling module is configured to input the preliminary screening target position information into a double-channel decoupling model, perform multi-scale context perception and spatial-spectral joint analysis on a visible light channel layer, perform thermal contour extraction and thermal radiation intensity calibration on an infrared channel layer, and output a target decoupled feature vector. A probability quantization module is configured to perform manifold mapping on the target decoupled feature vector to obtain a multi-source evidence body, input the multi-source evidence body into a D-S evidence fusion engine, perform conflict factor calculation and weighted aggregation, and generate a target fusion probability value. A detection judgment module is configured to perform binarization determination on the target fusion probability value to obtain a target binary mask, perform morphological closing operation and connected domain analysis on the target binary mask, and form a target fusion detection result.
[0014] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, any step of the visible light infrared multi-source image target fusion detection method according to the first aspect of the present application is implemented.
[0015] In a fourth aspect, the present application provides a computer readable storage medium having stored thereon a computer program, wherein the computer program, when executed by a processor, implements any step of the visible light and infrared multi-source image target fusion detection method according to the first aspect of the present application.
[0016] The present application has the beneficial effects that: the visible light image and the infrared image are independently modeled and deeply analyzed by the dual-channel decoupling model, key modal features can be more fully retained, redundant information interference can be inhibited, the discriminative ability of target feature expression can be significantly improved, and thus the accuracy and processing efficiency of target detection are improved. The conflict problem existing in the fusion process of multi-modal information is effectively solved by using the D-S evidence fusion engine, the misjudgment phenomenon caused by the conflict between modalities is avoided, and thus the accuracy of target detection is enhanced. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0018] Fig. 1 It is a flowchart of the visible light and infrared multi-source image target fusion detection method.
[0019] Fig. 2 It is a schematic diagram of the visible light and infrared multi-source image target fusion detection system.
[0020] Fig. 3 It is a flowchart of the dual-channel decoupling model processing process.
[0021] Fig. 4 It is a flowchart of the working principle of the D-S evidence fusion engine. DETAILED DESCRIPTION
[0022] In order to make the above-mentioned purposes, features and advantages of the present application more apparent and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification.
[0023] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present application, therefore the present application is not limited to the specific embodiments disclosed below.
[0024] Second, the "one embodiment" or "an embodiment" referred to herein can include a particular feature, structure, or characteristic. The various embodiments are not mutually exclusive, but a single embodiment can be selected from a plurality of mutually exclusive embodiments.
[0025] Referring to Figs. 1-4 For one embodiment of the present application, the embodiment provides a visible light infrared multi-source image target fusion detection method, comprising the following steps: S1, performing radiation correction and spatial registration on the multi-source image set to form an aligned multi-image sequence, and using a PROSAC optimization algorithm for coordinate solution and position association to obtain preliminary screening target position information.
[0026] The specific operation is as follows, S1.1, collect the multi-source image set and perform preprocessing, further, the multi-source image set includes visible light images, infrared images and multispectral images, the visible light images are collected by a visible light imaging sensor, the infrared images are collected by an infrared thermal imager, and the multispectral images are collected by a multispectral camera.
[0027] Next, the multi-source image set is preprocessed, for the visible light image, Gaussian filtering is applied for smoothing and denoising to suppress noise interference, and histogram equalization is used for contrast enhancement to improve detail clarity, and at the same time, bilateral filtering is used for edge preservation to retain the target contour information of the visible light image; For the infrared image, nonlinear stretching is used for dynamic range optimization to enhance the target thermal contrast in the infrared image, and at the same time, mean filtering is used for background stray light suppression to highlight the thermal target region in the infrared image; For the multispectral image, Min-Max standardization is used for dimension compression to reduce data redundancy, and at the same time, linear interpolation is used for spatial resolution unification to realize the spatial consistency of the multispectral image.
[0028] S1.2, homomorphic filtering, radiation correction and wavelet transform are performed on the visible light image, the infrared image and the multispectral image to output a multi-scale corrected image, in the specific operation, a Butterworth filter is used for homomorphic filtering of the visible light image, further, logarithmic transformation is performed on the visible light image to obtain a logarithmic luminance domain representation; the logarithmic luminance domain representation is migrated and modulated in the frequency domain to form a filtered frequency spectrum; the filtered frequency spectrum is converted in the spatial domain and restored in brightness to output a homomorphic filtering enhanced image.
[0029] The infrared image is radiometrically corrected by an atmospheric transmission attenuation function, and further, radiometric calibration is performed on the infrared image to generate a radiometric calibration image; the radiometric calibration image is compensated for atmospheric transmission to output a corrected radiance brightness image; the corrected radiance brightness image is subjected to local contrast balance to obtain a contrast-optimized image; It should be noted that radiometric calibration refers to the process of converting the radiation value and eliminating the dark current noise of the infrared image, and atmospheric transmission compensation refers to the process of radiometric attenuation correction and pollution effect compensation on the radiometric calibration image.
[0030] The multispectral image is subjected to wavelet transform using a Daubechies wavelet basis function, and further, multiscale decomposition is performed on the multispectral image to form low-frequency approximation coefficients and high-frequency detail coefficients, and the low-frequency approximation coefficients and high-frequency detail coefficients are inversely transformed and reconstructed to obtain a multiscale fusion image; the multiscale fusion image is subjected to spectral feature preservation and enhancement to output a spectral preservation image.
[0031] The homomorphic filter-enhanced image, the contrast-optimized image, and the spectral preservation image are subjected to weighted fusion to obtain a multiscale corrected image.
[0032] S1.3, the multiscale corrected image is subjected to bicubic interpolation resampling and spatial registration to form an aligned multi-image sequence, in specific operations, bicubic convolution method is applied for bicubic interpolation resampling, and further, the neighborhood pixel points of the multiscale corrected image are extracted and subjected to weighted interpolation to obtain a weighted transition image; the weighted transition image is resampled to ensure that the resolution is improved while the geometric consistency is maintained, and edge-directed filtering is simultaneously performed to retain effective edge information of the image and eliminate the sawtooth effect, outputting a resampled optimized image; It should be noted that edge-directed filtering refers to the process of smoothing and denoising and sharpening and enhancing the weighted transition image.
[0033] SIFT (Scale Invariant Feature Transform) method is used for spatial registration, and further, the resampled optimized image is subjected to difference decomposition to extract scale-invariant feature points and perform feature description to generate feature descriptors; the feature descriptors are subjected to bidirectional K-nearest neighbor matching to obtain a matching point pair set, and the matching point pair set is subjected to pixel position mapping to output an aligned multi-image sequence; It should be noted that bidirectional K-nearest neighbor matching refers to the process of bidirectional nearest neighbor search on two groups of feature descriptors and matching according to the distance ratio of the nearest neighbor and the second nearest neighbor.
[0034] S1.4, use the PROSAC optimization algorithm to screen feature points and project spatial coordinates of the aligned multi-image sequence, and generate a preliminary screening target positioning vector. In the feature point screening stage, the PROSAC optimization algorithm is used to gradually optimize and screen the aligned multi-image sequence. Further, the feature similarity of the aligned multi-image sequence is compared, and the feature descriptor distance is obtained. When the feature descriptor distance is lower than the distance threshold, the current feature descriptor is defined as a high-confidence feature point. The high-confidence feature point is used as an initial sample, and the initial sample is progressively sampled to obtain an inlier sample. When the number of progressive sampling exceeds the preset sampling number (such as 2000 times), the current inlier sample is output as an optimal feature point. It should be noted that the distance threshold is defined based on the nearest neighbor distance ratio of the feature descriptor, and the exemplary value range is 0.7-0.9.
[0035] In the spatial coordinate projection stage, the least squares method is used to fit the optimal feature points into a preliminary screening target positioning vector. Further, the optimal feature points are normalized in coordinates to obtain normalized spatial coordinates, and the normalized spatial coordinates are subjected to geometric transformation and projection to generate a spatial position estimate. The spatial position estimate is subjected to amplitude range limitation and vector recombination to output a preliminary screening target positioning vector. The preliminary screening target positioning vector contains the spatial position (latitude and longitude or pixel coordinates) of the target in the aligned multi-image sequence, providing a basis for subsequent target tracking.
[0036] S1.5, perform spatial buffer analysis on the preliminary screening target positioning vector, obtain position correlation data, and fuse the position correlation data to generate preliminary screening target position information. In the spatial buffer analysis stage, the spatial coordinates in the preliminary screening target positioning vector are extracted, and a homogeneous coordinate transformation is performed simultaneously to construct a three-dimensional projection space. A buffer radius (such as 50 meters or 50 pixels) is set for the three-dimensional projection space to generate a circular buffer area for each target. According to the circular buffer area of each target, a polygon intersection algorithm is used for spatial superposition analysis. Further, the intersection area of each target's circular buffer area is counted to obtain overlapping region data. The overlapping region data is converted according to the unit area ratio (defined based on the total area of the circular buffer area) to obtain the intersection area ratio. When the intersection area ratio exceeds the ratio threshold (such as 30%), it is determined that the target positions are correlated, and a position correlation set is output. It should be noted that the ratio threshold is determined based on the statistical distribution rate of the historical intersection area ratio.
[0037] The position correlation set is then fused in confidence, further, the position correlation set is integrated in coordinates by weighted average method to obtain fused coordinate data, the fused coordinate data is superimposed in confidence to generate a comprehensive position estimate; the comprehensive position estimate is spatially smoothed by bilateral filtering to eliminate isolated noise points to form preliminary screening target position information.
[0038] S2, input the preliminary screening target position information into a dual-channel decoupling model, the visible light channel layer performs multi-scale context perception and spatial-spectral joint analysis, the infrared channel layer performs thermal profile extraction and thermal radiation intensity calibration, and outputs a target decoupling feature vector.
[0039] The specific operation is as follows, S2.1, construct and train the dual-channel decoupling model, in the specific operation, in the PyTorch framework, call the residual network through the nn.Module parameter, embed the residual network with the hollow pyramid convolution, set the hollow rate to [1, 6, 12], set the output channel number to 256, and set the activation function to ReLU; connect the attention mechanism after the residual network to weight the multi-band features, so as to enhance the visible light texture features, and standardize the features through L2 normalization to complete the construction of the visible light channel layer; call the DCNv2 (deformable convolution) architecture through the nn.conv2d function; set the DCNv2 architecture to an offset learning rate of 0.1, a convolution kernel size of 3x3, and a normalization method of Sigmoid; connect the full connection layer after the DCNv2 architecture to perform nonlinear mapping of the thermal radiation value, so as to eliminate the environmental temperature interference, and limit the amplitude through gradient clipping to complete the construction of the infrared channel layer; The visible light channel layer and the infrared channel layer are interacted in cross-level features by using the bidirectional gate mechanism to obtain cross-modal fusion features, the cross-modal fusion features are decoupled in multi-scale features to generate modal-specific feature representation; the spatial information of the modal-specific feature representation is aggregated by using global average pooling (GAP) to obtain channel weight; according to the channel weight, the visible light channel layer and the infrared channel layer are stacked in levels, and the gradient propagation optimization is performed through residual connection to complete the construction of the dual-channel decoupling model; Next, the dual-channel decoupling model is trained. Further, the historical initial screening target location information is divided into a sample set, a training set, and a validation set. On the sample set, data augmentation strategies are used for random rotation and flipping, and data normalization is performed through standardization to form augmented samples. On the training set, the Adam optimizer is used to update the parameters of the augmented samples, and a gradient inversion layer (GRL) is simultaneously applied to quantize the modality discrimination loss and obtain multi-task loss values. On the validation set, an early stopping mechanism is used to monitor the multi-task loss values to obtain convergence state parameters. When the convergence state parameters exceed the convergence threshold for several consecutive rounds (e.g., 5 times), training terminates, and the trained dual-channel decoupling model is output simultaneously. It should be noted that the convergence threshold is defined based on the smoothness change rate of the multi-task loss value, with an exemplary value range of [0.01, 0.05].
[0040] S2.2. The visible light channel layer applies dilated pyramid convolution to perform multi-scale context awareness and spatial-spectral joint analysis to obtain high-resolution texture feature maps. Specifically, the initial target location information is input into the visible light channel layer of the dual-channel decoupled model via the Input interface. Dilated pyramid convolutions with different dilation rates (e.g., 1, 6, 12) are used to perform multi-scale context awareness on the initial target location information. Furthermore, the branch with a dilation rate of 1 uses a 1×1 convolution kernel to perform local feature extraction on the initial target location information to obtain basic texture features, and then performs edge detection on the basic texture features. Gradient enhancement is used to extract local detail features. The branch with a dilation rate of 6 performs context expansion on the initial target location information through mid-range receptive field convolution to form region association features. Semantic capture is performed on the region association features to generate mid-range context features. The branch with a dilation rate of 12 samples large receptive field convolution to integrate global context information on the initial target location information to obtain scene structure features. Long-range dependency association is performed on the scene structure features to capture long-range semantic features. The local detail features, mid-range context features, and long-range semantic features are unified in terms of channel number and feature concatenation to output multi-scale fusion features. Simultaneously, an attention mechanism is used to perform joint spatial-spectral analysis of the multi-scale fusion features. Furthermore, global average pooling is used to compress the spectral dimension of the multi-scale fusion features, generating compressed vectors for each spectral channel. Two fully connected layers and a sigmoid activation function are used to weight and dynamically adjust the compressed vectors for each spectral channel, calculating the spectral importance weight values. The specific mathematical format is as follows. ; in, Indicates the spectral channel index. Indicates the first The spectral importance weight values for each spectral channel. Indicates activator. This represents the weights of the second fully connected layer. This represents the weights of the first fully connected layer. Indicates the first The compressed vector of each spectral channel. This represents the bias term of the first fully connected layer. This represents the bias term of the second fully connected layer; It should be noted that the activation factor is defined based on the nonlinear response intensity of the compression vector of each spectral channel, and the exemplary value is [0.3, 0.7]; the weight of the fully connected layer is defined based on the characteristic distribution characteristics of the compression vector of each spectral channel.
[0041] By performing element-wise product of spectral importance weights and initial target location information using 3×3 convolution, spatial-spectral synergistic enhancement is achieved, forming a weighted enhancement feature. The multi-scale fusion feature and the weighted enhancement feature are then fused element-wise, and the original high-frequency details are preserved through skip connections to generate a cross-modal feature representation. The ReLU function is used to perform non-linear activation on the cross-modal feature representation, and a fully connected function is used to project the feature dimension, outputting a high-resolution texture feature map. High-resolution texture feature maps preserve both local texture details (such as edges and corners) and global semantic context features (such as target shape and scene layout), and can maintain stable feature representation capabilities even under complex lighting conditions (such as shadows and backlighting).
[0042] S2.3 The infrared channel layer extracts thermal contours and calibrates thermal radiation intensity through deformable convolution, generating a thermal response feature map. Specifically, a multi-layer DCNv2 architecture is used to extract thermal contours from the initial target location information. Further, the first layer extracts the geometric deformation features of the initial target location information through 1×1 deformable convolution and activates the geometric deformation features with the Tanh function to obtain a deformation feature vector. The second layer uses 3×3 deformable convolution to dynamically adjust the position and transform the feature space of the deformation feature vector, forming a deformation adaptation feature map. The third layer performs sub-pixel sampling on the deformation adaptation feature map through bilinear interpolation to ensure accurate matching of the actual contours (such as vehicle edges and human body contours) in the deformation adaptation feature map, thereby extracting thermal contour features that characterize geometric deformation.
[0043] Subsequently, thermal radiation intensity calibration is performed on the thermal profile features. Further, the thermal profile features are normalized to the [0,1] interval using the Sigmoid activation function to obtain normalized radiation values. These normalized radiation values are then combined with ambient temperature values (collected via a temperature sensor) to calibrate the thermal radiation intensity, outputting the calibrated radiation value. The specific mathematical formula is as follows. ; in, Indicates the calibration radiation value. Represents the normalized radiation value. Indicates the temperature compensation coefficient. Indicates the ambient temperature value. Indicates the reference ambient temperature value; It should be noted that the temperature compensation coefficient is defined based on the temperature drift rate of historical calibration radiation values, with an exemplary value of 0.01±0.002 / ℃; the reference ambient temperature value is defined based on the temperature sensor calibration conditions.
[0044] The calibrated radiation values and thermal profile features are concatenated by channels and then fused using 3×3 convolution to obtain a thermal response feature map. The thermal response feature map retains accurate target geometry and radiation intensity information, supporting subsequent cross-modal fusion and target detection tasks.
[0045] S2.4. The high-resolution texture feature map and the thermal response feature map are fused bidirectionally across modally to form a joint spectral feature representation. Specifically, a 1×1 convolution is used to adjust the number of channels in the high-resolution texture feature map and the thermal response feature map to a uniform dimension, eliminating dimensionality differences. Then, a bidirectional cross-attention mechanism is used to achieve bidirectional cross-modal fusion. Further, spatial structure features are extracted from the high-resolution texture feature map, and radiative intensity features are simultaneously extracted from the thermal response feature map. The spatial structure features are used as the query vector, and the radiative intensity features as the key vector. Feature interaction and weight allocation are performed on the query vector and the key vector to ensure that the features of the two modalities are fully complementary, outputting cross-modal collaborative features. The cross-modal collaborative features, the high-resolution texture feature map, and the thermal response feature map are then joined using residual connections to preserve low-level detail information, resulting in enhanced collaborative features. Finally, a 3×3 convolution is used to locally smooth and reduce the dimensionality of the enhanced collaborative features, outputting a joint spectral feature representation. The joint spectral feature representation incorporates both high-frequency texture details of visible light and infrared thermal radiation characteristics, achieving effective complementarity in both spatial and spectral dimensions, thus providing a robust feature foundation for subsequent target detection and recognition tasks.
[0046] S2.5. Global average pooling is used to perform multi-level spatial information aggregation on the spectral joint feature representation, outputting a target decoupled feature vector. Specifically, multi-level global average pooling is used to aggregate spatial information on the spectral joint feature representation. Further, the first level uses standard global average pooling to compress the spectral joint feature representation and extract global context, obtaining a basic feature vector while preserving scene-level semantic characteristics. The second level segments the basic feature vector, dividing the spectral joint feature representation into a 2×2 grid, and independently performing local average pooling on each grid to generate four regional feature vectors, preserving the mid-scale spatial distribution characteristics. The third level further subdivides the regional feature vectors, dividing the subdivided regional feature vectors into a 4×4 grid, and simultaneously performing local average pooling, outputting 16 local feature vectors to represent fine-grained local spatial information.
[0047] Then, the basic feature vector, regional feature vector, and local feature vector are concatenated along the channel dimension to form a multi-scale spatial aggregated feature. Next, the multi-scale spatial aggregated feature is subjected to channel dimensionality reduction and feature recombination through 1×1 convolution. At the same time, the LayerNorm layer is used for normalization to eliminate the dimensional differences between features of different scales and output the target decoupling feature vector. The target decoupling feature vector contains multi-level spatial context information from global to local, and strengthens the correlation between different scales through feature recombination, supporting subsequent tasks to understand the target at multiple granularities.
[0048] S3. Perform manifold mapping on the target decoupling feature vector to obtain multi-source evidence. Input the multi-source evidence into the DS evidence fusion engine to calculate the conflict factor and perform weighted aggregation to generate the target fusion probability value.
[0049] The specific steps are as follows. S3.1. The t-SNE method is used to perform manifold embedding and nonlinear mapping on the target decoupling feature vector to generate a manifold coordinate representation. Selective pseudo-labeling is applied to the manifold coordinate representation to obtain multi-source evidence. Specifically, high-dimensional feature points of the target decoupling features are extracted and manifold embedding is performed on the high-dimensional feature points. Furthermore, Gaussian kernel transformation is performed on the high-dimensional feature points to obtain a high-dimensional similarity distribution. Spatial structure transformation is performed on the high-dimensional similarity distribution to obtain an initial low-dimensional representation. Gradient descent is then performed on the initial low-dimensional representation to generate a low-dimensional manifold embedding representation. It should be noted that Gaussian kernel transformation refers to the process of performing distance measurement and similarity assignment on high-dimensional feature points to reflect the local structural characteristics of high-dimensional feature points.
[0050] Next, the low-dimensional manifold embedding representation is nonlinearly mapped. Furthermore, neighborhood enhancement and coordinate projection are performed on the low-dimensional manifold embedding representation to obtain an optimized coordinate distribution. Spatial density equalization is performed on the optimized coordinate distribution to generate a stable manifold structure. The coordinates of the stable manifold structure are normalized to output the manifold coordinate representation. It should be noted that spatial density equalization refers to the process of adjusting the density and balancing the regions of the optimized coordinate distribution to eliminate the deviations caused by sparse or overly dense regions in the optimized coordinate distribution and achieve a uniform expression of the manifold space.
[0051] Subsequently, the Density Peak Clustering (DPC) algorithm is applied to selectively pseudo-label the manifold coordinate representation. Further, density peak statistics are performed on the manifold coordinate representation to obtain candidate density centers. These candidate density centers are used as candidate pseudo-label samples, and the local density of each candidate pseudo-label sample is calculated using the core formula of DPC. The specific mathematical formula is as follows. ; in, Indicates the index of the candidate pseudo-label sample. Indicates the first Local density of candidate pseudo-label samples, Represents the neighborhood sample index. Indicates candidate pseudo-label samples and neighborhood samples Euclidean distance, Indicates the cutoff distance; It should be noted that the neighborhood sample refers to the spatial nearest point of the candidate pseudo-label sample, which is obtained by performing a k-nearest neighbor search on the candidate pseudo-label sample; the cutoff distance refers to the neighborhood radius calculated by the local density, which is obtained by performing median statistics on the distance distribution of the candidate pseudo-label sample.
[0052] The local density is normalized and sorted in descending order to obtain the minimum distance between samples. The top U candidate pseudo-label samples with the largest product of local density and minimum distance are selected as high-confidence pseudo-labels. Finally, the high-confidence pseudo-labels and the target decoupled feature vector are normalized to ensure dimensional consistency and then concatenated to form a multi-source evidence body. The high-confidence pseudo-labels in the multi-source evidence body serve as category prior information, and the target decoupled feature vector serves as observational evidence, providing input with both topological preservation and semantic discriminative properties for subsequent DS evidence fusion.
[0053] S3.2 The DS evidence fusion engine performs basic probability allocation and weighted aggregation on multi-source evidence to generate weighted evidence. Specifically, in the basic probability allocation stage, the DS evidence fusion engine models the uncertainty of multi-source evidence. Furthermore, it projects the multi-source evidence into a feature space to generate evidence response features and normalizes the intensity of the evidence response features to obtain an evidence strength matrix. It then projects the evidence strength matrix into a probability matrix to obtain evidence probability weights. Finally, it converts the evidence probability weights into numerical values and outputs basic probability values to quantify the support differences between evidence.
[0054] In the weighted aggregation stage, the multi-source evidence is weighted according to the basic probability value. For example, when the basic probability value is greater than the probability threshold, the multi-source evidence is assigned a high weight (e.g., 0.9); when the basic probability value is less than the probability threshold, the multi-source evidence is assigned a low weight (e.g., 0.1). Based on the assigned weights, the multi-source evidence is probabilistically weighted and aggregated to generate a weighted evidence body. The weighted evidence body retains the semantic information of the original evidence and eliminates intermodal conflicts through evidence fusion. It should be noted that the probability threshold is defined based on the consistency of the category distribution of multi-source evidence, and the exemplary value range is [0.6, 0.8].
[0055] S3.3 Calculate the conflict factor for the weighted evidence body and output the overall BPA (Basic Probability) score. The specific mathematical formula is as follows: ; in, This represents the overall BPA score. Indicates the number of modes. Indicates modal index, Indicates the first The weights of each modality Indicates the first The basic probability values of each modality Indicates conflict factors; It should be noted that the conflict factor is defined based on the intermodal contradiction rate of historical weighted evidence, with an exemplary value range of [0.2, 0.6].
[0056] The overall BPA score is subjected to a Pignistic probability transformation (decision probability transformation) to generate a target fusion probability value. The specific mathematical formula is as follows. ; in, This represents the target fusion probability value. Indicates the category index for identification. Represents the set of identification categories. This indicates the current recognition category to be calculated. This represents the overall BPA score. Indicates the category of recognition The overall BPA score; It should be noted that the identification category refers to the category to which the target belongs. For example, in the scenario of identifying vehicles, pedestrians and buildings through visible light and infrared multi-source images, the identification category includes a single category (such as "vehicle") and a composite category (such as "vehicle or pedestrian").
[0057] S4. Based on the target fusion probability value, perform binarization judgment to obtain the target binary mask, perform morphological closing operation and connected component analysis on the target binary mask, and form the target fusion detection result.
[0058] The specific steps are as follows. S4.1. The target fusion probability value is binarized using the Poisson probability method to obtain the target binary mask. Specifically, the target fusion probability value is intensity-mapped to obtain the Poisson parameter. The Poisson parameter and the target fusion probability value are jointly sampled to obtain the probability sampling distribution. The probability sampling distribution is spatially projected to form a binary candidate region. The binary candidate region is binarized using a decision threshold. Furthermore, the intensity of the binary candidate region is statistically analyzed to obtain the region intensity value. When the region intensity value is higher than the decision threshold, it is determined to be a target pixel and marked with a foreground mask. When the region intensity value is lower than the decision threshold, it is determined to be a background pixel and marked with a background mask. The marked foreground mask and background mask are integrated to generate the target binary mask. It should be noted that the judgment threshold is defined based on the regional intensity ratio of the binarized candidate region, and the exemplary value range is [0.4, 0.6].
[0059] S4.2 Perform morphological closing operation on the target binary mask to generate an optimized mask. Perform connected component labeling and contour extraction on the optimized mask to form the target fusion detection result. In the specific operation, during the morphological closing operation stage, a 3×3 rectangular structuring element is used to perform a dilation operation on the target binary mask to fill the small holes inside the binary candidate region caused by sampling noise or threshold segmentation, and output the dilated mask. Then, a 3×3 rectangular structuring element is used to perform an erosion operation on the dilated mask to restore the original boundary of the target and eliminate the edge expansion effect caused by dilation, thus obtaining the optimized mask.
[0060] Next, the optimized mask is labeled with connected components. Further, according to the neighborhood connection rule, the optimized mask is labeled with regional attributes to obtain labeled connected regions. Geometric features of the labeled connected regions are extracted and aggregated to generate candidate target regions. Then, conditional constraint filtering is performed on the candidate target regions. Further, the shape description parameters in the candidate target regions are extracted and linearly normalized to obtain the candidate target score. When the candidate target score is within the effective conditional constraint range (e.g., [0.7, 1.0]), it is a candidate effective target region. The candidate effective target regions are then optimized for regional boundaries to output the effective target regions. It should be noted that the neighborhood connection rule is based on the spatial continuity definition of the historical optimization mask; the effective condition constraint range is based on the probability distribution characteristics of the historical candidate target scores.
[0061] The Suzuki85 contour tracking algorithm is used to extract the external contour points in the effective target area. Then, polygon approximation is performed on the effective target area to obtain a simplified contour. The simplified contour is then smoothed at the edges to form a regularized contour. The regularized contour is then dynamically rendered in 3D to output the target fusion detection result. The target fusion detection result is stored in the form of effective target area and regularized contour, which can be directly used for target tracking or classification tasks. It should be noted that polygon approximation refers to the process of extracting contour points and removing redundant points from the effective target area.
[0062] This embodiment also provides a visible light infrared multi-source image target fusion detection system, including: a location screening module, which is used to perform radiometric correction and spatial registration on the multi-source image set to form an aligned multi-image sequence, and use the PROSAC optimization algorithm to calculate latitude and longitude coordinates and position association to obtain the location information of the screened target; The feature decoupling module is used to input the initial target location information into the dual-channel decoupling model. The visible light channel layer performs multi-scale context awareness and spatial-spectral joint analysis, while the infrared channel layer performs thermal contour extraction and thermal radiation intensity calibration, and outputs the target decoupling feature vector. The probability quantization module is used to perform manifold mapping on the target decoupling feature vector to obtain multi-source evidence. The multi-source evidence is then input into the DS evidence fusion engine to perform conflict factor calculation and weighted aggregation to generate the target fusion probability value. The detection and judgment module is used to perform binarization judgment based on the target fusion probability value, obtain the target binary mask, perform morphological closing operation and connected component analysis on the target binary mask, and form the target fusion detection result.
[0063] This embodiment also provides a computer device applicable to the visible light infrared multi-source image target fusion detection method, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the visible light infrared multi-source image target fusion detection method proposed in the above embodiment.
[0064] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0065] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the visible light and infrared multi-source image target fusion detection method proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0066] In summary, this invention employs a dual-channel decoupling model to independently model and perform depth analysis on visible light and infrared images. This approach more fully preserves key modal features, suppresses redundant information interference, and significantly improves the discriminative ability of target feature representation, thereby enhancing the accuracy and processing efficiency of target detection. Furthermore, the use of the DS evidence fusion engine effectively resolves the conflict issues that exist during the fusion of multimodal information, avoiding misjudgments caused by intermodal conflicts and thus enhancing the accuracy of target detection.
[0067] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for target fusion detection using visible light and infrared multi-source images, characterized in that: include, Radiometric correction and spatial registration are performed on the multi-source image set to form an aligned multi-image sequence. The PROSAC optimization algorithm is then used to calculate coordinates and correlate positions to obtain the initial target location information. The initial target location information is input into the dual-channel decoupling model. The visible light channel layer performs multi-scale context awareness and spatial-spectral joint analysis, while the infrared channel layer performs thermal contour extraction and thermal radiation intensity calibration, and outputs the target decoupling feature vector. Manifold mapping is performed on the target decoupling feature vector to obtain multi-source evidence. The multi-source evidence is then input into the DS evidence fusion engine to calculate the conflict factor and generate the target fusion probability value. Binarization is performed based on the target fusion probability value to obtain the target binary mask. Morphological closing operation and connected component analysis are then performed on the target binary mask to form the target fusion detection result.
2. The visible light infrared multi-source image target fusion detection method as described in claim 1, characterized in that: The multi-source image set includes visible light images, infrared images, and multispectral images.
3. The visible light infrared multi-source image target fusion detection method as described in claim 2, characterized in that: The process of obtaining the initial target location information includes the following steps: Homomorphic filtering, radiometric correction, and wavelet transform are applied to visible light images, infrared images, and multispectral images to output multiscale corrected images. Bicubic interpolation resampling and spatial registration are performed on the multi-scale corrected images to form an aligned multi-image sequence; The PROSAC optimization algorithm is used to align multiple image sequences, perform feature point filtering and spatial coordinate projection, and generate initial target localization vectors. Spatial buffer analysis is performed on the initial screening target location vector to obtain location association data, and the location association data is fused with confidence to generate the initial screening target location information.
4. The visible light infrared multi-source image target fusion detection method as described in claim 3, characterized in that: The output target decoupled feature vector specifically includes the following steps. Construct visible light channel layer and infrared channel layer, and perform multi-scale feature decoupling and layer stacking to build a dual-channel decoupled model; The initial target location information is input into the dual-channel decoupled model. The visible light channel layer applies dilated pyramid convolution to perform multi-scale context awareness and spatial-spectral joint analysis to obtain high-resolution texture feature maps. The infrared channel layer uses deformable convolution to extract thermal profiles and calibrate thermal radiation intensity, generating thermal response feature maps. High-resolution texture feature maps and thermal response feature maps are fused bidirectionally across modes to form a joint spectral feature representation; Global average pooling is used to perform multi-level spatial information aggregation on the joint spectral feature representation, and the target decoupled feature vector is output.
5. The visible light infrared multi-source image target fusion detection method as described in claim 4, characterized in that: The acquisition of multi-source evidence specifically includes the following steps. The t-SNE method is used to perform manifold embedding and nonlinear mapping on the target decoupled feature vector to generate a manifold coordinate representation; Selective pseudo-labeling is applied to the manifold coordinate representation to obtain multi-source evidence.
6. The visible light infrared multi-source image target fusion detection method as described in claim 5, characterized in that: The generation of the target fusion probability value specifically includes the following steps. The DS evidence fusion engine performs basic probability allocation and weighted aggregation on multi-source evidence to generate weighted evidence. The conflict factor is calculated on the weighted evidence body, and the comprehensive BPA score is output; the comprehensive BPA score is then subjected to a Pignistic probability transformation to generate the target fusion probability value.
7. The visible light infrared multi-source image target fusion detection method as described in claim 6, characterized in that: The formation of the target fusion detection result specifically includes the following steps. The target fusion probability value is binarized and determined using the Poisson probability method to obtain the target binary mask; A morphological closing operation is performed on the binary mask of the target to generate an optimized mask. Connected component labeling and contour extraction are performed on the optimized mask to form the target fusion detection result.
8. A visible light infrared multi-source image target fusion detection system, based on the visible light infrared multi-source image target fusion detection method according to any one of claims 1 to 7, characterized in that: include, The location screening module is used to perform radiometric correction and spatial registration on a multi-source image set to form an aligned multi-image sequence, and uses the PROSAC optimization algorithm to calculate latitude and longitude coordinates and associate locations to obtain the location information of the initial screening target. The feature decoupling module is used to input the initial target location information into the dual-channel decoupling model. The visible light channel layer performs multi-scale context awareness and spatial-spectral joint analysis, while the infrared channel layer performs thermal contour extraction and thermal radiation intensity calibration, and outputs the target decoupling feature vector. The probability quantization module is used to perform manifold mapping on the target decoupling feature vector to obtain multi-source evidence. The multi-source evidence is then input into the DS evidence fusion engine to perform conflict factor calculation and weighted aggregation to generate the target fusion probability value. The detection and judgment module is used to perform binarization judgment based on the target fusion probability value, obtain the target binary mask, perform morphological closing operation and connected component analysis on the target binary mask, and form the target fusion detection result.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the visible light infrared multi-source image target fusion detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the visible light infrared multi-source image target fusion detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
IR and visible light image fusion method by fuzzy measure and morphology alternating operators
CN104574334A
Railway track damage detection method
CN120314456A
Real-time fusion method and system of infrared image and visible light image
CN120852186A
Cradle head dynamic target anti-interference tracking method fusing multi-sensor data
CN120876540A
Deep generative modeling of smooth image manifolds for multidimensional imaging
US20220222781A1
Cited By
Infrared optical image fusion method and system for manifold adaptive scale decomposition
CN122265055A
An infrared optical image fusion method and system based on manifold adaptive filtering
CN122265055B