An image lossless fusion method and system based on deep feature extraction

Through the lossless fusion method based on deep feature extraction, the multi-layer pulsed neuron network and adaptive weight adjustment mechanism are used to solve the modal conflict problem between infrared and visible images, and efficient image fusion in dynamic scenes is achieved, image quality and reliability are improved.

CN119904726BActive Publication Date: 2025-07-01IVIDEA CULTURAL CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510398463.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-01
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

Traditional image fusion methods are difficult to adapt to complex and changeable dynamic scenes, especially the modal conflict problem between infrared images and visible light images, resulting in the fused image being unable to quickly reach the optimal state.

Method used

The image lossless fusion method based on deep feature extraction is adopted, and the pulse distribution characteristics of biological neurons are simulated using a multi-layer pulsed neuron network, and dynamic weight allocation and confidence-weighted fusion are carried out to achieve efficient fusion of images.

Benefits of technology

Real-time and fidelity of images are achieved in dynamic scenarios, improving the quality and reliability of the fused images, and quickly achieving the best results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904726B_ABST
    Figure CN119904726B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image optimization technology, and discloses an image lossless fusion method and system based on deep feature extraction. The method includes the following steps: collecting multi-modal input data from different sensors, preprocessing each modality in the multi-modal input data. The multi-modal input data includes visible light images, infrared images, and depth maps. Using a multi-layer spiking neuron network to simulate the spiking characteristics of biological neurons, performing temporal feature extraction on the multi-modal input data to capture feature representations at different levels. The present invention performs image fusion by extracting deep features through the SNN algorithm, which can effectively solve the problems of feature mutual exclusion and modality conflict faced by traditional methods when processing infrared and visible light images. It can be converted from concrete images to abstract features, and then from abstract features to the end-to-end conversion process of fused images, meeting the real-time and fidelity requirements in dynamic scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image optimization, and more specifically, to an image lossless fusion method and system based on depth feature extraction. Background Art

[0002] With the rapid development of computer vision and deep learning, image fusion technology has been widely applied in many fields, such as medical imaging, remote sensing, security monitoring, and autonomous driving.

[0003] Traditional image fusion methods usually rely on manually designed feature extraction and fusion strategies, and it is difficult to adapt to complex and changeable dynamic scenarios. Especially for infrared images and visible light images, the mutual exclusivity between the temperature-sensitive features of infrared images and the texture features of visible light images causes a trust crisis due to environmental interference caused by sensor modality conflicts, making the fused image unable to quickly reach the optimal state. Summary of the Invention

[0004] The present invention provides an image lossless fusion method and system based on depth feature extraction to solve the technical problems in the related art.

[0005] The present invention provides an image lossless fusion method based on depth feature extraction, including the following steps:

[0006] S100, Input data preparation: Collect multi-modal input data from different sensors, and preprocess each modality in the multi-modal input data. The multi-modal input data includes visible light images, infrared images, and depth maps;

[0007] S200, Bio-inspired feature extraction: Use a multi-layer spiking neuron network to simulate the spiking characteristics of biological neurons, extract temporal features from the multi-modal input data, and capture feature representations at different levels;

[0008] S300, Dynamic weight adjustment mechanism: Introduce an adaptive weight adjustment mechanism to dynamically adjust the weights according to the performance of the sensor modality in a specific environment;

[0009] S400, Modality conflict detection: Implement a modality conflict detection algorithm to identify conflicts between different sensor modalities, and use the algorithm to analyze the consistency of the modality outputs to determine whether there are conflicts;

[0010] S500, Feature fusion: Apply an improved fusion strategy. During the fusion process, use dynamic weight allocation and confidence-weighted fusion to maximize the influence of important modalities on the final result;

[0011] S600, Output Generation: Decode the compressed features into preliminary image data, reconstruct the image using the pulse timing information, and perform detail enhancement and color correction, then output the final image.

[0012] Furthermore, the multi-layer pulse neuron network includes an L1 layer and an L2 layer. The L1 layer is used for simulating retinal bipolar cells, and the L2 layer is used for simulating ganglion cells.

[0013] Furthermore, in the L1 layer, the membrane potential equation is as follows:

[0014] ;

[0015] Where represents the pulse feature output, represents the short-term memory constant, = 10 ms, represents the edge detection convolution kernel, , represents the membrane potential of the neurons in the L1 layer, represents the rate of change of the membrane potential with time in the L1 layer, represents the convolution operation;

[0016] Pulse emission condition:

[0017] ;

[0018] Where represents the high-sensitivity threshold, , represents the pulse output of the L1 layer at time t, represents the membrane potential of the L1 layer at time t;

[0019] In the L2 layer, the dynamic synapse equation is as follows:

[0020] ;

[0021] Where represents the synaptic decay constant, = 30 ms, where represents the pulse response intensity, = 0.8, represents the synaptic current of the L2 layer, represents the number of received pulses, represents the timestamp of the k-th pulse, represents the current time;

[0022] The feature output is as follows:

[0023] ;

[0024] where represents the feature output of layer L2, and represents the synaptic current of layer L2 at time t.

[0025] Furthermore, in S200, when performing temporal feature extraction on multi-modal input data, STDP weight adjustment is required, and its weight update rule is as follows:

[0026] ;

[0027] ;

[0028] where and , where represents the learning rate, = 20 ms, where represents the time-related window, where represents the pulse timing difference, represents the weight change amount, represents the pulse output of the pre-neuron i, represents the pulse output of the post-neuron j, represents the pulse time of the pre-neuron i, represents the pulse time of the post-neuron j;

[0029] Constraint:

[0030] ;

[0031] where represents the weight upper limit, = 2.5, represents the connection weight between neurons i and j, represents the operation of restricting the weight within a specified range.

[0032] Furthermore, in S200, using the pulse density gating mechanism, the calculation formula for its pulse density is as follows:

[0033] ;

[0034] ;

[0035] where represents the adjustment coefficient, = 1.5, represents the reference density, = 0.6, represents the Sigmoid function, represents the time window length, = 8, represents the total number of neurons, , represents the dynamic fusion weight, represents the spike output of the nth neuron in layer L1 at time t, represents the spike density of layer L1;

[0036] ;

[0037] where represents the fused feature, represents the feature output of layer L1; represents the feature output of layer L2;

[0038] Through downsampling spike pooling:

[0039] ;

[0040] where represents the pyramid level, , where represents the scale decay coefficient, , represents the feature after downsampling in the lth layer, represents the max pooling operation, represents the pooling kernel size;

[0041] Through upsampling feature concatenation:

[0042] ;

[0043] where represents the channel concatenation operation, represents the feature pyramid output, represents the upsampling operation, represents the upsampling scale factor.

[0044] Furthermore, in S300, it specifically includes the following steps:

[0045] S310, Sensor Output Confidence Monitoring: Calculate the confidence according to the output of each sensor modality;

[0046] S320, Initial Weight Allocation: Allocate initial weights to each modality according to the historical data of sensor performance;

[0047] S330, Dynamic Weight Update: Update the weights in real time based on the sensor output confidence;

[0048] S340, Bio-inspired Feedback Regulation Mechanism: Introduce feedback regulation based on conflict detection to optimize the weight oscillation problem.

[0049] Further, in S400, the calculation formula of the modal conflict detection algorithm is as follows:

[0050] Conflict determination condition:

[0051] ;

[0052] where represents the weight difference threshold, = 0.4, is the weight of the j-th neuron at time t, is the weight of the i-th neuron at time t, represents the conflict determination result;

[0053] Construct a pulse gating function:

[0054] ;

[0055] where represents the pulse density of the i-th neuron, where represents the pulse density of the j-th neuron, represents the anti-zero constant, represents the conflict adjustment coefficient.

[0056] Further, in S500, the calculation formula of dynamic weight allocation is as follows:

[0057] ;

[0058]

[0059] where represents the corrected dynamic weight, represents the total number of neurons;

[0060] The calculation formula of confidence weighted fusion is as follows:

[0061] ;

[0062] ;

[0063] where represents the enhancement function for features, represents the learnable convolution kernel, , represents the bias term, , represents the fused feature, where represents the attention weighting;

[0064] ;

[0065] where represents the affine transformation network, represents the parameters learned in STDP, represents the spike feature output of the i-th neuron, , represents the feature output of the i-th neuron after alignment.

[0066] Furthermore, in S600, the calculation formula of the encoding network is as follows:

[0067] ;

[0068] where represents the deconvolution kernel, , where = 3 corresponds to the RGB channels, represents the decoding bias term, , represents the ReLU function, represents the decoded feature, represents the finally output compressed feature, represents the convolution operation;

[0069] The calculation formula for image reconstruction is as follows:

[0070] ;

[0071] ;

[0072] where represents the time decay function, represents the decay duration, = 25ms, represents the reconstruction time step, = 15, represents the reconstructed image;

[0073] The calculation formula for detail enhancement is as follows:

[0074] ;

[0075] where represents the Laplacian sharpening kernel, represents the enhancement intensity, = 0.3, represents the enhanced image;

[0076] The calculation formula for color correction is as follows:

[0077] ;

[0078] Among them represents the color correction matrix, , represents the image after color correction;

[0079] The calculation formula of the final image is as follows:

[0080] ;

[0081] Among them represents the finally output image.

[0082] The present invention also proposes an image lossless fusion system based on deep feature extraction, which executes the steps in the foregoing image lossless fusion method based on deep feature extraction, including:

[0083] Multimodal acquisition and preprocessing module: Collect multimodal input data from different sensors and preprocess each modality in the multimodal input data;

[0084] Biological pulse feature encoding module: A hierarchical pulse neuron network realizes retinal bionic encoding, extracts a multi-scale spatio-temporal feature pyramid, and constructs a cross-layer feature association gated by pulse density;

[0085] Dynamic weight optimization module: Introduce an adaptive weight adjustment mechanism to dynamically adjust the weight according to the performance of the sensor modality in a specific environment;

[0086] Modal conflict detection module: Implement a modal conflict detection algorithm to identify conflicts between different sensor modalities, and use the algorithm to analyze the consistency of modal outputs to determine whether there are conflicts;

[0087] Cross-modal fusion decision module: Apply an improved fusion strategy. During the fusion process, use dynamic weight allocation and confidence-weighted fusion to maximize the impact of important modalities on the final result;

[0088] Pulse inverse mapping reconstruction module: Decode the compressed features into preliminary image data, reconstruct the image using pulse timing information, and after detail enhancement and color correction, output the final image.

[0089] The beneficial effects of the present invention are as follows:

[0090] The present invention performs image fusion by extracting deep features through the SNN algorithm, which can effectively solve the problems of feature mutual exclusion and modal conflict faced by traditional methods when processing infrared and visible light images. It can convert from concrete images to abstract features and then to the end-to-end conversion process of fused images, meeting the real-time and fidelity requirements in dynamic scenarios, improving the quality and reliability of fused images, and ensuring rapid achievement of the best effect in dynamic complex scenarios. Description of the Drawings

[0091] Figure 1 is the flowchart of an image lossless fusion method based on deep feature extraction proposed by the present invention;

[0092] Figure 2 is the structural block diagram of an image lossless fusion system based on deep feature extraction proposed by the present invention;

[0093] Figure 3 is the key data flow verification table in the example proposed by the present invention.

[0094] In the figure: 101, multi-modal acquisition and preprocessing module; 102, biological pulse feature encoding module; 103, dynamic weight optimization module; 104, modal conflict detection module; 105, cross-modal fusion decision-making module; 106, pulse inverse mapping reconstruction module. Detailed implementation manners

[0095] Now, the subject matter described herein will be discussed with reference to example embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described relative to some examples can also be combined in other examples.

[0096] As Figure 1 shown, an image lossless fusion method based on deep feature extraction includes the following steps:

[0097] S100, input data preparation: Collect multi-modal input data from different sensors (including visible light images, infrared images, depth maps, etc.).

[0098] It is also necessary to perform preprocessing on each modality, including normalization, denoising, and feature extraction.

[0099] In an embodiment of the present invention, it specifically includes the following steps:

[0100] S110, multi-modal data acquisition: Input visible light image , infrared image , depth map ;

[0101] Perform time synchronization error on the above input images, where is the lowest sensor frame rate, and the spatial resolution is unified to ;

[0102] S120, modality-specific normalization

[0103] Visible light image:

[0104] Execution formula:

[0105] ;

[0106] Wherein and respectively represent the mean and standard deviation of channel , and and respectively represent the abscissa and ordinate of the image, represents the normalized visible light image;

[0107] Output range: ;

[0108] Infrared image:

[0109] Dynamic range compression:

[0110] ;

[0111] Wherein represents the range of the infrared sensor, wherein , and respectively represent the maximum and minimum values of the range of the infrared sensor, represents the pixel value of the infrared image at position ;

[0112] Depth map:

[0113] Distance normalization:

[0114] ;

[0115] Wherein represents the maximum detection distance, represents the scaling factor, , represents the normalized depth value, represents the depth value at position (x, y);

[0116] S130, Noise suppression processing

[0117] Visible light image:

[0118] Gaussian-bilateral joint filtering:

[0119] ;

[0120] Wherein represents the Gaussian kernel, , Denotes bilateral filtering, where the spatial domain , Denotes the weight coefficient, Denotes the denoised visible light image;

[0121] Infrared image:

[0122] Non-local means denoising:

[0123] ;

[0124] ;

[0125] Where Denotes the denoised infrared image, Denotes the position The pixel value of the infrared image at the position, where Denotes centered on A 7×7 image block, Denotes the smoothing parameter, , Denotes the normalization constant to ensure that the weight sum is 1, Denotes the set of pixels within the search window, Denotes the position The weight of the pixel at the position;

[0126] S140, Pulse feature pre-extraction

[0127] Pulse coding network (PCN):

[0128] Membrane potential equation:

[0129] ;

[0130] Where Denotes the membrane time constant, , Denotes the coding convolution kernel, , Denotes the neuron membrane potential, Denotes the rate of change of the membrane potential with time, Denotes the denoised input image;

[0131] Pulse trigger condition:

[0132] ;

[0133] Where Denotes the pulse trigger threshold, =0.8, Denotes the pulse output at time t, Denotes the membrane potential at time t;

[0134] Output feature:

[0135] ;

[0136] where represents the time decay factor, , represents the time step, = 10, represents the pulse feature output;

[0137] S200, Bio-inspired feature extraction: Use a spiking neural network (SNN) for feature extraction.

[0138] where the SNN can simulate the spiking characteristics of biological neurons, extract temporal features, and design a multi-layer spiking neuron network to capture feature representations at different levels.

[0139] In one embodiment of the present invention, it specifically includes the following steps:

[0140] S210, Hierarchical pulse feature encoding

[0141] Input data: Preprocessed pulse features ;

[0142] Level division:

[0143] Layer L1 (simulating retinal bipolar cells):

[0144] Membrane potential equation:

[0145] ;

[0146] where represents the short-term memory constant, = 10 ms, represents the edge detection convolution kernel, , represents the membrane potential of neurons in layer L1, represents the rate of change of the membrane potential over time in layer L1, represents the convolution operation;

[0147] Spiking condition:

[0148] ;

[0149] where represents the high-sensitivity threshold, , represents the pulse output of layer L1 at time t, represents the membrane potential of layer L1 at time t;

[0150] Layer L2 (ganglion cell simulation):

[0151] Dynamic synapse equation:

[0152] ;

[0153] where represents the synaptic decay constant, = 30 ms, where represents the impulse response intensity, = 0.8, represents the synaptic current of layer L2, represents the number of received impulses, represents the timestamp of the k-th impulse, represents the current time;

[0154] Feature output:

[0155] ;

[0156] where represents the feature output of layer L2, represents the synaptic current of layer L2 at time t;

[0157] S220, time-sequence dependent feature enhancement:

[0158] Through STDP weight adjustment, its weight update rule is as follows:

[0159] ;

[0160] ;

[0161] where where represents the learning rate, = 20 ms, where represents the time-related window, where represents the impulse time difference, represents the weight change amount, represents the impulse output of the pre-neuron i, represents the impulse output of the post-neuron j, represents the impulse time of the pre-neuron i, represents the impulse time of the post-neuron j;

[0162] Constraint conditions:

[0163] ;

[0164] where Represents the weight upper limit, = 2.5, Represents the connection weight between neurons i and j, Represents the operation of restricting the weight within a specified range;

[0165] S230, Cross - level feature fusion:

[0166] Utilize the pulse density gating mechanism, and the calculation formula for its pulse density is as follows:

[0167] ;

[0168] ;

[0169] Where Represents the adjustment coefficient, = 1.5, Represents the reference density, = 0.6, Represents the Sigmoid function, Represents the time window length, = 8, Represents the total number of neurons, , Represents the dynamic fusion weight, Represents the pulse output of the nth neuron in layer L1 at time t, Represents the pulse density of layer L1;

[0170] Final fusion output:

[0171] ;

[0172] Where Represents the feature after cross - level fusion, , Represents the feature output of layer L1; Represents the feature output of layer L2;

[0173] S240, Multi - scale feature pyramid construction:

[0174] Down - sampling pulse pooling:

[0175] ;

[0176] Where Represents the pyramid level, Where Represents the scale decay coefficient, , Represents the feature after down - sampling in the lth layer, Represents the max - pooling operation, Indicates the pooling kernel size;

[0177] Upsampling feature concatenation:

[0178] ;

[0179] Where Indicates the channel concatenation operation, Indicates the output of the feature pyramid, , Indicates the upsampling operation, Indicates the upsampling scale factor;

[0180] S300, Dynamic weight adjustment mechanism: An adaptive weight adjustment mechanism is introduced to dynamically adjust the weights according to the performance of the sensor modality in a specific environment;

[0181] Among them, the confidence of the monitoring sensor output is monitored, and the weights are updated in real time based on the relative importance of the modalities. A biologically inspired mechanism, such as feedback regulation based on conflict detection, is used to optimize the weight oscillation problem.

[0182] In an embodiment of the present invention, it specifically includes the following steps:

[0183] S310, Sensor output confidence monitoring: Calculate the confidence according to the output of each sensor modality;

[0184] The calculation formula is as follows:

[0185] ;

[0186] Where, Is the confidence of the i-th sensor modality, Is the output score of the i-th modality, Is the output score of the j-th modality, Is the total number of modalities.

[0187] S320, Initial weight assignment: Assign initial weights to each modality according to the historical data of sensor performance;

[0188] The calculation formula is as follows:

[0189] ;

[0190] Where, Is the initial weight of the i-th sensor, Is the confidence of the i-th sensor modality, Is the confidence of the k-th sensor modality, Is the total number of modalities;

[0191] S330, Dynamic weight update: Update the weights in real time based on the confidence of sensor outputs;

[0192] The calculation formula is as follows:

[0193] ;

[0194] Where is the weight of the i-th sensor at time t, is the learning rate, which controls the update amplitude, is the weight of the i-th sensor at time t-1;

[0195] S340, Bio-inspired feedback regulation mechanism: Introduce feedback regulation based on conflict detection to optimize the weight oscillation problem;

[0196] The calculation formula is as follows:

[0197] ;

[0198] ;

[0199] Where is the conflict metric of the i-th modality, is the intensity coefficient of feedback regulation.

[0200] S400, Modal conflict detection: Implement a modal conflict detection algorithm to identify conflicts between different sensor modalities;

[0201] Use statistical methods or deep learning algorithms to analyze the consistency of modal outputs and determine whether there are conflicts.

[0202] In an embodiment of the present invention, it specifically includes the following steps:

[0203] S410, Feature space consistency metric

[0204] Input data: Cross-level fusion features from S230 ;

[0205] Define the feature similarity matrix between modalities , where:

[0206] ;

[0207] ;

[0208] ;

[0209] Where and represent the impulse features extracted by different modalities, Denote the pulse feature output of the $i$-th modality, Denote the pulse feature output of the $j$-th modality, $N$ represents the total number of modalities (example: $N = 3$ corresponding to RGB / IR / Depth), Denote the feature similarity between the $i$-th modality and the $j$-th modality;

[0210] S420, Pulse timing synchronization analysis

[0211] Input data: Pulse sequence from S210

[0212] Synchronization index:

[0213] ;

[0214] where and respectively denote the pulse times of modalities $a$ and $b$ at the $i$-th and $j$-th neurons, $= 15\text{ms}$, denotes the synchronization tolerance window, denotes the time window length;

[0215] S430, Dynamic confidence conflict detection

[0216] Input parameter: Dynamic weight from S330 ;

[0217] Conflict determination condition:

[0218] ;

[0219] where denotes the weight difference threshold, $= 0.4$, is the weight of the $j$-th neuron at time $t$, denotes the conflict determination result;

[0220] S440, Conflict resolution based on pulse competition inhibition:

[0221] Construct a pulse gating function:

[0222] ;

[0223] where denotes the pulse density of the $i$-th neuron, where denotes the pulse density of the $j$-th neuron, denotes the anti-zero constant, denotes the conflict adjustment coefficient;

[0224] S500, Feature Fusion: Apply improved fusion strategies, such as weighted average, feature concatenation, or neural network-based fusion methods.

[0225] During the fusion process, use dynamic weights and modal confidence for weighting to ensure that the influence of important modalities on the final result is maximized.

[0226] In an embodiment of the present invention, it specifically includes the following steps:

[0227] S510, Dynamic Weight Allocation

[0228] Input parameters: The dynamic weight of S330 , the conflict adjustment coefficient of S440 ;

[0229] Weight correction formula:

[0230] ;

[0231] Where represents the corrected dynamic weight;

[0232] Constraint conditions:

[0233] ;

[0234] Where represents the total number of neurons;

[0235] S520, Spatiotemporal Feature Alignment: Input the multi-scale feature pyramid of S240 and the spike feature of S140 ;

[0236] Alignment transformation:

[0237] ;

[0238] Where represents the affine transformation network, represents the parameters obtained by STDP learning, represents the spike feature output of the i-th neuron, , represents the feature output of the i-th neuron after alignment;

[0239] S530, Confidence-Weighted Fusion:

[0240] Fusion formula:

[0241] ;

[0242] ;

[0243] Among them represents the enhancement function of the feature, represents the learnable convolution kernel, , represents the fusion bias term, , represents the fused feature;

[0244] S540, Pulse Feedback Optimization:

[0245] Input signal: Pulse emission sequence of S210

[0246] Optimization objective function:

[0247] ;

[0248] Among them and respectively represent the first and second balance coefficients, = 0.8, = 0.2, represents the true fusion target collected, represents the gradient operator, represents the multi-scale network;

[0249] S550, Multimodal Feature Compression

[0250] Dimensionality reduction operation:

[0251] ;

[0252] Among them represents the compression matrix, , = 128, Output dimensionality compression ratio , represents the compressed feature of the final output, , represents the flattening operation, represents the dimensionality reduction weight parameter;

[0253] S600, Output Generation: Decode the compressed feature into preliminary image data, reconstruct the image using the pulse timing information, and perform detail enhancement and color correction, then output the final image;

[0254] In an embodiment of the present invention, it specifically includes the following steps:

[0255] S610, Pulse Feature Decoding: Input the compressed feature of S550 into the decoding network;

[0256] The calculation formula of the decoding network is as follows:

[0257]

[0258] wherein represents the deconvolution kernel, , wherein = 3 corresponds to the RGB channels, represents the decoding bias term, , represents the ReLU function, represents the decoded feature, represents the convolution operation;

[0259] S620, performs image reconstruction through pulse inverse mapping;

[0260] wherein the calculation formula for image reconstruction is as follows:

[0261] ;

[0262] ;

[0263] wherein represents the time decay function, represents the decay duration, = 25 ms, represents the reconstruction time step, = 15, represents the reconstructed image;

[0264] S630, image post - processing: performs detail enhancement and color correction on the reconstructed image;

[0265] Detail enhancement:

[0266] ;

[0267] wherein represents the Laplacian sharpening kernel, represents the enhancement intensity, = 0.3, represents the enhanced image, wherein represents the attention weighting;

[0268] Color correction:

[0269] ;

[0270] wherein represents the color correction matrix, , represents the color - corrected image;

[0271] S640, Output the final image: The output format of the final image is a lossless PNG image;

[0272] ;

[0273] where represents the finally output image, and the output format of the image is a lossless PNG image (H×W×3, 8bit / channel), represents the Sigmoid function, represents the operation of restricting the weight within a specified range.

[0274] The above finally output image also needs quality verification:

[0275] ;

[0276] where represents brightness consistency, , represents contrast retention, , and are respectively the squares of the standard deviations of the pixel values x in the image and the pixel values y in the image, and respectively represent the means of the pixel values x in the image and the pixel values y in the image, represents covariance.

[0277] Through the above image lossless fusion method, the following SF-Net full process example is given: Multi-modal fusion in a night vision driving scenario;

[0278] Step 1, Input data preparation:

[0279] Input visible light image IRGB (1280×720 pixels, 30fps), infrared image IIR (640×480 pixels, 25fps), lidar point cloud D (100,000 points / frame);

[0280] Preprocessing:

[0281] Visible light normalization:

[0282] =( −[125.3,116.2,110.1]) / [63.0,61.5,65.2];

[0283] Infrared dynamic compression: =( −500) / 3500;

[0284] Depth map generation: =Point Cloud Depth(D, 50m), where Point Cloud Depth represents point cloud extraction;

[0285] Step 2, Bio-inspired Feature Extraction:

[0286] Pulse Coding:

[0287] Output pulse sequence of Layer L1 ;

[0288] Membrane potential parameters: = 8ms, = 1.5;

[0289] Feature Pyramid:

[0290] Multi-scale features ;

[0291] Step 3, Dynamic Weight Adjustment:

[0292] Initial weights: = [0.45(RGB), 0.35(IR), 0.20(D)];

[0293] After dynamic update: = [0.62, 0.28, 0.10] (temporary failure of infrared sensor);

[0294] Step 4, Modal Conflict Detection:

[0295] Feature similarity: = 0.32 (significant difference under low light);

[0296] Conflict determination: ∣ ∣ = 0.34 > = 0.3, then it is determined that a conflict exists;

[0297] Step 5, Feature Fusion:

[0298] Corrected weights: = [0.58, 0.32, 0.10] (regulated by pulse competition inhibition);

[0299] Fusion output: ;

[0300] Step 6, Output Generation: Comparison of night vision fusion effects (left: visible light input, middle: infrared input, right: fusion result);

[0301] Quality index: SSIM = 0.927 (benchmark value 0.913).

[0302] In one embodiment of the present invention, the result of the key data flow verification in the above example is as Figure 3 shown. In this example, pedestrian recognition at a distance of 200 meters was successfully achieved in real vehicle testing (false negative rate < 0.1%), verifying the effectiveness of the process.

[0303] As Figure 2 shown, the present invention also proposes an image lossless fusion system based on deep feature extraction, including the following modules:

[0304] Multimodal acquisition and preprocessing module 101: Collect multimodal input data from different sensors and preprocess each modality in the multimodal input data;

[0305] Biological pulse feature encoding module 102: Implement retinal bionic encoding with a hierarchical pulse neuron network, extract a multi-scale spatio-temporal feature pyramid, and construct cross-layer feature associations gated by pulse density;

[0306] Dynamic weight optimization module 103: Introduce an adaptive weight adjustment mechanism to dynamically adjust the weights according to the performance of the sensor modality in a specific environment;

[0307] Modal conflict detection module 104: Implement a modal conflict detection algorithm to identify conflicts between different sensor modalities, and use the algorithm to analyze the consistency of the modal outputs to determine whether there are conflicts;

[0308] Cross-modal fusion decision module 105: Apply an improved fusion strategy. During the fusion process, use dynamic weight allocation and confidence-weighted fusion to maximize the influence of important modalities on the final result;

[0309] Pulse inverse mapping reconstruction module 106: Decode the compressed features into preliminary image data, reconstruct the image using the pulse timing information, and after detail enhancement and color correction, output the final image.

[0310] One embodiment of the present invention also proposes a storage medium storing non-transitory computer-readable instructions for executing one or more steps in the foregoing image lossless fusion method based on deep feature extraction.

[0311] The computer program can be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with other hardware or as part of other hardware, but can also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.

[0312] The embodiments of the present invention have been described above. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms, all of which fall within the protection scope of the present invention.

Claims

1. A lossless image fusion method based on deep feature extraction, characterized in that: The following steps are involved: S100, input data preparation: collecting multimodal input data from different sensors and preprocessing each modality in the multimodal input data, where the multimodal input data includes a visible light image, an infrared image, and a depth map; S200, Bio-Inspired Feature Extraction: Use a multi-layer spiking neural network to simulate the spiking characteristics of biological neurons, extract temporal features from multi-modal input data, and capture feature representations at different levels; S300, dynamic weight adjustment mechanism: introduces an adaptive weight adjustment mechanism to dynamically adjust the weight according to the performance of the sensor modality in a specific environment; S400, modal conflict detection: implementing a modal conflict detection algorithm to identify conflicts between different sensor modalities, and using the algorithm to analyze the consistency of modal outputs to determine whether there is a conflict; S500, feature fusion: Apply an improved fusion strategy. During the fusion process, dynamic weight allocation and confidence-weighted fusion are used to maximize the impact of important modalities on the final result. S600, output generation: decoding the compressed features into preliminary image data, reconstructing the image using the pulse timing information, and outputting the final image after performing detail enhancement and color correction.

2. The image lossless fusion method based on deep feature extraction according to claim 1 is characterized in that: The multi-layer spiking neuron network includes an L1 layer and an L2 layer, wherein the L1 layer is used for simulating retinal bipolar cells and the L2 layer is used for simulating ganglion cells.

3. The image lossless fusion method based on deep feature extraction according to claim 2 is characterized in that: In the L1 layer, the membrane potential equation is as follows: ; in Represents the pulse characteristic output, represents the short-term memory constant, =10ms, represents the edge detection convolution kernel, , represents the membrane potential of neurons in layer L1, represents the rate of change of membrane potential in layer L1 over time, Represents the convolution operation; Pulse emission conditions: ; in represents the high sensitivity threshold, , represents the pulse output of L1 layer at time t, represents the membrane potential of L1 layer at time t; In the L2 layer, the dynamic synaptic equation is as follows: ; in represents the synaptic decay constant, =30ms, where represents the impulse response strength, =0.8, represents the synaptic current of layer L2, Indicates the number of pulses received, represents the timestamp of the kth pulse, Indicates the current time; The feature output is as follows: ; in represents the feature output of the L2 layer, represents the synaptic current of L2 layer at time t.

4. The image lossless fusion method based on deep feature extraction according to claim 3 is characterized in that: In S200, when extracting time series features from multimodal input data, STDP weight adjustment is required, and the weight update rule is as follows: ; ; in ,in represents the learning rate, =20ms, where represents a time-dependent window, where Indicates the pulse timing difference, represents the weight change, represents the pulse output of the front neuron i, represents the pulse output of neuron j, represents the spike time of the front neuron i, represents the spike time of the later neuron j; Constraints: ; in Indicates the upper limit of weight, =2.5, represents the connection weight between neurons i and j, Represents the operation of restricting weights to within a specified range.

5. The image lossless fusion method based on deep feature extraction according to claim 4 is characterized in that: In S200, the pulse density gating mechanism is used, and the pulse density is calculated as follows: ; ; in represents the adjustment coefficient, =1.5, represents the base density, =0.6, represents the Sigmoid function, represents the time window length, =8, represents the total number of neurons, , represents the dynamic fusion weight, represents the pulse output of the nth neuron in the L1 layer at time t, represents the pulse density of L1 layer; ; in represents the fused features, Represents the feature output of the L1 layer; Represents the feature output of the L2 layer; Pooling by downsampling spikes: ; in represents the pyramid level, ,in represents the scale attenuation coefficient, , represents the features after downsampling at the lth layer, represents the maximum pooling operation, Indicates the pooling kernel size; Concatenate by upsampling features: ; in Indicates the channel splicing operation, represents the feature pyramid output, represents the upsampling operation, Indicates the upsampling scale factor.

6. The image lossless fusion method based on deep feature extraction according to claim 5 is characterized in that: In S300, the following steps are specifically included: S310, sensor output confidence monitoring: calculating the confidence of each sensor modal output; S320, initial weight allocation: assigning an initial weight to each modality based on historical data of sensor performance; S330, dynamic weight update: updating weights in real time based on sensor output confidence; S340, Bio-inspired Feedback Regulation Mechanism: Introducing conflict detection based feedback regulation to optimize weight oscillation problem.

7. The image lossless fusion method based on deep feature extraction according to claim 6 is characterized in that: In S400, the calculation formula of the modal conflict detection algorithm is as follows: Conflict determination conditions: ; in represents the weight difference threshold, =0.4, is the weight of the jth neuron at time t, is the weight of the ith neuron at time t, Indicates the conflict determination result; Based on the conflict determination results, the conflict is resolved by constructing a pulse gating function: ; in represents the pulse density of the ith neuron, where represents the pulse density of the jth neuron, Indicates the zero-proof constant, Represents the conflict adjustment coefficient.

8. The image lossless fusion method based on deep feature extraction according to claim 7 is characterized in that: In S500, the calculation formula for dynamic weight allocation is as follows: ; ; in represents the modified dynamic weight, represents the total number of neurons; The calculation formula of confidence weighted fusion is as follows: ; ; in represents the enhancement function of the feature, represents the learnable convolution kernel, , represents the bias term, , represents the fused features, where represents attention weighting; ; in represents the affine transformation network, represents the parameters learned in STDP, represents the pulse characteristic output of the i-th neuron, , Represents the feature output of the ith neuron after alignment.

9. The image lossless fusion method based on deep feature extraction according to claim 8, characterized in that: In S600, the calculation formula for decoding the compressed features into preliminary image data is as follows: ; in represents the deconvolution kernel, ,in =3 corresponds to RGB channels, represents the decoding bias term, , represents the ReLU function, represents the decoded features, represents the compressed features of the final output, Represents the convolution operation; The calculation formula for image reconstruction is as follows: ; ; in represents the time decay function, Indicates the decay time. =25ms, represents the reconstruction time step, =15, represents the reconstructed image; The calculation formula for detail enhancement is as follows: ; in represents the Laplace sharpening kernel, Indicates increased strength, =0.3, represents the enhanced image; The color correction formula is as follows: ; in represents the color correction matrix, , represents the color corrected image; The final image is calculated as follows: ; in Represents the final output image.

10. A lossless image fusion system based on deep feature extraction, characterized in that: Executing the steps in the method for lossless image fusion based on deep feature extraction as described in any one of claims 1 to 9, comprising: Multimodal acquisition and preprocessing module (101): collects multimodal input data from different sensors and preprocesses each mode in the multimodal input data; Biological pulse feature coding module (102): hierarchical pulse neuron network realizes retinal bionic coding, extracts multi-scale spatiotemporal feature pyramid, and constructs pulse density-gated cross-layer feature association; Dynamic weight optimization module (103): introduces an adaptive weight adjustment mechanism to dynamically adjust the weight according to the performance of the sensor modality in a specific environment; Modal conflict detection module (104): implements a modal conflict detection algorithm to identify conflicts between different sensor modalities, and uses the algorithm to analyze the consistency of modal outputs to determine whether there is a conflict; Cross-modal fusion decision module (105): applies an improved fusion strategy. During the fusion process, dynamic weight allocation and confidence-weighted fusion are used to maximize the impact of important modalities on the final result. Pulse inverse mapping reconstruction module (106): decodes the compressed features into preliminary image data, reconstructs the image using pulse timing information, performs detail enhancement and color correction, and then outputs the final image.

Citation Information

Patent Citations

  • Multi-modal visual information fusion method and system based on modal feature constraint

    CN117456332A

  • Time-space domain feature dynamic target identification method based on spiking neural network

    CN118823484A