A Smart Fusion Method for Multispectral Endoscopic Images
By constructing a multi-expert encoder based on the U-Net network and a physics-driven deep learning fusion method, the problem of insufficient fusion of multispectral endoscopic image information was solved, and efficient diagnosis of early lesions was achieved.
Patent Information
- Application Number
- CN202511212857.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Existing multispectral endoscopic imaging technology has difficulty effectively fusing image information from different wavelengths, resulting in insufficient accuracy and clarity in the diagnosis of early lesions.
A deep learning fusion network is constructed using a multi-expert encoder based on the U-Net network architecture, a cross-channel attention mechanism enhanced by medical priors, a lesion-adaptive spectral attention decoder, and a physically driven reflectivity-illuminance decoupling module. This network is then used to intelligently fuse multispectral endoscopic images.
It significantly improves the visual recognition and diagnostic specificity of early lesions. By fusing deep vascular and mucosal texture features across scales, it provides high-contrast and high-fidelity feature input, thereby improving the lesion detection rate.
Smart Images

Figure CN120746866B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing technology, and in particular to a method for intelligent fusion of multispectral endoscopic images. Background Technology
[0002] In multispectral endoscopic imaging, different wavelengths of light exhibit differentiated diagnostic value due to their specific interactions with biological tissues:
[0003] Red light (600-700 nm) has a longer wavelength and a deeper tissue penetration ability (about 1-2 mm), which can highlight deep vascular structures and bleeding areas, and is especially suitable for the identification of chronic ulcers or varicose veins.
[0004] Green light (500-600 nm) is located near the absorption peak of hemoglobin (540-580 nm). It enhances the contrast of the superficial capillary network through selective absorption and is often used to detect vascular abnormalities in inflammation or early tumors.
[0005] Blue light (400-500 nm) has high scattering characteristics due to its short wavelength, and is mainly reflected on the surface of the mucosa (<0.5 mm), which can enhance the fine texture features such as gland openings, micro-erosions or intestinal metaplasia.
[0006] White light, as a broadband composite light, provides baseline color and morphological information of anatomical structures. Through multispectral collaborative analysis (such as spectral reflectance modeling), it can achieve cross-scale fusion of tissue pathological features. For example, combining the deep blood vessel distribution of red light with the surface morphological changes of blue light can improve the specificity of early vascular and mucosal lesion diagnosis. Summary of the Invention
[0007] This disclosure provides a four-channel multispectral endoscopic image intelligent fusion method that combines the images of these four types of light, utilizing their respective advantages to provide more comprehensive information and improve the accuracy and clarity of diagnosis.
[0008] Specifically, the following steps are included:
[0009] S1, Determine the objectives of the image fusion network design;
[0010] S2, based on the U-Net network architecture, introduces a multi-expert encoder, a cross-channel attention mechanism with medical prior enhancement, a lesion-adaptive spectral attention decoder, and a physics-driven reflectivity-illuminance decoupling module to construct a deep learning fusion network architecture with pathological interpretability.
[0011] S3, jointly optimizes image quality and medical prior knowledge to construct a loss function;
[0012] S4 acquires multiple sets of four-channel endoscopic images and performs phased progressive training and optimization, including: physical decoupling module pre-training, expert encoder differential training, and end-to-end fine-tuning of the entire network.
[0013] Furthermore, step S1 specifically includes:
[0014] The U-Net variant network uses four-channel images (white, red, green, and blue) as input and outputs a fused enhanced image. The input image size is H×W×4, and the output image size is H×W×3 in RGB space.
[0015] Furthermore, the deep learning fusion network architecture in step S2 specifically includes:
[0016] (1) Multi-expert encoder:
[0017]
[0018] Each expert encoder contains a domain-specific set of convolutional kernels and an attention module.
[0019] E anatomy Used for extracting anatomical structures of white light channels;
[0020] E vessel Used for vascular analysis in red and green light channels;
[0021] E mucosa : Used for mucosal texture feature analysis in the blue light channel;
[0022] E integration Used for cross-expert feature integration analysis;
[0023] The outputs from each expert are dynamically fused through a gating network:
[0024]
[0025] Among them, E i For the i-th expert encoder, I input For the input image, G i For global feature F global Dynamic gating function:
[0026] ;
[0027] W i Let F be the weight matrix of the i-th expert, and MLP(F) be the weight matrix of the i-th expert. global () represents the global features after processing by the multilayer perceptron. As a normalization factor, sum the index values of all four experts (from j=1 to j=4) to ensure that the sum of all gate values is 1;
[0028] (2) Medical a priori enhanced cross-channel attention:
[0029] An attention enhancement mechanism based on medical priors is introduced, enabling the network to utilize medical prior knowledge to selectively focus on the most diagnostically valuable features in different spectral channels; specifically:
[0030] In addition to the standard Query-Key-Value interaction, it also incorporates an organization-spectral coupling matrix. :
[0031]
[0032] Where Q is the query matrix, K is the key matrix, and V is the value matrix. In the standard attention mechanism, similarity is calculated, where d is the feature dimension. This is to scale the similarity values and prevent gradient vanishing; α is a balance coefficient used to control the influence of medical prior knowledge in attention calculation; softmax() is a normalization function that converts the similarity into a probability distribution (all weights sum to 1).
[0033] Coupling matrix Used to reflect the distinguishability of different tissue types in each spectral channel:
[0034] ;
[0035] Where P(tissue) i |spectral j ): The conditional probability of observing tissue type i given spectral channel j; P(tissue) i ) represents the prior probability of organization type i;
[0036] (3) Lesion-adaptive spectral attention decoder:
[0037] In the decoder, adaptive spectral channel attention that takes into account lesion type is employed:
[0038]
[0039] F c AvgPool(F) is a feature map representing the feature representations of all spatial locations in the current layer. c ) is for feature map F cGlobal average pooling is performed to compress the spatial dimensions (height and width) into a single vector, extracting global contextual information. FC1 is the first fully connected layer, which performs a linear transformation on the pooled feature vector for dimensionality reduction and feature transformation. ReLU is a modified linear unit activation function that introduces non-linearity to enhance the model's expressive power while maintaining computational simplicity and efficiency. FC2 is the second fully connected layer, which further transforms the output of the first layer to generate the initial values for channel attention. σ is the sigmoid activation function, which compresses the output to between 0 and 1, making it suitable as attention weights.
[0040] Among them, D lesion (c) is the channel bias function related to lesion type:
[0041]
[0042] Among them, E lesion W is the lesion encoding vector. lesion Let be the learning mapping matrix, and β be the intensity coefficient;
[0043] This decoder automatically adjusts the importance weight of each spectral channel for different lesion types, including:
[0044] Inflammation: Characterized by capillary dilation and increased weighting of the green light channel;
[0045] Intestinal metaplasia: characterized by changes in the morphology of surface glands and increased weighting of blue light channels;
[0046] Early-stage tumors: Changes occur in both blood vessels and surface morphology, balancing red and blue light channels;
[0047] This allows the decoder to dynamically adjust its feature fusion strategy based on the type of lesion detected, thereby improving the detection sensitivity of specific lesions;
[0048] (4) Physically driven reflectivity-light decoupling module:
[0049] By incorporating physical optics theory into deep learning networks, images can be separated into reflectivity and illumination components through a learnable physically parameterized decomposition model.
[0050]
[0051] I observed (x,y,λ) represents the image intensity observed at position (x,y) and wavelength λ; R(x,y,λ) represents the reflectance component at position (x,y) and wavelength λ, representing the optical properties of the tissue itself; L(x,y,λ) represents the illumination component at position (x,y) and wavelength λ, characterizing the illumination conditions; ⊙ is the element-wise multiplication operator.
[0052] The reflectivity component is generated through a multi-scale feature aggregation network:
[0053]
[0054] F reflectance (x,y) is the reflectance feature map, which is a multi-scale feature representation extracted by a deep network; Conv is a wavelength-specific convolution operation.
[0055] The illumination component is estimated using a model based on local consistency constraints:
[0056]
[0057] in, I is a solver for the Poisson equation. initial (x,y,λ) represents the initial image estimate, M smoothness Structure-aware smoothness mask:
[0058]
[0059] ∇I(x,y) represents the magnitude of the image gradient, with parameters γ and α set to 10.0 and 0.8 respectively, controlling the balance between edge preservation and smoothness;
[0060] Furthermore, a wavelength correlation constraint based on spectral physics is introduced to ensure that the decoupled reflectance physically conforms to the spectral characteristics of tissue, thus avoiding deviations between the mathematical solution and the physical solution.
[0061]
[0062] Where R(x,y,λ) i ) represents the position (x, y) at wavelength λ. i and λ j Reflectance estimation under the following conditions; R ref (x,y,λ i ), R ref (x,y,λ j () is a reference reflectance diagram.
[0063] Furthermore, step S3 specifically includes:
[0064] S31, Perception-based structural similarity loss:
[0065] We introduce medical feature weighting on the traditional structural similarity loss (SSIM) model to adapt to the characteristics of endoscopic images:
[0066]
[0067] Among them, I pred(x,y) represents the value of the image predicted by the model at position (x,y), i.e., the enhanced endoscopic image output by the network; I target (x,y) represents the value of the target image at position (x,y), serving as a reference standard for model learning; w med The weighting function dynamically allocates weights based on the regional medical importance, prioritizing the structural fidelity of key diagnostic regions.
[0068]
[0069] In the formula, the coefficients α, β and γ are optimized using labeled data and set to 0.4, 0.35 and 0.25 respectively;
[0070] Among them, w vessel w texure w edge These are all weighting mapping functions used to assign different levels of importance to different types of medical feature regions in SSIM calculations:
[0071] w vessel (x,y) is the vascular region weight map, which assigns higher weights to vascular structures in the image;
[0072] w texture (x,y) is the organization texture region weight map, which assigns higher weights to regions with important texture information;
[0073] w edge (x,y) is the edge region weight map, which assigns higher weights to the structural edges of tissues and organs;
[0074] S32, Multi-scale gradient difference loss:
[0075] Gradient-aware loss is constructed based on multi-directional gradients to ensure accurate reconstruction of anisotropic structures and preserve microvascular and glandular boundaries:
[0076]
[0077] Among them, G θ For the directional gradient operator, I s For an image at scale s, η s is the scale weight; s represents the scale level of the image, and S represents the total number of scale levels, usually 4, where s=1 is the original resolution, s=2 is downsampling once, s=3 is downsampling twice, s=4 is downsampling three times, and so on.
[0078] S33, Medical Prior Fusion Loss:
[0079] Using pre-trained pathology models to guide feature learning, including:
[0080] a. Loss of vascular topological preservation:
[0081]
[0082] Where F vessel For pre-training the vascular extractor, I pred For the enhanced image predicted by the model, I red For the original red channel image, V pred V red The vascular skeletons were extracted from the predicted image and the red light image, respectively, λ. topo D represents the weighting coefficients of the topology loss. topo The topological distance metric for blood vessels is calculated based on the Hausdorff distance of the refined skeleton.
[0083] b. Loss of mucosal texture consistency:
[0084]
[0085] Among them, I blue The image is the original blue light channel image. GM is the mucosal texture gradient matrix extractor, which ensures that the enhanced image retains the fine structural features of the original blue light channel.
[0086] c. Emphasis on loss in the lesion area:
[0087]
[0088] Among them, M lesion ω is the mask for the lesion area. lesion To ensure accurate representation of key areas of disease, a magnification factor is used.
[0089] S34, Overall Loss Function:
[0090]
[0091] The weighting coefficients [λ1,λ2,λ3,λ4,λ5] are initially set to [0.35, 0.25, 0.2, 0.1, 0.1], and are gradually adjusted using a decay strategy.
[0092]
[0093] Where, λ i (0) represents the initial weights, [0.35, 0.25, 0.2, 0.1, 0.1] mentioned above. The parameter δ... i The weight change rate is controlled, and T is the total training period; this ensures that the network first learns the basic structure and then gradually optimizes the medical feature details.
[0094] Furthermore, the phased approach in step S4 specifically includes:
[0095] (1) Pre-training of the physical decoupling module:
[0096] Only the reflectivity-lighting decoupling module is trained, using synthetic data and physical constraints:
[0097]
[0098] Where R is the reflectivity component, L is the illumination component, and I... input The input image is ⊙, where ⊙ is the element-wise multiplication operator and λ is the input image. s L is the illumination smoothness regularization coefficient. spectral λ is the spectral uniformity regularization term. phys These are the physical constraint weighting coefficients.
[0099] Using the Adam optimizer with a learning rate of 1e-4, the module was trained until it could stably separate tissue characteristics from lighting conditions.
[0100] (2) Differentiated training of expert encoders:
[0101] Freeze the decoupling module and train four expert encoders using a task-specific loss:
[0102] E anatomy Using anatomical structures to segment auxiliary tasks;
[0103] E vessel : Using vessel enhancement and segmentation tasks;
[0104] E mucosa : Using mucosal classification and edge detection tasks;
[0105] E integration : Use MSE loss to integrate features;
[0106] Each expert employs a different learning rate strategy to ensure domain specialization:
[0107]
[0108] Wherein, parameter κ i A learning rate modulation factor specific to each expert;
[0109] lr base The base learning rate represents the initial learning rate baseline value. In the differentiated training strategy, the actual learning rate of each expert encoder is dynamically adjusted based on this base learning rate through a specific modulation factor κi.
[0110] (3) End-to-end fine-tuning across the entire network:
[0111] Unfreeze all modules and perform end-to-end optimization using the full loss function:
[0112]
[0113] A cosine annealing learning rate strategy is adopted, with an initial learning rate of 5e-5 and a minimum learning rate of 1e-6, and several training rounds are performed.
[0114] To address the scarcity and complexity of medical image data and prevent overfitting, a hybrid regularization strategy is introduced:
[0115]
[0116] Simultaneously, a dynamic loss balancing mechanism is used to prevent any one loss term from dominating the training process, ensuring that various medical features receive appropriate learning attention.
[0117]
[0118] Where, σ i η is the moving average of the loss statistic, and η is an adjustment factor of 0.5;
[0119] In the above formula: ||W||2 is L2 norm regularization, also known as weight decay, which is used to control the sum of squares of network weights and prevent the weight values from being too large; ||W||1 is L1 norm regularization, which is used to make the network weights sparse, and some weights will become zero, which is equivalent to feature selection.
[0120] DropPath represents regularization applied to the entire path or layer, rather than a single neuron; rate=0.2 indicates that 20% of the paths are randomly dropped during training; α, β, and γ are coefficients that control the relative importance of various regularization techniques; W represents the network's weight parameters.
[0121] L i σ is the original value of the i-th loss term; i The moving average estimate of the standard deviation or range of variation of the loss term, used for normalization; μ i The moving average estimate of the mean of this loss term serves as a baseline; η is an adjustment factor used to control the strength of the mean shift; exp(-η·(L i -μ i )) is an adaptive term used when the loss value L i Below its mean μ i To reduce the impact of this loss item, while L i Higher than μ i This will enhance its influence.
[0122] Furthermore, the method also includes:
[0123] S0 is the step of performing feature enhancement on each channel image before inputting it into the fusion network.
[0124] Compared with existing technologies, the beneficial effects of this disclosure are: ① The network combines multispectral feature fusion capability with physical interpretability, and can achieve cross-scale fusion of deep blood vessels (red light) and mucosal texture (blue light) through a cross-attention mechanism, significantly improving the visual recognition and diagnostic specificity of early lesions; ② The channel-specific targeted enhancement strategy not only overcomes the suboptimal adaptation problem of a single algorithm to different spectral characteristics, but also provides high-contrast, high-fidelity feature input for subsequent multi-channel fusion through the synergistic optimization of physical and data-driven approaches, significantly improving the visual recognition and quantitative analysis accuracy of early lesions; ③ A composite loss function system is designed for the medical diagnostic characteristics of endoscopic images, enabling the network to first learn the basic structure and then gradually optimize medical feature details; ④ A three-stage progressive training strategy is designed to address the problem of scarce medical image data. Attached Figure Description
[0125] The above and other objects, features and advantages of this disclosure will become more apparent from the more detailed description of exemplary embodiments of this disclosure taken in conjunction with the accompanying drawings, in which the same reference numerals generally represent the same components.
[0126] Figure 1 This is a flowchart illustrating an exemplary embodiment of the present disclosure. Detailed Implementation
[0127] Preferred embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.
[0128] Since a single optical channel may not be able to capture all features, this disclosure provides a four-channel endoscopic image intelligent fusion method. Fusing these images allows doctors to see multiple pieces of information simultaneously, improving the detection rate of lesions, especially early-stage lesions. An exemplary procedure according to this disclosure is attached. Figure 1 As shown, the main steps include:
[0129] 1. Enhanced channel-specific directional features
[0130] The channel-specific targeted enhancement strategy overcomes the suboptimal adaptation problem of single algorithms to different spectral characteristics. Furthermore, through synergistic optimization of physical and data-driven approaches, it provides high-contrast, high-fidelity feature inputs for subsequent multi-channel fusion, significantly improving the visual recognition and quantitative analysis accuracy of early lesions. For the intelligent fusion enhancement of four-channel endoscopic images (white light, red light, green light, and blue light), a deep learning network based on an improved U-Net architecture achieves high-fidelity image enhancement through multimodal feature fusion and physical-driven modeling.
[0131] 2. Define the objectives of the converged network design:
[0132] The four-channel images (size H×W×4) of white light, red light, green light and blue light are used as input to the U-Net variant network, and the output is the fused enhanced image (H×W×3, RGB space).
[0133] 3. Converged network architecture design:
[0134] Based on the above spectral characteristic studies, this embodiment constructs a deep learning fusion network MOE-UNet (Mixture of Experts U-Net) with pathological interpretability based on the classic U-Net encoder-decoder structure. Its main distinguishing technical features include:
[0135] (1) Multi-expert encoder architecture:
[0136] Unlike a simple two-branch structure, this embodiment designs a multi-expert coding system with specialized division of labor:
[0137]
[0138] Each expert encoder contains domain-specific sets of convolutional kernels and attention modules:
[0139] E anatomy Focusing on the extraction of anatomical structures in white light channels, based on the ResNeXt backbone network;
[0140] Evessel: Focuses on blood vessel analysis in red and green light channels, introducing adaptive dilated convolution;
[0141] Emucosa: Focuses on the mucosal texture features of the blue light channel and integrates multi-scale residual modules;
[0142] Eintegration: Responsible for cross-expert feature integration, employing a channel-space dual attention mechanism;
[0143] The outputs from each expert are dynamically fused through a gating network:
[0144]
[0145] Among them G i For global feature F global Dynamic gating function:
[0146]
[0147] This expert routing mechanism can dynamically adjust the contribution weights of each expert network based on the content of the input image, adapting to the feature extraction needs of different lesion types.
[0148] (2) Medical a priori enhanced cross-channel attention:
[0149] In the cross-branch interaction attention module, an attention enhancement mechanism based on medical priors is introduced. Specifically, in addition to the standard Query-Key-Value interaction, a tissue-spectral coupling matrix is also incorporated. :
[0150]
[0151] Coupling matrix This reflects the distinguishability of different tissue types across various spectral channels:
[0152]
[0153] This attention mechanism enables the network to utilize prior medical knowledge to more effectively focus on the most diagnostically valuable features in different spectral channels.
[0154] (3) Lesion-adaptive spectral attention decoder:
[0155] In this embodiment, an adaptive spectral channel attention mechanism that considers lesion type is designed in the decoder:
[0156]
[0157] Where D lesion (c) is the channel bias function related to lesion type:
[0158]
[0159] E lesion W is the lesion encoding vector. lesion The learning mapping matrix is denoted by β, which is the intensity coefficient (0.5). The importance weights of each spectral channel are automatically adjusted for different lesion types.
[0160] Inflammation: Enhances the weight of green light channels (capillary dilation characteristic)
[0161] Intestinal metaplasia: Enhances blue light channel weight (changes in surface gland morphology)
[0162] Early-stage tumors: Balancing red and blue light channels (due to changes in both blood vessels and surface morphology)
[0163] This mechanism enables the decoder to dynamically adjust its feature fusion strategy based on the type of lesion detected, thereby improving the detection sensitivity of specific lesions.
[0164] (4) Physically driven reflectivity-light decoupling module:
[0165] This embodiment introduces physical optics theory into a deep learning network, using a learnable physical parameterization decomposition model to separate the image into reflectivity and illumination components:
[0166]
[0167] The reflectivity component is generated through a multi-scale feature aggregation network:
[0168]
[0169] The illumination component is estimated using a model based on local consistency constraints:
[0170]
[0171] in M is a solver for the Poisson equation. smoothness Structure-aware smoothness mask:
[0172]
[0173] The parameters γ and α are set to 10.0 and 0.8 respectively to control the balance between edge maintenance and smoothness.
[0174] In addition, wavelength correlation constraints based on spectral physics are introduced:
[0175]
[0176] This constraint ensures that the decoupled reflectance physically conforms to the tissue spectral characteristics, avoiding deviations between the mathematical solution and the physical solution.
[0177] 4. Loss Function Design
[0178] To address the medical diagnostic characteristics of endoscopic images, this embodiment designs a composite loss function system:
[0179] (1) Perception-based structural similarity loss: Improve the traditional SSIM to adapt to the characteristics of endoscopic images by introducing medical feature-weighted SSIM:
[0180]
[0181] Weight function wmed Based on the dynamic allocation of regional medical importance, priority is given to ensuring the structural fidelity of key diagnostic areas:
[0182]
[0183] The coefficients α, β, and γ were optimized using data annotated by medical experts and set to 0.4, 0.35, and 0.25, respectively.
[0184] w vessel w texure w edge These are all weighting mapping functions used to assign different levels of importance to different types of medical feature regions in SSIM calculations:
[0185] w vessel (x,y) is a vascular region weight map, which assigns higher weights to vascular structures in the image.
[0186] w texture (x,y) is a weighted map of tissue texture regions, which assigns higher weights to regions with important texture information (such as mucosal surface structures, gland openings, etc.).
[0187] w edge (x,y) is a weighted map of the edge region, which gives high weight to the edges of structures such as tissue boundaries and organ outlines.
[0188] (2) Multi-scale gradient difference loss: To accurately preserve the boundaries between microvessels and glands, a gradient sensing loss is designed:
[0189]
[0190] Among them, G θ For the directional gradient operator, I s For an image at scale s, η s The scale weights are [0.4, 0.3, 0.2, 0.1], and the multi-directional gradient ensures accurate reconstruction of anisotropic structures (such as blood vessels);
[0191] 's' represents the image scale level, and 'S' represents the total number of scale levels, typically set to 4. (s=1 is the original resolution, s=2 is downsampling once, s=3 is downsampling twice, s=4 is downsampling three times, and so on.)
[0192] (3) Medical prior fusion loss: Using pre-trained pathology models to guide feature learning, including:
[0193] a. Loss of vascular topological preservation:
[0194]
[0195] Where F vessel For pre-training the vascular extractor, D topo The vascular topological distance metric is calculated based on the Hausdorff distance of the refined skeleton.
[0196] b. Loss of mucosal texture consistency:
[0197]
[0198] GM is a mucosal texture gradient matrix extractor that ensures that the enhanced image retains the fine structural features of the original blue light channel.
[0199] c. Emphasis on loss in the lesion area:
[0200]
[0201] Where M lesion ω is the mask for the lesion area. lesion The magnification factor is 5.0 to ensure accurate representation of key areas of disease.
[0202] (4) Overall loss function:
[0203]
[0204] The weighting coefficients [λ1,λ2,λ3,λ4,λ5] are initially set to [0.35, 0.25, 0.2, 0.1, 0.1], and are gradually adjusted using a decay strategy.
[0205]
[0206] Parameter δ i By controlling the rate of change of weights, and with T being the total training period, this strategy ensures that the network first learns the basic structure and then gradually optimizes the details of medical features.
[0207] 5. Training and optimization strategies:
[0208] To address the scarcity of medical image data, this embodiment designs a three-stage progressive training strategy:
[0209] (1) Physical decoupling module pre-training: First, only the reflectivity-illuminance decoupling head is trained, using synthetic data and physical constraints:
[0210]
[0211] Using the Adam optimizer with a learning rate of 1e-4, train for 50 rounds until the decoupled module can stably separate tissue characteristics from lighting conditions.
[0212] (2) Differentiated training of expert encoders: Freeze the decoupled module and train four expert encoders separately, using task-specific loss:
[0213] E anatomy Using anatomical structures to segment auxiliary tasks
[0214] E vessel Using vessel enhancement and segmentation tasks
[0215] E mucosa Using mucosal classification and edge detection tasks
[0216] E integration : Using MSE loss integration features
[0217] Each expert employs a different learning rate strategy to ensure domain specialization:
[0218]
[0219] Parameter κ i This refers to the learning rate modulation factor specific to each expert. `lrbase` is the base learning rate, representing the initially set baseline learning rate. In differentiated training strategies, the actual learning rate of each expert encoder is dynamically adjusted based on this base learning rate using a specific modulation factor `κi`.
[0220] (3) End-to-end fine-tuning of the entire network: Finally, unfreeze all modules and perform end-to-end optimization using the complete loss function:
[0221]
[0222] A cosine annealing learning rate strategy is adopted, with an initial learning rate of 5e-5 and a minimum learning rate of 1e-6, and training for 100 epochs. To prevent overfitting, a hybrid regularization strategy is introduced:
[0223]
[0224] In this context, ||W||2 represents L2 norm regularization (also known as weight decay), which controls the sum of the squares of network weights to prevent excessively large weight values. ||W||1 represents L1 norm regularization, which makes network weights sparse, with some weights becoming zero, essentially equivalent to feature selection. DropPath(rate=0.2) is a regularization technique similar to Dropout, but applied to the entire path or layer, not just a single neuron; rate=0.2 indicates that 20% of the path is randomly dropped during training. α, β, and γ are coefficients controlling the relative importance of various regularization techniques. W represents the network weight parameters.
[0225] Note: This hybrid regularization strategy combines the advantages of different regularization methods, which can more effectively prevent overfitting and is particularly suitable for fields such as medical images where data is scarce and complex.
[0226] Simultaneously, a dynamic loss balancing mechanism is used to prevent any one loss term from dominating the training process:
[0227]
[0228] Where σ i η is the moving average of the loss statistic, and η is the adjustment factor (0.5).
[0229] L i This is the original value of the i-th loss term. σ i It is a moving average estimate of the standard deviation or range of variation of the loss term, used for normalization. μ i This is a moving average estimate of the mean of the loss term, serving as a baseline. η is an adjustment factor (set to 0.5), controlling the strength of the mean shift. exp(-η·(L i -μ i )) is an adaptive term, when the loss value L i Below its mean μ i When L increases, it reduces the impact of the loss term; i Higher than μ i At times, it will decrease or increase its impact.
[0230] Explanation: This mechanism can dynamically balance the contributions of multiple loss terms, preventing a single loss term from dominating the entire optimization process due to differences in magnitude or rate of change, and ensuring that various medical features (such as blood vessels, glandular structures, etc.) receive appropriate learning attention.
[0231] The fusion network in this embodiment combines multispectral feature fusion capability with physical interpretability. It can achieve cross-scale fusion of deep blood vessels (red light) and mucosal texture (blue light) through cross-attention mechanism. Relying on reflectivity-illuminance decoupling to suppress artifacts, it meets processing requirements under GPU acceleration, which can significantly improve the visual recognition and diagnostic specificity of early lesions.
[0232] The above technical solutions are merely exemplary embodiments of the present invention. For those skilled in the art, based on the application methods and principles disclosed in the present invention, it is easy to make various types of improvements or modifications, and not limited to the methods described in the specific embodiments of the present invention. Therefore, the methods described above are merely preferred and not restrictive.
Claims
1. A method for intelligent fusion of multispectral endoscopic images, characterized in that, Includes the following steps: S1, Determine the objectives of the image fusion network design; S2, based on the U-Net network architecture, introduces a multi-expert encoder, a cross-channel attention mechanism with medical prior enhancement, a lesion-adaptive spectral attention decoder, and a physically driven reflectivity-illuminance decoupling module to construct a deep learning fusion network architecture with pathological interpretability. S3, jointly optimizes image quality and medical prior knowledge to construct a loss function; S4 acquires multiple sets of four-channel endoscopic images and performs phased progressive training and optimization, including: physical decoupling module pre-training, expert encoder differential training, and end-to-end fine-tuning of the entire network. The deep learning fusion network architecture in step S2 specifically includes: (1) Multi-expert encoder: Each expert encoder contains a domain-specific set of convolutional kernels and an attention module. E anatomy Used for extracting anatomical structures of white light channels; E vessel Used for vascular analysis in red and green light channels; E mucosa : Used for mucosal texture feature analysis in the blue light channel; E integration Used for cross-expert feature integration analysis; The outputs from each expert are dynamically fused through a gating network: Among them, E i For the i-th expert encoder, I input For the input image, G i For global feature F global Dynamic gating function: ; W i Let F be the weight matrix of the i-th expert, and MLP(F) be the weight matrix of the i-th expert. global () represents the global features after processing by the multilayer perceptron. As a normalization factor, sum the index values of all four experts to ensure that the sum of all gate values is 1; (2) Medical a priori enhanced cross-channel attention: An attention enhancement mechanism based on medical priors is introduced, enabling the network to utilize medical prior knowledge to selectively focus on the most diagnostically valuable features in different spectral channels; specifically: In addition to the standard Query-Key-Value interaction, it also incorporates an organization-spectral coupling matrix. : Where Q is the query matrix, K is the key matrix, and V is the value matrix. In the standard attention mechanism, similarity is calculated, where d is the feature dimension. This is to scale the similarity values and prevent gradient vanishing; α is a balance coefficient used to control the influence of medical prior knowledge in attention calculation; softmax() is a normalization function that converts similarity into a probability distribution; Coupling matrix Used to reflect the distinguishability of different tissue types in each spectral channel: ; Where P(tissue) i |spectral j ): The conditional probability of observing tissue type i given spectral channel j; P(tissue) i ) represents the prior probability of organization type i; (3) Lesion-adaptive spectral attention decoder: In the decoder, adaptive spectral channel attention that takes into account lesion type is employed: F c AvgPool(F) is a feature map representing the feature representations of all spatial locations in the current layer. c ) is for feature map F c Global average pooling is performed to compress the spatial dimension into a single vector, extracting global contextual information. FC1 is the first fully connected layer, which performs a linear transformation on the pooled feature vector for dimensionality reduction and feature transformation. ReLU is a modified linear unit activation function that introduces non-linearity to enhance the model's expressive power while maintaining computational simplicity and efficiency. FC2 is the second fully connected layer, which further transforms the output of the first layer to generate the initial values for channel attention. σ is the sigmoid activation function, which compresses the output to between 0 and 1, making it suitable as attention weights. Among them, D lesion (c) is the channel bias function related to lesion type: Among them, E lesion W is the lesion encoding vector. lesion Let be the learning mapping matrix, and β be the intensity coefficient; This decoder automatically adjusts the importance weight of each spectral channel for different lesion types, including: Inflammation: Characterized by capillary dilation and increased weighting of the green light channel; Intestinal metaplasia: characterized by changes in the morphology of surface glands and increased weighting of blue light channels; Early-stage tumors: Changes occur in both blood vessels and surface morphology, balancing red and blue light channels; This allows the decoder to dynamically adjust the feature fusion strategy based on the type of lesion detected, thereby improving the detection sensitivity of specific lesions; (4) Physically driven reflectivity-light decoupling module: By incorporating physical optics theory into deep learning networks, images can be separated into reflectivity and illumination components through a learnable physically parameterized decomposition model. I observed (x,y,λ) represents the image intensity observed at position (x,y) and wavelength λ; R(x,y,λ) represents the reflectance component at position (x,y) and wavelength λ, representing the optical properties of the tissue itself; L(x,y,λ) represents the illumination component at position (x,y) and wavelength λ, characterizing the illumination conditions; ⊙ is the element-wise multiplication operator. The reflectivity component is generated through a multi-scale feature aggregation network: F reflectance (x,y) is the reflectance feature map, which is a multi-scale feature representation extracted by a deep network; Conv is a wavelength-specific convolution operation. The illumination component is estimated using a model based on local consistency constraints: in, For the Poisson equation solver, I initial (x,y,λ) represents the initial image estimate, M smoothness Structure-aware smoothness mask: ∇I(x,y) represents the magnitude of the image gradient, with parameters γ and α set to 10.0 and 0.8 respectively, controlling the balance between edge preservation and smoothness; Furthermore, a wavelength correlation constraint based on spectral physics is introduced to ensure that the decoupled reflectance physically conforms to the spectral characteristics of tissue, thus avoiding deviations between the mathematical solution and the physical solution. Where R(x,y,λ) i ) represents the position (x, y) at wavelength λ. i and λ j Reflectance estimation under the following conditions; R ref (x,y,λ i ), R ref (x,y,λ j () is a reference reflectance diagram.
2. The method according to claim 1, characterized in that, Step S1 specifically includes: The U-Net variant network uses four-channel images (white, red, green, and blue) as input and outputs a fused enhanced image. The input image size is H×W×4, and the output image size is H×W×3 in RGB space.
3. The method according to claim 1, characterized in that, Step S3 specifically includes: S31, Perception-based structural similarity loss: We introduce medical feature weighting on the traditional structural similarity loss (SSIM) model to adapt to the characteristics of endoscopic images: Among them, I pred (x,y) represents the value of the image predicted by the model at position (x,y), i.e., the enhanced endoscopic image output by the network; I target (x,y) represents the value of the target image at position (x,y), serving as a reference standard for model learning; w med The weighting function dynamically allocates weights based on the regional medical importance, prioritizing the structural fidelity of key diagnostic regions. In the formula, the coefficients α, β, and γ are optimized using labeled data and set to 0.4, 0.35, and 0.25, respectively; w vessel w texure w edge These are all weighting mapping functions used to assign different levels of importance to different types of medical feature regions in SSIM computation: w vessel (x,y) is the vascular region weight map, which assigns higher weights to vascular structures in the image; w texture (x,y) is the organization texture region weight map, which assigns higher weights to regions with important texture information; w edge (x,y) is the edge region weight map, which assigns higher weights to the structural edges of tissues and organs; S32, Multi-scale gradient difference loss: Gradient-aware loss is constructed based on multi-directional gradients to ensure accurate reconstruction of anisotropic structures and preserve microvascular and glandular boundaries: Among them, G θ For the directional gradient operator, I s For an image at scale s, η s is the scale weight; s represents the scale level of the image, and S represents the total number of scale levels, usually 4, where s=1 is the original resolution, s=2 is downsampling once, s=3 is downsampling twice, s=4 is downsampling three times, and so on. S33, Medical Prior Fusion Loss: Using pre-trained pathology models to guide feature learning, including: a. Loss of vascular topological preservation: Where F vessel For pre-training the vascular extractor, I pred For the enhanced image predicted by the model, I red For the original red channel image, V pred V red The vascular skeletons were extracted from the predicted image and the red light image, respectively, λ. topo D represents the weighting coefficients of the topology loss. topo The topological distance metric for blood vessels is calculated based on the Hausdorff distance of the refined skeleton. b. Loss of mucosal texture consistency: Among them, I blue The image is the original blue light channel image. GM is the mucosal texture gradient matrix extractor, used to ensure that the enhanced image retains the fine structural features of the original blue light channel. c. Emphasis on loss in the lesion area: Among them, M lesion ω is the mask for the lesion area. lesion To ensure accurate representation of key areas of disease, a magnification factor is used. S34, Overall Loss Function: The weighting coefficients [λ1,λ2,λ3,λ4,λ5] are initially set to [0.35, 0.25, 0.2, 0.1, 0.1], and are gradually adjusted using a decay strategy. Where, λ i (0) represents the initial weights, [0.35, 0.25, 0.2, 0.1, 0.1] mentioned above; parameter δ i The weight change rate is controlled, and T is the total training period; this ensures that the network first learns the basic structure and then gradually optimizes the medical feature details.
4. The method according to claim 3, characterized in that, The specific phases in step S4 include: (1) Pre-training of the physical decoupling module: Only the reflectivity-lighting decoupling module is trained, using synthetic data and physical constraints: Where R is the reflectivity component, L is the illumination component, and I... input The input image is ⊙, where ⊙ is the element-wise multiplication operator and λ is the input image. s L is the illumination smoothness regularization coefficient. spectral λ is the spectral uniformity regularization term. phys These are the physical constraint weighting coefficients; Using the Adam optimizer with a learning rate of 1e-4, the module was trained until it could stably separate tissue characteristics from lighting conditions. (2) Differentiated training of expert encoders: Freeze the decoupling module and train four expert encoders using a task-specific loss: E anatomy Using anatomical structures to segment auxiliary tasks; E vessel : Using vessel enhancement and segmentation tasks; E mucosa : Using mucosal classification and edge detection tasks; E integration : Use MSE loss to integrate features; Each expert employs a different learning rate strategy to ensure domain specialization: Wherein, parameter κ i A learning rate modulation factor specific to each expert; lr base The base learning rate represents the initial, benchmark learning rate. In a differentiated training strategy, the actual learning rate of each expert encoder is based on this base learning rate and modulated by a specific modulation factor κ. i Make dynamic adjustments; (3) End-to-end fine-tuning across the entire network: Unfreeze all modules and perform end-to-end optimization using the full loss function: A cosine annealing learning rate strategy is adopted, with an initial learning rate of 5e-5 and a minimum learning rate of 1e-6, and several training rounds are performed. To address the scarcity and complexity of medical image data and prevent overfitting, a hybrid regularization strategy is introduced: Simultaneously, a dynamic loss balancing mechanism is used to prevent any one loss term from dominating the training process, ensuring that various medical features receive appropriate learning attention. Where, σ i η is the moving average of the loss statistic, and η is an adjustment factor of 0.5; In the above formula: ||W||2 is L2 norm regularization, also known as weight decay, which is used to control the sum of squares of network weights and prevent the weight values from being too large; ||W||1 is L1 norm regularization, which is used to make the network weights sparse, and some weights will become zero, which is equivalent to feature selection. DropPath represents regularization applied to the entire path or layer, rather than a single neuron; rate=0.2 indicates that 20% of the paths are randomly dropped during training; α, β, and γ are coefficients that control the relative importance of various regularization techniques; W represents the network's weight parameters. L i σ is the original value of the i-th loss term; i The moving average estimate of the standard deviation or range of variation of the loss term, used for normalization; μ i The moving average estimate of the mean of this loss term serves as a baseline; η is an adjustment factor used to control the strength of the mean shift; exp(-η·(L i -μ i )) is an adaptive term used when the loss value L i Below its mean μ i To reduce the impact of this loss item, while L i Higher than μ i This will enhance its influence.
5. The method according to any one of claims 1-4, characterized in that, Also includes: S0 is the step of performing feature enhancement on each channel image before inputting it into the fusion network.
Citation Information
Patent Citations
Infrared visible image fusion method based on self-supervised feature decoupling
CN116468644A
Multispectral image and hyperspectral image fusion method and system
CN120339016A