Multi-modal pathological image fusion method based on cross-cancer data enhancement

Through the multimodal pathological image fusion method enhanced across cancer data, the original physical signals of H&E and MPM are directly fused, solving the problems of visual differences and fuzzy nuclear development in the existing technology, and achieving high-fidelity and multi-dimensional pathological image support, which is suitable for the clinical application of multi-photon microscopy technology.

CN120339085APending Publication Date: 2025-07-18FUJIAN NORMAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510424262.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively fuse the nuclear morphological information of hematoxylin-eosin (H&E) staining images with matrix functional information of multiphoton microscopy (MPM) images without relying on virtual staining, resulting in large visual differences, blurred nuclear development and additional nuclear staining steps, weakening the advantage of label-free.

Method used

A multimodal pathological image fusion method based on cross-cancer data augmentation is adopted, through an end-to-end deep learning framework, a dual-branch network architecture (U-Net+ dense connection layer) combined with a dynamic weighted attention mechanism is used to directly fuse the original physical signals of H&E and MPM, and the model generalization ability is improved through a multi-center data set across cancer types and a two-step training strategy.

Benefits of technology

It realizes that the matrix functional information of MPM images is embedded on the basis of retaining the development accuracy of H&E staining images, and generates multi-dimensional pathological image data. It is compatible with conventional FFPE production processes, lowers the threshold for clinical application, and improves the complementary characterization ability of tumor microenvironment information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339085A_ABST
    Figure CN120339085A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal pathological image fusion method based on cross-cancer data enhancement. The method comprises the following steps: acquiring a staining bright field Hamp of the same pathological tissue; e, dyeing an image and a multi-photon microscopic MPM image; respectively extracting the Hamps through a multi-scale feature extraction module; e, dyeing cell nucleus morphological characteristics of the image and matrix functional characteristics of the MPM image; based on a dynamic weight distribution method of sparsity and noise immunity evaluation, fusing the morphological features of the cell nucleus and the functional features of the matrix to generate a fused feature map; decoding the fusion feature map to generate a fusion image, wherein the fusion image retains Hamp; e, the cell nucleus developing precision of the dyeing image is improved, and matrix function information of the MPM image is embedded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image processing and computer vision, and particularly relates to a multi-modal pathological image fusion method based on a deep learning framework, aiming to enhance the complementary representation ability of nuclear morphology and stromal function information in the tumor microenvironment through cross-cancer type data augmentation and dynamic weighted feature fusion technologies. This solution focuses on the algorithm-level fusion of multi-modal images (H&E stained bright-field images and multi-photon microscopy images), and solves problems such as cross-modal information distortion and insufficient generalization ability in the prior art by improving the network architecture (U-Net + dense connection layer), optimizing the training strategy (two-step cross-cancer type learning), and loss function design (MSE + SSIM + TV), belonging to the category of artificial intelligence-driven medical image analysis technology. Background Art

[0002] Hematoxylin and eosin (H&E) staining is the gold standard for clinical pathological diagnosis, providing an intuitive basis for tumor morphological assessment through two-color contrast development (nuclei appear blue-purple, cytoplasm and stroma appear pink). However, H&E staining has limited ability to analyze extracellular matrix components in the tumor microenvironment, and special tumor types require immunohistochemistry or molecular detection for auxiliary diagnosis, significantly increasing time and economic costs.

[0003] Multiphoton microscopy (MPM), as a label-free and high-resolution non-linear optical imaging technology, can reveal three-dimensional functional information of the tumor microenvironment through endogenous fluorescence characteristics (such as second harmonic signals of collagen fibers and two-photon excited fluorescence of cell metabolites). Since the two-photon excited fluorescence microscopy technology was proposed in 1990, MPM has shown potential in fields such as intraoperative pathological assessment, but its clinical promotion still faces two major challenges: First, the visual difference between MPM images and H&E staining familiar to clinicians is significant, and clinicians need to re-adapt to the interpretation method; Second, MPM has insufficient imaging ability for cell nuclei, and the morphological characteristics of cell nuclei (such as atypia and mitotic figures) are the core basis for tumor diagnosis.

[0004] The prior art attempts to convert MPM images into an H&E style through virtual staining to bridge the visual gap, but due to the lack of nuclear information in the MPM signal, virtual staining easily leads to blurred nuclear imaging and requires an additional nuclear staining step, weakening its label-free advantage. Therefore, how to fuse the nuclear morphological information of H&E and the stromal function information of MPM without relying on virtual staining while maintaining the interpretation habits of clinicians has become a technical problem to be solved urgently. Summary of the Invention

[0005] In view of the defects and deficiencies existing in the prior art, the present invention provides a multi-modal pathological image fusion method based on cross-cancer type data augmentation. Through an end-to-end deep learning framework, the nuclear morphological features of hematoxylin-eosin (H&E) stained images are deeply fused with the tumor microenvironment matrix function information of multi-photon microscopy (MPM) images. This solution innovatively adopts a dual-branch network architecture (U-Net + dense connection layer) combined with a dynamic weighted attention mechanism to directly fuse the original physical signals of H&E and MPM, avoiding the signal distortion risk of virtual staining technology. At the same time, through a cross-cancer type multi-center data set (covering 15 types of malignant tumors) and a two-step training strategy (pre-training + fine-tuning), the generalization ability of the model to tumor heterogeneity is significantly improved. On the basis of retaining the imaging accuracy of H&E stained images, the fusion result embeds three-dimensional function information such as collagen fibers unique to MPM, generates multi-dimensional pathological image data, is compatible with the conventional FFPE section preparation process, provides high-fidelity and multi-dimensional information support for pathological analysis, and promotes the application of multi-photon microscopy technology in the field of medical image processing.

[0006] The technical solution specifically adopted by the present invention to solve its technical problems is as follows:

[0007] A multi-modal pathological image fusion method based on cross-cancer type data augmentation includes the following steps:

[0008] Obtain the hematoxylin-eosin (H&E) stained bright-field image and the multi-photon microscopy (MPM) image of the same pathological tissue;

[0009] Respectively extract the nuclear morphological features of the H&E stained image and the matrix function features of the MPM image through a multi-scale feature extraction module;

[0010] Based on a dynamic weight allocation method for sparsity and noise resistance evaluation, fuse the nuclear morphological features and the matrix function features to generate a fusion feature map;

[0011] Decode the fusion feature map to generate a fusion image, where the fusion image retains the nuclear imaging accuracy of the H&E stained image and embeds the matrix function information of the MPM image.

[0012] Further, the multi-scale feature extraction module includes:

[0013] A U-Net structure for extracting multi-scale context features;

[0014] A dense connection convolutional layer for strengthening feature transfer and fusion;

[0015] The feature extraction branches of the H&E stained image and the MPM image adopt a symmetric network structure with weight sharing.

[0016] Further, the dynamic weight allocation method includes:

[0017] Calculate the sparsity weight of the nuclear morphological features;

[0018] Calculate the noise resistance weight of the matrix functional features;

[0019] Weightedly fuse the sparsity weight and the noise resistance weight to generate a fused feature map.

[0020] Furthermore, the model is trained through the following steps:

[0021] Pre-training stage: Use H&E-MPM image pairs of a single cancer type for initial training;

[0022] Fine-tuning stage: Optimize the model on a cross-cancer dataset containing multiple malignant tumors to improve the generalization ability.

[0023] Furthermore, the training process of the fused image generation jointly optimizes the following loss functions:

[0024] Mean squared error MSE loss;

[0025] Structural similarity SSIM loss;

[0026] Total variation TV loss.

[0027] Furthermore, the H&E stained bright-field image and the MPM image are from the same formalin-fixed paraffin-embedded FFPE tissue section, and the fused image meets the clinical pathological interpretation criteria; the H&E stained bright-field image and the MPM image are obtained by co-localization imaging technology to ensure spatial alignment.

[0028] Furthermore, the fused image is used to replace immunohistochemistry IHC staining and special staining.

[0029] Furthermore, the H&E stained image and the MPM image are original physical signals and are not subjected to virtual staining processing;

[0030] The nuclei of the fused image are developed in accordance with the H&E staining standard and do not contain artifacts generated by virtual staining.

[0031] And, an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method described above when executing the program.

[0032] A non-transitory computer-readable storage medium, on which a computer program is stored, and the steps of the method described above are implemented when the computer program is executed by a processor.

[0033] Compared with the prior art, the present invention and its preferred solutions at least include the following beneficial effects:

[0034] Cross-modal physical fusion breakthrough: By directly fusing the original physical signals of H&E and MPM, the problem of distorted nuclear imaging in virtual staining is avoided. Based on the feature fusion framework and spatial registration strategy of the attention mechanism, while retaining the biological authenticity and imaging accuracy of the H&E stained image, the matrix function information of MPM is completely embedded, realizing the complementary representation of nuclear morphology and microenvironment function within a single section.

[0035] Improved cross-cancer generalization ability: Based on a cross-cancer multi-center dataset covering 15 types of malignant tumors and a two-step training strategy (pre-training + fine-tuning), the adaptability of the model to tumor heterogeneity is significantly enhanced, breaking through the generalization limitations of traditional single-cancer models.

[0036] End-to-end fusion accuracy optimization: The dual-branch network architecture (U-Net + dense connection layer) combined with the dynamic weighting mechanism is used to specifically extract and fuse the nuclear details of H&E and the matrix texture features of MPM, taking into account the capture of multi-scale context information and noise suppression, and improving the structural consistency of the fused image.

[0037] Seamless compatibility with clinical processes: Directly use conventional FFPE tissue samples for fusion processing, without additional staining or operation procedures, expanding multi-dimensional information while maintaining the doctor's reading habits and reducing the technical application threshold. Description of the Drawings

[0038] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:

[0039] Figure 1 It is the image acquisition flow chart of the embodiment of the present invention.

[0040] Figure 2 It is the image fusion flow chart of the embodiment of the present invention.

[0041] Figure 3 It is the image fusion network structure diagram of the embodiment of the present invention.

[0042] Figure 4 It is the model step-by-step training flow chart of the embodiment of the present invention.

[0043] Figure 5 It is the effect comparison diagram of different training strategies of the embodiment of the present invention.

[0044] Figure 6 It is the result comparison diagram of the fusion effect of the embodiment of the present invention and the existing model.

[0045] Figure 7 It is the application example diagram of the fusion result of the embodiment of the present invention. Specific Embodiments

[0046] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:

[0047] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.

[0048] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0049] In view of the defects and deficiencies of the prior art, the embodiments of the present invention propose an innovative solution: by image fusion technology, the MPM image is combined with the H&E stained image. The H&E stained image and the MPM image are complementary in imaging content. Through the image fusion method, the advantages of the MPM technology in label-free imaging can be fully utilized, while retaining the accuracy of the H&E staining in nuclear imaging. To verify the effectiveness of this method, this embodiment constructs a cross-cancer H&E and MPM dataset involving 15 types of cancers and develops a deep learning image fusion model. By fusing the H&E stained image and the MPM image, the information abundance obtained by pathologists from H&E stained sections can be greatly improved, providing convenience for pathologists without increasing the clinical slide preparation process and the learning cost of doctors, and promoting the clinical application of the MPM technology.

[0050] In this embodiment, all samples are prepared as conventional formalin-fixed paraffin-embedded (FFPE) tissue sections to facilitate their conversion into pathological sections, as Figure 1 shown. Each FFPE tissue sample is cut into 5-μm thick sections, and after H&E staining, standard bright-field microscopy imaging and MPM imaging are performed respectively.

[0051] Bright-field images of H&E-stained whole-slide images (WSIs) were scanned at 40× magnification using a commercial digital slide scanner. MPM images were acquired from a commercial upright laser confocal microscope (LSM880, Zeiss, Germany) equipped with an external titanium-sapphire femtosecond laser. The femtosecond laser had a repetition rate of 80 MHz, a pulse width of 140 fs, and an output wavelength tunable between 690 nm and 1064 nm. The plano-convex objective lens was (20×, numerical aperture NA = 0.8, Zeiss, Germany). To obtain the best imaging effect, the excitation wavelength of the laser was set to 810 nm. Depending on the tissue type, the excitation power range of TPEF (two-photon excited fluorescence) was approximately 5 mW to 80 mW, while that of SHG (second harmonic generation) was approximately 25 mW to 200 mW. The TPEF signal was collected in the range of 428 to 695 nm using a 32-channel GaAsP photomultiplier tube (PMT) array detector, and the SHG signal was collected in the range of 395 to 415 nm using a PMT. Images were acquired at a speed of 1.54 microseconds per pixel, and the acquisition time for a single-frame image (512×512 pixels) in the bidirectional scanning mode was 1.89 seconds. All images were automatically recorded and stitched by Zeiss software after averaging twice. The overlap rate between adjacent frames was 5%, and the pixel depth of the images was 12 bits. In addition, the TPEF images were encoded as red, and the SHG images were encoded as green to enhance the contrast.

[0052] Current image fusion mainly focuses on specific fusion tasks such as infrared and visible light image fusion, multi-focus image fusion, and PET and MRI image fusion. Existing fusion models are mainly designed for these specific tasks. Although there are some general models that can be compatible with these several fusion tasks, since the models are not designed and developed according to the characteristics of MPM images and H&E-stained images, these models perform poorly when fusing H&E-stained images and MPM images. Therefore, in this embodiment, a model specifically for the fusion task of MPM and H&E-stained images is developed. Figure 2 This is the image fusion flowchart of the present invention. This unsupervised multi-channel image fusion model includes three key parts: an encoder, a feature fusion, and a decoder, which work together to integrate morphological and functional information. The model is trained for three fusion tasks respectively: (Task 1: Fusion of H&E-stained images and MPM (image of the superposition mode of SHG and TPEF) images; Task 2: Fusion of H&E-stained images and SHG images; Task 3: Fusion of H&E-stained images and TPEF images).

[0053] The detailed structure diagram of the model proposed by the present invention is as Figure 3As shown, the model adopts a dual-branch structure, aiming to simultaneously utilize the advantages of U-Net in capturing multi-scale context information and detailed information, as well as the ability of dense connected convolutional layers in strengthening feature transfer and fusion. It separately processes the H&E stained image (I A ) and the MPM image (I B ), and merges the features of the two branches through a feature fusion layer, finally outputting a fused image. The input source images for the three tasks are shown in Table 1 below:

[0054] Table 1

[0055] Task Name <![CDATA[Input image I A > <![CDATA[Input image I B > Task 1 H&E MPM (Superimposed Image of SHG and TPEF) Task 2 H&E SHG Task 3 H&E TPEF

[0056] Channel A contains a U-Net structure and dense connected convolutional layers. In the U-Net structure, the U-Net encoder uses 3x3 convolutional layers and the LReLU activation function for feature extraction, and then performs downsampling through a 2x2 max pooling layer to capture multi-scale context information. Skip connections directly transfer the feature maps of the U-Net encoder to the U-Net decoder to preserve spatial information and enhance feature fusion. The U-Net decoder then upsamples through 2x2 bilinear interpolation and concatenates with the feature maps of the corresponding layers of the U-Net encoder, and finally outputs feature maps through 3x3 convolutional layers. The dense connected convolutional layer in Channel A starts from the C11 layer, uses 3x3 convolutional kernels and 16 output channels to extract shallow features, and the dense connected blocks composed of D11 - D41 layers strengthen feature transfer and fusion to capture more complex image features. The feature maps extracted by the U-Net structure and the dense connected convolutional layers are concatenated and sent to the feature fusion layer, further enhancing the model's comprehensive processing ability for multi-level image features.

[0057] Channel B has the same structure as Channel A, but reduces the model parameters and computational complexity through weight sharing.

[0058] After being processed by the feature extraction layer of the encoder, the features of I A and I B are respectively extracted as EI A and EI B . The feature fusion layer uses the L1 norm (L1-Norm) and the L2 norm (L2-Norm) to identify the key features in EI A and EI B , and effectively suppresses noise interference. This process determines the respective weights W1 and W2 by calculating the weighted average of the L1-Norm sum of EI A and EI B , and the calculation formula is shown in Formula (1 - 6). Subsequently, these weighted features generate fused features EI F, see (Equation 7). The fused feature F is fed into the main decoding layer, which consists of five convolutional layers from C2 to C6, and finally generates and outputs the fused image. The detailed structural parameters of each layer in the network are shown in Table 2, including information such as the convolutional kernel size, the number of input channels, the number of output channels, and the type of activation function.

[0059] L1 norm:

[0060] L2 norm:

[0061] x i represents the pixel value in the feature map, and n is the number of channels in the feature map:

[0062]

[0063]

[0064] Among them, EI A and EI B respectively represent the feature maps of I A and I B after passing through the encoder of the input image:

[0065]

[0066] EI F = W1 × EI A + W2 × EI B (Equation 7)

[0067] Table 2 Network structure parameters

[0068]

[0069]

[0070] Loss function

[0071] To enhance the quality of the images generated by the model and retain more details of the original images, this embodiment combines multiple loss functions in the model training. Specifically, the loss function of this embodiment includes three losses: MSE, SSIM, and TV. The MSE loss focuses on pixel-level exact matching, the SSIM loss focuses on visual perception quality assessment and can well retain the structural information of the images, and the TV loss focuses on the spatial smoothness of the images, which helps to reduce noise while maintaining edge and texture details. By combining these three loss functions, the pixel-level accuracy, visual perception quality, and image smoothness can be considered simultaneously during the training of the model, thus obtaining better image fusion results. The definition of the total loss function L is as follows:

[0072] L = Lmse +λ1L ssim +λ 1LTV (Formula 8)

[0073] The MSE loss ensures pixel-level fusion and is defined as follows:

[0074] L mse =||I F -I A ||2+||I f -I B ||2 (Formula 9)

[0075] The SSIM loss helps the model better learn structural information from images and is defined as:

[0076] L ssim =[1-SSIM(I F ,I A )]+[1-SSIM(I F ,I B )] (Formula 10)

[0077] The total variation loss is used to better preserve the gradients in the source image and further eliminate noise and is defined as follows:

[0078] L TV =TV(R1)+TV(R2) (Formula 11)

[0079] where R1(i,j)=I F (i,j)-I A (i,j), R2(i,j)=I F (i,j)-I B (i,j) (Formula 12)

[0080] where ||.|| represents the L2 norm, and R1 and R2 respectively represent the direct differences between the fused image I F and the original images I A and I B i and j represent the abscissa and ordinate of the image pixels.

[0081] The definitions and calculation formulas of the three losses are as follows:

[0082] The MSE loss is one of the most commonly used loss functions. It evaluates the image quality by calculating the average of the sum of the squares of the pixel differences between the predicted image and the real image. However, it is sensitive to outliers (such as noise) and does not consider the visual perception characteristics of the human eye.

[0083]

[0084] Among them, X and Y are the input image and the fused image respectively, and x i and y i are the intensity values of the two images at the i-th pixel point respectively, and n is the total number of pixel points.

[0085] The SSIM loss is a loss function based on the structural similarity index, which considers the similarity in three aspects: the brightness, contrast, and structure of the image. It is more in line with the human visual perception system and can better handle changes in brightness, contrast, and structure.

[0086]

[0087] Among them, X and Y are the input image and the fused image respectively, and u x and u y are the means of images X and Y, and are their variances, and σ xy is their covariance. C1 and C2 are small constants added to avoid the denominator being zero.

[0088] The TV loss is a regularization term of the total variation, which evaluates the smoothness of the image by calculating the differences between each pixel in the image and its adjacent pixels. It can reduce the influence of noise on the result while maintaining the edge and texture details of the image.

[0089]

[0090] Among them, X represents the difference between the input image and the fused image, and X i,j is the intensity value of X at the i, j-th pixel point, and X i+1,j is the intensity value of X at the adjacent pixel point below X i,j and X i,j+1 is the intensity value of X at the adjacent pixel point to the right of X i,j respectively.

[0091] Training of the model:

[0092] (1) Construction of a cross-cancer dataset: To improve the universality of the model, this embodiment constructs a cross-cancer dataset that includes two centers and covers 15 common cancer types. First, conventional H&E staining images are obtained through a digital pathology slide scanner. Then, two senior pathologists mark the target regions (about 2 mm × 2 mm) of the diagnostic feature distributions of each cancer type. Next, co-localized MPM images of these regions of interest (ROIs) are obtained using an MPM microscopy system. The present invention has collected a total of 227 pairs of co-localized ROI images of H&E staining images and MPM images. To reduce memory occupancy during training and expand the number of training samples, this embodiment cuts each image into image patches of 256×256 pixel size, and finally obtains a total of 17,697 image patches.

[0093] (2) Two-step training strategy of the model

[0094] To facilitate the model to quickly learn the complementary features of H&E staining images and MPM images, and at the same time improve the universality of the model on multiple tissues such as cross-cancer types. This embodiment proposes a two-step training strategy, as Figure 4 shown. First, the model is pre-trained in a small dataset of a single cancer type. Since the tissue types of a single cancer type are similar and the image features are relatively unified, it is convenient for the model to learn the complementary features of the two modalities faster. Then, the pre-trained model is fine-tuned in the large cross-cancer dataset for the second time to improve the model's learning ability for the image features of different types of tissue components, and finally an end-to-end automatic fusion model of H&E staining images and MPM images that is cross-cancer and has universality is obtained. During training, the sizes of λ1 and λ2 in the loss function (Formula 8) are both set to 100, the learning rates of the two training stages are both 0.001, and the size of the input image is 256*256. After training, in the model inference stage, the size of the input image can be an integer multiple of 256*256. Figure 5 is a comparison chart of the results of different training strategies. It can be seen from the figure that the fusion effect of the model trained in two steps is significantly better than that directly trained with the cross-cancer dataset.

[0095] Figure 6 is a comparison chart of this model and existing models. This embodiment compares mainstream models for infrared and visible light fusion, multi-focus image fusion, medical image fusion, etc. Among them, FusionDN and U2Fusion are designed for general fusion tasks. However, it can be seen from the figure that the model of the present invention can better retain the fiber features of the MPM image and the nuclear features of the H&E staining image, and the visual effect is also better. For example, although DDcGAN can better retain the nuclear features, the image of the collagen fiber will be distorted. And several other models cannot well retain the nuclear features.

[0096] Figure 7 Application demonstration in the fusion images of lung cancer and breast cancer images. The MPM technique can conveniently image elastic fibers and collagen fibers in lung tissue, and these two types of fibers play a key role in lung cancer diagnosis. Elastic fibers are widely distributed in the alveolar wall, around bronchi, and blood vessel walls, forming a reticular structure that supports the morphology and function of alveoli. Through the distribution image of elastic fibers, changes in the alveolar wall structure, such as the breakage and destruction of elastic fibers, can be observed, which is of great value for differentiating in-situ adenocarcinoma of the lung from invasive adenocarcinoma. Collagen fibers are mainly distributed in the alveolar septum, around bronchi, and around blood vessels, etc., and together with elastic fibers, they form the scaffold structure of alveoli, maintaining the morphology and function of alveoli. By fusing MPM images with H&E staining images, elastic fibers and collagen fibers in the tumor microenvironment of lung cancer can be co-localized and displayed together with tumor cells, providing more information for clinicians to diagnose difficult structures. As Figure 7 shown, the red signal of TPEF can perfectly correspond to the staining of elastic fibers, and elastic fiber staining can be avoided through the fusion image method. Myoepithelial immunohistochemistry plays a key role in breast cancer pathological diagnosis, especially in differentiating in-situ carcinoma, invasive carcinoma, and micro-invasive carcinoma. The characteristic of in-situ carcinoma is the intact myoepithelial layer, and the diagnosis of invasive carcinoma depends on the destruction of the myoepithelial layer, at which time myoepithelial markers are absent around cancer nests. In conventional H&E staining images, it is difficult to identify myoepithelial cells. Especially for micro-invasive carcinoma, since the tiny invasive foci are often hidden in the background of in-situ carcinoma, it is necessary to accurately identify the area where myoepithelial cells are absent. From Figure 7 it can be seen that by comparing the CK5 / 6 immunohistochemistry image and the fusion image, the myoepithelial positive positions corresponding to the CK5 / 6 staining image can be seen, and the red TPEF signal also has a signal emphasis and differentiation degree. In addition, from the H&E-SHG fusion image, it can be seen that the SHG signal can additionally display the signal of the basement membrane.

[0097] The references for the comparison models are as follows:

[0098] 1. DDcGAN infrared and visible light fusion model (Jiayi Ma, et.al., DDcGAN: A Dual-Discriminator Conditional Generative Adversarial Network for Multi-Resolution Image Fusion, IEEE Transactions on Image Processing, 2020, 29: 4980-4995)

[0099] 2. EMFusion Medical Image Fusion Model (Han Xu, Jiayi Ma, EMFusion: An unsupervised enhanced medical image fusion network, Information Fusion, 2021, 76: 177-186)

[0100] 3. FusionDN General Image Fusion Model (Han Xu, et.al., FusionDN: A Unified Densely Connected Network for Image Fusion, Proceedings of the AAAI Conference on Artificial Intelligence, 2020, 34: 12484-12491)

[0101] 4. U2Fusion General Image Fusion Model (Han Xu, et.al., U2Fusion: A Unified Unsupervised Image Fusion Network, IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44: 502-518)

[0102] 5. FusionGAN Infrared and Visible Light Fusion Model (Jiayi Ma, et.al., FusionGAN: A generative adversarial network for infrared and visible image fusion, Information Fusion, 2019, 48: 11-26).

[0103] Based on the same inventive concept, the present invention further provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions. Specifically, it is used to load and execute one or more instructions in the computer storage medium to implement the above method.

[0104] It should be further noted that, based on the same inventive concept, the present invention further provides a computer storage medium, on which a computer program is stored, and the computer program, when run by a processor, executes the above method. The storage medium may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or combined with an instruction execution system, apparatus, or device.

[0105] In the description of this specification, the description with reference to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0106] The above has shown and described the basic principles, main features and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements all fall within the scope of the present disclosure claimed.

[0107] This patent is not limited to the above best implementation mode. Anyone inspired by this patent can obtain various other forms of a multi-modal pathological image fusion method based on cross-cancer type data enhancement. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by this patent.

Claims

1. A multi-modal pathological image fusion method based on cross-cancer type data augmentation, characterized in that It includes the following steps: Obtain the hematoxylin and eosin (H&E) stained bright-field image and the multiphoton microscopy (MPM) image of the same pathological tissue; Extract the nuclear morphological features of the H&E stained image and the stromal functional features of the MPM image respectively through a multi-scale feature extraction module; Based on a dynamic weight assignment method for sparsity and noise resistance evaluation, fuse the nuclear morphological features and the stromal functional features to generate a fused feature map; Decode the fused feature map to generate a fused image, where the fused image retains the nuclear imaging accuracy of the H&E stained image and embeds the stromal functional information of the MPM image.

2. The multi-modal pathological image fusion method based on cross-cancer type data augmentation according to claim 1, wherein: The multi-scale feature extraction module includes: A U-Net structure for extracting multi-scale context features; A densely connected convolutional layer for strengthening feature transfer and fusion; The feature extraction branches of the H&E stained image and the MPM image adopt a symmetric network structure with weight sharing.

3. The multi-modal pathological image fusion method based on cross-cancer type data augmentation according to claim 1, wherein: The dynamic weight assignment method includes: Calculate the sparsity weight of the nuclear morphological features; Calculate the noise resistance weight of the stromal functional features; Weightedly fuse the sparsity weight and the noise resistance weight to generate a fused feature map.

4. The multi-modal pathological image fusion method based on cross-cancer type data augmentation according to claim 1, wherein: Train the model through the following steps: Pre-training stage: Use the H&E-MPM image pairs of a single cancer type for initial training; Fine-tuning stage: Optimize the model on a cross-cancer type data set containing multiple malignant tumors to improve the generalization ability.

5. The multi-modal pathological image fusion method based on cross-cancer type data augmentation according to claim 1, wherein: The training process of generating the fused image jointly optimizes the following loss functions: Mean squared error (MSE) loss; Structural similarity (SSIM) loss; Total variation (TV) loss.

6. A multi-modal pathological image fusion method based on cross-cancer type data augmentation according to claim 1, characterized in that: The H&E stained bright-field image and the MPM image are from the same formalin-fixed paraffin-embedded (FFPE) tissue section, and the fused image conforms to the clinical pathological interpretation standard; the H&E stained bright-field image and the MPM image are obtained through co-localization imaging technology to ensure spatial alignment.

7. A multi-modal pathological image fusion method based on cross-cancer type data augmentation according to claim 1, characterized in that: The fused image is used to replace immunohistochemistry (IHC) staining and special staining.

8. The multi-modal pathological image fusion method based on cross-cancer type data augmentation according to claim 1, wherein: The H&E stained image and the MPM image are original physical signals without virtual staining processing; The nuclear imaging of the fused image conforms to the H&E staining standard and does not contain artifacts generated by virtual staining.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program based on the method according to any one of claims 1-8.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1-8.

Citation Information

Cited By

  • Cell virtual staining method and system, computer equipment and storage medium

    CN121884337A