Digital pathology virtual hematoxylin and eosin staining method and system based on unsupervised learning DF-GAN model
Patent Information
- Application Number
- CN202610775658.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本发明所要解决的技术问题在于针对上述现有技术中的不足,提供一种无监督学习DF-GAN模型的数字病理虚拟HE染色方法及系统,用于解决现有非配对虚拟染色技术存在的配对样本获取成本高、染色保真度低(易出现颜色漂移、细胞核边缘模糊、核质对比度不足)、高层语义与局部几何结构一致性差(易出现结构错位、细胞核形态受损)、临床级WSI推理存在显存约束和接缝伪影、跨样本染色一致性差的技术问题
一种无监督学习DF-GAN模型的数字病理虚拟HE染色方法,从根源上突破了现有非配对虚拟染色技术的核心瓶颈:无需大量严格配对的临床病理样本,大幅降低了数据获取成本和伦理风险,通过病理先验引导的混合监督体系,解决了纯无监督训练初期颜色漂移严重、收敛不稳定的问题;针对性强化了细胞核等关键诊断结构的保真度,避免了传统方法中核边缘模糊、核质对比度不足的缺陷;同时引入跨模态语义约束,有效防止了结构错位和语义偏离。实现从双通道荧光图像到高保真H&E图像的端到端转换,生成结果能够直接用于临床病理诊断,为计算病理的临床落地提供了可靠的技术支撑。
Smart Images

Figure CN122820876A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical image processing technology, specifically relating to a digital pathological virtual HE staining method and system using an unsupervised learning DF-GAN model. Background Technology
[0002] Fluorescence pathological imaging is a novel non-invasive imaging technique that uses DAPI (4',6-diamidinyl-2-phenylindole) for specific staining of DNA and FITC (fluorescein isothiocyanate) for labeling of substrates such as proteins. It can clearly present intracellular molecular information and tissue structural features, possessing irreplaceable application value in early tumor diagnosis, molecular subtyping, efficacy assessment, and prognosis. Compared to traditional H&E staining, fluorescence pathological imaging provides more molecular-level information, helping to detect early microlesions and occult metastases. However, the visual characteristics of dual-channel fluorescence images differ significantly from those of hematoxylin-eosin (H&E) stained slides, which pathologists have long been accustomed to interpreting. Pathologists require extensive specialized training to interpret these images accurately, and the interpretation efficiency is far lower than that of H&E slides, making it difficult to meet the needs of large-scale clinical diagnosis. Furthermore, the boundary between tumor areas and normal tissue areas is often blurred in dual-channel fluorescence images, further increasing diagnostic difficulty and the risk of misdiagnosis, severely limiting its widespread clinical application. Therefore, transferring dual-channel fluorescence images to the H&E staining domain using virtual staining technology, while preserving molecular information and conforming to the reading habits of pathologists, has become a key approach to solving this problem.
[0003] Unpaired virtual staining technology can achieve cross-domain image conversion without strictly paired samples, significantly reducing the cost of clinical data acquisition and solving problems such as sample scarcity and pairing difficulties, making it a research hotspot in the current field of digital pathology. Mainstream unpaired virtual staining methods include CycleGAN, CUT, and their variants, but these methods all have fundamental limitations and cannot meet the high accuracy requirements of clinical diagnosis. As the most widely used model, CycleGAN achieves unsupervised mapping through a bidirectional generator, discriminator, and cycle consistency loss, but this model is susceptible to color drift, resulting in significant differences in hue and saturation between the generated images and real H&E staining images. Furthermore, the lack of targeted constraints on key pathological structures such as cell nuclei leads to blurred nucleus edges, insufficient nucleocytoplasmic contrast, and even missing nuclear structures in the generated images, severely affecting the accuracy of pathological diagnosis. The CUT method reduces structural drift through fragment-level contrast learning, but overly strong content alignment constraints limit the transferability of staining styles, resulting in pale staining effects and poor layering in the generated images, showing a significant difference from real H&E staining. While its improved version ASP can adapt to weakly paired scenarios, it fails to effectively utilize prior knowledge of pathological structures, cannot guarantee the fidelity of key diagnostic structures such as cell nuclei, and has weak model generalization ability, making it difficult to adapt to samples of different pathological tissue types and different staining batches.
[0004] Furthermore, most existing virtual staining techniques are designed for small image patches and do not consider the inference requirements of clinical-grade whole-slide images (WSI). Actual clinical data is typically stored in multi-resolution, multi-channel WSI format, with single images reaching hundreds of millions of pixels. Directly inputting these images into a model can lead to memory overflow, making inference impossible. Simple block-based inference can produce noticeable mesh artifacts at block boundaries, disrupting image integrity and continuity, and even causing cell nuclear structure breakage, affecting diagnostic accuracy. Simultaneously, due to differences in slide preparation processes, staining conditions, and imaging equipment, significant tonal and staining density drift exists between different samples. Existing technologies have failed to effectively address this issue, resulting in large differences in virtual staining results between different samples and poor consistency in diagnostic outcomes. These technical shortcomings severely restrict the clinical application of virtual staining technology, ensuring that the traditional H&E staining process remains cumbersome, time-consuming (up to several hours), irreversible, and costly, hindering the effective improvement of pathological diagnostic capabilities in primary healthcare institutions. Therefore, there is an urgent need to develop a digital pathology virtual HE staining method that does not require a large number of paired samples, offers high staining fidelity, good semantic and geometric consistency, and is adaptable to clinical-grade WSI inference. Summary of the Invention
[0005] The technical problem to be solved by this invention is to provide a digital pathology virtual HE staining method and system for unsupervised learning DF-GAN model, which addresses the shortcomings of the existing technology. This method solves the technical problems of existing unpaired virtual staining technology, such as high cost of obtaining paired samples, low staining fidelity (easily resulting in color drift, blurred cell nucleus edges, and insufficient nucleocytoplasmic contrast), poor consistency between high-level semantics and local geometric structure (easily resulting in structural misalignment and damaged cell nucleus morphology), memory constraints and seam artifacts in clinical-grade WSI inference, and poor consistency of staining across samples.
[0006] The present invention adopts the following technical solution: A digital pathology virtual HE staining method using an unsupervised learning DF-GAN model includes the following steps: S1. Acquire dual-channel fluorescence pathological images, wherein the dual-channel fluorescence pathological images include DAPI channel images and FITC channel images; S2. Construct a color lookup table (LUT) based on the statistical patterns of the input and target domains. The color lookup table LUT adopts a three-segment structure: the low grayscale segment is mapped to the background color, the medium grayscale segment is mapped to the pinkish-purple series, and the high grayscale segment is mapped to the dark purple to purplish-black series. The dual-channel fluorescent pathological image is mapped through the color lookup table LUT to generate pseudo-H&E staining images, forming a pseudo-paired dataset. S3. Construct a hybrid data system, which includes the pseudo-paired dataset and the unpaired dataset. The unpaired dataset contains dual-channel fluorescence images without corresponding H&E staining images and real H&E staining images without corresponding dual-channel fluorescence images. A phased training mechanism is used to train the DF-GAN model: in the warm-up phase, pseudo-paired data is used as the main training sample; after the warm-up phase, the pairing-guided sampling probability is linearly reduced to a preset lower limit according to the training rounds, and then the model enters the alignment phase, which is mainly composed of unpaired data. S4. Extract the cell nucleus segmentation mask from the DAPI channel of the dual-channel fluorescence pathology image, calculate the global statistics of the cell nucleus based on the cell nucleus segmentation mask, the global statistics include the normalized value of the number of cell nuclei, the proportion of nuclear aggregation, the maximum degree of aggregation, the average area of cell nuclei, and the proportion of cell tissue; calculate the weighting coefficients based on the global statistics, the weighting coefficients are restricted to a preset range; apply the weighting coefficients to the cycle consistency loss of the dual-channel fluorescence domain to construct a dynamic weighted cycle consistency loss; S5. Input the dual-channel fluorescence pathology image and the generated virtual H&E image into the pre-trained CLIP visual encoder respectively, and extract the global semantic features output by the top pooling layer and the intermediate feature representations output by the multiple hidden layers; calculate the semantic consistency loss based on the global semantic features using cosine distance, and calculate the geometric structure preservation loss based on the intermediate feature representations using mean square error; add the semantic consistency loss and the geometric structure preservation loss as joint constraints to the objective function of the DF-GAN model. S6. The trained DF-GAN model is used to infer the input dual-channel fluorescence pathological image and output a virtual H&E staining image.
[0007] Preferably, in step S2, the color lookup table (LUT) is constructed in the following way: statistical analysis is performed on the grayscale distribution of a large number of input DAPI channels and the color distribution of the target H&E image to determine several key grayscale segments and corresponding representative color points, and a complete 256-level color lookup table is generated in each segment using a smooth interpolation method.
[0008] Preferably, in step S3, the pairing guided sampling probability :
[0009] in, Indicates the training round index. Indicates the length of the preheating stage. This is the round at which linear decay ends. This represents the lower bound of the probability that paired guide samples will continue to participate in the update after entering the unpaired main training phase.
[0010] Preferably, in step S4, the cell nucleus segmentation mask is obtained in the following way: The DAPI single-channel image is binarized to obtain a coarse segmentation mask; the coarse segmentation mask is then subjected to 3×3 convolution kernel closing and opening operations to obtain an optimized cell kernel segmentation mask.
[0011] Preferably, in step S4, the weighting coefficient for:
[0012] in, For numerical clipping operators, For the nuclear aggregation ratio, To maximize the degree of reunion, This is the normalized value for the number of cell nuclei. This represents the percentage of cells and tissues.
[0013] Preferably, the dynamic weighted cyclic consistency loss of the dual-channel fluorescence domain for:
[0014] in, The weighting coefficients for the cycling loss of the dual-channel fluorescence domain are... Represents the mathematical expectation. It is an L1 norm. Input image for dual-channel fluorescence domain, To be Through generator Generate virtual H&E images Then, through the generator The reconstructed image is obtained by inverse mapping back to the dual-channel fluorescence domain.
[0015] Preferably, in step S5, the pre-trained CLIP visual encoder is a CLIP model pre-trained on the Quilt-1M pathological image dataset; before inputting into the CLIP visual encoder, random perspective transformation and random cropping data augmentation processing are applied to the dual-channel fluorescence pathological image and the virtual H&E image, and the image resolution is adjusted to the standard input size of the CLIP visual encoder.
[0016] Preferably, the objective function of the DF-GAN model for:
[0017] in, To combat the losses, For the cyclic reconstruction loss of the dual-channel fluorescence image domain, Cyclic consistency loss for H&E staining image domain For semantic consistency loss, Loss is preserved for the geometry.
[0018] Preferably, step S1 further includes the following preprocessing of the dual-channel fluorescence pathological image: channel conversion, region of interest extraction and size regularization of the original dual-channel fluorescence whole-slice image, removal of background noise and invalid regions, to obtain image data that meets the model input requirements.
[0019] Secondly, embodiments of the present invention provide a digital pathology virtual HE staining system based on an unsupervised learning DF-GAN model, comprising: The data acquisition module is used to acquire dual-channel fluorescence pathological images containing DAPI channel images and FITC channel images; The pseudo-paired data generation module is used to construct a three-segment color lookup table (LUT) based on the statistical rules of the input domain and the target domain, and to map the dual-channel fluorescence pathological image through the color lookup table LUT to generate pseudo-H&E staining images, thus forming a pseudo-paired dataset. The hybrid training module is used to construct a hybrid data system containing the pseudo-paired dataset and the unpaired dataset, and to train the DF-GAN model using a phased training mechanism. The phased training mechanism includes a warm-up phase with pseudo-paired data as the main component and an alignment phase with unpaired data as the main component is entered after the pairing-guided sampling probability decreases linearly to a preset lower limit according to the training rounds. The cell nucleus prior weighting module is used to extract the cell nucleus segmentation mask from the DAPI channel, calculate the global cell nucleus statistics including the normalized value of the number of cell nuclei, the proportion of nuclear aggregation, the maximum degree of aggregation, the average area of cell nuclei and the proportion of cell tissue, generate weighting coefficients limited to a preset range, and apply the weighting coefficients to the cycle consistency loss of the dual-channel fluorescence domain. The CLIP constraint module is used to extract global semantic features and intermediate feature representations of multi-layer hidden layer outputs using a pre-trained CLIP visual encoder. Based on the global semantic features, semantic consistency loss is calculated using cosine distance. Based on the intermediate feature representations, geometric structure preservation loss is calculated using mean square error. The two are then added as joint constraints to the objective function of the DF-GAN model. The inference output module is used to infer the input dual-channel fluorescence pathology image using the trained DF-GAN model and output a virtual H&E staining image.
[0020] Thirdly, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the digital pathology virtual HE staining method for the unsupervised learning DF-GAN model described above.
[0021] Fourthly, embodiments of the present invention provide a computer-readable storage medium including a computer program, which, when executed by a processor, implements the steps of the digital pathology virtual HE staining method described above for the unsupervised learning DF-GAN model.
[0022] Fifthly, a chip includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the digital pathology virtual HE staining method for the unsupervised learning DF-GAN model described above.
[0023] In a sixth aspect, embodiments of the present invention provide an electronic device, including a computer program, which, when executed by the electronic device, implements the steps of the digital pathological virtual HE staining method described above using an unsupervised learning DF-GAN model.
[0024] Compared with the prior art, the present invention has at least the following beneficial effects: A novel unsupervised learning DF-GAN model for digital pathology virtual HE staining fundamentally overcomes the core bottleneck of existing unpaired virtual staining techniques: it eliminates the need for a large number of strictly paired clinical pathology samples, significantly reducing data acquisition costs and ethical risks. Through a hybrid supervision system guided by pathological priors, it solves the problems of severe color drift and unstable convergence in the early stages of purely unsupervised training. It specifically enhances the fidelity of key diagnostic structures such as cell nuclei, avoiding the defects of blurred nuclear edges and insufficient nucleocytoplasmic contrast in traditional methods. Simultaneously, it introduces cross-modal semantic constraints to effectively prevent structural misalignment and semantic deviation. It achieves end-to-end conversion from dual-channel fluorescence images to high-fidelity H&E images, and the generated results can be directly used for clinical pathology diagnosis, providing reliable technical support for the clinical application of computational pathology.
[0025] Furthermore, through systematic statistical analysis of the grayscale distribution of a large number of DAPI channels and the color distribution of H&E images, combined with the prior pathological knowledge of the blue-purple nuclear region and pink cytoplasm, key grayscale segments and representative color points were identified, and a 256-level continuously differentiable color map was generated through smooth interpolation. A three-segment structure design of background-transition-nuclear region enhancement was adopted, with low grayscale segments suppressing background noise, medium grayscale segments representing the basic tissue hue, and high grayscale segments enhancing the color difference of cell nuclei. This approach can generate pseudo-paired data that is highly consistent with the semantics of H&E staining, providing stable local rule guidance for the early training of the model and significantly reducing color drift in the early stages of training.
[0026] Furthermore, the warm-up phase uses pseudo-paired data with 100% probability to quickly establish an accurate basic color mapping; the linear decay phase gradually reduces the sampling ratio of paired data to guide the model from local rule learning to global distribution alignment; the stabilization phase maintains paired guidance with a fixed lower bound probability to prevent semantic drift during training; this effectively solves the problems of slow convergence and poor stability in pure unsupervised training, while avoiding the dependence of supervised learning on a large number of paired samples, thus significantly improving training efficiency and generalization ability while ensuring model accuracy.
[0027] Furthermore, a coarse segmentation mask is obtained by binarizing the DAPI single-channel image, followed by optimization through sequential closing and opening operations of 3×3 convolution kernels: the closing operation effectively fills the tiny pores within the cell nucleus, preserving its morphological features; the opening operation removes background noise and minor impurities, improving segmentation accuracy. The global statistics of the cell nucleus extracted from this precise segmentation mask accurately reflect key features such as the density, aggregation, size distribution, and tissue proportion of cell nuclei in the sample. This provides a reliable biological prior for constructing the subsequent dynamically weighted cycle consistency loss, ensuring that the weight coefficients accurately reflect the pathological importance of different samples.
[0028] Furthermore, the weighting coefficients are positively correlated with cell nuclear density, aggregation degree, and tissue complexity, enabling automatic identification of sample regions with greater pathological significance. The coefficients are empirically optimal values optimized for liver tissue pathological data, and can be adapted to different pathological tissue types using a grid search method, demonstrating good versatility. By using the clip operator to restrict the weights to the range [0.7, 1.7], training oscillations and gradient vanishing / exploding problems caused by excessively high or low sample weights are effectively avoided, ensuring the stability of the training process.
[0029] Furthermore, the constraint strength is dynamically adjusted based on the characteristics of cell nuclei in the samples. Stronger cycle consistency constraints are applied to key pathological samples with densely distributed and significantly clustered cell nuclei, forcing the model to prioritize structural fidelity in these regions during optimization. Simultaneously, the unweighted cycle consistency loss of the H&E staining domain is retained, maintaining overall staining style consistency while strengthening key structures. This solves the core problems of blurred cell nucleus edges and insufficient nucleocytoplasmic contrast in existing technologies, significantly improving the accuracy and readability of key diagnostic structures in the generated images.
[0030] Furthermore, a CLIP model pre-trained on the Quilt-1M large-scale pathological image dataset, rather than a general CLIP model, is employed. This allows the model to learn specific semantic features and structural patterns from a large number of pathological images, making it more suitable for medical image processing tasks. A unified random perspective transformation and random cropping data augmentation strategy is applied to the input dual-channel fluorescence image and the generated virtual H&E image, improving the model's robustness to viewpoint changes, local differences, and noise in clinical images. Adjusting the image resolution to the CLIP standard input size ensures the accuracy and consistency of feature extraction, providing a high-quality feature foundation for subsequent calculations of semantic consistency loss and geometric structure preservation loss.
[0031] Furthermore, the objective function of multi-loss fusion achieves multi-dimensional collaborative optimization: adversarial loss ensures that the generated image statistically approximates the real H&E image; dynamically weighted cyclic loss prioritizes the fidelity of key structures such as the cell nucleus; and CLIP semantic-geometric joint loss simultaneously constrains high-level semantic consistency and local geometric stability. This solves multiple problems that cannot be addressed by a single loss in existing technologies, significantly improving the overall quality of virtual staining.
[0032] Furthermore, by parsing the DAPI and FITC channels using OpenSlide, a three-channel format adapted to the model input was constructed. Brightness normalization and 8-bit depth mapping eliminated brightness differences between different imaging devices, reducing computational overhead. In the HSV space, Otsu thresholding combined with 3×3 closing operations and connected component analysis accurately extracted effective tissue areas of interest (ROIs), eliminating over 90% of blank background and impurities, significantly reducing unnecessary computation. Finally, the ROIs were normalized to 1024×1024 image blocks, consistent with the model training size. These preprocessing operations significantly improved the efficiency and accuracy of model inference, providing high-quality input data for the subsequent virtual staining process.
[0033] It is understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.
[0034] In summary, this invention systematically solves the key defects of existing virtual staining technologies through three core innovations: a hybrid supervised training system, prior integration of cell nucleus location, and CLIP semantic geometric constraints. This enables high-fidelity, low-cost, and high-efficiency digital pathology virtual HE staining.
[0035] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0036] Figure 1 To generate a pseudo-pairing data graph using the method of this invention; Figure 2 To generate a pseudo-pairing data graph using the method of this invention; Figure 3 Diagram of the phased training mechanism; Figure 4 A diagram illustrating the prior weighting mechanism of cell nuclear structure; Figure 5 This is a diagram of the overall architecture of the inference terminal for WSI pathology images. Figure 6 Comparison of virtual staining results obtained by direct splicing and Poisson fusion; Figure 7 Images showing virtual staining results and real staining results at multiple scales; Figure 8 A comparative ablation visualization of each improved module in the virtual staining model; Figure 9 A schematic diagram of a computer device provided in an embodiment of the present invention; Figure 10 This is a block diagram of a chip provided according to an embodiment of the present invention.
[0037] Among them, 60. Computer equipment; 61. Processor; 62. Memory; 63. Computer program; 600. Electronic device; 610. Processing unit; 620. Storage unit; 6201. Random access memory unit; 6202. Cache memory unit; 6203. Read-only memory unit; 6204. Program / utility; 6205. Program module; 630. Bus; 640. Display unit; 650. Input / output interface; 660. Network adapter; 700. External device. Detailed Implementation
[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0039] In the description of this invention, it should be understood that the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0040] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0041] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this invention generally indicates that the preceding and following objects have an "or" relationship.
[0042] It should be understood that although terms such as first, second, third, etc., may be used in the embodiments of the present invention to describe the preset range, these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from one another. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0043] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."
[0044] The accompanying drawings illustrate various structural schematic diagrams according to embodiments disclosed in this invention. These drawings are not to scale, and some details have been enlarged for clarity, and some details may have been omitted. The shapes of the various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary and may deviate from reality due to manufacturing tolerances or technical limitations. Furthermore, those skilled in the art can design regions / layers with different shapes, sizes, and relative positions as needed.
[0045] This invention provides a digital pathology virtual HE staining method using an unsupervised learning DF-GAN model. It constructs a complete algorithm design for obtaining H&E images from dual-channel fluorescence images through virtual staining, and then detecting tumor regions on these H&E images. Specifically, it improves upon the Cyclegan model to obtain a DF-GAN virtual staining model (DAPI-FITC-GAN) for staining DAPI and FITC dual-channel fluorescence images. The DF-GAN model is innovatively improved in three aspects: a hybrid supervision system, prior knowledge of cell nucleus location, and semantic geometric constraints from a contrastive language-image pre-training model (CLIP). Multiple rounds of validating training are performed. Then, an end-to-end WSI inference process is designed to address memory constraints and seam artifacts in large-size image inference. A hybrid supervised training system is constructed using the DF-GAN model, combining a small amount of pseudo-paired data with a large amount of real unpaired data, and employing a phased training mechanism to achieve a smooth transition from paired guidance to unpaired alignment. The model combines prior knowledge of cell nucleus location and dynamically weights the data to ensure the fidelity of key cell nucleus structures. The model also incorporates a pre-trained CLIP model to construct joint semantic and geometric loss constraints, improving cell structure stability. It efficiently converts dual-channel fluorescence pathological images into virtual staining H&E images. Experimental analysis shows that the core evaluation metrics of the virtual staining results of this invention are reduced to 62 for FID and 0.0698 for KID, representing reductions of 32.74% and 41.6% respectively compared to the traditional virtual staining model CycleGAN.
[0046] Please see Figure 1 This invention discloses a digital pathology virtual HE staining method using an unsupervised learning DF-GAN model. The virtual staining module is responsible for converting the input DAPI and FITC dual-channel fluorescence pathology images into high-fidelity H&E staining images. This module comprises four core parts: the generator and discriminator structures of the baseline DF-GAN model, a hybrid supervised training system, a priori integration mechanism for cell nucleus locations, and CLIP semantic-geometric joint constraints. The method includes the following steps: S1. Input preprocessing: The original dual-channel fluorescence WSI image is analyzed, channel converted, ROI extracted and size normalized, and background noise and invalid regions are removed to obtain image data that meets the model input requirements. Preprocessing aims to convert dual-fluorescence WSI into input data adapted for DF-GAN and remove invalid regions. A three-channel input format is constructed by parsing the DAPI and FITC channels using OpenSlide; brightness normalization and 8-bit bit-depth mapping are applied to eliminate device brightness differences and reduce computational overhead; Otsu thresholding is used in HSV space, combined with 3×3 closing operations and connected component analysis to extract effective areas of interest (ROIs) and remove blank backgrounds and impurities; finally, the ROIs are normalized to 1024×1024 image patches to match the model training size.
[0047] S2, DF-GAN Model Training: Construct a DF-GAN virtual staining model that includes a hybrid supervised training system, a priori inclusion mechanism for cell nucleus location, and CLIP semantic-geometric joint constraints. Use a hybrid data system for phased training to optimize model parameters.
[0048] S201. Generate paired pseudo-data Pseudo-paired data was generated using a lookup table (LUT) pairing guidance method, which served as the primary data for the initial training of the DF-GAN model.
[0049] The construction of the lookup table (LUT) is not empirically specified, but rather based on a systematic design using statistical patterns of data from the input and target domains. Specifically, firstly, statistical analysis is performed on the grayscale distribution of a large number of input DAPI channels and the color distribution of the target H&E image. Combined with the prior knowledge in pathology that "the nuclear region is dominated by hematoxylin and appears blue-purple, while the tissue background and cytoplasm are dominated by eosin and appear pink," several key grayscale segments and corresponding representative color points are determined. Subsequently, a complete 256-level color lookup table is generated within each segment using smooth interpolation to ensure continuous, differentiable, and unnatural color transitions in the grayscale space. For example... Figure 2 As shown, this LUT adopts a three-segment structure of "background-transition-nuclear region enhancement": the low grayscale segment is mapped to a near-white background to suppress noise; the medium grayscale segment is mapped to a pinkish-purple hue to represent the basic color tone of tissues and cytoplasm; and the high grayscale segment is mapped to a dark purple to purplish-black hue to enhance the color difference in the nuclear region. Through segmented interpolation and smoothing constraints, the mapping result can provide a color starting point consistent with the semantics of H&E staining for subsequent unpaired learning while maintaining the nuclear region localization.
[0050] A LUT construction method based on the statistical regularities of input and target domain data is adopted. First, a large number of DAPI channel images and corresponding real H&E images are collected, and the grayscale distribution of DAPI channels and the color distribution of H&E images are statistically analyzed. Combining the prior knowledge in pathology that "the nuclear region is dominated by hematoxylin and appears blue-purple, while the tissue background and cytoplasm are dominated by eosin and appear pink," the DAPI grayscale values are divided into three key segments: low, medium, and high, and representative color points are determined for each segment. A complete 256-level color lookup table is generated within each segment using linear interpolation, forming a three-segment structure of "background-transition-nuclear region enhancement." The dual-channel fluorescence images are input into this LUT for color mapping to generate pseudo-H&E staining images, which, together with the original fluorescence images, constitute a pseudo-paired dataset.
[0051] This provides a stable color starting point for early model training, effectively reducing color drift in the early stages of training. Ablation experiments show that, after adding only the LUT pairing guidance module, the model's FID value decreased from 92.1791 at the baseline to 88.4174, and the KID value decreased from 0.1195 to 0.1162. The alignment between the generated image and the distribution of real H&E staining was significantly improved, and the training convergence speed was accelerated by approximately 30%.
[0052] S202, Constructing a Hybrid Data System The hybrid data system consists of two parts: paired pseudo-datasets and unpaired datasets. The roles and compositions of each part are as follows: Pseudo-paired dataset: Generated using the LUT pairing guidance method described above, it contains dual-channel fluorescence images and corresponding pseudo-H&E staining images, serving as local rule guidance signals for model training to complete the establishment of basic mapping.
[0053] Unpaired dataset: Contains a large number of dual-channel fluorescence images without corresponding H&E staining images and a large number of real H&E staining images without corresponding dual-channel fluorescence images, which are used as the global distribution learning signal for model training.
[0054] S203, Phased Training Mechanism (1) Sampling rules In the training implementation, the warm-up phase consists of a small number of training rounds, during which the training samples mainly come from pseudo-paired data to maximize the learning efficiency of the base mapping. After the warm-up phase ends, the training enters the alignment phase, which mainly uses unpaired samples, while allowing paired guide samples to participate in the update with a probability of not less than a very small amount.
[0055] The above mechanism can be abstracted into the following sampling rule: Let Indicates the training round index. Indicates the length of the preheating stage. This indicates the round in which linear decay ends. represents the lower limit of the participation probability of pairing guidance, then the pairing guidance sampling probability can be written as: (1) Wherein, represents the training round index, represents the length of the warm-up phase, represents the probability that the network is updated with paired pseudo data samples in the -th round, represents the lower limit of the probability that pairing guidance samples continue to participate in updating after entering the unpaired main training phase. Wherein, the linear decay phase starts from and continues to , the sampling probability decreases linearly from 1 to ; after that, the participation of pairing guidance samples is maintained with a fixed probability . If , the process completely switches to unpaired learning.
[0056] (2) Training strategy Refer to Figure 3 , group 1 is paired dual-channel fluorescence images and paired pseudo H&E images generated therefrom by the LUT method, which has a small data volume; group 2 is a large number of unpaired dual-channel fluorescence images and real H&E images, which respectively correspond to the input domain A and the target domain B.
[0057] In each parameter update step, the training process maintains both a pairing guidance data stream and an unpaired data stream, and determines which type of data is used as the input source for the current update step according to , so as to form a dynamic sampling strategy of "pairing guidance dominates in the early stage, and unpaired learning dominates in the later stage".
[0058] Construct a hybrid data system including a pseudo-paired dataset and an unpaired dataset, wherein the pseudo-paired dataset is used to provide local rule guidance, and the unpaired dataset is used to learn global distribution features; a three-stage staged training mechanism is designed: The first stage is the warm-up phase (t<T0), pseudo-paired data is sampled with a 100% probability to quickly establish a basic color mapping; The second stage is the linear decay phase (T0≤t<T1), the sampling probability of paired data decreases linearly from 1 to , which guides the model to gradually transition to unpaired learning; The third stage is the stabilization phase (t≥T1), pairing guidance is maintained with a fixed probability to prevent semantic drift during the training process.
[0059] A smooth transition from paired guidance to unpaired alignment was achieved, leveraging both the efficiency of pseudo-paired data and the global distribution learning advantage of unpaired data. The training process was more stable, without significant oscillations or semantic drift; the model's generalization ability was effectively improved, with the FID value on the independent test set reduced by approximately 25% compared to pure unsupervised training.
[0060] S204. Extract prior knowledge of cell nucleus location and construct a dynamically weighted cycle consistency loss. Prior information about cell nucleus location was extracted from the DAPI channel of a dual-channel fluorescence image, and a dynamically weighted cycle consistency loss function was constructed.
[0061] S2041, Prior extraction of cell nucleus location The nucleus location prior method processes the DAPI channel in a dual-channel fluorescence image because the DAPI channel directly reflects the spatial information of the cell nucleus. Binarizing the DAPI channel reveals the spatial distribution and morphological characteristics of the cell nucleus. Specifically, Fiji software is used to perform channel separation on the dual-channel fluorescence image, first separating the DAPI channel.
[0062] Then, for the DAPI single-channel image, the areas with high gray values correspond to the cell nuclei, and the areas with low gray values are the background, to obtain a coarse segmentation mask.
[0063] Next, the results were optimized through morphological processing: first, a 3×3 convolution kernel was closed to fill the tiny holes in the cell nucleus, and then an opening operation of the same size was performed to remove background noise, resulting in a more accurate segmentation mask. Based on the optimized cell nucleus segmentation mask, global statistics, spatial heatmaps, and instance representations of the cell nucleus were extracted.
[0064] S2042, Construction of Dynamically Weighted Cyclic Consistency Loss Considering the sensitivity of pathological images to subtle features, this invention retains the generator backbone structure of the recurrent consistency adversarial network and maps the global statistics of cell nuclei to sample-level weights of the recurrent consistency loss. By dynamically weighting the data, the optimization direction of the model is changed, so that the model applies stronger recurrent consistency constraints to samples with dense cell nucleus distribution or significant clustering features during cross-domain mapping.
[0065] Please see Figure 4 The geometric prior vector P directly affects the weighting coefficients. The weighting coefficients w are normalized values based on the number of nuclei. Nuclear aggregation ratio Maximum degree of reunion Average area of cell nuclei, percentage of cell tissue Its value is restricted to Within the interval, to avoid excessive gradient amplification leading to training instability, the calculation formula is as follows: (2) in, Numeric pruning operators are used to prune variables. The weights are limited to the interval [0.7, 1.7] to avoid training oscillations caused by excessively high or low sample weights. Empirical optimal values for various coefficients are 0.20, 0.10, etc., and the optimal coefficient combination can be determined using a grid search method for different pathological tissue types. The core logic of these weight coefficients is: the higher the nuclear density, the more pronounced the nuclear aggregation characteristics, and the more complex the tissue structure, the larger the weight coefficient, and the stronger the cycle consistency constraint, thus prioritizing the structural fidelity of these key regions.
[0066] Based on the aforementioned weighting coefficients, the dual-channel fluorescence image domain is denoted as domain A. The cyclic reconstruction loss term in domain A is constructed using a weighted approach, as shown in the following formula: (3) in, This represents the weighting coefficient for the cycling loss of the dual-channel fluorescence domain. Represents the mathematical expectation. This is the L1 norm, used to calculate the pixel-level differences between the reconstructed image and the original input image. For example... Figure 4 As shown, the core function of this loss term is to impose a stronger cycle consistency constraint on samples with dense cell nucleus distribution or significant clustering characteristics, forcing the generator model to prioritize the fidelity of key structures for pathological diagnosis during the optimization process. Input image for dual-channel fluorescence domain, To be Through generator Generate virtual H&E images Then, through the generator The reconstructed image is obtained by inverse mapping back to the dual-channel fluorescence domain.
[0067] The H&E staining image domain is denoted as the B domain, and its cycle consistency loss retains its unweighted form. The calculation formula is as follows: (4) in, The weighting coefficient for the H&E staining domain cyclic loss is preferably set to 10; Input image for H&E staining region; A generator for converting H&E staining images into dual-channel fluorescence images; A generator for converting dual-channel fluorescence images into H&E staining images.
[0068] S205. Construct CLIP semantic-geometric joint loss. We introduce the CLIP model pre-trained on the Quilt-1M pathological image dataset to extract global semantic features and multi-layer intermediate geometric features, and construct semantic consistency loss and geometric structure preservation loss.
[0069] The objective function of the basic CycleGAN relies on adversarial loss and cycle consistency loss, focusing on inter-domain distribution matching and reversible reconstruction. However, for cross-modal virtual staining tasks ranging from dual-channel fluorescence to H&E, these losses are still insufficient to constrain high-level semantic consistency and the preservation of local geometry, which may easily lead to phenomena such as similar colors but structural misalignment, realistic textures but damaged cell nucleus morphology. To address this issue, this invention introduces a CLIP model pre-trained on the large pathological image dataset Quilt-1M into the generator optimization, constructing a constraint framework that combines semantic and geometric losses to improve the semantic consistency and structural stability of virtual staining.
[0070] During the loss reconstruction process, the dual-channel fluorescence domain input image is processed. With the generated virtual HE image A unified multi-view data augmentation strategy is applied, including random perspective transformation and random cropping, and then the image resolution is adjusted to the standard input size of the CLIP visual encoder. Two types of feature representations are extracted from the visual encoder: global semantic features output from the top pooling layer, used to calculate semantic consistency loss; and intermediate feature representations output from multiple hidden layers, used to calculate geometric structure preservation loss.
[0071] Semantic consistency loss, calculated using cosine distance, is used to constrain the consistency between the generated and input images in the high-level semantic space. Geometric structure preservation loss, calculated using mean squared error, is used to constrain the stability of the generated image in terms of spatial layout, local shape, and structural relationships. The mathematical expressions for the two losses are as follows: (5) (6) in, This represents the feature extraction function of the Quilt-1M pre-trained CLIP visual encoder. This is the global semantic representation output from the top layer of the encoder. For the first The intermediate feature representation of the output of the hidden layer. For the set of hidden layers participating in geometric constraints, This is the cosine similarity calculation function. It is an L2 norm. Its core function is to reduce the risk of generated images that are visually similar but deviate from the pathological semantics. By aligning intermediate layer features, the stability of local structures such as cell nucleus boundaries and spatial distribution is ensured. The two losses are weighted by coefficients in the generator's overall objective function. and Adjust contribution level.
[0072] S206. Construct the final loss function and optimize the model parameters. The final objective function is based on adversarial loss, bidirectional cyclic consistency loss, and CLIP semantic-geometric joint loss, which optimizes the parameters of the DF-GAN model.
[0073] The final objective function of the DF-GAN model proposed in this invention is composed of adversarial loss, bidirectional dynamic weighted cycle consistency loss, and CLIP semantic-geometric joint loss. It comprehensively considers multiple optimization objectives such as inter-domain distribution matching, structural fidelity, and semantic consistency. The specific forms are summarized as follows: (7) The adversarial loss adopts the standard GAN adversarial loss form, which is used to constrain the image generated by the generator to realistically simulate the image distribution of the target domain (H&E staining domain), and at the same time constrain the discriminator to accurately distinguish between real images and generated images. The calculation formula is as follows: (8) (9) in, Representative generator With discriminator The losses in the confrontation, among which For source domain To the target domain A cross-domain mapping generator, specifically a cross-domain mapping generator from dual-channel fluorescence channel images to H&E images. For use in determining the target domain A discriminator for the authenticity of samples. For the expectation operator, it is approximated in the implementation by the empirical mean of a small batch of samples; Indicates samples in the target domain Take the expected value. For the empirical data distribution of the target domain, This represents image samples obtained from the target domain dataset. For the discriminator to analyze real H&E samples The first item is the output of the authenticity judgment. The constraint discriminator outputs a value close to 1 for real H&E samples; Indicates the source domain samples Take the expected value. For the empirical data distribution of the source domain, This represents image samples obtained from the source domain dataset. For the generator, source domain samples Generated samples mapped to the target domain, The second term is the discriminator's output for determining the authenticity of the generated sample. The constrained discriminator outputs near-zero values for generated samples. For the generator, its optimization objective is equivalent to causing... The goal is to get as close to 1 as possible, so that the generated virtual H&E statistically approximates the real H&E.
[0074] S3, Virtual staining: Input the preprocessed dual-channel fluorescence image into the trained DF-GAN model to generate a high-fidelity H&E staining image. Through seam optimization and staining normalization, artifacts and color shifts are eliminated. The trained DF-GAN model ultimately needs to be applied to large-scale clinical-grade pathological image inference tasks. However, real-world clinical data is typically stored in multi-resolution, multi-channel WSI format, which cannot be directly used as network input. Therefore, this invention designs an end-to-end WSI virtual staining inference workflow, covering four core stages: input preprocessing, block inference, suture optimization, and post-processing normalization. This effectively solves engineering problems in large-scale image inference, such as memory limitations, boundary artifacts, and cross-sample color bias. Its overall architecture is as follows: Figure 5 As shown.
[0075] S301, Block-based reasoning To address the issue of WSI's ultra-large format not being directly inferenced, an overlapping sliding window block segmentation strategy is adopted. Blocks are segmented with a size of 1024×1024 and a 25% overlap rate to ensure the continuity of tumor edge detection. Image blocks are input into a pre-trained DF-GAN to generate H&E staining results block by block. Dropout is disabled during inference to ensure result stability. Stained image blocks are temporarily stored at their original coordinates, and a position index table is established to provide positional basis for subsequent seam optimization and global stitching.
[0076] S302, Joint Optimization During block-based inference, slight differences exist between the boundaries of image blocks, leading to noticeable block boundary grid artifacts when directly stitched together, affecting image integrity and diagnostic accuracy. To address this, this invention designs a three-level transition strategy: "weighted fusion—seam region smoothing—Poisson fusion refinement," effectively suppressing seam artifacts.
[0077] First is weighted fusion, let the first... The virtual coloring output of each tile is Its weighted graph in the global canvas is The weights can be constructed using a Hann window, a feather window, or an edge fading window, and smoothly increase from the edge to the center within the overlap band. For any pixel location... fused images of overlapping regions Defined as a weighted average: (10) in, This is a numerically stable term. This method suppresses abrupt boundary changes and reduces mesh effects by smoothing weights. Next, to further suppress low-frequency color drift at the seams, this invention applies color space-based smoothing and soft alpha fusion to the seam mask neighborhood: first, the image is converted to... Space, for Bilateral filtering is applied to the chroma channel to smooth color discontinuities, while simultaneously... Apply a light Gaussian smooth to the luminance channel; then apply Gaussian blur to the mask to generate a soft boundary. Then, local mixing is performed. Finally, Poisson fusion is introduced for optional finishing of the seam tape. The core idea of Poisson fusion is in the seam area... Solve an output image internally To make its gradient field as close as possible to the source image At the same time at the boundary Above and target image Maintain consistency.
[0078] To visually verify the effect of Poisson fusion on suppressing artifacts at block boundaries, this invention presents a comparison of virtual staining results after direct stitching and refinement using Poisson fusion, such as... Figure 6 As shown in the figure, the first row shows the result of directly stitching together the blocks after virtual staining: clear block boundary grid artifacts are visible in the global view, and the magnified view of the local cell nuclei further shows that there are severe color discontinuities and nuclear structure breaks at the seams, which seriously affect the accuracy of pathological diagnosis; the second row shows the result after Poisson fusion: no visible grid effect is visible in the global view, the local cell nuclei structure is complete and continuous, and the color transition is natural, which fully verifies the effectiveness of Poisson fusion in aligning boundary gradients in the gradient domain, eliminating seam artifacts, and improving the structural fidelity of virtual staining images.
[0079] S303, staining normalization To suppress hue and staining intensity drift across samples, this invention introduces Vahadane staining normalization in the output stage. The core of this method is to decompose the image into a dye basis matrix. and concentration matrix Staining style normalization is achieved by aligning the concentration of the target image with the dye base of the reference image. The decomposition process can be represented as follows: (11) in, This is the image matrix in optical density space.
[0080] The algorithm is known for its accurate color reproduction and high efficiency. The implementation process is as follows: First, a dye basis matrix is fitted using a reference image. Then, the target image is decomposed to obtain... Finally, the normalized image is reconstructed. The image is then transformed back to RGB space. Finally, the system outputs the global results and necessary local ROI comparison views in the form of high-resolution images, facilitating subsequent pathological slide reading and quantitative analysis.
[0081] S4. Output Results: Output high-fidelity H&E staining images to provide a reference for pathological diagnosis.
[0082] Finally, the system outputs global results and necessary local ROI comparison views in the form of high-resolution images, which facilitates subsequent pathological slide reading and quantitative analysis.
[0083] S5. Verification of Multi-Scale Virtual Staining Effects To comprehensively verify the clinical-grade large-scale pathological image reasoning task proposed in this invention and the virtual staining fidelity of the DF-GAN model at different magnifications, this invention constructs a multi-scale visual comparison experiment. It systematically compares the input dual-channel fluorescence image, the virtual H&E image generated by DF-GAN, and the real H&E gold standard image from three core levels of pathological reading.
[0084] The results are as follows Figure 7 As shown. The three columns correspond to: the original dual-channel fluorescence input image, the virtual H&E image generated by DF-GAN, and the real H&E gold standard image, respectively; the three rows represent the 0.7x global whole-slice WSI view, the 10x medium-magnification tissue view, and the 40x high-magnification cell view, respectively. Each scale bar is labeled with a corresponding scale bar to quantify the spatial scale. White arrows mark typical nuclear structures in the high-magnification view for comparing nuclear morphology and staining fidelity.
[0085] At a 0.7× global scale, the virtual H&E image generated by DF-GAN is highly consistent with the real H&E image in terms of overall staining style, tissue outline distribution, and background region features, completely preserving the tissue region range in the original dual-channel fluorescence image. At a 10× medium magnification scale, the virtual H&E image accurately reproduces the mesoscopic pathological features such as tissue texture and stroma distribution in the real H&E image, completely corresponding to the tissue spatial structure of the original dual-channel fluorescence image. At a 40× high magnification scale, as can be seen from the nucleus region marked by the white arrow, the morphology, size, staining depth of the nucleus in the virtual H&E image generated by DF-GAN, as well as the eosin staining effect of the cytoplasm and stroma, are highly consistent with the real H&E image, verifying the model's high-fidelity reproduction capability of key pathological structures such as the nucleus.
[0086] Design an end-to-end WSI inference flow consisting of "input preprocessing - block inference - three-level seam optimization - coloring normalization": Input preprocessing: WSI images are parsed using OpenSlide to extract DAPI and FITC channels, and brightness normalization and 8-bit bit depth mapping are performed; Otsu thresholding combined with morphological processing is used to extract effective tissue ROIs and remove blank backgrounds; Block-based inference: A sliding window block-based strategy with a block size of 1024×1024 and an overlap rate of 25% is adopted. The H&E staining results are generated by inputting the block into the model one by one. Dropout is turned off during inference to ensure the stability of the results. Three-level seam optimization: The first level uses Hann window weighted fusion of overlapping areas; the second level performs bilateral filtering on the chroma channel and Gaussian smoothing on the luminance channel in Lab space; the third level uses Poisson fusion to refine the seam strip and aligns the boundary gradient in the gradient domain. Coloring normalization: The Vahadane method is used to decompose the image into a dye base matrix and a concentration matrix. Coloring style normalization is achieved by aligning the dye base with the reference image.
[0087] The memory limitation of large-size WSI inference was solved, and the inference time for a single 100,000×100,000 pixel WSI image was controlled within 5 minutes. After three-level seam optimization, the visibility of block boundary artifacts was reduced from 100% to 0%, and the cell nuclear structure breakage rate was reduced from 28% to 0%. Vahadane staining normalization reduced the color difference across samples by about 65%, which significantly improved the consistency of diagnostic results.
[0088] In another embodiment of the present invention, a digital pathology virtual HE staining system based on an unsupervised learning DF-GAN model is provided. This system can be used to implement the aforementioned digital pathology virtual HE staining method based on the unsupervised learning DF-GAN model. Specifically, the digital pathology virtual HE staining system based on the unsupervised learning DF-GAN model includes a data acquisition module, a pseudo-paired data generation module, a hybrid training module, a cell nucleus prior weighting module, a CLIP constraint module, and an inference output module.
[0089] The data acquisition module is used to acquire dual-channel fluorescence pathological images containing DAPI channel images and FITC channel images. The pseudo-paired data generation module is used to construct a three-segment color lookup table (LUT) based on the statistical rules of the input domain and the target domain, and to map the dual-channel fluorescence pathological image through the color lookup table LUT to generate pseudo-H&E staining images, thus forming a pseudo-paired dataset. The hybrid training module is used to construct a hybrid data system containing the pseudo-paired dataset and the unpaired dataset, and to train the DF-GAN model using a phased training mechanism. The phased training mechanism includes a warm-up phase with pseudo-paired data as the main component and an alignment phase with unpaired data as the main component is entered after the pairing-guided sampling probability decreases linearly to a preset lower limit according to the training rounds. The cell nucleus prior weighting module is used to extract the cell nucleus segmentation mask from the DAPI channel, calculate the global cell nucleus statistics including the normalized value of the number of cell nuclei, the proportion of nuclear aggregation, the maximum degree of aggregation, the average area of cell nuclei and the proportion of cell tissue, generate weighting coefficients limited to a preset range, and apply the weighting coefficients to the cycle consistency loss of the dual-channel fluorescence domain. The CLIP constraint module is used to extract global semantic features and intermediate feature representations of multi-layer hidden layer outputs using a pre-trained CLIP visual encoder. Based on the global semantic features, semantic consistency loss is calculated using cosine distance. Based on the intermediate feature representations, geometric structure preservation loss is calculated using mean square error. The two are then added as joint constraints to the objective function of the DF-GAN model. The inference output module is used to infer the input dual-channel fluorescence pathology image using the trained DF-GAN model and output a virtual H&E staining image.
[0090] This invention provides a terminal device comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to achieve a corresponding method flow or function. The processor described in this embodiment can be used for the operation of a digital pathology virtual HE staining method for unsupervised learning DF-GAN models, including: A dual-channel fluorescence pathology image is acquired, comprising a DAPI channel image and a FITC channel image. A color lookup table (LUT) is constructed based on the statistical regularities of the input and target domains. The LUT adopts a three-segment structure: low grayscale segments are mapped to the background color, medium grayscale segments to pinkish-purple tones, and high grayscale segments to dark purple to purplish-black tones. The dual-channel fluorescence pathology images are mapped using the LUT to generate pseudo-H&E staining images, forming a pseudo-paired dataset. A hybrid data system is constructed, comprising the pseudo-paired dataset and an unpaired dataset. The unpaired dataset contains dual-channel fluorescence images without corresponding H&E staining images and real H&E staining images without corresponding dual-channel fluorescence images. A phased training mechanism is used to train the DF-GAN model: in the warm-up phase, pseudo-paired data is used as the main training sample; after the warm-up phase, the pairing-guided sampling probability linearly decays to a preset lower limit according to the training rounds, transitioning to an alignment phase dominated by unpaired data. The DAPI channel of the dual-channel fluorescence pathology image is used as the training sample. The process involves extracting a nucleus segmentation mask from the DF-GAN model, calculating global nucleus statistics based on the mask, including normalized nucleus count, nuclear aggregation ratio, maximum aggregation degree, average nucleus area, and cell tissue proportion; calculating weighting coefficients within a preset range based on these global statistics; applying these weighting coefficients to the cyclic consistency loss of the dual-channel fluorescence domain to construct a dynamically weighted cyclic consistency loss; inputting the dual-channel fluorescence pathology image and the generated virtual H&E image into a pre-trained CLIP visual encoder to extract global semantic features from the top pooling layer and intermediate feature representations from multiple hidden layers; calculating semantic consistency loss using cosine distance based on the global semantic features, and calculating geometric structure preservation loss using mean square error based on the intermediate feature representations; and incorporating the semantic consistency loss and geometric structure preservation loss as joint constraints into the objective function of the DF-GAN model; and using the trained DF-GAN model to infer from the input dual-channel fluorescence pathology image to output a virtual H&E staining image.
[0091] Please see Figure 9 The terminal device is a computer device. In this embodiment, the computer device 60 includes a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When executed by the processor 61, the computer program 63 implements the digital pathology virtual HE staining method of the unsupervised learning DF-GAN model in this embodiment. To avoid repetition, details are omitted here. Alternatively, when executed by the processor 61, the computer program 63 implements the functions of each model / unit in the digital pathology virtual HE staining system of the unsupervised learning DF-GAN model in this embodiment. To avoid repetition, details are omitted here.
[0092] Computer device 60 can be a desktop computer, laptop, handheld computer, cloud server, or other computing device. Computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art will understand that... Figure 9 This is merely an example of computer device 60 and does not constitute a limitation on computer device 60. It may include more or fewer components than shown, or combine certain components, or different components. For example, computer device may also include input / output devices, network access devices, buses, etc.
[0093] The processor 61 may be a Central Processing Unit (CPU), or other general-purpose processors, graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0094] The memory 62 can be an internal storage unit of the computer device 60, such as a hard disk or memory of the computer device 60. The memory 62 can also be an external storage device of the computer device 60, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the computer device 60.
[0095] Furthermore, the memory 62 may include both internal storage units of the computer device 60 and external storage devices. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 can also be used to temporarily store data that has been output or will be output.
[0096] Please see Figure 10The terminal device is an electronic device 600, which is manifested in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including storage unit 620 and processing unit 610), a display unit 640, etc.
[0097] The storage unit stores program code, which can be executed by the processing unit 610 to perform the steps described in the method section of this specification according to various exemplary embodiments of the present invention. For example, the processing unit 610 can perform actions such as... Figure 1 The steps are shown in the figure.
[0098] Storage unit 620 may include readable media in the form of volatile storage units, such as random access memory (RAM) 6201 and / or cache memory 6202, and may further include read-only memory (ROM) 6203.
[0099] Storage unit 620 may also include a program / utility 6204 having a set (at least one) program module 6205, such program module 6205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.
[0100] Bus 630 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the multiple bus structures.
[0101] Electronic device 600 can also communicate with one or more external devices 700 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 600, and / or with any device that enables electronic device 600 to communicate with one or more other computing devices (e.g., router, modem). This communication can be performed via input / output interface 650. Furthermore, electronic device 600 can also communicate with one or more networks (e.g., local area network, wide area network, and / or public network, such as the Internet) via network adapter 660. Network adapter 660 can communicate with other modules of electronic device 600 via bus 630. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.
[0102] This invention also provides a storage medium, specifically a computer-readable storage medium, which is a memory device in a terminal device for storing programs and data. It is understood that the computer-readable storage medium here can include both built-in storage media in the terminal device and extended storage media supported by the terminal device; it can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor, which can be one or more computer programs (including program code). More specific examples of the computer-readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, random access memory, read-only memory, erasable programmable read-only memory, optical fiber, portable compact disk read-only memory, optical storage device, magnetic storage device, or any suitable combination thereof.
[0103] Computer-readable storage media also include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium can also be any readable medium other than a readable storage medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, radio frequency, etc., or any suitable combination thereof.
[0104] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0105] One or more instructions stored in a computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the digital pathology virtual HE staining method for the unsupervised learning DF-GAN model in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by the processor in the following steps: A dual-channel fluorescence pathology image is acquired, comprising a DAPI channel image and a FITC channel image. A color lookup table (LUT) is constructed based on the statistical regularities of the input and target domains. The LUT adopts a three-segment structure: low grayscale segments are mapped to the background color, medium grayscale segments to pinkish-purple tones, and high grayscale segments to dark purple to purplish-black tones. The dual-channel fluorescence pathology images are mapped using the LUT to generate pseudo-H&E staining images, forming a pseudo-paired dataset. A hybrid data system is constructed, comprising the pseudo-paired dataset and an unpaired dataset. The unpaired dataset contains dual-channel fluorescence images without corresponding H&E staining images and real H&E staining images without corresponding dual-channel fluorescence images. A phased training mechanism is used to train the DF-GAN model: in the warm-up phase, pseudo-paired data is used as the main training sample; after the warm-up phase, the pairing-guided sampling probability linearly decays to a preset lower limit according to the training rounds, transitioning to an alignment phase dominated by unpaired data. The DAPI channel of the dual-channel fluorescence pathology image is used as the training sample. The process involves extracting a nucleus segmentation mask from the DF-GAN model, calculating global nucleus statistics based on the mask, including normalized nucleus count, nuclear aggregation ratio, maximum aggregation degree, average nucleus area, and cell tissue proportion; calculating weighting coefficients within a preset range based on these global statistics; applying these weighting coefficients to the cyclic consistency loss of the dual-channel fluorescence domain to construct a dynamically weighted cyclic consistency loss; inputting the dual-channel fluorescence pathology image and the generated virtual H&E image into a pre-trained CLIP visual encoder to extract global semantic features from the top pooling layer and intermediate feature representations from multiple hidden layers; calculating semantic consistency loss using cosine distance based on the global semantic features, and calculating geometric structure preservation loss using mean square error based on the intermediate feature representations; and incorporating the semantic consistency loss and geometric structure preservation loss as joint constraints into the objective function of the DF-GAN model; and using the trained DF-GAN model to infer from the input dual-channel fluorescence pathology image to output a virtual H&E staining image.
[0106] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0107] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0108] Experimental setup This study trained and evaluated the model on a single workstation equipped with an NVIDIA RTX 4070 Ti graphics card. The implementation of this method is based on the PyTorch framework, and image processing utilizes common tool libraries such as OpenCV and PIL.
[0109] Experimental data The virtual staining module used liver tissue pathology slides provided by Xijing Hospital, comprising dual-channel fluorescence images and H&E staining images. The dual-channel fluorescence input consisted of both DAPI and FITC channels. In the preprocessing stage, ROI extraction was performed on the WSI whole-slice images to remove background areas and retain effective cellular tissue regions. Because the WSI images were too large to be directly input into the network, they were uniformly processed into 1024×1024 image blocks. After removing folded and contaminated areas and performing quality screening, the datasets were divided into three categories: a paired dataset containing 2055 dual-channel fluorescence and 2055 H&E images, used for model pairing and warm-up; an unpaired dataset containing 6969 dual-channel fluorescence images and 5445 H&E images, serving as the primary training source; and an independent test set containing 1000 dual-channel fluorescence and 1000 H&E images, which did not participate in any training and were used to objectively evaluate model performance.
[0110] Evaluation indicators Multiple evaluation indicators were selected for the virtual staining module to comprehensively evaluate the performance of the method of the present invention. The specific definitions and calculation formulas of each indicator are as follows.
[0111] 1. FID (Frechet Inception Distance): This metric measures the Fréchet distance between the true distribution and the generated distribution in the Inception feature space.
[0112] 2. KID (Kernel Inception Distance): Similar to KID, this term can be used in the feature space to measure the distributional difference between two sets of images. The KID score is based on the squared estimate of the maximum mean difference (MMD) in the Inception feature space.
[0113] 3. PHV (Peak Histogram Value): This invention uses minimum distance statistics based on perceptual hash (pHash) to implement this metric, which is used to measure the similarity between generated images and real images in terms of structure and edge patterns.
[0114] Virtual staining module experiment ablation experiment To verify the effectiveness of the various improved modules in the proposed method, we examined the contributions of LUT pairing guidance, prior weighting of cell nucleus location, pre-trained feature space constraints, and semantic-geometric joint constraints to the virtual staining effect. We conducted ablation experiments on this part, verifying each module step by step. All models were trained and evaluated using the same data partitioning, the same number of training rounds, and the same evaluation process to ensure that the performance improvements brought by different modules could be clearly explained and compared.
[0115] (1) Quantitative analysis Table 1 Results of the ablation experiment
[0116] As shown in Table 1, each row corresponds to one of the following: the baseline is the CycleGAN model; the improved model with LUT pairing guidance is Baseline+LUT; the model with LUT pairing guidance and semantic-geometric joint constraints is Baseline+LUT+CLIP; and the complete DF-GAN model proposed in this invention integrates LUT pairing guidance, prior weighting of cell nucleus location, and CLIP semantic-geometric joint constraints. Quantitative results show that adding these modules one by one has a significant positive impact on model performance.
[0117] Compared to the baseline model CycleGAN, the Baseline+LUT model with the LUT pairing guidance module reduced FID from 92.179 to 88.417 and KID mean from 0.1195 to 0.1162. This shows that the LUT pairing guidance module can indeed reduce color drift that occurs at the beginning of training and improve the alignment of color distribution between generated and real images. This is because pairing guidance optimizes the cycle consistency loss in the objective function.
[0118] Building upon the Baseline+LUT, adding joint semantic and geometric constraints yields the Baseline+LUT+CLIP, further reducing FID to 71.524 and the mean KID to 0.0875, demonstrating a more significant performance improvement. This indicates that joint semantic and geometric constraints can further uncover correlation information in the feature space, compensating for the shortcomings of previous modules and further enhancing model performance.
[0119] The final DF-GAN model of this invention further incorporates a dynamic weighted loss term, achieving the best results on the main metrics: FID decreased to 62.28, the mean KID decreased to 0.06986, and the mean PHV remained at a good level of 0.31401. This indicates that the prior weighted term for cell nucleus location can guide the model to pay attention to the structural information of the cell nucleus using existing knowledge in pathology, significantly improving the consistency between the generated distribution and the true distribution.
[0120] (2) Qualitative analysis To further verify the independent contributions and synergistic gain effects of each improved module from a visualization perspective, this invention selected three representative 40× high-magnification fields of view and conducted a comparative ablation visualization experiment. The results are as follows: Figure 8 As shown, the scale bar in the figure is 50. .
[0121] In the Baseline model, from Figure 8 (c) Obvious staining distribution shifts, blurred nuclear margins, and insufficient nuclear-cytoplasmic contrast are visible. Uneven staining and texture artifacts also appear in some areas, showing a significant gap in visual consistency with the true H&E gold standard (GT). After introducing a LUT-guided warm-start strategy and CLIP semantic-geometric joint constraints, the model's attention to key pathological structures is significantly improved. The nuclear morphology is more regular, the boundaries are clearer, and the tissue texture is closer to the true H&E staining, effectively suppressing local artifacts.
[0122] Furthermore, the complete model formed by integrating LUT pairing guidance, prior weighting of cell nucleus position, and CLIP semantic-geometric joint constraints on the basis of DF-GAN has achieved the best level in terms of visual effect: cell nucleus morphology, staining depth, and nucleoplasmic contrast are highly consistent with the real H&E gold standard, the tissue texture is natural and there are no obvious artifacts, and it can perfectly reproduce the core microscopic features on which pathological diagnosis depends. It also shows stable and high-fidelity staining effect in three different pathological scenarios.
[0123] The ablation experiments confirmed the necessity of the design of each module, including LUT pairing guidance, prior weighting of cell nucleus location, pre-trained feature constraints, and semantic-geometric joint constraints. Each module exhibited a synergistic gain effect in the virtual staining task, jointly optimizing the model's objective function and effectively improving the quality and accuracy of virtual staining.
[0124] Comparative experiment To verify the relative advantages of the proposed method in virtual coloring tasks, this section selects current mainstream unpaired image translation and coloring transfer methods as benchmarks. All models were trained and evaluated under the same data partitioning, the same number of training rounds, and the same evaluation process to ensure the fairness and comparability of the comparison results. The selected comparison models include traditional recurrent generative adversarial networks CycleGAN, CUT, ASP, and the proposed method as models to be verified.
[0125] Table 2 Results of the comparative experiment
[0126] As can be seen from the quantitative results in Table 2, the method of the present invention is significantly better than all the comparison baselines in terms of the core evaluation indicators.
[0127] Besides Cyclegan, which has been compared in ablation experiments, the FID values of CUT and ASP methods are as high as 117.63 and 120.42, respectively, and the mean KID values are 0.1659 and 0.1712, respectively, both significantly higher than the DF-GAN model. This indicates that the excessive content alignment of CUT limits the staining style transfer effect, while ASP, although adaptable to weakly paired scenarios, lacks prior pathological structures and semantic constraints, making it difficult to balance staining consistency and structural integrity. The method of this invention, through multi-module collaborative optimization, effectively solves the above problems and achieves better alignment between the generated distribution and the real distribution.
[0128] Based on the combined quantitative and qualitative analysis results, the advantages of the method of this invention mainly stem from the synergy of three mechanisms: LUT pairing guides stable early cross-modal mapping; nuclear priors improve the fidelity of key regions; and pre-trained feature constraints enhance semantic and geometric consistency, thereby improving the staining fidelity of the generated results in both statistical distribution and structure.
[0129] Experimental results show that the method of this invention significantly improves performance in both virtual staining and tumor detection tasks. The FID value of virtual staining is 62.0000 and the KID value is 0.0698, which are 32.4% and 41.6% lower than CycleGAN, respectively, and outperform mainstream methods such as CycleGAN, CUT, and ASP. Furthermore, the method of this invention has the following technical advantages: it does not require a large number of strictly matched samples, resulting in low sample acquisition costs; the virtual staining process is short, achieving staining within minutes, significantly improving the efficiency of pathological diagnosis; and the core technology is transferable, extending to immunohistochemical imaging and multiple disease detection scenarios, adapting to diverse clinical needs, and helping the healthcare industry upgrade towards high efficiency, precision, and intelligence.
[0130] In summary, this invention presents a digital pathology virtual HE staining method and system using an unsupervised learning DF-GAN model. Through multi-module collaborative optimization, it achieves a breakthrough performance improvement in digital pathology virtual HE staining tasks. Experimental results show that the core evaluation metrics of this invention are reduced to 62.0000 and KID to 0.0698, representing reductions of 32.4% and 41.6% respectively compared to the traditional CycleGAN model, significantly outperforming mainstream unpaired image translation methods such as CUT and ASP. In multi-scale visual verification, the generated virtual H&E images are highly consistent with real staining in terms of 0.7x global tissue contour, 10x mesoscopic stroma distribution, and 40x high-magnification cell nuclear morphology, perfectly reproducing all the core features relied upon for pathological diagnosis. This invention eliminates the need for a large number of strictly matched clinical samples, reducing data acquisition costs by more than 80%. The virtual staining process can be completed in minutes, improving efficiency by tens of times compared to traditional H&E staining. The core technology has good transferability and can be quickly extended to immunohistochemical imaging and multi-disease detection scenarios, providing primary healthcare institutions with a low-cost, high-efficiency pathological diagnosis solution with extremely high clinical application value and socio-economic benefits.
[0131] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0132] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0133] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0134] In the embodiments provided by this invention, it should be understood that the disclosed devices / terminals and methods can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0135] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0136] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0137] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random-access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0138] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0139] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0140] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0141] The above content is only for illustrating the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solution based on the technical concept proposed in this invention shall fall within the scope of protection of the claims of this invention.
Claims
1. A digital pathological virtual HE staining method using an unsupervised learning DF-GAN model, characterized in that, Includes the following steps: S1. Acquire dual-channel fluorescence pathological images, wherein the dual-channel fluorescence pathological images include DAPI channel images and FITC channel images; S2. Construct a color lookup table (LUT) based on the statistical patterns of the input and target domains. The color lookup table LUT adopts a three-segment structure: the low grayscale segment is mapped to the background color, the medium grayscale segment is mapped to the pinkish-purple series, and the high grayscale segment is mapped to the dark purple to purplish-black series. The dual-channel fluorescent pathological image is mapped through the color lookup table LUT to generate pseudo-H&E staining images, forming a pseudo-paired dataset. S3. Construct a hybrid data system, which includes the pseudo-paired dataset and the unpaired dataset. The unpaired dataset contains dual-channel fluorescence images without corresponding H&E staining images and real H&E staining images without corresponding dual-channel fluorescence images. A phased training mechanism is used to train the DF-GAN model: in the warm-up phase, pseudo-paired data is used as the main training sample; after the warm-up phase, the pairing-guided sampling probability is linearly reduced to a preset lower limit according to the training rounds, and then the model enters the alignment phase, which is mainly composed of unpaired data. S4. Extract the cell nucleus segmentation mask from the DAPI channel of the dual-channel fluorescence pathology image, calculate the global statistics of the cell nucleus based on the cell nucleus segmentation mask, the global statistics include the normalized value of the number of cell nuclei, the proportion of nuclear aggregation, the maximum degree of aggregation, the average area of cell nuclei, and the proportion of cell tissue; calculate the weighting coefficients based on the global statistics, the weighting coefficients are restricted to a preset range; apply the weighting coefficients to the cycle consistency loss of the dual-channel fluorescence domain to construct a dynamic weighted cycle consistency loss; S5. Input the dual-channel fluorescence pathological image and the generated virtual H&E image into the pre-trained CLIP visual encoder to extract the global semantic features output by the top pooling layer and the intermediate feature representations output by the multi-layer hidden layer. Based on the global semantic features, the semantic consistency loss is calculated using cosine distance; based on the intermediate feature representation, the geometric structure preservation loss is calculated using mean square error; the semantic consistency loss and the geometric structure preservation loss are added as joint constraints to the objective function of the DF-GAN model. S6. The trained DF-GAN model is used to infer the input dual-channel fluorescence pathological image and output a virtual H&E staining image.
2. The digital pathological virtual HE staining method using an unsupervised learning DF-GAN model according to claim 1, characterized in that, In step S2, the color lookup table (LUT) is constructed in the following way: statistical analysis is performed on the grayscale distribution of a large number of input DAPI channels and the color distribution of the target H&E image to determine several key grayscale segments and corresponding representative color points, and a complete 256-level color lookup table is generated in each segment using a smooth interpolation method.
3. The digital pathological virtual HE staining method using an unsupervised learning DF-GAN model according to claim 1, characterized in that, In step S3, the pairing guided sampling probability : in, Indicates the training round index. Indicates the length of the preheating stage. This is the round at which linear decay ends. This represents the lower bound of the probability that paired guide samples will continue to participate in the update after entering the unpaired main training phase.
4. The digital pathological virtual HE staining method using an unsupervised learning DF-GAN model according to claim 1, characterized in that, In step S4, the cell nucleus segmentation mask is obtained in the following way: The DAPI single-channel image is binarized to obtain a coarse segmentation mask; the coarse segmentation mask is then subjected to 3×3 convolution kernel closing and opening operations to obtain an optimized cell kernel segmentation mask.
5. The digital pathological virtual HE staining method using an unsupervised learning DF-GAN model according to claim 1, characterized in that, In step S4, the weighting coefficient for: in, For numerical clipping operators, For the nuclear aggregation ratio, To maximize the degree of reunion, This is the normalized value for the number of cell nuclei. This represents the percentage of cells and tissues.
6. The digital pathological virtual HE staining method using an unsupervised learning DF-GAN model according to claim 5, characterized in that, Dynamic weighted cycle consistency loss in dual-channel fluorescence domain for: in, The weighting coefficients for the cycling loss of the dual-channel fluorescence domain are... Represents the mathematical expectation. It is an L1 norm. Input image for dual-channel fluorescence domain, To be Through generator Generate virtual H&E images Then, through the generator The reconstructed image is obtained by inverse mapping back to the dual-channel fluorescence domain.
7. The digital pathological virtual HE staining method using an unsupervised learning DF-GAN model according to claim 1, characterized in that, In step S5, the pre-trained CLIP visual encoder is a CLIP model pre-trained on the Quilt-1M pathological image dataset. Before inputting the CLIP visual encoder, random perspective transformation and random cropping data augmentation processing are applied to the dual-channel fluorescence pathological image and the virtual H&E image, and the image resolution is adjusted to the standard input size of the CLIP visual encoder.
8. The digital pathological virtual HE staining method using an unsupervised learning DF-GAN model according to claim 7, characterized in that, Objective function of DF-GAN model for: in, To combat the losses, For the cyclic reconstruction loss of the dual-channel fluorescence image domain, Cyclic consistency loss for H&E staining image domain For semantic consistency loss, Loss is preserved for the geometry.
9. The digital pathological virtual HE staining method using an unsupervised learning DF-GAN model according to claim 1, characterized in that, Step S1 also includes the following preprocessing of the dual-channel fluorescence pathological image: channel conversion, region of interest extraction and size regularization of the original dual-channel fluorescence whole-slice image, removal of background noise and invalid regions, to obtain image data that meets the model input requirements.
10. A digital pathology virtual HE staining system based on an unsupervised learning DF-GAN model, characterized in that, include: The data acquisition module is used to acquire dual-channel fluorescence pathological images containing DAPI channel images and FITC channel images; The pseudo-paired data generation module is used to construct a three-segment color lookup table (LUT) based on the statistical rules of the input domain and the target domain, and to map the dual-channel fluorescence pathological image through the color lookup table LUT to generate pseudo-H&E staining images, thus forming a pseudo-paired dataset. The hybrid training module is used to construct a hybrid data system containing the pseudo-paired dataset and the unpaired dataset, and to train the DF-GAN model using a phased training mechanism. The phased training mechanism includes a warm-up phase with pseudo-paired data as the main component and an alignment phase with unpaired data as the main component is entered after the pairing-guided sampling probability decreases linearly to a preset lower limit according to the training rounds. The cell nucleus prior weighting module is used to extract the cell nucleus segmentation mask from the DAPI channel, calculate the global cell nucleus statistics including the normalized value of the number of cell nuclei, the proportion of nuclear aggregation, the maximum degree of aggregation, the average area of cell nuclei and the proportion of cell tissue, generate weighting coefficients limited to a preset range, and apply the weighting coefficients to the cycle consistency loss of the dual-channel fluorescence domain. The CLIP constraint module is used to extract global semantic features and intermediate feature representations of multi-layer hidden layer outputs using a pre-trained CLIP visual encoder. Based on the global semantic features, semantic consistency loss is calculated using cosine distance. Based on the intermediate feature representations, geometric structure preservation loss is calculated using mean square error. The two are then added as joint constraints to the objective function of the DF-GAN model. The inference output module is used to infer the input dual-channel fluorescence pathology image using the trained DF-GAN model and output a virtual H&E staining image.