A claudin 18.2 expression prediction method and system based on a drift diffusion mechanism and graph multi-instance learning, and a storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 金凤实验室
- Filing Date
- 2026-03-13
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies for CLDN18.2 expression detection suffer from problems such as easy manipulation of cell topology, lack of medical semantic alignment, and inability to output intuitive evidence and spatial distribution interpretation, resulting in low diagnostic accuracy and efficiency.
By employing a drift-diffusion mechanism and a graph multi-instance learning method, high-fidelity virtual IHC images are generated and diagnostic reports are automatically written through intelligent preprocessing, cross-modal feature cross-coupling, and spatial-graph context multi-instance learning.
It achieves high fidelity of cell membrane, preservation of non-destructive strength features, cross-modal collaborative verification, and spatial heterogeneity perception, improving the accuracy and efficiency of diagnosis and forming a closed-loop intelligent workflow.
Smart Images

Figure CN122453701A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of medical image processing, computer vision and artificial intelligence-assisted diagnosis and treatment technology, and specifically relates to a CLDN18.2 expression prediction method, system and storage medium based on drift diffusion mechanism and graph multi-instance learning. Background Technology
[0002] Currently, CLDN18.2 is a key biomarker for targeted therapy of gastric cancer, and its detection heavily relies on immunohistochemical (IHC) staining, a process that is time-consuming and costly. To achieve label-free or automated pathological interpretation, recent technological explorations have mainly focused on the following three directions: Direction 1: End-to-end pathological diagnosis based on deep learning. Chinese patent CN116386902A discloses an AI-assisted pathological diagnosis system for colorectal cancer based on deep learning, employing a general convolutional neural network (CNN) for segmentation, feature analysis, and recognition of digital pathological images. This represents the mainstream "black box" method of directly mapping H&E or single-modality images to diagnostic results. Direction 2: Pathological image staining normalization based on deep generative models (GAN / VAE). To eliminate slide preparation differences, researchers have attempted to unify image styles using generative models. Chinese patent CN111028923B discloses a digital pathological image staining normalization method, system, electronic device, and storage medium, objectively reviewing the evolution of existing technologies and pointing out that many institutions have begun using generative adversarial networks (GAN) or variational autoencoders (VAE) for style transfer and staining normalization of pathological images. Direction 3: Virtual immunohistochemistry (IHC) staining based on GAN, targeting IHC... Due to the high cost, some technologies attempt to directly generate virtual IHC images from H&E images. Chinese patent CN113256617B discloses a virtual immunohistochemical staining method and system for pathological sections, proposing a scheme to achieve virtual staining through a network model, attempting to simulate the style of real DAB staining.
[0003] While the aforementioned technologies have made some progress in routine pathological analysis, for CLDN18.2, a special biomarker whose interpretation criteria are "weak cell membrane expression (1+)" and "membrane structural continuity," existing technologies reveal the following insurmountable fundamental defects: 1. Black-box classification loses high-frequency details, lacks clinical evidence, and end-to-end CNN... During multiple pooling downsampling processes, classification models irreversibly smooth out high-frequency texture information. The weak positive (1+) expression of CLDN18.2 is extremely pale and closely adheres to the cell membrane, making it easily confused with background noise by CNNs. More importantly, pure black-box models cannot provide doctors with intuitive "virtual staining evidence," resulting in extremely low clinical trust. 2. Traditional generative models are prone to structural manipulation, leading to edge blurring due to the loss of high-frequency textures. As acknowledged in the background section of Chinese patent CN111028923B, existing staining normalization methods, such as adversarial generative networks, are prone to generating uncontrollable noise points, potentially altering the image structure. For CLDN18.2, any change in cell membrane topology or artifacts will directly affect the distinction between 1+ and 2+. 3. Lack of medical semantic-level spatial alignment and report output. Existing virtual IHC generation technology usually focuses on the conversion of "color style" and lacks strict medical semantic alignment. In addition, existing systems often only output a single classification and quantification result, lacking interpretive descriptions of spatial distribution (such as "positive" or "2+") and explanations of the distribution of each intensity region. Pathologists still need to manually write cumbersome diagnostic reports, and a closed-loop clinical workflow has not been formed. Summary of the Invention
[0004] To address the aforementioned shortcomings in existing technologies, this invention provides a CLDN18.2 expression prediction method, system, and storage medium based on drift diffusion mechanism and graph multi-instance learning, which solves the problem that existing traditional GANs are prone to tampering with cell topology.
[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0006] One of the objectives of this invention is to provide a CLDN18.2 expression prediction method based on drift diffusion mechanism and graph multi-instance learning, which can ensure that the cell topology is not tampered with and generate a more accurate diagnostic report.
[0007] To achieve the above objectives, the technical solution of the present invention is as follows:
[0008] A CLDN18.2 expression prediction method based on drift-diffusion mechanism and graph multi-instance learning includes the following steps:
[0009] S1. Intelligent preprocessing and adaptive topological feature extraction: extract effective tissue from the original digital pathology whole slice image and perform style alignment without destroying the potential biological strength.
[0010] S2. Virtual IHC image synthesis based on the Drifting mechanism: High-fidelity cross-modal conversion from the H&E morphological domain to the IHC expression domain is achieved by utilizing the Drifting mechanism based on stochastic differential equations.
[0011] S3, cross-modal feature cross-coupling and tile-level interpretation, simulates the pathologist's slide reading logic, breaks down independent modal barriers, and achieves deep synergy between morphological and expressive features;
[0012] S4, spatial-graph contextual multi-instance learning full-slice aggregation, restores discrete map features to the spatial topology of the whole slice, evaluates the overall expression of the tumor microenvironment, and outputs numerical interpretation results and attention heat distribution;
[0013] S5. Intelligent pathology report generation based on visual large language model: Utilizing the numerical interpretation results and attention heat distribution in S4, the visual large language model is driven to generate natural language diagnostic reports that conform to clinical standards.
[0014] Further optimization of the design, step S1 includes: in the process of extracting effective tissue from the original digital pathology whole slice image, a multi-resolution pyramid strategy is used to cut the original digital pathology whole slice image, and the local gray-level variance is calculated for each generated patch. When the variance is less than the threshold, it is determined to be a pure glass background and discarded, and only effective patches containing tissue structures are retained.
[0015] To further optimize the design, step S1 also includes: when projecting with frequency domain style invariance, the patch is transformed to the frequency domain using a two-dimensional discrete fast Fourier transform, low-frequency amplitudes are extracted and averaged for alignment, high-frequency texture structure is strictly preserved and the image is reconstructed.
[0016] Further optimization of the design, step S2 includes: using the forward perturbation and inverse generation physical process of stochastic differential equations, generating a virtual IHC image by guiding conditional drift, and introducing an edge constraint penalty term to ensure that the virtual IHC signal is strictly anchored on the physical cell membrane boundary.
[0017] Further design optimization involves training a scoring network that approximates the true conditional score function during conditional drift-guided generation, and using the H&E morphology to guide random noise toward the corresponding DAB positive signal.
[0018] Further optimization of the design, step S3 includes: constructing a dual-branch heterogeneous visual Transformer encoder to extract spatial structure features and expression intensity features respectively, and using a cross-modal collaborative attention mechanism to fuse the two features.
[0019] Further optimization of the design, step S4 includes: calculating the Euclidean distance based on the physical coordinates of the tiles in the original digital pathological whole slice image, constructing a spatial topology graph, and applying a graph convolutional network to perform message passing in the local microenvironment, thereby obtaining the node features after perceiving the local microenvironment.
[0020] Further optimization of the design, step S5 includes: using the Qwen3VL-3B model, which has been fine-tuned in advance using a dataset containing tens of thousands of clinical desensitized pathology questions and answers and public labeled reports, inputting the multimodal context into the model deployed on the vLLM engine for inference, and generating a structured natural language diagnostic report.
[0021] The second objective of this invention is to provide a CLDN18.2 expression prediction system based on drift diffusion mechanism and graph multi-instance learning, for running the above-mentioned CLDN18.2 expression prediction method on a computer.
[0022] To achieve the above objectives, the technical solution of the present invention is as follows:
[0023] A CLDN18.2 expression prediction system based on drift-diffusion mechanism and graph multi-instance learning includes a data acquisition and processing module, a high-fidelity generation and inference module, a cross-modal coupling and graph aggregation module, an intelligent report generation module, and a display and editing module.
[0024] The data acquisition and processing module is configured to: acquire digital pathology whole slice images for analysis and processing, cut them into image blocks, and perform frequency domain fast Fourier transform preprocessing operations;
[0025] The high-fidelity generation inference module is configured to: be equipped with a high-performance GPU array, deploy an SDE-based Drifting generation model, and use an Euler-Maruyama numerical solver to perform virtual IHC image synthesis in step S2 in real time;
[0026] The cross-modal coupling and graph aggregation module is configured to be responsible for performing ViT cross-attention calculation in step S3 and spatial topology graph construction and graph convolution inference in step S4.
[0027] The intelligent report generation module is configured to deploy a vLLM inference engine and a Qwen3VL-3B visual language large model finely tuned to a public pathology dataset, which is responsible for receiving multimodal context and performing high-speed inference in step S5.
[0028] The display editing module is configured to: show users the original H&E, high-fidelity virtual IHC and full-screen high-brightness heatmap, and display a structured diagnostic report generated by the Qwen3VL-3B visual language large model, supporting users to add, delete and modify the report and print it for archiving with one click.
[0029] A third objective of this invention is to provide a computer-readable storage medium for implementing the aforementioned CLDN18.2 expression prediction method.
[0030] To achieve the above objectives, the technical solution of the present invention is as follows:
[0031] A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the CLDN18.2 expression prediction method and optimization design as described in one of the objectives of this invention.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] 1. This invention has extremely high cell membrane fidelity, overcomes the pain point of artifacts, and innovatively adopts a drifting generation mechanism. Combined with "edge overlap constraint", the generated DAB signal can be precisely anchored on the cell membrane at the pixel level, eliminating the fatal problem of traditional GANs that are prone to tampering with the cell topology structure, and providing absolutely reliable underlying visual evidence for the generation of large model reports.
[0034] 2. This invention features non-destructive intensity characteristics, improves the detection rate of weak positive samples, and adopts frequency domain pattern invariant projection technology to retain the phase spectrum that determines spatial texture. It completely eliminates the defects of conventional optical density domain normalization that forcibly changes the concentration, and greatly protects the true color development depth of 1+ weak positive samples.
[0035] 3. This invention follows the clinical logic of cross-modal collaborative verification. The cross-modal feature cross-coupling mechanism perfectly simulates the reading logic of pathologists comparing double-stained continuous sections. It not only retains the advantages of each modality, but also greatly enhances the model's ability to distinguish between broken membranes and continuous membranes.
[0036] 4. This invention provides accurate whole-slice diagnosis with spatial heterogeneity perception. The SGC-MIL framework effectively filters out false positive signals from necrotic and stromal regions by constructing a physically adjacent graph convolutional network and utilizing the spatial dependence of the tumor microenvironment.
[0037] 5. This invention features a closed-loop workflow that improves work efficiency. By introducing the vLLM engine and the Qwen3VL-3B model, it directly transforms complex deep learning multidimensional outputs (numerical values, probabilities, heatmaps, and generated images) into structured clinical natural language reports. It automatically quantifies intensity ratios and writes morphological descriptions of key regions, completely eliminating the tedious report writing work for users and realizing full-process intelligentization from "image reading" to "report issuance". Attached Figure Description
[0038] Figure 1 This is a schematic diagram of the overall process of a CLDN18.2 expression prediction method based on drift diffusion mechanism and graph multi-instance learning, provided in an embodiment of the present invention.
[0039] Figure 2 This invention provides a schematic diagram of the frequency domain style invariant projection preprocessing flow in a CLDN18.2 representation prediction method based on drift diffusion mechanism and graph multi-instance learning, as an embodiment of the present invention.
[0040] Figure 3 This invention provides a virtual IHC image generation architecture based on the Drifting mechanism in the CLDN18.2 expression prediction method based on drift diffusion mechanism and graph multi-instance learning.
[0041] Figure 4 This invention provides a cross-modal feature cross-coupling network structure diagram in the CLDN18.2 expression prediction method based on drift diffusion mechanism and graph multi-instance learning, as an embodiment of the present invention.
[0042] Figure 5 This invention provides a schematic diagram of the Spatial-Graph Contextual Multi-Instance Learning (SGC-MIL) aggregation model in the CLDN18.2 expression prediction method based on drift diffusion mechanism and graph multi-instance learning, as an embodiment of the present invention.
[0043] Figure 6 This invention provides a flowchart of a pathology report generation process based on vLLM and a large visual model in the CLDN18.2 expression prediction method based on drift diffusion mechanism and graph multi-instance learning.
[0044] Figure 7 This is a schematic diagram of a pathology report output by a CLDN18.2 expression prediction method based on drift diffusion mechanism and graph multi-instance learning, provided as an embodiment of the present invention (for output example only). Detailed Implementation
[0045] The present invention will be further described below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims.
[0046] To better explain the technical solution, the following technical terms are explained:
[0047] H&E imaging: A commonly used histological staining technique that uses two dyes, hematoxylin and eosin, to make the cell nucleus (hematoxylin staining) and cytoplasm (eosin staining) appear in different colors, making it easier to observe tissue structures. This is a commonly used pathological staining method.
[0048] IHC imaging: a technique for detecting and locating proteins or antigens in tissues or cells. It uses specific antibodies to label antigens and combines them with chromogenic reagents to visualize specific proteins. It is often used in pathology to detect the expression levels of specific proteins.
[0049] CLDN18.2: A tight junction protein widely found in the tight junctions of epithelial cells such as the stomach and pancreas. In certain types of cancer, such as gastric cancer and pancreatic cancer, CLDN18.2 expression is upregulated, thus it can be used as a biomarker for diagnosis and treatment.
[0050] WSI: Whole-slice imaging technology in digital pathology, which uses a high-resolution scanner to scan pathological slides into digital images, facilitating remote consultation, storage, and analysis.
[0051] SGC-MIL: A multi-instance learning method based on graph neural networks (GNNs) that aggregates multi-instance features by constructing a spatial topology graph. It is suitable for spatial context-aware feature learning, such as local microenvironment features in pathological images.
[0052] Qwen3VL-3B: A visual large language model for converting multimodal inputs (such as images and text) into structured natural language reports. This model has been fine-tuned to adapt to the task of generating pathology reports.
[0053] vLLM: A visual large language model that supports the generation of natural language descriptions, such as pathology reports, from multimodal inputs (e.g., pathology images and text). The model is fine-tuned for specific tasks to improve the accuracy and structure of the generated reports.
[0054] Stochastic differential equations: a mathematical model used to describe continuous-time dynamic systems containing stochastic factors. In this application, it is used to generate virtual IHC images, ensuring that the generated images are morphologically similar to actual IHC images.
[0055] Dual-branch heterogeneous visual Transformer encoder: A visual transformer model that contains two distinct branches, one for extracting spatial structure features and the other for extracting expression intensity features, which are fused through a cross-modal collaborative attention mechanism.
[0056] This application provides a CLDN18.2 expression prediction method based on drift diffusion mechanism and graph multi-instance learning. The overall process includes five main stages: intelligent preprocessing, drifting virtual generation, cross-modal cross-coupling, graph multi-instance full-piece aggregation, and intelligent report generation based on a visual large language model. Specifically, it includes the following steps:
[0057] S1. Intelligent preprocessing and adaptive topological feature extraction extract effective tissue from the original digital pathology whole slice image and perform style alignment without destroying the potential biological strength.
[0058] Specifically, step S1 includes:
[0059] S1.1, Full-slice image acquisition and dynamic scale segmentation:
[0060] Input: Full slice image of H&E from a patient with original gastric cancer.
[0061] Parameters: Magnification set to Extracting tile size pixels, overlap step set to Pixel; Local grayscale variance threshold .
[0062] Output: The original set of H&E tiles containing the effective organization regions. and the corresponding physical coordinate matrix .
[0063] Processing logic: A multi-resolution pyramid strategy is used for tile cutting. The local grayscale variance is calculated for each generated tile; when the variance is less than a threshold... If the background is identified as pure glass, it will be discarded, and only valid tiles containing organizational structures will be retained.
[0064] S1.2, Frequency Domain Pattern Invariant Projection:
[0065] Input: Set of tiles any tile in Preset standard reference blocks .
[0066] Parameters: Frequency domain low-pass mask cutoff frequency radius (That is, only extract data in the very low frequency domain).
[0067] Output: Standardized H&E tiles with consistent style and intact structure .
[0068] Processing logic: The image tiles are transformed to the frequency domain using a two-dimensional discrete fast Fourier transform (2D-FFT). Decompose the frequency domain signal into an amplitude spectrum. Phase spectrum Define the low-pass center mask. Extract global color style components, use The low-frequency amplitudes are averaged and aligned to strictly preserve the phase spectrum representing high-frequency texture structures (such as cell membrane edges) in the current image. The reconstructed amplitudes... as follows: The standardized tiles are obtained through inverse Fourier transform (IFFT):
[0069] S2. Virtual IHC image synthesis based on the Drifting mechanism: High-fidelity cross-modal conversion from the H&E morphological domain to the IHC expression domain is achieved by utilizing the Drifting mechanism based on stochastic differential equations.
[0070] Specifically, step S2 includes:
[0071] S2.1 Definition of drift-diffusion physical process:
[0072] Input: Real IHC data distribution .
[0073] Parameter: Total diffusion time Drift coefficient diffusion coefficient .
[0074] Processing logic: Construct a continuous-time generative model. Define the forward SDE to perturb the real IHC as prior Gaussian noise. The inverse process is used to generate the target image from the noise:
[0075] S2.2 Conditional drift-guided generation:
[0076] Input: Pure Gaussian noise patch Condition variables .
[0077] Output: Initially synthesized virtual tile representation.
[0078] Processing logic: Training the scoring network An approximate true conditional score function is used to guide random noise toward the corresponding DAB positive signal by utilizing the H&E morphology.
[0079] S2.3, Membrane Structure Fidelity Constraints and Output:
[0080] Parameter: Edge constraint penalty weight The Sobel operator is used as the gradient extractor. .
[0081] Output: High-fidelity virtual IHC tiles .
[0082] Processing logic: Introduce a "high-frequency edge overlap penalty term" to force the generation of virtual IHC signals. Strictly anchored to the physical cell membrane boundary, the variance of the gradient distribution between the standardized H&E patch and the generated virtual IHC patch at the same coordinates is calculated. When the consistency of the edge gradient directions of the two is lower than a preset threshold, a loss function is applied. Backpropagation optimization is performed to achieve pixel-level shape locking, eliminating shape artifacts.
[0083] S3, cross-modal feature cross-coupling and tile-level interpretation, simulates the pathologist's slide reading logic, breaks down independent modal barriers, and achieves deep synergy between morphological and expressive features.
[0084] Specifically, step S3 includes:
[0085] S3.1, Heterogeneous Modal Independent Feature Encoding:
[0086] Output: Structural feature vector ; Representation feature vector .
[0087] Processing logic: Construct a dual-branch heterogeneous visual Transformer encoder. The morphological branch extracts spatial structural features; the expression branch extracts expression intensity features.
[0088] S3.2 Cross-modal collaborative attention mechanism:
[0089] Processing logic: Calculate the attention matrix using morphological features as the query and virtual IHC features as the key and value. This step uses cell morphology as an "anchor" to precisely capture weak positive signals.
[0090] S3.3, Patch-level Coupling Feature Output:
[0091] Patch-level feature vectors and the probability of local expression intensity (Corresponding to 0 / 1+ / 2+ / 3+).
[0092] S4, the full-slice aggregation of spatial-graph contextual multi-instance learning, restores discrete map features to the spatial topology of the entire slice, evaluates the overall expression of the tumor microenvironment, and outputs numerical interpretation results and attention heat distribution.
[0093] Specifically, step S4 includes:
[0094] S4.1 Organization Topology Diagram Construction:
[0095] Processing logic: Calculate the Euclidean distance based on the physical coordinates of the tile in the original tile, and find the nearest one. Neighbors construct graph edges Generate a spatial topology map .
[0096] S4.2 Spatial Graph Convolution Aggregation:
[0097] Processing logic: A two-layer graph convolutional network (GCN) is applied to perform message passing on the local micro-environment of the topology graph, and node features are obtained after perceiving the local micro-environment. .
[0098] S4.3 Global Gated Attention Pooling and Output Judgment:
[0099] Output: Slice-level final diagnostic score (category distribution) ) and the distribution of attention weights for all nodes in the graph .
[0100] Processing logic: A gated attention mechanism is used to calculate the node contribution weight. The WSI-level representation vector is obtained by weighted summation of all nodes in the graph, and then fed into the Softmax classifier to output the final classification.
[0101] S5. Intelligent pathology report generation based on visual large language model: Utilizing the numerical interpretation results and attention heat distribution in S4, the visual large language model is driven to generate natural language diagnostic reports that conform to clinical standards.
[0102] Specifically, step S5 includes:
[0103] S5.1 Multimodal diagnostic context assembly:
[0104] Input: Full-scale diagnostic score The set of intensity probabilities for the entire image patch Patch attention weights and the original With virtual image.
[0105] Parameter: Number of key regions to extract .
[0106] Output: Structured multimodal contextual cue words.
[0107] Processing logic:
[0108] 1. Statistically calculate the proportion of tumor tissue area (patches) in the entire slice for each expression intensity classified as 0, 1+, 2+, and 3+. 2. Based on attention weights... Sort in descending order and extract the top results that contribute the most to the diagnosis. H&E and virtual IHC images corresponding to key areas. 3. Assemble intensity statistics with image pairs, with an accompanying instruction template: "As a senior pathologist, please generate a standardized pathological diagnosis report based on the quantitative statistics and representative field of view in CLDN18.2 below."
[0109] S5.2 Report Inference Based on vLLM Engine:
[0110] parameter:
[0111] 1. Base model: The Qwen3VL-3B model is pre-tuned using tens of thousands of clinical desensitized pathology question-and-answer and public labeled report datasets (such as PathVQA and open source text);
[0112] 2. Inference engine: vLLM (configured with PagedAttention mechanism);
[0113] 3. Generation parameter: Temperature coefficient nuclear sampling To ensure the rigor of the generated text.
[0114] Output: The original generated natural language pathology diagnosis paragraph.
[0115] Processing logic: The multimodal prompts constructed by S5.1 are input into the Qwen3VL-3B model deployed on the vLLM engine. The high throughput of vLLM ensures that the visual large language model can simultaneously analyze the local microscopic image features (the integrity of membrane staining) and macroscopic statistical data of the input with extremely low latency.
[0116] S5.3, Output of Structured Pathology Report:
[0117] Output: Final formatted clinical "CLDN18.2 Immunohistochemistry Automated Analysis Report".
[0118] Processing logic: Parse the output text of the large model and map it to a fixed report template, clearly defining three core sections:
[0119] Expression intensity ratio: Precisely quantifies and displays the relative proportion of 0, 1+, 2+, and 3+ tumor cells.
[0120] Microscopic description of key regions: Describe the morphological features of the Top-K regions in natural language (e.g., "Region A shows weak, discontinuous basal-lateral membrane staining of some tumor cells").
[0121] Final conclusion: Based on a comprehensive assessment of the case's CLDN18.2 positive / negative status, and in accordance with the latest guidelines, we provide eligibility recommendations for targeted therapy.
[0122] This embodiment provides a CLDN18.2 expression prediction system based on drift-diffusion mechanism and graph multi-instance learning, including a data acquisition and processing module, a high-fidelity generation and inference module, a cross-modal coupling and graph aggregation module, an intelligent report generation module, and a display and editing module.
[0123] The data acquisition and processing module is configured to acquire digital pathology whole-slice images for analysis and processing, cut them into image blocks, and perform frequency domain fast Fourier transform preprocessing operations.
[0124] The high-fidelity generative inference module is configured as follows: equipped with a high-performance GPU array, deploying an SDE-based Drifting generative model, and using an Euler-Maruyama numerical solver to perform virtual IHC image synthesis in step S2 in real time;
[0125] The cross-modal coupling and graph aggregation module is configured to be responsible for performing ViT cross-attention calculation in step S3 and spatial topology graph construction and graph convolution inference in step S4.
[0126] The intelligent report generation module is configured to deploy a vLLM inference engine and a Qwen3VL-3B visual language large model finely tuned to a public pathology dataset, which is responsible for receiving multimodal context and performing high-speed inference in step S5.
[0127] The display editing module is configured to show users the original H&E, high-fidelity virtual IHC, and full-screen high-brightness heatmap, as well as a structured diagnostic report generated by the Qwen3VL-3B visual language large model. Users can add, delete, and modify the report and then print and archive it with one click.
[0128] This embodiment provides a computer-readable storage medium storing computer-executable instructions. When executed by a processor, these instructions are used to implement the CLDN18.2 expression prediction method based on drift diffusion mechanism and graph multi-instance learning.
[0129] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A CLDN18.2 expression prediction method based on drift-diffusion mechanism and graph multi-instance learning, characterized in that, Includes the following steps: S1. Intelligent preprocessing and adaptive topological feature extraction: extract effective tissue from the original digital pathology whole slice image and perform style alignment without destroying the potential biological strength. S2. Virtual IHC image synthesis based on the Drifting mechanism: High-fidelity cross-modal conversion from the H&E morphological domain to the IHC expression domain is achieved by utilizing the Drifting mechanism based on stochastic differential equations. S3, cross-modal feature cross-coupling and tile-level interpretation, simulates the pathologist's slide reading logic, breaks down independent modal barriers, and achieves deep synergy between morphological and expressive features; S4, spatial-graph contextual multi-instance learning full-slice aggregation, restores discrete map features to the spatial topology of the whole slice, evaluates the overall expression of the tumor microenvironment, and outputs numerical interpretation results and attention heat distribution; S5. Intelligent pathology report generation based on visual large language model: Utilizing the numerical interpretation results and attention heat distribution in S4, the visual large language model is driven to generate natural language diagnostic reports that conform to clinical standards.
2. The CLDN18.2 expression prediction method based on drift-diffusion mechanism and graph multi-instance learning as described in claim 1, characterized in that, Step S1 includes: in the process of extracting effective tissue from the original digital pathology whole slice image, a multi-resolution pyramid strategy is used to cut the original digital pathology whole slice image, and the local gray-level variance is calculated for each generated patch. When the variance is less than the threshold, it is determined to be a pure glass background and discarded, and only effective patches containing tissue structures are retained.
3. The CLDN18.2 expression prediction method based on drift-diffusion mechanism and graph multi-instance learning as described in claim 2, characterized in that, Step S1 further includes: during frequency domain pattern invariant projection, the patch is transformed to the frequency domain using a two-dimensional discrete fast Fourier transform, low-frequency amplitudes are extracted and averaged for alignment, high-frequency texture structure is strictly preserved and the image is reconstructed.
4. The CLDN18.2 expression prediction method based on drift-diffusion mechanism and graph multi-instance learning as described in claim 1, characterized in that, Step S2 includes: using the forward perturbation and reverse generation physical process of stochastic differential equations, generating a virtual IHC image by conditional drift, and introducing an edge constraint penalty term to ensure that the virtual IHC signal is strictly anchored on the physical cell membrane boundary.
5. The CLDN18.2 expression prediction method based on drift-diffusion mechanism and graph multi-instance learning as described in claim 4, characterized in that: When generating conditional drift-guided signals, the training scoring network approximates the true conditional score function, and uses the H&E morphology to guide random noise to converge toward the corresponding DAB positive signal.
6. The CLDN18.2 expression prediction method based on drift-diffusion mechanism and graph multi-instance learning as described in claim 1, characterized in that, Step S3 includes: constructing a dual-branch heterogeneous visual Transformer encoder to extract spatial structure features and expression intensity features respectively, and using a cross-modal collaborative attention mechanism to fuse the two features.
7. The CLDN18.2 expression prediction method based on drift-diffusion mechanism and graph multi-instance learning as described in claim 1, characterized in that, Step S4 includes: calculating the Euclidean distance based on the physical coordinates of the plot in the original digital pathological whole slice image, constructing a spatial topology graph, and applying a graph convolutional network to perform message passing in the local microenvironment, thereby obtaining the node features after perceiving the local microenvironment.
8. The CLDN18.2 expression prediction method based on drift-diffusion mechanism and graph multi-instance learning as described in claim 1, characterized in that, Step S5 includes: using the Qwen3VL-3B model, which has been fine-tuned in advance using a dataset containing tens of thousands of clinical desensitized pathology questions and answers and public labeled reports, inputting the multimodal context into the model deployed on the vLLM engine for inference, and generating a structured natural language diagnostic report.
9. A CLDN18.2 expression prediction system based on drift-diffusion mechanism and graph multi-instance learning, comprising a data acquisition and processing module, a high-fidelity generation and inference module, a cross-modal coupling and graph aggregation module, an intelligent report generation module, and a display and editing module, characterized in that, The data acquisition and processing module is configured to: acquire digital pathology whole slice images for analysis and processing, cut them into image blocks, and perform frequency domain fast Fourier transform preprocessing operations; The high-fidelity generation inference module is configured as follows: a Drifting generation model based on SDE, using an Euler-Maruyama numerical solver to perform virtual IHC image synthesis in step S2 in real time; The cross-modal coupling and graph aggregation module is configured to perform ViT cross-attention calculation in step S3 and spatial topology graph construction and graph convolution inference in step S4. The intelligent report generation module is configured as follows: based on the vLLM inference engine and the Qwen3VL-3B visual language large model fine-tuned by the public pathology dataset, it is used to receive multimodal context and perform high-speed inference in step S5. The display editing module is configured to show users the original H&E, high-fidelity virtual IHC, and full-screen high-brightness heatmap, and to display a structured diagnostic report generated by the Qwen3VL-3B visual language large model, supporting users to add, delete, and modify the report before printing and archiving it with one click.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-8.
Citation Information
Patent Citations
CN111028923B
CN113256617B
CN116386902A