Cross-modal graph learning tumor TRG discrimination method, apparatus and device, and storage medium

By combining a cross-modal knowledge distillation framework and graph neural networks, the problems of cross-modal information fusion and small target structure extraction in tumor pathology image analysis are solved, thereby improving the objectivity and accuracy of TRG assessment and making it suitable for clinical application in medical institutions at all levels.

CN121885113APending Publication Date: 2026-04-17CHONGQING UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING UNIV
Filing Date
2026-03-20
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies fail to fully utilize information related to the immune microenvironment in the analysis of pathological images for tumor prognosis. Traditional TRG assessment methods are subjective and ignore TLS microstructures. Cross-modal information acquisition and fusion are difficult, and small target structures are difficult to extract, resulting in assessment results that are not objective and accurate enough.

Method used

A cross-modal knowledge distillation framework is constructed, which combines student and teacher models with graph neural networks to achieve feature fusion and dynamic graph construction of H&E and multiple immunofluorescence images. Pseudo-labels and attention distillation are used to improve model performance, and adaptive gating fusion strategy is used to enhance model robustness and accurately capture TLS spatial features.

Benefits of technology

It achieves objective interpretation of TRG assessment, reduces human error, integrates immune microenvironment information, improves the comprehensiveness and accuracy of prognostic judgment, lowers the application threshold, and adapts to the clinical needs of medical institutions at all levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121885113A_ABST
    Figure CN121885113A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical image analysis and artificial intelligence, and discloses a cross-modal graph learning tumor TRG discrimination method, device and equipment and a storage medium, and the method comprises the steps: obtaining a full-slice pathological image, including Hamp; an E dyeing image and a second modal dyeing pathological image; a cross-modal knowledge distillation framework is constructed, the framework comprises a student model and a teacher model, and the teacher model only comprises Hamp; e, generating a pseudo label and attention distribution from training data of the dyed image, and performing training optimization on the student model by taking the pseudo label and the attention distribution as supervision signals; the method comprises the following steps of: analyzing to-be-analyzed Hamp; and E, inputting the dyed image into the optimized student model, and outputting a TRG classification result and index information of the key region. According to the method, on the basis of multi-instance learning, cross-modal knowledge distillation and a space-semantic dynamic graph representation technology are fused, and more accurate and interpretable tumor TRG intelligent discrimination is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of medical image analysis and artificial intelligence technology, and more specifically, to a cross-modal graph learning method for tumor TRG discrimination. Background Technology

[0002] Cancer, especially malignant cancer, is a major disease with high morbidity and mortality rates worldwide, resulting in a significant disease burden in some regions. For various locally advanced solid tumors, neoadjuvant therapies such as neoadjuvant chemotherapy, radiotherapy, or immunotherapy have become important comprehensive treatment strategies. The efficacy evaluation results have significant clinical guiding significance for the formulation of postoperative adjuvant therapy plans, the assessment of patient recurrence risk, and individualized diagnosis and treatment management.

[0003] Currently, the pathological assessment of neoadjuvant therapy efficacy in clinical practice primarily uses the Tumor Regression Grade (TRG) as the core reference indicator. Existing TRG assessment methods largely rely on pathologists manually estimating the residual degree of tumor tissue on hematoxylin and eosin (H&E) stained sections. This manual assessment method is highly dependent on the pathologist's clinical experience and subjective judgment, resulting in difficulties in standardizing assessment criteria and limited inter-observer consistency. More importantly, traditional manual TRG assessment only focuses on morphological changes such as tumor cells and fibrous stroma, failing to fully reflect the differences in immune status within the tumor microenvironment (TME). The tumor immune microenvironment is closely related to patient treatment response, tumor recurrence and metastasis, and long-term survival. Relying solely on morphological information for TRG interpretation leads to an insufficiently comprehensive characterization of treatment efficacy and patient prognosis.

[0004] Recent studies have confirmed that tertiary lymphoid structures (TLS) in the tumor microenvironment are key immunological indicators influencing treatment response and survival prognosis in cancer patients. TLS are lymph node-like lymphocyte aggregates that form locally in the tumor, containing structures such as T-cell regions, B-cell follicles, and germinal centers. Their presence, quantity, and maturity are closely related to the immunotherapy response and long-term survival of cancer patients, serving as independent prognostic factors. Mature TLS, containing germinal centers, can induce and maintain local anti-tumor immune responses; higher TLS density generally indicates a more favorable immune status and a better prognostic trend. Therefore, incorporating the analysis of immune microenvironment information such as TLS into TRG assessment can improve the objectivity and accuracy of neoadjuvant therapy efficacy evaluation while retaining traditional morphological interpretation criteria, providing more comprehensive pathological evidence for prognostic judgment in cancer patients.

[0005] However, in actual digital pathology diagnosis, achieving accurate quantitative analysis of TLS still faces two major technical challenges: First, the acquisition and fusion of cross-modal information is difficult. Identifying the cellular components and maturity of TLS usually requires the support of specially stained pathological images. Multiplex immunofluorescence (mIF) technology can simultaneously label multiple immune cell types and is the gold standard for identifying TLS. However, the imaging equipment for this technology is expensive and the operation is complex, making it difficult to apply on a large scale to all clinical cases. On the other hand, H&E stained sections are inexpensive and are a routine clinical detection method. However, TLS in its images only show clusters of lymphocytes and lack cell phenotypic information. Immature TLS is difficult to distinguish morphologically from general inflammatory infiltration. The characteristics of the germinal centers inside TLS are also atypical. There are significant differences in the information of different modal pathological images. How to use a small sample of data from a limited number of mIF modalities to transfer the rich knowledge of the immune microenvironment to the analysis of a large number of cases with only H&E images has become an urgent cross-modal fusion problem to be solved. Secondly, the extraction of small target structures is difficult. In whole-slide images (WSI), lymphocytes and lymphocytes (TLS) occupy only a tiny area compared to tens of thousands of tumor cells, typically less than 1% of the total area. They are sparse small targets. Under the weakly supervised multiple instance learning (MIL) framework, the feature signals of TLS are easily overwhelmed by the massive background areas such as normal tissue, necrosis, and fibrosis. At the same time, most existing MIL algorithms ignore the spatial topological constraints of pathological tissues when constructing WSI feature representations. They often use simple feature aggregation strategies such as averaging and max pooling, without explicitly considering the spatial distribution characteristics of the tissue. This can easily lead to the model incorrectly associating lymphocytes that are physically far apart but look similar, and misjudging discretely distributed TLS as a whole structure, resulting in ineffective attention to non-key areas. How to effectively extract and amplify the feature signals of sparse small structures like TLS and avoid spatial relationship confusion is another major challenge faced by existing technologies.

[0006] In summary, current technologies in tumor prognostic pathological image analysis fail to fully utilize information related to the immune microenvironment. Existing TRG assessment methods not only exhibit strong subjectivity but also neglect microstructures such as TLS, which are crucial for prognostic judgment. Therefore, there is an urgent need for a novel intelligent analysis method that can integrate information from different pathological imaging modalities, inject molecular phenotypic knowledge of TLS into the H&E image analysis model, and accurately capture the spatial aggregation characteristics of TLS and highlight its characteristic signals, thereby improving the objectivity and accuracy of TRG classification assessment. Summary of the Invention

[0007] To address the aforementioned problems, a first aspect of this invention provides a cross-modal graph learning-based tumor TRG discrimination method, comprising: Obtain whole-section pathological images, including H&E stained images and second-modality stained pathological images; Construct a cross-modal knowledge distillation framework, which includes a student model and a teacher model, wherein: The student model is used to receive H&E stained images, divide the input H&E stained images into blocks and extract features to obtain image block feature sets; construct a dynamic graph based on the semantic association and spatial proximity relationship between the image block feature sets, and aggregate the dynamic graph through a graph neural network to obtain its whole-patch feature representation and the attention weight distribution of the image blocks, and perform TRG classification prediction based on the whole-patch feature representation; The teacher model is used to receive paired H&E stained images and their corresponding second-modality stained pathological images. For the H&E stained images, the dynamic graph is constructed and aggregated to obtain its full-slice-level feature representation and attention weight distribution. Feature extraction and representation are performed on the second-modality stained pathological images. After fusing the feature representations of the two modalities, TRG classification prediction is performed. The teacher model is used to generate pseudo-labels and attention distributions on training data containing only H&E stained images, and the student model is trained and optimized using the pseudo-labels and attention distributions as supervision signals. The H&E staining image to be analyzed is input into the optimized student model, and the TRG classification results and key region index information are output.

[0008] In one optional embodiment, the second modality staining pathological image is at least one of multiplex immunofluorescence staining images, immunohistochemical staining images, in situ hybridization staining images, or spatial transcriptomics data mapping images.

[0009] In one optional implementation, the construction of the dynamic graph based on the semantic association and spatial proximity relationship between image patch feature sets specifically includes: The image patch feature set is projected into a head representation and a tail representation through a learnable linear transformation, and the semantic relevance is obtained by calculating the inner product of the head representation and the tail representation. Spatial affinity is calculated using a Gaussian decay function based on the spatial coordinates of the image patch. The semantic relevance and spatial affinity are nonlinearly fused to generate adjacency weights. Based on these adjacency weights, the top K neighbor nodes with the largest adjacency weights are adaptively selected for each image block to construct the dynamic graph.

[0010] In one alternative implementation, spatial affinity is calculated using a Gaussian decay function, the formula of which is:

[0011] in, For spatial affinity, To control the neighborhood range using adjustable bandwidth parameters, d ij Let be the Euclidean distance between image patch i and image patch j.

[0012] In one optional implementation, the feature representations of the two modalities are fused using a channel-by-channel adaptive gating fusion strategy, specifically: Full-section feature representations were extracted from H&E stained images and second modality stained pathological images in the paired data, respectively. Through an adaptive gated fusion network, a channel-wise fusion weight vector is learned and feature fusion is performed using the following formula:

[0013] in, To concatenate the WSI feature vectors of the H&E mode and the mIF mode along the channel dimension, and The weight matrix is ​​a learnable matrix. To modify the activation function of the linear unit, It is the Sigmoid activation function. It is the output gating weight vector.

[0014] In one optional implementation, when training and optimizing the student model, the loss function used is a comprehensive loss function, the expression of which is:

[0015]

[0016]

[0017] in, For the comprehensive loss function, For cross-entropy loss, For the probability distribution distillation loss in knowledge distillation, For attention distillation loss, and For hyperparameters, and These are the unnormalized score vectors output by the teacher model and the student model for the same input sample, respectively. To obtain the probability distribution by taking the Softmax method on the unnormalized score vector, This is the distillation temperature coefficient. for Compared to KL divergence, For the student model to the first Attention weights calculated from each image patch For the teacher model to the first Attention weights calculated from each image patch This represents the total number of image patches for this sample.

[0018] In one alternative implementation, the indicator information of the key region is a saliency heatmap, which is generated by mapping the attention weight distribution output by the student model back to the corresponding position in the H&E staining image to be analyzed.

[0019] A second aspect of this invention provides a cross-modal graph learning tumor TRG discrimination device, the device comprising: The image acquisition module is used to acquire whole-section pathological images, including H&E staining images and second-modality staining pathological images; The cross-modal knowledge distillation framework building blocks include student model units and teacher model units; The student model unit is used to receive H&E stained images, divide the input H&E stained images into blocks and extract features to obtain image block feature sets; construct a dynamic graph based on the semantic association and spatial proximity relationship between the image block feature sets, and aggregate the dynamic graph through a graph neural network to obtain its whole-patch feature representation and the attention weight distribution of the image blocks, and perform TRG classification prediction based on the whole-patch feature representation. The teacher model unit is used to receive paired H&E stained images and their corresponding second-modality stained pathological images. For the H&E stained images, the dynamic graph is constructed and aggregated to obtain its full-slice-level feature representation and attention weight distribution; features are extracted and represented for the second-modality stained pathological images; and TRG classification prediction is performed after fusing the feature representations of the two modalities. The double distillation optimization module is used to generate pseudo-labels and attention distributions from training data containing only H&E staining images using the teacher model, and to train and optimize the student model using the pseudo-labels and attention distributions as supervision signals; and The TRG discrimination module is used to input the H&E stained image to be analyzed into the optimized student model and output the TRG classification results and key region index information.

[0020] A third aspect of the present invention provides an electronic device, characterized in that it includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a cross-modal graph learning tumor TRG discrimination method.

[0021] A fourth aspect of the present invention provides a computer-readable storage medium, characterized in that a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a processor, is a cross-modal graph learning method for tumor TRG discrimination.

[0022] This application has at least the following advantages or beneficial effects: (1) Achieve objective interpretation of TRG and reduce human error: Replace the manual visual assessment mode, establish a unified TRG discrimination standard through machine learning, effectively reduce the subjective judgment bias of doctors and human resource consumption, improve the consistency of assessment results, and help standardize clinical decision-making.

[0023] (2) Integrating immune microenvironment information to improve the criteria for judgment: Key features of the tumor microenvironment such as TLS are incorporated into the TRG assessment system to make up for the shortcomings of traditional methods that only focus on the morphological regression of tumor cells. This allows the assessment results to reflect both the degree of tumor regression and local immune activity, greatly improving the comprehensiveness and accuracy of prognosis.

[0024] (3) Cross-modal training for single-modal application to solve the problem of data scarcity: The CMKD cross-modal knowledge distillation framework is constructed, and the teacher model is trained using small sample high-information mIF data. The immune microenvironment discrimination knowledge is transferred to the student model. In actual deployment, only clinically routine and easily accessible H&E images are needed for interpretation. This not only improves model performance with the help of cross-modal information, but also ensures the convenience of clinical application.

[0025] (4) Pseudo-labels + double distillation to improve learning ability with small data: The teacher model generates high-confidence pseudo-labels for unlabeled H&E data to expand the training samples of the student model; combined with Logits probability distribution distillation and attention distillation, the deep knowledge of the teacher model is fully explored, so that the student model can still approach the discrimination performance of the teacher model under small sample and weak supervision conditions.

[0026] (5) Adaptive modal fusion to enhance model robustness: The teacher model adopts a channel-by-channel adaptive gating fusion strategy, which can automatically adjust the contribution weights of H&E and mIF modes according to the actual case. When the mIF signal is poor, missing or noisy, the model can automatically focus on the H&E mode to complete the interpretation, ensuring the stability of the prediction results.

[0027] (6) Spatial-semantic dynamic graph, accurately capturing TLS spatial features: The SCDG dynamic graph method combines the semantic similarity and spatial proximity of image blocks, and punishes long-distance erroneous connections through spatial Gaussian regularization constraints, effectively preserving the integrity of the local structure of TLS, avoiding misjudging discrete lymphocytes as pseudo-TLS structures, reducing false positive noise, and improving the accuracy of TLS localization and discrimination.

[0028] (7) Gated attention mechanism to enhance sparse small target signals: In view of the sparse characteristics of TLS accounting for less than 1%, the gated attention mechanism is used to assign learnable weights to each image patch, automatically filter the background area and amplify the feature signals of the key TLS area, which greatly improves the model’s sensitivity to TLS recognition; at the same time, the attention heat map is output to intuitively show the basis of the model’s decision and improve the interpretability of the model.

[0029] (8) Comprehensive improvement of performance and clinical applicability, and reduction of application threshold: Through the synergistic effect of each module, the accuracy and stability of TRG discrimination are significantly improved, and the prediction results are highly consistent with the actual prognosis of patients; the model only requires H&E image input, the inference speed meets the clinical timeliness requirements, and it can be adapted to medical institutions at all levels, especially enabling small hospitals and remote areas to obtain immunopathological assessment results close to the level of multiple staining, thus reducing the clinical application threshold of high-end diagnostic technology. Attached Figure Description

[0030] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 This is a flowchart of a cross-modal graph learning method for tumor TRG discrimination proposed in one embodiment of this application; Figure 2 This is a structural diagram of a cross-modal graph learning tumor TRG discrimination device proposed in one embodiment of this application; Figure 3 This is a schematic diagram of the cross-modal knowledge distillation (CMKD) framework proposed in one embodiment of this application; Figure 4 This is a comparison of the H&E and mIF modal TLS morphology of postoperative pathological images of gastric cancer according to an embodiment of this application; Figure 4 (a) is a multiplex immunofluorescence mIF image of a postoperative pathological WSI of a tumor; Figure 4 (b) For this WSI, hematoxylin-eosin H&E imaging is performed; Figure 4 (c) is an enlargement of the structure selected by the red box in the mIF image, which contains multiple tertiary lymphoid structures (TLS) at different stages of maturity; Figure 4 (d) is a magnified view of the structure highlighted in red in the H&E image, containing multiple tertiary lymphoid structures (TLS) at different stages of maturity. Figure 4 (c) is a positional correspondence; Figure 5This is a schematic diagram of an electronic device according to this application. Detailed Implementation

[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0033] Please refer to Figure 1 , Figure 1 This is a flowchart of a cross-modal graph learning-based tumor TRG discrimination method proposed in one embodiment of this application. Figure 1 As shown, a cross-modal graph learning method for tumor TRG discrimination includes: S100: Acquire whole-section pathological images, including H&E stained images and second-modality stained pathological images; Specifically, full-field digital pathological images (WSI) of tumor patients after surgery are acquired. The core is to acquire WSI images stained with conventional hematoxylin and eosin (H&E), while simultaneously collecting paired second-modality stained pathological images. The second-modality pathological images are at least one of multiplex immunofluorescence (mIF) stained images, immunohistochemical (IHC) stained images, in situ hybridization stained images, or spatial transcriptomics data mapping images. In this embodiment, multiplex immunofluorescence stained images that can label key immune cells of tertiary lymphoid structures (TLS) are preferred.

[0034] The acquired whole-slice pathological images were preprocessed. First, the images were binarized using the Otsu thresholding method to obtain a mask, completing tissue region segmentation and removing blank backgrounds to retain only the tissue-containing areas. Then, each WSI image was divided into several non-overlapping image blocks according to a fixed-size grid (e.g., 256×256 pixels). This division was repeated at multiple magnifications (5x, 10x, 20x, etc.) to obtain image block sequences at different resolution scales. For cases with paired second-modality stained pathological images, the VALIS algorithm was used to perform rigid and non-rigid registration of the H&E images and second-modality images of consecutive slices, establishing a one-to-one correspondence between them in the slice coordinate system to ensure image block matching of the same tissue region during subsequent feature extraction.

[0035] S200: Construct a cross-modal knowledge distillation framework, which includes a student model and a teacher model, wherein: In this embodiment, the present invention introduces a cross-modality knowledge distillation (CMKD) framework during the model training phase, the specific structure of which is as follows: Figure 3 As shown, a small number of paired mIF multiple staining slices are used as teacher modalities and trained together with corresponding H&E images to encode the immune phenotypic cues of TLS into the decision-making process. Then, the student model is guided to learn by distillation information such as soft labels and attention distribution. Finally, the student model can achieve judgment performance close to that of the teacher based solely on H&E features.

[0036] S210: The student model is used to receive H&E stained images, divide the input H&E stained images into blocks and extract features to obtain image block feature sets; construct a dynamic graph based on the semantic association and spatial proximity relationship between the image block feature sets, and aggregate the dynamic graph through a graph neural network to obtain its whole-patch feature representation and the attention weight distribution of image blocks, and perform TRG classification prediction based on the whole-patch feature representation; Specifically, the student model only accepts H&E-stained images as input, and completes feature extraction, graph construction, aggregation, and classification entirely based on single-modal data. The specific process is as follows: Image patch feature extraction: Preprocessed H&E image patches are input into a deep network (such as GigaPath, UNI, CONCH, etc.) pre-trained on large-scale pathological image data as a feature extractor, and each image patch is processed. Generate a compact feature vector in d dimensions (preferably 512 dimensions in this embodiment). The feature vectors of all image patches constitute the image patch feature set of this WSI. .

[0037] Spatial-Semantic Co-construction of Dynamic Graph: A dynamic graph is constructed based on the semantic associations and spatial proximity relationships between the image patch feature sets. The specific steps are as follows: Each feature vector in the image patch feature set is projected onto the head representation through two learnable linear transformations. With tail indication Calculate the inner product of the head representation and the tail representation to obtain the instance. and semantic relevance ; Record the center coordinates of each image patch within the slice. Based on the aforementioned spatial coordinates, the spatial affinity is calculated using a Gaussian decay function, as shown in the formula:

[0038] in, For spatial affinity, To control the neighborhood range using adjustable bandwidth parameters, d ijLet be the Euclidean distance between image patch i and image patch j. .

[0039] The semantic relevance and spatial affinity are nonlinearly fused to generate adjacency weights. The preferred fusion method is a weighted sum followed by Sigmoid activation, i.e.:

[0040] in, This is a balancing coefficient used to adjust the model's emphasis on global semantic patterns and local spatial structure. When... When the spatial distance is large, the inhibitory effect of the spatial distance on the edge weights is enhanced, which can effectively filter out spurious connections that are too far apart; when When the size is small, the model considers semantic similarity more to capture long-range associations. For example and examples semantic relevance, For example and examples The fusion adjacency weights in the final dynamic graph.

[0041] Based on the adjacency weight, the top K neighbor nodes with the largest adjacency weights are adaptively selected for each image patch to construct a sparse directed dynamic graph. The node set of the graph All image patches corresponding to WSI, edge set It reflects the spatial-semantic composite association between instances.

[0042] Graph Neural Network Aggregation and Attention Weight Calculation: A graph neural network (such as a graph convolutional network GCN) is used to perform message passing and feature updates on the dynamic graph. The initial features of each node are weighted and combined with the features of its neighboring nodes according to edge weights to obtain new node features that fuse local context. A gated attention aggregation mechanism is introduced, constructing a two-layer fully connected neural network as an attention submodule. Through dual nonlinear transformations and gating coefficient calculation, the attention weights of each image patch are obtained. After Softmax normalization, a normalized attention weight distribution is obtained. Then, the updated node features are weighted and summed according to this weight to aggregate and obtain the full-patch feature representation of WSI. The formula is:

[0043] in, For the full-piece feature representation of WSI, To normalize the attention weights, N is the total number of image patches. The updated node features are those of the graph neural network.

[0044] TRG classification prediction: The full-scale feature representation is input into the fully connected layer and the Softmax classifier to complete the TRG classification prediction of the H&E staining image.

[0045] S220: The teacher model is used to receive paired H&E staining images and their corresponding second-modality staining pathological images. For the H&E staining images, the dynamic graph is constructed and aggregated to obtain its full-slice-level feature representation and attention weight distribution. Feature extraction and representation are performed on the second-modality staining pathological images. After fusing the feature representations of the two modalities, TRG classification prediction is performed. Specifically, the teacher model receives paired H&E staining images and their corresponding second-modality staining pathology images as input. High-precision TRG discrimination is achieved through dual-modality feature fusion, providing distillation supervision signals to the student model. The specific process is as follows: Bimodal feature extraction and representation: For paired H&E stained images, the same steps as the student model are followed to complete image patch feature extraction, dynamic graph construction, and graph neural network aggregation to obtain their full-scale feature representation and attention weight distribution of image patches; for the corresponding second-modal stained pathological images, the same preprocessing, image patch feature extraction, dynamic graph construction, and aggregation process is performed to obtain the full-scale feature representation and attention weight distribution of the second-modal stained pathological images.

[0046] Dual-modal feature fusion: A channel-wise adaptive gating fusion strategy is used to fuse full-patch feature representations of two modalities, specifically: The WSI full-patch feature vectors of the H&E mode and the second mode are concatenated along the channel dimension to obtain the concatenated feature vectors. ; ]; The spliced ​​features are processed by an adaptive gated fusion network to learn a channel-wise fusion weight vector. The gated fusion network is a two-layer fully connected network, and the formula for calculating the fusion weight vector is as follows:

[0047] in, To concatenate the WSI feature vectors of the H&E mode and the mIF mode along the channel dimension, and The weight matrix is ​​a learnable matrix. To modify the activation function of the linear unit, It is the Sigmoid activation function. It is the output gating weight vector.

[0048] The dual-modal feature fusion is performed based on the fusion weight vector, and the formula is as follows:

[0049] in, This is an element-wise multiplication of vectors. This represents the full-scale feature representation after fusion.

[0050] TRG classification prediction: The fused full-scale feature representation is input into the fully connected layer and the Softmax classifier to complete the TRG level classification prediction. After training convergence, the teacher model can make full use of bimodal information to achieve high-precision TRG discrimination and output the predicted class probability distribution, the attention weight distribution of each of the two modalities and the gating fusion weight, as supervision information for subsequent knowledge distillation.

[0051] S300: Using the teacher model, generate pseudo-labels and attention distributions for training data containing only H&E stained images, and use the pseudo-labels and attention distributions as supervision signals to train and optimize the student model; Specifically, high-confidence pseudo-label generation: WSI data containing only H&E-colored images and no paired second modality images in the training set are input into the teacher model after training convergence for inference. At this time, since there is no second modality input, the fusion weight g of the teacher model automatically degenerates into an all-1 vector, and only relies on the H&E branch to complete the inference; the TRG class probability distribution output by the teacher model is extracted, and the class corresponding to the highest probability is selected as the soft label of the WSI, and the attention weight distribution of the H&E branch of the teacher model is extracted at the same time; a confidence threshold τ (e.g., 0.8) is set, and only samples with the maximum predicted probability greater than τ are selected to form a high-confidence pseudo-label sample set, and the soft label and attention weight distribution of the samples are saved.

[0052] Hybrid supervised training of the student model: The pseudo-labeled sample set is merged with the original H&E paired samples with real TRG labels to serve as the training set for the student model, which is then used to train and optimize the model. During training, a comprehensive loss function is used as the target function, expressed as:

[0053] in, Cross-entropy loss is used to measure the difference between the model's predicted TRG class and the true label (or pseudo-label), and its form is: ,in The true labels of the samples (represented by one-hot vectors) This represents the probability that the model predicts the label for this sample. Cross-entropy loss ensures that the student model learns basic classification and discrimination abilities, where N is the total number of image patches for the sample.

[0054] The probability distribution distillation loss in knowledge distillation is given by the following formula: ,in, and These are the unnormalized score vectors output by the teacher model and the student model for the same input sample, respectively. To obtain the probability distribution by taking the Softmax method on the unnormalized score vector, This is the distillation temperature coefficient. for Compared to The KL divergence. Transferring soft-label knowledge from the teacher model to the student model; The formula for attention distillation loss is: ,in, For the student model to the first Attention weights calculated from each image patch For the teacher model to the first Attention weights calculated from each image patch The total number of image patches in the sample is represented by this loss value, which helps the student model align its attention distribution with the teacher model, focusing on the key lesion areas that the teacher is paying attention to. (Coefficient) and This is a hyperparameter used to balance the proportion of the two distillation losses in the total loss.

[0055] Student model distillation fine-tuning: On a small amount of paired data of H&E and second mode, the pre-trained student model is distilled and fine-tuned. The initial weights of the student model are set to the weights obtained from the pre-training. The true TRG labels of the paired data, the soft labels output by the teacher model, and the attention distribution are used as supervision. The comprehensive loss function is recalculated and the model parameters are iteratively optimized so that the student model can fully absorb the knowledge essence of the teacher model and complete the final model shaping.

[0056] S400: Input the H&E staining image to be analyzed into the optimized student model, and output the TRG classification results and key region index information.

[0057] Specifically, the postoperative H&E stained whole-section pathological images of tumor patients to be analyzed in clinical practice are preprocessed according to the S100 standard to complete tissue segmentation and image block segmentation, and then input into the trained and optimized student model. The student model automatically completes image block feature extraction, dynamic graph construction, graph neural network aggregation and attention weight calculation according to the S210 process, and outputs the TRG classification result corresponding to the image, including the specific TRG level and prediction confidence.

[0058] Simultaneously, the system outputs key region indicator information, which is a saliency heatmap. This heatmap is generated by mapping the normalized attention weight distribution output by the student model back to the corresponding image block position in the H&E staining image to be analyzed. The highlighted areas in the heatmap are the key regions that the model focuses on during discrimination (such as TLS-enriched regions), which can intuitively present the decision basis of the model and provide a reference for pathologists' clinical diagnosis.

[0059] Please refer to Figure 2 , Figure 2 This is a structural diagram of a cross-modal graph learning tumor TRG discrimination device proposed in one embodiment of this application. Figure 2 As shown in the figure, this disclosure also provides a cross-modal graph learning tumor TRG discrimination device, the device comprising: an image acquisition module 201, a cross-modal knowledge distillation framework construction module 202, a double distillation optimization module 203, and a TRG discrimination module 204; wherein, The image acquisition module 201 is used to acquire whole-section pathological images, including H&E staining images and second-modality staining pathological images. The cross-modal knowledge distillation framework construction module 202 includes student model units and teacher model units; The student model unit is used to receive H&E stained images, divide the input H&E stained images into blocks and extract features to obtain image block feature sets; construct a dynamic graph based on the semantic association and spatial proximity relationship between the image block feature sets, and aggregate the dynamic graph through a graph neural network to obtain its whole-patch feature representation and the attention weight distribution of the image blocks, and perform TRG classification prediction based on the whole-patch feature representation. The teacher model unit is used to receive paired H&E stained images and their corresponding second-modality stained pathological images. For the H&E stained images, the dynamic graph is constructed and aggregated to obtain its full-slice-level feature representation and attention weight distribution; features are extracted and represented for the second-modality stained pathological images; and TRG classification prediction is performed after fusing the feature representations of the two modalities. The double distillation optimization module 203 is used to generate pseudo-labels and attention distributions from training data containing only H&E staining images using the teacher model, and to train and optimize the student model using the pseudo-labels and attention distributions as supervision signals; and TRG discrimination module 204 is used to input the H&E staining image to be analyzed into the optimized student model and output TRG classification results and key region index information.

[0060] This disclosure also provides an electronic device, please refer to... Figure 5 , Figure 5 This is a schematic diagram of an electronic device illustrated in an embodiment of this disclosure. For example... Figure 5 As shown, the electronic device 100 includes a memory 110 and a processor 120. The memory 110 and the processor 120 are connected via a bus for communication. The memory 110 stores a computer program that can run on the processor 120 to implement the steps in the cross-modal graph learning tumor TRG discrimination method disclosed in this embodiment.

[0061] The disclosed embodiments also provide a computer-readable storage medium that, when the instructions in the computer-readable storage medium are executed by a processor of a computer device, enables the computer device to perform steps in the cross-modal graph learning tumor TRG discrimination method as described in the embodiments of this disclosure.

[0062] Based on the aforementioned cross-modal graph learning method for tumor TRG discrimination, this section verifies the core performance of the method through a main experimental example (Example 1) and verifies the compatibility of the method under different parameter / data conditions through a parameter substitution experimental example (Example 2). It also extends and verifies the alternative implementation methods and equivalent substitutions of the method. The method framework and modal feature differences are visually illustrated with figures, and the model performance is quantified and key experimental parameters are labeled with tables. Specific details are as follows.

[0063] Example 1 This experiment uses gastric cancer as the application scenario and multiplex immunofluorescence (mIF) as the second modality to verify the core TRG discriminative performance of the method, serving as the basic implementation benchmark for the method.

[0064] (1) Core experimental parameters Experimental data: H&E-WSI of 400 patients after gastric cancer surgery, such as Figure 4 As shown, the H&E and mIF modalities of the collected postoperative pathological images of gastric cancer are expressed in different morphologies of tertiary lymphoid structures (TLS). The left side shows the mIF image, and the right side shows the H&E image. Among them, 70 cases were paired with six-label seven-color mIF-WSI, and were divided into training and test sets at 80% / 20%; Image preprocessing: Image blocks were segmented into 256×256 pixels under 20x magnification, cross-modal registration was performed using the VALIS algorithm, and tissue regions were segmented using the Otsu thresholding method; Feature extraction: The pre-trained model UNI extracts 512-dimensional image patch feature vectors; Dynamic graph construction: Semantic association is achieved using a Head-Tail projection mechanism, with Gaussian decay of spatial affinity (σ initially 500 pixels, γ=1), and the first 8 adjacent nodes are selected; Model training: Teacher model: Adam optimizer, epoch=50, initial learning rate 1e-4; Student model: pseudo-label confidence threshold τ=0.8, pre-training epoch=30, batch size=8, distillation fine-tuning epoch=30, learning rate 5e-5; Loss function: The integrated loss achieves dual distillation of labels and attention.

[0065] (2) Experimental results Teacher model: The TRG classification accuracy on the validation set is nearly 90%, and it can generate effective TLS region attention heatmaps; Student model: Test set accuracy 88.25%, AUC 94.36%, F1-score 88.14%, all indicators are better than traditional multi-instance learning methods; Clinical fit: The model accurately locates attention in TLS-rich regions, and the TRG discrimination results are highly consistent with the actual neoadjuvant therapy efficacy of patients.

[0066] The six-label, seven-color staining table for multiplex immunofluorescence images based on the gastric cancer dataset is shown in Table 1: Table 1

[0067] The model performance of this application method was compared with other multi-instance learning methods. The main comparison focused on the performance of different multi-instance learning methods on the TRG classification task of postoperative pathological images in the gastric cancer dataset of Example 1, considering three metrics: Accuracy, AUC, and F1-score. The results are shown in Table 2. Table 2

[0068] Example 2 This experiment verifies the adaptability of the method by replacing core elements such as the second modality, dynamic graph parameters, feature extraction network, and training hyperparameters. The parameter differences and experimental results of each experimental group compared with the main experimental case are as follows: (1) Second modality replacement: mIF→IHC stained image Parameter differences: The six-label seven-color mIF was replaced with CD20+CD3 double-label IHC. The dual-channel IHC image was obtained by staining adjacent slices and registering and overlaying. The teacher model replaced the IHC feature branch. The remaining parameters were the same as the main experimental example. Experimental results: The teacher model can improve the discrimination performance by utilizing the TLS location information of IHC, and the accuracy of the student model after distillation is similar to that of the main experimental case, proving that the method is compatible with low-cost and easily scalable immune labeling modalities.

[0069] (2) Dynamic graph space regularization strategy replacement: Gaussian decay → fixed radius truncation Parameter differences: The spatial affinity calculation of Gaussian decay is replaced with fixed radius truncation. =200μm, the distance between nodes is less than a certain threshold. hour If it exceeds The remaining composition parameters remain unchanged; Experimental results: The model performance is not significantly different from that of the Gaussian decay method, which can effectively avoid incorrect association of distant image patches; Gaussian decay is the preferred solution because its gradient is differentiable and it can automatically optimize bandwidth, resulting in more stable training convergence.

[0070] (3) Feature extraction network replacement: UNI→ViT model Parameter differences: The feature extractor is replaced with the Vision Transformer model, and the feature dimension of the image patch is adjusted from 512 dimensions to 256 dimensions. The rest of the model structure and training strategy remain unchanged. Experimental results: The student model still achieved TRG discriminant performance similar to that of the main experimental example after distillation, proving that the method has low dependence on different CNN / Transformer type deep models and can be flexibly integrated.

[0071] (4) Training hyperparameter tuning experiment False label confidence threshold: 0.8 → 0.5 Parameter differences: Lower the pseudo-label screening threshold, increase the pseudo-label sample size, and keep other training parameters unchanged; Experimental results: The accuracy of pseudo-labeled samples decreased significantly, and the classification accuracy of the student model declined slightly, confirming the key role of high-confidence pseudo-labels in model performance.

[0072] Attention distillation loss weights (normal values) →Increase / Set to 0 Will After increasing the attention alignment weight to a certain extent, the student model's attention map almost perfectly matched the teacher's, but the classification accuracy did not change significantly; When the value is reduced to 0 (without using attention distillation), the student model exhibits a larger attention bias and a decrease in accuracy, indicating that attention distillation contributes to both improving model interpretability and performance to a certain extent.

[0073] Experimental conclusion: The method of the present invention has good adaptability under various conditions. Regardless of the type of immunolabeled data (mIF or IHC) used, or the feature network or graph parameter settings employed, as long as the overall idea of ​​cross-modal distillation and spatial-semantic graph representation is followed, it can effectively enhance the TRG discrimination ability of H&E pathological images.

[0074] The core framework of this application's method is highly extensible, and key modules can be replaced or extended equivalently. The replacement schemes, parameter adjustments, and verification results for each module are as follows, all of which do not change the core idea of ​​the method and can achieve the same technical effect: Equivalent replacement of teacher modalities: In the cross-modal knowledge distillation framework described in this invention, the second modality used by the teacher model is not limited to multiplex immunofluorescence (mIF). Any pathological imaging or molecular data that can provide information on the tumor microenvironment or cellular composition can serve as a teacher modality, such as double / multiple immunohistochemical staining images, in situ hybridization staining images, or even spatial transcriptomics data mapping maps. Only appropriate adjustments need to be made to the data preprocessing and registration of different modalities to integrate them into the teacher model training, achieving knowledge transfer to the student model.

[0075] Multimodal Teacher Extension: In another variation, the teacher model can be extended to utilize data from more than two modalities for training, such as simultaneously fusing H&E, mIF, and gene expression profiles. By constructing multiple branch networks to learn different modal features and designing multi-channel weight fusion at the decision layer, the teacher model's ability to discriminate complex pathological patterns can be further improved. Although additional modalities may be difficult to obtain in practical applications, this extension can verify the universality of the framework for multi-source data under research conditions.

[0076] Replacement of Feature Extraction Network: While this embodiment uses a pre-trained deep model to extract image patch features, this is not the only option. Optionally, the feature extractor can be any convolutional neural network (CNN) or Vision Transformer model with image discrimination capabilities, such as ResNet, DenseNet, ViT, etc. It is even possible to train end-to-end, allowing the student model to learn feature representations independently, without relying on a pre-trained model. When there is sufficient training data, the end-to-end approach is expected to achieve results comparable to using pre-trained features.

[0077] Evolution of Dynamic Graph Construction Strategies: In the construction of dynamic graphs with spatial-semantic co-constraints, the fusion of semantic similarity and spatial regularization is not limited to the weighted summation form shown in the formula. Optionally, adjacency weights can be generated by multiplying the two or using other non-linear combinations; alternatively, hard neighborhoods can be set first based on spatial distance, and then edges can be connected within the neighborhoods based on feature similarity to achieve a similar effect. Furthermore, the balance coefficient... It can be set to update dynamically as the model is trained, or different balancing parameters can be used for different levels of graph convolution to flexibly characterize the local and global relationships.

[0078] Replacement of Graph Neural Network Type: The Graph Convolutional Network (GCN) used in the dynamic graph representation of this invention can be replaced with any equivalent graph neural network model. For example, GraphSAGE, Graph Attention Network (GAT), Transformer-type graph attention networks, or higher-order message passing mechanisms can be selected. These replacements can achieve similar information fusion effects without changing the topology. Only by adjusting the corresponding adjacency matrix utilization method can they be introduced into the scheme of this invention to achieve equivalent technical effects.

[0079] Replacement of Feature Aggregation Methods: Gated attention aggregation is the preferred WSI feature aggregation method in this invention, but it is not limited to this implementation. Other MIL aggregation strategies can also be used as equivalent replacements, such as: (a) using a multi-head attention mechanism to simultaneously focus on image patch features at different scales before fusion; (b) using hierarchical attention to first aggregate WSI features into sub-regions, and then aggregate them globally; (c) using Max-Min combined pooling, adaptive average pooling, and other methods to amplify the influence of a few significant features. These methods are essentially similar to attention mechanisms and are all equivalent replacements of conventional techniques.

[0080] Variations of knowledge distillation loss: The KL divergence and MSE loss used in the distillation process can be replaced with other equivalent measures. For example, JS divergence (Jensen-Shannon divergence) or cross-entropy can be used to measure the difference between the probability distributions of students and teachers; attention alignment can be achieved using cosine similarity loss or KL divergence instead of MSE. Any function that can measure the difference between the two distributions is applicable. Furthermore, the distillation signal is not limited to probability and attention; it can also include modal alignment of intermediate feature representations, such as applying L2 constraints to the hidden layer output of the student model in relation to the corresponding layer of the teacher model, which is also an equivalent implementation variation.

[0081] Expanding the Scope of Application: Although the main real-time example of this invention focuses on TRG assessment for neoadjuvant therapy of gastric cancer, its technical solution is universal and can be widely applied to other diseases and tasks. For example, it can be used for post-treatment pathological response grading of other solid tumors such as rectal cancer and breast cancer, requiring only corresponding training data for model transfer. Furthermore, when applied to pathological scoring models predicting the efficacy of immunotherapy, selectively labeling immune checkpoint-related modalities as teachers allows student models to learn information about tumor immune status. Moreover, in other tasks requiring the fusion of multimodal pathological information (such as lymph node metastasis detection and molecular feature prediction), the multi-instance learning method combining cross-modal distillation and dynamic graphs of this invention can also provide an effective solution. Therefore, the scope of protection of this invention is not limited to specific cancer types or specific scoring criteria, and is equally applicable to any similar pathological image analysis task.

[0082] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0083] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0084] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0085] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0086] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0087] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0088] The foregoing has provided a detailed description of a cross-modal graph learning method, apparatus, device, and storage medium for tumor TRG discrimination. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and its core ideas. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A cross-modality graph learning tumor TRG discrimination method, characterized in that, include: Obtain whole-section pathological images, including H&E stained images and second-modality stained pathological images; Construct a cross-modal knowledge distillation framework, which includes a student model and a teacher model, wherein: The student model is used to receive H&E stained images, divide the input H&E stained images into blocks and extract features to obtain image block feature sets; construct a dynamic graph based on the semantic association and spatial proximity relationship between the image block feature sets, and aggregate the dynamic graph through a graph neural network to obtain its whole-patch feature representation and the attention weight distribution of the image blocks, and perform TRG classification prediction based on the whole-patch feature representation; The teacher model is used to receive paired H&E stained images and their corresponding second-modality stained pathological images. For the H&E stained images, the dynamic graph is constructed and aggregated to obtain its full-slice-level feature representation and attention weight distribution. Feature extraction and representation are performed on the second-modality stained pathological images. After fusing the feature representations of the two modalities, TRG classification prediction is performed. The teacher model is used to generate pseudo-labels and attention distributions on training data containing only H&E stained images, and the student model is trained and optimized using the pseudo-labels and attention distributions as supervision signals. The H&E staining image to be analyzed is input into the optimized student model, and the TRG classification results and key region index information are output.

2. The cross-modal graph learning tumor TRG discrimination method according to claim 1, characterized in that, The second modality of staining pathological image is at least one of multiple immunofluorescence staining image, immunohistochemical staining image, in situ hybridization staining image, or spatial transcriptomics data mapping image.

3. The cross-modal graph learning-based tumor TRG discrimination method according to claim 1, characterized in that, The construction of the dynamic graph based on the semantic association and spatial proximity relationship between image patch feature sets specifically includes: The image patch feature set is projected into a head representation and a tail representation through a learnable linear transformation, and the semantic relevance is obtained by calculating the inner product of the head representation and the tail representation. Spatial affinity is calculated using a Gaussian decay function based on the spatial coordinates of the image patch. The semantic relevance and spatial affinity are nonlinearly fused to generate adjacency weights. Based on these adjacency weights, the top K neighbor nodes with the largest adjacency weights are adaptively selected for each image block to construct the dynamic graph.

4. The cross-modal graph learning tumor TRG discrimination method according to claim 3, characterized in that, Spatial affinity is calculated using the Gaussian decay function, and the formula is as follows: in, For spatial affinity, To control the neighborhood range using adjustable bandwidth parameters, d ij Let be the Euclidean distance between image patch i and image patch j.

5. The cross-modal graph learning tumor TRG discrimination method according to claim 1, characterized in that, The feature representations of the two modalities are fused using a channel-wise adaptive gating fusion strategy, specifically as follows: Full-section feature representations were extracted from H&E stained images and second modality stained pathological images in the paired data, respectively. Through an adaptive gated fusion network, a channel-wise fusion weight vector is learned and feature fusion is performed using the following formula: in, To concatenate the WSI feature vectors of the H&E and mIF modes along the channel dimension, the WSI feature vectors of both modes were obtained using the SCDG graph aggregation algorithm. and The weight matrix is ​​a learnable matrix. To modify the activation function of the linear unit, It is the Sigmoid activation function. It is the output gating weight vector.

6. The cross-modal graph learning tumor TRG discrimination method according to claim 1, characterized in that, When training and optimizing the student model, the loss function used is the comprehensive loss function, the expression of which is: in, For the comprehensive loss function, For cross-entropy loss, For the probability distribution distillation loss in knowledge distillation, For attention distillation loss, and For hyperparameters, and These are the unnormalized score vectors output by the teacher model and the student model for the same input sample, respectively. To obtain the probability distribution by taking the Softmax method on the unnormalized score vector, This is the distillation temperature coefficient. for Compared to KL divergence, For the student model to the first Attention weights calculated from each image patch For the teacher model to the first Attention weights calculated from each image patch This represents the total number of image patches in the sample.

7. The cross-modal graph learning tumor TRG discrimination method according to claim 1, characterized in that, The key region's indicator information is a saliency heatmap, which is generated by mapping the attention weight distribution output by the student model back to the corresponding position in the H&E staining image to be analyzed.

8. A cross-modal graph learning tumor TRG discrimination device, characterized in that, The device includes: The image acquisition module is used to acquire whole-section pathological images, including H&E staining images and second-modality staining pathological images; The cross-modal knowledge distillation framework building blocks include student model units and teacher model units; The student model unit is used to receive H&E stained images, divide the input H&E stained images into blocks and extract features to obtain image block feature sets; construct a dynamic graph based on the semantic association and spatial proximity relationship between the image block feature sets, and aggregate the dynamic graph through a graph neural network to obtain its whole-patch feature representation and the attention weight distribution of the image blocks, and perform TRG classification prediction based on the whole-patch feature representation. The teacher model unit is used to receive paired H&E stained images and their corresponding second-modality stained pathological images. For the H&E stained images, the dynamic graph is constructed and aggregated to obtain its full-slice-level feature representation and attention weight distribution. Features are extracted and represented for the second-modality stained pathological images. After fusing the feature representations of the two modalities, TRG classification prediction is performed. The double distillation optimization module is used to generate pseudo-labels and attention distributions from training data containing only H&E staining images using the teacher model, and to train and optimize the student model using the pseudo-labels and attention distributions as supervision signals; and The TRG discrimination module is used to input the H&E stained image to be analyzed into the optimized student model and output the TRG classification results and key region index information.

9. An electronic device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the cross-modal graph learning tumor TRG discrimination method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the cross-modal graph learning tumor TRG discrimination method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Tumor classification method and system based on multi-modal completion and knowledge distillation strategy

    CN118799659A

  • Mode-deficient brain tumor image segmentation method, device, equipment, medium and product

    CN120580241A

  • Classification method, device and equipment based on cross-modal and dynamic distillation and medium

    CN121234142A

Cited By

  • Breast cancer prediction method based on multi-modal knowledge distillation

    CN122091216A