Multi-mode chemotherapy response prediction system and method

Through the multimodal chemotherapy response prediction system with compact bilinear pooling and attention clustering module, the problem of insufficient utilization of cross-modal complementarity in chemotherapy response prediction is solved, and chemotherapy response prediction with high accuracy and low resource consumption is achieved.

CN120299620APending Publication Date: 2025-07-11TIANJIN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510410981.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-11

Smart Images

  • Figure CN120299620A_ABST
    Figure CN120299620A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode chemotherapy response prediction system and method. A data set is based on pathological image data and gene data related to chemotherapy data; the first data processing module acquires a histopathological image data block through multi-parameter collaborative acquisition of chemotherapy patient pathological image slice segmentation and hole elimination; the second data processing module queries a gene data set according to gene expression standardization based on chemotherapy patient genes to obtain a first pathological gene feature vector; the third data processing module is used for processing the tissue pathological image data block based on an ImageNet pre-trained ResNet-50 model to obtain a first pathological image feature vector; the bilinear pooling module combines the first pathological gene feature vector and the first pathological image feature vector based on a cross-modal interaction method to obtain a fused pathological image gene feature vector; the attention clustering module performs package-level prediction on the fused pathological image features and gene features based on multi-instance learning to obtain patient chemotherapy response; the method can more accurately and reliably predict the potential response of the patient after chemotherapy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - field of artificial intelligence and biomedicine, and specifically relates to a multimodal chemotherapy response prediction system and method thereof. Background Art

[0002] Neoadjuvant chemotherapy is mainly used for some patients with intermediate - stage tumors. First, the tumor is shrunk through chemotherapy, and then the tumor is cured through treatment methods such as surgery or radiotherapy. However, neoadjuvant chemotherapy also has risks, and there are significant individual differences in its efficacy. Some patients may experience tumor progression or lose the opportunity for radical cure due to poor chemotherapy response. Therefore, it is necessary to predict the state after chemotherapy, such as chemotherapy effect, recurrence, etc. Effective prognosis prediction can avoid the waste of medical resources, prevent over - treatment, reduce costs and provide reference for formulating treatment plans. Therefore, it is very important to develop an efficient and accurate chemotherapy response prediction method, which helps to better treat cancer patients and reduce costs.

[0003] Currently, many researchers have used technologies such as deep learning to conduct research in the field of whole - slide pathology image (WSI) classification. For example, in the paper "Classification and Mutation Prediction from Non - Small Cell Lung Cancer Histopathology Images using Deep Learning" by Coudray et al., for the classification task, a convolutional neural network (CNN) based on the Inception v3 architecture was trained from scratch and achieved an accuracy of 83% in non - small cell lung cancer biopsy samples. In the paper "Pathologist - level classification of histopathological melanoma images with deep neural networks" by Hekler et al., an automated analysis pipeline for multi - classification of lung adenocarcinoma histopathological images was developed, improving the efficiency of pathological diagnosis.

[0004] Although the above research has achieved good classification results using CNN-based deep learning methods, these works still have great limitations. Since the pathological key features often show a highly sparse distribution in WSI, traditional CNN methods mostly rely on fully supervised learning with a large number of noisy labels or manually annotated regions of interest (ROIs) by pathologists at high cost. LU et al. proposed a weakly supervised multi-instance learning (MIL) method called CLAM in the paper "Data-efficient and weakly supervised computational pathology on whole-slide images", which combines an attention mechanism and can accurately locate sub-regions with high diagnostic value and only requires WSI-level labels. In the paper "Dual-stream Multiple Instance Learning Network for Whole Slide Image Classification with Self-supervised Contrastive Learning" by Li et al., a new MIL method DSMIL was proposed, which uses a pyramid fusion mechanism to fuse multi-scale WSI features and further improves the accuracy of classification and localization.

[0005] The above methods are all based on the assumption that all instances in each bag are independent and identically distributed. However, when making a diagnostic decision, pathologists usually consider the context information around a single region and the correlation information between different regions. Therefore, Shao et al. modeled the global correlation between instances through the self-attention mechanism of Transformer and retained spatial information by combining positional encoding in the paper "TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image Classification". However, the feature representation of this method is still limited to the single modality of tissue morphological information. There is an association between the molecular features (such as gene expression) of the tumor microenvironment and the tissue pathological phenotype, but the existing methods have not effectively utilized this cross-modal complementarity, which restricts the generalization ability of the model in complex chemotherapy response prediction. Moreover, there are few existing methods for cancer datasets treated with chemotherapy drugs. Therefore, it is of great significance to develop a chemotherapy response prediction method that fuses digital pathology images and gene expression data. Summary of the Invention

[0006] Aiming at the problems existing in the prior art, the present invention provides a multi-modal chemotherapy response prediction system and method. The invention realizes cross-modal complementarity by introducing compact bilinear pooling (MCB), and significantly reduces the computational overhead while retaining the cross-modal second-order interaction information of the high-dimensional outer product calculation complexity. At the same time, the present invention solves the core pain points of the existing methods in cross-modal interaction, interpretability and computational resource consumption through multi-modal efficient fusion, pathological image adaptive processing and computational optimization design, providing a tool with both high accuracy and clinical practicability for chemotherapy response prediction.

[0007] To solve the problems of the prior art, the present invention adopts the following technical solutions:

[0008] A multi-modal chemotherapy response prediction system, the system includes a data set, a first data processing module, a second data processing module, a third data processing module, a bilinear pooling module and an attention clustering module. The bilinear pooling module is composed of a first feature vector dimensionality reduction unit, a second feature vector dimensionality reduction unit and a fused feature vector unit; the attention clustering module is composed of a gate attention unit, a clustering unit and a classifier; where:

[0009] The data set is based on pathological image data and gene data related to chemotherapy data;

[0010] The first data processing module obtains tissue pathological image data blocks through multi-parameter collaborative acquisition of chemotherapy patient pathological image slice segmentation and hole elimination;

[0011] The second data processing module standardizes the gene expression data to obtain a first pathological gene feature vector after screening out genes related to chemotherapy drugs;

[0012] The third data processing module processes the tissue pathological image data blocks based on the ResNet-50 model pre-trained on ImageNet to obtain a first pathological image feature vector;

[0013] The bilinear pooling module combines the first pathological gene feature vector and the first pathological image feature vector based on a cross-modal interaction method to obtain a fused pathological image gene feature vector;

[0014] The attention clustering module performs bag-level prediction on the fused pathological image features and the first pathological gene features based on a multi-instance learning method to obtain the chemotherapy response of the patient.

[0015] Further, the process of the bilinear pooling module combining the first pathological gene feature vector and the first pathological image feature vector based on a cross-modal interaction method to obtain a fused pathological image gene feature vector includes:

[0016] The first eigenvector dimensionality reduction unit calculates the second pathological image eigenvector data from the first pathological image eigenvector data through the Count Sketch mapping function, i.e.:

[0017] S wsi =Ψ(F wsi ) (1)

[0018] The second eigenvector dimensionality reduction unit calculates the second pathological gene eigenvector data from the first pathological gene eigenvector data through the Count Sketch mapping function, i.e.:

[0019]

[0020] where: S wsi , S gene are the Count Sketches of WSI and gene respectively, Ψ is the Count Sketch mapping function, which compresses the feature dimension through the hash and sign functions, d p is the low-rank projection dimension (d p << d w , d g );

[0021] The fused eigenvector unit calculates the fused pathological image gene eigenvector from the second pathological gene eigenvector data and the second pathological image eigenvector data through the following formula, i.e.: F fusion

[0022]

[0023] where: FO is the fast Fourier transform, which converts the time-domain signal to the frequency domain, ⊙ is the element-wise multiplication (Hadamard product), and FO -1 is the inverse fast Fourier transform.

[0024] Furthermore, the attention clustering module obtains the patient's chemotherapy response process based on the multi-instance learning method for the fused pathological image gene eigenvector, including:

[0025] The gate attention unit calculates the attention score for the fused pathological image gene eigenvector according to the following formula;

[0026]

[0027] where: h k is the eigenvector of the k-th instance, W n is the learnable weight matrix corresponding to the n-th class, is the attention score of the k-th instance to the n-th class, reflecting its pathological relevance;

[0028] The clustering unit obtains pseudo-labels according to the attention scores according to the following formula;

[0029]

[0030] Where: top k and bottom k respectively represent the k instances with the highest and lowest attention scores, is the pseudo binary label; The classifier makes a patient's chemotherapy response based on the pseudo binary label.

[0031] Furthermore, the attention visualization module is composed of an attention score extraction unit, a statistical normalization unit, and a spatial visualization unit;

[0032] The attention score extraction unit extracts the unnormalized attention scores of all tissue sections from the attention branch corresponding to the system prediction category;

[0033] The statistical normalization unit converts the original scores into percentile scales in the [0,1] interval using the quantile normalization method;

[0034] The spatial visualization stage: Establish a spatial association between the tissue microenvironment attention gradient and the system decision logic by mapping the normalized values to RGB intensity values.

[0035] The present invention can also be implemented through the following technical solutions: The method includes: a multi-modal chemotherapy response prediction system construction stage, a multi-modal chemotherapy response prediction system training and testing stage, where: the multi-modal chemotherapy response prediction system construction stage includes the following steps:

[0036] Step 101, construct a data set based on pathological image data and genes related to chemotherapy data;

[0037] Step 102, obtain tissue pathological image data blocks through multi-parameter collaborative acquisition of chemotherapy patient pathological image slice segmentation and hole elimination;

[0038] Step 103, after screening out genes related to chemotherapy drugs, perform standardization processing on gene expression data to obtain a first pathological gene feature vector; That is: d g is the dimension of the gene feature;

[0039] Step 104, use the ResNet-50 model pre-trained on ImageNet to process the tissue pathological image data blocks to obtain a first pathological image feature vector, that is: d w is the dimension of the pathological image feature;

[0040] Step 105: Process the first pathological gene feature vector and the first pathological image feature vector respectively through a hash mapping function to obtain a second pathological gene feature vector and a second pathological image feature vector;

[0041] S wsi =Ψ(F wsi ),

[0042]

[0043] where: S wsi , S gene are the Count Sketches of WSI and gene respectively, Ψ is the Count Sketch mapping function, which compresses the feature dimension through a hash and sign function, and d p is the low-rank projection dimension;

[0044] Step 106: Calculate the fused pathological image gene feature vector from the second pathological gene feature vector data and the second pathological image feature vector data through the following formula, that is: F fusion

[0045]

[0046] where: FO is the fast Fourier transform, which converts the time-domain signal to the frequency domain, ⊙ is the element-wise multiplication (Hadamard product), and FO -1 is the inverse fast Fourier transform;

[0047] Step 107: Calculate the attention score from the fused pathological image gene feature vector according to the following formula;

[0048]

[0049] where: h k is the feature vector of the k-th instance, W n is the learnable weight matrix corresponding to the n-th class, is the attention score of the k-th instance for the n-th class, reflecting its pathological relevance;

[0050] Step 108: Obtain the pseudo label according to the attention score according to the following formula;

[0051]

[0052] where: top k and bottom k represent the k instances with the highest and lowest attention scores respectively, is the pseudo binary label; Make the patient's chemotherapy response based on the pseudo binary label;

[0053] Further, in the training and testing phase of the multi-modal chemotherapy response prediction system, the following steps are included:

[0054] Step 201: Update the parameters of the multi-modal chemotherapy response prediction system through the Adam optimizer with a learning rate of 0.0001 and a weight decay coefficient of 0.00001;

[0055] Step 202: Determine the training set and the test set according to K-fold cross-validation:

[0056] Perform ten-fold cross-validation on the data set in steps 101 - 103. First, divide the data into 10 parts, select 9 of them as the training set Train, and the remaining 1 part as the test set Test. Train 10 classifiers, and take the average of the test set results of the 10 classifiers as the final evaluation result of the model;

[0057] Step 203: Determine the evaluation metrics.

[0058] Beneficial effects

[0059] 1. The advanced cross-modal interaction mechanism (MCB module) of multi-modal data fusion in the present invention overcomes the problem that existing methods (such as simple splicing, bitwise multiplication / addition) cannot effectively capture the non-linear correlation between gene expression and pathological images. The advanced cross-modal interaction mechanism of multi-modal data fusion in the present invention combines Count Sketch and FFT through compact bilinear pooling (MCB), reduces the computational complexity of high-dimensional outer product calculation, and significantly reduces the computational overhead while retaining cross-modal second-order interaction information; experiments prove that the fusion features of the MCB module can more accurately reflect the complex correlation between chemotherapeutic drugs and biomarkers (such as the interaction between platinum drugs and DNA repair genes). At the same time, the present invention adopts coordinate-indexed

[0060] HDF5 storage (only record the tile coordinates instead of directly storing), saving about 75% of the storage space compared with traditional full-tile storage

[0061] (Taking 256×256 pixel tiles as an example, the storage requirement is reduced from the MB level to the KB level).

[0062] 2. The present invention improves Otsu threshold segmentation and adaptive tile division for pathological image processing, and realizes accurate tissue region segmentation through multi-parameter collaborative optimization, eliminating holes and maintaining tissue continuity. In order to overcome the problem that existing MIL methods (such as max-

[0063] pooling) only focus on the most significant instances, the present invention realizes the accurate positioning of key diagnostic regions (such as the tumor microenvironment) through parallel attention branches, and generates pseudo-labels using attention scores to constrain the feature space under the condition of no instance-level labels, improving the chemotherapy response correlation by 82%. Brief Description of the Drawings

[0064] Figure 1 is a flowchart of the present invention;

[0065] Figure 2 is a box plot comparing the present method with other methods on the OV dataset in Example 1;

[0066] Figure 3 is a box plot comparing the present method with other methods on the COAD dataset in Example 1;

[0067] Figure 4 is a box plot comparing the present method with other methods on the BLCA dataset in Example 1;

[0068] Figure 5 is a ridge plot comparing the present method with other methods on the OV dataset in Example 1;

[0069] Figure 6 is a ridge plot comparing the present method with other methods on the COAD dataset in Example 1;

[0070] Figure 7 is a ridge plot comparing the present method with other methods on the BLCA dataset in Example 1;

[0071] Figure 8 is a parallel coordinates plot comparing the present method with other methods on the OV dataset in Example 1;

[0072] Figure 9 is a parallel coordinates plot comparing the present method with other methods on the COAD dataset in Example 1;

[0073] Figure 10 is a parallel coordinates plot comparing the present method with other methods on the BLCA dataset in Example 1; Detailed Description of the Invention

[0074] The following further describes the present invention in conjunction with the attached Figure 1 - attached Figure 10 drawings:

[0075] The present invention provides a multi-modal chemotherapy response prediction system, which includes a dataset, a first data processing module, a second data processing module, a third data processing module, a bilinear pooling module, and an attention clustering module. The bilinear pooling module consists of a first feature vector dimensionality reduction unit, a second feature vector dimensionality reduction unit, and a fused feature vector unit; the attention clustering module consists of a gated attention unit, a clustering unit, and a classifier; as Figure 1As shown, the present invention is based on the histopathological image data W and gene expression data G of the publicly available dataset TCGA database, and the patient data C of patients who received chemotherapy drugs from relevant literature;

[0076] The present invention provides a multi-modal chemotherapy response prediction method, including the following steps:

[0077] Step 1, construct a dataset based on the pathological image data related to chemotherapy patient data and based on the expression data

[0078] 1-1 Prepare the dataset. The dataset includes: the histopathological image data W and gene expression data G of the TCGA database, and the patient data C of patients who received chemotherapy from relevant literature;

[0079] In step 1-1, the histopathological image data and gene expression data are downloaded from https: / / portal.gdc.cancer.gov / , and the patient data of patients who received chemotherapy drugs are obtained from the literature "Predicting responses to platin chemotherapy agents with biochemically-inspired machine learning";

[0080] In step 1-1, the histopathological image data W can be obtained in the following way: first, according to the patient data of patients who received chemotherapy, find the corresponding histopathological images and gene expression data in the database: ovarian cancer patients (OV, received carboplatin treatment), colorectal cancer patients (COAD, received oxaliplatin treatment), and bladder urothelial cancer patients (BLCA, received cisplatin treatment), and download them using the official tool gdc-client. The histopathological image data are whole-slide digital slides (WSIs), and each WSI is in svs format; the gene expression data are mRNA-Seq and are all saved in csv format files.

[0081] In the embodiment of the present invention, the original dataset W downloaded from the database contains 1374 pathological image data of OV, 983 of COAD, and 469 of BLCA, and the size of each image is different, about hundreds of millions of pixels. The corresponding gene expression data of OV, COAD, and BLCA are 18503, 17519, and 60660 respectively.

[0082] 1-2 Screen and label the chemotherapy patient data C in step 1-1;

[0083] In step 1-2, the present invention obtained a clinical data table of patients from the literature "Predicting responses to platin chemotherapy agents with biochemically-inspired machine learning", which has three columns of important information, namely "RADIATION_TREATMENT_ADJUVANT", "PHARMACEUTICAL_TX_ADJUVANT", and "DFS_STATUS". The present invention screened out the data where "RADIATION_TREATMENT_ADJUVANT" is "NO" and "PHARMACEUTICAL_TX_ADJUVANT" is "YES", that is, patients who did not receive radiotherapy but received chemotherapy. After screening, the corresponding patients with OV, COAD, and BLCA are 410, 99, and 72 cases respectively. The value of "DFS_STATUS" being "1:Recurred / Progressed" indicates that the patient's disease has recurred; "0:DiseaseFree" indicates that the patient has no disease progression. According to the screening standard method, tissue pathological image data W and chemotherapy response labels for chemotherapy patients were obtained.

[0084] In an embodiment of the present invention, disease recurrence is encoded as 1; conversely, no disease progression is encoded as 0.

[0085] 1-3 Data augmentation: The present invention performed data augmentation on the original dataset W through operations such as rotation, Gaussian filtering, and adjustment of contrast and sharpness to obtain W'.

[0086] In an embodiment of the present invention, the data of the training set was amplified, and after amplification, the present invention has approximately 5,100 training images.

[0087] Step 2, data preprocessing

[0088] 2-1 Preprocessing tissue pathological image data: After obtaining the original dataset W, the first data preprocessing module cuts each WSI into several pathological image instance tiles (patches). Based on the improved Otsu threshold algorithm, precise pathological slice segmentation and hole elimination are achieved through multi-parameter collaborative optimization, while controlling the computational complexity while ensuring tissue continuity. Since the size of each image is different, the number of small tiles cut from each image ranges from hundreds to thousands. The present invention adopts a coordinate-indexed storage architecture - only the upper-left corner coordinate information of each small tile is recorded in the HDF5 file (each slice corresponds to a.h5 file), rather than directly storing the small tiles. In the feature extraction stage, the target area is dynamically loaded according to the coordinate metadata.

[0089] In an embodiment of the present invention, the size of each tile sliced from the WSI is 256x256 pixels, and the segmentation level uses the pyramid level closest to 64× downsampling.

[0090] 2-2 Preprocessing gene data: The second data processing module obtained genes related to the effectiveness or response of three chemotherapy drugs from the literature "Predicting responses to platinum chemotherapy agents with biochemically-inspired machine learning": for carboplatin, oxaliplatin, and cisplatin, 90, 288, and 179 genes are involved respectively. After excluding genes with missing data, Z-score normalization is performed to eliminate the dimensional difference across genes. The first pathological gene feature vector data is obtained according to this method, denoted as d g as the dimension of the gene feature.

[0091] In an embodiment of the present invention, after excluding genes with missing data, 89, 282, and 176 genes are involved respectively.

[0092] 2-3 Extracting pathological image features: The third data preprocessing uses the ResNet-50 model pre-trained on ImageNet as the basic feature extractor, and makes adaptive adjustments according to the characteristics of pathological images: retain the first three residual blocks (ResBlock 1-3 ), remove the subsequent fully connected layers, and introduce an adaptive average spatial pooling layer (Adaptive Average Spatial Pooling) at the end of the third residual block.

[0093] In an embodiment of the present invention, each small tile of 256×256 pixels is uniformly and efficiently mapped into a 1024-dimensional feature vector to obtain the first pathological image feature vector data, denoted as d w as the dimension of the gene feature.

[0094] Step 3, constructing a multimodal chemotherapy response prediction model MCB

[0095] Constructing the multimodal chemotherapy response prediction model (BiChemoCLAM) based on the attention mechanism and bilinear pooling includes: constructing an attention clustering module, constructing a bilinear pooling module, and constructing an attention visualization module, where,

[0096] 3-1 Constructing a bilinear pooling module, the bilinear pooling module includes a first feature vector dimensionality reduction unit, a second feature vector dimensionality reduction module, and a module for fusing pathological image gene feature vectors

[0097] The input data of the bilinear pooling module is the preprocessed first pathological gene feature vector data described in step 2-2 and the first pathological image feature vector data in step 2-3.

[0098] The bilinear pooling module constructs a cross-modal interaction mechanism to integrate the spatial morphological features of tissue pathological image data and the functional molecular features of the genome, improving the prediction accuracy of chemotherapy response. Traditional methods such as concatenation, element-wise product, element-wise sum, etc. cannot capture the non-linear association between genes and pathological morphology. In theory, the outer product operation enhances the cross-modal representation ability through second-order interaction. However, the computational complexity of the outer product operation is O(d w d g ), which will lead to unacceptable memory overhead when the feature dimension is high.

[0099] The present invention constructs an MCB module to map the result of the outer product into a low-dimensional space and does not require explicit calculation of the outer product: compressing the outer product through Count Sketch projection and fast Fourier transform (FFT) to reduce the computational complexity.

[0100] The first feature vector dimensionality reduction unit calculates the second pathological gene feature vector data from the first pathological gene feature vector data through the Count Sketch mapping function, that is:

[0101] S wsi = Ψ(F wsi ) (1)

[0102] The second feature vector dimensionality reduction unit calculates the second pathological image feature vector data from the first pathological image feature vector data through the Count Sketch mapping function, that is:

[0103]

[0104] where: S wsi , S gene are the Count Sketches of WSI and gene respectively, Ψ is the Count Sketch mapping function, compressing the feature dimension through a hash and sign function, d p is the low-rank projection dimension (d p << d w , d g );

[0105] The fused pathological image gene feature vector module calculates and obtains the fused pathological image gene feature vector from the second pathological gene feature vector data and the second pathological image feature vector data through the following formula, that is: F fusion is the feature after image and gene fusion:

[0106]

[0107] where: FO is the fast Fourier transform, which converts the time-domain signal to the frequency domain, and ⊙ is the element-wise multiplication (Hadamard product), and FO -1 is the inverse fast Fourier transform, which restores the frequency-domain signal to the time domain. This strategy reduces the computational complexity to O(d p (d w +d g ))), which can be regarded as a complementary operation that allows the present invention to more effectively capture the different interactions between the two modalities.

[0108] In an embodiment of the present invention, d p = 512, d w = 1024, d g = 176 (OV dataset).

[0109] 3-2 Construct an attention clustering module, including a gate attention unit, a clustering unit, and a classifier:

[0110] BiChemoCLAM is constructed based on the multi-instance learning (MIL) framework, and its core principle can be formally defined as: Given a WSI as a bag β = {x1, x2,..., x k}, where each instance x k corresponds to a local small patch in the tissue region. Traditional MIL follows the binary classification hypothesis: If there is at least one positive instance, then the bag label Y = 1 Conversely, if all instances are negative, then Y = 0

[0111] First, the gate attention unit generates attention scores for each small patch to locate sub-regions with high diagnostic value. Let the number of classes be N, and construct N parallel attention branches, and each branch generates a class-specific attention distribution:

[0112]

[0113] where: h k is the feature vector of the k-th instance, and W n is the learnable weight matrix corresponding to the n-th class, The attention score of the k-th instance for the n-th class, which reflects its pathological relevance. The patch-level representation is generated by weighted summation:

[0114]

[0115] where: h β,n represents the WSI-level representation aggregated according to the attention score distribution of the n-th class.

[0116] Secondly, for the purpose of the clustering unit, pseudo-labels are generated using the attention scores under the condition of no instance-level labels to constrain the feature space:

[0117]

[0118] where: top k and bottom k respectively represent the k instances with the highest and lowest attention scores, is the pseudo-binary label, which divides the instances into "key evidence" or "interference noise".

[0119] Finally, the classifier makes a patient's chemotherapy response according to the pseudo-binary label. Wherein: in the embodiment of the present invention, the number of classes N = 2, and for the k instances with the highest and lowest attention scores, k = 8.

[0120] 3-3 Construct an attention visualization module

[0121] The purpose of this module is to achieve interpretable visualization of WSI classification. It is specifically divided into three stages:

[0122] Attention score extraction stage: Extract the non-normalized attention scores of all tissue sections from the attention branch corresponding to the predicted class of the model (i.e., the original activation values before softmax probability conversion), and retain the relative importance relationship between regions; Statistical normalization stage: Use the quantile normalization method to convert the original scores into percentile scales in the [0,1] interval; Spatial visualization stage: Map the normalized values to RGB intensity values, where warm-colored regions represent high-attention features (discriminative morphological patterns that positively affect the classification decision), and cold-colored regions indicate low-attention regions (morphological patterns irrelevant to classification), thereby establishing the spatial association between the tissue microenvironment attention gradient and the model decision logic.

[0123] In the embodiment of the present invention, the warm color tone is the red series, and the cold color tone is the blue series. 1.0 represents the highest percentile of the attention intensity within the section.

[0124] Step 4, train the model and evaluate

[0125] 4-1 Train the classification model

[0126] Use the multi-modal chemotherapy response prediction model obtained in Step 3. Predicting chemotherapy response is a binary classification problem of a single label, that is, recurrence or non-recurrence after chemotherapy. The present invention makes full use of the spatial morphological information of WSI and gene expression data to construct a multi-modal chemotherapy response prediction model based on the attention mechanism and compact bilinear pooling. Input the preprocessed histopathological image data W into the network, and input it into the MCB module together with the gene feature F gene in the feature fusion stage to predict the chemotherapy response for different cancer data sets.

[0127] In the embodiment of the present invention, the model parameters are updated by the Adam optimizer with a learning rate of 0.0001 and a weight decay coefficient of 0.00001. The model is trained for at least 40 epochs, and at most 100 epochs if the early stopping criterion is not met. That is to say, the validation loss is monitored in each epoch, and when it does not decrease for more than 10 consecutive epochs, the early stopping method is used.

[0128] 4-2 Determine the training set and test set according to K-fold cross-validation

[0129] Perform ten-fold cross-validation on the data sets in Steps 1-3. First, divide the data into 10 parts, select 9 of them as the training set Train, and the remaining 1 part as the test set Test. Train 10 classifiers, and take the average of the test set results of the 10 classifiers as the final evaluation result of the model.

[0130] In the embodiment of the present invention, the data in the data set W' after data augmentation is used for the training set, and the data in the original data set W is used for the test set.

[0131] 4-3 Determine the evaluation metrics

[0132] The performance metrics for evaluating the model include accuracy (ACC), F1 value (F1-score), and area under the ROC curve (AUC). Since the present invention has a background in biomedical research, sensitivity (SEN) and specificity (SPE) commonly used in biomedical research are added to the evaluation metrics.

[0133] Example 1

[0134] Input the F obtained in Steps 2-3 wsi and the F obtained in Step 2-2 gene into the multi-modal chemotherapy response prediction model described in Step 3, train the model and evaluate it according to the method in Step 4 above. Retain the AUC, ACC, SEN, SPE, and F1-score of the test set evaluation results for each fold in Step 4-2, and take the average of the ten-fold cross-validation as the final result of the method BiChemoCLAM.

[0135] The calculation methods for the accuracy (ACC), sensitivity (SEN), specificity (SPE), and F1-score (F1-score) of each fold are as follows:

[0136]

[0137] Among them, TP, FP, TN, and FN represent true positive, false positive, true negative, and false negative, respectively.

[0138] According to the model evaluation method in step 4-2 above, ten-fold cross-validation is used to evaluate the model performance. Experimental data shows that for the chemotherapy response prediction problem, the method reaches 79.70%, 69.78%, and 70.83% on the OV, COAD, and BLCA cancer datasets, respectively. As shown in Table 1, the multi-modal chemotherapy response prediction method based on the attention mechanism and compact bilinear pooling has a high prediction accuracy.

[0139] Table 1 Model Evaluation Results

[0140]

[0141]

[0142] Although the present invention has been described above, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many variations without departing from the purpose of the present invention, and these all fall within the protection scope of the present invention.

Claims

1. A system for predicting multimodal chemotherapy response, characterized in that: The system includes a data set, a first data processing module, a second data processing module, a third data processing module, a bilinear pooling module, and an attention clustering module. The bilinear pooling module consists of a first feature vector dimensionality reduction unit, a second feature vector dimensionality reduction unit, and a fused feature vector unit; the attention clustering module consists of a gate attention unit, a clustering unit, and a classifier; where: The data set is based on pathological image data and gene data related to chemotherapy data; The first data processing module obtains tissue pathological image data blocks through multi-parameter collaborative acquisition of chemotherapy patient pathological image slice segmentation and hole elimination; The second data processing module standardizes gene expression data after screening out genes related to chemotherapy drugs to obtain a first pathological gene feature vector; The third data processing module processes tissue pathological image data blocks based on the ImageNet pre-trained ResNet-50 model to obtain a first pathological image feature vector; The bilinear pooling module combines the first pathological gene feature vector and the first pathological image feature vector based on a cross-modal interaction method to obtain a fused pathological image gene feature vector; The attention clustering module performs bag-level prediction on the fused pathological image features and the first pathological gene features based on the multi-instance learning method to obtain the patient's chemotherapy response. The process by which the bilinear pooling module combines the first pathological gene feature vector and the first pathological image feature vector based on a cross-modal interaction method to obtain a fused pathological image gene feature vector includes: The first feature vector dimensionality reduction unit calculates the second pathological image feature vector data from the first pathological image feature vector data through the Count Sketch mapping function, that is: The second feature vector dimensionality reduction unit calculates the second pathological gene feature vector data from the first pathological gene feature vector data through the Count Sketch mapping function, that is: The process by which the attention clustering module obtains the patient's chemotherapy response based on the multi-instance learning method from the fused pathological image gene feature vector includes: The gate attention unit calculates the attention score from the fused pathological image gene feature vector according to the following formula:

2. The system for predicting multi-modal chemotherapy response according to claim 1, wherein: The clustering unit obtains the pseudo-label according to the attention score according to the following formula: The classifier makes a patient's chemotherapy response based on the pseudo-binary label. S wsi = Ψ(F wsi ) (1) The multi-modal chemotherapy response prediction system further includes an attention visualization module; the attention visualization module consists of an attention score extraction unit, a statistical normalization unit, and a spatial visualization unit; Where: S wsi , S gene are the Count Sketches of WSI and gene respectively, Ψ is the Count Sketch mapping function, which compresses the feature dimension through the hash and sign functions, and d p is the low-rank projection dimension (d p << d w , d g ); The fusion feature vector unit calculates and obtains a fused pathological image gene feature vector from the second pathological gene feature vector data and the second pathological image feature vector data through the following formula, that is: F fusion Where: FO is the Fast Fourier Transform that converts a time-domain signal to the frequency domain, ⊙ is the element-wise multiplication (Hadamard product), and FO -1 is the Inverse Fast Fourier Transform.

3. The system for predicting multi-modal chemotherapy response according to claim 1, characterized in that: The attention score extraction unit extracts the unnormalized attention scores of all tissue sections from the attention branch corresponding to the system prediction category; The statistical normalization unit converts the original scores into percentile scales in the [0,1] interval using the quantile normalization method; where: h k is the feature vector of the k-th instance, and W n is the learnable weight matrix corresponding to the n-th class, is the attention score of the k-th instance for the n-th class, reflecting its pathological relevance; In the spatial visualization stage: a spatial association between the tissue microenvironment attention gradient and the system decision logic is established by mapping the normalized values to RGB intensity values. where: top k and bottom k represent the k instances with the highest and lowest attention scores respectively, which are pseudo binary labels; ​ 4. A system for predicting multimodal chemotherapy response according to any one of claims 1-3, characterized in that: ​ ​ ​ ​ 5. A method for predicting multimodal chemotherapy response, characterized in that, The method includes: the construction stage of the multi-modal chemotherapy response prediction system and the training and testing stage of the multi-modal chemotherapy response prediction system. Among them: In the construction stage of the multi-modal chemotherapy response prediction system, the following steps are included: Step 101: Construct a dataset based on pathological image data and genes related to chemotherapy data; Step 102: Obtain tissue pathological image data blocks through multi-parameter collaborative acquisition of chemotherapy patient pathological image slice segmentation and hole elimination; Step 103: After screening out genes related to chemotherapy drugs, perform normalization processing on gene expression data to obtain a first pathological gene feature vector; that is: d g is the dimension of the gene feature; Step 104: Use the ResNet-50 model pre-trained on ImageNet to process the tissue pathology image data block to obtain the first pathological image feature vector, that is: d w is the dimension of the pathological image feature; Step 105: Process the first pathological gene feature vector and the first pathological image feature vector respectively through a hash mapping function to obtain a second pathological gene feature vector and a second pathological image feature vector; S wsi = Ψ(F wsi ), Where: S wsi , S gene are the Count Sketches of WSI and gene respectively, Ψ is the Count Sketch mapping function, which compresses the feature dimension through the hash and sign functions, and d p is the low-rank projection dimension; Step 106. Calculate and obtain a fused pathological image gene feature vector from the second pathological gene feature vector data and the second pathological image feature vector data through the following formula, i.e., F fusion Where: FO is the Fast Fourier Transform that converts a time-domain signal to a frequency domain, ⊙ is the element-wise multiplication (Hadamard product), and FO -1 is the Inverse Fast Fourier Transform; Step 107: Calculate the attention score for the fused pathological image gene feature vector according to the following formula; where: h k is the feature vector of the k-th instance, and W n is the learnable weight matrix corresponding to the n-th class, is the attention score of the k-th instance for the n-th class, reflecting its pathological relevance; Step 108: Obtain the pseudo-label according to the attention score according to the following formula; Among them: top k and bottom k respectively represent the k instances with the highest and lowest attention scores, which are pseudo binary labels; Make a patient's chemotherapy response based on the pseudo-binary label.

6. The method for predicting multi-modal chemotherapy response according to claim 5, characterized in that, In the training and testing stage of the multi-modal chemotherapy response prediction system, the following steps are included: Step 201: Update the parameters of the multi-modal chemotherapy response prediction system through the Adam optimizer with a learning rate of 0.0001 and a weight decay coefficient of 0.00001; Step 202: Determine the training set and the test set according to K-fold cross-validation: Perform ten-fold cross-validation on the dataset in steps 101-103. First, divide the data into 10 parts, select 9 of them as the training set Train, and the remaining 1 part as the test set Test. Train 10 classifiers, and take the average of the test set results of the 10 classifiers as the final evaluation result of the model; Step 203: Determine the evaluation index.

Citation Information

Cited By

  • Clinical multi-mode cancer drug response prediction method based on feature reconstruction

    CN121528291A

  • A Feature Reconstruction-Based Method for Predicting Clinical Multimodal Cancer Drug Response

    CN121528291B