Prediction method and system for pan-cancer biomarkers based on multi-instance multi-task learning
Through the multi-instance multi-task learning method, the twin multi-instance fully convolutional network and linear attention mechanism are used to process the full-field digital pathology images of cancer patients, which solves the problem of insufficient multi-task learning in cancer research in the existing technology and realizes efficient prediction and early screening assistance for multiple biomarkers.
Patent Information
- Application Number
- CN202510041802.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-01-10
AI Technical Summary
Existing computational pathology methods in cancer research suffer from heavy workload, low result accuracy and reliability, high computational resource consumption, and insufficient multi-task learning, which limits the ability to simultaneously analyze multiple biomarkers.
A multi-instance multi-task learning method is adopted to preprocess and extract features of full-field digital pathology images of patients with solid tumors through a twin multi-instance fully convolutional network model and a linear attention mechanism, and biomarkers are predicted in combination with a multi-task classification prediction model.
It achieves simultaneous, efficient, and effective prediction of multiple pan-cancer biomarkers, assisting doctors in early screening and selecting clinical treatment plans, and optimizing the diagnosis and treatment process.
Smart Images

Figure CN119888349B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical pathology image technology, and specifically relates to a method and system for predicting pan-cancer biomarkers based on multi-instance multi-task learning. Background Art
[0002] Molecular profiling plays a crucial role in cancer research and treatment. It helps scientists better understand the mechanisms of tumorigenesis and provides a foundation for precision medicine. Precision medicine relies on the accurate diagnosis of genetic alterations, which are typically identified through molecular biology assays such as Sanger sequencing and fluorescence in situ hybridization. However, these assays have several limitations, such as long turnaround times, the need for specialized expertise, and high costs, which have limited their widespread adoption in oncology workflows.
[0003] Recent advances in computational pathology (CPath) have provided new perspectives for cancer research. Studies have shown that histological sections stained with hematoxylin and eosin (H&E) not only contain known biomarker information but may also harbor potential biomarkers. Using digital pathology slides (WSI) technology, scientists can infer molecular lineages directly from the sections. This approach has the advantage of not requiring complex molecular biology testing but instead leveraging technologies such as deep learning to extract useful information from tissue sections.
[0004] However, despite the enormous potential shown in computational pathology (CPath) research, certain limitations remain. The currently used manual feature engineering increases workload and is susceptible to subjective experience bias, which in turn reduces the accuracy and reliability of the results. The lightweight ShuffleNet model faces challenges in learning complex patterns or advanced functions, which limits its performance in tasks requiring high global correlation. Traditional global attention methods in CPath applications require calculating the similarity between a pixel and all other pixels, which consumes a large amount of time and computing resources, increasing research costs and limiting the scale and speed of research. Furthermore, the expression of different molecular markers in tissues is often not isolated; many are simultaneously over- or under-expressed. Current multi-task learning (MTL) approaches have been relatively underexplored in CPath research, limiting the ability to analyze tissue sections from multiple perspectives simultaneously. Summary of the Invention
[0005] The present invention provides a multi-instance multi-task learning method for predicting pan-cancer biomarkers. By preprocessing and extracting full-field digital pathology images of solid tumor patients to obtain pathological features, the pathological features are weighted and calculated based on a twin multi-instance fully convolutional network model in combination with a linear attention mechanism, and then classified and predicted. This method enables simultaneous, efficient, and effective prediction of multiple pan-cancer biomarkers based on full-field digital pathology images of solid tumor patients, thereby assisting doctors in clinically screening for cancer patients for markers, replicating clinical treatment options, and optimizing the diagnosis and treatment process.
[0006] A multi-instance multi-task learning method for predicting pan-cancer biomarkers, including:
[0007] Acquire full-field digital pathology images and pan-cancer biomarker expression data for patients with solid tumors;
[0008] Based on the full-field digital pathology images, preprocessing and feature extraction are performed to obtain a pathology feature dataset;
[0009] Based on the pathological feature dataset, it is input into the twin multi-instance fully convolutional network model for training to obtain low-dimensional pathological features. Combined with the linear attention mechanism, the attention score corresponding to each pathological feature is calculated;
[0010] By combining each pathological low-dimensional feature and the corresponding attention score, the pathological deep feature is calculated and input into the multi-task classification prediction model for classification prediction to obtain the prediction results of each biomarker in pan-cancer.
[0011] By preprocessing and extracting pathological features from the full-field digital pathology images of solid tumor patients, the pathological features are weighted and calculated based on the twin multi-instance fully convolutional network model and combined with the linear attention mechanism for classification and prediction. This makes it possible to simultaneously, efficiently and effectively predict multiple pan-cancer biomarkers based on the full-field digital pathology images of solid tumor patients, thereby assisting doctors in clinically screening for cancer patients' markers early, replicating clinical treatment plan selection, and optimizing the diagnosis and treatment process.
[0012] Furthermore, based on the full-field digital pathology images, preprocessing and feature extraction are performed to obtain a pathology feature dataset, including:
[0013] Based on the full-field digital pathology images, the Otsu threshold method was used for segmentation to obtain the tissue area of each full-field digital pathology image;
[0014] Downsampling is performed based on the tissue area of each full-field digital pathology image;
[0015] Based on the tissue area of the full-field digital pathology image after downsampling, combined with the preset minimum threshold of the RGB channel median, patches with lack of tissue or minimal tissue are excluded to obtain the tissue area of the preprocessed full-field digital pathology image;
[0016] Based on the tissue areas of the preprocessed full-field digital pathology images, a fine-tuned pre-trained ResNet50 model was used for feature extraction to obtain the pathological features of each full-field digital pathology image, and a pathological feature dataset of full-field digital pathology images was constructed.
[0017] Furthermore, the tissue area of the full-view digital pathology image is a square block, and the tissue areas of multiple full-view digital pathology images do not overlap; the edge length of each square block is , the size is Pixels.
[0018] Furthermore, based on the tissue regions of the preprocessed full-field digital pathology images, a fine-tuned pre-trained ResNet50 model was used to extract features to obtain the pathological features of each full-field digital pathology image. A pathological feature dataset of full-field digital pathology images was constructed, including:
[0019] The ResNET50 model is selected as the base model;
[0020] Cluster-guided contrastive learning is used to pre-train the ResNET50 model to obtain a pre-trained ResNET50 model;
[0021] Based on the pre-trained ResNET50 model, a fully connected layer is added at the end to fine-tune the model to obtain a fine-tuned pre-trained ResNet50 model;
[0022] Based on the multi-instance MIL framework, a fine-tuned pre-trained ResNet50 model was used to extract features of the tissue area of each full-field digital pathology image, obtain the pathological features of each full-field digital pathology image, and construct a pathological feature dataset of full-field digital pathology images.
[0023] By introducing a fine-tuned pre-trained ResNET model for feature extraction based on the multi-instance MIL framework, the time-consuming and tedious manual labeling is eliminated, and an information-rich embedding representation of pathological features can be obtained.
[0024] Furthermore, based on the pathological feature dataset, it is input into the twin multi-instance fully convolutional network model for training to obtain low-dimensional pathological features. Combined with the linear attention mechanism, the attention score corresponding to each pathological feature is calculated, including:
[0025] Based on a fully convolutional neural network, a twin-architecture two-stream subnet is constructed; the two-stream subnet includes two weight-sharing subnets; each subnet consists of one convolutional layer, two rectifier unit activation layers, one adaptive average pooling layer, and two stacked fully connected layers;
[0026] Based on the pathological feature dataset, it is input into the two-stream sub-network for training to obtain the pathological low-dimensional features;
[0027] The expression of the pathological low-dimensional feature is:
[0028] ;
[0029] Where, Indicates the The low-dimensional feature embedding representation corresponding to the pathological embedded features, , represents the total number of pathological features; represents the activation function; Represents the weight matrix of the first fully connected layer; Represents the bias vector of the first fully connected layer; Represents the weight matrix of the second fully connected layer; Represents the bias vector of the second fully connected layer; Indicates the The unmapped embedding vector, pathological features;
[0030] Based on the low-dimensional features of the pathology, they are input into the focused linear attention module for training to obtain the attention score corresponding to each pathology feature;
[0031] The expression of the attention score corresponding to each pathological feature is:
[0032] ;
[0033] Where, Indicates the first pathological features and Attention scores for each task; Indicates the pathological features and The query vector of each task; Indicates the The key vector of tokens; Indicates the The value vector of tokens; , Indicates the number of tokens in the input sequence, that is, the total number of elements that the model needs to process; represents the depth convolution function; Represents a mapping function or feature transformation function.
[0034] By designing a dual-stream subnet constructed by a multi-instance fully convolutional neural network and combining it with a focused linear attention network model, the multi-layer fully convolutional layer is combined with a nonlinear activation function, which can effectively reduce the number of model parameters and enable the model to better focus on important features; by introducing nonlinear transformations, it can help the model capture more complex feature relationships, thereby improving the model's expressiveness and performance.
[0035] Furthermore, by combining each pathology low-dimensional feature and the corresponding attention score, we calculated the pathology deep features. These deep features were then fed into a multi-task classification prediction model for classification prediction, yielding the prediction results for each pan-cancer biomarker, including:
[0036] Based on a given slide, the deep features of the given task are calculated by combining the pathological low-dimensional features and their corresponding attention scores;
[0037] The expression of the depth feature of the given task is:
[0038] ;
[0039] Where, Indicates the Deep feature representation of all blocks in a given slide for each task;
[0040] Based on the deep features of a given task, it is input into a multi-task classification prediction model for classification prediction, and a slice-level probability prediction score of each task is calculated to obtain a prediction result for each biomarker in pan-cancer; the multi-task classification prediction model includes multiple classifiers;
[0041] The expression of the slice-level probability prediction score of each task is:
[0042] ;
[0043] Where, Indicates the Slice-level probability prediction score for each task; Normalization function; Indicates the The weight matrix of each task in the classification layer; Indicates the The bias vector of each task in the classification layer.
[0044] By designing a multi-task classification prediction model containing multiple classifiers, it is possible to output predictions for different markers simultaneously.
[0045] Furthermore, the method further includes calculating a loss function based on the multi-task classification prediction model; the calculation of the loss function includes:
[0046] Based on each slide, the loss function of a single task is calculated using the standard cross entropy loss function;
[0047] Based on the loss function of a single task, the total loss function of all tasks is calculated by weighted sum;
[0048] The expression of the total loss function is:
[0049] ;
[0050] Where, represents the total loss of all tasks; Indicates the The weight of the task, , Indicates the total number of tasks, ; Indicates the The loss function for each task.
[0051] A system for predicting pan-cancer biomarkers using multi-instance multi-task learning, comprising:
[0052] A data acquisition module for acquiring full-field digital pathology images and pan-cancer biomarker expression data for patients with solid tumors;
[0053] A data processing module is used to perform preprocessing and feature extraction based on the full-field digital pathology image to obtain a pathology feature data set;
[0054] The calculation module is used to input the pathological feature dataset into the twin multi-instance fully convolutional network model for training to obtain low-dimensional pathological features. In combination with the linear attention mechanism, it calculates the attention score corresponding to each pathological feature.
[0055] The prediction module is used to combine each pathology low-dimensional feature and the corresponding attention score to calculate the pathology deep feature, and input the pathology deep feature into the multi-task classification prediction model for classification prediction to obtain the prediction results of each biomarker in pan-cancer.
[0056] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the above method when executing the computer program.
[0057] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0058] The beneficial effects of the present invention are:
[0059] The present invention obtains pathological features by preprocessing and extracting full-field digital pathology images of solid tumor patients. It then performs classification prediction after weighted calculation of the pathological features based on a twin multi-instance fully convolutional network model in combination with a linear attention mechanism. This enables simultaneous, efficient, and effective prediction of multiple pan-cancer biomarkers based on full-field digital pathology images of solid tumor patients, thereby assisting doctors in clinically conducting early screening for markers in cancer patients, replicating clinical treatment plan selection, and optimizing the diagnosis and treatment process. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 is a flow chart of the present invention;
[0061] Figure 2 This is the ROC curve of the test set;
[0062] Figure 3 Schematic diagram of the structure of the system of the present invention;
[0063] Figure 4 It is a structural diagram of the computer device in the present invention. DETAILED DESCRIPTION
[0064] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0065] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement an apparatus and / or practice a method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this apparatus and / or practice this method.
[0066] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood in specific situations.
[0067] Example 1
[0068] Figure 1 The method shown here is a multi-instance multi-task learning method for predicting pan-cancer biomarkers. It preprocesses and extracts pathological features from full-field digital pathology images of solid tumor patients. It then uses a twin multi-instance fully convolutional network model and a linear attention mechanism to weight the pathological features and perform classification predictions. This allows for simultaneous, efficient, and effective prediction of multiple pan-cancer biomarkers based on full-field digital pathology images of solid tumor patients. This method can assist physicians in early screening for cancer biomarkers, replicate clinical treatment options, and optimize the diagnosis and treatment process. Specifically, it includes:
[0069] S1: Acquire full-field digital pathology images (WSIs) and pan-cancer biomarker expression data for patients with solid tumors;
[0070] In this example, the full-field digital pathology images (WSI) and pan-cancer biomarker expression data of solid tumor patients required for the experiment were downloaded from The Cancer Genome Atlas (TCGA) and The Clinical ProteomicTumor Analysis Consortium (CPTAC) databases.
[0071] S2: Based on the full-field digital pathology image, preprocessing and feature extraction are performed to obtain a pathology feature dataset;
[0072] S11: Based on the full-field digital pathology image WSI, the Otsu threshold method is used for segmentation to obtain the tissue region patch of each full-field digital pathology image WSI;
[0073] In this embodiment, the tissue area of the full-view digital pathology image is a square block, and the tissue areas of multiple full-view digital pathology images do not overlap; the edge length of each square block is , the size is Pixels.
[0074] S12: performing downsampling processing based on the tissue region patch of each full-field digital pathology image WSI;
[0075] In this embodiment, 64× resolution is used for downsampling and reading into memory; in actual applications, other single resolutions or multiple resolutions can be selected for sampling to obtain hierarchical pathological information and thus obtain rich pathological features.
[0076] S13: Based on the tissue region patch of the full-field digital pathology image WSI after downsampling, combined with a preset minimum threshold of the RGB channel median, patches with a lack of tissue or minimal tissue are excluded to obtain a preprocessed tissue region patch of the full-field digital pathology image WSI;
[0077] In this embodiment, patches with little or minimal tissue are excluded by setting a minimum threshold of 50 for the median of each 8-bit RGB channel.
[0078] S14: Based on the preprocessed tissue region patches of the full-field digital pathology image WSI, a fine-tuned pre-trained ResNet50 model is used to perform feature extraction to obtain the pathological features of each full-field digital pathology image WSI, and a pathological feature dataset of the full-field digital pathology image WSI is constructed;
[0079] S141: Select the ResNET50 model as the base model;
[0080] S142: Pre-train the ResNET50 model using cluster-guided contrastive learning (CCL) to obtain a pre-trained ResNET50 model;
[0081] S143: Based on the pre-trained ResNET50 model, a fully connected layer is added to the end to fine-tune the model to obtain a fine-tuned pre-trained ResNet50 model;
[0082] S144: Based on the multi-instance MIL framework, a fine-tuned pre-trained ResNet50 model is used to extract features from the tissue region patches of each full-field digital pathology image (WSI). Pathological features of each full-field digital pathology image are obtained, and a pathological feature dataset of full-field digital pathology images is constructed.
[0083] In this embodiment, each full-field digital pathology image WSI is regarded as a package, in which each tissue region patch is regarded as an example, and after feature extraction, the image containing The WSI of the example, and then for the The patch-level embedding representation of the patch, that is, the pathological feature is Among them, the dimension of the pathological feature output is 2048, that is, .
[0084] S3: Based on the pathological feature dataset, it is input into the twin multi-instance fully convolutional network model for training to obtain low-dimensional pathological features. Combined with the linear attention mechanism, the attention score corresponding to each pathological feature is calculated;
[0085] S31: Based on the fully convolutional neural network (FCN), a twin-architecture two-stream subnet is constructed. The two-stream subnet consists of two weight-sharing subnets. Each subnet consists of one convolutional layer, two rectifier activation layers, one adaptive average pooling layer, and two stacked fully connected layers.
[0086] Each subnetwork is capable of learning a fully convolutional neural network (FCN) for information representation. The convolutional layer has a kernel size of 1×1 to change the number of channels in the feature map. Two rectifier unit activation layers are used to increase the model's nonlinearity. An adaptive average pooling layer is used to dynamically adjust the size of the pooling window. Two stacked fully connected layers are used to integrate information from the feature map for fine-tuning.
[0087] S32: Based on the pathological feature dataset, input it into the two-stream subnet for training to obtain the pathological low-dimensional features ;
[0088] Among them, the expression of pathological low-dimensional features is:
[0089] ;
[0090] Where, Indicates the The low-dimensional feature embedding representation corresponding to the pathological embedded features, , represents the total number of pathological features; represents the activation function; represents the weight matrix of the first fully connected layer, and ; represents the bias vector of the first fully connected layer, and ; represents the weight matrix of the second fully connected layer, and ; represents the bias vector of the second fully connected layer, and ; Indicates the The unmapped embedding vector, pathological features;
[0091] S33: Based on pathological low-dimensional features , input it into the focused linear attention module for training, and obtain the attention score corresponding to each pathological feature ;
[0092] In this embodiment, the pathological low-dimensional feature set , as the input of the focused linear attention module. All tasks are learned jointly and the model parameters in the second fully connected layer are shared. The focused linear attention model is trained for each patch and task Assign an attention score.
[0093] Among them, the expression of the attention score corresponding to each pathological feature is:
[0094] ;
[0095] Where, Indicates the first pathological features and Attention scores for each task; Indicates the pathological features and The query vector of each task; Indicates the The key vector of tokens; Indicates the The value vector of tokens; , Indicates the number of tokens in the input sequence, that is, the total number of elements that the model needs to process; represents the depth convolution function; Represents a mapping function or feature transformation function.
[0096] In this embodiment, according to the associative property of matrix multiplication, the calculation order is changed from Adjust to , thereby reducing the computational complexity of tokens from Restricted to , reducing the number of model parameters so that the model can better focus on important features; and by introducing nonlinear transformation , which can help the model capture more complex feature relationships, thereby improving the model's expressiveness and performance.
[0097] S4: Combine each pathological low-dimensional feature And the corresponding attention score , calculate the pathological deep features, and input the pathological deep features into the multi-task classification prediction model for classification prediction to obtain the prediction results of each biomarker in pan-cancer;
[0098] S41: Based on a given slide, combined with pathological low-dimensional features and its corresponding attention score , calculate the given task Deep features;
[0099] Among them, the given task The expression of the depth feature is:
[0100] ;
[0101] Where, Indicates the Deep feature representation of all blocks in a given slide for each task;
[0102] S42: Based on a given task Deep features , input it into the multi-task classification prediction model for classification prediction, and calculate the slice-level probability prediction score of each task , get the prediction results of each biomarker in Pan-Cancer;
[0103] In this embodiment, the multi-task classification prediction model includes multiple classifiers. By designing a multi-task classification prediction model including multiple classifiers, it is possible to facilitate simultaneous output predictions of different markers.
[0104] Among them, the slice-level probability prediction score of each task is The expression is:
[0105] ;
[0106] Where, Indicates the Slice-level probability prediction score for each task; Normalization function; Indicates the The weight matrix of each task in the classification layer; Indicates the The bias vector of each task in the classification layer.
[0107] In this embodiment, a total loss function is calculated based on the multi-task classification prediction model. The calculation process of the total loss function includes:
[0108] S51: Based on each slide, the loss function of a single task is calculated using the standard cross entropy loss function ;
[0109] In this embodiment, in each task The classification layer cls is composed of the weight matrix of the classification layer , and the bias vector of the classification layer defined, Indicates a task The number of classes.
[0110] S52: Based on the loss function of a single task, the total loss function of all tasks is calculated using weighted sum ;
[0111] Among them, the expression of the total loss function is:
[0112] ;
[0113] Where, represents the total loss of all tasks; Indicates the The weight of the task, , Indicates the total number of tasks, ; Indicates the The loss function for each task.
[0114] In this embodiment, the five-fold cross-validation method and the AUC method are also used to evaluate the performance of the multi-task classification prediction model.
[0115] Example 2
[0116] In this embodiment, a method for predicting pan-cancer biomarkers using multi-instance multi-task learning is provided, comprising:
[0117] T1: Acquire full-field digital pathology images (WSI) and pan-cancer biomarker expression data for patients with solid tumors;
[0118] T11: Based on the TCGA database, a solid tumor dataset was obtained and divided into a training set and a validation set;
[0119] In this embodiment, the TCGA database contains 12 solid tumor datasets, namely, a total of 4527 WSIs and corresponding pan-cancer biomarker expression data of 4257 patients; the pan-cancer type image data are shown in Table 1.
[0120] Table 1 Number of images and patients included in the pan-cancer multi-task study
[0121]
[0122] T12: Based on the CPTAC database, obtain the colorectal cancer dataset and use it as the test set;
[0123] In this embodiment, the CPTAC database is an external test set, in which the colorectal cancer dataset includes 221 WSIs of 105 patients and the corresponding pan-cancer biomarker expression data.
[0124] T2: Based on the full-field digital pathology image WSI, preprocessing and feature extraction are performed to obtain a pathology feature dataset;
[0125] T11: Based on the full-field digital pathology image WSI, the Otsu threshold method is used for segmentation to obtain the tissue region patch of each full-field digital pathology image WSI;
[0126] T12: Downsampling processing is performed based on the tissue area patch of each full-field digital pathology image WSI;
[0127] T13: Based on the tissue region patches of the downsampled full-field digital pathology image WSI, combined with the preset minimum threshold of 50 of the 8-bit RGB channel median, patches with lack of tissue or minimal tissue are excluded to obtain the preprocessed tissue region patches of the full-field digital pathology image WSI;
[0128] T14: Based on the preprocessed tissue region patches of the full-field digital pathology image (WSI), a fine-tuned pre-trained ResNet50 model is used to extract features to obtain the pathological features of each full-field digital pathology image (WSI). This constructs a pathological feature dataset for the full-field digital pathology image (WSI).
[0129] T141: Select the ResNET50 model as the base model;
[0130] T142: Use cluster-guided contrastive learning (CCL) to pre-train the ResNET50 model to obtain a pre-trained ResNET50 model;
[0131] T143: Based on the pre-trained ResNET50 model, add a fully connected layer at the end to fine-tune the model to obtain the fine-tuned pre-trained ResNet50 model;
[0132] T144: Based on the multi-instance MIL framework, a fine-tuned pre-trained ResNet50 model is used to extract features from tissue region patches of each full-field digital pathology image (WSI). Pathological features of each full-field digital pathology image are obtained, and a pathological feature dataset of full-field digital pathology images is constructed.
[0133] T3: Based on the pathological feature dataset, it is input into the twin multi-instance fully convolutional network model for training to obtain low-dimensional pathological features. Combined with the linear attention mechanism, the attention score corresponding to each pathological feature is calculated;
[0134] T31: Based on the fully convolutional neural network (FCN), a twin-architecture two-stream subnet is constructed. The two-stream subnet consists of two weight-sharing subnets. Each subnet consists of one convolutional layer, two rectifier unit activation layers, one adaptive average pooling layer, and two stacked fully connected layers.
[0135] T32: Based on the pathological feature dataset, input it into the two-stream subnet for training to obtain pathological low-dimensional features ;
[0136] T33: Based on pathological low-dimensional features , input it into the focused linear attention module for training, and obtain the attention score corresponding to each pathological feature ;
[0137] T4: Combine each pathological low-dimensional feature And the corresponding attention score , calculate the pathological deep features, and input the pathological deep features into the multi-task classification prediction model for classification prediction to obtain the prediction results of each biomarker in pan-cancer;
[0138] T41: Based on a given slide, combined with pathological low-dimensional features and its corresponding attention score , calculate the given task Deep features;
[0139] T42: Based on a given task Deep features , input it into the multi-task classification prediction model for classification prediction, and calculate the slice-level probability prediction score of each task and the total loss function to obtain the prediction results of each biomarker in Pan-Cancer;
[0140] T421: Based on each slide, the loss function of a single task is calculated using the standard cross entropy loss function , and based on the loss function of a single task, the total loss function of all tasks is calculated using weighted sum ;
[0141] T422: Based on a given task Deep features , input it into the multi-task classification prediction model for classification prediction, and calculate the slice-level probability prediction score of each task , and obtain the prediction results of each biomarker in Pan-Cancer.
[0142] The training process of the multi-task classification prediction model includes:
[0143] W11: Use the Adam optimizer, set the learning rate to 1e-2, and the weight decay to 4e-5;
[0144] W12: Add a Dropout layer after each hidden layer and set P=0.25 to avoid over-convergence;
[0145] W13: Set the number of iterations; the range of iterations is 50~200 epochs;
[0146] W14: Based on the preset number of iterations, training stop judgment is performed;
[0147] W141: When the total loss function of the validation set does not decrease after 20 consecutive epochs, training is stopped and the best multi-task classification prediction model is generated;
[0148] W142: When the total loss function of the validation set decreases within 20 consecutive epochs, continue training.
[0149] Before applying the multi-task classification prediction model, the model is trained on a test set and the AUC method is used to evaluate the model performance, including:
[0150] W21: Randomly sample with replacement based on the test set, generating a total of 1000 bootstrap samples;
[0151] W22: Calculate the AUC score based on each bootstrap sample;
[0152] In this embodiment, the closer the AUC score is to 1, the better the model performance is; the closer the AUC score is to 0.5, the worse the model performance is.
[0153] W23: Based on the calculated AUC scores, an AUC score distribution is generated. Using the percentile method, the 2.5% percentile is used as the lower limit of the confidence interval, and the 97.5% percentile is used as the upper limit of the confidence interval. This constitutes a confidence interval, which is used to evaluate the stability of the model's performance on the validation set. A narrow confidence interval indicates a more accurate evaluation result, whereas a narrow confidence interval indicates a greater degree of uncertainty.
[0154] Figure 2 Shown is the ROC curve of the test set, and the colorectal cancer instability detection result was obtained: AUC: 0.769 (95CI, 0.686~0.841).
[0155] Table 2 shows the average AUC performance and standard deviation for all cancer types.
[0156] Table 2. Average AUC performance and standard deviation of all implemented cancer types for multi-task learning
[0157]
[0158] Table 2 shows that across the 12 solid tumor datasets, KICH achieved the best overall performance, with a confidence interval of [0.850-0.050, 0.850+0.050]; PAAD achieved the worst overall performance, with a confidence interval of [0.557-0.056, 0.557+0.056]. The effectiveness and robustness of this method were confirmed by training and validation on internal solid tumor datasets and testing on an external colorectal cancer dataset.
[0159] Example 3
[0160] Based on the same technical concept, such as Figure 3 As shown, this embodiment also provides a multi-instance multi-task learning pan-cancer biomarker prediction system, including a data acquisition module, a data processing module, a calculation module and a prediction module.
[0161] Specifically, a data acquisition module, which is used to acquire full-field digital pathology images and pan-cancer biomarker expression data for patients with solid tumors;
[0162] Specifically, the data processing module is used to perform preprocessing and feature extraction based on the full-field digital pathology image to obtain a pathology feature dataset;
[0163] Specifically, the calculation module is used to input the pathological feature dataset into the twin multi-instance fully convolutional network model for training to obtain pathological low-dimensional features, and combine it with the linear attention mechanism to calculate the attention score corresponding to each pathological feature;
[0164] Specifically, the prediction module is used to combine each pathological low-dimensional feature and the corresponding attention score to calculate the pathological deep feature, and input the pathological deep feature into the multi-task classification prediction model for classification prediction to obtain the prediction results of each biomarker in pan-cancer.
[0165] Example 3
[0166] Based on the same technical concept, the embodiment of the present application also provides a computer device, including a memory 1 and a processor 2, such as Figure 4 As shown, the memory 1 stores a computer program, and the processor 2 implements any of the above methods when executing the computer program.
[0167] Memory 1 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, memory 1 can be an internal storage unit of the multi-instance multi-task learning pan-cancer biomarker prediction system, such as a hard disk. In other embodiments, memory 1 can also be an external storage device of the multi-instance multi-task learning pan-cancer biomarker prediction system, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, memory 1 can include both an internal storage unit and an external storage device of the multi-instance multi-task learning pan-cancer biomarker prediction system. Memory 1 can be used not only to store application software and various data installed in the multi-instance multi-task learning pan-cancer biomarker prediction system, such as the program code of the multi-instance multi-task learning pan-cancer biomarker prediction system, but also to temporarily store data that has been output or is about to be output.
[0168] In some embodiments, processor 2 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip, configured to execute program code stored in memory 1 or process data, such as a system program for predicting pan-cancer biomarkers using multi-instance multi-task learning.
[0169] The disclosed embodiments of the present invention further provide a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to perform the steps of the method described in the above method embodiments. The storage medium may be a volatile or non-volatile computer-readable storage medium.
[0170] The computer program product of the application page content refresh method provided in the disclosed embodiment of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the steps of the method described in the above method embodiment. For details, please refer to the above method embodiment and will not be repeated here.
[0171] The disclosed embodiments of the present invention further provide a computer program that, when executed by a processor, implements any of the methods of the aforementioned embodiments. This computer program product can be implemented in hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is embodied as a computer storage medium. In another alternative embodiment, the computer program product is embodied as a software product, such as a software development kit (SDK).
[0172] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.
[0173] It should be noted that, in the description of the present invention, the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "plurality" is at least two.
[0174] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0175] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0176] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0177] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.
[0178] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0179] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0180] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A multi-instance multi-task learning method for predicting pan-cancer biomarkers, characterized by: include: Acquire full-field digital pathology images and pan-cancer biomarker expression data for patients with solid tumors; Based on the full-field digital pathology images, preprocessing and feature extraction are performed to obtain a pathology feature dataset; Based on the pathological feature dataset, it is input into the twin multi-instance fully convolutional network model for training to obtain low-dimensional pathological features. Combined with the linear attention mechanism, the attention score corresponding to each pathological feature is calculated; Combining each pathology low-dimensional feature and the corresponding attention score, we calculate the pathology deep feature, which is then fed into a multi-task classification prediction model for classification prediction, yielding the prediction results for each pan-cancer biomarker. Based on the pathological feature dataset, the dataset is input into the twin multi-instance fully convolutional network model for training to obtain low-dimensional pathological features. In combination with the linear attention mechanism, the attention score corresponding to each pathological feature is calculated, including: Based on a fully convolutional neural network, a twin-architecture two-stream subnet is constructed; the two-stream subnet includes two weight-sharing subnets; each subnet consists of one convolutional layer, two rectifier unit activation layers, one adaptive average pooling layer, and two stacked fully connected layers; Based on the pathological feature dataset, it is input into the two-stream sub-network for training to obtain the pathological low-dimensional features; The expression of the pathological low-dimensional feature is: ; Where, Indicates the The low-dimensional feature embedding representation corresponding to the pathological embedded features, , represents the total number of pathological features; represents the activation function; Represents the weight matrix of the first fully connected layer; Represents the bias vector of the first fully connected layer; Represents the weight matrix of the second fully connected layer; Represents the bias vector of the second fully connected layer; Indicates the The unmapped embedding vector, pathological features; Based on the low-dimensional features of the pathology, they are input into the focused linear attention module for training to obtain the attention score corresponding to each pathology feature; The expression of the attention score corresponding to each pathological feature is: ; Where, Indicates the first pathological features and Attention scores for each task; Indicates the pathological features and The query vector of each task; Indicates the The key vector of tokens; Indicates the The value vector of tokens; , Indicates the number of tokens in the input sequence, that is, the total number of elements that the model needs to process; represents the depth convolution function; Represents a mapping function or feature transformation function.
2. The method for predicting pan-cancer biomarkers using multi-instance multi-task learning according to claim 1, characterized in that: Based on the full-field digital pathology images, preprocessing and feature extraction are performed to obtain a pathology feature dataset, including: Based on the full-field digital pathology images, the Otsu threshold method was used for segmentation to obtain the tissue area of each full-field digital pathology image; Downsampling is performed based on the tissue area of each full-field digital pathology image; Based on the tissue area of the full-field digital pathology image after downsampling, combined with the preset minimum threshold of the RGB channel median, patches with lack of tissue or minimal tissue are excluded to obtain the tissue area of the preprocessed full-field digital pathology image; Based on the tissue areas of the preprocessed full-field digital pathology images, a fine-tuned pre-trained ResNet50 model was used for feature extraction to obtain the pathological features of each full-field digital pathology image, and a pathological feature dataset of full-field digital pathology images was constructed.
3. The method for predicting pan-cancer biomarkers using multi-instance multi-task learning according to claim 2, characterized in that: The tissue area of the full-view digital pathology image is a square block, and the tissue areas of multiple full-view digital pathology images do not overlap; the edge length of each square block is , the size is Pixels.
4. The method for predicting pan-cancer biomarkers using multi-instance multi-task learning according to claim 2, wherein: Based on the tissue regions of the preprocessed full-field digital pathology images, a fine-tuned pre-trained ResNet50 model is used for feature extraction to obtain the pathological features of each full-field digital pathology image. A pathological feature dataset of full-field digital pathology images is constructed, including: The ResNET50 model is selected as the base model; Cluster-guided contrastive learning is used to pre-train the ResNET50 model to obtain a pre-trained ResNET50 model; Based on the pre-trained ResNET50 model, a fully connected layer is added at the end to fine-tune the model to obtain a fine-tuned pre-trained ResNet50 model; Based on the multi-instance MIL framework, a fine-tuned pre-trained ResNet50 model was used to extract features of the tissue area of each full-field digital pathology image, obtain the pathological features of each full-field digital pathology image, and construct a pathological feature dataset of full-field digital pathology images.
5. The method for predicting pan-cancer biomarkers using multi-instance multi-task learning according to claim 1, characterized in that: By combining each low-dimensional pathology feature and its corresponding attention score, we calculate the deep pathology feature. This deep pathology feature is then fed into a multi-task classification prediction model for classification prediction. This yields the following predictions for each pan-cancer biomarker: Based on a given slide, the deep features of the given task are calculated by combining the pathological low-dimensional features and their corresponding attention scores; The expression of the depth feature of the given task is: ; Where, Indicates the Deep feature representation of all blocks in a given slide for each task; Based on the deep features of a given task, it is input into a multi-task classification prediction model for classification prediction, and a slice-level probability prediction score of each task is calculated to obtain a prediction result for each biomarker in pan-cancer; the multi-task classification prediction model includes multiple classifiers; The expression of the slice-level probability prediction score of each task is: ; Where, Indicates the Slice-level probability prediction score for each task; Normalization function; Indicates the The weight matrix of each task in the classification layer; Indicates the The bias vector of each task in the classification layer.
6. The method for predicting pan-cancer biomarkers using multi-instance multi-task learning according to claim 5, characterized in that: The method further includes calculating a loss function based on the multi-task classification prediction model; the calculation of the loss function includes: Based on each slide, the loss function of a single task is calculated using the standard cross entropy loss function; Based on the loss function of a single task, the total loss function of all tasks is calculated by weighted sum; The expression of the total loss function is: ; Where, represents the total loss of all tasks; Indicates the The weight of the task, , Indicates the total number of tasks, ; Indicates the The loss function for each task.
7. A system for the multi-instance multi-task learning pan-cancer biomarker prediction method according to claim 1, characterized in that: include: A data acquisition module for acquiring full-field digital pathology images and pan-cancer biomarker expression data for patients with solid tumors; A data processing module is used to perform preprocessing and feature extraction based on the full-field digital pathology image to obtain a pathology feature data set; The calculation module is used to input the pathological feature dataset into the twin multi-instance fully convolutional network model for training to obtain low-dimensional pathological features. In combination with the linear attention mechanism, it calculates the attention score corresponding to each pathological feature. The prediction module is used to combine each pathology low-dimensional feature and the corresponding attention score to calculate the pathology deep feature, and input the pathology deep feature into the multi-task classification prediction model for classification prediction to obtain the prediction results of each biomarker in pan-cancer.
8. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Cancer pathology image classification system based on multi-instance learning, medium and equipment
CN117392468A
MHC-I type molecular neoantigen recognition method based on multi-instance learning
CN118553308A