Cancer treatment efficacy prediction method, device, equipment and storage medium
By combining multimodal information from radiomics, pathomics, transcriptomics, and clinical features to predict cancer treatment efficacy, this approach addresses the low prediction accuracy issue caused by single-modal features in existing technologies, achieving higher prediction accuracy and survival probability analysis.
Patent Information
- Application Number
- CN202510043274.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-01-10
AI Technical Summary
Existing machine learning-based radiomics methods for predicting cancer treatment efficacy rely on single-modality features, resulting in low prediction accuracy.
This study combines radiomics, pathomics, transcriptomics, and clinical features to predict the efficacy of cancer treatment. By acquiring CT scan images, pathological images, transcriptomics data, and clinical data, radiomics, pathomics, and transcriptomics features are extracted and evaluated using a pre-trained cancer treatment efficacy prediction model that integrates multiple features.
It improves the accuracy of predicting cancer treatment efficacy, enabling more precise predictions of cancer treatment outcomes and patient survival probabilities.
Smart Images

Figure CN119888351B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine learning technology, and in particular to a method, apparatus, device, and storage medium for predicting the efficacy of cancer treatment. Background Technology
[0002] Cancer, also known as malignant tumor, is caused by the malignant proliferation of cells and is invasive and metastatic.
[0003] In related technologies, cancer treatment methods mainly involve radiotherapy, chemotherapy, and surgery. Machine learning-based radiomics has shown great promise in identifying cancer types, predicting treatment outcomes, and forecasting disease progression. However, current machine learning-based radiomics methods for predicting cancer treatment efficacy rely on single-modality features, resulting in low prediction accuracy. Summary of the Invention
[0004] The purpose of this application is to provide a method, device, equipment, and storage medium for predicting the efficacy of cancer treatment. By combining radiomics features, pathomics features, transcriptomics features, and clinical features to predict the efficacy of cancer treatment, the accuracy of prediction can be improved.
[0005] This application provides a method for predicting the efficacy of cancer treatment, including:
[0006] Acquire CT scan images, pathological images, transcriptomics data, and clinical data;
[0007] Radiomics features are extracted from the CT scan images to obtain radiomics features;
[0008] The pathological images are subjected to attention-based extraction of pathological features of interest and local pathological feature extraction based on the results of the extraction of pathological features of interest, to obtain pathomic features;
[0009] Differential gene expression analysis and screening of differentially expressed genes based on the transcriptomics data were performed to obtain transcriptomics characteristics;
[0010] The fused features are obtained by integrating the radiomics features, the pathomics features, the transcriptomics features, and the clinical data;
[0011] The fusion features are input into a pre-trained cancer treatment efficacy prediction model to evaluate cancer treatment efficacy based on the fusion features, thereby obtaining cancer treatment efficacy prediction results.
[0012] In some embodiments, the step of extracting radiomics features from the CT scan image to obtain radiomics features includes:
[0013] The CT scan image is segmented into regions of interest to obtain several CT scan image blocks;
[0014] The CT scan image block is subjected to radiomics feature extraction to obtain the radiomics features; the radiomics feature extraction includes extracting first-order statistical features, region shape and size features and texture features from the CT scan image block.
[0015] In some embodiments, the extraction of pathological features of interest based on an attention mechanism and the extraction of local pathological features based on the results of the extraction of pathological features of interest to obtain pathomic features include:
[0016] The pathological image is segmented to obtain multiple pathological image blocks;
[0017] Based on the attention mechanism, the pathological features of interest in the pathological image blocks are perceived, and the pathological image blocks with the highest attention scores are selected to obtain a number of candidate image blocks; the pathological features of interest are pathological features that characterize the absence of tumor residue after cancer treatment.
[0018] Local pathological features are extracted from the candidate image blocks to obtain the pathomic features.
[0019] In some embodiments, the step of perceiving pathological features of interest in the pathological image blocks based on an attention mechanism to select a plurality of pathological image blocks with the highest attention scores, thereby obtaining a plurality of candidate image blocks, includes:
[0020] The pathological image blocks are encoded to obtain an encoded output vector;
[0021] The encoded output vector is subjected to fully connected dimensionality reduction processing to obtain a fully connected output vector;
[0022] Based on the importance of the pathological features of interest in the pathological image blocks, attention weights are calculated on the fully connected output vector to obtain several attention scores. The pathological image blocks with the highest attention scores are then selected to obtain several candidate image blocks.
[0023] In some embodiments, the step of performing differentially expressed gene analysis on the transcriptomics data and screening for differentially expressed genes based on the results of the differentially expressed gene analysis to obtain transcriptomics features includes:
[0024] The transcriptomics data and the reference genome data are compared to obtain the gene expression matrix corresponding to the transcriptomics data;
[0025] Differential expression analysis of the gene expression matrix was performed to target the efficacy of cancer treatment at the gene level, and several differentially expressed genes were screened based on the results of the differential expression analysis.
[0026] Linear regression analysis was performed on the gene expression values corresponding to the differentially expressed genes to screen out the differentially expressed genes that met the conditions for linear regression analysis, thereby obtaining the transcriptomics characteristics.
[0027] In some embodiments, the step of inputting the fused features into a pre-trained cancer treatment efficacy prediction model to evaluate cancer treatment efficacy based on the fused features and obtain cancer treatment efficacy prediction results includes:
[0028] Using a preset kernel function, the fusion features are mapped to a high-dimensional space and linearly transformed in the cancer treatment efficacy prediction model to obtain the high-dimensional mapping vector of the fusion features;
[0029] The loss value for each treatment efficacy category is calculated based on a pre-defined classification decision function;
[0030] Based on the loss value of the treatment efficacy category, the high-dimensional mapping vector is classified and predicted to obtain the cancer treatment efficacy prediction result.
[0031] In some embodiments, the cancer treatment efficacy prediction method further includes:
[0032] Based on the predicted results of cancer treatment efficacy, the patient's survival probability is calculated and a significance analysis is performed on the survival probability to obtain the survival analysis results.
[0033] This application also provides a cancer treatment efficacy prediction device, including:
[0034] The first module is used to acquire CT scan images, pathological images, transcriptomics data, and clinical data.
[0035] The second module is used to extract radiomics features from the CT scan images to obtain radiomics features;
[0036] The third module is used to extract pathological features of interest based on the attention mechanism and extract local pathological features based on the results of the extraction of pathological features of interest from the pathological images to obtain pathomic features.
[0037] The fourth module is used to perform differential gene expression analysis on the transcriptomics data and to screen differentially expressed genes based on the results of the differential gene expression analysis to obtain transcriptomics features;
[0038] The fifth module is used to fuse the radiomics features, the pathomics features, the transcriptomics features, and the clinical data to obtain fused features;
[0039] The sixth module is used to input the fusion features into a pre-trained cancer treatment efficacy prediction model, so as to evaluate the cancer treatment efficacy based on the fusion features and obtain the cancer treatment efficacy prediction result.
[0040] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method for predicting the efficacy of cancer treatment.
[0041] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for predicting the efficacy of cancer treatment.
[0042] The beneficial effects of this application are as follows: Patient CT scan images, pathological images, transcriptomics data, and clinical data are acquired. Corresponding feature extraction is used to obtain radiomics features, pathomics features, transcriptomics features, and clinical data. The fused features obtained by integrating these features are input into a cancer treatment efficacy prediction model. Cancer treatment efficacy is evaluated based on these fused features, resulting in a cancer treatment efficacy prediction result. Because it combines radiomics features, pathomics features, transcriptomics features, and clinical features, and predicts cancer treatment efficacy based on multiple modalities, the accuracy of cancer treatment efficacy prediction can be improved. Attached Figure Description
[0043] Figure 1 This is a flowchart of a cancer treatment efficacy prediction method provided in the embodiments of this application.
[0044] Figure 2 This is a flowchart of the specific method of step S103 provided in the embodiments of this application.
[0045] Figure 3 This is a flowchart of the specific method for step S202 provided in the embodiments of this application.
[0046] Figure 4 This is a flowchart of the specific method for step S104 provided in the embodiments of this application.
[0047] Figure 5 This is a flowchart of the specific method of step S106 provided in the embodiments of this application.
[0048] Figure 6 This is a schematic diagram of the structure of the cancer treatment efficacy prediction device provided in the embodiments of this application.
[0049] Figure 7This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0051] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and drawings are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0053] See Figure 1 , Figure 1 This is a flowchart of a cancer treatment efficacy prediction method provided in an embodiment of this application. In some embodiments, Figure 1 The methods include, but are not limited to, steps S101 to S106.
[0054] It should be noted that the execution subject of the cancer treatment efficacy prediction method provided in the embodiments of this specification can be applied to edge devices or cloud devices, and this embodiment does not limit it.
[0055] Step S101: Acquire CT scan images, pathological images, transcriptomics data, and clinical data.
[0056] The implementing entity can obtain CT scan images, pathological images, transcriptomics data, and clinical data. These images, along with the clinical data, are all obtained from corresponding tests performed on patients with the same type of cancer.
[0057] In some embodiments, the cancer may be esophageal squamous cell carcinoma.
[0058] CT scan images are generated by using X-rays to perform a tomographic scan of the human body and then processing the scan data using a computer. CT scan images are tomographic images that can display the tissue density distribution in a specific section of the body. They have high density resolution and can clearly show soft tissue and bone structures.
[0059] Pathological images are images created by scanning and fusing pathological tissues after processes such as dehydration, embedding, and sectioning using a microscopic scanning device. The creation of a pathological image begins with scanning the microscopic tissue images of the pathological slides using a scanning device. Then, the captured images are fused together. During the imaging process, adjacent images may overlap; through positioning and image registration, the fusion area is determined to obtain the pathological image.
[0060] Transcriptomics data refers to gene expression data obtained through transcriptome sequencing technology. Transcriptomics is the study of gene transcription and its regulatory mechanisms in cells, primarily examining gene expression at the RNA level. The transcriptome is the sum of all RNA that a living cell can transcribe, including mRNA, tRNA, rRNA, and various non-coding RNAs.
[0061] Clinical data refers to data obtained from medical institutions that is directly related to physicians' work with patients and addressing diseases. This data, concerning safety and performance during the clinical use of medical devices, may originate from pre- and post-marketing clinical trial information, relevant research data, and clinical experience data. In some embodiments, clinical data includes patient information such as age, sex, weight, lesion length, number of lymph nodes on CT scan, clinical stage, and pathological grade.
[0062] Step S102: Extract radiomics features from CT scan images to obtain radiomics features.
[0063] The executing entity can extract radiomics features from CT scan images to obtain radiomics features.
[0064] In practice, the implementing entity performs radiomics feature extraction on CT scan images. Radiomics feature extraction refers to extracting representative features from CT scan images. These features can describe the morphology, structure, and texture of regions of interest in CT scan images that are related to the disease state.
[0065] In some embodiments, step S102 specifically includes: segmenting the CT scan image into regions of interest to obtain several CT scan image blocks; and extracting radiomics features from the CT scan image blocks to obtain radiomics features. The radiomics feature extraction includes extracting first-order statistical features, region shape and size features, and texture features from the CT scan image blocks.
[0066] In practice, region of interest (ROI) segmentation of CT scan images can be performed manually using 3D-Slicer software. Oncologists with more than 5 years of experience delineate the ROI for each layer of the CT scan image. The ROI contains the tumor. Based on the segmentation results, the CT scan image is further segmented to obtain several CT scan image blocks containing the ROI. Radiomics features are then extracted from the CT scan image blocks, including first-order statistical features, region shape and size features, and texture features. Texture features can include 100 radiomics features such as gray-level co-occurrence matrix (GLCM), gray-level dependency matrix (GLDM), gray-level free length matrix (GLRLM), and gray-level size region matrix (GLSZM). Finally, the extracted features are fitted to obtain the radiomics features.
[0067] Step S103: Extract pathological features of interest based on attention mechanism and extract local pathological features based on the results of the extraction of pathological features of interest from the pathological image to obtain pathomic features.
[0068] The executing entity can extract pathological features of interest based on the attention mechanism and extract local pathological features based on the results of the extraction of pathological features of interest from pathological images to obtain pathomic features.
[0069] In practice, the implementing entity first extracts pathological features of interest from the pathological images based on an attention mechanism. The pathological features of interest are those that characterize the absence of tumor residue after cancer treatment. Based on the attention mechanism, several local image regions containing the pathological features of interest with the highest attention scores are selected as the results of the extraction of pathological features of interest. Then, local pathological features are extracted from the local image regions to extract multiple different types of local pathological features. Finally, the extracted local pathological features are fitted to obtain pathomic features.
[0070] Figure 2 This is a flowchart illustrating the specific method of step S103 provided in the embodiments of this application. (See attached document.) Figure 2 In some embodiments, the method includes, but is not limited to, steps S201 to S203.
[0071] Step S201: Segment the pathological image to obtain multiple pathological image blocks.
[0072] In some embodiments, pathological image segmentation can be based on an adaptive threshold selection algorithm. This algorithm maximizes the inter-class variance between the foreground and background of the pathological image to find the optimal threshold, thereby segmenting the pathological image into image blocks of a preset size, resulting in multiple pathological image blocks. For example, the Ostu threshold can be used to automatically segment the tissue region of each pathological image, read it into memory at a 20x magnification, and divide the pathological image into multiple 256×256 pixel pathological image blocks.
[0073] Step S202: Based on the attention mechanism, the pathological features of interest in the pathological image blocks are perceived, and the pathological image blocks with the highest attention scores are selected to obtain a number of candidate image blocks.
[0074] Among them, the pathological features of interest are those that characterize the absence of tumor residue after cancer treatment.
[0075] In some embodiments, pathological image blocks can be input into a pre-trained pathological feature extraction model, which perceives pathological features of interest in the pathological image blocks based on an attention mechanism, evaluates the importance of the perceived pathological features of interest and assigns attention scores, and outputs several pathological image blocks with the highest attention scores to obtain several candidate image blocks.
[0076] Figure 3 This is a flowchart illustrating the specific method of step S202 provided in the embodiments of this application. See also... Figure 3 In some embodiments, the pathological feature extraction model of interest includes an encoder layer, a fully connected layer, and an attention perception layer. The method includes, but is not limited to, steps S301 to S303.
[0077] Step S301: Encode the pathological image block to obtain the encoded output vector.
[0078] Step S302: Perform fully connected dimensionality reduction processing on the encoded output vector to obtain a fully connected output vector.
[0079] Step S303: Based on the importance of the pathological features of interest in the pathological image blocks, perform attention weighting on the fully connected output vector to obtain several attention scores, select the pathological image blocks with the highest attention scores, and obtain several candidate image blocks.
[0080] In the encoder layer, pathological image patches are encoded to obtain encoded output vectors. Specifically, a pre-trained ResNet50 is used as the main component to construct the encoder layer. The encoder layer is trained through cluster-guided contrastive learning. Pathological image patches are input into the encoder layer for encoding, resulting in encoded output vectors.
[0081] In the fully connected layer, the encoded output vector is subjected to fully connected dimensionality reduction processing to obtain a fully connected output vector. In specific implementation, a convolutional network with 512 neurons is used as the main body to form the fully connected layer. The encoded output vector is input into the fully connected layer to perform fully connected dimensionality reduction processing on the encoded output vector to obtain a fully connected output vector.
[0082] In the attention perception layer, attention weights are calculated on the fully connected output vector based on the importance of the pathological features of interest in the pathological image patches, resulting in several attention scores. The pathological image patches with the highest attention scores are then selected as candidate image patches. Specifically, the attention perception layer is constructed by stacking two attention networks, each consisting of three stacked fully connected networks. These two attention networks serve as part of a shared attention backbone for two categories. The first two attention networks are then divided into two parallel attention branches, and these parallel independent classifiers are used to score the slide-level representation specific to each category, yielding corresponding attention scores. After obtaining the attention scores, the pathological image patches are sorted according to these scores, and the pathological image patches with the highest attention scores are selected as candidate image patches.
[0083] The formula for calculating the attention weight score is:
[0084] ,
[0085] in, Attention weight score , and These are attention weights, and These are all elements in the fully connected output vector, where N is the number of pathological image patches. For Hadamard product, it means element-wise multiplication; tanh() is the hyperbolic tangent function; and sigm() is the sigmoid activation operation.
[0086] Step S203: Extract local pathological features from candidate image blocks to obtain pathological omics features.
[0087] In some embodiments, local pathological feature extraction of candidate image blocks can be performed using a pre-trained Mask R-CNN model. For each candidate image block, local pathological features such as area, convex region, eccentricity, range, filling area, major axis length, minor axis length, direction, perimeter, rigidity, proportion of tumor cells, proportion of lymphocytes, proportion of stromal cells, proportion of macrophages, proportion of nuclear rupture, and proportion of red blood cells are extracted. Finally, the average value of each local pathological feature of each candidate image block is calculated, and the local pathological features obtained by the average value calculation are fitted to obtain pathomic features.
[0088] Step S104 involves performing differential gene expression analysis on the transcriptomics data and screening for differentially expressed genes based on the results of the differential gene expression analysis to obtain transcriptomics characteristics.
[0089] The implementing entity can perform differentially expressed gene analysis on transcriptomics data and screen for differentially expressed genes based on the results of the differentially expressed gene analysis to obtain transcriptomics characteristics.
[0090] In practice, the implementing entity first preprocesses the transcriptomics data to remove low-quality gene expression values. Then, it performs differential gene expression analysis on the preprocessed transcriptomics data to select differentially expressed genes that are relevant to the efficacy of cancer treatment at the gene level. Next, it uses linear regression to learn the gene expression values corresponding to the differentially expressed genes and further screens the selected differentially expressed genes to identify the corresponding differentially expressed genes and obtain transcriptomics features.
[0091] Figure 4 This is a flowchart illustrating the specific method of step S104 provided in the embodiments of this application. See also... Figure 4 In some embodiments, the method includes, but is not limited to, steps S401 to S403.
[0092] Step S401: Compare the transcriptomics data with the reference genome data to obtain the gene expression matrix corresponding to the transcriptomics data.
[0093] In some embodiments, appropriate methods are selected based on the characteristics of the biological samples to perform sample quality testing. For samples that pass the testing, mRNA enrichment, reverse transcription, library construction, and sequencing are performed to obtain raw sequencing data. The raw sequencing data undergoes FastQC quality control to check the sequencing data quality, and then adapters, low-quality sequences, and contaminants are removed to obtain high-quality transcriptomics data.
[0094] In some embodiments, comparing transcriptomics data with reference genome data can be done by using the STAR tool to align transcriptomics data with a reference genome to obtain a gene expression matrix of the transcriptomics data.
[0095] Step S402: Perform differential expression analysis on the gene expression matrix to target the efficacy of cancer treatment at the gene level, and screen out several differentially expressed genes based on the results of the differential expression analysis.
[0096] In some embodiments, differential expression analysis of the gene expression matrix is performed to assess the efficacy of cancer treatment at the gene level. Based on the results of the differential expression analysis, several differentially expressed genes are screened out. This can be achieved by comparing the gene expression matrix with a reference gene expression matrix assigned to the PCR group (no tumor residue after cancer treatment) and a reference gene expression matrix assigned to the non-PCR group (tumor residue after cancer treatment), respectively. By setting appropriate threshold conditions, differentially expressed genes in the gene expression matrix are screened out. For example, by setting the threshold conditions of P < 0.05 and |log(2)FC| ≥ 1 (P is the differential expression threshold, and FC is the fold change), 791 differentially expressed genes are screened out from the gene expression matrix, of which 501 are upregulated and 290 are downregulated.
[0097] Step S403: Perform linear regression analysis on the gene expression values corresponding to differentially expressed genes to screen out differentially expressed genes that meet the conditions for linear regression analysis and obtain transcriptomics characteristics.
[0098] In some embodiments, linear regression analysis of gene expression values corresponding to differentially expressed genes can be performed using LASSO (Least Absolute Shrinkage and Selection Operator) regression to screen for more refined differentially expressed genes, thereby identifying differentially expressed genes that meet the conditions for linear regression analysis and obtaining transcriptomic features. LASSO regression is a linear regression method that employs L1 regularization. L1 regularization causes the feature weights of some learned, unimportant differentially expressed genes to be zero, thus selecting differentially expressed genes whose feature weights are not zero after LASSO regression, achieving the purpose of sparsification and feature selection. The expression for LASSO regression is as follows:
[0099] ,
[0100] in, Here, n represents the objective function value of the LASSO regression, and n is the number of differentially expressed genes. This represents the true value of the i-th differentially expressed gene. Let i be the feature value of the i-th differentially expressed gene. Let be the regression coefficient of the j-th differentially expressed gene. denoted as regularization strength, and p represents the number of features of differentially expressed genes.
[0101] Step S105: Integrate radiomics features, pathomics features, transcriptomics features, and clinical data to obtain fused features.
[0102] The implementing entity can integrate radiomics features, pathomics features, transcriptomics features, and clinical data to obtain fused features.
[0103] In practice, the implementing entity vectorizes radiomics features, pathomics features, transcriptomics features, and clinical data before concatenating them to obtain fused features in the form of an input vector that conforms to the cancer treatment efficacy prediction model. The expression for obtaining the fused features is as follows:
[0104] ,
[0105] in, For the vector representation of fused features, This is a vector representation of radiomics features. This is a vector representation of pathomic features. Vector representation of transcriptomic features. A vector representation of clinical data. This is for splicing operations.
[0106] Step S106: Input the fused features into the pre-trained cancer treatment efficacy prediction model to evaluate the cancer treatment efficacy based on the fused features and obtain the cancer treatment efficacy prediction results.
[0107] The implementing entity can input the fusion features into a pre-trained cancer treatment efficacy prediction model to evaluate the efficacy of cancer treatment based on the fusion features and obtain the cancer treatment efficacy prediction results.
[0108] In practice, the implementing entity inputs the fused features into the cancer treatment efficacy prediction model for high-dimensional mapping to obtain the high-dimensional mapping vector of the fused features. Then, based on the preset classification decision function, the high-dimensional mapping vector of the fused features is classified and predicted to obtain the cancer treatment efficacy prediction result.
[0109] In some embodiments, the cancer treatment efficacy prediction model is a deep neural network model based on support vector machines (SVMs), trained using CT scan sample images, pathological sample images, transcriptomics sample data, and clinical sample data. Support vector machines (SVMs) are supervised learning models and related learning algorithms used in classification and regression analysis, and each SVM contains support vector machine layers. The cancer treatment efficacy prediction model performs high-dimensional mapping by inputting fused features into the SVM layers, mapping the fused features to a Hilbert space to obtain a high-dimensional mapped vector of the fused features. Here, the Hilbert space is an abstract space composed of several independent coordinates, referring to a complete inner product space. Mapping the fused features to the Hilbert space and performing inner product calculations yields the high-dimensional mapped vector of the fused features.
[0110] Figure 5 This is a flowchart illustrating the specific method of step S106 provided in the embodiments of this application. See also... Figure 5 In some embodiments, the method includes, but is not limited to, steps S501 to S503.
[0111] Step S501: Using a preset kernel function, the fused features are mapped to a high-dimensional space and linearly transformed in the cancer treatment efficacy prediction model to obtain a high-dimensional mapping vector of the fused features.
[0112] In some embodiments, a preset kernel function is used to determine the projection relationship between the fused features and the support vectors in the high-dimensional space by calculating the similarity between the fused features and the preset support vectors. Combining the similarity measure and the position of the support vectors in the high-dimensional space, the fused features are mapped to the high-dimensional space to obtain the high-dimensional mapping vector of the fused features.
[0113] In some embodiments, a kernel function can be preset. Specifically, the preset kernel function may include, but is not limited to, one or more combinations of the following: polynomial kernel function, radial basis kernel function, Laplace kernel function, Sigmoid kernel function, etc.
[0114] Preferably, a radial basis function is selected as the preset kernel function.
[0115] Step S502: Calculate the loss value for each treatment efficacy category based on the preset classification decision function.
[0116] The treatment efficacy is categorized into two types: PCR and non-PCR. The PCR category indicates no tumor residue after cancer treatment, while the non-PCR category indicates tumor residue after cancer treatment.
[0117] Step S503: Based on the loss value of the treatment efficacy category, classify and predict the high-dimensional mapping vector to obtain the cancer treatment efficacy prediction result.
[0118] In practice, when the high-dimensional mapping vector is classified and predicted as a PCR category based on the loss value of the treatment efficacy category, a cancer treatment efficacy prediction result representing the current patient's condition has been alleviated is obtained. When the high-dimensional mapping vector is classified and predicted as a non-PCR category based on the loss value of the treatment efficacy category, a cancer treatment efficacy prediction result representing the current patient's condition has not been alleviated is obtained.
[0119] In some embodiments, the above method further includes: calculating the patient's survival probability based on the cancer treatment efficacy prediction results and performing a significance analysis on the survival probability to obtain survival analysis results.
[0120] In practice, after obtaining the predicted results of cancer treatment efficacy, the survival probability of patients is calculated using the formula for survival probability, and Kaplan-Meier survival curves are plotted. The log-rank test is used to measure whether the difference in survival distribution between the PCR and non-PCR groups is statistically significant (P-value < 0.05). Finally, based on the shape of the KM curves and the inter-group comparison results, the survival status of the pCR and non-pCR groups is interpreted, providing clinical decision support.
[0121] The formula for calculating the survival probability is:
[0122] ,
[0123] in, Let be the survival probability at time point k. This represents the survival probability at time point k-1. Let the historical death toll at time point k be . Let be the sum of the historical number of deaths and the historical number of survivors at the k-th time point.
[0124] See Figure 6 This application also provides a cancer treatment efficacy prediction device, which can implement the above-mentioned cancer treatment efficacy prediction method. The device includes:
[0125] The first module 601 is used to acquire CT scan images, pathological images, transcriptomics data and clinical data;
[0126] The second module 602 is used to extract radiomics features from CT scan images to obtain radiomics features.
[0127] The third module 603 is used to extract pathological features of interest based on the attention mechanism and extract local pathological features based on the results of the extraction of pathological features of interest from pathological images to obtain pathomic features.
[0128] Module 4, 604, is used for differential gene expression analysis of transcriptomics data and screening of differentially expressed genes based on the results of differential gene expression analysis to obtain transcriptomics features.
[0129] Module 5, 605, is used to fuse radiomics features, pathomics features, transcriptomics features, and clinical data to obtain fused features;
[0130] The sixth module 606 is used to input the fusion features into the pre-trained cancer treatment efficacy prediction model, so as to evaluate the cancer treatment efficacy based on the fusion features and obtain the cancer treatment efficacy prediction results.
[0131] The specific implementation of this cancer treatment efficacy prediction device is basically the same as the specific implementation of the cancer treatment efficacy prediction method described above, and will not be repeated here.
[0132] Figure 7 This is a block diagram illustrating an electronic device according to an exemplary embodiment.
[0133] The following reference Figure 7 To describe an electronic device 700 according to such an embodiment of the present disclosure. Figure 7 The electronic device 700 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.
[0134] like Figure 7 As shown, the electronic device 700 is presented in the form of a general-purpose computing device. The components of the electronic device 700 may include, but are not limited to: at least one processing unit 710, at least one storage unit 720, a bus 730 connecting different system components (including storage unit 720 and processing unit 710), a display unit 740, etc.
[0135] The storage unit stores program code, which can be executed by the processing unit 710, causing the processing unit 710 to perform the steps described in the above-described section of the cancer treatment efficacy prediction method according to various exemplary embodiments of this disclosure.
[0136] Storage unit 720 may include a readable medium in the form of a volatile storage unit, such as random access memory (RAM) 7201 and / or cache memory 7202, and may further include a read-only memory (ROM) 7203.
[0137] The storage unit 720 may also include a program / utility 7204 having a set (at least one) program module 7205, such program module 7205 including but not limited to: an operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.
[0138] Bus 730 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0139] Electronic device 700 can also communicate with one or more external devices 700' (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 700, and / or with any device that enables electronic device 700 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 750. Furthermore, electronic device 700 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 760. Network adapter 760 can communicate with other modules of electronic device 700 via bus 730. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0140] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for predicting the efficacy of cancer treatment.
[0141] The cancer treatment efficacy prediction method, apparatus, device, and storage medium provided in this application acquire patient CT scan images, pathological images, transcriptomics data, and clinical data. Through corresponding feature extraction, radiomics features, pathomics features, transcriptomics features, and clinical data are obtained. The fused features obtained by integrating these features are input into a cancer treatment efficacy prediction model to evaluate cancer treatment efficacy based on the fused features, thus obtaining a cancer treatment efficacy prediction result. Because it combines radiomics features, pathomics features, transcriptomics features, and clinical features, and predicts cancer treatment efficacy based on multiple modalities, the accuracy of cancer treatment efficacy prediction can be improved.
[0142] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, or network device, etc.) to execute the methods described above according to the embodiments of this disclosure.
[0143] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0144] Computer-readable storage media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0145] Those skilled in the art will understand that the above modules can be distributed in the device as described in the embodiments, or they can be modified accordingly and placed in one or more devices that are unique to this embodiment. The modules in the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.
[0146] Exemplary embodiments of this disclosure have been specifically shown and described above. It should be understood that this disclosure is not limited to the detailed structures, arrangements, or implementations described herein; rather, this disclosure is intended to cover various modifications and equivalent arrangements contained within the spirit and scope of the appended claims.
Claims
1. A method for predicting the efficacy of cancer treatment, characterized in that, include: Acquire CT scan images, pathological images, transcriptomics data, and clinical data; Radiomics features are extracted from the CT scan images to obtain radiomics features; The pathological images are subjected to attention-based extraction of pathological features of interest and local pathological feature extraction based on the results of the extraction of pathological features of interest, to obtain pathomic features; Differential gene expression analysis and screening of differentially expressed genes based on the transcriptomics data were performed to obtain transcriptomics characteristics; The fused features are obtained by integrating the radiomics features, the pathomics features, the transcriptomics features, and the clinical data; The fusion features are input into a pre-trained cancer treatment efficacy prediction model to evaluate cancer treatment efficacy based on the fusion features and obtain cancer treatment efficacy prediction results. The differentially expressed gene analysis of the transcriptomics data and the screening of differentially expressed genes based on the results of the differentially expressed gene analysis to obtain transcriptomics features include: The transcriptomics data and the reference genome data are compared to obtain the gene expression matrix corresponding to the transcriptomics data; Differential expression analysis of the gene expression matrix was performed to target the efficacy of cancer treatment at the gene level, and several differentially expressed genes were screened based on the results of the differential expression analysis. Linear regression analysis was performed on the gene expression values corresponding to the differentially expressed genes to screen out the differentially expressed genes that met the conditions for linear regression analysis, thereby obtaining the transcriptomics characteristics.
2. The method for predicting the efficacy of cancer treatment according to claim 1, characterized in that, The step of extracting radiomics features from the CT scan images to obtain radiomics features includes: The CT scan image is segmented into regions of interest to obtain several CT scan image blocks; The CT scan image block is subjected to radiomics feature extraction to obtain the radiomics features; the radiomics feature extraction includes extracting first-order statistical features, region shape and size features and texture features from the CT scan image block.
3. The method for predicting the efficacy of cancer treatment according to claim 1, characterized in that, The process of extracting pathological features of interest from the pathological image based on an attention mechanism and extracting local pathological features based on the results of the extraction of pathological features of interest, to obtain pathomic features, includes: The pathological image is segmented to obtain multiple pathological image blocks; Based on the attention mechanism, the pathological features of interest in the pathological image blocks are perceived, and the pathological image blocks with the highest attention scores are selected to obtain a number of candidate image blocks; the pathological features of interest are pathological features that characterize the absence of tumor residue after cancer treatment. Local pathological features are extracted from the candidate image blocks to obtain the pathomic features.
4. The method for predicting the efficacy of cancer treatment according to claim 3, characterized in that, The attention-based mechanism perceives pathological features of interest within the pathological image blocks, selecting the pathological image blocks with the highest attention scores to obtain several candidate image blocks, including: The pathological image blocks are encoded to obtain an encoded output vector; The encoded output vector is subjected to fully connected dimensionality reduction processing to obtain a fully connected output vector; Based on the importance of the pathological features of interest in the pathological image blocks, attention weights are calculated on the fully connected output vector to obtain several attention scores. The pathological image blocks with the highest attention scores are then selected to obtain several candidate image blocks.
5. The method for predicting the efficacy of cancer treatment according to claim 1, characterized in that, The step of inputting the fused features into a pre-trained cancer treatment efficacy prediction model to evaluate cancer treatment efficacy based on the fused features and obtain cancer treatment efficacy prediction results includes: Using a preset kernel function, the fusion features are mapped to a high-dimensional space and linearly transformed in the cancer treatment efficacy prediction model to obtain the high-dimensional mapping vector of the fusion features; The loss value for each treatment efficacy category is calculated based on a pre-defined classification decision function; Based on the loss value of the treatment efficacy category, the high-dimensional mapping vector is classified and predicted to obtain the cancer treatment efficacy prediction result.
6. The method for predicting the efficacy of cancer treatment according to claim 1, characterized in that, Also includes: Based on the predicted results of cancer treatment efficacy, the patient's survival probability is calculated and a significance analysis is performed on the survival probability to obtain the survival analysis results.
7. A device for predicting the efficacy of cancer treatment, characterized in that, include: The first module is used to acquire CT scan images, pathological images, transcriptomics data, and clinical data. The second module is used to extract radiomics features from the CT scan images to obtain radiomics features; The third module is used to extract pathological features of interest based on the attention mechanism and extract local pathological features based on the results of the extraction of pathological features of interest from the pathological images to obtain pathomic features. The fourth module is used to perform differential gene expression analysis on the transcriptomics data and to screen differentially expressed genes based on the results of the differential gene expression analysis to obtain transcriptomics features; The fifth module is used to fuse the radiomics features, the pathomics features, the transcriptomics features, and the clinical data to obtain fused features; The sixth module is used to input the fusion features into a pre-trained cancer treatment efficacy prediction model, so as to evaluate the cancer treatment efficacy based on the fusion features and obtain the cancer treatment efficacy prediction result. The differentially expressed gene analysis of the transcriptomics data and the screening of differentially expressed genes based on the results of the differentially expressed gene analysis to obtain transcriptomics features include: The transcriptomics data and the reference genome data are compared to obtain the gene expression matrix corresponding to the transcriptomics data; Differential expression analysis of the gene expression matrix was performed to target the efficacy of cancer treatment at the gene level, and several differentially expressed genes were screened based on the results of the differential expression analysis. Linear regression analysis was performed on the gene expression values corresponding to the differentially expressed genes to screen out the differentially expressed genes that met the conditions for linear regression analysis, thereby obtaining the transcriptomics characteristics.
8. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the cancer treatment efficacy prediction method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the cancer treatment efficacy prediction method according to any one of claims 1 to 6.
Citation Information
Patent Citations
System for predicting curative effect of PD-1 / PD-L1 monoclonal antibody treatment in advanced cancer patient
CN115440383A
System for predicting treatment resistance in rectal cancer and molecular mechanism thereof before treatment
WO2023193390A1