CNN-based multi-cancer drug combination effectiveness prediction method and system

Through the combination of KPCA and CNN, the complex interaction relationship between drugs in combination with multiple cancer drugs was captured, and the problem of insufficient accuracy in evaluating the effectiveness of multiple cancer drugs in the prior art was solved, and more efficient prediction accuracy and the formulation of personalized treatment plans were achieved.

CN120148901APending Publication Date: 2025-06-13FOSHAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510049106.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to accurately evaluate the effectiveness of multiple cancer combination drugs, especially when processing complex biomedical data, and it is difficult to capture nonlinear relationships and potential modes of action between drugs.

Method used

The effectiveness prediction method of multiple cancer combination drugs based on CNN is used to capture the complex interaction relationship between drugs through nonlinear dimensionality reduction of KPCA and deep feature learning of CNN.

Benefits of technology

It improves the accuracy of predicting the effectiveness of multiple cancer combination drugs, provides an important reference for the formulation of clinical personalized drug regimens, optimizes multi-drug treatment strategies for complex diseases, improves treatment effects and reduces the risk of adverse reactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148901A_ABST
    Figure CN120148901A_ABST
Patent Text Reader

Abstract

The invention discloses a CNN-based multi-cancer drug combination effectiveness prediction method, and the method comprises the steps: obtaining the feature data of a to-be-predicted drug combination, and carrying out the preprocessing of the feature data of the to-be-predicted drug combination, and obtaining initial data; performing dimension reduction processing on the initial data through a KPCA model to obtain dimension-reduced features; performing feature reconstruction on the dimension-reduced features to obtain reconstructed features; and inputting the reconstructed features into a CNN model to obtain a predicted value. According to the method, through nonlinear dimensionality reduction of KPCA and deep feature learning of CNN, a complex interaction relationship among drugs can be accurately captured, so that the accuracy of effectiveness prediction of multi-cancer drug combination is improved, an important reference is provided for formulation of a clinical personalized medication scheme, a multi-drug treatment strategy of complex diseases is effectively optimized, and a good application prospect is achieved. The treatment effect is improved, the adverse reaction risk is reduced, and remarkable technical advantages and practical value are shown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of predicting the effectiveness of drug use, and particularly to a method and system for predicting the effectiveness of combined drug use for multiple cancers based on CNN. Background Art

[0002] As a highly heterogeneous disease, cancer shows significant differences in the responses of different patients to drugs, and single-drug treatment often fails to achieve the best results. Therefore, combined drug use, as an effective treatment strategy, can enhance the anti-tumor effect through multiple mechanisms. In cancer treatment, combined drug use, as an effective strategy, can enhance the anti-tumor effect through multiple mechanisms. However, due to the complex interactions between drug combinations, evaluating their effectiveness and safety has become a major challenge. Traditional methods for evaluating combined drug use often rely on linear models or simple statistical methods, which have limitations in dealing with complex biomedical data and are difficult to accurately capture the non-linear relationships and potential action patterns between drugs. Therefore, how to efficiently and accurately evaluate the effectiveness of combined drug use for multiple cancers has become an urgent technical problem to be solved currently. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method and system for predicting the effectiveness of combined drug use for multiple cancers based on CNN. Through the non-linear dimensionality reduction of KPCA and the deep feature learning of CNN, the complex interaction relationships between drugs can be accurately captured, thereby improving the accuracy of predicting the effectiveness of combined drug use for multiple cancers, providing an important reference for the formulation of clinical personalized drug use plans, effectively optimizing the multi-drug treatment strategies for complex diseases, improving the treatment effect and reducing the risk of adverse reactions, and showing significant technical advantages and practical values.

[0004] To solve the above technical problem, in the first aspect of the present invention, a method for predicting the effectiveness of combined drug use for multiple cancers based on CNN is disclosed, and the method includes:

[0005] Obtain the feature data of the drug combination to be predicted and preprocess the feature data of the drug combination to be predicted to obtain initial data;

[0006] Perform dimensionality reduction processing on the initial data through a KPCA model to obtain the dimensionality-reduced features;

[0007] Perform feature reconstruction on the dimensionality-reduced features to obtain reconstructed features;

[0008] Input the reconstructed features into a CNN model to obtain a predicted value.

[0009] As an optional implementation manner, in the first aspect of the present invention, the feature data of the drug combination to be predicted includes drug feature information, drug concentration information, and cell line gene expression data.

[0010] As an alternative implementation, in the first aspect of the present invention, obtaining the characteristic data of the drug combination to be predicted and preprocessing the characteristic data of the drug combination to be predicted to obtain initial data includes:

[0011] Performing oversampling on the characteristic data of the drug combination to be predicted to obtain balanced data;

[0012] Performing normalization processing on the balanced data to obtain initial data.

[0013] As an alternative implementation, in the first aspect of the present invention, performing normalization processing on the balanced data to obtain initial data includes:

[0014] When the characteristic distribution characteristic of the balanced data is concentration, performing min-max normalization processing on the balanced data;

[0015] When the characteristic distribution characteristic of the balanced data is discreteness, performing logarithmic normalization processing on the balanced data.

[0016] As an alternative implementation, in the first aspect of the present invention, performing dimensionality reduction processing on the initial data through the KPCA model to obtain the dimensionality-reduced features includes:

[0017] The KPCA model uses a Gaussian kernel function to perform non-linear mapping on the initial data and performs PCA dimensionality reduction in the high-dimensional feature space to obtain the dimensionality-reduced features.

[0018] As an alternative implementation, in the first aspect of the present invention, performing feature reconstruction on the dimensionality-reduced features to obtain the reconstructed features includes:

[0019] Performing normalization processing and reconstruction on the dimensionality-reduced features to obtain a two-dimensional feature image with pixels of 120*120, and using the two-dimensional feature image as the reconstructed features.

[0020] As an alternative implementation, in the first aspect of the present invention, the CNN model includes a first convolutional block and a second convolutional block. The number of convolutional kernels in the convolutional layer of the first convolutional block is 8, and the number of convolutional kernels in the convolutional layer of the second convolutional block is 16.

[0021] The second aspect of the present invention discloses a multi-cancer combined drug use effectiveness prediction system based on CNN. The system includes:

[0022] A data acquisition module, which is used to acquire the characteristic data of the drug combination to be predicted and preprocess the characteristic data of the drug combination to be predicted to obtain initial data;

[0023] A feature dimensionality reduction module, which is used to perform dimensionality reduction processing on the initial data through a KPCA model to obtain the dimensionality-reduced features;

[0024] A feature reconstruction module, which is used to perform feature reconstruction on the dimensionality-reduced features to obtain the reconstructed features;

[0025] A prediction module, which is used to input the reconstructed features into a CNN model to obtain a predicted value.

[0026] As an optional implementation manner, in the second aspect of the present invention, the characteristic data of the drug combination to be predicted includes drug characteristic information, drug concentration information, and cell line gene expression data.

[0027] As an optional implementation manner, in the second aspect of the present invention, the data acquisition module acquires the characteristic data of the drug combination to be predicted and preprocesses the characteristic data of the drug combination to be predicted to obtain the initial data, including:

[0028] Performing oversampling on the characteristic data of the drug combination to be predicted to obtain balanced data;

[0029] Performing standardization processing on the balanced data to obtain the initial data.

[0030] As an optional implementation manner, in the second aspect of the present invention, the data acquisition module performs standardization processing on the balanced data to obtain the initial data, including:

[0031] When the characteristic distribution characteristic of the balanced data is concentration, performing min-max normalization processing on the balanced data;

[0032] When the characteristic distribution characteristic of the balanced data is discreteness, performing logarithmic normalization processing on the balanced data.

[0033] As an optional implementation manner, in the second aspect of the present invention, the feature dimensionality reduction module performs dimensionality reduction processing on the initial data through a KPCA model to obtain the dimensionality-reduced features, including:

[0034] The KPCA model uses a Gaussian kernel function to perform non-linear mapping on the initial data and performs PCA dimensionality reduction in the high-dimensional feature space to obtain the dimensionality-reduced features.

[0035] As an optional implementation manner, in the second aspect of the present invention, the feature reconstruction module performs feature reconstruction on the dimensionality-reduced features to obtain the reconstructed features, including:

[0036] Normalize and reconstruct the dimension-reduced features to obtain a two-dimensional feature image with 120 * 120 pixels, and use the two-dimensional feature image as the reconstructed feature.

[0037] As an alternative implementation, in the second aspect of the present invention, the CNN model includes a first convolutional block and a second convolutional block. The number of convolutional kernels in the convolutional layer of the first convolutional block is 8, and the number of convolutional kernels in the convolutional layer of the second convolutional block is 16.

[0038] The third aspect of the present invention discloses another CNN-based multi-cancer combination drug efficacy prediction device, which includes:

[0039] A memory storing executable program code;

[0040] A processor coupled to the memory;

[0041] The processor calls the executable program code stored in the memory and executes some or all of the steps in the CNN-based multi-cancer combination drug efficacy prediction method disclosed in the first aspect of the embodiments of the present invention.

[0042] The fourth aspect of the embodiments of the present invention discloses a computer storage medium storing computer instructions, which are used to execute some or all of the steps in the CNN-based multi-cancer combination drug efficacy prediction method disclosed in the first aspect of the embodiments of the present invention when the computer instructions are called.

[0043] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0044] Through the non-linear dimensionality reduction of KPCA and the deep feature learning of CNN, the complex interaction relationships between drugs can be accurately captured, thereby improving the accuracy of multi-cancer combination drug efficacy prediction, providing an important reference for the formulation of clinical personalized drug treatment plans, effectively optimizing the multi-drug treatment strategies for complex diseases, improving the treatment effect and reducing the risk of adverse reactions, showing significant technical advantages and practical value. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0046] Figure 1 It is a schematic flowchart of a CNN-based multi-cancer combination drug efficacy prediction method disclosed in the embodiments of the present invention;

[0047] Figure 2 It is a structural block diagram of a multi-cancer combined drug use effectiveness prediction system based on CNN disclosed in an embodiment of the present invention;

[0048] Figure 3 It is a schematic structural diagram of a multi-cancer combined drug use effectiveness prediction device based on CNN disclosed in an embodiment of the present invention. Specific embodiments

[0049] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0050] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, device, product or terminal including a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or terminals.

[0051] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0052] The present invention discloses a method and system for predicting the effectiveness of multi-cancer combined drug use based on CNN. Through the non-linear dimensionality reduction of KPCA and the deep feature learning of CNN, it is possible to accurately capture the complex interaction relationships between drugs, thereby improving the accuracy of predicting the effectiveness of multi-cancer combined drug use, providing an important reference for the formulation of clinical personalized drug use plans, effectively optimizing the multi-drug treatment strategies for complex diseases, improving the treatment effect and reducing the risk of adverse reactions, showing significant technical advantages and practical value. The following will be described in detail respectively.

[0053] Embodiment 1

[0054] Please refer to Figure 1 ,Figure 1 It is a schematic flowchart of a method for predicting the effectiveness of combined drug use for multiple cancers based on CNN disclosed in an embodiment of the present invention. Among them, Figure 1 The described method is applied to a device for predicting the effectiveness of combined drug use for multiple cancers based on CNN. This prediction device can be a corresponding prediction terminal, prediction device, or server, and this server can be a local server or a cloud server, which is not limited in the embodiments of the present invention. As Figure 1 shown, the method for predicting the effectiveness of combined drug use for multiple cancers based on CNN may include the following operations:

[0055] 101. Obtain the feature data of the drug combination to be predicted and preprocess the feature data of the drug combination to be predicted to obtain initial data.

[0056] In an embodiment of the present invention, the feature data of the drug combination to be predicted includes drug feature information, drug concentration information, and cell line gene expression data.

[0057] 102. Perform dimensionality reduction processing on the initial data through a KPCA model to obtain the dimensionality-reduced features.

[0058] 103. Perform feature reconstruction on the dimensionality-reduced features to obtain reconstructed features.

[0059] 104. Input the reconstructed features into a CNN model to obtain a predicted value.

[0060] In an optional embodiment, in step 101, the obtaining the feature data of the drug combination to be predicted and preprocessing the feature data of the drug combination to be predicted to obtain initial data includes:

[0061] Perform oversampling on the feature data of the drug combination to be predicted to obtain balanced data;

[0062] Perform normalization processing on the balanced data to obtain initial data.

[0063] In an embodiment of the present invention, first, the SMOTE algorithm is used to perform oversampling on the feature data of the drug combination to be predicted to solve the problem of class imbalance and generate a synthetic balanced data distribution. Then, normalization processing is performed on the balanced data.

[0064] In an optional embodiment, the performing normalization processing on the balanced data to obtain initial data in the above steps includes:

[0065] When the feature distribution characteristic of the balanced data is concentration, perform min-max normalization processing on the balanced data;

[0066] When the characteristic distribution characteristic of the balanced data is discreteness, perform logarithmic normalization processing on the balanced data.

[0067] In the embodiment of the present invention, when the characteristic distribution characteristic of the balanced data is centrality, use min-max normalization to process and map the balanced data to the interval [0,1]. The formula is: where X is the input feature, X min and X max are the minimum and maximum values of the feature respectively. When the characteristic distribution characteristic of the balanced data is discreteness, use logarithmic normalization to perform logarithmic transformation processing on the balanced data. The formula is: where X is the input feature, X min and X max are the minimum and maximum values of the feature respectively.

[0068] In an alternative embodiment, in step 102, the dimensionality reduction processing of the initial data by the KPCA model to obtain the dimensionality-reduced features includes:

[0069] The KPCA model uses a Gaussian kernel function to perform non-linear mapping on the initial data and performs PCA dimensionality reduction in the high-dimensional feature space to obtain the dimensionality-reduced features.

[0070] In an alternative embodiment, in step 103, the feature reconstruction of the dimensionality-reduced features to obtain the reconstructed features includes:

[0071] Perform normalization processing and reconstruction on the dimensionality-reduced features to obtain a two-dimensional feature image with pixels of 120*120, thereby establishing a mapping relationship from the original high-dimensional data to the low-dimensional image representation, and using the two-dimensional feature image as the reconstructed features.

[0072] In an alternative embodiment, the CNN model adopts a double convolutional block structure, including a first convolutional block and a second convolutional block. The number of convolutional kernels in the convolutional layer of the first convolutional block is 8, which is used to extract basic features. The number of convolutional kernels in the convolutional layer of the second convolutional block is 16, which is used to extract advanced features. Both the first convolutional block and the second convolutional block are equipped with a batch normalization layer, a ReLU activation function, and a max pooling layer with a stride of 2.

[0073] In this method, first, the non-linear dimensionality reduction technique of KPCA is used to map the data into a high-dimensional feature space through the kernel trick, and then principal component analysis is performed in this space. Compared with traditional linear PCA, KPCA has the ability to capture the non-linear features of data. In multi-omics data analysis, this characteristic of KPCA is particularly important because biological data often has complex non-linear relationships. By mapping high-dimensional gene expression data and copy number variation data into two-dimensional images, KPCA not only achieves data dimensionality reduction but also retains the key feature information in the data, laying a foundation for subsequent deep learning analysis. Then, CNN is used to perform deep feature learning on these dimension-reduced images. Through its powerful feature extraction ability, it can automatically identify and learn the key features and potential action patterns in drug combinations. Especially in the complex medical scenario of combination drug use, where there are multi-level and multi-dimensional interactions between drugs, the combination of KPCA and CNN can better characterize the synergistic or antagonistic effects between drugs, providing more reliable decision-making basis for clinicians, ultimately improving the treatment effect and reducing the risk of adverse reactions. This innovative method not only has sufficient interpretability in theory but also shows good effects in practical applications, providing a new research paradigm and solution ideas for complex medical prediction problems.

[0074] This combination drug use prediction method is based on multi-dimensional dimensionality reduction and deep learning technologies, showing significant technical advantages and practical value.

[0075] Through image processing combined with deep learning technology, a new drug data analysis framework is constructed. Using the deep feature extraction ability of CNN, it can effectively identify complex patterns and potential laws in drug interaction data, significantly improving the accuracy and efficiency of feature recognition. This method shows excellent data processing capabilities and model performance, can efficiently process large-scale drug combination data, and adapt to the diverse prediction task requirements. It not only provides a reliable analysis tool for drug screening and development but also has good scalability and adaptability, can flexibly integrate new data sources and prediction requirements, providing strong technical support for medical research and drug development, and has important theoretical significance and application value.

[0076] Example 2

[0077] This example is about the training methods of the KPCA model and the CNN model in Example 1. The training methods include the following steps:

[0078] Regarding the acquisition of training sample data, first, obtain the cancer combination drug synergy score data through the O'neil public database, and perform data screening according to the scores, selecting the data subsets with higher scores in each cancer type; integrate multi-modal data such as drug structure, administration concentration, and cell line characteristics to construct an initial feature set.

[0079] The training process of the KPCA model is as follows:

[0080] First, a five-fold cross-validation strategy is adopted to divide the dataset, and KPCA dimensionality reduction is performed on each fold of data respectively. The specific implementation process is as follows: To address the problem of uneven class distribution in the training sample data, the SMOTE algorithm is used to oversample the drug features, concentration, and gene expression data to generate synthetic samples and achieve data balance; then, the sample features are standardized, and different normalization methods are selected according to the feature distribution characteristics of the sample features. For features with concentrated distributions, Norm-1 normalization is used to map them to the [0, 1] interval, and for features with outliers, Norm-2 normalization is used for logarithmic transformation processing.

[0081] Among them, the Norm-1 normalization method is min-max normalization: where X is the input feature, X min and X max are the minimum and maximum values of the feature respectively; the Norm-2 normalization method is logarithmic normalization: where X is the input feature, X min and X max are the minimum and maximum values of the feature respectively.

[0082] Then, a KPCA model based on the Gaussian kernel function is constructed, and the balanced samples are mapped to a high-dimensional feature space. The main eigenvectors are obtained by eigenvalue decomposition of the kernel matrix to achieve the extraction and dimensionality reduction of non-linear features; finally, the dimensionality-reduced features are normalized and reconstructed into a two-dimensional feature image of 120×120 pixels to realize the conversion of data dimensionality reduction and feature representation, thereby establishing a mapping relationship from the original high-dimensional data to the low-dimensional image representation. The entire process continuously adjusts the feature extraction parameters through iterative optimization, and finally obtains a standardized feature image suitable for subsequent CNN model processing.

[0083] The training process of the CNN model is as follows:

[0084] Based on the feature image processed by KPCA, a CNN network model is constructed, and a double convolutional block structure is used for feature learning. The first convolutional block uses 8 3×3 convolutional kernels to extract basic features, and the second convolutional block uses 16 3×3 convolutional kernels to extract advanced features. Each convolutional block is equipped with a batch normalization layer, a ReLU activation function, and a max-pooling layer with a stride of 2.

[0085] The convolution operation can be expressed as: Conv(X) = W * X + b, where * represents the convolution operation, W is the convolutional kernel weight, and b is the bias term. The batch normalization operation can be expressed as: where μ B and are the mean and variance of the batch data, γ and β are learnable parameters, and ε is a small constant to prevent division by zero.

[0086] The CNN model uses the mean squared error loss function: where y i is the true drug response value of the i-th sample, is the predicted value, and n is the total number of samples. The model performance is evaluated by PCC: where and are the means of the true values and predicted values, respectively.

[0087] Different optimization strategies are adopted for different cancer types: for breast cancer and colon cancer data, training is performed using the Adam optimizer with a learning rate of 1e-4 and a batch size of 64; for lung cancer and melanoma data, training is performed using the Adam optimizer with a learning rate of 1e-3 and a batch size of 256. To comprehensively evaluate the model performance, the present invention conducts five independent training experiments on each cancer dataset, sets the maximum number of iterations to 400 epochs for each experiment, and calculates and records the MSE and PCC metric values after each epoch, and evaluates the model performance by continuously monitoring the changes in the evaluation metrics, and finally selects the model parameters with the best prediction effect.

[0088] Finally, the accuracy of the drug combination effect prediction model constructed by the present invention is evaluated by calculating the MSE and PCC metrics. At the same time, to verify the effectiveness of the CNN model of the present invention, the feature images after KPCA dimensionality reduction processing are respectively input into five classic methods such as DNN, FCNN, Linear, SVM, and Decision Tree for comparative experiments. The comparison results of the prediction performance of each method on different cancer types are shown in Table 1.

[0089] Table 1 Evaluation results of the prediction performance of each model on different cancer types

[0090]

[0091]

[0092] Example 3

[0093] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a multi-cancer combined drug use effectiveness prediction system based on CNN disclosed in an embodiment of the present invention. Among them, Figure 2 the described device can be applied to the corresponding prediction terminal, prediction device or server, and the server can be a local server or a cloud server, which is not limited in the embodiments of the present invention. As Figure 3As shown, the device may include:

[0094] A data acquisition module 201, which is used to acquire the characteristic data of the drug combination to be predicted and preprocess the characteristic data of the drug combination to be predicted to obtain initial data;

[0095] A feature dimensionality reduction module 202, which is used to perform dimensionality reduction processing on the initial data through a KPCA model to obtain the dimensionality-reduced features;

[0096] A feature reconstruction module 203, which is used to perform feature reconstruction on the dimensionality-reduced features to obtain reconstructed features;

[0097] A prediction module 204, which is used to input the reconstructed features into a CNN model to obtain a predicted value.

[0098] In an optional embodiment, the data acquisition module 201 acquires the characteristic data of the drug combination to be predicted and preprocesses the characteristic data of the drug combination to be predicted to obtain initial data, including:

[0099] Performing oversampling on the characteristic data of the drug combination to be predicted to obtain balanced data;

[0100] Performing standardization processing on the balanced data to obtain initial data.

[0101] In the embodiment of the present invention, first, the SMOTE algorithm is used to perform oversampling on the characteristic data of the drug combination to be predicted to solve the problem of class imbalance and generate a synthetic balanced data distribution. Then, the balanced data is subjected to standardization processing.

[0102] In an optional embodiment, the data acquisition module 201 performs standardization processing on the balanced data to obtain initial data, including:

[0103] When the characteristic distribution characteristic of the balanced data is centrality, performing min-max normalization processing on the balanced data;

[0104] When the characteristic distribution characteristic of the balanced data is discreteness, performing logarithmic normalization processing on the balanced data.

[0105] In the embodiment of the present invention, when the characteristic distribution characteristic of the balanced data is centrality, min-max normalization is used to process the balanced data and map it to the [0,1] interval. The formula is: Where X is the input feature, X min and X max are the minimum and maximum values of the feature respectively. When the characteristic distribution characteristic of the balanced data is discreteness, logarithmic normalization is used to perform logarithmic transformation processing on the balanced data. The formula is: where X is the input feature, X min and X max are the minimum and maximum values of the feature respectively.

[0106] In an optional implementation, the feature dimensionality reduction module 202 performs dimensionality reduction on the initial data through a KPCA model, and the obtained reduced-dimensional features include:

[0107] The KPCA model uses a Gaussian kernel function to perform a non-linear mapping on the initial data and performs PCA dimensionality reduction in the high-dimensional feature space to obtain the reduced-dimensional features.

[0108] In an optional implementation, the feature reconstruction module 203 performs feature reconstruction on the reduced-dimensional features, and the obtained reconstructed features include:

[0109] The reduced-dimensional features are normalized and reconstructed to obtain a two-dimensional feature image with pixels of 120*120, thereby establishing a mapping relationship from the original high-dimensional data to the low-dimensional image representation, and using the two-dimensional feature image as the reconstructed feature.

[0110] In an optional implementation, the CNN model adopts a double convolutional block structure, including a first convolutional block and a second convolutional block. The number of convolutional kernels in the convolutional layer of the first convolutional block is 8, which is used to extract basic features. The number of convolutional kernels in the convolutional layer of the second convolutional block is 16, which is used to extract advanced features. Both the first convolutional block and the second convolutional block are equipped with a batch normalization layer, a ReLU activation function, and a max pooling layer with a stride of 2.

[0111] In this embodiment, first, the non-linear dimensionality reduction technique of KPCA is utilized. Through the kernel trick, the data is mapped into a high-dimensional feature space, and then principal component analysis is performed in this space. Compared with traditional linear PCA, KPCA has the ability to capture the non-linear features of data. In multi-omics data analysis, this characteristic of KPCA is particularly important because biological data often has complex non-linear relationships. By mapping high-dimensional gene expression data and copy number variation data into two-dimensional images, KPCA not only realizes data dimensionality reduction but also retains the key feature information in the data, laying a foundation for subsequent deep learning analysis. Then, CNN is used to perform deep feature learning on these dimension-reduced images. Through its powerful feature extraction ability, it can automatically identify and learn the key features and potential action patterns in drug combinations. Especially in complex medical scenarios such as combination drug use, where there are multi-level and multi-dimensional interactions between drugs, the combination of KPCA and CNN can better characterize the synergistic or antagonistic effects between drugs, providing more reliable decision-making basis for clinicians, ultimately improving the treatment effect and reducing the risk of adverse reactions. This innovative method not only has sufficient interpretability in theory but also shows good effects in practical applications, providing a new research paradigm and solution ideas for complex medical prediction problems.

[0112] This combination drug use prediction method is based on multi-dimensional dimensionality reduction and deep learning technologies, demonstrating significant technical advantages and practical value.

[0113] Through image processing combined with deep learning technology, a new drug data analysis framework is constructed. Using the deep feature extraction ability of CNN, it can effectively identify complex patterns and potential laws in drug interaction data, significantly improving the accuracy and efficiency of feature recognition. This method shows excellent data processing capabilities and model performance, can efficiently process large-scale drug combination data, and adapt to the diverse requirements of prediction tasks. It not only provides a reliable analysis tool for drug screening and development but also has good scalability and adaptability, can flexibly integrate new data sources and prediction requirements, providing strong technical support for medical research and drug development, and has important theoretical significance and application value.

[0114] Embodiment 4

[0115] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of another CNN-based multi-cancer combination drug use effectiveness prediction device disclosed in the embodiments of the present invention. As Figure 3 shown, the device may include:

[0116] A memory 301 storing executable program code;

[0117] A processor 302 coupled to the memory 301;

[0118] The processor 302 calls the executable program code stored in the memory 301 and executes some or all of the steps in the method for predicting the effectiveness of combined drug use for multiple cancers based on CNN disclosed in Embodiment 1 of the present invention.

[0119] Embodiment 5

[0120] The embodiments of the present invention disclose a computer storage medium. When the computer instructions stored in the computer storage medium are called, they are used to execute some or all of the steps in the method for predicting the effectiveness of combined drug use for multiple cancers based on CNN disclosed in Embodiment 1 or Embodiment 2 of the present invention.

[0121] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0122] Through the above specific descriptions of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, and the storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data.

[0123] Finally, it should be noted that: The method and system for predicting the effectiveness of combined drug use for multiple cancers based on CNN disclosed in the embodiments of the present invention only disclose the preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, rather than limiting them; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A CNN-based method for predicting the effectiveness of combined drug therapy for multiple cancers, characterized in that: The method comprises: Acquiring characteristic data of the drug combination to be predicted and preprocessing the characteristic data of the drug combination to be predicted to obtain initial data; Performing dimensionality reduction processing on the initial data through a KPCA model to obtain features after dimensionality reduction; Reconstructing the features after dimension reduction to obtain reconstructed features; The reconstructed features are input into the CNN model to obtain the predicted values.

2. The CNN-based method for predicting the effectiveness of combined drug therapy for multiple cancers according to claim 1, characterized in that: The characteristic data of the drug combination to be predicted includes drug characteristic information, drug concentration information and cell line gene expression data.

3. The CNN-based method for predicting the effectiveness of combined drug therapy for multiple cancers according to claim 1, characterized in that: The acquiring characteristic data of the drug combination to be predicted and preprocessing the characteristic data of the drug combination to be predicted to obtain initial data comprises: Oversampling the characteristic data of the drug combination to be predicted to obtain balanced data; The equilibrium data are standardized to obtain initial data.

4. The CNN-based method for predicting the effectiveness of combined drug therapy for multiple cancers according to claim 3, characterized in that: The balance data is standardized to obtain initial data including: When the characteristic distribution characteristic of the balance data is centralization, performing minimum-maximum normalization processing on the balance data; When the characteristic distribution characteristic of the balance data is discrete, logarithmic normalization is performed on the balance data.

5. The CNN-based method for predicting the effectiveness of combined drug therapy for multiple cancers according to claim 1, characterized in that: The dimensionality reduction process of the initial data is performed by the KPCA model to obtain the features after dimensionality reduction, including: The KPCA model uses a Gaussian kernel function to perform nonlinear mapping on the initial data, and performs PCA dimensionality reduction in a high-dimensional feature space to obtain features after dimensionality reduction.

6. The method for predicting the effectiveness of combined drug therapy for multiple cancers based on CNN according to claim 1, characterized in that: The reconstructing the features after dimension reduction to obtain the reconstructed features comprises: The reduced-dimensional features are normalized and reconstructed to obtain a two-dimensional feature image with 120*120 pixels, and the two-dimensional feature image is used as the reconstructed feature.

7. The CNN-based method for predicting the effectiveness of combined drug therapy for multiple cancers according to claim 1, characterized in that: The CNN model includes a first convolution block and a second convolution block, the number of convolution kernels of the convolution layer in the first convolution block is 8, and the number of convolution kernels of the convolution layer in the second convolution block is 16.

8. A CNN-based multi-cancer combined drug effectiveness prediction system, characterized by: The system comprises: A data acquisition module, the data acquisition module is used to acquire characteristic data of the drug combination to be predicted and pre-process the characteristic data of the drug combination to be predicted to obtain initial data; A feature dimension reduction module, wherein the feature dimension reduction module is used to perform dimension reduction processing on the initial data through a KPCA model to obtain features after dimension reduction; A feature reconstruction module, wherein the feature reconstruction module is used to reconstruct the features after dimensionality reduction to obtain reconstructed features; A prediction module is used to input the reconstructed features into a CNN model to obtain a predicted value.

9. A CNN-based device for predicting the effectiveness of combined drug therapy for multiple cancers, characterized in that: The device comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the CNN-based multi-cancer combination drug effectiveness prediction method as described in any one of claims 1-7.

10. A computer storage medium, characterized in that: The computer storage medium stores computer instructions, which, when called, are used to execute the CNN-based multi-cancer combination drug effectiveness prediction method as described in any one of claims 1 to 7.