CDK4 / 6 inhibitor sensitivity scoring model

By constructing an AI-based CDK4/6 inhibitor sensitivity scoring model and using signaling pathway activity profiles to predict patient responses, this solves the problem of the inability to effectively screen CDK4/6 inhibitor-sensitive patients in existing technologies, and enables the screening of new indications for CDK4/6 inhibitors and improves treatment efficacy.

WO2026158513A1PCT designated stage Publication Date: 2026-07-30PHIL RIVERS TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
PHIL RIVERS TECH LTD
Filing Date
2026-01-23
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Current technologies are not yet able to effectively screen breast cancer patients who are sensitive to or resistant to CDK4/6 inhibitors, resulting in poor treatment outcomes. There is a lack of reliable biomarkers to predict patient response to CDK4/6 inhibitors.

Method used

A CDK4/6 inhibitor sensitivity scoring model based on artificial intelligence deep learning algorithm was constructed. By analyzing the tumor genome data of patients to generate signaling pathway activity profiles, the scoring model was trained using machine learning methods to predict the sensitivity of patients to CDK4/6 inhibitors.

Benefits of technology

It enables accurate assessment of CDK4/6 inhibitor sensitivity, helps select patient populations that may respond to the drug, expands new indications for CDK4/6 inhibitors, and improves treatment efficacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2026074411_30072026_PF_FP_ABST
    Figure CN2026074411_30072026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to a CDK4 / 6 inhibitor sensitivity scoring model. Provided is a method for constructing a CDK4 / 6 inhibitor sensitivity scoring model. The constructed CDK4 / 6 inhibitor sensitivity scoring model is an artificial intelligence scoring model. The model can be used to score the sensitivity of cancer patients to CDK4 / 6 inhibitors. A training set initially used for the model comes from sensitivity samples of breast cancer patients to CDK4 / 6 inhibitors. The model has been validated to be applicable in scoring the sensitivity of breast cancer patients to CDK4 / 6 inhibitors. As the model is extended to score the sensitivity of patient populations of more tumor types and subtypes thereof to CDK4 / 6 inhibitors, new indications of CDK4 / 6 inhibitors can be expanded. For example, a therapeutic response of a patient to a CDK4 / 6 inhibitor may be predicted, a determination may be made as to whether to treat a patient with a CDK4 / 6 inhibitor, and patient populations that are likely to respond to CDK4 / 6 inhibitor treatment may also be selected.
Need to check novelty before this filing date? Find Prior Art

Description

CDK4 / 6 inhibitor sensitivity scoring model

[0001] Cross-citation of related applications

[0002] This application claims priority to Chinese patent applications CN 202510114105.8, filed on January 24, 2025, and CN 202510846125.4, filed on June 23, 2025, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the application of CDK4 / 6 inhibitors in cancer treatment and artificial intelligence technology, and more specifically to the use of artificial intelligence algorithms and models to predict patients' sensitivity to CDK4 / 6 inhibitors. Background Technology

[0004] Both domestically and internationally marketed cyclin-dependent kinase 4 (CDK4 / 6) inhibitors have been approved for use as monotherapy or in combination therapy for the treatment of hormone receptor (HR)-positive, human epidermal growth factor receptor 2 (HER2)-negative locally advanced or metastatic breast cancer. The approved indications for these marketed CDK4 / 6 inhibitors are concentrated in HR+ and HER2- breast cancer, indicating that these drug molecules, which target and inhibit CDK4 / 6, have similar efficacy in treating tumors. The similar efficacy achieved by these CDK4 / 6 inhibitors in HR+ and HER2- breast cancer patients suggests that these CDK4 / 6 inhibitors exert their therapeutic effect through the same mechanism; that is, these CDK4 / 6 inhibitors could also achieve similar efficacy in appropriate new indications.

[0005] On the other hand, these CDK4 / 6 inhibitors have poor clinical efficacy against triple-negative (HR- / HER2-) subtypes of breast cancer. Therefore, the differences in signaling pathway functional characteristics among breast cancer subtypes may largely represent the differences in patients' sensitivity to and resistance to CDK4 / 6 inhibitors.

[0006] In clinical practice of cancer treatment, CDK4 / 6 inhibitors only benefit a subset of cancer patients. There have been explorations to screen suitable patients using biomarkers sensitive to CDK4 / 6 inhibitors. For example, the PALOMA-1 trial attempted to use CCND1 amplification or p16 deletion as biomarkers. However, clinical studies have shown that patients with this biomarker did not have a significant improvement in progression-free survival (PFS) compared to all ER+ (estrogen receptor-positive) / HER2- patients (Finn, RS, et al., The cyclin-dependent kinase 4 / 6 inhibitor palbociclib in combination with letrozole versus letrozole alone as first-line treatment of oestrogen receptor-positive, HER2-negative, advanced breast cancer (PALOMA-1 / TRIO-18): a randomised phase 2 study. Lancet Oncol, 2015.16(1):p.25-35).

[0007] Turner used a dichotomy to distinguish between high and low CCNE1 expression in patients compared to the overall median, indicating that patients with low CCNE1 expression were more likely to benefit from CDK4 / 6 inhibitors than those with high expression (Turner, NC, et al., Cyclin E1 Expression and Palbociclib Efficacy in Previously Treated Hormone Receptor-Positive Metastatic Breast Cancer. J Clin Oncol, 2019, 37(14): p. 1169-1178). However, there is currently no clear clinical practice standard to predict whether a patient will benefit based on this characteristic.

[0008] Herrera-Abreu pointed out that loss of function of the RB1 gene leads to congenital resistance to CDK4 / 6 inhibitors (Herrera-Abreu, MT, et al., Early Adaptation and Acquired Resistance to CDK4 / 6 Inhibition in Estrogen Receptor-Positive Breast Cancer. Cancer Res, 2016, 76(8): p. 2301-13.). However, after a large-scale analysis of thousands of patients in clinical trials, Bertucci's team found that only about 7% (26 cases) of patients carried this variant trait (Hortobagyi, GN, et al., Updated results from MONALEESA-2, a phase III trial of first-line ribociclib plus letrozole versus placebo plus letrozole in hormone receptor-positive, HER2-negative advanced breast cancer. Ann Oncol, 2018. 29(7): p. 1541-1547.; Slamon, DJ, et al., Phase III Randomized Study of Ribociclib and Fulvestrant in Hormone Receptor-Positive, Human Epidermal Growth Factor Receptor 2-Negative Advanced Breast Cancer: MONALEESA-3. J Clin Oncol, 2018. 36(24): p. 2465-2472; Bertucci, F., et al., Genomic characterization of metastatic breast cancer). Cancers. Nature, 2019, 569(7757):p.560-564. Therefore, using RB1 gene mutation as an exclusion criterion cannot effectively screen for patients suitable for CDK4 / 6 inhibitors.

[0009] Li pointed out that loss of FAT1 gene function leads to resistance to CDK4 / 6 inhibitors (Li, Z., et al., Loss of the FAT1 Tumor Suppressor Promotes Resistance to CDK4 / 6 Inhibitors via the Hippo Pathway. Cancer Cell, 2018.34(6):p.893-905e8.). However, Razavi's analysis of 1,501 patients with ER+ hormone receptor-positive breast cancer in situ and metastasis found that only 2% and 6% of patients, respectively, carried FAT1 gene mutations, and only one-third of these mutations led to loss of FAT1 gene function (Razavi, P., et al., The Genomic Landscape of Endocrine-Resistant Advanced Breast Cancers. Cancer Cell, 2018.34(3):p.427-438e6). Therefore, the FAT1 gene cannot be widely used as a biomarker for screening patients.

[0010] Different teams have reported that the amplification of FGFR1 and FGFR2 reduces patients’ sensitivity to CDK4 / 6 inhibitors. However, in the MONALESSA-2 clinical trial, only 4.7% of the 427 patients carried this feature (Formisano, L., et al., Aberrant FGFR signaling mediates resistance to CDK4 / 6 inhibitors in ER+ breast cancer. Nat Commun, 2019.10(1):p.1373.).

[0011] Asghar points out that a clinically feasible patient screening method cannot yet be established based on drug resistance characteristics (Asghar, US, et al., Systematic Review of Molecular Biomarkers Predictive of Resistance to CDK4 / 6 Inhibition in Metastatic Breast Cancer. JCO Precis Oncol, 2022.6:p.e2100002).

[0012] Chinese patent application CN202010086208.5 (publication number CN111755133A) proposes to determine sensitivity to CDK4 / 6 inhibitors by detecting the expression value of at least one CDK, CDK4 or CDK6. However, studies have shown that the expression of most CDK4 / 6 signaling-related proteins is not different between the two types of cells that are sensitive to CDK4 / 6 inhibitors (CDK4 / 6i-S) and resistant to CDK4 / 6 inhibitors (CDK4 / 6i-R) (Wu, X., et al., Distinct CDK6 complexes determine tumor cell response to CDK4 / 6 inhibitors and degraders. Nat Cancer, 2021.2(4):p.429-443.).

[0013] To date, various research methods, including genomic variations, gene expression data, and circulating biomarkers, have attempted to identify potential characteristics of patient sensitivity or resistance to CDK4 / 6 inhibitors, but none have demonstrated clinical efficacy (Migliaccio, I., et al., CDK4 / 6 inhibitors: A focus on biomarkers of response and post-treatment therapeutic strategies in hormone receptor-positive HER2-negative breast cancer. Cancer Treat Rev, 2021.93:p.102136). A reliable and effective method for identifying CDK4 / 6 inhibitor tumor subtype indications remains an unmet need.

[0014] In order to obtain a model that can characterize the potential mechanism features and be used to evaluate whether subjects may be sensitive to CDK4 / 6 inhibitors, we hope to establish a CDK4 / 6 inhibitor sensitivity scoring model. Summary of the Invention

[0015] According to at least one embodiment of this application, based on the applicant's original digital twin technology, and by combining the data-driven algorithm DAGM (SCI journal: EBioMedicine. 2021 Jul; 69:103446; Chinese authorized patent number: ZL201880003025.3, authorized announcement number CN111602201B, the entire contents of which are incorporated herein by reference), a scoring model for assessing the sensitivity of subjects to CDK4 / 6 inhibitors has been successfully constructed.

[0016] In the method provided in this application, based on the artificial intelligence deep learning computation method DAGM, the overall signaling pathway activity profile (APSP) of patients is obtained by starting from the tumor genome information of many patients through DAGM. Based on the APSP, an artificial intelligence deep learning scoring model for judging the efficacy of CDK4 / 6 inhibitors against different cancers is trained and established using machine learning methods.

[0017] In addition to providing a CDK4 / 6 inhibitor sensitivity scoring model, another objective of this application is to provide a method for scoring the sensitivity of CDK4 / 6 inhibitors using this scoring model. This allows for the prediction of cancer patient response to CDK4 / 6 inhibitor treatment, determination of whether to treat cancer patients with CDK4 / 6 inhibitors, and selection of patient populations or cancer subtypes that may respond to CDK4 / 6 inhibitor treatment.

[0018] According to a first aspect of this application, a method for constructing a CDK4 / 6 inhibitor sensitivity scoring model is provided. The method may include: acquiring a training set comprising multiple training samples, each training sample including patient genomic data and a label indicating the patient's sensitivity to CDK4 / 6 inhibitors; generating a signaling pathway activity profile of the patient's tumor cells based on the patient genomic data in each training sample; selecting signaling features from the signaling pathway activity profile as model features; and training the model using an artificial intelligence algorithm based on the model features and the label indicating sensitivity to CDK4 / 6 inhibitors to obtain a CDK4 / 6 inhibitor sensitivity scoring model.

[0019] The patient genomic data mentioned herein may include the patient's germline genomic data (i.e., normal genomic data) and tumor genomic data, or may directly include the patient's genomic variation data.

[0020] In the method according to the first aspect of this application, preferably, generating a signaling pathway activity profile of the patient's tumor cells based on the patient's genomic data in each training sample may include: extracting genomic variation information based on the differences between the patient's germline genomic data and tumor genomic data in each training sample; and converting the genomic variation information into a corresponding signaling pathway activity profile.

[0021] Furthermore, the extraction of genomic variation information based on the differences between the germline genome data and tumor genome data of patients in each training sample may include: performing quality control and variation detection on the germline genome data and the tumor genome data to obtain variation sites.

[0022] Furthermore, the conversion of genomic variation information into a corresponding signaling pathway activity profile may include: calculating the genomic variation information using the DAGM algorithm to obtain the signaling pathway activity profile.

[0023] On the other hand, if the patient genome data in the training samples is directly genome variation data, then this variation data can be directly used as variation information to convert and obtain the signaling pathway activity profile.

[0024] In the method according to the first aspect of this application, preferably, the step of selecting signal features from the signal pathway activity spectrum as model features may include: performing feature engineering on the signal pathway activity spectrum to select signal features from the signal pathway activity spectrum as model features.

[0025] Furthermore, the step of selecting signal features from the signal pathway activity spectrum as model features may include: calculating the Z-score of each signal feature in the signal pathway activity spectrum; and selecting signal features with an absolute Z-score of not less than 3 as model features.

[0026] In the method according to the first aspect of this application, preferably, the selection of signal features from the signaling pathway activity spectrum as model features may include: selecting signal features related to cell cycle management and immune response from the signaling pathway activity spectrum as model features.

[0027] In the method according to the first aspect of this application, preferably, the step of training the model using an artificial intelligence algorithm based on the model features and the label of sensitivity to CDK4 / 6 inhibitors to obtain a CDK4 / 6 inhibitor sensitivity scoring model may further include: inputting the model features into the artificial intelligence model being trained to obtain probabilistic features of CDK4 / 6 inhibitor sensitivity; obtaining a classification result of sensitivity based on the probabilistic features of CDK4 / 6 inhibitor sensitivity; and adjusting the parameters of the artificial intelligence model using the label of sensitivity to CDK4 / 6 inhibitors as the output target so that the obtained classification result of sensitivity is consistent with the output target.

[0028] In the method according to the first aspect of this application, preferably, the artificial intelligence algorithm may include a machine learning algorithm and a deep learning algorithm. Preferably, the machine learning algorithm is a random forest algorithm.

[0029] According to a second aspect of this application, a method for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model is provided. The method may include: generating a signaling pathway activity profile of the patient's tumor cells based on the patient's genomic data; selecting signal features from the signaling pathway activity profile as model features; inputting the model features into a CDK4 / 6 inhibitor sensitivity scoring model obtained according to the method described in the first aspect of this application, and outputting a sensitivity score of the patient to the CDK4 / 6 inhibitor. Here, based on the signal features selected for model training during the construction of the CDK4 / 6 inhibitor sensitivity scoring model, corresponding signal features are selected from the signaling pathway activity profile as model features.

[0030] The patient genomic data mentioned herein may include the patient's germline genomic data (i.e., normal genomic data) and tumor genomic data, or may directly include the patient's genomic variation data.

[0031] In the method according to the second aspect of this application, preferably, generating a signaling pathway activity profile of the patient's tumor cells based on the patient's genomic data may include: extracting genomic variation information based on the differences between the patient's germline genomic data and tumor genomic data; and converting the genomic variation information into a corresponding signaling pathway activity profile.

[0032] Furthermore, the extraction of genomic variation information based on the differences between the patient's germline genome data and tumor genome data may include: performing quality control and variation detection on the germline genome data and the tumor genome data to obtain variation sites.

[0033] Furthermore, the conversion of genomic variation information into a corresponding signaling pathway activity profile may include: calculating the genomic variation information using the DAGM algorithm to obtain the signaling pathway activity profile.

[0034] On the other hand, if the patient's genomic data is directly genomic variation data, then this variation data can be directly used as variation information to obtain the signaling pathway activity profile.

[0035] In the method according to the second aspect of this application, preferably, the sensitivity score of the patient to the CDK4 / 6 inhibitor output by the CDK4 / 6 inhibitor sensitivity scoring model is a probability score. Therefore, the method according to the second aspect of this application may further include: making a judgment about the patient's sensitivity to the CDK4 / 6 inhibitor based on the sensitivity score.

[0036] In the method according to the second aspect of this application, preferably, the artificial intelligence model includes a machine learning model and a deep learning model. Preferably, the machine learning model is a random forest model.

[0037] According to a third aspect of this application, a method is provided for constructing a sensitivity scoring model for breast cancer patients to CDK4 / 6 inhibitors. The method may include: acquiring a training set comprising multiple training samples, each training sample including genomic data of a breast cancer patient and a label indicating the patient's sensitivity to CDK4 / 6 inhibitors; generating a signaling pathway activity profile of the patient's tumor cells based on the patient's genomic data in each training sample; selecting signaling features from the signaling pathway activity profile as model features; and training the model using an artificial intelligence algorithm based on the model features and the label indicating sensitivity to CDK4 / 6 inhibitors to obtain a sensitivity scoring model for breast cancer patients to CDK4 / 6 inhibitors.

[0038] In the method according to the third aspect of this application, preferably, the plurality of training samples correspond to two breast cancer subtypes: HR+ / HER2- and HR- / HER2-.

[0039] According to a fourth aspect of this application, a method is provided for scoring the sensitivity of breast cancer patients to CDK4 / 6 inhibitors using an artificial intelligence model. The method may include: generating a signaling pathway activity profile of the patient's tumor cells based on the patient's genomic data; selecting signal features from the signaling pathway activity profile as model features; inputting the model features into a sensitivity scoring model for breast cancer patients to CDK4 / 6 inhibitors obtained according to the method described in the third aspect of this application, and outputting a sensitivity score. Here, based on the signal features selected for model training during the construction of the sensitivity scoring model for breast cancer patients to CDK4 / 6 inhibitors, corresponding signal features are selected from the signaling pathway activity profile as model features.

[0040] According to a fifth aspect of this application, a method for predicting a patient's treatment response to a CDK4 / 6 inhibitor is provided. The method may include: obtaining a sensitivity score for the patient to the CDK4 / 6 inhibitor using an artificial intelligence model as described in a second aspect of this application; and predicting the patient's treatment response to the CDK4 / 6 inhibitor based on the obtained sensitivity score.

[0041] In the method according to the fifth aspect of this application, if the sensitivity score is greater than 0.55, greater than 0.60, greater than 0.65, greater than 0.70, greater than 0.75, greater than 0.8, greater than 0.85, or greater than 0.9, then the patient is predicted to respond to treatment with a CDK4 / 6 inhibitor; preferably, if the sensitivity score is greater than 0.70 or greater than 0.75, then the patient is predicted to respond to treatment with a CDK4 / 6 inhibitor.

[0042] According to a sixth aspect of this application, a method is provided for determining whether a patient should be treated with a CDK4 / 6 inhibitor. The method may include: obtaining a sensitivity score for the patient to a CDK4 / 6 inhibitor using an artificial intelligence model as described in a second aspect of this application; and determining whether to treat the patient with a CDK4 / 6 inhibitor based on the obtained sensitivity score.

[0043] In the method according to the sixth aspect of this application, if the sensitivity score is greater than 0.55, greater than 0.60, greater than 0.65, greater than 0.70, greater than 0.75, greater than 0.8, or greater than 0.85, or greater than 0.9, the patient will receive treatment with a CDK4 / 6 inhibitor; preferably, if the sensitivity score is greater than 0.70 or greater than 0.75, the patient will receive treatment with a CDK4 / 6 inhibitor.

[0044] According to a seventh aspect of this application, a method is provided for selecting a patient population that may respond to CDK4 / 6 inhibitor treatment. The method may include: obtaining a sensitivity score for the patient population to the CDK4 / 6 inhibitor using an artificial intelligence model as described in a second aspect of this application; and selecting the patient population to receive the CDK4 / 6 inhibitor if the obtained sensitivity score is higher than a threshold.

[0045] In the method according to the seventh aspect of this application, the threshold is 0.65, 0.70, 0.75, 0.8, 0.85, or 0.9; preferably, the threshold is 0.70 or 0.75.

[0046] In the method according to the seventh aspect of this application, preferably, the patient population is a patient population classified into one or more cancer subtypes.

[0047] According to an eighth aspect of this application, a system is provided for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model. The system may include: a data acquisition module for acquiring patient genomic data; a model feature generation module for generating a signaling pathway activity profile of the patient's tumor cells based on the patient's genomic data, and selecting signal features from the signaling pathway activity profile as model features; and a CDK4 / 6 inhibitor sensitivity scoring model obtained according to the method described in the first aspect of this application, used to output a sensitivity score of the patient to CDK4 / 6 inhibitors based on the model features. The model feature generation module may be further configured to select corresponding signal features from the signaling pathway activity profile as model features based on the signal features selected for model training during the construction of the CDK4 / 6 inhibitor sensitivity scoring model.

[0048] According to a ninth aspect of this application, a non-transitory computer-readable storage medium is provided for storing a computer program. The computer program includes instructions that, when executed by a processor of an electronic device, cause the electronic device to perform at least one of the following methods: a method for constructing a CDK4 / 6 inhibitor sensitivity scoring model as described in a first aspect of this application; a method for scoring the sensitivity of a CDK4 / 6 inhibitor using an artificial intelligence model as described in a second aspect of this application; a method for constructing a CDK4 / 6 inhibitor sensitivity scoring model for breast cancer patients as described in a third aspect of this application; a method for scoring the sensitivity of breast cancer patients to CDK4 / 6 inhibitors using an artificial intelligence model as described in a fourth aspect of this application; a method for predicting a patient's treatment response to a CDK4 / 6 inhibitor as described in a fifth aspect of this application; a method for determining whether a patient should be treated with a CDK4 / 6 inhibitor as described in a sixth aspect of this application; or a method for selecting a patient population that may respond to CDK4 / 6 inhibitor treatment as described in a seventh aspect of this application.

[0049] According to a tenth aspect of this application, a computer system is provided. The computer system may include: a processor; a memory; and a computer program. The computer program is stored in the memory and configured to be executed by the processor, the computer program including instructions for implementing at least one of the following methods: a method for constructing a CDK4 / 6 inhibitor sensitivity scoring model as described in a first aspect of this application; a method for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model as described in a second aspect of this application; a method for constructing a CDK4 / 6 inhibitor sensitivity scoring model for breast cancer patients as described in a third aspect of this application; a method for scoring the sensitivity of breast cancer patients to CDK4 / 6 inhibitors using an artificial intelligence model as described in a fourth aspect of this application; a method for predicting a patient's treatment response to a CDK4 / 6 inhibitor as described in a fifth aspect of this application; a method for determining whether a patient should be treated with a CDK4 / 6 inhibitor as described in a sixth aspect of this application; or a method for selecting a patient population that may respond to CDK4 / 6 inhibitor treatment as described in a seventh aspect of this application.

[0050] According to embodiments of this application, various events within a cell are essentially the activity states of various signaling pathways, and these pathways communicate and regulate each other to varying degrees. The state of the cell is equivalent to the combination of the activity states of these signaling pathways. Using the deep learning computational method DAGM developed by the applicant, which infers the set of deterministic events within a cell based on global genetic information or a subset thereof, the activity profile of signaling pathways can be derived from the tumor genomic information of numerous patients. This application utilizes this method to obtain the overall signaling pathway activity profile (APSP) of patients by analyzing breast cancer tumor genomic data, and based on the APSP, a scoring model for judging the efficacy of CDK4 / 6 inhibitors in cancer is trained using machine learning methods. This application uses this scoring model to obtain pan-cancer tumor genomic data analysis results, discovering multiple subtypes of multiple cancer types as new indications for CDK4 / 6 inhibitors, and verifying from multiple perspectives that the new indications screened in this application have clinical efficacy.

[0051] Therefore, this application also provides a method for expanding the indications for drugs, particularly for CDK4 / 6 inhibitors. By establishing a CDK4 / 6 inhibitor response mechanism (sensitivity score) model, this model can be used to screen for new indications. Attached Figure Description

[0052] This application will be more fully understood through the following detailed description and in conjunction with the accompanying drawings, wherein similar elements or components are numbered in a similar manner, wherein:

[0053] Figure 1 shows a flowchart of a method for constructing a CDK4 / 6 inhibitor sensitivity scoring model according to an embodiment of this application.

[0054] Figure 2 shows a flowchart of a method for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model according to an embodiment of this application.

[0055] Figure 3 illustrates a schematic diagram relating the method for constructing a CDK4 / 6 inhibitor sensitivity scoring model according to an embodiment of this application to the method for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model.

[0056] Figure 4 is the ROC curve of the scoring model according to an embodiment of this application using leave-one-out cross-validation on the TCGA training set.

[0057] Figure 5 is the ROC curve of the scoring model according to an embodiment of this application on a Chinese patient cohort (internal data).

[0058] Figure 6 is the ROC curve of the scoring model according to an embodiment of this application on an external dataset.

[0059] Figure 7 shows key APSP features in the scoring model according to an embodiment of this application.

[0060] Figure 8 shows a schematic block diagram of a system for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model according to an embodiment of this application. Detailed Implementation

[0061] The technical solutions of this application will be further described in detail below through embodiments and in conjunction with the accompanying drawings. Those skilled in the art will understand that the various embodiments described below are merely exemplary and should not be considered as limiting the scope of this application.

[0062] Unless otherwise stated, the scientific and technical terms used in this application have the meanings commonly understood by one of ordinary skill in the art to which this application pertains.

[0063] The "DAGM" algorithm mentioned in this application is an algorithm developed by the applicant, which has been published in the literature (SCI journal: EBioMedicine. 2021 Jul; 69:103446) and has been granted a Chinese patent (Chinese patent ZL201880003025.3, authorization announcement number CN111602201B). Its entire contents are incorporated herein by reference. The input of this algorithm is human genome sequence data, and the output is the pathway activity profile (APSP) of the subject (see further explanation below). The DAGM algorithm is an artificial intelligence-based method for assessing genomic mutation damage. It employs eQTL (expression Quantitative Trait Loci, a type of genetic locus that can influence gene expression) analysis technology to establish a correlation between "gene expression" as a trait and gene mutation. This algorithm can efficiently transform complex, discrete, and high-dimensional genomic variation information into continuous, lower-dimensional, and highly correlated pathway activity profiles (APSPs). Through this process, the DAGM algorithm takes genomic data as input and, with the help of multiple technologies such as the algorithm's built-in data model, deep learning model, and knowledge model, ultimately outputs the overall cellular signaling pathway functional profile of the subject, providing a comprehensive assessment of cellular biological function. The specific process first requires acquiring the subject's whole-exome sequencing (WES) data (Fast file), and after rigorous quality control, extracting rare variant (MAF < 0.01) sites. After processing by the DAGM algorithm, the VCF file of the variant sites generates a set of continuous variables representing the functional characteristics of the subject's cells, namely the signaling pathway activity profile (APSP). The APSP is further transformed into a point in a high-dimensional space, where each dimension represents a cellular function. This process transforms each biological individual into a digital entity, forming a functional digital twin of the subject. By integrating these digital entities to form a topological structure, a digital twin model representing a group with similar characteristics or diseases can be constructed. Within this high-dimensional space composed of specific digital twins, this application further develops an artificial intelligence-driven mechanism model that can effectively distinguish between drug-responsive and non-responsive patients. This mechanism model, by encoding a generalizable hyperplane (scoring model), enables computer simulation of clinical trials, thereby discovering new indications for CDK4 / 6 inhibitors for various subtypes of many tumor diseases, such as traditional chordoma.

[0064] APSP, or Activity Profiles of Signaling Pathways, is the output of the DAGM algorithm. As a feature profile of signaling pathways, APSP itself represents signaling pathway characteristics; therefore, models built upon APSP can reveal underlying mechanisms.

[0065] The term “treatment” refers to a therapeutic approach. When a specific condition is involved, treatment means: (1) alleviating one or more biological manifestations of the disease or condition; (2) interfering with (a) one or more points in a biological cascade that causes or precipitates the condition or (b) one or more biological manifestations of the condition; (3) improving one or more symptoms, effects or side effects associated with the condition, or one or more symptoms, effects or side effects associated with the condition or its treatment; or (4) slowing the development of the condition or one or more biological manifestations of the condition.

[0066] The terms "patient" or "subject" refer to any animal, preferably a mammal, that has been or is about to receive administration of the compound according to embodiments of this application, with humans being the most preferred. The term "mammal" includes any mammal. Examples of mammals include, but are not limited to, cattle, horses, sheep, pigs, cats, dogs, mice, rats, rabbits, guinea pigs, monkeys, and humans, with humans being the most preferred.

[0067] As used in this application, the term "sensitive" refers to something that works, has a therapeutic effect, or responds to treatment. Sensitivity refers to the degree of sensitivity; for example, if expressed numerically, a sensitivity of 0 indicates insensitivity, while a sensitivity of 1 indicates sensitivity. Terms opposite to "sensitive" may include: resistance, inhibition, ineffectiveness, lack of therapeutic effect, and no response to treatment.

[0068] As used in this application, the term "Z-score," also known as a standard score, is a statistical indicator used to measure the relative position of a data point to the mean of a dataset. Essentially, it is the process of standardizing raw data to a "standard normal distribution" (mean of 0, standard deviation of 1), facilitating comparisons between different datasets or outlier detection. In this application, the Z-score can be used to indicate whether the mean of a small sample (e.g., chordoma patients) drawn under specific conditions from the entire sample (e.g., all cancer patients in the TCGA database and a local database) has a statistically significant change relative to a random sample for a certain characteristic value (e.g., the model score in this paper). Typically, in biomedical research, a Z-score greater than or equal to 3 (i.e., not less than 3) means that the mean score of a certain cancer type is 3 standard deviations or more higher than the mean score of the randomly sampled cancer patients, indicating that the mean score of that cancer type is significantly higher than that of the random sample (Dowdy, S., Wearden, S., & Childo, D. (2011). Statistics for research. John Wiley & Sons.).

[0069] As used herein, the term "about" or "approximately" refers to an acceptable error in a particular value as determined by a person skilled in the art, depending in part on how the value is measured or determined. In some embodiments, the term "about" or "approximately" means within 1, 2, 3, or 4 standard deviations. In some embodiments, the term "about" or "approximately" means within 50%, 20%, 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, or 2%, or 1%, 0.5%, or 0.05% of a given value or range.

[0070] When a range of values ​​is listed here, it is intended to cover every value and subrange within that range. For example, “1-5mg” or “about 1mg to about 5mg” is intended to cover 1mg, 2mg, 3mg, 4mg, 5mg, 1-2mg, 1-3mg, 1-4mg, 1-5mg, 2-3mg, 2-4mg, 2-5mg, 3-4mg, 3-5mg, and 4-5mg.

[0071] The terms “including,” “comprising,” and “having” are used in an inclusive, open sense, meaning that additional elements may be included in addition to those specified. The terms “such as” and “e.g.” as used herein are non-limiting and are for illustrative purposes only. “Including” and “including but not limited to” are used interchangeably.

[0072] In this document, the term "and / or" as used in phrases such as "A and / or B" is intended to include both A and B; A or B; (alone) A; and (alone) B. Similarly, the term "and / or" as used in phrases such as "A, B and / or C" is intended to cover each of the following embodiments: A, B and C; A, B or C; A or C; A or B; B or C; A and C; A and B; B and C; (alone) A; (alone) B; and (alone) C.

[0073] As used in this application, the phrase "at least one" in relation to a list of one or more elements should be understood to mean at least one element selected from any one or more elements in the element list, but not necessarily including at least one of each element specifically listed in the element list, and does not exclude any combination of elements in the element list. This definition also allows for the optional presence of elements other than those specifically identified in the element list referred to by the phrase "at least one," whether related to or unrelated to those specifically identified elements. Thus, as a non-limiting example, in one implementation, "at least one of A and B" (or equivalently, "at least one of A or B," or equivalently, "at least one of A and / or B") can refer to at least one, optionally including more than one, referring to A and not having B (and optionally including elements other than B); in another implementation, it refers to at least one, optionally including more than one, referring to B and not having A (and optionally including elements other than A); in yet another implementation, it refers to at least one, optionally including more than one, referring to A, and at least one, optionally including more than one, referring to B (and optionally including other elements); and so on.

[0074] Construction of the scoring model

[0075] At its most basic level, this application proposes a method for constructing a CDK4 / 6 inhibitor sensitivity scoring model.

[0076] Figure 1 shows a flowchart of a method for constructing a CDK4 / 6 inhibitor sensitivity scoring model according to an embodiment of this application.

[0077] As shown in Figure 1, the method 100 for constructing a CDK4 / 6 inhibitor sensitivity scoring model according to an embodiment of this application begins at step S110, i.e., obtaining a training set. The training set should include multiple training samples. Each training sample should include patient genomic data and a label indicating the patient's sensitivity to CDK4 / 6 inhibitors.

[0078] In embodiments according to this application, sample data from patients with two subtypes of breast cancer are used to construct training samples. These two subtypes are HR-positive / HER2-negative breast cancer and triple-negative (HR-negative / HER2-negative) breast cancer. The sensitivity of these two subtypes to CDK4 / 6 inhibitors exhibits significant polarization. HR-positive / HER2-negative patients are sensitive to CDK4 / 6 inhibitors, while triple-negative patients are insensitive. Therefore, the sensitivity labels for CDK4 / 6 inhibitors in these two scenarios can be set to "sensitive" or "insensitive," with values ​​corresponding to "1" and "0," respectively. These values ​​can correspond to the predicted scores in the scoring model.

[0079] Those skilled in the art should understand and fully recognize that, although the embodiments described in detail below use sample data from patients with the two breast cancer subtypes as described above to constitute the training samples, with the continuous discovery of new indications for CDK4 / 6 inhibitors and insensitive tumor subtypes, sample data from patient populations using these tumors and their subtypes can also be used to construct scoring models applicable to a wider range of tumor patients or patients with more specific cancer types. The scoring models proposed in this application, except where explicitly stated, are applicable to a wide range of tumor types and are not intended to be limited to only a specific cancer type.

[0080] In step S120, based on the patient's genomic data in each training sample, a signaling pathway activity profile of the patient's tumor cells is generated.

[0081] The patient genomic data mentioned in steps S110 and S120 may include the patient's germline genomic data (i.e., normal genomic data) and tumor genomic data, or may directly include the patient's genomic variation data.

[0082] A core idea behind this application for constructing the scoring model is to project discrete, high-dimensional genomic variations onto continuous, low-dimensional signaling pathway functional characteristics. This provides a reliable biological basis for subsequent analysis and modeling. Therefore, any method capable of generating a signaling pathway activity profile of a patient's tumor cells based on the patient's tumor genomic data falls within the scope of the scoring model construction method claimed in this application.

[0083] Based on the aforementioned core ideas, this application also proposes a preferred implementation scheme. As mentioned earlier, the DAGM algorithm developed by the applicant is a novel artificial intelligence algorithm for assessing genomic mutation damage. It employs eQTL analysis to associate "gene expression" as a trait with gene mutations, projecting discrete, high-dimensional, and multivariately correlated genomic variation information onto continuous, low-dimensional, and convergently correlated signaling pathway activity characteristics. The DAGM algorithm takes genomic data as input and, through established data models, deep learning models, and knowledge models, ultimately outputs the overall signaling pathway activity profile (APSP) of the subject's cells. For example, the specific process is as follows: First, DNA samples from the patient's normal and tumor tissues are collected to obtain whole-exome sequencing (WES) results (Fastq files). After quality control (QC) processing, rare variant (MAF<0.01) sites are retrieved from these files, including alignment, duplication labeling, InDel realignment, and base quality fraction correction. The variant site format (VCF) file is calculated by the DAGM algorithm to generate the APSP of the patient's tumor. This application utilizes this method to extract the potential biological mechanism characteristics of the response of breast cancer patients to CDK4 / 6 inhibitors, specifically quantitative information on the significant activation and inhibition of several signaling pathways.

[0084] In a preferred embodiment according to this application, specifically in step S120, firstly, genomic variation information is extracted based on the differences between the germline genome data and tumor genome data of patients in each training sample. For example, quality control and variation detection are performed on the germline genome data and the tumor genome data to obtain variation sites. Then, the genomic variation information is converted into the corresponding pathway activity profile (APSP). As described in the embodiments of this application, the genomic variation information is calculated using the DAGM algorithm to obtain the APSP. For details on the DAGM algorithm and APSP, please refer to the explanation above and the more detailed embodiments below.

[0085] However, those skilled in the art should understand and fully recognize that step S120 can also be implemented using methods other than the DAGM algorithm, as long as they conform to the core idea of ​​this application as described above—projecting discrete, high-dimensional genomic variations onto continuous, low-dimensional signaling pathway functional features.

[0086] On the other hand, if the patient genome data in the training samples is directly genome variation data, then this variation data can be directly used as variation information to convert and obtain the APSP.

[0087] Next, in step S130, signal features are selected from the APSP as model features.

[0088] In a preferred embodiment according to this application, specifically in step S130, feature engineering is performed on the APSP to select signal features from the APSP as model features.

[0089] More specifically, the Z-score of each signal feature in the APSP can be calculated. Signal features with an absolute Z-score of not less than 3 are selected as model features.

[0090] On the other hand, in another preferred embodiment of this application, signal features related to cell cycle management and immune response can also be directly selected from APSP as model features.

[0091] Finally, in step S140, based on the aforementioned model features and the aforementioned labels on sensitivity to CDK4 / 6 inhibitors, the model is trained using an artificial intelligence algorithm to obtain a CDK4 / 6 inhibitor sensitivity scoring model.

[0092] In a preferred embodiment, step S140 can be further interpreted as the following process. First, the model features selected in step S130 are input into the AI ​​model being trained to obtain a probability feature regarding CDK4 / 6 inhibitor sensitivity. For example, this probability feature is a value between 0 and 1. Then, a classification result regarding sensitivity is obtained based on this probability feature. For example, the classification result can be 0 or 1, i.e., insensitive or sensitive. For example, when the value of the probability feature is greater than 0.5 (or other thresholds), the classification result is sensitive; otherwise, it is insensitive. Using the label of CDK4 / 6 inhibitor sensitivity as the output target, the parameters of the AI ​​model are adjusted so that the obtained classification result regarding sensitivity is consistent with the output target. Depending on the AI ​​algorithm used, the method and steps for adjusting the parameters of the AI ​​model will also differ, and will not be elaborated here.

[0093] The artificial intelligence algorithm used here can be a machine learning algorithm or a deep learning algorithm. In the preferred embodiment of this application, the artificial intelligence algorithm used is a machine learning algorithm, more specifically, the random forest algorithm.

[0094] In short, the scoring model proposed in this application, including its initial model and the prediction model built after training, is actually a classifier (e.g., a binary classifier). The algorithms used in classifiers are generally divided into machine learning methods and deep learning methods. Commonly used machine learning methods include: Linear Regression, Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Logistic Regression (LR), Decision Tree, K-Means, Random Forest, Naive Bayes, Gradient Boosting, and Ensemble Learning. Deep learning methods generally include: Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Recursive Neural Networks (RNNs), and Long Short-Term Memory (LSTM) networks.

[0095] Although the final scoring model was constructed using the Random Forest algorithm and its initial model in the embodiments of this application described in more detail below, those skilled in the art should understand and fully recognize that other algorithms or initial models besides the Random Forest algorithm and model, whether listed or not, may also be used when constructing the scoring model described in this application.

[0096] Using a rating model for scoring

[0097] This application proposes a method for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model.

[0098] Figure 2 shows a flowchart of a method for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model according to an embodiment of this application.

[0099] As shown in Figure 2, the method 200 for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model according to an embodiment of this application may begin at step S210, i.e., generating a signaling pathway activity profile of the patient's tumor cells based on the patient's genomic data.

[0100] The patient genomic data mentioned herein may include the patient's germline genomic data (i.e., normal genomic data) and tumor genomic data, or may directly include the patient's genomic variation data.

[0101] In a preferred embodiment of this application, specifically in step S210, firstly, genomic variation information is extracted based on the differences between the patient's germline genome data and tumor genome data. For example, quality control and variation detection are performed on the germline genome data and the tumor genome data to obtain variation sites. Then, the genomic variation information is converted into a corresponding pathway activity profile (APSP). As described in the embodiments of this application, the genomic variation information is calculated using the DAGM algorithm to obtain the APSP. For details on the DAGM algorithm and APSP, please refer to the explanation above and the more detailed embodiments below.

[0102] On the other hand, if the patient's genomic data is directly genomic variation data, then this variation data can be directly used as variation information to convert and obtain the APSP.

[0103] In step S220, signal features are selected from the signal pathway activity spectrum as model features.

[0104] Next, in step S230, the model features are input into the CDK4 / 6 inhibitor sensitivity scoring model obtained according to method 100 in Figure 1, and the sensitivity score of the patient to the CDK4 / 6 inhibitor is output.

[0105] Here, based on the signal features selected for model training during the construction of the CDK4 / 6 inhibitor sensitivity scoring model, in step S220, corresponding signal features are selected from the signal pathway activity spectrum as model features.

[0106] In a preferred embodiment of this application, the sensitivity score of the patient to the CDK4 / 6 inhibitor output by the CDK4 / 6 inhibitor sensitivity scoring model can be a probability score. Therefore, method 200 further includes: making a judgment about the patient's sensitivity to the CDK4 / 6 inhibitor based on the sensitivity score.

[0107] The artificial intelligence algorithm used in method 200 can be a machine learning algorithm or a deep learning algorithm. In a preferred embodiment of this application, the artificial intelligence algorithm used is a machine learning algorithm, more specifically a random forest algorithm.

[0108] Figure 3 illustrates a schematic diagram linking the method for constructing a CDK4 / 6 inhibitor sensitivity scoring model according to an embodiment of this application with the method for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model. In the schematic diagram shown in Figure 3, the method for constructing a CDK4 / 6 inhibitor sensitivity scoring model is shown to the left of the dashed line, while the method for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model is shown to the right of the dashed line.

[0109] The processes on the left and right sides of the dashed line in Figure 3 are essentially the methods shown in Figures 1 and 2, respectively. As can be seen from Figure 3, both model construction and the actual scoring process require a model feature extraction step: generating a signaling pathway activity profile of the patient's tumor cells based on the patient's genomic data; then, selecting signal features from the signaling pathway activity profile as model features. The training samples include the patient's genomic data and the patient's sensitivity label to CDK4 / 6 inhibitors. As mentioned earlier, the patient's sensitivity label to CDK4 / 6 inhibitors can be a value of 0 or 1, representing insensitivity and sensitivity, respectively. Therefore, during model construction, the genomic data from the training samples is used to extract model features. The scoring model (i.e., the classifier) ​​is fully trained using the model features as input data and the patient's CDK4 / 6 inhibitor sensitivity label from the training samples as output data, thus obtaining the final CDK4 / 6 inhibitor sensitivity scoring model. The constructed scoring model can then be used in the actual work of scoring (i.e., predicting) the sensitivity to CDK4 / 6 inhibitors, as shown by the process on the right side of the dashed line in Figure 3.

[0110] Scoring models and methods for breast cancer patients

[0111] Similar to Figure 1, a method for constructing a sensitivity scoring model for CDK4 / 6 inhibitors in breast cancer patients can be proposed. This method, specifically applicable to breast cancer, may include: acquiring a training set comprising multiple training samples, each sample including genomic data of a breast cancer patient and a label indicating the patient's sensitivity to CDK4 / 6 inhibitors; generating a signaling pathway activity profile of the patient's tumor cells based on the patient's genomic data in each training sample; selecting signaling features from the signaling pathway activity profile as model features; and training the model using an artificial intelligence algorithm based on the model features and the CDK4 / 6 inhibitor sensitivity label to obtain a sensitivity scoring model for breast cancer patients to CDK4 / 6 inhibitors.

[0112] In a preferred embodiment, the plurality of training samples correspond to two breast cancer subtypes: HR+ / HER2- and HR- / HER2- (triple negative).

[0113] Accordingly, similar to Figure 2, a method for scoring the sensitivity of breast cancer patients to CDK4 / 6 inhibitors using an artificial intelligence model is proposed. This method is also specifically applicable to breast cancer and may include: generating a signaling pathway activity profile of the patient's tumor cells based on the patient's genomic data; selecting signal features from the signaling pathway activity profile as model features; inputting the model features into a breast cancer patient CDK4 / 6 inhibitor sensitivity scoring model obtained according to the method described above, which is specifically applicable to constructing a sensitivity scoring model for breast cancer patients to CDK4 / 6 inhibitors, and outputting a sensitivity score. Here, based on the signal features selected for model training during the construction of the breast cancer patient CDK4 / 6 inhibitor sensitivity scoring model, corresponding signal features are selected from the signaling pathway activity profile as model features.

[0114] Using scoring models to predict treatment response

[0115] A method for predicting a patient's treatment response to a CDK4 / 6 inhibitor is proposed. The method may include: obtaining a sensitivity score for the patient to the CDK4 / 6 inhibitor using an artificial intelligence model as shown in Figure 2; and predicting the patient's treatment response to the CDK4 / 6 inhibitor based on the obtained sensitivity score.

[0116] In this method, a patient is predicted to respond to CDK4 / 6 inhibitor treatment if the obtained sensitivity score is greater than 0.55, 0.60, 0.65, 0.70, 0.75, 0.8, 0.85, or 0.9. Preferably, the threshold for the sensitivity score can be 0.70 or 0.75, that is, if the sensitivity score is greater than 0.70 or greater than 0.75, a patient is predicted to respond to CDK4 / 6 inhibitor treatment.

[0117] Determine whether to treat the patient with a CDK4 / 6 inhibitor.

[0118] A method is proposed to determine whether a patient should be treated with a CDK4 / 6 inhibitor. The method may include: obtaining a sensitivity score for the patient to a CDK4 / 6 inhibitor using an artificial intelligence model, as shown in Figure 2; and determining whether to treat the patient with a CDK4 / 6 inhibitor based on the obtained sensitivity score.

[0119] In this method, if the obtained sensitivity score is greater than 0.55, greater than 0.60, greater than 0.65, greater than 0.70, greater than 0.75, greater than 0.8, greater than 0.85, or greater than 0.9, then the patient will receive CDK4 / 6 inhibitor treatment. Preferably, the threshold for the sensitivity score can be 0.70 or 0.75, that is, if the sensitivity score is greater than 0.70 or greater than 0.75, then the patient will receive CDK4 / 6 inhibitor treatment.

[0120] Select a patient population that is likely to respond to CDK4 / 6 inhibitor treatment.

[0121] A method is proposed for selecting a patient population that may respond to CDK4 / 6 inhibitor treatment. The method may include: obtaining a sensitivity score for the patient population to CDK4 / 6 inhibitors using an artificial intelligence model, as shown in Figure 2; if the obtained sensitivity score is higher than a threshold, then selecting that patient population to receive the CDK4 / 6 inhibitor.

[0122] The threshold value can be 0.65, 0.70, 0.75, 0.8, 0.85, or 0.9. Preferably, the threshold value is 0.70 or 0.75.

[0123] In this method, the patient population is a group of patients classified into one or more cancer subtypes.

[0124] Example 1

[0125] CDK4 / 6 inhibitors have different mechanisms of action for different breast cancer subtypes, resulting in significant differences in clinical efficacy.

[0126] For HR (hormone receptor) positive and HER2 (human epidermal growth factor receptor 2) negative breast cancer (HR+ / HER2-, including Luminal A type (or luminal A type) and Luminal B (HER2 negative) type (or luminal B type)), the mechanism of action of CDK4 / 6 inhibitors is mainly the direct inhibition of tumor cell growth (Goel, S., et al., CDK4 / 6 Inhibition in Cancer: Beyond Cell Cycle Arrest. Trends Cell Biol, 2018.28(11):p.911-925), which enables most patients with this breast cancer subtype to achieve clinical benefits.

[0127] However, CDK4 / 6 inhibitors have limited prognostic benefits for patients with triple-negative breast cancer (HR- / HER2-, which is either triple-negative or basaloid). Only a small percentage of patients benefit from the protective inhibition of bone marrow-derived cells before treatment with CDK4 / 6 inhibitors, allowing the immune system to continue functioning after treatment (Goel, S., et al., CDK4 / 6 inhibition triggers antitumor immunity. Nature, 2017. 548(7668): p. 471-475). Currently, there is still a lack of strong clinical evidence for the efficacy of CDK4 / 6 inhibitors in treating triple-negative breast cancer (Mitri, Z., et al., A phase 1 study with dose expansion of the CDK inhibitor dinaciclib (SCH 727965) in combination with epirubicin in patients with metastatic triple-negative breast cancer. Invest New Drugs, 2015. 33(4): p. 890-4).

[0128] Unbound by theoretical constraints, this application proposes that a high proportion of patients with certain cancer types or subtypes can benefit from treatment by directly inhibiting tumor cell growth through CDK4 / 6 inhibitors. This application establishes a response mechanism model by comparing the differences in molecular biological characteristics between HR+ / HER2- and HR- / HER2- subtypes of breast cancer. This model can screen for specific cancer types or subtypes that benefit from treatment with CDK4 / 6 inhibitors.

[0129] This application establishes the CDK4 / 6 inhibitor screening model through the following operations.

[0130] The model is established in two steps: (1) The variants obtained on the tumor genomic DNA through in vivo microevolution are systematically annotated into a set of cell function and signaling pathway activity profiles using the DAGM algorithm; (2) Using the aforementioned activity profiles as input, based on the label information of different patient groups (i.e., HR+ / HER2- classification of CDK4 / 6 inhibitor approved indications and HR- / HER2- classification of CDK4 / 6 inhibitors with poor clinical efficacy), machine learning algorithms are used to perform regression modeling, and finally a scoring model that can predict patient groups is obtained. This scoring model can be used to screen whether any cancer population is suitable for CDK4 / 6 inhibitors.

[0131] To obtain initial data for model training, this application categorized 980 breast cancer patients from the Cancer Genome Atlas (TCGA) into four subtypes based on the expression levels of ER (estrogen receptor), PR (progesterone receptor), and HER2 (217 of the 980 patients could not be definitively classified):

[0132] -HR+ / HER2- type (i.e., Luminal A and Luminal B (HER2 negative) types, totaling 454 cases)

[0133] -HR+ / HER2+ type (HER2 positive (HR positive type, a total of 58 cases)

[0134] -HR- / HER2+ type (HER2 positive (HR negative) type, a total of 33 cases)

[0135] -HR- / HER2- (triple negative or basaloid type, 127 cases in total) (CSCO Breast Cancer Diagnosis and Treatment Guidelines 2022).

[0136] Using genomic data (whole exome sequencing data of patient tumors) from breast cancer patients with HR+ / HER2- (positive control) and HR- / HER2- (negative control) subtypes obtained from the TCGA database in the previous step, pathway activity profiles (APSPs) were obtained using the DAGM algorithm. Through machine learning methods, a predictive model capable of distinguishing between HR+ / HER2- and HR- / HER2- subtypes of breast cancer was trained. Leave-one-out cross-validation was employed.

[0137] Feature engineering is a crucial step in machine learning, involving the extraction, transformation, and selection of features from raw data to improve model performance. The following are the feature engineering methods involved in this application:

[0138] 1. Missing value handling: Handling missing values ​​is part of feature engineering. This application uses a method that includes the mean, median and interpolation to fill in missing values.

[0139] 2. Standardization and normalization: Scaling features to a similar range to avoid some features having an excessive impact on the model. This application achieves this by subtracting the mean and dividing by the standard deviation, while normalization is achieved by scaling to a specific range, such as [0,1].

[0140] 3. Feature Interaction: Creating new features by combining two or more features to better capture the relationships between features.

[0141] 4. Dimensionality Reduction: Using dimensionality reduction techniques, this application employs Principal Component Analysis (PCA), Linear Discriminant Analysis (LDA), and t-distributed stochastic neighbor embedding (t-SNE) to reduce the number of features while retaining important information in the data.

[0142] 5. Outlier handling: Detect and handle outliers to prevent them from affecting the model.

[0143] 6. Feature Selection: Selecting features that have a stronger predictive power for the target variable. This involves determining which features are most relevant and valuable for the modeling task.

[0144] The feature engineering described in this application can be accomplished using the basic functionality and functions of various versions of tools as described at https: / / www.R-project.org / . Those skilled in the art should understand and fully recognize that any other suitable tools and functions can also be used to perform the feature engineering described herein.

[0145] The following are the feature selection techniques involved in this application:

[0146] 1. Filtering Methods: Using certain statistics of features to filter before training the model. This application adopts variance threshold (by checking the variance of features, removing features with low variance, and setting the median of the variance of all features as the threshold), correlation coefficient (calculating the correlation between features and the target variable, selecting features with high correlation, the correlation is a specific value between 0 and 1, generally 0.6 or 0.8 is selected as the threshold for high and low correlation), and mutual information (measuring the information gain between features and the target variable, selecting features with high information gain, the value range of mutual information is [0, +∞], the larger the value, the closer the relationship between the feature and the target variable. Generally, there is no specific threshold. The specific steps of using mutual information for feature filtering are: (1) Calculate the mutual information value between each feature and the target variable (2) Sort each feature according to the mutual information value (3) Select the top k features with large mutual information values ​​as the filtering results).

[0147] 2. Wrapper Methods: Evaluating the quality of different feature subsets based on model performance. This application employs Recursive Feature Elimination (RFE; repeatedly training the model and eliminating features with the least impact on the model), forward selection (starting from an empty feature set and gradually adding features that contribute the most to model performance), and backward elimination (starting from a set containing all features and gradually deleting features with the least impact on model performance).

[0148] 3. Embedded Methods: Feature selection is performed during model training, embedding the importance of features into the model's training process. This application employs regularization methods (penalizing model coefficients to reduce the weights of irrelevant features to zero), decision tree algorithms (calculating feature importance scores using a decision tree model and then selecting highly important features), and model-based selection (using feature importance from the model training process to select the most relevant features).

[0149] 4. Stability Selection: This is a method that combines filtering, wrapping, and embedding to determine which features are the most stable and important by training the model on different subsets of data and features.

[0150] This application, through feature engineering, ultimately selected APSPs with an absolute Z-score ≥3 for the average APSP between HR+ / HER2- and HR- / HER2- subtypes as features to serve as the basis for subsequent model training. The Z-score is a commonly used statistical indicator for calculating sample variability.

[0151] The feature selection described in this application can be accomplished using the basic functions and features of various versions of tools as described at https: / / www.R-project.org / . Those skilled in the art should understand and fully recognize that any other suitable tools and functions can also be used to perform the feature selection described herein.

[0152] In this application, the input APSP of the scoring model is a set of continuous variables representing a sample, and the output is the discrete class label of that sample. To effectively handle the problem of continuous variable input and discrete class label output, and to obtain a scoring model with good performance, this embodiment selects a random forest method as the scoring model. The advantages of this method are fast training speed, accuracy, and strong generalization ability across different populations. The random forest method of this application is implemented through the following steps:

[0153] (1) Data splitting: The dataset is split into a training dataset (80%) and a test dataset (20%);

[0154] (2) Use trainControl to define the parameter tuning method: use repeated sampling cross-validation (method parameter set to "repeatedcv"), set the number of repetitions to 3 (repeats parameter set to 3), use 10-fold cross-validation (number parameter set to 3), specify the search method as random search (search = "random"), specify the number of iterations of random search to 30 (tuneLength = 30), generate a set of mtry and ntree combinations in each iteration, search for a total of 30 rounds, test a set of random parameter combinations in each round, and select the optimal parameters for modeling;

[0155] (3) Establish a random forest model: Use the train function and method="rf" to specify the random forest algorithm.

[0156] (4) Evaluate the model: Evaluate the model using the test dataset. Calculate metrics such as confusion matrix and accuracy.

[0157] (5) Using models: making predictions on new data.

[0158] The final random forest model (i.e., the scoring model) gives a score between 0 and 1. The closer the score is to 1, the closer the tumor profile of this type of patient is to that of HR+ / HER2- breast cancer patients, and the more suitable this type of patient is to use CDK4 / 6 inhibitors.

[0159] Conversely, if the score is closer to 0, the tumor profile of this type of patient is closer to that of HR- / HER2- breast cancer patients, and this type of patient is not suitable for CDK4 / 6 inhibitors.

[0160] Since the predictive model is based on biological characteristics (signaling pathways), this scoring model is essentially a mechanism model of how patients respond to CDK4 / 6 inhibitors.

[0161] The model training in this application was accomplished using the basic functionalities and various features of the tools described in the versions available at https: / / www.R-project.org / . Those skilled in the art should understand and fully recognize that any other suitable tools and functions can also be used to perform the model training described herein.

[0162] More specifically, to obtain a model that can characterize the underlying mechanisms and assess whether subjects are likely to be sensitive to CDK4 / 6 inhibitors, a CDK4 / 6 inhibitor sensitivity scoring model can be established based on APSP. The two steps described above can be further refined as follows:

[0163] The first step was to obtain APSP data. Tumor genomic data were collected from patients with HR-positive / HER2-negative breast cancer (positive population) and triple-negative breast cancer (negative population), respectively. The DAGM algorithm was used to convert the genomic variation information of the tumor samples into a cell signaling pathway activity profile (APSP). Table 1 below shows the tumor APSP results of a real patient. This application uses this method to extract the potential biological mechanism characteristics of the response of breast cancer patients to CDK4 / 6 inhibitors, specifically the quantitative information on the significant activation and inhibition of several signaling pathways.

[0164] Table 1: Tumor APSP in Real Patients

[0165] The second step involves using machine learning algorithms for regression modeling. By comparing the differences in APSP between the two breast cancer subtypes, a scoring model can be established to evaluate the potential sensitivity of subjects or indication groups to CDK4 / 6 inhibitors.

[0166] The model training process can be further described as follows:

[0167] 1. First, genomic variation data were obtained from the publicly available database TCGA (Cancer Genome Atlas) for two categories of breast cancer patients: HR-positive / HER2-negative (positive group, 454 cases) and triple-negative (negative group, 127 cases). Quality control was then performed on these genomic data.

[0168] 2. Next, these variation information are input into the DAGM algorithm and converted into the corresponding pathway activity profile (APSP).

[0169] 3. Next, calculate the Z-score of the APSP difference between HR-positive / HER2-negative (positive population) and triple-negative (negative population) breast cancer patients (as a feature engineering method), and select APSP with an absolute Z-score greater than or equal to 3 (i.e. not less than 3) as model features.

[0170] 4. Next, the random forest algorithm was used to train the model to distinguish between the two types of breast cancer patients. In the training set (TCGA), leave-one-out cross-validation was used, and the model AUC reached 0.7883 (as shown in Figure 4).

[0171] This embodiment primarily uses the AUC metric for evaluation. The specific definitions of the relevant metrics are as follows:

[0172] True Positive (TP)

[0173] False positive (FP)

[0174] True Negative (TN)

[0175] False negative (FN)

[0176] Where TP (True Positive) and TN (True Negative) represent the number of correctly classified positive and negative samples, FP (False Positive) and FN (False Negative) represent the number of incorrectly classified positive and negative samples, and P and N represent the number of positive and negative samples.

[0177] AUC (Area under the Receiver Operator Characteristic Curve) is a performance metric for machine learning models to distinguish between classes. It shows the degree to which the model can differentiate between classes. AUC is in the range of 0-1. The closer the AUC is to 1, the better the model's performance in distinguishing between two classes of patients (Zhou, ZH (2021). Machine learning. Springer Nature).

[0178] In addition to validation using the training set, this application also validated the established mechanism model using local patient information. Specifically, tumor genomic data (whole exome sequencing data of patient tumors) of 256 local breast cancer patients (i.e., local patient information not from the TCGA database, HR+ / HER2- (171 cases) and HR- / HER2- (85 cases)) were input into the model for scoring. The results showed that in this independent validation cohort, the model's AUC value was 0.7563 (as shown in Figure 5), which not only confirmed that the model's scoring could effectively distinguish between HR+ / HER2- (171 cases) and HR- / HER2- (85 cases) breast cancer patients, but also indicated that the model had strong generalization ability across different populations.

[0179] Next, an external patient dataset evaluating the efficacy of CDK4 / 6 inhibitors, published in the journal Cancer Discovery (Wander SA et al., Cancer Discov 2020; 10(8):1174-1193), was used to validate the model's ability to distinguish between patients' tumor sensitivity and resistance to CDK4 / 6 inhibitors. This dataset contained 59 breast cancer patient samples, of which 18 were sensitive to CDK4 / 6 inhibitors and 41 were resistant (28 cases of primary resistance and 13 cases of acquired resistance). The scoring model established in the previous steps was used to predict the efficacy of the drug on this dataset, achieving an AUC of 0.7737 (as shown in Figure 6), demonstrating the model's ability to differentiate patient efficacy.

[0180] Example 2

[0181] The scoring model, once trained, provides scores between 0 and 1. A score closer to 1 indicates that the patient's tumor mechanism is closer to that of HR+ / HER2- breast cancer patients (positive population), and thus CDK4 / 6 inhibitors are more suitable for this type of patient. Conversely, a score closer to 0 indicates that the patient's tumor mechanism is closer to that of triple-negative breast cancer patients (negative population), and thus CDK4 / 6 inhibitors are not suitable for this type of patient.

[0182] Chordoma is a rare type of cancer. Based on the 2021 NCCN Clinical Practice Guidelines for Bone Tumors in the United States, chordoma is classified into three histological subtypes: conventional chordoma, chondroid chordoma, and dedifferentiated chordoma. Currently, surgical treatment is the primary approach for most chordoma cases.

[0183] Treatment of chordoma remains challenging, with surgery being the primary approach, while radiotherapy and chemotherapy have limited effectiveness. Regarding drug therapy, although numerous clinical trials are underway for drugs such as cetoximab, lapatinib, and imatinib, the most advanced trials so far have only reached Phase II (according to information from clinicaltrials.gov), and no drugs have yet been successfully marketed.

[0184] This model further revealed that for the rare tumor chordoma (a cohort of 149 Chinese patients), the model score was 0.86 (compared to 0.80 for the 206 HR+ / HER2- patients in the Chinese cohort), with a median score of 0.89 (0.81 for HR- / HER2-), indicating a condition sensitive to CDK4 / 6 inhibitors. Clinical practice confirmed an ORR (objective response rate) of 33% and a DCR (disease control rate) of 83% (see Table 2 below), significantly higher than previous clinical trials for chordoma. In Table 2, PR indicates partial remission, SD indicates stable disease, and PD indicates disease progression. ORR is the proportion of PR+CR (complete remission) out of all cases, and DCR is the proportion of cases excluding PD (i.e., CR+PR+SD).

[0185] Through systematic analysis of this scoring model, we discovered that traditional chordoma is a new potential indication for CDK4 / 6 inhibitors. Clinical trials have shown partial remission or disease stabilization, providing a strong theoretical basis and practical guidance for precision treatment.

[0186] Table 2: Efficacy outcomes of CDK4 / 6 inhibitors in patients with chordoma

[0187] Therefore, for patients with chordoma, the scoring model constructed according to the method in Figure 1 can also be used, and the scoring (prediction) can be performed according to the method shown in Figure 2, to determine whether the patient is sensitive to CDK4 / 6 inhibitors, that is, whether CDK4 / 6 inhibitors can be used for treatment.

[0188] Example 3

[0189] The scoring model proposed in this application is based on the differential characteristics of tumor cell signaling pathway activity profiles. Therefore, it is essentially a mechanistic model that can not only be used to predict the potential efficacy of CDK4 / 6 inhibitors in patients, but also help to deepen the understanding of the differences in the molecular mechanisms by which CDK4 / 6 inhibitors act on different breast cancer subtypes. As shown in Figure 7, the key features in the model are mainly concentrated in six biological functions, including cell cycle management, immune response, cell metabolism, cell growth and differentiation, apoptosis and programmed cell death, and stress response.

[0190] The applicant discovered that, in addition to the functions of CDK4 / 6 itself in the cell cycle, sensitivity to CDK4 / 6 inhibitors is associated with the tumor's immune response. Tumors resistant to CDK4 / 6 inhibitors (HR- / HER2- breast cancer) exhibit immunosuppressive characteristics, such as activation of the strongly immunosuppressive ERBB2 / ERBB3 pathway. Tumors sensitive to CDK4 / 6 inhibitors (HR+ / HER2- breast cancer) exhibit a stronger immune response, with significant activation of pathways such as IL-1 / LBTR observed.

[0191] In step S130 of the model construction method 100 in Figure 1, signal features need to be selected from the APSP as model features. Similarly, in step S220 of the scoring method 200 in Figure 2, signal features also need to be selected from the signal pathway activity spectrum as model features. Here, based on the signal features selected for model training in step S120 of the CDK4 / 6 inhibitor sensitivity scoring model construction process, corresponding signal features are selected from the signal pathway activity spectrum as model features in step S220.

[0192] Therefore, based on the above discussion, as an alternative to feature engineering, signal features related to cell cycle management and immune response can be directly selected from APSP as model features.

[0193] Example 4

[0194] Furthermore, based on various experiments conducted by the applicant, the scoring model proposed in this application can be widely used for patients with various types of tumors. In addition to breast cancer and chordoma mentioned in the above examples, the scoring model has shown that patients with certain tumor types and subtypes (classifications) are sensitive to CDK4 / 6 inhibitors, thus achieving good therapeutic effects with CDK4 / 6 inhibitors. These tumor types and subtypes (classifications) include at least those mentioned in the following examples. However, those skilled in the art should understand and fully recognize that not only can the scoring model of this application be widely applied to patient populations with more tumor types and subtypes not mentioned, but it can also be used to discover more unmentioned tumor types and subtypes as indications for CDK4 / 6 inhibitors.

[0195] The following example illustrates how scoring models can be used to discover these new indications.

[0196] As described above, methods for predicting patient response to CDK4 / 6 inhibitors, determining whether a patient should be treated with a CDK4 / 6 inhibitor, and selecting patient populations that may respond to CDK4 / 6 inhibitor treatment can be performed as described above. In all these methods, the patient's sensitivity to CDK4 / 6 inhibitors is first scored using the scoring model proposed in this application. A threshold can be set for the sensitivity score to evaluate the sensitivity. For example, the threshold can be 0.55, 0.60, 0.65, 0.70, 0.75, 0.8, 0.85, or 0.9, preferably 0.70 or 0.75. If the sensitivity score exceeds this threshold, it can be determined that the patient has responded to CDK4 / 6 inhibitor treatment, that the patient can be treated with a CDK4 / 6 inhibitor, or that the patient population (tumor type or subtype) may respond to CDK4 / 6 inhibitor treatment.

[0197] This application analyzed 9616 patients across 32 cancer types from the TCGA database and a local database using the scoring model provided above and the more specific DAGM algorithm. The results showed significant differences in the average scores among the various cancer types. For example, lung squamous cell carcinoma (LUSC) had the lowest average score (solid line), and less than 10% of the patient population had a score above the benefit threshold (0.75, dashed line).

[0198] Numerous trial results regarding inhibitors demonstrate that the scoring model in this application can indeed distinguish between indications where high and low benefits can be obtained from CDK4 / 6 inhibitors. Furthermore, the proportion of patients benefiting from CDK4 / 6 inhibitors varies significantly across different indications. Therefore, further analysis is needed to determine whether, similar to breast cancer, different subtypes of patients exhibit varying responses to CDK4 / 6 inhibitors within each indication, in order to specifically identify potential new indications (subtypes) for CDK4 / 6 inhibitors.

[0199] To select clinically feasible new indications for CDK4 / 6 inhibitors from various indications, this application categorized 9316 patients from 32 cancer types in the TCGA database and local databases into 466 subtypes based on primary tumor pathological diagnosis (according to the World Health Organization (WHO) International Classification of Cancer (ICD-O)) and AJCC pathological staging (according to the 6th and 7th editions of the AJCC manual). The tumor genomic data of these patients were then analyzed and scored using this model, and new indications for CDK4 / 6 inhibitors were screened based on the following steps:

[0200] Step 1: Calculate the average score of patients with each type of cancer. If the score is significantly higher than the average score of cancer patients randomly sampled from the TCGA database and the local database (Z score ≥ 3), the cancer type is selected as a candidate cancer type.

[0201] Step 2: Based on the pathological diagnosis of the primary tumor and the AJCC pathological staging, the candidate cancer types are divided into different subtypes;

[0202] Step 3: Calculate the average patient score for each subtype;

[0203] Step 4: If the subtype meets any of the following inclusion criteria (1), (2), and (3), then the subtype is listed as a new indication for CDK4 / 6 inhibitors:

[0204] Inclusion criteria (1): The average score of patients in this subtype is ≥0.75;

[0205] Inclusion criteria (2): The average score of patients in this subtype is <0.75, but more than 25% of patients in this subtype have a score ≥0.75;

[0206] Based on the model and screening steps described above, this application selects new indications.

[0207] To examine the model's effectiveness in screening patients for CDK4 / 6 inhibitors, this application selected various relevant trials with published results and scored the corresponding enrolled patients using the model.

[0208] It should be noted that the approved indications for various marketed CDK4 / 6 inhibitors are all concentrated in HR+ / HER2- breast cancer, and different combination regimens have been used in approved clinical trials. This indicates that these drug molecules, which target CDK4 / 6 and exert inhibitory effects, share a certain response mechanism, and this mechanism exists in HR+ / HER2- breast cancer. The model in this application is constructed using HR+ / HER2- breast cancer as a positive control, with the aim of extracting this shared response mechanism. Therefore, the model provided in this application is applicable for prediction of drug molecules with CDK4 / 6 inhibitory effects. If the CDK4 / 6 inhibitor is effective in the trial patients, all groups should show a certain degree of significant benefit; however, different groups may show slightly different benefits due to different combination regimens.

[0209] The embodiments shown above have demonstrated that the model of this application can indicate whether a specific patient population can benefit from CDK4 / 6 inhibitors: on the one hand, patients with indications who have high model scores can indeed benefit from CDK4 / 6 inhibitors in clinical practice; on the other hand, patients with indications who have low model scores cannot obtain clinical utility from CDK4 / 6 inhibitors.

[0210] For example, scoring models can be used to predict the sensitivity of patients with renal papillary carcinoma to CDK4 / 6 inhibitors. Clinical trial data have shown that birocilib, lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on renal papillary carcinoma, with patients achieving pathological complete response (pCR), complete response (CR), stable disease (SD), or partial response (PR) in response to the drugs.

[0211] For example, scoring models can be used to predict the sensitivity of patients with renal chromophobe carcinoma to CDK4 / 6 inhibitors. Clinical trial data have shown that birocilib, lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on renal chromophobe carcinoma, with patients achieving pathological complete remission (pCR), complete remission (CR), stable disease (SD), or partial remission (PR) in response to the drugs.

[0212] For example, scoring models can be used to predict the sensitivity of fibromyxoid sarcoma patients to CDK4 / 6 inhibitors. Clinical trial data have shown that Birocilib, Lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on fibromyxoid sarcoma, with patients achieving pathological complete remission (pCR), complete remission (CR), stable disease (SD), or partial remission (PR) in response to the drug.

[0213] For example, scoring models can be used to predict the sensitivity of leiomyosarcoma patients to CDK4 / 6 inhibitors. Clinical trial data have shown that Birocilib, Lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on leiomyosarcoma, with patients achieving pathological complete remission (pCR), complete remission (CR), stable disease (SD), or partial remission (PR) in response to the drug.

[0214] For example, scoring models can be used to predict the sensitivity of patients with malignant fibrous histiocytoma to CDK4 / 6 inhibitors. Clinical trial data have shown that Birocilib, Lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on malignant fibrous histiocytoma, with patients achieving pathological complete remission (pCR), complete remission (CR), stable disease (SD), or partial remission (PR) in response to the drug.

[0215] For example, scoring models can be used to predict the sensitivity of patients with malignant peripheral schwannomas to CDK4 / 6 inhibitors. Clinical trial data have shown that Birocilib, Lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on malignant peripheral schwannomas, with patients achieving pathological complete remission (pCR), complete remission (CR), stable disease (SD), or partial remission (PR) in response to the drugs.

[0216] For example, scoring models can be used to predict the sensitivity of synovial sarcoma patients to CDK4 / 6 inhibitors. Clinical trial data have shown that Birocilib, Lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on synovial sarcoma, with patients achieving pathological complete remission (pCR), complete remission (CR), stable disease (SD), or partial remission (PR) in response to the drug.

[0217] For example, scoring models can be used to predict the sensitivity of undifferentiated sarcoma patients to CDK4 / 6 inhibitors. Clinical trial data have shown that Birocilib, Lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on undifferentiated sarcoma, with patients achieving pathological complete remission (pCR), complete remission (CR), stable disease (SD), or partial remission (PR) in response to the drug.

[0218] For example, scoring models can be used to predict the sensitivity of patients with malignant epithelioid mesothelioma to CDK4 / 6 inhibitors. Clinical trial data have shown that Birocilib, Lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on malignant epithelioid mesothelioma, with patients achieving pathological complete remission (pCR), complete remission (CR), stable disease (SD), or partial remission (PR) in response to the drugs.

[0219] For example, scoring models can be used to predict the sensitivity of patients with malignant bipolar mesothelioma to CDK4 / 6 inhibitors. Clinical trial data have shown that Birocilib, Lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on malignant bipolar mesothelioma, with patients achieving pathological complete response (pCR), complete response (CR), stable disease (SD), or partial response (PR) in response to the drug.

[0220] For example, scoring models can be used to predict the sensitivity of testicular embryonic carcinoma patients to CDK4 / 6 inhibitors. Clinical trial data have shown that Birocilib, Lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on testicular embryonic carcinoma, with patients achieving pathological complete remission (pCR), complete remission (CR), stable disease (SD), or partial remission (PR) in response to the drugs.

[0221] For example, scoring models can be used to predict the sensitivity of patients with mixed testicular germ cell tumors to CDK4 / 6 inhibitors. Clinical trial data have shown that Birocilib, Lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on mixed testicular germ cell tumors, with patients achieving pathological complete remission (pCR), complete remission (CR), stable disease (SD), or partial remission (PR) in response to the drugs.

[0222] For example, scoring models can be used to predict the sensitivity of testicular seminoma patients to CDK4 / 6 inhibitors. Clinical trial data have shown that Birocilib, Lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on testicular seminoma, with patients achieving pathological complete remission (pCR), complete remission (CR), stable disease (SD), or partial remission (PR) in response to the drugs.

[0223] For example, scoring models can be used to predict the sensitivity of adrenocortical carcinoma patients to CDK4 / 6 inhibitors. Clinical trial data have shown that Birocilib, Lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on adrenocortical carcinoma, with patients achieving pathological complete response (pCR), complete response (CR), stable disease (SD), or partial response (PR) in response to the drugs.

[0224] For example, scoring models can be used to predict the sensitivity of patients with epithelioid uveal melanoma to CDK4 / 6 inhibitors. Clinical trial data have shown that Birocilib, Lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on epithelioid uveal melanoma, with patients achieving pathological complete remission (pCR), complete remission (CR), stable disease (SD), or partial remission (PR) in response to the drugs.

[0225] For example, scoring models can be used to predict the sensitivity of patients with mixed epithelioid and spindle cell uveal melanoma to CDK4 / 6 inhibitors. Clinical trial data have shown that Birocilib, Lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on mixed epithelioid and spindle cell uveal melanoma, with patients achieving pathological complete remission (pCR), complete remission (CR), stable disease (SD), or partial remission (PR) in response to the drug.

[0226] For example, scoring models can be used to predict the sensitivity of spindle cell uveal melanoma patients to CDK4 / 6 inhibitors. Clinical trial data have shown that Birocilib, Lerociclib, FCN-437c, TY-302, compounds of formula I, II, and III have good therapeutic effects on spindle cell uveal melanoma, with patients achieving pathological complete remission (pCR), complete remission (CR), stable disease (SD), or partial remission (PR) in response to the drug.

[0227] In contrast, the scores of various tumor types and subtypes, such as gastric adenocarcinoma subtype of gastric cancer, esophageal squamous cell carcinoma subtype of esophageal cancer, non-small cell lung adenocarcinoma subtype of lung cancer, Ewing sarcoma subtype of sarcoma, and hepatocellular carcinoma, did not reach the sensitivity threshold in the scoring model, and the experimental data showed that known CDK4 / 6 inhibitors did not have good therapeutic effects on these types.

[0228] Rating system

[0229] Those skilled in the art will understand that, based on the method for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model according to this application, a system for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model can be developed. Figure 8 shows a schematic block diagram of a system for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model according to an embodiment of this application. Specifically, the system 800 for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model includes a data acquisition module 810, a model feature generation module 820, and a CDK4 / 6 inhibitor sensitivity scoring model 830.

[0230] The data acquisition module 810 is used to acquire patient genomic data. The patient genomic data mentioned here may include the patient's germline genomic data (i.e., normal genomic data) and tumor genomic data, or it may directly include the patient's genomic variation data.

[0231] The model feature generation module 820 is used to generate a signaling pathway activity profile of the patient's tumor cells based on the patient's genomic data, and select signal features from the signaling pathway activity profile as model features.

[0232] The CDK4 / 6 inhibitor sensitivity scoring model 830 was obtained according to the method for constructing a CDK4 / 6 inhibitor sensitivity scoring model of this application (e.g., Figure 1). The scoring model 830 is used to output a patient's sensitivity score to CDK4 / 6 inhibitors based on model features.

[0233] It should be noted here that the model feature generation module 820 selects the corresponding signal features from the signal pathway activity spectrum as model features based on the signal features selected for model training during the construction of the CDK4 / 6 inhibitor sensitivity scoring model 830.

[0234] Those skilled in the art will understand that the system described above for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model can be implemented by a computer. For example, a computer program can be executed to perform the following operations: acquire patient genomic data; generate a signaling pathway activity profile of the patient's tumor cells based on the patient's genomic data, select signal features from the signaling pathway activity profile as model features; and use a CDK4 / 6 inhibitor sensitivity scoring model to output a sensitivity score for the patient to CDK4 / 6 inhibitors based on the model features.

[0235] In other words, although the various operational functions corresponding to the scoring method are described as modules or sub-modules in the system for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model according to this application, those skilled in the art should understand that such modules or sub-modules may be circuit components or other physical components, or they may not be composed of circuit components or other physical components, but rather functional modules constructed by computer programs.

[0236] Examples related to computer programs

[0237] Furthermore, those skilled in the art should recognize that all methods of this application can be implemented as computer programs. As described above in conjunction with the accompanying drawings, the methods of the above embodiments are executed by one or more programs, the instructions in which cause a computer or processor to execute the algorithms described in conjunction with the drawings. These programs can be stored and provided to a computer or processor using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (such as floppy disks, magnetic tapes, and hard disk drives), magneto-optical recording media (such as magneto-optical disks), CD-ROMs (Compact Disc Read-Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (such as ROMs, PROMs (Programmable ROMs), EPROMs (Erasable and Writable PROMs), flash memory ROMs, and RAMs (Random Access Memory)). Further, these programs can be provided to a computer using various types of transient computer-readable media. Examples of transient computer-readable media include electrical signals, optical signals, and electromagnetic waves. Transient computer-readable media can be used to provide programs to a computer via wired or wireless communication paths such as wires and optical fibers.

[0238] For example, according to one embodiment of this application, a non-transitory computer-readable storage medium may be provided for storing a computer program, the computer program including instructions that, when executed by a processor of an electronic device, cause the electronic device to perform at least one of the following methods: a method for constructing a CDK4 / 6 inhibitor sensitivity scoring model as described above; a method for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model as described above; a method for constructing a CDK4 / 6 inhibitor sensitivity scoring model for breast cancer patients as described above; a method for scoring the sensitivity of breast cancer patients to CDK4 / 6 inhibitors using an artificial intelligence model as described above; a method for predicting a patient's treatment response to a CDK4 / 6 inhibitor as described above; a method for determining whether a patient should be treated with a CDK4 / 6 inhibitor as described above; or a method for selecting a patient population that may respond to CDK4 / 6 inhibitor treatment as described above.

[0239] Additionally, according to the disclosure of this application, a computer system can also be provided, comprising: a processor; a memory; and a computer program. The computer program is stored in the memory and configured to be executed by the processor. The computer program includes instructions for implementing at least one of the following methods: a method for constructing a CDK4 / 6 inhibitor sensitivity scoring model as described above; a method for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model as described above; a method for constructing a CDK4 / 6 inhibitor sensitivity scoring model for breast cancer patients as described above; a method for scoring the sensitivity of breast cancer patients to CDK4 / 6 inhibitors using an artificial intelligence model as described above; a method for predicting a patient's treatment response to CDK4 / 6 inhibitors as described above; a method for determining whether a patient should be treated with a CDK4 / 6 inhibitor as described above; or a method for selecting a patient population that may respond to CDK4 / 6 inhibitor treatment as described above.

[0240] This application relates to a CDK4 / 6 inhibitor sensitivity scoring model. A method for constructing a CDK4 / 6 inhibitor sensitivity scoring model is provided. The constructed CDK4 / 6 inhibitor sensitivity scoring model is an artificial intelligence scoring model. Using this model, the sensitivity of cancer patients to CDK4 / 6 inhibitors can be scored. The initial training set used for this model came from samples of CDK4 / 6 inhibitor sensitivity in breast cancer patients, and it has been validated for use in scoring the sensitivity of breast cancer patients to CDK4 / 6 inhibitors. As this model is extended to score the sensitivity of patient groups with more tumor types and subtypes to CDK4 / 6 inhibitors, new indications for CDK4 / 6 inhibitors can be expanded. For example, it can predict patient response to CDK4 / 6 inhibitor treatment, determine whether to treat patients with CDK4 / 6 inhibitors, and select patient groups that are likely to respond to CDK4 / 6 inhibitor treatment.

[0241] Therefore, the technical solution provided in this application includes a method for expanding new indications for drugs, particularly a method for expanding new indications for CDK4 / 6 inhibitors. This method includes establishing a CDK4 / 6 inhibitor response mechanism model and using this model to screen for new indications. The establishment of the CDK4 / 6 inhibitor response mechanism model includes model training and validation, which includes model preparation and foundation, modeling and model prediction, and model validation. The screening for new indications includes the selection of thresholds for benefiting patients and pan-cancer analysis and screening.

[0242] According to this application, in addition to breast cancer and chordoma, CDK4 / 6 inhibitors can be predicted to have good inhibitory effects in the treatment of intrahepatic cholangiocarcinoma, invasive fibromatosis, renal clear cell adenocarcinoma, renal papillary carcinoma, malignant epithelioid mesothelioma, malignant mesothelioma, malignant biphasic mesothelioma, cervical adenocarcinoma, cervical squamous cell carcinoma, cervical endometrioid adenocarcinoma, astrocytoma, spindle cell uveal melanoma, and other diseases.

[0243] The implementation methods of this application are not limited to those described in the above embodiments. Without departing from the spirit and scope of this application, those skilled in the art can make various changes and improvements to this application in form and detail, and these are all considered to fall within the protection scope of this application.

Claims

1. A method for building a CDK4 / 6 inhibitor sensitivity score model, characterized in that, The method includes: Obtain a training set, which includes multiple training samples, each of which includes patient genomic data and a label of the patient's sensitivity to CDK4 / 6 inhibitors; Based on the patient's genomic data in each training sample, a spectrum of signaling pathway activity in the patient's tumor cells is generated. Signal features are selected as model features from the activity spectrum of the signaling pathway; Based on the aforementioned model features and the aforementioned labels for CDK4 / 6 inhibitor sensitivity, an artificial intelligence algorithm is used to train the model and obtain a CDK4 / 6 inhibitor sensitivity scoring model.

2. The method according to claim 1, characterized in that, The generation of a signaling pathway activity profile of a patient's tumor cells based on the patient's genomic data in each training sample includes: Genomic variation information is extracted based on the differences between germline genome data and tumor genome data of patients in each training sample; Transform genomic variation information into corresponding signaling pathway activity profiles.

3. The method according to claim 2, characterized in that, The extraction of genomic variation information based on the differences between germline genomic data and tumor genomic data of patients in each training sample includes: performing quality control and variation detection on the germline genomic data and the tumor genomic data to obtain variation sites.

4. The method according to claim 2, characterized in that, The process of converting genomic variation information into a corresponding signaling pathway activity profile includes: calculating the genomic variation information using the DAGM algorithm to obtain the signaling pathway activity profile.

5. The method according to claim 1, characterized in that, The method of selecting signal features from the signal pathway activity spectrum as model features includes: performing feature engineering on the signal pathway activity spectrum to select signal features from the signal pathway activity spectrum as model features.

6. The method according to claim 5, characterized in that, The selection of signal features from the activity spectrum of the signal pathway as model features includes: Calculate the Z-score of each signal feature in the activity spectrum of the signal pathway; Signal features with an absolute Z-score of not less than 3 were selected as model features.

7. The method according to claim 1, characterized in that, The selection of signal features from the signaling pathway activity spectrum as model features includes: selecting signal features related to cell cycle management and immune response from the signaling pathway activity spectrum as model features.

8. The method according to claim 1, characterized in that, The method of obtaining a CDK4 / 6 inhibitor sensitivity scoring model by training the model using an artificial intelligence algorithm based on the aforementioned model features and the aforementioned labels for CDK4 / 6 inhibitor sensitivity further includes: The aforementioned model features are input into the AI ​​model being trained to obtain probabilistic features regarding the sensitivity to CDK4 / 6 inhibitors. The classification results regarding sensitivity are obtained based on the probabilistic characteristics of CDK4 / 6 inhibitor sensitivity described above; Using the aforementioned label of sensitivity to CDK4 / 6 inhibitors as the output target, the parameters of the artificial intelligence model are adjusted so that the obtained classification results regarding sensitivity are consistent with the output target.

9. The method according to claim 1, characterized in that, The artificial intelligence algorithms include machine learning algorithms and deep learning algorithms.

10. The method of claim 9, wherein, The machine learning algorithm is the random forest algorithm.

11. A method for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model, characterized in that, The method includes: Based on the patient's genomic data, a spectrum of signaling pathway activity in the patient's tumor cells was generated. Signal features are selected as model features from the activity spectrum of the signaling pathway; The model features are input into the CDK4 / 6 inhibitor sensitivity scoring model obtained by the method according to any one of claims 1-10, and the sensitivity score of the patient to the CDK4 / 6 inhibitor is output. Specifically, based on the signal features selected for model training during the construction of the CDK4 / 6 inhibitor sensitivity scoring model, corresponding signal features are selected from the signal pathway activity spectrum as model features.

12. The method according to claim 11, characterized in that, The generation of a signaling pathway activity profile of the patient's tumor cells based on the patient's genomic data includes: Genomic variation information is extracted based on the differences between the patient's germline genome data and tumor genome data; Transform genomic variation information into corresponding signaling pathway activity profiles.

13. The method of claim 12, wherein, The extraction of genomic variation information based on the differences between the patient's germline genome data and tumor genome data includes: performing quality control and variation detection on the germline genome data and the tumor genome data to obtain variation sites.

14. The method of claim 12, wherein, The process of converting genomic variation information into a corresponding signaling pathway activity profile includes: calculating the genomic variation information using the DAGM algorithm to obtain the signaling pathway activity profile.

15. The method of claim 11, wherein, The sensitivity score of a patient to a CDK4 / 6 inhibitor, output by the CDK4 / 6 inhibitor sensitivity scoring model, is a probability score. The method further includes: making a judgment on the patient's sensitivity to CDK4 / 6 inhibitors based on the sensitivity score.

16. The method of claim 11, wherein, The artificial intelligence model includes machine learning models and deep learning models.

17. The method of claim 16, wherein, The machine learning model is a random forest model.

18. A method for constructing a breast cancer patient sensitivity score model to CDK4 / 6 inhibitors, characterized in that, The method includes: Obtain a training set, which includes multiple training samples, each of which includes genomic data of a breast cancer patient and a label of the patient's sensitivity to CDK4 / 6 inhibitors; Based on the patient's genomic data in each training sample, a spectrum of signaling pathway activity in the patient's tumor cells is generated. Signal features are selected as model features from the activity spectrum of the signaling pathway; Based on the aforementioned model features and the aforementioned labels on sensitivity to CDK4 / 6 inhibitors, an artificial intelligence algorithm is used to train the model and obtain a CDK4 / 6 inhibitor sensitivity scoring model for breast cancer patients.

19. The method of claim 18, wherein, The multiple training samples correspond to two breast cancer subtypes respectively: HR+ / HER2- HR- / HER2-.

20. A method of scoring sensitivity of a breast cancer patient to a CDK4 / 6 inhibitor using an artificial intelligence model, characterized in that, The method includes: Based on the genomic data of a breast cancer patient, a spectrum of signaling pathway activity in the patient's tumor cells was generated. Signal features are selected as model features from the activity spectrum of the signaling pathway; The model features are input into the CDK4 / 6 inhibitor sensitivity scoring model for breast cancer patients obtained according to the method described in claim 18 or 19, and the sensitivity score is output. Specifically, based on the signal features selected for model training during the construction of the CDK4 / 6 inhibitor sensitivity scoring model for breast cancer patients, corresponding signal features are selected from the signal pathway activity spectrum as model features.

21. A method of predicting a patient's therapeutic response to a CDK4 / 6 inhibitor, characterized in that, The method includes: The patient's sensitivity score to CDK4 / 6 inhibitors is obtained by the method of scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model as described in any one of claims 11-17. Based on the obtained sensitivity scores, the patient's response to CDK4 / 6 inhibitors can be predicted.

22. The method according to claim 21, characterized in that, If the sensitivity score is greater than 0.55, greater than 0.60, greater than 0.65, greater than 0.70, greater than 0.75, greater than 0.8, greater than 0.85, or greater than 0.9, then the patient is predicted to respond to treatment with a CDK4 / 6 inhibitor; preferably, if the sensitivity score is greater than 0.70 or greater than 0.75, then the patient is predicted to respond to treatment with a CDK4 / 6 inhibitor.

23. A method for determining whether a patient should be treated with a CDK4 / 6 inhibitor, characterized in that, The method includes: The patient's sensitivity score to CDK4 / 6 inhibitors is obtained by the method of scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model as described in any one of claims 11-17. Based on the obtained sensitivity score, it is determined whether to treat the patient with a CDK4 / 6 inhibitor.

24. The method according to claim 23, characterized in that, If the sensitivity score is greater than 0.55, greater than 0.60, greater than 0.65, greater than 0.70, greater than 0.75, greater than 0.8, or greater than 0.85, or greater than 0.9, the patient will receive treatment with a CDK4 / 6 inhibitor; preferably, if the sensitivity score is greater than 0.70 or greater than 0.75, the patient will receive treatment with a CDK4 / 6 inhibitor.

25. A method of selecting a patient population likely to respond to CDK4 / 6 inhibitor therapy, comprising: The method includes: The sensitivity score of the patient population to CDK4 / 6 inhibitors is obtained by the method of scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model as described in any one of claims 11-17. If the obtained sensitivity score is higher than the threshold, then the patient population is selected to receive CDK4 / 6 inhibitors.

26. The method according to claim 25, characterized in that, The threshold is 0.65, 0.70, 0.75, 0.8, 0.85, or 0.9; preferably, the threshold is 0.70 or 0.

75.

27. The method of claim 25, wherein, The patient population refers to those classified into one or more cancer subtypes.

28. A system for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model, characterized in that, The system includes: The data acquisition module is used to acquire the patient's tumor genomic data; The model feature generation module is used to generate a signaling pathway activity profile of the patient's tumor cells based on the patient's genomic data, and select signal features from the signaling pathway activity profile as model features; The CDK4 / 6 inhibitor sensitivity scoring model obtained by the method according to any one of claims 1-10 is used to output a sensitivity score of the patient to CDK4 / 6 inhibitors based on the model features. The model feature generation module is further configured to select corresponding signal features from the signal pathway activity spectrum as model features based on the signal features selected for model training during the construction of the CDK4 / 6 inhibitor sensitivity scoring model.

29. A non-transitory computer-readable storage medium for storing a computer program, the computer program comprising instructions that, when executed by a processor of an electronic device, cause the electronic device to perform at least one of the following methods: The method for constructing a CDK4 / 6 inhibitor sensitivity scoring model as described in any one of claims 1-10; The method for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model as described in any one of claims 11-17; The method for constructing a sensitivity scoring model for CDK4 / 6 inhibitors in breast cancer patients as described in claim 18 or 19; The method for scoring the sensitivity of breast cancer patients to CDK4 / 6 inhibitors using an artificial intelligence model as described in claim 20; The method for predicting a patient's response to a CDK4 / 6 inhibitor as described in claim 21 or 22; The method for determining whether a patient should be treated with a CDK4 / 6 inhibitor as described in claim 23 or 24; or The method of selecting a patient population that may respond to CDK4 / 6 inhibitor treatment as described in any one of claims 25-27.

30. A computer system, the computer system comprising: processor; Memory; and A computer program, wherein the computer program is stored in the memory and configured to be executed by the processor, the computer program including instructions for implementing at least one of the following methods: The method for constructing a CDK4 / 6 inhibitor sensitivity scoring model as described in any one of claims 1-10; The method for scoring the sensitivity of CDK4 / 6 inhibitors using an artificial intelligence model as described in any one of claims 11-17; The method for constructing a sensitivity scoring model for CDK4 / 6 inhibitors in breast cancer patients as described in claim 18 or 19; The method for scoring the sensitivity of breast cancer patients to CDK4 / 6 inhibitors using an artificial intelligence model as described in claim 20; The method for predicting a patient's response to a CDK4 / 6 inhibitor as described in claim 21 or 22; The method for determining whether a patient should be treated with a CDK4 / 6 inhibitor as described in claim 23 or 24; or The method of selecting a patient population that may respond to CDK4 / 6 inhibitor treatment as described in any one of claims 25-27.