Method for selection of antitumor therapy for breast cancer using genomic data
Molecular genetic clustering with a logistic regression model improves breast cancer treatment prediction by considering pathway activation and resistance, enhancing treatment efficacy and alignment with clinical standards.
Patent Information
- Authority / Receiving Office
- RU · RU
- Patent Type
- Patents
- Current Assignee / Owner
- OBSHCHESTVO S OGRANICHENNOI OTVETSTVENNOSTIU ONKO ANALITIKA
- Filing Date
- 2025-08-29
- Publication Date
- 2026-06-29
AI Technical Summary
Current methods for predicting breast cancer treatment efficacy are limited by focusing on driver mutations without considering pathway activation, lack of resistance prediction, and inconsistency with modern classifications, leading to suboptimal therapeutic strategies.
A method using molecular genetic clustering based on gene mutations, expression, and copy number variations, combined with a pre-trained logistic regression model, to identify four stable clusters that predict therapeutic response and resistance, aligning with clinical guidelines.
Enhances treatment selection accuracy by predicting response to various therapies, including targeted, hormonal, and immune treatments, while identifying potential resistance, and aligning with modern breast cancer classifications.
Smart Images

Figure 00000005 
Figure 00000006 
Figure 00000013
Abstract
Description
[0001] The invention relates to the field of medical diagnostics and oncology, in particular to methods for personalized selection of antitumor therapy based on molecular genetic analysis of tumor tissue samples from patients with breast cancer.
[0002] Breast cancer is a heterogeneous group of malignant neoplasms that arise from the accumulation of genetic damage in breast cells, leading to loss of cell cycle control, unrestrained proliferation, and the acquisition of the ability to invade and metastasize. This pathology is among the leading causes of cancer-related mortality among women both in our country and globally.
[0003] Currently, although tools exist in clinical practice for predicting a specific patient's sensitivity to various types of anticancer treatment based on the individual molecular profile of their disease, they are imperfect. Patients are screened for a driver mutation and, for example, a targeted drug is prescribed based on it, but the activation of an entire signaling pathway, which can also be targeted by a drug in the absence of a driver mutation, is not taken into account. This results in the vast majority of patients receiving therapy based on traditional clinical and pathological characteristics: stage of the disease, histological type of tumor, degree of differentiation, lymph node status, and expression of individual biomarkers. This approach often fails to achieve an optimal therapeutic effect, leading to treatment failure and disease progression.
[0004] The development of personalized breast cancer therapy strategies based on a comprehensive analysis of the molecular biological characteristics of a specific patient's tumor is one of the most pressing challenges in modern oncology, which this invention is designed to address.
[0005] An analysis of existing solutions revealed that most commercial products, such as OncoBox, FoundationOne, Genomed, and Genetico, focus on analyzing a limited number of genetic markers or driver mutations, often without considering the functional activity of genes and systemic interactions. Furthermore, driver mutations are detected in only 20% of patients. These companies' approaches are primarily aimed at targeted therapy, relying on a "one mutation, one targeted drug" approach, and are limited in predicting response to hormonal and cytostatic treatment. However, the fundamental problem is the inconsistency of this approach: the absence of a given mutation coupled with the presence of resistance does not mean that the use of a targeted drug is inappropriate. Resistance may be caused by the activation of other elements of a signaling pathway or multiple pathways, which can also lead to the development of resistance.
[0006] Furthermore, current patient clustering based on the clinical and biological characteristics of the tumor does not fully reflect all the nuances of various therapeutic approaches, generating only strict drug combinations, which does not allow for the individual characteristics of each treatment. Below, we will examine in detail the patent applications for each solution and highlight the key differences from this method.
[0007] A method and system for assessing the clinical efficacy of targeted drugs, OncoBox [RU 2741703, 2018-03-01, C12N-005 / 09, G01N-033 / 15, G16B-050 / 30, G16H-010 / 40], are known. The patent describes a method for integrating gene expression and mutational status to predict drug sensitivity. The method allows for assessing the presence and severity of a positive effect from the use of a particular drug (efficacy), and also provides information on the possibility of developing resistance to a given drug.
[0008] However, this approach does not allow for the formation of stable, therapeutically significant clusters that would correspond to clinically and biologically verified tumor groups. This limits the applicability of this technology within the framework of practical oncology standards and complicates its implementation into routine clinical practice due to its incompatibility with the current classification and the need for whole-genome or whole-exome sequencing. Furthermore, the method does not focus on a specific nosology (e.g., breast cancer only), which may reduce the accuracy of predictions: for each tumor location, the same molecular biological pathways may influence oncogenesis differently. This may create additional obstacles to selecting the optimal treatment strategy.
[0009] A known method for determining the activity of cellular signaling pathways using probabilistic modeling of target gene expression [WO / 2013 / 011479, 19.07.2012, C12Q 1 / 68, G06F 19 / 24], is based on the analysis of the activity of tumor cell signaling pathways based on the expression level of a set of key genes. The assessment is carried out using a probabilistic model, in particular, a Bayesian network, which takes into account the relationship between the level of gene expression, the activity of transcription factors and the functional state of the signaling pathway (e.g., Wnt, ER, Hedgehog, AR). For each pathway, a set of genes is defined, the expression indicators of which are used as markers of pathway activity. This approach is capable of predicting sensitivity to drugs that affect a specific pathway, and also serving as a recommendation for the appointment of targeted therapy if the activity of certain elements of the pathway is pathologically altered.
[0010] However, this approach has several limitations. Firstly, the model does not take into account drug resistance, focusing primarily on drug sensitivity. This limits its application in real-world oncology practice, where it is important to consider both the positive and negative effects of therapy. Secondly, it lacks consistency with modern classifications, such as PAM50, complicating the implementation of the technology into routine practice. It also does not allow for the creation of a molecular profile (ie, comprehensive information on expression changes or the presence of mutations) of the tumor, which could negatively impact the appropriateness of therapy selection.
[0011] Systems and methods for evaluating the effectiveness of a drug are known [US 12266426, 2018-12-03, C12Q-001 / 6886, G06F-017 / 15, G06F-017 / 16, G06F-017 / 18, G16B-040 / 20, G16B-045 / 00, G16B-050 / 20, G16B-050 / 30, G16H-050 / 20, G16H-050 / 50]. The patent describes a method for analyzing the expression of genes and proteins associated with the immune response in the context of immunotherapy of oncological diseases. Specifically, the invention utilizes molecular data to assess the activity of immune checkpoints (such as CTLA4, PD1, PDL1, LAG3, TIGIT, and others) responsible for regulating T-cell activity and forming an immunosuppressive tumor microenvironment. This method allows for determining a patient's sensitivity to immune therapy and its effectiveness; however, it is not possible to determine the effectiveness of other therapies using this method.
[0012] A known method for determining the risk of recurrence of breast cancer [RU 2626603, 31.12.2015, GO IN 33 / 574, C12Q 1 / 68], allows for predicting the risk of breast cancer recurrence based on the quantitative determination of the expression of three genes: ELOVL5, IGFBP6, TXNDC9. This application describes an approach to determining the probability of tumor recurrence of the Luminal A subtype of breast cancer, which is its serious limitation. The method does not provide for the ability to determine the effectiveness of therapy or the development of resistance, and is also not applicable to other subtypes of breast cancer and other types of oncological processes.
[0013] A system, method, and software for improving the efficacy and safety of drugs in a patient are known [US 20160132632, 2016-05-12, C12Q-001 / 68 G16B-005 / 00 G16B-050 / 00; US20170193176, 2017-07-06, G06F-019 / 00 G06N-005 / 04 /
[0014] The closest to the proposed technical solution in terms of the set of essential features and the result obtained is a system, method and software for predicting the clinical outcome of drug treatment for breast cancer in a patient [US 20160224739, 2016-08-04, G16Z-099 / 00]
[0015] By the patent applications US 20160132632, US 20170193176, and US 20160224739, a method for predicting the clinical outcome of antitumor therapy for breast cancer is described, based on the analysis of signaling pathway activity and the use of machine learning. The technology is based on the OncoFinder software tool, which calculates a Pathway Activation Score (PAS) based on gene expression data. One of the key advantages of the application is the use of a systems approach to modeling signaling pathways, which allows for consideration of not only individual mutations or gene expression levels but also their functional impact on biological processes. The ability to automate analysis and integrate with clinical data is also claimed, opening up prospects for application in personalized medicine.
[0016] However, this approach has a number of limitations.
[0017] Firstly, there is a lack of consistency with modern classifications such as RAM50, which makes it difficult to implement the technology into routine practice.
[0018] Secondly, although the system can assess sensitivity to therapy, it does not focus on identifying patients who are obviously resistant to treatment, which reduces its practical value in a clinical setting where the prognosis of both the effectiveness and ineffectiveness of drugs is important.
[0019] Third, the system requires whole genome sequencing or expression data of tens of thousands of genes, making it labor-intensive and expensive to implement, especially in resource-limited healthcare settings.
[0020] The objective of the invention is to create an effective method for selecting antitumor therapy for breast cancer using genomic and transcriptomic data, which overcomes the shortcomings of analogues.
[0021] The technical result is an increase in the accuracy of selection of antitumor therapy for breast cancer and, consequently, the effectiveness of treatment.
[0022] Increased efficiency is achieved through:
[0023] - the ability to calculate the probability of achieving a complete clinical response to tumor response to various types of treatment, including targeted, hormonal, immune and chemotherapy;
[0024] - identifying not only sensitivity to therapy, but also resistance;
[0025] - the use of genes that have high prognostic value and biological interpretability, comparability with clinical and pathomorphological data of patients.
[0026] Due to the fact that the condition of consistency with modern classifications such as RAM50, St. Galen classification of breast cancer is taken into account, ease of implementation is also achieved.
[0027] Modern approaches to personalized medicine are becoming an integral part of oncology, increasing the accuracy of predicting the effectiveness of anticancer treatments. In this regard, we have developed a cutting-edge solution that implements a molecular genetic profiling system for breast cancer patients, enabling us to calculate the probability of achieving a complete clinical response based on the tumor's response to various treatments, including targeted, hormonal, immune, and chemotherapy.
[0028] The key difference of the proposed invention is the use of a unique system of molecular genetic clustering of breast tumors based on gene mutations, which complements the modern classification of tumors and is consistent with clinical guidelines (ESMO, NCCN, PAM50, etc.).
[0029] The technology is based on the identification of four stable molecular genetic clusters (hereinafter also referred to as “cluster”).
[0030] The cluster is defined by the presence and combination of mutations in a set of genes and is described by changes in the expression and copy number of genes associated with both sensitivity and resistance to different classes of anticancer drugs. Furthermore, the genes used have high prognostic value and biological interpretability, making the model reliable and ready for use in clinical practice.
[0031] The presented method allows not only to select the most effective antitumor drugs, but also to exclude in advance therapies to which the patient may have primary resistance.
[0032] The method allows the use of a wide range of molecular genetic data obtained from tumor tissue samples on somatic mutations, variations in the number of gene copies (amplifications and deletions), and the expression level for key genes obtained by targeted high-throughput sequencing.
[0033] The method also uses information on the therapeutic outcomes of various antitumor drugs and the molecular characteristics of breast cancer subtypes.
[0034] The method uses a pre-trained logistic regression model with a weight matrix W of dimension 4×26, where 4 is the number of molecular genetic clusters, 26 is the number of genes studied, and a displacement vector of length 4, and an algorithm for determining the probability of a breast tumor belonging to one of four molecular genetic clusters.
[0035] This method can be fully automated, eliminating potential errors associated with manual interpretation of genomic data and allowing for the consideration of patient-specific changes. This method will allow for the selection of a therapeutic strategy for the patient based on the analysis of objective, individual genomic changes occurring in the tumor tissue.
[0036] The method for selecting antitumor therapy is a sequential multi-stage process that begins with obtaining the patient's biological material and ends with the determination of the therapeutic strategy.
[0037] A method for selecting antitumor therapy for breast cancer using genomic data is proposed, which includes the following steps:
[0038] Step 1: Obtaining a sample of the patient's biological material in which the tumor cell content is at least 20%, the degradation index is DIN ≥ 6.0, and the DNA concentration is at least 10 ng / μl.
[0039] Stage 2: Molecular genetic research to obtain a structured set of molecular genetic data containing a vector of binary values of gene mutations consisting of values 1 and 0, where 1 means there is a mutation, 0 means there is no mutation, including the following steps:
[0040] 1) Extraction of nucleic acids by isolating DNA from samples with DNA quality control;
[0041] 2) preparation of libraries for NGS sequencing by sequential execution:
[0042] - restoration of ends,
[0043] - adenylation,
[0044] - ligation of adapters,
[0045] - PCR amplification,
[0046] - food cleaning;
[0047] 3) High-throughput sequencing to generate genomic data array using prepared libraries for NGS sequencing by using high-throughput sequencing platforms with a coverage depth of at least 100× for the tumor sample;
[0048] 4) comprehensive bioinformatics processing of the data obtained from sequencing to extract clinically significant information about the mutational status of the tumor by sequentially performing:
[0049] - read quality control;
[0050] - alignment to the reference genome;
[0051] - remove duplicates;
[0052] - detection of somatic mutations using artifact filtering;
[0053] - filtering somatic mutations by reading depth (DP)> 5, mutant allele frequency (VAF)> 0.05;
[0054] - mutation annotations (ANNOVAR, VEP) with functional impact determination.
[0055] According to the invention, to generate an array of genomic data, targeted sequencing of a panel of genes is performed, including
[0056] A set of genes relevant for molecular genetic clustering of breast cancer tumors: PIK3CA, TP53, TTN, CDH1, MUC16, GATA3, MAP3K1, IGDCC4, PTEN, KMT2C, RYR2, SYNE1, NEB, HMCN1, RYR3, DST, SPTA1, FLG, CSMD2, DMD, CSMD3, OBSCN, MUC5B, TBX3, NCOR1,
[0057] for which the presence of mutations is determined,
[0058] and also for detailed characteristics of the cluster
[0059] set of genes: IL19, IL20, IL24, MYC, PTEN, for which copy number variation (CNV) is determined, and
[0060] set of genes: GATA3, MLPH, KRT5, KRT14, ESR1, PGT, FOX1A, ERBB2, for which the expression level is determined.
[0061] According to the invention, targeted sequencing is performed with a minimum coverage depth of at least 500×, paired reads of at least 150 bp in length.
[0062] Step 3: Determine the molecular genetic cluster to which the patient's tumor belongs using the pre-trained logistic regression model containing a 4×26 gene weight matrix W = [[3.518883, -3.68215, -3.64135, -0.0379, 0.352602, - 0.03618, 0.494444, 0.20169, 0.158118, 0.105565, -0.1241, 0.051482, -0.06702, 0.022258, 0.142287, 0.105093, 0.01345, 0.049623, -0.05263, -0.01629, 0.027151, 0.022393, -0.06247, -0.01593, 0.200114, 0.02975], [-0.05177, -3.15651, 3.689056, 0.016745, -0.41153, -0.09685, -0.21961, -0.24982, 0.088314, -0.05926, -0.04971, -0.03935, 0.006928, -0.02346, 0.070376, -0.02617, 0.040066, 0.038742, 0.148841, 0.059149, -0.00621, 0.005337, 0.062735, 0.064656, -0.08761, 0.046106], [-0.20097, 3.720687, -3.07543, -0.12832, 0.269606, 0.018014, 0.092575, 0.252909, -0.08035, -0.02177, 0.172667, -0.13865, 0.080254, -0.13274, -0.06047, 0.051832, -0.10939, 0.036242, -0.00127, 0.026676, -0.07229, -0.08424, 0.076258, -0.0006, -0.10111,-0.00236], [-3.26614, 3.117965, 3.027726, 0.149473, -0.21068, 0.115015, -0.3674, -0.20478, -0.16609, -0.02453, 0.001146, 0.126522, -0.02016, 0.133943, -0.1522, -0.13075, 0.055872, -0.12461, -0.09494, -0.06954, 0.051354, 0.056514, -0.07652, -0.04813, -0.0114, -0.0735]], where 4 is the number of molecular genetic clusters, 26 is the free intercept member of length 4 and the weights of 25 genes studied: PIK3CA, TP53, TTN, CDH1, MUC16, GATA3, MAP3K1, IGDCC4, PTEN, KMT2C, RYR2, SYNE1, NEB, HMCN1, RYR3, DST, SPTA1, FLG, CSMD2, DMD, CSMD3, OBSCN, MUC5B, TBX3, NCOR1.,
[0063] According to the invention, a vector of binary values of mutations of a set of 25 studied genes, obtained at the stage of molecular genetic research, is fed to the input of the logistic regression model, then the following is carried out:
[0064] - calculation of linear logits for each molecular genetic cluster k as the sum of the products of gene weights by the corresponding values of gene mutations plus the bias (the value of the free term intercept) for a given molecular genetic cluster;
[0065] - transformation of the obtained logits into probabilities using the softmax function, which ensures the normalization of values in the range from 0 to 1 with the sum of probabilities equal to 1, and the softmax function is calculated as P(k) = exp(S k ) / [exp(S1) + exp(S2) + exp(S3) + exp(S4)], where S k - logit of the molecular genetic cluster k, where k = 1, 2, 3, 4;
[0066] - assignment of the patient to a molecular genetic cluster with the highest probability.
[0067] Step 4: Defining a personalized therapeutic strategy and clinical recommendations for the obtained cluster.
[0068] According to the invention, the determination of a personalized therapeutic strategy and clinical recommendations for the found molecular genetic cluster is carried out using a reference information table containing for each molecular genetic cluster: data on genomic changes for a set of genes: PIK3CA, TP53, TTN, CDH1, MUC16, GATA3, MAP3K1, IGDCC4, PTEN, KMT2C, RYR2, SYNE1, NEB, HMCN1, RYR3, DST, SPTA1, FLG, CSMD2, DMD, CSMD3, OBSCN, MUC5B, TBX3, NCOR1;
[0069] Transcriptome gene change data for the gene set: IL19, IL20, IL24, MYC, PTEN, and the gene set: GATA3, MLPH, KRT5, KRT14, ESR1, PGT, FOX1A, ERBB2;
[0070] Description of associated histological and molecular genetic subtypes and clinical characteristics of breast tumor; description of personalized therapeutic strategy and clinical recommendations.
[0071] Genomic alteration data includes mutation type data such as single nucleotide polymorphisms, indels, coverage depth, and allele frequency in the tumor.
[0072] Transcriptome alteration data includes gene copy number and alteration type data for each gene in the gene set: IL19, IL20, IL24, MYC, PTEN, and expression level data for each gene in the gene set: GATA3, MLPH, KRT5, KRT14, ESR1, PGT, FOX1A, ERBB2.
[0073] The description of the associated histological and molecular genetic subtypes and clinical characteristics of breast tumors was performed for four breast tumor subtypes: luminal subtype A, a hybrid phenotype between basal-like and HER2-positive subtype, luminal subtype A with an immunohistochemical profile, and HER2-positive subtype represented by HER2-enriched or luminal B / HER2 variants.
[0074] Description of the invention
[0075] A patient's tumor tissue sample, obtained by biopsy or surgical removal, is used for the analysis. Patient tumor tissue samples must meet the following quality criteria:
[0076] • tumor cell content of at least 20%,
[0077] • absence of significant DNA degradation (DIN degradation index ≥ 6.0),
[0078] • DNA concentration is not less than 10 ng / μl.
[0079] After receiving high-quality biological material, the research proceeds to the molecular genetic stage, which includes the extraction of nucleic acids and their preparation for high-throughput sequencing.
[0080] DNA extraction from samples is performed using standard extraction methods such as phenol-chloroform extraction or commercial nucleic acid extraction kits.
[0081] DNA quality control is performed using agarose gel electrophoresis or microfluidic analyzers (e.g., Bioanalyzer, TapeStation). The resulting high-quality DNA is then subjected to specialized processing to generate libraries optimized for sequencing target genomic regions. Library preparation includes end repair, adenylation, adapter ligation, PCR amplification, and product purification.
[0082] Library preparation is carried out for targeted sequencing.
[0083] The prepared libraries are then submitted to high-throughput sequencing to generate the genomic dataset required for subsequent analysis. Sequencing is performed on high-throughput sequencing platforms (e.g., Illumina NovaSeq, HiSeq, or MiSeq) with a coverage depth of at least 100× for the tumor sample.
[0084] The proposed method examines two types of genetic markers with different functions.
[0085] The first set of genes included in the diagnostic panel is analyzed for mutations - this data is used to determine the tumor's belonging to one of four molecular genetic clusters. The resulting mutational profile serves as the basis for classification. The set of genes analyzed for mutations (hereinafter Set-1): PIK3CA, TP53, TTN, CDH1, MUC16, GATA3, MAP3K1, IGDCC4, PTEN, KMT2C, RYR2, SYNE1, NEB, HMCN1, RYR3, DST, SPTA1, FLG, CSMD2, DMD, CSMD3, OBSCN, MUC5B, TBX3, NCOR1.
[0086] The second and third gene sets, also included in the diagnostic panel, are used to further characterize the identified cluster. These genes are analyzed for copy number variations (CNVs) and gene expression levels. This expanded analysis allows for a more complete molecular picture of the tumor within a specific cluster. The gene set analyzed for copy number variations (hereinafter referred to as Set 2) includes IL19, IL20, IL24, MYC, and PTEN. The gene set analyzed for expression levels (hereinafter referred to as Set 3) includes GATA3, MLPH, KRT5, KRT14, ESR1, PGT, FOX1A, and ERBB2.
[0087] This additional information has important practical implications, as it helps establish correlations between molecular features and the clinical and morphological characteristics of the tumor, providing a deeper understanding of the biology of a particular case. Furthermore, this approach significantly expands the range of potential therapeutic options, as it allows for the identification of additional targets for targeted therapy that may be effective in a given molecular subtype of breast cancer.
[0088] Thus, a diagnostic panel including three sets of genes is used for targeted sequencing.
[0089] Targeted sequencing is performed with a minimum coverage depth of at least 500×, paired-end reads of at least 150 bp in length.
[0090] The output from the sequencer is three files:
[0091] 1) Copy Number Variable (CNV) file, which is a table containing information about gene amplifications and deletions in patients.
[0092] The file is in text format (e.g. TSV or CSV) and includes the following required fields: Gene (gene name), Copy_Number (estimated number of gene copies) and Status (change type, e.g. "amplification" or "deletion".
[0093] 2) A gene expression data file containing a table listing the expression level of each gene in patient samples. The file is in text format (e.g., TSV or CSV) and includes the following mandatory fields: Gene (gene name or identifier, e.g., Ensembl ID), Sample_ID (sample identifier), and Expression_Value (quantitative expression metric, e.g., FPKM, TPM, or RNA-seq counts).
[0094] 3) A mutation file in VCF (Variant Call Format) version 4.2 or higher. This file contains data on genetic variants, including single-nucleotide polymorphisms, indels, and other types of mutations detected during sequencing.
[0095] The mutation file consists of two main parts: meta rows and data rows. Meta rows include information about the format version, the reference genome used, for example, GRCh38, and descriptions of additional fields, such as mutation type, allele frequency, and quality parameters. The data table header contains column headings, including sample identifiers. Sample data includes the genotype, read depth, and other parameters necessary for variant interpretation. Each data row describes a single genetic variant (allele) and includes the following fields: CHROM (chromosome), POS (chromosome position), ID (variant identifier, if known), REF (reference allele), ALT (alternative allele), QUAL (variant quality), FILTER (filter results), INFO (additional variant information), FORMAT (data format for samples), and data for each sample.
[0096] The INFO field can contain key information critical for cancer mutation analysis, such as mutation type, coverage depth, and allele frequency in the tumor.
[0097] Mutation, CNV, and gene expression data from the corresponding files are stored in an anonymized patient database for subsequent matching with the reference information table to differentiate therapeutic options within a cluster based on additional data (CNV and expression).
[0098] Next, gene sequencing data relevant for molecular genetic clustering of breast cancer tumors is extracted from the VCF file containing mutations. The resulting data is subjected to comprehensive bioinformatics processing to extract clinically relevant information about the tumor's mutation status. The binary results are fed into a logistic regression model to determine the molecular genetic cluster.
[0099] Complex bioinformatics processing is carried out manually and includes:
[0100] 1. Read quality control (FastQC, MultiQC);
[0101] 2. alignment to the reference genome (BWA, Bowtie2);
[0102] 3. removing duplicates (Picard, Sambamba);
[0103] 4. detection of somatic mutations (MuTect2, VarScan, Strelka) using artifact filtering (using databases of known artifacts, such as dbSNP, COSMIC);
[0104] 5. filtering somatic mutations by parameters:
[0105] • reading depth (DP)> 5,
[0106] • mutant allele frequency (VAF)> 0.05;
[0107] • mutation annotation (ANNOVAR, VEP) with functional impact determination (synonymous, missense, nonsense, splice site mutations).
[0108] The result of this stage is a structured set of molecular genetic data, which contains, among other things, a vector of binary values reflecting the presence (1) or absence (0) of mutations for each of the studied genes of Set-1, the copy number values (CNV) of genes for each of the studied genes of Set-2, and the level of gene expression for each of the studied genes of Set-3.
[0109] The following stages are automated and, to determine the tumor's affiliation with one of four molecular genetic clusters and select antitumor therapy, a specially developed software package is used, including:
[0110] - matrix of weights W of genes of Set-1, which make up molecular genetic clusters of breast cancer,
[0111] - pre-trained logistic regression model,
[0112] - softmax function,
[0113] - reference information table,
[0114] - software component for assigning a patient to a molecular genetic cluster and clinical recommendations.
[0115] The matrix of weights W of genes that make up molecular genetic clusters of breast cancer, with dimensions of 4×26, where 4 is the number of molecular genetic clusters of breast cancer, 26 is the number of genes studied, and the intercept free term (bias for a given cluster) of length 4, is presented in Table 1.
[0116] The algorithm for assigning a patient to a molecular genetic cluster is implemented using a pre-trained logistic regression model that uses the gene weight matrix W to determine the logit of each cluster.
[0117] The input of the pre-trained logistic regression model is a vector of binary values reflecting the presence (1) or absence (0) of mutations for each of the studied genes in a specific patient.
[0118] For each molecular genetic cluster k, the linear combination (logit) S is calculated k , as the sum of the products of gene weights by the corresponding mutation values (1 - there is a mutation, 0 - there is no mutation) plus the bias for a given cluster.
[0119] The calculation of the logit of each molecular genetic cluster is performed using formulas (1), (2), (3), (4).
[0120]
[0121]
[0122] The resulting logits for molecular genetic clusters are then converted into probabilities using the softmax function, which ensures that values are normalized to a range between 0 and 1 with the sum of probabilities equal to 1. The softmax function for cluster k is calculated as follows: P(k) = exp(S k ) / [exp(S1) + exp(S2) + exp(S3) + exp(S4)].
[0123] After determining the probability for each molecular genetic cluster, the patient belongs to the cluster with the highest probability.
[0124] The intercept parameter in a logistic regression model is an intercept term in the equation that determines the value of the target variable under the condition that all predictors are zero.
[0125] Thus, in the presented model for clustering breast cancer patients, classification is carried out by calculating the individual probability (score) for each molecular genetic cluster and then assigning the patient to the molecular genetic cluster with the highest probability value.
[0126] If there are no mutations in all analyzed genes, the patient's score for each molecular genetic cluster is determined solely by the intercept value, resulting in the patient being classified into the cluster with the highest value of this parameter. Analysis of the weight matrix shows that the highest intercept value is observed in the first cluster (3.518883), while in the second cluster it is -0.05177, in the third -0.20097, and in the fourth it reaches a minimum value of -3.26614. This pattern has important clinical significance, as it indicates that patients without detectable mutations in key genes such as TP53, PIK3CA, and other analyzed loci will be systematically classified as belonging to the first cluster.This reflects the baseline probability of belonging to a particular molecular subtype of breast cancer in the absence of specific genetic markers and may indicate the existence of a phenotypically and clinically significant group of patients characterized by a relatively low mutational load in the genes studied.
[0127] Based on the specific cluster affiliation of the tumor, the transition to the final stage of the method is carried out - the formation of a personalized therapeutic strategy and clinical recommendations.
[0128] The therapeutic strategy and clinical recommendations are determined using a reference information table containing for each molecular genetic cluster: genomic alteration data, mutation type, such as single nucleotide polymorphisms, indels, coverage depth and allele frequency in the tumor, for Gene Set-1; transcriptomic alteration data, for Gene Set-2, gene copy number, for Gene Set-3, gene expression level; description of the associated histological and molecular genetic subtypes and clinical characteristics of the breast tumor; description of the personalized therapeutic strategy and clinical recommendations.
[0129] Table 2 of the reference information is given as an example, which presents the integrative classification of breast cancer.
[0130] The method is implemented through a sequential four-step process. The first two steps are performed using standard methods, instruments, and equipment. The steps of determining the patient's molecular genetic cluster and selecting antitumor therapy are performed using a system that includes:
[0131] 1. one or more processors,
[0132] 2. RAM,
[0133] 3. Long-term data storage devices with a specially developed software package, including:
[0134] - matrix of gene weights W of dimension 4×26,
[0135] - reference information table,
[0136] - pre-trained logistic regression model,
[0137] - softmax function,
[0138] - a software component for assigning a patient to a molecular genetic cluster and selecting antitumor therapy,
[0139] 4. I / O interfaces,
[0140] 5. network communication tools.
[0141] The following were used in developing the claimed method:
[0142] - electronic database “Database of clinical and pathomorphological databases of patients with breast cancer, intended for machine learning algorithms” [Certificate of state registration of the database No. 2025622625, date of state registration 06 / 18 / 2025], which contains clinical, pathomorphological and therapeutic data in text format (names of drugs, classification of tumors, etc.), binary format (presence / absence of a certain characteristic in a patient) and numerical format (age);
[0143] - algorithm “System for personalization of breast cancer therapy based on the interpretation of whole-genome sequencing data using machine learning methods” [Certificate of state registration of computer program No. 2025665590, state registration date 04.06.2025], modifying and “preparing” data from the database for use in the algorithms of the software package;
[0144] - algorithm “System for determining the probability of achieving a complete therapeutic response in patients with breast cancer based on machine analysis of clinical parameters from databases” [Certificate of state registration of computer program No. 2025664447, date of state registration on 04.06.2025].
[0145] Implementation examples
[0146] Example 1. Determination of a molecular genetic cluster based on the mutational profile, copy number variation (CNV) and expression of selected genes, assessment of their impact on the therapeutic response for a patient with the identifier: xxxx-xxxx-xxxl.
[0147] The laboratory research data supplied to the input of the calculation model are presented in Tables 3, 4, 5.
[0148] Table 3 shows the detected single nucleotide substitutions (SNPs) and short insertions and deletions (indels).
[0149] Table 4 shows the detected copy number variants (CNVs).
[0150] Table 5 shows the gene expression level (TRM).
[0151]
[0152]
[0153] Based on the analysis of significant genetic features, the cluster assigned was: CLUSTER I. Luminal, lobular morphology.
[0154]
[0155] Cluster Description
[0156] Cluster I is characterized by frequent mutations in the GATA3 and CDH1 genes and the absence of mutations in the TP53 and PIK3CA genes, which corresponds to the luminal phenotype of the tumor.
[0157] The cluster exhibits significant focal amplifications of IL19, IL20 and IL24, indicating activation of cytokine signaling.
[0158] Gene expression signature - upregulated for GATA3, MLPH, TBC1D9, SLC39A6 and downregulated for ENO1, PKM, STMN1, YBX1 - consistent with well-differentiated luminal subtype and low proliferative / metabolic activity.
[0159] The effect on the therapeutic effect is shown in Table 6.
[0160]
[0161] Forecast
[0162] Cluster I lobular tumors tend to present as invasive lobular carcinoma (ILC). ILCs account for about 10% of breast cancer, are usually ER+, but have a unique clinical behavior - a more indolent course, often multicentricity and poor response to chemotherapy [Wilson N, Ironside A, Diana A and Oikonomidou O (2021) Lobular Breast Cancer: A Review. Front. Oncol. 10:591399. doi: 10.3389 / fonc.2020.591399]. ILCs greatly benefit from endocrine therapy and have long-term disease control with antiestrogen treatment [Wilson N, Ironside A, Diana A and Oikonomidou O (2021) Lobular Breast Cancer: A Review. Front. Oncol. 10:591399. doi: 10.3389 / fonc.2020.591399].
[0163] When a tumor is classified as Cluster I, it is advisable to prioritize endocrine therapy as the primary systemic treatment, avoiding chemotherapy whenever possible, especially in cases with low genomic risk. Thus, identifying a Cluster I tumor may spare the patient the unnecessary toxicity of chemotherapy in borderline cases.
[0164] Example 2. Determination of a molecular genetic cluster based on the mutational profile, copy number variation (CNV) and expression of selected genes, assessment of their impact on the therapeutic response for a patient with the identifier: xxxx-xxxx-xxx2.
[0165] The laboratory research data supplied to the input of the calculation model are presented in Tables 7, 8, 9.
[0166] Table 7 shows the detected single nucleotide substitutions (SNPs) and short insertions and deletions (indels).
[0167] Table 8 shows the detected copy number variants (CNVs).
[0168] Table 9 shows the gene expression level (TRM).
[0169]
[0170]
[0171]
[0172] Based on the analysis of significant genetic features, the cluster assigned was: CLUSTER II. Basal-like | HER2-enriched, ductal morphology.
[0173]
[0174] Cluster Description
[0175] Cluster II is characterized by the presence of a TP53 driver mutation and an AMPD1 mutation. A PIK3CA mutation is absent. The mutational profile of Cluster 2 (TP53^mut, PIK3CA^WT) is typical of basal-like and HER2-enriched tumors.
[0176] The cluster exhibits MYC amplification and PTEN gene deletion, which are frequently observed in basal-like breast cancer, indicating activation of the PI3K pathway despite the absence of PIK3CA mutation.
[0177] The gene expression signature of upregulation of basal cytokeratins (KRT5, KRT14, KRT17), EGFR, and proliferation / cell cycle genes (TOR2A, STMN1) and downregulation of luminal markers (ESR1, PGR, FOXA1, GATA3, MUC1) may signal increased proliferative activity and decreased tumor differentiation.
[0178] The effect on the therapeutic effect is shown in Table 10.
[0179]
[0180] Forecast
[0181] Cluster II corresponds to highly proliferative tumors caused by the TP53 mutation, which include basal-like and HER2-enriched breast cancers. This cluster corresponds to high-grade tumors with a significant risk of early recurrence.
[0182] Cluster II shows a significantly lower response to hormonal drugs - namely adjuvant Tamoxifen and Letrozole (p<0.01 for each) [Grote I, Bartels S, Kandt L, et al. TP53 mutations are associated with primary endocrine resistance in luminal early breast cancer. Cancer medicine (Maiden, MA). 2021; 10(23): 8581-8594. https: / / doi.org / 10.1002 / cam4.4376].
[0183] Cluster II has a lower response to capecitabine chemotherapy (p < 0.05). This may reflect the chemoresistance of some triple-negative tumors or be related to TP53 deficiency, which impairs the effectiveness of capecitabine as an antimetabolite. Therefore, standard chemotherapy (anthracyclines / taxanes) should be considered.
[0184] For HER2-positive cases in Cluster II, anti-HER2 therapy (trastuzumab + / - pertuzumab) is necessary. In case of detection of early triple-negative breast cancer, anti-PD-l / L1 therapy may be recommended as treatment [Schmid P. et al. Overall survival with pembrolizumab in early-stage triple-negative breast cancer / / New England Journal of Medicine. - 2024. - V. 391. - No. 21. - C. 1981-1991. https: / / doi.org / 10.1056 / NEJMoa2409932].
[0185] Example 3. Determination of a molecular genetic cluster based on the mutational profile, copy number variation (CNV) and expression of selected genes, assessment of their impact on the therapeutic response for a patient with the identifier: xxxx-xxxx-xxx3.
[0186] The laboratory research data supplied to the input of the calculation model are presented in Tables 11, 12, 13.
[0187] Table 11 shows the detected single nucleotide substitutions (SNPs) and short insertions and deletions (indels).
[0188] Table 12 shows the detected copy number variants (CNVs).
[0189] Table 13 shows the gene expression level (TRM).
[0190] Based on the analysis of significant genetic features, the cluster was assigned: CLUSTER III, Luminal, various morphology.
[0191]
[0192] Cluster Description
[0193] Cluster III is characterized by a PIK3CA driver mutation, as well as MAP3K1 and MED15 mutations. No TP53 mutations are observed in the cluster. This mutational profile (PIK3CA^mut / TP53^WT) is indicative of the luminal subtype of breast cancer.
[0194] Gene expression signature - upregulated for ESR1, PGR, GATA3, FOXA1, TFF1, XBP1, SCUBE2 and downregulated for STMN1, TOR2A, KPNA2 - suggests a possible enriched stromal signature or tumor-stroma interactions in these tumors, and is also an indicator of a well-differentiated, low-proliferative luminal status.
[0195] The effect on the therapeutic effect is shown in Table 14.
[0196]
[0197]
[0198] Forecast
[0199] Cluster III tumors demonstrate positive results with both endocrine therapy and chemotherapy. Cluster III demonstrated higher sensitivity to standard chemotherapy: the combination of doxorubicin + cyclophosphamide + paclitaxel had higher response rates than Cluster 3 (p<0.05 for each). Patients in Cluster III responded better to tamoxifen and exemestane (p<0.05).
[0200] For patients in this cluster, treatment de-escalation may be considered: given the good response to endocrine therapy, some may not require chemotherapy if diagnosed early. The addition of CDK4 / 6 inhibitors should also be considered in advanced disease. Given the presence of the PIK3CA mutation, PI3K inhibitors, which are a key adjunct in endocrine resistance, should be considered.
[0201] Example 4. Determination of a molecular genetic cluster based on the mutational profile, copy number variation (CNV) and expression of selected genes, assessment of their impact on the therapeutic response for a patient with the identifier: xxxx-xxxx-xxx4.
[0202] The laboratory research data supplied to the input of the calculation model are presented in Tables 15, 16, 17.
[0203] Table 15 shows the detected single nucleotide substitutions (SNPs) and short insertions and deletions (indels).
[0204] Table 16 lists the detected copy number variants (CNVs).
[0205] Table 17 shows the gene expression level (TRM).
[0206]
[0207] Based on the analysis of significant genetic features, the cluster assigned was: CLUSTER IV. Luminal B-like | HER2-enriched, ductal morphology.
[0208]
[0209] Cluster IV is characterized by co-mutations in PIK3CA and TP53 and also showed evidence of a higher overall mutational burden—mutations were detected in the MUC17, OFD1, RALGPS1, and POU3F2 genes. This mutational profile suggests an increased likelihood of tumor invasion or metastasis.
[0210] The effect on the therapeutic effect is shown in Table 18.
[0211] Forecast
[0212] Cluster IV tumors will have slightly lower complete response rates to HER2 blockade and potentially earlier recurrences, in particular due to the PIK3CA mutation [Kim, Ju Won, et al. "PIK3CA mutation is associated with poor response to HER2-targeted therapy in breast cancer patients." Cancer Research and Treatment: Official Journal of Korean Cancer Association 55.2 (2023): 531–541. https: / / doi.org / 10.4143 / crt.2022.2211]. In addition, they tend to be less sensitive to chemotherapy.
[0213] Cluster IV patients, including both Luminal B-like (ER+, HER2+) and HER2-enriched (ER-, HER2+) subtypes, may benefit from new therapy combinations such as HER2 blockade in combination with PI3K pathway inhibitors. Even in hormone-negative breast cancer, treatment escalation or the addition of PI3K or AKT / mTOR inhibitors to standard therapy may provide optimal disease control [Guarneri V. et al.: PIK3CA Mutation in the ShortHER Randomized Adjuvant Trial for Patients with Early HER2+ Breast Cancer: Association with Prognosis and Integration with PAM50 Subtype. Clin Cancer Res 15 November 2020; 26 (22): 5843-5851. Cluster IV patients require close monitoring due to their genomic predisposition to resistance.
Claims
1. A method for selecting antitumor therapy for breast cancer using genomic data, including the following steps: 1) obtaining a sample of the patient’s biological material in which the tumor cell content is at least 20%, the degradation index is DIN ≥ 6.0, and the DNA concentration is at least 10 ng / μl; 2) a molecular genetic study with the receipt of a structured set of molecular genetic data containing a vector of binary values of gene mutations, consisting of values 1 and 0, including: - extraction of nucleic acids by isolating DNA from samples with DNA quality control; - preparation of libraries for NGS sequencing by sequentially performing the following steps: end repair, adenylation, adapter ligation, PCR amplification, product purification; - high-performance targeted sequencing to generate an array of genomic data using prepared libraries for NGS sequencing with a minimum coverage depth of at least 500×, paired reads of at least 150 bp in length, and targeted sequencing is performed using a diagnostic panel of genes including a set of genes relevant for molecular genetic clustering of breast cancer tumors: PIK3CA, TP53, TTN, CDH1, MUC16, GATA3, MAP3K1, IGDCC4, PTEN, KMT2C, RYR2, SYNE1, NEB, HMCN1, RYR3, DST, SPTA1, FLG, CSMD2, DMD, CSMD3, OBSCN, MUC5B, TBX3, NCOR1, for each gene of which the presence of mutations is determined, as well as sets of genes for a detailed characterization of the cluster, namely a set genes: IL19, IL20, IL24, MYC, PTEN, for each gene of which the copy number is determined, and a set of genes: ATA3, MLPH, KRT5, KRT14, ESR1, PGT, FOX1A, ERBB2, for each gene of which the expression level is determined; - complex bioinformatics processing of data obtained as a result of sequencing to extract clinically significant information on the mutational status of the tumor by means of read quality control; alignment to the reference genome; removal of duplicates; detection of somatic mutations using artifact filtering; filtering of somatic mutations by read depth, DP> 5, mutant allele frequency, VAF> 0.05; annotation of mutations with determination of the functional impact; 3) determining the molecular genetic cluster for the patient using a pre-trained logistic regression model containing a 4×26 gene weight matrix W = [[3.518883, -3.68215, -3.64135, -0.0379, 0.352602, -0.03618, 0.494444, 0.20169, 0.158118, 0.105565, -0.1241, 0.051482, -0.06702, 0.022258, 0.142287, 0.105093, 0.01345, 0.049623, -0.05263, -0.01629, 0.027151, 0.022393, -0.06247, -0.01593, 0.200114, 0.02975], [-0.05177, -3.15651, 3.689056, 0.016745, -0.41153, -0.09685, -0.21961, -0.24982, 0.088314, -0.05926, -0.04971, -0.03935, 0.006928, -0.02346, 0.070376, -0.02617, 0.040066, 0.038742, 0.148841, 0.059149, -0.00621, 0.005337, 0.062735, 0.064656, -0.08761, 0.046106], [-0.20097, 3.720687, -3.07543, -0.12832, 0.269606, 0.018014, 0.092575, 0.252909, -0.08035, -0.02177, 0.172667, -0.13865, 0.080254, -0.13274, -0.06047, 0.051832, -0.10939, 0.036242, -0.00127, 0.026676, -0.07229, -0.08424, 0.076258, -0.0006, -0.10111, -0.00236], [-3.26614, 3.117965, 3.027726,0.149473, -0.21068, 0.115015, -0.3674, -0.20478, -0.16609, -0.02453, 0.001146, 0.126522, -0.02016, 0.133943, -0.1522, -0.13075, 0.055872, -0.12461, -0.09494, -0.06954, 0.051354, 0.056514, -0.07652, -0.04813, -0.0114, -0.0735]], where 4 is the number of molecular genetic clusters, 26 is the free intercept term of length 4 and the weights of twenty-five studied genes: PIK3CA, TP53, TTN, CDH1, MUC16, GATA3, MAP3K1, IGDCC4, PTEN, KMT2C, RYR2, SYNE1, NEB, HMCN1, RYR3, DST, SPTA1, FLG, CSMD2, DMD, CSMD3, OBSCN, MUC5B, TBX3, NCOR1, by feeding the vector of binary values of gene mutations obtained at the stage of molecular genetic research to the input of the logistic regression model; calculating the linear logit for each molecular genetic cluster k,as the sum of the products of gene weights by the corresponding values of gene mutations plus the bias for a given molecular genetic cluster; transformation of the obtained logits for all molecular genetic clusters into probabilities using the softmax function, which ensures normalization of values in the range from 0 to 1 with the sum of probabilities equal to 1, and the softmax function is calculated as P(k) = exp(S, k ) / [exp(S1) + exp(S2) + exp(S3) + exp(S4)], where S k - logit of molecular genetic cluster k, k = 1, 2, 3, 4; assignment of a patient to a molecular genetic cluster with the highest probability; 4) definition of a personalized therapeutic strategy and clinical recommendations for the identified molecular genetic cluster.
2. The method according to paragraph 1, characterized in that samples of the patient’s biological material are obtained by biopsy or surgically.
3. The method according to paragraph 1, characterized in that DNA is isolated from samples using standard extraction methods, for example, phenol-chloroform extraction or commercial kits for isolating nucleic acids.
4. The method according to paragraph 1, characterized in that DNA quality control is performed using electrophoresis in agarose gel or microfluidic analyzers.
5. The method according to paragraph 1, characterized in that the microfluidic analyzers are selected from the group: Bioanalyzer, TapeStation.
6. The method according to claim 1, characterized in that high-performance targeted sequencing is carried out on platforms such as Illumina NovaSeq, HiSeq or MiSeq.
7. The method according to paragraph 1, characterized in that the quality control of reads is performed using the FastQC, MultiQC programs.
8. The method according to claim 1, characterized in that the alignment to the reference genome is performed using the BWA, Bowtie2 programs.
9. The method according to paragraph 1, characterized in that the removal of duplicates is performed using the Picard and Sambam programs.
10. The method according to paragraph 1, characterized in that the detection of somatic mutations is performed using the tool: MuTect2, VarScan, Strelka, using artifact filtering using databases of known artifacts, such as dbSNP, COSMIC.
11. The method according to claim 1, characterized in that the annotation of mutations with the determination of the functional impact is performed using a tool for annotating genetic variants, for example, ANNOVAR, VEP.
12. The method according to claim 1, characterized in that the determination of a personalized therapeutic strategy and clinical recommendations for the found molecular genetic cluster is carried out using a reference information table containing for each molecular genetic cluster: data on genomic changes for each gene from the set of genes: PIK3CA, TP53, TTN, CDH1, MUC16, GATA3, MAP3K1, IGDCC4, PTEN, KMT2C, RYR2, SYNE1, NEB, HMCN1, RYR3, DST, SPTA1, FLG, CSMD2, DMD, CSMD3, OBSCN, MUC5B, TBX3, NCOR1; data on transcriptomic changes in genes for each gene from the set of genes: IL19, IL20, IL24, MYC, PTEN; Data on transcriptome changes in genes for each gene from the gene set: GATA3, MLPH, KRT5, KRT14, ESR1, PGT, FOX1A, ERBB2; description of the associated histological and molecular genetic subtypes and clinical characteristics of breast tumors; description of a personalized therapeutic strategy and clinical recommendations.
13. The method according to claim 12, characterized in that the data on genomic changes includes data on the type of mutations, for example, single nucleotide polymorphisms, indels, depth of coverage and allele frequency in the tumor.
14. The method according to claim 12, characterized in that the transcriptome change data includes data on the number of gene copies and the type of changes for each gene from the set of genes: IL19, IL20, IL24, MYC, PTEN, and data on the expression level for each gene from the set of genes: GATA3, MLPH, KRT5, KRT14, ESR1, PGT, FOX1A, ERBB2.
15. The method according to claim 12, characterized in that a table is used containing descriptions of four subtypes of breast tumor: luminal subtype A, a hybrid phenotype between basal-like and HER2-enriched subtypes, luminal subtype A with an immunohistochemical profile, and a HER2-positive subtype represented by HER2-enriched or luminal B / HER2 variants.