System for selection of antitumor therapy for breast cancer using genomic data

A molecular genetic profiling system for breast cancer addresses the limitations of current treatment prediction methods by clustering tumors based on gene mutations and signaling pathways, enhancing treatment accuracy and effectiveness.

RU2864781C1Active Publication Date: 2026-06-29OBSHCHESTVO S OGRANICHENNOI OTVETSTVENNOSTIU ONKO ANALITIKA +1
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
RU · RU
Patent Type
Patents
Current Assignee / Owner
OBSHCHESTVO S OGRANICHENNOI OTVETSTVENNOSTIU ONKO ANALITIKA
Filing Date
2025-09-01
Publication Date
2026-06-29

AI Technical Summary

Technical Problem

Current methods for predicting the effectiveness of breast cancer treatment are imperfect, often failing to account for the activation of signaling pathways beyond driver mutations and lack consistency with modern classifications, leading to treatment failures and resistance.

Method used

A molecular genetic profiling system for breast cancer patients that uses a unique system of clustering based on gene mutations, considering both sensitivity and resistance to various treatments, and is consistent with clinical guidelines, allowing for the selection of personalized therapeutic strategies.

Benefits of technology

The system increases the accuracy of selecting antitumor therapy by predicting both sensitivity and resistance, aligning with modern classifications, thereby improving treatment effectiveness and reducing unnecessary therapies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000001
    Figure 00000001
  • Figure 00000002
    Figure 00000002
  • Figure 00000003
    Figure 00000003
Patent Text Reader

Abstract

FIELD: medical diagnostics and oncology.SUBSTANCE: personalized selection of antitumor therapy based on molecular genetic analysis of tumor tissue samples from patients with breast cancer. A solution has been proposed that allows for the prediction of tumor response to various types of treatment, including targeted, hormonal, immune, and chemotherapy. The key difference of the invention is the use of a unique system of molecular genetic clustering of breast tumors based on gene expression profiles, which complements the modern classification of tumors and is consistent with clinical guidelines (ESMO, NCCN, PAM50, etc.). Moreover, the genes used have high prognostic value and biological interpretability, which makes the model reliable and ready for use in clinical practice. The presented system allows not only to select the most effective medications, but also to proactively exclude therapies to which the patient may have primary resistance.EFFECT: increase in the accuracy of selection of antitumor therapy and, consequently, the effectiveness of treatment.7 cl, 18 tbl, 4 ex
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to the field of medical diagnostics and oncology, in particular to the personalized selection of antitumor therapy based on molecular genetic analysis of tumor tissue samples from patients with breast cancer.

[0002] Breast cancer is a heterogeneous group of malignant neoplasms that arise from the accumulation of genetic damage (genetic aberrations, altered transcript levels) in breast cells. This leads to loss of cell cycle control, unrestricted proliferation, and the acquisition of invasive and metastatic capabilities. This pathology is among the leading causes of cancer-related mortality among women, both in our country and globally.

[0003] Currently, although tools exist in clinical practice for predicting a specific patient's sensitivity to various types of anticancer treatment based on the individual molecular profile of their disease, they are imperfect. Patients are screened for a driver mutation and, for example, a targeted drug is prescribed based on it, but the activation of an entire signaling pathway, which can also be targeted by a drug in the absence of a driver mutation, is not taken into account. This results in the vast majority of patients receiving therapy based on traditional clinical and pathological characteristics: stage of the disease, histological type of tumor, degree of differentiation, lymph node status, and expression of individual biomarkers. This approach often fails to achieve an optimal therapeutic effect, leading to treatment failure and disease progression.

[0004] The development of personalized breast cancer therapy strategies based on a comprehensive analysis of the molecular biological characteristics of a specific patient's tumor is one of the most pressing challenges in modern oncology, which this invention is designed to address.

[0005] An analysis of existing solutions revealed that most commercial products, such as OncoBox, FoundationOne, Genomed, and Genetico, focus on analyzing a limited number of genetic markers (driver mutations), often without considering the functional activity of genes and systemic interactions. Furthermore, driver mutations are detected in only 20% of patients. These companies' approaches are primarily aimed at targeted therapy, relying on a "one mutation, one targeted drug" approach, and are limited in predicting response to hormonal and cytostatic treatment. However, the main problem is the inconsistency of this approach: the absence of a given mutation coupled with the presence of resistance does not mean that the use of a targeted drug is inappropriate. Resistance may be caused by the activation of other elements of a signaling pathway or multiple pathways, which can also lead to the development of resistance.

[0006] Furthermore, current patient clustering based on the clinical and biological characteristics of the tumor does not fully reflect all the nuances of various therapeutic approaches, generating only strict drug combinations, which does not allow for the individual characteristics of each treatment. Below, we will examine in detail the patent applications for each solution and highlight the key differences from this method.

[0007] A method and system for assessing the clinical efficacy of targeted drugs, OncoBox [RU 2741703, 2018-03-01, C12N-005 / 09, G01N-033 / 15, G16B-050 / 30, G16H-010 / 40], are known. The patent describes a method for integrating gene expression and mutational status to predict drug sensitivity. The method allows for assessing the presence and severity of a positive effect from the use of a particular drug (efficacy), and also provides information on the possibility of developing resistance to a given drug.

[0008] However, this approach does not allow for the formation of stable, therapeutically significant clusters that would correspond to clinically and biologically verified tumor groups. This limits the applicability of this technology within the framework of practical oncology standards and complicates its implementation into routine clinical practice due to its incompatibility with the current classification and the need for whole-genome or whole-exome sequencing. Furthermore, the method does not focus on a specific nosology (e.g., breast cancer only), which may reduce the accuracy of predictions: for each tumor location, the same molecular biological pathways may influence oncogenesis differently. This may create additional obstacles to selecting the optimal treatment strategy.

[0009] A known method for determining the activity of cellular signaling pathways using probabilistic modeling of target gene expression [WO / 2013 / 011479, 19.07.2012, C12Q 1 / 68, G06F 19 / 24], is based on the analysis of the activity of tumor cell signaling pathways based on the expression level of a set of key genes. The assessment is carried out using a probabilistic model, in particular, a Bayesian network, which takes into account the relationship between the level of gene expression, the activity of transcription factors and the functional state of the signaling pathway (e.g., Wnt, ER, Hedgehog, AR). For each pathway, a set of genes is defined, the expression indicators of which are used as markers of pathway activity. This approach is capable of predicting sensitivity to drugs that affect a specific pathway, and also serving as a recommendation for the appointment of targeted therapy if the activity of certain elements of the pathway is pathologically altered.

[0010] However, this approach has several limitations. Firstly, the model does not take into account drug resistance, focusing primarily on drug sensitivity. This limits its application in real-world oncology practice, where it is important to consider both the positive and negative effects of therapy. Secondly, it lacks consistency with modern classifications, such as PAM50, complicating the implementation of the technology into routine practice. It also does not allow for the creation of a molecular profile (ie, comprehensive information on expression changes or the presence of mutations) of the tumor, which could negatively impact the appropriateness of therapy selection.

[0011] Systems and methods for evaluating the effectiveness of a drug are known [US 12266426, 2018-12-03, C12Q-001 / 6886, G06F-017 / 15, G06F-017 / 16, G06F-017 / 18, G16B-040 / 20, G16B-045 / 00, G16B-050 / 20, G16B-050 / 30, G16H-050 / 20, G16H-050 / 50]. The patent describes a method for analyzing the expression of genes and proteins associated with the immune response in the context of immunotherapy of oncological diseases. Specifically, the invention utilizes molecular data to assess the activity of immune checkpoints (such as CTLA4, PD1, PDL1, LAG3, TIGIT, and others) responsible for regulating T-cell activity and forming an immunosuppressive tumor microenvironment. This method allows for determining a patient's sensitivity to immune therapy and its effectiveness; however, it is not possible to determine the effectiveness of other therapies using this method.

[0012] A known method for determining the risk of recurrence of breast cancer [RU 2626603, 31.12.2015, G01N 33 / 574, C12Q 1 / 68], allows for predicting the risk of breast cancer recurrence based on the quantitative determination of the expression of three genes: ELOVL5, IGFBP6, TXNDC9. This application describes an approach to determining the probability of tumor recurrence of the Luminal A subtype of breast cancer, which is its serious limitation. The method does not provide for the ability to determine the effectiveness of therapy or the development of resistance, and is also not applicable to other subtypes of breast cancer and other types of oncological processes.

[0013] A system, method, and software for improving the efficacy and safety of drugs in a patient are known [US 20160132632, 2016-05-12, C12Q-001 / 68 G16B-005 / 00 G16B-050 / 00; US 20170193176, 2017-07-06, G06F-019 / 00 G06N-005 / 04 /

[0014] The closest to the proposed technical solution in terms of the set of essential features and the result obtained is a system, method and software for predicting the clinical outcome of drug treatment for breast cancer in a patient [US 20160224739, 2016-08-04, G16Z-099 / 00]

[0015] By the patent applications US 20160132632, US 20170193176, and US 20160224739, a method for predicting the clinical outcome of antitumor therapy for breast cancer is described, based on the analysis of signaling pathway activity and the use of machine learning. The technology is based on the OncoFinder software tool, which calculates a Pathway Activation Score (PAS) based on gene expression data. One of the key advantages of the application is the use of a systems approach to modeling signaling pathways, which allows for consideration of not only individual mutations or gene expression levels but also their functional impact on biological processes. The ability to automate analysis and integrate with clinical data is also claimed, opening up potential for application in personalized medicine.

[0016] However, this approach has a number of limitations.

[0017] Firstly, there is a lack of consistency with modern classifications such as RAM50, which makes it difficult to implement the technology into routine practice.

[0018] Secondly, although the system can assess sensitivity to therapy, it does not focus on identifying patients who are obviously resistant to treatment, which reduces its practical value in a clinical setting where the prognosis of both the effectiveness and ineffectiveness of drugs is important.

[0019] Third, the system requires whole genome sequencing or expression data of tens of thousands of genes, making it labor-intensive and expensive to implement, especially in resource-limited healthcare settings.

[0020] The objective of the invention is to create an effective system for selecting antitumor therapy for breast cancer using genomic and transcriptomic data, which overcomes the shortcomings of analogues.

[0021] The technical result is an increase in the accuracy of selection of antitumor therapy for breast cancer and, consequently, the effectiveness of treatment.

[0022] Increased efficiency is achieved through:

[0023] - the ability to calculate the probability of achieving a complete clinical response to tumor response to various types of treatment, including targeted, hormonal, immune and chemotherapy;

[0024] - identifying not only sensitivity to therapy, but also resistance;

[0025] - the use of genes that have high prognostic value and biological interpretability, comparability with clinical and pathomorphological data of patients.

[0026] Due to the fact that the condition of consistency with modern classifications such as RAM50, St. Galen classification of breast cancer is taken into account, ease of implementation is also achieved.

[0027] Modern approaches to personalized medicine are becoming an integral part of oncology, increasing the accuracy of predicting the effectiveness of anticancer treatments. In this regard, we have developed a cutting-edge solution that implements a molecular genetic profiling system for breast cancer patients, enabling us to calculate the probability of achieving a complete clinical response based on the tumor's response to various treatments, including targeted, hormonal, immune, and chemotherapy.

[0028] The key difference of the proposed invention is the use of a unique system of molecular genetic clustering of breast tumors based on gene mutations, which complements the modern classification of tumors and is consistent with clinical guidelines (ESMO, NCCN, PAM50, etc.).

[0029] The model is based on the identification of four stable molecular genetic clusters (hereinafter also referred to as "cluster"). A cluster is defined by the presence and combination of mutations in a set of genes and is described by changes in the expression and copy number of genes associated with both sensitivity and resistance to different groups of anticancer drugs. Furthermore, the genes used have high prognostic value and biological interpretability, making the model reliable and ready for use in clinical practice.

[0030] The presented system allows not only to select the most effective antitumor drugs, but also to exclude in advance therapies to which the patient may have primary resistance.

[0031] The system allows you to use information:

[0032] - based on a wide range of molecular genetic data obtained from tumor tissue samples on somatic mutations, gene copy number variations (amplifications and deletions), and expression levels for key genes obtained by targeted high-throughput (NGS) sequencing;

[0033] - on the therapeutic outcomes of various antitumor drugs and the molecular characteristics of breast cancer subtypes.

[0034] The system uses a pre-trained logistic regression model with a weight matrix W of size 4×26, where 4 is the number of molecular genetic clusters, 26 is the number of genes studied, and a bias vector of length 4, and an algorithm to determine the probability of a breast tumor belonging to one of the four molecular genetic clusters.

[0035] This system automates the process of determining tumor affiliation with one of four molecular genetic clusters and selecting antitumor therapy for breast cancer. This eliminates potential errors associated with manual interpretation of genomic data and allows for the consideration of patient-specific changes. This system enables the selection of a therapeutic strategy for a patient based on the analysis of objective, individual genomic changes occurring in the tumor tissue.

[0036] A system for selecting antitumor therapy for breast cancer using genomic data is proposed, which contains:

[0037] 1. one or more processors,

[0038] 2. RAM,

[0039] 3. Long-term data storage devices with a specially developed software package, including:

[0040] - matrix of gene weights W of dimension 4×26,

[0041] - reference information table,

[0042] - pre-trained logistic regression model,

[0043] - softmax function,

[0044] - a software component for assigning a patient to a molecular genetic cluster and selecting antitumor therapy,

[0045] 4. I / O interfaces,

[0046] 5. network communication tools.

[0047] According to the invention, the system comprises a gene weight matrix W of dimension 4×26 = [[3.518883, -3.68215, -3.64135, -0.0379, 0.352602, -0.03618, 0.494444, 0.20169, 0.158118, 0.105565, -0.1241, 0.051482, -0.06702, 0.022258, 0.142287, 0.105093, 0.01345, 0.049623, -0.05263, -0.01629, 0.027151, 0.022393, -0.06247, -0.01593, 0.200114, 0.02975], [-0.05177, -3.15651, 3.689056, 0.016745, -0.41153, -0.09685, -0.21961, -0.24982, 0.088314, -0.05926, -0.04971, -0.03935, 0.006928, -0.02346, 0.070376, -0.02617, 0.040066, 0.038742, 0.148841, 0.059149, -0.00621, 0.005337, 0.062735, 0.064656, -0.08761, 0.046106], [-0.20097, 3.720687, -3.07543, -0.12832, 0.269606, 0.018014, 0.092575, 0.252909, -0.08035, -0.02177, 0.172667, -0.13865, 0.080254, -0.13274, -0.06047, 0.051832, -0.10939, 0.036242, -0.00127, 0.026676, -0.07229, -0.08424, 0.076258, -0.0006, -0.10111, -0.00236], [-3.26614, 3.117965, 3.027726, 0.149473, -0.21068, 0.115015, -0.3674, -0.20478, -0.16609, -0.02453, 0.001146, 0.126522, -0.02016, 0,133943, -0.1522, -0.13075, 0.055872, -0.12461, -0.09494, -0.06954, 0.051354, 0.056514, -0.07652, -0.04813, -0.0114, -0.0735]], where 4 is the number of molecular genetic clusters, 26 is the free intercept term of length 4 and the weights of twenty-five studied genes: PIK3CA, TP53, TTN, CDH1, MUC16, GATA3, MAP3K1, IGDCC4, PTEN, KMT2C, RYR2, SYNE1, NEB, HMCN1, RYR3, DST, SPTA1, FLG, CSMD2, DMD, CSMD3, OBSCN, MUC5B, TBX3, NCOR1.,

[0048] According to the invention, the system comprises a reference information table containing for each molecular genetic cluster: data on genomic changes, including data on the type of mutations, for example, single nucleotide polymorphisms, indels, coverage depth and allele frequency in the tumor, for each gene from the set of genes: PIK3CA, TP53, TTN, CDH1, MUC16, GATA3, MAP3K1, IGDCC4, PTEN, KMT2C, RYR2, SYNE1, NEB, HMCN1, RYR3, DST, SPTA1, FLG, CSMD2, DMD, CSMD3, OBSCN, MUC5B, TBX3, NCOR1 (hereinafter Set-1);

[0049] transcriptome variation data, namely gene copy number (CNV) data and gene variation type (e.g. amplification, deletion), for each gene in the gene set: IL19, IL20, IL24, MYC, PTEN (hereinafter referred to as Set-2); transcriptome variation data, namely expression level data, for each gene in the gene set: GATA3, MLPH, KRT5, KRT14, ESR1, PGT, FOX1A, FOX1A, ERBB2 (hereinafter referred to as Set-3), namely the data;

[0050] Description of associated histological and molecular genetic subtypes and clinical characteristics of breast tumor; description of personalized therapeutic strategy and clinical recommendations.

[0051] According to the invention, the reference information table contains descriptions of four breast tumor subtypes: luminal subtype A, a hybrid phenotype between basal-like and HER2-enriched subtypes, luminal subtype A with an immunohistochemical profile, and a HER2-positive subtype represented by HER2-enriched or luminal B / HER2 variants.

[0052] According to the invention, the system comprises a software package for clustering breast tumors and assigning a patient to one of four molecular genetic clusters. The software package contains a pre-trained logistic regression model for calculating the linear logit for each molecular genetic cluster k, as the sum of the products of the gene weights by the corresponding values ​​of gene mutations fed to the input of the logistic regression model in the form of a vector of binary values ​​reflecting the presence (1) or absence (0) of mutations for each of the genes under study, plus an offset for a given molecular genetic cluster; a softmax function for converting the obtained logit for each molecular genetic cluster into a probability, which ensures normalization of values ​​in the range from 0 to 1 with the sum of the probabilities equal to 1 and is calculated as P(k) = exp(S k ) / [exp(S1) + exp(S2) + exp(S3) + exp(S4)], where S k- logit of the molecular genetic cluster k, k = 1, 2, 3, 4; a software element (algorithm) for assigning a patient to a molecular genetic cluster with the highest probability, with the determination of a personalized therapeutic strategy and clinical recommendations for this cluster based on the corresponding data in the reference information table.

[0053] Long-term data storage can include hard drives, solid-state drives, RAID arrays, or cloud storage.

[0054] Input / output interfaces enable interaction with the user and external systems. These include data input devices (keyboards, mice, touchpads), information display devices (monitors, printers), as well as interfaces for connecting laboratory equipment and medical information systems.

[0055] Networking tools allow the system to be integrated into the existing IT infrastructure of a medical institution, providing access to external databases, cloud computing resources, and remote workstations.

[0056] The software package can be conditionally divided into three functionally related blocks: a data input block, a calculation block, and a results output block.

[0057] The operator enters the following using the data entry device:

[0058] - patient information, including full name, basic demographic and clinical parameters, and information on the analytical equipment used to test the patient's samples;

[0059] - a structured set of molecular genetic data of the patient obtained at the stage of molecular genetic research, including a vector of binary values ​​of mutations of the genes of Set-1, values ​​of the expression level of the genes of Set-3, and the number of copies (CNV) of the genes of Set-2.

[0060] A vector of binary values ​​of patient gene mutations is passed to the calculation block, which is the input for the pre-trained logistic regression model.

[0061] The calculation block carries out:

[0062] - calculation by means of a logistic regression model using a vector of binary values ​​of gene mutations and a matrix of gene weights W of a linear logit for each molecular genetic cluster, as the sum of the products of gene weights by the corresponding values ​​of gene mutations plus an offset for a given molecular genetic cluster,

[0063] - transformations using the softmax function of the obtained logits for all molecular genetic clusters into probabilities,

[0064] - assignment of the patient to a molecular genetic cluster, the probability of which is maximum,

[0065] - definition for the resulting cluster of personalized therapeutic strategy and clinical recommendations in accordance with the reference information table,

[0066] - transfer of received information to the results output block.

[0067] The results output unit organizes the output of results to display devices (monitors, printers), as well as interfaces for connecting laboratory equipment and medical information systems. Any data on the patient being studied can be output, including research data fed to the input of the calculation model, such as those presented in Tables 3, 4, and 5; cluster number and characteristics; cluster description; impact on therapeutic effect; prognosis.

[0068] The system can be implemented in various configurations:

[0069] - as a standalone workstation based on a personal computer,

[0070] - as a server solution for serving multiple users,

[0071] - as a cloud service with a web interface, or

[0072] - as a module integrated into specialized medical information systems and laboratory complexes.

[0073] The input to the system is a structured set of molecular genetic data, which contains, among other things, a vector of binary values ​​reflecting the presence (1) or absence (0) of mutations for each gene from the studied genes of Set-1, the copy number values ​​(CNV) of genes for each gene from the studied genes of Set-2, and the level of gene expression for each gene from the studied genes of Set-3.

[0074] The preparation of a structured set of molecular genetic data is carried out as follows.

[0075] First, a patient biopsy or surgical removal of the tumor is performed to obtain a sample of biological material for analysis. Patient tumor tissue samples must meet quality criteria: tumor cell content of at least 20%, absence of significant DNA degradation (DIN degradation index ≥ 6.0), and DNA concentration of at least 10 ng / μL.

[0076] After obtaining high-quality biological material, the molecular genetic stage of the study begins, including nucleic acid extraction and preparation for high-throughput sequencing. DNA is isolated from samples using standard extraction methods, such as phenol-chloroform extraction or commercial nucleic acid extraction kits. DNA quality control is performed using agarose gel electrophoresis or microfluidic analyzers (e.g., Bioanalyzer, Tare Station).

[0077] The resulting high-quality DNA is then subjected to specialized processing to generate libraries optimized for sequencing target genomic regions. Library preparation includes end repair, adenylation, adapter ligation, PCR amplification, and product purification.

[0078] Library preparation is carried out for targeted sequencing.

[0079] Prepared libraries for targeted sequencing are then submitted to high-throughput sequencing to generate the genomic data set required for subsequent analysis. Sequencing is performed on high-throughput sequencing platforms (e.g., Illumina NovaSeq, HiSeq, or MiSeq) with a coverage depth of at least 100× for the tumor sample.

[0080] Targeted sequencing is performed with a minimum coverage depth of at least 500×, paired-end reads of at least 150 bp in length for a gene panel comprising three gene sets.

[0081] The first set of genes included in the diagnostic panel is analyzed for mutations - this data is used to determine the tumor's belonging to one of four molecular genetic clusters. The resulting mutational profile serves as the basis for classification. The set of genes analyzed for mutations (hereinafter Set-1): PIK3CA, TP53, TTN, CDH1, MUC16, GATA3, MAP3K1, IGDCC4, PTEN, KMT2C, RYR2, SYNE1, NEB, HMCN1, RYR3, DST, SPTA1, FLG, CSMD2, DMD, CSMD3, OBSCN, MUC5B, TBX3, NCOR1.

[0082] The second and third gene sets, also included in the diagnostic panel, are used to further characterize the identified cluster. These genes are analyzed for copy number variations (CNVs) and gene expression levels. This expanded analysis allows for a more complete molecular picture of the tumor within a specific cluster. The gene set analyzed for copy number variations (hereinafter referred to as Set 2) includes IL19, IL20, IL24, MYC, and PTEN. The gene set analyzed for expression levels (hereinafter referred to as Set 3) includes GATA3, MLPH, KRT5, KRT14, ESR1, PGT, FOX1A, and ERBB2.

[0083] This additional information has important practical implications, as it helps establish correlations between molecular features and the clinical and morphological characteristics of the tumor, providing a deeper understanding of the biology of a particular case. Furthermore, this approach significantly expands the range of potential therapeutic options, as it allows for the identification of additional targets for targeted therapy that may be effective in a given molecular subtype of breast cancer.

[0084] The output from the sequencer is three files:

[0085] 1) A copy number variance (CNV) file, which is a table containing information about gene amplifications and deletions in patients. The file is in text format (e.g., TSV or CSV) and includes the following required fields: Gene (gene name), Copy_Number (estimated gene copy number), and Status (the type of alteration, e.g., "amplification" or "deletion").

[0086] 2) A gene expression data file containing a table listing the expression level of each gene in patient samples. The file is in text format (e.g., TSV or CSV) and includes the following mandatory fields: Gene (gene name or identifier, e.g., Ensembl ID), Sample_ID (sample identifier), and Expression_Value (quantitative expression metric, e.g., FPKM, TPM, or RNA-seq counts).

[0087] 3) A mutation file in VCF (Variant Call Format) version 4.2 or higher. This file contains data on genetic variants, including single-nucleotide polymorphisms, indels, and other types of mutations detected during sequencing.

[0088] The mutation file consists of two main parts: meta rows and data rows. Meta rows include information about the format version, the reference genome used, for example, GRCh38, and descriptions of additional fields, such as mutation type, allele frequency, and quality parameters. The data table header contains column headings, including sample identifiers. Sample data includes the genotype, read depth, and other parameters necessary for variant interpretation. Each data row describes a single genetic variant (allele) and includes the following fields: CHROM (chromosome), POS (chromosome position), ID (variant identifier, if known), REF (reference allele), ALT (alternative allele), QUAL (variant quality), FILTER (filter results), INFO (additional variant information), FORMAT (sample data format), and data for each sample.

[0089] The INFO field can contain key information critical for cancer mutation analysis, such as mutation type, coverage depth, and allele frequency in the tumor.

[0090] Mutation, CNV, and gene expression data from the corresponding files are stored in an anonymized patient database for subsequent matching with the reference information table to differentiate therapeutic options within a cluster based on additional data (CNV and expression).

[0091] Next, gene sequencing data relevant for molecular genetic clustering of breast cancer tumors is extracted from the VCF file containing mutations. The resulting data is subjected to comprehensive bioinformatics processing to extract clinically relevant information about the tumor's mutation status. The binary results are fed into a logistic regression model to determine the molecular genetic cluster.

[0092] Complex bioinformatics processing includes:

[0093] 1. Read quality control (FastQC, MultiQC);

[0094] 2. alignment to the reference genome (BWA, Bowtie2);

[0095] 3. remove duplicates (Picard, Sambamba);

[0096] 4. detection of somatic mutations (MuTect2, VarScan, Strelka) using artifact filtering (using databases of known artifacts, such as dbSNP, COSMIC);

[0097] 5. filtering somatic mutations by parameters:

[0098] • reading depth (DP)> 5,

[0099] • mutant allele frequency (VAF)> 0.05;

[0100] • mutation annotation (ANNOVAR, VEP) with definition of functional impact (synonymous, missense mutations, nonsense mutations, splice site mutations).

[0101] The result of this stage is a structured set of molecular genetic data, which contains a vector of binary values ​​of mutations for each of the studied genes of Set-1, copy number values ​​(CNV) of genes for each of the studied genes of Set-2, and the level of gene expression for each of the studied genes of Set-3.

[0102] In the following stages, to determine the tumor's belonging to one of the four molecular genetic clusters and to select antitumor therapy, the proposed system is used with a specially developed software package, including:

[0103] - matrix of weights W of genes of Set-1, which make up molecular genetic clusters of breast cancer,

[0104] - pre-trained logistic regression model,

[0105] - softmax function,

[0106] - reference information table,

[0107] - a software component for assigning a patient to a molecular genetic cluster and determining a personalized therapeutic strategy and clinical recommendations.

[0108] The matrix of weights W of genes that make up molecular genetic clusters of breast cancer, with dimensions of 4×26, where 4 is the number of molecular genetic clusters of breast cancer, 26 is the number of genes studied, and the intercept free term (bias for a given cluster) of length 4, is presented in Table 1.

[0109]

[0110]

[0111] The algorithm for assigning a patient to a molecular genetic cluster is implemented using a pre-trained logistic regression model that uses the gene weight matrix W to determine the logit of each cluster.

[0112] The input of the pre-trained logistic regression model is a vector of binary values ​​reflecting the presence (1) or absence (0) of mutations for each of the studied genes in a specific patient.

[0113] For each molecular genetic cluster k, the linear combination (logit) S is calculated k as the sum of the products of gene weights by the corresponding mutation values ​​(1 - there is a mutation, 0 - there is no mutation) plus the bias for a given cluster.

[0114] The calculation of the logit of each molecular genetic cluster is performed using formulas (1), (2), (3), (4).

[0115]

[0116]

[0117]

[0118] The resulting logits for molecular genetic clusters are then converted into probabilities using the softmax function, which ensures that values ​​are normalized to a range between 0 and 1 with the sum of probabilities equal to 1. The softmax function for cluster k is calculated as follows: P(k) = exp(S k ) / [exp(S1) + exp(S2) + exp(S3) + exp(S4)].

[0119] After determining the probability for each molecular genetic cluster, the patient belongs to the cluster with the highest probability.

[0120] The intercept parameter in a logistic regression model is an intercept term in the equation that determines the value of the target variable under the condition that all predictors are zero.

[0121] Thus, for clustering breast cancer patients, classification is performed by calculating the individual probability (score) for each molecular genetic cluster and then assigning the patient to the molecular genetic cluster with the highest probability value.

[0122] If there are no mutations in all analyzed genes, the patient's score for each molecular genetic cluster is determined solely by the intercept value, resulting in the patient being classified into the cluster with the highest value of this parameter. Analysis of the weight matrix shows that the highest intercept value is observed in the first cluster (3.518883), while in the second cluster it is -0.05177, in the third -0.20097, and in the fourth it reaches a minimum value of -3.26614. This pattern has important clinical significance, as it indicates that patients without detectable mutations in key genes such as TP53, PIK3CA, and other analyzed loci will be systematically classified as belonging to the first cluster.This reflects the baseline probability of belonging to a particular molecular subtype of breast cancer in the absence of specific genetic markers and may indicate the existence of a phenotypically and clinically significant group of patients characterized by a relatively low mutational load in the genes studied.

[0123] Based on the specific cluster affiliation of the tumor, the transition to the final stage is carried out, the formation of a personalized therapeutic strategy and clinical recommendations.

[0124] The therapeutic strategy and clinical recommendations are determined using a reference information table containing for each molecular genetic cluster: genomic alteration data, mutation type, such as single nucleotide polymorphisms, indels, coverage depth and allele frequency in the tumor, for Gene Set-1; transcriptomic alteration data, for Gene Set-2, gene copy number, for Gene Set-3, gene expression level; description of the associated histological and molecular genetic subtypes and clinical characteristics of the breast tumor; description of the personalized therapeutic strategy and clinical recommendations.

[0125] Table 2 of the reference information is given as an example, which presents the integrative classification of breast cancer.

[0126]

[0127]

[0128]

[0129]

[0130]

[0131] The following were used in the development of the system’s software package:

[0132] - electronic database “Database of clinical and pathomorphological databases of patients with breast cancer, intended for machine learning algorithms” [Certificate of state registration of the database No. 2025622625, date of state registration 06 / 18 / 2025].

[0133] - computer program “System for personalizing breast cancer therapy based on the interpretation of whole-genome sequencing data using machine learning methods” [Certificate of state registration of computer program No. 2025665590, state registration date 04.06.2025], modifying and “preparing” data from the database for use in the algorithms of the software package

[0134] - computer program “System for determining the probability of achieving a complete therapeutic response in patients with breast cancer based on machine analysis of clinical parameters from databases” [Certificate of state registration of computer program No. 2025664447, date of state registration 04.06.2025].

[0135] Implementation examples.

[0136] Example 1. Determination of a molecular genetic cluster based on the mutational profile, copy number variation (CNV) and expression of selected genes, assessment of their impact on the therapeutic response for a patient with the identifier: xxxx-xxxx-xxx1.

[0137] The laboratory research data supplied to the input of the calculation model are presented in Tables 3, 4, 5.

[0138] Table 3 shows the detected single nucleotide substitutions (SNPs) and short insertions and deletions (indels).

[0139] Table 4 shows the detected copy number variants (CNVs).

[0140] Table 5 shows the gene expression level (TRM).

[0141]

[0142]

[0143] Based on the analysis of significant genetic features, the cluster assigned was: CLUSTER 1. Luminal, lobular morphology.

[0144]

[0145]

[0146] Cluster description.

[0147] Cluster I is characterized by frequent mutations in the GATA3 and CDH1 genes and the absence of mutations in the TP53 and PIK3CA genes, which corresponds to the luminal phenotype of the tumor.

[0148] The cluster exhibits significant focal amplifications of IL19, IL20 and IL24, indicating activation of cytokine signaling.

[0149] Gene expression signature - upregulated for GATA3, MLPH, TBC1D9, SLC39A6 and downregulated for ENO1, PKM, STMN1, YBX1 - consistent with well-differentiated luminal subtype and low proliferative / metabolic activity.

[0150] The effect on the therapeutic effect is shown in Table 6.

[0151]

[0152]

[0153] Forecast.

[0154] Cluster 1 lobular tumors tend to present as invasive lobular carcinoma (ILC). ILCs account for about 10% of breast cancer, are usually ER+, but have a unique clinical behavior - a more indolent course, often multicentricity and poor response to chemotherapy [Wilson N, Ironside A, Diana A and Oikonomidou O (2021) Lobular Breast Cancer: A Review. Front. Oncol. 10:591399. doi: 10.3389 / fonc.2020.591399]. ILCs greatly benefit from endocrine therapy and have long-term disease control with antiestrogen treatment [Wilson N, Ironside A, Diana A and Oikonomidou O (2021) Lobular Breast Cancer: A Review. Front. Oncol. 10:591399. doi: 10.3389 / fonc.2020.591399].

[0155] When a tumor is classified as Cluster 1, it is advisable to prioritize endocrine therapy as the primary systemic treatment, avoiding chemotherapy whenever possible, especially in cases with low genomic risk. Thus, identifying a Cluster 1 tumor may spare the patient the unnecessary toxicity of chemotherapy in borderline cases.

[0156] Example 2. Determination of a molecular genetic cluster based on the mutational profile, copy number variation (CNV) and expression of selected genes, assessment of their impact on the therapeutic response for a patient with the identifier: xxxx-xxxx-xxx2.

[0157] The laboratory research data supplied to the input of the calculation model are presented in Tables 7, 8, 9.

[0158] Table 7 shows the detected single nucleotide substitutions (SNPs) and short insertions and deletions (indels).

[0159] Table 8 shows the detected copy number variants (CNVs).

[0160] Table 9 shows the gene expression level (TRM).

[0161]

[0162] Based on the analysis of significant genetic features, the cluster assigned was: CLUSTER II. Basal-like | HER2-enriched, ductal morphology.

[0163]

[0164] Cluster description.

[0165] Cluster II is characterized by the presence of a TP53 driver mutation, as well as an AMPD1 mutation. The PIK3CA mutation is absent. The mutational profile of Cluster 2 (TP53 ∧ mut, PIK3CA ∧ WT) is typical for basal-like and HER2-enriched tumors.

[0166] The cluster exhibits MYC amplification and PTEN gene deletion, which are frequently observed in basal-like breast cancer, indicating activation of the PI3K pathway despite the absence of PIK3CA mutation.

[0167] The gene expression signature of upregulation of basal cytokeratins (KRT5, KRT14, KRT17), EGFR, and proliferation / cell cycle genes (TOR2A, STMN1) and downregulation of luminal markers (ESR1, PGR, FOXA1, GATA3, MUC1) may signal increased proliferative activity and decreased tumor differentiation.

[0168] The effect on the therapeutic effect is shown in Table 10.

[0169]

[0170]

[0171] Forecast.

[0172] Cluster 2 corresponds to highly proliferative tumors caused by the TP53 mutation, which include basal-like and HER2-enriched breast cancers. This cluster corresponds to high-grade tumors with a significant risk of early recurrence.

[0173] Cluster II shows a significantly lower response to hormonal drugs - namely adjuvant Tamoxifen and Letrozole (p<0.01 for each) [Grote I, Battels S, Kandt L, et al. TP53 mutations are associated with primary endocrine resistance in luminal early breast cancer. Cancer medicine (Maiden, MA). 2021; 10(23): 8581-8594. https: / / doi.org / 10.1002 / cam4.4376].

[0174] Cluster 2 has a lower response to capecitabine chemotherapy (p < 0.05). This may reflect the chemoresistance of some triple-negative tumors or be related to TP53 deficiency, which impairs the effectiveness of capecitabine as an antimetabolite. Therefore, standard chemotherapy (anthracyclines / taxanes) should be considered.

[0175] For HER2-positive cases in Cluster 2, anti-HER2 therapy (trastuzumab + / - pertuzumab) is necessary. In case of detection of early triple-negative breast cancer, the use of anti-PD-1 / L1 therapy may be recommended as treatment [Schmid P. et al. Overall survival with pembrolizumab in early-stage triple-negative breast cancer / / New England Journal of Medicine. - 2024. - V. 391. - No. 21. - C. 1981-1991. https: / / doi.org / 10.1056 / NEJMoa2409932].

[0176] Example 3. Determination of a molecular genetic cluster based on the mutational profile, copy number variation (CNV) and expression of selected genes, assessment of their impact on the therapeutic response for a patient with the identifier: xxxx-xxxx-xxx3.

[0177] The laboratory research data supplied to the input of the calculation model are presented in Tables 11, 12, 13.

[0178] Table 11 shows the detected single nucleotide substitutions (SNPs) and short insertions and deletions (indels).

[0179] Table 12 lists the detected copy number variants (CNVs).

[0180] Table 13 shows the gene expression level (TRM).

[0181]

[0182]

[0183] Based on the analysis of significant genetic features, the cluster was assigned: CLUSTER III, Luminal, various morphology.

[0184]

[0185]

[0186] Cluster description.

[0187] Cluster III is characterized by a PIK3CA driver mutation, as well as MAP3K1 and MED15 mutations. No TP53 mutations are observed in the cluster. This mutational profile (PIK3CA ∧ mut / TP53 ∧ WT) is an indicator of the presence of the luminal subtype of breast cancer.

[0188] Gene expression signature - upregulated for ESR1, PGR, GATA3, FOXA1, TFF1, XBP1, SCUBE2 and downregulated for STMN1, TOR2A, KPNA2 - indicates a possible enriched stromal signature or tumor-stroma interactions in these tumors, and is also an indicator of well-differentiated, low-proliferative luminal status.

[0189] The influence on the therapeutic effect is shown in Table 14.

[0190]

[0191]

[0192] Forecast.

[0193] Cluster III tumors demonstrate positive results with both endocrine therapy and chemotherapy. Cluster 3 showed higher sensitivity to standard chemotherapy: the combination of doxorubicin + cyclophosphamide + paclitaxel had higher response rates in Cluster 3 (p < 0.05 for each). Patients in Cluster 3 responded better to tamoxifen and exemestane (p < 0.05).

[0194] For patients in this cluster, treatment de-escalation may be considered: given their good response to endocrine therapy, some may not require chemotherapy if diagnosed early. The addition of CDK4 / 6 inhibitors should also be considered in advanced disease. Given the presence of the PIK3CA mutation, PI3K inhibitors, which are a key adjunct in endocrine resistance, should be considered.

[0195] Example 4. Determination of a molecular genetic cluster based on the mutational profile, copy number variation (CNV) and expression of selected genes, assessment of their impact on the therapeutic response for a patient with the identifier: xxxx-xxxx-xxxx4.

[0196] The laboratory research data supplied to the input of the calculation model are presented in Tables 15, 16, 17.

[0197] Table 15 shows the detected single nucleotide substitutions (SNPs) and short insertions and deletions (indels).

[0198] Table 16 lists the detected copy number variants (CNVs).

[0199] Table 17 shows the gene expression level (TRM).

[0200]

[0201]

[0202] Based on the analysis of significant genetic features, the cluster assigned was: CLUSTER IV. Luminal B-like | HER2-enriched, ductal morphology.

[0203]

[0204] Cluster Description:

[0205] Cluster IV is characterized by co-mutations in PIK3CA and TP53 and also showed evidence of a higher overall mutational burden—mutations were detected in the MUC17, OFD1, RALGPS1, and POU3F2 genes. This mutational profile suggests an increased likelihood of tumor invasion or metastasis.

[0206] The effect on the therapeutic effect is shown in Table 18.

[0207]

[0208]

[0209] Forecast.

[0210] Cluster 4 tumors will have slightly lower complete response rates to HER2 blockade and potentially earlier recurrences, in particular due to the PIK3CA mutation [Kim, Ju Won, et al. "PIK3CA mutation is associated with poor response to HER2-targeted therapy in breast cancer patients." Cancer Research and Treatment: Official Journal of Korean Cancer Association 55.2 (2023): 531–541. https: / / doi.org / 10.4143 / crt.2022.221]. In addition, they tend to be less sensitive to chemotherapy.

[0211] Cluster 4 patients, including both Luminal B-like (ER+, HER2+) and HER2-enriched (ER-, HER2+) subtypes, may benefit from new therapy combinations such as HER2 blockade in combination with PI3K pathway inhibitors. Even in hormone-negative breast cancer, treatment escalation or the addition of PI3K or AKT / mTOR inhibitors to standard therapy may provide optimal disease control [Guarneri V. et al.: PIK3CA Mutation in the ShortHER Randomized Adjuvant Trial for Patients with Early HER2+ Breast Cancer: Association with Prognosis and Integration with PAM50 Subtype. Clin Cancer Res 15 November 2020; 26 (22): 5843-5851. Cluster 4 patients require close monitoring due to their genomic predisposition to resistance.

Claims

1. A system for selecting antitumor therapy for breast cancer using genomic data, comprising a processor, random access memory, a long-term data storage device with a software package, input-output interfaces, and network communication means, characterized in that the system contains: 4×26 matrix of gene weights W = [[3.518883, -3.68215, -3.64135, -0.0379, 0.352602, -0.03618, 0.494444, 0.20169, 0.158118, 0.105565, -0.1241, 0.051482, -0.06702, 0.022258, 0.142287, 0.105093, 0.01345, 0.049623, -0.05263, -0.01629, 0.027151, 0.022393, -0.06247, -0.01593, 0.200114, 0.02975], [-0.05177, -3.15651, 3.689056, 0.016745, -0.41153, -0.09685, -0.21961, -0.24982, 0.088314, -0.05926, -0.04971, -0.03935, 0.006928, -0.02346, 0.070376, -0.02617, 0.040066, 0.038742, 0.148841, 0.059149, -0.00621, 0.005337, 0.062735, 0.064656, -0.08761, 0.046106], [-0.20097, 3.720687, -3.07543, -0.12832, 0.269606, 0.018014, 0.092575, 0.252909, -0.08035, -0.02177, 0.172667, -0.13865, 0.080254, -0.13274, -0.06047, 0.051832, -0.10939, 0.036242, -0.00127, 0.026676, -0.07229, -0.08424, 0.076258, -0.0006, -0.10111, -0.00236], [-3.26614, 3.117965, 3.027726, 0.149473, -0.21068, 0.115015, -0.3674, -0.20478, -0.16609, -0.02453, 0.001146, 0.126522, -0.02016, 0.133943, -0.1522, -0.13075, 0.055872, -0.12461,-0.09494, -0.06954, 0.051354, 0.056514, -0.07652, -0.04813, -0.0114, -0.0735]], where 4 is the number of molecular genetic clusters, 26 is the free intercept term of length 4 and the weights of twenty-five studied genes: PIK3CA, TP53, TTN, CDH1, MUC16, GATA3, MAP3K1, IGDCC4, PTEN, KMT2C, RYR2, SYNE1, NEB, HMCN1, RYR3, DST, SPTA1, FLG, CSMD2, DMD, CSMD3, OBSCN, MUC5B, TBX3, NCOR1;, a table of reference information containing for each molecular genetic cluster: data on genomic changes for each gene from the gene set: PIK3CA, TP53, TTN, CDH1, MUC16, GATA3, MAP3K1, IGDCC4, PTEN, KMT2C, RYR2, SYNE1, NEB, HMCN1, RYR3, DST, SPTA1, FLG, CSMD2, DMD, CSMD3, OBSCN, MUC5B, TBX3, NCOR1; data on transcriptome changes in genes for each gene from the gene set: IL19, IL20, IL24, MYC, PTEN, data on transcriptome changes in genes for each gene from the gene set: GATA3, MLPH, KRT5, KRT14, ESR1, PGT, FOX1A, FOX1A, ERBB2; description of associated histological and molecular genetic subtypes and clinical characteristics of breast tumors; description of personalized therapeutic strategy and clinical recommendations; a software package for clustering breast tumors and assigning a patient to a specific molecular genetic cluster, containing: a pre-trained logistic regression model for calculating the linear logit for each molecular genetic cluster k as the sum of the products of the gene weights by the corresponding values ​​of gene mutations fed to the input of the logistic regression model in the form of a vector of binary values ​​reflecting the presence, 1, or absence, 0, of mutations for each of the genes being studied, plus an offset for the given molecular genetic cluster; the softmax function to transform the logit obtained for each molecular genetic cluster into a probability, which ensures the normalization of values ​​in the range from 0 to 1 with the sum of probabilities equal to 1, and is calculated as P(k) = exp(S k ) / [exp(S₁) + exp(S₂) + exp(S₃) + exp(S₄)], where S K - logit of molecular genetic cluster k, k = 1, 2, 3, 4; a software element for assigning a patient to a molecular genetic cluster with the highest probability, with the determination of a personalized therapeutic strategy and clinical recommendations for this cluster based on the corresponding data in the reference information table.

2. The system according to claim 1, characterized in that the reference information table contains data on genomic changes, including data on the type of mutations, for example, single nucleotide polymorphisms, indels, coverage depth and allele frequency in the tumor.

3. The system according to claim 1, characterized in that the table contains data on transcriptome changes, including data on the number of gene copies and the type of changes for each gene from the set of genes: IL19, IL20, IL24, MYC, PTEN.

4. The system according to claim 1, characterized in that the table contains data on transcriptome changes, including data on the expression level for each gene from the set of genes: GATA3, MLPH, KRT5, KRT14, ESR1, PGT, FOX1A, FOX1A, ERBB2.

5. The system according to claim 1, characterized in that the reference information table contains descriptions of four subtypes of breast tumor: luminal subtype A, a hybrid phenotype between basal-like and HER2-enriched subtypes, luminal subtype A with an immunohistochemical profile, and a HER2-positive subtype represented by HER2-enriched or luminal B / HER2 variants.

6. The system according to paragraph 1, characterized in that monitors, printers, interfaces for connecting laboratory equipment and medical information systems are used as the output interface.

7. The system according to paragraph 1, characterized in that hard drives, solid-state drives, RAID arrays, or cloud data storage are used as long-term storage devices.