Mirna marker profiles of breast cancer subtypes from urine samples

The method uses miRNA markers from urine samples and machine learning algorithms to provide a non-invasive, cost-effective solution for diagnosing and classifying breast cancer subtypes, addressing the inefficiencies of current methods.

WO2025119847A1PCT designated stage expired Publication Date: 2025-06-12RWTH AACHEN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/084351
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-07
Filing Date
2024-12-02
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Current methods for diagnosing and classifying breast cancer are invasive, costly, and expose patients to radiation or strong magnetic fields, making them inefficient for early detection and subtype classification.

Method used

A method involving the identification of small-molecule markers, specifically miRNAs, from urine samples using sequencing and machine learning algorithms to form decision trees, allowing for non-invasive diagnosis and classification of breast cancer subtypes.

Benefits of technology

This method enables reliable, cost-effective, and non-invasive diagnosis and classification of breast cancer subtypes using urine samples, reducing the need for tissue biopsies and exposure to harmful radiation or fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000011_0001
    Figure IMGF000011_0001
  • Figure IMGF000029_0001
    Figure IMGF000029_0001
Patent Text Reader

Abstract

The present invention relates to a method for identifying small-molecule markers for diagnosing breast cancer and to a method for diagnosing breast cancer, as well as to a computer program for carrying out such methods and to a training dataset for training a computer program of this kind.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] miRNA marker profiles of breast cancer subtypes from urine samples

[0002] The present invention relates to a method for identifying small-molecule markers for diagnosing breast cancer and a method for diagnosing breast cancer, as well as a computer program for executing such a method and a training dataset for training such a computer program. Breast cancer is the most common cancer among women in Germany, accounting for approximately 30 percent of all cancer cases. Approximately 70,000 patients, predominantly women, are diagnosed with breast cancer annually. Men can also develop breast cancer, although men represent only approximately 1% of diagnosed breast cancer cases.

[0003] Currently, one in eight women will develop breast cancer in their lifetime. The risk increases with age. Younger women are rarely affected; the risk only increases from the age of 40 and particularly from the age of 50. The average age of diagnosis for breast cancer is around 64, several years younger than the average for all cancers, with one in four patients being younger than 55 and one in ten patients under 45. Although breast cancer is the most common type of cancer in women, it is not the most dangerous. If detected and treated early, most cases of breast cancer are curable. The mortality rate has been steadily declining for decades. Although over 18,000 women die of breast cancer each year, around 87 percent of all women diagnosed with breast cancer are still alive after five years.

[0004] This positive development is due, on the one hand, to improved early detection, which means that tumors are detected at an early, still easily treatable stage, but also to advances in therapy.

[0005] Early detection is therefore of paramount importance in the treatment of breast cancer. The earlier breast cancer is diagnosed, the greater the treatment success rate.

[0006] Early detection of breast cancer is currently based on various examination methods.

[0007] If breast cancer is initially suspected, a manual examination of the breasts and armpits is performed. While this is a suitable, inexpensive, and quick initial examination, it is a rather inaccurate method. A diagnosis can only be made with the help of additional procedures. Furthermore, classification into different breast cancer subtypes is not possible through a manual examination.

[0008] Imaging techniques are frequently used as additional examination methods. These often include X-rays, ultrasound, or magnetic resonance imaging. Tissue biopsies are also often taken and examined in the laboratory.

[0009] Although these imaging techniques are relatively precise, they require a significant amount of personnel, technology, and time, resulting in high costs. Furthermore, some patients are hesitant to undergo such examinations because they expose the body to X-rays or strong magnetic fields. Any metal in or on the body is also problematic when using magnetic fields, for example, in patients with metal-containing implants.

[0010] Furthermore, determining tumor size and spread alone often does not provide a sufficient estimate of the risk posed by a tumor. Molecular biological tests can help characterize the specific tumor and better assess its danger. They are also becoming increasingly important for selecting treatment and are leading to increasingly individualized treatment concepts.

[0011] Currently, breast cancer is divided into the subtypes Triple negative, Luminal A, Luminal B, and Her2 positive, which differ both in their clinical course and thus prognosis, but also in the type of therapy (chemotherapy, endocrine therapy, antibody-based anti-Her2 therapy and immunotherapy).

[0012] To classify a tumor into one of the breast cancer subtypes, a tissue biopsy is currently required, which involves immunohistochemical determination of estrogen, progesterone, Her2 receptor, and the proliferation marker Ki-67. Classification based on these methods also requires significant personnel, technical resources, and time, thus resulting in high costs.

[0013] Furthermore, breast cancer is characterized by high intratumoral heterogeneity with different expression patterns of the receptors, which complicates histological assessment.

[0014] However, for the treatment of breast cancer, it is essential that a diagnosis is made as early as possible and that as many patients undergo screening as possible. Therefore, there is a high demand for simple, cost-effective, and reliable screening methods that, if possible, avoid the need for tissue samples and do not expose patients to harmful radiation or strong magnetic fields.

[0015] Erbes et al., "Feasibility of urinary microRNA detection in breast cancer patients and its potential as an innovative non-invasive biomarker" BMC Cancer (2015) 15:193, DOI 10.1186 / s12885-015-1190-4, describe that individual miRNAs associated with breast cancer can be detected in urine and measured at different concentrations between breast cancer patients and healthy controls. Nine known breast cancer-associated miRNAs were used as a starting point, with four miRNAs described as suitable.

[0016] A few years later, Ritter et al., "Discovery of potential serum and urine-based microRNA as minimally invasive biomarkers for breast and gynecological cancer," Cancer Biomarkers 27 (2020) 225-242, DOI 10.3233 / CBM-190575, also described the possibility of detecting individual miRNAs as biomarkers for breast cancer in urine. The starting point was selected miRNAs. Individual miRNAs of these were described as suitable, but whether these will also prove to be biomarkers in large-scale studies remains to be validated.

[0017] Diagnosing breast cancer using urine samples would offer outstanding advantages in early detection and would be a simple examination method in which patients would only have to provide a urine sample and would not be exposed to any further external influences such as radiation or strong magnetic fields.

[0018] However, to date, only individual miRNAs have been described as suitable. However, individual miRNAs are far from sufficient for a reliable diagnosis, which also significantly influences further treatment. In the event of a misdiagnosis, patients would lose valuable time and reduce the success of treatment if the correct diagnosis is not made until years later.

[0019] Furthermore, a patient's individual therapy is strongly influenced by a comprehensive diagnosis. The selection of a specific treatment option, such as surgical removal of the tumor, radiation therapy, chemotherapy, anti-hormone therapy, or antibody therapy, cannot be made based on individual miRNAs. Rather, a multitude of different miRNAs, such as several hundred, would be necessary to make a reliable diagnosis.

[0020] The same considerations apply even more to the classification of breast cancer into one of the subtypes, since the classification influences further therapy and especially its success.

[0021] A primary object of the present invention was therefore to provide a simpler, cost-effective and at the same time reliable method for diagnosing breast cancer, in which, if possible, no tissue samples have to be taken and the patients to be examined are not exposed to further external influences such as radiation or strong magnetic fields.

[0022] A further object of the present invention was to provide a simpler, cost-effective and at the same time reliable method for classifying breast cancer into one of the subtypes, in which, if possible, no tissue samples have to be taken and the patients to be examined are not exposed to any further external influences such as radiation or strong magnetic fields.

[0023] The above objects are achieved according to the invention by a method for identifying small-molecule markers for diagnosing breast cancer, comprising the following steps: i) providing a plurality of urine samples from a plurality of subjects, wherein at least one of the plurality of subjects is suffering from breast cancer and preferably at least one urine sample is provided per subject, ii) sequencing the miRNAs from the plurality of urine samples to obtain a first plurality of miRNAs, wherein the first plurality of miRNAs comprises all the different miRNAs sequenced in the plurality of urine samples, iii) identifying a second plurality of miRNAs as small-molecule markers for diagnosing breast cancer from the first plurality of miRNAs, wherein the second plurality of miRNAs comprises fewer miRNAs than the first plurality of miRNAs, by forming and evaluating a plurality of decision trees,preferably through a machine learning algorithm.

[0024] Surprisingly, it was found that several hundred miRNAs could be identified as suitable markers for the diagnosis of breast cancer using the method according to the invention.

[0025] A particular advantage of the method according to the invention is that urine samples can be collected easily, eliminating the need for an invasive procedure such as biopsies. Another advantage is that the subject is not exposed to any further influences after the urine sample has been collected, and the resulting sample can be examined without the subject being present.

[0026] The term "small molecule markers," as used herein, describes biomarkers, i.e., measurable biological characteristics that can be used as reference points for health or disease. The small molecule markers described herein describe miRNA. The term "miRNA," as used herein, describes single-stranded ribonucleic acids, preferably non-coding ribonucleic acids, preferably with a length in the range of 10 to 50 nucleotides, preferably in the range of 12 to 45 nucleotides, particularly preferably in the range of 15 to 40 nucleotides, further preferably in the range of 18 to 35 nucleotides, especially preferably in the range of 19 to 30 nucleotides, most preferably in the range of 20 to 25 nucleotides.

[0027] Preferably, the term “subject” as used herein describes a human being, most preferably a woman.

[0028] The term “plurality of urine samples” as used herein preferably describes at least 2, preferably at least 3, preferably at least 4, preferably at least 5, preferably at least 6, preferably at least 6, preferably at least 7, preferably at least 8, preferably at least 9, preferably at least 10, preferably at least 12, preferably at least 15, preferably at least 20, preferably at least 25, preferably at least 30, preferably at least 40, preferably at least 50 urine samples.

[0029] The term "sequencing," as used herein, describes the determination of the nucleotide sequence of a molecule. Examples of methods for determining the nucleotide sequence of a molecule include Sanger sequencing; second-generation sequencing, such as pyrosequencing, hybridization sequencing, semiconductor sequencing, bridge synthesis sequencing, or two-base sequencing; or third-generation sequencing, such as nanopore sequencing.

[0030] The term "urine sample," as used herein, describes a quantity of urine collected for subsequent testing that contains miRNA. Preferably, one, the, or all urine samples are obtained by micturition. Preferably, the urine in one, the, or all urine samples is midstream urine. Preferably, the urine in one, the, or all urine samples is morning urine or spontaneous urine. Preferably, the urine in all urine samples is morning urine or spontaneous urine.

[0031] Preferably, the term "breast cancer," as used herein, describes breast cancer of a subtype selected from triple negative, luminal A, luminal B, and HER2 positive. Preferably, the term "triple negative," as used herein, describes breast cancer that lacks therapy-relevant expression of the estrogen receptor (ER), the progesterone receptor (PR), and the human epidermal growth factor receptor 2 (HER2). The defined thresholds for this are 1% positive tumor cells for ER and PR and 10% positive tumor cells for HER2, each determined by immunohistochemical methods.

[0032] Preferably, the term "luminal A," as used herein, describes breast cancer that exhibits therapy-relevant expression of the estrogen receptor (ER) and the progesterone receptor (PR), but no therapy-relevant expression of the human epidermal growth factor receptor 2 (HER2). The defined thresholds for this are 1% positive tumor cells for ER and PR and 10% positive tumor cells for HER2, each determined by immunohistochemical methods.

[0033] Preferably, the term "luminal B," as used herein, describes breast cancer that exhibits therapy-relevant expression of the estrogen receptor (ER) but no therapy-relevant expression of the human epidermal growth factor receptor 2 (HER2). Furthermore, breast cancer of the luminal B subtype exhibits high Ki-67 levels and / or no therapy-relevant expression of the progesterone receptor (PR). The defined thresholds for this are 1% positive tumor cells for ER and PR, 10% positive tumor cells for HER2, and 25% positive tumor cells for Ki-67, each determined by immunohistochemical methods.

[0034] Preferably, the term "Her2 positive," as used herein, describes breast cancer that exhibits therapy-relevant expression of the human epidermal growth factor receptor 2 (HER2). This is preferably determined using immunohistochemical methods and, where appropriate, fluorescence in situ hybridization (FISH), preferably with the immunohistochemical methods taking into account only staining signals of the cell membranes. The cutoff for Her2 is at least 10% Her2-positive tumor cells in immunohistochemical methods. From a cutoff of at least 30% Her2-positive tumor cells, a Her2-positive breast cancer (here a score of 3+) is diagnosed. In cases ranging from 10 to less than 30% (here a score of 2+), gene amplification is additionally detected by fluorescence in situ hybridization to establish a Her2-positive diagnosis.

[0035] The term "multiplicity of miRNAs," as used herein, describes a collection of different miRNAs. The term "different miRNAs," as used herein, refers to miRNA molecules with different sequences.

[0036] The term "wherein the first plurality of miRNAs comprises all different miRNAs sequenced in the plurality of urine samples," as used herein, describes that the first plurality of miRNAs comprises all different miRNAs identified in the urine samples as a whole, and not every miRNA in every urine sample needs to be identified. For example, the entire miRNA from each urine sample is sequenced, and the sequences obtained from all urine samples are aligned with each other, and any sequence that is not identical to any other sequence represents a different miRNA.

[0037] The term “identifying a second plurality of miRNAs [...] from the first plurality of miRNAs” as used herein describes that all miRNAs of the second plurality are present in the first plurality.

[0038] Preferably, sequencing the miRNAs from the plurality of urine samples in step ii) comprises extracting a plurality of miRNAs from the urine samples, wherein a sample containing extracted miRNAs is obtained from each of the urine samples. By extracting a plurality of miRNAs from the urine samples, a plurality of samples containing extracted miRNAs are thus obtained. Preferably, in this case, the miRNAs from the plurality of samples containing extracted miRNAs are sequenced to obtain a first plurality of miRNAs in step ii).

[0039] Methods for extracting miRNA from urine are familiar to those skilled in the art. Commercially available miRNA extraction kits can be used for this purpose, for example.

[0040] Identifying a second plurality of miRNAs from the first plurality of miRNAs is preferably done by forming and evaluating a plurality of decision trees.

[0041] These decision trees are preferably ordered, directed trees that serve to represent decision rules. The graphical representation as a tree diagram illustrates hierarchically successive decisions. A decision tree preferably consists of a root node, any number of inner nodes, and at least two leaves. There is preferably exactly one path between two nodes. Each node, including the root node, represents a logical rule, and each leaf represents an answer to the decision problem. To obtain a classification of an individual data object, one proceeds downwards along the tree from the root node. At each node, at least one attribute is queried, and a decision is made about the selection of the next node. This procedure is continued until a leaf is reached. The leaf corresponds to the classification.A tree preferably contains rules for answering exactly one question. The logical rule is preferably a mathematical operation relating to an attribute, in particular checking whether a certain attribute value is above or below a threshold. It is also possible for a logical rule to affect multiple attributes, i.e., to assign a data object to another node based on the attribute values ​​of several different attributes. It is also possible for the logical rule to encompass mathematical operations relating to multiple attributes in a single mathematical operation, and for the multiple attributes to be checked by the mathematical operation, for example, based on a threshold.

[0042] In a binary decision problem, there are only two answers for each logical rule, i.e., each node. However, the method according to the invention is not limited to binary decision trees; rather, each node of a decision tree can have any number of answers to the node's logical rule.

[0043] A data set is required to create a decision tree. This data set preferably contains a large number of individual data objects, each of which has one or more identical attributes with preferably different attribute values ​​and can be classified based on these attribute values. In addition, the classification of each data object is known. For each node of the decision tree, a logical rule is established based on one or more attributes of the data objects using the individual attribute values. Using this logical rule, the data objects of the data set can be divided into two or more states at each node. A state specifies a subset of data objects of the set of data objects contained in the data set. Thus, a state in turn describes a set of individual data objects, in particular one or more data objects.

[0044] To determine a logical rule for a node, it is advantageous to consider the information gained by dividing the data set by the logical rule. Entropy, for example, provides a calculable measure for this. Entropy is preferably defined as the expected value of the information content of a state:

[0045] Here H is the entropy, E the expected value, here from the information content / , where / (z) = - log2p z indicates the information content of an event z that occurs with probability p z occurs. Z is the set of all distinct events of a state.

[0046] An event z is preferably a specific attribute value of an attribute of a data object. To determine the probability p z, with which the event z occurs, the various attribute values ​​of this attribute of the data objects of a state are considered. The probability p z of an event z indicates how likely it is that, when randomly selecting a data object from the data objects of the state, a data object with the specific attribute value of the event z will be drawn. Thus, the entropy can be determined for a state comprising one or more data objects.

[0047] To calculate the information gain from dividing a set of individual data objects based on a logical rule, the entropy of the initial state of the set of individual data objects before the division based on a logical rule, as well as the individual entropies of the states resulting from the division, are considered. For example, the following formula provides a measure of the information gain:

[0048] IG = H (initial state) — Wi H(statei) zi=0

[0049] Here, IG is the information gain, H is the entropy, n is the number of all different states after dividing the set of individual data objects by applying the logical rule and w, a weighting of the state, based on the number of individual data objects of the state compared to the number of individual data objects of the initial state, in particular:

[0050] Number of individual data objects of state i

[0051] W; = -

[0052] Number of individual data objects in the initial state. This makes it possible to calculate a measure for a logical rule and compare different logical rules with each other. The greater the information gained after dividing the set of individual data objects by applying the logical rule, the better the logical rule.

[0053] By optimizing the information gain, a preferred logical rule can be determined for each node. In particular, the information gain for each of a multitude of possible logical rules is determined and compared. Optimizing the information gain when determining a logical rule for a node has a positive effect on the size of the decision tree and thus on the computing power required to traverse the decision tree. Optimized logical rules make a decision tree more compact and lead more quickly to a result for the classification of a new data object.

[0054] A decision tree according to the method of the invention thus comprises one or more logical rules by which a urine sample can be classified based on the miRNAs sequenced from it. This classification is preferably carried out into the categories: signs of breast cancer and no signs of breast cancer, in particular, categorization into one of the subtypes of breast cancer.

[0055] The aim is to assign a urine sample to a classification based on its miRNAs using a decision tree according to the method according to the invention.

[0056] "Machine learning" is a generic term for the "artificial" generation of knowledge from experience: An artificial system learns from examples and can generalize them after the learning phase is complete. To achieve this, algorithms in machine learning, also called "machine learning algorithms," build a statistical model based on training data and test it against test data. This means that the examples from the training data are not simply memorized, but patterns and regularities are recognized in the training data. This allows the system to evaluate even unknown data.

[0057] For the method according to the invention, a machine learning algorithm is preferably trained using a dataset. This dataset contains data on urine samples from preferably breast cancer patients and test subjects, in particular information on miRNAs sequenced from the urine samples, preferably the first plurality of miRNAs obtained in step ii) of the method according to the invention. The machine learning algorithm then preferably creates one or more decision trees based on the dataset with the aim of enabling a classification based on the information about the miRNAs into: signs of breast cancer and no signs of breast cancer, in particular a classification into the various subtypes of breast cancer.

[0058] It is a finding of the invention that such a method reveals miRNAs that are more relevant than others in the classification of urine samples. These miRNAs represent small-molecule markers for the diagnosis of breast cancer. Preferably, these miRNAs are or contain the second plurality of miRNAs identified in step iii) of the method according to the invention.

[0059] For this purpose, the logical rules in the nodes of the decision tree are examined in more detail. Over a large number of iterations, it was identified that certain miRNAs stand out, which are particularly frequently present in logical rules compared to other miRNAs and / or the logical rules in which these specific miRNAs are included exhibit a particularly high information gain. In particular, these specific miRNAs are or contain the miRNAs from the second plurality of miRNAs identified in step iii) of the method according to the invention.

[0060] Preferably, the first plurality of miRNAs comprises at least 3,000 different miRNAs, preferably at least 3,500 different miRNAs, particularly preferably at least 3,750 different miRNAs.

[0061] The second plurality of miRNAs preferably comprises a maximum of 20%, preferably a maximum of 15%, particularly preferably a maximum of 10% of the different miRNAs in the first plurality of miRNAs. The percentages refer to the different miRNAs, with 100% representing all different miRNAs in the first plurality of miRNAs. If the first plurality of miRNAs comprises 3,000 different miRNAs, specifying a maximum of 10% of the different miRNAs in the first plurality of miRNAs describes a maximum of 300 different miRNAs in the second plurality of miRNAs.

[0062] The second plurality of miRNAs preferably comprises at least 100 different miRNAs, particularly preferably at least 125 different miRNAs, further preferably at least 150 different miRNAs, especially preferably at least 175 different miRNAs, very particularly preferably at least 200 different miRNAs. The second plurality of miRNAs preferably comprises a maximum of 20%, preferably a maximum of 15%, particularly preferably a maximum of 10% of the different miRNAs of the first plurality of miRNAs and at least 100 different miRNAs, particularly preferably at least 125 different miRNAs, further preferably at least 150 different miRNAs, especially preferably at least 175 different miRNAs, very particularly preferably at least 200 different miRNAs.

[0063] Preferably, the machine learning algorithm is or includes a random forest algorithm.

[0064] A random forest algorithm is a machine learning algorithm that consists of several uncorrelated, in particular a large number of uncorrelated decision trees. In this context, uncorrelated means that the decision trees were formed independently of one another, in particular according to different logical rules and data sets. All decision trees grew under a certain type of randomization during the learning process. The learning process particularly includes the process of forming a decision tree. The individual trees are then combined to form an ensemble, the random forest. The results of the individual trees are summarized in the ensemble using an aggregation function. Preferably, an aggregation function consisting of: mean, median, or majority vote is used; however, the invention is not limited to this aggregation function, but also includes other aggregation functions.

[0065] For the learning process of the Random Forest algorithm, a dataset as described above, comprising information on miRNAs from urine samples, preferably from breast cancer patients and volunteers, is used. Preferably, each different miRNA represents an attribute, and its relative concentration, i.e., in relation to the total concentration of all miRNAs present, preferably all miRNAs from the first plurality of miRNAs, in a urine sample represents a corresponding attribute value.

[0066] In a preferably first step of the learning process of the Random Forest algorithm, a large number of second data sets, also called sub-data sets, are created from the original data set, the source data set. For this purpose, for each sub-data set, a number of individual data objects, for example individual urine samples, with the corresponding attribute values, for example their information about the corresponding miRNAs sequenced from the urine sample, are copied from the source data set. The data objects selected for the respective sub-data sets are determined randomly. It is advantageous if the total number of individual data objects in each of the sub-data sets matches the number of individual data objects in the source data set. In this case, a data object with corresponding attribute values ​​can be copied into one, several, or no sub-data sets.Preferably, each individual data object and its attribute values ​​are copied to at least one subset of the data set. It is also possible for a single data object and its attribute values ​​to be copied multiple times into the same subset of the data set. In particular, such a data processing process is a bootstrapping process.

[0067] In a preferably second step, the decision trees of the Random Forest algorithm are formed. A corresponding process was already explained above. The difference here is that the partial data sets are used to determine preferred logical rules, not the original data set. Preferably, an uncorrelated decision tree is formed for each partial data set. Particularly preferred is that not all individual data objects present in this partial data set are used for the decision tree learning process. The individual data objects used to determine preferred logical rules are preferably selected randomly. This reduces the probability of a possible correlation between the individual decision trees and has a positive effect on the accuracy of the algorithm.Assuming the same data objects are used to create all decision trees, the probability that the same preferred logical rules will occur in different decision trees increases, because the optimization process attempting to find a preferred logical rule uses the same dataset. This promotes correlation between the trees and leads to less accurate results, since the correlated decision trees give one possible classification path of the correlated decision trees more weight in the aggregation function than the many other paths of the uncorrelated decision trees.

[0068] Preferably, the number of individual data objects used to determine preferred logical rules is an integer in the range of the square root of the number of individual data objects in the source data set. "Square" here means that the absolute value of the square root is mathematically correctly rounded to the nearest integer, and the number of individual data objects used to determine preferred logical rules deviates from the rounded number by ±10, preferably by ±5, more preferably by ±3, and particularly preferably by ±1. Alternatively, it is also advantageous to use the logarithmic function instead of the square root function for calculating the number of individual data objects used to determine preferred logical rules.

[0069] To classify a new urine sample, the decision trees of the resulting multitude of uncorrelated decision trees are traversed using the information from the sequenced miRNAs of the new urine sample. Each individual decision tree classifies the new urine sample. As already described, the individual classifications of the individual decision trees are summarized and evaluated using an aggregation function.

[0070] It is a finding of the invention that, from the multitude of decision trees formed by such a random forest algorithm, certain miRNAs emerge that are more relevant than others in the classification of urine samples. These miRNAs represent small-molecule markers for the diagnosis of breast cancer. Preferably, these miRNAs are or contain the second plurality of miRNAs identified in step iii) of the method according to the invention.

[0071] For this purpose, the logical rules in the nodes of the plurality of decision trees are examined in more detail. Through a large number of iterations in which a random forest algorithm was used to generate ever-new plurality of decision trees, it was discovered that certain miRNAs stand out, which are particularly frequently present in logical rules compared to other miRNAs and / or the logical rules containing the specific miRNAs exhibit a particularly high information gain. In particular, these specific miRNAs are or contain the miRNAs from the second plurality of miRNAs identified in step iii) of the method according to the invention.

[0072] The terms "particularly frequent" and "particularly high information gain" are preferably evaluated based on an average of the respective values ​​of all sequenced miRNAs, especially the miRNAs of the first plurality of miRNAs. In particular, these terms denote an above-average occurrence in logical rules and an above-average information gain. However, it is also possible to make this selection based on defined thresholds. For example, a defined threshold of 75% means that a particular miRNA is contained in logical rules more often than 75% of all sequenced miRNAs and / or the average information gain of the logical rules in which this particular miRNA is contained is above the average information gain of the logical rules of 75% of all sequenced miRNAs.Such specific miRNAs that are particularly frequently included in the logical rules and / or have a particularly high, especially average, information gain of the logical rules in which they are included are also referred to as particularly relevant miRNAs.

[0073] This makes it particularly possible to limit the classification of a new urine sample to these particularly relevant miRNAs, or at least to restrict the number of possible miRNAs to be selected for logical rules. This has the advantage that miRNAs that are not among these particularly relevant miRNAs, i.e., miRNAs that do not stand out from other miRNAs due to a higher number of occurrences in logical rules and / or a higher information content of the logical rules in which they are included, are not considered in the selection of logical rules, thus reducing the duration and computational effort of a corresponding algorithm.

[0074] In addition to the Random Forest algorithm, it is also possible to carry out a method according to the invention using other machine learning algorithms that are also included in the invention. These other machine learning algorithms include, for example, logistic regression, decision trees and a support vector machine. It is also possible to carry out a method according to the invention with a deep-learning algorithm and / or a neural network. However, it is a finding of the invention that the Random Forest algorithm is advantageous over these other suitable machine learning algorithms. It has been found that the Random Forest algorithm is advantageous over the other suitable machine learning algorithms in particular with regard to the aspects of accuracy, hit rate, repeatability, precision and F-measure, a combination of accuracy and hit rate using the weighted harmonic mean.

[0075] Preferably, the step of identifying the second plurality of miRNAs in step iii) of the method according to the invention is carried out specifically for a subtype of breast cancer.

[0076] For the classification of a new urine sample into specific subtypes of breast cancer, the initial data set with which the machine learning algorithm, in particular the random forest algorithm, is trained comprises information about the subtypes of breast cancer for which classification is to be made possible. Preferably, the initial data set comprises information about the subtype of breast cancer for at least 10%, preferably at least 20%, preferably at least 30%, preferably at least 40%, preferably at least 50%, preferably at least 60%, preferably at least 70%, preferably at least 80%, preferably at least 90%, of the individual data objects. In particular, the initial data set comprises information about the subtype of breast cancer for each individual data object.

[0077] The machine learning algorithm, in particular the Random Forest algorithm, is trained in such a way that the algorithm no longer classifies a new urine sample into “breast cancer is present” and “breast cancer is not present”, but into the subtypes, for example “breast cancer subtype 1 is present”, “breast cancer subtype 2 is present” and “breast cancer is not present”.

[0078] It was recognized that different specific miRNAs emerge for different subtypes of breast cancer, which are particularly relevant for classification into these specific subtypes of breast cancer. As already described, the logical rules of the nodes of the plurality of decision trees were considered. In particular, it was found that in a large number of runs, in each of which a random forest algorithm creates a large number of decision trees, predominantly the same specific miRNAs emerged as particularly relevant for certain subtypes. Thus, the particularly relevant miRNAs determined by a method according to the invention, in particular the miRNAs of the second plurality of miRNAs, are particularly reproducible.

[0079] For each of the breast cancer subtypes studied, specific amounts of particularly relevant miRNAs were identified. In particular, it was found that these amounts differ in the number of miRNAs and / or the miRNAs contained in these amounts.

[0080] This finding makes it possible to specialize algorithms, especially random forest algorithms, for detecting a specific subtype of breast cancer, particularly by prioritizing the miRNAs particularly preferred for this subtype during the selection process of the logical rules of the nodes of the plurality of decision trees. The method according to the invention can thus be applied not only to the diagnosis of breast cancer but also to classify breast cancer into a specific subtype. As described herein, further treatment depends heavily on the respective breast cancer subtype.

[0081] Especially with the current classification into breast cancer subtypes, tissue biopsies of patients are necessary, which are then examined using methods such as immunohistochemistry. As described herein, such a procedure is laborious and complicated by high intratumoral heterogeneity.

[0082] Surprisingly, it was found that miRNAs in urine samples are suitable for classifying breast cancer into subtypes.

[0083] The method according to the invention can also be used to classify breast cancer into a specific subtype using a patient's urine sample. Such a method is particularly advantageous, as described herein, and represents a particularly simple procedure for patients.

[0084] Preferably, the method according to the invention comprises step iii) as described herein and for diagnosing breast cancer, and step iv) iv) identifying a third plurality of miRNAs as small molecule markers for diagnosing breast cancer from the first plurality of miRNAs, wherein the third plurality of miRNAs comprises fewer miRNAs than the first plurality of miRNAs, by forming and evaluating a plurality of decision trees, preferably by a machine learning algorithm, wherein the identifying is carried out specifically for a subtype of breast cancer.

[0085] Preferably, step iv) is carried out simultaneously with step iii).

[0086] Preferably, step iv) is carried out after step iii).

[0087] What is described herein for the second plurality of miRNAs also applies mutatis mutandis to the third plurality of miRNAs. In particular, everything described for one or more subtypes of breast cancer within the context of the second plurality of miRNAs also applies mutatis mutandis to the subtype(s) of breast cancer within the context of the third plurality of miRNAs. In particular, everything described for one or more process steps within the context of the second plurality of miRNAs also applies mutatis mutandis to the third plurality of miRNAs, where applicable.

[0088] Preferably, in the method according to the invention, the subtype of breast cancer is one of the following: triple negative luminal A, luminal B,

[0089] Her2 / new.

[0090] Preferably, the identification in step iii) or iv) of the method according to the invention is specific for breast cancer of the triple negative subtype.

[0091] Preferably, the identification in step iii) or iv) of the method according to the invention is specific for breast cancer of the luminal A subtype.

[0092] Preferably, the identification in step iii) or iv) of the method according to the invention is specific for breast cancer of the luminal B subtype.

[0093] Preferably, the identification in step iii) or iv) of the method according to the invention is specific for breast cancer of the Her2 / neu subtype.

[0094] Preferably, step iii) or iv) of the method according to the invention is carried out several times, wherein in each repetition of the step a further (i.e. third, fourth, fifth, sixth, seventh, etc.) plurality of miRNAs is identified, wherein the identifying is carried out specifically for a subtype of breast cancer, wherein the subtype in each repetition is different from the subtype(s) in the previous repetition(s). Preferably, step iii) comprises the steps

[0095] 111.1) Identifying a second plurality of miRNAs as small molecule markers for diagnosing breast cancer from the first plurality of miRNAs, wherein the second plurality of miRNAs comprises fewer miRNAs than the first plurality of miRNAs, by forming and evaluating a plurality of decision trees, preferably by a machine learning algorithm, wherein the identifying is carried out specifically for a subtype of breast cancer,

[0096] 111.2) Identifying a third plurality of miRNAs as small molecule markers for diagnosing breast cancer from the first plurality of miRNAs, wherein the second plurality of miRNAs comprises fewer miRNAs than the first plurality of miRNAs, by forming and evaluating a plurality of decision trees, preferably by a machine learning algorithm, wherein the identifying is carried out specifically for a subtype of breast cancer, wherein the subtype of breast cancer is different from the subtype from step iii.1),

[0097] 111.3) Identifying a fourth plurality of miRNAs as small molecule markers for diagnosing breast cancer from the first plurality of miRNAs, wherein the second plurality of miRNAs comprises fewer miRNAs than the first plurality of miRNAs, by forming and evaluating a plurality of decision trees, preferably by a machine learning algorithm, wherein the identifying is carried out specifically for a subtype of breast cancer, wherein the subtype of breast cancer is different from the subtypes from step iii.1) and iii.2),

[0098] 111.4) Identifying a fifth plurality of miRNAs as small molecule markers for diagnosing breast cancer from the first plurality of miRNAs, wherein the second plurality of miRNAs comprises fewer miRNAs than the first plurality of miRNAs, by forming and evaluating a plurality of decision trees, preferably by a machine learning algorithm, wherein the identifying is carried out specifically for a subtype of breast cancer, wherein the subtype of breast cancer is different from the subtypes from step iii.1), iii.2) and iii.3), iii.5) Identifying a sixth plurality of miRNAs as small molecule markers for

[0099] Diagnosis of breast cancer from the first plurality of miRNAs, wherein the second plurality of miRNAs comprises fewer miRNAs than the first plurality of miRNAs, by forming and evaluating a plurality of decision trees, preferably by a machine learning algorithm, wherein the identifying is performed specifically for a subtype of breast cancer, wherein the subtype of breast cancer is different from the subtypes of step iii.1), iii.2), iii.3) and iii.4).

[0100] Preferably, step iv) comprises the steps iv.1) identifying a third plurality of miRNAs as small molecule markers for diagnosing breast cancer from the first plurality of miRNAs, wherein the second plurality of miRNAs comprises fewer miRNAs than the first plurality of miRNAs, by forming and evaluating a plurality of decision trees, preferably by a machine learning algorithm, wherein the identifying is carried out specifically for a subtype of breast cancer, iv.2) identifying a fourth plurality of miRNAs as small molecule markers for diagnosing breast cancer from the first plurality of miRNAs, wherein the second plurality of miRNAs comprises fewer miRNAs than the first plurality of miRNAs, by forming and evaluating a plurality of decision trees, preferably by a machine learning algorithm, wherein the identifying is carried out specifically for a subtype of breast cancer, wherein the subtype of breast cancer is different from the subtype from step iv.1), iv.3) identifying a fifth plurality of miRNAs as small molecule markers for diagnosing breast cancer from the first plurality of miRNAs, wherein the second plurality of miRNAs comprises fewer miRNAs than the first plurality of miRNAs, by forming and evaluating a plurality of decision trees, preferably by a machine learning algorithm, wherein the identifying is carried out specifically for a subtype of breast cancer, wherein the subtype of breast cancer is different from the subtypes from step iv.1) and iv.2), iv.4) identifying a sixth plurality of miRNAs as small molecule markers for diagnosing breast cancer from the first plurality of miRNAs, wherein the second plurality of miRNAs comprises fewer miRNAs than the first plurality of miRNAs, by forming and evaluating a plurality of decision trees, preferably by a machine learning algorithm, wherein the identifying is carried out specifically for a subtype of breast cancer, wherein the subtype of breast cancer is different from the subtypes from step iv.1), iv.2) and iv.3), iv.5) Identifying a seventh plurality of miRNAs as small molecule markers for diagnosing breast cancer from the first plurality of miRNAs, wherein the second plurality of miRNAs comprises fewer miRNAs than the first plurality of miRNAs, by forming and evaluating a plurality of decision trees, preferably by a machine learning algorithm, wherein the identifying is performed specifically for a subtype of breast cancer, wherein the subtype of breast cancer is different from the subtypes from step iv.1), iv.2), iv.3) and iv.4).

[0101] What is described herein for the second plurality of miRNAs also applies accordingly to each further plurality of miRNAs, in particular the third, fourth, fifth, sixth and seventh plurality of miRNAs.

[0102] Preferably, steps iii.2) to iii.5) or iv.2) to iv.5) correspond to the repetition of steps iii) or iv) carried out multiple times, as described herein. The present invention further relates to a method for diagnosing breast cancer, comprising the following steps: a) providing a urine sample from a subject to be examined, b) extracting miRNAs from the urine sample to obtain a sample containing extracted miRNAs, c) determining the concentration of each of the miRNAs of the second plurality of miRNAs, identified according to a method according to the invention for identifying small molecule markers for diagnosing breast cancer, in the sample containing extracted miRNAs from step b), preferably the concentration in relation to the total concentration of all miRNAs present, preferably all miRNAs of the first plurality of miRNAs and

[0103] Traversing the plurality of decision trees formed according to a method according to the invention for identifying small molecular markers for diagnosing breast cancer in order to obtain a plurality of results based on the traversed decision trees, d) combining the plurality of results from step c) by majority decision or weighting of the results of the individual decision trees of the plurality of decision trees to obtain a breast cancer diagnosis.

[0104] What is described herein for the method for identifying small molecule markers for the diagnosis of breast cancer also applies mutatis mutandis to the method for the diagnosis of breast cancer, where applicable.

[0105] Preferably, the method according to the invention comprises step b.2) b.2) quantifying the various miRNAs from the sample containing extracted miRNA from step b).

[0106] Preferably, the method according to the invention comprises step b.2) b.2) sequencing and quantifying the various miRNAs from the sample containing extracted miRNA from step b). Preferably, the method according to the invention comprises steps bi) sequencing the various miRNAs from the sample containing extracted miRNA from step b), b.ii) quantifying the various miRNAs sequenced in step bi).

[0107] Methods for quantifying miRNA from urine are familiar to those skilled in the art. These can include, for example, arrays, preferably microarrays, qPCR methods, or commercially available miRNA assays.

[0108] Quantifying miRNA typically involves determining the amount of miRNA in a sample under test. The amount can be an absolute value or a relative value, i.e., compared to another quantified parameter, preferably compared to the amount of one or more other miRNAs.

[0109] If the quantification of a specific miRNA yields a relative value compared to several other miRNAs, the other miRNAs may also contain the specific miRNA. For example, the quantification of a specific miRNA in a sample may yield a relative value compared to the total miRNA content of the sample, where the total miRNA also contains the specific miRNA.

[0110] The term “determining the concentration of each of the miRNAs” as used herein preferably describes a single determination of the concentration for each of the miRNAs.

[0111] When determining the concentration of the miRNAs in relation to the total concentration of all miRNAs present, preferably all miRNAs of the first plurality of miRNAs, in step c), the concentration of the respective miRNA and the concentration of all miRNAs present, preferably all miRNAs of the first plurality of miRNAs, are determined and related to one another, preferably wherein the concentration of all miRNAs present, preferably all miRNAs of the first plurality of miRNAs, also includes the concentration of the respective miRNA. For example, the concentration of miRNA X and the concentration of all miRNAs of the first plurality of miRNAs are determined and related to one another, wherein the concentration of all miRNAs of the first plurality of miRNAs also includes the concentration of miRNA X.The term "merging by majority decision to obtain a breast cancer diagnosis" as used herein describes that the breast cancer diagnosis is made based on the majority of results in the plurality of results obtained in step c). If the breast cancer diagnosis involves determining whether or not breast cancer is present, for example, the decision is made based on the result that is obtained with more than 50% probability in the plurality of results obtained in step c). If the breast cancer diagnosis involves classifying the diagnosed breast cancer into a subtype of breast cancer, the decision is made based on the result that is obtained most frequently in the plurality of results obtained in step c), for example (relative majority).

[0112] If step iii) of the method according to the invention for identifying small-molecule markers for diagnosing breast cancer is repeated as described herein, step c) of the method according to the invention for diagnosing breast cancer is preferably also repeated for several or all of the multiplicity obtained in the repetition(s). If a third, fourth, fifth, or sixth multiplicity is identified in step iii) of the method according to the invention for identifying small-molecule markers for diagnosing breast cancer, the determination of the concentration and the traversal of the decision trees in step c) are preferably repeated for the third, fourth, fifth, or sixth multiplicity.

[0113] If the method according to the invention for identifying small molecular markers for diagnosing breast cancer comprises step iv) as described herein, step c) of the method according to the invention for diagnosing breast cancer is preferably also repeated for the identified third plurality.

[0114] If step iv) of the method according to the invention for identifying small molecular markers for diagnosing breast cancer is repeated as described herein, step c) of the method according to the invention for diagnosing breast cancer is preferably also repeated for several or all of the multiplicity obtained in the repeat(s). If a fourth, fifth, sixth, or seventh multiplicity is identified in step iv) of the method according to the invention for identifying small molecular markers for diagnosing breast cancer, the determination of the concentration and the running of the decision trees in step c) are preferably repeated for the fourth, fifth, sixth, or seventh multiplicity. The diagnosis of breast cancer preferably comprises the classification of the diagnosed breast cancer into a subtype of breast cancer.

[0115] Preferably the subtype of breast cancer is one of the following: triple negative luminal A, luminal B,

[0116] Her2 / new.

[0117] Preferably, for classifying the diagnosed breast cancer into a subtype of breast cancer, step iii) as described herein is repeated and step c) as described herein is repeated accordingly.

[0118] Preferably, for classifying the diagnosed breast cancer into a subtype of breast cancer, step iv) is carried out as described herein and repeated if necessary, and step c) is repeated accordingly as described herein.

[0119] The present invention also relates to a computer program for carrying out a method according to the invention for identifying small molecular markers for diagnosing breast cancer, as described herein.

[0120] The present invention also relates to a computer program for carrying out a method according to the invention for diagnosing breast cancer as described herein.

[0121] The present invention also relates to a training data set for training a computer program according to the invention as described herein.

[0122] Typically, such a training dataset contains information on concentrations, preferably relative concentrations, of various miRNAs in patients with breast cancer and healthy controls. Preferably, such a dataset contains information on concentrations, preferably relative concentrations, of various miRNAs in patients with luminal A breast cancer.

[0123] Preferably, such a dataset contains information on concentrations, preferably relative concentrations, of various miRNAs in patients suffering from Luminal B breast cancer.

[0124] Preferably, such a dataset contains information on concentrations, preferably relative concentrations, of various miRNAs in patients suffering from Her2 / neu breast cancer. Preferably, such a dataset contains information on concentrations, preferably relative concentrations, of various miRNAs in patients suffering from triple-negative breast cancer.

[0125] Preferably, such a dataset contains information from samples from patients of different ages. Preferably, all patients and all controls included in the dataset are female.

[0126] In the following, the present invention is explained in more detail using selected examples.

[0127] Examples

[0128] Example 1 : miRNA for the diagnosis of breast cancer

[0129] 81 urine samples from patients with different subtypes of breast cancer and healthy controls were collected.

[0130] The miRNA was extracted as described below:

[0131] RNA isolation was performed using the Qiagen miRNeasy Mini Kit (#217004) according to the instructions in the user manual. Samples were isolated from a 5 ml aliquot to ensure equal volume for each sample. After isolation, the isolated RNA was stored at -80 °C until cDNA transcription.

[0132] The extracted miRNA was then sequenced using next generation sequencing.

[0133] Using a random forest algorithm, the following 270 miRNAs were identified that are suitable for diagnosing breast cancer: hsa.let.7d.3p, hsa.miR.17.5p, hsa.miR.20a.5p, hsa.miR.22.3p, hsa.miR.23a.5p, hsa.miR.23a.3p, hsa.miR.26b.5p, hsa.miR.96.5p, hsa.miR.99a.3p, hsa.miR.197.3p, hsa.miR.129.5p, hsa.miR.30c.2.3p, hsa.miR.30d.3p, hsa.miR.10b.5p, hsa.miR.203a.3p, hsa.miR.30b.5p, hsa.miR.141 ,5p, hsa.miR.191 ,5p, hsa.miR.191 ,3p, hsa.miR.126.3p, hsa.miR.154.3p, hsa.miR.194.5p, hsa.miR.195.5p, hsa.nniR.200a.3p, hsa.miR.99b.5p, hsa.miR.296.5p, hsa.miR.130b.3p, hsa.miR.30e.3p, hsa.miR.365a.5p, hsa.miR.135b.5p, hsa.miR.324.3p, hsa.miR.339.5p, hsa.miR.424.5p, hsa.miR.429, hsa.miR.503.5p, hsa.miR.505.5p, hsa.miR.508.3p, hsa.miR.92b.3p, hsa.miR.574.3p, hsa.miR.583, hsa.miR.548c.5p, hsa.miR.653.3p, hsa.miR.320b, hsa.miR.1301.3p, hsa.miR.889.3p, hsa.miR.147b.3p, hsa.miR.190b.3p, hsa.miR.744.5p, hsa.miR.1179, hsa.miR.1234.3p, hsa.miR.1207.5p, hsa.miR.1285.3p, hsa.miR.1290, hsa.miR.1293, hsa.miR.1249.3p, hsa.miR.1250.5p, hsa.miR.1251 ,3p, hsa.miR.1258, hsa.miR.513b.5p, hsa.miR.1469, hsa.miR.1908.3p, hsa.miR.3125, hsa.miR.3128, hsa.miR.3149, hsa.miR.3065.5p, hsa.miR.514b.5p, hsa.miR.4259, hsa.miR.4271 , hsa.miR.3922.3p, hsa.miR.3937, hsa.miR.4450, hsa.miR.4454, hsa.miR.4488, hsa.miR.4709.5p, hsa.miR.4709.3p, hsa.miR.4733.3p, hsa.miR.122b.5p, hsa.miR.4773, hsa.miR.4785, hsa.miR.4800.3p, hsa.miR.5088.3p, hsa.miR.5189.3p, hsa.miR.5584.5p, hsa.miR.6724.5p, hsa.miR.6727.3p, hsa.miR.6735.5p, hsa.miR.6752.3p, hsa.miR.6766.3p, hsa.miR.6801.3p, hsa.miR.6810.3p, hsa.miR.6819.3p, hsa.miR.6826.3p, hsa.miR.6831 ,5p, hsa.miR.6869.5p, hsa.miR.6884.3p, hsa.miR.6891 ,3p, hsa.miR.7108.3p, hsa.miR.7843.5p, hsa.miR.7851 ,3p, hsa.miR.8060, hsa.miR.8069, hsa.miR.8485, hsa.miR.1843, hsa.miR.10399.3p, hsa.miR.12136, hsa.let.7a.1 , hsa.let.7b, hsa.mir.18a, hsa.mir.21 , hsa.mir.27a, hsa.mir.30a, hsa.mir.32, hsa.mir.106a, hsa.mir.139, hsa.mir.10a, hsa.mir.10b, hsa.mir.34a, hsa.mir.181 b.1 , hsa.mir.196a.2, hsa.mir.210, hsa.mir.218.2, hsa.mir.15b, hsa.mir.124.2, hsa.mir.130a, hsa.mir.144, hsa.mir.154, hsa.mir.184, hsa.mir.186, hsa.mir.190a, hsa.mir.375, hsa.mir.342, hsa.mir.339, hsa.mir.345, hsa.mir.448, hsa.mir.429, hsa.mir.329.2, hsa.mir.409, hsa.mir.485, hsa.mir.181d, hsa.mir.1260b, hsa.mir.4770, hsa.mir.520g, hsa.mir.3168, hsa.mir.4774, hsa.mir.506, hsa.mir.3180.1 , hsa.mir.5004, hsa.mir.564, hsa.mir.3180.2, hsa.mir.4444.2, hsa.mir.571 , hsa.mir.3180.3, hsa.mir.5571 , hsa.mir.575, hsa.mir.3181 , hsa.mir.5572, hsa.mir.582, hsa.mir.3194, hsa.mir.664b, hsa.mir.548a.2, hsa.mir.3195, hsa.mir.6077, hsa.mir.589, hsa.mir.4296, hsa.mir.6090, hsa.mir.550a.1 , hsa.mir.4312, hsa.mir.6503, hsa.mir.596, hsa.mir.4313, hsa.mir.6510, hsa.mir.605, hsa.mir.4319, hsa.mir.6719, hsa.mir.610, hsa.mir.4251 , hsa.mir.6730, hsa.mir.622, hsa.mir.4327, hsa.mir.6748, hsa.mir.626, hsa.mir.4267, hsa.mir.6762, hsa.mir.631 , hsa.mir.4268, hsa.mir.6765, hsa.mir.33b, hsa.mir.4275, hsa.mir.6778, hsa.mir.650, hsa.mir.4283.1 , hsa.mir.6781 , hsa.mir.658, hsa.mir.4291 , hsa.mir.6793, hsa.mir.1264, hsa.mir.1972.2, hsa.mir.6797, hsa.mir.671 , hsa.mir.4283.2, hsa.mir.6805, hsa.mir.675, hsa.mir.3610, hsa.mir.6819, hsa.mir.874, hsa.mir.3612, hsa.mir.6835, hsa.mir.190b, hsa.mir.3648.1 , hsa.mir.6837, hsa.mir.922, hsa.mir.3654, hsa.mir.6847, hsa.mir.935, hsa.mir.3659, hsa.mir.6873, hsa.mir.939, hsa.mir.3679, hsa.mir.6893, hsa.nnir.941.3, hsa.mir.548z, hsa.mir.7150, hsa.mir.944, hsa.mir.4431, hsa.mir.7159, hsa.mir.663b, hsa.mir.4476, hsa.mir.7847, hsa.mir.1208, hsa.mir.548ak, hsa.mir.941.5, hsa.mir.1285.2, hsa.mir.4496, hsa.mir.10393, hsa.mir.1293, hsa.mir.4497, hsa.mir.10395, hsa.mir.548m, hsa.mir.4511 , hsa.mir.10397, hsa.mir.1277, hsa.mir.4530, hsa.mir.10523, hsa.mir.548i.3, hsa.mir.3977, hsa.mir.11399, hsa.mir.1827, hsa.mir.4634, hsa.mir.12127, hsa.mir.1914, hsa.mir.4649, hsa.mir.9902.2, hsa.mir.548q, hsa.mir.4654, hsa.mir.2277, hsa.mir.4668, hsa.mir.3124, hsa.mir.4689, hsa.mir.3129, hsa.mir.4707, hsa.mir.3130.2, hsa.mir.4709, hsa.mir.3149, hsa.mir.4713, hsa.mir.3159, hsa.mir.4724, hsa.mir.3160.2, hsa.mir.4745, hsa.mir.3163, hsa.mir.4756.

[0134] Example 2: miRNA for the diagnosis of Luminal A breast cancer

[0135] From the extracted and sequenced miRNA from Example 1, the following 188 miRNAs suitable for diagnosing Luminal A breast cancer were identified using a Random Forest algorithm: hsa.miR.29b.1 ,5p, hsa. miR.103a.1 ,5p, hsa.miR.10b.3p, hsa.miR.216a.3p, hsa.miR.222.5p, hsa. miR.200b.3p, hsa. miR.135a.3p, hsa. miR.142.3p, hsa. miR.125a.5p, hsa. miR.125a.3p, hsa. miR.185.5p, hsa. miR.302b.3p, hsa. miR.378a.3p, hsa. miR.151 a.5p, hsa. miR.335.5p, hsa. miR.335.3p, hsa. miR.425.5p, hsa. miR.449a, hsa. miR.433.5p, hsa. miR.412.3p, hsa. miR.193b.3p, hsa. miR.524.5p, hsa. miR.503.5p, hsa. miR.504.5p, hsa. miR.513a.3p, hsa. miR.532.5p, hsa. miR.556.5p, hsa. miR.550a.5p, hsa. miR.550a.3p, hsa. miR.671.5p, hsa. miR.1468.5p, hsa.miR.449c.5p, hsa. miR.889.5p, hsa. miR.942.5p, hsa. miR.1180.3p, hsa. miR.1236.3p, hsa.miR.548j.5p, hsa. miR.664a.5p, hsa. miR.1915.3p, hsa. miR.2681 ,3p, hsa. miR.3121 ,5p, hsa. miR.3130.5p, hsa. miR.3156.5p, hsa. miR.3160.5p, hsa. miR.3183, hsa. miR.3190.3p, hsa. miR.3197, hsa. miR.514b.3p, hsa. miR.3612, hsa. miR.3670, hsa. miR.3150b.5p, hsa.miR.548ac, hsa.miR.4476, hsa.miR.4477b, hsa.miR.4484, hsa.miR.4492, hsa. miR.548am.5p, hsa.miR.4536.3p, hsa. miR.219b.5p, hsa. miR.4698, hsa.miR.4707.5p, hsa. miR.203b.5p, hsa.miR.4796.3p, hsa. miR.5189.3p, hsa. miR.5585.3p, hsa. miR.4666b, hsa. miR.6072, hsa. miR.6088, hsa.miR.378j, hsa.miR.6132, hsa.miR.6754.3p, hsa.miR.6756.5p, hsa. miR.6759.3p, hsa. miR.6763.5p, hsa. miR.6795.3p, hsa.miR.6801 ,5p, hsa. miR.6819.3p, hsa. miR.6826.3p, hsa. miR.6869.5p, hsa.miR.6870.5p, hsa. miR.6871 ,3p, hsa. miR.6872.5p, hsa. miR.7113.5p, hsa. miR.7152.5p, hsa. miR.7843.5p, hsa.miR.4433b.5p, hsa. miR.7974, hsa. miR.7977, hsa. miR.8069, hsa. miR.9500, hsa.let.7b, hsa.mir.15a, hsa.mir.107, hsa.mir.148a, hsa.mir.196a.2, hsa.mir.30b, hsa.mir.137, hsa.mir.146a, hsa.mir.149, hsa.mir.190a, hsa.mir.195, hsa.mir.128.2, hsa.mir.151a, hsa.mir.423, hsa.mir.483, hsa.mir.523, hsa.mir.524, hsa.mir.516a.1 , hsa.mir.564, hsa.mir.567, hsa.mir.574, hsa.mir.593, hsa.mir.603, hsa.mir.605, hsa.mir.609, hsa.mir.626, hsa.mir.638, hsa.mir.767, hsa.mir.151b, hsa.mir.940, hsa.mir.941.3, hsa.mir.941.4, hsa.mir.942, hsa.mir.944, hsa.mir.1180, hsa.mir.1285.2, hsa.mir.1249, hsa.mir.513b, hsa.mir.320d.1 , hsa.mir.1914, hsa. mir.3116.1 , hsa.mir.3145, hsa.mir.3074, hsa.mir.3160.2, hsa.mir.3166, hsa.mir.3180.1 , hsa.mir.3180.3, hsa.mir.4299, hsa.mir.4304, hsa.mir.4309, hsa.mir.4265, hsa.mir.4273, hsa.mir.4280, hsa.mir.1972.2, hsa.mir.3610, hsa.mir.3649, hsa.mir.3671 , hsa.mir.3942, hsa.mir.5480.2, hsa.mir.548ad, hsa.mir.548x.2, hsa.mir.3689d.1 , hsa.mir.4533, hsa.mir.4687, hsa.mir.1343, hsa.mir.4708, hsa.mir.4709, hsa.mir.4721 , hsa.mir.4737, hsa.mir.4765, hsa.mir.4766, hsa.mir.4782, hsa.mir.5193, hsa.mir.5579, hsa.mir.5695, hsa.mir.6502, hsa.mir.6722, hsa.mir.6737, hsa.mir.6749, hsa.mir.6756, hsa.mir.6797, hsa.mir.6802, hsa.mir.6804, hsa.mir.6829, hsa.mir.6837, hsa.mir.6838, hsa.mir.6857, hsa.mir.7106, hsa.mir.651 1 a.3, hsa.mir.7150, hsa.mir.7976, hsa.mir.8087, hsa.mir.8485, hsa.mir.6859.4, hsa.mir.9899, ​​hsa.mir.9851 , hsa.mir.12127, hsa.mir.12136.

[0136] Example 3: miRNA for the diagnosis of Luminal B breast cancer

[0137] From the extracted and sequenced miRNA from Example 1, the following 196 miRNAs suitable for diagnosing Luminal B breast cancer were identified using a Random Forest algorithm: hsa.let.7a.5p, hsa.let.7b.5p, hsa.let.7e.5p, hsa.let.7f.5p, hsa.miR.22.5p, hsa.miR.25.5p, hsa.miR.98.5p, hsa.miR.99a.5p, hsa.miR.100.5p, hsa.miR.107, hsa.miR.197.3p, hsa.miR.30d.3p, hsa.miR.139.3p, hsa.miR.34a.5p, hsa.miR.181c.3p, hsa.miR.205.3p, hsa.miR.222.3p, hsa.miR.223.3p, hsa.let.7i.5p, hsa.miR.30b.3p, hsa. miR.125b.1 ,3p, hsa.miR.141 ,3p, hsa. miR.152.3p, hsa. miR.188.5p, hsa.miR.29c.3p, hsa.miR.200a.3p, hsa. miR.130b.3p, hsa. miR.378a.3p, hsa. miR.338.3p, hsa.miR.181 d.5p, hsa. miR.499a.5p, hsa. miR.506.5p, hsa. miR.508.3p, hsa. miR.514a.5p, hsa. miR.574.5p, hsa. miR.574.3p, hsa. miR.597.3p, hsa. miR.548a.5p, hsa.miR.629.5p, hsa.miR.671 ,5p, hsa.miR.378d, hsa. miR.891a.5p, hsa. miR.888.5p, hsa. miR.942.5p, hsa. miR.548k, hsa. miR.5481 , hsa. miR.1269a, hsa. miR.1469, hsa. miR.1913, hsa. miR.3122, hsa. miR.3143, hsa. miR.3200.5p, hsa. miR.3201 , hsa. miR.378c, hsa. miR.4288, hsa. miR.3614.5p, hsa. miR.3691 ,5p, hsa. miR.3936, hsa.miR.4433a.5p, hsa. miR.3960, hsa. miR.4679, hsa. miR.4689, hsa. miR.3529.3p, hsa.miR.4726.3p, hsa. miR.4768.3p, hsa.miR.4769.5p, hsa. miR.4999.5p, hsa. miR.5706, hsa.miR.6076, hsa.miR.6499.3p, hsa.miR.6716.5p, hsa. miR.6734.5p, hsa.miR.6751 ,3p, hsa.miR.6781 ,3p, hsa. miR.6813.5p, hsa. miR.6830.5p, hsa.miR.6833.3p, hsa.miR.6780b.3p, hsa.miR.6869.5p, hsa. miR.6872.3p, hsa. miR.6882.5p, hsa.miR.6890.5p, hsa. miR.6891 ,5p, hsa. miR.6894.3p, hsa. miR.7159.3p, hsa. miR.7705, hsa. miR.7843.5p, hsa. miR.1273h.5p, hsa. miR.1273h.3p, hsa. miR.8060, hsa. miR.3059.5p, hsa.let.7e, hsa.let.7f.1 , hsa.mir.16.1 , hsa.mir.197, hsa.mir.30d, hsa.mir.181 b.1 , hsa.mir.182, hsa.mir.196a.2, hsa.mir.218.1 , hsa.mir.122, hsa.mir.125b.1 , hsa.mir.135a.2, hsa.mir.138.2, hsa.mir.144, hsa.mir.128.2, hsa.mir.200a, hsa.mir.101 .2, hsa.mir.26a.2, hsa.mir.361 , hsa.mir.376c, hsa.mir.532, hsa.mir.564, hsa.mir.568, hsa.mir.598, hsa.mir.33b, hsa.mir.639, hsa.mir.653, hsa.mir.550a.3, hsa.mir.889, hsa.mir.924, hsa.mir.938, hsa.mir.1228, hsa.mir.1231 , hsa.mir.1234, hsa.mir.1299, hsa.mir.1249, hsa.mir.1251 , hsa.mir.1265, hsa.mir.5481 .2, hsa.mir.1279, hsa.mir.1538, hsa.mir.320d.2, hsa.mir.1825, hsa.mir.1914, hsa.mir.3124, hsa.mir.548s, hsa.mir.3129, hsa.mir.3163, hsa.mir.1260b, hsa.mir.3178, hsa.mir.3180.1 , hsa.mir.4299, hsa.mir.4315.1 , hsa.mir.4318, hsa.mir.4285, hsa.mir.3658, hsa.nnir.3917, hsa.mir.4425, hsa.mir.4434, hsa.mir.4492, hsa.mir.4496, hsa.mir.4498, hsa.mir.1269b, hsa.mir.4531 , hsa.mir.4641 , hsa.mir.4671 , hsa.mir.4672, hsa.mir.4683, hsa.mir.4685, hsa.mir.1343, hsa.mir.4690, hsa.mir.4739, hsa.mir.4754, hsa.mir.4776.1 , hsa.mir.4782, hsa.mir.4785, hsa.mir.5000, hsa.mir.5186, hsa.mir.5192, hsa.mir.5685, hsa.mir.5695, hsa.mir.1199, hsa.mir.6079, hsa.mir.6134, hsa.mir.548ay, hsa.mir.6738, hsa.mir.6756, hsa.mir.6789, hsa.mir.6797, hsa.mir.6811 , hsa.mir.6819, hsa.mir.6847, hsa.mir.6856, hsa.mir.6891 , hsa.mir.6892, hsa.mir.6894, hsa.mir.7106, hsa.mir.6511 b.2, hsa.mir.6511 a.4, hsa.mir.7854, hsa.mir.8067, hsa.mir.8083, hsa.mir.8088, hsa.mir.8485, hsa.mir.10400.

[0138] Example 4: miRNA for the diagnosis of Her2 / neu breast cancer

[0139] From the extracted and sequenced miRNA from Example 1, the following 162 miRNAs suitable for diagnosing Luminal A breast cancer were identified using a Random Forest algorithm: hsa.miR.27a.3p, hsa.miR.31 ,5p, hsa.miR.103a.3p, hsa.miR.30d.5p, hsa.miR.199b.5p, hsa.miR.203a.3p, hsa.miR.222.5p, hsa.miR.200b.3p, hsa.miR.125b.2.3p, hsa.miR.361 ,5p, hsa.miR.365a.5p, hsa.miR.378a.5p, hsa.miR.423.3p, hsa.miR.495.5p, hsa.miR.525.5p, hsa.miR.517b.3p, hsa.miR.518d.5p, hsa.miR.516a.5p, hsa.miR.519a.2.5p, hsa.miR.551b.5p, hsa.miR.596, hsa.miR.619.5p, hsa.miR.653.3p, hsa.miR.654.3p, hsa.miR.675.5p, hsa.miR.744.5p, hsa.miR.1182, hsa.miR.1287.3p, hsa.miR.1246, hsa.miR.1262, hsa.miR.3128, hsa.miR.378b, hsa.miR.3143, hsa.miR.3148, hsa.miR.3160.5p, hsa.miR.3173.3p, hsa.miR.3187.3p, hsa.miR.3191 ,3p, hsa.miR.378c, hsa.miR.3617.5p, hsa.miR.3714, hsa.miR.550b.2.5p, hsa.miR.4421 , hsa.miR.4446.5p, hsa.miR.4446.3p, hsa.miR.4455, hsa.miR.4530, hsa.miR.4716.5p, hsa.miR.4767, hsa.miR.4771 , hsa.miR.4776.5p, hsa.miR.4783.5p, hsa.miR.4787.5p, hsa.miR.4790.3p, hsa.miR.4793.3p, hsa.miR.5190, hsa.miR.664b.5p, hsa.miR.5585.3p, hsa.miR.6088, hsa.miR.6125, hsa.miR.6127, hsa.miR.548ay.5p, hsa.miR.6510.5p, hsa.miR.6729.5p, hsa.miR.6754.3p, hsa.miR.6788.3p, hsa.miR.6793.5p, hsa.miR.6810.3p, hsa.miR.6814.3p, hsa.miR.6823.3p, hsa.miR.6836.3p, hsa.miR.6845.3p, hsa.miR.6870.3p, hsa.miR.7706, hsa.miR.8073, hsa.miR.10392.5p, hsa.mir.30a, hsa.mir.100, hsa.mir.103a.1 , hsa.mir.197, hsa.mir.181 b.1 , hsa.mir.223, hsa.mir.27b, hsa.mir.30b, hsa.mir.140, hsa.mir.145, hsa.mir.186, hsa.mir.376c, hsa.mir.380, hsa.mir.328, hsa.mir.376b, hsa.mir.518d, hsa.mir.503, hsa.mir.505, hsa.mir.582, hsa.mir.586, hsa.mir.589, hsa.mir.605, hsa.mir.658, hsa.mir.300, hsa.mir.922, hsa.mir.944, hsa.mir.1228, hsa.mir.1237, hsa.mir.1914, hsa.mir.1915, hsa.mir.718, hsa.mir.2861 , hsa.mir.3120, hsa.mir.3141 , hsa.mir.3158.2, hsa.mir.3178, hsa.mir.3180.2, hsa.mir.3194, hsa.mir.4321 , hsa.mir.3622b, hsa.mir.3648.1 , hsa.mir.3654, hsa.mir.3943, hsa.mir.548ag.2, hsa.mir.4452, hsa.mir.4457, hsa.mir.4467, hsa.mir.4472.1 , hsa.mir.4485, hsa.mir.4521 , hsa.mir.4533, hsa.mir.3976, hsa.mir.4672, hsa.mir.4674, hsa.mir.4700, hsa.mir.4436b.1 , hsa.mir.4800, hsa.mir.5581 , hsa.mir.5690, hsa.mir.5701 .1 , hsa.mir.6125, hsa.mir.6510, hsa.mir.6511 a.1 , hsa.mir.6738, hsa.mir.6742, hsa.mir.6772, hsa.mir.6778, hsa.mir.6801 , hsa.mir.6805, hsa.mir.6810, hsa.mir.6818, hsa.mir.6833, hsa.mir.6844, hsa.mir.6873, hsa.mir.6879, hsa.mir.6894, hsa.mir.7108, hsa.mir.7150, hsa.mir.1273h, hsa.mir.7973.2, hsa.mir.6859.2, hsa.mir.9718, hsa.mir.3648.2, hsa.mir.3085, hsa.mir.6529, hsa.mir.12127.

[0140] Example 5: miRNA for the diagnosis of triple negative breast cancer

[0141] From the extracted and sequenced miRNA from Example 1, the following 207 miRNAs were identified using a Random Forest algorithm, which are suitable for the diagnosis of triple negative breast cancer: hsa.miR.16.5p, hsa.miR.20a.5p, hsa.miR.98.5p, hsa.miR.100.5p, hsa.miR.196a.5p, hsa.miR.197.5p, hsa.miR.129.1 ,3p, hsa.miR.30d.3p, hsa.miR.181 c.5p, hsa.miR.204.3p, hsa.miR.211 ,5p, hsa.miR.1 ,3p, hsa.miR.122.5p, hsa.miR.125b.2.3p, hsa.miR.193a.5p, hsa.miR.200c.3p, hsa.miR.30c.1,3p, hsa.miR.200a.3p, hsa.miR.34c.3p, hsa.miR.99b.3p, hsa.miR.30e.3p, hsa.miR.378a.3p, hsa.miR.151a.5p, hsa.miR.151a.3p, hsa.miR.324.3p, hsa.miR.422a, hsa.miR.495.3p, hsa.miR.514a.5p, hsa.miR.532.5p, hsa.miR.455.3p, hsa.miR.574.3p, hsa.miR.596, hsa.miR.651 ,5p, hsa.miR.41 1 ,5p, hsa.miR.454.3p, hsa.miR.449c.5p, hsa.miR.378d, hsa.miR.874.3p, hsa.miR.708.5p, hsa.miR.190b.5p, hsa.miR.1229.3p, hsa.miR.1236.3p, hsa.miR.1237.5p, hsa.miR.3170, hsa.nniR.3065.3p, hsa.miR.3197, hsa.miR.514b.3p, hsa.miR.378c, hsa.miR.3605.3p, hsa.miR.3622b.3p, hsa.miR.3911 , hsa.miR.548aa, hsa.miR.4433a.5p, hsa.miR.4435, hsa.miR.4454, hsa.miR.4474.3p, hsa.miR.4488, hsa.miR.4492, hsa.miR.4537, hsa.miR.3972, hsa.miR.4634, hsa.miR.4637, hsa.miR.3529.3p, hsa.miR.4771 , hsa.miR.4794, hsa.miR.4800.3p, hsa.miR.5189.3p, hsa.miR.664b.5p, hsa.miR.1295b.3p, hsa.miR.5787, hsa.miR.6499.5p, hsa.miR.6499.3p, hsa.miR.6716.5p, hsa.miR.6756.5p, hsa.miR.6793.5p, hsa.miR.6807.3p, hsa.miR.6824.5p, hsa.miR.6826.3p, hsa.miR.6833.5p, hsa.miR.6833.3p, hsa.miR.6862.3p, hsa.miR.6873.3p, hsa.miR.6882.5p, hsa.miR.6888.5p, hsa.miR.6891 ,3p, hsa.mir.7113.5p, hsa.miR.9500, hsa.let.7c, hsa.mir.24.1 , hsa.mir.24.2, hsa.mir.100, hsa.mir.106a, hsa.mir.34a, hsa.mir.199b, hsa.mir.205, hsa.mir.212, hsa.mir.217, hsa.mir.124.3, hsa.mir.152, hsa.mir.126, hsa.mir.138.1 , hsa.mir.149, hsa.mir.150, hsa.mir.185, hsa.mir.190a, hsa.mir.302a, hsa.mir.367, hsa.mir.380, hsa.mir.382, hsa.mir.491 , hsa.mir.518d, hsa.mir.564, hsa.mir.577, hsa.mir.579, hsa.mir.580, hsa.mir.590, hsa.mir.609, hsa.mir.619, hsa.mir.626, hsa.mir.636, hsa.mir.652, hsa.mir.887, hsa.mir.941.3, hsa.mir.944, hsa.mir.1228, hsa.mir.1237, hsa.mir.1250, hsa.mir.1268a, hsa.mir.320d.2, hsa.mir.2117, hsa.mir.718, hsa.mir.3124, hsa.mir.3126, hsa.mir.1273c, hsa.mir.3150a, hsa.mir.3074, hsa.mir.3159, hsa.mir.3160.2, hsa.mir.3176, hsa.mir.5787, hsa.mir.3180.3, hsa.mir.6069, hsa.mir.3195, hsa.mir.6078, hsa.mir.4299, hsa.mir.6133, hsa.mir.4257, hsa.mir.6510, hsa.mir.4281 , hsa.mir.6722, hsa.mir.4282, hsa.mir.6742, hsa.mir.4285, hsa.mir.6749, hsa.mir.4284, hsa.mir.6756, hsa.mir.3609, hsa.mir.6762, hsa.mir.3613, hsa.mir.6764, hsa.mir.3662, hsa.mir.6778, hsa.mir.3663, hsa.mir.6800, hsa.mir.3665, hsa.mir.6838, hsa.mir.3674, hsa.mir.6870, hsa.mir.3685, hsa.mir.6887, hsa.mir.3180.5, hsa.mir.7106, hsa.mir.378f, hsa.mir.7848, hsa.mir.4429, hsa.mir.7977, hsa.mir.4434, hsa.mir.8052, hsa.mir.3689e, hsa.mir.8089, hsa.mir.3689f, hsa.mir.10398, hsa.mir.4488, hsa.mir.9851 , hsa.mir.4497, hsa.mir.12116, hsa.mir.4501 , hsa.mir.12118, hsa.mir.548an, hsa.mir.4537, hsa.mir.3960, hsa.mir.3973, hsa.mir.4681 , hsa.mir.4687, hsa.mir.4731 , hsa.mir.4753, hsa.mir.4767, hsa.mir.4773.1 , hsa.mir.4778, hsa.mir.4785, hsa.mir.4999, hsa.mir.5000, hsa.mir.5004, hsa.mir.5684, hsa.mir.5702, hsa.nnir.5708, hsa.mir.5701 .2.

Claims

Claims 1 . A method for identifying small molecule markers for diagnosing breast cancer, comprising the following steps: i) providing a plurality of urine samples from a plurality of subjects, wherein at least one of the plurality of subjects is suffering from breast cancer and preferably at least one urine sample is provided per subject, ii) sequencing the miRNAs from the plurality of urine samples to obtain a first plurality of miRNAs, wherein the first plurality of miRNAs comprises all the different miRNAs sequenced in the plurality of urine samples, iii) identifying a second plurality of miRNAs as small molecule markers for diagnosing breast cancer from the first plurality of miRNAs, wherein the second plurality of miRNAs comprises fewer miRNAs than the first plurality of miRNAs, by forming and evaluating a plurality of decision trees, preferably by a machine learning algorithm.

2. The method according to claim 1, wherein the first plurality of miRNAs comprises at least 3000 different miRNAs, preferably at least 3500 different miRNAs, particularly preferably at least 3750 different miRNAs.

3. The method according to claim 1 or 2, wherein the second plurality of miRNAs comprises a maximum of 20% of the different miRNAs of the first plurality of miRNAs, but in particular at least 100 different miRNAs.

4. The method according to any one of the preceding claims, wherein the machine learning algorithm is a random forest algorithm or comprises a random forest algorithm.

5. The method of any preceding claim, wherein the step of identifying the second plurality of miRNAs is performed specifically for a subtype of breast cancer.

6. The method of claim 5, wherein the subtype of breast cancer is one of the list: triple negative luminal A, luminal B, Her2 / new.

7. A method for diagnosing breast cancer, comprising the following steps: a) providing a urine sample of a subject to be examined, b) extracting miRNAs from the urine sample to obtain a sample containing extracted miRNAs, c) determining the concentration of each of the miRNAs of the second plurality of miRNAs, identified according to a method according to the invention for identifying small molecule markers for diagnosing breast cancer, in the sample containing extracted miRNAs from step b), preferably the concentration in relation to the total concentration of all miRNAs present, preferably all miRNAs of the first plurality of miRNAs and Traversing the plurality of decision trees formed according to a method according to the invention for identifying small molecular markers for diagnosing breast cancer in order to obtain a plurality of results on the basis of the traversed decision trees, d) combining the plurality of results from step c) by majority decision or weighting of the results of the individual decision trees of the plurality of decision trees in order to obtain a breast cancer diagnosis.

8. The method of claim 7, wherein diagnosing breast cancer comprises classifying the diagnosed breast cancer into a subtype of breast cancer.

9. The method of claim 8, wherein the subtype of breast cancer is one of the list: triple negative luminal A, luminal B, Her2 / new.

10. Computer program for carrying out a method comprising the steps of claim 1.

11. A computer program for carrying out a method comprising the steps of claim 7.

12. Training data set for training a computer program according to claim 10.

Citation Information

Patent Citations

  • Methods to determine breast cancer risk by detecting the expression level of microRNAs (miRNAs)

    CN107429295B

  • Circulating microrna panel for the early detection of breast cancer and methods thereof

    WO2023014297A2

  • Breast cancer diagnostic and treatment

    WO2023170659A1