Use of circulating microRNA profiles for identification of BRCA1 or BRCA2 mutations

By detecting specific circulating microRNAs in the blood and combining them with statistical models, the problem of high cost of genetic testing has been solved, enabling efficient screening of BRCA1/2 mutation carriers and improving the efficiency of early cancer detection and prevention.

CN121532525APending Publication Date: 2026-02-13DANA FARBER CANCER INSTITUTE INC +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480030158.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-02
Filing Date
2024-05-02
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing genetic technologies cannot effectively screen individuals without a family history of cancer for BRCA1/2 mutations, resulting in high costs for genetic testing, limiting its widespread application, and impacting early cancer detection and prevention.

Method used

By detecting the circulating microRNA profile in the blood, especially the amount of 10 microRNAs such as hsa-miR-20b-5p, hsa-miR-19b-3p, and hsa-let-7b-5p, and using statistical models to predict the presence of BRCA1 or BRCA2 mutations, combined with dimensionality reduction technology and machine learning algorithms, the screening efficiency is improved.

Benefits of technology

It enables efficient screening of BRCA1/2 mutation carriers, reduces testing costs, and improves the efficiency of early cancer detection and prevention, especially for breast cancer, ovarian cancer, pancreatic cancer, and prostate cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121532525A_ABST
    Figure CN121532525A_ABST
Patent Text Reader

Abstract

A method of predicting the lifelong risk of having one or more cancers in a subject suspected of having a BRCA1 or BRCA2 mutation, comprising: obtaining a sample collected from the subject; determining the amount of circulating microRNAs selected from the group consisting of hsa-miR-20b-5p (SEQ ID NO: 4), hsa-miR-19b-3p (SEQ ID NO: 3), hsa-let-7b-5p (SEQ ID NO: 1), hsa-miR-320b (SEQ ID NO: 8), hsa-miR-139-3p (SEQ ID NO: 6), hsa-miR-30d-5p, (SEQ ID NO: 5), hsa-miR-17-5p (SEQ ID NO: 2), hsa-miR-182-5p (SEQ ID NO: 7), hsa-miR-421 (SEQ ID NO: 9), and hsa-miR-375-3p (SEQ ID NO: 10), and determining the amount of circulating microRNAs Hsa-miRNA-106b-5p (SEQ ID NO: 11), which is shown in the description; the kit is characterized in that the kit is prepared from hsa-miRNA-134-5p (SEQ ID NO: 12), hsa-miRNA-493-5p (SEQ ID NO: 13), hsa-miRNA-500a-3p (SEQ ID NO: 14), hsa-miR-1273h-3p (SEQ ID NO: 15), hsa-miR-4433a-3p (SEQ ID NO: 16), hsa-miR-4433b-5p (SEQ ID NO: 17), hsa-miR-485-3p (SEQ ID NO: 18) and hsa-miR-1304-3p (SEQ ID NO: 19), and the kit is prepared from the following raw materials: Hsa-miRNA-134-5p, comparing the amount of circulating microRNA determined in step (b) with a statistical model; and identifying the presence of at least one mutation in the BRCA1 or BRCA2 gene.
Need to check novelty before this filing date? Find Prior Art

Description

Cross Reference to Related Applications

[0001] This application claims priority to European application EP 23461572.2, filed on 2 May 2023, the contents of which are incorporated herein by reference in their entirety. Statement as to Federally Sponsored Research or Development

[0002] This invention was made with government support, based on approvals granted by the National Institutes of Health (NIH) under Approvals No. 1P50CA240243-01A1, No. 1R03CA283252-01, and No. 5P50CA240243-04. The government holds certain rights in this invention. SEQUENCE LISTING

[0003] As a separate part of this disclosure, this application contains a sequence list in XML format. The file name is 129319.01021_SL_ST26.xml, created on April 30, 2024, and is 17,745 bytes long, which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0004] This disclosure provides a novel method for identifying cyclic miRNA profiles. BRCA1 or BRCA2 Methods for detecting mutations and predicting the lifetime risk of developing cancer. BACKGROUND

[0005] Hereditary breast and ovarian cancer (HBOC) is the most common hereditary cancer syndrome, and the two most common genes mutated in HBOC are... BRCA1 and BRCA2 Both play a key role in DNA repair mediated by homologous recombination (HR). See Shulman, (2010). BRCA1 or BRCA1 (BRCA1 / 2) Germline mutations account for 10%-15% of ovarian cancer, 5%-10% of breast cancer, and 3%-5% of pancreatic and prostate cancer. The loss of HR, known as HR deficiency (HRD), impairs the cell's ability to repair double-strand DNA breaks, making the cell susceptible to mutagenesis induced by ionizing radiation and oxidative stress. BRCA1 / 2 Identifying mutation carriers is an essential component of cancer risk reduction strategies and provides opportunities for cascade testing of other family members. Mutation carriers have several opportunities for cancer prevention or interception, including risk-reducing salpingo-oophorectomy or mastectomy, hormonal chemoprevention, and enhanced surveillance programs such as MRI-based breast cancer screening.

[0006] BRCA1 / 2 Prevention or early detection of related cancers is based on BRCA1 / 2 Identification of mutation carriers. Currently... BRCA1 / 2 Genetic testing is only recommended for individuals or families with a known history of breast cancer, ovarian cancer, fallopian tube cancer, or primary peritoneal cancer, or for descendants of populations with a high mutation rate (e.g., Ashkenazi Jews). However, more than half of all those with [a specific genetic condition] are [affected by the mutation]. BRCA1 / 2 Carriers of the mutation have no family history of cancer, which will prompt them to undergo genetic testing. In the United States, an estimated 1 million... BRCA1 / 2 Of the mutation carriers, only 10% know their carrier status.

[0007] Because the cost of genetic testing makes universal testing impractical, "BRCAness" functional screening can improve the efficiency of early cancer detection and prevention efforts, regardless of personal or family history. To this end, this disclosure addresses the challenge of stratifying individuals as potentially or unlikely to have [a specific genetic predisposition / condition] by using microRNAs (miRNAs) circulating in the blood. BRCA1 / 2 The need for effective methods for mutation. SUMMARY OF THE DISCLOSURE

[0008] One implementation described herein is a method for predicting the lifetime risk of a subject suspected of having cancer or a BRCA1 or BRCA2 mutation having one or more cancers, comprising: (a) obtaining a sample collected from the subject; and (b) determining the amount of circulating microRNAs consisting of: hsa-miR-20b-5p (SEQ ID NO: 4), hsa-miR-19b-3p (SEQ ID NO: 3), hsa-let-7b-5p (SEQ ID NO: 1), hsa-miR-320b (SEQ ID NO: 8), hsa-miR-139-3p (SEQ ID NO: 6), hsa-miR-30d-5p (SEQ ID NO: 5), hsa-miR-17-5p (SEQ ID NO: 2), hsa-miR-182-5p (SEQ ID NO: 7), hsa-miR-421 (SEQ ID NO: 9), and hsa-miR-375-3p (SEQ ID NO: 1). 10); hsa-miRNA-106b-5p (SEQ ID NO: 11); hsa-miRNA-134-5p (SEQ ID NO: 12), hsa-miRNA-493-5p (SEQ ID NO: 13), hsa-miRNA-500a-3p (SEQ ID NO: 14), hsa-miR-1273h-3p (SEQ ID NO: 15), hsa-miR-4433a-3p (SEQ ID NO: 16), hsa-miR-4433b-5p (SEQ ID NO: 17), hsa-miR-485-3p (SEQ ID NO: 18) and has-miR-1304-3p (SEQ ID NO: 19); (c) compare the amount of circulating microRNA determined in step (b) with a statistical model; and (d) identify BRCA1 or BRCA2 The presence of at least one mutation in the gene.

[0009] In one embodiment, step (b) includes determining the amount of 10 microRNAs in the sample: hsa-miR-20b-5p (SEQ ID NO: 4), hsa-miR-19b-3p (SEQ ID NO: 3), hsa-let-7b-5p (SEQ ID NO: 1), hsa-miR-320b (SEQ ID NO: 8), hsa-miR-139-3p (SEQ ID NO: 6), hsa-miR-30d-5p (SEQ ID NO: 5), hsa-miR-17-5p (SEQ ID NO: 2), hsa-miR-182-5p (SEQ ID NO: 7), hsa-miR-421 (SEQ ID NO: 9), and hsa-miR-375-3p (SEQ ID NO: 10).

[0010] In another implementation, the cancer includes one or more of breast cancer, ovarian cancer, pancreatic cancer, or prostate cancer.

[0011] In another embodiment, the sample is selected from blood samples. In yet another embodiment, the blood sample is selected from the group consisting of plasma, serum, and whole blood.

[0012] In another implementation, the statistical model includes one or more models selected from the group consisting of linear discriminant analysis, logistic regression, multivariate adaptive regression splines, Naive Bayes, neural networks, support vector machines, functional trees, LAD trees, Bayesian networks, elastic network regression, and random forests. In one implementation, the statistical model includes a logistic regression model. In another implementation described herein, steps (b) and / or step (d) are performed using RNA sequencing.

[0013] In another implementation, the statistical model includes dimensionality reduction techniques. In yet another implementation described herein, a joint lasso method can be used for dimensionality reduction. In yet another implementation, the model includes joint lasso dimensionality reduction and sparse machine learning techniques.

[0014] In another implementation of the method, the method further includes performing genetic testing or genetic counseling.

[0015] In another embodiment of the method, the method further includes administering treatment to the subject, wherein the treatment is selected from the group consisting of: surgery, chemotherapy, immunotherapy, radiotherapy, hormone therapy, and stem cell transplantation.

[0016] In one implementation described herein, the subject is female.

[0017] Another implementation described herein is a method for identifying subjects suspected of having BRCA mutations, the method comprising: obtaining a sample collected from the subject; determining the amount of circulating microRNAs consisting of: hsa-miR-20b-5p (SEQ ID NO: 4), hsa-miR-19b-3p (SEQ ID NO: 3), hsa-let-7b-5p (SEQ ID NO: 1), hsa-miR-320b (SEQ ID NO: 8), hsa-miR-139-3p (SEQ ID NO: 6), hsa-miR-30d-5p (SEQ ID NO: 5), hsa-miR-17-5p (SEQ ID NO: 2), hsa-miR-182-5p (SEQ ID NO: 7), hsa-miR-421 (SEQ ID NO: 9), hsa-miR-375-3p (SEQ ID NO: 10); hsa-miRNA-106b-5p (SEQ ID NO: 4). NO:11); hsa-miRNA-134-5p (SEQ ID NO: 12), hsa-miRNA-493-5p (SEQ ID NO: 13), hsa-miRNA-500a-3p (SEQ ID NO: 14), hsa-miR-1273h-3p (SEQ ID NO: 15), hsa-miR-4433a-3p (SEQ ID NO: 16), hsa-miR-4433b-5p (SEQ ID NO: 17), hsa-miR-485-3p (SEQ ID NO: 18), and has-miR-1304-3p (SEQ ID NO: 19); compare the amount of circulating microRNAs identified in step (b) with a statistical model; identify the presence of at least one mutation in the BRCA1 or BRCA2 gene; and perform genetic testing on the subjects. Subjects suspected of having BRCA mutations can refer to subjects at risk of having BRCA mutations, regardless of any symptoms the subject may have.

[0018] In one implementation, step (b) includes determining the amount of 10 microRNAs in the sample: hsa-miR-20b-5p (SEQ ID NO: 4), hsa-miR-19b-3p (SEQ ID NO: 3), hsa-let-7b-5p (SEQ ID NO: 1), hsa-miR-320b (SEQ ID NO: 8), hsa-miR-139-3p (SEQ ID NO: 6), hsa-miR-30d-5p (SEQ ID NO: 5), hsa-miR-17-5p (SEQ ID NO: 2), hsa-miR-182-5p (SEQ ID NO: 7), hsa-miR-421 (SEQ ID NO: 9), and hsa-miR-375-3p (SEQ ID NO: 10).

[0019] In another implementation, the method also includes monitoring the subject's BRCA Related cancers. In another implementation, BRCA The associated cancers include one or more of breast cancer, ovarian cancer, pancreatic cancer, or prostate cancer.

[0020] In another implementation, the method further includes administering treatment to the subject, wherein the treatment is selected from the group consisting of: surgery, chemotherapy, immunotherapy, radiotherapy, hormone therapy, and stem cell transplantation.

[0021] In another implementation described herein, treatment is given to patients suspected of having... BRCAMethods for treating patients with related cancers, comprising: (a) obtaining samples collected from subjects; (b) determining the amount of circulating microRNAs consisting of: hsa-miR-20b-5p (SEQ ID NO: 4), hsa-miR-19b-3p (SEQ ID NO: 3), hsa-let-7b-5p (SEQ ID NO: 1), hsa-miR-320b (SEQ ID NO: 8), hsa-miR-139-3p (SEQ ID NO: 6), hsa-miR-30d-5p (SEQ ID NO: 5), hsa-miR-17-5p (SEQ ID NO: 2), hsa-miR-182-5p (SEQ ID NO: 7), hsa-miR-421 (SEQ ID NO: 9), hsa-miR-375-3p (SEQ ID NO: 10); hsa-miRNA-106b-5p (SEQ ID NO: 10). 11); hsa-miRNA-134-5p (SEQ ID NO: 12), hsa-miRNA-493-5p (SEQ ID NO: 13), hsa-miRNA-500a-3p (SEQ ID NO: 14), hsa-miR-1273h-3p (SEQ ID NO: 15), hsa-miR-4433a-3p (SEQ ID NO: 16), hsa-miR-4433b-5p (SEQ ID NO: 17), hsa-miR-485-3p (SEQ ID NO: 18), and has-miR-1304-3p (SEQ ID NO: 19); (c) compare the amount of circulating microRNA determined in step (b) with a statistical model; (d) identify BRCA1 or BRCA2 The presence of at least one mutation in a gene; and (e) administering to the subject a treatment selected from the group consisting of: surgery, chemotherapy, immunotherapy, radiotherapy, hormone therapy, and stem cell transplantation.

[0022] In another embodiment, a kit is provided comprising at least one test probe capable of specifically hybridizing with microRNAs selected from the group consisting of: hsa-miR-20b-5p (SEQ ID NO: 4), hsa-miR-19b-3p (SEQ ID NO: 3), hsa-let-7b-5p (SEQ ID NO: 1), hsa-miR-320b (SEQ ID NO: 8), hsa-miR-139-3p (SEQ ID NO: 6), hsa-miR-30d-5p (SEQ ID NO: 5), hsa-miR-17-5p (SEQ ID NO: 2), hsa-miR-182-5p (SEQ ID NO: 7), hsa-miR-421 (SEQ ID NO: 9), hsa-miR-375-3p (SEQ ID NO: 10); hsa-miRNA-106b-5p (SEQ ID NO: 10). 11); hsa-miRNA-134-5p (SEQ ID NO: 12), hsa-miRNA-493-5p (SEQ ID NO: 13), hsa-miRNA-500a-3p (SEQ ID NO: 14), hsa-miR-1273h-3p (SEQ ID NO: 15), hsa-miR-4433a-3p (SEQ ID NO: 15) ID NO: 16), hsa-miR-4433b-5p (SEQ ID NO: 17), hsa-miR-485-3p (SEQ ID NO: 18) and has-miR-1304-3p (SEQ ID NO: 19).

[0023] In one embodiment of the kit, at least one probe contains a detectable label. In another embodiment, the kit includes reagents for reverse transcription of microRNA molecules.

[0024] In various implementation schemes, each of the foregoing implementation schemes may be used in combination with any other stated implementation scheme. BRIEF DESCRIPTION OF DRAWINGS

[0025] FIG. 1AThis is a table showing the six serum biomaterial libraries used by the research group, based on: Brigham and Women's Hospital (BWH; Boston, MA; N=87), Dana-Farber Cancer Institute (DFCI; Boston, MA; N=200), including separate sample sets from DFCI Cancer Genetics and Prevention Center (CCGP; Boston, MA; N=162), Tata Medical Center (DGO; Kolkata, India; N=20), Pomeranian Medical University (IHCC; Szczecin, Poland; N=52), and the University of Pennsylvania (UPenn; Philadelphia, PA; N=132).

[0026] FIG. 1B This is a flowchart showing the two forms of differential expression analysis performed.

[0027] FIG. 2A This is the PCA representation of samples from all evaluation cohorts (without batch adjustment).

[0028] FIG. 2B This is the PCA representation of samples from all evaluation cohorts after batch adjustment (UPenn cohorts are used as an unmodified reference).

[0029] FIG. 3A It is shown BRCA State has a strong influence FIG. 2A The PCA representation of the expression profile in the unadjusted batch is shown.

[0030] FIG. 3B The PCA representation of samples from all evaluation cohorts shows that BRCA status strongly influences FIG. 2B The expression profiles of the adjusted batches.

[0031] FIG. 3C It shows the results based on phylogenetic relationships by superimposing the results of two data preprocessing strategies (on the original data). BRCA1 / 2 Scatter plot of differentially expressed (DE) miRNAs with mutations.

[0032] FIG. 3D It involves overlaying the results from two data preprocessing strategies and then displaying the data based on phylogenetic patterns after batch adjustment. BRCA1 / 2 Scatter plot of differentially expressed (DE) miRNAs with mutations.

[0033] FIG. 3EA heatmap of miRNA expression values ​​is shown, with unsupervised hierarchical clustering of samples from all subjects across five groups used for miRNA selection and model development, illustrating sample-based miRNA expression. BRCA1 / 2 Mutation clustering.

[0034] FIG. 4 A heatmap showing the expression values ​​of miRNAs used in the classification model of the UPenn group is presented.

[0035] FIG. 5A The ROC curve of the final logistic regression model is shown, with an area under the curve of 0.89 (95% CI: 0.87–0.93).

[0036] FIG. 5B This shows the predictions based on the training, test, and validation sets using the reference mutation state. BRCA1 / 2 Box plot of mutation probability.

[0037] FIG. 5C It shows the prediction in the context of menopausal status. BRCA1 / 2 Box plot of mutation probability.

[0038] FIG. 5D This shows the prediction in the context of having ovaries at the time of testing. BRCA1 / 2 Box plot of mutation probability. FIG. 5C and FIG. 5D The results show that neither menopausal status nor ovarian absence affects the prediction. BRCA1 or BRCA2 Mutation probability. In a box plot, the median is marked by the center line, the box lines indicate the first and third quartiles, and the whisker lines represent 1.5x IQR. FIG. 5B , FIG. 5C and FIG. 5D All N=653 samples were presented, and statistics were derived using all samples in the corresponding subgroups.

[0039] FIG. 6A The ROC curves for BRCA1 / 2 classification using joint lasso are shown.

[0040] FIG. 6B The confusion matrix is ​​shown using the maximum probability threshold t corresponding to the Youden exponent, as highlighted on the ROC curve. In the confusion matrix, 0 corresponds to negative (e.g., non-BRCA) and 1 corresponds to positive (e.g., BRCA).

[0041] FIG. 7AThe figure shows a comparison of the ROC of the joint lasso model (blue, upper curve) with the metadata component of the joint lasso model (red, lower curve), i.e., model training using only metadata (e.g., BRCA family history), and the x-axis component of the central plot in Figure 6. This figure shows the results using the model with all subjects and the complete data.

[0042] FIG. 7B The figure shows a comparison of the ROC of the joint lasso model (blue, upper curve) with the metadata component of the joint lasso model (red, lower curve), i.e., model training using only metadata (e.g., BRCA family history), and the x-axis component of the central plot in Figure 6. This figure illustrates the results using BRCA test subjects and the full data model.

[0043] FIG. 7C The figure shows a comparison of the ROC of the joint lasso model (blue, upper curve) with the metadata component of the joint lasso model (red, lower curve), i.e., model training using only metadata (e.g., BRCA family history), and the x-axis component of the central plot in Figure 6. This figure illustrates the results using models with all subjects and limited data.

[0044] FIG. 7D The figure shows a comparison of the ROC of the joint lasso model (blue, upper curve) with the metadata component of the joint lasso model (red, lower curve), i.e., model training using only metadata (e.g., BRCA family history), and the x-axis component of the central plot in Figure 6. This figure illustrates the results using BRCA test subjects and a model with limited data.

[0045] FIG. 8A The BRCA prediction results using 20 miRNAs and 5 metadata features are shown.

[0046] FIG. 8B The confusion matrix of BRCA prediction results using 20 miRNAs and 5 metadata features is shown.

[0047] FIG. 8C The classification AUC (where k2=5) is shown with different numbers of miRNAs (1≤k1≤20).

[0048] FIG. 9A The model performance regarding cancer history is shown. Top row: ROC curve. Bottom row: Sensitivity (TPR) and specificity (TNR) scores in the subgroup using the BRCA probability of t = 0.04 used in Figure 6.

[0049] FIG. 9BThe model performance with respect to age is shown. Top row: ROC curve. Bottom row: Sensitivity (TPR) and specificity (TNR) scores in the subgroups using the BRCA probability of t = 0.04 used in Figure 6.

[0050] FIG. 9C The model performance regarding race is shown. Top row: ROC curve. Bottom row: Sensitivity (TPR) and specificity (TNR) scores in subgroups using the BRCA probability of t = 0.04 used in Figure 6.

[0051] FIG. 10A The model performance for cancer history stratification using a more limited set of 20 miRNAs and 5 metadata features is shown.

[0052] FIG. 10B The model performance with respect to age stratification is shown using a more limited set of 20 miRNAs and 5 metadata features.

[0053] FIG. 10C The model performance regarding racial stratification is shown using a more limited set of 20 miRNAs and 5 metadata features.

[0054] FIG. 11A The BRCA score distribution after 10-fold cross-validation is shown when both miRNA and metadata are used to train a joint lasso model.

[0055] FIG. 11B Box plots of mean BRCA scores stratified by germline status and history of breast or ovarian cancer are shown. In the box plots, the p-value indicates whether there is a statistically significant difference in mean BRCA scores. The mean and standard deviation values ​​in the box plots are shown above, along with the corresponding fold change (FC) and p-value for each comparison. Here, the label “cancer” indicates subjects with a prior diagnosis of ovarian or breast cancer on record.

[0056] FIG. 12A The TSNE plot of miRNAs is shown, along with the distribution of non-BRCA, BRCA1, and BRCA2 subjects. The plot demonstrates a significant linear separation between non-BRCA and BRCA (1 or 2) subjects in the miRNA data. There is almost no separation between BRCA1 and BRCA2.

[0057] FIG. 12B The TSNE plot of miRNAs and metadata is shown, along with the distribution of non-BRCA, BRCA1, and BRCA2 subjects. There is almost no separation between BRCA1 and BRCA2.

[0058] FIG. 13AThe external validation results using PLCO data with complete data are shown.

[0059] FIG. 13B External validation results for PLCO data using the finite feature model are shown.

[0060] FIG. 13C The external validation results using complete PLCO data are shown. This is shown in comparison with... FIG. 13A The confusion matrix corresponding to the Youden exponent of the ROC curve trained in the training process.

[0061] FIG. 13D External validation results for PLCO data using the finite feature model are shown. This is illustrated with... FIG. 13B The confusion matrix corresponding to the Youden index of the training ROC curve. Note: When the BRCA score > 0.7, the relative risk curve becomes noisier due to the small sample size (e.g., most PLCO subjects have a BRCA score < 0.7).

[0062] FIG. 13E The external validation results of the PLCO using complete data are shown. For FIG. 13E The observed trend of relative cancer risk increasing with BRCA score was R = 0.80 (p < 0.0001).

[0063] FIG. 13F External validation results for PLCO data using the finite feature model are shown. For FIG. 13F R = 0.92 (p < 0.0001). Note: When the BRCA score > 0.7, the relative risk curve becomes noisier due to the small sample size (e.g., most PLCO subjects have a BRCA score < 0.7).

[0064] FIG. 14 A schematic diagram of the classification procedure is shown. The central plot shows the results of joint lasso dimensionality reduction of the test samples after 10x cross-validation. This is done for visualization to show how BRCA and non-BRCA samples are separated. When validating the model, the classifier is trained using only the training samples in the final step. The scatter plot on the left... , and Represents the TSNE component. DETAILED DESCRIPTION

[0065] This disclosure pertains to predicting the lifetime risk of developing one or more cancers or identifying individuals in a given study. BRCA1 or BRCA2 Methods for identifying mutations, and methods for identifying suspected mutations. BRCAMethods for subjects with mutations.

[0066] Although various embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, modifications, and substitutions will occur to those skilled in the art without departing from the present disclosure. It should be understood that various alternatives may be adopted to the embodiments of the present disclosure described herein.

[0067] It should be understood that the methods described in this disclosure are not limited to the specific methods and experimental conditions disclosed herein; therefore, methods and conditions may vary. It should also be understood that the terminology used herein is for the purpose of describing specific embodiments only and is not intended to be restrictive.

[0068] Generally, the terminology used herein in relation to cell and tissue culture, molecular biology, immunology, microbiology, genetics, and protein and nucleic acid chemistry and hybridization is that which is well known and commonly used in the art. Unless otherwise indicated, the methods and techniques provided herein are generally performed according to conventional methods well known in the art and as described in the various general and more specific references cited and discussed throughout this specification. Enzymatic reactions and purification techniques are performed according to the manufacturer's instructions for use, as is commonly done in the art or as described herein. The terminology, experimental procedures, and techniques used herein in relation to analytical chemistry, synthetic organic chemistry, and medicinal and pharmaceutical chemistry are that which are well known and commonly used in the art. Standard techniques are used for chemical synthesis, chemical analysis, drug preparation, formulation and delivery, and patient treatment.

[0069] Furthermore, unless otherwise stated, the experiments described herein utilize routine molecular and cell biology and immunology techniques within the scope of the art. Such techniques are well known to those skilled in the art and are well explained in the literature. See, for example, Ausubel et al., eds., Current Protocols in Molecular Biology, John Wiley & Sons, Inc., NY, NY (1987–2008), including all supplements; M.R. Green and J. Sambrook, Molecular Cloning: A Laboratory Manual (4th ed.); and Harlow et al., Antibodies: A Laboratory Manual, Chapter 14, Cold Spring Harbor Laboratory, Cold Spring Harbor (2013, 2nd ed.).

[0070] Unless otherwise defined herein, the scientific and technical terms used herein have the meanings commonly understood by one of ordinary skill in the art. In the event of any potential ambiguity, the definitions provided herein take precedence over any dictionary or external definition. Unless the context requires otherwise, singular terms shall include plural forms, and plural terms shall include singular forms. Unless otherwise stated, the use of “or” means “and / or”. The use of the term “including” and other forms such as “includes” and “included” is not restrictive.

[0071] Unless otherwise defined, all technical terms, symbols, and other scientific terms or terminology used herein are intended to have the meaning commonly understood by one of ordinary skill in the art to which this application pertains. In some instances, for clarity and / or ease of reference, terms are defined herein with the meaning commonly understood, and the inclusion of such definitions herein should not necessarily be construed as indicating a material difference from the content commonly understood in the art. The following references provide general definitions for many of the terms used in this disclosure to those skilled in the art: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd edition, 1994); The Cambridge Dictionary of Science and Technology (ed., Walker, 1988); The Glossary of Genetics, 5th edition, R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The HarperCollins Dictionary of Biology (1991). As used herein, unless otherwise specified, the following terms have the meanings attributed to them.

[0072] Unless otherwise specified or obvious from the context, the term “or” as used herein is understood to be inclusive. Unless otherwise specified or obvious from the context, the terms “a”, “an”, and “the” as used herein are understood to be singular or plural.

[0073] Unless otherwise specified or obvious from the context, as used herein, the term "approximately" is understood to mean within the normal tolerance range in the field, such as within 2 standard deviations of the mean. "Approximately" can be understood as within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, or 0.01% of the stated value. Unless otherwise specified from the context, all numerical values ​​provided herein are modified by the term "approximately".

[0074] As used herein, “administration” means applying the agent directly to a subject. In the implementation plan, the drug is administered by ingestion, inhalation, infusion, injection, or any other means, whether self-administered or administered by a clinician or other medical professional.

[0075] As used herein, the term "nucleic acid" refers to a polymer of two or more nucleotides or nucleotide analogs capable of hybridizing with complementary nucleic acids (such as ribonucleic acids having a methylene bridge between the 2'-O and 4'-C atoms of the ribose ring). As used herein, the term includes, but is not limited to, DNA, RNA, LNA, and PNA.

[0076] As used herein, the term “microRNA” or “miRNA” refers to a small non-coding ribonucleic acid (RNA) gene product that is between 19 and 26 nucleotides long and forms a hairpin secondary structure. The microRNAs described herein are named using the nomenclature listed in Ambros et al., RNA, March 2003, 9(3):277-9, which is incorporated herein by reference, and the sequences can be found at mirbase.org.

[0077] As used herein, the phrase "definite quantity" refers to the quantification of an analyte (such as microRNA) in a biological sample (e.g., a blood sample) using one or more detection techniques (such as qPCR, microarray detection, etc.) and quantification using methods known in the art. An analyte detected in a biological sample using a detection technique is considered "present." An analyte not detected in a biological sample using a detection technique is considered "absent."

[0078] As used herein, the term “bind” or “binding” refers to a non-covalent or covalent interaction between two molecules, such as two complementary nucleic acids.

[0079] As used herein, the term "specific hybridization" refers to a non-covalent interaction between a first nucleic acid molecule (e.g., a nucleic acid probe having a specific nucleotide sequence) and a second nucleic acid molecule (e.g., a microRNA having a nucleotide sequence complementary to that specific nucleotide sequence of the nucleic acid probe). Hybridization conditions have been described in the art and are known to those skilled in the art. In some embodiments, the conditions used to detect hybridization are suitable conditions for nucleic acid detection assays (e.g., microarrays, RT-PCR, or RT-qPCR). The likelihood of hybridization between two nucleic acids is related to the complementary nucleotide sequences between the two nucleic acids.

[0080] As used herein, the term "hybridization" refers to the annealing of a first single-stranded nucleic acid with a second complementary single-stranded nucleic acid, wherein the complementary nucleotides of the first and second nucleic acids pair via hydrogen bonding.

[0081] As used herein, the phrase "detecting probe binding" refers to the use of a detection method that allows determination of non-covalent or covalent interactions between a probe (e.g., a nucleic acid probe) and a target molecule (e.g., a target nucleic acid in a sample). For example, detecting probe binding in qPCR may include optically detecting the fluorescence of a self-quenching probe after binding to a complementary sequence of a target nucleic acid in a sample. In some embodiments, detecting probe binding may include detecting a nucleic acid intercalating agent to detect amplified double-stranded nucleic acids, such as fluorescent intercalating agents used in qPCR.

[0082] As used herein, the term "probe" refers to a molecule or complex used to determine the presence or absence and / or amount of microRNA in a sample (e.g., a blood sample). In some embodiments, the probe comprises a nucleic acid moiety (e.g., DNA, modified DNA, or modified RNA) capable of specifically hybridizing with microRNA or its complementary DNA (cDNA). In some embodiments, the probe comprises a sequence of at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 consecutive nucleotides that are identical to or complementary to the microRNA. In some embodiments, the probe further comprises a detectable label covalently or non-covalently conjugated to the nucleic acid moiety. Exemplary detectable labels include, but are not limited to, fluorophores, small molecules (e.g., small molecules of the avidin family), enzymes, antibodies or antibody fragments, or nucleic acid sequences not present in the subject in the form of a linker to the microRNA (e.g., barcode sequences). Thus, the probe may be a nucleic acid fluorophoretically labeled with a nucleotide sequence complementary to the nucleotide sequence of the microRNA.

[0083] As used herein, the term "PCR" refers to a polymerase chain reaction used to amplify a specific amount of target DNA. PCR relies on thermal cycling, which consists of repeated cycles of heating and cooling of the reaction, used for DNA denaturation, annealing, and enzymatic elongation of the amplified DNA. First, DNA strands are separated at high temperatures during a process known as DNA melting or denaturation. Next, the temperature is lowered, allowing primers and the target DNA strands to selectively bind or anneal, producing a template for DNA polymerase to amplify the target DNA. Then, at the operating temperature of the DNA polymerase, template-dependent DNA synthesis occurs. These steps are repeated to produce numerous copies of the target DNA.

[0084] As used herein, a "primer" refers to a short, single-stranded DNA sequence that selectively binds to a target DNA sequence and enables DNA polymerase to add a new deoxyribonucleotide at the 3' end. According to certain embodiments, the length of a forward primer is 18-35, 19-32, or 21-31 nucleotides. The nucleotide sequence of the forward primer is not limited, provided that it specifically hybridizes to part or all of the target site and that its melting temperature ("Tm") value is within the range of 50°C to 72°C, particularly within the range of 58°C to 61°C, and within the range of 59°C to 60°C. The nucleotide sequence of the primer can be manually designed to confirm the Tm value using a primer Tm prediction tool. Primer nucleotides may include nucleotide analogs and / or modified nucleotides, such as LNA or PNA.

[0085] As used herein, the term "RT-PCR" refers to reverse transcription polymerase chain reaction, a process used to amplify RNA. RNA molecules are reverse transcribed into complementary DNA (cDNA) using reverse transcriptase, and the resulting cDNA is then amplified using PCR.

[0086] As used herein, the term "RT-qPCR" refers to quantitative reverse transcription polymerase chain reaction, a variant of RT-PCR in which the amplification of cDNA during the RT-PCR process is quantitatively detected in real time using probes that detect the amplified target DNA. For example, in some embodiments, self-quenching nucleic acid probes are added to the reaction mixture. Self-quenching nucleic acid probes fluoresce only when they bind to the target sequence. At the end of each PCR cycle, the self-quenching probe binds to the amplified DNA, unquenches, and fluoresces upon exposure to a light excitation source. As DNA amplifies, the increased binding of the probe to the target results in an increase in the fluorescence of the self-quenching nucleic acid probe. Detecting the fluorescent probe after each amplification cycle allows for real-time measurement of the amplification process as the amount of nucleic acid probe bound to the amplified target DNA increases and fluoresces. In some embodiments, an intercalating dye probe is added to the reaction mixture that fluoresces upon interaction with the double-stranded nucleic acid. The increase in dye fluorescence during the amplification process allows for real-time measurement of DNA amplification as the amount of dye probe intercalation increases with the amount of amplified target DNA.

[0087] As used herein, the term "normalize" or "normalizing" refers to adjusting a first measurement (e.g., the level of a housekeeping gene) relative to a second measurement (e.g., the level of a housekeeping gene), wherein the first and second measurements are taken from the same sample (e.g., different portions of the same homogeneous sample), and wherein the second measurement is related to the quantity and / or quality of the sample. Normalization allows for obtaining a relative quantity of the first value that is unaffected by the quantity and / or quality of the sample, which may vary from individual sample preparations.

[0088] As used herein, the term "normalized microRNA" refers to a microRNA known to have a stable amount in a sample (e.g., a blood sample) and used to normalize measurements of the test microRNA in the sample. A single normalized microRNA can be used to normalize the measurement of the target microRNA in the sample, or the average of multiple microRNAs can be used for normalization. In some embodiments, normalization can be calculated as: amplification cycle number (average of the normalized microRNAs) - amplification cycle number (miRNA of interest).

[0089] As used herein, the term "test microRNA" refers to a microRNA whose presence or absence and / or quantity is determined, for example, for diagnostic purposes (e.g., using an algorithm). In some embodiments, the presence or absence and / or quantity of one or more test microRNAs may also be used for normalization purposes.

[0090] As used herein, the term "normalized probe" refers to a probe used to determine the presence or absence and / or amount of normalized microRNA in a sample (e.g., a blood sample). In some embodiments, the normalized probe includes a nucleic acid motif (e.g., DNA, modified DNA, or modified RNA) capable of specifically hybridizing with normalized microRNA or its complementary DNA (cDNA).

[0091] As used herein, the term "test probe" refers to a probe used to determine the presence or absence and / or amount of a test microRNA in a sample (e.g., a blood sample). In some embodiments, the test probe includes a nucleic acid motif (e.g., DNA, modified DNA, or modified RNA) capable of specifically hybridizing with the test microRNA or its complementary DNA (cDNA).

[0092] The phrase “reagents for amplifying DNA sequences” includes, but is not limited to: (1) thermostable DNA polymerases; (2) deoxyribonucleotide triphosphates (dNTPs); (3) buffer solutions that provide a suitable chemical environment for optimal activity, binding kinetics, and stability of the DNA polymerases; (4) divalent cations such as magnesium or manganese ions; and (5) monovalent cations such as potassium ions. Reagents may be provided in the form of solutions, concentrated solutions, or powders.

[0093] The phrase “reagents for reverse transcription of RNA molecules” includes, but is not limited to: reverse transcriptase; RNase inhibitors; primers for hybridization with nucleic acid sequences (such as RNA or DNA); primers for hybridization with adenosine oligonucleotides; and buffer solutions that provide a suitable chemical environment for optimal activity, binding kinetics, and stability of reverse transcriptase. Reagents may be provided in solution, concentrated solution, or powder form.

[0094] As used herein, the term “blood sample” refers to a quantity of blood taken from a subject, such as whole blood, or a component portion of the subject’s blood, such as plasma that lacks the cells typically found in whole blood (e.g., red blood cells, white blood cells, and platelets), or serum that is plasma that lacks fibrinogen and some clotting factors.

[0095] As used herein, the term "nucleic acid detection method" encompasses any method that can be used to detect the presence of nucleic acids, including sequencing (e.g., Gilbert sequencing, Sanger sequencing, SMRT sequencing, or next-generation sequencing), microarray detection, PCR, RT-PCR, real-time qPCR, and real-time RT-qPCR methods.

[0096] As used herein, the term "next-generation sequencing" refers to high-throughput parallel sequencing of short fragments of single-stranded nucleic acids attached to a slide or bead, such as by technologies from Illumina, Roche (454 sequencing), or Ion Torrent, Thermofisher. The incorporation of a single nucleotide into a single-stranded nucleic acid can be detected optically (via the fluorescence of the incorporated nucleotide) or by detecting hydrogen ions released during nucleotide incorporation (e.g., ion semiconductor sequencing).

[0097] As used herein, the term "artificial neural network" refers to a predictive model based on a computer-simulated set of connected neural units that loosely mimics a simple mathematical model of the brain. Artificial neural networks allow the identification of complex, nonlinear relationships between their response variables and their predictor variables. An artificial neural network can have one or more hidden layers, each containing one or more neurons; given two or more variables, these neurons interact to produce a prediction.

[0098] As used in this article, the term "cancer" generally refers to a class of diseases or conditions in which abnormal cells divide uncontrollably and may invade nearby tissues.

[0099] As used herein, the terms “cancerous cell,” “cancer cell,” “tumor cell,” or variations thereof refer to a single cell of cancerous growth or cancerous tissue. A tumor generally refers to a swelling or lesion formed by the abnormal growth of cells, which can be benign, precancerous, or malignant. Most cancers form tumors, but some, such as leukemia, do not necessarily form tumors. For those cancers that do form tumors, the terms cancer (cells) and tumor (cells) are used interchangeably. The amount of tumor in an individual is called the “tumor burden,” which can be measured as the number, volume, or weight of the tumor.

[0100] As used in this article, the term " BRCA "Related cancers" refers to cancers associated with... BRCA Cancer is associated with gene mutations that make cells more likely to divide and change rapidly, thus leading to cancer.

[0101] As used in this article, the term "breast cancer" refers to a group of malignant tumors affecting breast tissue. Molecular classification of breast cancer has identified specific subtypes with clinical and biological significance, often referred to as "intrinsic" subtypes, including intrinsic luminal subtypes, intrinsic HER2-enriched subtypes (also known as HER2+ or ER- / HER2+ subtypes), and intrinsic basal-like breast cancer (BLBC) subtypes. This article considers all forms of breast cancer. Approximately 45%–72% are genetically inherited. BRCA1 or BRCA2 Women with this variant will develop breast cancer between the ages of 70 and 80.

[0102] As used in this article, the term "ovarian cancer" refers to a group of malignant tumors affecting the ovary that have developed from epithelial cells, sex cord-stromal cells (e.g., granulosa cells, theca cells, and phylum cells), or germ cells (e.g., oocytes). Approximately 60% of ovarian tumors are of epithelial origin and account for 90% of ovarian cancers. Such epithelial-derived ovarian cancers are heterogeneous in character, varying in tumor morphology, clinical symptoms, and genetic alterations. The World Health Organization (WHO) lists eight distinct tumor histologies, including serous, endometrioid, mucinous, clear cell, transitional cell, squamous cell, mixed epithelial, and undifferentiated. Tumors of each of these subtypes can be classified as benign (with low malignant potential and / or indolent), malignant, or borderline, and as low-grade (type I) or high-grade (type II). This article considers all forms of ovarian cancer.

[0103] As used in this article, the term "pancreatic cancer" refers to cancer that originates in the pancreas and is associated with poor prognosis and low survival. Treatment for pancreatic cancer includes surgery, chemotherapy, radiation therapy, and palliative care. Treatment options can vary depending on the stage of pancreatic cancer. BRCA1 and BRCA2 The mutation is associated with familiar pancreatic cancer.

[0104] As used in this article, the term "prostate cancer" refers to cancer of the prostate gland. Prostate cancer is a commonly diagnosed malignant tumor in men and the second leading cause of cancer death in Western populations after lung cancer. If detected early, prostate cancer can be cured surgically in approximately 90% of cases. However, once the tumor spreads beyond the glandular area and forms distant metastases, the disease becomes slowly fatal. Therefore, early detection and accurate staging are crucial for selecting the right treatment and should improve treatment success rates and reduce prostate cancer-related mortality. Up to 10% of all prostate cancers may be hereditary. BRCA It is related to gene mutation.

[0105] As used herein, unless the context otherwise requires, the words “comprise,” “comprises,” and “comprising” will be understood to mean including the stated steps, elements, or groups of steps or elements, but not excluding any other steps, elements, or groups of steps or elements. “consisting of” means including and limited to anything following the phrase “consisting of.” Therefore, the phrase “consisting of” indicates that the listed elements are necessary or mandatory, and no other elements may be present. “consisting essentially of” means including any elements listed following the phrase, and is limited to other elements that do not interfere with or contribute to the activity or function specified in this disclosure of the listed elements.

[0106] As used herein, the term "malignant" refers to cancer in which a group of tumor cells exhibits one or more of the following: uncontrolled growth (e.g., division beyond normal limits), invasion (e.g., invasion and destruction of adjacent tissues), and metastasis (e.g., spread to other locations in the body via the lymphatic or bloodstream). As used herein, the term "metastasis" refers to cancer spreading from one part of the body to another. Tumors formed from cells that have already spread are called "metastatic tumors" or "metastases." Metastatic tumors contain cells similar to those in the original (primary) tumor. As used herein, the terms "benign" or "non-malignant" refer to tumors that can grow larger but do not spread to other parts of the body. Benign tumors are self-limiting and typically do not invade or metastasize.

[0107] As used herein, the term "treating" or "treatment" means relieving, reducing, or alleviating at least one symptom in a subject or achieving a delay in disease progression. For example, treatment may be the reduction of one or more symptoms of a disorder or the complete eradication of the disorder, such as cancer. Within the meaning of this disclosure, the term "treatment" also means blocking and / or reducing the risk of disease exacerbation, or preventing at least one symptom associated with or caused by the prevented state, disease, or disorder. For example, treatment may relieve, reduce, or alleviate... BRCA At least one symptom of a related cancer (e.g., breast cancer, ovarian cancer, pancreatic cancer, or prostate cancer). Exemplary treatments for such cancers include immunotherapy, drug therapy, chemotherapy, and / or surgery.

[0108] As used herein, “immunotherapy” refers to the use of substances that can stimulate an immune response to prevent, improve, or treat a disease. Examples of immunotherapy include, but are not limited to, monoclonal antibodies or immune checkpoint inhibitors, nonspecific immunotherapy, oncolytic virus therapy, T-cell therapy, and cancer vaccines. As used herein, “pharmacological therapy” refers to the administration of any small molecule compound or combination thereof for a specific type of cancer or related symptoms. For example, one class of pharmacological therapy is anti-inflammatory agents or drugs. Exemplary anti-inflammatory agents or drugs include, but are not limited to, steroids and glucocorticoids (including betamethasone, budesonide, dexamethasone, hydrocortisone acetate, hydrocortisone, hydrocortisone, methylprednisolone, prednisolone, prednisone, triamcinolone), nonsteroidal anti-inflammatory drugs (NSAIDs) including aspirin, ibuprofen, naproxen, methotrexate, sulfasalazine, leflunomide, anti-TNF drugs, cyclophosphamide, and mycophenolate mofetil. Chemotherapy includes, but is not limited to, platinum-based chemotherapy (e.g., cisplatin or carboplatin and taxane). In some cases, if the cancer is resistant to platinum-based drugs, alone or in combination, other chemotherapeutic agents may be used, such as liposomal doxorubicin, paclitaxel, docetaxel, nab-paclitaxel, gemcitabine, etoposide, pemetrexed, cyclophosphamide, topotecan, vinorelbine, irinotecan, or PARP inhibitors. Surgical treatment may include, but is not limited to, debulking surgery (e.g., cytoreductive surgery to remove tumors, salpingo-oophorectomy, hysterectomy, mastectomy, pancreaticoduodenectomy, or radical prostatectomy), followed by chemotherapy. This article considers any appropriate treatment for a specific type of cancer, including any appropriate combination of treatments at appropriate doses and using appropriate regimens as prescribed by a physician.

[0109] This article describes the use of circulating miRNAs that collectively possess specific characteristics or profiles to identify suspected cases of [disease name missing]. BRCA Related cancers or having one or more BRCA Methods for individuals with mutations. These methods can also be used for individuals suffering from... BRCA Related cancers or having one or more BRCA Any individual at risk of mutation. The method described in this paper is applicable to individuals with... BRCA1 or BRCA2 Patients with a known family history of mutations and related cancers have a simple initial screening option before undergoing genetic counseling or genetic testing. Applying miRNA-based tests to identify patients at highest risk for these mutations offers the opportunity to reduce screening costs and make it widely available, and also provides the means to identify cancer before it develops. BRCA1 or BRCA2 The opportunity for patients who are mutation carriers allows individuals to seek pre-treatment or surgical intervention.

[0110] Therefore, a prediction is suspected to have BRCAA method for assessing the lifetime risk of developing one or more cancers in subjects with mutations includes: (a) obtaining samples collected from the subjects; (b) determining the amount of one or more miRNA sequences from a set of circulating microRNAs; (c) comparing the amount of circulating microRNAs determined in step (b) with a statistical model; and (d) identifying BRCA1 or BRCA2 The presence of at least one mutation in the gene.

[0111] Well-defined dysregulation of miRNA expression in cancer, along with the contribution of miRNAs to tumorigenesis, and BRCA1 / 2 The fact that mutation carriers possess genetic alterations present in all somatic cells provides a diverse profile or spectrum of circulating miRNAs. In one implementation, the miRNA profile includes one or more miRNA sequences from Table 1:

[0112] Table 1: miRNA sequences that can be used for miRNA profiling.

[0113] In another implementation, the miRNA profile includes one or more miRNA sequences from Table 2:

[0114] Table 2: miRNA sequences that can be used for miRNA profiling.

[0115] The miRNA spectrum described in this article provides BRCA1 or BRCA2 The information about the mutation, and consistently demonstrated that... BRCA The ability to distinguish the mutation from the wild type. Therefore, the presence or absence of one or more miRNA sequences in this group and subsequent quantification can identify patients as carriers of that mutation and can be used to predict the presence of one or more diseases. BRCA Lifetime risk of related cancers. BRCA Mutations can occur in both men and women, are hereditary, and are associated with several types of cancer. BRCA1 or BRCA1 Related. For example, BRCA1 Mutations are associated with an increased risk of breast cancer, including triple-negative breast cancer. BRCA2 It is associated with a higher risk of other cancers, including prostate cancer and pancreatic cancer. Therefore, in one embodiment, the cancers described herein include one or more of breast cancer, ovarian cancer, pancreatic cancer, or prostate cancer.

[0116] MicroRNAs are endogenous non-coding small RNA molecules that can be secreted into circulation and exist in a very stable form. Therefore, the use of circulating miRNAs is ideal for patient samples using blood that can be routinely drawn by a physician or clinic. In one embodiment, the blood sample used with the methods described herein includes plasma, serum, or whole blood. Any suitable blood sample from which circulating miRNAs can be extracted, detected, and / or quantified or measured can be used. MicroRNAs in biological samples (e.g., blood samples) can be detected and quantified using one or more detection techniques (such as qPCR, microarray detection, etc.) for detecting analytes and quantification using methods known in the art.

[0117] After obtaining the samples, the amount of circulating miRNA was compared with a statistical model to determine the concentration of miRNAs in samples containing... BRCA The sample is differentiated from the mutated sample. In one implementation, the statistical model used as described herein includes one or more models selected from the group consisting of linear discriminant analysis, logistic regression, multivariate adaptive regression splines, Naive Bayes, neural networks, support vector machines, functional trees, LAD trees, Bayesian networks, elastic network regression, and random forests. In another implementation, the statistical model includes a logistic regression model with the output parameters shown in Table 3.

[0118]

[0119] Table 3: Used for prediction BRCA The logistic regression model parameters for the state are defined. Estimates, odds ratios, and two-sided p-values ​​are estimated on the training set of samples. The final model is evaluated on a validation cohort with a cutoff value indicating a 50% probability of positive cells. The ability to identify patients with BRCA1 mutations, BRCA2 mutations, or both, or to predict the lifetime risk of a subject having one or more cancers, allows for continued screening and potential treatment interventions if cancer is present. Treatment options include all known treatment options for any cancer described herein, including immunotherapy, drug therapy, chemotherapy, surgical interventions, and combinations thereof. In one embodiment, the method described herein also includes administering treatment to the subject, wherein the treatment is selected from the group consisting of: surgery, chemotherapy, immunotherapy, radiation therapy, hormone therapy, and stem cell transplantation. Treatment can be prescribed and administered as advised by a physician.

[0120] This disclosure also considers kits containing reagents for detecting miRNA molecules in the methods described herein. In some embodiments, the kit may further include instructions for use according to the methods of this disclosure, including how to draw blood, process samples, and perform [the procedure]. BRCA1 or BRCA2Instructions for mutation detection. The kit may also include features for generating developmental milestones based on the detection results. BRCA Instructions for relevant cancer risk scores. Additionally, the kit may include references for subsequent genetic screening and / or genetic counseling and / or screening for potential cancers in patients requiring further diagnosis.

[0121] The practice of this disclosure offers numerous improvements and advantages compared to other technologies. The methods described herein include serum analytes (miRNAs) and clinical metadata (personal and family history) to improve screening tests for BRCAness or BRCA1 / 2 mutations. BRCAness refers to alternative biomarkers for cancer risk, such as recent ovarian cancer risk. The techniques described herein consistently perform well across numerous categories, including age, cancer type, and ethnicity. Furthermore, the methods described herein have unique applications for predicting ovarian cancer risk, including determining the 5-year risk of ovarian cancer. The methods described herein can be used as “screening tests.” In other words, a subject’s miRNAs and clinical metadata can be tested to determine whether it is beneficial to refer the subject for further genetic testing. This improves the cost and time required for testing for BRCA1 / 2 mutations.

[0122] The model's performance remained consistent across all age groups. The model incorporates clinical metadata and miRNA assessments of BRCAness.

[0123] This document may cite numerous patent and non-patent references. Each cited reference is incorporated herein by way of citation. In the event of any discrepancy between the definition of a term in this specification and its definition in a cited reference, the definition in this specification shall prevail.

[0124] The following embodiments are provided to provide those skilled in the art with a complete disclosure and description of how to prepare and use the methods and groups described herein, and are not intended to limit the scope of this disclosure. EMBODIMENT

[0125] GENERAL METHODS The following embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.

[0126] EMBODIMENT 1: The aim of this study was to develop serum-based miRNA genomes to identify miRNAs in healthy controls. BRCA1 / 2Mutation carriers. A diagnostic biomarker study was conducted based on serum samples collected from six international cohorts (BWH - Brigham and Women's Hospital; CCGP - DFCI Center for Cancer Genetics and Prevention; DGO - Department of Gynecologic Oncology, Tata Medical Center, Kolkata, India; IHCC - International Center for Hereditary Cancers, Medical University of Pomerania, Poland; DFCI - DFCI / BWH Biobank; UPenn - University of Pennsylvania). Mutation carriers were used in the analysis from 653 individuals with known mutations. BRCA1 and BRCA2 Serum samples from healthy participants with mutated states. Receiver operating characteristic area under the curve (AUC ROC), sensitivity, and specificity (assessed in an independent validation sample set) were used to evaluate the performance of the final classification model. In the study population, 350 (53.6%) subjects had... BRCA Mutations, and 303 cases (46.4%) were... BRCA1 / 2 Wild-type. miRNAs were isolated from all individuals and their expression was quantified using RNA sequencing. In the pooled, batch-adjusted cohort, variable selection based on differential expression analysis identified a group of individuals with... BRCA Nineteen miRNAs were identified as significantly associated with mutation carrier status, and ten of them were ultimately used for class separation using a logistic regression model.

[0127] The model achieved an AUC ROC of 0.89 (95% CI: 0.87–0.93), and in the validation group, the accuracy was 85.61%, the sensitivity was 93.88%, and the specificity was 80.72%. BRCA1 or BRCA2 Mutations, menopausal status, or prior oophorectomy performed before blood sample collection do not affect classification performance. Therefore, circulating microRNAs can be used to identify patients at high genetic risk for ovarian and breast cancer. BRCA1 or BRCA2 Mutations can facilitate inexpensive first-line screening for further genetic research.

[0128] sample The research group consists of six serum biomaterial libraries ( FIG. 1A The study included samples from Brigham and Women's Hospital (BWH; Boston, MA; N=87), Dana-Farber Cancer Institute (DFCI; Boston, MA; N=200), and separate sample sets from DFCI Cancer Genetics and Prevention Center (CCGP; Boston, MA; N=162), Tata Medical Center (DGO; Kolkata, India; N=20), Medical University of Pomerania (IHCC; Szczecin, Poland; N=52), and the University of Pennsylvania (UPenn; Philadelphia, PA; N=132). The study included samples from individuals with genetically confirmed...BRCA1 / 2 Samples were taken from patients with a history of ovarian cancer or diagnosed with other cancers within one year of sampling. Patients with benign adnexal masses were included. Missed diagnoses or... BRCA Patients with missing status. All study samples were collected under a locally approved institutional review board protocol after obtaining informed consent from the study participants.

[0129] Next-generation sequencing Total RNA was extracted, and then size selection, adaptor ligation, and library preparation were performed as described above. All miRNA sequencing data were mapped to a reference miRNA database (miRBase version 22.1) using nf-core / smrnaseq version 1.1.0, a unified and standardized bioinformatics pipeline developed and released as part of the Nextflow project. Reads not mapped to miRBase were subsequently mapped to the human genome GRCh38. The sequencing protocol was set to QIAseq and Illumina. ® Alternatively, Nextflex, depending on the sample set (QIAseq miRNA sequencing for BWH, IHCC, DFCI, and UPenn; Illumina miRNA sequencing for CCGP; and NEXTFLEX small RNA sequencing for DGO). All pipeline parameters were kept at the default values ​​recommended by the code authors to ensure reproducibility.

[0130] Data integration and miRNA selection MicroRNAs were filtered for detection in at least 33% of samples in each group using a minimum detection threshold of >= 10 transcripts / million (TPM). After filtering, 227 of the initial 2621 miRNAs were retained. Principal component analysis (PCA) was used to visualize the presence of batch effects (Figure 2). After Voom normalization and removal of mean-variance trends (Voom model formula: ~0 + ...), ... BRCA Data from all subject groups were pooled using ComBat, with the UPenn group used as a reference (ComBat model formula: ~brca status + ovarian). Because different techniques were used to quantify miRNA levels in different subject groups, using the empirical Bayesian framework (ComBat) was a necessary step to combine data from all subject groups while considering technological heterogeneity. However, to limit the potential confounding effects of ComBat on the effects of interest, two versions of differential expression analysis were performed (…). FIG. 1B): With and without batch adjustment, and comparing results to identify miRNAs detected in both variants. Differential expression analysis was performed using the limma linear model (bioconductor.org / packages / devel / bioc / vignettes / limma / inst / doc / usersguide.pdf) based on microarray and RNA-seq data. The limma model formula includes the following effects: BRCA1 / 2 The effects of mutation and prior bilateral salpingo-oophorectomy (~0 + BRCA status + ovarian presence). Visualization of samples in reduced-dimensional space was performed using Uniform Manifold Approximation Projection (UMAP). Settings were as follows: number of neighbors for representation: 10 for batch-adjusted data and 5 for unadjusted data; minimum distance: 0.2 for batch-adjusted data and 0.9 for unadjusted data; distance metric: Euclidean distance in both cases. Hierarchical clustering was performed, with Ward's method used for connections, Euclidean distance metric used for samples (columns), and relevance distance metric used for miRNAs (rows).

[0131] Model development and statistical analysis In this step, the dataset was divided into a training set (N=391, 75% of cases from all groups except UPenn, randomly split), a test set (N=130, 25% of cases from all groups except UPenn, randomly split), and a validation set (N=132, UPenn group only). Model development and validation were performed using the internal OmicSelector software (version 1.0; biostat.umed.pl / OmicSelector). In short, OmicSelector tested 94 feature selection methods based on 25 different variable selection methods. As described above, feature selection based on OmicSelector followed an initial consistency-based pre-selection. Feature sets with more than 10 miRNAs were filtered out. The selected feature sets were ranked using four modeling techniques (logistic regression, conditional decision tree, recursive splitting tree, and artificial neural network with one hidden layer), and hyperparameter optimization (a randomized set of 2000 hyperparameters) and hold-out validation were performed on the test set. The number of modeling techniques was reduced to ensure low complexity of the resulting model and thus reduce the chance of overfitting.

[0132] To evaluate model performance, the training area under ROC was analyzed, and models were selected based on the highest Youden exponent. BRCA Cutoff value for state prediction. This cutoff value is applied to predictions on both the test and validation sets. Calculate accuracy, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) for all sets. Under indicated conditions, the α level for statistical significance is set to < 0.05.

[0133] Researchers responsible for feature selection and model development know BRCA Status and other available clinical data. Use BRCA The model was developed based on the results of genetic testing of the state; therefore, the modeling results were unknown when these tests were performed. All hypothesis tests were two-sided. All analyses were performed in R (cran.r-project.org).

[0134] Study the characteristics of groups The characteristics of the research group are summarized in Table 4.

[0135]

[0136] Table 4: Clinical characteristics of the study group used in Example 1. a IQR - Interquartile Range A total of 653 study participants were sampled. Within the study population, 350 (53.6%) participants had [a specific condition / condition]. BRCA The mutation (BRCA-mt) was present in 303 cases (46.4%). BRCA1 / 2 Wild type (BRCA-wt). The clinical characteristics of each group are summarized in Table 5.

[0137]

[0138] Table 5: Clinical characteristics of all cohorts used in Example 1.

[0139] A small number of participants (75 / 653; 11.5%) had undergone a reduced-risk salpingo-oophorectomy prior to blood collection due to BRCA-related cancer risk, which was considered in the differential expression analysis.

[0140] Identification of miRNAs associated with germline BRCA mutations Use unsupervised linear and nonlinear dimensionality reduction (PCA and UMAP) to examine BRCA The effects of defect state and batch (Figure 2). The batch effect clearly separates the groups. However, in the two observed batches, BRCA State strongly influences expression profile ( FIG. 3A The goal is to determine the phylogenetic distribution of data by overlaying the results of two data preprocessing strategies. BRCA1 / 2 Mutation recognition of differentially expressed (DE) miRNAs, the two strategies being based on the original data ( FIG. 3C ) and after batch adjustment ( FIG. 3DRegardless of the data preprocessing strategy, the 19 miRNAs showed convergence (p<0.01, and |log2(FC)|>0.5, with consistent direction in both analysis variants, and the FC ratio from the two analysis variants ranging between 0.8 and 1.25). FIG. 3C and FIG. 3D (Gray markers in the diagram). Unsupervised hierarchical clustering of all subject samples from 5 groups used for miRNA selection and model development is shown, with samples based on... BRCA1 / 2 Mutation clustering ( FIG. 3E Notably, in the validation set containing UPenn samples, the 19 miRNAs were clearly distinguished. BRCA -mt and BRCA -wt samples confirmed the robustness of their selection, such as FIG. 4 The heatmap of the expression values ​​is shown.

[0141] Predicting BRCA mutation status using miRNA In the pre-selection, those with consistent separation BRCA -mt sample and BRCA Following the 19 miRNAs with -wt sample capacity, OmicSelector-based model development was used to distinguish based on batch-adjusted log2 (TPM) expression values. BRCA -mut and wild-type samples. The feature set derived from Table 6 of the training set was used for modeling using four different methods.

[0142]

[0143] Table 6. Variable selection procedure for identifying a subset of miRNAs with optimal class separation properties.

[0144] Based on the following 10 miRNAs: hsa-miR-20b-5p, hsa-miR-19b-3p, hsa-let-7b-5p, hsa-miR-320b, hsa-miR-139-3p, hsa-miR-30d-5p, hsa-miR-17-5p, hsa-miR-182-5p, hsa-miR-421, and hsa-miR-375-3p, the best predictive performance was obtained using a logistic regression model with the parameters shown in Table 7.

[0145]

[0146] Table 7: Confusion matrix of the final classification model used in Example 1.

[0147] The dataset was divided into three parts to determine what role each variable played on the training set and to provide a first list of parameters. Therefore, the validation set provides a small amount of "independent validation" because these samples were not used at all in the training process. The miRNA set was selected on a training set balanced with Synthetic Minority Class Oversampling Technique (SMOTE) using a ROC AUC-based feature ranking and minimum description length (MDL) discretization algorithm.

[0148] The final model achieved 82.35% accuracy, 84.51% sensitivity, and 79.39% specificity on the original training set. Training AUC ROC ( FIG. 5A The accuracy was 0.89 (95% CI: 0.87–0.93). The model achieved 84.62% accuracy, 95.33% sensitivity, and 83.64% specificity on the test set, and 85.61% accuracy, 93.88% sensitivity, and 80.72% specificity on the external validation set, including the UPenn group. Confusion matrices were obtained for separate sets. BRCA The predicted probability of BRCA-mt in the context of the state is FIG. 5B Presented in the text. Menopausal state ( FIG. 5C Or a pre-operative oophorectomy may be performed before blood samples are drawn. FIG. 5D This does not affect classification performance. The presented diagnostic performance is calculated against a cutoff value determined based on optimal accuracy. However, to better evaluate the utility of the proposed test, estimated positive and negative predictive values ​​(based on results from the entire patient cohort) at different thresholds are shown for populations with different mutation rates in genes associated with homologous recombination pathways related to DNA repair. Although 52% of participants provided accurate age data at the time of testing, this was predominantly from controls. The model's performance remained consistent across the entire age category.

[0149] EMBODIMENT 2: The following describes a method for classifying patients as potential BRCA mutation carriers or non-carriers using a combined lasso dimensionality reduction (DR) approach in conjunction with a linear classification model. This model incorporates both serum miRNAs and clinical metadata, providing a “BRCAness” score. BRCAness can be used as an alternative biomarker for near-term ovarian cancer risk. The model described below is novel in its inclusion of miRNA assessment of clinical metadata and BRCAness, and unique in its application for predicting ovarian cancer risk. This can be used as a clinical test to estimate an individual’s future risk of ovarian cancer and as a screening test before proceeding with genetic testing.

[0150] Hereditary breast and ovarian cancer syndrome (HBOC) is characterized by a significantly increased lifetime risk of breast and ovarian cancer, as well as melanoma, prostate cancer, and pancreatic cancer. The most commonly mutated gene in HBOC is the DNA homology repair (HR) gene. BRCA1 and BRCA2 Although mutations in many other genes with moderate penetrance involved in HR can present similar phenotypes, the only clinical indication for HBOC genetic testing is currently a personal history of HBOC-related cancer, a strong family history of cancer that may suggest HBOC, or a close family relationship with gene mutations known to be associated with HBOC. Routine testing for HBOC is not recommended. Therefore, among the estimated 1 million HBOC patients in the United States... BRCA1 or BRCA2 Of mutation carriers, only 10% are aware of their mutation status, and the percentage may be even lower among carriers of genes with moderate penetrance. This lack of awareness is exacerbated by routine... BRCA1 and BRCA2 The test group had historically performed poorly on genetic variants observed in non-white European populations, resulting in a higher proportion of variants of unknown significance (VUS) in non-white European racial and ethnic groups.

[0151] Recent data indicate that germline BRCA1 / 2 mutation carriers without cancer possess a different circulating microRNA (miRNA) profile compared to non-carriers. In this report, a model incorporating 10 miRNAs selected from next-generation sequencing (NGS) data achieved 93.9% sensitivity and 80.7% specificity in identifying mutation carriers in an independent validation set. However, NGS remains a resource-constrained technique and cannot be widely implemented in clinical practice. Furthermore, risk assessment is not performed in a vacuum detached from individual subject factors. In the following examples, additional miRNA groups are evaluated as… BRCA1 / 2 Mutation screening. This functional assessment of BRCAness was tested to determine whether it was clinically informative for estimating the 5-year risk of ovarian cancer using an independent prospective cohort.

[0152] result patient group The demographic characteristics of the study group are shown in Table 8. The study participants included... n = 1,831 women, the majority (95%) of whom lived within Massachusetts zip codes and received routine primary or specialist medical care within institutions belonging to Mass General Brigham (a large healthcare system in the New England region). Among the study participants, n =100 are known lineages BRCA1 / 2 Mutation carrier, and n= 1,731 are confirmed non-carriers ( n = 159) or no pedigree testing ( n = 1,572). BRCA1 / 2 Mutation carriers were more likely to be white or postmenopausal. Unsurprisingly, mutation carriers were more likely to have a personal or family history of breast or ovarian cancer and a higher incidence of benign breast disease. Mutation carriers were similar to non-carriers in terms of number of pregnancies, number of miscarriages, ectopic pregnancy, smoking history, BMI, and gynecological history. All variables listed in Table 8 were used to train the classification model, except for oral contraceptive use and hysterectomy. These variables were omitted from model training due to potential reverse causality. For example, in this study… BRCA1 The mutation carrier may have already undergone a hysterectomy as part of a risk-reduction procedure to minimize the risk of cancer.

[0153]

[0154] Table 8: Characteristics of the training group. The right-hand column contains... p The values ​​indicate whether a statistically significant difference exists based on a two-sample t-test. In each case, the number and proportion of subjects belonging to each category are given in parentheses. For some variables (e.g., number of births), data are missing for some subjects. The first column in parentheses provides the sample size (n) for which data for each variable was recorded.

[0155] Classification results The combined lasso dimensionality reduction technique was used to combine clinical metadata with miRNA expression data from a set of 179 serum miRNAs to classify study subjects into possible groups. BRCA1 / 2 Mutation carriers or non-carriers ( FIG. 6A For a complete description of the joint lasso model, see the Methods section below. A complete list of all 179 miRNAs considered in this study is shown in Table 9. The area under the receiver operating characteristic (ROC) curve (AUC) after 10-fold cross-validation was 0.98 (95% CI 0.94–1) (Note: all 10-fold predictions were concatenated and evaluated against the complete real BRCA tag set). The model had 151 non-zero coefficients in the miRNA component (84% of the miRNAs) and 11 non-zero coefficients in the metadata component (58% of the considered metadata variables).

[0156]

[0157] Table 9: A complete list of all 179 miRNAs considered in this embodiment.

[0158] Next, we will investigate the effect of further limiting the number of input variables. If the subjects' BRCA Mutation probability greater than Then classify it as BRCA Mutation carriers; and conversely, if BRCA The mutation probability is less than or equal to Then it is classified as non- BRCA Mutation carrier. FIG. 6B Showing the corresponding The confusion matrix; where These values ​​correspond to the Youden index. Different shading is used to indicate cases where the target and output classes are the same (e.g., 0,0 or 1,1) and different (e.g., 0,1 or 1,0). The model has a specificity of 98%, a sensitivity of 96%, and an overall classification accuracy of [missing value]. For example, for use BRCA First-pass test of mutations, where the goal is to identify of BRCA The carrier can be set as (like FIG. 6B (In this case, 98% of the control group (equivalent to 93% of the overall population) can be safely excluded from further testing because these subjects...) BRCA The probability of mutation is low (i.e., The remaining 7% of participants were more likely to be... BRCA1 / 2 Carriers can undergo confirmatory genetic testing. The positive predictive value for this population is 71%. Similarly, depending on the application, t It can be modified to improve specificity or sensitivity.

[0159] exist FIG. 7A To demonstrate the benefits of the combined model, the combined lasso model (which integrates miRNA and metadata) was compared with the separate metadata component as a baseline. After 10-fold cross-validation, the metadata model provided an AUC of 0.68 (95% CI 0.62–0.74), which was 30% lower than the AUC of 0.98 (95% CI 0.94–1) provided by the combined model, thus demonstrating a significant improvement in classification performance by including miRNA in the combined lasso model.

[0160] The impact of the number of input variables in different models When using cross-validation to select the lasso parameter, the co-lasso model retained 151 out of 179 miRNAs.

[0161] Based on the number of different miRNA inputs ( ) and metadata features ( To check model performance, a subset was selected: miRNA and Individual metadata characteristics, and their relationship with BRCA The states exhibit a strong linear relationship (as listed in Table 8). Therefore, the number of non-zero lasso components (e.g., non-zero entries for v1 and v2, see the Methods section) is limited to a maximum of [number missing]. and .exist and The model performance was measured after 10-fold cross-validation. This is an upper limit, because it is the total number of metadata features used to train the joint lasso model. Regarding the AUC score, when... Exceed After that, performance no longer improves, and AUC decreases. It increases monotonically as it increases. FIG. 8C Showing when In this case, how does the model performance (in terms of AUC) change? Changes. AUC score in Maximize. For example, when and At that time, the AUC score was (95% CI 0.95-0.98). FIG. 8A Showing the corresponding miRNA and ROC curve of the individual data feature model. Selected by the lasso model. miRNA and The individual metadata features are listed in Tables 10 and 11, respectively. FIG. 8B The confusion matrix corresponding to the Youden index is shown, such as FIG. 6A Therefore, by limiting... and In terms of AUC, there is almost no sacrifice in performance. In this case, the specificity provided by the model is 92%, and the sensitivity is 91%, with an overall accuracy (ACC) of 92%, indicating excellent performance. The model also has a positive predictive value of 38.9% (95% CI 33.4%–43.6%) and a negative predictive value of 99.4% (95% CI 99.0–99.7%).

[0162]

[0163] Table 10: Selections using lasso Each miRNA has a unique characteristic. This shows the non-... BRCA and BRCAThe mean (μ) and standard deviation (σ) of log2 expression in the subjects, as well as the fold change and p-value (indicating whether the difference in mean is significant). Fold change is defined as mean non-BRCA expression divided by mean BRCA expression; for example, 0.96 on row 1 = 5.95 / 6.2.

[0164]

[0165] Table 11: List of k2 = 5 metadata features selected by lasso. Statistics for these variables are included in Table 8.

[0166] Performance of BRCA1 / 2 classifiers across racial, age, or cancer status subgroups Subset analyses were performed, in which study participants were grouped by race, age, or cancer status to determine the stability of the combined lasso model across different subgroups. No differences were observed whether participants were separated by ovarian cancer or breast cancer, nor were any differences observed when a collective “cancer” classifier that also included non-HBOC cancers such as thyroid cancer, cervical cancer, colon cancer, and skin cancer was used. FIG. 9A Similarly, if the study cohort is divided into 10-year blocks based on age, the model performs similarly well across all age groups. FIG. 9B Finally, the impact of race and ethnicity on model performance was examined. Due to the small sample size of individual minority groups, race was categorized as either "non-Hispanic white" or "all others." The model performed similarly across racial groups. FIG. 9C Similar to larger models, the classification in these subsets remains consistent when used for classification with the more limited set of model inputs described above. FIG. 10A- FIG. 10C ).

[0167] Of the 1831 participants considered in this study, 259 underwent [the following procedure / treatment]. BRCA The genetic testing for mutations, while untested subjects were assumed to be non-mutant when training the joint lasso model. BRCA Mutation carriers. This can introduce bias or confounding effects into classification results, particularly by misclassifying true positives as assumed negatives. To address this issue, Figure 7 shows a subgroup analysis of the performance of the joint lasso model and metadata component of Figure 6, with the subjects limited to those with known mutations. BRCASubjects with results of mutational genetic testing. In patients who underwent genetic testing, the performance of the combined lasso model with complete data was largely preserved (AUC = 0.96, 95% CI 0.92–0.98). The baseline model (trained using only metadata) provided an AUC of 0.57 (95% CI 0.48–0.63) only in patients who underwent formal testing. Therefore, when evaluating the model against known tested subjects, after 10-fold cross-validation, the AUC score provided by the combined lasso model (96%) was 39% higher than that of the baseline model trained on metadata only (57%), highlighting the benefit of including miRNAs in the combined model. Similar effects were observed with the combined lasso model using limited data. See also FIG. 7C- FIG. 7D .

[0168] BRCA1 / 2 classifier analysis in non-carriers with breast or ovarian cancer Although germline lineages have been identified in 13%–15% of ovarian cancers and 3% of unscreened breast cancers. BRCA1 / 2 Mutations, but also carrier somatic cells in 5%-7% of ovarian cancers and 3% of breast cancers. BRCA1 / 2 Mutations. Therefore, the study investigated whether the miRNA profiles of non-mutation carriers with breast or ovarian cancer might be more similar to those of mutation carriers with or without cancer. Using a combined lasso model, mutation carriers and non-carriers had non-overlapping miRNA profiles. BRCA Mutation probability score ( FIG. 11A ).related BRCA For a detailed definition of the scoring, please refer to the Methods section. When considering mutation carriers and non-mutation carriers separately, cancer subjects and non-cancer subjects are indistinguishable. BRCA Rating (see FIG. 11B (Box plot). Overall, among all study participants, the mean BRCA score of non-cancer participants was 0.30, significantly lower than the mean BRCA score of cancer participants (0.41). miRNA profiling cannot distinguish between different RNAs. BRCA1 Carriers and BRCA2 carriers ( FIG. 12A ).

[0169] Clinical Application: The Relationship Between Model Prediction and Ovarian Cancer Risk While screening for BRCAness may improve the efficiency of genetic testing referrals, if the test does indicate susceptibility to ovarian cancer, this should be reflected in unknown... BRCACancer risk was observed in populations with mutated states. Therefore, the relationship between BRCAness and ovarian cancer was investigated. In this case, BRCAness was predicted by a combined lasso model; in data collected as part of the PLCO cancer screening trial... Ovarian cancer prediction was performed on samples from 259 participants who were subsequently diagnosed with ovarian cancer and 785 matched controls. Blood samples were drawn from cancer cases within a timeframe of 1 day to 18-14 days (up to 5 years) prior to cancer diagnosis. PLCO samples... BRCA The mutation status was unknown, thus testing the ability of the combined lasso model to directly assess the 5-year risk of ovarian cancer. See also FIG. 13A- FIG. 13F Using a joint lasso model with both complete and limited data, the external validation AUC scores were 0.74 (95% CI 0.71–0.78) and 0.71 (95% CI 0.67–0.75), respectively. Only cancer thresholds computed on the training data (e.g., For the complete data model, the sensitivity and specificity scores are 54% and 83%, and 51% and 78%, respectively, for the complete data model and the finite data model. For a complete breakdown of the classification scores, please see [link to relevant documentation]. FIG. 13C and Figure 13D The confusion matrix in [the dataset]. To calculate... Figure 13A and Figure 13B The AUC in the model will be compared with the predicted BRCA score output by the combined lasso model and the actual ovarian cancer diagnosis.

[0170] This explains the observed decrease in AUC score (i.e., from 0.98 to 0.74), as predicted. BRCA Mutation status is being used as a surrogate marker for ovarian cancer risk. Figure 13E and Figure 13F It shows BRCA How the score correlates with the relative risk of ovarian cancer. The comparison is shown using a full data model versus a finite data model. BRCA There was a strong positive correlation between the score and the 5-year relative risk of ovarian cancer (respectively...). and Cancer risk is calculated relative to the PLCO population. Therefore, the combined lasso model (both full data and limited data variants) can be successfully used to identify long-term cancer risk. For example, if subjects present... BRCA A score of 0.6, using a limited data model, indicates a 5-year relative risk of approximately 2, which would suggest the need for closer monitoring and more regular cancer screenings, while a score of 0.9 would indicate a 5-year relative risk of approximately 6, which could spark discussion about risk-reducing surgical options.

[0171] discuss This embodiment proposes a novel first-pass screening test for BRCA1 / 2 mutation detection that combines serum analytes (miRNA expression) with clinical metadata (personal and family history). A combined lasso model is used to reduce the dimensionality of the miRNA and metadata inputs to two dimensions. Subsequently, a shallow neural network is trained on the reduced-dimensional data to determine... BRCA1 / 2 Mutation probability. After 10-fold cross-validation, the proposed model provided an AUC score of 0.98, a significant improvement over the 0.89 AUC achieved using only miRNA reports, which were measured using next-generation sequencing in previous reports. Model performance was evaluated across diverse ethnic, age, and cancer status groups, and it has been shown that the model largely preserves its predictive power. Model performance was also tested with limited miRNAs and metadata, and it was found that the smaller input set almost replicated the full data model. This offers an advantage over previous work, where model accuracy was not tested for long-term cancer risk. Specifically, this work tested the ability of a combined lasso model to assess the long-term (5-year) cancer risk of an external population collected as part of the PLCO cancer screening trial. When used for direct prediction of ovarian cancer, the model provided an AUC score of 0.73, and the model output... BRCA The score is strongly positively correlated with relative cancer risk. , ).

[0172] This work differs from previous case-control studies in that it uses miRNA expression to distinguish between healthy subjects or subjects with benign tumors and subjects with cancer (see Chan, M., et al., 2013). Clin Cancer Res 19, 4477-4487; Yamamoto, Y. et al., 2020 Hepatol Commun 4, 284-297; Usuba, W. et al. 2019, Cancer Sci 110, 408-419; Elias, KM, et al., 2017, Elife 6). Such studies cannot identify early indicators of cancer risk because these models are trained on subjects who already have cancer, most of which are advanced-stage. Because the model outputs are binary (e.g., cancer vs. control), they cannot handle competing cancer risks with shared risk factors, such as breast, ovarian, and uterine cancer. In contrast, BRCA Mutations are binary outputs and therefore do not encounter the same competing risk issues. Identifying these high-risk groups for cancer, who can benefit most from risk-reducing surgery or intensive surveillance strategies, provides a more targeted approach to preventing cancer death.

[0173] These findings, added to the literature, prove BRCA Mutations are associated with changes in miRNA expression. Tumor profiling analysis has shown that both breast and ovarian cancer tissues from women with germline mutations exhibit changes. BRCA1 / 2 Unlike sporadic tumors, this work, in addition to identifying early indicators of cancer risk, also identifies circulating miRNA profiles in both healthy and women with cancer. This suggests that miRNAs play a crucial role in the phenotype of HBOC, either as a compensatory response to DNA repair defects or as a downstream effect of loss of homology repair. The current study has several notable advantages. First, it presents a large clinical dataset reflecting the clinical and demographic diversity of women in real-world clinical practice, rather than a concentrated subgroup of women participating in cancer prevention trials. This contrasts with other miRNA-based models in the literature, which focus on smaller sample sets and are trained on miRNA expression from a single cancer type. Second, the described model works equally well in women with and without HBOC-related cancers. This data includes healthy subjects and subjects with multiple cancers, such as breast cancer, ovarian cancer, skin cancer, and cervical cancer. Finally, miRNA expression techniques and metadata are used to predict… BRCA Mutations, compared to the conventional use of next-generation sequencing for BRCA Compared to genetic testing for mutations, first-pass screening offers significant efficiency and cost advantages. First-pass screening based on miRNA and metadata has very high sensitivity, which can narrow down a broader population to high-risk subgroups, who can then be transferred to further screening (e.g., routine genetic testing).

[0174] In summary, this work describes a highly robust and accurate method for... BRCA The mutation model, compared to conventional genetic testing, can be performed with significantly reduced costs and increased efficiency. This result can improve the ability to identify individuals at risk for HBOC and implement novel cancer prevention and risk management strategies.

[0175] method research group Serum samples were collected from study participants enrolled between 2012 and 2022 who participated in the Mass General Brigham Biobank. Participants were selected based on gynecologist visits recorded in their electronic health records. Samples were collected under Mass General Brigham IRB Protocol 2018P001680. Demographic characteristics and medical history were extracted from medical records using manual chart review. Race and ethnicity were self-identified in the medical records and, for analytical purposes, defined as White, Non-Hispanic, and Non-White. Mutation carriers were identified using germline genetic testing reports recorded in their electronic health records.

[0176] miRNA profiling analysis A miRNA profile of 179 different miRNA species was generated using Fireplex® probes (Abcam, Cambridge, MA) and measured at mean fluorescence units (MFI) using a Guava Easycyte 5HT flow cytometer (Luminex, Austin, TX) according to the manufacturer's instructions. This miRNA profile was optimized based on previous next-generation sequencing studies to capture serum miRNAs detectable in at least 50% of the samples. 25 μL of serum was used for each sample. Since this assay can analyze up to 68 miRNAs per well on a 96-well plate, each biological sample was distributed across three assay groups to construct a complete profile of 179 miRNAs, with some overlap between groups to allow for quality control. Each plate also included wells containing pooled human serum, a water control, and an incorporated reference miRNA. This group included miRNAs targeting *C. elegans* (C. elegans). C. Elegans In vitro control probes for miRNAs were used to establish background signal levels. Samples were processed using a STARlet liquid handling robot (Hamilton Robotics, Franklin, MA) and analyzed using FirePlex® analysis platform software (abcam.com / FireflyAnalysisSoftware). No technical replicates were performed, but the average coefficient of variation between individual miRNA values ​​within the same sample was less than 20%. Quality control was performed using in vitro miRNAs and reference miRNAs, and outlier samples were replicated for quality control purposes.

[0177] Joint Lasso Model The classification is performed using a combined lasso dimensionality reduction (DR) method and a linear classification model. BRCA and non BRCA. set up This is a matrix of normalized miRNA expression values, where n For the number of samples and p 1Let be the number of miRNAs, and set . A matrix of metadata, where p 2 This refers to the number of metadata variables. For example, X 2 The middle corresponds to BRCA The columns for family history (see Table 8) are binary vectors (i.e., their entries are either 0 or 1), where 0 indicates none. BRCA Family history, and 1 indication has BRCA Family history. (Settings) A binary vector labeled with categories, where 0 indicates non-categorical. BRCA And 1 indicates BRCA Lasso is used in conjunction with other methods to reduce the dimensionality of miRNAs and metadata. Specifically, the goal is to obtain: (A, 1) and , in express L 1 The norms β1 and β2 > 0 are regularization parameters, which control the sparsity levels of v1 and v2, respectively. The lasso model is fitted using the "lasso" Matlab function. Once v1 and v2 are determined, the miRNA and metadata are mapped to a two-dimensional space:

[0178] Then, the subjects were classified as BRCA or not BRCA ,exist X Train a linear classification model using the following equation:

[0179] in It is a category tag. It is a sample in the reduced-dimensional space (i.e., X (a line), and (a line) ) refers to the weights and biases to be trained. Here... y This represents the category label assigned to x. If The subjects were then classified as having BRCA mutations, among which... This is the BRCA threshold. The classifier was trained using the Matlab function "trainSoftmaxLayer". To validate the joint lasso model, 10-fold cross-validation was used. Nested 10-fold cross-validation was used for each training fold to select the hyperparameters β1 and β2.

[0180] “ BRCA The rating was discussed above. BRCA Rating definition Then translate and scale it to the range [0,1] for better interpretability.

[0181] Figure 14 A schematic diagram of the classification process is shown. The left side shows a 3-D t-distributed random neighborhood embedding (TSNE) plot of miRNAs and metadata to visualize the two datasets and illustrate... BRCA and non BRCA How the subjects were separated. These diagrams indicate... BRCA and non BRCA The categories show a slight linear separation. The center shows the results of joint lasso dimensionality reduction (i.e., X miRNA characteristics () ) is displayed on the y-axis, and metadata features ( On the x-axis. In the reduced-dimensional space, BRCA and non BRCA Significant linear separation was observed among the subjects. X The space is divided into two parts: one part is possible. BRCA Group, and another part is possible non- BRCA Group. For example... Figure 14 After projecting miRNA and metadata into 2D space as shown in the central scatter plot, the classifier (in...) Figure 14 (The right-hand diagram) As described above, the probability is assigned... BRCA (Right now, ) and non BRCA ( Then, define the threshold. And if the mutation probability of the subject is greater than t The subjects are then classified as BRCA Mutation carriers, and conversely, if BRCA The probability is less than or equal to t If the result is positive, the subject is classified as a non-carrier. Parameter t This allows us to decide whether to favor model specificity or sensitivity. For example, setting... t The closer a value is to zero, the greater the weight given to sensitivity, and vice versa.

Claims

1. A prediction is suspected of having BRCA1 or BRCA2 Methods for assessing the lifetime risk of developing one or more cancers in subjects with mutations include: (a) Obtaining samples collected from the subject; (b) Determine the amount of circulating microRNA selected from the following groups: hsa-miR-20b-5p (SEQ ID NO: 4), hsa-miR-19b-3p (SEQ ID NO: 3), hsa-let-7b-5p (SEQ ID NO: 1), hsa-miR-320b (SEQ ID NO: 8), hsa-miR-139-3p (SEQ ID NO: 6), hsa-miR-30d-5p (SEQ ID NO: 5), hsa-miR-17-5p (SEQ ID NO: 2), hsa-miR-182-5p (SEQ ID NO: 7), hsa-miR-421 (SEQ ID NO: 9), hsa-miR-375-3p (SEQ ID NO: 10); hsa-miRNA-106b-5p (SEQ ID NO: 11); hsa-miRNA-134-5p (SEQ ID NO: 12), hsa-miRNA-493-5p (SEQ ID NO: 13), hsa-miRNA-500a-3p (SEQ ID NO: 14), hsa-miR-1273h-3p (SEQ ID NO: 15), hsa-miR-4433a-3p (SEQ ID NO: 16), hsa-miR-4433b-5p (SEQ ID NO: 17), hsa-miR-485-3p (SEQ ID NO: 18) and has-miR-1304-3p (SEQ ID NO: 19); (c) Compare the amount of circulating microRNA determined in step (b) with a statistical model; and (d) Identification based on the amount of one or more circulating miRNAs BRCA1 or BRCA2 The presence of at least one mutation in a gene is used to predict the lifetime risk of the subject developing one or more types of cancer.

2. The method according to claim 1, wherein step (b) comprises determining the amount of 10 microRNAs comprising hsa-miR-20b-5p (SEQ ID NO: 4), hsa-miR-19b-3p (SEQ ID NO: 3), hsa-let-7b-5p (SEQ ID NO: 1), hsa-miR-320b (SEQ ID NO: 8), hsa-miR-139-3p (SEQ ID NO: 6), hsa-miR-30d-5p (SEQ ID NO: 5), hsa-miR-17-5p (SEQ ID NO: 2), hsa-miR-182-5p (SEQ ID NO: 7), hsa-miR-421 (SEQ ID NO: 9), and hsa-miR-375-3p (SEQ ID NO: 10) in the sample.

3. The method according to any one of claims 1 or 2, wherein the cancer includes one or more of breast cancer, ovarian cancer, pancreatic cancer, or prostate cancer.

4. The method according to any one of claims 1 or 2, wherein the sample is selected from blood samples.

5. The method according to any one of claims 1 or 2, wherein the blood sample is selected from the group consisting of plasma, serum, and whole blood.

6. The method according to any one of claims 1 or 2, wherein the statistical model comprises one or more models selected from the group consisting of linear discriminant analysis, logistic regression model, multivariate adaptive regression spline, Naive Bayes, neural network, support vector machine, functional tree, LAD tree, Bayesian network, elastic network regression and random forest.

7. The method of claim 6, wherein the statistical model includes a logistic regression model.

8. The method of claim 6, wherein the statistical model includes dimensionality reduction techniques.

9. The method of claim 6, wherein the statistical model comprises a combined lasso dimensionality reduction technique.

10. The method of claim 6, wherein the model comprises a combination of lasso dimensionality reduction and sparse machine learning techniques.

11. The method according to any one of claims 1 or 2, wherein step (b) and / or step (d) is performed using RNA sequencing.

12. The method according to any one of claims 1 or 2 further includes performing genetic testing or genetic counseling.

13. The method according to any one of claims 1 or 2, further comprising administering treatment to the subject, wherein the treatment is selected from the group consisting of: surgery, chemotherapy, immunotherapy, radiotherapy, hormone therapy, and stem cell transplantation.

14. The method according to any one of claims 1 or 2, wherein the subject is female.

15. A method for identifying subjects who possess... BRCA A mutation method, the method comprising: (a) Obtaining samples collected from the subject; (b) Determine the amount of circulating microRNA selected from the following groups: hsa-miR-20b-5p (SEQ ID NO: 4), hsa-miR-19b-3p (SEQ ID NO: 3), hsa-let-7b-5p (SEQ ID NO: 1), hsa-miR-320b (SEQ ID NO: 8), hsa-miR-139-3p (SEQ ID NO: 6), hsa-miR-30d-5p (SEQ ID NO: 5), hsa-miR-17-5p (SEQ ID NO: 2), hsa-miR-182-5p (SEQ ID NO: 7), hsa-miR-421 (SEQ ID NO: 9), hsa-miR-375-3p (SEQ ID NO: 10); hsa-miRNA-106b-5p (SEQ ID NO: 11); hsa-miRNA-134-5p (SEQ ID NO: 12), hsa-miRNA-493-5p (SEQ ID NO: 13), hsa-miRNA-500a-3p (SEQ ID NO: 14), hsa-miR-1273h-3p (SEQ ID NO: 15), hsa-miR-4433a-3p (SEQ ID NO: 16), hsa-miR-4433b-5p (SEQ ID NO: 17), hsa-miR-485-3p (SEQ ID NO: 18) and has-miR-1304-3p (SEQ ID NO: 19); (c) Compare the amount of circulating microRNA determined in step (b) with a statistical model; (d) Based on identification BRCA1 or BRCA2 The presence of at least one mutation in a gene is used to predict the lifetime risk of the subject developing one or more types of cancer; and (e) Perform genetic testing on the subject.

16. The method of claim 15, wherein step (b) comprises determining the amount of 10 microRNAs of hsa-miR-20b-5p (SEQ ID NO: 4), hsa-miR-19b-3p (SEQ ID NO: 3), hsa-let-7b-5p (SEQ ID NO: 1), hsa-miR-320b (SEQ ID NO: 8), hsa-miR-139-3p (SEQ ID NO: 6), hsa-miR-30d-5p (SEQ ID NO: 5), hsa-miR-17-5p (SEQ ID NO: 2), hsa-miR-182-5p (SEQ ID NO: 7), hsa-miR-421 (SEQ ID NO: 9), and hsa-miR-375-3p (SEQ ID NO: 10) in the sample.

17. The method according to any one of claims 15 or 16, further comprising monitoring the subject's... BRCA Related cancers.

18. The method according to any one of claims 15 or 16, wherein... BRCA The associated cancers include one or more of breast cancer, ovarian cancer, pancreatic cancer, or prostate cancer.

19. The method according to any one of claims 15-18, further comprising administering treatment to the subject, wherein the treatment is selected from the group consisting of: surgery, chemotherapy, immunotherapy, radiotherapy, hormone therapy, and stem cell transplantation.

20. The method according to any one of claims 15 or 16, wherein the statistical model comprises one or more models selected from the group consisting of linear discriminant analysis, logistic regression model, multivariate adaptive regression spline, Naive Bayes, neural network, support vector machine, functional tree, LAD tree, Bayesian network, elastic network regression and random forest.

21. The method of claim 20, wherein the statistical model comprises a logistic regression model.

22. The method of claim 20, wherein the statistical model includes dimensionality reduction techniques.

23. The method of claim 20, wherein the model comprises a combined lasso dimensionality reduction technique.

24. The method of claim 20, wherein the model comprises a combination of lasso dimensionality reduction and sparse machine learning techniques.

25. A treatment was suspected of causing... BRCA Methods for patients with relevant cancers, the methods including: (a) Obtaining samples collected from the subject; (b) Determine the amount of circulating microRNA selected from the following groups: hsa-miR-20b-5p (SEQ ID NO: 4), hsa-miR-19b-3p (SEQ ID NO: 3), hsa-let-7b-5p (SEQ ID NO: 1), hsa-miR-320b (SEQ ID NO: 8), hsa-miR-139-3p (SEQ ID NO: 6), hsa-miR-30d-5p (SEQ ID NO: 5), hsa-miR-17-5p (SEQ ID NO: 2), hsa-miR-182-5p (SEQ ID NO: 7), hsa-miR-421 (SEQ ID NO: 9), hsa-miR-375-3p (SEQ ID NO: 10); hsa-miRNA-106b-5p (SEQ ID NO: 11); hsa-miRNA-134-5p (SEQ ID NO: 12), hsa-miRNA-493-5p (SEQ ID NO: 13), hsa-miRNA-500a-3p (SEQ ID NO: 14), hsa-miR-1273h-3p (SEQ ID NO: 15), hsa-miR-4433a-3p (SEQ ID NO: 16), hsa-miR-4433b-5p (SEQ ID NO: 17), hsa-miR-485-3p (SEQ ID NO: 18) and has-miR-1304-3p (SEQ ID NO: 19); (c) Compare the amount of circulating microRNA determined in step (b) with a statistical model; (d) Based on identification BRCA1 or BRCA2 The presence of at least one mutation in a gene is used to predict the lifetime risk of the subject developing one or more types of cancer; and (e) Administering to the subject a treatment selected from the group consisting of: surgery, chemotherapy, immunotherapy, radiotherapy, hormone therapy, and stem cell transplantation.

26. A kit comprising at least one test probe capable of specifically hybridizing with microRNAs selected from the group consisting of: hsa-miR-20b-5p (SEQ ID NO: 4), hsa-miR-19b-3p (SEQ ID NO: 3), hsa-let-7b-5p (SEQ ID NO: 1), hsa-miR-320b (SEQ ID NO: 8), hsa-miR-139-3p (SEQ ID NO: 6), hsa-miR-30d-5p (SEQ ID NO: 5), hsa-miR-17-5p (SEQ ID NO: 2), hsa-miR-182-5p (SEQ ID NO: 7), hsa-miR-421 (SEQ ID NO: 9), hsa-miR-375-3p (SEQ ID NO: 10); hsa-miRNA-106b-5p (SEQ ID NO: 10); 11); hsa-miRNA-134-5p (SEQ ID NO: 12), hsa-miRNA-493-5p (SEQ ID NO: 13), hsa-miRNA-500a-3p (SEQ ID NO: 14), hsa-miR-1273h-3p (SEQ ID NO: 15), hsa-miR-4433a-3p (SEQ ID NO: 16), hsa-miR-4433b-5p (SEQ ID NO: 17), hsa-miR-485-3p (SEQ ID NO: 18) and has-miR-1304-3p (SEQ ID NO: 19).

27. The kit of claim 26, wherein at least one of the probes comprises a detectable marker.

28. The kit according to any one of claims 26 or 27 further includes reagents for reverse transcription of microRNA molecules.

29. The kit according to any one of claims 26-28, wherein the at least one test probe is capable of specifically hybridizing with microRNAs including hsa-miR-20b-5p (SEQ ID NO: 4), hsa-miR-19b-3p (SEQ ID NO: 3), hsa-let-7b-5p (SEQ ID NO: 1), hsa-miR-320b (SEQ ID NO: 8), hsa-miR-139-3p (SEQ ID NO: 6), hsa-miR-30d-5p (SEQ ID NO: 5), hsa-miR-17-5p (SEQ ID NO: 2), hsa-miR-182-5p (SEQ ID NO: 7), hsa-miR-421 (SEQ ID NO: 9), and hsa-miR-375-3p (SEQ ID NO: 10).