Markers for the early detection of colonic cell proliferation disorders

JP7926984B2Active Publication Date: 2026-09-30FREENOM HLDG INC
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2023520319
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-05
Filing Date
2021-09-30
Publication Date
2026-09-30
Estimated Expiration
2041-09-30

Smart Images

  • Figure 0007926984000006
    Figure 0007926984000006
  • Figure 0007926984000007
    Figure 0007926984000007
  • Figure 0007926984000008
    Figure 0007926984000008
Patent Text Reader

Abstract

The systems, liquid media, compositions, methods, and kits disclosed herein relate to panels of autoantibody biomarkers for the early detection of colon cell proliferative disorders, including colon cancer. For the autoantibody panels described herein, the presence or level of autoantibodies in a biological sample can be used as input for classifier generation and in machine learning models useful for classifying subjects in a population for the detection of colon cell proliferative disorders.
Need to check novelty before this filing date? Find Prior Art

Description

[[Technical Field]]

[0001] Cross-Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 835,281, filed on October 5, 2020, which is incorporated herein by reference in its entirety.

[0002] The present disclosure relates to biomarkers and methods for early detection of rectal cell proliferative disorders including progressive adenoma and colorectal cancer. [[Background Art]]

[0003] Colorectal cancer is a leading cause of cancer-related death in Western countries. Despite being one of the most well-characterized solid tumors, colorectal cancer remains one of the leading causes of death in developed countries due to late diagnosis. Among other reasons, delayed diagnosis in patients is caused by the excessively late timing of diagnostic tests such as colonoscopy. Deaths from colorectal cancer can be prevented through effective screening. [[Summary of the Invention]]

[0004] The present disclosure provides methods and systems directed to autoantibody profiling of biological samples associated with detection of colorectal cancer and disease progression.

[0005] In one aspect, the present disclosure provides a predetermined autoantibody panel characteristic of colonic cell proliferative disorders, the panel comprising autoantibodies against three or more antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, ASB9, NXN, ZBTB21, EYA1, GSPT1, MLIP, RBM38, ARMC5, TP53, BRD9, CDK4, PRMT6, PCOLCE, and SDCBP.

[0006] In some embodiments, the three or more autoantibodies are IgG autoantibodies, IgM autoantibodies, or a combination thereof.

[0007] In some embodiments, the panel is configured to differentiate between healthy subjects, subjects with benign colorectal polyps, subjects with progressive adenomas, or subjects with colorectal cancer.

[0008] In some embodiments, the panel is configured to show a progressive adenoma, and The antibody comprises: 1) IgM autoantibodies against at least three antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, and SDCBP; 2) IgM autoantibodies against at least one antigen selected from the group consisting of UBE2S, NME5, and CD20; 3) IgG autoantibodies against at least three antigens selected from the group consisting of ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, PCOLCE, and ASB9; 4) IgG autoantibodies against at least one antigen selected from the group consisting of ASB9, NAT6, Supt6h, and PRDM8; or a combination thereof.

[0009] In some embodiments, the panel is configured to show colorectal cancer, and 1) IgM autoantibodies against at least three antigens selected from the group consisting of PELO, CDK4, MTP1, PRMT6, ZBTB2, and PCOLCE; 2) IgM autoantibodies against at least one antigen selected from the group consisting of CDK4, MTCP1, and PCOLCE; 3) IgG autoantibodies against at least three antigens selected from the group consisting of TSSC4, BRD9, BCCIP, and TP53; 4) IgG autoantibodies against TP53, or a combination thereof.

[0010] In some embodiments, colonic cell proliferation disorders are selected from the group consisting of adenoma (adenomatous polyp), polyposis, Lynch syndrome, sessile serrated adenoma (SSA), progressive adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal cell tumor, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.

[0011] In another aspect, the Disclosure provides a classifier configured to differentiate between a population of healthy subjects and subjects having colonic cell proliferation disorder, the classifier comprising a set of measurements representing autoantibodies of a predetermined autoantibody panel characteristic of colonic cell proliferation disorder, the measurements obtained from autoantibody expression data of healthy subjects and subjects having colonic cell proliferation disorder, the measurements being used to generate a set of features corresponding to the characteristics of the autoantibodies, the set of features being input into a machine learning or statistical model, and the model providing a feature vector useful as a classifier capable of differentiating between a population of healthy subjects and subjects having colonic cell proliferation disorder.

[0012] In some embodiments, a predetermined autoantibody panel includes autoantibodies against three or more antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, ASB9, PRDM8, NXN, ZBTB21, EYA1, GSPT1, MLIP, RBM38, ARMC5, TP53, BRD9, CDK4, PRMT6, PCOLCE, and SDCBP.

[0013] In some embodiments, the three or more autoantibodies are IgG autoantibodies, IgM autoantibodies, or a combination thereof.

[0014] In some embodiments, the panel is configured to differentiate between healthy subjects, subjects with benign colorectal polyps, subjects with progressive adenomas, or subjects with colorectal cancer.

[0015] In some embodiments, the panel is configured to show a progressive adenoma, and The antibody comprises: 1) IgM autoantibodies against at least three antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, and SDCBP; 2) IgM autoantibodies against at least one antigen selected from the group consisting of UBE2S, NME5, and CD20; 3) IgG autoantibodies against at least three antigens selected from the group consisting of ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, PCOLCE, and ASB9; 4) IgG autoantibodies against at least one antigen selected from the group consisting of ASB9, NAT6, Supt6h, and PRDM8; or a combination thereof.

[0016] In some embodiments, the panel is configured to show colorectal cancer and includes: 1) IgM autoantibodies against at least three antigens selected from the group consisting of PELO, CDK4, MTP1, PRMT6, ZBTB2, and PCOLCE; 2) IgM autoantibodies against at least one antigen selected from the group consisting of CDK4, MTCP1, and PCOLCE; 3) IgG autoantibodies against at least three antigens selected from the group consisting of TSSC4, BRD9, BCCIP, and TP53; 4) IgG autoantibodies against TP53, or a combination thereof.

[0017] In some embodiments, colonic cell proliferation disorders are selected from the group consisting of adenoma (adenomatous polyp), polyposis, Lynch syndrome, sessile serrated adenoma (SSA), progressive adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal cell tumor, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.

[0018] In another aspect, the disclosure provides a system comprising a machine learning model classifier for detecting colonic cell proliferation disorders, the system comprising a computer-readable medium containing a classifier operable to classify subjects on at least partially basis to a predetermined autoantibody panel, and one or more processors for executing instructions stored in the computer-readable medium.

[0019] In some embodiments, the classifiers are loaded into the memory of a computer system, and the machine learning model is trained using training vectors obtained from training biological samples, in which a first subset of the training biological samples is identified as having colonic cell proliferation disorder, and a second subset of the training biological samples is identified as not having colonic cell proliferation disorder.

[0020] In some embodiments, a predetermined autoantibody panel includes autoantibodies against three or more antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, ASB9, PRDM8, NXN, ZBTB21, EYA1, GSPT1, MLIP, RBM38, ARMC5, TP53, BRD9, CDK4, PRMT6, PCOLCE, and SDCBP.

[0021] In some embodiments, the classifier is selected from the group consisting of deep learning classifiers, neural network classifiers, linear discriminant analysis (LDA) classifiers, quadratic discriminant analysis (QDA) classifiers, support vector machine (SVM) classifiers, random forest (RF) classifiers, K nearest neighbor classifiers, linear kernel support vector machine classifiers, first- or second-order polynomial kernel support vector machine classifiers, ridge regression classifiers, elastic network algorithm classifiers, sequential minimal problem optimization algorithm classifiers, naive Bayes algorithm classifiers, and principal component analysis classifiers.

[0022] In other embodiments, the Disclosure provides a method for determining an autoantibody profile of a subject, the method comprising the steps of obtaining a biological sample from the subject and measuring the amount of antibodies from a predetermined autoantibody panel containing autoantibodies against three or more antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, ASB9, NXN, ZBTB21, EYA1, GSPT1, MLIP, RBM38, ARMC5, TP53, BRD9, CDK4, PRMT6, PCOLCE, and SDCBP, in order to provide an autoantibody profile of the subject.

[0023] In some embodiments, the autoantibody profile is associated with colonic cell proliferation disorder and provides a classification of subjects as having colonic cell proliferation disorder.

[0024] In some embodiments, the biological sample obtained from the subject is selected from the group consisting of body fluids, feces, colonic effluent, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, tissue biopsy, and combinations thereof.

[0025] In some embodiments, colonic cell proliferation disorders are selected from the group consisting of adenoma (adenomatous polyp), polyposis, Lynch syndrome, sessile serrated adenoma (SSA), progressive adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal cell tumor, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.

[0026] In another aspect, the present disclosure provides a method for detecting a colorectal proliferative disorder in a subject, the method comprising the steps of: obtaining a biological sample from the subject; measuring the amount of antibodies from a predetermined autoantibody panel comprising autoantibodies against three or more antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, ASB9, NXN, ZBTB21, EYA1, GSPT1, MLIP, RBM38, ARMC5, TP53, BRD9, CDK4, PRMT6, PCOLCE, and SDCBP, to provide an autoantibody profile of the subject; and processing the autoantibody profile using a machine learning model trained to distinguish between healthy subjects and subjects having a colorectal proliferative disorder, to provide an output value associated with the presence of a colorectal proliferative disorder, thereby indicating the presence of a colorectal proliferative disorder in the subject.

[0027] In some embodiments, the autoantibody profile is associated with a colorectal proliferative disorder and provides classification of the subject as having a colorectal proliferative disorder.

[0028] In some embodiments, the method further comprises the step of detecting the methylation status of nucleic acid molecules in the biological sample to provide a methylation profile.

[0029] In some embodiments, the method further comprises the step of processing the methylation profile using the machine learning model, and the methylation profile is combined with the autoantibody profile in the machine learning model to distinguish between healthy subjects and subjects with a colorectal proliferative disorder.

[0030] In some embodiments, the method further comprises the step of measuring the amount of one or more proteins in the biological sample to provide a protein profile.

[0031] In some embodiments, the method further includes the step of processing protein profiles using a machine learning model, where the protein profiles are combined with autoantibody profiles in the machine learning model to differentiate between healthy subjects and subjects with colonic cell proliferation disorder.

[0032] In some embodiments, the biological sample obtained from the subject is selected from the group consisting of body fluids, feces, colonic effluent, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, tissue biopsy, and combinations thereof.

[0033] In some embodiments, colonic cell proliferation disorders are selected from the group consisting of adenoma (adenomatous polyp), polyposis, Lynch syndrome, sessile serrated adenoma (SSA), progressive adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal cell tumor, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.

[0034] In some embodiments, the panel is configured to show progressive adenoma and includes: 1) IgM autoantibodies against at least three antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, and SDCBP; 2) IgM autoantibodies against at least one antigen selected from the group consisting of UBE2S, NME5, and CD20; 3) IgG autoantibodies against at least three antigens selected from the group consisting of ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, PCOLCE, and ASB9; 4) IgG autoantibodies against at least one antigen selected from the group consisting of ASB9, NAT6, Supt6h, and PRDM8, or a combination thereof.

[0035] In some embodiments, the panel is configured to show colorectal cancer and includes: 1) IgM autoantibodies against at least three antigens selected from the group consisting of PELO, CDK4, MTP1, PRMT6, ZBTB2, and PCOLCE; 2) IgM autoantibodies against at least one antigen selected from the group consisting of CDK4, MTCP1, and PCOLCE; 3) IgG autoantibodies against at least three antigens selected from the group consisting of TSSC4, BRD9, BCCIP, and TP53; 4) IgG autoantibodies against TP53, or a combination thereof.

[0036] In some embodiments, the method further includes the step of treating a colonic cell proliferation disorder in the subject. In some embodiments, the treatment is selected from the group consisting of surgery, radiofrequency ablation, chemotherapy, radiotherapy, targeted therapy, and immunotherapy.

[0037] Further aspects and advantages of this disclosure will be readily apparent to those skilled in the art from the detailed description below, and only exemplary embodiments of this disclosure are shown and described here. As will be understood, this disclosure may also be possible in other embodiments and different embodiments, and various details thereof may be modified in various obvious ways without all departing from this disclosure. Accordingly, the drawings and description are intended to be illustrative and not limiting.

[0038] Embedding by citation All publications, patents, and patent applications referenced herein are incorporated herein by reference to the same extent as each individual publication, patent, or patent application is incorporated herein by reference specifically and individually. To the extent that any publications and patents or patent applications incorporated by reference are in conflict with the disclosures contained herein, this Specified Publication is intended to supersede and / or take precedence over such conflicting material. [Brief explanation of the drawing]

[0039] Novel features of the present invention are described in particular in the appended claims. The features and advantages of the present invention will be better understood by referring to the following detailed description illustrating exemplary embodiments in which the principles of the present invention are used, and to the following appended drawings (also referred to herein as “Figure” and “FIG.”). [Figure 1] To carry out the methods provided herein, a schematic diagram of a computer system programmed with machine learning models and classifiers, or otherwise configured, is provided. [Figure 2] This provides a graph showing the CV coefficients of the top 5 AAb targets for CRC classification. [Figure 3] This provides a graph showing recursive feature removal for CRC classification using cross-validation. [Figure 4] This provides a graph showing the CV coefficients of the top 10 AAb targets for the AA classification. [Figure 5] This provides a graph showing recursive feature removal for AA classification using cross-validation. [Figure 6] This provides a graph showing the CV coefficients of the top 5 AAb targets for NAA classification. [Figure 7] This provides a graph showing recursive feature removal for NAA classification using cross-validation. [Modes for carrying out the invention]

[0040] While various embodiments of the present invention are shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only as examples. It will be understood that numerous modifications, variations, and substitutions can be made without departing from the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be utilized.

[0041] Colorectal cancer is a leading cause of cancer-related death in Western countries. Although colorectal cancer is one of the most distinctive solid tumors, its delayed diagnosis contributes to its continued prevalence as a major cause of death in developed countries. Among other reasons, the late timing of diagnostic tests such as colonoscopy is a major factor in the delayed diagnosis of colorectal cancer. Deaths from colorectal cancer can be prevented through effective screening. Specific antibody responses to tumor-associated antigens have been observed in cancer patients. Because these antibody responses can be triggered by changes in the structure or expression of autologous proteins within tumor cells, the presence of certain antibodies may serve as immunological markers for cancer.

[0042] This disclosure generally relates to cancer detection and disease monitoring. More specifically, this disclosure relates to the detection of cancer-related autoantibodies and disease monitoring in colonic cell proliferation disorders such as early colorectal cancer. In particular, a circulating autoantibody signature panel and its applications are provided for identifying human subjects who have or are at risk of developing colonic cell proliferation disorders such as colorectal cancer (CRC) and / or colorectal adenoma (CA), e.g., advanced colorectal adenoma (AA).

[0043] This disclosure describes tumor antigen-associated autoantibodies ("tAAbs" or "autoantibodies") in subjects that indicate, for example, the presence of or high risk of developing colonic cell proliferation disorder in subjects with colorectal lesions. Cancer screening and monitoring improve survival rates because early detection allows cancer to be eliminated before it proliferates and spreads. For example, in colorectal cancer, colonoscopy plays a role in improving early diagnosis. Unfortunately, regular checkups are not recommended due to low patient attendance and the invasive nature of the procedure.

[0044] Described herein are methods for screening or identifying subjects with or at risk of having colonic cell proliferation disorder, based at least in part on the expression profile or abundance of autoantibodies whose expression is upregulated or overexpressed in subjects with colonic cell proliferation disorder. Further described herein are methods for obtaining data useful for the diagnosis of colonic cell proliferation disorder in subjects, for example, human subjects.

[0045] Colonic cell proliferation disorder may be at any tumor stage (e.g., TX, T0, Tis, T1, T2, T3, T4), any stage of regional lymph node or distant metastasis (e.g., NX, N0, N1, M0, M1), any stage (e.g., stage 0 (Tis, N0, M0), stage IA (T1, N0, M0), stage IIA (T3, N0, M0), stage IIB (T1-3, N1, M0), stage III (T4, any N, M0) or stage IV (any T, any N, M1)), resectable, locally advanced (unresectable), or metastatic.

[0046] Screening tools can be compromised by false positive and false negative results, as well as specificity and sensitivity. An ideal cancer screening tool should have a high positive predictive value (PPV) that minimizes unnecessary testing (low false positives) and detects the majority of cancers (low false negatives). Another important compromise is "detection sensitivity," which is different from test sensitivity. Detection sensitivity is the lower limit of detection based on tumor size. Allowing tumors to grow large enough to release detectable levels of circulating tumor markers defeats the purpose of early cancer detection and prevention of progression. Therefore, highly sensitive and effective blood-based screening is needed for the early diagnosis of colorectal cancer.

[0047] The detection of circulating tumor DNA, known as "liquid biopsy," enables non-invasive tumor detection and information gathering. Identifying tumor-specific mutations in these liquid biopsies has been used in the diagnosis of colorectal, breast, and prostate cancers. However, these techniques may have limited sensitivity due to the high background of circulating normal (i.e., non-tumor-derived) DNA. Therefore, there remains a need for more sensitive and specific screening tools to detect early or low-tumor-burden colorectal cancer tumor markers for recurrence screening and primary screening in at-risk populations. Circulating autoantibodies against tumor-associated antigens provide a valuable source of biomarkers in liquid biopsy samples that can be used in the machine learning models described herein.

[0048] This disclosure provides methods and systems for profiling circulating autoantibodies associated with colonic cell proliferation disorders and their progression, for example, colorectal cancer. These autoantibodies, indicating the presence of colonic cell proliferation disorders or a high risk of developing them, can be used, for example, to diagnose, treat, or prevent the progression of colonic cell proliferation disorders as early as possible, especially if the subject has only colorectal lesions. Further provided herein are kits and methods for diagnosing colonic cell proliferation disorders or assessing the risk of developing colonic cell proliferation disorders in subjects, particularly if the subject has colorectal lesions.

[0049] In one embodiment, the foregoing provides a method for using a panel of autoantibodies to differentiate samples from subjects based on disease status. In another embodiment, the foregoing provides methods, assays, and kits for detecting, identifying, and differentiating colonic cytoproliferative disorders using a panel of autoantibodies. Non-exclusive examples of colonic cytoproliferative disorders include adenomas (adenomatous polyps), polyposis disorders, Lynch syndrome, sessile serrated adenoma (SSA), progressive adenoma, dysplasia of the colon, adenomas of the colon, colorectal cancer, colon cancer, rectal cancer, colorectal cell tumor, colorectal adenocarcinoma, carcinoid tumors, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors (GIST), lymphoma, and sarcoma.

[0050] In some embodiments, the methods provided herein involve using one or more autoantibodies selected as markers for the identification, detection, and differentiation of colonic cell proliferation disorders.

[0051] definition As used herein and in the claims, the singular “a,” “and,” and “the” include multiple references unless otherwise explicitly indicated in the content. For example, the term “nucleic acid” includes multiple nucleic acids and mixtures thereof.

[0052] As used herein, the term “subject” means an entity or medium having testable or detectable genetic information. A subject may be a person, individual, or patient. A subject may be a vertebrate, such as a mammal. Non-limiting examples of mammals include humans, monkeys, domestic animals, sports animals, rodents, and pets. A subject may exhibit symptoms that indicate a health or physiological state or condition of the subject, such as a disease or disorder of the subject. Alternatively, a subject may be asymptomatic with respect to such a health or physiological state or condition.

[0053] As used herein, the term “sample” generally refers to a biological sample obtained from or derived from one or more subjects. A biological sample may be a cell-free or substantially cell-free biological sample, or may be processed or fractionated to produce a cell-free biological sample. For example, a cell-free biological sample may include cell-free ribonucleic acid (cfRNA), cell-free deoxyribonucleic acid (cfDNA), cell-free fetal DNA (cffDNA), proteins, autoantibodies, plasma, serum, urine, saliva, amniotic fluid, and derivatives thereof. Cell-free biological samples may be obtained from or derived from subjects using ethylenediaminetetraacetic acid (EDTA) collection tubes, cell-free RNA collection tubes (e.g., Streck® RNA Complete BCT®), or cell-free DNA collection tubes (e.g., Streck® Cell-Free BCT®). Cell-free biological samples may be derived from whole blood samples by fractionation (e.g., by differential centrifugation). A biological sample or its derivatives may contain cells. For example, the biological sample may be a blood sample or a derivative thereof (e.g., blood collected by blood collection tube or blood dropper).

[0054] As used herein, the term “cell-free sample” generally refers to a biological sample that is substantially free of intact cells. Cell-free samples may be obtained from biological samples that are substantially cell-free themselves, or from samples from which cells have been removed. Non-limiting examples of cell-free samples include those derived from blood, serum, plasma, urine, semen, saliva, feces, ductal exudate, lymph, and recovered lavage fluid.

[0055] As used herein, the term “colonic cell proliferative disorder” generally refers to a disease or condition involving a disorder or abnormality in the proliferation of cells of the colon or rectum. Non-exclusive examples of colonic cell proliferative disorders include adenomas (adenomatous polyps), polyposis disorders, Lynch syndrome, sessile serrated adenoma (SSA), progressive adenoma, dysplasia of the colon, adenoma of the colon, colorectal cancer, colon cancer, rectal cancer, colorectal cell tumor, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma. As used herein, the abbreviation “CRC” is used to identify a biological specimen of a subject diagnosed with colorectal cancer. As used herein, the abbreviation “AA” is used to identify a specimen of a subject diagnosed with at least one progressive adenoma. As used herein, the abbreviation “NAA” is used to identify a specimen of a subject diagnosed with a progressive adenoma or a benign tumor of the colon rather than colorectal cancer.

[0056] As used herein, the term "colon cancer" is a disease characterized by the cancerous transformation of cells in the intestinal tract below the small intestine (colon, e.g., cecum, ascending colon, transverse colon, descending colon, sigmoid colon).

[0057] As used herein, the term “colorectal adenoma” generally refers to a benign adenoma of the colorectal colon, also known as an adenomatous polyp, which is a precancerous stage of colorectal cancer. Colorectal adenomas may indicate a high risk of progression to colorectal cancer.

[0058] As used herein, the term “progressive colorectal adenoma” refers to an adenoma that is 10 mm or larger in size, or an adenoma that histologically has more than 20% high-grade dysplasia or villous components.

[0059] As used herein, the terms “risk of developing colonic cell proliferative disorder” or “high risk of developing colonic cell proliferative disorder” generally refer to an increased risk of developing colonic cell proliferative disorder in the near future compared to an individual who does not have colonic cell proliferative disorder or who is at low risk of developing it in the near future. As used herein, the term “near future” generally refers to a period of about one month to about two years, about six months to about eighteen months, or about one year.

[0060] As used herein, the terms “type” and “subtype” of cancer are generally used relative to each other. For example, one “type” of cancer, such as breast cancer, may be a “subtype” based on, for example, stage, morphology, histology, gene expression, receptor profile, mutation profile, aggressiveness, prognosis, and malignant characteristics. Similarly, “type” and “subtype” may be applied at a finer level. For example, one histological “type” may be identified as a “subtype” defined according to, for example, mutation profile or gene expression. Cancer “stage” is further used to refer to a classification of cancer types based on histological and pathological features related to disease progression.

[0061] The term “neoplasm” generally refers to any new, abnormal growth of tissue. Therefore, a neoplasm can be a pre-malignant neoplasm or a malignant neoplasm. The term “neoplasm-specific marker” refers to any biomaterial that can be used to indicate the presence of a neoplasm. Examples of biomaterials include, but are not limited to, nucleic acids, polypeptides, carbohydrates, fatty acids, cellular components (e.g., cell membranes and mitochondria), and whole cells. The term “colon neoplasm-specific marker” refers to any biomaterial that can be used to indicate the presence of a colon neoplasm (e.g., pre-malignant colon neoplasm; malignant colon neoplasm).

[0062] As used herein, the term “healthy” generally refers to subjects that do not have colonic cell proliferation disorders. Health is a dynamic state, but as used herein, the term refers to the pathological state of subjects that lack the disease condition mentioned in a particular statement. For example, when referring to a signature panel that can classify subjects having colorectal cancer, healthy individuals, healthy samples, or samples from healthy individuals refer to individuals that do not have colorectal cancer (CRC), progressive adenoma (AA), or benign adenoma (NAA). As used herein, the abbreviation “NAA” is used to identify samples from individuals that have been evaluated as negative for colorectal tumors, and therefore, in some embodiments, samples identified as NAA are included in the group of healthy samples. Other diseases or health conditions may be present in the subjects, but as used herein, the term “healthy” generally indicates the absence of the disease described for the purpose of comparison or classification between subjects with and without the disease condition discussed.

[0063] The term "minimal residual disease" or "MRD" generally refers to the small number of cancer cells remaining in a patient's body after cancer treatment. MRD testing may be performed to assess the effectiveness of cancer treatment and guide further treatment planning.

[0064] As used herein, the term “screening” generally refers to examining or testing a population of subjects at risk of developing colorectal cancer or colorectal adenoma, with the aim of distinguishing between healthy subjects and subjects with undiagnosed colorectal cancer or colorectal adenoma, or subjects at high risk of developing the aforementioned indications.

[0065] As used herein, the terms “less invasive biological sample” or “non-invasive sample” generally refer to any sample taken from a patient’s body without requiring any instruments other than a fine needle used to draw blood from the subject. In some embodiments, less invasive biological samples include samples of blood, serum, or plasma.

[0066] As used herein, the terms “upregulated” or “overexpressed” generally refer to an increase in expression levels of at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%, at least 110%, at least 120%, at least 130%, at least 140%, at least 150%, or 150% or more above a predetermined “threshold” or “cutoff value.”

[0067] As used herein, when referring to expression levels, the terms “threshold” or “cutoff value” generally refer to a reference expression level at which, if the expression level of a subject exceeds the threshold, cutoff, or reference level, the subject is likely to develop colorectal cancer or colorectal adenoma with predetermined sensitivity and specificity.

[0068] As used herein, the term “kit” generally includes, but is not limited to, any device suitable for carrying out the present invention, such as a microarray, bioarray, biochip, biochip array, or bead-based assay, without being limited to any specific device.

[0069] Sample assay Cell-free biological samples may be obtained from or derived from human subjects. Cell-free samples may be stored under various storage conditions before processing, such as different temperatures (e.g., room temperature, refrigerated or frozen conditions, e.g., 25°C, 4°C, -18°C, -20°C, or -80°C) or different suspensions (e.g., EDTA collection tubes, cell-free RNA collection tubes, or cell-free DNA collection tubes).

[0070] Cell-free biological samples can be obtained from subjects with cancer, subjects suspected of having cancer, or subjects that do not have cancer or are not suspected of having cancer.

[0071] Cell-free biological samples can be obtained before and / or after treatment of subjects with cancer. Cell-free biological samples may be obtained from subjects during treatment or during the treatment regimen. Multiple cell-free biological samples may be obtained from subjects to monitor the effects of treatment over time. Cell-free biological samples may be taken from subjects who are known or suspected to have cancer but have not received a definitive positive or negative diagnosis through clinical testing. Samples may be taken from subjects suspected to have cancer. Cell-free biological samples may be taken from subjects experiencing symptoms of unknown cause, such as fatigue, nausea, weight loss, pain and aches, weakness, or bleeding. Cell-free biological samples may be taken from subjects with the described symptoms. Cell-free biological samples may be taken from subjects at risk of developing cancer due to factors such as family history, age, hypertension or prehypertension, diabetes or prediabetes, overweight or obesity, environmental exposure, lifestyle risk factors (e.g., smoking, alcohol consumption, or drug use), or the presence of other risk factors.

[0072] Cell-free biological samples may include one or more assayable analytes, such as cell-free ribonucleic acid (cfRNA) molecules suitable for assays to generate transcriptomic data, cell-free deoxyribonucleic acid (cfDNA) molecules suitable for assays to generate genomic data, protein molecules (including autoantibodies) suitable for assays to generate proteomic data, or mixtures or combinations thereof.

[0073] After obtaining a cell-free biological sample from a subject, the cell-free biological sample may be processed to generate a dataset indicating impaired colonic cell proliferation in the subject. For example, this could involve the presence, absence, or quantitative evaluation of antibody molecules in the cell-free biological sample in a panel of autoantibodies. Processing a cell-free biological sample obtained from a subject may include (i) exposing the cell-free biological sample to conditions sufficient to separate, concentrate, or extract multiple autoantibodies, and (ii) assaying multiple autoantibody molecules to generate a dataset.

[0074] A biological sample may be used directly in an autoantibody assay to generate an autoantibody profile of the sample. In some embodiments, the biological sample may be concentrated for autoantibodies before the assay (e.g., using protein-conjugated microbeads). In one embodiment, the biological sample is a plasma sample, which is concentrated. The biological sample may be assayed using various laboratory methodologies to determine the presence and / or concentration or level of antibodies in the biological sample. In various embodiments, such approaches may include, but are not limited to, protein microarrays, high-density protein microarrays (e.g., CDI), ELISA, Meso Scale Discovery, bead-based immunoassays (e.g., Luminex® magnetic bead-based capture assays), secondary fluoroantibody assays, or combinations thereof for determining the autoantibody profile of a biological sample from a subject.

[0075] Signature Panel This disclosure provides a method and system for analyzing biological samples to obtain measurable features associated with the development of colonic cell proliferation disorder from combinations of autoantibody molecules identified in the samples. The collection of identified autoantibody molecules described herein has informational value in the creation of classifiers for detecting colonic cell proliferation disorder or its stages, and in the model thereof. While the identified autoantibody molecules may be individually informational and useful, the autoantibody molecules may be used in combinations described herein to form a signature panel in which the signatures are features of colonic cell proliferation disorder or its stages. Features from the signature panel may be processed using a trained algorithm (e.g., a machine learning model) to create a classifier configured to stratify populations of subjects having colonic cell proliferation disorder. This method is characterized by using one or more autoantibodies described in the signature panel. In one embodiment, a signature panel of at least three autoantibodies is useful for the classifier and method described herein.

[0076] The autoantibody signature panels described herein may enable rapid and specific analysis of specific autoantibodies associated with colonic cell proliferation disorders. The signature panels described and employed in the methods herein can be used to improve the diagnosis, prognosis, treatment selection, and monitoring (e.g., treatment monitoring) of colonic cell proliferation disorders.

[0077] This signature panel and method represent a significant improvement over current approaches to detecting early colonic cytoproliferative disorders from bodily fluid samples such as whole blood, plasma, or serum. Current methods used to detect and diagnose colonic cytoproliferative disorders include colonoscopy, sigmoidoscopy, and fecal occult blood testing for colon cancer. Compared to these methods, the method provided herein is far less invasive than colonoscopy and may be more sensitive, if not more sensitive, than sigmoidoscopy, fecal immunochemistry (FIT), and fecal occult blood testing (FOBT). The method provided herein offers significant advantages in terms of sensitivity and specificity through the favorable combination of using a gene panel and highly sensitive assay techniques.

[0078] This disclosure provides a method and system for autoantibody profiling of tumor antigen-associated autoantibodies ("tAAb" or "autoantibodies") associated with the detection and progression of colonic cell proliferative disorder (COLCPR). Specific embodiments of the present invention provide autoantibodies that are differentially abundant in samples of subjects with or at high risk of developing COCPR compared to corresponding samples of subjects without COCPR or at low risk of developing it. In one embodiment, each of the subjects at high risk of developing COCPR and the subjects at low risk of developing COCPR have non-invasive prodromal lesions (hereinafter, colorectal lesions) occurring in the colonic mucosa. Autoantibodies present in different amounts in samples of healthy subjects and subjects with COCPR may be used as biomarkers for the diagnosis, treatment, and / or prevention of COCPR.

[0079] To identify autoantibodies of informational value for the methods and classifiers described herein, plasma from patients with colonic cell proliferative disorder (COLC) and plasma from subjects without COC (control plasma or reference plasma) were examined to identify a signature panel of autoantibodies produced by patients with COC in response to COC and its respective reactive proteins. To this end, plasma from patients with COC and control plasma were examined using high-density protein microarrays. Protein microarrays offer a series of advantages compared to other approaches used for autoantibody identification: i) the proteins printed on the array are pre-known, thus preventing subsequent identification and eliminating the possibility of mimotope selection; and ii) all proteins are printed at similar concentrations, eliminating any predisposition to select specific proteins. This combination of factors allows for highly sensitive identification of biomarkers.

[0080] The autoantibodies identified herein can be used to differentiate subjects with colonic cell proliferation disorders from subjects without colonic cell proliferation disorders, to differentiate subjects at high risk of developing colonic cell proliferation disorders from subjects at low risk of developing them, or to identify subjects with precursors to colonic cell proliferation disorders. Thus, these autoantibodies can be used as an aid in guiding decisions regarding the monitoring, treatment, and management of colonic cell proliferation disorders.

[0081] In certain embodiments, disclosed herein is a panel of plasma tumor antigen-related autoantibodies (TAAbs) biomarkers useful for the early detection of colonic proliferative disorders and associated with the early detection of colorectal cancer.

[0082] In other embodiments, methods related to detection, diagnosis, and treatment are disclosed herein. Patient plasma is screened for tumor antigen-associated autoantibodies (TAAs) against tumor-derived proteins as an indicator of colonic proliferative disorder.

[0083] In one embodiment, the present disclosure provides an autoantibody panel characteristic of colonic cell proliferation disorders, comprising immunoglobulins against three or more antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, ASB9, NXN, ZBTB21, EYA1, GSPT1, MLIP, RBM38, ARMC5, TP53, BRD9, CDK4, PRMT6, PCOLCE, and SDCBP.

[0084] In one embodiment, the immunoglobulin is IgG, IgM, or a combination thereof.

[0085] In one embodiment, an autoantibody signature panel is useful for differentiating between healthy subjects, subjects with benign colorectal polyps, subjects with progressive adenoma, or subjects with colorectal cancer.

[0086] In one embodiment, the panel helps to indicate progressive adenoma and includes IgM autoantibodies against at least three antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, SDCBP, and CD20. In one embodiment, the panel includes IgM autoantibodies against UBE2, NME5, and CD20. In one embodiment, the panel helps to indicate progressive adenoma and includes IgG autoantibodies against at least three antigens selected from the group consisting of ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, and ASB9. In one embodiment, the panel includes IgG autoantibodies against ASB9, NAT6, Supt6h, and PRDM8.

[0087] In one embodiment, the panel helps to show a progressive adenoma and includes: The antibody comprises: 1) IgM autoantibodies against at least three antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, and SDCBP; 2) IgM autoantibodies against at least one antigen selected from the group consisting of UBE2S, NME5, and CD20; 3) IgG autoantibodies against at least three antigens selected from the group consisting of ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, PCOLCE, and ASB9; 4) IgG autoantibodies against at least one antigen selected from the group consisting of ASB9, NAT6, Supt6h, and PRDM8; or a combination thereof.

[0088] In one embodiment, the panel helps to show a sample of the subject having a benign polyp and includes IgG autoantibodies against at least three antigens selected from the group consisting of NXN, EYA1, GSPT1, and MLIP.

[0089] In one embodiment, the panel helps to show a sample of the subject having a benign polyp and contains an IgM autoantibody against ZBTB21.

[0090] In one embodiment, the panel helps to indicate colorectal cancer and includes IgM autoantibodies against at least three antigens selected from the group consisting of PELO, CDK4, MTCP1, PRMT6, PCOLCE, and ZBtb2. In one embodiment, the panel includes at least three IgG autoantibodies against TSSC4, BRD9, BCCIP, and TP53. In one embodiment, the panel includes IgM autoantibodies against CDK4, PRMT6, and MTCP1. In one embodiment, the panel includes IgG autoantibodies against TP53 and RBM38.

[0091] In one embodiment, the panel helps to indicate colorectal cancer and includes: 1) IgM autoantibodies against at least three antigens selected from the group consisting of PELO, CDK4, MTP1, PRMT6, ZBTB2, and PCOLCE; 2) IgM autoantibodies against at least one antigen selected from the group consisting of CDK4, MTCP1, and PCOLCE; 3) IgG autoantibodies against at least three antigens selected from the group consisting of TSSC4, BRD9, BCCIP, and TP53; 4) IgG autoantibodies against TP53, or a combination thereof.

[0092] In some embodiments, a predetermined set of autoantibodies includes autoantibodies against at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, or more antigens, such as those described herein.

[0093] In some embodiments, the autoantibodies in a predetermined panel are IgM and IgG autoantibodies. In one embodiment, the autoantibodies in a predetermined panel are IgM autoantibodies. In another embodiment, the autoantibodies in a predetermined panel are IgG autoantibodies.

[0094] Classifiers, machine learning models and systems Machine learning approaches are used to characterize autoantibody data derived from biological samples obtained from subjects and to identify panels of beneficial autoantibodies. Identified panels of autoantibodies beneficial to colonic cell proliferation disorders are useful for training classifier models that help differentiate between samples from healthy subjects and samples from subjects with colonic cell proliferation disorders.

[0095] Furthermore, described herein are machine learning model classifiers trained with autoantibodies described herein expressed in plasma samples from healthy subjects and plasma samples from subjects with colonic cell proliferation disorders. By training the machine learning model, a classifier is obtained having a predetermined set of autoantibody biomarkers ("autoantibody panel" or "signature panel") useful for classifying healthy subjects or subjects with colonic cell proliferation disorders. In one example, a method is provided for a blood-based, minimally invasive autoantibody assay that may be used in subjects with colorectal lesions to assess histological severity. In another embodiment, autoantibodies indicating colonic cell proliferation disorders are detected in cell-free samples from subjects, such as whole blood, plasma, or serum. Thus, the autoantibodies disclosed herein may be used to differentiate between the presence or absence of colonic cell proliferation disorders, high-risk colorectal lesions, or low-risk colorectal lesions requiring treatment such as surgical resection, immunotherapy, radiotherapy, or chemotherapy, and monitoring of low-risk colorectal lesions. Monitoring and confirmation of the presence of colonic cell proliferation disorders or lesions can be performed, for example, by colonoscopy, ultrasound, MM, or CT scan.

[0096] In various examples, autoantibody features are used as input datasets to trained algorithms (e.g., machine learning models or classifiers) to discover correlations between autoantibody profiles and patient groups. Examples of such patient groups include the presence, stage, subtype, responder vs. non-responder, and progression vs. non-progression. In various examples, feature matrices are generated to compare samples obtained from subjects with known diseases or features. In some embodiments, samples are obtained from healthy subjects or subjects with none of the known indications, and from patients known to have cancer.

[0097] As used herein, and in relation to machine learning and pattern recognition, the term “feature” generally refers to individual measurable properties or characteristics of an observed phenomenon. The concept of “feature” is related, for example, to the concept of explanatory variables used in statistical methods such as linear regression and logistic regression, but not limited to these. Features are usually numerical, but structural features such as strings and graphs are used in syntactic pattern recognition.

[0098] As used herein, the term “input feature” (or “feature”) generally refers to a variable that a trained algorithm (e.g., a model or classifier) ​​uses to predict the output classification (label) of a sample, such as state, autoantibody identity, antibody sequence content (e.g., mutation), suggested data acquisition operation, or suggested treatment. The values ​​of the variables are determined for the sample and may be used to determine the classification.

[0099] For multiple assays, the system identifies the set of features to input into a trained algorithm (e.g., a machine learning model or classifier). The system performs the assay for each molecular class and forms a feature vector from the measurements. The system inputs the feature vector into the machine learning model to obtain an output classification of whether or not the biological sample possesses a specific characteristic.

[0100] In some embodiments, the machine learning model provides a classifier capable of distinguishing between two or more groups or classes within the characteristics of an object or a group of objects. In some embodiments, the classifier is a trained machine learning classifier.

[0101] In some embodiments, biomarker-rich loci or features in cancer tissue are assayed to form profiles. Receiver-operating characteristic (ROC) curves can be generated by plotting the performance of specific features (e.g., any of the biomarkers described herein and / or any additional biomedical information) that differentiate two populations (e.g., subjects that respond to a therapeutic agent and individuals that do not). In some embodiments, feature data for the entire population (e.g., cases and controls) are sorted in ascending order based on the value of a single feature.

[0102] In various examples, the specified characteristics are selected from healthy vs. cancer, disease subtype, disease stage, progressive vs. non-progressive, and responder vs. non-responder.

[0103] In some embodiments, colonic cell proliferation disorders are selected from the group consisting of adenoma (adenomatous polyp), polyposis, Lynch syndrome, sessile serrated adenoma (SSA), progressive adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal cell tumor, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.

[0104] A. Data Analysis In some examples, this disclosure provides systems, methods, or kits having data analysis implemented by software applications, computing hardware, or both. In various examples, the analysis application or system includes at least a data receiving module, a data preprocessing module, a data analysis module (which can operate on one or more types of genomic data), a data interpretation module, or a data visualization module. In some embodiments, the data receiving module includes a computer system that connects laboratory hardware or equipment to a computer system that processes laboratory data. In some embodiments, the data preprocessing module includes a hardware system or computer software that performs operations on the data in preparation for analysis. Examples of operations that can be applied to the data by the preprocessing module include affine transformation, denoising operations, data cleaning, reformatting, or subsampling. The data analysis module may be specialized for analyzing genomic data from one or more genomic materials, for example, by incorporating assembled genomic sequences and performing probabilistic and statistical analyses to identify abnormal patterns associated with disease, pathology, condition, risk, illness, or phenotype. The data interpretation module can use analytical methods derived from, for example, statistics, mathematics, or biology to support an understanding of the relationship between identified anomaly patterns and health status, functional status, prognosis, or risk. The data visualization module can use mathematical modeling, computer graphics, or rendering methods to create a visual representation of the data that can facilitate the understanding or interpretation of the results.

[0105] In various examples, machine learning methods are applied to differentiate between samples within a population of samples. In some embodiments, machine learning methods are applied to differentiate between healthy samples and samples with progressive diseases (e.g., adenomas).

[0106] In some embodiments, one or more machine learning operations used to train the prediction engine include one or more of the following: generalized linear models, generalized additive models, nonparametric regression operations, random forest classifiers, spatial regression, Bayesian regression models, time series analysis, Bayesian networks, Gaussian networks, decision tree learning operations, artificial neural networks, recurrent neural networks, reinforcement learning operations, linear nonlinear regression operations, support vector machines, clustering operations, and genetic algorithm operations.

[0107] In various examples, computer processing methods are selected from a group consisting of logistic regression, multiple linear regression (MLR), dimensionality reduction, partial least squares (PLS) regression, principal component regression, autoencoders, variational autoencoders, singular value decomposition, Fourier bases, wavelets, discriminant analysis, support vector machines, decision trees, classification trees and regression trees (CART), tree-based methods, random forests, gradient boosted trees, logistic regression, matrix decomposition, multidimensional scaling (MDS), dimensionality reduction methods, t-distribution stochastic nearest neighbor embedding (t-SNE), multilayer perceptrons (MLP), network clustering, neuro-fuzzy, and artificial neural networks.

[0108] In some examples, the methods disclosed herein may include computational analysis of nucleic acid sequence data of samples from one or more subjects.

[0109] B. Classifier generation In one embodiment, the systems and methods disclosed herein provide a classifier generated based on characteristic information obtained from autoantibody analysis of a biological sample containing autoantibodies. The classifier forms part of a predictive engine for differentiating groups within a population based on characteristics identified in the biological sample, such as autoantibodies. The collective representation of autoantibody information in a biological sample is sometimes referred to as an autoantibody profile.

[0110] In some embodiments, the classifier is generated by normalizing autoantibody information by formatting similar portions of autoantibody information into a unified format and scale, storing the normalized autoantibody information in a column database, and training a predictive engine by applying one or more machine learning operations to the stored normalized autoantibody information, wherein the predictive engine maps one or more feature combinations to a particular population, applies the predictive engine to accessed field information to identify objects associated with a group, and classifies the objects into a group.

[0111] Specificity, as used herein, generally refers to "the probability of testing negative among people who do not have the disease." Specificity can be calculated by dividing the number of disease-free individuals who tested negative by the total number of disease-free subjects.

[0112] In various examples, the model, classifier, or predictive test has specificity of at least approximately 40%, at least approximately 45%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 95%, or at least approximately 99%.

[0113] As used herein, sensitivity generally refers to "the probability of a person with the disease testing positive." Sensitivity can be calculated by dividing the number of diseased individuals who test positive by the total number of diseased individuals.

[0114] In various examples, the model, classifier, or predictive test has a sensitivity of at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%.

[0115] C. Digital Processing Unit In some embodiments, what is described herein is a digital processing unit or a use thereof. In some examples, the digital processing unit may include one or more hardware central processing units (CPUs), graphics processing units (GPUs), or tensor processing units (TPUs) that perform the functions of the device. In some examples, the digital processing unit may include an operating system configured to execute executable instructions.

[0116] In some examples, the digital processing unit may optionally be connected to a computer network. In some examples, the digital processing unit may optionally be connected to the internet. In some examples, the digital processing unit may optionally be connected to a cloud computing infrastructure. In some examples, the digital processing unit may optionally be connected to an intranet. In some examples, the digital processing unit may optionally be connected to a data storage device.

[0117] Suitable digital processing devices include, but are not limited to, server computers, desktop computers, laptop computers, notebook computers, subnotebook computers, netbook computers, netpad computers, set-top computers, handheld computers, internet appliances, mobile smartphones, and tablet computers. Suitable tablet computers may include, for example, booklet, slate, and convertible configurations.

[0118] In some examples, a digital processing device may include an operating system configured to execute executable instructions. For example, an operating system may include software, including programs and data, that manage the device's hardware and provide services for running applications. Non-exclusive examples of operating systems include Ubuntu, FreeBSD, OpenBSD, NetBSD®, Linux®, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Non-exclusive examples of suitable personal computer operating systems include Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and UNIX®-like operating systems, such as GNU / Linux®. In some examples, an operating system may be provided by cloud computing, and cloud computing resources may be provided by one or more service providers.

[0119] In some examples, the device may include a storage device and / or memory device. A storage device and / or memory device may be one or more physical devices used to store data or programs on a temporary or permanent basis. In some examples, the device may be volatile memory and require power to maintain the stored information. In some examples, the device may be non-volatile memory and retain the stored information when the digital processing device is not powered. In some embodiments, non-volatile memory may include flash memory. In some examples, non-volatile memory may include dynamic random access memory (DRAM). In some examples, non-volatile memory may include ferroelectric random access memory (FRAM®). In some examples, non-volatile memory may include phase-change random access memory (PRAM).

[0120] In some examples, the device may be a storage device, including, for example, a CD-ROM, DVD, flash memory device, magnetic disk drive, magnetic tape drive, optical disk drive, and cloud computing-based storage device. In some examples, the storage device and / or memory device may be a combination of devices such as those disclosed herein. In some examples, the digital processing device may include a display for transmitting visual information to a user. In some examples, the display may be a cathode ray tube (CRT). In some examples, the display may be a liquid crystal display (LCD). In some examples, the display may be a thin-film transistor liquid crystal display (TFT-LCD). In some examples, the display may be an organic light-emitting diode (OLED) display. In some examples, the OLED display may be a passive-matrix OLED (PMOLED) or an active-matrix OLED (AMOLED) display. In some examples, the display may be a plasma display. In some examples, the display may be a video projector. In some examples, the display may be a combination of devices such as those disclosed herein.

[0121] In some examples, the digital processing device may include an input device for receiving information from a user. In some examples, the input device may be a keyboard. In some examples, the input device may be a pointing device including, for example, a mouse, trackball, trackpad, joystick, game controller, or stylus. In some examples, the input device may be a touchscreen or multitouchscreen. In some examples, the input device may be a microphone for capturing voice or other audio input. In some examples, the input device may be a video camera for capturing motion or visual input. In some examples, the input device may be a combination of devices such as those disclosed herein.

[0122] D. Non-temporary computer-readable storage media In some examples, the inventive features disclosed herein may include one or more non-temporary computer-readable storage media encoded with a program containing instructions executable by an operating system of a networked digital processing device, which is optional. In some examples, the computer-readable storage media may be a tangible component of the digital processing device. In some examples, the computer-readable storage media may be removable from the digital processing device, which is optional. In some examples, the computer-readable storage media may include, for example, CD-ROMs, DVDs, flash memory devices, solid memory, magnetic disk drives, magnetic tape drives, optical disk drives, cloud computing systems and services, etc. In some examples, the programs and instructions may be coded on the medium permanently, substantially permanently, semi-permanently, or non-temporarily.

[0123] E. Computer Systems This disclosure provides a computer system programmed to carry out the methods described herein. Figure 1 shows a computer system (101) programmed, or otherwise configured, to store, process, identify, or interpret patient data, biodata, biosequences, reference sequences, and autoantibody profiles. The computer system (101) may process various aspects of the patient data, biodata, biosequences, reference sequences, and autoantibody profiles of this disclosure. The computer system (101) may be an electronic device of the user or the computer system, and may be remotely located relative to the electronic device. The electronic device may be a mobile electronic device.

[0124] The computer system (101) includes a central processing unit (CPU) (105), which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system (101) further includes memory or storage locations (110) (e.g., random-access memory, read-only memory, flash memory), electronic storage devices (115) (e.g., hard disks), communication interfaces (120) (e.g., network adapters) for communicating with one or more other systems, and peripheral devices (125), e.g., caches, other memory, data storage devices, and / or electronic display adapters. The memory (110), storage devices (115), interfaces (120), and peripheral devices (125) communicate with the CPU (105) through a communication bus (solid line), such as a motherboard. The storage device (115) may be a data storage device (or data repository) for storing data. A computer system (101) may be operationally connected to a computer network ("Network") (130) with the help of a communication interface (120). The Network (130) may be the Internet and / or an extranet, an intranet and / or an extranet in communication with the Internet. In some examples, the Network (130) is a telecommunications and / or data network. The Network (130) may include one or more computer servers, which may enable distributed computing such as cloud computing. The Network (130) may, in some cases, implement a peer-to-peer network with the help of the computer system (101), thereby enabling devices connected to the computer system (101) to act as clients or servers.

[0125] The CPU(105) can execute a set of machine-readable instructions that can be integrated by a program or software. These instructions may be stored in a storage location such as memory(110). The instructions may be directed to the CPU(105), which can then be programmed or otherwise configured to perform the methods of this disclosure. Examples of operations performed by the CPU(105) include fetching, decoding, executing, and writing back.

[0126] The CPU (105) may be part of a circuit, such as an integrated circuit. One or more other components of the system (101) may be included in the circuit. In some examples, the circuit is an application-specific integrated circuit (ASIC).

[0127] The storage device (115) can store files such as drivers, libraries, and saved programs. The storage device (115) can also store user data, such as user preferences and user programs. The computer system (101) may include one or more additional data storage devices located outside the computer system (101), such as being located on a remote server that is in communication with the computer system (101) via an intranet or the internet, in some examples.

[0128] The computer system (101) can communicate with one or more remote computer systems via a network (130). For example, the computer system (101) can communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), slate or tablet PCs (e.g., Apple® iPad®, Samsung® Galaxy Tab), telephones, smartphones (e.g., Apple® iPhone®, Android-enabled devices, Blackberry®), or personal digital assistants. The user can access the computer system (101) via the network (130).

[0129] Methods described herein may be implemented by machine-executable code (e.g., computer processor) stored in an electronic storage location of a computer system (101), such as memory (110) or electronic storage device (115). The machine-executable code or machine-readable code may be provided in the form of software. During use, the code may be executed by the processor (105). In some examples, the code may be retrieved from storage device (115) and stored in memory (110) for easy access by the processor (105). In some examples, electronic storage device (115) may be omitted, and machine-executable instructions are stored in memory (110).

[0130] The code may be pre-compiled and configured for use with a machine having a processor suitable for executing the code, or it may be interpreted or compiled during execution time. The code may be supplied in a programming language that can be chosen to make the code executable in a pre-compiled, interpreted, or as-compiled form.

[0131] Embodiments of systems and methods provided herein, such as computer systems (101), can be integrated during programming. Various embodiments of this technology can typically be considered as “products” or “products of manufacture” in the form of machine (or processor) executable code and / or associated data carried on or embedded therein on some kind of machine-readable medium. Machine executable code can be stored in electronic storage devices such as memory (e.g., read-only memory, random-access memory, flash memory) or hard disks. “Storage” type media can include any or all of the tangible memory of a computer or processor, or its associated modules, such as various semiconductor memories, tape drives, disk drives, etc., which can provide a non-temporary recording medium for programming software at any time. All or part of the software is sometimes communicated over the Internet or various other telecommunication networks. Such communication can enable, for example, the loading of software from one computer or processor to another computer or processor, for example, from a management server or host computer to an application server computer platform. Therefore, other types of media that may have software elements include optical waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices via wired and optical terrestrial communication line networks and over various air-links. Physical elements that carry such waves, such as wired or wireless links and optical links, can also be considered media with software. Unless limited to non-temporary and tangible “storage” media as used herein, terms such as computer or machine “readable media” generally refer to media involved in providing instructions to a processor for execution.

[0132] Therefore, machine-readable media such as computer executable code may take many forms, but are not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include optical disks or magnetic disks, such as any storage device in any computer, which may be used to implement databases, etc. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include copper wires and optical fibers, such as coaxial cables and wires including buses in computer systems. Carrier wave transmission media may take the form of electrical signals or electromagnetic signals, or sound waves or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Therefore, common forms of computer-readable media include, for example: floppy disks, flexible disks, hard disks, magnetic tapes, other magnetic media, CD-ROMs, DVDs or DVD-ROMs, other optical media, punch cards, paper tapes, other physical storage media having hole patterns, RAM, ROMs, PROMs and EPROMs, FLASH-EPROMs, other memory chips or cartridges, carriers carrying data or instructions, cables or links transmitting such carriers, or other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media can be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0133] The computer system (101) may include, or be in communication with, an electronic display (135) having a user interface (UI) (140) for providing, for example, the analysis of nucleic acid sequences, concentrated nucleic acid samples, autoantibody profiles, expression profiles, and RNA expression profiles. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.

[0134] The methods and systems of this disclosure may be implemented by one or more algorithms. The algorithms may be implemented by software after execution by a central processing unit (105). The algorithms can store, process, identify, or interpret, for example, patient data, biometric data, biometric sequences, and reference sequences.

[0135] While specific examples of methods and systems have been shown and described herein, those skilled in the art will understand that these are provided for illustrative purposes only and are not intended to limit the scope herein. Numerous variations, alterations, and substitutions will be conceived by those skilled in the art without departing from the scope described herein. Furthermore, all aspects of the methods and systems described herein are not limited to specific descriptions, configurations, or relative proportions described herein, which depend on various conditions and variables, and it will be understood that this description is intended to include such alternatives, modifications, variations, or equivalents.

[0136] In some embodiments, the inventive features disclosed herein may include at least one computer program, or the use of a computer program. A computer program can be a sequence of instructions written to perform a specific task, executable on the CPU, GPU, or TPU of a digital processing unit. Computer-readable instructions may be implemented as program modules, such as functions, objects, application programming interfaces (APIs), or data structures, that perform a specific task or implement a specific extracted data type. In light of the disclosures provided herein, a computer program may be written in various versions of various languages.

[0137] The functionality of computer-readable instructions can be combined or distributed as needed in various environments. In some examples, a computer program may contain one instruction sequence. In some examples, a computer program may contain multiple instruction sequences. In some examples, a computer program may be provided from one location. In some examples, a computer program may be provided from multiple locations. In some examples, a computer program may contain one or more software modules. In some examples, a computer program may, in part or in whole, contain one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plugins, extensions, add-ins, or add-ons, or a combination thereof.

[0138] In some examples, the computer processing may be methods from statistics, mathematics, biology, or any combination thereof. In some examples, the computer processing methods include, for example, dimensionality reduction methods including logistic regression, dimensionality reduction, principal component analysis, autoencoders, singular value decomposition, Fourier-based methods, wavelets, discriminant analysis, support vector machines, tree-based methods, random forests, gradient-boosted trees, logistic regression, matrix factorization, network clustering, and neural networks.

[0139] In some examples, computer processing methods include supervised machine learning methods, such as regression, support vector machines, tree-based methods, and networks.

[0140] In some examples, the computer processing methods are unsupervised machine learning methods, including, for instance, clustering, networking, principal component analysis, and matrix factorization.

[0141] F. Database In some examples, the inventive features disclosed herein may include one or more databases, or the use thereof for storing patient data, biometric data, biometric sequences, reference sequences, or autoantibody profiles. Reference sequences may be derived from the databases. With regard to the disclosures provided herein, many databases may be suitable for storing and retrieving sequence information. In some examples, suitable databases may include, for example, relational databases, non-relational databases, object-oriented databases, object databases, entity-relational model databases, associative databases, and XML databases. In some examples, the databases may be internet-based. In some examples, the databases may be web-based. In some examples, the databases may be cloud computing-based. In some examples, the databases may be based on one or more local computer storage devices.

[0142] In one embodiment, the Disclosure provides a non-temporary computer-readable medium containing instructions that direct a processor to perform the methods disclosed herein.

[0143] In one embodiment, the disclosure provides a computing device including a computer-readable medium.

[0144] In other embodiments, the Disclosure provides a system for performing the classification of biological samples, the system is a) A receiving container for receiving multiple training samples, each of which training samples has multiple classes of molecules, and each of which training samples contains one or more known labels, b) A feature module that identifies a set of features corresponding to an assay that can be manipulated to be input into a machine learning model for each of a set of training samples, wherein the set of features corresponds to the properties of molecules in the set of training samples, and for each of the set of training samples, the system can be manipulated to subject multiple classes of molecules in the training sample to multiple different assays to obtain a set of measurements, where each set of measurements is a value from one assay applied to a class of molecules in the training sample, and multiple sets of measurements are obtained for multiple training samples, the feature module, c) An analysis module for analyzing a set of measurements to obtain a training vector for a training sample, wherein the training vector comprises N sets of feature values ​​of features of a corresponding assay, each feature value corresponds to a feature and comprises one or more measurements, and the training vector is formed using at least one feature from at least two features of N sets of features corresponding to a first subset of multiple different assays, d) A labeling module that notifies the system of the training vector using the parameters of the machine learning model in order to obtain output labels for multiple training samples. e) A comparator module that compares the output label with a known label of the training sample. f) A training module that iteratively searches for optimal parameter values ​​as part of training a machine learning model based on a comparison of the output label with known labels of the training sample, and g) Includes parameters for the machine learning model and an output module for providing a set of features for the machine learning model.

[0145] Methods for classifying objects within a group The disclosed method aims to identify parameters of autoantibody expression associated with colonic cell proliferation disorder through the analysis of expressed autoantibodies in subjects. This method is intended for use in improving the diagnosis, treatment, and monitoring of colonic cell proliferation disorder, more specifically by enabling improved identification and differentiation of the stage or subclass of the disorder and genetic predisposition to the disorder.

[0146] In some embodiments, the method includes the step of analyzing the differential expression of autoantibodies in biological samples from subjects in a population.

[0147] Generally, this disclosure provides a method for detecting colonic cell proliferation disorders that can be applied to cell-free samples, for example, a method for detecting the presence and characterization of autoantibodies between subjects with and without colonic cell proliferation disorders, or between different colonic cell proliferation disorders. The method utilizes the detection of autoantibodies as a basic "positive" or "negative" for colonic cell proliferation disorder signals compared to healthy subjects without colonic cell proliferation disorders.

[0148] In some embodiments, colonic cell proliferation disorders are selected from the group consisting of adenoma (adenomatous polyp), polyposis, Lynch syndrome, sessile serrated adenoma (SSA), progressive adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal cell tumor, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.

[0149] In a third aspect, the Disclosure provides a method for determining the autoantibody profile of a biological sample, the method being: a) A step of obtaining a biological sample containing autoantibodies from a subject, and b) To provide a target autoantibody profile, the process includes measuring the amount of antibodies from a predetermined autoantibody panel containing autoantibodies against three or more antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, ASB9, NXN, ZBTB21, EYA1, GSPT1, MLIP, RBM38, ARMC5, TP53, BRD9, CDK4, PRMT6, PCOLCE, and SDCBP.

[0150] In some embodiments, the autoantibody profile is associated with colonic cell proliferation disorder and provides a classification of subjects as having colonic cell proliferation disorder.

[0151] In some embodiments, the biological sample obtained from the subject is selected from the group consisting of body fluids, feces, colonic effluent, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, and combinations thereof.

[0152] In some embodiments, colonic cell proliferation disorders are selected from the group consisting of adenoma (adenomatous polyp), polyposis, Lynch syndrome, sessile serrated adenoma (SSA), progressive adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal cell tumor, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.

[0153] In some embodiments, colonic cell proliferation disorders are selected from the group consisting of stage 1 colorectal cancer, stage 2 colorectal cancer, stage 3 colorectal cancer, and stage 4 colorectal cancer.

[0154] In some embodiments, the progressive adenoma is a tubular adenoma, tubular villous adenoma, chorioadenoma, adenocarcinoma, or hyperplastic polyp.

[0155] In a fourth aspect, the Disclosure provides a method for detecting colonic cell proliferation disorder in a subject, the method being: a) A step of obtaining a biological sample containing autoantibodies from a subject, b) A step of measuring the amount of antibodies from a predetermined autoantibody panel containing autoantibodies against three or more antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, ASB9, NXN, ZBTB21, EYA1, GSPT1, MLIP, RBM38, ARMC5, TP53, BRD9, CDK4, PRMT6, PCOLCE, and SDCBP, in order to provide the target autoantibody profile, and c) The process includes processing autoantibody profiles using a machine learning model trained to differentiate between healthy subjects and subjects with colonic cell proliferation disorder, provide output values ​​associated with the presence of colonic cell proliferation disorder, and thereby indicate the presence of colonic cell proliferation disorder in the subjects.

[0156] In some embodiments, the biological sample obtained from the subject is selected from the group consisting of body fluids, feces, colonic effluent, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, and combinations thereof.

[0157] In other embodiments, the present invention relates to a method for detecting the binding of autoantibodies to proteins in order to generate an autoantibody profile of a sample, the method being: a) A step of contacting a biological sample with a protein or fragment thereof that is easily recognized by an autoantibody, and b) A step to provide an autoantibody profile of a sample, comprising detecting the formation of an antibody-protein complex formed by the binding of an antibody to a protein or a fragment thereof, wherein the protein is selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, ASB9, NXN, ZBTB21, EYA1, GSPT1, MLIP, RBM38, ARMC5, TP53, BRD9, CDK4, PRMT6, PCOLCE, and SDCBP.

[0158] In other embodiments, the present invention relates to a method for obtaining data in a biological sample from a subject, comprising the step of detecting at least three autoantibodies against proteins, wherein the at least three autoantibodies are selected from the group consisting of autoantibodies against UBE2S protein, CD20 protein, ASB9 protein, PRDM8 protein, CDK4 protein, MTCP1 protein, and TP53 protein. In some embodiments, the method further comprises the step of determining the level of autoantibodies in the sample.

[0159] In some embodiments, colonic cell proliferation disorders are selected from the group consisting of adenoma (adenomatous polyp), polyposis, Lynch syndrome, sessile serrated adenoma (SSA), progressive adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal cell tumor, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.

[0160] In some embodiments, colonic cell proliferation disorders are selected from the group consisting of stage 1 colorectal cancer, stage 2 colorectal cancer, stage 3 colorectal cancer, and stage 4 colorectal cancer.

[0161] In other embodiments, the Disclosure provides a method for determining the autoantibody profile of a biological sample, the method being: a) A step of obtaining a biological sample containing autoantibodies from a subject, and b) To provide a target autoantibody profile, the process includes measuring the amount of antibodies from a predetermined autoantibody panel containing autoantibodies against three or more antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, ASB9, NXN, ZBTB21, EYA1, GSPT1, MLIP, RBM38, ARMC5, TP53, BRD9, CDK4, PRMT6, PCOLCE, and SDCBP.

[0162] In another aspect, the Disclosure provides a method for detecting colonic cell proliferation disorder in a subject, the method being a) A step of obtaining a biological sample containing autoantibodies from a subject, b) A step of measuring the amount of antibodies from a predetermined autoantibody panel containing autoantibodies against three or more antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, ASB9, NXN, ZBTB21, EYA1, GSPT1, MLIP, RBM38, ARMC5, TP53, BRD9, CDK4, PRMT6, PCOLCE, and SDCBP, in order to provide the target autoantibody profile. c) Processing the autoantibody profiles of subjects using a machine learning model trained to differentiate between subjects without colonic cell proliferation disorder and subjects with colonic cell proliferation disorder, and d) The process includes determining values ​​associated with subjects having impaired colonic cell proliferation using a machine learning model that is at least partially based on autoantibody profiles, thereby detecting impaired colonic cell proliferation in the subjects.

[0163] In other embodiments, the Disclosure provides a method for monitoring minimal residual disease in subjects who have been treated for a disease, the method comprising the steps of: determining an autoantibody profile of a biological sample from a subject using an autoantibody panel containing autoantibodies against antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, ASB9, NXN, ZBTB21, EYA1, GSPT1, MLIP, RBM38, ARMC5, TP53, BRD9, CDK4, PRMT6, PCOLCE, and SDCBP, thereby generating a baseline autoantibody state; and, after generating the baseline autoantibody state, determining an autoantibody profile of a biological sample obtained from the subject at one or more time points, thereby generating a current autoantibody state, wherein the change between the baseline autoantibody state and the current autoantibody state indicates a change in the subject's minimal residual disease.

[0164] The trained machine learning methods, models, and discriminant classifiers described herein may be applied to a variety of medical applications, including cancer detection, diagnosis, and treatment response. Since the models can be trained on metadata of the subject and features derived from the analyte, the application can be tailored to stratify subjects within a population and guide treatment decisions accordingly.

[0165] diagnosis The methods and systems provided herein may perform predictive analytics using an artificial intelligence-based approach to analyze data obtained from subjects (patients) to generate a diagnostic output for subjects having cancer (e.g., colorectal cancer). For example, an application may apply a predictive algorithm to the obtained data to generate a diagnosis for subjects having cancer. The predictive algorithm may include an artificial intelligence-based predictor, such as a machine learning-based predictor, configured to process the obtained data to generate a diagnosis for subjects having cancer.

[0166] A machine learning predictor may be trained by using a dataset, for example, a dataset generated by performing autoantibody assays using the signature panel described herein on biological samples of a subject from one or more sets of patients with cancer, as input, and using the known diagnostic results of the subject (e.g., staging and / or tumor rate) as output to the machine learning predictor.

[0167] A training dataset (for example, a dataset generated by performing autoantibody assays on biological samples of a subject using the signature panel described herein) may be generated from, for example, one or more sets of subjects having common characteristics (features) and results (labels). The training dataset may include a set of features and labels corresponding to diagnostic features. Features may include, for example, a specific range or category of autoantibody assay measurements, such as the presence or characteristics of autoantibodies in biological samples obtained from healthy individuals and diseased individuals. For example, a set of features collected from predetermined subjects at a given time point may collectively function as a diagnostic signature, indicating the identified cancer of the subject at that predetermined time point. Features may further include labels indicating the diagnostic outcome of the subject, such as one or more cancers.

[0168] The label may include results such as the outcome of a known diagnosis of the subject (e.g., staging and / or tumor fractionation). The results may include characteristics associated with the subject's cancer. For example, characteristics may indicate that the subject has one or more cancers.

[0169] The training set (e.g., the training dataset) may be selected by random sampling of sets of data corresponding to one or more sets of subjects (e.g., retrospective and / or prospective cohorts of patients with or without one or more cancers). Alternatively, the training set (e.g., the training dataset) may be selected by proportional sampling of sets of data corresponding to one or more sets of subjects (e.g., retrospective and / or prospective cohorts of patients with or without one or more cancers). The training set may be balanced across sets of data corresponding to one or more sets of subjects (e.g., patients from different clinical sites or trials). The machine learning predictor may be trained until certain predetermined conditions regarding accuracy or performance are met, such as having a minimum desired value corresponding to a measure of diagnostic accuracy. For example, the measure of diagnostic accuracy may correspond to the prediction of a diagnosis, staging, or tumor fraction of one or more cancers in the subjects.

[0170] Examples of measures of diagnostic accuracy may include sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), precision, and area under the receiver operating characteristic (ROC) curve (AUC) corresponding to the diagnostic accuracy of detecting or predicting cancer (e.g., colorectal cancer).

[0171] In one aspect, the disclosure provides a method for using a classifier that can identify a target population, the method being a) A step of obtaining a biological sample containing autoantibodies from a subject, b) A step of measuring the amount of antibodies from a predetermined autoantibody panel containing autoantibodies against three or more antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, ANKHD1, TXNL1, NAT6, Supt6h, OTUD5, PNKP, SRSF7, ASB9, PRDM8, NXN, ZBTB21, EYA1, GSPT1, MLIP, RBM38, ARMC5, TP53, BRD9, CDK4, PRMT6, PCOLCE, and SDCBP, in order to provide the target autoantibody profile. c) The step of processing the target autoantibody profile using a machine learning model trained to differentiate between two or more populations, and d) A step of using a machine learning model based on at least a portion of the autoantibody profile to determine values ​​associated with a population, thereby differentiating the target population.

[0172] In other embodiments, the Disclosure provides a method for identifying a target cancer, the method comprising the steps of: a) obtaining a biological sample containing autoantibodies from the target; b) A step of measuring the amount of antibodies from a predetermined autoantibody panel containing autoantibodies against three or more antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, ASB9, NXN, ZBTB21, EYA1, GSPT1, MLIP, RBM38, ARMC5, TP53, BRD9, CDK4, PRMT6, PCOLCE, and SDCBP, in order to provide the target autoantibody profile, and c) The process includes processing autoantibody profiles using a machine learning model trained to indicate the presence of colonic cell proliferation disorder in a subject to differentiate between a healthy subject and a subject with colonic cell proliferation disorder, and to provide output values ​​associated with the presence of colonic cell proliferation disorder, thereby generating the likelihood that the subject will develop the cancer.

[0173] Various statistical and mathematical methods may be used to establish expression thresholds or cutoff levels. Threshold or cutoff expression levels for a particular biomarker may be selected based on data from receiver operating characteristic (ROC) plots, for example, as described in the examples and figures disclosed herein. Those skilled in the art will understand that these threshold or cutoff expression levels can be varied, for example, by shifting along the ROC plot for a particular biomarker or combination thereof, to obtain different values ​​of sensitivity or specificity, thereby affecting the overall assay performance. For example, if the goal is a clinically robust diagnostic method, high sensitivity should be prioritized. However, if the goal is a cost-effective method, high specificity should be prioritized. The best cutoff refers to the value obtained from the ROC plot for a particular biomarker that yields the highest sensitivity and specificity. Sensitivity and specificity values ​​are calculated over a range of thresholds (cutoffs). Therefore, the threshold or cutoff value may be selected such that the sensitivity and / or specificity is at least about 50%, and may be, for example, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% in at least 60% of the patient population assayed, or at least 65%, at least 70%, at least 75%, or at least 80% of the patient population assayed.

[0174] As a result, some embodiments of the present invention are performed by determining the presence and / or level of at least the autoantibodies cited above in a minimally invasive sample isolated from a subject to be diagnosed or screened, and comparing the presence and / or level of the autoantibodies to a predetermined threshold or cutoff value, where the predetermined threshold or cutoff value corresponds to the level of autoantibody expression that correlates with the highest specificity with a desired sensitivity in an ROC curve calculated at least in part based on the expression levels of the autoantibodies determined in a patient population at risk of developing colorectal cancer or colorectal adenoma, where the overexpression of at least one of the autoantibodies relative to the predetermined cutoff value indicates that the subject has colorectal cancer or colorectal adenoma with the desired sensitivity.

[0175] As another example, such a given condition may be that the specificity of the prediction of colonic cell proliferation disorder includes, for example, values ​​of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0176] As another example, such a given condition may be that the positive predictive value (PPV) for predicting colonic cell proliferation disorder includes, for example, values ​​of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0177] As another example, such a given condition may be that the negative predictive value (NPV) for predicting colonic cell proliferation disorder includes, for example, values ​​of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0178] As another example, such a given condition may include an area under the curve (AUC) of the receiver operating characteristic (ROC) curve predicting colonic cell proliferation impairment, which includes values ​​of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.

[0179] Colorectal cancer monitoring After processing the dataset using a trained algorithm, colorectal cancer may be identified or monitored in the subjects. Identification may be at least partially based on quantitative measurement of sequence reads of the dataset in a panel of colorectal cancer-related autoantibodies.

[0180] In some embodiments, the methods disclosed herein may be applied to monitor and / or predict tumor load.

[0181] In some embodiments, the methods disclosed herein may be applied to detect and / or predict residual tumors after surgery.

[0182] In some embodiments, the methods disclosed herein may be applied to detect and / or predict minimal residual disease after treatment.

[0183] In some embodiments, the methods disclosed herein may be applied to detect and / or predict recurrence.

[0184] In one embodiment, the method disclosed herein may be applied as a secondary screening.

[0185] In one embodiment, the method disclosed herein may be applied as a primary screening.

[0186] In one embodiment, the methods disclosed herein may be applied to monitor the development of cancer.

[0187] In one embodiment, the methods disclosed herein may be applied to monitor and / or predict the risk of cancer.

[0188] Colorectal cancer can be identified in subjects with an accuracy of at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 81%, at least approximately 82%, at least approximately 83%, at least approximately 84%, at least approximately 85%, at least approximately 86%, at least approximately 87%, at least approximately 88%, at least approximately 89%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, at least approximately 99%, or higher. The accuracy of colorectal cancer identification by a trained algorithm may be calculated as the percentage of independent test samples that are accurately identified or classified as having or not having colorectal cancer (e.g., subjects known to have colorectal cancer, or subjects with negative clinical test results for colorectal cancer).

[0189] Colorectal cancer may also be identified in subjects with a positive predictive value (PPV) of at least approximately 5%, at least approximately 10%, at least approximately 15%, at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 81%, at least approximately 82%, at least approximately 83%, at least approximately 84%, at least approximately 85%, at least approximately 86%, at least approximately 87%, at least approximately 88%, at least approximately 89%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, at least approximately 99%, or higher. The PPV for colorectal cancer identification using a trained algorithm may be calculated as the proportion of cell-free biological samples identified or classified as having colorectal cancer that actually have colorectal cancer.

[0190] Colorectal cancer may also be identified in subjects with a negative predictive value (NPV) of at least approximately 5%, at least approximately 10%, at least approximately 15%, at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 81%, at least approximately 82%, at least approximately 83%, at least approximately 84%, at least approximately 85%, at least approximately 86%, at least approximately 87%, at least approximately 88%, at least approximately 89%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, at least approximately 99%, or higher. The NPV for colorectal cancer identification using a trained algorithm may be calculated as the proportion of cell-free biological samples identified or classified as not having colorectal cancer that actually do not have colorectal cancer.

[0191] Colorectal cancer occurs in at least approximately 5%, at least approximately 10%, at least approximately 15%, at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 81%, at least approximately 82%, at least approximately 83%, at least approximately 84%, at least approximately 85%, at least approximately 86%, at least approximately 87%, at least approximately 88%, at least approximately 89%, at least approximately 90%, and a small percentage. It may be identified in subjects with a clinical sensitivity of at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, at least approximately 99%, at least approximately 99.1%, at least approximately 99.2%, at least approximately 99.3%, at least approximately 99.4%, at least approximately 99.5%, at least approximately 99.6%, at least approximately 99.7%, at least approximately 99.8%, at least approximately 99.9%, at least approximately 99.99%, at least approximately 99.999%, or higher. Using a trained algorithm, the clinical sensitivity for identifying colorectal cancer may be calculated as the proportion of independent test samples associated with the presence of colorectal cancer (e.g., subjects known to have colorectal cancer) that are accurately identified or classified as having colorectal cancer.

[0192] Colorectal cancer occurs in at least approximately 5%, at least approximately 10%, at least approximately 15%, at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 81%, at least approximately 82%, at least approximately 83%, at least approximately 84%, at least approximately 85%, at least approximately 86%, at least approximately 87%, at least approximately 88%, at least approximately 89%, at least approximately 90%, and a small percentage. It may be identified in subjects with a clinical sensitivity of at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, at least approximately 99%, at least approximately 99.1%, at least approximately 99.2%, at least approximately 99.3%, at least approximately 99.4%, at least approximately 99.5%, at least approximately 99.6%, at least approximately 99.7%, at least approximately 99.8%, at least approximately 99.9%, at least approximately 99.99%, at least approximately 99.999%, or at least 99.999% or higher. Using a trained algorithm, the clinical sensitivity for identifying colorectal cancer may be calculated as the proportion of independent test samples associated with the absence of colorectal cancer (e.g., subjects with negative results in clinical trials for colorectal cancer) that are accurately identified or classified as not having colorectal cancer.

[0193] In some embodiments, a trained algorithm may determine that a subject is at risk of colorectal cancer of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more.

[0194] The trained algorithms are at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 81%, at least approximately 82%, at least approximately 83%, at least approximately 84%, at least approximately 85%, at least approximately 86%, at least approximately 87%, at least approximately 88%, at least approximately 89%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, and at least approximately 94%. It may be possible to determine that a subject is at risk for colorectal cancer with an accuracy of at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, at least approximately 99%, at least approximately 99.1%, at least approximately 99.2%, at least approximately 99.3%, at least approximately 99.4%, at least approximately 99.5%, at least approximately 99.6%, at least approximately 99.7%, at least approximately 99.8%, at least approximately 99.9%, at least approximately 99.99%, at least approximately 99.999%, or higher.

[0195] If a subject is identified as having colorectal cancer, they may be offered a therapeutic intervention (e.g., prescription of an appropriate treatment plan to address their colorectal cancer) on an optional basis. Therapeutic interventions may include prescription of an effective dose of medication, further examination or evaluation of colorectal cancer, further monitoring of colorectal cancer, or a combination thereof. If the subject is currently receiving treatment for colorectal cancer under a treatment plan, the therapeutic intervention may include a subsequent different treatment plan (e.g., to enhance the effectiveness of treatment resulting from the ineffectiveness of the current treatment plan).

[0196] Therapeutic interventions may include recommending secondary clinical tests to confirm the diagnosis of colorectal cancer. These secondary clinical tests may include imaging tests, blood tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free cytology, FIT tests, FOBT tests, or combinations thereof.

[0197] Target colorectal cancers may be monitored by monitoring the treatment strategy for target colorectal cancers. Monitoring may include evaluating target colorectal cancers at two or more time points. Evaluation may be based on quantitative measurements of autoantibodies in a dataset in a panel of colorectal cancer-related autoantibodies, including quantitative measurements of the panel of colorectal cancer-related autoantibodies determined at at least two or more time points.

[0198] In some embodiments, the difference in quantitative measurements of sequence reads of a dataset containing quantitative measurements of a panel of colorectal cancer-related autoantibodies, determined between two or more time points, may indicate one or more clinical indicators, such as (i) the diagnosis of colorectal cancer in question, (ii) the prognosis of colorectal cancer in question, (iii) an increased risk of colorectal cancer in question, (iv) a decreased risk of colorectal cancer in question, (v) the effectiveness of a treatment strategy for treating colorectal cancer in question, and (vi) the ineffectiveness of a treatment strategy for treating colorectal cancer in question.

[0199] In some embodiments, a difference in quantitative autoantibody measurements, including quantitative measurements of a panel of colorectal cancer-related autoantibodies determined between two or more time points, may indicate a diagnosis of colorectal cancer in the subject. For example, if colorectal cancer was not detected in the subject at an earlier time point but was detected at a later time point, the difference indicates a diagnosis of colorectal cancer in the subject. Based at least in part on this indication of a diagnosis of colorectal cancer in the subject, clinical actions or decisions may be made, for example, by prescribing a new therapeutic intervention to the subject. Clinical actions or decisions may include recommending the subject for secondary clinical examinations to confirm the diagnosis of colorectal cancer. These secondary clinical examinations may include imaging studies, blood tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free cytology, FIT tests, FOBT tests, or combinations thereof.

[0200] In some embodiments, the difference in quantitative autoantibody measurements in a dataset containing quantitative measurements of a panel of colorectal cancer-related autoantibodies determined between two or more time points may indicate the prognosis of the colorectal cancer in question.

[0201] In some embodiments, a difference in quantitative autoantibody measurements in a dataset, including quantitative measurements of a panel of colorectal cancer-related autoantibodies determined between two or more time points, may indicate that a subject has an increased risk of colorectal cancer. For example, if colorectal cancer is detected in the subject at both an earlier and a later time point, and the difference is positive (e.g., the quantitative autoantibody measurements in the dataset of the panel of colorectal cancer-related autoantibodies increased from the earlier to the later time point), the difference may indicate that the subject has an increased risk of colorectal cancer. Based at least in part on this indication of increased risk of colorectal cancer, clinical actions or decisions may be made, such as prescribing a new therapeutic intervention for the subject or switching therapeutic interventions (e.g., discontinuing the current treatment and prescribing a new treatment). Clinical actions or decisions may include recommending the subject for secondary clinical examination to confirm the increased risk of colorectal cancer. These secondary clinical tests may include imaging tests, blood tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free cytology, FIT tests, FOBT tests, or combinations thereof.

[0202] In some embodiments, a difference in quantitative measurements of sequence reads from a panel of colorectal cancer-related autoantibodies, including quantitative measurements of colorectal cancer-related autoantibodies determined between two or more time points, may indicate that a subject has a reduced risk of colorectal cancer. For example, if colorectal cancer is detected in a subject at both an earlier and a later time point, and the difference is negative (e.g., the quantitative measurement of autoantibodies in a panel of colorectal cancer-related autoantibodies, including quantitative measurements of colorectal cancer-related autoantibodies, decreased from the earlier to the later time point), the difference may indicate that the subject has a reduced risk of colorectal cancer. Clinical actions or decisions (e.g., continuation or termination of current therapeutic interventions) may be made based at least in part on this indication of a subject's reduced risk of colorectal cancer. Clinical actions or decisions may include recommending secondary clinical testing for the subject to confirm the reduced risk of colorectal cancer. These secondary clinical tests may include imaging tests, blood tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free cytology, FIT tests, FOBT tests, or combinations thereof.

[0203] In some embodiments, a difference in quantitative measurements of sequence reads from a panel dataset of colorectal cancer-associated autoantibodies, including quantitative measurements of colorectal cancer-associated autoantibodies determined between two or more time points, may indicate the effectiveness of a treatment strategy for treating the subject's colorectal cancer. For example, if colorectal cancer was detected in the subject at an earlier time point but not at a later time point, the difference may indicate the effectiveness of a treatment strategy for treating the subject's colorectal cancer. Based at least in part on this indication of the effectiveness of a treatment strategy for treating the subject's colorectal cancer, clinical actions or decisions may be made, for example, to continue or terminate the current therapeutic intervention for the subject. Clinical actions or decisions may include recommending the subject for secondary clinical testing to confirm the effectiveness of a treatment strategy for treating colorectal cancer. These secondary clinical tests may include imaging tests, blood tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free cytology, FIT tests, FOBT tests, or combinations thereof.

[0204] In some embodiments, a difference in quantitative measurements of autoantibodies in a panel dataset of colorectal cancer-related autoantibodies, including quantitative measurements of colorectal cancer-related autoantibodies determined between two or more time points, may indicate the ineffectiveness of a treatment strategy for treating the colorectal cancer in the subject. For example, if colorectal cancer is detected in the subject at both an earlier and a later time point, and the difference is positive or zero (e.g., quantitative measurements of autoantibodies in a panel dataset of colorectal cancer-related autoantibodies, including quantitative measurements of colorectal cancer-related autoantibodies, increased from the earlier to the later time point or remained at a constant level), and an effective treatment was shown at the earlier time point, the difference may indicate the ineffectiveness of a treatment strategy for treating the colorectal cancer in the subject. Based at least in part on this indication of ineffectiveness of a treatment strategy for treating the colorectal cancer in the subject, clinical action or decision may be to discontinue the current therapeutic intervention and / or switch to a different new therapeutic intervention for the subject (e.g., prescribe). Clinical actions or decisions may include recommending a patient for secondary clinical testing to confirm the ineffectiveness of the treatment plan for colorectal cancer. This secondary clinical testing may include imaging studies, blood tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free cytology, FIT tests, FOBT tests, or combinations thereof.

[0205] kit This disclosure provides a kit for identifying or monitoring a target cancer. The kit may include probes or primers for identifying a quantitative measure (e.g., presence, absence, or relative amount) of a sequence in each of several cancer-associated autoantibodies in a target cell-free biological sample. A quantitative measure (e.g., presence, absence, or relative amount) of a panel of autoantibodies in a cell-free biological sample may indicate one or more cancers. The probes may be selective for autoantibodies in the cell-free biological sample. The kit may include instructions for processing the cell-free biological sample with the probes to generate a dataset showing a quantitative measure (e.g., presence, absence, or relative amount) of autoantibodies in the target cell-free biological sample.

[0206] The probes in the kit may be selective for sequences in multiple cancer-associated autoantibodies in a cell-free biological sample. The probes in the kit may be configured to selectively enrich autoantibody molecules corresponding to multiple cancer-associated autoantibodies. The probes in the kit may be proteins that are recognized by autoantibodies and tagged to allow separation after binding to the autoantibodies in the biological sample.

[0207] Kit instructions may include instructions for assaying cell-free biological samples using probes that are selective for cancer-associated autoantibodies in the cell-free biological sample. Quantitative measures (e.g., presence, absence, or relative amount) of sequences in each of multiple cancer-associated autoantibodies in the cell-free biological sample may indicate one or more cancers.

[0208] The kit instructions may include instructions for measuring and interpreting assay readouts that can be quantified by one or more of the multiple cancer-associated autoantibodies, in order to generate a dataset showing quantitative measures (e.g., indicating presence, absence, or relative amount) of the sequences of each of the multiple cancer-associated autoantibodies in a cell-free biological sample. [Examples]

[0209] Example 1: Analysis of autoantibodies in patient plasma samples.

[0210] In cancer, autoantibodies against novel cancer antigens or canonical protein antigens represent a source of biomarkers for the early diagnosis of colorectal cancer. Autoantibodies are produced in response to protein overexpression or mutations in cancer patients. Several autoantibodies have been identified as being associated with breast cancer, prostate cancer, colorectal cancer, lung cancer, ovarian cancer, and others.

[0211] To identify autoantibodies useful for the methods and classifiers described herein, plasma from patients with colonic cell proliferative disorder (CLL) and plasma from subjects without CLL (control plasma or reference plasma) were examined to identify a signature panel of autoantibodies produced by patients with CLL in response to CLL and their respective reactive proteins. Therefore, plasma from patients with CLL and control plasma were examined using high-density protein microarrays. Protein microarrays offer a series of advantages compared to other approaches used to identify autoantibodies: i) the proteins printed on the array are pre-known, preventing subsequent identification and eliminating the possibility of mimotope selection; and ii) all proteins are printed at similar concentrations, eliminating any tendency to select specific proteins. This combination of factors allows for highly sensitive identification of biomarkers.

[0212] The identified antibody panel made it possible to differentiate between plasma from patients with colonic cell proliferation disorder and plasma from healthy patients.

[0213] method Sample classification To detect autoantibodies in plasma samples, high-density protein microarrays expressing thousands of candidate tumor antigens were then probed in plasma collected from subjects identified as having colorectal cancer (CRC), advanced adenoma (AA), benign polyp (NAA), or none of the above (NEG). Binding immunoglobulins were measured by the intensity of a fluorescently labeled secondary antibody (anti-IgG / IgM).

[0214] Plasma samples were obtained from a general population of controls matched for age, sex, and location, using a standardized serum collection protocol, and stored at -80°C until use. Individuals with a personal history of cancer were excluded from the control group. Written informed consent was obtained from all participants with the approval of the Institutional Review Board (IRB).

[0215] A description of the study cohort is provided in Table 1, which shows a classification model of the number (stage, sex, and age) of healthy and cancerous samples used in the CRC experiment.

[0216] [Table 1]

[0217] The primary objective of this study was to identify serum TAAb biomarkers that differentiate colorectal cancer from advanced adenoma, benign disease, and healthy controls, thereby improving the sensitivity of existing biomarkers and guiding clinical decision-making. We implemented a sequential screening strategy to identify a panel of TAAb biomarkers from over 21,000 (replica-measured) human proteins.

[0218] Plasma was isolated from NEG, CRC, AA, and NAA control populations and screened using protein arrays. A total of 42,390 features were identified among the NEG, CRC, AA, and NAA control populations, and differential expression between plasma from subjects with colonic cell proliferation disorder and plasma from healthy subjects was investigated.

[0219] Image analysis and quantification were performed using standard methodologies pre-described for the protein array platform. In short, slides were scanned with a 2-channel microarray scanner, and the intensity of the foreground (spot region) and background (periphery of the spot) was measured. The raw intensity values ​​were normalized using the following procedure. 1) Remove the background by subtracting the median background intensity from the foreground intensity of all spots on the array (background correction intensity). 2) Using the foreground intensity of the negative control spot, estimate the parameters (mean, variance) of the foreground / background normal + exponential convolution model (assuming the raw value represents the sum of background and foreground contributions). 3) Subtract the average intensity and coefficient of variation (variance of the control spot divided by the average of the protein foreground) of the background reinforcement. 4) Report the background-corrected protein intensity as the maximum likelihood estimate of the convolutional model.

[0220] Filtering of raw feature values ​​is performed in advance.

[0221] The raw feature values ​​for both IgG and IgM channels were concatenated into a single feature matrix for all cohort samples. This includes a total of 42,390 features across 941 samples (including unclassifiable samples).

[0222] After preprocessing (background correction, IQR median normalization, outlier trimming, and batch normalization), the feature space was narrowed down to include only those samples with a raw foreground intensity greater than 2000 (range: 0–64000 rfu) in 10 or more samples. 16,570 proteins / antigens met this criterion.

[0223] For each representation (CRC, AA, NAA, vs. NEG), stratified cross-validation was performed (within each layer) with 4 layers using 5 random seeds, with the following feature selection criteria. A) The normalized values ​​were binarized based on whether they were at least two standard deviations from the mean of the features. B) If the binarized chi2 p-value is less than 0.01, the feature was preserved (binarization is used only for chi2 comparisons). C) The retained features were subjected to recursive feature removal using logistic regression weights (logreg weights), and the top 100 features were used for each cross-validation (CV) classification.

[0224] Features were ranked by the number of times each layer / seed was selected (up to 20 times), and by the mean and sum of the weights for all logistic regressions.

[0225] result CRC vs. NEG A total of 28 proteins (supplementary) were selected from over 50% of the entire layer. Table 2 shows the top 5 AAb targets of the 28 proteins according to the CRC classification.

[0226] [Table 2]

[0227] Figure 2 provides a graph showing the CV coefficients of the top 5 AAb targets selected for potential development for CRC classification.

[0228] Figure 3 provides a graph showing the performance of recursive feature removal for CRC classification using cross-validation (CV).

[0229] AA vs NEG A total of 23 proteins (supplementary) were selected from over 50% of the entire layer. Table 3 shows the top 5 AAb targets of the 23 proteins according to their AA classification.

[0230] [Table 3]

[0231] Figure 4 provides a graph showing the CV coefficients of the top 10 AAb targets for the AA classification.

[0232] Figure 5 provides a graph showing the performance of recursive feature removal for AA classification using cross-validation.

[0233] NAA vs. NEG Thirteen targets met the selection criteria in more than 50% of all layers. Table 4 shows the top five AAb targets of the 13 proteins according to the NAA classification.

[0234] [Table 4]

[0235] Figure 6 provides a graph showing the CV coefficients of the top 5 AAb targets for the NAA classification.

[0236] Figure 7 provides a graph showing the performance of recursive feature removal for NAA classification using cross-validation.

[0237] In both cases, the results provide a list of AAb biomarkers for the classification of CRC, AA, and NAA, as shown in Table 5.

[0238] [Table 5]

[0239] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only as examples. The present invention is not intended to be limited by any specific examples provided herein. Although the present invention is described in relation to the foregoing specification, the descriptions and examples of embodiments herein are not intended to be constrained. Those skilled in the art will be able to conceive of many modifications, variations, and substitutions without departing from the present invention. Furthermore, it will be understood that all aspects of the present invention are not limited to any specific descriptions, configurations, or relative proportions described herein, depending on various conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be used in carrying out the invention of this disclosure. Therefore, the present invention is intended to extend to any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the present invention, and it is intended that methods and structures within the scope of these claims and their equivalents are encompassed thereby.

Claims

1. A method for detecting a target colorectal cancer, wherein the method is: a) A step of measuring the amount of autoantibodies from a predetermined autoantibody panel obtained from a blood sample from the subject in order to obtain the autoantibody profile of the subject, wherein the predetermined autoantibody panel is (i) IgM autoantibodies against NME5 and at least two antigens selected from the group consisting of USP16, UBE2S, RNF41, CD20, and SDCPP, (ii) IgM autoantibodies against NME5 and at least two antigens selected from the group consisting of PELO, CDK4, MTP1, PRMT6, ZBTB2, MTCP1, and PCOLCE. (iii) IgG autoantibodies against NME5 and at least two antigens selected from the group consisting of ANKHD1, TXNL1, NAT6, Sup6h, PRDM8, OTUD5, PNKP, SRSF7, PCOLCE, and ASB9, or (iv) IgG autoantibodies against NME5 and at least two antigens selected from the group consisting of TSSC4, TP53, BRD9, and BCCIP, or (v) A process including combinations of (i) to (iv), and b) Processing the autoantibody profile using a machine learning model that is trained to indicate the presence of colorectal cancer in a subject, thereby enabling differentiation between subjects without colorectal cancer and subjects with colorectal cancer, and providing output values ​​related to the presence of colorectal cancer in the subject. Methods that include...

2. The method according to claim 1, wherein the autoantibody profile is associated with colorectal cancer and provides classification of the subject as having colorectal cancer.

3. The method according to claim 1 or 2, wherein the blood sample is selected from the group consisting of plasma, serum, whole blood, isolated blood cells, cells separated from blood, and combinations thereof.

4. The method according to claim 1 or 2, wherein the colorectal cancer is selected from the group consisting of Lynch syndrome, colon cancer, rectal cancer, colonic cell tumor, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.

5. The method according to any one of claims 1 to 4, further comprising the step of detecting the methylation status of one or more nucleic acid molecules in the blood sample in order to provide the methylation profile of the target.

6. The method of claim 5, further comprising the step of processing the methylation profile using the machine learning model.

7. The method according to any one of claims 1 to 6, further comprising the step of measuring the amount of one or more proteins in the blood sample in order to provide the target protein profile.

8. The method according to claim 7, further comprising the step of processing the protein profile using the machine learning model.

9. The method according to claim 1, wherein the machine learning model includes a classifier selected from the group consisting of deep learning classifiers, neural network classifiers, linear discriminant analysis (LDA) classifiers, quadratic discriminant analysis (QDA) classifiers, support vector machine (SVM) classifiers, random forest (RF) classifiers, K nearest neighbor classifiers, linear kernel support vector machine classifiers, first- or second-order polynomial kernel support vector machine classifiers, ridge regression classifiers, elastic net algorithm classifiers, sequential minimal problem optimization algorithm classifiers, naive Bayes algorithm classifiers, and principal component analysis classifiers.

10. The method according to claim 1, wherein the predetermined autoantibody panel further comprises (i) IgM autoantibodies against at least three antigens selected from the group consisting of NME5, USP16, UBE2S, RNF41, CD20, and SDCPP.

11. The method according to claim 1, wherein the predetermined autoantibody panel further comprises IgG autoantibodies against at least three antigens selected from the group consisting of (iii) ANKHD1, TXNL1, NAT6, Supt6h, PRDM8, OTUD5, PNKP, SRSF7, PCOLCE, and ASB9.

12. The method according to claim 1, wherein the predetermined autoantibody panel further comprises (ii) IgM autoantibodies against at least three antigens selected from the group consisting of PELO, CDK4, MTP1, PRMT6, ZBTB2, MTCP1, and PCOLCE.

13. The method according to claim 1, wherein the predetermined autoantibody panel further comprises IgG autoantibodies against at least three antigens selected from the group consisting of (iv)TSSC4, TP53, BRD9, and BCCIP.

14. The method according to claim 13, wherein the predetermined autoantibody panel further comprises an IgG autoantibody against TP53.

15. The method according to claim 3, wherein the blood sample is plasma.

Citation Information

Patent Citations

  • Methods for the diagnosis or prognosis of colorectal cancer

    EP2444811A1

  • Serological autoantibodies as biomarker for colorectal cancer

    EP3193173A1

  • A method for diagnosing or determining the prognosis of colorectal cancer (crc) using novel autoantigens: gene expression guided autoantigen discovery

    GB2494741A

  • Cancer diagnostic agents

    JP2012521552A

  • cancer biomarkers

    JP2013511728A