Cancer diagnostic method based on quantitative biomarkers and database thereof
By using QDB and mass spectrometry to determine the expression levels of biomarkers in FFPE samples, an individual cancer profile database was constructed, which solved the problem of insufficient utilization of FFPE samples in existing technologies and enabled accurate cancer diagnosis and prognosis as well as personalized treatment.
Patent Information
- Application Number
- CN202080076711.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-11
- Filing Date
- 2020-11-02
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2040-11-02
AI Technical Summary
Existing technologies cannot effectively utilize biomarker data in stored formalin-fixed-paraffin-embedded (FFPE) samples, resulting in insufficient accuracy in clinical diagnosis and prognosis. In particular, the subjectivity and inherent inconsistencies of IHC methods make it difficult to distinguish individual differences at the population level.
By using quantitative dot immunoblotting (QDB) to determine the expression levels of biomarkers in FFPE samples as absolute and continuous variables, combined with mass spectrometry and enzyme-linked immunosorbent assay (ELISA), a retrospective cancer database was constructed. A unique 'fingerprint' for each sample was established using multiple absolutely quantitative protein biomarker combinations. Individual cancer profile information (ICP) was constructed by combining clinical records, and sample points were located in three-dimensional space for personalized diagnosis and prognosis.
It enables efficient utilization of FFPE samples, improves the accuracy of cancer diagnosis and prognosis, distinguishes individual samples from a large number of stored samples, provides personalized treatment plans, and enhances the ability to share and merge data.
Smart Images

Figure BDA0003625092680000181 
Figure BDA0003625092680000191 
Figure BDA0003625092680000192
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims all benefits based on U.S. Provisional Patent Application No. 62 / 929,396, filed November 1, 2019, the entire contents of which are hereby incorporated by reference. Technical Field
[0003] This invention relates to methods, systems, and software for diagnosing, predicting, and prognosing cancer patients based on the quantitative levels of a set of biomarkers. Specifically, this invention relates to diagnosing, predicting, and prognosing cancer patients by referring to the quantitative levels of the same set of biomarkers in a database. Background Technology
[0004] For most cancer patients, their tumor tissue, after surgical removal, is preserved in hospitals or other medical facilities in the form of formalin-fixed paraffin-embedded (FFPE). Thus, the millions of stored FFPE samples, along with associated detailed case records including treatment protocols and clinical outcomes, accumulate as a vast and underutilized medical resource. The sheer number of these FFPE samples is sufficient to cover diverse molecular identities at the individual level. Combined with their known clinical outcomes, these samples become an unparalleled resource for research into personalized clinical treatment. Here, we can identify known cases with similar molecular characteristics for any cancer patient in the world.
[0005] In clinical practice, immunohistochemical analysis (IHC) is widely used to assess biomarkers at the protein level for clinical diagnostic and prognostic purposes. Typical IHC reports for biomarkers are expressed as "+" or "-", or further categorized as "0, 1+, 2+, 3+". For example, analyzing the expression level of human epidermal growth factor receptor 2 (Her2), a biomarker commonly used in breast cancer diagnosis, using IHC methods can determine whether Her2-dependent therapy should be included in treatment. IHC results are divided into three groups: 0 and 1+, 2+ and 3+. Results in the 0 and 1+ groups are considered negative, the 3+ group is considered positive, and the 2+ group is considered ambiguous.
[0006] Combining IHC results for diagnosis and prognosis is a routine procedure in clinical practice. For example, in breast cancer patients, four tumor biomarkers, including estrogen receptor (ER), progesterone receptor (PR), Ki67, and Her2, are used to classify patients into four subtypes: luminal A, luminal B, Her2, and triple-negative. IHC results for Her2 and ER / PR are used to classify patients into luminal, Her2, and triple-negative types, while Ki67 expression levels further subdivide the luminal subtype into luminal A and luminal B.
[0007] IHC classification results are also difficult to use in further clinical practice. For example, although there are significant differences among positive individual patients, they are often grouped into the same category in clinical practice. Therefore, results from IHC analysis are difficult to use for adequate data analysis to provide more accurate and predictive diagnoses and prognoses.
[0008] IHC methods are also severely limited by inherent subjectivity and inconsistency. The heterogeneity of tumor tissue further complicates the diagnostic process.
[0009] Many scientists have dedicated themselves to measuring biomarkers as absolute and continuous variables at the tissue level. For example, enzyme-linked immunosorbent assay (ELISA) can be used to measure the expression levels of biomarkers in fresh and frozen tissues. However, this method cannot be used to measure the expression levels of biomarkers in FFPE samples, thus severely limiting its application in clinical diagnosis and prognosis.
[0010] Quantitative dot blot (QDB) assays can measure biomarker expression levels in fresh, frozen, and FFPE samples as absolute and continuous variables in a high-throughput manner. When standard proteins are introduced, whether in recombinant or purified form, this method can be easily converted into an absolute quantitative method for the absolute quantification of a specific protein at the cellular or tissue level.
[0011] Prior to the application of the QDB method, millions of stored FFPE samples could not be processed using existing protein technologies because these technologies could not effectively distinguish individual FFPE samples at the population level. Currently used protein analysis methods, including immunohistochemistry (IHC), Western blot analysis, reversed-phase protein array analysis (RPPA), and mass spectrometry, have been used to analyze FFPE samples. However, these methods are insufficient for analyzing the vast number of FFPE samples.
[0012] For IHC and Western imprint analysis, their qualitative intrinsic characteristics can obscure inter-individual differences at the population level.
[0013] Other methods can quantify protein expression levels to reveal individual differences at the population level, but they provide relative results, thus limiting the scale of studies. We can better illustrate this limitation with an example. Protein expression levels can be expressed as absolute values (e.g., nmole / g) or as relative values (percentage of reference protein B). While it is easy to compare protein levels expressed in absolute form across multiple analyses (in this article, "a plurality of" specifically refers to "more than two"), it is difficult to compare results based on different reference protein B levels across multiple analyses.
[0014] MS and RPPA both fall into this category. Their results are expressed relative to a reference protein, the content of which may vary in each study [Boellner, et al, Microarrays, 4(2):98-114, 2015, DeSouza, et al, Clin. Biochem. 46:421-431, 2013]. Therefore, the scale of studies based on these methods is limited by the number of samples in each study, and cannot be expanded by merging results from other studies. Real-time quantitative PCR (RT-PCR) datasets also suffer from the same problem.
[0015] On the other hand, the QDB method offers a way to handle large datasets of FFPE samples that possess both absoluteness and continuity. Regarding continuity, quantified results are needed to reflect subtle differences between individuals at the population level. Regarding absoluteness, the quantification of each protein should be consistent regardless of changes in time, location, etc., thereby ensuring that data can be shared, cross-validated, and merged to meet the urgent need for dataset growth to handle the massive number of FFPE samples.
[0016] This invention provides methods, systems, and software to assist in patient diagnosis by utilizing the world's widely stored FFPE samples. In clinical practice, this method can significantly improve treatment effectiveness, enabling personalized care. Summary of the Invention
[0017] This invention provides a method for diagnosing, predicting, and prognosing cancer using three or more biomarkers as continuous variables. The evaluation of biomarkers in this invention is based on quantitative rather than current, universally accepted qualitative measurements, and is then expressed as absolute units that can be easily incorporated into existing databases.
[0018] The experimental sample can be tissue from an individual. In one embodiment of the invention, tissue refers to biopsy tissue. In another embodiment of the invention, tissue refers to a formalin-fixed and paraffin-embedded sample (FFPE sample).
[0019] The individual can be a patient. More specifically, the individual can be a cancer patient. In one embodiment of the invention, the individual can be a breast cancer patient.
[0020] Absolutely quantifiable protein biomarker levels can be used to develop a retrospective cancer profile database (RC), or more precisely, a database of different cancer types (including breast cancer, colorectal cancer, or prostate cancer) to make full use of the massive stored FFPE samples (in this article, "cancer profile" is the English equivalent of "cancer profile," and is sometimes referred to as "cancer profile information" or simply "profile information" or "information" depending on the context).
[0021] By combining multiple absolutely quantitative protein biomarkers, it is possible to distinguish individual FFPE samples from millions of stored FFPE samples. In a sense, the combination of these protein biomarkers becomes a unique "fingerprint" for each FFPE sample in the database.
[0022] Using this unique "fingerprint" as the core, it can be combined with matching clinical records, including traditional clinicopathological parameters, treatment plans, and corresponding clinical outcomes, to provide comprehensive information for each FFPE sample.
[0023] All the information from these different aspects constitutes the Individual Cancer Profile (ICP) for each FFPE sample in the database. Any other clinically relevant traits can be included in these cancer profiles. For example, genetic information, including single nucleotide variants (SNVs), chromosomal translocations, and scores from various gene prediction analyses, can be included in the cancer profile.
[0024] The database's absolute characteristics ensure its continuous growth. Although the sources of ICPs differ, their absolute characteristics allow for effective integration. New cancer profiles can also be added and updated over time. In the future, the database is expected to house a significant number of stored FFPE sample profiles to support clinical diagnosis based on "big data."
[0025] Furthermore, the above method further includes a method for constructing a retrospective cancer (RC) database to provide cancer diagnosis, prediction, and prognosis, the method comprising: providing multiple individuals with cancers of known clinical outcomes; generating an ICP for each of the multiple individuals, wherein the ICP includes: i) multiple protein biomarkers determined by absolute quantification, and ii) the known clinical outcome of the cancer, and storing the generated ICPs of the multiple individuals in the database.
[0026] In one embodiment of the present invention, the expression level of a biomarker can be measured as an absolute and continuous variable.
[0027] In one embodiment of the present invention, the expression level of a biomarker can be determined by mass spectrometry.
[0028] In one embodiment of the present invention, the expression level of the biomarker can be determined by enzyme-linked immunosorbent assay (ELISA).
[0029] In one embodiment of the present invention, the expression level of the biomarker can be determined by quantitative dot immunoblotting (QDB) analysis.
[0030] In one embodiment of the present invention, the protein expression levels of three or more biomarkers can be determined by any combination of ELISA, QDB and mass spectrometry.
[0031] The quantitative expression levels of ICP biomarkers in the database can be combined with their associated clinical information for mathematical analysis used in medicine. For example, the potential relationship between the absolute levels of biomarkers and disease-free survival (DFS) can be explored, providing predictive clinical outcomes for patients.
[0032] In one embodiment of the present invention, the amounts of various biomarkers from ICP, expressed as continuous variables, can be combined with relevant clinical information (including but not limited to disease-free survival, overall survival (OS), side effects, age, and different stages of disease progression) to find potential causal relationships, and these causal relationships can be used for cancer diagnosis and prognosis purposes.
[0033] In one embodiment of the invention, the absolute amounts of three or more biomarkers from ICP can be used as (x, y, z) coordinates to locate an individual in a space determined by the X, Y, and Z axes (location point). The sample's location point is combined with relevant clinical information (including but not limited to disease-free survival, overall survival, side effects, age, and different stages of disease progression) to find spatial correlations, and these correlations can be used for diagnostic and prognostic purposes.
[0034] In one embodiment of the invention, more than one ICP location point in the space determined by the X, Y, and Z axes can be divided into clinical subgroups related to clinical diagnosis, prediction, and prognosis.
[0035] Another aspect of the invention relates to a reference database for cancer diagnosis based on quantitative analysis of more than one biomarker in a patient's biopsy sample. The reference database comprises multiple ICPs, each constructed through the following steps: (a) obtaining a biopsy sample with a known clinical diagnosis from a cancer patient; (b) determining the levels of three or more said biomarkers in the biopsy sample as absolute and continuous variables; (c) spatially locating each ICP (location point) using the three biomarker levels as x, y, z coordinate points; and (d) associating each ICP with the known clinical diagnosis, prediction, and prognosis of the cancer patient according to its location point, thereby obtaining a reference profile based on spatial location.
[0036] On the other hand, the present invention provides a method for diagnosing cancer in patients. The method includes the following steps: (i) providing the aforementioned reference space database; (ii) obtaining a biopsy sample from the patient; (iii) measuring three biomarkers in the biopsy sample, the results of which are expressed as absolute and continuous variables, the measurement results being continuous variables of the three biomarkers in the biopsy sample; (iv) locating the sample in the reference space database using the levels of the three biomarkers as (x, y, z) coordinates; and (v) identifying the reference space profile information that best matches the patient sample in the reference space database, and outputting a known clinical diagnosis associated with the identified reference profile information.
[0037] Another aspect of the invention relates to a database for providing diagnosis, prediction, or prognosis of cancer, the database containing multiple ICPs generated by individuals with known cancer clinical outcomes. Furthermore, the ICPs include: i) multiple clinical parameters quantitatively measured from FFPE samples from stored individuals, and ii) known clinical outcomes of the cancer. Each of the multiple clinical parameters represents a quantitative measurement of a biomarker. Moreover, the quantitative measurement results are continuous and represent the absolute amount of the biomarker in the sample.
[0038] Another aspect of the present invention relates to a method for providing diagnosis, prediction, or prognosis of cancer in a patient. The method includes: 1) collecting an FFPE sample from the patient; 2) obtaining i) stored ICPs and ii) a set of clinical parameters (i.e., a set of clinical parameters) used in the database from the aforementioned database; 3) comparing the quantitative levels of the set of clinical parameters in each ICP in the database with the quantitative levels of the same set of clinical parameters in the patient's FFPE sample; 4) identifying the ICP in the database that best matches the patient through comparison; and 5) outputting the clinical outcome of the ICP identified in the database.
[0039] In the above method, the comparison aims to determine the maximum similarity between the set of clinical parameters in the ICP and the same set of clinical parameters measured from the patient's FFPE sample.
[0040] In one embodiment of the invention, similarity can be determined by comparing the absolute levels of a set of protein biomarkers. ICPs within a preset range of the same biomarker are identified by using the quantitative levels of the patient's biomarkers. An ICP whose level in each biomarker of a set of biomarkers is within the preset range of the corresponding biomarker in the same set of biomarkers in the patient is considered similar to that patient.
[0041] In one embodiment of the present invention, the preset range of each biomarker in a set of biomarkers may be the same.
[0042] In another embodiment of the invention, the preset ranges of different biomarkers in a set of protein biomarkers used to assess the similarity between ICP and patients can be different.
[0043] In one embodiment of the invention, the similarity between the ICP and the patient can be calculated based on the Euclidean distance between two sets of quantitative clinical parameters.
[0044] Multiple ICPs can be identified from a database based on similarity. Furthermore, personalized prognoses can be provided to patients through mathematical analysis of their clinical outcomes.
[0045] Multiple ICPs can be identified from the database based on similarity. Furthermore, mathematical analysis of their treatment plans and clinical outcomes can be used to provide patients with the best prognostic treatment option.
[0046] In another embodiment of the invention, other clinical manifestations (including conventional clinical parameters such as age, tumor size, tumor grade, and lymph node status) can be used to further improve the similarity between ICP and the patient.
[0047] Details of the invention are also reflected in the accompanying drawings and set forth in the description below. Other features, objects, and advantages of the invention will become apparent to those skilled in the art upon reading these drawings and description, as well as the appended claims.
[0048] Brief description of the attached figures
[0049] Figure 1A three-dimensional scatter plot is shown, constructed using the expression levels of PR, ER, and Her2 from 1049 breast cancer samples as coordinates. The expression levels of ER, PR, and Her2 were determined using the QDB method, and their values were used to construct the three-dimensional scatter plot using Origin software, with PR represented by the X-axis, ER by the Y-axis, and Her2 by the Z-axis. The distribution of localization points from each patient divides the space into different regions, including the hormone group (samples are completely distributed on the plane presented by the X and Y axes), the Her2 group (samples are clustered around the Z-axis), and the corner group (samples are piled up at the intersection of the X, Y, and Z axes). The corner group includes triple-negative and normal-like groups.
[0050] Figure 2 This paper presents a comparison of overall survival (OS) with the corresponding clinical subtypes for five presumptive patients identified using Kaplan-Meier survival analysis based on absolute levels of ER, PR, Her2, and Ki67. An IHC-based alternative analysis categorized profiles in the database into luminal A-like, luminal B-like, Her2-positive, and triple-negative (TNBC) subtypes, and their OS was used as a reference. Comparisons with the five presumptive patients' similar groups were performed using the Log-Rank test, with p < 0.05 considered statistically significant. (a) OS comparison between the similar groups and the TNBC subtype for patients #1388 and #1843; (b) OS comparison between the similar group and the Her2-positive subtype for patient #1445; (c) OS comparison between the similar group and the luminal A-like subtype for patient #1807; and (d) OS comparison between the similar group and the luminal B-like subtype for patient #1519*. Levels of biomarkers below twice the limit of quantitation (LOQ) are considered indistinguishable to increase the amount of profiling information available for analysis.
[0051] Figure 3The Kaplan-Meier survival analysis was used to compare the overall survival of a similar group of five hypothetical patients identified based on absolute levels of ER, PR, Her2, Ki67, and cyclin D1 with the overall survival of their respective clinical subtypes. An IHC-based alternative analysis categorized profile information from the database into luminal A-like, luminal B-like, Her2-positive, and triple-negative (TNBC) subtypes, and their overall survival was used as a reference. Comparisons with the similar groups of the five hypothetical patients were performed using the Log-Rank test, with p < 0.05 considered statistically significant. (a) Overall survival of the similar groups of patients #1388 and #1843 was compared with the TNBC subtype; (b) Overall survival of the similar group of patient #1445 was compared with the Her2-positive subtype; (c) Overall survival of the similar group of patient #1807 was compared with the luminal A-like subtype; and (d) Overall survival of the similar group of patient #1519* was compared with the luminal B-like subtype. Biomarkers within twice the limit of quantitation (LOQ) are considered indistinguishable in order to increase the amount of profiling information available for analysis.
[0052] Figure 4 This paper presents a comparison of overall survival using Kaplan-Meier survival analysis of profiles of five hypothetical patients identified based on absolute values of ER, PR, Her2, and Ki67 within a similar group who received different clinical treatments. Profiles with unclear treatment regimens were not included in the analysis. (a) Overall survival analysis of profiles of patient #1388 receiving chemotherapy (Chemo), endocrine therapy (ET), and combined chemotherapy and endocrine therapy (CET: also referred to as "C&E" or "C+E" in this specification and its accompanying figures, hereinafter the same) within a similar group; (b) Overall survival analysis of profiles of patient #1843 receiving chemotherapy (Chemo), endocrine therapy (ET), and combined chemotherapy and endocrine therapy (CET) within a similar group; (c) Overall survival analysis of profiles of patient #1807 receiving chemotherapy (Chemo) within a similar group; (d) Overall survival analysis of profiles of patient #1519* receiving chemotherapy (Chemo) and combined chemotherapy and endocrine therapy (CET) within a similar group. Biomarker levels within twice the limit of quantitation (LOQ) are considered equivalent to increase the amount of profiling information available for analysis. Detailed Implementation
[0053] Before describing the method of the present invention, it should be stated that the present invention may vary in specific applications and is not limited to the methods and apparatus described herein. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit these embodiments. The scope of protection of the present invention is limited only by the appended claims.
[0054] Unless otherwise stated, all technical terms used in this specification are understood by those skilled in the art.
[0055] This invention relates primarily to clinical diagnostic purposes. Therefore, in this specification, the terms "determining," "measuring," "assessing," and "analyzing" are used interchangeably, and encompass both quantitative and qualitative measurement methods. These terms can specify both quantitative and semi-quantitative quantities; thus, "determining," "analyzing," "measuring," and similar descriptions can be used interchangeably. When describing quantitative measurement, the phrase "determining the amount of analyte" or a similar description is used. When describing quantitative or semi-quantitative measurement, the phrase "determining the level of analyte" or "detecting the analyte" is used.
[0056] The “quantitative” analysis described in this specification typically provides information on the relative levels of the analyte in the sample to the reference (control), usually presented numerically, where a “0” value can be specified as the amount of analyte below the limit of detection (LOD).
[0057] The terms “subject,” “host,” “patient,” and “individual” may be used interchangeably in this instruction manual to refer to any mammalian (especially human) subject requiring diagnosis or treatment.
[0058] "Spatial" and "three-dimensional" (3D) can be used interchangeably to describe the location of a spot representing a sample in three-dimensional space, where the intensity of the spot represents the amount of a fourth biomarker as a continuous variable.
[0059] Quantitative determination of the protein expression level of a biomarker at the tissue level can be achieved by any method. The method should be considered in its broadest context, including any method where the expression level of the biomarker can be quantified as a continuous variable. These methods include, but are not limited to, mass spectrometry or immunoassay, and combinations of both.
[0060] The results obtained in this invention can be relative or absolute, expressed using standard proteins. The terms "relative" and "absolute" refer to both methods and should be used within their broadest context. A relative determination compares one substance to another, while an absolute determination uses standard units to measure a known level of a substance. Perhaps the most significant difference between these two methods lies in their applicability. Relative results are only meaningful under the same experimental conditions, while absolute results can be compared across many different analyses, even if the analyses were performed at completely different times and locations.
[0061] In this specification, the terms "sample," "patient sample," "specimen," and "biological sample" generally refer to a sample that can be used to measure specific molecules, preferably specific biomarker molecules associated with biological characteristics, such as biomarkers described below. Samples may include, but are not limited to, peripheral blood cells, CNS fluid, serum, plasma, oral swabs, urine, saliva, tears, pleural effusion, and the like. In this invention, "sample" generally refers to tissue.
[0062] The terms "marker" and "biomarker" should be defined in their broadest context. Used interchangeably here, "marker" and "biomarker" generally refer to a molecule (e.g., a peptide) that is different in a sample from one phenotype (e.g., a patient) compared to one from another phenotype (e.g., not diseased or with a different disease). The establishment of a biomarker is based on its different expression in two different phenotypes, that is, when its mean or median level in the first phenotype is statistically significantly different from its level in the second phenotype.
[0063] In this invention, a biomarker refers to a measurable molecule associated with an organism or disease state. It can be a well-established clinical diagnostic biomarker (e.g., a clinical biomarker for immunohistochemical analysis) or a newly discovered in vitro diagnostic biomarker.
[0064] In this specification, "reference" or "control" is used interchangeably to refer to known data or sets of known data that can be used to compare data with observed data. Known data represents a known relationship between two parameters, such as the relationship between the expression level of a biomarker and its associated phenotype. Here, known data constitutes reference profile information in the reference database.
[0065] Accordingly, the reference database can store multiple reference profiles for diagnostic purposes, each of which includes the level of a biomarker in a sample obtained from an individual with a known diagnosis or known clinical efficacy after treatment.
[0066] In one embodiment, the present invention relates to a method for constructing a RC database to provide diagnosis, prediction, and prognosis for cancer patients, the method comprising: providing multiple individuals with known clinical outcomes of cancer; constructing an ICP for each of the multiple individuals, the ICP including i) multiple absolutely quantitative protein markers and ii) known clinical outcomes of cancer; and storing the generated ICPs for the multiple individuals in a database.
[0067] Each ICP includes not only multiple protein biomarkers, but also other clinical manifestations (including age, tumor size, tumor grade, and lymph node status). Other results from clinical analysis (including levels of blood biomarkers and various enzymes) can also be included in the ICP.
[0068] The present invention also relates to a method for identifying one or more reference ICPs that best match a patient's profile information from a retrospective cancer database. The method includes: (a) comparing the expression levels of a set of protein markers of the patient with the expression levels of each ICP in the database on a suitable programmable computer; (b) identifying ICPs that are highly similar to the patient on a suitable programmable computer; and (c) outputting the maximum similarity or related phenotype of the ICPs in the reference database that best match the patient's profile information to a user interface device, a computer-readable storage medium, or a regional or remotely accessible computer system, or displaying it directly.
[0069] There are several methods for comparing the similarity between patients and ICPs in a database, including: assessment based on mathematical analysis of pre-defined protein biomarkers, or stepwise screening of ICPs by the expression levels of a set of biomarkers.
[0070] The expression levels of biomarkers can be corrected during the process of comparing the similarity between ICP and patients through mathematical analysis, or they can be left uncorrected.
[0071] The expression levels of biomarkers can be weighted or unweighted when comparing the similarity between ICP and patients through mathematical analysis.
[0072] In one implementation, similarity is achieved by calculating the Euclidean distance between the ICP and the patient based on a set of biomarkers.
[0073] ICPs similar to the patient can also be identified through stepwise screening based on a set of biomarker expression levels. This method includes: a) screening all ICPs whose biomarker a is within a predefined expression range of the patient's biomarker a; b) further screening from the screened ICPs for all ICPs whose biomarker b is within a predefined expression range of the patient's biomarker b; c) among the further screened ICPs, further screening for all ICPs whose biomarker c is within a predefined expression range of the patient's biomarker c; ... n) among the further screened ICPs, further screening for all ICPs whose biomarker n is within a predefined expression range of the patient's biomarker n.
[0074] Among the pre-defined biomarkers, the preset ranges for each biomarker can be the same or different.
[0075] In one embodiment of the present invention, the ICPs selected through the above process can be further screened based on other clinical manifestations, including age, sex, tumor size, and tumor grade. For example, for a 59-year-old male lung cancer patient with a tumor size of grade 2, a tumor grade of 3, and a lymph node status of N2, similar ICPs based on protein markers can be further narrowed down to ICPs of similar age (55 to 60 years), male, with a tumor size of grade 2, a tumor grade of 3, and a lymph node status of N2 to improve the accuracy of clinical diagnosis of the patient.
[0076] A spatial relationship-based approach may include: (a) measuring a sample with three or more biomarkers as continuous variables; (b) using the values of the three biomarkers (A, B, C) as coordinates (x, y, z) to locate a point (locating point) representing the sample in a space defined by the X, Y, and Z axes; and (c) assessing the patient based on the spatial location, particularly step (c) which includes the diagnosis and prognosis of the patient's cancer. Examples of the diagnosis and prognosis of the patient's cancer include disease-free survival, overall survival, or cancer treatment prediction.
[0077] The above method may further include: (d) a fourth biomarker (D) can be used to replace one of the three biomarkers mentioned above, such as (A, B, D), to establish the location of the sample in the new space; then (e) the patient is further evaluated based on the location in the new space, in particular step (e) includes the diagnosis and prognosis of the patient's cancer. Examples of the diagnosis and prognosis of the patient's cancer include disease-free survival, overall survival, or cancer treatment prediction.
[0078] Furthermore, the above method may also include: (d) the use of a fourth and fifth biomarker (D and E) to replace one of the three biomarkers mentioned above, such as (A, D, E), to establish the location of the sample in the new space; and (e) further evaluation of the patient based on the location in the new space, particularly including the diagnosis and prognosis of the patient's cancer in step (e). Examples of the diagnosis and prognosis of the patient's cancer include disease-free survival, overall survival, or prediction of cancer treatment.
[0079] Furthermore, the above methods may also include: (d) the use of fourth, fifth, and sixth biomarkers (D, E, and F) to locate the sample in the new space; and (e) further evaluation of the patient based on the location in the new space, particularly step (e) which includes the diagnosis and prognosis of the patient's cancer. Examples of the diagnosis and prognosis of the patient's cancer include disease-free survival, overall survival, or prediction of cancer treatment.
[0080] The spatial location determined by A, B, and C, and the new location determined by A, B, and D, can be used sequentially to further evaluate the patient, including the diagnosis and prognosis of the patient's cancer. Examples of cancer diagnosis and prognosis include disease-free survival, overall survival, or cancer treatment prediction.
[0081] The spatial location determined by A, B, and C, as well as the new location determined by A, B, and D, can be used simultaneously for further patient evaluation, including the diagnosis and prognosis of the patient's cancer. The diagnosis and prognosis of the patient's cancer include disease-free survival, overall survival, or prediction of cancer treatment outcomes.
[0082] The spatial location determined by A, B, and C, and the new location determined by A, D, and E, can be used sequentially for further patient evaluation, including cancer diagnosis and prognosis. Cancer diagnosis and prognosis include disease-free survival, overall survival, or cancer treatment prediction.
[0083] The spatial location determined by A, B, and C, as well as the new location determined by A, D, and E, can be used simultaneously for further patient evaluation, including the diagnosis and prognosis of the patient's cancer. The diagnosis and prognosis of the patient's cancer include disease-free survival, overall survival, or prediction of cancer treatment.
[0084] The spatial location points determined by A, B, and C, and the new location points determined by D, E, and F, can be used sequentially for further patient evaluation, including the diagnosis and prognosis of the patient's cancer. The diagnosis and prognosis of the patient's cancer include disease-free survival, overall survival, or prediction of cancer treatment.
[0085] The spatial location determined by A, B, and C, as well as the new location determined by D, E, and F, can be used simultaneously for further patient evaluation, including the diagnosis and prognosis of the patient's cancer. The diagnosis and prognosis of the patient's cancer include disease-free survival, overall survival, or prediction of cancer treatment.
[0086] In one embodiment, the invention further includes a method for determining a reference spatial profile (sub-group) of a reference spatial database that best matches the patient's spatial profile information. This method includes the steps of: (a) comparing the spatial locations (location points) of a patient sample using the levels of three biomarkers as coordinates on a suitable programmable computer; and (b) comparing the location points with the sub-group of reference spatial profile information in the reference spatial database to determine the closeness to the sub-group; (c) identifying, on a suitable programmable computer, the sub-group of reference spatial profile information in the reference database that is closest to the location point; and (d) outputting the maximum similarity or related phenotype of the sub-group of reference spatial profile information in the reference database that best matches the patient's spatial profile information to a user interface device, a computer-readable storage medium, or a regional or remotely accessible computer system, or displaying it directly.
[0087] Mathematical methods are used to explore the presumed relationship between patient location points and a clinical presentation. Examples include, but are not limited to, the relationship between location points and disease-free survival in a three-dimensional scatter plot composed of ER, PR, and Her2 expression levels. This information can provide prognostic guidance for other patients in similar analyses.
[0088] The aforementioned RC database can be used to explore the relationship between ICP and clinical outcomes. Mathematical analysis is used to investigate the causal relationship between known clinical outcomes associated with each ICP and the expression levels of biomarkers.
[0089] Clinical “trait” and clinical “information” are interchangeable and should be considered in their broadest context. Clinical trait includes, but is not limited to, age, sex, blood pressure, glucose level, cancer grade, disease-free survival, or any information relevant to a patient’s diagnosis, prevention, and treatment.
[0090] Clinical outcomes from multiple ICPs in the database can be aggregated for statistical analysis to provide personalized diagnosis, prediction, and prognosis for patients. Statistical methods include univariate survival analysis, multivariate survival analysis, C-index analysis, Kaplan-Meier survival analysis, and log-rank survival analysis.
[0091] Unlike current classification methods used for breast and prostate cancer diagnosis, this invention utilizes a large number of stored FFPE samples widely available worldwide to identify a group of ICPs highly similar to a single patient, and analyzes their clinical outcomes to provide personalized diagnosis, prediction, and prognosis for each cancer patient.
[0092] Clearly, the more FFPE samples available for storage, the more accurate the diagnosis, prediction, and prognosis can be for patients.
[0093] In short, existing diagnostic methods involve drawing several circles and placing each cancer patient into these circles to provide precision treatment. In contrast, this invention uses each patient as the center to draw a personalized circle containing ICPs (Intracancer Chronic Diseases) similar to theirs. As a result, there are as many circles as there are cancer patients, allowing for personalized diagnosis, prediction, and prognosis for each patient using the RC (Reliable Cancer Database).
[0094] It should be noted that the exemplary embodiments described herein are currently preferred embodiments and should therefore be considered as descriptions only and not as limitations thereto. The descriptions of features or solutions in each embodiment should generally be considered as applicable to other similar features or solutions in other embodiments.
[0095] Example 1
[0096] Materials and methods
[0097] Formalin-fixed paraffin-embedded (FFPE) breast cancer tissue sections of individuals and human cell lines, along with their clinical information, were obtained from local hospitals (Yantai Yuhuangding Hospital and Binzhou Medical University Affiliated Hospital, both located in Yantai, Shandong).
[0098] All general reagents used for cell culture, including cell culture media and culture dishes, were purchased from Thermo Fisher Scientifics (Waltham, Massachusetts, USA). Protease inhibitors were purchased from Sigma Aldrich (St. Louis, Missouri, USA). All other chemicals were purchased from Sinopharm Chemical Reagent Co., Ltd. (Beijing, China).
[0099] Preparation of lysis buffer: Two 2×15 μm formalin-fixed paraffin-embedded tissue sections were collected into Eppendorf centrifuge tubes. The sections were deparaffinized and sonicated for 2 minutes in 300 μl of lysis buffer (50 mM HEPES, 137 mM NaCl, 5 mM EDTA, 1 mM MgCl2, 10 mM Na2P2O7, 1% Triton X-100, 10% glycerol) containing protease inhibitors (2 μg / ml leuprolide, 2 μg / ml aprotinin, 1 μg / ml pepsin, 2 mM PMSF, 2 mM NaF). The sections were then centrifuged at 12000×g for 5 minutes. The supernatant was collected for Western blot analysis. Total protein concentration was measured using the Pierce BCA Protein Assay Kit according to the manufacturer's instructions.
[0100] The linear ranges of the QDB assay-specific antibodies (EP3 or 4B5 clones against Her2, MIB1 clone against Ki67, SP1 against estrogen receptor (ER) and 1E2 against progesterone receptor, and SA38-08 against cyclin D1) were determined using a mixture of lysates from several patients who were positive for each of these biomarkers. First, equal volumes of 3 to 4 tissue lysates from breast cancer tissue were mixed, and the mixed lysates were serially diluted from 0 to 2 μg to determine the linear range for the QDB assay. Standard proteins, whether commercially available or expressed and purified in-house, were also serially diluted from 0 pg to 500 pg to determine the linear range for the QDB assay.
[0101] Samples were added to QDB plates in triplicate at 2 μl / unit and processed as described in the prior art. 100 μl of primary antibody was added to each well and incubated overnight at 4°C. Donkey anti-rabbit or donkey anti-mouse secondary antibody was then incubated with the QDB plates at room temperature for 4 hours. The QDB plates were first rinsed twice with TBST, then washed five times for 10 minutes each time, and then placed in white 96-well plates pre-filled with 100 μl / well of ECL solution prepared according to the manufacturer's instructions for 3 minutes. The chemiluminescence signal from each well of the recombinant plate was then quantitatively measured using a Tecan Infiniti 200pro microplate reader by selecting "Lid Plate" on the user interface.
[0102] The obtained readings were used to determine the biomarker levels in FFPE samples using control standard proteins (protein standards). The measured biomarker levels of PR, ER, Her2, and Ki67 were entered into the database. Samples were divided into three groups for QDB analysis. Six samples from each group (two strongly expressed, two weakly expressed, and two moderately expressed) were selected and measured in the same experiment to verify the consistency of the results.
[0103] These results were used to create a 3D scatter plot using OriginPro 9.1 software.
[0104] Example 1 describes how to use the protein levels of ER, PR, and Her2 to create a 3D scatter plot and use the scatter plot to determine a patient's treatment plan.
[0105] The protein levels of PR, ER, and Her2 were determined as absolute and continuous variables using the QDB method, and the results were entered into the QDB database.
[0106] Input the results of more samples into the QDB database to ensure its growth.
[0107] A 3D scatter plot was created using the ER, PR, and Her2 levels in the QDB database, and the scatter plot was continuously adjusted to ensure its accuracy and comprehensiveness. Figure 1 ).
[0108] A 3D scatter plot where each location point represents a sample is defined as a reference spatial database.
[0109] Mathematical analysis was used to correlate clinical information (including DFS and OS) with each location in the reference spatial database.
[0110] The levels of ER, PR, and Her2 in FFPE samples from patients were determined using the QDB method.
[0111] The location point representing the patient will be located in the reference spatial database.
[0112] The reference space overview information is identified based on the patient's location on the reference space distribution map, and the clinical information derived from this reference space overview information is analyzed for the patient's diagnosis, prediction, and prognosis.
[0113] Alternatively, the reference spatial profile information can be divided into different subgroups based on spatial positioning.
[0114] Clinical information (including DFS and OS) is correlated with each subgroup of the reference space profile information. In this case, this refers to the hormone group, the Her2 group, and the corner group.
[0115] The levels of ER, PR, and Her2 in the FFPE samples of patients were measured using the QDB method, and the localization points were located in a certain subgroup through spatial localization.
[0116] Clinical diagnosis, prediction, and prognosis are provided to patients based on the subgroup to which the location point is located.
[0117] Example 2
[0118] A detailed description of the materials and methods is given in Example 1.
[0119] This example demonstrates how to use 3D models based on clinical research to provide patients with diagnosis, prediction, and prognosis.
[0120] By analyzing a large amount of research spatial information with matching clinical information, a 3D model was constructed that associates spatial location with clinical information (including DFS and OS).
[0121] The protein levels of three biomarkers in patients were measured as absolute and continuous variables.
[0122] In a 3D model supported by instruments or software, patients are spatially located based on the expression levels of three biomarkers.
[0123] Using 3D models, patients can be spatially located based on the levels of three biomarkers, providing diagnosis, prediction, or prognosis.
[0124] Example 3
[0125] A detailed description of the materials and methods is given in Example 1.
[0126] This example demonstrates how to use two 3D scatter plots sequentially for clinical diagnosis, prediction, and prognosis of sub-sub-group patients.
[0127] The levels of six biomarkers, ER, PR, Her2, ki67, PCNA, and p53, were determined in FFPE samples from two patients using the QDB method.
[0128] In a 3D scatter plot of a reference spatial database with ER, PR, and Her2 expression levels as the X, Y, and Z axes, the two patients were assigned to the same subgroup based on the spatial location of the localization points determined by their ER, PR, and Her2 expression levels.
[0129] The two localization points determined based on the expression levels of Ki67, PCNA, and p53 in these two patients were located in a 3D scatter plot with Ki67, PCNA, and p53 as the X, Y, and Z axes.
[0130] Location point 201 is located on the sidewall constructed by Ki67 and PCNA, and p53 is not expressed, while location point 202 floats in space due to the strong expression of Ki67, PCNA and p53.
[0131] Therefore, although patients 201 and 202 belong to the same subgroup (luminal group A) in the 3D scatter plots of ER, PR, and Her2, they belong to different subgroups in the 3D scatter plots of Ki67, PCNA, and p53.
[0132] Example 4
[0133] The study used a breast cancer profiling database developed from 427 locally collected FFPE samples. Clinical outcomes were limited to overall survival (OS) of patients in this study. While the optimal set of protein biomarkers for assessing similarity remains to be explored, absolute quantification of several commonly used breast cancer biomarkers (ER, PR, Her2, Ki67, and cyclin D1) in these FFPE samples was performed using the QDB method. The results of these protein biomarker measurements, combined with documented clinicopathological parameters (age, tumor size, tumor grade, lymph node status), treatment regimens, and obtained clinical outcomes, created an ICP for each sample in this primary cancer profiling database.
[0134] Profiles of five FFPE samples (#1388, #1843, #1445, #1807, and #1519) randomly selected from the cancer profile database, showing significant differences in at least one biomarker level, were used as presupposed new patients (Table 1). Additionally, these samples also represent four clinical subtypes based on the IHC alternative typing (analysis): #1388 and #1843 are triple-negative (TNBC) subtypes, #1445 is Her2-positive subtype, #1807 is luminal A-like subtype, and #1519 is luminal B-like subtype.
[0135] In clinical practice, ER, PR, Her2, and Ki67 biomarkers were initially used to define patient clinical subtypes. In this study, these four biomarkers were also used to assess the similarity between different hypothetical patients and each cancer profile in the database. The Euclidean distance between each hypothetical patient and each cancer profile in the database was calculated and ranked from lowest to highest. Biomarker levels were considered to be equivalent if all levels were below the limit of quantitation (LOQ). Furthermore, due to the small size of the database, cancer profiles were rejected if one or more biomarker levels were less than 50% or more than twice that of the hypothetical patient. For example, if the hypothetical patient #1519 had an ER level of 2 nmole / g, any cancer profile in the database with an ER level less than 1 nmole / g or greater than 4 nmole / g was considered dissimilar to patient #1519 and rejected.
[0136] In this small database, 18, 35, 10, and 14 similar cancer profiles were found for hypothetical patients #1388, #1843, #1445, and #1807, respectively, while only 3 cancer profiles were similar to patient #1519 (Table 1). Clearly, the size of the database significantly limits the analytical capabilities of users or researchers, further emphasizing the need to expand the database to tens or even millions of cancer profiles to fully realize the potential of this approach.
[0137] OS was compared between the similar cancer profile information of each group (similar group) and the OS of the hypothetical patient with the corresponding clinical subtype. Figure 2 ).like Figure 2 As shown in Section a, the 10-year survival rate (10-year SP) of the similar group to #1388 (also referred to as the "#1388 similar group" or "#1388 group" in this specification and its accompanying figures, and so on) and the similar group to #1843 was slightly higher than that of the TNBC subtype. However, these differences did not reach statistical significance (p = 0.057). For the similar group to #1445, which is a Her2-positive subtype, the overall survival (OS) was slightly higher than that of the entire Her2-positive group, with the 10-year SP increasing from 72% to 88% (p = 0.27). For the 14 cancer profiles similar to #1807, which is a luminal A-like subtype, the 10-year SP was highly similar to that of the entire luminal A-like subtype (p = 0.91).
[0138] Due to the small size of the database, only three cancer profiles were similar to #1519, which is a luminal B-like subtype. To extract potential indicative information from this dataset, the similarity criteria were slightly relaxed to include cancer profiles with biomarker levels below 2 × LOQ. For example, if the LOQ for Her2 was 0.15 nmole / g, any profile with Her2 levels less than 0.3 nmole / g was considered to have the same Her2 level. Under this relaxed condition, 16 cancer profiles were identified as similar to #1519, and this similarity group showed a significantly worse 10-year SP compared to the entire luminal B-like subtype (p = 0.0096).
[0139] As is conceivable, the more biomarkers included in the assessment, the higher the level of similarity between the profile information and the new patients. Therefore, given that cyclin D1 is a biomarker independent of Ki67 in predicting overall survival in patients with the luminal subtype, it was included in the assessment to identify the similarity groups for each of the five hypothetical patients. The Euclidean distance between each profile information in the database and these five hypothetical patients was calculated separately and sorted according to the distance (Table 2). Any cancer profile information that was less than 50% or more than twice the level of the hypothetical patient's biomarker was not included in the similarity group.
[0140] As expected, even fewer cases were found to have similar cancer profiles to the presumed patients. The number of similar cases decreased from 18 to 7 in group #1388, from 35 to 20 in group #1843, from 10 to 7 in group #1445, from 14 to 5 in group #1807, and from 16 to 9 in group #1519. The overall survival (OS) of each similar group was compared with the OS of the presumed patients for the corresponding clinical subtype using the Log Rank test. Figure 3 Surprisingly, the addition of cyclin D1 resulted in statistically significant differences between groups #1843, #1388, and the TNBC group (p = 0.023). The 10-year SP (100%) of group #1843 was significantly higher than that of group #1388, while the 10-year SP (57%) of group #1388 was significantly lower than that of the TNBC subtype (75%). Figure 3 (a). On the other hand, the addition of cyclinD1 did not significantly improve the prognosis of group #1445. Figure 3 (b) The addition of cyclin D1 resulted in a worse prognosis for group #1807 compared to the luminal A-like subtype, but this difference did not reach statistical significance (p = 0.095). For group #1519, the overall survival (OS) remained worse than that of the luminal B-like subtype, with a Log Rank test result of p = 0.034.
[0141] The greatest advantage of this method is that it provides personalized treatment recommendations for new patients. Therefore, profiling information within each similar group was further subdivided based on the treatment received, and overall survival (OS) analysis was performed to determine the optimal treatment regimen. Due to the relatively small database size, treatment modalities were simply categorized into chemotherapy (CT), endocrine therapy (ET), and combined chemo-endocrine therapy (CET). This study used similar groups identified based on four biomarkers: ER, PR, Her2, and Ki67, to incorporate more profiling information for analysis.
[0142] For the #1388 similar group ( Figure 4 (a) The 5-year survival rate (5-year SP) was 92% for patients who received CT alone (n=12), compared to 67% for patients who received CET alone (n=3). In group #1843, all patients survived regardless of the treatment received. For group #1445, no relevant information could be obtained because all patients in this group received only CT. In group #1807, the 10-year SP was 100% for patients who received CT alone (n=4), compared to 71% for patients who received CET alone (n=8). In group #1519, although patients who received CET (n=6) showed an advantage in 5-year SP compared to those who received CT alone (n=9), this advantage disappeared later in the treatment.
[0143] Therefore, this paper demonstrates the feasibility of a proximity-based diagnostic approach to provide personalized overall survival (OS) prediction for five presumed patients. The 10-year SP (spinal paralysis) of #1388, #1843, and #1519 differed significantly from the 10-year SP of their corresponding clinical subtypes. OS analysis also assessed the effectiveness of various treatment modalities in similar groups of the five presumed patients. However, due to database size limitations, their differences were not statistically significant and could not provide guidance for these presumed patients.
[0144] It is noteworthy that although #1843 belongs to the worst-prognostic clinical subtype of TNBC, all patients in its similar groups survived at the end of the study regardless of the treatment received. Figure 3 a and Figure 4 (b). This raises suspicion that #1843 belongs to the normal-like subtype (also known as "quasi-normal subtype") described in the original molecular typing studies.
[0145] Clearly, as the concepts and technologies mature, the rate-limiting step in applying this novel diagnostic method in routine clinical practice is to develop a cancer profile database much larger than the one used in this study. For example, with more than 10,000 breast cancer profiles, based on five biomarkers, approximately 500 profiles (10,000 / 400 × 20) could be identified as highly similar to #1843, 125 profiles as highly similar to #1807, and 175 profiles as highly similar to #1388 and #1445, thus providing reliable guidance for these five presumptive patients.
[0146] However, even at this scale, it remains difficult to identify a sufficient number of cancer profiles similar to #1519, as only 75 profiles can be identified based on four biomarkers, and even fewer can be identified if five biomarkers are used to assess similarity. It is conceivable that patients like #1519 are more likely to be overtreated or undertreated in current clinical practice, as they are unlikely to receive adequate care in clinical trials involving hundreds or thousands of cases.
[0147] Similarly, due to the limited size of the database, this study categorized treatment modalities into CT, ET, and CET. However, regarding chemotherapy alone, there are at least four types of drugs: alkylating agents, antitumor antibiotics, antimetabolites, and mitotic inhibitors. Within each type, several different drugs are available. Clearly, only by expanding the database can the exact drugs with the best therapeutic outcomes for new patients be identified.
[0148] All these considerations underscore the necessity of extensively expanding the database. Fortunately, due to the absolutely quantitative nature of QDB technology, the aforementioned cancer profile database is a continuously growing database. As this approach gains global acceptance, it is expected to grow exponentially in the near future.
[0149] The five biomarkers used in this study may not be the optimal choice for assessing similarity among breast cancer patients. However, they can serve as a basis for identifying the optimal number and combination of biomarkers from all existing and candidate biomarkers to assess similarity among breast cancer patients. The minimum amount of profile information on similarity required to provide reliable guidance to patients needs to be determined through collaborative efforts by oncologists and statisticians worldwide.
[0150] Admittedly, the role of Euclidean distance in this study is limited because the number of cancer profiles with all biomarker levels greater than 50% and less than 200% of the hypothetical patient is finite. However, its enormous potential for significantly expanded databases can be imagined, allowing the establishment of a cut value based on Euclidean distance to identify cancer profiles with the highest similarity to new patients.
[0151] It is also worth mentioning that, although the similarity assessment in this study was based entirely on the expression levels of a set of protein markers, other clinicopathological parameters (such as age, tumor size, lymph node status, RS score based on Oncotype, and ROR score based on PAM50) could be incorporated into the assessment to improve the level of similarity. Similarly, although this study was limited to OS analysis, other clinical outcomes, including recurrence, could be used for assessment when needed.
[0152] Materials and methods
[0153] Individuals and Human Cell Lines: 490 formalin-fixed paraffin-embedded (FFPE) breast cancer tissue sections (2×15μm / case) were provided jointly by the Yantai Affiliated Hospital of Binzhou Medical University and the Yantai Yuhuangding Hospital Affiliated to Qingdao University (Yantai, China), of which 427 cases had overall survival (OS) data. All studies, including sample collection and research protocols, complied with the Declaration of Helsinki and were approved by the Ethics Committee of the Yantai Affiliated Hospital of Binzhou Medical University (Approval No.: #20191127001 – Hao Junmei) and the Ethics Committee of Yuhuangding Hospital (Approval No.:
[2017] 76 – Yu Guohua). Informed consent was waived because the individuals in these two studies were anonymized FFPE samples with retrospective clinical data stored in large quantities by the hospitals.
[0154] Except for biomarker levels, which were determined using the QDB method, all clinical information was obtained from medical records. The QDB procedure is also described in detail in other literature.
[0155] Universal reagents: All universal reagents are described elsewhere. ER(SP1) rabbit monoclonal antibody was purchased from Abcam Inc., PR(1E2) rabbit monoclonal antibody from Roche Diagnostics GmbH, and Her2(EP3) rabbit monoclonal antibody, Ki67(MIB1) mouse monoclonal antibody, and cyclinD1(EP12) rabbit monoclonal antibody were all purchased from ZS Biotech Co., Ltd. (Beijing, China, www.zsbio.com). HRP-labeled donkey anti-rabbit IgG secondary antibody was purchased from Jackson ImmunoResearch Laboratory (PA, Pike West Grove, USA). QDB plates were provided by Quanticision Diagnostics Inc. (NC, RTP, USA).
[0156] Euclidean distance calculation: Calculate the distances between sample a (ER, PR, Her2, Ki67) and sample b (ER', PR', Her2', Ki67') and sample c (ER, PR, Her2, Ki67, and cyclinD1) and sample d (ER', PR', Her2', Ki67', cyclinD1') using the following formulas:
[0157] d_ab=√((ER-ER')^2+(PR-PR')^2+(Her2-Her2')^2+(Ki67-Ki67')^2)
[0158] d_cd=√((ER-ER')^2+(PR-PR')^2+(Her2-Her2')^2+(Ki67-Ki67')^2+(cyclinD1-cyclinD1')^2).
[0159] The QDB results for each biomarker in all samples were pre-standardized (via z-score conversion). Limits of quantitation (LOQ): Her2 0.15 nmol / g, ER 0.1 nmol / g, PR 0.25 nmol / g, Ki67 1.3 nmol / g.
[0160] Survival analysis: Overall survival in different subgroups was displayed using the Kaplan-Meier method and compared using the log-rank test. P < 0.05 was considered statistically significant. All statistical analyses were performed using R 4.0.1 (http: / / www.r-project.org).
[0161] It is believed that certain combinations and sub-combinations of features specifically pointed out in the appended claims, derived from the invention disclosed above, possess novelty and inventiveness. Inventions embodied in other combinations and sub-combinations of features, functions, elements, or attributes can be protected by amending these claims or by filing new claims in this application or related applications. These amended or new claims, whether they pertain to different or the same invention, and whether they differ in scope, are broader, narrower, or identical to the original claims, are considered to be included within the scope of the disclosure of this invention.
[0162] Table 1: Profile information groups of similar cancers based on absolute quantitative levels of ER, PR, Her2, and Ki67 proteins (similar groups)
[0163] Table 1-1#1388 Similar Groups
[0164]
[0165] Table 1-2 #1843 Similar Groups
[0166]
[0167] Table 1-3#1445 Similar Groups
[0168]
[0169] Table 1-4#1807 Similar Groups
[0170]
[0171] Table 1-5#1519* Similar Groups
[0172]
[0173] Abbreviations: All clinicopathological factors meet the American Joint Committee on Cancer (AJCC) definition. C: Chemotherapy; E: Endocrine therapy; LumA and LumB: Lum A-like and Lum B-like clinical subtypes; TNBC: Triple-negative breast cancer; Her2: Her2-positive subtype; Survival outcome: 0 - alive, 1 - dead.
[0174] Note: All biomarker levels within the limit of quantitation (LOQ) are considered identical. Samples with at least one biomarker level that is significantly different from the presumed patient level (<50% or >200%) are discarded. *: Biomarker levels within the 2×LOQ range are considered identical. Samples with at least one biomarker level that is significantly different from the presumed patient level (<50% or >200%) are discarded.
[0175] Table 2: Summary information groups of similar cancers based on absolute quantitative levels of ER, PR, Her2, Ki67, and cyclin D1 proteins (similar groups)
[0176] Table 2-1#1388 Similar Groups
[0177]
[0178] Table 2-2#1843 Similar Groups
[0179]
[0180] Table 2-3#1445 Similar Groups
[0181]
[0182] Table 2-4#1807 Similar Groups
[0183]
[0184] Table 2-5#1519* Similar Groups
[0185]
[0186] Abbreviations: All clinicopathological factors meet the American Joint Committee on Cancer (AJCC) definition. C: Chemotherapy; E: Endocrine therapy; LumA and LumB: Lum A-like and Lum B-like clinical subtypes; TNBC: Triple-negative breast cancer; Her2: Her2-positive subtype; Survival outcome: 0 - alive, 1 - dead.
[0187] Note: All biomarker levels within the limit of quantitation (LOQ) are considered identical. Samples with at least one biomarker level that is significantly different from the presumed patient's level (<50% or >200%) are discarded. *: Biomarker levels within the 2×LOQ range are considered identical. Samples with at least one biomarker level that is significantly different from the presumed patient's level (<50% or >200%) are discarded.
Claims
1. A method for generating a database for cancer diagnosis, prediction, and prognosis, comprising the following steps: Provide multiple individuals, each with a known clinical outcome of cancer; Individual cancer profile information (ICP) is generated for each of the plurality of individuals. This ICP includes i) a set of clinical parameters to distinguish each individual and ii) known clinical outcomes of the cancer, wherein each clinical parameter in the set represents a quantitative measurement of a different biomarker, the quantitative measurement being a continuous and absolute quantity of the biomarker from a stored formalin-fixed paraffin-embedded FFPE sample; and The individual cancer profile information of multiple individuals generated is stored in a database. The database is structured as follows: The clinical parameter set in the database can be combined in a way that is sufficient to distinguish individual FFPE samples from the stored FFPE samples. The combination of absolute quantitative determination results of the different biomarkers forms a unique combination for each of the stored FFPE samples in the database, and can be combined with matching clinical records around this unique combination.
2. The method as described in claim 1, wherein, The biomarker in question is a protein biomarker.
3. The method as described in claim 1, wherein, The quantitative determination was performed using the Quantitative Dot Immunoblot (QDB) method.
4. The method of claim 1, wherein, The cancer is breast cancer, and the biomarker is estrogen receptor ER, progesterone receptor PR, Ki67, p53, cyclin D1, or Her2.
5. A database for providing cancer diagnosis, prediction, and prognosis, the database comprising multiple individual cancer profiles, each of which is generated by an individual with a known clinical outcome of cancer, wherein: The individual cancer profile information includes i) a set of clinical parameters quantitatively determined from the individual’s stored FFPE samples to distinguish each individual and ii) the known clinical outcomes of the cancer; Each of the clinical parameters represents the quantitative measurement results of different biomarkers, and The quantitative determination result is a continuous and absolute amount of the biomarker in the sample; The database is structured as follows: The clinical parameter set in the database can be combined in a way that is sufficient to distinguish individual FFPE samples from the stored FFPE samples. The combination of absolute quantitative determination results of the different biomarkers forms a unique combination for each of the stored FFPE samples in the database, and can be combined with matching clinical records around this unique combination.
6. The database as described in claim 5, wherein, The quantitative determination was achieved using the Quantitative Dot Immunoblot (QDB) method.
7. The database as described in claim 5, wherein, The cancer in question is breast cancer, and the biomarkers are estrogen receptor ER, progesterone receptor PR, Ki67, p53, cyclin D1, or Her2.
8. An apparatus for diagnosing cancer in a patient, the apparatus comprising the database of claim 5.
9. A kit for diagnosing cancer in a patient, the kit comprising the database of claim 5.
10. A method for providing cancer diagnosis, prediction, and prognosis to patients, comprising the following steps: 1) Collect and store FFPE samples from patients; 2) Obtain the database of claim 5, and obtain from it i) stored individual cancer profile information and ii) the set of clinical parameters used in the database; 3) Quantitative levels of clinical parameters identical to those used in the clinical parameter set in the database are determined from the FFPE samples of the patient; 4) Compare the quantitative levels of the set of clinical parameters measured from the FFPE samples of the patient with the quantitative levels of the same clinical parameters in the database; 5) The comparison is used to identify the individual cancer profile information in the database that best matches the patient; 6) Output clinical outcomes from individual cancer profile information identified from the database.
11. The method of claim 10, wherein, The comparison is used to determine the maximum similarity between the set of clinical parameters measured from the patient's FFPE sample and the same set of clinical parameters in the individual's cancer profile information.
12. The method of claim 11, wherein, When determining the similarity, the absolute level of each biomarker is within a preset range of the absolute levels of the same biomarkers in the same set of clinical parameters measured from the FFPE sample of the patient.
13. The method of claim 11, wherein, The similarity is calculated based on the Euclidean distance between two sets of quantitatively identical clinical parameters.
14. The method of claim 10, wherein, The quantitative levels of the clinical parameter set from the FFPE samples were achieved using the Quantitative Dot Immunoblot (QDB) method.
15. The method of claim 10, wherein, The cancer is breast cancer, and the set of clinical parameters includes estrogen receptor ER, progesterone receptor PR, Ki67, p53, cyclin D1, Her2, or combinations thereof.
Citation Information
Patent Citations
Apparatus and method for absolute quantification of biomarkers for solid tumor diagnosis
WO2019006156A1