Method and System for Predicting Malignancy of Indeterminate Pulmonary Nodules
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ABBOTT LAB INC
- Filing Date
- 2023-07-13
- Publication Date
- 2026-07-17
AI Technical Summary
Current lung cancer screening methods, such as LDCT, fail to accurately distinguish between malignant and benign indeterminate pulmonary nodules (IPNs), leading to high false-positive rates and unnecessary invasive procedures, especially in high-risk populations.
A method and system using machine learning algorithms to analyze subject values including pack-year smoking history, IPN size, and biomarker concentrations (CA125, CEA, HE4, Cyfra 21-1, NSE, SCC, ProGRP, total IgG, IgA, IgM, IgE, kappa and lambda light chains) to generate a score for determining the likelihood of malignancy, compared against a reference score.
Improves the accuracy of distinguishing malignant IPNs from benign, reducing unnecessary procedures by providing a reliable predictive tool for early-stage lung cancer detection.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] Information on Related Applications This application claims priority to U.S. Application Publication No. 63 / 389,077, filed on July 14, 2023, the content of which is incorporated herein by reference.
[0002] The present disclosure relates to a method and system for determining whether at least one indeterminate pulmonary nodule (IPN) identified in a subject is likely to be malignant or not likely to be malignant. The methods and systems of the present disclosure utilize certain subject values including (i) the subject's pack-year smoking value and (ii) a measured value of the size of the IPN in the subject, and (iii) (a) at least one of the subject's cancer antigen 125 (CA125) concentration, the subject's carcinoembryonic antigen (CEA) concentration, the subject's human epididymis protein 4 (HE4) concentration, the subject's cytokeratin fragment 21-1 (Cyfra 21-1) concentration, the subject's neuron-specific enolase (NSE) concentration, the subject's squamous cell carcinoma antigen (SCC) concentration, the subject's progastrin-releasing peptide (ProGRP) concentration, or any combination thereof, and (b) at least one of the subject's total IgG concentration, the subject's IgA concentration, the subject's IgM concentration, the subject's IgE concentration, the subject's free kappa light chain concentration, the subject's free lambda light chain concentration, or any combination thereof, derived from a biological sample obtained from the subject. The subject values and assay values are applied to at least one machine learning algorithm used to create or output a score (a machine learning score) for the subject. The machine learning score for the subject is compared to a reference score to determine whether the IPN is likely to be malignant or not likely to be malignant.
Background Art
[0003] Lung cancer is the second most common cancer worldwide. In 2020, an estimated 2.2 million cases of lung cancer and over 1.7 million deaths were seen globally (see Sung H, Ferlay J, Siegel R et al., Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries, CA Cancer J Clin 71(2021)209 - 249). In the United States, the most common form of lung cancer is non - small cell lung cancer (NSCLC), which accounts for approximately 80 - 85% of all lung cancers. Furthermore, more than half of all NSCLC patients are also diagnosed with progressive and / or metastatic disease (NIH Surveillance, Epidemiology, and End Results (SEER Program)). Currently, 57% of all lung cancers are diagnosed at a late stage, significantly reducing the chances of survival. This data is consistent with current survival statistics for patients diagnosed with metastatic lung cancer, as recent epidemiological studies have shown that the 5 - year survival rate for patients with metastatic lung cancer is less than 10% (NIH Surveillance, Epidemiology, and End Results). For early - stage localized lung cancer, the cure rate is significantly higher compared to cancer that has spread throughout the body. Therefore, early diagnosis can play an important role in improving the outcome of cancer by early treatment intervention, subsequently increasing the cure rate and survival rate.
[0004] To identify and treat lung cancer patients in the early stages of the disease, the National Lung Screening Trial (NLST) of the National Cancer Institute established that annual low-dose computed tomography (LDCT) screening in certain high-risk groups reduces lung cancer mortality (Wendler R, Fontham E, Barrera E et al., American Cancer Society Lung Cancer Screening Guidelines, CA Cancer J Clin 63, 2 (2013) 106-117). The high-risk lung cancer population eligible to receive the benefits of this screening is defined by the American Cancer Society as apparently healthy patients aged 55-75 years with at least a 20-pack-year smoking history who are currently smoking or who quit smoking within the past 15 years (Smith R, Andrews K, Brooks D et al., Cancer Screening in the United States, 2018: A Review of Current American Cancer Society Guidelines and Current Issues in Cancer Screening, CA Cancer J Clin 68, 4 (2018) 297-316). However, LDCT does not detect all lung cancers and can generate a high frequency of false-positive findings. LDCT false-positive findings can lead to further non-invasive and / or invasive procedures being performed to determine whether the observed image abnormalities are truly cancerous (Wendler R, Fontham E, Barrera E et al., American Cancer Society Lung Cancer Screening Guidelines, CA Cancer J Clin 63, 2 (2013) 106-117). Among the image abnormalities found with the increasing LDCT screening of the high-risk lung cancer population are indeterminate pulmonary nodules (IPNs).An IPN is a non-calcified pulmonary nodule (usually 7 - 20 mm in size) that requires further diagnostic workup for the risk of malignancy (Maission P, Walker R., Indeterminate Pulmonary Nodules: Risk for Having or for Developing Lung Cancer, Cancer Prev Res (Phila) 7, 12 (2014) 1173 - 1178). The IPN size itself correlates with risk, with larger nodule sizes having a higher risk of being cancerous: IPN < 5 mm: 0% - 1%; 5 - 10 mm: 6% - 28%; 11 - 20 mm: 37% - 64%; > 20 mm: 64% - 82% (Maission P, Walker R., Indeterminate Pulmonary Nodules: Risk for Having or for Developing Lung Cancer, Cancer Prev Res (Phila) 7, 12 (2014) 1173 - 1178).
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
[0006] Since most of the high-risk lung cancer population is classified as IPN (55-76%), further tools are needed to distinguish malignant tumors from benign conditions that exceed the nodule size (Maission P, Walker R., Indeterminate Pulmonary Nodules: Risk for Having or for Developing Lung Cancer, Cancer Prev Res (Phila) 7, 12 (2014) 1173-1178). [Means for Solving the Problems]
[0007] [Abstract] In one embodiment, the present disclosure relates to a method for determining whether at least one indeterminate pulmonary nodule (IPN) identified in a subject is likely to be malignant. The method comprises a) providing a subject value for the subject, wherein the subject value is the following: i) the subject's smoking pack-year value, ii) identification of biological sex, iii) identification of ethnicity, iv) identification of the type of nodules, and v) a measurement of the size of the IPN in the subject including at least one of; b) providing at least two assay values, wherein the at least two assay values are i) at least one of the concentration of cancer antigen 125 (CA125) of the subject, the concentration of carcinoembryonic antigen (CEA) of the subject, the concentration of human epididymis protein 4 (HE4) of the subject, the concentration of cytokeratin fragment 21-1 (Cyfra 21-1) of the subject, the concentration of neuron-specific enolase (NSE) of the subject, the concentration of squamous cell carcinoma antigen (SCC) of the subject, the concentration of progastrin-releasing peptide (ProGRP) of the subject, or any combination thereof, derived from a biological sample obtained from the subject, and ii) at least one of the total IgG concentration of the subject, the IgA concentration of the subject, the IgM concentration of the subject, the IgE concentration of the subject, the free kappa light chain concentration, the free lambda light chain concentration, or any combination thereof, derived from a biological sample obtained from the subject including; c) providing a processing system including a computer processor, a database, and a non-transitory computer memory including at least one machine learning algorithm, wherein the at least one machine learning algorithm is configured to process the object values and the assay values, and further, the processing system i) is configured to apply the at least one machine learning algorithm to the assay values and the object values to output a machine learning score for the subject, ii) is configured to report the machine learning score for the subject, and iii) is configured to provide a reference score for comparison with the machine learning score including; and d) determining that the IPN is likely to be malignant if the machine learning score is higher than the reference score, and is likely to be non-malignant if the machine learning score is the same as or lower than the reference score including.
[0008] In one aspect of the above method, the target smoking pack-year value is the number of tobacco packs smoked by the target per year multiplied by the number of years the target has smoked.
[0009] In another aspect of the above method, the method includes obtaining an assay value including the target's CA125 concentration, the target's total IgG concentration, the target's IgA concentration, the target's IgM concentration, the target's IgE concentration, the kappa free light chain concentration, and the lambda free light chain concentration from a biological sample obtained from the target.
[0010] In another aspect of the above method, the method includes obtaining an assay value including the target's CA125 concentration, the target's CEA concentration, the target's HE4 concentration, the target's Cyfra 21-1 concentration, the target's NSE concentration, the target's SCC concentration, the target's ProGRP concentration, the target's total IgG concentration, the target's IgA concentration, the target's IgM concentration, the target's IgE concentration, the kappa free light chain concentration, and the lambda free light chain concentration from a biological sample obtained from the target.
[0011] In some aspects of the above method, at least one machine learning algorithm is an Adaptive Index Modeling (AIM) algorithm that generates an AIM score. In other aspects, the machine learning algorithm is a random forest algorithm that generates a random forest score. In still other aspects, the machine learning algorithm is a logistic regression algorithm that generates a logistic regression score. In some aspects, the machine learning algorithm uses an AIM algorithm, a random forest algorithm, or any combination of random forest algorithms.
[0012] In another aspect of the above method, the reference score is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100.
[0013] In a further aspect of the above method, the biological sample is a whole blood sample, a serum sample or a plasma sample.
[0014] In yet another aspect of the above method, obtaining the target value, the assay value, or the target value and the assay value includes receiving the target value, the assay value, or the target value and the assay value from a laboratory, from the subject, from an analytical testing system, from a handheld or point-of-care testing device, or from any combination thereof.
[0015] In a further aspect of the above method, obtaining the target value, the assay value, or the target value and the assay value includes electronically receiving the target value.
[0016] In a further aspect of the above method, the method further includes manually or automatically inputting the target value, the assay value, or the target value and the assay value into the processing system.
[0017] In a further aspect of the above method, the processing system compares a machine learning score for the subject to the reference score.
[0018] In a further aspect of the above method, a determination as to whether the IPN is likely to be malignant or not likely to be malignant is displayed on the device.
[0019] In a further aspect of the above method, the subject is human.
[0020] In a further aspect of the above method, at least one of the CA125 concentration of the subject, the CEA concentration of the subject, the HE4 concentration of the subject, the Cyfra 21-1 concentration of the subject, the NSE concentration of the subject, the SCC concentration of the subject, the ProGRP concentration of the subject or any combination thereof is determined using an immunoassay.
[0021] In a further aspect of the above method, at least one of the total IgG concentration of the subject, the IgA concentration of the subject, the IgM concentration of the subject, the IgE concentration of the subject, the free kappa light chain concentration of the subject, the free lambda light chain concentration of the subject or any combination thereof is determined using a clinical chemistry assay.
[0022] In another embodiment, the present disclosure is a. a subject value for a subject, comprising i) the subject's pack-year value, and ii) a measured value of the size of the IPN in the subject comprising the subject value; b. i) at least one of the concentration of cancer antigen 125 (CA125) of the subject, the carcinoembryonic antigen (CEA) concentration of the subject, the human epididymis protein 4 (HE4) concentration of the subject, the cytokeratin fragment 21-1 (Cyfra 21-1) concentration of the subject, the neuron-specific enolase (NSE) concentration of the subject, the squamous cell carcinoma antigen (SCC) concentration of the subject, the pro-gastrin-releasing peptide (ProGRP) concentration of the subject or any combination thereof, derived from a biological sample obtained from the subject, and ii) at least one of the total IgG concentration of the subject, the IgA concentration of the subject, the IgM concentration of the subject, the IgE concentration of the subject, the free kappa light chain concentration, the free lambda light chain concentration or any combination thereof, derived from a biological sample obtained from the subject for measuring one or more assays; c. a device comprising a processing system, the processing system comprising a computer processor and a non-transitory computer memory comprising a database and at least one machine learning algorithm, At least one machine learning algorithm is configured to process a target value and an assay value to generate a machine learning score for the target. Furthermore, the processing system is i) configured to apply at least one machine learning algorithm to the target value and the assay value to output a machine learning score for the target, ii) configured to report the machine learning score for the target, and iii) configured to provide a reference score for comparison with the machine learning score and is a device that displays that the IPN is (1) likely to be malignant when the machine learning score is higher than the reference score, or (2) less likely to be malignant when the machine learning score is the same as or lower than the reference score. relates to a system that includes.
[0023] In one aspect of the above system, the target smoking pack-year value is the number of tobacco packs smoked by the target per year multiplied by the number of years the target has smoked.
[0024] In another aspect of the above system, the assay value includes the target's CA125 concentration, the target's total IgG concentration, the target's IgA concentration, the target's IgM concentration, the target's IgE concentration, the kappa free light chain concentration, and the lambda free light chain concentration derived from a biological sample obtained from the target.
[0025] In another aspect of the above method, the method includes obtaining an assay value that includes the target's CA125 concentration, the target's CEA concentration, the target's HE4 concentration, the target's Cyfra 21-1 concentration, the target's NSE concentration, the target's SCC concentration, the target's ProGRP concentration, the target's total IgG concentration, the target's IgA concentration, the target's IgM concentration, the target's IgE concentration, the kappa free light chain concentration, and the lambda free light chain concentration derived from a biological sample obtained from the target.
[0026] In some aspects of the above system, at least one machine learning algorithm is an Adaptive Index Modeling (AIM) algorithm that generates an AIM score. In other aspects, the machine learning algorithm is a random forest algorithm that generates a random forest score. In still other aspects, the machine learning algorithm is a logistic regression algorithm that generates a logistic regression score. In some aspects, the machine learning algorithm uses the AIM algorithm, the random forest algorithm, or any combination of the random forest algorithms.
[0027] In yet another aspect of the above system, the reference score is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100.
[0028] In yet another aspect of the above system, the biological sample is a whole blood sample, a serum sample, or a plasma sample.
[0029] In a further aspect of the above system, the target value, the assay value, or the target value and the assay value are received from a laboratory, from the subject, from an analytical testing system, from a handheld or point-of-care testing device, or from any combination thereof.
[0030] In yet another aspect of the above system, the target value, the assay value, or the target value and the assay value are received electronically.
[0031] In yet another aspect of the above system, the target value, the assay value, or the target value and the assay value are input into the processing system manually or automatically.
[0032] In yet another aspect of the above system, the subject is a human.
[0033] In yet another aspect of the above system, at least one of the subject's CA125 concentration, the subject's CEA concentration, the subject's HE4 concentration, the subject's Cyfra 21-1 concentration, the subject's NSE concentration, the subject's SCC concentration, the subject's ProGRP concentration, or any combination thereof is determined using an immunoassay.
[0034] In yet another aspect of the above system, at least one of the subject's total IgG concentration, the subject's IgA concentration, the subject's IgM concentration, the subject's IgE concentration, the subject's free kappa light chain concentration, the subject's free lambda light chain concentration, or any combination thereof is determined using a clinical chemistry assay.
[0035] In yet another aspect, the present disclosure a) measuring, in one or more diagnostic assays configured to measure assay values including (a) at least one of the subject's cancer antigen 125 (CA125) concentration, the subject's carcinoembryonic antigen (CEA) concentration, the subject's human epididymis protein 4 (HE4) concentration, the subject's cytokeratin fragment 21-1 (Cyfra 21-1) concentration, the subject's neuron-specific enolase (NSE) concentration, the subject's squamous cell carcinoma antigen (SCC) concentration, the subject's progastrin-releasing peptide (ProGRP) concentration, or any combination thereof, and (b) at least one of the subject's total IgG concentration, the subject's IgA concentration, the subject's IgM concentration, the subject's IgE concentration, free kappa light chain concentration, free lambda light chain concentration, or any combination thereof, derived from a biological sample obtained from the subject; b) providing a processing system including a computer processor and a non-transitory computer memory including a database and at least one machine learning algorithm, At least one machine learning algorithm is configured to process assay values and subject values for a subject, including the subject's smoking pack-year value and a measurement of the size of the IPN in the subject, where the subject's smoking pack-year value and the measurement of the size of the IPN in the subject are pre-entered into a database. A processing system is i) configured to apply at least one machine learning algorithm to the assay values and subject values to output a machine learning score for the subject, ii) configured to report the machine learning score for the subject, and iii) configured to provide a reference score for comparison with the machine learning score in step which is included in a method.
[0036] In some aspects of the above method, at least one machine learning algorithm is an Adaptive Index Modeling (AIM) algorithm that generates an AIM score. In other aspects, the machine learning algorithm is a random forest algorithm that generates a random forest score. In still other aspects, the machine learning algorithm is a logistic regression algorithm that generates a logistic regression score. In some aspects, at least one machine learning algorithm uses an AIM algorithm, a random forest algorithm, or any combination of random forest algorithms.
[0037] In yet another embodiment, the present disclosure a) At least one of the concentration of cancer antigen 125 (CA125) of the subject derived from a biological sample obtained from the subject, the concentration of carcinoembryonic antigen (CEA) of the subject, the concentration of human epididymis protein 4 (HE4) of the subject, the concentration of cytokeratin fragment 21-1 (Cyfra 21-1) of the subject, the concentration of neuron-specific enolase (NSE) of the subject, the concentration of squamous cell carcinoma antigen (SCC) of the subject, the concentration of progastrin-releasing peptide (ProGRP) of the subject, or any combination thereof, and (b) at least one of the total IgG concentration of the subject derived from a biological sample obtained from the subject, the IgA concentration of the subject, the IgM concentration of the subject, the IgE concentration of the subject, the kappa free light chain concentration, the lambda free light chain concentration, or any combination thereof, and one or more diagnostic assays set to measure the assay values containing these; b) A processing system including a computer processor, a database, and a non-transitory computer memory containing at least one machine learning algorithm, wherein at least one machine learning algorithm is set to process assay values and subject values related to the subject, including the smoking pack-year value of the subject and the measured value of the size of the IPN in the subject, and the smoking pack-year value of the subject and the measured value of the size of the IPN in the subject are pre-entered into the database, and the processing system is i) set to apply at least one machine learning algorithm to the assay values and subject values to output a machine learning score related to the subject, ii) set to report the machine learning score related to the subject, and iii) set to provide a reference score for comparison with the machine learning score the processing system relating to a system including this.
[0038] In some aspects of the above system, at least one machine learning algorithm is an Adaptive Index Modeling (AIM) algorithm that generates an AIM score. In other aspects, the machine learning algorithm is a random forest algorithm that generates a random forest score. In still other aspects, the machine learning algorithm is a logistic regression algorithm that generates a logistic regression score. In some aspects, the machine learning algorithm uses the AIM algorithm, the random forest algorithm, or any combination of the random forest algorithms.
Brief Description of the Drawings
[0039]
Figure 1
Modes for Carrying Out the Invention
[0040] The present disclosure relates to a method and system for determining whether at least one indeterminate pulmonary nodule (IPN) identified in a subject is likely to be malignant or not likely to be malignant. The methods and systems described herein utilize certain subject values including (i) the subject's pack-year smoking value and (ii) a measured value of the size of the IPN in the subject, and (iii) (a) at least one of the subject's cancer antigen 125 (CA125) concentration, the subject's carcinoembryonic antigen (CEA) concentration, the subject's human epididymis protein 4 (HE4) concentration, the subject's cytokeratin fragment 21-1 (Cyfra 21-1) concentration, the subject's neuron-specific enolase (NSE) concentration, the subject's squamous cell carcinoma antigen (SCC) concentration, the subject's progastrin-releasing peptide (ProGRP) concentration, or any combination thereof, and (b) at least one of the subject's total IgG concentration, the subject's IgA concentration, the subject's IgM concentration, the subject's IgE concentration, kappa free light chain concentration, lambda free light chain concentration, or any combination thereof, from a biological sample obtained from the subject. The subject values and assay values are applied to at least one machine learning algorithm used to create a score (a machine learning score) for the subject. The machine learning score for the subject is compared to a reference score to determine whether the IPN is likely to be malignant or not likely to be malignant. Specifically, if the subject's machine learning score is higher than the reference score, the IPA is likely to be malignant. If the subject's machine learning score is the same as or less than (e.g., less than) the reference score, the IPA is likely not to be malignant.
[0041] 1. Definitions Unless otherwise noted, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art. In case of conflict, the present document, including definitions, will control. Preferred methods and materials are described below, but methods and materials similar or equivalent to those described in this specification can be used in the practice or testing of the present disclosure. All publications, patent applications, patents, and other references mentioned in this specification are hereby incorporated by reference in their entirety. The materials, methods, and examples disclosed in this specification are for illustrative purposes only and are not intended to be limiting.
[0042] The terms "comprising," "including," "having," "has," "can," "containing," and variations thereof, as used in this specification, are intended to be open-ended transitional phrases, terms, or words that do not preclude the possibility of further acts or structures. The singular forms ("a," "an," and "the") include the plural referents unless the context clearly dictates otherwise. The present disclosure also contemplates other embodiments that "comprise," "consist of," and "consist essentially of" the embodiments or elements presented herein, whether or not explicitly described.
[0043] Regarding the recitation of numerical ranges in this specification, each number therebetween is explicitly intended with the same degree of precision. For example, for the range of 6 to 9, in addition to 6 and 9, the numbers 7 and 8 are intended, and for the range of 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are explicitly intended.
[0044] As used interchangeably herein, the “biological sample” or “sample” includes, but is not necessarily limited to, blood-related samples (e.g., whole blood (e.g., including capillary whole blood samples), serum, plasma, and other blood-derived samples), urine, cerebrospinal fluid, bronchoalveolar lavage fluid, and other body fluids. Another example of a biological sample is a tissue sample. The biological sample may be fresh or stored (e.g., blood or blood fraction stored in a blood bank). The biological sample may be a body fluid obtained expressly for use in the assays described herein or a body fluid obtained for another purpose that can be subsampled for use in the assays described herein. In certain embodiments, the biological sample is whole blood. Whole blood can be obtained from a subject using standard clinical procedures. In other embodiments, the biological sample is plasma. Plasma can be obtained from a whole blood sample by centrifugation of anticoagulated blood. Such a process provides a buffy coat of leukocyte components and a supernatant of plasma. In certain embodiments, the biological sample is serum. Serum can be obtained by centrifugation of a whole blood sample collected in a tube without anticoagulant. The blood may be allowed to clot before centrifugation. The yellowish fluid obtained by centrifugation is serum. In another embodiment, the sample is urine. The sample may be pretreated by dilution in a suitable buffer solution as needed, heparinized, concentrated if desired, or fractionated by any number of methods including, but not limited to, ultracentrifugation, fractionation by fast protein liquid chromatography (FPLC), or precipitation of apolipoprotein B-containing proteins with dextran sulfate or other methods. Any of a number of standard buffered aqueous solutions at physiological pH, such as phosphate, Tris, etc., can be used. In further embodiments, the biological sample is a whole blood sample and the subject is human. In further embodiments, the biological sample is a plasma sample and the subject is human. In further embodiments, the biological sample is a serum sample and the subject is human. In further embodiments, the biological sample is a capillary blood sample and the subject is human.
[0045] As used interchangeably herein, "decentralize," "decentralized," or "decentralization" refers to the performance of one or more medical tests and / or assays outside of traditional medical settings (e.g., hospitals, clinics, independent research facilities, etc.) at one or more locations such as emergency medical clinics, retail clinics, pharmacies, grocery stores or convenience stores, residences (e.g., homes, apartments, etc.), workplaces, and / or government agencies (e.g., the U.S. Transportation Security Administration) in the context of an examination. "Hybrid decentralized" or "hybrid decentralized" refers to a situation where a subject or patient avoids a specialized collection facility (e.g., a hospital, clinic, or independent sample collection or research facility) and collects a sample at a residence and / or workplace and sends the sample to a laboratory.
[0046] As used interchangeably herein, "higher throughput assay analyzer" or "non-point-of-care device" refers to a device that is not a point-of-care device or a disposable device. A higher throughput assay analyzer or non-point-of-care device refers to any device that does not meet any of the limitations of a point-of-care or disposable device as defined herein. In some embodiments, a "higher throughput assay analyzer" or "non-point-of-care device" can be a device that is (a) a relatively large device compared to a handheld point-of-care device, such as a device ranging in size from that of a benchtop device (e.g., typically considered low throughput or medium throughput) to a large room-sized or multiple room-sized device (e.g., typically considered high throughput), (b) not a handheld device, (c) capable of performing assays simultaneously on more than one clinical sample, and (d) any combination of (a)-(c). A higher throughput assay analyzer can be a clinical chemistry analyzer, an immunoassay analyzer, or a combination thereof. Exemplary higher throughput assay analyzers or non-point-of-care devices include, for example, the ARCHITECT or Alinity platforms manufactured by Abbott Laboratories.
[0047] A "point-of-care device" refers to a device used to provide medical diagnostic tests in or near the point-of-care situation (i.e., usually outside the laboratory), at the time and place of patient care (e.g., in a hospital, clinic, emergency or other healthcare facility, patient's home, nursing home and / or long-term care or hospice facility). Examples of point-of-care devices include those manufactured by Abbott Laboratories (Abbott Park, IL) (e.g., i-STAT and i-STAT Alinity), Universal Biosensors (Lower Hutt, New Zealand) (see US2006 / 0134713), Axis-Shield PoC AS (Oslo, Norway), and Clinical Lab Products (Los Angeles, USA).
[0048] A "reference score" as used herein refers to a value used to evaluate diagnosis, prognosis or treatment effectiveness and is associated with or related to various clinical parameters herein (e.g., presence of a disease (such as malignant vs. non-malignant), stage of the disease, severity of the disease, progression, non-progression or improvement of the disease, etc.).
[0049] As used interchangeably herein, "subject" and "patient" refer to any vertebrate including, but not limited to, mammals (e.g., cows, pigs, camels, llamas, horses, goats, rabbits, sheep, hamsters, guinea pigs, cats, dogs, rats, and mice, non-human primates (e.g., monkeys, e.g., cynomolgus or rhesus monkeys, chimpanzees, etc.) and humans). In some embodiments, the subject can be human or non-human. In some embodiments, the subject is human. In some embodiments, the subject is biologically male, female, or otherwise. In some embodiments, the subject is identified ethnically / racially, alone or in combination as American Indian or Alaska Native, Asian, Black or African American, Native Hawaiian or Other Pacific Islander, White, or otherwise. The subject or patient can have a tumor(s) that is benign, malignant, or a combination thereof. The subject or patient may be undergoing other forms of treatment.
[0050] "Treat", "treating", or "treatment" are each used interchangeably herein and are described as reversing, alleviating, or inhibiting the progression of a disease and / or injury or one or more symptoms of such a disease. Depending on the state of the subject, the term also refers to preventing a disease, including preventing the onset of the disease or preventing the symptoms associated with the disease. Treatment can be carried out in either an acute or chronic manner. The term also refers to reducing the severity of a disease or the symptoms associated with such a disease prior to contracting the disease. Such prevention or reduction of the severity of a disease prior to contracting refers to the administration of a pharmaceutical composition to a subject who does not have the disease at the time of administration. "Preventing" also refers to preventing the recurrence of a disease or one or more symptoms associated with such a disease. "Treatment" and "therapeutically" refer to the act of treating as "treating" is defined above.
[0051] Unless otherwise noted herein, scientific and technical terms used in connection with the present disclosure shall have the meanings as commonly understood by those of ordinary skill in the art. For example, any nomenclature and techniques used in connection with cell and tissue culture, molecular biology, immunology, microbiology, genetics, and protein and nucleic acid chemistry and hybridization described in this specification are well-known and commonly used in the art. The meanings and scopes of the terms should be clear, but in case of any potential ambiguity, the definitions provided in this document shall prevail over any dictionary or extrinsic definition. Further, unless the context or other circumstances require otherwise, singular terms shall include the plural, and plural terms shall include the singular.
[0052] 2. Method and apparatus for determining whether at least one indeterminate pulmonary nodule identified in a subject is malignant In one embodiment, the present disclosure relates to a method and system for determining whether at least one indeterminate pulmonary nodule (IPN) identified in a subject is likely to be malignant or unlikely to be malignant. Indeterminate pulmonary nodules can be identified in a subject, particularly a high-risk subject (e.g., a clearly healthy patient aged 55 to 75 years with a smoking history of at least 30 pack-years and currently smoking or having quit smoking within the past 15 years), using routine techniques known in the art such as low-dose helical computed tomography (LDCT) screening (Smith R, Andrews K, Brooks D et al., Cancer Screening in the United States, 2018: A Review of Current American Cancer Society Guidelines and Current Issues in Cancer Screening, CA Cancer J Clin 68, 4 (2018) 297-316).
[0053] In one aspect, the methods and systems of the present disclosure include obtaining certain target values and assay values for a subject for which at least one IPN has been identified. Target values to be obtained include (a) the smoking pack-year value of the subject and (b) a measured value of the size of the IPN in the subject. Assay values to be obtained include (a) the concentration of cancer antigen 125 (CA125), carcinoembryonic antigen (CEA), human epididymis protein 4 (HE4), cytokeratin fragment 21-1 (Cyfra 21-1), neuron-specific enolase (NSE), squamous cell carcinoma antigen (SCC), pro-gastrin releasing peptide (ProGRP) or at least one of any combination thereof from a biological sample obtained from the subject and (b) the total IgG concentration, IgA concentration, IgM concentration, IgE concentration, free kappa light chain concentration, free lambda light chain concentration or at least one of any combination thereof from a biological sample obtained from the subject.
[0054] Biological samples obtained from the subject for determining the concentrations of (i) one or more of CA125, CEA, HE4, Cyfra 21-1, NSE, SCC, ProGRP or any combination thereof and (ii) one or more of total IgG, IgA, IgM, IgE, free kappa light chain, free lambda light chain or any combination thereof can be obtained using techniques known to those of ordinary skill in the art and the samples can be used directly as obtained from the subject or after pretreatment to alter the characteristics of the sample. Such pretreatment can include, for example, preparation of plasma from blood, dilution of viscous fluids, filtration, precipitation, dilution, distillation, mixing, concentration, inactivation of interfering components, addition of reagents, lysis, etc. In some aspects, the biological sample is a whole blood sample, serum sample, plasma sample or capillary blood sample. In other aspects, the same or different biological samples can be used to determine the concentrations of (i) one or more of CA125, CEA, HE4, Cyfra 21-1, NSe, SCC, ProGRP or any combination thereof and (ii) one or more of total IgG, IgA, IgM, IgE, free kappa light chain, free lambda light chain or any combination thereof in the subject.
[0055] The source of the target value and the assay value is not important. For example, the target value and the assay value can be obtained from an analytical testing system, a handheld or point-of-care testing device, a high-throughput analyzer, or any combination thereof, in a clinic, hospital or other medical facility, laboratory, or decentralized setting.
[0056] The target pack-year value can be obtained by determining the number of tobacco packs the subject smoked in a year and then multiplying that number by the number of years the subject smoked. For example, if a subject smoked 2 packs per day for 20 years, the subject's pack-year value would be 40.
[0057] The measured value of the size of the IPN in a subject can be determined using routine techniques known in the art. For example, LDCT can be used to detect and measure the size of the IPN identified in a subject. It is known that larger IPN nodule sizes correlate with a higher risk that the IPN is cancerous: IPN < 5 mm: 0% - 1%; 5 - 10 mm: 6% - 28%; 11 - 20 mm: 37% - 64%; > 20 mm: 64% - 82% (Maission P, Walker R., Indeterminate Pulmonary Nodules: Risk for Having or for Developing Lung Cancer, Cancer Prev Res (Phila) 7, 12 (2014) 1173 - 1178).
[0058] As mentioned previously, the method and system also require obtaining, providing, and / or determining the concentration of one or more of CA125, CEA, HE4, Cyfra 21-1, NSE, SCC, or ProGRP, or any combination thereof, in a biological sample obtained from a subject. In another aspect, the method and system also require obtaining, providing, and / or determining the concentration of two or more of CA125, CEA, HE4, Cyfra 21-1, NSE, SCC, ProGRP, or any combination thereof, in a biological sample obtained from a subject. In another aspect, the method and system also require obtaining, providing, and / or determining the concentration of three or more of CA125, CEA, HE4, Cyfra 21-1, NSE, SCC, ProGRP, or any combination thereof, in a biological sample obtained from a subject. In another aspect, the method and system also require obtaining, providing, and / or determining the concentration of four or more of CA125, CEA, HE4, Cyfra 21-1, NSE, SCC, ProGRP, or any combination thereof, in a biological sample obtained from a subject. In another aspect, the method and system also require obtaining, providing, and / or determining the concentration of five or more of CA125, CEA, HE4, Cyfra 21-1, NSE, SCC, ProGRP, or any combination thereof, in a biological sample obtained from a subject. In another aspect, the method and system also require obtaining, providing, and / or determining the concentration of six or more of CA125, CEA, HE4, Cyfra 21-1, NSE, SCC, and ProGRP, or any combination thereof, in a biological sample obtained from a subject. In another aspect, the method and system also require obtaining, providing, and / or determining the concentration of CA125, CEA, HE4, Cyfra 21-1, and NSE, SCC, and ProGRP, in a biological sample obtained from a subject.
[0059] Any assay known in the art for determining the concentration of CA125, CEA, HE4, Cyfra 21-1, NSE, SCC, ProGRP, or any combination thereof in a biological sample can be provided and / or used in the methods and systems of the present disclosure. For example, immunoassays, clinical chemistry, radioimmunoassays, immunoradiometric assays, etc. can be used or provided in the methods and systems of the present disclosure. In some embodiments, the concentration of at least one of CA125, CEA, HE4, Cyfra 21-1, NSE, SCC, ProGRP, or any combination thereof in a biological sample obtained from a subject can be determined using a high-throughput analyzer or a point-of-care device. For example, the Abbott Laboratories CA125 chemiluminescent assay for use on the ARCHITECT® i2000 automated immunoassay platform can be used in the methods and systems described herein.
[0060] The method and system further require determining at least one concentration of the total IgG concentration, IgA concentration, IgM concentration, IgE concentration, LFLC concentration of the subject, or any combination thereof, in a biological sample obtained from the subject. In some embodiments, the method and system further require determining at least two concentrations of the total IgG concentration, IgA concentration, IgM concentration, IgE concentration, LFLC concentration of the subject, or any combination thereof. In further embodiments, the method and system further require determining at least three concentrations of the total IgG concentration, IgA concentration, IgM concentration, IgE concentration, LFLC concentration of the subject, or any combination thereof. In further embodiments, the method and system further require determining at least four concentrations of the total IgG concentration, IgA concentration, IgM concentration, IgE concentration, LFLC concentration of the subject, or any combination thereof. In further embodiments, the methods and systems of the present disclosure require determining the total IgG concentration, IgA concentration, IgM concentration, IgE concentration of the subject, and the kappa free light chain (KFLC) and lambda free light chain (LFLC) concentrations of the subject. Any assay known in the art for determining the total IgG concentration, IgA concentration, IgM concentration, IgE concentration, LFLC concentration of the subject, or any combination thereof can be provided and / or used in the methods and systems of the present disclosure. For example, immunoassays, clinical chemistry, radioimmunoassays, immunoradiometric assays can be used. In some embodiments, the total IgG concentration, IgA concentration, IgM concentration, IgE concentration, LFLC concentration of the subject, or any combination thereof, in a biological sample obtained from the subject can be determined using a high throughput analyzer or a point-of-care device.For example, using the Abbott Laboratories ARCHITECT c8000 automated clinical chemistry platform utilizing a turbidimetric assay format (either using antiserum or latex enhanced antibody coated particles), the total IgG concentration of a subject, the IgA concentration of the subject, the IgM concentration of the subject, the IgE concentration of the subject, the LFLC concentration of the subject, or any combination thereof can be determined in the methods and systems described herein.
[0061] Once the subject value and the assay value are obtained, the values are input and processed by a processing system. The processing system can include a computer processor and a non-transitory computer memory including one or more computer programs and databases. In some embodiments, the subject value, the assay value, and the subject value and the assay value are manually input into the processing system. In other embodiments, the subject value, the assay value, or the subject value and the assay value are automatically input into the processing system. In further embodiments, the subject value, the assay value, or the subject value and the assay value are received electronically, such as via email. In further embodiments, the processing system further includes a handheld or point-of-care testing device. In further embodiments, the processing system further includes a high-throughput analyzer.
[0062] In other aspects, at least one of the computer programs embodied in the processing system is one or more machine learning algorithms. Any machine learning algorithm can be used in the methods and systems of the present disclosure. In some aspects, at least one machine learning algorithm is an Adaptive Index Modeling (AIM) algorithm. The processing system can contain any AIM algorithm known in the art. For example, the AIM algorithm described in Tian L, Tibshirani R, Adaptive index models for marker-based risk stratification, Biostatistics 12, 1 (2011), pp. 68-86 and AIM: AIM: Adaptive index model. R package version 1.01 can be used. In other aspects, at least one machine learning algorithm is a random forest algorithm. Any random forest algorithm known in the art can be used. In yet other aspects, at least one machine learning algorithm is a logistic regression algorithm. Any machine learning algorithm known in the art can be used. In some aspects, at least one machine learning algorithm uses the AIM algorithm, the random forest algorithm, or any combination of the random forest algorithms.
[0063] In some embodiments, the processing system is configured to apply at least one machine learning algorithm to the assay value and the target value to create, generate, or output a score for the target (e.g., a machine learning score such as an AIM score, a random forest algorithm score, a logistic regression algorithm score, or any combination thereof). In some embodiments, the processing system is further configured to communicate (e.g., report) the machine learning score, and the machine learning score is communicated (e.g., reported) for further analysis, interpretation, processing, and / or display. The machine learning score for the target can be communicated (e.g., reported) by the processing system (e.g., a computer) in a document and / or spreadsheet, on a mobile device (e.g., a smartphone), on a website, in an email, or any combination thereof.
[0064] In yet other aspects, the machine learning score for a subject is compared to a reference score. In some aspects, a clinician or other healthcare provider can compare the machine learning score for a subject to the reference score. The reference score can be provided in a product insert or other publication, or on a website, or on a mobile device (e.g., via an app, etc.). In other aspects, the processing system is configured to provide a reference score for comparison to the machine learning algorithm score. The reference score can be determined using conventional techniques known in the art. For example, the reference score can be determined by a machine learning algorithm in the processing system. In some aspects, the reference score is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100. In some aspects, the reference score is 1. In other aspects, the reference score is 2. In further aspects, the reference score is 3. In further aspects, the reference score is 4. In further aspects, the reference score is 5. In further aspects, the reference score is 6. In further aspects, the reference score is 7. In further aspects, the reference score is 8. In further aspects, the reference score is 9. In further aspects, the reference score is 10. In further aspects, the reference score is 11. In further aspects, the reference score is 12. In further aspects, the reference score is 13. In further aspects, the reference score is 14. In further aspects, the reference score is 15. In further aspects, the reference score is 16. In further aspects, the reference score is 17. In further aspects, the reference score is 18. In further aspects, the reference score is 19.In a further aspect, the reference score is 20. In a further aspect, the reference score is 21. In a further aspect, the reference score is 22. In a further aspect, the reference score is 23. In a further aspect, the reference score is 24. In a further aspect, the reference score is 25. In a further aspect, the reference score is 26. In a further aspect, the reference score is 27. In a further aspect, the reference score is 28. In a further aspect, the reference score is 29. In a further aspect, the reference score is 30. In a further aspect, the reference score is 31. In a further aspect, the reference score is 32. In a further aspect, the reference score is 33. In a further aspect, the reference score is 34. In a further aspect, the reference score is 35. In a further aspect, the reference score is 36. In a further aspect, the reference score is 37. In a further aspect, the reference score is 38. In a further aspect, the reference score is 39. In a further aspect, the reference score is 40. In a further aspect, the reference score is 41. In a further aspect, the reference score is 42. In a further aspect, the reference score is 43. In a further aspect, the reference score is 44. In a further aspect, the reference score is 45. In a further aspect, the reference score is 46. In a further aspect, the reference score is 47. In a further aspect, the reference score is 48. In a further aspect, the reference score is 49. In a further aspect, the reference score is 50. In a further aspect, the reference score is 51. In a further aspect, the reference score is 52. In a further aspect, the reference score is 53. In a further aspect, the reference score is 54. In a further aspect, the reference score is 55. In a further aspect, the reference score is 56. In a further aspect, the reference score is 57. In a further aspect, the reference score is 58. In a further aspect, the reference score is 59. In a further aspect, the reference score is 60. In a further aspect, the reference score is 61. In a further aspect, the reference score is 62. In a further aspect, the reference score is 63. In a further aspect, the reference score is 64. In a further aspect, the reference score is 65. In a further aspect, the reference score is 66.In a further aspect, the reference score is 67. In a further aspect, the reference score is 68. In a further aspect, the reference score is 69. In a further aspect, the reference score is 70. In a further aspect, the reference score is 71. In a further aspect, the reference score is 72. In a further aspect, the reference score is 73. In a further aspect, the reference score is 74. In a further aspect, the reference score is 75. In a further aspect, the reference score is 76. In a further aspect, the reference score is 77. In a further aspect, the reference score is 78. In a further aspect, the reference score is 79. In a further aspect, the reference score is 80. In a further aspect, the reference score is 81. In a further aspect, the reference score is 82. In a further aspect, the reference score is 83. In a further aspect, the reference score is 84. In a further aspect, the reference score is 85. In a further aspect, the reference score is 86. In a further aspect, the reference score is 87. In a further aspect, the reference score is 88. In a further aspect, the reference score is 89. In a further aspect, the reference score is 90. In a further aspect, the reference score is 91. In a further aspect, the reference score is 92. In a further aspect, the reference score is 93. In a further aspect, the reference score is 94. In a further aspect, the reference score is 95. In a further aspect, the reference score is 96. In a further aspect, the reference score is 97. In a further aspect, the reference score is 98. In a further aspect, the reference score is 99. In a further aspect, the reference score is 100.
[0065] Based on the comparison between the target machine learning score and the reference score, a determination is made as to whether the IPN is likely to be malignant or not likely to be malignant. Specifically, when the target machine learning score is higher than the reference score, it is determined that the IPN in the target is likely to be malignant. When the target machine learning score is the same as or lower than the reference score, it is determined that the IPN in the target is less likely to be malignant. In some embodiments, the determination of whether the target IPN is likely to be malignant or less likely to be malignant can be communicated (e.g., reported) for further display. Specifically, this determination of whether the target IPN is likely to be malignant or not likely to be malignant can be communicated (e.g., reported) by a processing system (e.g., a computer) in a document and / or spreadsheet, on a mobile device (e.g., a smartphone), on a website, in an email, or in any combination thereof.
[0066] Subjects identified as having an IPN that is likely to be malignant based on the methods and systems described herein may be treated, monitored, or both treated and monitored. In some embodiments, a surgical or non-surgical biopsy can be performed to further evaluate an IPN identified as likely to be malignant. In other embodiments, a portion of the lung containing the IPN can be surgically removed or excised from the subject using techniques known in the art such as, for example, video-assisted thoracic surgery or thoracotomy. In yet other embodiments, the subject may receive one or more pharmaceutical or biopharmaceutical treatments. For example, in some embodiments, the subject can be treated with chemotherapy, radiation, budesonide, fluticasone, or any combination thereof. In further embodiments, the subject can also be monitored. For example, the IPN in the subject can be monitored using one or more of computed tomography (including LDCT) scans, positron emission tomography (PET) scans, bronchoscopy, or any combination thereof. In some embodiments, the subject can be monitored before, during, and / or after any biopsy and / or treatment.
[0067] Other suitable modifications and adaptations of the methods of the present disclosure described herein are readily applicable and recognizable and can be made using suitable equivalents without departing from the scope of the present disclosure or the aspects and embodiments disclosed herein, which will be readily apparent to those skilled in the art. Having described the present disclosure in detail, the present disclosure will be more clearly understood by reference to the following examples, which are intended to merely illustrate some aspects and embodiments of the present disclosure and should not be regarded as limiting the scope of the present disclosure. All academic journal references, U.S. patents and publications disclosures mentioned herein are incorporated herein by reference in their entirety.
Example
[0068] Materials and Methods 1. Study Population The cohort of patient samples examined consisted of a total of 141 patients with the following conditions: 36 patients with benign indeterminate pulmonary nodules (median nodule size 14.9 mm) and 105 patients with malignant indeterminate pulmonary nodules (median nodule size 19 mm). Approximately 74.5% of the examined cohort was classified as having stage I NSCLC lung cancer. Clinical and pathological details regarding malignant cases were obtained from the medical record system. The criteria for inclusion in the study in the malignant NSCLC cohort were broad (consisting of surgical resection with lymph node sampling and accompanying pathological examination) and were not limited by any demographic or clinical factors. The benign cohort with malignant pulmonary nodules consisted of patients with granulomas, pneumonitis or pneumonia. These patients underwent an anatomical resection for suspected malignancy. The benign and malignant samples collected represent real clinical collections to which histological selection criteria were not applied.
[0069] The demographic variable of pack-years of smoking was defined by multiplying the number of tobacco packs smoked per year by the number of years the individual smoked. For this pilot study, pack-years of smoking was calculated for any patient case who underwent a CT scan based on risk assessment or as an incidental finding.
[0070] 2. Sample Collection, Handling, and Storage Specimens were obtained with full written informed consent under a protocol approved by the Rush UMC Institutional Review Board (IRB). Peripheral blood collected at Rush UMC was obtained from each patient immediately prior to the start of treatment using standard venipuncture techniques. The start of treatment for IPN could be surgical removal, biopsy, or further radiographic evaluation. All specimens were handled in the same manner and processed into EDTA plasma. The time interval from sample collection to processing was less than 90 minutes. All EDTA vacutainer collection tubes were centrifuged at 750 RCF for 20 minutes to isolate the plasma layer. The subsequent plasma layer was transferred to a second tube and recentrifuged to remove particulates. The specimens after recentrifugation were aliquoted into 0.75 mL aliquots, stored immediately, and frozen at -80 degrees Celsius. For this study, specimens were not subjected to more than two thaw cycles. Samples were coded with the basic demographic and clinical parameters provided to the investigator for the purposes of this study.
[0071] All samples were collected, and the EDTA plasma levels of CA-125, SCC, CEA, HE4, ProGRP, NSE, Cyfra 21-1, and ferritin were determined using the Abbott ARCHITECT® i2000 automated immunoassay platform with a two-step dual monoclonal chemiluminescent immunoassay (Table 1; see Quinn, F.A., The Immunoassay Handbook, 3rd edition, Wild, D; Elsevier Ltd.: United Kingdom, 2005. Chapter 34: ARCHITECT i2000 and i2000SR Analyzers). Additionally, hs-CRP, total IgG, IgG1, IgG2, IgG3, IgG4, IgE, IgM, IgA, free kappa light chain, and free lambda light chain were determined using the Abbott ARCHITECT® c8000 automated clinical chemistry platform with a nephelometric assay format (using either antiserum or latex-enhanced antibody-coated microparticles) (Table 1; see Clinical Chemistry Learning Guide Series 2020. [https: / / www.corelaboratory.abbott / sal / learningGuide / ADD-00061345_ClinChem_Learning_Guide.pdf (corelaboratory.abbott)], 2020). The IA and CC assay determinations were performed at Abbott Laboratories (IL, USA). The test results were included in the database along with demographic data and other clinical parameter data.
[0072]
Table 1
[0073] Assays for IgG1, IgG2, IgG3, and IgG4 for research use only (RUO) in clinical chemistry were independently validated against the commercially available Abbott Total IgG assay. The sum of the IgG1-IgG4 test results is comparable to the total IgG results with a Passing-Bablok slope of 0.97 and a correlation coefficient of 0.96.
[0074] 3. Statistical methods Patient samples were stratified based on the IPN nodule category. Descriptive analyses were performed to show the distribution of demographic and clinical variables by IPN nodule category. For variables that were approximately normally distributed, the mean and standard deviation were reported, and for variables that were not normally distributed, the median was reported along with the minimum and maximum values. The significance level for this study was set at α = 0.05. For IPN nodule category prediction, the individual profiles of biomarkers were examined using distribution plots, receiver operating characteristic (ROC) curves, and the area under the curve (AUC). Also, the Wilcoxon rank sum test was used to compare whether there was a difference in the distribution of biomarkers between the benign and nodule categories. In the subsequent multivariate analysis, single imputation of the median was performed for variables containing missing values: CEA (9.93% missing in the total population), pack-years of smoking (3.55%), lambda free light chain (2.13%), IgE (1.42%), IgG (0.71%), and IgA (0.71%).
[0075] To examine the multivariate relationship between biomarkers and IPN nodule categories, Adaptive Index Modeling (AIM) was applied (Tian L, Tibshirani R, Adaptive index models for marker-based risk stratification, Biostatistics 12, 1 (2011), pp. 68-86; AIM: AIM: Adaptive Index Model. R package version 1.01). This method uses the concept of an index predictor defined as a binary rule based on the value of a predictor variable, e.g., whether a patient's age is 55 years or older or less than 55 years. After providing a set of variables as potential predictors, AIM adaptively searches for individual cutoffs for each variable to build an overall model. The index predictors are selected by maximizing the score test statistic up to a pre-specified number of total predictors (Tian L, Tibshirani R, Adaptive index models for marker-based risk stratification, Biostatistics 12, 1 (2011), pp. 68-86; see, e.g., Table 2). In this study, the maximum number of index predictors was set to 8 to avoid having models that are unlikely to be implemented and clinically adopted based on the cost and / or complexity of the algorithm. After the adaptive selection process, 5-fold cross-validation (due to the small sample size) was used to select the model with the optimal number of index predictors.
[0076] Once the optimal AIM model is selected, each individual subject can be scored according to the values of these index predictors for the binary rule cut-off. Table 2 shows an example of the scoring process for AIM. Each subject has a score from 0 to n, where n represents the total number of index predictors. Subsequently, model performance can be evaluated by creating a binary outcome variable based on score cut-offs of >0, >1, ..., >n-1. Performance metrics were evaluated for the entire dataset, including AUC, accuracy, sensitivity, specificity, positive predictive value, and negative predictive value, across this range of possible cut-offs. The positive predictive value and negative predictive value were calculated based on the prevalence of the disease in the study (74.5% malignancy). The final score cut-off for each AIM model was selected by choosing the maximum AUC that provides an average balance of sensitivity and specificity.
[0077] In this study, four possible models were developed using AIM.
[0078] 1. Demographic variables only: age + gender + smoking pack-year + nodule size 2. ARCHITECT immunoassay biomarkers 3. ARCHITECT clinical chemistry biomarkers 4. ARCHITECT immunoassay biomarkers + clinical chemistry biomarkers + demographic variables
[0079] In the absence of a baseline model with clinical features such as the Mayo clinical model for IPN prediction, models using only demographic variables were treated as baseline models to determine whether biomarkers provide additional predictive value (The R Project for Statistical Computing. [https: / / www.R-project.org / ]). The relative classification of patients by the status of IPN nodules was shown using a bar graph, and the predictions of the optimal AIM model versus the baseline model were compared. All statistical analyses were performed using R version 4.0.5 (The R Project for Statistical Computing; Yang B, Jhun BW, Shin SH, Jeong B-H, Um S-W, Zo JI et al., 2018. Comparison of four models predicting the malignancy of pulmonary nodules: A single-center study of Korean adults, PLoS ONE 13, 7, e0201242).
[0080] In Table 2, the score of the Adaptive Index Model (AIM) is equivalent to the number of criteria met, and in the following example, the score is 3. The higher the AIM score, the higher the risk of the model predicting a medical condition such as lung cancer.
[0081]
Table 2
[0082] Description: The bold text meets the AIM cut-off criteria The non-bold text does not meet the AIM cut-off criteria
[0083] Results In this study, to better understand the clinical factors and biomarker distributions across various individuals with pulmonary nodules requiring anatomical resection, the inclusion criteria for the patient population were not limited to the "high-risk" population as defined by the NLST study. There were several cases where IPNs were identified as incidental findings on radiographs. None of these cases had annotations indicating symptoms in these electronic records at the time of sample collection. Based on p-values less than 0.05, the demographic variables of age, nodule size, race, and pack-years of smoking were statistically significant between benign and malignant IPN nodules (see Table 3).
[0084]
Table 3
[0085] * For continuous variables, the t-test was used to compare groups in the case of a normal distribution, and the Mann-Whitney test was used in the case of a non-normal distribution. For categorical variables, the Pearson chi-square test without continuity correction was used when all expected cell counts were greater than 5, and Fisher's exact probability test was used in other cases.
[0086] All blood-based biomarkers were individually evaluated for their effectiveness in risk stratification of this indeterminate risk pulmonary nodule population. Plasma values for CA-125, SCC, CEA, HE4, ProGRP, NSE, Cyfra 21-1, hs-CRP, ferritin, total IgG, IgG1, IgG2, IgG3, IgG4, IgE, IgM, IgA, kappa free light chain, and lambda free light chain are summarized in Table 4 along with their calculated AUCs for each. Individually, none of these biomarkers would be a compelling case for risk stratification as estimated by their AUC values (range 0.433 - 0.594). As a result, the biomarkers were combined using the AIM methodology to determine whether any modeled combination could be useful in the practice of risk stratification, with or without demographic / clinical data.
[0087]
Table 4
[0088] Table 5 shows the results of implementing the AIM statistical methodology for four possible combinations of biomarkers and demographic variables. Of the models, the one with the best performance was the model that included the ARCHITECT immunoassay and clinical chemistry biomarkers and demographic variables (Smith R, Andrews K, Brooks D et al., Cancer Screening in the United States, 2018: A Review of Current American Cancer Society Guidelines and Current Issues in Cancer Screening, CA Cancer J Clin 68, 4 (2018) 297–316). This model was composed of IgG, IgM, IgE, IgA, lambda free light chain, CA-125, pack-years of smoking, and nodule size, and resulted in an AUC of 0.819 (95% CI 0.730–0.899) as well as a sensitivity of 0.971 and a specificity of 0.667. The AIM model (Smith R, Andrews K, Brooks D et al., Cancer Screening in the United States, 2018: A Review of Current American Cancer Society Guidelines and Current Issues in Cancer Screening, CA Cancer J Clin 68, 4 (2018) 297–316) showed a statistically significant improvement over individual biomarker predictions for the IPN nodule category (CI range 0.318–0.710).The AIM model (Smith R, Andrews K, Brooks D, et al., Cancer Screening in the United States, 2018: A Review of Current American Cancer Society Guidelines and Current Issues in Cancer Screening, CA Cancer J Clin 68, 4 (2018) 297-316) also had a relative improvement in AUC compared to the AIM model using only demographic variables (Sung H, Ferlay J, Siegel R, et al., Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries, CA Cancer J Clin 71 (2021) 209-249) (AUC 0.699, 95% CI 0.614-0.784). Figure 1 shows the relative classification of the AIM models (Sung H, Ferlay J, Siegel R, et al., Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries, CA Cancer J Clin 71 (2021) 209-249) and (Smith R, Andrews K, Brooks D, et al., Cancer Screening in the United States, 2018: A Review of Current American Cancer Society Guidelines and Current Issues in Cancer Screening, CA Cancer J Clin 68, 4 (2018) 297-316) by IPN nodule category.The AIM model (Smith R, Andrews K, Brooks D et al., Cancer Screening in the United States, 2018: A Review of Current American Cancer Society Guidelines and Current Issues in Cancer Screening, CA Cancer J Clin 68, 4 (2018) 297 - 316) accurately classified 102 / 105 malignant samples and 24 / 36 benign samples, while the AIM model (Sung H, Ferlay J, Siegel R et al., Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries, CA Cancer J Clin 71 (2021) 209 - 249) accurately classified 68 / 105 malignant samples and 27 / 36 benign samples.
[0089]
Table 5
[0090] Results The addition of the exploratory blood-based biomarkers described in this example, which were narrowed down to the important benign-versus-malignant tumor discriminative biomarkers CA125, IgG, IgM, IgA, IgE, and lambda free light chain in conjunction with smoking pack-years and nodule size, demonstrated an improvement over demographics / nodule size data alone in helping to assess the malignancy risk for IPN nodules (see Tables 5 and 1). The developed AIM algorithm can potentially be used to stratify the IPN risk of malignancy as follows: an AIM score above a cutoff of 4 has a high likelihood of the IPN being malignant, whereas an AIM score at or below the cutoff has a low likelihood of the IPN being malignant. For example, to confirm the results of a high AIM score, a suspicious IPN can be further imaged with contrast MRI to assess for spiculation (spiky extensions from the nodule), which can be a high-risk indicator of cancer. If an IPN is determined to be high risk based on the AIM score and MRI imaging, the nodule may potentially be biopsied, and the resulting pathology may require aggressive treatment such as surgical removal of the IPN, radiation, and / or chemotherapy. Alternatively, IPNs presenting with low AIM scores by the current algorithm and confirmed imaging consistent with smooth calcified nodules generally have a low risk of cancer.
[0091] This study also took up several additional items of interest using immune-related clinical chemistry assays. The immune biomarkers tested were not specific to lung cancer, but in this sample set, the immune biomarkers tested appeared to be highly sensitive to the biological changes associated with cancer. The concentrations of IgG, IgG1, IgG2, IgG4, and IgE appeared to be suppressed in cancer patients compared to patients with benign nodules, potentially indicating downregulation of the immune system (see Table 4). This opens the possibility that the immune system profile may be useful in predicting benign disease.
[0092] It should be understood that the foregoing detailed description and accompanying examples are merely illustrative and are not to be construed as limiting the scope of the disclosure, which is defined only by the appended claims and their equivalents.
[0093] Various changes and modifications to the disclosed embodiments will be apparent to those skilled in the art. Such changes and modifications, including but not limited to those related to the chemical structures, substituents, derivatives, intermediates, syntheses, compositions, formulations, or methods of use of the present disclosure, can be made without departing from the spirit and scope of the present disclosure.
[0094] For reasons of completeness, the various aspects of the present disclosure are described in the numbered clauses below: Clause 1. A method for determining whether at least one indeterminate pulmonary nodule (IPN) identified in a subject is likely to be malignant, comprising: a) providing a subject value for the subject, the subject value comprising at least one of the following: i) the subject's smoking pack-year value, ii) identification of biological sex, iii) identification of race, iv) identification of nodule type, and v) a measured value of the size of the IPN in the subject; b) providing at least two assay values, the at least two assay values comprising: i) at least one of the subject's cancer antigen 125 (CA125) concentration, the subject's carcinoembryonic antigen (CEA) concentration, the subject's human epididymis protein 4 (HE4) concentration, the subject's cytokeratin fragment 21-1 (Cyfra 21-1) concentration, the subject's neuron-specific enolase (NSE) concentration, the subject's squamous cell carcinoma antigen (SCC) concentration, the subject's progastrin-releasing peptide (ProGRP) concentration, or any combination thereof, obtained from a biological sample from the subject, and ii) at least one of the total IgG concentration of the subject, the IgA concentration of the subject, the IgM concentration of the subject, the IgE concentration of the subject, the kappa free light chain concentration, the lambda free light chain concentration, or any combination thereof, derived from a biological sample obtained from the subject a step comprising; c) providing a processing system comprising a computer processor and a non - transitory computer memory including a database and at least one machine - learning algorithm, wherein the at least one machine - learning algorithm is configured to process target values and assay values, and further, the processing system is i) configured to apply the at least one machine - learning algorithm to the assay values and the target values to output a machine - learning score for the subject, ii) configured to report the machine - learning score for the subject, and iii) configured to provide a reference score for comparison with the machine - learning score a step; and d) determining that the IPN is likely to be malignant if the machine - learning score is higher than the reference score and is likely to be non - malignant if the machine - learning score is the same as or lower than the reference score a method comprising.
[0095] Clause 2. The method of Clause 1, wherein the smoking pack - year value of the subject is the number of tobacco packs smoked by the subject per year multiplied by the number of years the subject has smoked.
[0096] Clause 3. The method of Clause 1 or Clause 2, comprising obtaining assay values including the CA125 concentration of the subject, the total IgG concentration of the subject, the IgA concentration of the subject, the IgM concentration of the subject, the IgE concentration of the subject, the kappa free light chain concentration, and the lambda free light chain concentration, derived from a biological sample obtained from the subject.
[0097] Article 4. The method according to any one of Articles 1 to 3, wherein the reference score is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100.
[0098] Article 5. The method according to any one of Articles 1 to 4, wherein the biological sample is a whole blood sample, a serum sample or a plasma sample.
[0099] Article 6. The method according to any one of Articles 1 to 5, wherein the step of providing the target value, the assay value, or the target value and the assay value includes receiving the target value, the assay value, or the target value and the assay value from a laboratory, from the subject, from an analytical testing system, from a handheld or point-of-care testing device, or from any combination thereof.
[0100] Article 7. The method according to any one of Articles 1 to 6, wherein the step of providing the target value, the assay value, or the target value and the assay value includes electronically receiving the target value.
[0101] Article 8. The method according to any one of Articles 1 to 7, further comprising manually or automatically inputting the target value, the assay value, or the target value and the assay value into the processing system.
[0102] Article 9. The method according to any one of Articles 1 to 8, wherein the processing system compares a machine learning score regarding the subject with the reference score.
[0103] Article 10. The method of Article 9, wherein a determination of whether the IPN is likely to be malignant or not likely to be malignant is displayed on the device.
[0104] Article 11. The method of any one of Articles 1 to 10, wherein the subject is a human.
[0105] Article 12. The method of any one of Articles 1 to 11, wherein at least one of the subject's CA125 concentration, the subject's CEA concentration, the subject's HE4 concentration, the subject's Cyfra 21-1 concentration, the subject's NSE concentration, the subject's SCC concentration, the subject's ProGRP concentration, or any combination thereof is determined using an immunoassay.
[0106] Article 13. The method of any one of Articles 1 to 9, wherein at least one of the subject's total IgG concentration, the subject's IgA concentration, the subject's IgM concentration, the subject's IgE concentration, the subject's free kappa light chain concentration, the subject's free lambda light chain concentration, or any combination thereof is determined using a clinical chemistry assay.
[0107] Article 14. a. A target value for the subject, i) the subject's pack-year smoking value, and ii) a measured value of the size of the IPN in the subject comprising the target value; b. i) at least one of the subject's cancer antigen 125 (CA125) concentration, the subject's carcinoembryonic antigen (CEA) concentration, the subject's human epididymis protein 4 (HE4) concentration, the subject's cytokeratin fragment 21-1 (Cyfra 21-1) concentration, the subject's neuron-specific enolase (NSE) concentration, the subject's squamous cell carcinoma antigen (SCC) concentration, the subject's progastrin-releasing peptide (ProGRP) concentration, or any combination thereof, derived from a biological sample obtained from the subject, and ii) at least one of the subject's total IgG concentration, the subject's IgA concentration, the subject's IgM concentration, the subject's IgE concentration, free kappa light chain concentration, free lambda light chain concentration, or any combination thereof, derived from a biological sample obtained from the subject for one or more assays for measuring; c. A device including a processing system, the processing system including a computer processor and a non-transitory computer memory including a database and at least one machine learning algorithm, wherein the at least one machine learning algorithm is configured to process a target value and an assay value to generate a machine learning score for the target, and further, the processing system is i) configured to apply the at least one machine learning algorithm to the target value and the assay value to output a machine learning score for the target, ii) configured to report the machine learning score for the target, and iii) configured to provide a reference score for comparison with the machine learning score and the device is configured such that an IPN indicates that (1) when the machine learning score is higher than the reference score, there is a high likelihood of malignancy, or (2) when the machine learning score is the same as or lower than the reference score, there is a low likelihood of malignancy. A system comprising the device.
[0108] Clause 15. The system of Clause 14, wherein the target smoking pack-year value is the number of tobacco packs smoked by the target per year multiplied by the number of years the target has smoked.
[0109] Clause 16. The system of Clause 14 or Clause 15, wherein the assay value includes the target's CA125 concentration, the target's total IgG concentration, the target's IgA concentration, the target's IgM concentration, the target's IgE concentration, the kappa free light chain concentration, and the lambda free light chain concentration, derived from a biological sample obtained from the target.
[0110] Clause 17. A system according to any one of Clauses 14 to 16, wherein the reference score is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100.
[0111] Clause 18. A system according to any one of Clauses 14 to 17, wherein the biological sample is a whole blood sample, a serum sample or a plasma sample.
[0112] Clause 19. A system according to any one of Clauses 14 to 18, wherein the target value, the assay value, or the target value and the assay value are received from a laboratory, from the subject, from an analytical test system, from a handheld or point-of-care testing device, or from any combination thereof.
[0113] Clause 20. A system according to any one of Clauses 14 to 19, wherein the target value, the assay value, or the target value and the assay value are received electronically.
[0114] Clause 21. A system according to any one of Clauses 14 to 20, wherein the target value, the assay value, or the target value and the assay value are manually or automatically input into the processing system.
[0115] Clause 22. A system according to any one of Clauses 14 to 21, wherein the subject is human.
[0116] Clause 23. Any one of the methods of Clauses 14 to 22, wherein at least one of the CA125 concentration of the subject, the CEA concentration of the subject, the HE4 concentration of the subject, the Cyfra 21-1 concentration of the subject, the NSE concentration of the subject, the SCC concentration of the subject, the ProGRP concentration of the subject, or any combination thereof is determined using an immunoassay.
[0117] Clause 24. Any system of Clauses 14 to 23, wherein at least one of the total IgG concentration of the subject, the IgA concentration of the subject, the IgM concentration of the subject, the IgE concentration of the subject, the free kappa light chain concentration of the subject, the free lambda light chain concentration of the subject, or any combination thereof is determined using a clinical chemistry assay.
[0118] Clause 25. a) Providing one or more diagnostic assays configured to measure assay values including (a) at least one of the concentration of cancer antigen 125 (CA125) of the subject, the concentration of carcinoembryonic antigen (CEA) of the subject, the concentration of human epididymis protein 4 (HE4) of the subject, the concentration of cytokeratin fragment 21-1 (Cyfra 21-1) of the subject, the concentration of neuron-specific enolase (NSE) of the subject, the concentration of squamous cell carcinoma antigen (SCC) of the subject, the concentration of progastrin-releasing peptide (ProGRP) of the subject, or any combination thereof, and (b) at least one of the total IgG concentration of the subject, the IgA concentration of the subject, the IgM concentration of the subject, the IgE concentration of the subject, the free kappa light chain concentration, the free lambda light chain concentration, or any combination thereof, derived from a biological sample obtained from the subject; b) Providing a processing system including a computer processor, a database, and a non-transitory computer memory including at least one machine learning algorithm, wherein at least one machine learning algorithm is configured to process assay values and subject values related to the subject, including the subject's smoking pack-year value and the measured value of the size of the IPN in the subject, and the subject's smoking pack-year value and the measured value of the size of the IPN in the subject are pre-entered into the database, and the processing system, i) applying at least one machine learning algorithm to the assay values and the target values to output a machine learning score for the target; ii) reporting the machine learning score for the target; and iii) providing a reference score for comparison with the machine learning score configured steps including a method.
[0119] Clause 26.a) (a) at least one of the concentration of cancer antigen 125 (CA125) of the subject, the concentration of carcinoembryonic antigen (CEA) of the subject, the concentration of human epididymis protein 4 (HE4) of the subject, the concentration of cytokeratin fragment 21-1 (Cyfra 21-1) of the subject, the concentration of neuron-specific enolase (NSE) of the subject, the concentration of squamous cell carcinoma antigen (SCC) of the subject, the concentration of progastrin-releasing peptide (ProGRP) of the subject or any combination thereof, and (b) at least one of the total IgG concentration of the subject, the IgA concentration of the subject, the IgM concentration of the subject, the IgE concentration of the subject, the free kappa light chain concentration, the free lambda light chain concentration or any combination thereof, derived from a biological sample obtained from the subject, and one or more diagnostic assays configured to measure the assay values; b) a processing system including a computer processor and a non-transitory computer memory including a database and a machine learning algorithm, wherein at least one machine learning algorithm is configured to process assay values and target values for the subject, including the subject's smoking pack-year value and the measured value of the size of the IPN in the subject, and the subject's smoking pack-year value and the measured value of the size of the IPN in the subject are pre-entered in the database; the processing system i) applying at least one machine learning algorithm to the assay values and the target values to output a machine learning score for the target; ii) reporting the machine learning score for the target; and iii) providing a reference score for comparison with the machine learning score configured processing system A system comprising