Cancer-staging method using impedance spectroscopy

The impedance spectroscopy-based cancer-staging method addresses the invasiveness and inefficiencies of current techniques by using machine learning to determine cancer stage from nucleic acid impedance properties, enhancing accessibility and reducing costs.

WO2025202842A1PCT designated stage Publication Date: 2025-10-02ASIMA HEALTH INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/053033
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-26
Filing Date
2025-03-21
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Current cancer staging methods are invasive, costly, time-consuming, and inefficient, leading to patient discomfort, increased healthcare costs, and disparities in access, particularly in underserved communities.

Method used

A cancer-staging method using impedance spectroscopy that applies test signals at varying frequencies to nucleic acids isolated from a subject, calculates impedance properties, and uses a machine learning model to determine cancer stage based on reference impedance properties.

Benefits of technology

Provides a non-invasive, efficient, and cost-effective method for cancer staging that minimizes patient discomfort and resource consumption, improving accessibility and reducing healthcare disparities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025053033_02102025_PF_FP_ABST
    Figure IB2025053033_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a cancer-staging method using impedance spectroscopy to measure the electrical conductivity of a test sample in response to test signals at various frequencies. The method includes applying test signals at more than frequency to a test sample comprising DNA isolated from a subject with cancer. Impedance properties of the test sample are detected in response to the respective test signals. To determine the stage of cancer in the subject, the impedance properties are compared to training data.
Need to check novelty before this filing date? Find Prior Art

Description

Docket No. P12901PC00 CANCER-STAGING METHOD USING IMPEDANCE SPECTROSCOPY CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to US Provisional application 63 / 569,902, filed March 26, 2024. The entire contents of the foregoing are incorporated herein by reference. FIELD

[0002] The present specification is directed to medical diagnostics, and particularly methods for staging cancer in subjects. BACKGROUND

[0003] Invasiveness is a significant deficiency in the current state of the art for cancer staging. Traditional methods, such as surgical biopsies, require physical intrusion into the body, which can lead to complications, pain, and a longer recovery period for the patient. This approach can be particularly distressing for individuals already dealing with the psychological stress of a cancer diagnosis. Moreover, invasive procedures carry inherent risks, including infection, bleeding, and adverse reactions to anesthesia. As a result, there is a need for less invasive, yet equally effective, methods for cancer staging that can minimize patient discomfort and risk.

[0004] The lack of early detection, time and resource consumption, and cost are additional shortcomings. Early detection is crucial for effective cancer treatment, yet many existing staging techniques fail to identify tumors at their initial stages when they are most treatable. This delay in diagnosis can lead to worse patient outcomes and more complex, expensive treatments. Furthermore, current staging processes are often time-consuming and resource-intensive, requiring multiple medical appointments and long waits for results, which can increase patient anxiety and reduce healthcare system efficiency. The high costs associated with advanced imaging and biopsy procedures can also limit access to necessary staging, particularly in underserved communities and developing countries, exacerbating health disparities and impacting overall treatment success. These factors underscore the urgent need for advancements in cancer staging technologies that are efficient, accessible, and cost-effective.Docket No. P12901PC00 SUMMARY

[0005] An aspect of the specification provides a cancer-staging method using impedance spectroscopy including: applying test signals at a plurality of frequencies to a test sample including a nucleic acid isolated from a subject with cancer; calculating impedance properties of the test sample, each of the impedance properties responsive to one of the respective test signals; retrieving training data including reference impedance properties measured in response to the test signals, each reference impedance property associated with a stage of cancer and further associated with the frequency of the respective test signal; comparing the training data to the test data, the test data including the detected impedance properties, each of the detected impedance properties associated with the frequency of the respective test signal; and determining the stage of the cancer based on a comparison of the training data to test data.

[0006] An aspect of the specification provides a cancer-staging method wherein the nucleic acid includes cell-free DNA.

[0007] An aspect of the specification provides a cancer-staging method wherein the impedance properties include magnitude and phase angle.

[0008] An aspect of the specification provides a cancer-staging method wherein comparing the detected impedance properties to reference impedance properties includes applying a cancer staging model to determine the stage of the cancer.

[0009] An aspect of the specification provides a method for training a cancer staging model, including: (a) Collecting plasma samples from a plurality of subjects, the plasma samples including samples from subjects with cancer at different stages and from healthy subjects; (b) Isolating cell-free DNA (cfDNA) from the plasma samples; (c) Characterizing the cfDNA to determine concentration and purity, wherein the characterization includes quantifying cfDNA concentration using at least one of UV spectrophotometry or fluorescence-based quantification; (d) Preparing a test sample by suspending the characterized cfDNA in a zwitterionic buffer solution; (e) Applying a plurality of test signals at different frequencies to the test sample, wherein each test signal has a distinct frequency and amplitude; (f) Measuring impedance properties of the test sample inDocket No. P12901PC00 response to the test signals, wherein the impedance properties include magnitude and phase angle; (g) Calculating differential impedance properties by comparing the measured impedance properties of the test sample to impedance properties of a reference solution; (h) Normalizing the calculated impedance properties and cfDNA concentration data to prepare input features for a machine learning model; (i) Training a machine learning model using the normalized input features and labeled data representing cancer stages; (j) Storing the trained model and the associated input features in a memory device.

[0010] An aspect of the specification provides a method, wherein the zwitterionic buffer solution of step (d) includes a Good's buffer selected from the group consisting of HEPES, POPSO, and EPPS.

[0011] An aspect of the specification provides a method, wherein the plurality of test signals applied in step (e) have frequencies ranging from about 1 Hz to about 100 MHz.

[0012] An aspect of the specification provides a method, wherein normalizing the calculated impedance properties in step (h) includes performing one-hot encoding for categorical features and min-max scaling for continuous features to improve model convergence during training.

[0013] An aspect of the specification provides a method, wherein the machine learning model trained in step (i) is selected from the group consisting of decision trees (dt), support vector machines (svm), k-nearest neighbors (knn), and extra trees (et) classifiers.

[0014] An aspect of the specification provides a method, further including applying Synthetic Minority Over-sampling Technique (SMOTE) to the training data prior to step (i) to address class imbalance among different cancer stages.

[0015] An aspect of the specification provides a method, wherein measuring impedance properties in step (f) includes calculating the magnitude (|Z|) and phase angle (θ) of impedance using the following equations: ∣Z∣ = sqrt{(Z')2+ (Z'')2} θ = tan(Z′′ / Z′) where Z′ represents the real part of the impedance and Z′′ represents the imaginary part of the impedance.

[0016] An aspect of the specification provides a method, wherein the cfDNA isolatedDocket No. P12901PC00 in step (b) includes circulating tumor DNA (ctDNA).

[0017] An aspect of the specification provides a method, further including adjusting the pH of the zwitterionic buffer solution in step (d) to about 7.4 to optimize the stability of the cfDNA during impedance measurements.

[0018] An aspect of the specification provides a method, wherein training the machine learning model in step (i) includes performing cross-validation using a k-fold technique, where k is between 5 and 10, to validate model performance and prevent overfitting.

[0019] An aspect of the specification provides a method for predicting a cancer stage using a trained cancer staging model, including: (a) Receiving a test sample including cfDNA isolated from plasma of a subject; (b) Preparing the test sample by suspending the cfDNA in a zwitterionic buffer solution; (c) Applying a plurality of test signals at different frequencies to the test sample, wherein each test signal has a distinct frequency and amplitude; (d) Measuring impedance properties of the test sample in response to the test signals, wherein the impedance properties include magnitude and phase angle; (e) Calculating differential impedance properties by comparing the measured impedance properties of the test sample to impedance properties of a reference solution; (f) Normalizing the calculated impedance properties to prepare input features for the trained cancer staging model; (g) Inputting the normalized features into the trained cancer staging model stored in a memory device; (h) Generating a prediction of a cancer stage for the test sample using the trained cancer staging model; (i) Outputting the predicted cancer stage. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Embodiments are described with reference to the following figures.

[0021] Figure 1 is a schematic diagram of a system for cancer staging.

[0022] Figure 2 is a block diagram of a cancer-staging method, using the system of Figure 1.

[0023] Figure 3 is a block diagram of retraining a cancer staging model using the system of Figure 1.Docket No. P12901PC00

[0024] Figure 4 shows a method for building cancer staging models, using the system of Figure 1.

[0025] Figure 5 shows the 5-fold cross-validation model used for training. (Ref. https: / / scikit-learn.org / stable / modules / cross_validation.html)

[0026] Figure 6 shows absolute impedance values plotted as a function of frequency spectra, shown for a representative set of healthy and cancer subjects as indicated.

[0027] Figure 7 shows differential impedance values plotted as a function of frequency spectra, shown for a representative set of healthy and cancer subjects as indicated.

[0028] Figure 8 shows absolute phase with an inset at the 100 kHz range plotted as a function of frequency spectra, shown for a representative set of healthy and cancer subjects as indicated.

[0029] Figure 9 shows differential phase plotted as a function of frequency spectra, shown for a representative set of healthy and cancer subjects as indicated.

[0030] Figure 10 shows distribution of classifications based on different features (after outlier removal on the single frequency dataset). The graphs based on ∆Z are enlarged as well.

[0031] Figure 11 shows confusion matrices showing predicted versus true label for each type of classifier for the 100 kHz measurements, using accuracy as the evaluation metric.

[0032] Figure 12 shows a grid search CV-tuned decision tree classifier trained using accuracy as the evaluation metric.

[0033] Figure 13 shows learning and validation loss curves for 2 of the 4 learners tuned with grid search CV. The shaded zone shows the + / - 1 standard deviation of the metric around the mean (corresponding opaque line). Ideally both opaque lines should converge showing asymptotic behaviour while the standard deviation should reduce with more training examples.

[0034] Figure 14 shows confusion matrices showing predicted versus true label for each type of classifier for the multi-frequency measurements, using accuracy as theDocket No. P12901PC00 evaluation metric.

[0035] Figure 15 shows learning and validation loss curves for 2 of the 4 learners tuned with grid search CV for multi-frequency data. The translucent zone shows the + / - 1 standard deviation of the metric around the mean (corresponding opaque line). Ideally both opaque lines should converge showing asymptotic behaviour while the standard deviation should reduce with more training examples.

[0036] Figure 16 shows a schematic diagram of a non-limiting example of a computing device. DETAILED DESCRIPTION TABLE OF ABBREVIATIONS

[0037] The following abbreviations are used herein: cfDNA Cell free DNA cDNA Complementary DNA ctDNA Circulating tumor DNA EPPS 4-(2-Hydroxyethyl)-1-piperazinepropanesulfonic acid gDNA genomic DNA gRNA guide RNA HEPES (4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid) IDEs Interdigitated electrodes LOD Limit of detection M Molar mg Milligram min Minutes ml MilliliterDocket No. P12901PC00 mM Millimolar mtDNA Mitochondrial DNA μg Microgram μl Microlitre μm Micrometer μM Micromolar nm Nanometer PDMS Polydimethylsiloxane PIPES Piperazine-N,Nʹ-bis(2-ethanesulfonic acid PLA Polylactic acid PLGA Polylactic-co-glycolic acid PMMA Poly(methyl methacrylate) POPSO Piperazine-1,4-bis(2-hydroxypropanesulfonic Acid) TNM Tumor-node-metastasis UV Ultraviolet Z Impedance

[0038] Definitions:

[0039] The following definitions are used herein:

[0040] “About” herein refers to a range of ±20% of the numerical value that follows. In one example, the term “about” refers to a range of ±10% of the numerical value that follows. In one example, the term “about” refers to a range of ±5% of the numerical value that follows.

[0041] “Biomarker” herein refers to the biological marker which is a measurable indicator of some biological state or condition.Docket No. P12901PC00

[0042] “Impedance” herein refers to the effective resistance of an electric current or component to an altering current arising from combined effects of ohmic resistance and reactance.

[0043] “Zwitterionic” herein refers to a molecule that possesses both a positive and a negative electrical charge. The terms “zwitterion”, “inner salt” or “dipolar ion” may be used interchangeably herein to describe a zwitterionic molecule.

[0044] CANCER-STAGING method

[0045] A characteristic feature of epigenetic remodeling in cancer DNA is the differential hypomethylation in the coding and intergenic regions and differential hypermethylation in the CpG-rich regulatory regions such that the overall genomic landscape is markedly hypomethylated. Global hypomethylation has also been correlated with cancer stage and grade.

[0046] The present disclosure provides a method for staging cancer based on the effect of methylation patterns of nucleic acids which can be detected with impedance spectroscopy. By applying test signals with varying electrical properties, unique patterns of impedance signals can be detected and correlated to the stage of the subject’s cancer. These unique patterns likely result from the differences in the solvation properties of nucleic acids at different cancer stages because of structural variations in methylation patterns. These structural dissimilarities can vary significantly in genomic DNA and accordingly in shorter DNA molecules such as cfDNA, giving rise to differences in their dielectric properties.

[0047] Figure 1 is a schematic diagram of a system 100 for cancer staging, according to one embodiment. The system 100 includes a biosensor 102 comprising a microchamber 104 configured to hold a test sample. The microchamber 104 encloses a plurality of electrodes 106 positioned to contact the test sample. An impedance analyzer 110 is operably connected to the microchamber 104. The impedance signals are measured by placing the test sample on the plurality of electrodes 106 enclosed inside the microchamber 104. The impedance analyzer 110 may be connected to a computing device 112 for processing measurements received from the impedance analyzer 110.Docket No. P12901PC00

[0048] The microchamber 104 may comprise an opening for receiving the test sample. The opening of the microchamber 104 may be covered with a seal 108. Any suitable material may comprise the seal 108 including but not limited to glass, film, pressure-sensitive adhesive tapes, polymers, and the like. In specific non-limiting examples, the seal 108 comprises a glass coverslip, Parafilm™, Scotch™ Tape, or a transparency slide. In some examples, the seal 108 is removably attachable to the microchamber 104. In further examples, the seal 108 rests on the microchamber 104. The seal 108 may reduce evaporation and contamination of the test sample.

[0049] The microchamber 104 may comprise a material selected from but not limited to polydimethylsiloxane (PDMS), polylactic acid (PLA), polylactic-co-glycolic acid (PLGA), polyether ether ketone, silicone, nitrile, polyurethane, soft vinyl chloride resin, polypropylene, polyamide, polyethylene, polycarbonate, acrylonitrile butadiene styrene (ABS) resin, polystyrene, and poly(methyl methacrylate) (PMMA).

[0050] The plurality of electrodes 106 may comprise any suitable type of electrode including, but not limited to, interdigitated micro electrodes (IDEs), microdisk electrodes, microband electrodes, and a three-dimensional microelectrode array. The electrodes may comprise any suitable inert metal including, but not limited to platinum, and a combination thereof. The plurality of electrodes may be printed onto a substrate. The substrate may comprise any suitable material including but not limited to glass, FR4, polyimide, and the like. The plurality of electrodes 106 are positioned to contact the test sample. The plurality of electrodes 106 are configured to apply voltage to the test sample and receive a response signal from the test sample.

[0051] The impedance analyzer 110 may comprise a power source for applying an electrical signal to the plurality of electrodes 106. The impedance analyzer 110 may include or be connected to a programmable controller and a digital processing unit to receive information from the plurality of electrodes 106. As explained below, the impedance analyzer 110 is configured to detect impedance properties of the test sample based on response signals received by the electrodes 106.

[0052] The computing device 112 may be connected to the impedance analyzer 110 directly or via a network. The computing device 112 may include one or more processingDocket No. P12901PC00 units, volatile memory (i.e. random-access memory), persistent memory (i.e. hard disk devices), and a network interface, all of which are interconnected by a bus. In some examples, the computing device 112 is a virtual computing device. The computing device 112 may further include an input device for receiving inputs from a user. The computing device 112 may further include an output device such as a display. The computing device 112 may be configured to control the impedance analyzer 110 to apply electrical signals to the electrodes. As will be explained in detail below, the computing device 112 is configured to receive impedance properties from the impedance analyzer 110.

[0053] Figure 2 shows a method 200 of cancer staging according to one embodiment. In the embodiment described herein, method 200 is performed by system 100, however method 200 is not particularly limited. The method is not particularly limited to the order shown in Figure 2, and the blocks may be performed in other orders.

[0054] Block 204 comprises applying test signals to the test sample at more than one frequency. In system 100, block 204 is performed by a power source which applies electrical signals to the test sample via at least one of the electrodes 106.

[0055] The test sample comprises nucleic acid suspended in a liquid. As part of block 204, the method 200 may include collecting a biological sample from a subject with cancer and extracting a nucleic acid fraction from said biological sample. The biological sample may comprise a biological fluid selected from but not limited to whole blood, plasma, platelets, saliva, white blood cells, serum, urine, saliva, cerebrospinal fluid, amniotic fluid, bone marrow, and synovial fluid.

[0056] In embodiments where the biological sample comprises plasma, plasma may be obtained by separating a blood sample. In particular examples, the plasma is separated from blood via double centrifugation at 1500 to 2500g and preferably 2000g for about 7 to 12 minutes and preferably 10 min at about 4 °C. This is followed by plasma fractionation at about 3000g to about 3500g and preferably 3200 g for about 10 to 20 minutes, and preferably 15 minutes at about 4°C to obtain platelet poor plasma. If the platelet poor plasma is not used immediately, the platelet poor plasma may be stored at about -80° for future use. Prior to nucleic acid extraction, the platelet poor plasma may be subject to further centrifugation at about 12,000g to 18,000g and preferably 16,000g forDocket No. P12901PC00 8 to 12 minutes and preferably about 10 minutes at about 4°C to separate any residual cell debris.

[0057] The nucleic acid fraction may comprise any suitable type of DNA, including but not limited to, genomic DNA (gDNA), cell free DNA (cfDNA), mitochondrial DNA (mtDNA), complementary DNA (cDNA), guide RNA (gRNA) and combinations thereof.

[0058] Any suitable means known in the art may be used to extract the nucleic acid fraction from the biological sample, including but not limited to phenol-chloroform, cetyltrimethylammonium bromide, silica columns, anion exchange columns, Chelex™ resin, salting out, solid-phase reversible immobilization (SPRI), and combinations thereof. In specific examples where the nucleic acid fraction comprises cfDNA, the nucleic acid fraction may be extracted from the biological sample with a MagMax™ Cell-Free DNA Isolation Kit (Thermo Fisher Scientific; Waltham, Massachusetts).

[0059] In some examples, block 204 further includes amplifying the nucleic acid. Any suitable method of amplification may be used including but not limited to polymerase chain reaction (PCR), isothermal amplification, multiple displacement amplification, and ligase chain reaction.

[0060] Block 204 may further include suspending the nucleic acid fraction in a liquid to prepare the test sample. The volume of liquid may be selected such that the concentration of nucleic acid in the test sample is equal or approximately equal to the concentration of the nucleic acid fraction in the biological sample. In other examples, the volume is selected such that the concentration of nucleic acid in the test sample is equal or approximately equal to a pre-determined standard. As part of this process, the concentration of nucleic acid in the nucleic acid fraction may be quantified. The method of quantifying the nucleic acid is not particularly limited. In specific examples, the nucleic acid may be quantified with fluorometric quantification (Qubit™, Thermo Fisher Scientific) or UV-Visible spectroscopy (Nanodrop™ Lite, Thermo Fisher Scientific).

[0061] The liquid in which the nucleic acid fraction is suspended may include but is not limited to, water, a zwitterionic buffer, and combinations thereof. In particular examples, where the liquid is water, the liquid may comprise ultrapure water, Milli-Q™Docket No. P12901PC00 water (MilliporeSigma; Burlington Massachusetts), deionized water, or the like.

[0062] In examples where the liquid comprises a zwitterionic buffer, the zwitterionic buffer may comprise a Good’s buffer. Suitable examples of Good’s buffers include but are not limited to (4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid) (HEPES), piperazine-1,4-bis(2-hydroxypropanesulfonic acid) (POPSO) and 4-(2-Hydroxyethyl)-1- piperazinepropanesulfonic acid (EPPS).

[0063] The concentration of the zwitterionic buffer may be between about 1 mM and about 100 mM. In specific examples, the concentration of the zwitterionic buffer is about 5 mM. In specific examples, the concentration of the zwitterionic buffer is about 10 mM. In further examples, the concentration of the zwitterionic buffer is about 15 mM. In yet further examples, the concentration of the zwitterionic buffer is about 20 mM.

[0064] Block 204 may further comprise adjusting the pH of the liquid. In specific non- limiting examples, the pH of the liquid is adjusted to about 7.4.

[0065] As a further part of block 204, the method 200 may include loading the test sample over the plurality of electrodes 106 in the biosensor 102. The test sample may be added into the microchamber 104 such that the test sample covers at least two of the plurality of electrodes 106. After loading the test sample into the microchamber 104, the microchamber 104 may be enclosed with the seal 108 to prevent evaporation of the test sample.

[0066] Any suitable number of test signals may be applied to the test sample. In certain embodiments, the test signals may differ in frequency and / or amplitude. The test signals may further differ in amplitude. In certain embodiments, the amplitude of the test signals may be between about 1 mV and about 500 mV. In other embodiments, the amplitude may be higher but should not cause Joule heating or electrolysis.

[0067] The frequencies of the respective test signals may be between about 1 Hz to about 100 MHz. In particular examples, the frequency of one of the test signals may be about 0.1 MHz. Generally, the frequencies of the test signals are selected to maximize the difference between measurements for cancerous and non-cancerous samples.

[0068] In some examples, the test signals are applied sequentially. In a specific non-Docket No. P12901PC00 limiting example, a first signal is applied to the test sample, the first signal having a first frequency, then a second signal is applied to the test sample, the second signal having a second frequency that is different from the first frequency. A skilled person will understand that any suitable number of test signals may be similarly applied.

[0069] In some examples, two or more of the test signals are applied simultaneously using a first set of electrodes and a second set of electrodes.

[0070] Block 208 comprises calculating impedance properties of the test sample. In system 100, block 208 is performed by the impedance analyzer 110 which measures the current and voltage of response signals using at least two of the electrodes 106. Each of the response signals is responsive to a respective one of the test signals. Based on the current and voltage of the response signal and the frequency of the test signal, the impedance analyzer 110 calculates an impedance property of the test sample.

[0071] The impedance properties of the test sample may be calculated sequentially in response to the application of the test signals. In a specific non-limiting example, a first impedance property is calculated in response to a first test signal applied at block 204 and a second impedance property is calculated in response to a second test signal applied at block 204.

[0072] In some examples, the test signals are applied to a first one of the electrodes 106 and the conductivity is measured in a second one of the electrodes 106. In other examples, the test signals are applied to the same electrode in which the response signals are measured. It is not strictly necessary for each of the test signals to be applied by the same one of the electrodes 106, nor is it necessary for each of the response signals to be measured at the same electrodes 106. In some examples, the test signals are applied to various electrodes 106. In some examples, the response signals are measured at various electrodes 106.

[0073] As part of block 208, the impedance analyzer 110 calculates impedance properties based on the response signals detected at the electrodes 106. To calculate the impedance properties, the impedance analyzer 110 compares respective test signals to their respective response signals. The impedance properties may include phase shiftDocket No. P12901PC00 (Φ), phase angle (°), and magnitude (Z0). (Note that the symbol for phase angle (°) may also referred to elsewhere as θ.)

[0074] As a further part of block 208, the computing device 112 may store test data in memory. The test data comprises the detected impedance properties and the frequency of the respective test signals. The detected impedance properties may be stored in association with the frequency of the respective test signals. In some examples, the test data further includes the amplitude of the respective test signal, and the detected impedance properties are stored in association with the amplitude of the respective test signals. In some examples, the test data further comprises medical data, and the impedance properties are stored in association with the medical data. The medical data may include any suitable information about the subject including demographic data, medical history, cancer type, cell of origin, genetic mutation, treatment information, lifestyle factors, psychosocial information, laboratory test data, imaging results, and combinations thereof. The medical data may be input at the computing device 112 or received via the network. The test data may further comprise concentrations of nucleic acids for the respective test samples. The concentrations of nucleic acid may be stored in association with the detected impedance properties.

[0075] In a specific, non-limiting example, the test data is stored in a database, as represented below in Table A. In Table A, the impedance properties are phase angle and magnitude. In Table A, the test sample is associated with a unique identifier, T-1, T-2 as indicated in the “Test Sample Identifier” column. In Table A, the impedance properties are stored in association with the frequency and amplitude of the respective test signal. Table A Magnitude [Nucleic Test Frequency Amplitude of acid] Sample of Test of Test Phase Impedanc (ng / µL) Identifier Signal (Hz) Signal (V) Angle (°) e (Ohms) T-1 1 kHz 0.5 -0.5 700 1.4 T-1 10 kHz 0.1 -1 600 1.4Docket No. P12901PC00 T-1 100 kHz 0.5 -4 500 1.4 T-1 1 MHz 0.1 -3.5 400 1.4 T-2 1 kHz 0.5 -5 525 3.6 T-2 10 kHz 0.1 -2 325 3.6 T-2 100 kHz 0.5 -1.5 225 3.6 T-2 1 MHz 0.1 -2 125 3.6

[0076] Block 212 comprises retrieving training data. In system 100, block 212 is performed by the computing device 112 which retrieves a portion of the training data stored in memory at the computing device 112.

[0077] The training data includes reference impedance properties. Each of the reference impedance properties include, but are not limited to, phase shift (Φ), phase angle (°), magnitude (Z0), conductance (G), admittance (Y), resistance (R), reactance (X), susceptance (B), and the like. The training data further includes cancer stages, and each of the reference impedance properties is associated with a cancer stage. The cancer stage may be indicated according to TNM (tumor-node-metastasis) system, Ann Arbor system, Dukes staging system, Roman Numeral Staging, or any other suitable staging system. In examples where the cancer stage is indicated with the Roman Numeral System, the cancer stage may be stage 0, stage I, stage II, stage III, or stage IV. In examples where the cancer stage is indicated with the TNM system, the cancer stage may be TX, T0, Tis, T1-T4, NX, N0, N1-3, M0, or M1.

[0078] The training data further includes applied frequencies of test signals, and each of the reference impedance properties may be associated with the frequency of a respective test signal. The training data may further include applied amplitudes of test signals, and each of the reference impedance properties may be associated with the amplitude of a respective test signal.

[0079] The training data may further include cancer types, and each of the reference impedance properties may be associated with one of the cancer types. The cancer types may include, but are not limited to, carcinoma, sarcoma, lymphoma, leukemia, melanoma,Docket No. P12901PC00 the like, and combinations thereof. The training data may further include cells of origin, and each of the reference impedance properties may be associated with one of the cells of origin. The cells of origin may include, but are not limited to, lung, breast, prostate, colorectal, skin, pancreas, kidney, liver, stomach, brain, thyroid, the like, and combinations thereof. The training data may further include genetic mutations, and each of the reference impedance properties may be associated with one of the genetic mutations. The genetic mutation may include, but is not limited to, BRCA1, BRCAs, TP53, EGFR (epidermal growth factor receptor), KRAS, HER2, APC, the like and combinations thereof. The training data may further include reference impedance properties for non- cancerous subjects. The reference impedance properties for non-cancerous subjects are not associated with a cancer type, cell of origin, genetic mutation, or cancer stage.

[0080] The training data may further comprise concentrations of nucleic acids for the respective test samples. The concentrations of nucleic acid may be stored in associated with the detected impedance properties.

[0081] In a specific, non-limiting example, the training data is stored in a database, as represented below in Table B. In Table B, the reference impedance properties are phase angle and magnitude. In Table B, the control samples are associated with a unique identifier C-1, C-2, as indicated in the “Control Sample Identifier” column. In Table B, each of the reference impedance properties is associated with a cancer stage, which has been independently determined using a diagnostic method such as physical examination, computed tomography (CT), magnetic resonance imaging (MRI), X-ray, positron emission tomography (PET), ultrasound, biopsy, endoscopy, surgical reports, autopsy, the like, or a combination thereof. In Table B, the reference impedance properties are associated with the frequency and amplitude of the respective test signal. Table B Frequenc Magnitude [Nucleic Control Cance y of Test Amplitud Phase of acid] Sample r Signal e of Test Angle Impedanc Identifier Stage (Hz) Signal (V) (°) e (Ohms) (ng / µl) C-1 IV 1 kHz 0.5 -3.5 1101.4C-1 IV 10 kHz 0.1 -4 2201.4Docket No. P12901PC00 C-1 IV 100 kHz 0.5 -4.5 3001.4C-1 IV 1 MHz 0.1 -1 5901.4C-2 II 1 kHz 0.5 0 1003.6C-2 II 10 kHz 0.1 -1.5 2003.6C-2 II 100 kHz 0.5 -4 3003.6C-2 II 10 MHz 0.1 -2 6003.6

[0082] In the non-limiting example shown in Table B, each of the control samples is further associated with a cancer type. Both C-1 and C-2 are associated with lung cancer.

[0083] Generally, the training data are obtained by detecting impedance properties of control samples. Each control sample comprises nucleic acid obtained from a control subject, for whom the stage of cancer has been otherwise assessed. In some examples, the training data includes impedance properties for healthy subjects. In some examples, one or more of the reference impedance properties represent the impedance properties of a plurality of control samples. In examples where one of the reference impedance properties represents a plurality of control samples, the reference impedance property may comprise a range of values, a minimum, or a maximum. Generally, the test data is measured using a similar method and biosensor that was used to obtain the impedance properties of the control samples. The dimensions of the microchamber 104 and other variables may affect the electrical conductivity determined through the method 200. A skilled person will appreciate that the accuracy of the method 200 will depend in part on the congruity of the methods used to obtain the test data and the training data.

[0084] In some examples, the computing device 112 retrieves all the training data from memory. In other examples, the computing device 112 retrieves only a portion of the training data.

[0085] In examples where the computing device 112 retrieves a portion of the training data from memory, the retrieved training data corresponds with one or more property of the test data. The retrieved training data may include electrical properties of the test signal that correspond with the test signals applied at block 204. In particular examples, the retrieved training data includes frequencies that correspond with the frequencies of theDocket No. P12901PC00 test signals applied at block 204. The retrieved training data may include impedance properties that correspond to the impedance properties detected at block 208.

[0086] Block 216 comprises comparing the retrieved training data to the test data. In system 100, block 216 is performed by the computing device 112 which selects a portion of the retrieved training data based on a comparison to the test data. The selection is based at least on the frequency of the respective test signals and the respective impedance properties. The selection may be further based on the amplitude of the test signal, the cancer type, the cell of origin, the genetic mutation, the medical data, the like, and combinations thereof.

[0087] As part of block 216, the computing device 112 selects reference impedance properties which are associated with the same or similar frequencies as the detected impedance properties.

[0088] In some examples, the selected training data include reference impedance properties that most closely match the detected impedance properties. In examples where the reference impedance properties comprise a range of values, the selected training data include ranges that encompass the detected impedance property. In examples where the reference impedance properties comprise a minimum, the selected training data include a minimum that is below the detected impedance property. In examples where the reference impedance properties comprise a maximum, the selected training data includes a maximum that is above the detected impedance property.

[0089] In a specific non-limiting example, the computing device 112 selects one of the reference impedance properties on the basis that both the frequency of the respective test signal and the phase angle are the same as the frequency and phase angle for one of the detected impedance properties. Ideally, the selected training data represents control samples with similar impedance properties as the test samples.

[0090] In some embodiments, block 216 further comprises applying a cancer staging model. The cancer staging model may comprise a machine learning algorithm, deep- learning-based algorithm, neural network, or the like, which is trained to recognize patterns in the test data and training data. The cancer staging model may be configuredDocket No. P12901PC00 to identify similarities between the training data and the test data to improve the selection at block 216.

[0091] Block 220 comprises determining the stage of the subject’s cancer based on the comparison at block 216. In system 100, block 220 is performed by the computing device 112 which determines the cancer stage for the test sample based on the cancer stage associated with the selected training data. In some examples, the computing device 112 is configured to determine that the test sample is non-cancerous.

[0092] If the training data further includes a cancer type, the computing device 112 may further determine the cancer type based on the comparison. If the training data further includes a cell of origin, the computing device 112 may further determine the cell of origin based on the comparison. If the training data further includes a genetic mutation, the computing device 112 may further determine the genetic mutation based on the comparison.

[0093] The computing device 112 may be further configured to calculate a probability that the determination is correct.

[0094] In an exemplary performance of blocks 212 to 220, the computing device 112 identifies that the test data includes an impedance property which was responsive to the test signal having a frequency of 0.1 MHz. In this example, the computing device 112 retrieves from memory a portion of the training data that is associated with test signals having a frequency of 0.1 MHz. The computing device 112 compares the retrieved training data to the detected impedance property and selects one or more of the training data based on the comparison. In particular, the computing device 112 selects the training data that corresponds to the detected impedance properties. The computing device 112 repeats the procedure by identifying a detected impedance property which was responsive to the test signal having a frequency of 1 MHz and selecting the training data that corresponds to the detected impedance property. The computing device 112 may repeat this process for each of the detected impedance properties in order to select a plurality of training data.

[0095] As part of block 220, the computing device 112 may control the output deviceDocket No. P12901PC00 to display the determination. The output can be formatted to display some or all of the following: the cancer stage, the cancer type, the cell of origin, the genetic mutation, a probability that the determination is correct, and electrical properties of the test signals.

[0096] A person of skill in the art will understand that the cancer staging model may be trained prior to the performance of method 200 to recognize patterns in the training data, and in particular, to identify relationships between the cancer stage and impedance properties for control samples. In these examples, blocks 212 and 216 may be omitted, and block 220 applies the cancer staging model to the test data to determine the cancer stage.

[0097] In examples where the method 200 includes applying a cancer staging model, the model may be retrained with confirmatory data. Figure 3 is a block diagram showing an exemplary method 300 for retraining the cancer staging model.

[0098] In the example shown in Figure 3, method 300 is performed after determining the cancer stage at block 208, however method 300 is not particularly limited. Method 300 may be performed at any suitable time. In some examples, method 300 is performed in response to receiving confirmatory data.

[0099] At block 304, the computing device 112 receives confirmatory data. In system 100, block 304 is performed by computing device 112 which receives confirmatory data input at the computing device 112 or received via the network. In embodiments where the confirmatory data is received via the network, the confirmatory data may be generated at one or more biosensors 102 which are connected via the network and configured to share data. [000100] Confirmatory data may comprise the result of a diagnostic test such as physical examination, computed tomography (CT), magnetic resonance imaging (MRI), X-ray, positron emission tomography (PET), ultrasound, biopsy, endoscopy, surgical reports, autopsy, the like, or a combination thereof. In general, the confirmatory data is selected to verify the cancer stage of the test sample or the control samples, however the confirmatory data is not particularly limited. The confirmatory data may further include the cancer type, cell of origin, genetic mutation, or other information about the cancer. ADocket No. P12901PC00 person of skill in the art will understand that the confirmatory data is preferably the product of a reliable and accurate method of assessing the cancer stage or other information about the cancer. [000101] At block 308, the computing device 112 compares the confirmatory data with the cancer stage determined at block 220. In system 100, block 308 is performed by computing device 112. [000102] At block 312, the computing device 112 updates the training data based on the comparison at block 308. In system 100, block 312 is performed by computing device 112 which updates the training data stored in memory. The updated training data may comprise the confirmatory data and the associated test data. The updated training data is stored in memory at the computing device 112. In some examples, the updated training data may be stored in association with a unique identifier associated with the test sample. In some examples, the updated training data may be stored in association with the subject’s medical data. [000103] A person of skill in the art will understand that the input of confirmatory data can improve the accuracy of the cancer stage determined at block 220. As the computing device 112 collects confirmatory data, the model may perform cancer staging that is more tailored to a particular subject or demographic. The model may further adjust the electrical properties of the test signals to improve accuracy of the cancer staging. [000104] In an exemplary performance of method 300, block 220 determines that the test sample was isolated from a subject with stage II lung cancer. At block 304, the computing device 112 receives confirmatory data which comprises the results of a biopsy indicating that the subject has stage I lung cancer. At block 308, the computing device 112 identifies that the type of cancer was correct, but the stage was incorrect. At block 312, the cancer training model is retrained based on this correction. [000105] Figure 4 provides an overall process flow block diagram for building a cancer staging method in accordance with another embodiment, as described below. [000106] 1. Experimental Data Collection [000107] 1.1 Plasma fractionation:Docket No. P12901PC00 [000108] Block 1.1 comprises fractionating plasma from subjects, including healthy subjects and cancer subjects between stages I and IV). According to a sample dataset, a total of 45 colorectal cancer plasma samples were collected from Ontario Tumor Biobank (Ontario, Canada) and PrecisionMed (Carlsbad, CA, USA). A total of 30 healthy plasma samples were collected from consenting adult volunteers under an approved IRB protocol (#2023-3365-16254-6). From each participant, 4-6 ml whole blood was collected in a K2 / K3 EDTA tube via routine venous phlebotomy and the plasma was isolated from blood by centrifugation at 2000g for 10 min at 4 °C, within 4 h of blood collection. The resulting plasma was again centrifuged at 3200g for 15 min at 4 °C, and the supernatant was collected and transferred into cryovials before immediately freezing at -80 °C. [000109] 1.2. cfDNA isolation: [000110] Block 1.2 comprises isolating cfDNA from the plasma from block 1.1. In a present embodiment, 1 mL of plasma aliquots at -80 °C were sequentially thawed, first at -20 °C for 16 h and then at RT for 30 min. The cfDNA was extracted using the Thermofisher MagMaxTM cfDNA isolation kit following a slightly modified protocol from the manufacturer. Briefly, plasma was mixed with magnetic beads and the lysis / binding buffer. The cfDNA-bound magnetic beads were then separated using a magnetic stand followed by washing with the wash solution and final washing with 80% ethanol. The cfDNA was finally eluted in 20 μl of elution buffer (henceforth referred to as the “stock”) and used for preparing all the test samples. [000111] 1.3 cfDNA characterization: [000112] Block 1.3 comprises characterizing the cfDNA. In an embodiment, stock concentration was used. cfDNA concentration in the stock was quantified using UV spectroscopy (Nanodrop One, Thermo Fisher Scientific) and Qubit 4 fluorometer (Invitrogen, ThermoScientific, Q33226) using Qubit dsDNA high sensitivity assay kit (Q32851, Thermo Fisher Scientific). [000113] 1.4 HEPES preparation: [000114] Block 1.4 comprises preparing a HEPES buffer. In a present embodiment, 15 mM of HEPES (4-(2-Hydroxyethyl) piperazine-1- ethanesulfonic acid, Sigma Aldrich,Docket No. P12901PC00 H3375) was prepared in milliQ water (conductivity < 0.6 μS / cm). The pH of the buffer was set to 7.4 using 1 M sodium hydroxide. The buffer was then routinely stored at 4 °C and filtered through a 0.22 µm nylon syringe filter prior to use. [000115] 1.5 Reference solution preparation: [000116] Block 1.5 comprises preparing a reference solution. The reference solution was prepared by mixing 2% (v / v) of the elution buffer (from Magmax cfDNA extraction kit) with filtered HEPES buffer. It was prepared at the time of sample preparation and used for calculation of differential sample impedance. [000117] 1.6 Test sample preparation: [000118] Block 1.6 comprises preparing a test sample. 2 µl of eluted cfDNA (in the MagMaxTMelution buffer) was added to 98 µl of 15 mM HEPES buffer at pH 7.4, and mixed vigorously. Then the sample was incubated for 3 h at 4 °C, followed by 1.5 h at room temperature (the 4.5 h reading) before impedance measurement. The samples were kept in the refrigerator following measurement and measured for impedance again after overnight incubation at 4 °C followed by 1.5 h at room temperature (the overnight reading). [000119] 1.7 Impedance measurement: [000120] Block 1.7 comprises measuring impedance. For each subject, the impedance of HEPES, reference and cfDNA sample were measured (in that order) by placing a 20 µL volume of liquid in a chamber enclosed on top of 100 x 100 µm interdigitated platinum electrodes (IDE). The IDE in turn was connected to an impedance analyzer (Sciospec ISX-3, Germany). A fixed voltage of 200 mVpp was then applied across a wide frequency range of 20 Hz to 30 MHz and data were collected as 301 discrete data points to generate a frequency sweep. The sample chamber was kept covered throughout the duration of the measurement to minimize evaporation. After each sample measurement, the chamber was rinsed thrice with HEPES buffer, and after every 20 measurements or so, it was cleaned with an extra wash with 1 M NaOH, followed by milliQ water. All measurements were performed at room temperature. [000121] 2. Data Recording and ReportingDocket No. P12901PC00 [000122] 2.1 Recording clinical data: [000123] Block 2.1 comprises recording of clinical data. In a present embodiment, the sex, age, cancer stage, and cancer self-history of each subject were recorded as shown in Table 1. [000124] 2.2 cfDNA concentration data collation: [000125] Block 2.2 comprises recording and collating the cfDNA concentration data. According to a sample data set, the test sample cfDNA concentrations were obtained for each subject using their respective stock cfDNA concentrations, which were determined using two methods: (i) UV spectrophotometry with the Nanodrop One instrument (Thermo Fisher Scientific), which measures DNA concentration based on light absorption at 260 nm, and (ii) fluorescence-based quantification with the Qubit 4 fluorometer (Invitrogen, Thermo Fisher Scientific), which uses a dye that selectively binds to double-stranded DNA. The sample concentration data obtained from these methods were then collated as shown in Table 1. Table 1. Participant details used as features in our AI model.Docket No. P12901PC00 Sample ID Sex Age range Cancer Cancer Qubit plasma Nanodrop plasma (column ignored (yrs) stage self- concentration concentration during training) history (ng / ml) (ng / ml) CRC007 M 65-69 I No 5.24 51.04 CRC009 M 60-64 II No 2.82 44.00 CRC010 F 50-54 II No 7.00 42.24CRC001 F 65-69 III Yes 6.23 38.72 CRC004 M 70-74 III No 3.63 54.56 CRC005 M 50-59 III No 3.45 45.76CRC006 F 60-64 III No 9.13 54.56 CRC002 F 45-49 IV No 7.64 44.00CRC003 M 65-69IVYes 9.33 67.20 CRC011 M 60-64IVNo 9.44 240.00 CRC012 F 65-69 II No 4.61 40.32CRC013 F 50-54 III No 12.77 45.60 CRC014 M 65-69 IV Yes 32.80 52.00 CRC015 M 65-69 III No 10.84 43.91CRC016 M 50-54 III No 27.36 55.10 CRC017 F 65-69 I No 13.21 51.84 CRC018 M 60-64 III No 37.06 96.90CRC019 F 55-59 II No 4.45 37.38 CRC020 F 70-74 I Yes 12.36 62.00 CRC021 F 50-54 IV No 26.91 50.91CRC022 M 55-59 IV No 7.30 57.60 CRC023 F 55-59 II No 8.42 46.00Docket No. P12901PC00CRC024 F 75-79 IV No 21.71 58.88CRC025 F 60-64 III No 25.02 60.72 CRC026 F 60-64 I No 11.30 40.70CRC027 F 60-64 I Yes 9.68 59.20CRC028 F 50-54 II No 19.10 98.00 CRC029 M 70-74 III No 9.86 84.00CRC030 M 70-74 III Yes 42.54 134.54CRC031 M 60-64 III No 8.23 45.90 CRC032 M 70-74 II No 22.60 114.00 CRC033 M 55-59 II No 104.04 154.80 CRC034 M 60-64 II Yes 53.40 174.00 CRC035 F 55-59 IV No 430.00 728.00 CRC036 F 60-64 II No 8.91 57.80 CRC037 F 75-79 I No 7.78 46.00 CRC038 F 50-54 IV No 15.72 65.00 CRC039 M 65-69 II No 12.92 39.60 CRC040 M 75-79 IV No 15.12 46.00 CRC041 M 70-74 IV No 12.52 50.00 CRC042 M 65-69 IV No 12.94 48.60 CRC043 M 70-74 II No 31.00 88.00 CRC044 F 60-64 I No 20.18 81.40 CRC045 M 55-59 III No 36.50 115.00 CRC046 F 45-49 II No 29.40 90.00 H001 M 55-59 Healthy No 3.60 115.00Docket No. P12901PC00H002 F 60-64 Healthy No 4.80 66.00H003 F 65-59 Healthy No 8.40 96.00 H004 F 55-59 Healthy No 6.70 76.00H005 M 55-59 Healthy No 16.20 78.00H006 F 50-54 Healthy No 5.70 64.00 H007 M 45-49 Healthy No 4.50 58.00H008 M 60-64 Healthy No 7.40 62.00H009 M 55-59 Healthy No 9.60 62.00 H011 M 70-74 Healthy No 7.20 72.00H012 F 55-59 Healthy No 8.42 52.00H013 F 50-54 Healthy No 4.92 76.00 H014 M 60-64 Healthy No 9.62 68.00H015 F 45-49 Healthy No 7.95 87.50H016 F 60-64 Healthy No 4.94 30.00 H017 M 60-64 Healthy No 9.94 104.00H018 F 45-49 Healthy No 4.34 80.00H019 F 50-54 Healthy No 4.83 120.00H020 M 50-54 Healthy No 7.74 58.00H021 F 50-54 Healthy No 8.34 34.00 H022 M 55-59 Healthy No 8.70 48.00H023 F 40-45 Healthy No 5.94 60.00H024 F 35-39 Healthy No 6.38 78.00 R001 F 60-64 Healthy Yes 13.44 52.00R002 F 50-54 Healthy Yes 54.40 98.00Docket No. P12901PC00 R003 F 55-59 Healthy Yes 18.70 162.00R004 M 70-74 Healthy Yes 30.60 60.00 R005 M 60-64 Healthy Yes 27.80 74.00 R006 M 60-64 Healthy Yes 27.40 114.00R007 M 70-74 Healthy Yes 23.20 62.00 [000126] 2.3 Impedance data: [000127] Block 2.3 comprises measuring and recording the impedance data. Impedance data were collected for HEPES, reference solution and cfDNA samples for each subject across a frequency spectrum. Measurements were taken at 4.5 h after preparation, and once overnight after preparation. The measured impedance output was collected as the real part of Z (Z′) and the imaginary part (Z′′), where Z′ indicates the system resistance and Z′′ indicates the system reactance (capacitance plus inductance). All the measured values in our experiments yielded a negative value of Z′′ which indicated that our system behaved as a capacitor, not an inductor. The relative contributions of Z’ and Z′′ were frequency-dependent as expected. The measured Z’ and Z′′ values were used to calculate the magnitude (∣Z∣) and phase angle (θ) of impedance using Equations 1 & 2: [000128] ∣Z∣ = Sqrt ((Z′)2+(Z′′)2) Eq. 1 [000129] θ = tan−1(Z′′ / Z′) Eq. 2 [000130] The Z values were then reported as the mod of differential impedance, |∆Z|, where ∆Z = Zsample - Zreference. Similarly, the phase angles were reported as the mod of differential phase angle |∆θ|, where ∆θ = θsample - θreference. The impedance values including the magnitude and phase angle, obtained at 4.5 h and overnight incubation, are illustrated in Tables 2 to 4. (Note that the symbol used for phase angle (θ) may also referred to elsewhere as °.)Docket No. P12901PC00 Table 2. Impedance magnitude and phase values obtained at 100 kHz for a subset of five representative subjects listed in Table 1. These data were taken at the 4.5 h mark after sample preparation. Sample ID Z HEPES θ HEPES Z Ref 4.5 h θ Ref 4.5 h Z Sample θ Sample |∆Z| 4.5 h |∆θ| 4.5 h 4.5 h 4.5 h (deg) (ohms) (deg) 4.5 h 4.5 h (deg) (ohms) (ohms) (ohms) (deg) CRC001 768.02 -1.57 771.1 -1.56 751.39 -1.58 19.71 0.02 CRC002 765.54 -1.59 772.57 -1.58 746.62 -1.58 25.95 0.0 CRC003 767.1 -1.56 773.62 -1.57 748.78 -1.56 24.84 0.01 H001 757.37 -1.59 761.94 -1.59 709.44 -1.58 52.50 0.01 H002 757.27 -1.58 760.20 -1.58 715.03 -1.59 45.17 0.01 Table 3. Overnight impedance values at 100 kHz obtained for the same subjects shown in Table 2. Sample ID Z HEPES θ HEPES Z Ref θ Ref Z Sample θ Sample |∆Z| |∆θ| Overnight Overnight Overnight Overnight Overnight Overnight Overnight Overnight (ohms) (deg) (ohms) (deg) (ohms) (deg) (ohms) (deg) CRC001 754.57 -1.53 761.22 -1.53 734.79 -1.57 26.430.04CRC002 755.94 -1.54 763.63 -1.53 731.82 -1.57 31.810.04CRC003 759.34 -1.53 763.85 -1.54 734.79 -1.52 29.060.02H001 754.33 -1.55 759.7 -1.55 709.44 -1.56 50.260.01H002 755.36 -1.54 759.83 -1.55 714.99 -1.54 44.840.01Table 4. Frequency sweep data shown for limited range for a representative cancer patient sample (CRC001) taken at the 4.5 h mark after sample preparation. Frequency HEPES 4.5 h Reference Sample 4.5 h 4.5 h |∆Z| HEPES 4.5 h Reference Sample 4.5 4.5 h |∆θ| (deg) (Hz) Z (Ohm) 4.5 h Z Z (Ohm) (Ohm) θ (deg) 4.5 h θ h θ (deg) (Ohm) (deg) 20.0 21308.2 21457.07 21753.91 296.84 -83.86 -84.41 -83.49 0.92 20.97 20558.02 20624.32 20736.28 111.96 -81.92 -82.15 -81.91 0.24Docket No. P12901PC00 21.99 20010.12 20187.26 19936.02 251.234 -83.66 -83.28 -83.62 0.34 23.06 19180.53 19266.77 19195.31 71.46 -82.92 -82.06 -82.56 0.5 24.17 17905.6 18174.43 18132.85 41.58 -82.66 -83.52 -83.02 0.5 25.35 17105.66 17035.24 17221.34 186.10 -83.21 -83.48 -83.58 0.10 26.58 16619.01 16714.27 16631.52 82.75 -82.02 -81.87 -82.1 0.23 27.87 15833.58 15863.64 15931.5 67.86 -81.88 -82.01 -82.25 0.24 29.22 14962.84 14956.77 14999.91 43.14 -81.38 -81.54 -81.39 0.15 30.64 14295.87 14410.4 14430.52 20.12 -81.0 -81.21 -81.17 0.04 … 24818376.64 236.66 236.56 237.12 0.56 -47.14 -47.32 -46.94 0.38 26023178.59 229.34 229.35 229.85 0.50 -46.33 -46.49 -46.18 0.31 27286467.38 223.55 223.45 224.04 0.59 -46.03 -46.24 -45.91 0.33 28611082.12 217.45 217.31 217.86 0.55 -46.27 -46.44 -46.12 0.32 30000000.03 207.68 207.53 208.08 0.55 -46.28 -46.48 -46.16 0.32 Table 5. Frequency sweep data shown for limited range for a representative healthy sample (H001) taken at the 4.5 h mark after sample preparation. Frequency HEPES 4.5 h HEPES 4.5 h Reference Reference Sample 4.5 h Sample 4.5h 4.5 h |∆Z| 4.5 h |∆θ| (Hz) Z (Ohm) θ (deg) 4.5 h Z (Ohm) 4.5 h θ (deg) Z (Ohm) θ (deg) (ohm) (deg) 20 21363.91 -82.88 21563.34 -83.79 21401.36 -83.03 161.98 0.76 20.97 20872.32 -82.05 20779.67 -82.17 20704.57 -82.5 75.10 0.33 21.99 19818.95 -82.89 19905.17 -82.6 19607.42 -81.91 297.75 0.69 23.06 18793.23 -83.47 18903.44 -83.78 18653.02 -84.15 250.42 0.37 24.17 18098.42 -83.03 18181.32 -81.91 18015.14 -82.5 166.18 0.59 25.35 17544.14 -81.19 17442.27 -81.31 17260.6 -81.3 181.67 0.01 26.58 16767.24 -81.54 16760.86 -82.65 16577.87 -81.53 182.990 1.12 27.87 15923.98 -82.82 15908.33 -83.49 15946.91 -82.88 38.578 0.61 29.22 15201.37 -82.58 15154.76 -81.84 14926.67 -82.05 228.09 0.21 30.64 14727.44 -80.8 14666.71 -80.54 14545.12 -80.94 121.59 0.40 …Docket No. P12901PC00 26023178.59 227.54 -46.1 227.51 -46.18 228.53 -45 1.02 1.18 27286467.38 224.18 -45.93 224.23 -46 225.32 -44.91 1.09 1.09 28611082.12 215.86 -46.21 215.84 -46.3 217.01 -45.22 1.17 1.08 30000000.03 206.42 -45.99 206.54 -46.06 207.59 -45.01 1.05 1.05 26023178.59 227.54 -46.1 227.51 -46.18 228.53 -45 1.02 1.18 [000131] 3. Feature Transformation for Multi-classifier Input Training Data [000132] Block 3.1 comprises one-hot encoding of categorical recordings obtained from Block 2.1. In a present example embodiment, the one-hot encoding transforms each category into a distinct binary vector to enable processing by multi-class classifiers, thereby preventing the misinterpretation of categorical relationships by the classifiers. [000133] Block 3.2 comprises normalizing the measurements from blocks 2.2 and 2.3. Before input to the classifier training model, the numerical values referred to in Tables 1 to 5 were normalized to fit the ranges shown in Table 6. This was done to improve algorithm convergence by transforming data to comparable scales, reducing steep cost function oscillations, and enhancing gradient descent efficiency. The scaling of features prevents overshadowing, mitigates outlier influence, ensures consistent feature contribution, and enhances learning for gradient-based algorithms by addressing vanishing / exploding gradient problems and improving predictions. [000134] Scalar transformation was performed according to the following functions: [000135] X_std = (X - Xmin) / (Xmax- Xmin) Eq. 3 [000136] X_scaled = X_std * (max - min) + min Eq. 4 [000137] where, the standard deviation was derived for the given value according to Eq. 3. Here, Xmax and Xmin define the maximum and minimum datapoints attainable for that value. The standard deviation was then used to scale the value based on the maximum and minimum (max, min) values for each range as defined in Eq. 4. Table 6. Feature data normalization for AI model training of both the multi-frequency data and the 100 kHz-only impedance data. In the latter, outliers (sample IDs CRC011, CRC033Docket No. P12901PC00 and CRC035) were removed. Feature Actual range of Adjusted range of Minimum Maximum value measurements from measurements after value used for used for experiments (multi- removing outliers normalization normalization frequency dataset) (100 kHz-only dataset) cfDNA concentration 30 - 728 30 - 174 0 1000 via Nanodrop (ng / ml) cfDNA concentration 2.82 - 430 2.82 - 54.4 0 100 via Qubit (ng / ml) Absolute value of 4 - 70 4 - 70 0 100 differential impedance |∆Z| (ohms) Absolute value of 0.0 - 0.14 0.0 - 0.1 0 0.5 differential phase angle |∆θ| (degrees) [000138] The predicted class was label-encoded as follows: Cancer stage: Healthy, I, II, III, IV represented as 0, 1, 2, 3, 4, respectively. [000139] The remaining input variables were one-hot encoded versus being ordinally encoded, i.e., instead of being encoded as numeric values, they were encoded as vectors where each vector's dimensionality was equal to the number of categories. These were sparse vectors where all but one of the values was 0, and the non-zero value (of 1) was given to the dimension that matched the category. For example, M and F were [0, 1] and [1, 0], respectively. More specifically, the input categorical values referred to in Tables 1 to 5 were one-hot encoded as follows: [000140] Sex: M / F were split into separate binary-valued attributes “Donor Sex_M” and “Donor Sex_F”. For example, a data point corresponding to a male patient would have these attributes valued as 1 and 0 respectively in the transformed input data used to train the models. [000141] Cancer self-history: Yes / No were split into separate binary-valued attributes “Cancer history-self_Yes” and “Cancer history-self_No”. [000142] Age: Quantized age ranges of roughly 5 years (e.g. 45-49, 50-54, etc.) wereDocket No. P12901PC00 split into separate binary-valued attributes “Donor Age Range (yrs)_45-49”, “Donor Age Range (yrs)_50-54” etc. In all, there were 10 such attributes, corresponding to the age ranges 35-39, 40-44, 45-49, 50-54, 50-59, 55-59, 60-64, 65-69, 70-74 and 75-79. The overlapping age range 50-59 had to be included because of its presence as such in the dataset. [000143] Taking impedance measurements across a frequency spectrum is a classic example of what is considered high-dimensional data (https: / / www.sciencedirect.com / topics / computer-science / high-dimensional-data) in statistical machine learning. The number of features in the multi-frequency data (after all transformations outlined in the previous section) stood at 617, whereas the available sample size was only 75. [000144] In such situations, high-dimensional techniques may be used to reduce the dimensionality of the data without losing the underlying variance. Alternately, function approximation techniques may be used to represent the full frequency sets with a reduced number of parameters. For this dataset, we did not make any assumptions about the nature of the underlying physical phenomena that results in the observed values at each frequency in the spectrum. So instead of using function approximation techniques, we focused instead on dimensionality reduction. The techniques considered were PCA (https: / / www.ibm.com / think / topics / principal-component- analysis#:~:text=Principal%20component%20analysis%2C%20or%20PCA,of%20variab les%2C%20called%20principal%20components), UMAP (https: / / umap- learn.readthedocs.io / en / latest / ) and Tucker decomposition (https: / / en.wikipedia.org / wiki / Tucker_decomposition). [000145] PCA could be applied over all 617 features. The top 10 principal components accounted for 89.54% of the variance in the data, with no outlier removal. [000146] UMAP could also be applied over all 617 features. With some tuning, the following combination of UMAP parameters was chosen to achieve the results outlined in further sections: [000147] No. of components: 5Docket No. P12901PC00 [000148] No. of neighbours: 5 [000149] Minimum distance between components: 0.6 (range of 0.0 to 0.99) [000150] Distance metric: Euclidean [000151] Tucker decomposition applies to tensors. The data were modeled as a tensor (samples x dimensions x measures = 75 x 301 x 2). Out of the few experiments tried with Tucker decomposition with different low-rank values (e.g., 8 and 10) for the dimensions’ factor, this technique was dropped because of favourable results obtained with UMAP. It is possible that higher values of the rank could do better. Tucker may also be better suited than UMAP when using more than just the differential impedance and phase angle features, effectively increasing the dimensionality of the 3rd factor (measures) in the tensor. [000152] The best results were obtained using UMAP for feature reduction, and then training classifiers (as described below) on the reduced dimension data. [000153] Since UMAP was used finally for dimensionality reduction, outlier removal was not necessary. UMAP is less sensitive to the presence of extreme outliers compared to, for example, PCA which makes linearity assumptions when reducing dimensions. [000154] 4. Multi-classifier Training [000155] For the single frequency dataset, 3 data points with extreme outlier values (on more than one of the attributes) were removed from the available dataset of 75 patients. As a result, the single and multi-frequency datasets had a slight class imbalance. The split of classes was as follows for the single frequency dataset (72 samples after outlier removal): [000156] Class 0 (Healthy): 30 samples [000157] Class 1 (Cancer Stage 1): 7 samples [000158] Class 2 (Cancer Stage 2): 12 samples [000159] Class 3 (Cancer Stage 3): 13 samples [000160] Class 4 (Cancer Stage 4): 10 samplesDocket No. P12901PC00 [000161] For the multi-frequency dataset, the class counts were as follows: [000162] Class 0 (Healthy): 30 samples [000163] Class 1 (Cancer Stage 1): 7 samples [000164] Class 2 (Cancer Stage 2): 13 samples [000165] Class 3 (Cancer Stage 3): 13 samples [000166] Class 4 (Cancer Stage 4): 12 samples [000167] Given the paucity of training data and the imbalance within it, SMOTE (https: / / imbalanced- learn.org / stable / references / generated / imblearn.over_sampling.SMOTE.html) was used to oversample the training split. Techniques such as SMOTE affect only the training split. The choice of evaluation metric (on the test split) for the trained model has to take this into consideration. [000168] 4.1 Training procedure: [000169] Block 4.1 comprises pre-processing input data by shuffling the input data with stratification, transforming the input data, and splitting the input data into a training set and a testing set in an 80:20 ratio. The data were randomly shuffled, stratified, transformed and then split into an 80:20 ratio such that 80% of the data were used for training using the 5-fold cross-validation (Fig. 5). The remaining 20% was the test (or hold-out) set used to evaluate the final model performance. [000170] 4.2 Training models studied: [000171] Block 4.2 comprises training at least one machine learning model. A number of standard learners / training models shown below were studied to evaluate their performance: [000172] (i) decision tree (DT) - https: / / scikit- learn.org / stable / modules / tree.html#classification [000173] (ii) support vector machine (SVM) - https: / / scikit- learn.org / stable / modules / svm.html#classificationDocket No. P12901PC00 [000174] (iv) k-nearest neighbors (KNN) - https: / / scikit- learn.org / stable / modules / neighbors.html#nearest-neighbors-classification [000175] (v) extra trees (ET) - https: / / scikit- learn.org / stable / modules / generated / sklearn.ensemble.ExtraTreesClassifier.html [000176] (https: / / scikit-learn.org / stable / modules / cross_validation.html) [000177] (vi) grid search for learner hyperparameter optimization (https: / / scikit- learn.org / stable / modules / grid_search.html#exhaustive-grid-search). [000178] 4.3 Evaluation of the goodness of training: [000179] The models were validated using the following quantitative scores before proceeding to testing: [000180] Accuracy - selected as a well-rounded baseline. [000181] F1 (macro) - selected because of class imbalance and to balance the importance of both sensitivity and specificity. [000182] Recall (macro) - selected because of class imbalance and because it is important to reduce the number of false negatives (not detecting the presence of any stage of cancer) at the cost of higher number of false positives [000183] Cancer-focused weighted recall - a custom metric that ignores recall for healthy samples and focuses instead only on recall for cancer stage classes. Within the cancer classes, it progressively accords higher weightage to correct predictions on classes corresponding to later cancer stages. In particular, the ratio of weights for the 4 stages are respectively 1, 1.2, 1.5 and 2. This metric aims to select models which are increasingly better at late-stage cancer detection sensitivity. [000184] The choice of evaluation metric did not impact training and hyperparameter tuning of results. [000185] References: https: / / scikit-learn.org / stable / modules / model_evaluation.html# [000186] https: / / scikit-learn.org / stable / modules / model_evaluation.html#string-name- scorersDocket No. P12901PC00 [000187] 5. Output of Testing Multi-classifier [000188] 5.1 Cancer classification (Healthy, I, II, III, IV): [000189] Block 5.1 comprises classifying each subject based on cancer staging. The classification for each subject was based on the correlation of staging to primary input features: ∆Z at 4.5 h, ∆Z overnight, and secondary input features: age, sex and cancer history. [000190] 5.2 Evaluation [000191] Block 5.2 comprises evaluation of testing. Evaluation was based on accuracy, recall, F1-score and a cancer-focused weighted recall score for each model tested. [000192] RESULTS [000193] Our previous research demonstrated that genome-wide hypomethylation alters the solvation behavior of cancer-derived cfDNA across more than fifteen cancer types, enabling its differentiation from healthy cfDNA based solely on its electro- physicochemical properties (Ref: https: / / www.biorxiv.org / content / 10.1101 / 2024.05.10.593096v1.abstract). [000194] This distinctive behavior using impedance spectroscopy was used in our present study to assess the global hypomethylation status of cfDNA for rapid colorectal cancer detection. The assay exhibited slow kinetics, with cfDNA aggregation progressing gradually over time and reaching saturation after approximately 10–12 h. To achieve optimal differentiation between healthy and cancer cfDNA, impedance measurements were carried out at an early time point of 4.5 h and a later time point following overnight incubation, once equilibrium had been reached. [000195] The frequency spectra of impedance for healthy and cancer cfDNA samples showed an exponential drop in magnitude (Fig.6), with the highest signal difference seen around 100 kHz (ΔZ most negative) as shown in Fig. 7. Therefore, this frequency was prioritized for our initial AI modeling. [000196] At 100 kHz, the impedance magnitude exhibited a plateau and low phase angles (< -2o) (Fig. 8), suggesting that the system behaves mainly as a resistor aroundDocket No. P12901PC00 this frequency range. The healthy cfDNA was found to be more conductive than cancer cfDNA (lower Z value) at all frequencies (Fig. 9). [000197] The effect of voltage was found to be invariant in the 5 to 500 mV range, so it was dropped as one of the features for analysis. [000198] Together, these figures demonstrated the distinct electro-physicochemical properties of cfDNA associated with cancer status, reinforcing the selection of 100 kHz as a key frequency for AI model input. [000199] AI results with 100 kHz data [000200] The classifier test model was initially evaluated at only 100 kHz frequency rather than the multi-frequency sweep data to check the performance of the training model at reduced dimensionality. This frequency was specifically chosen because it exhibited the greatest difference in impedance magnitudes between cancerous and healthy samples as discussed above. The training was performed for both 4.5 h as well as for the overnight data. The results as well as their error margins obtained with this approach are shown as box plots in Fig. 10. [000201] Confusion matrices were also produced for the 100 kHz measurements to determine the precision and recall, where the true label was set against the predicted label (Fig. 11). In particular, confusion matrices showed predicted versus true label for each type of classifier, using accuracy as the evaluation metric. [000202] No. healthy (0) samples = 6 [000203] No. stage I (1) samples = 1 [000204] No. stage II (2) samples = 3 [000205] No. stage III (3) samples = 3 [000206] No. stage IV (4) samples = 2 [000207] The quantitative scores generated to assess the performance of the classifiers are summarized in Table 7. The best results were obtained with SVM, having an accuracy of 0.533, precision of 0.467, recall of 0.433, F1 of 0.447 and cancer-focused weighted recall of 0.600.Docket No. P12901PC00 Table 7. A summary of quantitative scores for different classifiers studied at 100 kHz. CLASSIFIER dt svm knn et Accuracy 0.400 0.533 0.533 0.533 Precision 0.330 0.467 0.425 0.447 Recall 0.367 0.433 0.433 0.433 F1 0.314 0.447 0.423 0.381 Cancer focused 0.600 0.600 0.600 0.625 weighted recall [000208] The et classifier also deserves a mention as it exhibited an accuracy of 0.533, precision of 0.447, recall of 0.433 and F1 score of 0.381, but its cancer-focused weighted recall score was 0.625. Upon visual inspection, the extra tees classifier did a better job of predicting the cancer stage 3 class closer to its actual values, while losing out on the accuracy of predicting cancer stages 1 and 2. This illustrates the difference between the standard precision / recall / F1 metrics and the custom metric aimed at reducing false negatives in cancer stage prediction. If closeness to the ground truth class is also important (in addition to simple progressive weighting of later stages), the evaluation metric can be reconfigured to grade learners accordingly. [000209] Fig.12 shows an example of a search CV-tuned decision tree classifier trained using accuracy as the evaluation metric. [000210] Interpretation of AI results with the 100 kHz data [000211] Inspection of the training and validation loss curves for the dt and SVM classifiers highlights the broad theme that the learners are over-fitting (low bias, high variance) the training split, but are learning well on the validation split (i.e., hitherto unseen data) (Fig.13). This means that the models are either too complex, or that more training data would help. The exception to this is the et classifier, which didn’t over-fit the training data, but still showed evidence of being able to learn more, provided there is more training data provided to it. [000212] Considering that the majority class constitutes 41% of the dataset and a brute-Docket No. P12901PC00 force “classifier” that simply classifies every incoming sample as “Healthy” can only be 41% accurate at best, these results imply that the accuracy of 0.533 achieved by the SVM classifier could be considered as the baseline to assess learner performance as more data becomes available. [000213] The upward trend in both validation and cross validation scores indicates that these learners have the ability to get better if more, especially varied, data is provided to them. It also means that there is scope for incorporating higher-dimensional data (full frequency sweeps) along with high-dimension reduction techniques (e.g. smoothing splines, low-rank decomposition) to potentially achieve even better prediction performance. [000214] AI results with multi-frequency sweep data [000215] The classifier test model was finally evaluated across the multi-frequency dataset. This was done on a logarithmic frequency spectrum of 301 points between 20 Hz and 30 MHz. Confusion matrices were produced for this spectrum to determine the precision and recall (Fig.14), where the true label was set against the predicted label for each type of classifier as shown below: [000216] No. healthy (0) samples = 6 [000217] No. stage I (1) samples = 1 [000218] No. stage II (2) samples = 3 [000219] No. stage III (3) samples = 3 [000220] No. stage IV (4) samples = 2 [000221] The quantitative scores generated for assessing the performance of the classifiers are summarized in Table 8. Cancer-focused weighted recall was ignored for this dataset. Table 8. Quantitative scores for different classifiers for multi-frequency data CLASSIFIER dt svm knn et Accuracy 0.267 0.600 0.467 0.733Docket No. P12901PC00 Precision 0.275 0.525 0.417 0.7 Recall 0.200 0.567 0.5 0.7 F1 0.219 0.476 0.4 0.647 [000222] The et classifier performed the best on this dataset. Considering that a brute- force classifier which only classifies every incoming sample as “Healthy” will only achieve a 41% classification accuracy at best, the 0.733 accuracy returned by the tuned extra trees classifier is a remarkable result, especially for such high-dimensional data. As such, this should be the baseline upon which future performance improvements must be measured. [000223] Interpretation of AI results with multi-frequency sweep data [000224] The impact of UMAP in addressing the high dimensionality of the dataset is readily apparent. UMAP does a good job of preserving both local and global structures in the data and modelling non-linear relationships between dimensions, which gives it an advantage over methods such as PCA, which are limited to capturing global structure and linear relationships between dimensions. [000225] Given that no assumptions were made about the underlying phenomenon that result in the observed values for differential impedance and phase angle at different frequencies, UMAP was the better choice to represent the high-dimensional data. [000226] Inspection of the training and validation loss curves (Fig. 15) for each of the classifiers also mimics the theme that the learners are over-fitting the training split, but are learning well on the validation split, meaning that the models are either too complex, or that more training data would help. They aren’t over-fitting the training data, but they still show evidence of being able to learn more provided there is more training data provided to them. [000227] [000228] Potential architectures for computing device [000229] Referring now to Fig. 16, it is useful to give a brief overview of potentialDocket No. P12901PC00 architectures for computing device 112 or other computing devices contemplated herein. [000230] Computing device 112 can include at least one input device 204a. Input device 204a may include traditional input mechanisms like keyboards or mice. Output device 212a may be a display or could be a speaker for providing real-time feedback. Input device 204a can include impedance analyzer 110. [000231] Processor 208a may be implemented as a plurality of processors, one or more multi-core processors, or specialized hardware accelerators, such as Graphics Processing Units (GPUs), Tensor Processing Units (TPUs), or similar parallel processing units optimized for handling large-scale data and complex machine learning models. In certain embodiments, processor 208a may include a memory-centric or near-memory compute architecture, where processing elements are co-located with memory cells to minimize data transfer latency and power consumption, which is beneficial for applications involving high-speed inference. Processor 208a may be configured to execute different programming instructions, including those optimized for AI and machine learning tasks, responsive to input received via one or more input devices 204a and to control one or more output devices 212a to generate output on those devices. [000232] To fulfill its programming functions, processor 208a is configured to communicate with one or more memory units, including non-volatile memory 216a and volatile memory 220a. Non-volatile memory 216a can be based on any persistent memory technology, such as an Erasable Electronic Programmable Read-Only Memory (“EEPROM”), flash memory, solid-state hard disk (SSD), other types of hard disks, or combinations thereof. In embodiments using a memory-centric architecture, memory 216a may include processing units embedded within or adjacent to memory cells to reduce latency and improve energy efficiency in data handling, supporting high- performance inference tasks. Non-volatile memory 216 may also be described as a non- transitory computer-readable medium, with multiple types of non-volatile memory 216a provided as needed. [000233] Volatile memory 220a can be based on any random access memory (RAM) technology, such as Double Data Rate (DDR) Synchronous Dynamic Random-Access Memory (SDRAM). In certain embodiments, volatile memory 220a may incorporate high-Docket No. P12901PC00 speed, low-latency configurations, supporting rapid data access and processing required for real-time inference tasks. Other types of volatile memory 220a, such as Low-Power DDR (LPDDR) or High-Bandwidth Memory (HBM), are also contemplated to optimize performance in high-computation environments. [000234] Processor 208a can also connect to a network 108a via a network interface 232a. Network interface 232a can also be used to connect another computing device that has an input and output device, thereby obviating the need for input device 204a and / or output device 212a altogether. [000235] Programming instructions in the form of applications 224a are typically maintained, persistently, in non-volatile memory 216a and used by the processor 208 which reads from and writes to volatile memory 220 during the execution of applications 224a. Various methods discussed herein can be coded as one or more applications 224a. One or more tables or databases 228a are maintained in non-volatile memory 216a for use by applications 224a. Applications 224a include machine learning models or NLP modules configured for training and deployment of virtual interaction profiles. Applications 224a also include executable code for the various methods, and their variants, described herein. Databases 228a can include the various machine learning models developed, trained and maintained herein. Applications 224a can implement method 200 and method 300 and the method 400 of Figure 4. [000236] By the same token, a plurality of computing device 112 may be provided. Overall, the computing device 112 may be implemented using cloud computing platforms such as Microsoft Azure™ or Amazon Web Services (AWS)™. [000237] Furthermore, computing device 112 may be implemented as virtual machines and / or with mirror images to provide load balancing. [000238] It will now be apparent to a person of skill in the art that the present specification affords certain advantages over the prior art. The above-described impedance spectroscopy methods offer a label-free cancer-staging solution for a wide array of cancer types, thanks to the prevalence of aberrant methylation across all cancers. This approach contrasts with existing multi-cancer staging methods that predominantlyDocket No. P12901PC00 rely on assays or DNA sequencing, presenting a faster, more cost-effective alternative. Patients will benefit from quicker, non-invasive, and less expensive staging results, while healthcare providers will be empowered to stage more frequently. Early cancer staging is crucial for successful cancer treatment, and the increased screening frequency facilitated by this method will increase the chances of identifying late-stage cancers sooner. Moreover, the label-free nature of these methods contributes to reduced waste production in clinical laboratories, aligning with sustainable practices in medical testing. [000239] Importantly, the performance metrics obtained from the AI models indicate that the method is particularly effective at identifying later-stage cancers. This includes a recall score of 0.625 and 0.7 for the extra trees single and multi-frequency trained model, respectively. The upward trend in learning curves also indicate that with further training data and refinement using these techniques, it is expected these scores will increase. Regardless, these results are sufficient to support the method's utility as an accelerated prescreening tool given the ability to identify potential cancer cases with a relatively high sensitivity for later stages makes this method valuable for streamlining the diagnostic pipeline in busy healthcare systems. By reducing the need for immediate, more invasive diagnostic procedures for all patients, this prescreening approach can help prioritize cases that warrant further investigation, thereby optimizing resource allocation and reducing patient burden. The relatively simple and cost-effective nature of the impedance spectroscopy technique further enhances its potential as a prescreening tool that can be deployed widely, including in resource-limited settings. Thus, the method not only provides a practical solution for early and accessible cancer staging but also addresses a critical gap in current diagnostic workflows by enabling more efficient triaging of patients based on their risk profiles. [000240] The features and advantages of the invention are apparent from the detailed specification and, thus, it is intended by the appended claims to cover all such features and advantages of the invention that fall within the true spirit and scope of the invention. Further, since numerous modifications and changes will readily occur to those skilled in the art, it is not desired to limit the invention to the exact construction and operation illustrated and described, and accordingly all suitable modifications and equivalents mayDocket No. P12901PC00 be resorted to, falling within the scope of the invention. [000241] Below is the full log of results from the training script corresponding to the results shown in Figs.10 to 12 and Table 7. Original class distribution: Counter({np.int64(0): 24, np.int64(3): 10, np.int64(2): 9, np.int64(4): 8, np.int64(1): 6}) Balanced class distribution: Counter({np.int64(1): 24, np.int64(2): 24, np.int64(4): 24, np.int64(3): 24, np.int64(0): 24}) Tuning dt using accuracy... Hyper-parameter exploration space: {'max_depth': [5, 7, 9, 11, 13], 'min_samples_split': [2, 3, 5, 7], 'min_samples_leaf': [1, 2, 3], 'criterion': ['gini', 'entropy']} Fitting 5 folds for each of 120 candidates, totalling 600 fits Best parameters for dt: {'criterion': 'gini', 'max_depth': 7, 'min_samples_leaf': 1, 'min_samples_split': 3} Best cross-validation score: 0.808 Test set metrics for dt: accuracy: 0.400 precision_macro: 0.330 recall_macro: 0.367 f1_macro: 0.314 cancer_focused_recall: 0.600 Tuning svm using accuracy... Hyper-parameter exploration space: {'C': [10, 50, 100, 500], 'kernel': ['rbf', 'linear', 'poly', 'sigmoid'], 'gamma': ['scale', 'auto', 0.1, 0.01], 'degree': [2, 3, 4, 5, 6]} Fitting 5 folds for each of 320 candidates, totalling 1600 fits Best parameters for svm: {'C': 100, 'degree': 2, 'gamma': 'scale', 'kernel': 'rbf'} Best cross-validation score: 0.758 Test set metrics for svm: accuracy: 0.533 precision_macro: 0.467 recall_macro: 0.433 f1_macro: 0.447 cancer_focused_recall: 0.600 Tuning knn using accuracy... Hyper-parameter exploration space: {'n_neighbors': [2, 3, 5, 7], 'weights': ['distance'], 'metric': ['manhattan', 'minkowski'], 'p': [1, 2], 'algorithm': ['auto', 'ball_tree', 'kd_tree', 'brute']} Fitting 5 folds for each of 64 candidates, totalling 320 fitsDocket No. P12901PC00 Best parameters for knn: {'algorithm': 'auto', 'metric': 'manhattan', 'n_neighbors': 5, 'p': 1, 'weights': 'distance'} Best cross-validation score: 0.708 Test set metrics for knn: accuracy: 0.533 precision_macro: 0.425 recall_macro: 0.433 f1_macro: 0.423 cancer_focused_recall: 0.600 Tuning et using accuracy... Hyper-parameter exploration space: {'n_estimators': [50, 100, 250, 500], 'max_depth': [2, 3, 4], 'min_samples_split': [4, 5, 6], 'min_samples_leaf': [3, 4, 5], 'criterion': ['gini', 'entropy', 'log_loss']} Fitting 5 folds for each of 324 candidates, totalling 1620 fits Best parameters for et: {'criterion': 'entropy', 'max_depth': 4, 'min_samples_leaf': 3, 'min_samples_split': 4, 'n_estimators': 100} Best cross-validation score: 0.775 Test set metrics for et: accuracy: 0.533 precision_macro: 0.447 recall_macro: 0.433 f1_macro: 0.381 cancer_focused_recall: 0.625 [000242] Below is the full log of results from the training script corresponding to the results shown in Figs.14 and 15 and Table 8. Loaded Healthy Cohort3: 7 samples Loaded Cancer Cohort4: 16 samples Loaded Cancer Cohort3: 5 samples Loaded Cancer Cohort1: 10 samples Loaded Cancer Cohort2: 14 samples Loaded Healthy Cohort1: 23 samples UMAP Information: Number of neighbors: 5 Minimum distance: 0.6 Metric used: euclideanDocket No. P12901PC00 Original class distribution: Counter({'Healthy': 24, 'IV': 10, 'II': 10, 'III': 10, 'I': 6}) Balanced class distribution: Counter({'IV': 24, 'II': 24, 'Healthy': 24, 'III': 24, 'I': 24}) Tuning dt using accuracy... Hyper-parameter exploration space: {'max_depth': [5, 7, 9, 11, 13], 'min_samples_split': [2, 3, 5, 7], 'min_samples_leaf': [1, 2, 3], 'criterion': ['gini', 'entropy'], 'class_weight': ['balanced']} Fitting 5 folds for each of 120 candidates, totalling 600 fits Best parameters for dt: {'class_weight': 'balanced', 'criterion': 'gini', 'max_depth': 13, 'min_samples_leaf': 1, 'min_samples_split': 2} Best cross-validation score: 0.575 Test set metrics for dt: accuracy: 0.267 precision_macro: 0.275 recall_macro: 0.200 f1_macro: 0.219 cancer_focused_recall: 0.000 Tuning svm using accuracy... Hyper-parameter exploration space: {'C': [10, 50, 100, 500, 1000], 'kernel': ['rbf', 'linear', 'poly', 'sigmoid'], 'gamma': ['scale', 'auto', 0.1, 0.01], 'degree': [1, 2, 3, 4, 5, 6, 7], 'class_weight': ['balanced']} Fitting 5 folds for each of 560 candidates, totalling 2800 fits Best parameters for svm: {'C': 500, 'class_weight': 'balanced', 'degree': 1, 'gamma': 0.1, 'kernel': 'rbf'} Best cross-validation score: 0.642 Test set metrics for svm: accuracy: 0.600 precision_macro: 0.525 recall_macro: 0.567 f1_macro: 0.476 cancer_focused_recall: 0.000 Tuning knn using accuracy... Hyper-parameter exploration space: {'n_neighbors': [2, 3, 5, 7], 'weights': ['distance'], 'metric': ['manhattan', 'minkowski'], 'p': [1, 2], 'algorithm': ['auto', 'ball_tree', 'kd_tree', 'brute']} Fitting 5 folds for each of 64 candidates, totalling 320 fits Best parameters for knn:Docket No. P12901PC00 {'algorithm': 'auto', 'metric': 'manhattan', 'n_neighbors': 2, 'p': 1, 'weights': 'distance'} Best cross-validation score: 0.575 Test set metrics for knn: accuracy: 0.467 precision_macro: 0.417 recall_macro: 0.500 f1_macro: 0.400 cancer_focused_recall: 0.000 Tuning et using accuracy... Hyper-parameter exploration space: {'n_estimators': [50, 100, 250, 500], 'max_depth': [2, 3, 4], 'min_samples_split': [4, 5, 6], 'min_samples_leaf': [3, 4, 5], 'max_features': ['sqrt', 'log2'], 'criterion': ['gini', 'entropy', 'log_loss'], 'bootstrap': [True], 'class_weight': ['balanced']} Fitting 5 folds for each of 648 candidates, totalling 3240 fits Best parameters for et: {'bootstrap': True, 'class_weight': 'balanced', 'criterion': 'entropy', 'max_depth': 4, 'max_features': 'sqrt', 'min_samples_leaf': 3, 'min_samples_split': 4, 'n_estimators': 250} Best cross-validation score: 0.500 Test set metrics for et: accuracy: 0.733 precision_macro: 0.700 recall_macro: 0.700 f1_macro: 0.647 cancer_focused_recall: 0.000

Claims

Docket No. P12901PC00 What is claimed is:

1. A cancer-staging method using impedance spectroscopy comprising: applying test signals at a plurality of frequencies to a test sample comprising a nucleic acid isolated from a subject with cancer; calculating impedance properties of the test sample, each of the impedance properties responsive to one of the respective test signals; retrieving training data comprising reference impedance properties measured in response to the test signals, each reference impedance property associated with a stage of cancer and further associated with the frequency of the respective test signal; comparing the training data to the test data, the test data comprising the detected impedance properties, each of the detected impedance properties associated with the frequency of the respective test signal; and determining the stage of the cancer based on a comparison of the training data to test data.

2. The cancer-staging method of claim 1 wherein the nucleic acid comprises cell-free DNA.

3. The cancer-staging method of claim 2 wherein the impedance properties include magnitude and phase angle.

4. The cancer-staging method of claim 3 wherein comparing the detected impedance properties to reference impedance properties includes applying a cancer staging model to determine the stage of the cancer.

5. A method for training a cancer staging model, comprising: ● (a) Collecting plasma samples from a plurality of subjects, the plasma samples including samples from subjects with cancer at different stages and from healthy subjects; ● (b) Isolating cell-free DNA (cfDNA) from the plasma samples;Docket No. P12901PC00 ● (c) Characterizing the cfDNA to determine concentration and purity, wherein the characterization includes quantifying cfDNA concentration using at least one of UV spectrophotometry or fluorescence-based quantification; ● (d) Preparing a test sample by suspending the characterized cfDNA in a zwitterionic buffer solution; ● (e) Applying a plurality of test signals at different frequencies to the test sample, wherein each test signal has a distinct frequency and amplitude; ● (f) Measuring impedance properties of the test sample in response to the test signals, wherein the impedance properties include magnitude and phase angle; ● (g) Calculating differential impedance properties by comparing the measured impedance properties of the test sample to impedance properties of a reference solution; ● (h) Normalizing the calculated impedance properties and cfDNA concentration data to prepare input features for a machine learning model; ● (i) Training a machine learning model using the normalized input features and labeled data representing cancer stages; ● (j) Storing the trained model and the associated input features in a memory device.

6. The method of claim 5, wherein the zwitterionic buffer solution of step (d) comprises a Good’s buffer selected from the group consisting of HEPES, POPSO, and EPPS.

7. The method of claim 5, wherein the plurality of test signals applied in step (e) have frequencies ranging from about 1 Hz to about 100 MHz.

8. The method of claim 5, wherein normalizing the calculated impedance properties in step (h) comprises performing one-hot encoding for categorical features and min-max scaling for continuous features to improve model convergence during training.

9. The method of claim 5, wherein the machine learning model trained in step (i) is selected from the group consisting of decision trees, support vector machines, k-nearest neighbors, and extra trees classifiers.Docket No. P12901PC00 10. The method of claim 5, further comprising applying Synthetic Minority Over-sampling Technique (SMOTE) to the training data prior to step (i) to address class imbalance among different cancer stages.

11. The method of claim 5, wherein measuring impedance properties in step (f) comprises calculating the magnitude (|Z|) and phase angle (θ) of impedance using the following equations: ● ∣Z∣ = sqrt{(Z')2+ (Z'')2} ● θ = tan−1(Z′′ / Z′) where Z′ represents the real part of the impedance and Z′′ represents the imaginary part of the impedance.

12. The method of claim 1, wherein the cfDNA isolated in step (b) comprises circulating tumor DNA (ctDNA).

13. The method of claim 5, further comprising adjusting the pH of the zwitterionic buffer solution in step (d) to about 7.4 to optimize the stability of the cfDNA during impedance measurements.

14. The method of claim 5, wherein training the machine learning model in step (i) includes performing cross-validation using a k-fold technique, where k is between 5 and 10, to validate model performance and prevent overfitting.

15. A method for predicting a cancer stage using a trained cancer staging model, comprising: (a) Receiving a test sample comprising cfDNA isolated from plasma of a subject; (b) Preparing the test sample by suspending the cfDNA in a zwitterionic buffer solution; (c) Applying a plurality of test signals at different frequencies to the test sample, wherein each test signal has a distinct frequency and amplitude; (d) Measuring impedance properties of the test sample in response to the test signals, wherein the impedance properties include magnitude and phase angle;Docket No. P12901PC00 (e) Calculating differential impedance properties by comparing the measured impedance properties of the test sample to impedance properties of a reference solution; (f) Normalizing the calculated impedance properties to prepare input features for the trained cancer staging model; (g) Inputting the normalized features into the trained cancer staging model stored in a memory device; (h) Generating a prediction of a cancer stage for the test sample using the trained cancer staging model; (i) Outputting the predicted cancer stage.

Citation Information

Patent Citations

  • Use of multi-frequency impedance cytometry in conjunction with machine learning for classification of biological particles

    US20200333235A1

  • Deep learning-based methods, devices, and systems for prenatal testing

    US20210020314A1

  • Machine learning techniques for identifying malignant b- and t-cell populations

    US20220223227A1

  • Methods, compositions, and devices for rapid analysis of biological markers

    US20220325350A1

  • Method and reagent composition for performing leukocyte differential counts on fresh and aged whole blood samples, based on intrinsic peroxidase activity of leukocytes

    US5639630A