Method and system for analysing a sample containing DNA by variational spectroscopy
Vibrational spectroscopy and machine learning are used to efficiently analyze DNA fragment sizes and methylation patterns, addressing the challenge of distinguishing ctDNA from cfDNA, enabling rapid and cost-effective cancer diagnosis.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-04-02
AI Technical Summary
Current diagnostic methods for cancer, particularly those involving liquid biopsies, face challenges in accurately distinguishing low concentration circulating tumour DNA (ctDNA) from healthy cell-free DNA (cfDNA) due to their similar base-pair lengths, leading to difficulties in early detection and prolonged, costly sequencing processes.
Utilizing vibrational spectroscopy, specifically mid-infrared and Raman spectroscopy, to analyze DNA fragment sizes and methylation patterns, combined with machine learning algorithms, to quickly and cost-effectively determine the distribution of base-pair lengths and methylation status of DNA fragments, enabling the detection and quantification of ctDNA.
This approach provides rapid, accurate, and affordable detection of ctDNA, enhancing diagnostic capabilities for cancer by distinguishing DNA fragment sizes and methylation patterns, facilitating early diagnosis and personalized treatment strategies.
Smart Images

Figure EP2025077627_02042026_PF_FP_ABST
Abstract
Description
[0001] METHOD AND SYSTEM FOR ANALYSING A SAMPLE CONTAINING DNA
[0002] The present disclosure relates to analysing samples containing DNA.
[0003] Cancer is a major global cause of mortality and morbidity, exacerbated by deficiencies in current diagnostic methods, which can have low accuracy, high invasiveness, and prolonged waiting times for test results. The intrinsic heterogeneity of tumours further contribute to treatment failures post-diagnosis.
[0004] A promising solution has emerged in the form of liquid biopsies and the analysis of circulating tumour DNA (ctDNA) therein. It is assumed that ctDNA is released into the blood by tumour tissue through various cell death mechanisms. ctDNA serves as a valuable biomarker, offering sensitivity for early cancer diagnosis, effectiveness in detecting residual disease, monitoring response to therapy and diagnostic capabilities for posttherapy recurrence.
[0005] While the concentration of ctDNA in the blood increases as the size of the tumour increases, the differentiation between low concentration ctDNA that is diluted by its healthy cell free DNA (cfDNA) counterpart is more difficult during early stages of a malignant neoplasm. Analyses of DNA fragment sizes in terms of base-pair lengths of the DNA fragments can improve detection of ctDNA due to observed differences in the basepair length distributions between tumour derived DNA fragments and healthy DNA fragments in the blood. Note that unless stated to the contrary, the term cfDNA in the present disclosure is intended to refer specifically to healthy cell free DNA (i.e., DNA from non-cancer cells) and not to any DNA that is outside of cells.
[0006] Measurement of base-pair lengths of DNA fragments can be useful in other contexts. For example, measurement of DNA fragment base-pair length is useful for enriching foetal DNA in maternal plasma, as foetal fragments are shorter on average . The enrichment may improve the performance of sequencing based aneuploidy testing.
[0007] Accurate measurement of base-pair lengths of DNA fragments contributes generally to advancements in understanding the genetic basis of diseases, as well as aiding in their detection.
[0008] Measurement of base-pair lengths of DNA fragments can be useful in other contexts in molecular biology. For example, accurate quantification of DNA length is essential in the preparation of certain DNA sequencing experiments. Furthermore, DNA fragment length is often quantified after using technologies to generate DNA fragments using polymerase chain reactions as well as DNA cloning.
[0009] Detecting DNA fragment size can be challenging. In the context of detecting ctDNA, for example, measurements can be complicated by the presence of healthy cell- free DNA molecules, which typically significantly outnumbers ctDNA. Such challenges can be overcome by performing analyses of detected DNA fragment size length, sequence composition and methylation status, but such analyses often require expensive and timeconsuming measurement modalities, such as DNA sequencing.
[0010] While various sequencing methods exist, they share common fundamental principles in how they operate. To gauge the length of DNA fragments isolated from blood, the DNA is processed following a series of molecular reactions, adapters are affixed to the fragment ends to generate a library which can then be sequenced. In one common approach, the sequence of nucleotides is unveiled by detecting fluorescent light emitted from each nucleotide, subsequently integrated into the full DNA strand. The sequencing data is then aligned to a reference genome, enabling the calculation of fragment lengths based on read positions in the reference. Notably, the sequencing process is a prolonged and costly endeavour^
[0011] Agarose gel electrophoresis also provides a process in which DNA fragments are separated based on size within an agarose gel matrix. This involves loading DNA molecules onto a gel, subjecting them to an electric field, and allowing migration through the gel. Smaller DNA fragments move more swiftly, resulting in distinct bands associated with varying fragment sizes. Automated components, including cameras and image analysis software, can be used to visualize and analyse the separated DNA bands, facilitating size selection and analysis. However, while platforms of this type offer faster results compared to sequencing, they involve higher error rates, are subject to user interpretation and it is harder to use them to quantify lower abundances of DNA at a particular molecular weight.
[0012] Pang, D., Thierry, A.R. and Dritschilo, A., 2015. DNA studies using atomic force microscopy: capabilities for measurement of short DNA fragments. Frontiers in molecular biosciences, 2, p.l. discloses an alternative approach for measuring short DNA fragments using atomic force microscopy (AFM). However, this approach is relatively complex, time consuming, and difficult to implement in a clinical setting.
[0013] It is an object the present disclosure to provide alternative and / or improved ways of analysing samples containing DNA.
[0014] According to an aspect of the disclosure, there is provided a method of analysing a sample containing DNA, comprising: performing vibrational spectroscopy on the sample to obtain vibrational spectroscopy data, the vibrational spectroscopy comprising directing electromagnetic radiation onto the sample and detecting electromagnetic radiation emitted from the sample after interaction with the sample; and analysing the vibrational spectroscopy data to obtain information about a distribution of base-pair lengths of DNA fragments from the obtained vibrational spectroscopy data.
[0015] Thus, a method is provided which uses vibrational spectroscopy data to obtain information about a distribution of base-pair lengths of DNA fragments. The inventors have demonstrated that vibrational spectroscopy is effective for providing information about distributions of base-pair lengths of DNA fragments, even in complex mixtures of fragments. Vibrational spectroscopy can be performed quickly and can be accomplished using significantly simpler and cheaper equipment than alternative approaches. The increased speed and lower cost facilitate use at the point of care, such as in smaller clinics and / or hospitals and / or at patients’ bedsides. Vibrational spectroscopy data can be obtained using handheld devices (e.g., handheld Raman or Mid-IR spectrometers), which further facilitate use at the point of care. Vibrational spectroscopy data can be obtained from samples that are held in disposable (e.g., single use) sample receivers, which may be implemented using on-chip silicon technology for miniaturization. Use of disposable sample receivers promotes hygiene and maintaining of sample integrity.
[0016] Optionally, the obtained information includes a median of the distribution. The median can be efficiently obtained from vibrational spectrometry data and has been found to be reliably correlated with information of interest, such as the presence and / or concentration of ctDNA in a sample.
[0017] Optionally, the method comprises detecting the presence of ctDNA using the obtained information about the distribution of base-pair lengths of DNA fragments. The inventors have found that the method is particularly effective for detecting ctDNA. Detecting ctDNA with increased speed and / or reduced cost greatly improves diagnostic workflows.
[0018] Optionally, the method comprises determining a concentration of ctDNA using the obtained information about the distribution of base-pair lengths of DNA fragments. The inventors have found that the vibrational spectrometry data is able to reliably yield information about the concentration of ctDNA. Obtaining information about not only the presence of ctDNA but also the concentration of ctDNA provides richer diagnostic information. In some implementations, the determining of the concentration uses a predetermined calibration relationship between the median of the distribution and the concentration of ctDNA.
[0019] Optionally, the method further comprises analysing the vibrational spectroscopy data to obtain information about methylation of DNA fragments in the sample. Obtaining information about methylation further enriches the diagnostic information obtainable by the method while maintaining high speed and low cost, thereby further distinguishing over alternative techniques for obtaining comparable information about a sample.
[0020] Optionally, the information is obtained at least partly using a trained machine learning model. The inventors have demonstrated that machine learning allows high quality information to be obtained efficiently. In some implementations, the trained machine learning model comprises partial least squares regression and / or principal component regression, which the inventors have found to be particularly effective.
[0021] Optionally, using the trained machine learning model comprises preprocessing the vibrational spectroscopy data and inputting the preprocessed vibrational spectroscopy data into the trained machine learning model. The preprocessing comprises applying a second- order derivative filter to the vibrational spectroscopy data. Second derivatives help reveal the fine structure of spectral bands, making subtle features, such as shoulders and peaks, more apparent, and thereby improving the detection of overlapping bands and weak signals.
[0022] Optionally, the vibrational spectroscopy comprises infrared spectroscopy. Optionally, the vibrational spectroscopy comprises attenuated total reflection spectroscopy. Optionally, the vibrational spectroscopy comprises Raman spectroscopy. Optionally, the vibrational spectroscopy comprises both infrared spectroscopy and Raman spectroscopy. Infrared (e.g. mid-infrared) spectroscopy is particularly useful in scenarios where the sample is provided in dried form, because this form of spectroscopy has high efficacy in non-water absorption prone conditions at specific wavelengths. Conversely, Raman spectroscopy is more resilient to such environmental factors, making it a preferred choice for analysis of samples provided in aqueous form. Using both techniques (infrared spectroscopy and Raman spectroscopy) allows the method of the disclosure to be used effectively for a wide range of sample types and measurement conditions.
[0023] According to an alternative aspect of the disclosure, there is provided a method of analysing a sample containing DNA, comprising: performing vibrational spectroscopy on the sample to obtain vibrational spectroscopy data, the vibrational spectroscopy comprising directing electromagnetic radiation onto the sample and detecting electromagnetic radiation emitted from the sample after interaction with the sample; and analysing the vibrational spectroscopy data to obtain information about methylation of DNA fragments in the sample.
[0024] According to an alternative aspect of the disclosure, there is provided a system for analysing a sample containing DNA, comprising: a sample receiver configured to receive the sample; a spectrometer configured to perform vibrational spectroscopy on the sample and output vibrational spectroscopy data, the vibrational spectroscopy comprising directing electromagnetic radiation onto the sample and detecting electromagnetic radiation emitted from the sample after interaction with the sample; and a data processor configured to analyse the vibrational spectroscopy data to obtain information about a distribution of base-pair lengths of DNA fragments from the obtained vibrational spectroscopy data.
[0025] According to an alternative aspect of the disclosure, there is provided a system for analysing a sample containing DNA, comprising: a sample receiver configured to receive the sample; a spectrometer configured to perform vibrational spectroscopy on the sample and output vibrational spectroscopy data, the vibrational spectroscopy comprising directing electromagnetic radiation onto the sample and detecting electromagnetic radiation emitted from the sample after interaction with the sample; and a data processor configured to analyse the vibration spectroscopy data to obtain information about methylation of DNA fragments in the sample. Embodiments of the disclosure will be further described by way of example only with reference to the accompanying drawings.
[0026] Figure 1 schematically depicts a system for analysing a sample containing DNA.
[0027] Figure 2 schematically depicts attenuated total reflection (ATR) in a spectrometer configured to operate in a multi-bounce mode involving multiple interactions between incident light and the sample at multiple respective “hot spots”. Multi -bounce modes can provide increased signal intensity and sensitivity but typically require a larger ATR crystal surface area and / or a larger sample.
[0028] Figure 3 schematically depicts ATR in a spectrometer configured to operate in a single-bounce mode. Single-bounce modes can be implemented more compactly and / or with smaller samples.
[0029] Figure 4 is a flow chart showing a protocol for demonstrating performance of a method of analysing a sample containing DNA.
[0030] Figure 5 is a graph depicting averaged spectral data of five different DNA fragment lengths.
[0031] Figures 6 and 7 are graphs depicting regression performance for implementations respectively using partial least squares regression (PLSR) and principal component regression (PCR) to obtain information from spectral data.
[0032] Figure 8 is a graph showing results from spectroscopic measurements conducted on mixtures containing two different lengths of DNA fragments.
[0033] Figure 9 is a graph showing spectral fingerprints for different states of DNA molecules (unmethylated, methylated, and hydroxym ethylated).
[0034] Figure 10 is a graph showing how improved sensitivity is achieved using a multibounce spectrometer compared to the results shown in Figures 6 and 7.
[0035] Figures 11-15 show prediction plots for proportions of different respective DNA fragment lengths in multiple mixtures.
[0036] Figures 16 and 17 are graphs remodelling the data of Figures 11-15 in a box plot (Figure 16) and violin plot (Figure 17).
[0037] Figure 18 is a box plot depicting grouping of datapoints below 150bp (corresponding to ctDNA) and above 150bp (corresponding to healthy cfDNA). Figure 19 is two graphs showing model generated distribution plots for mixtures tested containing base-pair length DNA fragments enriched below and above 15Obp respectively.
[0038] Figure 20 is a scatter plot of datapoints showing the true percentage of methylation of adenomatous polyposis coli (APC) DNA fragments in a sample against that predicted by a PLSR model.
[0039] Figure 21 is a scatter plot of datapoints showing the true percentage of hydroxymethylation of APC DNA fragments in a sample against that predicted by a PLSR model.
[0040] Figure 22 are scatter plots of datapoints showing true percentage of methylation and hydroxymethylation of APC DNA fragments in a complex sample against that predicted by a PLSR model.
[0041] Figure 23 is a set of plots displaying the predicted proportions of individual fragment lengths generated by a PLSR model compared to the true proportions of base pair lengths in a sample of DNA.
[0042] Figure 24 schematically depicts an example of an ensemble model architecture according to the present disclosure.
[0043] Figure 25 is a graph showing mean squared error for the feedforward neural network against epoch number.
[0044] Figures 26a and 26b show a set of scatter plots of datapoints displaying the predicted proportions of base-pair lengths of DNA fragments in a sample compared to the true proportions for each of the base learners of the ensemble model.
[0045] Figure 27 is graphs showing the simulation of realistic variability in spectral measurements.
[0046] Figure 28 illustrates an example model architecture of a CNN for use in methods of the disclosure.
[0047] Figure 29 shows a set of scatter plots of datapoints displaying the predicted proportions of base-pair lengths of DNA fragments in a sample by a CNN model.
[0048] Figure 30 is a pair of graphs showing means squared error and mean absolute error for a CNN model during training. Figures 31 and 32 are flowcharts illustrating example steps for data processing for traditional machine learning models and a CNN.
[0049] Figure 33 is a fused plot of concatenated FTIR and Raman spectra, shown as intensity versus concatenated wavenumber axis for differing base-pair lengths of DNA fragments.
[0050] Figure 34 is a scatter plot of datapoints showing predicted proportions of base-pair lengths of DNA fragments in a sample compared to the true proportions for a PLSR model trained on FTIR spectroscopy data.
[0051] Figure 35 is a scatter plot of datapoints showing predicted proportions of base-pair lengths of DNA fragments in a sample compared to the true proportions for a PLSR model trained on Raman spectroscopy data.
[0052] Figure 36 is a scatter plot of data datapoints showing predicted proportions of basepair lengths of DNA fragments in a sample compared to the true proportions for a PLSR model trained on combined FTIR and Raman spectroscopy data.
[0053] The present disclosure provides methods and systems for analysing samples containing DNA. The methods and systems provide an alternative to established techniques (e.g., sequencing, agarose gel electrophoresis, AFM, etc.) for obtaining information from samples containing DNA. The approach of the present disclosure uses vibrational spectroscopy. The vibrational spectroscopy may comprise either or both of mid-infrared spectroscopy and Raman spectroscopy.
[0054] The method comprises performing vibrational spectroscopy on the sample to obtain vibrational spectroscopy data. The vibrational spectroscopy comprises directing electromagnetic radiation onto the sample and detecting electromagnetic radiation emitted from the sample after interaction with the sample. The vibrational spectroscopy data is analysed to obtain information about a distribution of base-pair lengths of DNA fragments from the obtained vibrational spectroscopy data. The vibrational spectroscopy data may take any form that is suitable for allowing the information to be obtained. For example, the vibrational spectroscopy data may represent a variation of absorbance as a function of frequency, wavelength, and / or wavenumber.
[0055] Figure 1 schematically depicts a system 2 for performing the method. The system 2 comprises a sample receiver 4 for receiving the sample 6, a spectrometer 8 for performing the vibrational spectroscopy on the sample 6, and a data processor 10 for analysing vibrational spectroscopy data output by the spectrometer 8.
[0056] In one class of implementation, the vibrational spectroscopy comprises infrared spectroscopy. The infrared spectroscopy may, for example, comprise mid-infrared spectroscopy. The infrared spectroscopy probes molecular vibrations of entities in the sample by causing the infrared radiation to interact with the sample. As molecules absorb the infrared radiation, the molecules undergo vibrational transitions between distinct energy states, resulting in characteristic absorption peaks in the infrared (e.g., midinfrared) spectrum.
[0057] As exemplified in Figures 2 and 3 and referring also to Figure 1, the vibrational spectroscopy, e.g. infrared spectroscopy, may comprise attenuated total reflection spectroscopy (ATR). The ATR may be implemented by directing infrared light 16 onto a boundary 5 between an ATR crystal element 12 and the sample 6 in the sample receiver 4. The light 16 is directed onto the boundary 5 at an angle that is such as to cause total internal reflection within the ATR crystal element 12. An exponentially decaying evanescent field penetrates the sample 6. In implementations where the light 16 undergoes multiple “bounces”, as depicted in Figure 2, the evanescent field will penetrate the sample 6 at multiple corresponding localised penetration regions (which may be referred to as “hotspots”). In implementations where the light 16 undergoes a single bounce, as depicted in Figure 3, the evanescent field 18 will penetrate the sample 6 at a single localised region (hotspot). The spectrometer 8 detects light 16 after the light has interacted with the sample 6 via the evanescent field. Spectrometer software analyses the detected light to obtain absorption spectra containing information about DNA in the sample 6. These spectra can be processed further, for example using machine learning, to segregate the spectra based on base-pair lengths of corresponding DNA fragments.
[0058] Alternatively or additionally, the vibrational spectroscopy may comprise Raman spectroscopy. Raman spectroscopy uses a focused laser beam to illuminate the sample 6. The scattered light is captured and analysed to obtain information about molecular composition via determination of characteristic Raman wavenumber shifts that correspond to molecular variations. Machine learning algorithms can be used to assist with deciphering the Raman wavenumber shifts and allow a range of DNA base-pair lengths to be distinguished from each other.
[0059] Mid-infrared spectroscopy is particularly useful in scenarios where the sample 6 is provided in dried form, because this form of spectroscopy has high efficacy in non-water absorption prone conditions at specific wavelengths. Conversely, Raman spectroscopy is more resilient to such environmental factors, making it a preferred choice for analysis of samples 6 provided in aqueous form. In comparison to IR spectroscopy, which may be sensitive to vibrations involving changes in dipole moment, Raman spectroscopy depends on changes in polarizability, making it especially effective for modes that are weak or inactive in IR, and thus providing complementary structural and chemical information. Using both techniques (infrared spectroscopy and Raman spectroscopy) allows the method of the disclosure to be used effectively for a wide range of sample types and measurement conditions.
[0060] The sample 6 may be provided in various forms. The sample 6 may comprise any tissue and / or bodily fluid. The sample 6 may for example comprise or be derived from a blood sample and / or tissue sample and / or saliva and / or cerebrospinal fluid and / or urine and / or seminal fluid. Typically, a liquid biopsy will be taken from a patient and processed to provide the sample 6. The processing may comprise separating plasma or serum from a blood sample (e.g. ,by known techniques such as centrifugation), followed by extraction of cell-free DNA to provide the sample 6. The sample 6 may comprise DNA obtained from plasma or DNA obtained from serum, separated from a blood sample. In other words, the plasma or serum may be first separated from the blood sample before the DNA is obtained.
[0061] Samples may be held in a disposable sample receiver, which may for example be intended for single-use. Disposable sample receivers (or other sample receivers) may be implemented using on-chip silicon technology for efficient manufacturing and / or miniaturisation.
[0062] The analysing of the vibrational spectroscopy data to obtain information about the distribution of base-pair lengths of DNA fragments may be performed as part of a workflow for detecting ctDNA for the purposes of cancer diagnosis in place of sequencing.
[0063] For example, a workflow may comprise taking a blood sample from a suspected cancer patient, separating the plasma and extracting the circulating DNA, sequencing the resulting sample and processing the results to produce a distribution depicting relative abundancies of different DNA fragment lengths. If the peak in the distribution extends between 90 and 150bp, which is representative for the length of ctDNA, this may be taken as preliminary evidence that the sample may contain DNA from a cancer. Further analyses may be performed for a more definitive diagnosis, for example to identify cancer-specific mutations. Embodiments of the present disclosure provide similar or better information about DNA fragment size distribution without the need for sequencing, thereby saving time and cost.
[0064] The analysing of the vibrational spectroscopy data to obtain information about the distribution of base-pair lengths of DNA fragments from the obtained vibrational spectroscopy data may be performed in various ways as long as the obtained information is detailed enough to be practically useful, for example detailed enough to provide a reliable prediction about whether the sample being analysed contains DNA derived from cancer cells.
[0065] Methods of analysing a sample containing DNA according to the present disclosure may comprise detecting the presence of ctDNA using the obtained information about the distribution of base-pair lengths of DNA fragments.
[0066] In one class of embodiment, the obtained information includes a median of the distribution of base-pair lengths of DNA fragments. Conditions that involve the presence of base-pair lengths of DNA fragments having a distribution of sizes with a median that is different to that expected for healthy cell free DNA will cause a shift in the median of the distribution of base-pair lengths of DNA fragments in the obtained information. The median can thus be used as a metric to detect abnormalities in the DNA in the sample. In the specific case of ctDNA, it is known that fragments of ctDNA tend to be smaller than healthy cfDNA, such that the presence of ctDNA in a sample may lower the median DNA fragment size. A lower median of base-pair lengths of DNA fragments can thus be indicative of the presence of ctDNA and, therefore, of the presence of cancer that should be investigated further. This insight may be used to implement the detecting of the presence of circulating tumour DNA using the method mentioned above. The detecting of the presence of circulating tumour DNA may comprise determining whether the median of the distribution is below a predetermined threshold for example. ctDNA fragments are typically enriched in the base-pair length range of 100-150bp, whereas healthy cfDNA fragments typically have a median base-pair length of 167bp. The predetermined threshold may therefore be in the range of 130-220bp for example, optionally in the range of 150- 200bp. Selecting a predetermined threshold nearer to the upper limit of these ranges will promote higher sensitivity but may also involve a higher number of false positives. Selecting a lower predetermined threshold will lower false positives but may also lower sensitivity.
[0067] In some implementations, the method comprises determining a concentration of ctDNA using the obtaining information about the distribution of base-pair lengths of DNA fragments. In implementations where the obtained information includes a median of the distribution, the determining of the concentration may be performed using a predetermined calibration relationship between the median of the distribution and the concentration. The median will vary as a function of the concentration of ctDNA, typically being pulled down monotonically as a function of increasing concentration of ctDNA due to the lower average fragment length of ctDNA compared with healthy cfDNA.
[0068] The method may further comprise analysing the vibrational spectroscopy data to obtain information about methylation of DNA fragments in the sample. The information about methylation may comprise information about the presence and / or level of unmethylated, methylated, and / or hydroxymethylated characteristics in DNA fragments in the sample.
[0069] In some implementations, the information obtained by analysing the vibrational spectroscopy data is obtained at least partly using a trained machine learning model. The trained machine learning model may comprise any one or more of the following (or others): regression-based algorithms (e.g., PLSR, PCR, Ridge, Lasso, ElasticNet); support vector regression (SVR); decision tree based methods (e.g. decision trees, random forests); ensemble methods (e.g., using models such as stacking models, bootstrap aggregation models, or gradient boosting for example using models such as XGBoost, Adaboost, Catboost); neural networks (e.g., artificial neural networks, convolutional neural networks); and deep learning models. As shown in the section below headed “DEMONSTRATIONS,” particularly good performance has been demonstrated with PLSR, PCR, Ridge and Lasso. Similar or better performance may be possible using other models when fully optimised. Such optimisation may be achieved through hyperparameter tuning and spectral pre-processing (e.g., feature engineering) during the cross-validation stage.
[0070] The use of the trained machine learning model may comprise preprocessing the vibrational spectroscopy data and inputting the preprocessed vibrational spectroscopy data into the trained machine learning model. The preprocessing may, for example, comprise applying a smoothing filter with second-order derivatives to the vibrational spectroscopy data. The primary use of the smoothing filter is to reduce noise while preserving the shape and features of the signal. Second-order derivatives help reveal the fine structure of spectral bands. They make subtle features, such as shoulders and peaks, more apparent, improving the detection of overlapping bands and weak signals. Alternatively or additionally, the preprocessing may comprise baseline correction and / or normalisation. Examples of normalisation include vector normalisation, standard normal variate, min max normalisation, dimensionality reduction of the data through algorithms for example using Principal Component Analysis (PCA) or deep learning architectures such as autoencoders or variational autoencoders, and in some embodiments representation learning or synthetic data generation may be achieved using generative adversarial networks (GANs). Examples of baseline correction include the correction of systematic variations in the spectral data using Extended Multiplicative Scatter Correction (EMSC), rubber-band correction, asymmetric least squares (ALS) or any other variations of scatter correction algorithms.
[0071] DEMONSTRATIONS
[0072] Experiments were performed to demonstrate performance of the method of the disclosure.
[0073] These experiments used synthetic DNA samples including five different fragmented lengths: 50bp, lOObp, 150bp, 200bp and 300bp. Concentration of the samples were 0.5 pg / pL with a sample volume of 20pL each.
[0074] The protocol depicted in Figure 4 was used with a spectrometer 8 configured to perform mid-infrared ATR.
[0075] Step SI comprised cleaning of a surface of an ATR crystal of the spectrometer 8. This preparation of the ATR crystal aimed to eliminate any potential contaminants that could compromise the quality or interpretability of the spectra. This involved a sequential treatment with a DNA remover solution followed by an alcohol cleaning agent.
[0076] In step S2 a preliminary spectrometric single beam scan was performed with no sample on the ATR crystal. Subsequently, a preliminary single-beam background scan of the surroundings was conducted without introducing a sample. This baseline scan was later subtracted from subsequent scans to yield an almost featureless spectrum.
[0077] In step S3, a spectroscopic scan of the air background was taken three times. To ensure the absence of any residual cleaning agents on the ATR crystal, three additional scans were performed without introducing a sample.
[0078] In step S4, a 3 / / L volume of sample 6 was pipetted onto the ATR crystal and allowed to dry. Drying is desirable because of the Mid-IR regime's susceptibility to water absorption, which could mask spectral regions if drying were not performed.
[0079] In step S5, a set of 15 spectroscopic scans of the sample 6 were performed.
[0080] Step S6 comprised repeating steps S1-S5 three times for each of the different DNA fragment lengths. Repeating three times for each DNA fragment length promotes data integrity for the machine learning model (see below).
[0081] In step S7, the spectral data obtained in the preceding steps was exported and pre- processed for the construction and evaluation of the machine learning model. The preprocessing in this example comprised application of a minimum-maximum baseline correction, combined with spectral smoothing and taking the second-order derivative using the Savitzky-Golay filter.
[0082] The averaged spectral data of the five different DNA fragment lengths are shown in Figure 5. Figure 5 is a graph depicting absorbance in arbitrary units on the vertical axis and wavenumber on the horizontal axis. 225 spectra from five different fragment DNA base-pair lengths are shown, with different groups being indicated by different greyscale shading and labelled as follows: 50bp -“1050”; lOObp - “1100”; 150bp - “1150”; 200bp - “1200”; 300bp - “1300”. Table 1 shows assignments for absorption peaks in the spectra shown in Figure 5. Table 1:
[0083] Peak Assignment
[0084] 810 - 900 cm'1Deoxyribose ring vibration and main S-type sugar marker
[0085] 960 cm1O-P-O bending
[0086] 1080 cm1Symmetric stretching / ’Of
[0087] 1220 cm1Anti-symmetric stretching PO2, B-form double helix
[0088] 1400 cm'1Ring stretching vibrations, CH in plane bending
[0089] 1530 cm1Cytosine, Adenine
[0090] 1600 c Adenine
[0091] 1706 c ymine
[0092] 2370 c >
[0093] 2889 cm"'1C-H stretching
[0094] 2940 cm1( ’-II stretching
[0095] 3201 - 3216 cm"1Stretc hing of N-II symmetric. (Ml symmetric stretching
[0096] Boxes Data 1 and Data 2 in Figure 4 depict splitting of the exported and pre- processed data into a testing portion comprising 20% of the data (Data 1) and a training portion comprising 80% of the data (Data 2). In step S8, the machine learning model was created based on the training data (Data 2). Step S9 comprised cross-validation. Step S10 comprised testing of the model created in Step S8 using the testing data (Data 1).
[0097] The inventors assessed multiple machine learning algorithms for effectiveness at obtaining the requirement information from the spectral data. The assessed machine learning algorithms included random forest (RF), decision tree, partial least squares regression (PLSR), principal component regression (PCR), support vector regression (SVR), K-nearest neighbours (KNN) and artificial neural networks (ANN). The inventors determined from this assessment that PLSR and PCR consistently delivered superior results and that ANNs produced promising results in predicting the base-pair length of randomly selected DNA fragments when presented to the models. It is believed that the success may be due to the inherent linearity of the relationship between the data, which complements well with the ease of interpretability, speed and simplicity.
[0098] Figures 6 and 7 depict regression performance respectively for PLSR and PCR, with R2and MSE based on the predicted DNA base-pair length. These graphs show that the PLSR model performed extremely well, with high precision (R2= 0.997, RMSE = 16.71bp), followed closely by the PCR model (R2= 0.995, RMSE = 23.09bp). Spectroscopic measurements using the method of the present disclosure were also conducted on mixtures containing two different lengths of DNA fragments. Example results from such measurements are shown in Figure 8. This graph depicts prediction % of molecules with a fragmentation length (vertical axis) against the % split of DNA molecules between 150bp fragments (top horizontal axis) and lOObp fragments (bottom horizontal axis). Points corresponding to lOObp fragments are shaded as per the example labelled
[0099] 3001. Points corresponding to the 150bp fragments are shaded as per the example labelled
[0100] 3002. The measurements demonstrate the ability to predict proportions of fragments with varying lengths within complex mixtures resembling the diverse distribution of DNA fragment lengths typically found in blood samples.
[0101] DNA methylation plays an important role in the progression of cancer by intricately modulating gene expression patterns. Specifically, in cancer cells, the levels of methylation on genes governing cell growth regulation are altered, resulting in increased methylation on genes that suppress cell growth and decreased methylation on genes that facilitate uncontrolled proliferation. These alterations in DNA methylation status can be effectively detected in blood samples through vibrational spectroscopy techniques.
[0102] The inventors have demonstrated that DNA molecules exhibiting distinct states of unmethylated, methylated, and hydroxymethylated characteristics exhibit unique spectral fingerprints, as illustrated in Figure 9 (where curve 4001 represents hy dr oxym ethylated, curve 4002 represents methylated, and curve 4003 represents unmethylated). Notably, within the spectral range of 3300-2800cm-1, which corresponds to the vibration of C-H and O-H bonds, significant structural disparities are discernible. Furthermore, a discernible peak at 1660cm-1, associated with cytosine, experiences a distinct redshift (1657 / 1654cm-1) when cytosine undergoes methylation or hydroxymethylation. Given that DNA methylation levels vary between tumour and normal cells, these findings highlight the potential of multivariate spectroscopy as a powerful tool for detecting tumour-specific DNA methylation patterns in blood samples. Furthermore, by using the unique spectral signatures associated with differing methylation states, this approach offers promise for non-invasive and sensitive detection of cancer-related DNA methylation alterations, thereby facilitating early diagnosis and personalised treatment strategies. The inventors have developed a method able to quantify an amount of methylated DNA in a sample. For example, the method can quantify the methylation percentage of cfDNA derived from the adenomatous polyposis coli (APC) promoter region. Using a similar machine learning protocol to that discussed above, the inventors trained a machine learning model using fully methylated, fully hydroxymethylated and unmethylated versions of cfDNA, creating controlled mixtures to simulate intermediate methylation percentages. This allowed accurate detection and quantification of the level of methylation (i.e., number of 5mCs, 5hmCs) in the cfDNA.
[0103] The tables below detail the specific proportions used to create each mixture: Table 2:
[0104] Table 3:
[0105] Fourier Transform Infrared spectroscopy (FTIR) measurements were collected using the Agilent Cary 670 spectrometer equipped with mercury cadmium telluride (MCT) detector. The primary ATR attachment was the MIRacle diamond coated ZnSe crystal with a nine bounce configuration. All spectra were collected at a resolution of 4cm- 1, averaged over 64 accumulations for each sample scan. The spectral range was set from 6000cm- 1 to 700 cm-1.
[0106] For each sample measurement, 4pL of DNA was pipetted directly onto the crystal and left to dry under ambient laboratory conditions. Drying was monitored by observing changes in the spectral region around -3400 cm corresponding to water absorption. A stable spectrum, indicative of complete drying, was consistently achieved after approximately 15 minutes. Once dried, nine consecutive scans were recorded per sample, and each sample was measured in triplicate, resulting in 162 spectra for the methylated DNA experiment and 162 spectra for the hydroxy methylated experiment.
[0107] A PLSR model was selected for the quantification of both methylation and hydroxymethylation percentage, with the dataset split into 80% for training and 20% for testing. To enhance generalization and reduce the risk of overfitting, 10-fold cross- validation was employed during model evaluation and Van der Voet’s F-test was used to determine the optimal number of components.
[0108] Figures 20 and 21 illustrate the predictive performance of the PLSR model in independently quantifying the methylation and hydroxymethylation percentages respectively in a APC cfDNA sample. For methylation quantification, the model achieved an R2of 0.9938 with a RMSE of 2.618%. In the case of hydroxymethyl ati on, the model yielded an R2of 0.9927 and an RMSE of 2.9223%, demonstrating high predictive accuracy for both modifications.
[0109] The experiment was subsequently extended to include DNA mixtures containing both methylated and hydroxymethylated cfDNA. The unmethylated portion consisted of the previously used cfDNA, supplemented with small amounts of DNA fragments ranging from 50 to 300 bp. Table 4 below presents the proportions used in the experimental evaluation:
[0110] Spectral data from previous experiments, in which methylation and hydroxymethylation levels were independently varied, were integrated with the spectra collected from the DNA mixtures where both levels vary in the same sample. The objective was to develop a model capable of quantifying both methylation and hydroxymethylation percentages in mixtures containing varying proportions of each.
[0111] Figure 22 illustrates the quantification of both methylation and hydroxymethylation percentages in complex mixtures containing varying amounts of each component. The model’s performance metrics demonstrate its ability to accurately evaluate these percentages in samples designed to closely resemble real world conditions.
[0112] This approach may be usefully applied to cancer diagnostics, as APC gene hypermethylation (5mC) is a well-established biomarker in colorectal, breast, and gastric cancers. Importantly, hydroxymethylation (5hmC) is known to play a role in epigenetic regulation and may be altered in tumours, particularly in early stage or low grade cancers. By resolving both epigenetic marks simultaneously, the method of the present disclosure may offer deeper insight into methylation regulation in cfDNA
[0113] Figure 10 demonstrates how improved sensitivity is achieved using a multi -bounce spectrometer 8 (configured to use 9 bounces in this particular example) compared to the one-bounce spectrometer configuration used to obtain the results discussed above with reference to Figures 6 and 7.
[0114] Building on the demonstrations discussed above that used DNA fragments of specific lengths — 50bp, lOObp, 150bp, 200bp, and 300bp, the inventors have also created DNA mixtures containing the same synthetic fragments in various proportions. Using a similar machine learning protocol to that discussed above, this model was trained to predict not only fragment length but also the proportion within the mixture. The mixtures created had the following proportions:
[0115] Base-Pair 50bp lOObp ISObp 200bp 300bp
[0116] Length
[0117] Mixture 1 0.2 0.3 0.1 0.4 0.0
[0118] Mixture 2 0.1 0.1 0.4 0.2 0.2
[0119] Mixture 3 0.3 0.2 0.3 0.1 0.1
[0120] Mixture 4 0.25 0.25 0.25 0.25 0.25
[0121] Mixture 5 0.1 0.4 0.1 0.3 0.1
[0122] Figures 11-15 show prediction plots for the proportions of each DNA fragment length in the mixtures 1-5: Figure 11 corresponds to 50bp; Figure 12 corresponds to lOObp; Figure 13 corresponds to 150bp; Figure 14 corresponds to 200bp; and Figure 15 corresponds to 300bp.
[0123] Figures 16 and 17 remodel the data of Figures 11-15 in box plots and violin plots to provide further insight into data characteristics.
[0124] Additionally, grouping the data between 50-150 bp as <150bp and 200-300 bp as >150bp could be favoured since not all base-pair lengths in an unknown cohort will be identified, they can only be presumably predicted as a certain length. In this case 150bp is set as a threshold of interest, as typically ctDNA detection is enriched at lengths between 90-150bp. Figure 18 depicts a box plot with the grouping of datapoints collected below 150bp and above 150bp.
[0125] Figure 19 depicts base-pair length distributions generated by a model from tested DNA mixtures, wherein datapoints in the range of 50-150bp are grouped as <150bp and datapoints beyond 150bp as >150bp, the two groups being modelled as distribution plots (plot 5001 corresponding to <150bp and plot 5002 corresponding to >150bp), the outputs showing that the <150bp group is shifted towards shorter base-pair length, while the >150bp group tends towards longer base-pair lengths. Building on the demonstrations discussed above that used DNA mixtures containing the same synthetic fragments in various proportions, the inventors demonstrated the methods of the disclosure on an increased number of mixtures containing various proportions of samples. Using a similar machine learning protocol to that discussed above, models were again trained to predict fragment length and the proportions within the mixture. The mixtures created had proportions as shown in Table 5.
[0126] Table 5:
[0127] Mixtures with IDs 1-26 were used to train and test the machine learning models. To avoid data leakage and ensure a reliable evaluation, the data was split by repeat sets: spectra from the 1st and 2nd repeats were used only for training, while spectra from the 3rd repeat were reserved exclusively for testing. Within these sets, it was ensured that the training to test data split was maintained for every code run at 80% to 20% respectively.
[0128] In addition, mixtures with IDs Ul, U2, U3, and U4 were created to be used as an independent test set. These were specifically designed to challenge the model’s ability to quantify DNA fragment proportions in previously unseen combinations. Samples Ul and U3 were “top-heavy,” containing a higher proportion of longer fragments (200 bp and 300 bp), >150bp cutoff . Conversely, samples U2 and U4 were “bottom -heavy,” enriched with 150 bp and below, <150bp cutoff.
[0129] To assess the generalization performance of each machine learning model and diagnose issues such as overfitting or underfitting, 10-fold cross-validation was employed. This involved partitioning the training data into ten subsets, iteratively training on nine and validating on the remaining one, ensuring robust performance evaluation across the spectral feature space.
[0130] To identify the optimum number of PLS components needed for the model, Van der Voet' s F test was used. This method compares the CV prediction errors between consecutive models built with different PLS components. The test calculates an F-statistic by comparing the squared difference in PRESS (Prediction Error Sum of Squares) values between models with k and k+1 components, normalized by the estimated variance of this difference across CV folds. This approach accounts for the correlation between cross- validation folds, providing a more statistical based assessment than an inspection of validation curves. The optimal number of components is determined at the first nonsignificant comparison (p > 0.05), indicating that adding an additional component does not provide a statistically meaningful improvement in predictive performance. This hypothesis testing framework helps prevent overfitting while ensuring that the selected model complexity is justified by the data.
[0131] In this test, a PLSR model served as a baseline for evaluating the performance of fragment length quantification in DNA mixtures. Figure 23 shows a plot displaying the predicted proportions of individual fragment lengths generated by the model versus their corresponding true proportions in the test set. A strong correlation is shown for the predictions in each fragment length. These plots show clear ability of the PLSR model to both distinguish and quantify varying fragment compositions across a range of DNA mixtures, providing a benchmark for comparison against more complex models.
[0132] A stacked ensemble approach was also tested by the inventors. This comprised multiple regression models as base learners and a feedforward neural network as the meta- leamer, achieving improved performance on the dataset. In the architecture tested, each base learner generates predictions on the original data. These predictions were then concatenated to form a meta feature matrix, which serves as input to the neural network. The feedforward network introduces additional non-linear transformations, learning to combine the base predictions and produce the final output in its last layer. This is schematically depicted in Figure 24.
[0133] In an implementation of the feedforward network, the spectral input assumes the relationship for the input matrix XGRn x dwhere n = no. of samples and d = no. of features. The target labels are y E Rn. The base learners Mi, M2, M3, M4, . . . MK are then trained on X and y. For each base learner a prediction is made Pi = Mi(X) E Rn, P2 = M2(X) E Rn, P3 = Ms(X) E Rn, P4 = Mi(X) E RnPK = MK(X) E Rn. These features are concatenated into a stacked input feature matrix Z = E Rn x k. Finally, a meta model Mmeta makes predictions denoted by y = Mmeta(Z) E Rn.
[0134] Each base learner in the stacked ensemble was evaluated using 10-fold cross- validation, repeated multiple times to minimize the risk of overfitting and ensure robust performance estimates. For XGBoost and SVR, moderate hyperparameter tuning was performed using a grid search strategy to optimize model performance.
[0135] The inventors believe that PLSR, PCR, ridge regression and lasso regression may capture linear based relationships in the data, while SVR and RF may capture non-linear relationships in the data. Several different gradient boosting ensemble algorithms were also tested such as XGBoost, Adaboost and CatBoost utilizing decision trees to capture the features of the data.
[0136] In the feedforward neural network used as the meta-leamer, a validation split of 20% was applied to the meta-feature matrix. This allowed monitoring of both training and validation loss, with mean squared error (MSE) used as the loss function. Validation curves were closely tracked during training to ensure the model did not overfit and to guide early stopping or model checkpointing as needed. The meta learner’s loss can be seen in Figure 25, with the purpose of the plot showing the convergence between the training loss and the validation loss indicating the model is learning well and not overfitting.
[0137] The individual prediction results for each of the base learners, as well as the stacked ensemble are shown in Figures 26a and 26b.
[0138] These data illustrate that the stacked ensemble model outperformed the standalone PLSR model in predicting DNA fragment length proportions. While the PLSR model achieved an R2of 0.8557 and an RMSE of 7.2750%, the stacked ensemble improved these metrics to an R2of 0.9115 and a reduced RMSE of 5.5874%, indicating enhanced predictive accuracy and lower error.
[0139] Other ensemble methods such as stacking different linear based regressors such as PLSR, Ridge and Lasso also showed an improvement in the performance metrics R2 and MSE compared to the use of the models in a standalone format specifically for the DNA mixtures containing all five DNA fragment lengths. Convolutional neural networks (CNN) and artificial neural networks (ANN) showed strong regression performance for the mixtures that contained all five DNA fragments.
[0140] The inventors have also demonstrated that a ID convolutional neural network (CNN) can be used for of quantifying DNA fragment length proportions in a mixture according to methods of the disclosure. The inventors have found that CNNs are particularly well-suited for this task due to their ability to automatically extract meaningful features from sequential data such as infrared spectra, leveraging local receptive fields via convolutional filters. Unlike traditional models that depend on extensive manual preprocessing or feature engineering, CNNs may be able to learn to highlight the most discriminative spectral patterns directly from raw input, capturing subtle variations in absorbance associated with different fragment lengths. Furthermore, the convolutional architecture effectively reduces noise while preserving important local dependencies across the spectral domain, which is essential for precise quantification. By combining multiple convolutional and pooling layers followed by fully connected dense layers, the network learns complex nonlinear mappings between spectral signatures and fragment proportions. This makes CNNs especially advantageous in clinical or high throughput settings, where minimal preprocessing is needed and rapid, automated analysis is required.
[0141] The inventors have found that, using their adaptability, deep learning models such as ANNs and CNNs excel in contrast to PLSR and PCR in capturing nonlinear relationships between features by autonomously discerning and learning the characteristics present in datasets. This proficiency is highly desirable, particularly in scenarios requiring the direct integration of additional objectives, such as analysing methylation percentages, measuring cancer biomarker concentrations and detecting mutations within patient samples. Deep learning models such as ANNs and CNNs also provide good performance when used on more complex mixtures containing multiple DNA fragment lengths, as demonstrated below.
[0142] To accommodate the increased complexity of the neural network and ensure the model could effectively learn underlying patterns, the training dataset was expanded through data augmentation. Three augmentation strategies were applied, increasing the dataset size sevenfold. As a result, the input spectra matrix grew from Xoriginai to XAugmented, where XAugmented Xoriginai + Uaugmentations (Xoriginai).
[0143] Figure 27 shows the simulation of realistic variability in spectral measurements. In order to make the training data reflect realistic data the following augmentation strategies were applied: baseline shift, intensity scaling, and noise addition. These are illustrated in Figure 27. Baseline shifts were introduced by adding a random offset drawn from a uniform distribution within a specified range, effectively shifting the entire spectrum vertically to mimic experimental baseline drift. Intensity scaling was applied by multiplying each spectrum by a random factor sampled from a uniform range, representing fluctuations in signal strength due to concentration differences or instrument sensitivity. Additionally, random noise was added to each point in the spectrum using values drawn uniformly from a symmetric range, emulating the presence of measurement noise. These transformations were applied independently and randomly to each spectrum.
[0144] The model architecture shown in Figure 28 is an example CNN for use in methods of the disclosure. In some implementations, the CNN may be a deep or shallow ID CNN tailored for spectral data analysis and the quantification of DNA fragment length proportions. The CNN of Figure 28 begins with an initial ConvlD layer, followed by two stacked residual blocks, each composed of multiple ConvlD layers with ELU activation functions and skip connections. These residual connections help retain early spectral features and improve gradient flow during training. Max pooling layers are used between blocks to reduce spatial dimensionality while preserving important information. The model incorporates an attention mechanism through a Squeeze-and-Excitation (SE) block, which adaptively reweights channel wise feature responses. This is achieved by first applying global average pooling to each feature map (squeeze), then passing the result through a small fully connected bottleneck network (excitation) that learns the importance of each channel. The learned weights are used to scale the original feature maps, allowing the model to emphasize relevant spectral patterns and suppress less the informative ones. After attention, a multi-scale feature extraction module inspired by the Inception architecture captures spectral features at varying resolutions. The extracted features are flattened and passed through three fully connected dense layers, producing five output nodes that represent the predicted proportions of DNA fragments at 50 bp, 100 bp, 150 bp, 200 bp, and 300 bp.
[0145] During CNN training, the dataset was further divided to include a validation set. Performance metrics, including loss and mean absolute error (MAE), were tracked across training epochs and stored as part of the model’s training history. As shown in Figure I, both training and validation loss plateau around 70 epochs, suggesting a point of convergence suitable for early stopping. However, the validation MAE continues to improve up to approximately 200 epochs. This reveals a trade off between training efficiency by stopping at the loss convergence and maximum predictive performance by training until MAE convergence.
[0146] Figure 29 shows the predictive performance for the CNN model to predict the proportion in the DNA mixtures tested.
[0147] During CNN training, the dataset was further divided to include a validation set. Performance metrics, including loss and mean absolute error (MAE), were tracked across training epochs and stored as part of the model’s training history. As shown in Figure 30, both training and validation loss plateau around 70 epochs, suggesting a point of convergence suitable for early stopping. However, the validation MAE continues to improve up to approximately 200 epochs. This reveals a trade off between training efficiency by stopping at the loss convergence and maximum predictive performance by training until MAE convergence.
[0148] The flowcharts of Figures 31 and 32 illustrates the practical advantage of using a CNN model in clinical environments, where complex preprocessing can introduce dataset dependent parameters and reduce generalizability. Figure 31 shows an example of stages of processing which may be used for a CNN model, and Figure 32 shows the stages of processing for traditional machine learning models. Traditional machine learning models may perform well when training and testing conditions closely match, but this assumption often fails in real world scenarios. In contrast, CNNs are more robust to experimental variability, such as artifacts and distortions commonly encountered during testing, while still accurately extracting the underlying features needed for reliable predictions.
[0149] In contrast utilizing the traditional machine learning models, either in the ensemble or individually the workflow is shown in Figure 32, where an additional preprocessing step is required which is dataset dependent.
[0150] The inventors have also demonstrated that using both infrared spectroscopy (for example, FTIR) and Raman spectroscopy in tandem can provide improved characterisation. Infrared spectroscopy is particularly sensitive to vibrations involving dipole moment changes, while Raman excels at detecting modes associated with changes in polarizability. By using both methods in tandem, the inventors have found that a more comprehensive vibrational fingerprint of the sample may be captured, covering functional groups and structural changes that might otherwise be missed by either method alone. This combined approach may enhance both the sensitivity and specificity of molecular analysis, which may be especially advantageous for complex biological samples where overlapping signals, background fluorescence, or matrix effects can obscure subtle but diagnostically relevant features.
[0151] Initially, averaged Raman spectra for each DNA fragment length were obtained by computing the mean count for scans from all the repeats within that group. The spectra then from each modality were independently loaded and filtered to the biologically relevant region of 600-1800 cm then normalized to ensure comparable intensity scales. The IR spectra were retained on their physical axis, while the Raman spectra were shifted along the wavenumber axis by an amount equal to the fusion window plus any specified seam gap, to prevent spectral overlap. The two datasets were then concatenated into a single fused spectrum, with IR followed seamlessly by the shifted Raman segment. To mitigate data leakage, the IR training replicates were fused with the averaged Raman training spectra, while the IR test replicates were fused with the averaged Raman test spectra. For traceability, the original Raman shift values were preserved in a separate column. The fused spectra were visualized as concatenated plots in Figure 33, yielding unified datasets that capture complementary vibrational information from both modalities.
[0152] Machine learning models were trained to perform the fragment length quantification on the FTIR and Raman datasets alone, and the combined datasets. The same 80% training and 20% testing split was retained for consistency across all analyses. A fusion strategy was implemented through concatenating the preprocessed IR and Raman spectral blocks into a single feature matrix used to train the model. Other fusion strategies may also be employed, including mid-level fusion, in which latent features extracted from each modality are combined, and high-level fusion, in which outputs from separately trained models are integrated. The optimal number of latent variables was determined through k-fold cross-validation and f-test on the training set. Model performance was subsequently evaluated on the independent 20% test set to assess the predictive performance.
[0153] Figures 34 and 35 show the predictive performance for the trained models at predicting the proportions of DNA base-pair lengths in the mixtures tested for the individual FTIR and Raman datasets respectively. Figure 36 shows the predictive performance of the model trained on the combined dataset. While the individual PLSR models constructed from IR and Raman spectra demonstrated good predictive capability for DNA fragment length, with R2 = 0.9268 and RMSE = 24.14 bp for IR, and R2 = 0.9113 and RMSE = 24.44 bp for Raman. While both modalities captured relevant spectral variations associated with fragment size, each exhibited moderate variance around the regression line. In comparison, the fused model integrating IR and Raman spectra achieved significantly improved performance with R2 = 0.9773 and RMSE = 13.45 bp. This performance increase reflects the complementary nature of the two vibrational techniques: IR provides sensitivity to backbone and functional group vibrations, whereas Raman captures conformational and base related features. Therefore, in combination they deliver a more comprehensive molecular fingerprint, resulting in enhanced robustness and accuracy for DNA fragment length quantification. Beyond PLSR, the application of alternative machine learning approaches, such as methods that utilise complex nonlinear dependencies within and between the modalities can also be applied for model performance gains. Additionally, more advanced architectures, including deep learning models such as CNNs, are particularly well suited to automatically extract hierarchical and cross-modality features that may not be captured by linear methods. Using such approaches may detect subtle interactions between Raman and IR signals that contribute to the underlying signatures of DNA fragment length.
[0154] The following numbered clauses define embodiments of the present disclosure:
[0155] 1. A method of analysing a sample containing DNA, comprising: performing vibrational spectroscopy on the sample to obtain vibrational spectroscopy data, the vibrational spectroscopy comprising directing electromagnetic radiation onto the sample and detecting electromagnetic radiation emitted from the sample after interaction with the sample; and analysing the vibrational spectroscopy data to obtain information about a distribution of DNA fragment sizes from the obtained vibrational spectroscopy data.
[0156] 2. The method of clause 1, wherein the obtained information includes a median of the distribution.
[0157] 3. The method of clause 1 or 2, comprising detecting the presence of circulating tumour DNA using the obtained information about the distribution of DNA fragment sizes.
[0158] 4. The method of clause 3, wherein: the obtained information includes a median of the distribution; and the detecting of the presence of circulating tumour DNA comprises determining whether the median of the distribution is below a predetermined threshold.
[0159] 5. The method of any preceding clause, comprising determining a concentration of circulating tumour DNA using the obtained information about the distribution of DNA fragment sizes.
[0160] 6. The method of clause 5, wherein: the obtained information includes a median of the distribution; and the determining of the concentration uses a predetermined calibration relationship between the median of the distribution and the concentration.
[0161] 7. The method of any preceding clause, further comprising analysing the vibrational spectroscopy data to obtain information about methylation of DNA fragments in the sample.
[0162] 8. A method of analysing a sample containing DNA, comprising: performing vibrational spectroscopy on the sample to obtain vibrational spectroscopy data, the vibrational spectroscopy comprising directing electromagnetic radiation onto the sample and detecting electromagnetic radiation emitted from the sample after interaction with the sample; and analysing the vibrational spectroscopy data to obtain information about methylation of DNA fragments in the sample.
[0163] 9. The method of any preceding clause, wherein the information is obtained at least partly using a trained machine learning model.
[0164] 10. The method of clause 9, wherein the trained machine learning model comprises one or more of the following: partial least squares regression; principal component regression.
[0165] 11. The method of clause 9 or 10, wherein the vibrational spectroscopy data represents a variation of absorbance as a function of frequency, wavelength, or wavenumber.
[0166] 12. The method of clause 11, wherein: using the trained machine learning model comprises preprocessing the vibrational spectroscopy data and inputting the preprocessed vibrational spectroscopy data into the trained machine learning model; and the preprocessing comprising applying a smoothing filter with second-order derivatives to the vibrational spectroscopy data.
[0167] 13. The method of any preceding clause, wherein the vibrational spectroscopy comprises infrared spectroscopy.
[0168] 14. The method of any preceding clause, wherein the vibrational spectroscopy comprises attenuated total reflection spectroscopy.
[0169] 15. The method of any preceding clause, wherein the vibrational spectroscopy comprises Raman spectroscopy.
[0170] 16. The method of any preceding clause, wherein the sample comprises or is derived from a blood sample and / or tissue sample and / or saliva and / or cerebrospinal fluid and / or urine.
[0171] 17. The method of clause 16, wherein the sample comprises plasma DNA or serum DNA, separated from a blood sample.
[0172] 18. The method of any preceding clause, wherein the sample is provided in a disposable sample receiver.
[0173] 19. A system for analysing a sample containing DNA, comprising: a sample receiver configured to receive the sample; a spectrometer configured to perform vibrational spectroscopy on the sample and output vibrational spectroscopy data, the vibrational spectroscopy comprising directing electromagnetic radiation onto the sample and detecting electromagnetic radiation emitted from the sample after interaction with the sample; and a data processor configured to analyse the vibrational spectroscopy data to obtain information about a distribution of DNA fragment sizes from the obtained vibrational spectroscopy data.
[0174] 20. A system for analysing a sample containing DNA, comprising: a sample receiver configured to receive the sample; a spectrometer configured to perform vibrational spectroscopy on the sample and output vibrational spectroscopy data, the vibrational spectroscopy comprising directing electromagnetic radiation onto the sample and detecting electromagnetic radiation emitted from the sample after interaction with the sample; and a data processor configured to analyse the vibration spectroscopy data to obtain information about methylation of DNA fragments in the sample.
[0175] BIBLIOGRAPHIC DETAILS
[0176] This application claims priority to GB 2414133.5, which is incorporated by reference herein.
Claims
CLAIMS1. A method of analysing a sample containing DNA, comprising: performing vibrational spectroscopy on the sample to obtain vibrational spectroscopy data, the vibrational spectroscopy comprising directing electromagnetic radiation onto the sample and detecting electromagnetic radiation emitted from the sample after interaction with the sample; and analysing the vibrational spectroscopy data to obtain information about a distribution of base-pair lengths of DNA fragments from the obtained vibrational spectroscopy data.
2. The method of claim 1, wherein the obtained information includes a median of the distribution.
3. The method of claim 1 or 2, comprising detecting the presence of circulating tumour DNA using the obtained information about the distribution of base-pair lengths of DNA fragments.
4. The method of claim 3, wherein: the obtained information includes a median of the distribution; and the detecting of the presence of circulating tumour DNA comprises determining whether the median of the distribution is below a predetermined threshold.
5. The method of any preceding claim, comprising determining a concentration of circulating tumour DNA using the obtained information about the distribution of base-pair lengths of DNA fragments.
6. The method of claim 5, wherein: the obtained information includes a median of the distribution; and the determining of the concentration uses a predetermined calibration relationship between the median of the distribution and the concentration.
7. The method of any preceding claim, further comprising analysing the vibrational spectroscopy data to obtain information about methylation of DNA fragments in the sample.
8. A method of analysing a sample containing DNA, comprising: performing vibrational spectroscopy on the sample to obtain vibrational spectroscopy data, the vibrational spectroscopy comprising directing electromagnetic radiation onto the sample and detecting electromagnetic radiation emitted from the sample after interaction with the sample; and analysing the vibrational spectroscopy data to obtain information about methylation of DNA fragments in the sample.
9. The method of any preceding claim, wherein the information is obtained at least partly using a trained machine learning model.
10. The method of claim 9, wherein the trained machine learning model comprises one or more of the following: partial least squares regression; principal component regression.
11. The method of claim 9 or 10, wherein the vibrational spectroscopy data represents a variation of absorbance as a function of frequency, wavelength, or wavenumber.
12. The method of claim 11, wherein: using the trained machine learning model comprises preprocessing the vibrational spectroscopy data and inputting the preprocessed vibrational spectroscopy data into the trained machine learning model; and the preprocessing comprising applying a smoothing filter with second-order derivatives to the vibrational spectroscopy data.
13. The method of any preceding claim, wherein the vibrational spectroscopy comprises infrared spectroscopy.
14. The method of any preceding claim, wherein the vibrational spectroscopy comprises attenuated total reflection spectroscopy.
15. The method of any preceding claim, wherein the vibrational spectroscopy comprises Raman spectroscopy.
16. The method of any preceding claim, wherein the sample comprises or is derived from a blood sample and / or tissue sample and / or saliva and / or cerebrospinal fluid and / or urine.
17. The method of claim 16, wherein the sample comprises DNA obtained from plasma or DNA obtained from serum, separated from a blood sample.
18. The method of any preceding claim, wherein the sample is provided in a disposable sample receiver.
19. A system for analysing a sample containing DNA, comprising: a sample receiver configured to receive the sample; a spectrometer configured to perform vibrational spectroscopy on the sample and output vibrational spectroscopy data, the vibrational spectroscopy comprising directing electromagnetic radiation onto the sample and detecting electromagnetic radiation emitted from the sample after interaction with the sample; and a data processor configured to analyse the vibrational spectroscopy data to obtain information about a distribution of base-pair lengths of DNA fragments from the obtained vibrational spectroscopy data.
20. A system for analysing a sample containing DNA, comprising: a sample receiver configured to receive the sample; a spectrometer configured to perform vibrational spectroscopy on the sample and output vibrational spectroscopy data, the vibrational spectroscopy comprising directingelectromagnetic radiation onto the sample and detecting electromagnetic radiation emitted from the sample after interaction with the sample; and a data processor configured to analyse the vibration spectroscopy data to obtain information about methylation of DNA fragments in the sample.
Citation Information
Patent Citations
Method and system for analysing a sample containing dna
GB202414133D0
Apparatus and method for early cancer detection and cancer prognosis using a nanosensor with raman spectroscopy
US20240132967A1
KR20240012517A