Detection of pathogens in biological fluid samples using spectroscopic techniques in conjunction with analytical models
By combining spectral imaging with machine learning, FTIR and Raman spectroscopy techniques are used to analyze non-invasive biological fluid samples, solving the problems of high invasiveness and long detection time in cervical cancer screening. This enables rapid and accurate pathogen detection, especially the identification of hr-HPV.
Patent Information
- Application Number
- CN202480046429.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-15
- Filing Date
- 2024-07-05
- Publication Date
- 2026-02-13
AI Technical Summary
Existing cervical cancer screening methods are highly invasive and time-consuming, while non-invasive sample testing methods lack accuracy and sensitivity, making it difficult to quickly and effectively detect pathogens such as high-risk human papillomavirus (hr-HPV).
This method combines non-invasive biological fluid samples, such as urine and saliva, with spectral imaging and machine learning techniques. Spectral data is generated through FTIR and Raman spectroscopy, and machine learning algorithms are used to analyze the spectral data to quickly identify specific pathogens, including human papillomavirus.
It enables high-precision pathogen detection in a short time without sample pretreatment, can identify single or multiple pathogens, improves the sensitivity and specificity of detection, and is suitable for resource-limited field care environments.
Smart Images

Figure CN121532640A_ABST
Abstract
Description
Cross-reference to related applications
[0001] This application claims priority to U.S. Provisional Application No. 63 / 525,697, filed July 9, 2023, and U.S. Provisional Application No. 63 / 553,862, filed February 15, 2024, both of which are incorporated herein by reference. Technical Field
[0002] The disclosed principles involve the detection of pathogens in biological fluid samples. More specifically, they involve combining spectral imaging with machine learning techniques on samples obtained from non-invasive biological fluids (such as urine, saliva, blood, etc.) to quickly and accurately identify the presence of specific pathogens. Background Technology
[0003] Cervical cancer is the fourth most common cancer among women worldwide. It is caused by high-risk human papillomavirus (hr-HPV). The World Health Organization (WHO) claims that cervical cancer can be eliminated through its three-pillar strategy: HPV vaccination; screening and treatment of precancerous lesions; and treatment of invasive diseases. HPV vaccination prevents infection but is ineffective for those already exposed to the virus. Therefore, screening will continue to play a crucial role in eliminating cervical cancer for the foreseeable future. In the UK, cervical screening has helped reduce cervical cancer deaths by 70% since its introduction. However, despite its success, participation in cervical cancer screening has declined over the years. In 2021, only 70.2% of eligible individuals in the UK were screened, a record low.
[0004] Barriers to cervical cancer screening include embarrassment, inconvenience, and discomfort. To address some of these barriers, vaginal self-sampling for HPV testing has been developed. Currently, in at least 17 countries, this has become an alternative method for cervical cancer screening, both for under-screened populations and as a primary screening option. The effectiveness of vaginal sampling in improving screening acceptance rates has varied across different studies, ranging from 6% to 30%. Urine self-testing procedures have been shown to have similar diagnostic accuracy to vaginal self-testing procedures, approaching that of cervical samples collected by clinicians. Urine testing is non-invasive, requiring no sample collection by healthcare professionals and no need for healthcare facilities to set up clinical examination rooms to provide cervical samples. Women perceive urine testing as less "invasive" than vaginal self-testing.
[0005] Improving the laboratory diagnostic capabilities of healthcare systems at the national and community levels is a crucial component of managing viral and bacterial infections. As discussed in detail below, spectroscopic techniques, including infrared (IR) and Raman spectroscopy, are expected to allow for the analysis of bacteria and viruses with good specificity and sensitivity.
[0006] Infrared spectroscopy (IR) is a vibrational spectroscopy method that relies on the absorbance, transmittance, or reflectance of IR light when a sample is illuminated. A specific type of IR spectroscopy, Fourier transform IR (FTIR) spectroscopy, is a preferred method because it can simultaneously collect data across the entire range of IR wavelengths of interest. FTIR is considered a powerful tool for chemical analysis because it provides detailed information at the molecular level about the chemical composition of components such as proteins, nucleic acids, carbohydrates, and lipids. Essentially, it provides a biochemical fingerprint of the chemical and biomolecular structures in various environments, such as biological fluids. IR spectroscopy utilizes the fact that chemical bonds or functional groups within molecules vibrate at characteristic frequencies. IR spectroscopy primarily deals with the IR region of the electromagnetic spectrum, focusing mainly on absorption, i.e., absorption occurs when the frequency of IR radiation matches the vibrational frequency of the bonds. The wavelengths at which a sample absorbs light, represented as bands in the spectrum, are characteristic of its molecular structure.
[0007] The use of IR methods in clinical practice is increasing due to their speed, cost-effectiveness, and accurate disease prediction. To date, IR spectroscopy of biological fluids has demonstrated high sensitivity and specificity in the diagnosis of dementia, brain cancer, and endometrial cancer in large cross-sectional studies. Other studies have demonstrated the potential of spectroscopy of liquid-based cytology samples for triage, stratifying HPV-positive women. Early proof-of-concept studies have suggested similar potential for urine in the detection of gynecological cancers.
[0008] Raman spectroscopy is a slightly different technique used to measure the relative frequencies of scattered radiation from biological fluid samples. Raman spectroscopy can be used directly on biological fluid samples in liquid form, while IR spectroscopy requires the biological fluid sample to be dried before analysis.
[0009] Although these types of spectroscopic techniques have been used to detect a variety of viruses (including hr-HPV) based on cervical (i.e., invasive) samples, there is still a need to develop methods to apply these techniques to non-invasive samples (such as urine) to detect viruses such as, but not limited to, hr-HPV, and to achieve rapid, point-of-care testing turnaround times. Summary of the Invention
[0010] This paper discloses a system and method for performing pathogen detection by subjecting a sample of biological fluid to one or more spectral processes and then utilizing machine learning to rapidly obtain high-precision results. An exemplary system for obtaining non-invasive biological fluid samples and performing spectral analysis can be embodied in a portable component that allows for immediate results to be provided to a patient within minutes. The sample preparation process eliminates the need for subsequent sample processing (which could introduce contamination); instead, the collected biological fluid is used directly to perform the spectral procedures.
[0011] Machine learning algorithms are trained on labeled data, which allows for the analysis of submitted samples to identify specific biomarkers in pathogens of interest. Machine learning models can be embedded in components or accessed via secure communication links. Advantageously, data collected at each stage can subsequently be added (without patient IDs) to a database used during the machine learning process. Fourier transform infrared (FTIR) spectroscopy or Raman spectroscopy (or both) can be used to collect spectral data from IR-irradiated samples.
[0012] An exemplary embodiment of the present invention may take the form of a method for performing analysis on spectral data from a presented biological fluid sample to detect a specific pathogen, the method comprising the steps of: a) performing a non-invasive sample collection process to obtain selected biological fluids from a patient; b) preparing a controlled amount of the selected biological fluids on a cartridge for further processing; c) irradiating the prepared sample with a spectrometer and generating spectral data therefrom; and d) applying one or more machine learning models trained for detecting pathogens in biological fluids to analyze the spectral data to detect the presence or absence of a specific pathogen in the generated spectral data.
[0013] Another exemplary embodiment may include a system diagnostic system for analyzing spectral data from biological fluid samples to detect specific pathogens. The system includes: (1) a sample collection and preparation apparatus for collecting biological fluid samples; (2) a spectral apparatus for irradiating the acquired biological fluid samples and creating spectral data from the irradiation; and (3) a machine learning model processing unit responsive to the acquired biological fluid samples, wherein the machine learning model processing unit includes one or more trained models to detect the presence or absence of selected pathogens in the acquired spectral data.
[0014] According to some embodiments, the system can be configured as a portable component of a sample acquisition and preparation device combined with a spectrometer, using the computer processor functionality in the portable system or a local computer processor (possibly a tablet computer) to perform diagnostics on the collected samples.
[0015] According to some implementation methods, several samples can be collected simultaneously and analyzed independently to reduce the probability of errors.
[0016] According to some embodiments, machine learning models are collected and stored in a non-transitory computer-readable storage medium that also stores instructions. When executed by the system, these instructions cause the system to process the collected spectra and then execute a selected machine learning model to analyze the presented spectral data.
[0017] Other and further embodiments and features of the invention may become apparent during the following discussion and with reference to the accompanying drawings cited therein. Attached Figure Description
[0018] Now refer to the attached diagram,
[0019] Figure 1 A simplified diagram of a system for pathogen detection according to the principles of the present invention;
[0020] Figure 2 This is a block diagram of a spectrometer assembly according to the present invention for collecting and preparing biological fluid samples, acquiring spectral data, and applying machine learning models to the spectral data to perform pathogen detection;
[0021] Figure 3 It is relative to Figure 2 The block diagram shows an alternative spectrometer assembly to the spectrometer assembly shown, wherein, in this case, a Raman spectroscopy process is used instead of the one described above. Figure 2 Device-related FTIR spectroscopy processes;
[0022] Figure 4 It is a partial least squares (PLS) decomposition graph of the spectral signal obtained from urine samples with known hr-HPV status without derivation, based on a selected set of machine learning models.
[0023] Figure 5 Indicates and Figure 4 The same data is shown, but in this case, the data is presented in the form of the third derivative, which produces improved information about the trend relative to the original data.
[0024] Figure 6 It is a graph of the processed and enhanced spectral signal used as input for training machine learning models;
[0025] Figure 7 It shows in Figure 6 The training of a one-dimensional convolutional neural network (CNN) iteratively on the fingerprint region of the data shown; and
[0026] Figure 8 A block diagram of an example computer system is shown, illustrating some embodiments of a machine learning model that can be used to implement the detection of one or more pathogens in a presented sample. Detailed Implementation
[0027] Rapid and accurate detection and identification of pathogens, including viruses such as human papillomavirus (HPV), is crucial for infection management. Current methods are invasive and require relatively long detection times. Furthermore, the specificity and sensitivity of existing methods are relatively poor.
[0028] According to the principles of the present invention, as discussed in detail below, specific pathogens can be detected and identified from biological fluids (e.g., blood, saliva, urine, serum, plasma, etc.) obtained in a relatively non-invasive manner. The entire process, from sample acquisition to presentation of results to the patient, is expected to take less than thirty minutes. For example, this method can be used to examine the presence of hr-HPV in first-gap, self-collected urine samples via its biomolecular, cellular, and metabolic composition using an FTIR spectrometer that can be coupled with attenuated total reflectance (ATR), a sampling attachment, and possibly also Raman spectroscopy.
[0029] The method conceived according to the principles of this invention is considered a simple, non-invasive analytical technique that can characterize the biochemical properties of any biological fluid without requiring extensive sample preparation. The biological fluid itself undergoes FTIR and / or Raman spectroscopy to create spectral data. This data may contain certain resonance peaks associated with identified biomarkers, which will be identified by machine learning algorithms. In other words, the system and method of this invention allow for efficient and simple sample preparation, presenting the sample to a spectrometer, and, in conjunction with machine learning algorithms, analyzing the biomolecules and cells present in the biological fluid, thereby allowing for accurate, rapid, and simple detection of identified pathogens in biological fluid samples.
[0030] Furthermore, by employing machine learning algorithms, this technique can identify specific biomarker spectral peaks associated with the region of interest and can be used to analyze one or more biological fluids according to the principles of this invention. Moreover, for a single biological fluid sample, multiple machine learning algorithms can be used to identify multiple pathogens from a single spectrum.
[0031] The method of this invention can identify, for example, a single HPV strain, and can also identify other viruses or pathogens that may exist alone or in combination. For example, the identification of other viruses such as monkeypox, herpes simplex virus types 1 and 2, syphilis, and varicella-zoster virus can be performed using the system and method of this invention.
[0032] As described in detail below, the method of the present invention for spectral analysis of biological fluid samples is considered unique and novel because it emphasizes the identification of biomarker spectral peaks based on the biomolecular, cellular, and metabolic composition of the biological fluid sample, which may be composed of several known pathogens, each of which may differently affect the composition of the patient's biological fluids.
[0033] Furthermore, the method of this invention is considered unique because the machine learning model is used to select relevant signals in the spectrum (created by the spectral process) associated with a specific pathogen from among many other biochemical peaks in the spectral data that helps to recover from the patient's bodily fluids. The bodily fluids are sampled directly, without the need for biomolecular or cellular extraction, and the focus is on identifying specific peak features of infection, distinguishing them from normal, uninfected bodily fluid samples. A spectral database of various types of infected and normal bodily fluid samples is used as a supervised dataset for developing and training models to rapidly and accurately identify samples of unknown pathogens. Chemometric analysis utilizing neural networks and machine learning, coupled with the automation of the bodily fluid collection process, provides rapid and accurate detection and identification.
[0034] The method of this invention is considered to surpass methods that compare test spectra with a database of known spectra because the machine learning model employed is able to learn subtle signals in the spectrum that directly indicate the presence of a specific pathogen in a patient's bodily fluid sample, without requiring further separation of the sample to the cellular level before obtaining the spectrum. In other words, using samples directly from bodily fluids, rather than first isolating the biomolecule of interest, means the model needs to be able to find signals among the peaks of many other biomolecules in the bodily fluid sample. This can be done more accurately and efficiently by a machine learning model that learns from a large dataset of diverse sample spectra. The machine learning model also allows for improvement over time as more data (potentially from a wide variety of sources) is used for additional training, enabling it to learn more patterns from samples collected in various settings. Otherwise, this would be impossible, or more time-consuming, to simply compare test spectra with a database of known spectra.
[0035] The following discussion will describe the overall system and methods for performing pathogen detection, and then describe in detail the individual components, including: (1) the sample preparation process; (2) the spectral data collection process; and (3) the analysis of the collected spectral data using machine learning algorithms.
[0036] Figure 1 This is a simplified diagram of these components presented as a flowchart. The process begins with the collection of a biological fluid sample (step 100), where the biological fluid to be used is envisioned to be collected in a relatively non-invasive manner. In particular, the biological fluids preferably used for analysis include, but are not limited to, urine, saliva, blood, serum, plasma, etc. In fact, in many cases, the patient is able to collect the sample themselves, and the sample is stored in a contaminated container (step 110). Once the sample is obtained, a small portion (a known controlled volume) is loaded (step 120) and may be dried in the manner known for preparing slides for spectroscopic analysis (step 130).
[0037] In practice, the sample analysis performed in step 200 can take the form of Raman spectroscopy (step 210a) and / or FTIR spectroscopy (step 210b). These processes are complementary in the methods used to acquire spectral data, and therefore, using both techniques may further improve the reliability of the results. As discussed in detail below, one difference between the two methods is that FTIR spectroscopy requires a dried sample, while Raman spectroscopy can be used directly for liquid forms of biological fluids (as well as dried samples on a base).
[0038] As shown in step 220, the digital spectral image data generated by the spectral processing in step 210 is then presented to a set of machine learning models (MLMs) for analysis. The flowchart identifies a selected number of example MLMs, including Support Vector Machines (SVMs), Linear Discriminant Analysis (LDA), Logistic Regression, and Artificial Neural Networks (ANNs). These are just a selection of various MLMs that may be used; in fact, several different models can be used to analyze the same spectral data. As will be discussed below, the use of these MLMs allows for the evaluation of the presented data and provides results quickly and efficiently (as shown in step 300); in many cases, this takes approximately five minutes. Figure 1 As shown in step 400, the analysis results can be stored (after deleting any patient identification data) and used for further updating and training of the dataset.
[0039] Now, referring to the details of sample preparation and spectroscopic procedures, Figure 2 This is a simplified block diagram of a spectrometer assembly 10 that can be used to prepare and analyze biological body fluid samples according to the present invention. Figure 2 The image shows a standard biological fluid container 12 used by patients or healthcare professionals to collect samples. The biological fluid container 12 is inserted into a cartridge 14 (typically reusable) that engages with the rest of the spectrometer assembly 10. The assembly 10 is intended to handle multiple such cartridges, providing a compact arrangement. The cartridge 14 is configured to automatically dispense a precise volume of biological fluid (e.g., five microliters) onto a clean, sterile glass slide 16. The slide 16 may be made of CaF2, and the sample is spread into a small circle. In one embodiment, this process is repeated to create multiple sampling points on the slide 16 for multiple measurements on the same slide (minimizing error and improving the signal-to-noise ratio). Alternatively, the sample can be deposited on a diamond crystal for ATR analysis.
[0040] The box system includes an internal container for collecting biohazardous materials, allowing for the easy discharge of biological fluids directly into the biohazard chamber. The box itself incorporates identifiable features, such as laser-etched QR codes, to ensure traceability of samples and related data. The system scans the QR code to associate spectral measurements with the biological fluid sample being studied.
[0041] like Figure 2 As indicated by the forward arrow, the slide 16 containing biological fluid is then transported to a drying chamber 18 within the system, which is capable of drying the sample (e.g., using heat). Once dried, the slide 16 is transported to the detection and analysis module 20 within the spectrometer assembly 10. Preferably, the transport of the slide 16 is also automated, thus eliminating human intervention in the process and ensuring that the sample remains contaminated from its initial introduction into the assembly until the analytical results are obtained.
[0042] In one embodiment, the drying time is automatically determined by measuring the water content in the biological fluid sample, and then if the detected water content exceeds a certain threshold, the sample slide is submitted for further drying. In another method, this can be achieved through direct spectroscopic measurements, since a known water peak exists at 3400 cm⁻¹. -1 If the peak value is too high compared to the reference value, the sample slide may be automatically returned to the drying chamber for a period of time.
[0043] In such Figure 2 In the specific embodiment shown, the detection and analysis module 20 includes an FTIR spectrometer system 22, and the slide 16 is presented in optical alignment with the FTIR system 22. Preferably, an automatic calibration is performed on a known sample within the measurement range of the biological fluid sample before measurements are initiated on the presented sample. For example, sample measurements and control measurements can be performed multiple times on a set of individual slides prepared as described above. Using multiple measurements and multiple slides allows for the collection of more sample data, minimizing the possibility of errors and improving the sensitivity of the system.
[0044] FTIR spectroscopy can be used alone (in transmission mode) or at 4000–400 cm⁻¹. -1 It can be used in conjunction with attenuated total reflectance (ATR) sampling in the spectral range (also known as the mid-IR region). Similarly, in the 4000–400 cm⁻¹ range... -1 The mid-IR spectral range, 1800-600 cm⁻¹ -1 The fingerprint area and 4000-2500cm -1 Useful spectral information for detecting identified pathogens is collected in the high-wavelength region. Measurements are performed in specific fingerprint regions of the pathogen being detected. For example... Figure 2As shown, the IR beam output from the FTIR system 22 passes through the dried sample on the slide 16, and the characteristics of the beam as it passes through the sample S are collected by the OSA 24. The spectral data generated by the OSA 24 is then used as input to the MLM system 30, as discussed in detail below, which is used to analyze the presented data and generate pathogen detection results.
[0045] like Figure 2 As shown, one aspect of the invention is that the slides can be sterilized and reused, which is a significant advantage when the equipment is used in remote, poorly equipped facilities where adequate medical care is often difficult to provide. In this example, after the FTIR analysis is completed, the slide 16 is automatically transferred to a sterilization room, where it can be cleaned and sterilized for reuse.
[0046] Figure 3 An alternative configuration of the sample preparation and spectroscopic system, depicted as 10A, is shown. The difference in this case is the use of Raman spectroscopy to obtain spectral image data. Typically, sample collection is performed in the same manner as described above, with biological fluid samples placed in sterile containers 12 loaded in a cartridge 14. In this case, cartridge 14 is used to dispense a predetermined volume of biological fluid into vials 32 (or similar containers), which are then automatically conveyed to a detection and analysis system 34. In this case, the detection and analysis system 34 includes a Raman spectrometer 36, which includes a detachable probe 38. As shown, the probe 38 (which may contain a bundle of multiple optical fibers as known in the art) is inserted into the vial and used to illuminate the sample. Back-reflected light passes through the probe 38 and enters the Raman spectrometer 36.
[0047] Raman spectroscopy can utilize laser devices operating at appropriate wavelengths (e.g., 532 nm, 780 nm, etc.) for biomarker detection as illumination. That is, one or more of these laser sources can be used for illumination at wavelengths of 4000–1000 nm. -1 Analysis can be performed within a certain range. Miniature Raman spectrometers capable of producing spectra and portable Raman fibers can also be used. In particular, this facilitates the detection and identification of spectra covering 4000-400 cm⁻¹. -1 Spectral range, 1800-400cm -1 The fingerprint area and 4000-2600cm -1 Spectral information of various pathogens in the high-wavelength region. Measurements are performed in specific fingerprint regions of the pathogens being detected.
[0048] and Figure 2Similar to the setup, the spectral data collected by the detection and analysis system 34 is presented to the machine learning model 30 for use in performing the pathogen detection process. Once the analysis is complete, the probe 38 can be detached from the Raman spectrometer 36 and transferred to the cleaning and sterilization room 40 so that it can be reused with another patient.
[0049] Based on the above discussion and references Figures 1-3 It is evident that the system of the present invention is configured to provide a hygienic, contaminant-free flow of the presented sample (on a slide or in a vial) for spectral data measurement and analysis. Furthermore, although the machine learning model 30 in Figure 2 and Figure 3 While these models are presented as separate components in the system, it should be understood that they can be downloaded to a memory component within the device (or accessed via a communication link) so that the pathogen detection process can be used as a field care system. The ability to provide rapid and accurate results using automated systems at field care locations is an important aspect of this invention, allowing for rapid turnaround measurement and diagnosis of a wide variety of pathogens.
[0050] In this way, the system can not only detect various pathogens in the presented biological fluid samples, but also monitor a given patient's response to treatment for infection or disease by detecting the presence of pathogens over time.
[0051] Based on the description of devices that can be used to perform pathogen detection, the discussion now turns to the details of implementing machine learning models to evaluate spectral data. Importantly, these models are trained to detect unique pathogen signals in the spectral data associated with a specific sample preparation method. As described above, in one embodiment of the invention, multiple biological fluid sample points can be distributed on a single slide to minimize errors and better detect relevant signals, or a set of slides (or vials from the same sample) can be processed together as a unit. In another related embodiment, each of these multiple points (or vials) can represent a slightly different sample preparation method used, such as drying to different degrees, or a diluted version from the same biological fluid sample.
[0052] Spectral analysis of the raw data can be processed and validated using various chemometric analyses, including but not limited to (1) c-parameter support vector machines (SVM), nu-SVM, or stochastic gradient descent SVM algorithms and logistic regression, as well as PCA; or (2) artificial neural networks (ANN) and / or CNNs as described above. All raw spectral data were processed and validated for pathogen detection. Analysis can be performed using various software packages, such as Unscrambler, Python, Simca, and MatLab.
[0053] In the methods described above, models trained and / or developed on labeled spectral datasets obtained from patient bodily fluids are used to evaluate unknown urine samples (bodily fluids) with HPV infection (or other viral infections) prepared in the same manner. As another example, saliva or oral tissue samples can be used to assess the presence of precancerous or cancerous biomarkers in the oral cavity. At least one model is then used for each type of sample preparation. In this way, multiple models can be used to capture subtle differences in sample signals, and the accuracy of the final diagnosis can be improved through model combination.
[0054] More specifically, in order to capture subtle differences in spectral signals attributable to specific molecular variations associated with known pathogens, appropriate preprocessing of the spectra (normalization, scattering correction, etc.) may be necessary. In particular, by using machine learning tools that can sift through patterns in spectral data, a robust set of models can be developed to learn the unique spectral signals associated with specific pathogens.
[0055] In an exemplary process using urine as a biological fluid and hr-HPV as the pathogen to be detected, a preliminary small-scale study was conducted on 100 urine samples with known hr-HPV status using a well-validated GP5+ / 6+ PCR-EIA test. This test included 29 hr-HPV positive urine samples and 71 hr-HPV negative urine samples.
[0056] At the Lancaster University Bioengineering Laboratory, observer-blind statistical analysis of the spectral signals from these urine samples was performed using principal component analysis-linear discriminant analysis (PCA-LDA) and partial least squares-discriminant analysis (PLS-DA). When the data were preprocessed using the second and third derivatives of the raw data, the sample predictions performed using these methods achieved 100% sensitivity and 92% specificity in detecting hr-HPV DNA. Figure 4 and Figure 5 The graph illustrates the improvement in classification as the data undergoes these additional rounds of differentiation. Specifically, Figure 4 This is a scatter plot of the original data. Figure 5 This illustrates the improvements that can be provided in data analysis when using the third derivative of the original data (in this case). The results are summarized in the table below (Table 1):
[0057] According to the present invention, other effective machine learning methods can be used to analyze biological fluid spectral data, particularly those known for their classification robustness. For example, convolutional neural networks (CNNs) can be used to classify a set of biological fluid spectral data. One-dimensional convolutional neural networks (1D-CNNs) are commonly used for classification tasks of time series and frequency series data. While other machine learning tools such as LDA may benefit from standalone feature extractors such as PCA, feature extraction is embedded within a trained CNN. Because FTIR spectra within a specific wavelength range have characteristic shapes and amplitudes that reflect the chemical properties of the sample, 1D-CNNs, which learn “spatial” information from the data, are well-suited for extracting relevant but subtle properties of FTIR signals in a given real-world context. For example, some convolutions that can be applied to spectral data by 1D-CNNs include derivatives, smoothing functions, and / or variable selection. Furthermore, techniques such as data augmentation can be used to improve CNN performance compared to other techniques such as PCA and / or LDA, where data augmentation may not be helpful.
[0058] In an exemplary method of this invention, relating to machine learning details related to CNNs, a training dataset consisting of 221 spectra is first augmented by a factor of 100. Augmentation helps increase the size of the dataset while introducing random variations and noise into the spectra to simulate variations in data collection, such as baseline shifts. Subsequently, multiplicative scattering correction is applied to the augmented data to renormalize the spectra while preserving residual noise. Then, using the average of the training dataset as a reference, the multiplicative scattering correction is applied to the validation and test sets. The processed and augmented labeled spectra are as follows: Figure 6 As shown.
[0059] The fingerprint region of this data (i.e., the region containing the signal of the biomarker being sought, 800-1800 cm) -1 Training a 1-D CNN on it, Figure 7 The training and validation accuracy during the training cycle are shown. Specifically, on the unenhanced test set, the method achieved an accuracy of 93.7% (91.2% specificity and 100% sensitivity), which included 34 hrHPV negative cases and 14 hrHPV positive cases (as shown in Table 2 below). The test set was drawn from the same dataset but was removed before training and was not seen by the model during training.
[0060] By utilizing a larger urine spectral dataset, this 1-D CNN model can be further extended, making it more robust to variability. Essentially, a more comprehensive spectral database allows this CNN-based model to learn more from the dataset and better extract relevant biomarker signals from noise and irrelevant physiological changes that may exist in samples across various settings. After training with sufficient data, the machine learning model can also be extended to perform or assist in genotyping.
[0061] Figure 8 A block diagram of an example computer system 800 is shown, which can be used for sample analysis of pathogen detection using one or more machine learning models of the system of the present invention described herein. The computing system 800 may include one or more computer hardware processors 802 and non-transitory computer-readable storage media (e.g., memory 804 and one or more non-volatile storage devices 806). The processor 802 can control the writing of data to and reading data from (1) memory 804 and (2) non-volatile storage devices 806. To perform any of the machine learning functions described herein, the processor 802 can execute one or more processor-executable instructions stored in one or more non-transitory computer-readable storage media (e.g., memory 804), which may serve as non-transitory computer-readable storage media storing the processor-executable instructions executed by the processor 802.
[0062] As used herein, the terms "program" or "software" refer to any type of computer code or processor-executable instruction set that can be used to program a computer or other processor (physical or virtual) to implement the various aspects of the embodiments discussed above. Furthermore, according to one aspect, when performing the methods of this disclosure, one or more computer programs do not need to reside on a single computer or processor, but can be distributed in a modular manner across different computers or processors to implement the various aspects of the disclosure provided herein.
[0063] Processor-executable instructions can take many forms, such as program modules executed by one or more computers or other devices. Typically, program modules include routines, programs, objects, components, data structures, etc., that perform tasks or implement abstract data types. The functionality of program modules can usually be combined or distributed.
[0064] More generally, while preferred embodiments of the invention have been shown and described herein, those skilled in the art will understand that these embodiments are provided by way of example only. Many variations, modifications, and substitutions can be identified and used by those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. Indeed, it is intended that the claims define the scope of the invention, and that methods and structures within the scope of these claims and their equivalents are encompassed.
Claims
1. A diagnostic system for analyzing spectral data from biological body fluid samples to detect specific pathogens, comprising: A sample collection and preparation device for collecting biological body fluid samples; A spectroscopic device for irradiating the acquired biological fluid sample and creating spectral data from the irradiation; and A machine learning model processing unit, which responds to the acquired biological fluid sample, wherein... The machine learning model processing unit includes one or more trained models to detect whether one or more selected pathogens are present in the acquired spectral data.
2. The diagnostic system according to claim 1, wherein, The collected biological fluids were selected from the following collections: blood, urine, saliva, serum, and plasma.
3. The diagnostic system according to claim 1, wherein, The biological fluids are collected using a non-invasive process.
4. The diagnostic system according to claim 1, wherein, The sample collection and preparation device utilizes a hygienic flow configuration to provide sample movement from collection to preparation, minimizing the introduction of any contaminants.
5. The diagnostic system according to claim 1, wherein, The sample collection and preparation device includes a drying chamber for drying biological fluid samples loaded on glass slides.
6. The diagnostic system according to claim 5, wherein, The spectroscopic device is used to scan the water peaks in biological fluid samples mounted on glass slides to determine whether additional drying is required.
7. The diagnostic system according to claim 1, wherein, The sample collection and preparation device is configured to prepare multiple samples from the collected biological fluids for independent analysis.
8. The diagnostic system according to claim 7, wherein, At least two different preparation methods are used for a single sample from the plurality of samples.
9. The diagnostic system according to claim 1, wherein, The spectroscopic device includes a Fourier transform infrared spectrometer.
10. The diagnostic system according to claim 1, wherein, The spectroscopic device includes a Raman spectrometer.
11. The diagnostic system according to claim 1, wherein, The machine learning model processing unit includes a non-transitory computer-readable storage medium for storing instructions that, when executed by circuitry included within the processing unit, cause the circuitry to perform analysis of the presented spectral data using one or more machine learning models stored in the storage modules of the processing unit.
12. The diagnostic system according to claim 11, wherein, The machine learning models stored in the storage unit include SVM-based models, logistic regression-based models, and one or more of PCA, ANN, and CNN, wherein one or a combination of these different models is developed for a given set of pathogens under study.
13. The diagnostic system according to claim 11, wherein, The machine learning model provides the ability to detect different strains of the selected pathogen.
14. The diagnostic system according to claim 1, wherein, At least the sample acquisition and preparation device and the spectroscopic device are formed as a single component, providing a portable diagnostic system.
15. The diagnostic system according to claim 14, wherein, The single component also includes the machine learning model processing unit, which contains multiple downloaded machine learning models for pathogen detection.
16. The diagnostic system according to claim 14, wherein, The machine learning model processing unit includes a separate component configured to communicate with the individual component.
17. A method for analyzing spectral data from a presented biological fluid sample to detect a specific pathogen, the method comprising the steps of: a) Perform a non-invasive sample collection procedure to collect selected biological fluids from the patient; b) Prepare a controlled amount of the selected biological fluid from the box for further processing; c) Irradiate the prepared sample with a spectrometer and generate spectral data from it; as well as d) Apply one or more machine learning models trained for detecting pathogens in biological fluids to analyze spectral data to detect the presence of specific pathogens in the generated spectral data.
18. The method of claim 17, further comprising the step of repeating steps a)-d) over a period of time to monitor the response to treatment of the detected pathogen.
19. The method of claim 17, wherein, The selected biological fluids include urine, and the selected pathogens are one of hr-HPV and Ir-HPV.
20. The method of claim 17, wherein, The selected biological fluids include saliva, and the selected pathogens are either core precancerous lesions or oral cancer.