Pathogen detection in biological fluids using spectroscopic techniques in conjunction with analytical models.

JP2026529048APending Publication Date: 2026-08-27NSV INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026501095
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-15
Filing Date
2024-07-05
Publication Date
2026-08-27

AI Technical Summary

Benefits of technology

【0010】 本明細書中には、試料の生体液に1つ以上の分光プロセスを施し、その後に機械学習を利用して結果を高い正確度で取得することによって病原体検出を実行するシステム及び方法を開示する。非侵襲的生体液試料を取得して分光分析を施すための好適なシステムをポータブル(携帯)アセンブリ内で具体化することができ、このポータブル·アセンブリは、結果を患者に対して接点で表現することを、ものの数分で可能にする。試料準備プロセスは、(汚染物質を導入し得る)試料の後続処理を必要とせず;むしろ、収集した生体液を直接用いて分光手順を実行する。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026529048000001_ABST
    Figure 2026529048000001_ABST
Patent Text Reader

Abstract

A system and method for performing pathogen detection from collected biological fluid samples are disclosed. One or more spectroscopic processes are applied to the biological fluid, generating spectral data that may include specific biomarkers. These spectral data are then analyzed using a machine learning model trained on the detection of specific pathogens to rapidly obtain highly accurate results. This system enables the results to be presented to the patient on the spot within minutes. Sample preparation does not require subsequent processing of the sample (which may introduce contaminants); rather, the spectroscopic procedure is performed directly on the collected biological fluid. A machine learning algorithm is trained on labeled data, enabling the analysis of submitted samples for the search of pathogens of interest. The machine learning model can be incorporated within the assembly or accessed via a secure communication link.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related application Chris reference This application claims priority by U.S. Provisional Patent Application No. 63 / 525697, filed July 9, 2023, and U.S. Provisional Patent Application No. 63 / 553862, filed February 15, 2024, which are incorporated herein by full reference.

[0002] Technical field The disclosed principle is aimed at detecting pathogens in biological fluid samples, and more specifically, at rapidly and accurately identifying the presence of specific pathogens by using spectroscopy imaging of samples obtained from non-invasive biological fluids (such as urine, saliva, blood, etc.) in combination with machine learning models. [Background technology]

[0003] Background of the Invention Cervical cancer is the fourth most common cancer among women worldwide and is caused by high-risk human papillomavirus (hr-HPV). The World Health Organization (WHO) argues that cervical cancer can be eradicated through its three-pronged strategy: HPV vaccination; screening and treatment of precancerous lesions; and treatment of invasive disease. HPV vaccination prevents infection, but not infection in those who have already been exposed to the virus. Therefore, screening will continue to play a crucial role in eradicating cervical cancer in the foreseeable future. In the UK, cervical screening has helped reduce deaths from cervical cancer by 70% since its introduction. However, despite its success, cervical screening has been declining for many years. In 2021, only 70.2% of eligible individuals in the UK were screened, the lowest figure to date.

[0004] Obstacles to cervical screening include confusion, inconvenience, and discomfort. To address some of these obstacles, vaginal self-screening for HPV detection has been developed. Vaginal self-screening is currently available in at least 17 countries as an alternative to cervical screening, both for screening populations and as a primary screening option. The effectiveness of vaginal sampling in improving screening inclusion varies across different studies, ranging from 6% to 30%. Urine self-testing procedures have been demonstrated to have similar diagnostic accuracy to vaginal self-testing procedures and approach the diagnostic accuracy of vaginal samples obtained by clinicians. Urine testing is non-invasive, does not require a healthcare professional to take a sample, and does not require healthcare facilities to accommodate a clinical laboratory for collecting vaginal samples. Women recognize that urine testing is less "invasive" than vaginal self-testing.

[0005] The need to improve the clinical diagnostic capabilities of healthcare systems at both the national and community levels is a key part of managing viral and bacterial infections. As will be detailed below, spectroscopic techniques (including both infrared and Raman spectroscopy) enable the analysis of bacteria and viruses with good specificity and sensitivity.

[0006] Infrared (IR) spectroscopy is a form of vibrational spectroscopy that relies on the absorbance, transmittance, or reflectance of infrared light when a sample is illuminated. A specific type of IR spectroscopy, Fourier Transform IR (FTIR) spectroscopy, is a preferred process because it allows for the simultaneous collection of data across the entire IR wavelength range of interest. FTIR is considered a powerful tool for chemical analysis due to its ability to provide detailed information at the molecular level about the chemical composition of constituent substances such as proteins, nucleic acids, carbohydrates, and lipids. Essentially, this spectroscopy provides a biochemical fingerprint of chemical and biomolecular structures in diverse environments, such as biological fluids. IR spectroscopy primarily deals with the IR region of the electromagnetic spectrum, most commonly focusing on absorption; that is, absorption occurs when the frequency of the IR emission is the same as the vibrational frequency of the coupling. The wavelengths absorbed by the sample are represented as spectral bands and are characteristic of the molecular structure.

[0007] The use of IR spectroscopy in clinical practice is increasing due to the rapid, cost-effective, and accurate disease prediction capabilities of this technology. To date, IR spectroscopy of biological fluids has demonstrated high sensitivity and specificity for the diagnosis of dementia, brain tumors, and endometrial cancer in large-scale cross-sectional studies. Other studies have demonstrated the potential of spectroscopy of liquid-based cell samples as a triage (severity assessment test) for stratifying women who are HPV-positive. Early proof-of-concept studies have demonstrated similar potential for urine in the detection of gynecological cancers.

[0008] Raman spectroscopy is a somewhat different technique used to measure the relative frequencies at which biological fluid samples scatter radiation. Raman spectroscopy can be used directly with biological fluid samples in their liquid state, in contrast to IR spectroscopy, which requires biological fluid samples to be dried before analysis. [Overview of the project] [Problems that the invention aims to solve]

[0009] While these types of spectroscopic techniques have been used in the detection of several viruses (including hr-HPV) based on cervical (i.e., invasive) samples, there remains a need to develop methods to utilize these techniques with non-invasive samples such as urine to detect viruses like (but not limited to) hr-HPV in rapid point-of-care turnaround time. [Means for solving the problem]

[0010] This specification discloses a system and method for performing pathogen detection by subjecting a biological fluid sample to one or more spectroscopic processes and then using machine learning to obtain the results with high accuracy. A suitable system for obtaining and performing spectroscopic analysis on non-invasive biological fluid samples can be embodied in a portable assembly, which enables the presentation of results to the patient on the spot in minutes. The sample preparation process does not require subsequent processing of the sample (which may introduce contaminants); rather, the spectroscopic procedure is performed directly on the collected biological fluid.

[0011] Machine learning algorithms are trained on labeled data, enabling the analysis of submitted samples to search for specific biomarkers (biological indicators) within the pathogen of interest. The machine learning models can be integrated into the assembly described above or accessed via a secure communication link. It is advantageous to subsequently add the data collected during each step to a database used in the machine learning process (without patient ID (identification)). Either Fourier transform infrared (FTIR) spectroscopy or Raman spectroscopy can be used to collect spectral data from IR-illuminated samples.

[0012] A preferred example of the present invention may take the form of a method for performing the analysis of spectral data from a provided biological fluid sample, the method comprising: a) performing a non-invasive sample collection process to obtain a selected biological fluid from a patient; b) preparing the selected biological fluid in a controlled quantity on a cartridge for further processing; c) illuminating the prepared sample using a spectrometer to generate spectral data from the sample; and d) analyzing the spectral data by applying one or more machine learning models trained on detecting pathogens in biological fluids to detect the presence or absence of specific pathogens in the generated spectral data.

[0013] Other preferred examples include a diagnostic system that analyzes spectral data of a biological fluid sample to detect a specific pathogen. This system includes: (1) a sample acquisition and preparation device for collecting a biological fluid sample; (2) a spectrometer for illuminating the acquired biological fluid sample and generating spectral data from the light illuminating the biological fluid sample; and (3) a machine learning model processing device that responds to the acquired biological fluid sample, the machine learning model processing device including one or more trained models to detect the presence or absence of a selected pathogen in the generated spectral data.

[0014] According to some preferred examples, the system can be configured as a portable assembly combining a sample collection and preparation device with a spectrometer, and diagnostics on the collected samples are performed using the capabilities of the computer processor in this portable system, or the processor of a local computer (possibly a tablet computer).

[0015] According to some preferred examples, multiple samples can be collected simultaneously and analyzed independently to reduce the probability of errors.

[0016] According to one preferred example, machine learning models are collected and stored in a non - transient computer - readable storage medium, and this non - transient computer - readable storage medium can also be used to store instructions. When these instructions are executed by the system, the system is made to process the collected spectra, and then the selected machine learning model is executed to analyze the provided spectral data.

[0017] Additional preferred examples and features of the present invention will become apparent in the following description by referring to the drawings referred to in the description.

[0018] In the following, the following drawings are referred to.

Brief Description of the Drawings

[0019] [Figure 1] It is a figure including a schematic diagram of a system used for pathogen detection according to the principle of the present invention. [Figure 2] It is a block diagram of a spectrometer assembly useful in collecting and preparing a biological fluid sample, obtaining spectral data, and applying a machine learning model to the spectral data to perform pathogen detection according to the present invention. [Figure 3] It is a block diagram of a type of alternative spectrometer assembly with respect to that shown in FIG. 2, in which case a Raman spectroscopy process is used instead of the FTIR spectroscopy procedure related to the configuration of FIG. 2. [Figure 4] It is a plot diagram of the partial least - squares (PLS) decomposition without differentiation of the spectral signal obtained from a urine sample of a known hr - HPV state based on a selected group of machine learning models. [Figure 5] It is a figure showing the same data as the data shown in FIG. 4, but in this case, in the form of the third derivative of the data, generating improved information with respect to the trend of the raw data. <0000​​​​A diagram showing the training of a one-dimensional convolutional neural network (CNN), which iterates over the entire fingerprint region of the data shown in FIG. 6. [Figure 8] A block diagram of an example of a computer system that can be used to implement some embodiments of a machine learning model utilized in detecting one or more pathogens in a provided sample.

Embodiments for Carrying Out the Invention

[0020] Detailed Description of the Invention (For example, including viruses such as human papillomavirus) The rapid and precise detection and identification of pathogens is the key to managing infections. The currently used methods are invasive and require a relatively long period to detect infections. In addition, the specificity and sensitivity by using existing methods are relatively poor.

[0021] According to the principle of the present invention, and as will be described in detail below, specific pathogens can be detected and identified from a biological fluid (e.g., blood, saliva, urine, serum, plasma, etc.) obtained by a relatively non-invasive method. The complete process from obtaining the sample to presenting the result to the patient is expected to take less than 30 minutes. For example, using the above method, The first urine (initial urine) collected by oneself can be examined for the presence of hr-HPV by a FTIR spectrometer according to its biomolecular, cellular, and metabolic composition, and the FTIR spectrometer can be coupled to an attenuated total reflectance (ATR) sampling accessory, and in some cases, Raman spectroscopy can also be used for examination.

[0022] The method anticipated by the principles of the present invention is considered to be a simple, non-invasive analytical technique that can characterize the biochemical profile of any biological fluid without the preparation of numerous samples. The biological fluid itself is subjected to FTIR spectroscopy and / or Raman spectroscopy to generate spectral data. These data may include specific resonance peaks associated with identified biomarkers, which are identified by machine learning algorithms. In other words, the system and method of the present invention enable efficient and simple sample preparation and provision of the sample to a spectrometer combined with a machine learning model to analyze biomolecules and cells present in biological fluids, thereby enabling accurate, rapid, and simple detection of pathogens identified in biological fluid samples.

[0023] In addition, according to the principles of the present invention, by using machine learning algorithms, the above technique can identify spectral peaks of specific biomarkers related to the area of ​​interest, and can be used to analyze one or more biological fluids. Furthermore, using a single biological fluid sample, multiple pathogens can be identified from a single spectrum using multiple machine learning algorithms.

[0024] The methods of the present invention can, for example, identify individual HPV strains and can also identify other viruses or pathogens that may be present, individually or in combination. For example, the identification of other viruses such as monkeypox, herpes simplex virus types 1 and 2, syphilis, and varicella-zoster virus can all be performed using the systems and methods of the present invention.

[0025] As will be explained in detail below, the method for spectroscopic analysis of biological fluid samples according to the present invention is considered unique and novel because it emphasizes the identification of peaks in the spectra of biomarkers based on the biomolecular, cellular, and lipid composition of a biological fluid sample that may consist of multiple known pathogens, each of which may have a different effect on the composition of a patient's biological fluid.

[0026] In addition, the method of the present invention is considered unique in that it uses a machine learning model to select relevant signals in the spectrum (generated by the spectroscopic process) that are associated with a specific pathogen from among numerous other biochemical peaks contributing to the spectral data recovered from the patient's biological fluid. The biological fluid is directly sampled without the extraction of biomolecules or cells, and these peak characteristics are distinguished from normal, uninfected biological fluid samples, with a focus on identifying specific peak characteristics of infection. A spectral database of various types of infected and uninfected biological fluid samples is used as a training dataset to develop and train a model for rapid and accurate identification of unknown pathogen samples. Chemometric analysis utilizing neural networks and machine learning in conjunction with the automation of the biological fluid collection process results in rapid and accurate detection and identification.

[0027] The method of the present invention is considered to go beyond simply comparing a test spectrum to a database of known spectra because the machine learning model used can learn minute signals present in the spectrum, which directly represent the presence of specific pathogens in a patient's biofluid sample, eliminating the need to further separate the sample to the cellular level before obtaining the sample spectrum. In other words, using a sample directly from biofluid, rather than first separating it into biomolecules of interest, means that the model needs to be able to find the aforementioned signals among peaks originating from numerous other biomolecules in the biofluid sample. This can be done more accurately and efficiently by a machine learning model trained on a large dataset of spectra from diverse samples. The machine learning model can also be improved over time because it can learn a greater number of patterns from samples obtained in diverse settings by using more data (from a greater source of diversity) for additional learning. This is something that would otherwise be impossible or much more time-consuming if the test spectrum were simply compared to a database of known spectra.

[0028] The following description outlines the overall system and method for performing pathogen detection, followed by detailed descriptions of the following individual components: (1) the sample preparation process; (2) the spectroscopic data acquisition process; and (3) the analysis of the collected spectral data using machine learning algorithms.

[0029] Figure 1 is a schematic diagram of these components, presented in flowchart format. The process begins with the collection of a biological fluid sample (step 100), in which it is intended that the biological fluid to be used be collected in a relatively non-invasive manner. In particular, biological fluids preferred for analysis include, but are not limited to, urine, serum, plasma, etc. In practice, in many cases, the patient can perform the collection themselves, and the sample is stored in a contaminant-free container (step 110). Once the sample is obtained, a small portion (of a known, controlled volume) is loaded (step 120), and perhaps the slide is dried in a well-known manner to prepare it for spectroscopic analysis (step 130).

[0030] In practice, sample analysis is then performed in step 200, and the sample analysis can take the form of Raman spectroscopy (step 210a) and / or FTIR spectroscopy (step 210b). These processes are complementary in the methods used to obtain spectral data, and therefore, the use of both techniques tends to further increase the reliability of the results. As mentioned above and as will be explained in detail below, one difference between these two methods is that FTIR spectroscopy requires the use of a dry sample, while Raman spectroscopy can be used directly on liquid biological fluids (as well as dry samples on a mount).

[0031] The digital spectral image data generated by the spectroscopic process in step 210 is then provided to a set of machine learning models (MLMs), as shown in step 220. The flowchart identifies selection numbers for example MLMs, which include Support Vector Machines (SVMs), linear Discriminant Analysis (LDA), logistic regression, and artificial neural networks (ANNs). These are merely selections of the wide variety of MLMs that can be used; in fact, the same spectral data can be analyzed using multiple different models. As described below, the use of these MLMs allows for the evaluation of the provided data and the provision of results quickly and efficiently (in many cases, on the order of approximately 5 minutes) (as shown in step 300). As shown by step 400 in Figure 1, the results of the analysis (with any data that identifies a patient removed) can be stored and used in the future to further update and train the dataset.

[0032] Referring here to the details of the sample preparation and spectroscopy process, Figure 2 shows a simplified block diagram of the spectrometer assembly 10 according to the present invention, which can be used to prepare and analyze biological fluid samples. Figure 2 shows a standard biological fluid container 12, which is used by either the patient or a healthcare professional to collect the sample. The biological fluid container 12 is inserted into a (generally reusable) cartridge 14, which engages with the remaining components that form the spectrometer assembly 10. The assembly 10 is intended to be configured to handle multiple such cartridges to provide a compact configuration. The cartridge 14 is configured to automatically dispense a precise volume (e.g., 5 microliters) of biological fluid onto a clean, sterile slide 16. The slide 16 may be made of CaF2, and the sample is spread to form a small circle. In one embodiment, this is repeated to generate multiple sample spots on the slide 16, allowing multiple measurements to be performed on the same slide (minimizing error and increasing the signal-to-noise ratio). Alternatively, the sample can be placed on a diamond crystal for ATR analysis.

[0033] The cartridge system includes a linear container used to collect biohazardous materials, allowing for easy draining of the biological fluid and direct disposal into a biohazard bin. The cartridge itself has identifiable features, such as a laser-etched QR (Quick Response) code, enabling traceability of the sample and associated data. The system scans the QR code to associate spectroscopic measurements with the biological fluid sample under study.

[0034] As shown by the upward arrow in Figure 2, the slide 16 containing the biological fluid is then transported to the desiccation chamber 18 within the system, which allows for the drying of the sample (e.g., using heat). Once dried, the slide 16 is transported to the detection and analysis module 20 within the spectrometer assembly 10. The transport of the slide 16 is also automated, so that human intervention is not involved in the process and the sample remains uncontaminated from its initial introduction into the assembly to the acquisition of analytical results.

[0035] In one embodiment, the water content in a biological fluid sample is measured, and then, if the detected water content exceeds a specific threshold, the drying time is automatically determined by continuing the drying process on the sample slide. One method to achieve this is by direct spectroscopic measurement, where the known water content peak is 3400 cm⁻¹. -1 This is because it exists at [location]. If this peak is too high compared to the reference value, the sample slide can be automatically returned to the drying chamber for a different period.

[0036] In the specific embodiment shown in Figure 2, the detection and analysis module 20 includes an FTIR spectrometer system 22, and slide 16 is presented so as to be optically aligned with the FTIR system 22. Before commencing measurement of the presented sample, it is preferable to perform an automated calibration on a known sample within the range of biological fluid sample measurement. For example, the sample measurement can be performed multiple times on a set of separate slides, with control, all of which are prepared in the manner described above. The use of multiple measurements and multiple slides provides the ability to collect a larger amount of sample data, minimizing the possibility of error and increasing the sensitivity of the system.

[0037] FTIR spectroscopy, used alone (in transmission mode) or in combination with sampling accessory total internal reflection (ATR), can be used in the mid-infrared region, also known as 4000-400 cm⁻¹. -1 It can be used within the spectral range of 4000-400 cm². Here too, spectral information useful for the detection of identified pathogens is available within the range of 4000-400 cm².-1 Within the mid-infrared spectral range, 1800-600 cm -1 Within the fingerprint range, and 4000-2500cm -1 The data is collected within the high-wavelength region. The measurement is performed within a specific fingerprint region where the pathogen is detected. As shown in Figure 2, the IR beam output from the FTIR system 22 passes through the dry sample on slide 16, and the characteristics of the beam as it passes through the sample S are collected by the OSA (Optical Spectrum Analyzer) 24. Next, as will be described in detail below, the spectral data generated by the OSA 24 is applied as input to the MLM system 30, which functions to analyze the provided data and generate pathogen detection results.

[0038] As shown in Figure 2, the ability to disinfect and reuse slides is one aspect of the present invention, which is a significant advantage when the device is used in remote, poorly equipped facilities where providing appropriate medical care is often difficult. In this example, after completing the FTIR analysis, slide 16 is automatically transferred to a disinfection chamber, where it is cleaned and disinfected for reuse.

[0039] Figure 3 shows an alternative configuration of the sample preparation and spectroscopy system, designated as 10A. The difference is that in this case, spectral image data is acquired using a Raman spectroscopy process. Generally, sample collection is performed in the same manner as described above, with the biological fluid sample placed in a sterile container 12 loaded in a cartridge 14. In this case, the cartridge 14 is used to dispense a predetermined volume of biological fluid into a vial (small medicine bottle) 32 (or a similar type of container), and then the vial 32 is automatically transported into the detection and analysis system 34. In this case, the detection and analysis system 34 includes a Raman spectrometer 36, which includes a detachable probe 38. As shown in the figure, the probe 38 (which may include a bundle of multiple optical fibers, as is known in the present art) is inserted into the vial, and the sample is illuminated using the probe 38. The back-reflected light passes through the probe 38 and enters the Raman spectrometer 36.

[0040] Raman spectroscopy can utilize a laser device operating at wavelengths suitable for biomarker detection (e.g., 532, 780 nm, etc.) as illumination. That is, one or more of these laser sources can be used for analysis within the range of 4000 - 100 cm -1 . A micro-Raman spectrometer capable of creating a spectral map, as well as a portable Raman optical fiber, can also be used. In particular, spectral information useful for detecting and identifying various pathogens covers the spectral range of 4000 - 400 cm -1 , the fingerprint region of 1800 - 400 cm -1 , and the high-wavelength region of 4000 - 2600 cm -1 . The measurement is performed within a specific fingerprint region where the pathogen is detected.

[0041] Similar to the configuration of FIG. 2, the spectral data collected by the detection and analysis system 34 is provided to the machine learning model 30, and the machine learning model 30 uses these spectral data in performing the pathogen detection process. Once the analysis is completed, the probe 38 can be detached from the Raman spectrometer 36 and transferred to the cleaning and disinfection chamber 40, thereby enabling the probe 38 to be reused on other patients.

[0042] From the above description and by referring to Figures 1-3, it is clear that the system of the present invention is configured to provide a hygienic and contamination-free flow of samples provided (on slides or in vials) for measurement and analysis of spectral data. Furthermore, although the machine learning model 30 is shown as a separate component in the system in Figures 2 and 3, it should be understood that the model can be downloaded to a memory element in the device (or accessed via a communication link), thereby allowing the pathogen detection process to be used as a point-of-care system. The ability to use an automated system at a point-of-care location to provide immediate and accurate results is a key aspect of the present invention, enabling rapid turnaround measurement and diagnosis of a variety of different pathogens.

[0043] In this way, the above system can be used not only to detect a variety of different pathogens in a provided biological fluid sample, but also to monitor the therapeutic response to infection or disease in a given patient by detecting the presence of pathogens over a long period of time.

[0044] Having moved on from this description of the apparatus that can be used to perform pathogen detection, we now move on to a detailed explanation of implementing machine learning models to evaluate spectral data. It is important to train the model to detect unique pathogen signals in spectral data associated with a particular method of sample preparation. As described above, in one embodiment of the present invention, multiple spots of a biological fluid sample can be dispensed onto a single slide to minimize errors and better detect the relevant signals, or a set of slides (or a set of vials derived from the same sample) can be processed together as a single unit. In other relevant embodiments, these multiple spots (or vials) may represent subtly different sample preparation methods, for example, used to dry to different degrees or derived from diluted versions of the same biological fluid sample.

[0045] The spectral analysis of raw data can be processed and validated using different chemometric analyses, which include (1) C-parameter (C-value) support vector machines (SVM), Nu-SVM, or stochastic gradient descent SVM algorithms and logistic regression, PCA (Principal Component Analysis); or (2) the aforementioned artificial neural networks (ANN) and / or convolutional neural networks (CNN). All spectral raw data is processed and validated for pathogen detection. Different software packages such as Unscrambler®, Phyton®, Simca®, and Matlab® can be used for the analysis.

[0046] Using the methods described above, unknown samples of urine (biofibreeding fluid) from patients infected with HPV (or other viral infections), prepared in the same manner, are evaluated using models trained on labeled spectral datasets obtained from patient biofluids, or models developed from these datasets. As another example, saliva or oral tissue samples can be used to evaluate the presence of biomarkers for oral precancerous or oral cancer. Therefore, there is at least one model for each type of sample preparation. In this way, multiple models can be used to capture different nuances (subtle differences) in sample signals, contributing to a more accurate final diagnosis derived from the combination of models.

[0047] More specifically, proper spectral preprocessing (normalization, scattering correction, etc.) may be necessary to capture subtle differences in spectral signals resulting from specific molecular changes associated with known pathogens. In particular, by using machine learning tools that can extract patterns in spectral data, it is possible to develop a set of robust models that learn unique spectral signals associated with specific pathogens.

[0048] In one preferred process using urine as a biological fluid and hr-HPV as the pathogen to be detected, an initial small-scale test of 100 urine samples with known hr-HPV status was performed using a well-validated GPS5+ / 6+PCR (Polymerase Chain Reaction)-EIA (Enzyme Immunoassay) assay. This test included 29 hr-HPV-positive urine samples and 71 hr-HPV-negative urine samples.

[0049] Observer-blinded statistical analysis was performed on the spectral signals from these urine samples at the Bioengineering Lab at Lancaster University using principal component analysis-linear discriminant analysis (PCA-LDA) and partial least squares-discriminant analysis (PLS-DA). Preprocessing the data with the second and third derivatives of the raw data resulted in sample predictions using the methods described above achieving an overall sensitivity of 100% and a specificity of 92% for detecting hr-HPV DNA (Deoxyribonucleic Acid). Plots in Figures 4 and 5 show the improvement in classification when these additional derivatives are applied to the data. In particular, Figure 4 is a plot of the scattering of the raw data, and Figure 5 shows the improvement that can be brought to data analysis when the (in this case) third derivative of the raw data is used. The following table (Table 1) summarizes the results:

[0050] [Table 1]

[0051] According to the present invention, other effective machine learning methods, particularly those known for their robust classification, can be used to analyze spectral data of biological fluids. For example, a convolutional neural network (CNN) can be used to classify a set of spectral data of biological fluids. One-dimensional convolutional neural networks (1D-CNNs) are often used for classification tasks on time-series and frequency-series data. Other machine learning tools, such as LDA, can benefit from separate feature extractors like PCA, but feature extraction is built into the trained CNN. Since FTIR spectra within a specific wavelength range have characteristic shapes and amplitudes that reflect the chemical properties of the sample, a 1D-CNN that learns "spatial" information from the data is well-suited to the task of extracting relevant, but subtle, attributes of the FTIR signal for a given ground truth (verified true data). For example, some of the convolutions that can be applied to spectral data, such as a 1D-CNN, include derivatives, smoothing functions, and / or variable selection. Furthermore, techniques such as data augmentation can be used to improve the performance of CNNs, in contrast to other techniques like PCA and / or LDA, where data augmentation may prove ineffective.

[0052] In one preferred method of the present invention relating to machine learning specifications for CNNs, the training dataset consisting of 221 spectra was initially expanded to 100 folds (units of division). The expansion serves to increase the size of the dataset while also introducing random variationalities and noise into the spectra that simulate the variance in data acquisition, such as baseline offsets. Subsequently, a multiplicable scattering correction was applied to the expanded data to renormalize the spectra while retaining residual noise. Next, the multiplicable scattering correction was applied to validation and check sets using the mean of the training dataset as a reference. The processed, expanded, and labeled spectra are shown in Figure 6.

[0053] 1D-CNN is used to analyze the fingerprint region of this data (i.e., the region containing the biomarker signal being pursued) (800-1800 cm²). -1 The model was trained on the above, and this training is shown in Figure 7 for training accuracy and validation accuracy over the entire training epoch. In particular, this method was found to produce an accuracy of 93.7% (specificity 91.2% and sensitivity 100%) on an unexpanded test set containing 34 hr-HPV-negative cases and 14 hr-HPV-positive cases (shown in Table 2 below). The test set originates from the same batch of data, but is acquired before training and is not visible to the model during training.

[0054] [Table 2]

[0055] This 1D-CNN model can be further extended by leveraging a larger dataset of urine spectra to make the model more robust to variability. Essentially, a more comprehensive spectral database allows these CNN-based models to learn more from the dataset and better extract relevant biomarker signals from noise and irrelevant physiological variability that may be present in the sample under various settings. With sufficient data training, the machine learning model can be extended to perform or assist in genotyping.

[0056] Figure 8 shows a block diagram of an example of a computer system 800, which can be used to implement several embodiments of sample analysis for pathogen detection using one or more machine learning models of the system of the present invention as described herein. The computer system 800 may include one or more computer hardware processors 802 and non-temporary computer-readable storage media (e.g., memory 804 and one or more non-volatile storage devices 806). The processor 802 can control (1) writing data to and reading data from memory 804 and (2) non-volatile storage devices 804. To perform any of the machine learning functions described herein, the processor 802 can execute one or more computer-executable instructions stored in one or more non-temporary computer-readable storage media (e.g., memory 804), which may function as non-temporary computer-readable storage media that store processor-executable instructions for execution by the processor 802.

[0057] In this specification, “program” or “software” is used in a general sense to refer to any kind of computer code or set of instructions that a processor can execute, which can be used to program a computer or other processor (physical or virtual) to implement various aspects of the embodiments described above. In addition, according to one or more embodiments, one or more computer programs that, when executed, perform the methods of the disclosures provided herein do not need to reside on a single computer or processor, but can be distributed in a modular manner across different computers or processors to perform various aspects of the disclosures provided herein.

[0058] The instructions that a processor can execute can take many forms, such as program modules, which are executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform tasks or implement abstract data types. Generally, the functions of program modules can be combined or distributed.

[0059] More generally, preferred embodiments of the present invention have been illustrated and described herein, but it will be apparent to those skilled in the art that these embodiments are provided only as examples. A great many variations, modifications, and alternatives can be identified and used by those skilled in the art without departing from the present invention. In carrying out the present invention, it should be understood that various alternatives to the embodiments of the present invention described herein can be used. In fact, the following claims are intended to define the scope of the present invention, and that methods and structures within the scope of the claims and their equivalents are covered by the claims.

Claims

1. A diagnostic system that analyzes biological fluid samples to detect specific pathogens, A sample acquisition and preparation device for collecting the aforementioned biological fluid sample, A spectrometer that illuminates the acquired biological fluid sample and generates spectral data from the light illuminating the biological fluid sample, The system comprises a machine learning model processing device that responds to the acquired biological fluid sample, The machine learning model processing device is a diagnostic system that includes one or more trained models and detects the presence or absence of one or more selected pathogens in the generated spectral data.

2. The diagnostic system according to claim 1, wherein the collected biological fluid sample is selected from a set including blood, urine, saliva, serum, and plasma.

3. The diagnostic system according to claim 1, wherein the biological fluid is collected using a non-invasive process.

4. The diagnostic system according to claim 1, wherein the sample acquisition and preparation device utilizes a hygienic flow structure to provide the movement of the biological fluid sample from collection to preparation, thereby minimizing the introduction of any contaminants.

5. The diagnostic system according to claim 1, wherein the sample acquisition and preparation device includes a drying chamber for drying the biological fluid sample mounted on a slide.

6. The diagnostic system according to claim 5, wherein the spectrometer is used to perform a scan to search for the water peak in the biological fluid sample mounted on the slide, and to determine whether or not additional drying is necessary.

7. The diagnostic system according to claim 1, wherein the sample acquisition and preparation device is configured to prepare a plurality of samples from the collected biological fluid for independent analysis.

8. The diagnostic system according to claim 7, wherein at least two different preparation methods are used for individual samples among the plurality of samples.

9. The diagnostic system according to claim 1, wherein the spectrometer uses a Fourier transform infrared spectrometer.

10. The diagnostic system according to claim 1, wherein the spectroscopic device includes a Raman spectrometer.

11. The diagnostic system according to claim 1, wherein the machine learning model processing device comprises a non-temporary computer-readable storage medium for storing instructions, and when the instructions are executed by a circuit contained within the machine learning model processing device, the circuit is caused to perform an analysis of the provided spectral data using one or more machine learning models stored in the memory module of the machine learning model processing device.

12. The diagnostic system according to claim 11, wherein the machine learning model stored in the memory module includes one or more models based on SVM and logistic regression models, together with PCA, ANN and CNN, and one or a combination of these different models has been developed for a predetermined set of pathogens under study.

13. The diagnostic system according to claim 11, wherein the machine learning model provides the ability to detect different strains of a selected pathogen.

14. The diagnostic system according to claim 1, wherein at least the sample acquisition and preparation device and the spectrometer are formed as a single assembly to provide a portable diagnostic system.

15. The diagnostic system according to claim 14, wherein the single assembly further comprises the machine learning model processing device, the machine learning model processing device embodies a plurality of downloaded models useful for pathogen detection.

16. The diagnostic system according to claim 14, wherein the machine learning model processing device comprises a separate assembly configured to communicate with the single assembly.

17. A method for detecting a specific pathogen by performing an analysis of spectral data from a provided biological fluid sample, a) A step of performing a non-invasive sample collection process to obtain selected biological fluids from the patient, b) The step of preparing the selected biological fluid from a cartridge for additional processing in a controlled quantity, c) Using a spectrometer, illuminate the prepared sample and generate spectral data from the sample; d) The step of analyzing the spectral data by applying one or more machine learning models trained on detecting pathogens in the biological fluid, and detecting the presence or absence of a specific pathogen in the generated spectral data. A method that includes this.

18. The method according to claim 17, further comprising the step of repeating steps a) to d) over a period of time to monitor the response of the detected pathogen to treatment.

19. The method according to claim 17, wherein the selected biological fluid includes urine, and the specific pathogen is one of hr-HPV and lr-HPV.

20. The method according to claim 17, wherein the selected biological fluid includes saliva, and the specific pathogen is either oral precancerous disease or oral cancer.