Systems and methods for high-throughput chemical analysis
The high-throughput chemical analysis system enhances the efficiency of liquid biopsy analysis by using a gantry-mounted scanning head with machine learning algorithms to identify droplets and regions of interest, facilitating rapid disease prediction through Raman spectroscopy.
Patent Information
- Application Number
- PCT/US2025/018073
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-01
- Filing Date
- 2025-03-03
- Publication Date
- 2025-09-04
AI Technical Summary
Conventional methods for analyzing liquid biopsies, such as those used in cancer detection, are time-intensive due to the small field of view required for identifying molecular features, limiting the efficiency of sample analysis.
A high-throughput chemical analysis system utilizing a gantry-mounted scanning head with optical and spectroscopy units, combined with machine learning algorithms, to identify droplets and regions of interest within biological specimens, enabling rapid spectral analysis.
The system significantly increases the speed and efficiency of analyzing liquid biopsies by accurately identifying droplets and regions of interest, allowing for rapid disease prediction based on Raman spectroscopy.
Smart Images

Figure US2025018073_04092025_PF_FP_ABST
Abstract
Description
SYSTEMSAND METHODS FOR HIGH-THROUGHPUT CHEMICAL ANALYSISCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 560,274, filed on March 1 , 2024, the entire contents of which are incorporated herein by reference.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0002] This invention was made with government support under Grant No. R01 CA241666 awarded by the National Institutes of Health. The government has certain rights in the invention.BACKGROUND
[0003] Raman spectroscopy has emerged as a powerful analytical technique for characterizing molecular compositions based on the inelastic scattering of light. Traditionally, Raman spectroscopy has been employed for detailed chemical analysis, providing valuable insights into vibrational modes and molecular structures.SUMMARY
[0004] In some aspects, the techniques described herein relate to an optical system for high- throughput chemical analysis including: a gantry: a scanning head mounted to the gantry, wherein the scanning head includes an optical imaging unit and a spectroscopy unit; and a controller including a processor and a memory operably coupled to the processor, the memory having computer-executable instructions stored thereon, that when executed by the processor, cause the processor to: receive, from the optical imaging unit, an image of a biological specimen including a number of droplets; analyze, using a computer vision module, the image to identify a location of at least one droplet of the number of droplets within the biological specimen; control the scanning head to position the spectroscopy unit in proximity to the location of the at least one droplet; analyze, using the computer vision module, a magnified portion of the image including the at least one droplet to identify a region of interest within the at least one droplet; and control the spectroscopy unit to acquire a spectrum associated with the region of interest within the at least one droplet.
[0005] In some aspects, the techniques described herein relate to an optical system, wherein the region of interest includes a feature likely to have a high-value spectrum. In some aspects, the techniques described herein relate to an optical system, wherein analyzing, using the computer vision module, the image to identify the location of the at least one droplet of the number of droplets within the biological specimen includes using a first trained machine learning model to detect the at least one droplet in the image. In some aspects, the techniques described herein relate to an optical system, wherein the first trained machine learning model is a deep learning model. In some aspects, the techniques described herein relate to an optical system, wherein analyzing, using the computer vision module, the image to identify the location of the at least one droplet of the number of droplets within the biological specimen includes using a first ensemble of machine learning models to detect the at least one droplet in the image. In some aspects, the techniques described herein relate to an optical system, wherein analyzing, using the computer vision module, the image to identify the location of the at least one droplet of the number of droplets within the biological specimen includes using an edge detection algorithm to detect a boundary of the at least one droplet in the image.
[0006] In some aspects, the techniques described herein relate to an optical system, wherein analyzing, using the computer vision module, the magnified portion of the image including the at least one droplet to identify the region of interest within the at least one droplet includes using a second trained machine learning model to detect the region of interest in the magnified portion of the image. In some aspects, the techniques described herein relate to an optical system, wherein the second trained machine learning model is a deep learning model. In some aspects, the techniques described herein relate to an optical system, wherein analyzing, using the computer vision module, the magnified portion of the image including the at least one droplet to identify the region of interest within the at least one droplet includes using a second ensemble of machine learning models to detect the region of interest in the magnified portion of the image. In some aspects, the techniques described herein relate to an optical system, wherein the memory has further computer-executable instructions stored thereon, that when executed by the processor, cause the processor to control the spectroscopy unit to acquire a low-resolution spectra associated at least a portion of the at least one droplet, and analyze the low-resolution spectra to identify a region of interest within the at least one droplet. In some aspects, the techniques described herein relate to an optical system, wherein the memory has further computer-executable instructions stored thereon, that when executed by the processor, cause the processor to predict, using a thirdtrained machine learning model, a disease associated with the biological specimen based on the spectrum associated with the region of interest.
[0007] In some aspects, the techniques described herein relate to an optical system, wherein the memory has further computer-executable instructions stored thereon, that when executed by the processor, cause the processor to predict, using a third ensemble of machine learning models, a disease associated with the biological specimen based on the spectrum associated with the region of interest. In some aspects, the techniques described herein relate to an optical system, wherein the spectroscopy unit is configured to deliver and collect light. In some aspects, the techniques described herein relate to an optical system, further including a spectroscopy light source and a detector. In some aspects, the techniques described herein relate to an optical system, wherein the spectroscopy unit includes at least one of the spectroscopy light source and the detector. In some aspects, the techniques described herein relate to an optical system, wherein the optical imaging unit includes a widefield imaging device. In some aspects, the techniques described herein relate to an optical system, further including a sample stage configured to hold a substrate including the biological specimen. In some aspects, the techniques described herein relate to an optical system, wherein the spectroscopy unit is configured for Raman spectroscopy. In some aspects, the techniques described herein relate to an optical system, wherein the spectroscopy unit is configured for surface enhanced Raman spectroscopy (SERS). In some aspects, the techniques described herein relate to an optical system, wherein the biological specimen is a biofluid specimen. In some aspects, the techniques described herein relate to an optical system, wherein the biological specimen is a cell or tissue lysate.
[0008] In some aspects, the techniques described herein relate to a computer-implemented method for controlling an optical system for high-throughput chemical analysis including: receiving, from an optical imaging unit, an image of a biological specimen including a number of droplets; analyzing the image to identify a location of at least one droplet of the number of droplets within the biological specimen; controlling a scanning head including a spectroscopy unit to position the spectroscopy unit in proximity to the location of the at least one droplet; analyzing a magnified portion of the image including the at least one droplet to identify a region of interest within the at least one droplet; and controlling the spectroscopy unit to acquire a spectrum associated with the region of interest within the at least one droplet.
[0009] In some aspects, the techniques described herein relate to a computer-implemented method, wherein the region of interest includes a feature likely to have a high-value spectrum. In someaspects, the techniques described herein relate to a computer-implemented method, wherein analyzing the image to identify the location of the at least one droplet of the number of droplets within the biological specimen includes using a first trained machine learning model to detect the at least one droplet in the image. In some aspects, the techniques described herein relate to a computer-implemented method, wherein the first trained machine learning model is a deep learning model. In some aspects, the techniques described herein relate to a computer-implemented method, wherein analyzing the image to identify the location of the at least one droplet of the number of droplets within the biological specimen includes using a first ensemble of machine learning models to detect the at least one droplet in the image. In some aspects, the techniques described herein relate to a computer-implemented method, wherein analyzing the image to identify the location of the at least one droplet of the number of droplets within the biological specimen includes using an edge detection algorithm to detect a boundary of the at least one droplet in the image.
[0010] In some aspects, the techniques described herein relate to a computer-implemented method, wherein analyzing the magnified portion of the image including the at least one droplet to identify the region of interest within the at least one droplet includes using a second trained machine learning model to detect the region of interest in the magnified portion of the image. In some aspects, the techniques described herein relate to a computer-implemented method, wherein the second trained machine learning model is a deep learning model. In some aspects, the techniques described herein relate to a computer-implemented method, wherein analyzing the magnified portion of the image including the at least one droplet to identify the region of interest within the at least one droplet includes using a second ensemble of machine learning models to detect the region of interest in the magnified portion of the image. In some aspects, the techniques described herein relate to a computer-implemented method, further including controlling the spectroscopy unit to acquire a low-resolution spectra associated at least a portion of the at least one droplet, and analyzing the low-resolution spectra to identify a region of interest within the at least one droplet. In some aspects, the techniques described herein relate to a computer-implemented method, further including predicting, using a third trained machine learning model, a disease associated with the biological specimen based on the spectrum associated with the region of interest. In some aspects, the techniques described herein relate to a computer-implemented method, further including predicting, using a third ensemble of machine learning models, a disease associated with the biological specimen based on the spectrum associated with the region of interest.
[0011] In some aspects, the techniques described herein relate to a method including: receiving a number of images, each of the number of images capturing a respective biological specimen including a number of droplets: receiving a number of respective spectrum associated with a number of respective regions of interest, wherein each respective region of interest is associated with one of the number of droplets within one of the number of images; identifying a set of the number of respective regions of interest having a high-value spectrum; creating a labeled dataset including a set of the number of images including the set of the number of respective regions of interest having the high-value spectrum: and training a model using the labeled dataset, wherein the trained model is configured to detect a region of interest having the high-value spectrum.
[0012] In some aspects, the techniques described herein relate to a method, wherein the model is a machine learning model. In some aspects, the techniques described herein relate to a method, wherein the machine learning model is a deep learning model. In some aspects, the techniques described herein relate to a method, wherein the machine learning model is an ensemble of machine learning models. In some aspects, the techniques described herein relate to a method, further including acquiring the number of images using a widefield imaging device. In some aspects, the techniques described herein relate to a method, further including acquiring the number of respective spectrum using a Raman spectrometer. In some aspects, the techniques described herein relate to a method, further including, prior to training the machine learning model using the labeled dataset, pre-training the machine learning model to perform a task.
[0013] In some aspects, the techniques described herein relate to a method including: receiving an image of a biological specimen including a number of droplets; analyzing, using a computer vision model, the image to identify the number of droplets within the biological specimen; analyzing, using the computer vision model, a number of respective magnified portions of the image including a number of regions of interest within the number of droplets; acquiring a number of respective spectrum associated with the number of regions of interest; and predicting, using a trained model, a disease associated with the biological specimen based on the number of respective spectrum associated with the number of regions of interest. In some aspects, the techniques described herein relate to a method, further including providing a diagnosis or prognosis for a subject based on the disease associated with the biological specimen. In some aspects, the techniques described herein relate to a method, further including administering a treatment to a subject based on the disease associated with the biological specimen. In some aspects, thetechniques described herein relate to a method, wherein the trained model is a trained machine learning model.BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The above and other aspects and features of the present disclosure will become more apparent to those skilled in the art from the following detailed description of the example embodiments with reference to the accompanying drawings.
[0015] FIG. 1 is a block diagram illustrating an example optical system for high-throughput chemical analysis according to an implementation described herein.
[0016] FIG. 2 is a flow diagram illustrating example operations for controlling an optical system for high-throughput chemical analysis according to an implementation described herein.
[0017] FIG. 3 is a flow diagram illustrating example operations for training a machine learning model according to an implementation described herein.
[0018] FIG. 4 is an example computing device.
[0019] FIG. 5 A is an example ~10x magnification image of two droplets.
[0020] FIG. 5B is an example high-magnification image showing a portion of the edge of a droplet of a biofluid specimen.
[0021] FIG. 6 is a flow diagram illustrating example operations for controlling an optical system for high-throughput chemical analysis including model training aspects.
[0022] FIGS. 7A-7C illustrate an example optical system for high-throughput chemical analysis.DETAILED DESCRIPTION
[0023] Referring generally to the FIGURES, described herein are systems and methods of high- throughput chemical analysis. Speaking generally, in many contexts it may be necessary or desirable to analyze a liquid biopsy (e.g., plasma, saliva, urine, etc.) to identify one or more features of the liquid biopsy. For example, in the context of cancer detection / diagnosis, it may be necessary to identify white blood cells within a biological sample. In some contexts, analyzing a liquid biopsy may be prohibitively time intensive. For example, it may take a substantial amount of time to analyze / characterize a sample using a conventional microscope imaging system due tothe small field of view required to identify molecules within the sample as well as other factors. Systems and methods of the present disclosure may offer benefits over conventional systems by increasing a speed and / or efficiency of analyzing liquid biopsies. An extended example follows: a high-throughput chemical analysis system may receive a number of samples. For example, the system may receive a biological sample including hundreds of droplets. The system may image the biological sample using a first field of view and may analyze the image using machine vision to identify one or more droplets of the hundreds of droplets. Additionally or alternatively, the system may identify a specific portion of a droplet such as an outer “coffee ring” of the droplet caused by evaporation. Marangoni flow inside the droplet may cause size-based separation of particles / molecules within the droplet. The system may identify a specific portion of a droplet based on a size of a particle / molecule it is trying to identify (e.g., white blood cells may cluster in a region of the droplet due to Marangoni flow). The system may image the specific portion of the droplet (e.g., using a spectroscopy unit, etc.) using a second field of view that may be less than the first field of view to identify one or more features (e.g., acquire a spectrum associated with a region of interest that includes a particle / molecule, etc.) within the droplet.
[0024] Referring now to FIG. 1, a block diagram of an optical system 100 for high-throughput chemical analysis is shown. In some implementations, the optical system 100 is optionally a Raman spectroscopy system. This disclosure contemplates that the Raman spectroscopy system can be configured for one or more of the following Raman modalities: transmission Raman Spectroscopy, reflection Raman Spectroscopy, stimulated Raman Spectroscopy, Surface- Enhanced Raman Spectroscopy (SERS), or Coherent Anti-Stokes Raman Spectroscopy (CARS). In other implementations, the optical system 100 is optionally an infrared (1R) spectroscopy system, a Fourier transform IR (FTIR) spectroscopy system, or Surface-Enhanced Infrared Absorption Spectroscopy (SEIRAS) system. In other words, it should be understood that Raman spectroscopy is only provided as an example technique used for high-throughput chemical analysis and that other techniques may be used. Accordingly, this disclosure contemplates that the “imaging” of the biological specimens as described herein should not be limited to conventional brightfield imaging but should include any technique that forms an “image,” including, but not limited to, fluorescence, super-resolution, reflectance, spectroscopy, etc.
[0025] As shown in FIG. 1, the optical system 100 can include a gantry 102, a scanning head 104, a sample stage 106, and a controller 108. The scanning head 104 is mounted to the gantry 102, which provides stability and precise movement over a defined workspace. For example, the gantry102 can be equipped with one or more motorized stages, allowing for automated positioning of the scanning head 104 over the sample stage 106. In some implementations, the scanning head 104 is movable relative to at least one axis (e.g. an x-, y-, and / or z-axis). In some implementations, the scanning head 104 is optionally movable relative to the x- and y-axes. In some implementations, the scanning head 104 is optionally movable relative to the x-, y-, and z-axes.
[0026] The scanning head 104 includes the optical elements and detectors. For example, as shown in FIG. 1, the scanning head 104 includes an optical imaging unit 120, which is configured to acquire images, and a spectroscopy unit 130, which is configured to deliver and collect light. For example, the optical system 100 can include a spectroscopy light source and a detector. In a Raman spectroscopy implementation, the light source is a laser and the detector is a spectrometer. The laser source emits a narrow wavelength range suitable for Raman spectroscopy, and the laser beam is directed onto a sample, inducing Raman scattering. The spectrometer captures the scattered light, facilitating the analysis of molecular vibrations characteristic of the sample. In some implementations, at least one of the spectroscopy light source (e.g. laser) and the detector (e.g. spectrometer) is integrated into the spectroscopy unit 130. In other implementations, both the spectroscopy light source (e.g. laser) and the detector (e.g. spectrometer) are integrated into the spectroscopy unit 130. Additionally, the optical imaging unit 120 can be an imaging device (e.g., camera), for example optionally a widefield imaging device. In various embodiments, the scanning head 104 is configured to capture images having different field of views. For example, the scanning head 104 may capture a first image having a first field of view (e.g., using a first imaging device, etc.) and may capture a second image (e.g., using the first imaging device and / or a second imaging device, etc.) having a second field of view that is less than the first field of view (e.g., is zoomed in relative to the first image).
[0027] The sample stage 106 is configured to hold a substrate including a biological specimen. In other words, the sample stage 106 holds the substrate on which the samples (e.g. biological specimens) are placed for analysis. This disclosure contemplates that the substrate can be made of various materials including, but not limited to, glass, quartz, calcium fluoride, etc. and optionally SERS substrate materials. The sample stage 106 is designed for compatibility with various sample types. The sample stage 106 incorporates precise sample positioning mechanisms to ensure accurate alignment with the scanning head 104. In some implementations, the sample stage 106 is movable relative to at least one axis (e.g. an x-, y-, and / or z-axis). In some implementations, the sample stage 106 is optionally movable relative to the z-axis. In some implementations, the samplestage 106 is optionally movable relative to the x-, y-, and z-axes. Optionally, the sample stage 106 can include environmental control features to accommodate different temperature and humidity conditions.
[0028] In some aspects, the biological specimen is a biofluid specimen. The biofluid specimen may be blood, serum, plasma, saliva, urine, cerebrospinal fluid, synovial fluid, tears, breast milk, amniotic fluid, semen, ascites, bronchial lavage, or combinations thereof. Alternatively, in some aspects, the biological specimen may be material isolated from a biological sample, for example, exosomes and related extracellular vesicles, lipoprotein, metabolites, or other biological materials. Optionally, the biofluid specimen is a dried biofluid specimen. Alternatively, the biofluid specimen is a wet biofluid specimen. Alternatively, the biofluid specimen is between the wet and dry states (e.g. portions are wet and portion are dry). Alternatively, the biological specimen is another type of deposited biological sample, including, but not limited to, a cell or tissue lysate.
[0029] The controller 108 is configured to provide an interface that allows users to define scanning parameters, analyze results in real-time, and generate comprehensive reports. This disclosure contemplates that the controller 108 includes at least a processor and memory (e.g. the most basic configuration illustrated by box 402 in FIG. 4). As described herein, the controller 108 is configured to execute machine learning algorithms for pattern recognition and automated identification of chemical compositions. For example, the controller 108 may execute a machine vision algorithm to identify one or more “coffee rings” of one or more droplets. The controller 108 can be operably connected to the gantry 102, for example, through one or more communication links. In particular, the controller 108 can be operably connected to one or more actuators that facilitate movement of the gantry 102. As described herein, the controller 108 can be configured to control the relative position of the scanning head 104 relative to the sample stage 106. This disclosure contemplates that the communication links described above are any suitable communication link. For example, a communication link may be implemented by any medium that facilitates data exchange including, but not limited to, wired, wireless and optical links. Accordingly, the controller 108 can exchange data with the gantry 102. The controller 108 can therefore be configured to control the positioning of the scanning head 104 relative to the sample stage 106.
[0030] Additionally, the controller 108 can optionally be operably connected to the sample stage 106, for example, through one or more communication links. As described herein, the controller 108 can be configured to control the relative position of the scanning head 104 relative to thesample stage 106. Accordingly, the controller 108 can exchange data with the sample stage 106. The controller 108 can therefore be configured to control the positioning of the scanning head 104 relative to the sample stage 106. Alternatively or additionally, the sample stage 106 can be operated manually by a user. Additionally, the controller 108 can be operably connected to the scanning head 104, for example, through one or more communication links. Optionally, the controller 108 is operably connected to the optical imaging unit 120 through one or more first communication links. Optionally, the controller 108 is operably connected to the spectroscopy unit 130 through one or more second communication links. Accordingly, the controller 108 can exchange data with the optical imaging unit 120 and / or the spectroscopy unit 130. The controller 108 can therefore be configured to control respective operations of the optical imaging unit 120 and the spectroscopy unit 130 and / or receive imaging data.
[0031] The controller 108 is configured to receive an image of a biological specimen including a plurality of droplets. An example ~10x magnification image of two droplets is shown in FIG. 5 A. The image is acquired by the optical imaging unit 120 and transmitted to the controller 108. Optionally, as described above, the optical imaging unit 120 is a widefield imaging device. Additionally, as described above, the biological specimen is optionally a biofluid specimen such as blood, serum, plasma, saliva, urine, cerebrospinal fluid, synovial fluid, tears, breast milk, amniotic fluid, semen, ascites, bronchial lavage, or combinations thereof. This disclosure contemplates that the biofluid specimen is in a wet, dry, or intermediate (e.g. between wet and dry state). Accordingly, as used herein, the term “droplets” is intended to describe wet spots, dry spots, or drops in an intermediate state (e.g. between wet and dry). In the example described below, the biological specimen captured in the image acquired by the optical imaging unit 120 is a dried biofluid specimen. It should be understood that a dried biofluid specimen is provided only as an example biological specimen.
[0032] The controller 108 is also configured to analyze the image to identify a location of at least one droplet of the plurality of droplets within the dried biofluid specimen. This can be accomplished using a computer vision module. As used herein, a computer vision module is configured to interpret and analyze visual information from the images acquired by the optical imaging unit 120. For example, the computer vision module executes algorithms and processes that enable a machine or machines to recognize and understand the content of image data. Such algorithms and processes may include, but are not limited to, object detection, image segmentation, edge detection, and image classification. Optionally, the computer vision module executesmachine learning models, which may include, but are not limited to, supervised learning models such as support vector machines (S VMs), random forests, and ANNs. The computer vision module facilitates extraction of meaningful insights from visual data. In some implementations, the computer vision module uses an image segmentation algorithm such as a convolutional neural network (CNN), for example the U-Net architecture, when identifying the location of the at least one droplet in the image. Alternatively or additionally, the computer vision module uses an object detection algorithm, for example the YOLO (You Only Look Once) framework or the SSD (Single Shot Multibox Detector) framework, when identifying the location of the at least one droplet in the image. Alternatively or additionally, the computer vision module uses an edge detection algorithm to detect a boundary of the at least one droplet in the image when identifying the location of the at least one droplet in the image. As described above, the dried biofluid specimen includes a plurality of droplets. Thus, in some implementations, the controller 108 is configured to analyze the image to identify respective locations of each of a plurality of droplets within the dried biofluid specimen. For example, the operations described below can be repeated for each of a plurality of droplets within the dried biofluid specimen.
[0033] As described above, the computer vision module may use a trained machine learning model (e.g. first trained machine learning model) to detect the at least one droplet in the image. Such machine learning model can be trained to perform image segmentation, objection detection, or edge detection. It should be understood that the computer vision module may employ one or more trained machine learning models, each of which is trained to perform a different task. For example, a machine learning algorithm can be trained to distinguish droplets from background and / or other portions of the dried biofluid specimen. For example, supervised machine learning models can be trained using a labeled image dataset, where labels notate droplets. The labels serve as the ground truth for training so that the trained model is capable of detecting droplets. Optionally, the trained machine learning model is a deep learning model. For example, the trained machine learning model may be a CNN, which is a type of deep neural network that has been applied, for example, to image analysis applications. Unlike a traditional neural networks, each layer in a CNN has a plurality of nodes arranged in three dimensions (width, height, depth). CNNs can include different types of layers, e.g., convolutional, pooling, and fully-connected (also referred to herein as “dense”) layers. A convolutional layer includes a set of filters and performs the bulk of the computations. A pooling layer is optionally inserted between convolutional layers to reduce the computational power and / or control overfitting (e.g., by downsampling). A fully-connected layerincludes neurons, where each neuron is connected to all of the neurons in the previous layer. The layers are stacked similar to traditional neural networks. Optionally, the computer vision module may use an ensemble of trained machine learning models (e.g. a first ensemble) to detect the at least one droplet in the image. An ensemble is a meta-classifier that combines a plurality of machine learning models for classification via majority voting. In other words, the ensemble’s final prediction (e.g., class label) is the one predicted most frequently by the member machine learning models.
[0034] The controller 108 is also configured to control the scanning head 104 to position the spectroscopy unit 130 in proximity to the location of the at least one droplet. For example, once the location of the at least one droplet is detected, the controller 108 sends commands to one or more actuators that facilitate movement of the gantry 102 in order to properly position the scanning head 104 relative to the sample stage 106, which facilitates spectral analysis of the at least one droplet. The gantry 102 moves (e.g. repositions) in response to the commands.
[0035] The controller 108 is also configured to analyze a magnified portion of the image including the at least one droplet to identify a region of interest within the at least one droplet. An example high magnification image is shown in FIG. 5B. The image of FIG. 5B shows the drying edge 502 of a droplet. The region of interest includes a feature likely to have a high-value spectrum. As used herein, a high-value spectrum is associated with disease (e.g. cancer). In some implementations, a machine learning algorithm can be trained to distinguish between a region of interest having a high-value spectrum (e.g. one associated with disease) and a region of interest having a low-value spectrum (e.g. one not associated with disease). For example, supervised machine learning models can be trained using a labeled image dataset, where labels notate features with high-value spectra. The labels serve as the ground truth for training so that the trained model is capable of distinguishing regions of interest having high-value spectra from regions of interest having low- value spectra.
[0036] As described above, the computer vision module may use a trained machine learning model (e.g. second trained machine learning model) to identify a region of interest within the at least one droplet. Such machine learning model can be trained to perform image segmentation, objection detection, or edge detection. It should be understood that the computer vision module may employ one or more trained machine learning models, each of which is trained to perform a different task. Optionally, the trained machine learning model is a deep learning model. For example, the trained machine learning model may be a CNN, which is a type of deep neural network that has beenapplied, for example, to image analysis applications. Optionally, the computer vision module may use an ensemble of trained machine learning models (e.g. a second ensemble) to identify a region of interest within the at least one droplet.
[0037] Alternatively or additionally, the controller 108 can optionally be configured to analyze spectral features to identify a region of interest within the at least one droplet. For example, the controller 108 can be configured to control the spectroscopy unit 130 to acquire low quality, fast spectra taken at low resolution. The controller 108 can be configured to analyze such low- resolution spectra to identify the region of interest within the at least one droplet. Optionally, a machine learning model can be employed to analyze the low-resolution spectrum.
[0038] After identifying the region of interest, the controller 108 is also configured to control the spectroscopy unit 130 to acquire a spectrum associated with the region of interest within the at least one droplet. For example, once the region of interest within the at least one droplet is identified, the controller 108 sends commands to the spectroscopy unit 130, for example, to direct a light source (e.g. laser) such that it scans the region of interest. In other words, the controller 108 is configured to systematically control the light source in a precise manner. A detector (e.g. spectrometer) collects the scattered light. An example region of interest 504 that is interrogated by the spectroscopy unit 130 is shown in FIG. 5B.
[0039] Optionally, in some implementations, the controller 108 is also configured to predict a disease (e.g. cancer) associated with the dried biofluid specimen based on the spectrum associated with the region of interest. The spectrum is acquired using the spectroscopy unit 130 as described above. It should be understood that presence of disease is assessed based on the spectrum (or spectra) associated with the region (or regions) of interest. It is not possible to predict disease by analyzing the image of the dried biofluid specimen as a whole, i.e. the information contained in the image as a whole is too general. In contrast, by detecting droplets and then identifying regions of interest for further spectrometry analysis as described herein, it is possible to obtain information (e.g. spectra for regions of interest) that is useful for prediction of disease.
[0040] In some implementations, a machine learning algorithm can be trained to distinguish between a spectrum associated with disease and a spectrum not associated with disease. For example, such supervised machine learning model can be trained using a labeled dataset, where labels notate spectra associated with disease. The labels serve as the ground truth for training so that the trained model is capable of distinguishing spectra associated with disease from spectra notassociated with disease. It should be understood that using machine learning to distinguish spectra associated with disease from spectra not associated with disease is provided only as an example. This disclosure contemplates using machine learning to analyze spectra for other applications. Non- limiting examples include, but are not limited to, environmental analysis (e.g., nanoplastics within spots of dried waste water).
[0041] As described above, a trained machine learning model (e.g. third trained machine learning model) may be used to predict a disease (e.g. cancer) associated with the dried biofluid specimen based on the spectrum associated with the region of interest. Such machine learning models can be trained to perform classification. It should be understood that one or more trained machine learning models, each of which is trained to perform a different task, may be employed. Optionally, the trained machine learning model is an SVM or ANN. Optionally, the trained machine learning model is a deep learning model such as a deep ANN. Optionally, an ensemble of trained machine learning models (e.g. a third ensemble) may be used to predict the disease associated with the dried biofluid specimen based on the spectrum associated with the region of interest.
[0042] EXAMPLE METHODS
[0043] Referring now to FIG. 2, a flowchart of an example method for controlling an optical system for high-throughput chemical analysis is shown. This disclosure contemplates that the optical system may be the optical system 100 shown in FIG. 1. Additionally, this disclosure contemplates that the logical operations of FIG. 2 can be performed using a computing device (e.g., controller 108 of FIG. 1 and / or computing device 400 of FIG. 4).
[0044] At step 210, an image of a biological specimen is received, for example at a computing device. The image is received from an optical imaging unit (e.g., optical imaging unit 120 of FIG. 1). The biological specimen includes a plurality of droplets. Optionally, in some implementations, the biological specimen is a dried biofluid specimen. Biological specimens and images thereof are described in detail above.
[0045] At step 220, the image is analyzed, for example by the computing device, to identify a location of at least one droplet of the plurality of droplets within the biological specimen. Optionally, the computing device uses a computer vision module to analyze the image. Optionally, the computer vision module uses one or more trained machine learning models to identify the location of at least one droplet within the biological specimen. The computer vision module is described in detail above.
[0046] In some embodiments, step 220 includes preprocessing the image. For example, the image may be preprocessed to enhance the accuracy of droplet identification by addressing sources of variability in Raman imaging. Preprocessing may include: (1) cosmic ray removal (e.g., applying a statistical outlier rejection algorithm, such as median filtering or wavelet-based denoising, to suppress transient high-intensity artifacts associated with cosmic ray interference, etc.); (2) baseline correction for non-uniform backgrounds (e.g., implementing adaptive polynomial fitting, such as asymmetric least squares regression and / or a rolling ball correction algorithm, to compensate for background fluorescence and stray light, etc.); (3) spectral denoising pnor to feature extraction (e.g., employing filtering techniques, such as a Savitzky-Golay smoothing filter and / or a discrete wavelet transform, to enhance signal quality while preserving fine spectral features, etc.); and / or (4) applying hybrid edge detection and / or deep learning ensemble (e.g., identifying droplet boundaries using an ensemble approach, which may include a U-Net-based segmentation model in conjunction with an active contour algorithm, to improve segmentation robustness under varying illumination and substrate conditions, etc.).
[0047] In some embodiments, a self-supervised model may be trained to predict spectral variance across different samples (e.g., enabling the self-supervised model to generalize across a range of experimental conditions). This training process may include, for example, (1) unlabeled spectral data augmentation: generating contrastive learning pairs by artificially perturbing Raman spectral images, allowing the model to learn invariant representations of droplet morphology and composition; and / or (2) biofluid heterogeneity normalization: implementing a domain-invariant training objective, wherein the model learns to extract spatial and spectral features that remain stable across different biofluid compositions, such as serum, plasma, or saliva samples.
[0048] To minimize inter-experiment batch effects, in some embodiments, the system may employ fine-tuning methods using domain-adversarial training, allowing it to align feature distributions across different acquisition conditions. This may include, for example, (1) feature distribution alignment across Raman instruments: implementing domain- adaptive feature matching to correct for spectral shifts introduced by differences in excitation wavelengths, detector sensitivity, or laser stability between different Raman systems; (2) cross-batch normalization for sample consistency: applying batch normalization or adaptive instance normalization techniques to reduce variance between sample runs, thereby improving segmentation and classification consistency; and / or (3) transfer learning for model adaptation: fine-tuning a pre-trained deep learning model using small,task-specific datasets, wherein the system gradually adapts learned representations from a broad Raman dataset to a specific experimental setup.
[0049] At step 230, a scanning head (e.g. scanning head 104 of FIG. 1) comprising a spectroscopy unit (e.g. spectroscopy unit 130 of FIG. 1) is controlled, for example by the computing device, to position the spectroscopy unit in proximity to the location of the at least one droplet, which is disposed on a sample stage (e.g. sample stage 106 of FIG. 1). Control of the scanning head, for example using a gantry (e.g. gantry 102 of FIG. 1), relative to the sample stage is described in detail above.
[0050] At step 240, a magnified portion of the image including the at least one droplet is analyzed, for example by the computing device, to identify a region of interest within the at least one droplet. Optionally, the computing device uses a computer vision module to analyze the magnified portion of the image. Optionally, the computer vision module uses one or more trained machine learning models to identify a region of interest within the at least one droplet. The computer vision module is described in detail above. Optionally, the computing device analyzes spectral features to identify a region of interest within the at least one droplet. For example, the computing device can control the spectroscopy unit to acquire low-resolution spectra as described in detail above and then analyze the same to identify the region of interest. Optionally, one or more trained machine learning models may be used to analyze low-resolution spectra.
[0051] In some embodiments, in addition to or rather than statically defining regions of interest (ROIs) based on pre-selected features, the system dynamically refines ROIs using a spectral analysis-driven feedback loop. Dynamically refining ROIs may include an iterative spectral feature selection step, in which an initial set of low-resolution Raman spectra is acquired across the droplet to assess biochemical variability.
[0052] Dynamically refining ROIs may further include classifying potential regions of interest using a trained spectral analysis model based on spectral features associated with biochemical relevance, such as peak shifts indicating molecular interactions (e.g., protein aggregation, lipid oxidation, or metabolic byproducts), fluorescence background trends that may correlate with cellular degradation or biomarker concentration, and / or spectral clustering techniques (e.g., principal component analysis (PCA) or t-SNE) to identify outliers indicative of biologically significant regions.
[0053] In some embodiments, in response to identifying key spectral features, the system applies a spatial feedback mechanism that maps extracted spectral information onto the full optical image. This mapping may be performed using a feature-weighted spatial interpolation method, such as Gaussian process regression or Kriging interpolation, to estimate high-value spectral regions. Additionally, the imaging model may refine its search space by dynamically adjusting priority scoring, wherein regions exhibiting strong Raman scattering intensity and / or unique spectral signatures are assigned higher sampling weight. Additionally or alternatively, spatial-spectral correlation metrics, such as normalized spectral similarity or gradient-based attention maps, may be used to fine-tune ROI boundaries (e.g., in real time or near-real time, etc.).
[0054] In some embodiments, the system includes adaptive scanning with spectroscopy model feedback to dynamically re-prioritize sampling locations. Adaptive scanning may increase scanning efficiency. In various embodiments, the spectroscopy model updates (e.g., continuously or semi-continuously, etc.) spatial-spectral probability maps based on newly acquired spectra and / or deprioritizes low-information or redundant acquisitions using entropy-based scanning strategies. The system may select high-value regions for additional high-resolution spectral scans, focusing on areas that exhibit diagnostically significant spectral patterns. In various embodiments, this allows the system to “self- improve” (e.g., integrate iterative adjustments between the imaging model and spectral analysis module, progressively refining spectral maps across the droplet and increasing the probability of identifying high-value spectra while minimizing redundant acquisitions, etc.).
[0055] At step 250, the spectroscopy unit is controlled, for example by the computing device, to acquire a spectrum associated with the region of interest within the at least one droplet. Control of the spectrometry unit, for example to scan the region of interest with a light source and collect scattered light, is described in detail above. It should be understood that steps 220-250 can be repeated such that each region of interest in each droplet is subjected to the spectroscopy analysis.
[0056] Referring now to FIG. 3 , a flowchart of an example method for training a machine learning model is shown. This disclosure contemplates that the logical operations of FIG. 3 can be performed using a computing device (e.g. computing device 400 of FIG. 4).
[0057] At step 310, a plurality of images are received, for example by a computing device. Images can optionally be acquired using a widefield imaging device. Each of the plurality of images captures a respective biological specimen comprising a plurality of droplets. Optionally, in someimplementations, the biological specimen is a dried biofluid specimen. Biological specimens and images thereof are described in detail above. It should be understood the plurality of images form a dataset. Such a dataset may include a plurality of high-resolution images of biological specimens, a plurality of magnified portions of high-resolution images of biological specimens, or combinations thereof.
[0058] At step 320, a plurality of respective spectrum associated with a plurality of respective regions of interest are received, for example by a computing device. Each respective region of interest is associated with one of the plurality of droplets within one of the plurality of images. Spectra can be acquired using spectrometer. Optionally, spectra can optionally be acquired using a Raman spectrometer. It should be understood that some of the spectra would be of high value (e.g. associated with disease) and some of the spectra would be of low value (e.g. not associated with disease).
[0059] At step 330, a set of the plurality of respective regions of interest having a high-value spectrum are identified. This disclosure contemplates differentiating between high-value and low- value spectra manually (e.g. expert review), with algorithms, or combinations thereof.
[0060] At step 340, a labeled dataset comprising a set of the plurality of images including the set of the plurality of respective regions of interest having the high-value spectrum is created. In other words, regions of interest high-value spectra are labeled. As described herein, the labels serve as ground truth for model training. Additionally, the labeled dataset may include a plurality of high- resolution images of biological specimens with labels for high-value spectrum droplet regions, a plurality of magnified portions of high-resolution images of biological specimens with labels for high-value spectrum droplet regions, or combinations thereof.
[0061] At step 350, a model is trained using the labeled dataset. The model is trained by maximizing or minimizing an objective function. For example, the objective function may be an error between the model’s output (e.g. a prediction) and ground truth (e.g. a label). In some implementations, the model is a mathematical model such as a regression. In some implementations, the model is a machine learning model. Optionally, the machine learning model is a deep learning model. Optionally, the machine learning model is an ensemble of machine learning models. After completion of training, the trained model is configured to detect a region of interest having the high-value spectrum.
[0062] Optionally, in some implementations, transfer learning is employed. According to transfer learning, the model is pre- trained (e.g. prior to step 350) using a dataset to perform a task. Typically, the dataset and / or the task for pre-training is different than the labeled dataset and / or task (i.e. detecting a region of interest having the high-value spectrum) for training at step 350. The objective of transfer learning is to leverage knowledge gained from one task (i.e. the pretraining task) to improve the learning and performance of a related, but different, task (i.e. step 350). Thus, transfer learning involves training the model on one dataset and then applying the knowledge gained to a different but related dataset. This approach is particularly useful for the application described herein because the labeled dataset (i.e. step 340) may be limited and / or expensive to acquire.
[0063] In some aspects, the techniques described herein relate to a method including: receiving an image of a biological specimen including a plurality of droplets; analyzing, using a computer vision model, the image to identify the plurality of droplets within the biological specimen; analyzing, using the computer vision model, a plurality of respective magnified portions of the image including a plurality of regions of interest within the plurality' of droplets; acquiring a plurality of respective spectrum associated with the plurality of regions of interest; and predicting, using a trained model, a disease associated with the biological specimen based on the plurality of respective spectrum associated with the plurality of regions of interest.
[0064] In some aspects, the method further includes providing a diagnosis or prognosis for a subject based on the disease associated with the biological specimen.
[0065] In some aspects, the method further includes administering a treatment to a subject based on the disease associated with the biological specimen.
[0066] In some aspects, the trained model is a trained machine learning model.
[0067] EXAMPLE COMPUTING DEVICE
[0068] It should be appreciated that the logical operations described herein with respect to the various figures may be implemented (1) as a sequence of computer implemented acts or program modules (i.e., software) running on a computing device (e.g., the computing device described in FIG. 4), (2) as interconnected machine logic circuits or circuit modules (i.e., hardware) within the computing device and / or (3) a combination of software and hardware of the computing device. Thus, the logical operations discussed herein are not limited to any specific combination of hardware and software. The implementation is a matter of choice dependent on the performanceand other requirements of the computing device. Accordingly, the logical operations described herein are referred to variously as operations, structural devices, acts, or modules. These operations, structural devices, acts and modules may be implemented in software, in firmware, in special purpose digital logic, and any combination thereof. It should also be appreciated that more or fewer operations may be performed than shown in the figures and described herein. These operations may also be performed in a different order than those described herein.
[0069] Referring to FIG. 4, an example computing device 400 upon which the methods described herein may be implemented is illustrated. It should be understood that the example computing device 400 is only one example of a suitable computing environment upon which the methods described herein may be implemented. Optionally, the computing device 400 can be a well-known computing system including, but not limited to, personal computers, servers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, network personal computers (PCs), minicomputers, mainframe computers, embedded systems, and / or distributed computing environments including a plurality of any of the above systems or devices. Distributed computing environments enable remote computing devices, which are connected to a communication network or other data transmission medium, to perform various tasks. In the distributed computing environment, the program modules, applications, and other data may be stored on local and / or remote computer storage media.
[0070] In its most basic configuration, computing device 400 typically includes at least one processing unit 406 and system memory 404. Depending on the exact configuration and type of computing device, system memory 404 may be volatile (such as random access memory (RAM)), non-volatile (such as read-only memory (ROM), flash memory, etc.), or some combination of the two. This most basic configuration is illustrated in FIG. 4 by box 402. The processing unit 406 may be a standard programmable processor that performs arithmetic and logic operations necessary for operation of the computing device 400. The computing device 400 may also include a bus or other communication mechanism for communicating information among various components of the computing device 400.
[0071] Computing device 400 may have additional features / functionality. For example, computing device 400 may include additional storage such as removable storage 408 and non-removable storage 410 including, but not limited to, magnetic or optical disks or tapes. Computing device 400 may also contain network connection(s) 416 that allow the device to communicate with other devices. Computing device 400 may also have input device(s) 414 such as a keyboard, mouse,touch screen, etc. Output device(s) 412 such as a display, speakers, printer, etc. may also be included. The additional devices may be connected to the bus in order to facilitate communication of data among the components of the computing device 400. All these devices are well known in the art and need not be discussed at length here.
[0072] The processing unit 406 may be configured to execute program code encoded in tangible, computer-readable media. Tangible, computer-readable media refers to any media that is capable of providing data that causes the computing device 400 (i.e., a machine) to operate in a particular fashion. Various computer-readable media may be utilized to provide instructions to the processing unit 406 for execution. Example tangible, computer-readable media may include, but is not limited to, volatile media, non-volatile media, removable media and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. System memory 404, removable storage 408, and non-removable storage 410 are all examples of tangible, computer storage media. Example tangible, computer-readable recording media include, but are not limited to, an integrated circuit (e.g., field-programmable gate array or application-specific IC), a hard disk, an optical disk, a magneto-optical disk, a floppy disk, a magnetic tape, a holographic storage medium, a solid-state device, RAM, ROM, electrically erasable program read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices.
[0073] In an example implementation, the processing unit 406 may execute program code stored in the system memory 404. For example, the bus may carry data to the system memory 404, from which the processing unit 406 receives and executes instructions. The data received by the system memory 404 may optionally be stored on the removable storage 408 or the non-removable storage 410 before or after execution by the processing unit 406.
[0074] It should be understood that the various techniques described herein may be implemented in connection with hardware or software or, where appropriate, with a combination thereof. Thus, the methods and apparatuses of the presently disclosed subject matter, or certain aspects or portions thereof, may take the form of program code (i.e., instructions) embodied in tangible media, such as floppy diskettes, CD-ROMs, hard drives, or any other machine-readable storage medium wherein, when the program code is loaded into and executed by a machine, such as a computing device, the machine becomes an apparatus for practicing the presently disclosed subject matter. In the case of program code execution on programmable computers, the computing device generallyincludes a processor, a storage medium readable by the processor (including volatile and nonvolatile memory and / or storage elements), at least one input device, and at least one output device. One or more programs may implement or utilize the processes described in connection with the presently disclosed subject matter, e.g., through the use of an application programming interface (API), reusable controls, or the like. Such programs may be implemented in a high level procedural or object-oriented programming language to communicate with a computer system. However, the program(s) can be implemented in assembly or machine language, if desired. In any case, the language may be a compiled or interpreted language and it may be combined with hardware implementations.
[0075] Examples
[0076] The following examples are put forth so as to provide those of ordinary skill in the art with a complete disclosure and description of how the compounds, compositions, articles, devices and / or methods claimed herein are made and evaluated, and are intended to be purely exemplary and are not intended to limit the disclosure. Efforts have been made to ensure accuracy with respect to numbers (e.g., amounts, temperature, etc.), but some errors and deviations should be accounted for. Unless indicated otherwise, parts are parts by weight, temperature is in DC or is at ambient temperature, and pressure is at or near atmospheric.
[0077] This example describes a scanning optical device with built-in machine vision for high- throughput chemical analysis of liquid biopsy (e.g., plasma, saliva, urine) samples. The scanning head is mounted to a gantry and coupled to a robotic device frame via mirrors or optical fibers, and features: a Raman optical unit (both illumination and collection optics), and a microscope optical imaging system (e.g., CMOS camera). FIG. 6 illustrates example operations for controlling such an optical system including model training aspects. An example system is shown in Figs. 7A- 7C.
[0078] The scanning head performs two associated functions for high-throughput analysis of small volumes (e.g., droplets of < 10 pL) of liquid biopsy specimens. The microscope imaging system directs the location of the scanning head to areas of interest within each sample droplet via machine vision, while the Raman optical unit interrogates those areas via chemical spectroscopic measurements. This setup permits rapid, high-area scanning across, for example, hundreds of deposited liquid biopsy specimens for rapid diagnostics application. The system uses machine vision to scan within samples according to select microscopic features of interest. Furthermore, 1current Raman spectroscopy platforms are not compatible with robotic technologies, instead featuring fixed optical heads. The scanning head described herein is designed to be integrated into a robotic platform, permitting high-throughput operation amenable to high-volume clinical lab environments.
[0079] The system includes integrated modular Raman scanning head within a robotic controlled unit, compatible with multiple sources of illumination and detection units. The illumination and detection optics are modular, multiple types of units can be easily integrated into the scanning head design. In one implementation, the illumination source is a laser probe mounted to the scanning head. In another implementation, the illumination is coupled via mirrors mounted to the device frame. Similarly, collection optics can be coupled to an external detector unit (e.g., CMOS camera, photodiode) via integrated fiber optics or mirror mounted to the device frame. This flexibility in design permits focus on the rapid movement and large area of view compared to current systems.
[0080] The system includes automated sampling within a droplet driven by via machine vision using a microscopic imaging system integrated to the scanning head. This aspect describes the integration of a widefield camera adjacent to the Raman unit on the scanning head. This imaging system enables automatic recognition of the samples to be interrogated by machine learning training (also known as computer vision). In one implementation, the sample of interest would feature hundreds of droplets (few microliters in volume) arrayed onto the sample stage beneath the scanning unit head. This imaging system would automatically detect and register the exact shape and location of each droplet and drive the robotic system to iterate Raman measurements at desired spots within each droplet.
[0081] The design enables rapid scanning over large areas in automated fashion. A stage containing the samples of interest moves in one dimension (e.g. z-direction), while the scanning head operates in the other two dimensions (e.g. x- and y-direction). This design enables a large area to be imaged in a very short time. More generally, the scanning optical heads can be combined in succession (e.g. multiple scanning heads coupled to a single robotic unit) enabling vast numbers of samples to be interrogated at once.
[0082] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure. As used in the specification, and in the appended claims, the singular forms “a,” “an,”“the” include plural referents unless the context clearly dictates otherwise. The term “comprising” and variations thereof as used herein is used synonymously with the term “including” and variations thereof and are open, non-limiting terms. The terms “optional” or “optionally” used herein mean that the subsequently described feature, event or circumstance may or may not occur, and that the description includes instances where said feature, event or circumstance occurs and instances where it does not. Ranges may be expressed herein as from "about" one particular value, and / or to "about" another particular value. When such a range is expressed, an aspect includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent "about," it will be understood that the particular value forms another aspect. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint.
[0083] As used herein, the terms "about" or "approximately" when referring to a measurable value such as an amount, a percentage, and the like, is meant to encompass variations of ±20%, ±10%, ±5%, or ±1% from the measurable value.
[0084] “Administration” of “administering” to a subject includes any route of introducing or delivering to a subject an agent. Administration can be carried out by any suitable means for delivering the agent. Administration includes self- administration and the administration by another.
[0085] The term “subject” is defined herein to include animals such as mammals, including, but not limited to, primates (e.g., humans), cows, sheep, goats, horses, dogs, cats, rabbits, rats, mice and the like. In some embodiments, the subject is a human.
[0086] The term “artificial intelligence” is defined herein to include any technique that enables one or more computing devices or comping systems (i.e., a machine) to mimic human intelligence. Artificial intelligence (Al) includes, but is not limited to, knowledge bases, machine learning, representation learning, and deep learning. The term “machine learning” is defined herein to be a subset of Al that enables a machine to acquire knowledge by extracting patterns from raw data. Machine learning techniques include, but are not limited to, logistic regression, support vector machines (SVMs), decision trees, Naive Bayes classifiers, and artificial neural networks. The term “representation learning” is defined herein to be a subset of machine learning that enables a machine to automatically discover representations needed for feature detection, prediction, orclassification from raw data. Representation learning techniques include, but are not limited to, autoencoders. The term “deep learning” is defined herein to be a subset of machine learning that that enables a machine to automatically discover representations needed for feature detection, prediction, classification, etc. using layers of processing. Deep learning techniques include, but are not limited to, artificial neural network or multilayer perceptron (MLP).
[0087] Machine learning models include supervised, semi-supervised, and unsupervised learning models. In a supervised learning model, the model learns a function that maps an input (also known as feature or features) to an output (also known as target or targets) during training with a labeled data set (or dataset). In an unsupervised learning model, the model learns patterns (e.g., structure, distribution, etc.) within an unlabeled data set. In a semi-supervised model, the model learns a function that maps an input (also known as feature or features) to an output (also known as target or target) during training with both labeled and unlabeled data.
[0088] One example machine learning model is an artificial neural network (ANN). An ANN is a computing system including a plurality of interconnected neurons (e.g., also referred to as “nodes”). This disclosure contemplates that the nodes can be implemented using a computing device (e.g., a processing unit and memory as described herein). The nodes can be arranged in a plurality of layers such as input layer, output layer, and optionally one or more hidden layers. An ANN having hidden layers can be referred to as deep neural network or multilayer perceptron (MLP). Each node is connected to one or more other nodes in the ANN. For example, each layer is made of a plurality of nodes, where each node is connected to all nodes in the previous layer. The nodes in a given layer are not interconnected with one another, i.e., the nodes in a given layer function independently of one another. As used herein, nodes in the input layer receive data from outside of the ANN, nodes in the hidden layer(s) modify the data between the input and output layers, and nodes in the output layer provide the results. Each node is configured to receive an input, implement an activation function (e.g., binary step, linear, sigmoid, tanH, or rectified linear unit (ReLU) function), and provide an output in accordance with the activation function. Additionally, each node is associated with a respective weight. ANNs are trained with a dataset to maximize or minimize an objective function. In some implementations, the objective function is a cost function, which is a measure of the ANN’s performance (e.g., error such as LI or L2 loss) during training, and the training algorithm tunes the node weights and / or bias to minimize the cost function. This disclosure contemplates that any algorithm that finds the maximum or minimum of the objective function can be used for training the ANN. Training algorithms for ANNs include,but are not limited to, backpropagation. It should be understood that an artificial neural network is provided only as an example machine learning model. It should be understood that an artificial neural network is provided only as an example machine learning model. This disclosure contemplates that the machine learning model can be any supervised learning model, semisupervised learning model, or unsupervised learning model.
[0089] In some embodiments, machine learning models are enhanced to improve spectral analysis, optimize region selection, and / or refine longitudinal data interpretation. For time-dependent measurements, such as tracking metabolic shifts in biofluid samples over time, the system may incorporate recurrent neural networks (RNNs) for longitudinal spectral trends. Examples include Long Short-Term Memory (LSTM) and / or Gated Recurrent Unit (GRU) architectures, which enable the model to retain spectral context across different time points. By preserving temporal dependencies, the system can identify subtle spectral variations that could be overlooked in singlepoint analyses, thereby improving the detection of evolving biochemical trends.
[0090] In some embodiments, the system improves spectral feature discrimination by using contrastive learning for spectral similarity identification, wherein a contrastive loss framework is used to train models to group spectrally similar regions while distinguishing dissimilar features. Contrastive learning may enable the model to leam robust spectral representations, thereby improving generalization across diverse sample types and mitigating variations introduced by experimental noise. Additionally or alternatively, contrastive learning may improve the system’s ability to recognize diagnostically relevant spectral signatures even in the presence of background variability.
[0091] Additionally, the system may integrate multi-resolution learning for spatial-spectral coherence, wherein models are trained on spectral and imaging datasets at varying resolutions to maintain fine-scale molecular feature analysis and / or a broader morphological context. In some embodiments, fusing high-resolution spectral information with lower resolution widefield imaging data ensures that detailed biochemical signals are interpreted within the structural and spatial context of the droplet. This multi-scale approach may improve the accuracy of ROI selection and / or spectral classification, facilitating more precise diagnostic predictions while preserving the integrity of spatial relationships within the sample.
[0092] As utilized herein with respect to numerical ranges, the terms “approximately,” “about,” “substantially,” and similar terms generally mean+ / -10% of the disclosed values, unless specifiedotherwise. As utilized herein with respect to structural features (e.g., to describe shape, size, orientation, direction, relative position, etc.), the terms “approximately,” “about,” “substantially,” and similar terms are meant to cover minor variations in structure that may result from, for example, the manufacturing or assembly process and are intended to have a broad meaning in harmony with the common and accepted usage by those of ordinary skill in the art to which the subject matter of this disclosure pertains. Accordingly, these terms should be interpreted as indicating that insubstantial or inconsequential modifications or alterations of the subject matter described and claimed are considered to be within the scope of the disclosure as recited in the appended claims.
[0093] It should be noted that the term “exemplary” and variations thereof, as used herein to describe various embodiments, are intended to indicate that such embodiments are possible examples, representations, or illustrations of possible embodiments (and such terms are not intended to connote that such embodiments are necessarily extraordinary or superlative examples).
[0094] The term “coupled” and variations thereof, as used herein, means the joining of two members directly or indirectly to one another. Such joining may be stationary (e.g., permanent or fixed) or moveable (e.g., removable or releasable). Such joining may be achieved with the two members coupled directly to each other, with the two members coupled to each other using a separate intervening member and any additional intermediate members coupled with one another, or with the two members coupled to each other using an intervening member that is integrally formed as a single unitary body with one of the two members. If “coupled” or variations thereof are modified by an additional term (e.g., directly coupled), the generic definition of “coupled” provided above is modified by the plain language meaning of the additional term (e.g., “directly coupled” means the joining of two members without any separate intervening member), resulting in a narrower definition than the generic definition of “coupled” provided above. Such coupling may be mechanical, electrical, or fluidic.
[0095] References herein to the positions of elements (e.g., “top,” “bottom,” “above,” “below”) are merely used to describe the orientation of various elements in the figures. It should be noted that the orientation of various elements may differ according to other exemplary embodiments, and that such variations are intended to be encompassed by the present disclosure.
[0096] The present disclosure contemplates methods, systems, and program products on any machine-readable media for accomplishing various operations. The embodiments of the presentdisclosure may be implemented using existing computer processors, or by a special purpose computer processor for an appropriate system, incorporated for this or another purpose, or by a hardwired system. Embodiments within the scope of the present disclosure include program products comprising machine-readable media for carrying or having machine-executable instructions or data structures stored thereon. Such machine-readable media can be any available media that can be accessed by a general purpose or special purpose computer or other machine with a processor. By way of example, such machine-readable media can comprise RAM, ROM, EPROM, EEPROM, or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to carry or store desired program code in the form of machine-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer or other machine with a processor. Combinations of the above are also included within the scope of machine-readable media. Machine-executable instructions include, for example, instructions and data which cause a general-purpose computer, special purpose computer, or special purpose processing machines to perform a certain function or group of functions.
[0097] Although the figures and description may illustrate a specific order of method steps, the order of such steps may differ from what is depicted and described, unless specified differently above. Also, two or more steps may be performed concurrently or with partial concurrence, unless specified differently above. Such variation may depend, for example, on the software and hardware systems chosen and on designer choice. All such variations are within the scope of the disclosure. Likewise, software implementations of the described methods could be accomplished with standard programming techniques with rule-based logic and other logic to accomplish the various connection steps, processing steps, comparison steps, and decision steps.
[0098] The term “client or “server” include all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations, of the foregoing. The apparatus may include special purpose logic circuitry, e.g., a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC). The apparatus may also include, in addition to hardware, code that creates an execution environment for the computer program in question (e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a crossplatform runtime environment, a virtual machine, or a combination of one or more of them). Theapparatus and execution environment may realize various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures.
[0099] The systems and methods of the present disclosure may be completed by any computer program. A computer program (also known as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program may be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0100] The processes and logic flows described in this specification may be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows may also be performed by, and apparatus may also be implemented as, special purpose logic circuitry (e.g., an FPGA or an ASIC).
[0101] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random-access memory or both. The essential elements of a computer are a processor for performing actions in accordance with instractions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data (e.g., magnetic, magneto-optical disks, or optical disks). However, a computer need not have such devices. Moreover, a computer may be embedded in another device (e.g., a vehicle, a Global Positioning System (GPS) receiver, etc.). Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-opticaldisks; and CD ROM and DVD-ROM disks). The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0102] To provide for interaction with a user, implementations of the subject matter described in this specification may be implemented on a computer having a display device (e.g. , a CRT (cathode ray tube), LCD (liquid crystal display), OLED (organic light emitting diode), TFT (thin-film transistor), or other flexible configuration, or any other monitor for displaying information to the user. Other kinds of devices may be used to provide for interaction with a user as well; for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback).
[0103] Implementations of the subject matter described in this disclosure may be implemented in a computing system that includes a back-end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer) having a graphical user interface or a web browser through which a user may interact with an implementation of the subject matter described in this disclosure, or any combination of one or more such back end, middleware, or front end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a LAN and a WAN, an inter-network (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
Claims
WHAT IS CLAIMED IS:
1. An optical system for high-throughput chemical analysis comprising: a gantry; a scanning head mounted to the gantry, wherein the scanning head comprises an optical imaging unit and a spectroscopy unit; and a controller comprising a processor and a memory operably coupled to the processor, the memory having computer-executable instructions stored thereon, that when executed by the processor, cause the processor to: receive, from the optical imaging unit, an image of a biological specimen comprising a plurality of droplets; analyze, using a computer vision module, the image to identify a location of at least one droplet of the plurality of droplets within the biological specimen; control the scanning head to position the spectroscopy unit in proximity to the location of the at least one droplet; analyze, using the computer vision module, a magnified portion of the image including the at least one droplet to identify a region of interest within the at least one droplet; and control the spectroscopy unit to acquire a spectrum associated with the region of interest within the at least one droplet.
2. The optical system of claim 1 , wherein the region of interest comprises a feature likely to have a high-value spectrum.
3. The optical system of claim 1, wherein analyzing, using the computer vision module, the image to identify the location of the at least one droplet of the plurality of droplets within the biological specimen comprises using a first trained machine learning model to detect the at least one droplet in the image.
4. The optical system of claim 3, wherein the first trained machine learning model is a deep learning model.
5. The optical system of claim 1 , wherein analyzing, using the computer vision module, the image to identify the location of the at least one droplet of the plurality of droplets within the biological specimen comprises using a first ensemble of machine learning models to detect the at least one droplet in the image.
6. The optical system of claim 1 , wherein analyzing, using the computer vision module, the image to identify the location of the at least one droplet of the plurality of droplets within the biological specimen comprises using an edge detection algorithm to detect a boundary of the at least one droplet in the image.
7. The optical system of claim 1 , wherein analyzing, using the computer vision module, the magnified portion of the image including the at least one droplet to identify the region of interest within the at least one droplet comprises using a second trained machine learning model to detect the region of interest in the magnified portion of the image.
8. The optical system of claim 7, wherein the second trained machine learning model is a deep learning model.
9. The optical system of claim 1 , wherein analyzing, using the computer vision module, the magnified portion of the image including the at least one droplet to identify the region of interest within the at least one droplet comprises using a second ensemble of machine learning models to detect the region of interest in the magnified portion of the image.
10. The optical system of claim 1 , wherein the memory has further computer-executable instructions stored thereon, that when executed by the processor, cause the processor to control the spectroscopy unit to acquire a low-resolution spectra associated at least a portion of the at least one droplet, and analyze the low-resolution spectra to identify a region of interest within the at least one droplet.
11. The optical system of claim 1 , wherein the memory has further computer-executable instructions stored thereon, that when executed by the processor, cause the processor to predict, using a third trained machine learning model, a disease associated with the biological specimen based on the spectrum associated with the region of interest.
12. The optical system of claim 1 , wherein the memory has further computer-executable instructions stored thereon, that when executed by the processor, cause the processor to predict, using a third ensemble of machine learning models, a disease associated with the biological specimen based on the spectrum associated with the region of interest.
13. The optical system of claim 1, wherein the spectroscopy unit is configured to deliver and collect light.
14. The optical system of claim 13, further comprising a spectroscopy light source and a detector.
15. The optical system of claim 14, wherein the spectroscopy unit comprises at least one of the spectroscopy light source and the detector.
16. The optical system of claim 15, wherein the optical imaging unit comprises a widefield imaging device.
17. A computer- implemented method for controlling an optical system for high-throughput chemical analysis comprising: receiving, from an optical imaging unit, an image of a biological specimen comprising a plurality of droplets; analyzing the image to identify a location of at least one droplet of the plurality of droplets within the biological specimen; controlling a scanning head comprising a spectroscopy unit to position the spectroscopy unit in proximity to the location of the at least one droplet; analyzing a magnified portion of the image including the at least one droplet to identify a region of interest within the at least one droplet; and controlling the spectroscopy unit to acquire a spectrum associated with the region of interest within the at least one droplet.
18. The computer-implemented method of claim 17, wherein the region of interest comprises a feature likely to have a high-value spectrum.
19. The computer-implemented method of claim 18, wherein analyzing the image to identify the location of the at least one droplet of the plurality of droplets within the biological specimen comprises using a first trained machine learning model to detect the at least one droplet in the image.
20. A method comprising: receiving a plurality of images, each of the plurality of images capturing a respective biological specimen comprising a plurality of droplets; receiving a plurality of respective spectrum associated with a plurality of respective regions of interest, wherein each respective region of interest is associated with one of the plurality of droplets within one of the plurality of images; identifying a set of the plurality of respective regions of interest having a high-value spectrum;creating a labeled dataset comprising a set of the plurality of images including the set of the plurality of respective regions of interest having the high-value spectrum; and training a model using the labeled dataset, wherein the trained model is configured to detect a region of interest having the high-value spectrum.
Citation Information
Patent Citations
Methods for identifying viral infections and for analyzing exosomes in liquid samples by raman spectroscopy
US20230015302A1
Systems and methods for disease diagnosis
US20230018494A1