Adaptive raman spectra sampling with swept-source raman

WO2026206318A1PCT designated stage Publication Date: 2026-10-01MASSACHUSETTS INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/021552
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2026-10-01

Smart Images

  • Figure US2025021552_01102026_PF_FP_ABST
    Figure US2025021552_01102026_PF_FP_ABST
Patent Text Reader

Abstract

System and methods for minimizing spectral sampling with swept source Raman (SSR) spectroscopy are disclosed. The systems and methods disclosed herein may improve data acquisition with SSR spectrometers and / or reduce spectral sampling time with SSR spectrometers. The method may include VIP scores and regression coefficients calculated by a regression model, including Partial Least Squares Regression (PLSR), to minimize sampling time of a SSR spectrometer.
Need to check novelty before this filing date? Find Prior Art

Description

PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01ADAPTIVE RAMAN SPECTRA SAMPLING WITH SWEPT-SOURCE RAMANGOVERNMENT SUPPORT

[0001] This invention was made with government support under U01 FD006751 awarded by Food and Drug Administration. The government has certain rights in the invention.BACKGROUND

[0002] Raman spectroscopy is a technique used to gain information about the chemical composition of a material and the state of that material. In this technique, the sample is illuminated with a bright light source and the wavelength distribution of the scattered light is measured. ‘Fingerprint’ chemical spectra are acquired when a laser excites different molecular vibrations. These molecular vibrations extract some energy from the excitation light resulting in scattered light with a longer wavelength. While information rich, this scattered light, which is the Raman signal, is very weak as only one out of about a billion laser photons excite molecular vibrations.SUMMARY

[0003] A conventional Raman system uses a spectrometer to determine the distribution of the scattered light. Because these spectrometers typically disperse colors of light in different directions and rely on free space propagation for their spectral separation and detection, they exhibit tradeoffs between spectral resolution, sensitivity, and device size. Raman spectrometers that can provide the spectral resolution and sensitivity desired for monitoring continuous flow manufacturing processes tend to be bulky, power hungry, and expensive.

[0004] Conventionally, there are two primary modes for Raman spectroscopic analysis -dispersive Raman using a grating-based spectrometer or Fourier Transform Raman (FT-Raman) spectroscopy that uses a Michelson interferometer. A third approach to Raman spectroscopy is Swept Source Raman (SSR) spectroscopy (SSRS), which utilizes a tunable laser and a low-cost, fixed- wavelength detector instead of a spectrometer or interferometer. SSRS offers the throughput (light-gathering) advantage of FT-Raman spectroscopy without bulky optics or moving mirrors. SSRS is readily scaled to many measurement ports, making it more suitable for monitoring continuous flow manufacturing processes, in part because it usesPCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01small, inexpensive fixed-wavelength detector(s) instead of spectrometers. SSR sensor networks may also present an advantage in regards to number and range of Raman sensors that can be deployed and operated.

[0005] Despite these advantages, SSRS systems have other drawbacks. One drawback is sampling time since spectral acquisition with SSRS systems is done sequentially. Raman spectra acquisition often requires both a considerable integration time (e.g., about 1-30 seconds) and multiple (e.g., 10-50) acquisition repetitions to enhance the signal -to-noise ratio (SNR). The throughput enhancement of SSRS systems allow the integration time to be shortened (or alternatively the excitation power to be reduced) but total acquisition time is still considerably longer than with dispersive Raman because sampling is repeated for each spectral datapoint to enhance the SNR. For example, dispersive benchtop systems have anywhere between 1300-3600 pixels (resolvable spectral bins) and acquiring spectra in the same spectral density with a SSRS system would take days to complete for each spectrum. This challenge is exacerbated when considering the use of a single laser which is time-shared between multiple sensors.

[0006] Disclosed herein are systems and methods to reduce total spectral acquisition time for SSRS systems.

[0007] In some aspects, the techniques described herein relate to a method for reducing spectral acquisition time for a sample using a swept source Raman (SSR) spectrometer, the method including acquiring a full Raman spectrum of the sample using the SSR spectrometer, wherein the sample includes a plurality of analytes, identifying at least one region of interest in the Raman spectrum for each of the plurality of analytes, training a model on the at least one region of interest for each of the plurality of analytes using the full Raman spectrum of the sample, selecting a reduced number of spectral datapoints less than a number of spectral datapoint in the full Raman spectrum, and acquiring a reduced Raman spectrum of the sample with the reduced number of spectral datapoints using the SSR spectrometer based on the model, wherein the sample includes the plurality of analytes.

[0008] In some aspects, the techniques described herein relate to a method further including determining a concentration of at least one analyte of the plurality of analytes in the sample based on the reduced Raman spectrum.

[0009] In some aspects, the techniques described herein relate to a method wherein the sample is a cell culture and determining the concentration of the at least one analyte of the plurality ofPCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01analytes in the sample includes determining the concentration of the at least one analyte of the plurality of analytes in the sample at multiple points of a manufacturing process.

[0010] In some aspects, the techniques described herein relate to a method wherein the reduced number of spectral datapoints is less than 100.

[0011] In some aspects, the techniques described herein relate to a method wherein the reduced number of spectral datapoints is less than 50.

[0012] In some aspects, the techniques described herein relate to a method wherein acquiring the reduced Raman spectrum includes acquiring highest ranking spectral datapoints for the reduced number of spectral datapoints.

[0013] In some aspects, the techniques described herein relate to a method wherein the model includes a partial least squares regression (PLSR) model.

[0014] In some aspects, the techniques described herein relate to a method wherein the model reduces an acquisition time of the reduced Raman spectrum by a factor of at least five.

[0015] In some aspects, the techniques described herein relate to a method wherein the model reduces the acquisition time of the reduced Raman spectrum by a factor of at least seven.

[0016] In some aspects, the techniques described herein relate to a method wherein training the model further includes training the model on Raman spectra of each analyte of the plurality of analytes of the sample.

[0017] In some aspects, the techniques described herein relate to a method wherein the selecting the number of spectral datapoints includes selecting the number of spectral datapoints based on the Raman spectra of at least one of the plurality of analytes.

[0018] In some aspects, the techniques described herein relate to a system for measuring Raman signals of a sample, the system including a tunable laser to emit a tunable excitation beam, an optical fiber to transmit the tunable excitation beam to a measurement site, an optical probe, in optical communication with the optical fiber, to illuminate the measurement site with the tunable excitation beam and to collect the Raman signal emitted from the measurement site in response to the tunable excitation beam, a spectrally selective detector, in optical communication with the optical probe, to sense the Raman signal from the measurement site, and a processor, operably coupled to the optical probe, configured to run a model, wherein the model is configured to identify at least one region of interest in the Raman signal of the sample and to cause the tunable laser to emit the tunable excitation beam to acquire a reduced RamanPCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01signal of the least one region of interest in the Raman signal of the sample using a selected number of spectral datapoints.

[0019] In some aspects, the techniques described herein relate to a system wherein the model includes a partial least squares regression (PLSR) model.

[0020] In some aspects, the techniques described herein relate to a system wherein the sample includes a cell culture including a plurality of analytes and the model is further configured to determine a concentration of at least one of the plurality of analytes using the reduced Raman signal.

[0021] In some aspects, the techniques described herein relate to a method for swept-source Raman spectroscopy, the method including acquiring swept-source Raman spectra of a sample over a predetermined band of excitation wavelengths, training a machine-learning model on the swept-source Raman spectra, identifying, with the machine-learning model, a subset of excitation wavelengths corresponding to a region of interest in the swept-source Raman spectra, and acquiring a swept-source Raman spectrum of the sample for only the subset of excitation wavelengths.

[0022] In some aspects, the techniques described herein relate to a method further including determining a concentration of at least component of the sample based on the swept-source Raman spectrum of the sample for the subset of excitation wavelengths.

[0023] In some aspects, the techniques described herein relate to a method wherein acquiring the swept-source Raman spectrum of the sample for only the subset of excitation wavelengths has a lower acquisition time than acquiring the swept-source Raman spectra of the sample over the predetermined band of excitation wavelengths.

[0024] In some aspects, the techniques described herein relate to a method wherein identifying, with the machine-learning model, the subset of excitation wavelengths includes identifying an importance score for the region of interest based on a signal intensity the swept-source Raman spectra.

[0025] In some aspects, the techniques described herein relate to a method wherein acquiring the swept-source Raman spectrum of the sample further includes acquiring the swept-source Raman spectrum for a selected number of spectral datapoints for the subset of excitation wavelengths.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0026] All combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are part of the inventive subject matter disclosed herein. The terminology used herein that also may appear in any disclosure incorporated by reference should be accorded a meaning most consistent with the particular concepts disclosed herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The skilled artisan will understand that the drawings primarily are for illustrative purposes and are not intended to limit the scope of the inventive subject matter described herein. The drawings are not necessarily to scale; in some instances, various aspects of the inventive subject matter disclosed herein may be shown exaggerated or enlarged in the drawings to facilitate an understanding of different features. In the drawings, like reference characters generally refer to like features (e.g., functionally similar and / or structurally similar elements).

[0028] FIG. 1 is an illustration of a dispersive Raman system.

[0029] FIG. 2A is a block diagram of a swept source Raman (SSR) spectroscopy (SSRS) system with a tunable laser, Talbot wavemeter, band-edge filter, and narrowband detector.

[0030] FIG. 2B shows an SSRS sensor network for multi-point monitoring of a continuous manufacturing process.

[0031] FIG. 3 is a diagram illustrating the operation of a Swept Source Raman Spectrometer (SSRS) compared to a dispersive Raman system.

[0032] FIG. 4A is a diagram of a bioreactor with Chinese Hamster Ovary (CHO) cells showing the main components of the upstream section of the CHO-cell continuous bioreactor which produced samples of spent media and supernatant for measurement with three different Raman systems.

[0033] FIG. 4B is chronological description of CHO perfusion run R2 in the bioreactor shown in FIG. 4A, illustrating the complexity of the process and steps for maintaining the culture.

[0034] FIG. 4C is a chronological description of CHO perfusion run R3 in the bioreactor shown in FIG. 4B.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0035] FIG. 5 shows Raman spectra obtained with a Kaiser spectrometer (Kaiser spectra for short) of three CHO runs (Rl, R2, R3), focusing on the fingerprint region with an inlay of the entire spectral range (100-3425 cm '). These spectra were acquired using 400 mW and 10-second integrations repeated 75 times each.

[0036] FIG. 6A shows simulation results in which the Kaiser spectra were sampled in increasingly larger intervals (decimation) and the effect of this decimation on the Root Mean Squared Error of Estimation (RMSEE) of the Partial Least Squares Regression (PLSR) model for all 7 analytes (normalized to no decimation data results).

[0037] FIG. 6B shows simulation results in which the Kaiser spectra were sampled in increasingly larger intervals (decimation) and the effect of this decimation on the Root Mean Square Error of Prediction (RMSEP) of the PLSR model for all 7 analytes (normalized to no decimation data results).

[0038] FIG. 7A shows a comparison of 14 spectra from days in run R3 for the Biomod (top), Kaiser (middle) and SSRS (bottom). The Biomod (150 mW, 10-second integrations repeated 50 times each) and SSRS spectra (5-6 mW, 5-second integrations repeated 12 times each for 276 data points) are of supernatant samples while the Kaiser spectra is of the in-situ culture (400 mW, 10-second integrations repeated 75 times each). All spectra were smoothed but no additional processing was performed.

[0039] FIG. 7B shows an SSRS spectrum of one of the CHO supernatant before (raw data) and after (smoothed data) smoothing with a Savitzky-Golay filter order 2, 3rddegree.

[0040] FIG. 7C shows a normalized spectra comparison of samples from day 8 of the R3 run, with the Kaiser in-situ (middle spectrum), Biomod supernatant (top spectrum) and SSRS supernatant (bottom spectrum) with acquisition parameters as mentioned before.

[0041] FIG. 7D shows a normalized spectra of samples from Day 8 of the CHO R3 run from the SSRS supernatant (top spectrum), the Kaiser in-situ (bottom spectrum), Biomod supernatant (middle spectrum) (as in FIG. 7C) after background removal using an empty sample holder subtraction and a 6thorder polynomial Lieber fit.

[0042] FIG. 7E shows a comparison of 14 samples from different days in CHO run R3 for the Biomod (top), Kaiser (middle) and SSRS (bottom) after background removing techniques and cross-section correction have been applied (see FIGS. 7A-7D).PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0043] FIG. 8 shows SSRS spectra of 50mg / ml glucose (top), R3 run supernatant from Days 5,6 and 8 (middle), and lOmg / mL of NIST mAb (bottom). All spectra were acquired with 506mW, 5 seconds integration and 12 repetitions with 276 wavelengths.

[0044] FIG. 9A shows SSRS direct peak analysis (DPA) of the main glucose peak (at 1125 cm ') with the SSRS signal values (circles) presented on the left Y axis and the NOVA metabolite analyzer (diamonds) on the right Y axis for 12 samples of the CHO R3 run, showing good correlation of values.

[0045] FIG. 9B shows SSRS DPA for 12 samples of the CHO R3 run of mAb peak (1250 cm ') with the SSRS signal values (circles) presented on the left Y axis and the Octet mAb titer (diamonds) on the right Y axis

[0046] FIG. 9C shows an SSRS DPA for 12 samples of the CHO R3 run of lactate Raman peak (850 cm ') with SSRS signal values (circles) on the left and Nova metabolite analyzer on the right (diamonds).

[0047] FIG. 9D shows Glucose SSRS probe 3 c limit of detection (LOD) results for two tunable sources (Superlum with 6 mW (left)), Ti:Sapph with 40 mW (center)) and a dispersive benchtop Raman (right) using 150 mW at 830 nm excitation with a 7-bundle fiber collection. All spectra were acquired with 10 seconds integration repeated 50 times.

[0048] FIG. 10 shows example spectra illustrating sample degradation with extended acquisition time. Spectra of the Day 8 sample from run R3 acquired 4 consecutive times with each acquisition lasting 5 hours, showing sample degradation and varying fluorescent levels indicating denaturation.

[0049] FIG. 11A shows PLSR regression coefficients (P) computed from Kaiser Raman spectra of the CHO cell culture, estimating the glucose concentration.

[0050] FIG. 1 IB shows the VIP scores computed for each of the spectral data points, indicating their significance on the overall explained variance of the output.

[0051] FIG. 12A shows lactate regression coefficients with a dashed line at regression coefficient of 0.

[0052] FIG. 12B shows regression VIP scores showing the highly correlated spectral points with the prediction of lactate concentrations. The VIP scores may also include regions which are negatively correlated with the regression coefficients.

[0053] FIG. 13 illustrates selective VIP-informed spectral sampling and validation with SSRS.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0054] FIG. 14 shows PLSR VIP scores for all 7 analytes in the full spectral range between 810 cm1and 1670 cm1where thicker lines correspond to VIP scores greater than 1.

[0055] FIG. 15A shows a comparison of two ranking spectral data point significance using the sum of all analyte VIP scores (top spectrum) and the sum of products of VIP and regression coefficients (bottom spectrum).

[0056] FIG. 15B shows Raman spectra with the 50 spectral points selected in both the VIP method (upper spectrum) and the VIPxBeta method (lower spectrum).

[0057] FIG. 16 shows the percentage of variance explained in Kaiser spectra for glucose concentration as a function of number of partial least squares (PLS) components for both a fullrange spectra (bottom) and only 50 data points (top) showing the number of components changes for the partial spectra.

[0058] FIG. 17A shows Kaiser spectra PLSR model results showing the estimated and predicted analyte concentration results for glucose. The shading indicates data from the three different CHO perfusion runs that were monitored using the Kaiser system.

[0059] FIG. 17B shows Kaiser spectra PLSR model results showing the estimated and predicted analyte concentration results for lactate. The shading indicates data from the three different CHO perfusion runs that were monitored using the Kaiser system.

[0060] FIG. 17C shows Kaiser spectra PLSR model results showing the estimated and predicted analyte concentration results for total CD. The shading indicates data from the three different CHO perfusion runs that were monitored using the Kaiser system.

[0061] FIG. 17D shows Kaiser spectra PLSR model results showing the estimated and predicted analyte concentration results for mAb with only 40 datapoints ranked in the VIPxBeta method. The shading indicates data from the three different CHO perfusion runs that were monitored using the Kaiser system.

[0062] FIG. 18 shows SSRS PLSR results with Kp = 40 points for 12 CHO supernatant samples from Run R3, reflecting the model’s RMSEE estimation errors of the data. The empty circles represent the Nova analyzer values and the stars represent the PLSR estimation results.

[0063] FIG. 19 shows a method of reducing spectral acquisition time for a sample using a swept source Raman (SSR) spectrometer.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01DETAILED DESCRIPTIONRaman Spectroscopy Systems

[0064] FIG. 1 shows an example of a dispersive Raman spectroscopy system 100. The Raman spectroscopy system 100 includes a fixed laser 105 and a spectrometer 106. The spectrometer 106 includes an input slit 107, concave mirrors 108a and 108b, a grating 109, and a detector 114. In operation, the fixed laser 105 illuminates a sample with a Raman pump beam. This causes the sample to emit light, which enters the spectrometer 106 through the input slit 107 and then is collimated using the first concave mirror 108a to hit the grating 109. The second mirror 108b refocuses the light diffracted by the grating 109 onto the detector 114, which detects the diffracted light as the Raman spectrum of the sample.

[0065] FIG. 2 A shows an example of a swept source Raman (SSR) spectroscopy (SSRS) system 200. The SSR spectroscopy system 200 includes a tunable laser 210, high-throughput collection optics 230, a band-edge filter 250, and a bandpass filter 260 in front of a large-area detector 214, such as a photon-counting amplified detector, CCD, or CMOS array. The system 200 may also include a beam splitter 216 that couples the tunable laser 210 to a wavelength sensor, such as a spectrometer or, here, a Talbot wavemeter 220. The tunable laser 210 can be implemented in a photonic integrated circuit (PIC) along with the wavemeter’s passive components, the beam splitter 216, the band-edge filter 250, and electronics for calibrating the wavelength and amplitude of the tunable laser’s output. The collection optics 230 can include off-the-shelf components, such as bulk lenses, and / or custom nano-structured devices, such as photonic crystals and meta-surfaces, with enhanced collection efficiency.

[0066] In operation, the tunable laser 210 emits a swept or chirped pump or excitation beam 211, which may be discontinuous (e.g., due to mode hops) and / or nonlinear. The tunable laser 210 may sweep the pump beam’s wavelength over a tuning range of 850 nm to 1200 nm during a period of 1 second to 100 seconds (e.g., about 10, 20, 30, 40, or 50 seconds). Shorter, faster wavelength tuning is also possible; for instance, the tunable laser 210 may sweep the pump beam 211 over a range of 35-50 nm in milliseconds to seconds.

[0067] The pump beam 211 propagates through the beam splitter 216 and the band-edge filter 250, which rejects amplified spontaneous emission (ASE) light from the tunable laser 210. The collection optics 230 focus the filtered pump beam 211 onto a sample 21, which responds to the pump beam 211 by emitting Raman light 213 isotropically. The collection optics 230 collect and collimate a portion of the Raman light 213. This collimated Raman light reflects off thePCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01band-edge filter 250 and propagates through the bandpass filter 260, which reflects and / or attenuates light at the wavelength of the pump beam 211. The passband of the bandpass filter 260 should overlap with the stopband of the band-edge filter 250, which should have a passband that overlaps with the tuning range of the tunable laser 210.

[0068] The photodetector 214 senses the portion of the Raman light 213 transmitted by the bandpass filter 250. The photodetector 214 has a large area (e.g., an area of at least 50 pm x 50 pm) and intrinsic gain, it can count photons from the ultraviolet (UV) to the infrared (IR) regions of the electromagnetic spectrum to improve the system’s sensitivity. Generally, the photodetector 214 and bandpass filter 260 should allow for detection of a Raman shift of 200-4000 cm1(e.g., 200-1800 cm1or 3000-3500 cm ').

[0069] The Talbot wavemeter 220 measures the (absolute) wavelength and power of the pump beam 211 while the photodetector 214 detects the filtered Raman light 213. A processor 290 coupled to the Talbot wavemeter 220 and the photodetector 214 determines the sample’s Raman spectrum from the detected Raman signal and the wavelength and power measurements made by the Talbot wavemeter 220. In short, the processor maps each Raman signal measurement made by the photodetector 214 to the corresponding wavelength and power measurement made by the Talbot wavemeter 220. This accounts for wavelength discontinuities and other nonlinearities, if any, in the wavelength sweep of the tunable laser 210.

[0070] The SSR spectrometer 200 can be built with optical filters other than a narrowband bandpass filter 260. The narrowband bandpass filter 260 simplifies the estimation of Raman spectrum because the tunable SSR spectrometer 200 collects Raman signal at only one color or wavelength channel at a time. It is also possible to use a filter with another type of spectrally selective filer, such as a long-pass filter, a band-edge filter, or a multi-peak filter. Any spectrally selective filter should work, so long as the SSR spectrometer’ s transfer function can be inverted to estimate the Raman spectrum from the time series data representing the integrated optical power passed through the filter. (The time series data is collected as the laser is tuned across its tuning range.)

[0071] For example, a long-pass filter can be used instead of the bandpass filter 260. In this case, the SSR spectrometer detects the sum (or integral) of all colors beyond the cutoff wavelength of the filter. As the tunable laser’ s output 211 is tuned towards shorter wavelengths, less of the Raman light 213 passes through the filter, causing the integrated detected power toPCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01decrease. By taking the derivative of the time-series data collect by the SSR spectrometer, the processor 290 can recover the spectrum of the Raman light 213.

[0072] FIG. 2B shows an example of a SSR spectroscopy system 200a configured with a plurality of fiber probes 140. The SSR spectroscopy system 200a may allow for measurements in challenging conditions such as in-situ inside a bioreactor or integrated into a catheter or endoscope for in-vivo measurements. The SSRS system 200a used for multi-point monitoring of a continuous manufacturing line 10. This SSRS system 200a includes a Raman pump source, which may include one or more lasers 210, and monitors six remote measurement locations within the continuous manufacturing process via a 1 x N optical switch 120 (here, N = 6) and excitation fibers 130a-130f (collectively, excitation fibers 130). The lasers 210 may be tunable lasers with overlapping, contiguous, or non-contiguous tuning ranges for probing different Raman bands. The Raman pump may also include lasers that emit at different discrete wavelengths for target Raman measurements.

[0073] The excitation fibers 130 are coupled to respective probes 140a-140f, which are located at different points or sites along the continuous manufacturing line 10. In this example, the manufacturing line 10 includes a media tank 11, perfusion cell culture tank 12, harvest tank 13, first chromatography site 14, second chromatography site 15, third chromatography site 16, filtration system 17, and drug substance tank 18. The first probe 140a monitors the perfusion cell culture 12; the second probe 140b monitors the harvest 13; the third, fourth, and fifth probes 140c-140e monitor the outputs of the chromatography sites 14-16; and the sixth probe 140f monitors the output of the filtration system 17.

[0074] Each remote measurement location has an associated spectrally selective filter 142a-142f, detector 250a-250f, and analog-to-digital converter 152a-152f. These components can either be located near the tunable laser 210 or at the corresponding remote measurement location as shown in FIG. 2A. There may be one detector 150 per measurement site as in FIG.2A, either collocated with the corresponding probe or remotely coupled to the measurement site with an optical fiber. Alternatively, the system 100 may include a single filter, detector, and analog-to-digital converter that monitor all of the measurement sites or a subset of measurement sites, with a fiber network and fiber-coupled switch that switches among measurement sites in tandem with the switch 120 coupled to the Raman pump source.

[0075] Each detector 150 can be a photodiode, charge-coupled device (CCD) array, complementary metal-oxide-semiconductor (CMOS) array, photomultiplier tube (PMT),PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01single-photon avalanche diode (SPAD), avalanche photodiode (APD), or any other device that can detect photons in the desired wavelength band, which may be defined by a narrowband filter. The spectrally selective detectors 150 can either be located together or separately along with any filters and / or associated electronic components. In either case, the corresponding optical fiber 130 may collect Raman emissions or signals from a given measurement site and couple them to the corresponding spectrally selective detector 150.

[0076] The exact timing of the pump beam wavelength tuning and switching may depend on the process being monitored. For example, the laser 210 and switch 120 may operate in a roundrobin fashion, switching among measurement sites in a sequential fashion, with the laser wavelength selected based on the analyte(s) expected at each measurement site. The laser 210 and switch 120 may also scan and switch Raman pump beam in other sequences, e.g., monitoring one or more sites more frequently than other sites. The laser 210 can be tuned continuously (e.g., in repeated chirps), switched among discrete wavelengths, or tuned nonlinearly, depending on the measurement sequence and the analytes / Raman transitions to be probed at each site.

[0077] The SSR spectroscopy system 200a can support multiple measurement sites or Raman probes (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more measurement sites). Because SSR spectroscopy system 200a is not constrained by the collection power of the spectrometer (namely, by the spectrometer’ s entrance slit and numerical aperture), it can use collection optical fibers with enhanced light gathering capabilities.

[0078] The spectroscopy system 200a may be coupled to a computing device 290a through a wired and / or wireless (e.g., Bluetooth, WiFi, etc.) connection. The computing device 290a may be a local computing device (e.g., a workstation, a desktop, a laptop, a tablet, a smartphone, etc.) configured to process data acquired by the probes 140a-140f. In some embodiments, the computing device 290a may provide computational resources (e.g., memory, processor cycles, etc.). For instance, the computing device 290a may include a workstation having one or more special-purpose processors, such as a graphics processing unit (GPU) and / or a neural processing unit (NPU). The computing device 290a may also be operably connected to a network, e.g., an a wide-area network, such as the Internet, a local-area network, a metropolitan-area network, or another type of electronic communication network. The network may include wired and / or wireless data links. A variety of communications protocols may be used in the network including, but not limited to, Wi-Fi, Ethernet, Transport Control ProtocolPCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01(TCP), Internet Protocol (IP), Hypertext Transfer Protocol (HTTP), SOAP, remote procedure call protocols, and / or other types of communications protocols.

[0079] The computing device 290a may include at least one processor or processing unit and a system memory. The processor is a device configured to process a set of instructions. The system memory may be a component of processor or separate from the processor. Depending on the exact configuration and type of computing device, the system memory may be volatile (such as Random Access Memory (RAM)), non-volatile (such as Read-Only Memory (ROM), flash memory, etc.) or some combination of the two. System memory typically includes an operating system suitable for controlling the operation of the computing device 290a. The system memory may also include one or more software applications and may include program data (e.g., program data for the spectroscopy system 200a). The computing device 290a may be configured to store in memory instructions for implementing the various operations, methods and functions disclosed herein. When the instructions are executed by the computing device 290a, the instructions cause the computing device 290a to perform one or more of the operations or methods disclosed herein.

[0080] The computing device 290a may further be configured to communicate with one or more other computing devices (not shown) via a computer network (e.g., the internet). The computing device 290a and / or these other computing devices may include one or more servers, which may be configured to perform data analysis for the spectroscopy system 200a. For example, data obtained by the spectroscopy system 200a may be transferred to an external computing device via the computing device 290a for further processing.

[0081] FIG. 3 shows a comparison of the operation of SSR vs. dispersive Raman. With a dispersive Raman spectrometer (e.g., the Raman spectrometer shown in FIG. 1), spectra are obtained at several fixed excitation wavelengths from 750 nm to 830 nm. With an SSR spectrometer (e.g., Raman spectrometers shown in FIGS. 2A and / or 2B), a different segment of the Raman spectrum is detected as the laser excitation tunes and the scattered photons are detected through a narrow optical filter with a fixed wavelength. In other words, tuning of the excitation with the SSR spectrometer the complete spectrum is acquired (shown by the overlay SSR spectrum in FIG. 3).

[0082] Further examples of SSRS, applications of SSRS, and SSRS probes may be found in U.S. Patent No. 10,656,012, entitled “Swept-Source Raman Spectroscopy Systems and Methods,” and U.S. Patent No. 11,698,301, entitled “Multiplexed Sensor Network UsingPCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01Swept Source Raman Spectroscopy,” which are hereby incorporated by reference in their entireties for all purposes.Reducing Spectral Acquisition Time using Minimal Spectral Sampling

[0083] The systems and methods disclosed herein may improve data acquisition with SSR spectrometers and / or reduce spectral acquisition time for SSR spectrometers. Herein, samples from a CHO perfusion culture were measured using an SSR spectrometer (e.g., spectrometer 200 or 200a) and also two dispersive Raman systems (e.g., spectrometer 100). The SSR spectrometer achieved comparable spectra to the dispersive systems and enabled monitoring of multiple metabolites, including glucose and lactate, and also cell density and mAb titer using direct Raman peak analysis.

[0084] Data was acquired from a CHO-cell continuous manufacturing testbed and the obtained spectra were used to minimize total spectral acquisition time. Three different methods to reduce total spectral acquisition time are disclosed herein: Down-sampling, Raman peak sampling, and VIP -informed sampling. These methods were evaluated through a simulation using dispersive spectra and SSRS to validate them. These methods may be used to train one or more models that may be used to acquire a reduced Raman spectrum of a sample. The trained model is then used to evaluate analyte quantities from the smaller number of acquired spectral data points.

[0085] The results disclosed herein show that VIP-informed sampling may reduce acquisition time by at least a factor of 5, for example, a factor of 7, with minimal or no increase to estimation errors. Informed sampling using a model trained on VIP and regression coefficients can support the use of a SSR spectrometer as a utility for obtaining accurate information (e.g., concentration, component, quantity) about a sample and its components with lower sampling time — for example, when the SSRS system is integrated into a well-defined process, e.g. a chemical manufacturing facility or pharmaceutical production where the process monitored has a set of analytes to indicate the process success.

[0086] The systems and methods disclosed herein may utilize VIP scores and regression coefficients calculated by PLSR to reduce the sampling time of a SSR spectrometer. Initially, a full spectrum of a sample may be acquired by a SSR spectrometer. The sample may include one or more analytes (e.g., a protein, a sugar, a metabolite, a drug, etc.). The full spectrum of the sample may be used to train a model (e.g., a machine learning model and / or a PLSR model). The model may identify one or more portions of the full spectrum that are significant usingPCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01VIP scores and regression coefficients calculated by PLSR. A user may then select a number of data points to be sampled when obtaining a reduced Raman spectrum of the sample. The user may also select one or more analytes of the sample for which selecting sampling should be performed. Then a reduced spectrum of the sample may be obtained with the SSR spectrometer based on the model (also referred to herein as “VIP-informed sampling”). The reduced spectrum of the sample may be obtained using the selected number of data points and thus may be collected with a reduced sampling time. The reduced spectrum of the sample may then be used to obtain information (e.g., concentration) about a component (e.g., analyte) of the sample. For example, the model may be used to accurately determine a quantity of a component (e.g., analyte) in the sample based on a reduced number of spectral data points, thereby decreasing the acquisition time of the SSRS.

[0087] In one example, the use of Raman spectra to build prediction models may be made with a well-characterized process to reduce significant deviations from the training dataset to reduce model errors. For example, the model may be trained with a compilation library of predetermined Raman spectra for known analytes (e.g., analytes from a well -characterized process). Training the model with a compilation library can improve model estimations and predictions. Training with a compilation library may replace or augment training the model with an initial full spectrum of a sample by a SSR spectrometer. For example, if all of the analytes in the sample are known and the compilation library contains spectra for all of the analytes. The resulting model trained on the compilation library may then be used to obtain information (e.g., a concentration and / or quantity) about one or more components (e.g., analytes) of the sample using an SSRS. This may improve the model’s predictive capabilities by reducing errors that may be introduced with the initial full spectrum of the sample (e.g., due to noise and / or contamination of the sample).

[0088] Instead of, or in addition to, training the model on pre-determined Raman spectra for known analytes, the model may also be calibrated using high dynamic ranges of analytes (e.g., analytes from a well -characterized process and / or analytes that may be expected to be in a sample). For example, the model may be calibrated using Raman spectra for these ranges of analytes. The model may be calibrated using a pure analyte sample to enhance modeling performance. This may further improve the model’s predictive capabilities by reducing errors that may be introduced with the initial full spectrum of the sample (e.g., due to noise and / or contamination of the sample) and / or increasing the number of analytes that may be predicted using the model and SSRS.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0089] The methods disclosed herein may also use data acquisition and predictive-model building to model non-ideal scenarios, for example, in the development of an unknown and / or uncharacterized process and / or when a malfunction occurs. Informed sampling using VIP and regression coefficients may also support monitoring of random and / or unrelated sample(s) following obtaining the initial full spectrum of the sample and training the model prior on that initial full spectrum. Informed sampling with SSRS and the resulting model may then be used to obtain information about the random and / or unrelated sample, including, but not limited to concentration, quantity, type, and / or identification of one or more components (e.g., analytes) of the sample.

[0090] In another example, the model may also be trained using dynamic modeling. For example, if the SSRS is used to obtain spectra from a continuous manufacturing line (e.g., continuous manufacturing line 10) and / or a cell culture, the model may be updated throughout the manufacturing process. As SSR spectra are obtained from the continuous manufacturing line and / or cell culture, the model may classify the spectra as good or bad. For example, a bad spectrum, may result from an unexpected event, such as foaming, changes in media composition, clogging, calibration issues, contamination etc. The model may be trained to discard the bad spectra such that these spectra do not impair the model’ s predictive capabilities.

[0091] CHO-cell Continuous Testbed

[0092] FIG. 4A shows a diagram of the main components in the upstream section of an integrated CHO-cell continuous testbed 470 (e.g., a bioreactor, such as a 3-liter Applikon bioreactor) producing Adalimumab (also commercially named Humira) a monoclonal antibody (mAb) used for treating various forms of arthritis. The cell culture was maintained in the testbed 470. An automated MAST® autosampler system 480 (Millipore Sigma, USA) drew samples daily and transported them to an automated analyzer 490 (Nova, FLEX2) measuring key metabolites including glucose, lactate, glutamine, glutamate, ammonium, as well as ions, gasses, cell densities, and pH. Online analytics were integrated into the testbed 470 including a Viable Cell Density (VCD) sensor (Aber Futura capacitance sensor), an off-gas sensor (BlueSens), a Kaiser Raman RXn2-785nm probe 440a which may be operable coupled to a Kaiser Raman system (e.g., system 100), and an SSRS probe 440, which may be operably coupled to a SSRS system (e.g., SSRS system 200 or 200a) which measured the spectra inside the bioreactor 470 using 400-500mW of power (10 second integration repeated 75 times). When mAb was produced it was measured off-line using an Octet RED96e biolayerPCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01interferometry system (Sartorius, USA). A sample from the bioreactor 470 may be measured with the Biomod Raman spectrometer (e.g., system 100).

[0093] Table 1: Details regarding the three-perfusion cell culture runs conducted in the CHO testbed 470 in FIG. 4A.

[0094] Three perfusion runs were conducted and are described in Table 1. As the testbed 470 was also being developed as the runs were taking place, the runs were quite different from one another in regards to length, culture parameters and unexpected events which affected the inline Raman spectra. As an example, FIGS. 4B and 4C illustrate the progression of the first and second perfusion runs (R2 (FIG. 4B), R3 (FIG. 4C)) and some of the unexpected events such as foaming and changes in media composition. Run R2 predates the automation using the MAST® system 480 so samples were drawn manually once a day from the perfusate and measured in the Nova metabolite analyzer while spectra were acquired hourly.

[0095] Samples were drawn daily, approximately at noon, manually or automatically. A small volume was analyzed with the Nova analyzer 490 while the rest was spun down to remove the cells, and the supernatants were measured for mAb titer. The supernatant samples were frozen and kept in a -80 °C freezer. Fourteen samples from run R3, each 1.5 ml in volume, were aliquoted and measured with the Biomod and the SSRS probe 440 using the Superlum laser (5-second integration, repeated 12 times each with 5-6 mW of power). Each sample was measured with three different Raman spectrometers (Kaiser, Biomod, and SSRS). As shown in FIG. 4A, one or more probes (e.g., probes 440 and 440a) may be integrated into the bioreactor 470 to measure spectra inside the bioreactor 470 with one or more Raman spectrometers (e.g., SSRS 200 / 200a or Spectrometer 100).PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0096] In order to train models based on Raman data from the integrated Kaiser probe 440a, both spectra and ground truth measured using the Nova analyzer 490 are used. To make sure the spectra reflect the most accurate data, only spectra that were acquired up to an hour from the Nova analyzer 490 measurement were used to build the models disclosed herein (this mostly affected data for R2 since sampling was not automated). Some of the Nova measurements suffered from errors due to clogging or calibration issues and could not be used. Table 1 shows the total number of Raman spectra and Nova measurements acquired during each run while noting the available dataset for model training as “Usable sets”.

[0097] A-priori Information from Kaiser Spectra

[0098] FIG. 5 shows spectra acquired with the Kaiser probe 440a in-situ for all three runs. Since most of the relevant data for key analytes and the data we can currently acquire with the SSRS is in the fingerprint region (600-1800 cm ') the remainder of analysis will only use data from the Superlum spectral range between 810-1670However, Spectra in the 400-810 cm1and 2700-3600 cm1ranges may also be acquired since these ranges may also hold important information regarding water, ammonium, and other compounds.

[0099] The R2 spectra (top panel in FIG. 5) shows sharp peaks that may be attributed to roomlight leaking into the bioreactor 470, which was corrected for future runs. All of the spectra exhibit an increase in the fluorescence background with time, which is a sign of increasing biomass as cells expand in the bioreactor 470. The output data from the Kaiser system is given in 1 cm1intervals (the system resolution is 4 cm '). The Kaiser system may artificially extrapolate data to provide data in 1 cm1intervals.

[0100] To establish a baseline for comparison, we performed PLSR for 7 analytes: glucose, glutamine, glutamic acid, lactate, ammonium, Cell Density (CD) and mAb in the 810-1670 cm’1spectral range which includes 861 data points. Spectra from all three runs were then used for the creation of the regression models disclosed herein.

[0101] One of the regression models may include PLSR. PLSR does not rely on background removal, which is significant in these spectra as seen from FIG. 5. Additionally, PLSR may perform better than principal component analysis (PCA) for highly colinear spectra, which has been shown to be the case for many cell-culture processes. However, additional linear regression methods, including, but not limited to, linear regression (LR), multivariate linear regression (MLR), principal component regression based on principal component analysis (PCA) may be used. Alternatively, the model may include a machine-learning model,PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01including, but not limited to, a neural network, including but not limited to, a recurrent neural networks (RNN), artificial neural network (ANN), or a convolutional neural networks (CNN) may also be used.

[0102] The models disclosed herein may also be trained. For example, the models may be trained using uniformly sampled training data. In this example, (also referred to herein as “uniformly sampled”), the training set is chosen to be half of all usable dataset but uniformly distributed throughout the runs’ duration. This creates an evenly sampled process and the model may represent all stages of the cell culture. Alternatively, the models may be trained using a portion of sampled training data. In this example, (also referred to herein as “first half’) the training set similarly used half of the usable dataset but chosen to be only at the beginning of the perfusion culture. The validation (in this case also the prediction) may then be performed on the rest of the dataset, not including any of the data used for training. The number of PLSR components, Ncomp, is chosen automatically for each analyte separately to correspond to about 83-90% percent of cumulative variance (see details in Table 2)

[0103] Table 2: PLS model parameters and results from the two types of training methods, Uniformly Sampled and First Half.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0104] The above PLSR results on the obtained spectra show that preferably a process should be sampled uniformly throughout its duration in order to improve predictive modeling performance. However, in some embodiments, the model(s) disclosed herein may also be trained to perform regions of hyper sampling (e.g., for a region of interest in the spectrum). All future PLSR models described herein use the uniformly sampled training.

[0105] Method 1 - Down Sampling

[0106] The Kaiser spectra includes 860 data points in the Region of Interest (ROI). It would take approximately 24 hours to acquire equivalent spectra with the SSRS with 10-second integration times and 10 repetitions for each spectral point. As the signal intensity cannot be changed with the Superlum, reducing the integration period or decreasing the number of repetitions would likely compromise the sensitivity. Instead, we reduce the number of acquired data points to decrease sampling time.

[0107] First, we explore uniform down-sampling of the spectra and identifying the largest interval that allows us to reduce the number of data points while trying to reduce the effect on our signal resolution and predictive model performance. To evaluate this method, we created a simulation in which the data from the Kaiser system was down-sampled in increasingly larger intervals between 2 and 20cm'1. The performance was evaluated by performing PLSR on the original (see Table 2) and decimated spectra and comparing the RMSEE and RMSEP values Table 2.

[0108] FIGS. 6 A and 6B show the simulation results for all 7 analytes with RMSEE in FIG. 6 A and RMSEP in FIG. 6B. The results are normalized to the RMSEE and RMSEP values when no decimation is performed. Dashed lines mark 5% and 10% increase in error. Decimation by a factor of 4 had minimal effect on errors, resulting in only about 1% increase for RMSEE and about 1.2% for RMSEP. This result is expected as the resolution of the Kaiser system is 4cm'1, meaning that each data point contains information from a spectral range of 4cm'1. Adjacent pixels may contain some overlapping information that can be recovered from sampling at larger intervals that overlap with the system resolution.

[0109] A factor of four reduction in sampling rate is equivalent to sampling in intervals of 4 cm'1, corresponding to about 0.24 nm (at 770 nm) and about 0.27 nm (at 825 nm). Since the Superlum tunes in 0.1 nm intervals it was decided to sample spectra at 0.2 nm intervals, which would also provide some margin for laser wavelength instability.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0110] The suggested sampling interval results in about 276 spectral datapoints in the Superlum spectral range and still requires significant sampling time. In order to acquire spectra in a somewhat reasonable time frame, each spectra data point was acquired 12 times for 5 seconds each time to establish the mean and standard deviation (STD). The SSRS spectra acquired and the signal processing pipeline for the spectra is described below in the following section.

[0111] SSRS Spectra Signal Processing

[0112] After the minimal sampling parameters for the SSRS were found, spectra were acquired for 14 CHO-cell supernatant samples in a quartz cuvette using 5-6 mW of power. Spectra of the same samples were acquired with the Biomod system using 150 mW, 10-second integrations repeated 50 times each. FIG. 7A shows the spectra for 12 of the 14 days using all three systems.

[0113] Two of the SSRS spectra (corresponding to days 12 and 27 of the run) suffered from severe power fluctuations which was traced back to issues with the AC power supply and are excluded from the analysis.

[0114] The Kaiser spectra are presented after some internal processing (e.g., median filtering and smoothing) and has not gone through any additional processing steps. The Biomod data had been smoothed with a 3rdorder median filter and a 2ndorder, 11thdegree Savitzky-Golay filter.

[0115] The SSRS spectra may also involve some signal processing steps to enhance SNR and remove system artifacts. First, since the Single Photon Avalanche Detector (SPAD) is a single-point detector, it may be less prone to cosmic rays than a charge-coupled device (CCD) array. A median filter may be if extreme peaks are present in the signal (e.g., where the peak value is over 10 standard deviations above the mean signal for a specific wavelength).

[0116] Second, due to the laser power fluctuations and wavelength dependency a Savitzky-Golay filter, 2ndorder, 3rddegree, may be applied to the spectra (see FIG. 7B). These filter parameters were chosen to filter out laser fluctuations which are below the SSRS resolution while minimally attenuating the Raman peaks.

[0117] To better illustrate the differences of spectra acquired with all three systems, we inspect spectra acquired on Day 8 of the R3 run. These spectra are shown in FIG. 7C.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0118] FIG. 7C shows normalized spectra from three different Raman systems where the Kaiser data was acquired in-situ (middle spectrum) and includes both the supernatant and cells, while the SSRS (bottom spectrum) and Biomod (top spectrum) show spectra of the supernatant sample drawn that same day, frozen, and then thawed and measured by both systems. All spectra have been smoothed as described above, but no additional processing steps were performed.

[0119] Notably, the slow-varying fluorescence background curve is similar for the dispersive Biomod and Kaiser systems, (830 and 785nm excitation, respectively), where the higher fluorescence is closer to the excitation (0 cm'1). For the SSRS, however, the short wavenumber region is acquired at longer excitation wavelengths and the long wavenumber region is acquired with the shorter excitation wavelengths. This results in a “reverse” fluorescence curve.

[0120] Furthermore, the effective Raman cross-section is different for each spectral datapoint because each is acquired with a different excitation wavelength and is proportional to 1 / A4ex. In the case of photon counting detection (and not optical power), the dependency becomes 1 / A3exdue to the factor of hv (photon energy) in the calculation. This further exacerbates the background slant since the shorter wavenumbers suffer from a lower crosssection.

[0121] In order to account for the Raman cross-section wavelength dependency, each spectral datapoint j, using excitation wavelength Ay is multiplied by a correction factor given in Equation 1, whereois the shortest wavelength used (e.g., 770nm).A3WJ ■ = E- to -o /

[0122] Spectral background subtraction of an empty sample holder (possible for SSRS and the Biomod) and a following Lieber algorithm (6thorder) were used to remove residual fluorescence (for all systems) are used to reach the final spectra of Day 8 for all systems in FIG. 7D.

[0123] The SSRS, Biomod, and Kaiser spectra all have distinct similarities with the amide III region (around 1350 cm '), glucose (around 1125, 1060 cm ') and Glutamine (around 850 cm1) clearly visible, while the Kaiser spectra is expected to have some differences due to the presence of cells. The SSRS spectra has better peak contrast in the lower wavenumberPCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01region due to reduced fluorescence; however, there are some unexplained peaks in the 900-1000 cm1region, which do not appear in the Biomod or Kaiser data.

[0124] FIG. 7E shows the same 12 spectra presented in FIG. 7 A after background removal and cross-section correction. The final spectra of all systems may be sensitive to background removal, making direct spectra analysis without the use of training data challenging. The next sections explore analyte concentrations analysis based on both DPA and PLSR.

[0125] Method 2 - Direct Peak Analysis (DPA)

[0126] DPA is a method of detecting trends in analyte (e.g., protein, sugar, metabolite, drug, etc.) concentrations from Raman spectra where the locations of the Raman peaks are known in advance and the Raman signal is strong enough, e.g., surpasses the background shotnoise, to be detected reliability. One benefit of DPA lies in its simplicity and the fact that no training dataset is required for the analysis, though a-priori information may be used. Most DPA methods require the spectra to have visible Raman peaks, and some a-priori information regarding the peaks in the sample which are relevant for analysis, e.g., the target analyte pure spectrum or general curve fitting parameters. The relative increase or decrease of peak height or of the area under a peak can give both qualitative and quantitative estimations on analyte concentration.

[0127] Often, libraries of pure analytes expected to be in the sample are pre-compiled and the sample spectrum is then decomposed and fitted to the various components while estimating the relative analytes concentrations. Other, more computational demanding methods may use deconvolution (blind or based on prior knowledge) to deconvolve the spectrum into Gaussian, Lorentzian and Voigt functions, finding the likeliest ratios of components in a mixture.

[0128] However, DPA may have some disadvantages. First, it uses a single peak even if the analyte has several known Raman peaks. Second, when SNR is low, peaks may not be visible above the noise or background signal, rendering DPA useless. Furthermore, in complex samples, multiple overlapping Raman peaks can create ambiguity as to the analytes affecting the spectral variation. In these situations, Bayesian modeling often complements DPA by estimating both the peaks and background signals using likelihood probabilities, often improving model performance.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0129] FIG. 8 shows an example for DPA of glucose and mAb. The top panel shows spectra of 50 mg / ml glucose, the middle panel shows the spectra of Days 5, 6, and 8 in run R3 and the bottom spectra is of 10 mg / ml NIST mAb (RM 8671, IgGlx). While the NIST mAb is not the mAb produced in the CHO cells, it shares similar structure in the amide I, III regions. All spectra were acquired in a quartz cuvette with the SSRS using 5 second integration repeated 12 times. The middle shaded region is the location of glucose’ Raman peaks (about 1060, 1125 cm’1) and the outer two shaded regions are significant spectral regions for mAb (about 995, 1250 cm’1).

[0130] FIG. 8 also illustrates one of the challenges of DPA. The mAb spectra has other significant peaks, for example at about 1650 cm’1which is the Amide I band. However, it overlaps with the O-H stretching peak in water making it extremely difficult to detect with the excess noise. Similarly, glutamine has a Raman peak at about 850cm'1while lactate has a primary peak at about 860 cm’1, making them very hard to distinguish. Furthermore, with varying fluorescence background levels, as is the case for cell culture (see FIG. 7A) prevent DPA from being useful without extremely accurate and consistent background removal methods.

[0131] FIG. 9A shows the glucose peak (about 1125cm'1) DPA as a function of time for the R3 run. The left Y axis shows the SPAD counts and the right Y axis shows the ground truth measured with the Nova analyzer.

[0132] Observing the SSRS data error bars in the glucose plot, we see a STD (Is) of approximately Ig / L. Recalling the glucose 3 c limit of detection (LOD) measurement (see FIG.9D) which was also Ig / L (but with 10 second integration and 50 repetitions), we reach good agreement, expecting the SNR to be a factor of about 2.8 lower in the CHO glucose measurement due to half the integration time and only a quarter of repetitions: ( 0.5 ■ 0.25 = 0.35 = 1 / 2.8.

[0133] Another two examples are given in FIGS. 9B and 9C, where the mAb peak (about 1250 cm'1) and lactate peak (about 850cm'1) are monitored with DPA. While general trends of the lactate Raman peak are similar to those measured with the Nova analyzer, they do not have the same granularity and resolution but may still provide useful information regarding the process without any training data. For mAb, however, the DPA may fail to track the titer and may not be correlated with the Octet data.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0134] Despite the successful trend detection shown above for glucose and partially for mAh and lactate, these results may not transfer to other analytes that may have low concentrations, have lower Raman cross sections, or lack observable peaks in the region of interest (ROI). However, DPA may be used to prioritize certain analytes.

[0135] One challenge around acquisition time may not be resolved by using decimation and DPA. The total acquisition time for all Raman systems was 75*10seconds =750 seconds for the Kaiser (about 3400 data points), 50^10 seconds=500 seconds for the Biomod (about 1340 data points) and an overwhelming 16,560 seconds for the SSRS (5 seconds repeated 12 times each for each of the 276 datapoints). With the laser tuning overhead, the total measurement took about 5 hours to complete. Such a long acquisition time makes the measurements more susceptible to drift due to the laser and / or also to changing environmental and sample dynamics, emphasizing the benefits of reducing the acquisition time.

[0136] FIG. 10 shows the effect of extended measurement time on a supernatant sample. The sample of Day 8 from run R3 was measured repeatedly with the SSRS four consecutive times. The spectra show clear changes in the fluorescence background level, particularly for longer wavenumbers, and also changes in peak structure. Importantly, the sample placement in a quart cuvette, without the constant mixing which occurs in bioreactors, contributes to the degradation of the sample due to extended exposure to the excitation radiation.

[0137] If we attempt to perform a limited Raman peak acquisition, the fluorescence background may be a limiting factor. Polynomial fitting methods may be used to remove the background. However, these methods use a significant number of data points, particularly for complex spectra with many overlapping peaks and if we acquire only the peak region, we risk misinterpreting the data. Thus, signal processing and spectral analysis methods that do not rely on background removal may have a distinct advantage particularly for bioprocess monitoring.

[0138] Method 3 - Informed Sampling by Variable Importance in Projection (VIP)

[0139] VIP is a metric by which the importance of variables (e.g., spectral data points) on the output variance (e.g., analyte estimation) may be measured. The derivation of the VIP scoring is provided below.

[0140] Once VIP scores are calculated, a threshold is set (for example, 1) and only variables above the threshold comprise the final model. FIGS. 11A and 11B show thePCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01regression coefficients and VIP scores, respectively, for a PLSR model estimating the glucose concentration in a cell culture dataset.

[0141] Herein we explore the use of a-priori Raman data for minimal spectral sampling, in order to further reduce the SSRS acquisition time. Our adaptive and selective sampling technique samples only high-significance areas of the spectrum. The assignment of significance to spectral data points can be done in many ways. In one example, the assignment of significance to spectral data points is obtained using PLSR and VIP scores. This approach supports the SSRS utility model for applications where a known set of analytes are used as metrics for a process progress or success. However, if random samples are measured each time with different target analytes, a full spectrum may be acquired first.

[0142] PLSR is a widely used spectral analysis technique that does not rely on fluorescent background removal. PLSR (described below) uses mean-centered spectra, X, as inputs, e.g., the smoothed spectra which includes the fluorescent background shifted to have mean=0. The mean value of each spectrum is not lost but used to build the regression coefficients). The regression coefficient vector, ft, given in Equation 2 and repeated here for convenience, provides the “weight” each data point is given for the final analyte prediction, C:

[0143] In addition to the regression coefficients, we can calculate the VIP for each analyte (see FIGS. 11 A and 1 IB). VIP allows us to map the useful parts of the spectra for the modeling of each analyte but does not include information regarding the direction of correlation. As an example, we look at the regression coefficients and VIP scores for lactate which were calculated above and given in FIGS. 12A and 12B. If a single analyte model is used (e.g., PLSR1), the VIP may change between analytes. If a matrix model is used (e.g., PLSR2), the analyte(s) may be grouped together as a matrix and a single VIP score may be provided for all of the analyte(s). The models disclosed herein may utilize PLSR1, PLSR2, and / or a combination of PLSR1 and / or PLSR2.

[0144] The regression coefficient and VIP show resemblance to lactate spectra with a dominant peak in 860 cm’1. This peak is marked as an area of importance by VIP, with a value larger than 1, and also with a ? > 0, indicating this peak positively contributes to the lactate concentration estimation. On the other hand, different regions marked as significant in VIPPCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01map (910, 1130, 1348, 1542, 1635 cm’1) have significantly lower p values, some of which are even negative.

[0145] For spectra containing hundreds of data points, Equations 13 and 14 describe an ill-posed problem and so feature selection methods are used to reduce the dimensionality of the system. For the sake of predictive modeling, it may be worthwhile to select high importance VIP regions as all the spectra had already been sampled, even regions with negative p values. Low VIP areas which are likely to add noise may be ignored. However, the same cannot be said for selective sampling.

[0146] With SSRS, the objective is to minimize the number of acquired points and so it is worth considering avoiding sampling points with negative P values all together. Thus, the models disclosed herein may further reduce spectral sampling time by not sampling negatively correlated areas of a spectrum. In order to test this hypothesis, a simulation based on Kaiser data was constructed as described below.

[0147] VIP -Informed Sampling Simulation

[0148] FIG. 13 illustrates a simulation and validation process used to assess the selective VIP sampling method. In the first step all data (861 data points) in the 810-1670 cm1spectral range was used to compute VIP scores and regression coefficients for all 7 analytes (described above). The RMSEE and RMSEP values (see Table 2) may be used to benchmark the performance for future iterations. FIG. 14 shows the VIP scores in the first step, where the colored markers indicate the significant VIP regions for which VIP >1.

[0149] The user then inputs the analytes (e.g., a protein, a sugar, a metabolite, a drug, etc.) for which selective sampling should be performed so only their respective VIP scores and P values are considered. In the simulation results presented below, all 7 analytes were considered.

[0150] The second step assigns a value for each data point. Two methods may be used to evaluate the significance of each data point. In the first method, the VIP scores of all analytes (FIG. 14) are summed, creating a shared “VIP heat map” (FIG. 15 A), where the top spectrum markers indicate where the cumulative VIP score was greater than 1 for all analytes. In the second method, termed VIPxBeta, the products of the VIP scores and normalized p values are calculated for each analyte. Normalized p values allow us to compare analytes without accounting for their different units of measurement. FIG. 15A shows the data point ranking of the second method in the bottom spectrum.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0151] Spectral Data Point Selection

[0152] After the VIP heat maps are created, the user and / or the model may select the desired number of spectral datapoints, Kp, and they are ranked according to their value. For example, the model may select a Kp value based on the VIP heat maps. Instead of, or in addition to, the model selecting the Kp value, a user may select and / or modify the Kp value (e.g., to adjust for the desired acquisition duration). The Kp value impacts the length of time required to acquire a spectrum. Thus, the Kp value may be adjusted (e.g., by the user and / or model) to reduce spectral acquisition time while still maintaining accuracy. The Kp value may also be adjusted (e.g., by the user and / or model) based on the analyte(s) being measured. For example, some analytes (e.g., glucose) may allow for a lower Kp value since the analyte may have strong Raman peaks. As a result, it may be possible to provide information about analyte(s) with strong Raman peaks with a low Kp value (e.g., less than 100). In other words, analyte(s) with strong Raman peaks may be identified with fewer number of spectral datapoints. In contrast, other analyte(s) with weaker Raman peaks may (e.g., mAb), may require a higher Kp value (e.g., greater than 100). The number of spectral datapoints, Kp, may range from about 30 to about 900, including all values in between. For example, the number of spectral datapoints may be 30, 35, 40, 35, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, 500, 600, 700, 800, or 900, including all values in between. The number of spectral datapoints, Kp, may vary based on the analyte(s) being measured. For a sample including multiple analytes, the model may allow for different Kp values based on the analyte. For example, the first analyte may have a Kp value of 100 and the second analyte may have a Kp value of 200, or vice versa.

[0153] FIG. 15B shows an example of selecting 50 data points in both methods. The top spectrum dots show the position of the highest-ranking scores in the VIP method and the bottom spectrum shows the ranking in the VIPxBeta method.

[0154] In the third and final step of the simulation, the PLSR is computed again on the reduced spectra with Kp data points. The PLSR models were computed as before with about 50% of the CHO data used for training and about 50% for prediction. The number of principal components, Ncomp, was recalculated for each analyte to guarantee at least 83% explained variance based on the new partial spectra. For glucose, lactate, CD and mAb, Ncomp remained fairly consistent, changing by about 1 or 2 components even for 30 data points.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0155] An example of this is given in FIG. 16 that shows the variance explained in glucose concentration estimation as a function of number of PLS components for both the full spectrum and a 50-point reduced spectrum. However, for glutamine, glutamic acid and ammonium, which require a larger number of components, there was a significant increase and particularly for Kp<60, the maximal variance explained did not reach about 80% even for the maximal number of components. This is also attributed to the limited spectral range we examined in this simulation.

[0156] Simulation Results

[0157] The simulation was run for Kp values between 215 and 30 for all 7 analytes. Table 3 and Table 4 show the results in the VIP and VIPxBeta methods respectively. The results show that for both methods, it is possible to significantly reduce the number of spectral data points and have a relatively small effect on the RMSEE and RMSEP values (e.g., without reducing the accuracy of the determined analyte concentration). The VIPxBeta method may perform slightly better. As an example, for KP=80, which signifies an order of magnitude reduction in acquisition time, results show only an average about 19.6% increase in RMSEE and about 12.4% increase in RMSEP (VIPxBeta).

[0158] Table 3: Simulation results for analyte concentrations with various number of spectral data points with the VIP ranking method.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0159] FIGS. 17A-17D show the predicted PLSR values versus the measured values for Kp =40 data points (40 minutes per spectrum) for glucose, lactate, total CD and mAh in the VIPxBeta method. The plots include data from different CHO runs that were monitored using the Kaiser system (see FIG. 5).

[0160] Table 4: Simulation results for analyte concentrations with various number of spectral data points with the VIPxBeta ranking method.

[0161] The above results suggest that it is possible to significantly reduce acquisition time from approximately 5 hours to about 40 minutes while still maintaining good prediction values using VIP-informed sampling. The following section will put this theory to the test by selectively sampling with the SSRS.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0162] The Kaiser simulation data showed that based on VIP or VIPxBeta sampling, estimation and prediction errors were consistent, supporting the use of this method for analyte prediction with reducing spectral acquisition time. While the RMSEE errors were larger than those of SSRS (e.g., glucose had an RMSEE value of 0.6-0.7 g / L and SSRS only 0.4-0.5 g / L), the Kaiser data was monitoring the testbed development processes that suffered from large variability in culture conditions. SSR Spectrometry results, particularly in regards to validation errors, may be limited due to the small dataset but still show that essential metabolite and mAb values may be captured using informed sampling with as little as 40 data points.

[0163] FIG. 19 illustrates a method 1900 for minimizing spectral sampling with a SSR spectrometer. In the first step 1902, a first Raman spectrum of a first sample is acquired. Preferably, the first Raman spectrum is an entire Raman spectrum of the sample obtained using an SSR spectrometer. The sample may include at least one analyte and / or a plurality of analytes. For example, the sample may include one analyte, two analytes, three analytes, four analytes, five analytes, six analytes, seven analytes, eight analytes, nine analytes, and so on. In the second step 1904, one or more region(s) of interest in the first Raman spectrum are identified. The models disclosed herein may identify the one or more region(s) of interest in the first Raman spectrum. The models disclosed herein may further output the region of interest(s) in the first Raman spectrum In the third step 1906, a model may be trained on the region of interest of the first Raman spectrum. The model may also be trained on the entire first Raman spectrum. In the fourth step 1908, a number of spectral datapoints is selected. The number of spectral datapoints may range from about 30 to about 900, including all values in between. For example, the number of spectral datapoints may be 30, 35, 40, 35, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, 500, 600, 700, 800, or 900, including all values in between. In the fifth step 1910, a reduced Raman spectrum may be acquired of a second sample based on the model and the number of spectral datapoints. The first and second samples may be the same sample. Alternatively, the first and second samples may be different samples. Preferably, the first and second samples include at least one analyte in common. For example, the first and second samples may have one analyte, two analytes, three analytes, four analytes, five analytes, six analytes, seven analytes, eight analytes, nine analytes, and so on in common. For example, the first and second analytes may share a plurality of analytes in common.

[0164] The method 1900 may be used by the SSR spectrometer 200 or 200a. For example, the SSR spectrometer 200a may use the method 1900 with manufacturing line 10 toPCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01determine a presence and / or concentration of at least one analyte. In operation, using the method 1900, the SSR spectrometer 200a may obtain at least one reduced Raman spectrum using probes 140a-140f at a selected site along the continuous manufacturing line 10. The computing device 290a may store in memory instructions for implementing the method 1900. When the instructions are executed by the computing device 290a, the instructions cause the computing device 290a to perform the method 1900. For example, the computing device 290a may cause the SSR spectrometer 200a to obtain the at least one reduced Raman spectrum using probes 140a-140f at a selected site along the continuous manufacturing line 10. The computing device 290a may also be configured to determine the presence and / or concentration of at least one analyte in the continuous manufacturing line 10. Preferably, the time required to obtain the at least one reduced Raman spectra is less than the time required to obtain a full Raman spectrum.Validation by SSRS Measurement

[0165] It is generally quite difficult to transfer PLSR model results between different runs, let alone different Raman systems. Any Raman spectrum includes specific system spectral features from the lenses, filters, or optical windows, which are captured along with the sample under inspection. Particularly for the SSRS and Kaiser, which have different hardware and were also inspecting different samples (supernatants versus media with cells) it is unlikely one model can be used to inform another. However, we can use the acquired SSRS spectra of the 12 CHO samples from run R3 (described above) to evaluate the VIP -informed sampling strategy.

[0166] PLSR parameters for SSRS Spectra

[0167] Since only 12 CHO samples from run R3 were available for SSRS measurements, the data was not split into equal data sets of training and validation but rather a “leave-2-out” cross validation was performed iteratively. In each iteration, 10 of the 12 spectra were used for training and 2 spectra were used for validation until all permutations were exhausted (12*11 / 2=66). The regression coefficients and final RMSEE and RMSEP are the mean values computed in all iterations.

[0168] The number of PLSR components for each analyte were determined automatically to be the lowest number for which the explained variance would exceed about 83%. Preferably, the variance does not exceed 85% to reduce overfitting.

[0169] SSRS PLSR resultsPCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0170] The reduced number of samples limits the PLSR ability to predict analyte concentrations from previously unseen spectra and also limits the number of possible PLS components, further hindering performance. Additionally, as was established before, the SNR of the SSRS spectra was lower compared with the LOD measurement due to short integration times (e.g., about 5 seconds) and a low number of repetitions (12). Table 5 and Table 6 show the measurement results for VIP informed sampling and VIPxBeta informed sampling, respectively.

[0171] Overall, the RMSEE values for both VIP informed sampling and VIPxBeta informed sampling were quite good and remained so even for Kp<70 for both selection methods.

[0172] Table 5: SSRS measurement results for analyte concentrations with various number of spectral data points with the VIP ranking method.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0174] The RMSEP errors are significantly higher than RMSEE, attributed to the limited capacity of the PLSR model with such a limited sample number. However, the prediction errors remain similar even for a reduced number of Kp spectral points, indicating the usefulness of VIP informed sampling.

[0175] A check was performed by sampling Kp random data points and comparing to the VIP-informed methods (Table 7). The random RMSEE errors were quite low, but RMSEP values were significant for all analytes and all values of Kp. Thus, the random selection method may produce larger errors compared with the VIP-informed methods, further indicating the value in VIP- informed sampling.

[0176] The large differences between RMSEE and RMSEP values may be indicative of over-fitting, due to the sample size.

[0177] Table 7: SSRS measurement results for analyte concentrations with various number of randomly selected spectral data points.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0178] The SSRS PLSR still allows us to track trends in analyte concentrations, despite errors. FIG. 18 shows the SSRS PLSR results for Kp=40 and the ground truth of the Nova analyzer.VIPxBeta

[0179] PLSR

[0180] The following is a PLSR formulation for a general case where C is a response matrix, e.g., including J analytes for each sample in the learning dataset that include NLsamples and sized [NLx J] . In most PLSR Raman spectroscopy applications, C is a single analyte vector (i.e. J=l). However, the models disclosed herein include a response matrix that can perform on any J as long as J<NL.

[0181] PLSR is an iterative process likened to “peeling an onion” - each iteration identifies a significant spectral component based on maximal correlation in latent space and removes it. The process is repeated until we can account for a certain percentage of the correlation in the response based on the input. For example, a threshold may be set to do a certain number of iterations until at least 80% of the response correlation is explained to avoid overfitting but provide a solid model estimation.

[0182] The below explains the projection of both input (e.g., spectra) and response (e.g., analyte concentration) to latent spaces in a way that maximizes the covariance, where covariance is the correlation between the input and response.

[0183] For iteration of i we assume there are a pair of unit vectors, in the input space and Wi in the output space so that:PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01rt= w Ct (4) where ztand rtare called the input and output scores, respectively. We wish to find Vt and wtwhich maximize the correlation of ztand rt, which is equivalent to maximizing the covariance of the mean-centered matrices Xband C):

[0184] By using Equations 3-5 we get the following expression:

[0185] There are several methods that may used for finding these unit vectors which maximize the scores correlation.

[0186] After finding the scores, we wish to predict the concentration C£from our score Zi using the output loading vector q£: this is the actual model step where we use these results to create an estimation of the analyte:

[0187] By minimizing the estimation error (Equation 8) a solution for the optimal loading vector q0 ;is found (Equation 9). This may be solved by formulating Equation 9 as a minimum seeking problem and finding the minima with Lagrange multipliers and / or other numerical methods. This has an analytical solution (Equation 9):>

[0188] In order to find the next latent component, we first remove the first latent component from both the predictor (input) and response (output), a process called “deflation”. The deflated output, Ci+1, can be found from Equations 7 and 9:Ci+1= Ci - Ci = Ci - qOtiViTXi (10)

[0189] The derivation for the deflated input, X£+1can be shown to be:PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0190] Where the input loading vector, p0 bis defined as:_ C0V(Xt,Xt)vt(12)Po ivlTC0V(Xl,Xl)vl

[0191] And finally, after Ncomp latent variables have been computed, the final regression coefficients vector, p = H, is given by Equation C.10 that creates a regression vector Beta from contributions of each iteration:

[0192] The VIP (Variable importance in the Projection) score can therefore be described by Equations 13-14, where the index z stands for the index of latent variable so that i = [1,2, —,Ncomp] and the index m = [1,2, ...,M], stands for the variable (wavenumber) index.

[0193] Each wavenumber’s normalized contribution to the overall response estimation is:

[0194] The PLSRxBeta selection process

[0195] For the PLSRxBeta selection process, we multiply the VIP vector with a normalized regression vector to account for different units using Equation 15:

[0196] We then choose N highest values of VIP ■ p.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0197] We then repeat the PLSR process only with these variables to create a much smaller model.Conclusion

[0198] While various inventive embodiments have been described and illustrated herein, those of ordinary skill in the art will readily envision a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the inventive teachings is / are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific inventive embodiments described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, inventive embodiments may be practiced otherwise than as specifically described and claimed. Inventive embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the inventive scope of the present disclosure.

[0199] Also, various inventive concepts may be embodied as one or more methods, of which an example has been provided. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.

[0200] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.

[0201] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01

[0202] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.

[0203] As used herein in the specification and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of’ or “exactly one of,” or, when used in the claims, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e. “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” “Consisting essentially of,” when used in the claims, shall have its ordinary meaning as used in the field of patent law.

[0204] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at leastPCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.

[0205] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03.

Claims

PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01CLAIMS1. A method for reducing spectral acquisition time for a sample using a swept source Raman (SSR) spectrometer, the method comprising:acquiring a full Raman spectrum of the sample using the SSR spectrometer, wherein the sample comprises a plurality of analytes;identifying at least one region of interest in the Raman spectrum for each of the plurality of analytes;training a model on the at least one region of interest for each of the plurality of analytes using the full Raman spectrum of the sample;selecting a reduced number of spectral datapoints less than a number of spectral datapoint in the full Raman spectrum; andacquiring a reduced Raman spectrum of the sample with the reduced number of spectral datapoints using the SSR spectrometer based on the model, wherein the sample comprises the plurality of analytes.

2. The method of claim 1, further comprising:determining a concentration of at least one analyte of the plurality of analytes in the sample based on the reduced Raman spectrum.

3. The method of claim 2, wherein the sample is a cell culture and determining the concentration of the at least one analyte of the plurality of analytes in the sample comprises determining the concentration of the at least one analyte of the plurality of analytes in the sample at multiple points of a manufacturing process.

4. The method of claim 1, wherein the reduced number of spectral datapoints is less than 100.

5. The method of claim 4, wherein the reduced number of spectral datapoints is less than 50.

6. The method of claim 1, wherein acquiring the reduced Raman spectrum comprises acquiring highest ranking spectral datapoints for the reduced number of spectral datapoints.

7. The method of claim 1, wherein the model comprises a partial least squares regression (PL SR) model.PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO018. The method of claim 1, wherein the model reduces an acquisition time of the reduced Raman spectrum by a factor of at least five.

9. The method of claim 8, wherein the model reduces the acquisition time of the reduced Raman spectrum by a factor of at least seven.

10. The method of claim 1, wherein training the model further comprises training the model on Raman spectra of each analyte of the plurality of analytes of the sample.

11. The method of claim 10, wherein the selecting the number of spectral datapoints comprises selecting the number of spectral datapoints based on the Raman spectra of at least one of the plurality of analytes.

12. A system for measuring Raman signals of a sample, the system comprising:a tunable laser to emit a tunable excitation beam;an optical fiber to transmit the tunable excitation beam to a measurement site; an optical probe, in optical communication with the optical fiber, to illuminate the measurement site with the tunable excitation beam and to collect the Raman signal emitted from the measurement site in response to the tunable excitation beam;a spectrally selective detector, in optical communication with the optical probe, to sense the Raman signal from the measurement site; anda processor, operably coupled to the optical probe, configured to run a model, wherein the model is configured to identify at least one region of interest in the Raman signal of the sample and to cause the tunable laser to emit the tunable excitation beam to acquire a reduced Raman signal of the least one region of interest in the Raman signal of the sample using a selected number of spectral datapoints.

13. The system of claim 12, wherein the model comprises a partial least squares regression (PLSR) model.

14. The system of claim 12, wherein the sample comprises a cell culture comprising a plurality of analytes and the model is further configured to determine a concentration of at least one of the plurality of analytes using the reduced Raman signal.

15. A method for swept-source Raman spectroscopy, the method comprising:PCT / US25 / 21552 26 March 2025 (26.03.2025)Attorney Docket No. MIT-26528WO01acquiring swept-source Raman spectra of a sample over a predetermined band of excitation wavelengths;training a machine-learning model on the swept-source Raman spectra; identifying, with the machine-learning model, a subset of excitation wavelengths corresponding to a region of interest in the swept-source Raman spectra; andacquiring a swept-source Raman spectrum of the sample for only the subset of excitation wavelengths.

16. The method of claim 15, further comprising:determining a concentration of at least component of the sample based on the swept-source Raman spectrum of the sample for the subset of excitation wavelengths.

17. The method of claim 15, wherein acquiring the swept-source Raman spectrum of the sample for only the subset of excitation wavelengths has a lower acquisition time than acquiring the swept-source Raman spectra of the sample over the predetermined band of excitation wavelengths.

18. The method of claim 15, wherein identifying, with the machine-learning model, the subset of excitation wavelengths comprises identifying an importance score for the region of interest based on a signal intensity the swept-source Raman spectra.

19. The method of claim 15, wherein acquiring the swept-source Raman spectrum of the sample further comprises acquiring the swept-source Raman spectrum for a selected number of spectral datapoints for the subset of excitation wavelengths.