Module and method for aligning measurements between a plurality of spectrometers

The method aligns spectral measurements across different spectrometers using interpolation, preprocessing, and machine learning to encode spectra into embeddings, addressing inconsistencies and ensuring accurate, device-independent spectral analysis.

WO2025178561A1PCT designated stage Publication Date: 2025-08-28PROFILEPRINT PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/SG2025/050100
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-23
Filing Date
2025-02-14
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing spectral analysis methods in chemometrics are limited in their ability to generalize across different spectrometers, leading to inconsistencies and biases in quality control and product discovery applications due to spectrometer-induced variations, which affect the accuracy and objectivity of spectral comparisons.

Method used

A method and module that aligns measurements across spectrometers using interpolation, optimal preprocessing techniques, and a trained machine learning model to encode spectra into numerical vectors and embeddings, minimizing correlation loss through metric learning and triplet mining to ensure consistency and accuracy.

Benefits of technology

The method effectively reduces inter-spectrometer variability, enabling unbiased and consistent spectral analysis, ensuring accurate and device-independent comparisons of sample quality and similarity across different spectrometers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SG2025050100_28082025_PF_FP_ABST
    Figure SG2025050100_28082025_PF_FP_ABST
Patent Text Reader

Abstract

This document describes a module and method for aligning measurements performed on a sample by a spectrometer with measurements performed on the same sample by other spectrometers. Additionally, the module and method is configured to detect, in real-time, the quality of a sample using two different spectrometers.
Need to check novelty before this filing date? Find Prior Art

Description

MODULE AND METHOD FOR ALIGNING MEASUREMENTS BETWEEN A PLURALITY OF SPECTROMETERSCROSS REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of priority to Singapore patent application no. 10202400493P which was filed on 23 February 2024, the contents of which are hereby incorporated by reference in its entirety for all purposes.TECHNICAL FIELD

[0002] This application relates to a module and method for aligning measurements performed on a sample by a spectrometer with measurements performed on the same sample by other spectrometers. Additionally, the module and method is configured to detect, in realtime, the quality of a sample using two different spectrometers.BACKGROUND

[0003] The development of chemometrics has significantly advanced the application of spectral analysis in the fields of agricultural and farm products. Chemometrics is widely employed for both qualitative and quantitative analysis of infrared and Raman spectra, enabling precise identification and measurement of substances. Conventional chemometric methods involve two main steps: spectral pre-processing and model building of a calibration model. Spectral pre-processing plays a crucial role in removing noise from the spectra and improving prediction accuracy with each method addressing different types of distortions in spectral measurements. However, selecting an optimal combination of pre-processing techniques through trial and error can be time-consuming. Additionally, when the type of instrument used is changed, or when there are variations in the measured sample, this results in variations in the noise distribution which in turn reduces the effectiveness of pre-existing pre-processing techniques and degrades the performance of the model.

[0004] To address these challenges, those skilled in the art have explored the feasibility of combining deep learning methods with spectral analysis techniques. Deep learning models, such as convolutional neural networks (CNNs), provide a robust approach for spectral analysis by utilizing local connectivity and weight sharing to efficiently extract and learn hierarchical features directly from raw spectra. Although CNNs have been applied to spectral analysis,existing models still rely on spectral pre-processing to improve accuracy. Hence, this approach remains limited in its ability to generalize across varying spectrometers, as they do not address the inherent variability in spectral measurements across different devices.

[0005] One major challenge in spectral analysis is that spectra from the same sample can vary significantly depending on the spectrometer used to measure them. This variability poses critical challenges in commercial applications, particularly in quality control (QC) and product discovery applications. For example, in a QC workflow, a sample may be scanned at the producer's facility using one spectrometer and later analyzed at the buyer’s location using a different spectrometer. Any discrepancies in the spectra should reflect differences in the sample itself, not inconsistencies introduced by the instruments. If spectrometer-induced variations are not corrected, they may lead to false failures in quality control.

[0006] Similarly, in product discovery applications, users frequently compare a target spectrum against a database of reference spectra to identify the closest match. The comparison process should be independent of the spectrometer used, ensuring that the results are true comparisons of spectral similarities rather than biases introduced by spectrometer variations. For example, a company comparing its inventory with multiple suppliers should obtain objective results, rather than being biased toward its own inventory simply because the same spectrometer was used for those measurements.

[0007] To ensure unbiased and consistent spectral analysis, there is a need for a robust, generalizable method to align the spectra across different spectrometers in a device- independent manner. Traditional calibration transfer methods often rely on pre-trained models, which ties the alignment process to a specific model rather than to the spectra itself. This can limit adaptability and reduce generalization across instruments. Hence, despite the efforts of those skilled in the art, it is still a challenge to find an effective solution that is able to align raw spectra directly to ensure consistency across different spectrometers and measurement environments so that accurate and unbiased decisions may be made based on the measured spectral comparisons.SUMMARY

[0008] In one aspect of the present disclosure, a method for aligning measurements performed by a first spectrometer with measurements performed by a second and a third spectrometer is disclosed. The method comprises the steps of receiving, using a computing device, first spectra associated with a sample as measured by the first spectrometer, wherein the first spectra comprises a plurality of peaks and valleys measured at unevenly spaced wavelengths and applying, using the computing device, an interpolation process to the first spectra to align the first spectra to a common wavelength range. The method then comprises then steps of preprocessing, using the computing device, the interpolated first spectra by selecting and applying an optimal preprocessing technique, thereby encoding the data into first numerical vectors and generating, using the computing device, aligned first embeddings by applying a trained machine learning model to the first numerical vectors. In embodiments of the disclosure, the trained machine learning model is trained, using the computing device, by receiving spectra associated with a training set of samples as measured by the second spectrometer and the third spectrometer. For each of the received spectra, the computing device then applies the interpolation process to the spectra measured by the second and third spectrometers to generate interpolated second and third spectra, and applies the optimal preprocessing technique to the interpolated second and third spectra to encode the interpolated second and third spectra into corresponding second and third numerical vectors, correlates the second numerical vectors with the third numerical vectors by applying a metric learning technique to second and third embeddings generated by the machine learning model for the corresponding second and third numerical vectors, adjusts one or more parameters of the machine learning model to minimize a correlation loss between the second and third embeddings, repeats, using the computing device, at least the correlating step and the adjusting step until it is determined that the correlation loss between the second and third embeddings is less than a predetermined loss threshold. The computing device then sets the parameters of the trained machine learning model based on a final adjusted one or more parameters of the machine learning model.

[0009] In a further embodiment of this aspect, the step of selecting and applying the optimal preprocessing technique comprises the steps of applying a plurality of preprocessing techniques to the interpolated first spectra, wherein each preprocessing technique modifies the interpolated first spectra to enhance signal quality and remove baseline noise, evaluating the pre-processed first spectra using a predefined metric to determine effectiveness of each of the preprocessingtechniques, selecting the optimal preprocessing technique based on the evaluation, and applying the selected preprocessing technique to the interpolated first spectra

[0010] In a further embodiment of this aspect, a method for determining quality of a subject sample comprises the steps of generating, using the method according to this aspect, aligned subject first embeddings for the subject sample. The method then receives, using the computing device, a reference spectra associated with the subject sample measured by a reference spectrometer, then applies, using the computing device, the interpolation process to the reference subject spectra, and the optimal preprocessing technique to the interpolated reference subject spectra to encode the interpolated reference subject spectra into reference subject numerical vectors. The method then generates, using the computing device, reference subject embeddings for the reference subject numerical vectors using the trained machine learning model, computes, using the computing device, a similarity score based on a Euclidean distance, a Manhattan distance, a Mahalanobis distance or a Cosine distance between the aligned subject first embeddings and the reference subject embeddings; and determines based on the similarity score if the quality of the subject sample meets a predetermined quality threshold.

[0011] In another aspect of the present disclosure, a computing module for aligning measurements performed by a candidate spectrometer with measurements performed by a target spectrometer is disclosed. The disclosed computing module comprises a processing unit, and a non-transitory media readable by the processing unit, the media storing instructions that when executed by the processing unit causes the processing unit to receive first spectra associated with a sample as measured by the first spectrometer, wherein the first spectra comprises a plurality of peaks and valleys measured at unevenly spaced wavelengths, apply an interpolation process to the first spectra to align the first spectra to a common wavelength range, preprocess the interpolated first spectra by selecting and applying an optimal preprocessing technique, thereby encoding the data into first numerical vectors, and generate aligned first embeddings by applying a trained machine learning model to the first numerical vectors. In embodiments of the disclosure, the trained machine learning model is trained by receiving spectra associated with a training set of samples as measured by the second spectrometer and the third spectrometer, for each of the received spectra, applying the interpolation process to the spectra measured by the second and third spectrometers to generate interpolated second and third spectra, and applying the optimal preprocessing technique to the interpolated second and third spectra to encode the interpolated second and third spectra into corresponding second andthird numerical vectors, correlating the second numerical vectors with the third numerical vectors by applying a metric learning technique to second and third embeddings generated by the machine learning model for the corresponding second and third numerical vectors, adjusting one or more parameters of the machine learning model to minimize a correlation loss between the second and third embeddings, repeating at least the correlating step and the adjusting step until it is determined that the correlation loss between the second and third embeddings is less than a predetermined loss threshold, setting the parameters of the trained machine learning model based on a final adjusted one or more parameters of the machine learning model.

[0012] Tn a further embodiment of this aspect, a computing system for determining quality of a subject sample is disclosed. The disclosed system comprises the computing module being configured to generate aligned subj ect first embeddings for the subj ect sample according to this aspect. The computing module then receives, using the computing device, a reference spectra associated with the subject sample measured by a reference spectrometer, applies the interpolation process to the reference subject spectra, and the optimal preprocessing technique to the interpolated reference subject spectra to encode the interpolated reference subject spectra into reference subject numerical vectors, generates reference subject embeddings for the reference subject numerical vectors using the trained machine learning model, computes a similarity score based on a Euclidean distance, a Manhattan distance, a Mahalanobis distance or a Cosine distance between the aligned subject first embeddings and the reference subject embeddings, and determines based on the similarity score if the quality of the subject sample meets a predetermined quality threshold.BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Various embodiments of the present disclosure are described below with reference to the following drawings:Figure 1 illustrates a block diagram representative of a system for determining the quality of a sample using two spectrometers in accordance with embodiments of the present disclosure;Figure 2 illustrates a block diagram representative of a processing system for performing embodiments of the present disclosure;Figure 3 illustrates a block diagram representative of a system for training a machine learning model in accordance with embodiments of the present disclosure;Figure 4 illustrates a block diagram representative of a workflow of a triplet mining process in accordance with embodiments of the present disclosure;Figure 5 illustrates a flowchart showing a process for aligning measurements performed by a spectrometer with measurements performed by other spectrometers in accordance with embodiments of the disclosure; andFigures 6a and 6b illustrate the clusters of various coffee types, where spectra of the various coffee types as measured by different spectrometers are compared with the corresponding embeddings generated by a trained machine learning model for the measurements of the different spectrometers.DETAILED DESCRIPTION

[0014] The following detailed description is made with reference to the accompanying drawings, showing details and embodiments of the present disclosure for the purposes of illustration. Features that are described in the context of an embodiment may correspondingly be applicable to the same or similar features in the other embodiments, even if not explicitly described in these other embodiments. Additions and / or combinations and / or alternatives as described for a feature in the context of an embodiment may correspondingly be applicable to the same or similar feature in the other embodiments.

[0015] In the context of various embodiments, the articles “a”, “an” and “the” as used with regard to a feature or element include a reference to one or more of the features or elements.

[0016] Tn the context of various embodiments, the term “about” or “approximately” as applied to a numeric value encompasses the exact value and a reasonable variance as generally understood in the relevant technical field, e.g., within 10% of the specified value.

[0017] As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.

[0018] As used herein, “comprising” means including, but not limited to, whatever follows the word “comprising”. Thus, use of the term “comprising” indicates that the listed elements arc required or mandatory, but that other elements arc optional and may or may not be present.

[0019] As used herein, “consisting of’ means including, and limited to, whatever follows the phrase “consisting of’. Thus, use of the phrase “consisting of’ indicates that the listed elements arc required or mandatory, and that no other elements may be present.

[0020] As used herein, the terms "first," "second," “third, “fourth” and the like in the description, in the claims, and in the figures arc used for distinguishing between similar objects and not necessarily for describing a particular sequential or chronological order.

[0021] Further, one skilled in the art will recognize that certain functional units in this description have been labelled as modules, sub-modules or sets of processing elements throughout the specification. The person skilled in the art will also recognize that a module, a sub-module or a set of processing elements may be implemented as circuits, logic chips or any sort of discrete component. Still further, one skilled in the art will also recognize that a module, a sub-module or a set of processing elements may be implemented in software which may then be executed by a variety of processor architectures. In embodiments of the disclosure, a module, a sub-module or a set of processing elements may also comprise computer instructions, computations or executable code that may instruct a computer processor to carry out a sequence of events based on instructions received. The choice of the implementation of the modules, the sub-modules or the sets of processing elements is left as a design choice for a person skilled in the art and docs not limit the scope of the claimed subject matter in any way.

[0022] The method described in embodiments of the disclosure facilitates the transformation of the spectra for a sample or the same ingredient collected on different equipment by embedding the data into a high-dimensional space, thereby reducing inter- spcctromctcr variability. This method may be applied across similar equipment models as well as different brands and models, provided the data ranges overlap and the data formats are compatible. Additionally, this approach is able to accommodate the spectra collected through various ranges of the electromagnetic spectrum, including microwave, X-ray, ultraviolet (UV), visible (Vis), near-infrared (NIR), and mid-infrared (MIR).

[0023] In addition to the above, the disclosed method also accounts for different interaction mechanisms between electromagnetic waves and the sample, such as diffuse reflectance, specular reflectance, attenuated total reflectance, absorbance, transmission, fluorescence, andRaman scattering. A preferred embodiment involves the use of spectra in the visible and nearinfrared range (400-1700 nm) obtained through diffuse reflectance. Another preferred embodiment focuses on spectra in the 400-1100 nm range, also using diffuse reflectance as the primary interaction mechanism. It should be noted that this method may be applied to a wide range of samples, including but not limited to, food ingredients, agricultural produce, and intermediate products. Further, it is also suitable to be used with raw or processed agricultural products such as coffee beans, cocoa beans, cereals, legumes, tea leaves, tobacco leaves, spices, and herbs. The preferred samples include unroasted green coffee beans of any variety, as well as intact or powdered dry tea leaves of various types, including green, red, yellow, black, oolong, and white teas. Additionally, the method can be employed for intermediate food ingredients such as flour, starch, chocolate, and alcohol.

[0024] Figure 1 illustrates system 100 for evaluating the quality of subject sample 101 by aligning and analyzing spectra collected from spectrometer 102 and spectrometer 104. Spectrometer 102 generates spectra 103 for sample 101, while spectrometer 104 produces spectra 105 for the same sample 101. Spectra 103 and 105 are then provided to computing module 106 through wired and / or wireless means. Computing module 106 is then configured to process the received spectra to address any inter- spectrometer variability that may affect spectra 103 and 105. This ensures that any differences between the two datasets are actual representations of sample 101’s actual characteristics rather than discrepancies introduced by any one of spectrometers 102 or 104.

[0025] Within computing module 106, spectra 103 undergoes interpolation and an optimal preprocessing technique to create subject numerical vectors. These numerical vectors are then transformed into subject embeddings using a trained machine learning model provided within computing module 106. Spectra 105 is similarly interpolated and transformed using the trained machine learning model to generate subject embeddings. The detailed workings of the interpolation model, the preprocessing techniques and the trained machine learning model will be described in detail in the subsequent sections.

[0026] Computing module 106 then computes a similarity score for spectrometers 102 and 104 by comparing the embeddings associated with spectra 103 with the embeddings associated with spectra 105. In embodiments of the disclosure, the similarity score may be calculatedusing one of several distance metrics, such as, but arc not limited to Euclidean distance, Manhattan distance, Mahalanobis distance, or Cosine distance. The measurement of the distance metrics between the two embeddings quantifies the alignment of the spectra and highlights any genuine differences in sample 101’s spectral characteristics. Computing module 106 then utilizes the similarity score to assess whether the quality of the sample meets a predefined quality threshold.

[0027] In accordance with embodiments of the present disclosure, a block diagram representative of components of processing system 200 that may be provided within computing module 106 and the various modules contained therein to carry out the digital signal processing functions or computations in accordance with embodiments of the disclosure, or any other modules or sub-modules of the system is illustrated in Figure 2. One skilled in the art will recognize that the exact configuration of each processing system provided within these modules or sub-modules may be different and the exact configuration of processing system200 may vary and the arrangement illustrated in Figure 2 is provided by way of example only.

[0028] In embodiments of the disclosure, processing system 200 may comprise controller201 and user interface 202. User interface 202 is arranged to enable manual interactions between a user and the computing module as required and for this purpose includes the input / output components required for the user to enter instructions to provide updates to each of these modules. A person skilled in the ait will recognize that components of user interface202 may vary from embodiment to embodiment but will typically include one or more of display 240, keyboard 235 and optical device 236.

[0029] Controller 201 is in data communication with user interface 202 via bus 215 and includes memory 220, processing unit, processing element or processor 205 mounted on a circuit board that processes instructions and data for performing the method of this embodiment, an operating system 206, an input / output (I / O) interface 230 for communicating with user interface 202 and a communications interface, in this embodiment in the form of a network card 250. Network card 250 may, for example, be utilized to send data from these modules via a wired or wireless network to other processing devices or to receive data via the wired or wireless network. Wireless networks that may be utilized by network card 250 include, but are not limited to, Wireless -Fidelity (Wi-Fi), Bluetooth, Near Field Communication (NFC),cellular networks, satellite networks, telecommunication networks, Wide Area Networks (WAN) and etc.

[0030] Memory 220 and operating system 206 are in data communication with processor 205 via bus 210. The memory components include both volatile and non-volatile memory and more than one of each type of memory, including Random Access Memory (RAM) 223, Read Only Memory (ROM) 225 and a mass storage device 245, the last comprising one or more solid-state drives (SSDs). One skilled in the art will recognize that the memory components described above comprise non-transitory computer-readable media and shall be taken to comprise all computer-readable media except for a transitory, propagating signal. Typically, the instructions are stored as program code in the memory' components but can also be hardwired. Memory 220 may include a kernel and / or programming modules such as a software application that may be stored in either volatile or non-volatile memory.

[0031] Herein the term “processor” is used to refer generically to any device or component that can process such instructions and may include: a microprocessor, a processing unit, a plurality of processing elements, a microcontroller, a programmable logic device or any other type of computational device. That is, processor 205 may be provided by any suitable logic circuitry for receiving inputs, processing them in accordance with instructions stored in memory and generating outputs (for example to the memory components or on display 240). In this embodiment, processor 205 may be a single core or multi-core processor with memory addressable space. In one example, processor 205 may be multi-corc, comprising — for example — an 8 core CPU. In another example, it could be a cluster of CPU cores operating in parallel to accelerate computations.

[0032] Training the machine learning model

[0033] In embodiments of the disclosure, a machine learning model may be trained to transform numerical vectors into subject embeddings based on workflow 300 as illustrated in Figure 3. Initially, a first spectrometer and a second spectrometer may be used to measure and obtain raw training spectra associated with training samples. The measured raw spectra is illustrated as spectra 302 in Figure 3. Workflow 300 may then be applied to each spectra of spectra 302. For example, the training samples may comprise various different types of coffeebeans or any similar food product. One skilled in the art will recognize that the training samples are not limited to just these categories and may also comprise any agricultural and consumable goods that can produce measurable spectra. In embodiments of the disclosure, the first and second spectrometers may be configured to capture raw spectra from, but are not limited to, 192 distinct spectra, where each spectrum consists of 2048 reflectance values.

[0034] Spectra 302 is then provided to interpolation module 304. In embodiments of the disclosure, interpolation module 304 corrects and normalizes spectra 302 through the application of a spectral correction procedure. These corrections normalize spectra 302 by calibrating the spectra against predefined reference points. Following the normalization step, interpolation module 304 performs spectral alignment to address wavelength discrepancies between the two spectrometers. This may be done by interpolating spectra 302 to a common wavelength range. For example, while the candidate spectrometer may measure spectra at slightly offset wavelengths (e.g., 412.15 nm and 412.67 nm), the target spectrometer may possibly measure the spectra at different intervals (e.g., 412.33 nm and 412.89 nm). interpolation module 304 addresses this issue by applying linear interpolation to interpolate the spectra to a common wavelength range, e.g. 400-1100 nm. The detailed workings of interpolation module 304 is omitted for brevity as it is known to one skilled in the art.

[0035] The resulting interpolated spectra 306 is then provided to preprocessing module 308. Preprocessing module 308 attempts to enhance the quality of interpolated spectra 306 by reducing noise, correcting baseline shifts, and addressing scattering effects, to ensure that interpolated spectra 306 are suitable for downstream analysis. Preprocessing module 308 achieves this by applying a range of spectral preprocessing techniques to the interpolated spectra. Each preprocessing technique minimizes unwanted variability in the spectra while preserving meaningful sample differences. These techniques include, but are not limited to, standard normal variate (SNV), mean scatter correction (MSC), Savitzky-Golay filtering, and first or second derivatives of the spectra. Each of these techniques targets specific distortions in the spectra, such as multiplicative scatter effects, noise reduction, or feature enhancement, and after preprocessing of the spectra has been completed, the pre-processed spectra may then be represented as numerical vectors where each element of the numerical vectors corresponds to a processed value for a specific wavelength or feature of the spectra.

[0036] In embodiments of the disclosure, preprocessing module 308 systematically explores multiple preprocessing techniques by applying each technique to interpolated spectra 306 and generating corresponding prcproccsscd datasets. These datasets arc then evaluated for their effectiveness in improving the quality of the spectra and this evaluation may be performed by applying a Mean Average Precision at Rank (MAP@R) process to each of these datasets with the spectra measured by either the candidate or target spectrometers being used as the anchor spectra. The MAP@R process measures how well each preprocessing technique enables clustering of spectra from the same sample while maintaining separation from spectra of different samples. It should be noted that higher MAP@R scores indicate better alignment and feature preservation, ensuring that spectra of the same sample are grouped closely together in subsequent analysis. This systematic scoring process allows preprocessing module 308 to objectively assess and rank the performance of each preprocessing technique. The detailed workings of the MAP @R process is omitted for brevity as it is known to one skilled in the art.

[0037] Preprocessing module 308 then selects, based on the MAP@R scores of the various datasets, the optimal preprocessing technique that produces the best alignment and clustering performance for interpolated spectra 306. The selected technique is then applied uniformly across all spectra to ensure consistency in the preprocessing pipeline, thereby transforming interpolated spectra 306 into numerical vectors 310.

[0038] Numerical vectors 310 are then provided as the inputs for machine learning model 312. To recap, at this step, numerical vectors 310 would comprise numerical vectors that are associated with the spectra measured by the first and second spectrometers, i.e., the corresponding numerical vectors are associated with the first and second spectra respectively, where numerical vectors 310 may comprise, but is not limited to, between 50 and 200 features.

[0039] Machine learning model 312 receives numerical vectors 310 and generates embeddings 314 based on numerical vectors 310. These embeddings are representations of the spectra in a transformed, high-dimensional space, where the alignment of spectra is improved for samples from the same substance, irrespective of the spectrometer used. In embodiments of the disclosure, machine learning model 312 may comprise, but is not limited to, neural network architectures, such as fully connected or convolutional neural networks. During the initial training stage, machine learning model 312's parameters and hyperparameters such asthe learning rate, batch size, and margin parameter in the loss function, arc carefully initialized to optimize alignment accuracy. The machine learning model 312's parameters are initialized randomly or may comprise pre-trained values, and they arc subsequently iteratively updated through backpropagation to minimize the loss function.

[0040] Embeddings 314 arc then provided to decision module 315. Decision module 315 is configured to evaluate a correlation loss computed by correlation loss module 318 to decide whether machine learning model 312 should be further trained. If the loss value meets a predefined convergence criterion, such as falling below a threshold, the training process is terminated. Otherwise, decision module 315 initiates another iteration of embedding generation, metric learning, and loss computation. It should be noted that while decision module 315 does not compute the correlation loss directly, it relies on the loss values from correlation loss module 318 to make an informed decision about the progress of the training process.

[0041] If decision module 315 determines that the training process is to continue, embeddings 314 are then provided to metric learning module 316. Metric learning module 316 evaluates embeddings 314 by measuring the relationship between the embeddings using metric learning techniques such as, but are not limited to, Euclidean distance, Manhattan distance, Mahalanobis distance, or Cosine distance.

[0042] In embodiments of the disclosure, the metric learning technique may comprise a triplet loss function that is employed with Euclidean distance to learn a transformation z — / “(x), where the triplet loss function, £ is defined as:. equation (1) where x“, x“ are defined as the spectrum of the same sample measured by the target and candidate spectrometers respectively, xf is defined as the spectrum of a different sample measured by the target spectrometer, d(u, v) =11 u — v II 2 is defined as the Euclidean distance and a is defined as the margin parameter that enforces separation between positive and negative pairs. For simplicity, a single triplet is utilized resulting in:...equation (2)defined as the anchor-positive embedding and is defined as the anchor-negativeembedding.

[0043] To ensure efficient training of machine learning model 312, a triplet mining process is used to select triplets that violate the margin condition These triplets,referred to as hard triplets, contribute non-zero gradients to the loss, making them essential for updating the model parameters. Other triplet mining methods, such as hard positive, hard negative, and semi-hard mining, may also be employed to select samples based on varied conditions of margin violation. By focusing on these hard triplets, metric learning module 316 ensures that embeddings of spectra from the same sample may be brought closer together, while embeddings of different samples are pushed farther apart.

[0044] The triplets that violate the margin condition are forwarded to correlation loss module 318 for further processing. In these cases, the triplet loss is greater than zero (£ > 0) necessitating the computation of gradients to optimize the parameters of machine learning model 312. The gradients may be computed through partial derivative operations and the gradient computations for the anchor positive and negative embeddings may bederived as follows:Gradient with respect to the anchor embedding...equation (3)Gradient with respect to the positive embedding• • • equation (4)Gradient with respect to the negative embedding...equation (5)

[0045] These computed gradients arc then back propagated through the network to compute gradients with respect to the parameters of machine learning model 312. As each embedding is a function of machine learning model 312’s weights W,a chain rule may then be used to compute the gradient of the loss with regard to machine learning model 312’s weights W as:...equation (6)

[0046] It should be noted that Equation (6) is used to determine how the model weights should be adjusted to minimize the correlation loss. Hence, once the gradients have beencomputed, these gradients arc provided to parameter updating module 320.

[0047] Parameter updating module 320 then proceeds to update machine learning model 312’s parameters using an optimization algorithm such as an Adam optimizer or a Stochastic Gradient Descent (SGD) iterative process, which is defined as:... equation (7) where η is defined as the learning rate. The updated parameters are then used by machine learning model 312 to generate refined embeddings 314. With each iteration, machine learning model 312 progressively enhances the alignment and separation of embeddings, ensuring that spectra from the same sample remain closer while those from different samples are distinctly separated. When decision module 315 determines that the correlation loss value computed by correlation loss module 318 has met a predefined convergence criterion, decision module 315 will then terminate the training process and the latest set of parameters will be adopted by machine learning model 312. In certain trials, it was found that the correlation loss value arrives at a predefined convergence criterion after 10-30 epochs.

[0048] Figure 4 illustrates an exemplary implementation of triplet mining process 400 in accordance with embodiments of the present disclosure. The triplet mining process is a step in metric learning that involves the selection of triplets of embeddings. These triplets are then used for training the machine learning model 312 by ensuring that the model learns of the relationships between similar and dissimilar samples. Typically, each triplet comprises ananchor embedding a positive embedding (representing the same sample but measuredon a different device), and a negative embedding (representing a different sample). The goal is to ensure that the anchor-positive distance is smaller than the anchor-negative distanceby at least a predefined margin a. The triplet mining process identifies hard triplets which violates this margin condition as these hard triplets contribute non-zerogradients during loss computation.

[0049] To form the triplets, triplet mining process 400 first identifies similar pairs and dissimilar pairs from the embeddings. A similar pair 402 is formed by embeddings of the same sample measured on different spectrometers, such as Lot A which was measured using a candidate spectrometer and Lot A which was measured using a target spectrometer. A dissimilar pair 404 is formed by embeddings of different samples which were measured using the same spectrometer, such as Lot A which was measured using the target spectrometer and Lot B which was measured using the target spectrometer. In such a situation, under the assumption that Lot A (e.g., coffee beans) was scanned using two different spectrometers, the embeddings from these scans are expected to be close in the embedding space and as such would form a similar pair. Meanwhile, if Lot A (e.g., coffee beans) and Lot B (e.g., tea leaves), were to be scanned using the same spectrometer, it would be expected that the resulting embeddings would be located far apart thereby forming a dissimilar pair.

[0050] Once similar 402 and dissimilar 404 pairs have been identified, triplet mining process 400 then computes at step 406, the distances between the embeddings for similar' 402 and dissimilar 404 pairs using a metric such as Euclidean distance. For example, an anchorpositive distance is computed where this distance daprepresents the distance between theembedding of Lot A as measured by the candidate spectrometer and the embedding of Lot A as measured by the target spectrometer. An anchor-negative distanceis computed where this distance danrepresents the distance between the embedding of Lot A as measured by the target spectrometer and the embedding of Lot B as measured by the target spectrometer. The difference between this distance, i.c., , is then compared against a margin α.

[0051] If triplet mining process 400 determines at step 410 that the difference between this distance exceeds the margin a, process 400 then uses these corresponding triplets for further training. These triplets are passed to the correlation loss module (i.e., correlation loss module318 in Figure 3), where the triplet loss is subsequently calculated. By focusing on hard triplets, the machine learning model 312 learns to minimize the anchor-positive distance while maximizing the anchor-negative distance, improving the separation between similar and dissimilar samples in the embedding space. For example, embeddings of Lot A measured on different spectrometers will be pulled closer together, while embeddings of Lot A and Lot B are pushed further apart.

[0052] A process for aligning measurements performed by a candidate spectrometer with measurements performed by a target spectrometer is illustrated in Figure 5 whereby process 500 may be carried out by a computing module that is communicatively coupled to a target and a candidate spectrometer in accordance with embodiments of the disclosure.

[0053] Process 500 begins at step 502 with the computing module receiving first spectra associated with a sample measured by the first spectrometer, wherein the first spectra comprises a plurality of peaks and valleys measured at unevenly spaced wavelengths. At step 504, process 500 then proceeds to apply an interpolation process to the first spectra to align the first spectra to a common wavelength range. Process 500 then preprocesses the interpolated first spectra by selecting and applying an optimal preprocessing technique, thereby encoding the data into first numerical vectors. This takes place at step 506. At step 508, process 500 then generates aligned first embeddings by applying a trained machine learning model to the first numerical vectors. In embodiments of the disclosure, the trained machine learning model is trained by receiving spectra associated with each of the samples in the set of samples as measured by the second and third spectrometers. For each of the received spectra, the interpolation process is applied to the spectra measured by the second and third spectrometers to generate interpolated second and third spectra, and by applying the optimal preprocessing technique to the interpolated second and third spectra to encode the interpolated second and third spectra into corresponding second and third numerical vectors. Once this is done, the second numerical vectors are correlated with the third numerical vectors by applying a metric learning technique to second and third embeddings generated by the machine learning model for the corresponding second and third numerical vectors. The one or more parameters of the machine learning model are then adjusted to minimize a correlation loss between the second and third embeddings. At least the correlating step and the adjusting step are iteratively repeated until it is determined that the correlation loss between the second and thirdembeddings is less than a predetermined loss threshold. When this happens, the parameters of the trained machine learning model are set based on a final adjusted one or more parameters of the machine learning model.

[0054] In accordance with embodiments of the disclosure, process 500 selects and applies the optimal preprocessing technique by applying a plurality of preprocessing techniques to the interpolated first spectra, wherein each preprocessing technique modifies the interpolated first spectra to enhance signal quality and remove baseline noise, by evaluating the prcproccsscd first spectra using a predefined metric to determine effectiveness of each of the preprocessing techniques, by selecting the optimal preprocessing technique based on the evaluation, and by applying the selected preprocessing technique to the interpolated first spectra.

[0055] In accordance with embodiments of the disclosure, the plurality of preprocessing techniques comprise spectral correction, Standard Normal Variate, mean scatter correction, Savitzky-Golay filter and first or second derivatives of the interpolated candidate spectra while the predefined metric comprises a Mean Average Precision at Rank R (MAP@R) metric.

[0056] In accordance with embodiments of the disclosure, the quality of a subject sample may be determined by a computing device generating, using process 500, aligned subject first embeddings for the subject sample. The computing device may then receive a reference spectra associated with the subject sample measured by a reference spectrometer and then may apply the interpolation process to the reference subject spectra, and the optimal preprocessing technique to the interpolated reference subject spectra to encode the interpolated reference subject spectra into reference subject numerical vectors. The computing device then generates reference subject embeddings for the reference subject numerical vectors using the trained machine learning model and computes a similarity score based on a Euclidean distance, a Manhattan distance, a Mahalanobis distance or a Cosine distance between the aligned subject first embeddings and the reference subject embeddings. The computing device then determines based on the similarity score if the quality of the subject sample meets a predetermined quality threshold.

[0057] Figure 6a illustrates the clusters of various coffee types, where spectra measured by spectrometers labeled as D1, D2, D3 and D4 are compared with the corresponding embeddingsgenerated by a trained machine learning model. Specifically, spectra plot 602 illustrates the raw or preprocessed spectral data of different coffee samples as measured by spectrometers D1, D2, D3 and D4. Each symbol corresponds to a specific coffee sample (c.g., Coffee 1, Coffee 2, Coffee 3, Coffee 4, Coffee 5), with different shapes representing distinct coffee lots. Spectral plot 602 illustrates that when different spectrometers are used to measure a similar type of sample, e.g., Coffee 01, each of the spectrometers will produce different results. As a result, spectral plot 602 may show clusters where different types of samples have been incorrectly grouped together. For example, cluster 603 shows how the spectral measurements for Coffee 1, Coffee 2, Coffee 3, Coffee 4 and Coffee 5 as measured by spectrometer D2 have all been incorrectly clustered together when instead, all measurements for the same type of coffee, e.g., Coffee 1 , should have been clustered together.

[0058] Embedding plot 604 illustrates the transformed embeddings of the same spectral data after the spectral data has been processed by the trained machine learning model. The embeddings provide a refined representation of the data, where samples of the same coffee type, e.g., Coffee 4 are grouped closer together, such as cluster 605, even though the measurements were done using different spectrometers, c.g. spectrometers D1, D2, D3 and D4. Furthermore, it is shown that the samples of different coffee types are more distinctly separated into their individual distinct clusters.

[0059] Figure 6b illustrates the clusters of various spectrometers, where spectra measured by the various spectrometers for various types of samples, c.g., different types of coffee such as C1 , C2, C3, C4 and C5, are compared with the corresponding embeddings generated by a trained machine learning model. Spectra plot 606 illustrates the raw or preprocessed spectral data of various different samples C1 , C2, C3, C4 and C5 as measured by the different spectrometers (e.g., Device 01, Device 02, Device 03, Device 04). Spectral plot 606 illustrates that when a particular spectrometer, e.g., Device 01, is used to measure different types of samples, e.g. C1-C5, the spectrometer may end up incorrectly clustering all the different samples together. As a result, spectral plot 606 may show clusters where different types of samples have been incorrectly grouped together. For example, cluster 607 of measurements by Device 02 for Coffee C1 , C2, C3, C4 and C5 show how these measurements have been incorrectly clustered together when instead, these measurements should have been clustered by coffee type, regardless of the spectrometer used. In summary, spectra plot 606 showsinconsistencies between the measurements from different devices, highlighting inter-device variability.

[0060] Embedding plot 608 illustrates the transformed embeddings of the spectral data in spectral plot 606 after the spectral data has been processed by the trained machine learning model, where the same coffee types are properly clustered together, regardless of the spectrometer (or device) that performed the measurements as illustrated by cluster 609. In particular, cluster 609 shows how embeddings related to Coffee C4 arc clustered together even though the measurements were carried out using different spectrometers, i.e., Device 01, Device 02, Device 03 and Device 04. The embeddings demonstrate a marked improvement in alignment, with data points from the same sample grouping closely together, regardless of the device used for measurement.

[0061] Numerous other changes, substitutions, variations, and modifications may be ascertained by the skilled in the art and it is intended that the present application encompass all such changes, substitutions, variations, and modifications as falling within the scope of the appended claims.

Claims

CLAIMS:

1. A method for aligning measurements performed by a first spectrometer with measurements performed by a second and a third spectrometer, the method comprising: receiving, using a computing device, first spectra associated with a sample as measured by the first spectrometer, wherein the first spectra comprises a plurality of peaks and valleys measured at unevenly spaced wavelengths; applying, using the computing device, an interpolation process to the first spectra to align the first spectra to a common wavelength range; preprocessing, using the computing device, the interpolated first spectra by selecting and applying an optimal preprocessing technique, thereby encoding the data into first numerical vectors; generating, using the computing device, aligned first embeddings by applying a trained machine learning model to the first numerical vectors, wherein the trained machine learning model is trained, using the computing device, by: receiving spectra associated with a training set of samples as measured by the second spectrometer and the third spectrometer; for each of the received spectra, applying the interpolation process to the spectra measured by the second and third spectrometers to generate interpolated second and third spectra, and applying the optimal preprocessing technique to the interpolated second and third spectra to encode the interpolated second and third spectra into corresponding second and third numerical vectors; correlating the second numerical vectors with the third numerical vectors by applying a metric learning technique to second and third embeddings generated by the machine learning model for the corresponding second and third numerical vectors; adjusting one or more parameters of the machine learning model to minimize a correlation loss between the second and third embeddings; repeating, using the computing device, at least the correlating step and the adjusting step until it is determined that the correlation loss between the second and third embeddings is less than a predetermined loss threshold;setting, using the computing device, the parameters of the trained machine learning model based on a final adjusted one or more parameters of the machine learning model.

2. The method according to claim 1, wherein the step of selecting and applying the optimal preprocessing technique comprises: applying a plurality of preprocessing techniques to the interpolated first spectra, wherein each preprocessing technique modifies the interpolated first spectra to enhance signal quality and remove baseline noise; evaluating the preprocessed first spectra using a predefined metric to determine effectiveness of each of the preprocessing techniques; selecting the optimal preprocessing technique based on the evaluation; and applying the selected preprocessing technique to the interpolated first spectra.

3. The method according to claim 2, wherein the plurality of preprocessing techniques comprise spectral correction, Standard Normal Variate, mean scatter correction, Savitzky- Golay filter and first or second derivatives of the interpolated candidate spectra.

4. The method according to claim 1, wherein the metric learning technique comprises a triplet mining process.

5. The method according to claim 1, wherein the machine learning model comprises a onedimensional Convolutional Neural Network (1D-CNN).

6. The method according to claim 1, wherein the adjusting the one or more parameters of the machine learning model comprises the step of applying a Stochastic Gradient Descent optimization algorithm to update the one of more parameters to minimize the correlation loss between the second and third embeddings.

7. A method for determining quality of a subject sample, the method comprising: generating, using the method according to claim 1, aligned subject first embeddings for the subject sample; receiving, using the computing device, a reference spectra associated with the subject sample measured by a reference spectrometer;applying, using the computing device, the interpolation process to the reference subject spectra, and the optimal preprocessing technique to the interpolated reference subject spectra to encode the interpolated reference subj ect spectra into reference subj ect numerical vectors; generating, using the computing device, reference subject embeddings for the reference subject numerical vectors using the trained machine learning model; computing, using the computing device, a similarity score based on a Euclidean distance, a Manhattan distance, a Mahalanobis distance or a Cosine distance between the aligned subject first embeddings and the reference subject embeddings; and determining based on the similarity score if the quality of the subject sample meets a predetermined quality threshold.

8. A computing module for aligning measurements performed by a candidate spectrometer with measurements performed by a target spectrometer, the computing module comprising: a processing unit; and a non-transitory media readable by the processing unit, the media storing instructions that when executed by the processing unit causes the processing unit to: receive first spectra associated with a sample as measured by the first spectrometer, wherein the first spectra comprises a plurality of peaks and valleys measured at unevenly spaced wavelengths; apply an interpolation process to the first spectra to align the first spectra to a common wavelength range; preprocess the interpolated first spectra by selecting and applying an optimal preprocessing technique, thereby encoding the data into first numerical vectors; generate aligned first embeddings by applying a trained machine learning model to the first numerical vectors, wherein the trained machine learning model is trained by: receiving spectra associated with a training set of samples as measured by the second spectrometer and the third spectrometer; for each of the received spectra, applying the interpolation process to the spectra measured by the second and third spectrometers to generate interpolated second and third spectra, and applying the optimal preprocessing technique to the interpolated second andthird spectra to encode the interpolated second and third spectra into corresponding second and third numerical vectors; correlating the second numerical vectors with the third numerical vectors by applying a metric learning technique to second and third embeddings generated by the machine learning model for the corresponding second and third numerical vectors; adjusting one or more parameters of the machine learning model to minimize a correlation loss between the second and third embeddings; repeating at least the correlating step and the adjusting step until it is determined that the correlation loss between the second and third embeddings is less than a predetermined loss threshold; setting the parameters of the trained machine learning model based on a final adjusted one or more parameters of the machine learning model.

9. The computing module according to claim 8, wherein the instructions to instruct the processing unit to select and apply the optimal preprocessing technique comprises instructions for directing the processing unit to: apply a plurality of preprocessing techniques to the interpolated first spectra, wherein each preprocessing technique modifies the interpolated first spectra to enhance signal quality and remove baseline noise; evaluate the preprocessed first spectra using a predefined metric to determine effectiveness of each of the preprocessing techniques; select the optimal preprocessing technique based on the evaluation; and apply the selected preprocessing technique to the interpolated first spectra.

10. The computing module according to claim 9, wherein the plurality of preprocessing techniques comprise spectral correction, Standard Normal Variate, mean scatter correction, Savitzky-Golay filter and first or second derivatives of the interpolated candidate spectra.

11. The computing module according to claim 8, wherein the metric learning technique comprises a triplet mining process.

12. The computing module according to claim 8, wherein the machine learning model comprises a one-dimensional Convolutional Neural Network (1D-CNN).

13. The computing module according to claim 8, wherein the instructions to direct the processing unit to adjust the one or more parameters of the machine learning model comprises instructions for directing the processing unit to apply a Stochastic Gradient Descent optimization algorithm to update the one of more parameters to minimize the correlation loss between the second and third embeddings.

14. A computing system for determining quality of a subject sample, the system comprising: the computing module according to claim 8 being configured to generate aligned subject first embeddings for the subject sample; the computing module according to claim 8 being configured to receive a reference spectra associated with the subject sample measured by a reference spectrometer; the computing module according to claim 8 being configured to apply the interpolation process to the reference subject spectra, and the optimal preprocessing technique to the interpolated reference subject spectra to encode the interpolated reference subject spectra into reference subject numerical vectors; the computing module according to claim 8 being configured to generate reference subject embeddings for the reference subject numerical vectors using the trained machine learning model; the computing module according to claim 8 being configured to compute a similarity score based on a Euclidean distance, a Manhattan distance, a Mahalanobis distance or a Cosine distance between the aligned subject first embeddings and the reference subject embeddings; and the computing module according to claim 8 being configured to determine based on the similarity score if the quality of the subject sample meets a predetermined quality threshold.