Generation of chemometric models

The method addresses the inefficiencies in generating chemometric models by using a computer-implemented approach that reduces data requirements and expert input, resulting in fast and accurate model generation for spectrometers.

WO2025125102A1PCT designated stage expired Publication Date: 2025-06-19TRINAMIX GMBH

Patent Information

Application Number
PCT/EP2024/085007
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-12
Filing Date
2024-12-06
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing methods for generating chemometric models for spectrometers require large training datasets, are time-consuming, and often need expert input, making them inefficient and inaccessible to non-professionals.

Method used

A computer-implemented method that receives spectra and object data, augments the data, generates multiple candidate chemometric models, trains them using the augmented data, determines their prediction accuracy, and selects the most accurate model for output.

Benefits of technology

This method allows for the generation of high-accuracy chemometric models with reduced training data requirements, is fast, and can be used by non-experts, facilitating quick updates and broad applicability in various use cases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000026_0000
    Figure 00000026_0000
  • Figure 00000026_0001
    Figure 00000026_0001
  • Figure 00000027_0000
    Figure 00000027_0000
Patent Text Reader

Abstract

The invention is in the field of generation of chemometric models for spectrometers. It relates to a computer- implemented method for generating a chemometric model for a spectrometer comprising: a) receiving spectra of an object and object data associated with a physical or chemical characteristic of the object, b) augmenting the provided spectra and the associated object data, c) generating at least two different candidate chemometric models comprising d) training the candidate chemometric models using the augmented spectra and the object data, e) determining the prediction accuracy of the trained candidate chemometric models, f) selecting a chemometric model of the trained candidate chemometric models using the prediction accuracy and g) outputting the selected chemometric model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Generation of Chemometric Models

[0002] The invention is in the field of generation of chemometric models for spectrometers. The invention relates to a computer-implemented method for generating a chemometric model for a spectrometer, the use of chemometric model obtained from the method for a spectrometer, a spectrometer configured to use the chemometric model obtained from the method, a system for generating a chemometric model for a spectrometer and a non-transient computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform the method.

[0003] Background

[0004] Spectroscopy is a valuable non-destructive analytical technique for various applications, like quality control, quantitative and qualitative examination of agricultural products. A spectrum contains physical or chemical information about the recorded object but cannot be directly interpreted by its nature. For this purpose, a chemometric model is required which derives a physical or chemical characteristic of the measured object from its spectrum. To create a chemometric model, spectra of objects for which the characteristic of interest is known, for example from a laboratory analysis, are required. Such analyses are often time-consuming and sometimes not sufficiently many objects are available to perform such analyses. However, the higher the requirements for the chemometric model, the more such data is necessary. There is a need to facilitate the generation of chemometric models.

[0005] US 2023 / 0009725 A1 discloses a genetic algorithm to identify a processing pipeline that transforms spectra into a form usable to generate predicted characteristics of corresponding samples. However, the training data is used as received. Hence, an appropriate chemometric model can only be found if sufficient training data is available. Summary

[0006] The objective of the present invention was to provide a method for generating a chemometric model for a spectrometer with high accuracy and reliability without the need for large training datasets. The method should be fast and easy to be used by non-professionals. The method should be applicable for a broad variety of use cases without sophisticated adjustments.

[0007] In one aspect the invention relates to a computer-implemented method for generating a chemometric model for a spectrometer comprising: a) receiving spectra of an object and object data associated with a physical or chemical characteristic of the object, b) augmenting the provided spectra and the associated object data, c) generating at least two different candidate chemometric models, d) training the candidate chemometric models using the augmented spectra and the object data, e) determining the prediction accuracy of the trained candidate chemometric models, f) selecting a chemometric model of the trained candidate chemometric models using the prediction accuracy and g) outputting the selected chemometric model.

[0008] In another aspect the invention relates to a computer-implemented method for generating a chemometric model for a spectrometer comprising: a) receiving spectra of an object and object data associated with a physical or chemical characteristic of the object, b) generating at least two different candidate chemometric models, wherein a candidate chemometric model is configured to determine at least two physical or chemical characteristics from a spectrum, c) training the candidate chemometric models using the spectra and the object data, d) determining the prediction accuracy of the trained candidate chemometric models, e) selecting a chemometric model of the trained candidate chemometric models using the prediction accuracy and f) outputting the selected chemometric model.

[0009] In another aspect the present invention relates to a use of the chemometric model obtained from the method of any of the previous claims for a spectrometer.

[0010] In another aspect the present invention relates to a spectrometer configured to use the chemometric model obtained from the method according to the present invention.

[0011] In another aspect the present invention relates to a system for generating a chemometric model for a spectrometer comprising: a) an input for receiving spectra of an object and object data associated with a physical or chemical characteristic of the object, b) a processor for augmenting the provided spectra and the associated object data, generating at least two different candidate chemometric models, training the candidate chemometric models, determining the prediction accuracy of the trained candidate chemometric models using the augmented spectra and the object data, and selecting a chemometric model of the trained candidate chemometric models using the prediction accuracy and c) an output for outputting the selected chemometric model.

[0012] In another aspect the invention relates to a system for generating a chemometric model for a spectrometer comprising: a) an input for receiving spectra of an object and object data associated with a physical or chemical characteristic of the object, b) a processor for generating at least two different candidate chemometric models, wherein a candidate chemometric model is configured to determine at least two physical or chemical characteristics from a spectrum, training the candidate chemometric models using the spectra and the object data, determining the prediction accuracy of the trained candidate chemometric models, and selecting a chemometric model of the trained candidate chemometric models using the prediction accuracy and c) an output for outputting the selected trained chemometric model.

[0013] In another aspect the present invention relates to a non-transient computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform a method comprising: a) receiving spectra of an object and object data associated with a physical or chemical characteristic of the object, b) augmenting the provided spectra and the associated object data, c) generating at least two different candidate chemometric models, d) training the candidate chemometric models using the augmented spectra and the object data, e) determining the prediction accuracy of the trained candidate chemometric models, f) selecting a chemometric model of the trained candidate chemometric models using the prediction accuracy and g) outputting the selected chemometric model.

[0014] In another aspect the invention relates to a non-transient computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform a method comprising: a) receiving spectra of an object and object data associated with a physical or chemical characteristic of the object, b) generating at least two different candidate chemometric models, wherein a candidate chemometric model is configured to determine at least two physical or chemical characteristics from a spectrum, c) training the candidate chemometric models using the spectra and the object data, d) determining the prediction accuracy of the trained candidate chemometric models, e) selecting a chemometric model of the trained candidate chemometric models using the prediction accuracy and f) outputting the selected chemometric model.

[0015] The generation of chemometric models of the present invention can be automized, for example in a computer system. The required amount of training data is significantly lowered, so chemometric models of high accuracy with low amount of training data are accessible. No expert input is required. The automation makes the method very fast and easy to use. No experienced and skilled personal is required, different user will achieve the same result. The obtained chemometric model is reliably accurate. The method further enables a quick update of the model with new reference data for which different settings for the chemometric model selection may yield better results. The method can be implemented in an easy-to-use frontend in which the user only has to upload the reference spectra and associated object data, for example analytical data, and can receive the most accurate chemometric model with a short amount of time. The method allows the quick employment of spectroscopy for new use cases.

[0016] The term "spectrometer” may refer to an instrument which is capable of recording spectra of an object. The spectrometer may be portable or stationary, for example a laboratory device. A portable spectrometer may be a hand-held spectrometer or a module which is integrated into a portable device like a smartphone, a tablet or a wearable like a smartwatch. A portable spectrometer may be communicatively coupled to a computer device, for example a cloud computer or a smartphone. Such computer device may be configured to execute a chemometric model, for example the chemometric model obtained by the method of the invention. The computer device may further be configured to receive spectra from the spectrometer. The computer device may store such spectra, send it to a system for generating a chemometric model according to the present invention or execute the method of the present invention using the received and / or stored spectra. The computer system may further be configured to use the output of a chemometric model to derive instructions for action, for example to drink water if the skin was found to contain a low water content.

[0017] A spectrometer may comprise:

[0018] - an optical element configured for separating incident optical radiation provided by the measurement object into a spectrum of constituent wavelength components;

[0019] - a photosensor comprising at least one photosensitive region configured for receiving the optical radiation from the optical element, wherein the photosensor is configured for generating at least one photosensor signal dependent on an illumination of the photosensitive region by the optical radiation;

[0020] - a processor to process the photosensor signals into a spectrum.

[0021] The term "optical element” may refer to an arbitrary element configured for influencing optical radiation. The optical element may be configured for at least one of at least partially dispersing the optical radiation, at least partially filtering the optical radiation, at least partially reflecting the optical radiation, e.g. diffusely or directly, at least partially deflecting the optical radiation, at least partially transmitting the optical radiation and at least partially absorbing the optical radiation. The optical element may comprise at least one of a prism, a grating, a beam splitter, or an interferometer, for example a Michelson interferometer. The optical element may be configured for being used in mobile applications, for example for being used in handheld spectrometer devices and / or in spectrometer devices comprised by electronic communication devices, such as a smartphone or a tablet. As another example, the optical element may comprise at least one optical filter element. The optical filter element may be configured for filtering the optical radiation or more specifically at least one selected spectral range of the optical radiation. The optical filter element may be positioned in a radiation path before the photosensor. As an example, the portable spectrometer may comprise a plurality of a photosensors, for example 5 to 20, such as 8 to 12. The photosensors may be arranged as pixels in an array or in a matrix. The portable spectrometer may comprise a plurality of optical filter elements. An optical filter element may be positioned in a beam path before a photosensor. The optical filter elements may be transmissive at different wavelengths of different wavelength regions. For example, each photosensor may be positioned behind an optical filter with regard to the beam path, wherein each optical filter is transmissive at different wavelength or different wavelength region to the other optical filters.

[0022] The spectrometer may comprise one or more than one photosensor. The photosensor may comprise at least one photosensitive region. The photosensitive region may be configured for receiving the optical radiation from the optical element. The photosensor may be configured for generating at least one photosensor signal dependent on an illumination of the photosensitive region by the optical radiation. The term "sensor” may refer to a device configured for detecting at least one condition or for measuring at least one measurement variable. The sensor may be capable of generating at least one signal, such as a measurement signal, which is a qualitative or quantitative indication of the measurement variable and / or measurement property, e.g. of an illumination of the sensor or a part of the sensor. The signal may be or comprise an electrical signal, such as a current, specifically a photocurrent. The term "photosensor” may refer to a sensor or a detector configured for detecting or measuring optical radiation, such as for detecting an illumination and / or a radiation spot generated by at least one radiation beam, e.g. by using the photoelectric effect. The photodetector may comprise at least one substrate. As an example, a single photosensor may be a substrate with at least one single photosensitive region, which generates a physical response, e.g. an electronic response, to the illumination for a given wavelength range.

[0023] The term "photosensitive region” may refer to a unit of the photosensor, specifically to a spatial area or volume being part of the photosensor, configured for being illuminated, or in other words for receiving optical radiation, and for generating at least one signal, such as an electronic signal, in response to the illumination. The photosensitive region may be located on a surface of the photosensor. The photosensitive region may specifically be a single, closed, uniform photosensitive region. However, other options may also be feasible.

[0024] The spectrometer may comprise a radiation or light emitting element configured for emitting illumination light for illuminating the object in order to generate detection light from the object. The light emitting element may be an incandescent lamp, for example a tungsten filament lamp or a tungsten halogen lamp, a light-emitting diode (LED), a laser diode, a gas-discharge lamp, for example a xenon lamp, a mercury vapor lamp, or a deuterium lamp.

[0025] The term "light” or "radiation” may refer to electromagnetic radiation in one or more of the infrared, the visible and the ultraviolet spectral range. The term "ultraviolet spectral range” may refer to electromagnetic radiation having a wavelength of 1 nm to 380 nm, preferably of 100 nm to 380 nm. Further, in partial accordance with standard ISO- 21348 in a valid version at the date of this document, the term "visible spectral range” may refer to a spectral range of 380 nm to 760 nm. The term "spectral range” (IR) may refer to electromagnetic radiation of 760 nm to 1000 m, wherein the range of 760 nm to 1 .5 m is usually denominated as "near spectral range” (NIR) while the range from 1.5 pi to 15 m is denoted as "mid spectral range” (MidlR) and the range from 15 pm to 1000 pm as "far spectral range” (FIR). Preferably, light used for the typical purposes of the present invention is light in the infrared (IR) spectral range, more preferred, in the near infrared (NIR) and / or the mid spectral range (MidlR), especially the light having a wavelength of 1 pm to 5 pm, preferably of 1 pm to 3 pm, as these wavelength regions are particularly suitable for obtaining material properties of an object.

[0026] The spectrometer may comprise a processor to process the photosensor signals into a spectrum. The processor may output the spectrum, for example to an interface for further processing or to a user interface. The processor may further be configured to apply a chemometric model and output the object data obtained by the chemometric model. The processor may be configured to output both the spectrum and the object data. The spectrometer may further comprise a memory. The memory may be configured to store the chemometric model. The memory may be configured to store spectra.

[0027] The term "spectrum” may refer to a data structure in which several intensity values or values derived thereof such as absorbance or reflectance of radiation are associated with wavelengths or wavelength ranges of the radiation. The data structure may be a vector or matrix, wherein each element represents an intensity and the position in the vector or matrix represents a certain wavelength or wavelength range, so the value at a certain position represents the intensity of that wavelength or wavelength range. The data structure may be a vector or matrix containing value pairs, wherein one value represents the wavelength or wavelength range and the other value the intensity at this wavelength or wavelength range. A spectrum may be recorded by a spectrometer. The spectrum recorded by the spectrometer may be corrected by calibration coefficients to compensate for sensor imperfections or drifts. The spectrum may represent the absorbance or transmittance of radiation after having penetrated an object.

[0028] The term "chemometric model” may refer to a model which uses spectra as input and outputs object data. Hence, the chemometric model may translate a spectrum of an object into corresponding object data. The object data obtained from the chemometric model may be associated with one or more than one physical or chemical characteristics, for example at least two or at least three physical or chemical characteristics. A chemometric model which outputs more than one physical or chemical characteristic may be referred to as multilabel chemometric model. A multilabel chemometric model may contain several partial chemometric models, wherein each partial chemometric model may be configured to translate a spectrum to one physical or chemical characteristic. The multilabel chemometric model may be configured to merge the object data obtained from the at least two partial chemometric models, for example to output one vector comprising values for each physical or chemical characteristic.

[0029] A chemometric model may comprise a pre-processing method and a machine learning model to obtain the object data. A chemometric model may comprise a pre-processing method, a feature selection filter and a machine learning model. If the chemometric model comprises two or more partial chemometric models, each partial chemometric model may comprise a separate pre-processing method, a feature selection filter and a machine learning model. Alternatively, the partial models may use the same pre-processing method or feature selection filter.

[0030] A chemometric model may be configured to determine fitness and health information from a spectrum of the skin of a person, for example the hydration level of the skin or the blood glucose or lactose content of the skin. The result may be used to make recommendations to the user, for example to drink water according to the hydration level of the skin or adjust the training program according to the lactose level in the blood.

[0031] A chemometric model may be configured to analyze agricultural products, food or feed from a spectrum of an agricultural product, for example a fruit, vegetable, meat, dairy product; food, for example bread, sauces, sausages, sweets; feed like fodder or forage. For example the chemometric model may determine a relevant content, for example the sugar or protein content. The result may be used to make recommendations, for example the expected date for harvesting to a farmer or a suitable recipe for a cook.

[0032] A chemometric model may be configured to determine information related to recycling from a spectrum of an object to be recycled, for example a plastic object like a bottle or a metal object like a can. For example, it may be determined which material the object is made of, for example polyethyleneterephthalate (PET) for a bottle. The result may be used to make recommendations, for example the most suitable way for recycling the object or the most convenient place put the object so it can be recycled.

[0033] A chemometric model may be configured to determine product quality in a production process. Determining product quality may mean determining parameters of a specification, for example the concentration of an ingredient. Product quality may refer to the quality of the input material, for example as received from a supplier, of an intermediate, i.e. a material which as undergone some processing steps and will undergo further processing steps, and an output material, for example before it is packaged and shipped to a customer.

[0034] The term "object” may refer to any object which can be measured by a spectrometer. An object can be a living body or a non-living object. An object may refer to a complete object or a small piece thereof, for example a sample extracted from the object. The object may have any aggregate state, e.g. solid, liquid, gaseous, or supercritical. The object may be homogeneous, i.e. it does not contain internal phase boundaries between domains of the size of the illuminated radiation or larger, for example a plastic bottle or a blood sample. The object may be heterogeneous, it contains internal phase boundaries between domains of the size of illuminated radiation or larger, for example a soil sample containing sand.

[0035] The term "object data” may refer to data associated with a physical or chemical characteristic of the object. Object data may hence contain parameters representing physical or chemical characteristic of the object. Physical characteristics may comprise thermal characteristics, for example the temperature, the thermal conductivity or the specific heat capacity; the macroscopic or microscopic state of matter, for example the degree of crystallinity; optical characteristics, for example refractive index, optical conductivity or absorption coefficients; electrical characteristics, for example electrical conductivity or dielectric constant. Chemical characteristics of an object typically refer to the chemical composition of the object comprising the material composition, for example the type and the concentration of certain chemical compounds such as the water content; the molecular structure of an object, for example the presence of functional groups like carbonyl groups or hydrogen bonds; the molecular weight, for example the degree of polymerization of a polymer; the state of matter, for example the degree of crystallinity or the aggregate state, for example in a disperse phase. Object data can comprise one physical or chemical characteristic of the object or more than one, for example the concentration of two chemical compounds in a sample. A physical or chemical characteristic can be variable, it can assume any value in a certain range, for example a concentration, or it can be categorial, for example the type of fiber in a fabric.

[0036] Object data may further contain metadata, wherein metadata may be associated with factors which may have an influence on the physical or chemical characteristic of the object but are not the physical or chemical characteristic themselves. Metadata may be associated with factors related to the object, the spectrometer, or the measurement environment. Examples of factors related to the object may be a sample identifier, its geographic origin, a time information such as the time a sample was taken from a larger object. Examples of factors related to the spectrometer may be a spectrometer identifier, the type of spectrometer or spectrometer settings. Examples of factors related to the measurement environment may be a time information such as the time the measurement was performed, a geographic information related to where the measurement was taken or information about the analytical technique used to determine the physical or chemical characteristic of the object.

[0037] The method of the present invention comprises receiving spectra and object data. Receiving may mean obtaining the data through an interface, for example an interface to a storage device or a communication interface, for example to a cloud computing system. The spectra and object data may also be received through an interface to a user interface, for example a graphical user interface. For example, spectra and object data may be uploaded by a user through a user interface, for example by using a web browser. The spectra and object data may be received directly, for example via a communication interface, for example via the internet. Alternatively, the spectra and object data may be stored on a storage device after upload through the user interface from where the data is obtained. In this way, a user can first collect all spectra and object data over time and can then indicate completeness, for example by clicking a button labelled "completed” or "generate chemometric model”. Such indication may trigger the method of the present invention.

[0038] The received spectra of an object and object data associated with a physical or chemical characteristic of the object are associated to each other, i.e. a spectrum and corresponding object data may be linked. For example, both the spectrum and the object data may comprise an identifier of the object the spectrum is taken from and the object data relates to. In this way, it can be made sure that each element of the object data can be associated to the corresponding spectrum.

[0039] The object data to be determined by the chemometric model may be determined from the provided object data. If the object data contains only values associated with one physical or chemical characteristic, it can be easily determined that the chemometric model shall determine this object data. If the object data contains values associated with more than one physical or chemical characteristics, the determination may involve a predefined condition, for example the first characteristic appearing in the object data. Alternatively, the object data may comprise an identifier indicating which physical or chemical characteristic shall be determined by the chemometric model. The identifier may further indicate if the physical or chemical characteristic is associated with a variable value, for example for a concentration a decimal number value, or a categorial value, for example for distinguishing between a given set of compounds, for example plastic types like PET or cotton in cloths.

[0040] The method of the present invention may comprise design of experiments (DOE) for measurements directed to obtain spectra and object data received by the method of the present invention. The term "design of experiments” may refer to instructions for systematically measuring spectra on a set of objects with specified object data. Design of experiments may comprise:

[0041] - receiving an identifier indicating which physical or chemical characteristic shall be determined by the chemometric model,

[0042] - determining instructions for measurements on objects which different values for the identified physical or chemical characteristics by using the identifier,

[0043] - outputting the instructions, and

[0044] - in response to outputting the instructions receiving spectra of an object and object data associated with a physical or chemical characteristic of the object according to the instructions.

[0045] Design of experiments may involve a DOE model which uses the identifier as input. The identifier may further indicate a measurement range for the physical or chemical characteristic. The output may be generated by the DOE model. The DOE model may be adjusted to the candidate chemometric models, i.e. take into account the requirements for training the chemometric models. For example, if the candidate chemometric models comprise an artificial neural network, more measurements may be required in comparison to a multivariate regression.

[0046] The DOE model may comprise full factorial design, i.e. varying all possible combinations of physical or chemical characteristics and their levels to enable systematic evaluation their impact on the spectrometer measurements. The DOE model may comprise fractional factorial design, i.e. selecting a subset of the physical or chemical characteristics and their levels to enable systematic evaluation their impact on the spectrometer measurements. Selecting a subset may involve randomization, statistical replication, blocking, or orthogonality. The received spectra and object data may be compared to the instructions. The comparison may be designed to make sure that the instructions have sufficiently complied with. The comparison may yield a score indicating how many spectra and object data correspond to the instructions. In case the comparison yields an insufficient compliance, for example when the score is below a preset threshold, a warning may be output indicating the insufficiency of the received data. A user may be invited to provide further spectra and object data. The user may have the option to request continuation of the method, for example if external factors like inaccessibility of required objects hinder the recording of the requested data.

[0047] The method of the present invention may comprise data wrangling of the received spectra or the object data. The term "data wrangling” may refer to transforming and cleaning the data of the received spectra or the object data into data which can be used for training the candidate chemometric models. Sometimes data wrangling is also referred to as data munging. Data wrangling may comprise data parsing. Data parsing may involve natural language processing, for example large language models employing transformers. As an example, the object data may contain the word high density polyethylene and its abbreviation HDPE. Data parsing may converge both terms into one consistent term for polyethylene. Data wrangling may comprise data cleaning. For example, the decimal separator in numeric values may be a dot or a comma depending on the data source. Data cleaning may uniformly use one specific decimal separator. Data wrangling may comprise data format adjustment. For example, spectra may be received as comma-separated file and object data as json file, while training the candidate chemometric models requires both spectra and object data in vector format. Data format adjustment may hence transform a source format into a format required for training the chemometric models. Data wrangling may comprise data transformation. Data transformation may refer to changing values while retaining their meaning. For example, object data may comprise the concentration of a compound in mol / l, however, the chemometric model shall determine the concentration in g / l. Hence, data transformation may comprise convert one value given in one unit into a value of a different unit.

[0048] The method of the present invention comprises generating at least two different candidate chemometric models. A candidate chemometric model may be a chemometric model which is not yet trained, but contains the algorithms which are ready to be trained. The term "at least two” may mean two or more, for example 10 to 10 000 or 50 to 1000. The term "different” may mean fully or partially different, i.e. two candidate chemometric models differ in at least one aspect. The candidate chemometric model may differ in one of the pre-processing or machine learning method or in both. Two candidate chemometric models may differ in one or more parameters used for preprocessing or machine learning method. Two candidate chemometric models should be sufficiently different, i.e. a significant difference in their prediction accuracy can be expected. The choice of the candidate models may involve results of previous chemometric model generations, for example those combinations of parameters and methods which yielded satisfactory results for similar problems in the past. The choice of the candidate models may involve a design of experiment approach, i.e. selection of a limited number of combinations of methods and parameters optimized to systematically span the option space. The candidate chemometric model of the present invention may comprise a pre-processing method. The term "preprocessing” may refer to a method to reduce or eliminate interferences from a spectrum such as stray light, noise or baseline drift to enhance the subsequent machine learning. Hence, the pre-processing method may be applied before the machine learning method. Pre-processing may include one or more of baseline correction, scatter correction, smoothing, scaling, aggregation.

[0049] Baseline correction may be chosen from linear base line correction, polynomial baseline correction, logarithmic baseline correction, spline baseline correction, Savitzky-Golay baseline correction, Kaiser-Bessel baseline correction, Gaussian baseline correction, Lorentzian baseline correction, derivative baseline correction including first-order derivation or second-order derivation, continuous wavelet transform (CWT), or baseline correction by local regression.

[0050] Scatter correction may be chosen from multiplicative scatter correction (MSC), standard normal variate (SNV), background subtraction, dark current subtraction, scatter correction using a scatter map or a scatter model, wavelet denoising.

[0051] Smoothing may be chosen from Savitzky-Golay filtering (SGF), Kaiser-Bessel filtering (KBF), Gaussian filtering (GF), median filtering (MF), moving average filtering (MAF), exponential smoothing (ES), or Kalman filtering (KF).

[0052] Scaling may be chosen from linear scaling, polynomial scaling, logarithmic scaling, mean centering (MC), auto scaling (AS), Pareto scaling (PS) or Max-Min scaling (MMS).

[0053] Aggregation may be chosen from creation of median, spatial median, or k-medoids center across a set of spectra, for example all measurements of one sample measured by one device.

[0054] Pre-processing can comprise one pre-processing method or a combination of more than one, for example two, three, four, or five. For example, pre-processing may comprise a method for baseline correction, a method for scatter correction, a method for smoothing, a method for scaling, and a method for aggregation. In case pre-processing comprises more than one pre-processing method, the pre-processing method may be applied in various order, wherein changing the order is regarded as different pre-processing in the context of the present invention. Usually, pre-processing is applied before the machine learning method in a chemometric model.

[0055] The candidate chemometric model of the present invention may comprise a machine learning method. The term "machine learning method” may refer to a model which translates spectra into corresponding object data. The machine learning method hence may use a spectrum as input and derive object data therefrom. The machine learning method may be considered as an integral part of the chemometric model. Machine learning methods may be supervised, semi-supervised or unsupervised. Machine learning methods may include multivariate calibration, classification, pattern recognition, clustering, ensemble methods, neural nets and deep learning, or multivariate curve resolution.

[0056] The selection of machine learning methods may depend on the object data to be determined. Hence, selecting a machine learning model for a candidate chemometric model may involve the object data which the chemometric model shall output. If a variable, for example a concentration, shall be determined, multivariate calibration methods may be selected. If a categorial value, for example the kind of compound, shall be determined, classification models or spatial clustering models may be selected. Selecting a machine learning model may involve an identifier comprised in the object data indicating which object data the chemometric model shall output.

[0057] Multivariate calibration methods may be classical or inverse. A classical multivariate calibration model may be solved such that it optimally describes the spectrum. An inverse multivariate calibration method may be solved to optimally predict the object data. Hence, inverse multivariate calibration methods are preferred. Multivariate calibration may be chosen from partial-least squares regression (PLS), principal component regression (PCR), penalized linear regression (Lasso, Ridge, Elastic Net), random forest regression (RFR), support vector regression (SVR), gradient boosting regressor (GBR), vector regression (VR), matrix regression (MR), artificial neural networks (ANN), multivariate adaptive regression splines (MARS).

[0058] Classification models may be chosen from linear discriminant analysis (LDA), soft independent modeling of class analogy (SIMCA), penalized linear regression (Lasso, Ridge, Elastic Net), artificial neural networks (ANN), Naive Bayes, random forest classification (RFC), support vector machines (SVM).

[0059] Spatial clustering models may be chosen from K-means cluster analysis (KMCA), agglomerative hierarchical cluster analysis (AHCA), principal component analysis (PCA), fuzzy C means cluster analysis (FCMCA), vertex component analysis (VGA), divisive correlation cluster analysis (DCCA).

[0060] The candidate chemometric model of the present invention may comprises a feature selection filter. The term "feature selection filter” may refer to a method to select those parts of the spectrum with a correlation to the object data. A feature selection filter may facilitate the machine learning method of the chemometric model and thus avoid overfitting and reduce the number of required training datasets. A feature selection filter may use a spectrum as input, remove all unselected parts and output a spectrum with only the selected parts left. Hence, the output of the feature selection filter may be a spectrum in form of a vector of lower dimensionality than the input vector. The output of the feature selection filter can be used as input for the machine learning method. Hence, the feature selection filter may be applied before the machine learning method. The input of the feature selection filter may be the received spectrum or it may be the pre-processed spectrum, preferably the pre-processed spectrum. Hence, the feature selection filter may be applied after the pre-processing method. The feature selection filter may be obtained from a feature selection method, i.e. a method to determine the relevant part of the spectra which the feature selection filter selects. A feature selection method may be univariate, i.e. it selects those parts of the spectra which correlate with the object data, typically because the physical or chemical characteristic of the object gives rise to such parts. The feature selection method may be sequential, i.e. it ranks parts of the spectra in order and pairs the parts of the spectra in a forward or backward progression. A part of a spectrum may refer a single intensity value at a certain wavelength or to multiple intensity values at different wavelength, for example for a range of wavelengths. The feature selection method may be multivariate, for example multivariate linear regression (MLR), principle component regression (PCR), partial least square regression (PLSR), interactive variable selection, uninformative variable elimination (UVE), interval PLS (iPLS), significance tests of model parameters, or genetic algorithms (GAs).

[0061] A chemometric model which outputs more than one physical or chemical characteristic may comprise different preprocessing methods for different physical or chemical characteristics. A chemometric model which outputs more than one physical or chemical characteristic may comprise different machine learning methods for different physical or chemical characteristics. A chemometric model which outputs more than one physical or chemical characteristic may comprise different feature selection filters for different physical or chemical characteristics.

[0062] A candidate chemometric model may be generated by selecting a pre-processing method, a machine learning method and, optionally, a feature selection filter and merging them to a candidate chemometric model. The selection may be limited by the use case. For example, for a determination of the concentration of a certain compound in an object, only regression calibration methods may be selected as machine learning method while for discriminating certain materials, for example different types of plastics, only classification methods may be selected as machine learning method. The selection may be limited by a user input or a configuration file.

[0063] Two candidate chemometric models may be different in at least one of the pre-processing method, machine learning method, or feature selection filter. The candidate chemometric models may comprise at least two candidate chemometric models which comprise a different pre-processing method. The candidate chemometric models may comprise at least two candidate chemometric models which comprise a different machine learning method. The candidate chemometric models may comprise at least two candidate chemometric models which comprise a different feature selection filter. The candidate chemometric models may comprise at least two candidate chemometric models which comprise a different pre-processing method and at least two candidate chemometric models which comprise a different machine learning method. The candidate chemometric models may comprise at least two candidate chemometric models which comprise a different pre-processing method and at least two candidate chemometric models which comprise a different feature selection filter. The candidate chemometric models may comprise at least two candidate chemometric models which comprise a different machine learning method and at least two candidate chemometric models which comprise a different feature selection filter. The candidate chemometric models may comprise at least two candidate chemometric models which comprise a different pre- processing method and at least two candidate chemometric models which comprise a different machine learning method and at least two candidate chemometric models which comprise a different feature selection filter. A method may be regarded as different if it comprises at least one different parameter or parameter value.

[0064] The method of the present invention comprises training the candidate chemometric models. Training may comprise adjusting parameters of the chemometric models such that the output of the chemometric model most closely fits to the provided object data associated with the spectra used as input for the chemometric model. Often, training comprises minimizing a loss or cost function, for example a least mean square value of chemometric model output to provided object data. The complete set of received spectra and associated object data may be used for training or parts thereof. Parts of the received spectra and associated object data may be used for training and the remainder may be used for determining the prediction accuracy of the trained candidate chemometric models. Alternatively, cross-validation can be applied, for example K-fold cross-validation, leave-one-out cross-validation, stratified cross- validation.

[0065] The received spectra and associated object data may be subject to a data analysis. The term "data analysis” may refer to an analysis to determine the suitability of the received spectra and the object data for training a chemometric model before training. The analysis may involve anomaly detection, outlier detection, or signal-to-noise ratio analysis. Those spectra exceeding a preset threshold may be qualified as unsuitable and hence disregarded. The analysis may involve the object data. For example, the analysis may involve checking if the provided object data contains the required information for the chemometric model. In this way, accidentally added spectra from other use cases can be disregarded. Another example may be to check within the object data if all spectra are recorded with the same or comparable spectrometers to ensure proper training of the chemometric model designed for a certain spectrometer or type of spectrometer. It may further be checked if all object data is within the desired range. The spectra may be checked with respect to their value range, for example to identify if the spectrometer was at its limit, i.e. outside its intended measurement range.

[0066] Unsuitable spectra, i.e. spectra which have been identified to be unsuitable, may be disregarded. All unsuitable spectra may be disregarded or parts thereof. The unsuitable spectra may be displayed to a user. The user may in response provide an input indicating which unsuitable spectra should be disregarded. Disregarding unsuitable spectra may hence involve such user input. If the analysis for suitability yields that the number of suitable spectra is too low, the user may be invited to provide further data, for example by outputting such request to a user interface.

[0067] The method of the present invention may comprise data augmentation for augmenting the provided spectra and the associated object data. The term "data augmentation” may refer to generating additional spectra and associated object data starting from the provided spectra and associated object data. By using data augmentation, the chemometric model can be trained with more training data which may reduce overfitting. Hence, the candidate chemometric models may be trained using the augmented spectra and the object data, i.e. the received spectra and the object data and the spectra and object data obtained from data augmentation. Furthermore, the chemometric model may be applicable more generally, i.e. for measurement ranges for which scare or no training data is available. Hence, by using data augmentation a robust chemometric model is accessible with decreased need for large datasets containing spectra and associated object data.

[0068] Data augmentation typically involves a generative model which generates spectra and associated object data. The generative model may derive the correlations between spectra associated object data from the provided spectra associated object data. Hence, the generative model may be trained with the provided spectra and associated object data.

[0069] Data augmentation may involve the identifier for indicating which physical or chemical characteristic the chemometric model shall determine contained in the object data. The identifier may be used to restrict the training of the generative model to only the object data indicated by the identifier. This may help to avoid underfitting. Alternatively, a user input or a configuration file may be used for this purpose.

[0070] The generative model may generate new spectra and associated object data by adding noise to the provided spectra and object data, for example Gaussian noise. For example, the model may determine the standard deviation for each wavelength in the spectra, select a noise scaled according to the standard deviation, and add Gaussian noise to the spectra at each individual wavelength.

[0071] The generative model may generate new spectra and associated object data by blending the provided spectra and object data. Blending may refer to generating linear combinations of two or more spectra and associated object data. The new data hence constitutes interpolations of the provided data.

[0072] The generative model may generate new spectra and associated object data by synthetic minority oversampling technique (SMOTE). SMOTE may involve selecting similar spectra, for example those which are close to each other in a feature space, and generating new spectra along a trajectory in the feature space between the selected spectra.

[0073] The generative model may contain an artificial neural network. The artificial neural network may be trained with the provided spectra and associated object data and subsequently output new spectra and associated object data. Examples for generative models containing an artificial neural network include learning augmentation policies from data (AutoAugment), Fast AutoAugment, Faster AutoAugment, generative adversarial networks (GAN), variational autoencoders (VAE), conditional variation autoencoder (CVAE), Deterministic Regularized Autoencoder (RAE), or normalizing flow-based generative models.

[0074] The generative model may be selected manually, so the same generative model is used any time the method of the present invention is executed. Alternatively, a couple of different generative models may be tested for a specific chemometric model, wherein for each generative model a change of prediction accuracy of the chemometric model when using the augmented dataset for training in comparison to using only the provided dataset for training may be determined. The generative model may then be selected involving the determination of the change of prediction accuracy caused by using the augmented dataset for training the chemometric model. For example, the generative model causing the highest increase in prediction accuracy of the chemometric model may be selected.

[0075] In some cases, using different datasets augmented by different generative models may result in different prediction accuracies for different trained candidate chemometric models. Hence, multiple candidate generative models may be used for different candidate chemometric models and the prediction accuracy may be determined for each combination. The selection of the generative model and the trained candidate chemometric model may involve the thus determined prediction accuracy. Hence, the selection of the trained candidate chemometric model may depend on the choice of the generative model and vice versa.

[0076] The method of the present invention comprises determining the prediction accuracy of the trained candidate chemometric models. The prediction accuracy may be a value indicating the difference between output of the trained candidate chemometric model and measured data, for example a validation set of the received object data. Generally, the higher the prediction accuracy the better the fit of the chemometric model output to the reference object data. The value indicating the difference may be or may be derived from slope, bias, intercept, R-square, standard error of calibration (SEC), standard error of prediction (SEP), root mean square error of calibration (RMSEC), mean square error (MSE), root mean square error of prediction (RMSEP), mean absolute error (MAE), ratio of performance deviation (RPD), or range error ratio (RER). The prediction accuracy of chemometric models comprising a binomial classifier may be determined from the area under the curve (AUG) in a receiver operating characteristic curve (ROC), sensitivity and specificity based on a dedicated cut-off, according to the Youden index, or others). The prediction accuracy of chemometric models comprising a multinomial classifier may be determined from Sensitivity, Specificity, Precision, F1-Score and Accuracy per class and as averaged and weighted average across all classes.

[0077] The method of the present invention comprises selecting a chemometric model of the trained candidate chemometric models using the prediction accuracy. This may mean picking one of the trained candidate chemometric models which will become the selected chemometric model. It is possible to select the trained candidate chemometric model with the highest prediction accuracy. However, in some cases it may be advantageous to select a trained candidate chemometric model with a lower prediction accuracy. For example, a trained candidate chemometric model with the highest prediction accuracy below a preset threshold may be selected. The preset threshold may be chosen to exclude exceptionally high prediction accuracies. Exceptionally high prediction accuracies may indicate an overfit trained candidate chemometric model. Alternatively, a number of trained candidate chemometric models may be output to a user interface, for example two to ten or three to five, to a user interface and request a user selection input. The trained candidate chemometric models may be output with their prediction accuracy. The output trained candidate chemometric models may be the trained candidate chemometric models with the highest prediction accuracy. The number of trained candidate chemometric models to be output to the user interface may be a fraction of the number of generated trained candidate chemometric model, for example 1 to 5 %. The selection of a trained candidate chemometric model may then involve the user input.

[0078] The method of the present invention comprises outputting the selected trained chemometric model. The term "outputting” may refer to writing the selected trained chemometric model to a non-transitory data storage medium, for example into a file or database, display it on a user interface, for example a screen, or both. It is also possible to output the selected trained chemometric model through an interface to a cloud system for storage and / or further processing. The selected trained chemometric model may further be output through an interface to a spectrometer which may use the selected trained chemometric model for spectrometric measurements. The selected trained chemometric model may be output together with its prediction accuracy. In this case, a user may estimate the success of the method and decide if the output chemometric model shall be employed or not.

[0079] The present invention further relates to a system for generating a chemometric model for a spectrometer. The system comprises an input for receiving spectra of an object and object data associated with a physical or chemical characteristic of the object. The input may comprise an interface for receiving the data. The input may receive the data locally or remotely, for example via an interface to a telecommunication system, such as the internet. The input may receive the data directly from a spectrometer, or via a programmable logic controller, or a storage medium including a cloud service.

[0080] The system further comprises a processor configured to execute steps of the method of the present invention. The processor may be a local processor comprising a central processing unit (CPU) and / or a graphics processing unit (GPU) and / or an application specific integrated circuit (ASIC) and / or a tensor processing unit (TPU) and / or a field- programmable gate array (FPGA). The processor may also be an interface to a remote computer system such as a cloud service.

[0081] The system further comprises an output for outputting the selected trained chemometric model. The output may comprise an interface for outputting the selected trained chemometric model. The output may send the selected trained chemometric model locally or remotely, for example via an interface to a telecommunication system, such as the internet. The output may send the selected trained chemometric model to a programmable logic controller, or a storage medium including a cloud service.

[0082] The present invention further relates to a non-transient computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to the present invention. The term "computer-readable data medium” may refer to any suitable data storage device or computer readable memory on which is stored one or more sets of instructions (for example software) embodying any one or more of the methodologies or functions described herein. The instructions may also reside, completely or at least partially, within the main memory and / or within the processor during execution thereof by the computer, main memory, and processing device, which may constitute computer-readable storage media. The instructions may further be transmitted or received over a network via a network interface device. Computer-readable data medium include hard drives, for example on a server, USB storage device, CD, DVD or Blue-ray discs. The computer program may contain all functionalities and data required for execution of the method according to the present invention or it may provide interfaces to have parts of the method processed on remote systems, for example on a cloud system.

[0083] Brief Description of the Figures

[0084] Figure 1 illustrates the acquisition of object data and spectra from an object.

[0085] Figure 2 illustrates an embodiment of the method of the present invention.

[0086] Figure 3 illustrates an example of how candidate chemometric models may be generated.

[0087] Figure 4 illustrates a potential feature selection filter of an spectra.

[0088] Figure 5 illustrates an example for a chemometric model.

[0089] Figure 6 illustrates an embodiment of the system of the present invention.

[0090] Figure 7 shows an exemplary user interface for the system of the invention.

[0091] Description of Embodiments

[0092] Figure 1 illustrates the acquisition of object data and spectra from an object. A spectrometer 101 may illuminate the object 103, for example a nutrient, a pharmaceutical or a body part, with a light beam 102. The spectrometer may contain a light source, for example an incandescent lamp or an LED, and illumination optics. The light beam 102 may contain light in the infrared wavelength range, for example 780 nm to 3000 nm. The light beam may hit the object and be reflected back to the spectrometer. This measurement mode is often referred to as reflectance mode.

[0093] Alternatively, the light beam may penetrate the object and be directed back to the spectrometer, for example with mirrors. This measurement mode is often referred to as transmission mode. The spectrometer 101 may comprise light collection optics to capture the light beam 102 travelling from the object 103 to the spectrometer 101. The spectrometer 101 may comprise an optical element which separates different wavelengths of the light beam 102 travelling from the object 103 to the spectrometer 101 in space or in time. The spectrometer 101 may contain one or multiple optical sensors which generate an electrical signal depending on the intensity of light impinging on the optical sensor. The intensity data may be associated with the wavelength data, for example by a processor or microcontroller in the spectrometer. The spectrometer 101 may hence output an spectra 104, for example as vector of wavelength values and an associated vector of intensity values. The object 103 may be analyzed in a laboratory 105 to obtain object data 106, for example the content of a compound in the object 103 as determined by gas chromatography. The spectra 104 and the object data 106 may be associated, for example by labelling both with an object identifier.

[0094] Figure 2 illustrates an embodiment of the method of the present invention. Spectra 201 and object data 202 may be received, for example through a web interface. The object data 202 is associated with the spectra 201, for example by an identifier representing the sample a spectrum originates from and the object data is associated with. The received spectra 201 and object data 202 may be analyzed 203 for the suitability of training a chemometric model. Those IR spectra 201 and object 202 found to be unsuitable may be discarded, i.e. not used for training. The remaining IR spectra 201 and object 202 may be augmented by data augmentation 204. Data augmentation 204 may involve a generative model which is trained with the spectra 201 and object data 202 and which generates additional spectra and object data, for example those spectra and object data which are similar, but different to the spectra 201 and object data 202.

[0095] A candidate chemometric model may be generated 210. The generation may comprise selecting a pre-processing method 211, for example a baseline correction method, a scatter correction method, a smoothing method and / or a scaling method. The generation may further comprise selecting a machine learning method 212, for example a multivariate calibration. The generation may further comprise selecting a feature selection filter 213, for example a deletion operation removing everything except a certain set of peaks of the spectra 201 . The generating may comprise merging the selected pre-processing method, the selected machine learning method and feature selection filter, for example by connecting the output of the pre-processing method with the input of the feature selection filter and the output of the feature selection filter with the input of the machine learning model. A candidate chemometric model may be obtained which is untrained, i.e. it contains parameters which are not yet adjusted to fit the spectra 201 and the object data 202. Generating a candidate chemometric model 210 may be repeated several times to obtain several different candidate chemometric models, for example 200 or 400 different candidate chemometric models.

[0096] The candidate chemometric models may be trained 221, wherein for the training the complete or parts of the spectra 201 and object data 202 may be used. If data augmentation 203 has been performed, the augmented data sets are used, i.e. the provided spectra 201 and object data 202 and the newly generated spectra and object data. Once the candidate chemometric models are trained, their prediction accuracy may be determined 222, for example by determining a root mean square difference of the predicted object data to a validation dataset of object data, for example a part of the provided object data 202. Alternatively, cross validation may be applied, for example by K-fold cross validation. Subsequently, a trained candidate chemometric model may be selected 223, for example the trained candidate chemometric model with the highest prediction accuracy. The selected chemometric model may be output 224, for example by writing it to a data storage device, sending it to a spectrometer through a communication interface or by providing it for download through a web interface.

[0097] Figure 3 illustrates an example of how candidate chemometric models may be generated. The candidate chemometric models may contain a pre-processing method 310, a machine learning method 320 and a filter for the relevant parts of the spectra 330. Pre-processing 310 may comprise a baseline correction 311, a scatter correction 312, smoothing 313 and scaling 314. As an example, the baseline correction 311 may be effected by choosing from the first derivative method 311a and the continuous wavelet transform method 311b or none. Scatter correction 312 may be effected by choosing from standard normal variate method 312a, multiplicative scatter correction method 312b or none. Smoothing 313 may be effected by choosing from moving average filtering 313a, exponential smoothing 313b or none. Scaling may be effected by choosing from mean centering 314a, Pareto scaling 314b or none. Machine learning 320 may comprise multivariate calibration 321, classification 322, spatial clustering 323. As an example, the multivariate calibration 321 may be effected by choosing from partial-least squares regression 321a, principal component regression 321b, or none. Classification 322 may be effected by choosing from artificial neural networks 322a, linear discriminant analysis 322b, or none. Spatial clustering 323 may be effected by choosing from K-means cluster analysis 323a, agglomerative hierarchical cluster analysis 323b or none. A feature selection filter 330 may leave the complete spectra as it is 331, select a first wavelength range 332, for example 1.1 -1.4 pm, select a second wavelength range 333, for example 1 .6-1 .8 pm, select a third wavelength range 334, for example 1 .9-2.3 pm, or a combination thereof.

[0098] As an illustrative example, three candidate chemometric models may be generated. A first candidate chemometric model 341 may comprise the first derivative model 311a as baseline correction 311, no scatter correction 312, exponential smoothing 313b as smoothing 313, mean centering 314a as scaling 314, principal component regression 321b as multivariate calibration 321, and the full spectrum 331 as feature selection filter 330, meaning essentially no feature selection filter is used. A second candidate chemometric model 342 may comprise the continuous wavelet transform method 311b as baseline correction 311, multiplicative scatter correction method 312b as scatter correction 312, no smoothing 313, Pareto scaling 314b as scaling 314, linear discriminant analysis 322b as classification 322, and a selection of part 332 and 333 of the spectra as feature selection filter 330. A third candidate chemometric model 343 may comprise no baseline correction 311, standard normal variate method 312a as scatter correction 312, moving average filtering 313a as smoothing 313, no scaling 314, artificial neural networks 322a as multivariate calibration 321, and the selection of part 332 and 334 of the spectra as feature selection filter 330.

[0099] Figure 4 illustrates a potential feature selection filter of a spectra. The spectrum 400 plots the intensity of the radiation 402 against the wavelength 401 . In this example, a wavelength range of 1 to 3 pm is plotted. The spectrum 400 contains three maxima or peaks, one at 1.6 pm, one at 2.2 pm and one at 2.8 pm. Each peak can potentially correlate to the object data. Hence, a range around the peaks may be selected as potentially relevant parts of the spectrum, in this example first part 410, second part 420 and third part 430. Feature selection filter can be selected removing everything from the spectrum except one or more than one of the first part 410, second part 420 and third part 430. Another feature selection filter may be formed which leaves all or almost all of the spectrum unaffected. Another possibility is to form a feature selection filter which removes the parts of the spectra for which the intensity 402 does not exceed a predefined value, for example 0.1. In this way, low intensity regions may be discarded which usually do not contain relevant information. More sophisticated methods for feature selection filter are available as described above, but in essence they work as illustrated in figure 4.

[0100] Figure 5 illustrates an example for a chemometric model. The chemometric model 510 may contain a pre-processing method 511, a feature selection filter 513 and a machine learning method 515. A spectrum 501 may be provided as input to the chemometric model 510. The spectrum 501 may be a vector of multiple intensity values at different wavelength as obtained from a spectrometer. The spectrum 501 may be first pre-processed by pre-processing 511 . Pre-processing may comprise one or more of baseline correction, scatter correction, smoothing, and scaling. The pre-processing outputs a pre-processed spectrum 512. The pre-processed spectrum may be a vector of multiple intensity values at different wavelength, wherein the values may have been modified by the pre-processing in comparison to the spectrum 501 . The pre-processed spectrum 512 may be subject to a feature selection filter 513. The feature selection filter 513 may remove some intensity values which have been identified as hardly or not correlating to the desired object data the chemometric model shall determine, in this example l'2(A2). The feature selection filter 513 hence outputs relevant parts of the pre-processed spectrum 514. The spectrum 514 may be a vector with a lower dimensionality than spectrum 501 or 512. Spectrum 514 may be input to machine learning method 515 which in turn outputs object data 520. In this example, the output object data may be the concentration of an analyte compound to be 0.23 g / l.

[0101] Figure 6 illustrates an embodiment of the system of the present invention. The system for generating a chemometric model 620 may be a cloud-based service. The system 620 may receive spectra 611 and object data 612 from a web interface, for example via an upload functionality. The IR spectra 611 may have been recorded with the spectrometer 601 or with a comparable one. The object data 612 may have been obtained by a laboratory analysis. The system 620 may execute the method of the present invention using the spectra 611 and the object data 612. The system 620 may output the chemometric model 630, for example by a communication interface directly to the spectrometer 601 or to a smartphone 605 which is communicatively coupled to the spectrometer 601 . The smartphone 605 may employ the chemometric model 630 by triggering the spectrometer 601 to measure a spectrum an object 603. The spectrometer 601 may transfer the spectrum to the smartphone 605. The processor of the smartphone 605 may execute code which uses the chemometric model 630 to translate the spectrum into object data. The object data may be displayed on the display of the smartphone 605. Alternatively or additionally, the processor of the smartphone may execute an application configured to further translate the object data into a recommendation for an action, for example if the object 603 is the skin of a person and the chemometric model 630 is adjusted to detect the moisture level of the skin, the recommendation may be to apply a certain amount of a skin care product to the skin.

[0102] Figure 7 shows an exemplary user interface for the system of the invention. The user interface may be an application on a computer or a web interface. The interface may display a window for a chemometric model generator 700. The interface may have an upload section, in which the user can upload spectra 711 and object data 713. The interface may provide buttons for browsing for respective files 712, 714. Once the upload is successfully finished, the user may have the option to trigger the chemometric model generation method by clicking on generate model 721. The generated model may be stored in a buffer memory from where it can be downloaded by clicking on a button labelled download model 722.

[0103] The present disclosure has been described in conjunction with preferred embodiments and examples as well. However, other variations can be understood and effected by those persons skilled in the art and practicing the claimed invention, from the studies of the drawings, this disclosure and the claims.

[0104] Any steps presented herein can be performed in any order. The methods disclosed herein are not limited to a specific order of these steps. It is also not required that the different steps are per-formed at a certain place or in a certain computing node of a distributed system, i.e. each of the steps may be performed at different computing nodes using different equipment / data processing.

[0105] As used herein ..determining" also includes ..initiating or causing to determine", "generating" also includes ..initiating and / or causing to generate" and "providing” also includes "initiating or causing to determine, generate, select, send and / or receive”. "Initiating or causing to perform an action” includes any processing signal that triggers a computing node or device to perform the respective action.

[0106] In the claims as well as in the description the word "comprising” does not exclude other elements or steps and the indefinite article "a” or "an” does not exclude a plurality. A single element or other unit may fulfill the functions of several entities or items recited in the claims. The mere fact that certain measures are recited in the mutual different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation. In the claims as well as in the description the word "comprising” or "including” or similar wording does not exclude other elements or steps and shall not be construed limiting to the elements or steps lined out. The indefinite article "a” or "an” does not exclude a plurality. A single element or other unit may fulfill the functions of several entities or items recited in the claims. The mere fact that certain measures are recited in the mutual different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation or further elements may be included. Providing in the scope of this disclosure may include any interface configured to provide data. This may include an application programming interface, a human-machine interface such as a display and / or a software module interface. Providing may include communication of data or sub-mission of data to the interface, in particular display to a user or use of the data by the receiving node, entity or interface.

[0107] Various units, circuits, entities, nodes or other computing components may be described as "con-figured to” perform a task or tasks. Configured to shall recite structure meaning "having circuitry that” performs the task or tasks on operation. The units, circuits, entities, nodes or other computing components can be configured to perform the task even when the unit / circuit / component is not operating. The units, circuits, entities, nodes or other computing components that form the structure corresponding to "configured to” may include hardware circuits and / or memory storing program instructions executable to implement the operation. The units, circuits, entities, nodes or other computing components may be described as performing a task or tasks, for convenience in the description. Such descriptions shall be interpreted as including the phrase "configured to.” Any recitation of "configured to” is expressly intended not to invoke 35 U.S.C. § 112(f) interpretation.

[0108] In general, the methods, apparatuses, systems, computer elements, nodes or other computing components described herein may include memory, software components and hardware components. The memory can include volatile memory such as static or dynamic random-access memory and / or nonvolatile memory such as optical or magnetic disk storage, flash memory, programmable read-only memories, etc. The hardware components may include any combination of combinatorial logic circuitry, clocked storage devices such as flops, registers, latches, etc., finite state machines, memory such as static random-access memory or embedded dynamic random-access memory, custom designed circuitry, programmable logic arrays, etc.

[0109] Any disclosure and embodiments described herein relate to the methods, the systems, apparatuses, devices, chemicals, materials, computer program elements lined out above and vice versa. Advantageously, the benefits provided by any of the embodiments and examples equally apply to all other embodiments and examples and vice versa. All terms and definitions used herein are understood broadly and have their general meaning.

Claims

Claims1 . A computer-implemented method for generating a chemometric model for a spectrometer comprising: a) receiving spectra of an object and object data associated with a physical or chemical characteristic of the object, b) augmenting the provided spectra and the associated object data, c) generating at least two different candidate chemometric models, d) training the candidate chemometric models using the augmented spectra and the object data, e) determining the prediction accuracy of the trained candidate chemometric models, f) selecting a chemometric model of the trained candidate chemometric models using the prediction accuracy and g) outputting the selected chemometric model.

2. The method according to claim 1, wherein augmenting the provided spectra and the associated object data involves an artificial neural network as generative model.

3. The method according to claim 1 or 2, wherein the at least two different candidate chemometric models comprise a pre-processing method and a machine learning method.

4. The method according to any of the claims 1 to 3, the at least two different candidate chemometric models comprise a feature selection filter.

5. The method according to any of the claims 1 to 4, wherein the spectra represent the absorbance or transmittance of radiation of a wavelength of 760 nm to 3 pm.

6. The method according to any of the claims 1 to 5, wherein the object data comprises data associated with the chemical composition of the object.

7. The method according to any of the claims 1 to 6, wherein the received spectra and associated object data is analyzed for their suitability before training.

8. Use of the chemometric model obtained from the method of any of the previous claims for a spectrometer.

9. A spectrometer configured to use the chemometric model obtained from the method of any of the claims 1 to 7.

10. The spectrometer according to claim 9, wherein the spectrometer is a portable spectrometer.11 . The spectrometer according to claim 10, wherein the spectrometer is communicatively coupled to a computer device configured to execute the chemometric model.

12. A system for generating a chemometric model for a spectrometer comprising: a) an input for receiving spectra of an object and object data associated with a physical or chemical characteristic of the object, b) a processor for augmenting the provided spectra and the associated object data, generating at least two different candidate chemometric models and training the candidate chemometric models, determining the prediction accuracy of the trained candidate chemometric models using the augmented spectra and the object data, and selecting a chemometric model of the trained candidate chemometric models using the prediction accuracy and c) an output for outputting the selected chemometric model.

13. A non-transient computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform a method comprising: a) receiving spectra of an object and object data associated with a physical or chemical characteristic of the object, b) augmenting the provided spectra and the associated object data, c) generating at least two different candidate chemometric models d) training the candidate chemometric models using the augmented spectra and the object data, e) determining the prediction accuracy of the trained candidate chemometric models, f) selecting a chemometric model of the trained candidate chemometric models using the prediction accuracy and g) outputting the selected chemometric model.

14. The non-transient computer-readable medium according to claim 13, wherein the candidate chemometric models further comprise a feature selection filter.

15. The non-transient computer-readable medium according to claim 13 or 14, wherein the method further comprises data augmentation for augmenting the provided spectra and the associated object data.

Citation Information

Patent Citations

  • Chemometrics for near infrared spectral analysis

    US20130080070A1

  • Use of genetic algorithms to determine a model to identity sample properties based on raman spectra

    US20230009725A1

Cited By

  • Near infrared spectrogram analysis model, construction method and device thereof and computer equipment

    CN120977423A