Cross-validation-based calibration of spectroscopic model

Cross-validation techniques enhance spectroscopic model accuracy by merging datasets to adapt models across different spectrometers, addressing inefficiencies and cost issues in spectroscopic model deployment.

JP2025169262APending Publication Date: 2025-11-12VIAVI SOLUTIONS INC(US)
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025123886
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-06-21
Filing Date
2025-07-24
Publication Date
2025-11-12

AI Technical Summary

Technical Problem

Existing spectroscopic models become inaccurate over time due to changes in raw materials, and different spectrometers require individual calibration, leading to inefficiencies and increased costs in deployment.

Method used

Implement cross-validation techniques to update and transition spectroscopic models by merging data from a master dataset with a target dataset, using partial least squares (PLS) factors to improve accuracy and reduce the need for individual spectrometer-specific data collection.

Benefits of technology

Enhances the accuracy of spectroscopic models by reducing the necessity of acquiring a master dataset for each spectrometer, thereby lowering deployment costs and improving model precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025169262000001_ABST
    Figure 2025169262000001_ABST
Patent Text Reader

Abstract

To provide a device, a method, and a non-transitory computer-readable medium for generating a spectroscopic model using a training data set for spectroscopically determining material identification.SOLUTION: A device with a processor may: receive (410) a master data set for a first spectroscopic model; receive a target data set for a target population associated with the first spectroscopic model to update (420) the first spectroscopic model; generate (430) a training data set that includes the master data set and first data from the target data set; generate (440) a validation data set that includes second data from the target data set but not the master data set; generate (450), using cross-validation and using the training data set and the validation data set, a second spectroscopic model that is an update of the first spectroscopic model; and provide (460) the second spectroscopic model.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [Related Applications] This application is a joint venture under 35 U.S.C. § 119 of the "Near Infrared ( Updating calibration models based on NIR spectra No. 62 / 692,248, entitled "NEAR-INFRARED (NIR) SPECTRA" Priority is claimed, the contents of which are incorporated herein by reference in their entirety. [Background technology]

[0002] Raw material identification can be used for quality control of pharmaceuticals. For example, raw material identification can be used for medical materials. The medical materials are manufactured in accordance with the following criteria: Similarly, it is possible to perform ingredient quantification and determine whether the product corresponds to the packaging label. Spectroscopy can be used to determine the concentration of a particular chemical in a particular sample. Non-destructive raw material identification and characterization with less preparation and shorter data collection time than metric methods and / or may facilitate quantification. Summary of the Invention [Means for solving the problem]

[0003] According to some embodiments, the device comprises one or more memories and one or more memory one or more processors communicatively coupled to the first spectroscopic model master; receiving a dataset and target data for a target population associated with a first spectroscopic model; The first spectral model is updated by receiving the master data set and the target data set. generating a training dataset including first data from the target dataset; A validation dataset containing secondary data from the master dataset and no master data set. Generate a training dataset and a validation dataset using cross-validation. generating a second spectroscopic model that is an update of the first spectroscopic model using the spatial data set; and a processor configured to provide a second spectroscopic model.

[0004] According to some embodiments, the method further comprises the step of: generating a target associated with a first spectroscopic model by the device; receiving a target data set of the target population; and obtaining a master data set of the first spectral model based on receiving the set; Steps to find optimal partial least squares (PLS) factors using cross-validation with The optimal PLS factors are Multiple training datasets, each containing the entire dataset, and a target dataset Multiple validation runs that include parts of the data set and do not include data from the master data set and a step of determining the target data set and the merging the master data set and the new data set to generate a merged data set; The merged data set and the optimal PLS factors are used to generate a second spectroscopic model. The second spectral model is an update of the first spectral model; and providing a second spectroscopic model to replace the first spectroscopic model.

[0005] According to some embodiments, a non-transitory computer readable medium stores one or more instructions. The one or more instructions may be executed by one or more processors of the device. Upon receiving the first spectroscopic model, the one or more processors receive a master data set of the first spectroscopic model. receiving a target data set for a target population associated with the first spectroscopic model; The first spectral model is updated based on the master data set and the target data set. Generate multiple training datasets and select a master dataset based on the target dataset. Generate multiple validation datasets that do not contain data from the dataset, and Based on a training dataset and multiple validation datasets and cross-balanced Use reduction to find the model settings, the target dataset, and and generating a second spectral model based on the master data set, and providing the second spectral model. It can be done. [Brief explanation of the drawings]

[0006] [Figure 1A] 1 is a schematic diagram of an exemplary embodiment described herein. [Figure 1B] 1 is a schematic diagram of an exemplary embodiment described herein. [Figure 1C] 1 is a schematic diagram of an exemplary embodiment described herein. [Figure 1D] 1 is a schematic diagram of an exemplary embodiment described herein. [Figure 1E] 1 is a schematic diagram of an exemplary embodiment described herein. [Figure 2] FIG. 1 illustrates an example environment in which the systems and / or methods described herein may be implemented. [Figure 3] FIG. 3 is a diagram of example components of one or more devices of FIG. 2. [Figure 4] 1 is a flowchart of an exemplary process for cross-validation-based calibration of a spectroscopic model. [Figure 5] 1 is a flowchart of an exemplary process for cross-validation-based calibration of a spectroscopic model. [Figure 6] 1 is a flowchart of an exemplary process for cross-validation-based calibration of a spectroscopic model. DETAILED DESCRIPTION OF THE INVENTION

[0007] In the following detailed description of the embodiments, reference is made to the accompanying drawings, in which like reference numerals in different figures indicate: In the following description, a spectrometer is used as an example. The calibration principles, procedures, and methods described herein may be applied to, but are not limited to, other optical and spectroscopic sensors. The present invention can be used with any sensor, including a

[0008] A raw material identification (RMID) identifies the components (e.g., components) of a particular sample for identification, verification, etc. For example, RMID can be used to identify pharmaceutical materials. It can be verified that the ingredients correspond to the set of ingredients identified on the label. Similarly, raw material quantification involves determining the concentration of a specific material in a specific sample. A spectrometer is used to measure the concentration of a sample (e.g., a pharmaceutical material) Spectroscopy can be performed on the sample to determine the components of the sample, the concentration of the components in the sample, etc. The meter can obtain a set of measurements of the sample and interpret the set of measurements as spectroscopic interpretations. Spectral classification methods (e.g., classification The fire) can facilitate the determination of the components of a sample based on a set of measurements of the sample. .

[0009] A spectroscopic model is used to analyze one or more of the unknown samples to perform spectroscopic classification or quantification. The measurements can be evaluated. For example, the controller can evaluate one or more measurements of the unknown sample. corresponding to a particular class of spectral model, a particular level and / or amount associated with the spectral model, etc. However, if the raw materials change over time, As a result, spectral models can be inaccurate. For example, spectral classifications applied to agricultural products Regarding the results, different harvest years and associated yields may produce different spectra. in the master data set (e.g., the initial spectroscopic measurement set of the initial population at the initial time point) as The trained spectroscopic model is then applied to the target dataset (e.g., a subsequent population at a later time point). This may be inaccurate when applied to subsequent spectroscopic measurements of the same set.

[0010] In other cases, a master data set for each spectrometer is used to train the spectroscopic model for each spectrometer. As a result, the controller must maintain a master data set. train a single spectroscopic model and then scale that single spectrometer for use with many different spectrometers However, different spectrometers may have different associated calibration and / or operational models. As a result, the master spectroscopic measurement performed by the first spectrometer may be different. The spectroscopic model trained using the dataset is then used to analyze the data performed by the second spectrometer. It may be inaccurate when applied to a target data set of light measurements.

[0011] Some embodiments described herein utilize cross-validation techniques to validate spectroscopic models. Enables calibration updates and calibration transitions of the data from the target dataset. merged with data from a master dataset to enable the generation of new spectroscopic models In this case, data from the master data set can be used to trace the spectroscopic model. The data from the target dataset is used as the training set for the training Used for both the training set and the validation set for the spectroscopic model validation. Thus, compared to other approaches for model generation and / or model updating, The accuracy of the optical model is improved. Furthermore, the accuracy of the transferred spectral model is improved. Reducing the need to acquire a master data set for each spectrometer reduces the costs associated with spectrometer deployment. This reduces costs.

[0012] 1A-1E are diagrams of an exemplary embodiment 100 described herein. As shown, the exemplary embodiment 100 includes a first spectrometer 102 and a first controller 104. .

[0013] As further shown in FIG. 1A at 150, the first controller 104 controls the first spectrometer 12. 0 to cause the first spectrometer 120 to perform a set of spectroscopic measurements on the master population 152. For example, the first control device 104 may use a classification model to classify each class. The first spectrometer 120 measures the amount of the sample to be quantified using the quantification model. The classes of the classification model are lactose substances (in the pharmaceutical context), Fructose substances, acetaminophen substances, ibuprofen substances, aspirin substances, etc. It can refer to a grouping of similar materials that share one or more characteristics. The materials used in the manufacturing process and for which raw material identification is to be performed using a classification model are those of interest. They may be referred to as materials of interest.

[0014] As further shown in FIG. 1A by reference numerals 154 and 156, the first spectrometer 102 a first controller configured to execute a set of spectroscopic measurements and to transmit the set of spectroscopic measurements to the first controller for processing; For example, the first spectrometer 102 may provide a master population 152 A spectrum of each sample is obtained, and the first controller 104 calculates the unknown sample as a material of interest for a quantification model. Classification as one of the categories or as having a particular quantity in terms of a quantification model This may allow for the generation of a set of

[0015] As further shown in FIG. 1A at 158, the first control device 104 may include a master data For example, the first control device 104 may generate a first spectral model based on the set. generating a first spectroscopic model based on the set of spectroscopic measurements using a particular determination technique; In some embodiments, the first control unit 104 may use a support vector machine ( Generate a quantitative model using SVM (e.g., machine learning method for information judgment) Additionally or alternatively, the first controller 104 may be configured to perform another type of quantitative measurement. A quantification model can be generated using a quantification technique.

[0016] Quantification models are concerned with assigning specific spectra to specific quantity classes of the material of interest. In some embodiments, the quantification model may include information related to a particular quantity class. The first control device 1 may include information relating to the identification of the type of material of interest associated with the first control device 1. 04 is based on assigning the spectrum of an unknown sample to a specific quantity class of a quantification model. The output of the spectroscopy can provide information that identifies the amount of material in the unknown sample.

[0017] As shown in FIG. 1B by reference numeral 160, the second control device 104 controls the first spectral model. For example, the second control device 104 can receive information related to the first spectral model, The second controller may receive a star data set, etc. 104 may be associated with a different spectrometer than the first controller 104. For example, in the case of a calibration transition The second controller 104 is used in conjunction with the second spectrometer 102 (e.g., the target spectrometer). and receiving information related to the first spectroscopic model to operate the first spectrometer 102 ( For example, it can allow for calibration transfer from a master spectrometer to a second spectrometer 102. In this case, a second controller 104 and a second spectrometer, as described in more detail herein, 102 can perform measurements of the target population to generate a second spectral model. Alternatively, for calibration updates, the first spectroscopic model may be updated as described in more detail herein. Instead of transferring the signal to the second controller 104, the first controller 104 and the first spectrometer 102 Measurements of the target population may be performed to generate a second spectral model.

[0018] As further shown in FIG. 1B at 162, the second controller 104 controls the second spectrometer 10. 2 to send a command to the second spectrometer 120 to perform a set of spectroscopic measurements of the target population 164. For example, the second control device 104 may determine the first spectral model based on the reception of the first spectral model. 2. The spectrometer 102 can be caused to perform spectroscopic measurements of the target population 164. In some embodiments, the second controller 104 may determine an update or calibration of the first spectral model. and triggering the second spectrometer 102 to perform a set of spectroscopic measurements. In this case, the second control device 104 controls the second spectral model to enable generation of the second spectral model. 1 can communicate with the controller 104 to obtain information identifying the master data set. do.

[0019] In some embodiments, the target population 164 corresponds to the master population 152. For example, the target population 164 may contain the same classes as those included in the master population 152. In this case, the target population 164 may be an additional sample from the group of samples collected or measured. The sample may differ from the master population 152 with respect to time, location, environmental conditions, etc. Or alternatively, the target population 164 may be measured using a different spectrometer (e.g., For example, the master population 152 is measured by the second spectrometer 102 rather than the first spectrometer 102. The population may differ from the master population 152 based on the number of samples (measured).

[0020] As further shown in FIG. 1B by reference numerals 166 and 168, the second spectrometer 102 The spectroscopic measurement set can be executed, and information identifying the spectroscopic measurement set can be transmitted to the second control device. For example, the second spectrometer 102 may provide a target population 164 and (e.g., as a target data set) Information identifying the spectroscopic measurement may be provided to the second controller 104 for processing.

[0021] As shown in FIG. 1C at 170, the second controller 104 calculates the overall performance metric ( For example, the second control device 104 can calculate a total performance metric. The data is divided into multiple folds and multiple performance metrics are calculated for the multiple folds. Therefore, the performance metrics are aggregated to obtain a root mean square error (RMSE) value, and the RM The partial least squares (PLS) factors were optimized to minimize the SE value (referred to as optimal PLS factors). Based on the results of the analysis, a comprehensive performance metric can be calculated. Evaluate the accuracy of candidate models on the training set and prediction data to generate a model A subgroup of data for cross-validation containing a validation set for In another example, the second controller 104 may be configured to calculate principal component regression (PCR) factors, support Other types of optimized models, such as support vector regression (SVR) and modeling with respect to factors, In some embodiments, the second controller 104 may determine the pre-processing settings. For example, the second controller 104 may use the The optimized pre-processing parameters can be obtained.

[0022] In some embodiments, the second controller 104 may use training data for each fold. For example, the second control device 1 can be assigned to a test set or a validation set. 04 is a set of training sets 1 to N for N folds and a set of training sets 1 to N for N folds. For each code, multiple corresponding validation sets 1 to N can be obtained. In some embodiments, the training set comprises a master data set and a target data set. It may include merged data generated by merging sets. A training set (e.g., training set 1) is a set of all data from the master dataset. data (e.g., MDS) and a portion of the data from the target dataset (e.g., TDS1 ,TS ) In this case, the corresponding validation set may include the target data The corresponding part of the data from the set (e.g., TDS 1,TS ) and the master data set The corresponding validation set does not contain data from the training set. Data obtained from repeated scans of the same physical sample contained within the sample can be omitted.

[0023] Based on the allocation of data to multiple fields, the second control unit 104 For example, the second controller 104 may determine a performance metric for each field. The performance metrics for each field can be aggregated to obtain an overall performance metric. For example, The second control unit 104 can calculate the PLS factors for each fold and RMSE can be calculated for each PLS factor for each fold. Based on determining the RMSE value, the second control unit 104 can determine a total RMSE value. For example, the second control unit 104 may calculate the PLS coefficients as a function of all the PLS factors of all the folds. In this case, the RMSE value can be calculated based on the total RMSE value. The second control unit 104 selects the optimal PLS factors, which may be the PLS factors having the lowest RMSE value. You can ask for it.

[0024] In this case, the master data is added to the N-fold training set during cross-validation. The corresponding validation set contains the target dataset and the validation set. The accuracy of the second spectroscopic model is higher than that of other methods, based on the fact that it contains only a single dataset. For example, such a method can use the first spectral model without updating, To obtain PLS performance metrics using only the target dataset, All data in the master dataset is merged to create a merged dataset. Generate a merged dataset for both the training and validation sets. This can improve accuracy compared to using division of the grid.

[0025] As shown in FIG. 1D by reference numeral 172, the second controller 104 generates a second spectral model. For example, the second control unit 104 may store a master data set (MDS), Generate a second spectroscopic model using the target data set (TDS) and the optimal PLS factors. In this way, the second control device 104 can control the calibration spectroscopic model, the updated spectroscopic model, and the This may enable the generation of spectral models, transitional spectroscopic models, etc.

[0026] In some embodiments, the second controller 104 may be configured to generate a master data set and a target data set. Merge the data sets to create a merged data set (e.g., a second spectral model) For example, the second control device 1 can generate a final training set for training. 04 is the merged data that aggregates the master data set and the target data set. Based on the generation of the merged data set, the second control device 104 selects the merged dataset and the optimal PLS factors (e.g., the one with the lowest RMSE value). For example, the second control device 104 can generate the second spectral model by using A quantitative model generation method is used to generate a merged dataset (e.g., a second spectroscopic model) Generate a second spectral model in relation to the optimal PLS factors (which may be a training set). In this way, it is possible to find the optimal PLS factors without using a merged data set. By combining the optimal PLS factors from with the merged data set, the controller 104 , obtaining more accurate spectroscopic models than other methods.

[0027] In some embodiments, the second controller 104 determines a second spectral model based on the generation of the second spectral model. For example, the second control unit 104 may provide a spectral model via a data structure. and providing a second spectroscopic model for subsequent storage, deployment on one or more other spectrometers, etc. Additionally or alternatively, the second control device 104 may determine whether the second spectral model is based on the generated second spectral model. For example, the second spectral model may be provided based on the second spectral model. As will be described in the second section, the second controller 104 controls the second spectroscopic model to analyze the unknown sample. Based on the use, it can provide information to quantify unknown samples.

[0028] As shown in FIG. 1E by reference numeral 174, the second controller 104 commands the second spectrometer 102. A command can be sent to cause the second spectrometer 102 to perform a set of spectroscopic measurements on the unknown sample 176. For example, after generating the second spectroscopic model, the second control device 104 can Spectroscopic measurements can be performed on the unknown sample 176 .

[0029] As further shown in FIG. 1E at 178 and 180, the second spectrometer 102 The spectroscopic measurement set can be executed, and information identifying the spectroscopic measurement set can be transmitted to the second control device. For example, the second spectrometer 102 may provide a spectrum of the unknown sample 176. The signal can be measured and the spectrum can be identified for classification and / or quantification. The information can be provided to the second controller 104 .

[0030] As further shown in FIG. 1E at 182, the second controller 104 controls the second spectral model For example, the second control device 10 can be used to perform spectroscopic analysis of the spectroscopic measurement set. 4. Classifying the unknown sample 176 and / or quantifying the unknown sample 176 using the second spectroscopic model. In this case, the second control device 104 can determine the classification and / or quantification. In this way, the second control device 104 can provide an output that distinguishes between the second spectral model and the A second spectroscopic model is used based on the generation of the spectral

[0031] As noted above, FIGS. 1A-1E are provided only as one or more examples. Other examples may differ from those described with respect to Figures 1A-1E.

[0032] FIG. 2 is a diagram of an example environment 200 in which the systems and / or methods described herein may be implemented. As shown in FIG. 2, the environment 200 includes a control device 210, a spectrometer 220, a network The devices in the environment 200 may include wired, wireless, or both wired and wireless connections. The interconnections may be made via a combination of connections.

[0033] The controller 210 may store, process, and / or route information related to spectral classification. For example, the control device 210 may include one or more devices for generating the training set measurements. Generate and validate a spectroscopic model (e.g., a classification model or a quantification model) based on the constant set. Validate the spectroscopic model based on a measurement set of the inversion set and / or utilize the spectroscopic model and a server, computer, or software to perform spectroscopic analysis based on a measurement set of unknown samples. These may include wireless devices, cloud computing devices, etc. In some embodiments, the controller 210 may be associated with a particular spectrometer 220. , the controller 210 may be associated with multiple spectrometers 220. In some embodiments, the controller The device 210 receives information from and / or receives information from another device in the environment 200, such as a spectrometer 220. Information can be sent to each of them.

[0034] Spectrometer 220 includes one or more devices capable of performing spectroscopic measurements on a sample. For example, the spectrometer 220 may be a spectrometer that measures light emitted by a spectrometer (e.g., near-infrared (NIR) spectroscopy, mid-infrared spectroscopy (MIIR) spectroscopy, or other optical or optical techniques. Some spectroscopic instruments may perform spectroscopy (such as d-IR, Raman spectroscopy, etc.). In an embodiment, the spectrometer 220 is incorporated into a wearable device, such as a wearable spectrometer. In some embodiments, the spectrometer 220 may be connected to an environment 2, such as a control device 210. It can receive information from and / or send information to other devices within the network.

[0035] Network 230 may include one or more wired and / or wireless networks. For example, network 230 may be a cellular network (e.g., a Long Term Evolution LTE networks, 3G networks, Code Division Multiple Access (CDMA) networks Networks, Public Land Mobile Networks (PLMNs), Local Area Networks LAN, Wide Area Network (WAN), Metropolitan Area Network Networks (MAN), telephone networks (e.g., public switched telephone networks (PSTN)), private networks Network, ad-hoc network, intranet, internet, fiber optic network fiber optic-based networks, cloud computing networks, etc., and / or combinations of these or other types of networks.

[0036] The number and arrangement of devices and networks shown in FIG. 2 are provided as one or more examples. In practice, additional devices and / or networks may be used, or fewer devices may be used, than those shown in FIG. No devices and / or networks, different devices and / or networks, or different configurations Furthermore, two or more of the devices shown in FIG. The present invention may be implemented in a single device, or the single device shown in FIG. 2 may be implemented as multiple distributed devices. Additionally or alternatively, a set of devices in the environment 200 (e.g., one or A plurality of devices) may be implemented by one or more devices of the environment 200. may perform multiple functions.

[0037] 3 is a diagram of exemplary components of the device 300. The device 300 includes a control device 2 10 and / or spectrometer 220. In some embodiments, controller 210 and and / or spectrometer 220 may be connected to one or more devices 300 and / or one or more of devices 300. As shown in FIG. 3, the device 300 includes a bus 310, a processor 320, and a 320, memory 330, storage component 340, input component 35 0, output component 360, and communication interface 370.

[0038] Bus 310 is a component bus that allows communication between multiple components of device 300. The processor 320 may be implemented in hardware, firmware, and / or hardware. The processor 320 is implemented as a combination of hardware and software. processing unit (CPU), graphics processing unit (GPU), Advanced Processing Unit (APU), microprocessor, microcomputer controllers, digital signal processors (DSPs), field programmable gate arrays (FPGA), application specific integrated circuit (ASIC), or another type of processing component In some embodiments, the processor 320 is programmed to perform functions. The memory 330 includes one or more processors capable of running the program. RAM), read-only memory (ROM), and / or other memory for use by processor 320 Another type of dynamic or static storage for storing information and / or instructions devices (e.g., flash memory, magnetic memory, and / or optical memory).

[0039] The storage component 340 stores information and / or For example, the storage component 340 may be a hard disk drive. disks (e.g., magnetic disks, optical disks, and / or magneto-optical disks), solid-state Drive (SSD), Compact Disc (CD), Digital Versatile Disc (DVD) , floppy disk, cartridge, magnetic tape, and / or another type of non-transitory The computer readable medium may be included along with a corresponding drive.

[0040] Input component 350 allows device 300 to receive user input (e.g., a touchscreen device). Displays, keyboards, keypads, mice, buttons, switches, and / or microphones It also includes components that allow you to receive information via a mobile phone, etc. Or alternatively, the input component 350 may be a component that determines location (e.g., , Global Positioning System (GPS) components) and / or sensors (e.g., accelerometers , gyroscopes, actuators, other types of position or environmental sensors, etc.) An output component 360 outputs information from the device 300 (e.g., a display, (e.g., via speakers, haptic feedback components, auditory or visual indicators) Contains the components that it provides.

[0041] The communication interface 370 may be configured to allow the device 300 to communicate via a wired connection, a wireless connection, or both. Components such as transceivers that allow communication with other devices via a combination of A communication interface includes a separate component (e.g., a transceiver, a separate receiver, a separate transmitter, etc.). The source 370 allows the device 300 to receive information from and / or provide information to another device. For example, the communication interface 370 may be an Ethernet interface. Interface, optical interface, coaxial interface, infrared interface, radio frequency ( RF interface, Universal Serial Bus (USB) interface, Wi-Fi This may include an i-interface, a cellular network interface, etc.

[0042] The device 300 may perform one or more of the processes described herein. The processor 320 may include non-uniform memory components such as memory 330 and / or storage components 340. Based on executing software instructions stored on a temporary computer-readable medium These processes can be performed by a computer readable medium. The term refers to a non-transitory memory device. A memory device is a single physical storage device. This includes memory space within a device or memory space spread across multiple physical storage devices.

[0043] The software instructions may be transmitted over a communications interface from another computer-readable medium or another device. 370 into the input memory 330 and / or storage component 340 When executed, the memory 330 and / or storage component 340 The software instructions stored in the processor 320 cause the processor 320 to perform one or more of the functions described herein. Additionally or alternatively, hardware circuitry may be implemented as software. may be used in place of or in combination with software instructions to implement the functions described herein. The embodiments described herein may therefore include one or more of the following processes: , is not limited to any specific combination of hardware circuitry and software.

[0044] The number and arrangement of components shown in Figure 3 are given as an example. The device 300 may include additional components and fewer components than those shown in FIG. It may contain additional components, different components, or components in different arrangements. Alternatively, a set of components of the device 300 (e.g., one or more component) is described as being performed by another set of components of the device 300. The device may perform one or more of the functions described above.

[0045] FIG. 4 illustrates an example process 400 for cross-validation-based calibration of a spectroscopic model. In some embodiments, one or more blocks in FIG. In some embodiments, the first embodiment of FIG. One or more blocks may be separate from or include a control device, such as a spectrometer (e.g., This may be performed by another device or devices, such as a photometer 220.

[0046] As shown in FIG. 4, the process 400 includes receiving a master data set for a first spectroscopic model. For example, as described above, the control device may include (e.g., For example, a processor 320, a memory 330, a storage component 340, an input component component 350, output component 360, communication interface 370, etc.), A master data set for one spectroscopic model may be received.

[0047] As further shown in FIG. 4, the process 400 includes generating a target matrix associated with the first spectral model. receiving a target data set of the population to update the first spectroscopic model. (Block 420). For example, as described above, the control unit (e.g., processor 320, memory 330, storage component 340, input component 350, output component component 360, communication interface 370, etc.), related to the first spectroscopic model A target data set for the target population may be received to update the first spectroscopic model.

[0048] As further shown in FIG. 4, the process 400 includes a master data set and a target data set. generating a training data set including first data from the training data set; For example, as described above, the controller (e.g., processor 32) 0, memory 330, storage component 340, input component 350, output component 360, communication interface 370, etc.), master data set and a first data set from the target dataset. obtain.

[0049] As further shown in FIG. 4, the process 400 may include extracting second data from the target data set. Steps to generate a validation dataset that contains the master data but does not include the master dataset For example, as described above, the controller may include a processor (e.g., 320, memory 330, storage component 340, input component 35 0, output component 360, communication interface 370, etc.), Validation data that contains second data from the master dataset and does not include the master dataset. A dataset can be generated.

[0050] As further shown in FIG. 4, the process 400 uses cross-validation and The first spectroscopic model was updated using the training and validation datasets. The method may include generating a new second spectral model (block 450). As described above, the control device (e.g., processor 320, memory 330, storage components, etc.) component 340, input component 350, output component 360, communication interface Using cross-validation and training data sets (e.g., using Face370), Using the test and validation dataset, the second spectroscopic model was developed as an update of the first spectroscopic model. A del can be generated.

[0051] As further shown in FIG. 4, the process 400 includes providing a second spectroscopic model. For example, as described above, the controller (e.g., processor 3) 20, memory 330, storage component 340, input component 350, output a second spectral model (using the input component 360, the communication interface 370, etc.) It can be provided.

[0052] Process 400 may be implemented using one or more of the methods described below and / or elsewhere herein. Any single embodiment or any combination of embodiments described in relation to any number of other processes. Further embodiments may include combinations, etc.

[0053] In a first embodiment, the process 400 includes steps of receiving spectroscopic measurements and transmitting a second spectroscopic measurement to a second spectroscopic monitor. performing a spectroscopic determination using the model; and providing an output identifying the spectroscopic determination. Includes the following:

[0054] In a second embodiment, alone or in combination with the first embodiment, the training data set are multiple training datasets, and the validation dataset is multiple validators. The process 400 includes a plurality of training datasets and generating a plurality of performance metrics based on a plurality of validation data sets; determining an overall performance metric based on a plurality of performance metrics; A step of finding the optimal partial least squares (PLS) factors based on the algorithm, and a step of finding the optimal PLS factors and and determining a second spectroscopic model based on the merged data set, The read dataset includes a master dataset and a target dataset.

[0055] In a third embodiment, alone or in combination with one or both of the first and second embodiments, The first spectroscopic model and the second spectroscopic model are quantification models.

[0056] In a fourth embodiment, alone or in combination with one or more of the first to third embodiments, The target data set is based on a first set of spectroscopic measurements performed by the master spectrometer. The target data set was run with a target spectrometer that was different from the master spectrometer. Based on the second set of spectroscopic measurements.

[0057] In a fifth embodiment, alone or in combination with one or more of the first to fourth embodiments, The target data set is based on a first set of spectroscopic measurements performed by a particular spectrometer. The get data set is based on a second set of spectroscopic measurements performed by a particular spectrometer.

[0058] While FIG. 4 illustrates exemplary blocks of process 400, in some embodiments , process 400 may include additional blocks, fewer blocks, and different Additionally or alternatively, the block may include blocks of different configurations. Therefore, two or more of the blocks of process 400 may be performed in parallel.

[0059] FIG. 5 illustrates an example process 500 for cross-validation-based calibration of a spectroscopic model. In some embodiments, one or more of the process blocks of FIG. may be performed by a controller (e.g., controller 210). One or more block processes in FIG. 5 may be separate from or include the control device. This may be performed by another device or devices, such as a photometer (eg, spectrometer 220).

[0060] As shown in FIG. 5, the process 500 includes: The method may include receiving a target data set (block 510). As described above, the control device (e.g., processor 320, memory 330, storage components, etc.) component 340, input component 350, output component 360, communication interface 370, etc.), to obtain the target data of the target population associated with the first spectroscopic model. The data set may be received.

[0061] As further shown in FIG. 5, the process 500 may include, based on receipt of the target data set: Obtaining a master data set for the first spectroscopic model (block 520 For example, as described above, the control unit (e.g., processor 320, memory 330, Storage component 340, input component 350, output component 360 , using a communications interface 370, etc., based on receipt of the target data set. A master data set for the spectroscopic model may be obtained.

[0062] As further shown in FIG. 5, the process 500 may use cross-validation to find the optimal part. This is the step to find the least squares (PLS) factors. The optimal PLS factors are Multiple training data sets, including portions of the master dataset and the entire master dataset, are used. and a master dataset, each containing a portion of the target dataset. and multiple validation datasets that do not include the data of For example, as described above, the controller (e.g., a processor) may include: 320, memory 330, storage component 340, input component 350, output component 360, communication interface 370, etc.), The optimal partial least squares (PLS) factors may be determined using the eigenvalue analysis. The LS factors are the sum of the parts of the target dataset and the whole of the master dataset. Multiple training datasets containing parts of the target dataset, each of which is a part of the target dataset Multiple validation datasets that contain data from the master dataset and do not contain data from the master dataset. It is required based on the following.

[0063] As further shown in FIG. 5, the process 500 includes a target data set and a master data set. The method may include merging the data sets to generate a merged data set (block For example, as described above, the control device (e.g., processor 320, memory 3 30, storage component 340, input component 350, output component 360, communication interface 370, etc.), the target data set and the master data set. The data sets may be merged to generate a merged data set.

[0064] As further shown in FIG. 5, the process 500 includes a merged dataset and an optimal PLS using the factors to generate a second spectral model that is an update of the first spectral model. For example, as described above, the controller (e.g., processor 320) , memory 330, storage component 340, input component 350, output component component 360, communication interface 370, etc.), the merged data set and the optimal PLS factors may be used to generate a second spectroscopic model. The optical model is an update of the first spectral model.

[0065] As further shown in FIG. 5, the process 500 provides a second spectral model to compare the first spectral model. For example, as described above, the control device may include replacing the The components include (e.g., processor 320, memory 330, storage component 340, input input component 350, output component 360, communication interface 370, etc. In this case, a second spectral model can be provided to replace the first spectral model.

[0066] Process 500 may be implemented using one or more of the methods described below and / or elsewhere herein. Any single embodiment or any combination of embodiments described in relation to any number of other processes. Further embodiments may include combinations, etc.

[0067] In the first embodiment, the step of determining the optimized PLS factors is performed using a plurality of training data A partial minimum test is performed for each of the sets and each of the multiple validation data sets. A step of obtaining PLS performance metrics and an overall P based on the PLS performance metrics. determining a PLS performance metric and determining a second spectroscopic model based on the overall PLS performance metric; and optimizing the PLS factors of the model.

[0068] In a second embodiment, alone or in combination with the first embodiment, the overall PLS performance metric is related to the root mean square error (RMSE) value, and the step of optimizing the PLS factors is to Optimizing the PLS factors to minimize the E value is included.

[0069] In a third embodiment, alone or in combination with one or both of the first and second embodiments, multiple The multiple validation datasets are targeted at different targets than the multiple training datasets. Contains data from the NET dataset.

[0070] In a fourth embodiment, alone or in combination with one or more of the first to third embodiments, The step of determining the PLS performance metrics is performed by aggregating the PLS performance metrics. include.

[0071] In a fifth embodiment, alone or in combination with one or more of the first to fourth embodiments, The get dataset is a set of target measurements performed after the measurement associated with the master dataset. Relates to a set of measurements of a population.

[0072] In a sixth embodiment, alone or in combination with one or more of the first to fifth embodiments, The spectroscopic model is a calibration update model of the first spectroscopic model.

[0073] In a seventh embodiment, alone or in combination with one or more of the first to sixth embodiments, The acquired data set is one or more data sets that have performed measurements relative to the master data set. It refers to a set of measurements performed by a particular spectrometer, different from the spectrometer.

[0074] In an eighth embodiment, alone or in combination with one or more of the first to seventh embodiments, The spectroscopic model is a calibrated transition model of the first spectroscopic model.

[0075] In a ninth embodiment, alone or in combination with one or more of the first to eighth embodiments, The step of providing a spectroscopic model includes providing a spectroscopic model for use in connection with subsequent measurements by a particular spectrometer. providing a second spectroscopic model of

[0076] FIG. 5 illustrates exemplary blocks of process 500, although in some embodiments , the process 500 may include additional blocks, fewer blocks, and different Additionally or alternatively, the block may include blocks of different configurations. Therefore, two or more of the blocks of process 500 may be performed in parallel.

[0077] FIG. 6 illustrates an example process 600 for cross-validation-based calibration of a spectroscopic model. In some embodiments, one or more process blocks of FIG. This may be performed by a controller (e.g., controller 210). 6. One or more process blocks are separate from or include a control device, The measurement may be performed by another device or devices, such as a spectrometer (e.g., spectrometer 220).

[0078] As shown in FIG. 6, the process 600 includes receiving a master data set for a first spectroscopic model. a target data set for the target population associated with the first spectroscopic model. and updating the first spectral model by receiving the master data set and the target generating a plurality of training datasets based on the target dataset; Multiple validations based on a target dataset and not including data from the master dataset and generating a data set (block 610). The controller (e.g., processor 320, memory 330, storage components 340, input component 350, output component 360, communication interface 3 70, etc.), receives the master data set of the first spectroscopic model, and receiving a target data set for an associated target population to update the first spectral model; , multiple training data based on the master dataset and the target dataset Generate a set based on the target dataset and include the data from the master dataset. In some embodiments, multiple validation datasets may be generated. The data set does not contain the data from the master data set.

[0079] As shown in FIG. 6, the process 600 includes receiving a master data set for a first spectroscopic model. For example, as described above, the control device may include (e.g., For example, a processor 320, a memory 330, a storage component 340, an input component component 350, output component 360, communication interface 370, etc.), A master data set for one spectroscopic model may be received.

[0080] As shown in FIG. 6, a process 600 includes: The method may include receiving a target data set and updating the first spectral model (block For example, as described above, the control unit (e.g., processor 320, memory 330, storage component 340, input component 350, output component 360, communication interface 370, etc.), a target associated with the first spectral model A target data set of the target population may be received to update the first spectroscopic model.

[0081] As shown in FIG. 6, the process 600 includes a master data set and a target data set. generating a plurality of training data sets based on the set (block For example, as described above, the control unit (e.g., processor 320, memory 330) 0, storage component 340, input component 350, output component 360, communication interface 370, etc.), the master data set and the target A plurality of training data sets may be generated based on the data set.

[0082] As shown in FIG. 6, a process 600 generates a master data set based on a target data set. generating multiple validation datasets that do not include data from the original dataset. For example, as described above, the controller (e.g., processor 3) 20, memory 330, storage component 340, input component 350, output (using the input component 360, the communication interface 370, etc.) In some embodiments, multiple validation data sets may be generated based on the set. The numerical validation dataset does not include data from the master dataset.

[0083] As further shown in FIG. 6, the process 600 includes multiple training data sets and multiple Model building based on multiple validation datasets and using cross-validation For example, as described above, the controller may determine (block 650). For example, a processor 320, a memory 330, a storage component 340, an input component component 350, output component 360, communication interface 370, etc.) A model is constructed based on multiple training datasets and multiple validation datasets. The rule settings can be obtained.

[0084] As further shown in Figure 6, the process involves model configuration, target dataset, and generating a second spectral model based on the star data set (Block 6 60). For example, as described above, the control unit (e.g., processor 320, memory 330, , storage component 340, input component 350, output component 3 60, communication interface 370, etc.), model configuration, target data set, and a second spectroscopic model may be generated based on the master data set.

[0085] As further shown in FIG. 6, the process 600 includes providing a second spectroscopic model. For example, as described above, the controller (e.g., processor 3) may 20, memory 330, storage component 340, input component 350, output a second spectral model (using the input component 360, the communication interface 370, etc.) It can be provided.

[0086] Process 600 may be implemented using one or more of the methods described below and / or elsewhere herein. Any single embodiment or any combination of embodiments described in relation to any number of other processes. Further embodiments may include combinations, etc.

[0087] In the first embodiment, the model setting is performed by using partial least squares (PLS) factors, principal components, and The amount of components in the PCR model and the SVR performance of the SVR model parameters, or preprocessing settings.

[0088] In a second embodiment, alone or in combination with the first embodiment, the process 600 comprises: Each of the training datasets and the corresponding one of the validation datasets generating a plurality of partial performance metrics for the validation data set; aggregating a plurality of partial performance metrics to generate an overall performance metric; and finding a model configuration that minimizes the error value of the performance metric.

[0089] Third embodiment alone or in combination with one or both of the first and second embodiments In step 600, the process performs a spectroscopic determination based on the measurements and using a second spectroscopic model. and providing an output identifying the spectroscopic determination.

[0090] FIG. 6 illustrates exemplary blocks of a process 600, although in some embodiments , process 600 may include additional blocks, fewer blocks, and different Additionally or alternatively, the block may include blocks of different configurations. Therefore, two or more of the blocks of process 600 may be performed in parallel.

[0091] While the above disclosure has been illustrated and described, it is not intended to be exhaustive or to disclose embodiments. Nor are they intended to be limited to the precise forms disclosed. Modifications and variations are possible in light of the above disclosure. It is possible to do so or can be obtained from the practice of the embodiments.

[0092] As used herein, the term "component" refers to hardware, firmware, or any combination of hardware and software. That is why.

[0093] As used herein, meeting a threshold may mean, depending on the context, that a value exceeds a threshold, that a value is below a threshold, or that a value is below a threshold. Greater than, higher than threshold, equal to or greater than threshold, less than threshold, It can refer to being less than, lower than a threshold, being equal to or less than a threshold, or being equal to a threshold. .

[0094] The systems and / or methods described herein may be implemented in different forms of hardware, firmware, It is clear that the present invention may be implemented in software and / or a combination of hardware and software. The actual specific control hardware used to implement these systems and / or methods may vary. The hardware or software code does not limit the embodiments. The operations and behavior of the systems and / or methods described herein are expressed as specific software code. Regardless of the software hardware, the system and / or method described herein It is understood that the present invention can be designed to implement the above.

[0095] Specific feature combinations may be claimed and / or disclosed herein. However, this is not intended to limit the disclosure of embodiments in which these combinations are possible. In some cases, many of these features are specifically claimed and / or disclosed herein. Each of the appended dependent claims may be combined in ways not shown. Although the disclosure of various embodiments may depend directly only on the subject matter of the invention, each dependent claim should be considered to be a separate and distinct part of the invention. It includes any combination with any other claim in the claim set.

[0096] Any element, act, or instruction used herein is abbreviated to "unclear" unless expressly stated as such. These should not be construed as critical or essential. The definite articles "a" and "an" are intended to include one or more items, and "one or more" is used interchangeably. Furthermore, in this specification, the term "set" A word can refer to one or more items (e.g., related items, unrelated items, related and unrelated items). "one or more" is intended to include "one or more" and "one or more combinations" Where only one item is intended, the phrase "only one" or similar language shall be used. In addition, as used herein, the terms "has," "have," and "having" and similar terms are intended to be open-ended terms. The phrase "based at least in part on" is intended to mean "based at least in part on" unless otherwise specified. Also, in this specification, when the term "or" is used in a list, Cases are intended to be inclusive and unless otherwise specified (e.g., "any" or (when used in conjunction with "only one of") is used interchangeably with "and / or" It is possible.

Claims

1. one or more memories; one or more processors communicatively coupled to the one or more memories; receiving a master data set for a first spectroscopic model; receiving a target data set for a target population associated with the first spectral model; updating the first spectroscopic model; a first data set from the master data set and a first data set from the target data set; Generate a training dataset, including second data from the target data set and including the master data set. Generate a validation dataset that does not contain Using cross-validation, the training data set and the validator and generating a second spectral model that is an update of the first spectral model using the spatial data set. And providing the second spectroscopic model; and a processor configured to A device comprising:

2. 10. The apparatus of claim 1, wherein the one or more processors: receiving spectroscopic measurements; performing a spectral determination using the second spectral model; and providing an output identifying said spectroscopic determination; The apparatus is configured to:

3. 10. The apparatus of claim 1, wherein the training data set comprises a plurality of training a validation dataset, the validation dataset being a plurality of validation datasets; It is a The one or more processors, when generating the second spectral model, the plurality of training data sets and the plurality of validation data sets; generating a plurality of performance metrics based on the determining an overall performance metric based on the plurality of performance metrics; determining optimal partial least squares (PLS) factors based on the overall performance metric; and determining the second spectroscopic model based on the optimal PLS factors and the merged data set; R the merged data set is a sum of the master data set and the target data set. A device containing a set of data.

4. 2. The apparatus according to claim 1, wherein the first spectroscopic model and the second spectroscopic model are quantified. The device that is the model.

5. 10. The apparatus of claim 1, wherein the master data set is generated by a master spectrometer. Based on the first set of spectroscopic measurements performed, the target data set is The apparatus is based on a second set of spectroscopic measurements performed by a target spectrometer different from the spectrometer.

6. 10. The apparatus of claim 1, wherein the master data set is generated by a specific spectrometer. Based on the first set of spectroscopic measurements performed, the target data set is The apparatus is based on a second set of spectroscopic measurements performed by the spectrometer.

7. receiving a target data set of a target population associated with a first spectral model by the apparatus; The step of believing, The apparatus calculates the mass of the first spectral model based on receiving the target data set. obtaining a target data set; The device uses cross-validation to find optimal partial least squares (PLS) factors. This is the step The optimal PLS factors are each a part of the target dataset and the matrix. a plurality of training datasets including all of the star datasets, each of which is A plurality of data sets each including a portion of the target data set and excluding data from the master data set. and a validation data set of Merging the target data set and the master data set by the device. generating a merged dataset; The apparatus performs a second spectroscopic analysis using the merged data set and the optimal PLS factors. A step of generating a model, the second spectroscopic model being an update of the first spectroscopic model; providing the second spectroscopic model by the device to replace the first spectroscopic model. Pu and A method comprising:

8. 8. The method of claim 7, wherein the step of determining the optimal PLS factors comprises: Each of the plurality of training data sets and the plurality of validation data determining a partial least squares (PLS) performance metric for each of the sets; determining an overall PLS performance metric based on the PLS performance metrics; determining the optimal PLS factors of the second spectral model based on the overall PLS performance metric; optimizing the PLS factors to A method comprising:

9. 9. The method of claim 8, wherein the overall PLS performance metric is the root mean square error (R MSE) value, The step of optimizing the PLS factors comprises: optimizing the PLS factors to minimize the RMSE value. A method comprising:

10. 9. The method of claim 8, wherein the plurality of validation data sets are the target dataset includes data different from the training dataset of 。

11. 9. The method of claim 8, wherein the step of determining the overall PLS performance metric comprises: aggregating the PLS performance metrics A method comprising:

12. 8. The method of claim 7, wherein the target data set is the master data. a method relating to a set of measurements of said target population performed after measurements relating to said set 。

13. 13. The method of claim 12, wherein the second spectral model is a calibration update of the first spectral model. How to be a new model.

14. 8. The method of claim 7, wherein the target data set is the master data. performed by a particular spectrometer that is different from the spectrometer or spectrometers that performed the measurements associated with the set The method associated with the set of measurements performed.

15. 15. The method of claim 14, wherein the second spectroscopic model is a calibration shift of the first spectroscopic model. How to be a row model.

16. 15. The method of claim 14, wherein the step of providing a second spectral model comprises: providing said second spectroscopic model for use in connection with subsequent measurements by said particular spectrometer; Steps to take A method comprising:

17. A non-transitory computer-readable medium storing instructions, the instructions comprising: When executed by one or more processors, the one or more processors receiving a master data set for a first spectroscopic model; receiving a target data set for a target population associated with the first spectral model; and updating the first spectroscopic model by A plurality of training data sets based on the master data set and the target data set Generate a training dataset, based on the target dataset and not including data from the master dataset Generate multiple validation datasets, Based on the plurality of training data sets and the plurality of validation data sets, and use cross-validation to find the model settings. Based on the model settings, the target dataset, and the master dataset generating a second spectral model based on the The second spectroscopic model is provided. A non-transitory computer-readable medium containing one or more instructions.

18. 20. The non-transitory computer-readable medium of claim 17, wherein the model configuration comprises: Partial least squares (PLS) factors of the PLS model, the amount of components in the principal component regression (PCR) model, the SVR parameters of a support vector regression (SVR) model, or Preprocessing settings a non-transitory computer-readable medium, the non-transitory computer-readable medium being at least one of:

19. 18. The non-transitory computer-readable medium of claim 17, wherein the one or more processors The one or more instructions that cause a processor to determine the model settings include the one or more to the processor Each of the plurality of training data sets and the plurality of validation data Generate multiple partial performance metrics for the corresponding validation datasets of the set Let, aggregating the plurality of partial performance metrics to generate an overall performance metric; and Finding the model configuration that minimizes the error value of the overall performance metric Non-transitory computer-readable medium.

20. 20. The non-transitory computer-readable medium of claim 17, wherein the one or more instructions The instructions, when executed by the one or more processors, cause the one or more processors to Receive the measurement value, performing a spectral determination based on the measurements and using the second spectral model; and providing an output identifying said spectroscopic determination; Non-transitory computer-readable medium.

Citation Information

Patent Citations

  • Sample composition determination method based on increment partial least square method

    CN105092519A

  • Apparatus and method for non-invasive measurement of blood components

    JP2005537891A

  • Diagnosis method of concrete

    JP2008014779A

  • Advanced analytical infrastructure for machine learning

    JP2017004509A

  • Identification using spectrometry

    JP2017049246A