Multicomponent regression / multicomponent analysis of time and / or spatial series files
The combination of multicomponent regression and search methods automates the analysis of evolving spectral data, overcoming time-consuming challenges in spectral matching by estimating pure components and correlating them with reference spectra, achieving rapid and consistent results.
Patent Information
- Application Number
- DE112012004324
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2011-10-17
- Filing Date
- 2012-10-17
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2032-10-17
AI Technical Summary
Existing spectral analysis methods for time-dependent data, such as chemical reaction monitoring and thermal analysis, are laborious and time-consuming, especially when comparing large reference libraries and performing quantitative analysis, often taking hours or days even with high-speed processors.
An automated method using multicomponent regression (MCR) to extract linearly independent spectra, followed by multicomponent search (MCS), which simplifies and accelerates the analysis by estimating pure components and correlating them with reference spectra, providing a rapid and complete analysis of evolving samples.
The method significantly reduces computational time and eliminates the need for user expertise, allowing for consistent, high-quality analysis of complex spectral data sets, revealing detailed time-dependent information on sample components.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND OF THE INVENTIONField of the invention
[0001] The present invention relates to the field of spectral analysis and, in particular, to the automatic identification of evolving time series spectra using multicomponent regression in combination with multicomponent spectral matching. Discussion of the state of the art
[0002] A molecular spectrometer (sometimes called a spectroscope) is an instrument in which a solid, liquid, or gaseous sample is illuminated, often with non-visible light, such as light in the infrared region of the spectrum. The light from the sample is then detected and analyzed to obtain information about the characteristics of the sample. For example, a sample may be illuminated with infrared light of a known intensity over a range of wavelengths, and the light transmitted through and / or reflected from the sample can then be detected for comparison with the light source. Examination of the detected spectra can then illustrate the wavelengths at which the illuminating light was absorbed by the sample.The spectrum, and in particular the positions and amplitudes of the peaks within it, can be compared with libraries of previously obtained reference spectra to obtain information about the sample, such as its composition and characteristics. The spectrum essentially serves as a "fingerprint" of the sample and the substances within it, and by comparing the fingerprint with one or more known fingerprints, the identity of the sample can be determined.
[0003] However, there are numerous cases where time-dependent data are collected using previously described methods, such as in chemical reaction monitoring (kinetics) or thermal analysis with gas emission (TGA-IR) or chromatography (GC-IR). The most laborious step in this analysis is the extraction of independent spectra from the concatenated series of spectra, followed by analysis of these individual spectra. In GC-IR, the spectra typically refer to pure components—the GC performs the separation—but in TGA-IR, the individual spectra can also be mixtures themselves.
[0004] It can thus be seen that if one wishes to compare a time series number of spectra of an evolving sample against all possible combinations of one or more reference spectra, this can typically exceed a large number, particularly when a large reference library may have tens of thousands of entries. The computational time required to perform these comparisons can be further increased if quantitative analysis is to be performed as well as qualitative analysis, i.e., if the relative proportions of component spectra within the unknown spectrum as well as their identities are to be determined. Such quantitative analysis may require regression between a combination of reference spectra against the time series of spectra to determine the weighting that each reference spectrum should have to result in a combination that represents a best fit.As a result, exhaustive spectral matching can sometimes take hours—or even days—even when using dedicated computers or other machines with high-speed processors.
[0005] Background information on a method for spectral matching of an unknown spectrum using multicomponent analysis, which is incorporated herein by reference in its entirety, is described in U.S. Patent No. 7,698,098 B2, entitled "EFFICIENT SPECTRAL MATCHING, PARTICULARLY FOR MULTICOMPENENT SPECTRA," issued April 13, 2010 to Ritter et al., and includes the following: "[an] unknown spectrum obtained by infrared or other spectroscopy can be compared with spectra in a reference library to find the best matches. The best-matching spectra can then be combined with the reference spectra, with the combinations also being screened for best matches against the unknown spectrum. These resulting best matches can then also be subjected to the preceding combining and comparing steps.The process can be repeated in this way until a suitable stopping point is reached, for example, when a desired number of best matches are identified, when some specified number of iterations have been performed, etc. This methodology can return the best-matching spectra (and combinations of spectra) with far fewer computational steps and greater speed than if all possible combinations of reference spectra are considered."
[0006] Background information on a method for component spectral analysis is described and claimed in U.S. Patent No. US 7,072,771 B1 entitled "METHOD FOR IDENTIFYING COMPONENTS OF A MIXTURE VIA SPECTRAL ANALYSIS", issued July 4, 2006 to Schweitzer et al., and includes the following: "[t]he present invention relates generally to the field of spectral analysis and, more particularly, to an improved method for identifying unknown components of a mixture from a set of spectra collected from the mixture using a spectral library including potential candidates.For example, the present method relates to identifying components of a mixture by the steps of obtaining a set of spectral data for the mixture, thereby defining a mixture data space; rank-ordering a plurality of library spectra of known elements according to their projection angle into the mixture data space; calculating a corrected correlation coefficient for each combination of the top y-ranked library spectra; and selecting the combination with the highest corrected correlation coefficient, wherein the known elements of the selected combination are identified as the components of the mixture."
[0007] US 2004 / 0 220 760 A1 relates to a method and a system for efficiently determining grating profiles using dynamic learning in a library generation process, as well as a method and a system for searching and matching experimental grating profiles to determine shape, profile and spectrum data information associated with an actual grating profile.
[0008] US 7 072 770 B1 relates to the field of spectral analysis and, in particular, to a method for identifying unknown components of a mixture from a set of spectra collected from the mixture using a spectral library containing potential candidates.
[0009] EP 2 245 443 B1 deals with the identification of unknown spectra obtained from spectrometer measurements and, in particular, identification by comparing the unknown spectra with reference spectra. SUMMARY OF THE INVENTION
[0010] The present invention relates to an automated method for analyzing a series data file resulting from one or more evolving samples. Through analysis, particularly using MCR or multicomponent regression, a series of linearly independent spectra can first be extracted. Technically, an MCR result is referred to as a "factor," and MCR often produces a set of factors, and these factors, when recombined, can reproduce the original dataset. MCR, as disclosed herein, can then be conducted to pass the factors to a multicomponent search (MCS) routine, which can deconvolute the factors for which provided databases have been searched. The end result of such a process allows the identification of each component that was present in the original dataset.
[0011] The routine can complete the analysis by performing a spectral correlation of the components identified with the original dataset. This is essentially done by comparing the component spectra with those in the original dataset and providing a value that indicates how much of that component is present at that time. The aggregation of this dataset, developed over time, produces a profile that represents the time history of each component's presence. This ultimately results in a sequence of profiles that demonstrates the time dependence of each component.
[0012] The final report can often be customized, if desired, to consist of the extracted spectra, the search results, and the profiles for each identified component. This overcomes several problems with existing technology: • All spectra in the database are processed for information extraction. • The user does not need to have any prior knowledge of the sample. • The user does not need to have any knowledge of the analysis software. • The speed of the final analysis is significantly accelerated.
[0013] A first aspect of the present application therefore includes a method for analyzing spectra from a developing sample, including: using a spectrometer to obtain a time-series set of spectra; estimating one or more qualitative and quantitative components of each of the time-series set of spectra using a computer via a regression method;and using a computer to pass the estimated one or more qualitative and quantitative components from each of the time series set of spectra to a multi-component search (MCS) algorithm configured to iteratively correlate one or more comparison spectra stored in one or more spectral libraries with each of the estimated time series sets of spectra represented as one or more respective qualitative and quantitative components, the result being an iteratively determined best-matching time series set of one or more candidate spectra.;
[0014] A second aspect of the present application includes a system for analyzing spectra from a developing sample, including: a spectrometer configured to generate a time-series set of spectra; and a computer configured to estimate one or more qualitative and quantitative components of each of the time-series set of spectra using a regression method;wherein the computer passes the estimated one or more qualitative and quantitative components from each of the time series set of spectra to a multi-component search (MCS) algorithm configured to iteratively correlate one or more comparison spectra stored in one or more spectral libraries with each of the estimated time series sets of spectra represented as one or more respective qualitative and quantitative components, the result being an iteratively determined best-matching time series set of one or more candidate spectra; BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1A shows a spectrum of a given time point of a serial time file of an exemplary sample. Fig. Figure 1B shows MCR-estimated pure component absorption spectra for carbon dioxide, ammonia, isocyanic acid, and water, obtained from the deconvolution of the spectrum from Fig. 1A result. Fig. Figure 1C shows quantification time profiles for the exemplary estimated pure components used in Fig. 1B are illustrated. Fig. Figure 2 generally illustrates an exemplary series time file of estimated pure components P1, P2, and P3 resulting from multicomponent regression for later comparison with reference spectra L1, L2, and L3 obtained from one or more spectral libraries. Fig. 3 shows a more detailed version of the Fig. 2. The estimated pure component spectra (labeled P1, P2, and P3) are compared with reference spectra (labeled L1, L2, and L3) to determine the degree to which the pure component spectra match the library spectra. If the pure component spectra match the library spectrum to a desired degree, the comparison library spectrum is selected as the candidate spectrum (B i ) is viewed. Fig. 4 shows a flowchart illustrating the matching methodology of Fig. 3, where box 400 corresponds to step 200 of Fig. 3 is equivalent, box 430 to steps 210 and 220 of Fig. 2 (as well as future iterations of these steps), and the condition box 440 applies a stop condition for reporting candidate spectra to a user (in box 450). Fig. Figure 5 shows an example output report of candidate spectra that could be presented to a user after the MCR and / or MCR-MCS matching methodology has been performed on the estimated pure-component time series spectra. DETAILED DESCRIPTION
[0015] In the description of the invention herein, it is assumed that a word appearing in the singular includes its plural counterpart, and that a word appearing in the plural includes its singular counterpart, unless implicitly or explicitly understood or stated otherwise. It is further understood that generally for any specified component or embodiment described herein, any of the possible candidates or alternatives listed for that component may be used individually or in combination with one another, unless implicitly or explicitly understood or stated otherwise. It is also to be understood that the figures shown herein are not necessarily drawn to scale, and some of the elements may have been shown merely to illustrate the invention.Reference numerals may also be repeated among the various figures to indicate corresponding or analogous elements. It is further understood that any list of such candidates or alternatives is merely illustrative, not limiting, unless implicitly or explicitly understood or stated otherwise. Furthermore, unless otherwise indicated, numbers indicating amounts of ingredients, constituent materials, reaction conditions, and so forth used in the specification and claims are to be understood as modified by the term "about."
[0016] Therefore, unless otherwise stated, the numerical parameters recited in the description and the appended claims are approximations that may vary depending on the desired properties to be obtained by the subject matter presented herein. Lastly, and not in the spirit of limiting the application of the doctrine of equivalents to the scope of the claims, each numerical parameter should be constructed at least in light of the recited significant figures and by applying ordinary rounding techniques. Although the numerical ranges and parameters indicating the general scope of the subject matter presented herein are approximations, the numerical values reported in the specific examples are reported as accurately as possible. However, any numerical values will contain certain errors necessarily resulting from the standard deviation found in their respective test measurements. General description
[0017] The most laborious step in analyzing a series data file (e.g., a time series of spectra) is the single, sequential extraction followed by the analysis of the individual spectra, which may themselves be mixtures. Such an analysis methodology is time-consuming and requires some skill and "art" to be performed effectively. Such a single, sequential extraction method limits the user to analyzing small regions of a file identified as "interesting" to the user.To overcome this burden in a novel way, the embodiments disclosed herein include an automated process using multicomponent regression (MCR) that estimates the pure components in the sample under investigation, often followed by a multicomponent search (MCS) method that (when appropriately configured) utilizes unbound search criteria from one or more spectral libraries. One such MCS method is described in U.S. Patent No. 7,698,098 B2, entitled "EFFICIENT SPECTRAL MATCHING, PARTICULARLY FOR MULTICOMPENENT SPECTRA," issued April 13, 2010, to Ritter et al., incorporated herein by reference.
[0018] An MCR-MCS combination method according to the invention thus provides a user with an advantageous and novel tool that not only simplifies but automates a useful process, and ensures consistency from user to user. In particular, the MCR-MCS methodologies disclosed herein can provide full and complete analysis of the data set, so that even small details that might be overlooked by conventional methods can now be recognized and thus usefully interpreted by the user. One advantageous use of the present embodiments is, for example, the overlaying of profiles showing the temporal behavior of the various components. Such a result provides what a customer is seeking, i.e., an in-depth investigation of how the data evolves during the time-bound event.
[0019] For the end user, this means that a rapid, complete story can be told. For example, the profiles (what and when) for two or more materials that differ only in some additive can be compared, telling the user what is different. In cases where the same materials are present but the overall process has differed, the time evolution plots can illustrate how the different production process affects the materials. Importantly, the inventive methods are available to every skill level of the user, meaning that pharmaceutical laboratories without specialist knowledge in, for example, FT-IR analysis of materials, or the regular analytical laboratory with less skilled users, can now obtain high-quality results. Special description
[0020] The aspect of Multivariate Component Resolution (MCR) disclosed herein relates to a mathematical method of regressively extracting a set of concentration-time profiles and estimated spectra of pure components from a time series set of unknown mixture spectra without any prior knowledge of the mixture contained in the evolving sample under investigation. It can therefore be seen that the automated processing mode of the present application begins with MCR so as to extract a series of linearly independent factors from the sequence of acquired spectral data. The factors essentially represent a distillation of the series of spectra to their constituent parts, i.e., spectra that, when combined, describe the data. As a non-limiting illustration, such a time series data set of the MCR method as disclosed herein can be used to obtain estimated "pure components" (e.g.,fluorophores) of a fluorescent sample along with their respective relative concentrations to provide the quantitative contributions of such individual estimated “pure” components.
[0021] As a working procedure, absorption spectra measured against time are first obtained by using any number of means known to those of ordinary skill in the art, such as, but not limited to, thermogravimetric analysis (TGA), to produce a time series set of spectral data (spectra acquired from a developing sample) similar to that shown in Fig. 1A. The initial goal is to estimate the “pure components” that make up the time series set of spectra.
[0022] Accordingly, although multicomponent regression (MCR) can extract the desired series of linearly independent spectra through the analysis process, it should be noted that MCR software cannot distinguish between spectra with one component or ten, but can only extract spectra that show independent temporal evolution. For example, if ammonia and water evolve from a sample at the same time, MCR software, as used here, can extract the spectrum of ammonia plus water, but not the separate spectra of ammonia and water. On the other hand, if isocyanate also evolves, but at a different time, the result may show ammonia plus water and isocyanate, even if the resulting spectra overlap with the spectra of ammonia plus water.
[0023] If we specifically Fig. 1A, Fig. 1B and Fig. Turning to Figure 1C, the illustrated figures show exemplary data of carbon dioxide, ammonia, isocyanic acid, and water trapped in a developing epoxy sample, as obtained by instrumentation and subsequently extracted using the multivariate curve resolution (MCR) method step of the invention. Fig. In particular, the spectra shown in Figure 1A show a snapshot in time of the absorption spectra obtained by thermogravimetric analysis (TGA) of the sample. Users can acquire such data with a dedicated front-end, producing a series-time file of spectra similar to Fig. 1A is produced.
[0024] Fig. Figure 1B shows the absorption spectra of estimated pure components (e.g., carbon dioxide, ammonia, isocyanic acid, and water) as a result of MCR analysis of the obtained series time file of the spectra, one of which is shown as an example in Fig. 1A can be seen. Fig. Figure 1C finally shows MCR-produced time profiles for the Fig. Estimated components shown in Figure 1B.
[0025] As a more general, yet more detailed description of the MCR algorithm disclosed here, a set of absorption spectra, similar to Fig. 1A, but measured against time, recorded by means known to those of ordinary skill in the art. The MCR-embedded software retrieves the set of absorption spectra S (spectra x number of data points). Note that the first goal of the MCR software package is to estimate the "pure components" that make up the set of spectra. To start, the pure components are denoted by P (pure components x number of data points), and C (spectra x pure components) denotes the amount of each pure component in each spectrum.
[0026] For given actual spectra of pure component matrix S, where each row correlates with a spectrum of a mixture, the following form is produced as a result: S=PC
[0027] Here, P and C are the vector matrices, where P, as stated above, are the "pure components" (i.e., pure components x number of data points), and C is the amount of each pure component in each spectrum (spectra x pure components). It should also be noted that the "pure components" (i.e., pure components x number of data points) are desirably approximately the same as the total number of estimated components resulting from the series time file. The correlated spectrum resulting from Equation 1 above thus desirably produces best estimates of how the most dominant individual component intensities are changing in the evolving sample(s).
[0028] It should also be noted that the steps of the MCR method disclosed here also advantageously utilize constraints, such as unimodality constraints, but more often non-negativity constraints. Often, a non-negativity constraint is chosen as the preferred constraint based on specific knowledge of the data; e.g., that absorbance measurements should be positive, thus providing improved intensities and sample concentrations in the data, which can often be plagued by measurement ambiguity. The use of non-negativity constraints further constrains C and P to both be non-negative, i.e., c(i,j) >= 0 and p(j,k) >= 0; (where i corresponds to the number of samples spectroscopically measured k times at wavelength j).
[0029] To begin the iterative process, MCR must initially guess the number of components. Strategies have been proposed to estimate the number of components, but ultimately, each strategy involves a certain degree of arbitrariness. The technique must estimate both the pure-component spectra and the concentrations from a time-series set of measured spectra or a spatial collection of spectra. This is done in an iterative procedure called alternating least squares. The first step is to arbitrarily guess the shape of either the pure-component spectra or the concentration profiles.
[0030] If you guess the pure component spectra arbitrarily, you solve the least squares problem S = PC for C with the boundary condition that all c jk>= 0. This is done using an iterative procedure called non-negative least squares (NNLS). This leads to an estimate of C. This estimate of C, the concentrations for the spectra, is then used to make a new estimate of the pure component spectra, P. That is, the problem S = PC is solved by NNLS for P. The fact that the technique is NNLS ensures that all p ij >= 0. The steps of resolving for C and then resolving for P continue until the solution converges. This occurs after several iterations. The result is a least-squares solution for the pure-component spectra, P, and the concentrations for the spectra, C, producing the collection of measured spectra S.
[0031] It should be noted that the pure-component estimate is an approximation and has not been proven to match the spectrum of any real physical material. However, it is a meaningful starting point for an MCS (multicomponent search) analysis.
[0032] MCR can then provide the user with the estimated components and concentrations in diagrams or plots to show the time dependence, as similarly described in Fig. 1B (i.e., yields estimated pure components) and 1C (i.e., yields concentration-time profiles for each estimated pure component).
[0033] However, as indicated above, it can be seen that the advantageous aspect of the present invention lies in the ability to integrate the MCR analysis methodology with the MCS (multicomponent search) algorithm, which is similarly described in U.S. Patent No. 7,698,098 B2, entitled "EFFICIENT SPECTRAL MATCHING, PARTICULARLY FOR MULTICOMPENENT SPECTRA," issued April 13, 2010, to Ritter et al. Such an MCS process generally deconvolves the individual spectra for which provided databases are searched, as described in more detail below. MCS thus provides identification of each of the estimated components resulting from MCR by performing spectral correlation that correlates the individual spectra with original data sets. The overall advantageous result is the production of often improved, accurately estimated components and time profiles similar to those in Fig. 1A and Fig. 1C, i.e. to provide the user with even more reliable time-dependent information on each component in a developing sample.
[0034] Fig. 2 is now shown schematically to provide a general understanding of the integrated novel aspect of MCR-MCS. P1, P2, P3...., as in Fig. 2 specifically denote estimated pure-component time series spectra obtained from a spectrometer using MCR software. Such estimated spectral information, i.e., P1, P2, P3..., as provided by MCR, is then passed to the MCS software aspect for comparison with the previously obtained reference comparison spectra (denoted L1, L2, L3...).
[0035] Fig. 3 details how the estimated time series pure component spectra P1, P2, P3.... of Fig. 2 are compared with some of the reference (library) spectra to determine the degree to which the pure spectra correspond to the reference library spectra L1, L2, L3... (now in step 200 of Fig. 3 illustrated).
[0036] In particular, after an estimated time series of pure component spectra P1, P2, P3..., as illustrated in step 200, has been obtained from an optical instrument (e.g., a spectrometer), a database, or any source known to those skilled in the art, and thereafter processed by MCR as already discussed, comparison library spectra, e.g., L1, L2, L3, can be identified in the following manner.
[0037] Initially, comparison spectra, i.e., one or more reference spectra for comparison, are accessed from one or more spectral libraries or other sources. The one or more estimated pure-component time series spectra P1, P2, P3, ... extracted using MCR are then compared with at least some of the comparison spectra to determine the degree to which the time series of spectra correspond to the one or more comparison spectra. If the estimated pure-component time series spectra P1, P2, P3, ... correspond to one or more comparison spectra to some degree, such as by meeting or exceeding some user-defined or preselected correspondence threshold, the one or more comparison spectra are considered to be represented by one or more candidate spectra B(1)1, B(1)2, ... B(1) Midentified as long as the correspondence threshold is not set too high. If no candidate spectra are identified, the correspondence threshold can be set to a lower value.
[0038] Next, the possibility that any of the estimated pure-component time series spectra may have emerged from a multicomponent mixture is considered. New comparison spectra are generated, with each comparison spectrum being a combination of one of the previously identified candidate spectra and one of the comparison spectra from the spectral libraries or other sources. The estimated one or more pure-component time series spectra are then compared again with at least some of these new comparison spectra to determine the degree to which the estimated pure-component time series spectra correspond to the new comparison spectra. This step is described in 210 in Fig. 3 schematically illustrates, where any number of the estimated pure component time series spectra P1, P2, P3...., is compared with new reference spectra: B(1)1+L1,B(1)1+L2,…B(1)1+LN (ie the first of the previously identified candidate spectra from step 200 in Fig. 3 combined with each of the comparison spectra from the spectral libraries or other sources): B(1)2+L1,B(1)2+L2,…B(1)2+LN (i.e., the second of the previously identified candidate spectra from step 200 combined with each of the comparison spectra from the spectral libraries or other sources) and so on until the estimated pure-component time series spectra are compared with new comparison spectra: B(1)M+L1,B(1)M+L2,…B(1)M+LN (i.e., the last of the previously identified candidate spectra from step 200 combined with each of the comparison spectra from the spectral libraries or other sources).
[0039] If, from these comparisons, it emerges that, for example, any of the new comparison spectra has a desired degree of correspondence with the estimated pure-component time series spectra P1, P2, P3... (such as by meeting or exceeding the correspondence threshold), the new comparison spectrum is considered a new candidate spectrum. These new candidate spectra are in Fig. 3 in step 210 as B(2)1, B(2)2, ... B(2) M (Note that, if desired, M in step 210 need not be equal to M in step 200, i.e., the number of candidate spectra in step 210 need not be the same as the number of candidate spectra in step 200.) Here, each candidate spectrum represents B(2)1, B(2)2, ... B(2) M two components, i.e. two combined reference spectra obtained from a spectral library or other source.
[0040] The previous step can then be repeated once or several times in an unconstrained manner, if desired, with each repetition serving to generate new comparison spectra using the candidate spectra identified in the previous step. Examples of this can be found in step 220 in Fig. 3, where the candidate spectra B(2)1, B(2)2, ... B(2) M from step 210 in combination with the comparison spectra L1, L2, ... L N from the spectral libraries or other sources to generate new comparison spectra. Comparison of the estimated pure-component time series spectra P1, P2, P3,..., with these new comparison spectra in turn identifies new candidate spectra B(3)1, B(3)2, ... B(3) M(where M again need not be equal to M in steps 210 and / or 200). The iteration may end when the candidate spectra include any desired number of components, e.g., after the new comparison spectra include a desired number of combined comparison / reference spectra obtained from a spectral library or other source.
[0041] This condition is shown in the flowchart of Fig. 4, where step 400 is transferred to step 200 of Fig. 3 is equivalent, step 430 to steps 210 and 220 of Fig. 3 (as well as future iterations of these steps), and the condition box 440 evaluates the number of components c in the candidate spectra and terminates the iteration after some maximum number C is reached. The iteration may alternatively or additionally terminate when any desired number of candidate spectra is identified; when one or more candidate spectra are identified that match(es) the unknown spectrum in at least some qualifying correspondence value (where the qualifying correspondence value is greater than the threshold correspondence value), as discussed below; or when other suitable conditions appear.
[0042] At least some of the candidate spectra may then be presented to a user, wherein the candidate spectra are preferably presented to the user in rank order such that those candidate spectra with a greater correspondence to the unknown spectrum are presented first (as in step 450 in Fig. 4). An example format for an output list of candidate spectra that can be presented to a user is shown in Fig. 5. Details about the unknown spectrum are given in the output list heading, followed by details about the candidate spectra. The first candidate spectrum listed—listed with a rank / index of 1—is a spectrum for polystyrene film and has a match metric (roughly equivalent to a 'percent match') of 99.58 against the unknown spectrum. The spectral library or other source of this candidate spectrum is also listed (here, 'User Sample Library'), as is its position within the library / source (in 'Source Index' #2, i.e., this is the second spectrum provided in the 'User Sample Library').The second candidate spectrum listed is, in fact, a combination of three spectra from spectral libraries or other sources—a spectrum of toluene (transmission cell), a spectrum of an ABS plastic (ATR-corrected), and a polytetrafluoroethylene film spectrum—where, when combined in appropriate proportions (as discussed subsequently), these spectra yield a match metric of 68.97 with the unknown spectrum. Their cumulative match metrics are also presented, with toluene yielding a match metric of 56.96, toluene and ABS together yielding a match metric of 68.92, and toluene, ABS, and polytetrafluoroethylene together yielding the match metric of 68.97. The libraries or other sources for these spectra are again provided, along with an indication of the position of each spectrum within its library / source.
[0043] Additional metrics are preferably also provided with the output list, in particular the weight of each comparison spectrum (each component / reference spectrum) within the candidate spectrum, i.e., the scaling factor used to adjust each comparison spectrum to obtain the best match to the unknown spectrum. For example, the first listed candidate spectrum (polystyrene film) has a weight of 5.4195, which means that, according to estimates, the unknown spectrum has 5.4195 times the polystyrene content of the sample from which the candidate spectrum was obtained. The second listed candidate spectrum contains different weights of toluene, ABS, and polytetrafluoroethylene, these weights being determined by regression analysis of the comparison spectra against the unknown spectrum during the above-mentioned comparison step (i.e.,The different component / reference spectra within a comparison spectrum are weighted proportionally to achieve the best match with the unknown spectrum during comparison. The user can thus be provided with at least an approximate quantification of the components within the unknown spectrum.
[0044] The above methodology can be described as finding the reference spectra with "best match", combining the best-matching spectra with other reference spectra, and then further identifying best-matching spectra from these combinations (with the methodology continuing iteratively with the previous combination step). It can therefore be seen that instead of comparing all possible combinations of reference spectra L1, L2, ... L Ncan consider far fewer combinations, fundamentally by removing the reference spectra that bear less similarity to the unknown spectrum. The methodology returns high-quality matches in far less time than methods that consider all combinations, particularly when large numbers of reference spectra are used and when the unknown spectrum is tested for larger combinations of component / reference spectra—in some cases returning results in minutes where previously it took hours.
[0045] Before performing the above-mentioned comparisons between the estimated pure-component time series spectra and comparison spectra, the invention may perform one or more transformations on one or both of the estimated pure-component time series spectra and comparison spectra to accelerate and / or increase the accuracy of the comparison process or otherwise improve data processing. The invention may, as examples, perform one or more of data smoothing (noise reduction), peak discrimination, rescaling, domain transformation (e.g., transformation to vector format), differencing, or other spectral transformations.The comparison itself can also take a variety of forms, such as by simply comparing intensities / amplitudes over similar wavelength ranges between unknown and reference spectra, by converting the unknown and reference spectra into vectorial forms and comparing the vectors, or by other forms of comparison.
[0046] The methodology described above can also be modified to further advance the identification of candidate spectra. As an example of such a modification, when generating a new comparison spectrum by combining a previously identified candidate spectrum and a comparison spectrum obtained from a spectral library or other source, the combination can be skipped or discarded (i.e., deleted or not counted as a potential new candidate spectrum) if the candidate spectrum already contains the comparison spectrum.
[0047] For a more specific illustration, consider the situation where comparison spectrum L1, obtained from a spectral library, is used in step 200 ( Fig. 3) is chosen as B(1)1 due to sufficient agreement with unknown spectra. In the next iteration in step 210, the new comparison spectrum B(1)1 + L1 can be skipped or discarded because it is equivalent to L1 + L1 (i.e., reference spectrum L1 combined with itself, which simply results in L1 again). By avoiding the generation and / or use of comparison spectra that have redundant component spectra, the methodology can reserve computation time for comparison spectra that are more likely to yield new candidate spectra.
[0048] As another example of a modification that can be implemented to accelerate the identification of candidate spectra, if a candidate spectrum matches the unknown spectrum to a greater degree than or equal to some "qualifying" correspondence value—where this qualifying correspondence value is greater than the threshold correspondence value—the comparison spectra therein (i.e., their component spectra) can be excluded from any subsequent generation of new comparison spectra. This measure essentially takes the approach that if a candidate spectrum is already a very good match for an unknown spectrum (e.g., if it has a qualifying correspondence value above 95%), this may be sufficient, and there is no significant need to determine whether the match can be further improved when the candidate spectrum is combined with other spectra.
[0049] A further modification that can be made to accelerate the identification of candidate spectra applies to the special case when one or more of the components of the unknown spectrum are known – for example, when monitoring the output of a process in which a material with known components is to be generated in a fixed quantity. In this case, during the first comparison round (step 200 in Fig. 3, step 400 in Fig. 4) the candidate spectra B(1)1, B(1)2, ... B(1) M simply be placed on the spectra for the known components. Performing the remaining procedure then serves to identify any additional components (i.e., impurities) that may be present, as well as the relative proportions of the various components.
[0050] As stated above, if the correspondence threshold—that is, the degree of agreement required between the estimated one or more pure-component time series spectra and a reference spectrum for the reference spectrum to be considered a candidate spectrum—is set too high, the result may be that no candidate spectra are obtained. A value of 90% correspondence is typically appropriate for the correspondence threshold, although this value may be better set higher or lower depending on the details of the spectra considered.
[0051] It is also possible to set the correspondence threshold to zero (or a value close to zero), in which case a candidate spectrum can result from each comparison spectrum. If the correspondence threshold in step 200 of Fig. 3-4 is set to zero, for example, M = N and B(1)1, B(1)2, ... B(1) Mthen correspond to one of L1, L2, ... L N . Some of the candidate spectra in this case may actually be poor candidates because they poorly match the unknown spectrum. It is then useful to rank the candidate spectra from highest match to lowest match and then consider those candidate spectra with the highest match first when performing any subsequent steps. To reduce computational effort, it may be useful in this case to discard the candidate spectra with the lowest match when performing any subsequent steps. For example, one could keep only the top 10%, 25%, or 50% of the candidate spectra with the highest match and use these in subsequent steps.
[0052] It is expected that the invention can be implemented in spectral identification software for use in computers or other systems (e.g., spectrometers) that receive and analyze spectral data. Such systems may include portable / handheld computers, field measurement devices, application-specific integrated circuits (ASICs) and / or programmable logic devices (PLDs) provided in environmental, industrial, or other monitoring equipment, and any other systems where the invention may prove useful.
[0053] In another embodiment, the following non-limiting example illustrates an advantageous aspect of a user output interface for use with the methods disclosed herein. It will be appreciated that a closely related problem potentially solved with the present embodiments involves the analysis of two similar materials. Two example scenarios: In a first, a seal or O-ring from one batch fails, while a corresponding part from another batch performs well. In a second, Competitor B has introduced a product chemically similar to a product manufactured by Competitor A, and A wants to know the processing differences. In both cases, TGA-IR is often an insightful, advantageous method to implement, providing qualitative and quantitative data.
[0054] Thus, advantageously, a "lightbox" (i.e., digitally superimposed (or presented side by side)) may additionally be provided, an extension of the invention, which involves performing a coupled analysis (i.e., not sequentially, but simultaneously) of the two data sets. The final result may be a sequence of composition information and profile information. The output interface may provide views of the search results and views of the temporal evolution profiles of these components. An important aspect is the differences between these comparisons.
[0055] If analyses are configured to run sequentially, the ranking of search results and the number of components found can potentially differ, making comparisons more complex. By performing the analysis in a coupled fashion, the results are linked by both the composition and ranking of the search results. This enables the "lightbox" approach, where the results are digitally overlaid (or presented side by side) for easy comparison.
[0056] Returning to the two scenarios, in the first case, the overlay view could reveal that a component is missing—a formulation error—or that the temperature evolution profile for one or more components is shifted between the two—a processing error. In the second case, the deformulation profiles allow the comparison of the known product with known characteristics from Company A with the unknown material from Company B, again revealing compositional or processing differences. This ultimately represents the "final answer" that the entire analysis was striving for—what is different about these two samples.
[0057] Although the invention has been generally described as being applicable in the context of spectral alignment for molecular spectrometers, it may alternatively or additionally be used in mass spectroscopy, X-ray spectroscopy, or other forms of spectroscopy. It may also be useful in other forms of measurement analysis in which signals are measured against reference values, where such signals and reference values may be considered "spectra" in the context of the invention.
[0058] It should be understood that the features described with respect to the various embodiments contained herein may be mixed and combined in any combination without departing from the spirit and scope of the invention. Although various selected embodiments have been illustrated and described in detail, it should be understood that they are exemplary and that various substitutions and modifications are possible without departing from the spirit and scope of the present invention.
Claims
[1] A method for analyzing spectra from an evolving or changing sample, comprising: Using a spectrometer to obtain a time series set of spectra; estimating one or more qualitative and quantitative components from each of the time series sets of spectra using a computer by means of a regression procedure; and Using a computer to pass the estimated one or more qualitative and quantitative components from each of the time series set of spectra to a multi-component search (MCS) algorithm configured to iteratively correlate one or more comparison spectra stored in one or more spectral libraries with each of the estimated time series set of spectra represented as one or more respective qualitative and quantitative components, the result being an iteratively determined best-matching time series set of one or more candidate spectra, wherein the regression method within the estimation step comprises a multi-component regression (MCR) algorithm, wherein the MCS algorithm, which iteratively correlates one or more candidate spectra, is further configured to: Generating (210) one or more new comparison spectra (B(2)1, B(2)2, ... B(2) M ), where each of the new comparison spectra (B(2)1, B(2)2, ... B(2) M ) a combination of one of a previously identified candidate spectrum (B(1)1, B(1)2, ... B(1) M ) and one of a reference spectrum (L1, L2, L3) from a spectral library source, and Comparing the estimated one or more qualitative and quantitative components from each of the time series sets of spectra with the selected new comparison spectrum (B(2)1, B(2)2, ... B(2) M ) to determine a degree of correspondence; and Repeating the above generation and comparison steps until a desired number of the sets of one or more candidate spectra is identified, or when the set of one or more candidate spectra is identified which matches the estimated one or more qualitative and quantitative components from each of the time series set of spectra according to at least some qualifying correspondence value, or after a maximum number of components in the selected time series set of one or more candidate spectra is reached. [2] The method of claim 1, further comprising: presenting the iteratively determined best-matching time series set of one or more candidate spectra ranked and / or as temporal evolution profiles of the one or more qualitative and quantitative components. [3] The method of claim 1, wherein the multicomponent regression (MCR) algorithm comprises a unimodality constraint on the obtained time series set of spectra. [4] The method of claim 1, wherein the multicomponent regression (MCR) algorithm comprises a non-negativity constraint on the obtained time series set of spectra. [5] The method of claim 4, wherein the multicomponent regression (MCR) algorithm comprises an iterative non-negative least squares (NNLS) procedure to provide the one or more qualitative and quantitative components from each of the time series set of spectra. [6] The method of claim 1, wherein one or more transformations are performed on at least one of the estimated one or more qualitative and quantitative components from each of the time series sets of spectra and the one or more comparison spectra. [7] The method of claim 1, further comprising skipping or discarding the comparison step if the previously identified candidate spectra already include one of a comparison spectrum (L1, L2, L3) from the spectral library source. [8] A system for analyzing spectra from an evolving or changing sample, comprising: a spectrometer configured to generate a time series set of spectra; and a computer configured to estimate one or more qualitative and quantitative components of each of the time series set of spectra using a regression method;wherein the regression method within the estimation step comprises a multi-component regression (MCR) algorithm, the computer passing the estimated one or more qualitative and quantitative components from each of the time series set of spectra to a multi-component search (MCS) algorithm configured to iteratively correlate one or more comparison spectra stored in one or more spectral libraries with each of the estimated time series sets of spectra represented as one or more respective qualitative and quantitative components, the result being an iteratively determined best-matching time series set of one or more candidate spectra, the computer further configured to:; a. Generate one or more new comparison spectra (B(2)1, B(2)2, ... B(2) M ), where each of the new comparison spectra (B(2)1, B(2)2, ... B(2)M ) a combination of a previously identified candidate spectrum (B(1)1, B(1)2, ... B(1) M ) and one of the comparison spectra (L1, L2, L3) from a spectral library source, b. Comparing the estimated one or more qualitative and quantitative components from each of the time series sets of spectra with the selected new comparison spectrum (B(2)1, B(2)2, ... B(2) M ) to determine a degree of correspondence; and c. Repeating the above generation and comparison steps until a desired number of sets of one or more candidate spectra is identified, or when the set of one or more candidate spectra is identified which matches the estimated one or more qualitative and quantitative components from each of the time series set of spectra according to at least some qualifying correspondence value, or after a maximum number of components in the selected time series set of one or more candidate spectra is reached. [9] The system of claim 8, wherein the computer is further configured to present the iteratively determined best-matching time series set of one or more candidate spectra ranked and / or temporal evolution profiles of the one or more qualitative and quantitative components. [10] The system of claim 8, wherein the computer skips the comparing if the previously identified candidate spectrum already contains one of a comparison spectrum (L1, L2, L3) from the spectral library source.
Citation Information
Patent Citations
Efficient spectral matching, particularly for multicomponent spectra
EP2245443B1
Method and system of dynamic learning through a regression-based library generation process
US20040220760A1
Method for identifying components of a mixture via spectral analysis
US7072770B1
Efficient spectral matching, particularly for multicomponent spectra
US7698098B2