Data Processing Method and Data Processing System
By calculating the wavelength region similarity of the spectral data and setting the object range for matrix decomposition, the problem of overlapping multiple components in liquid chromatography is solved, and high-precision peak separation and quantification are achieved.
Patent Information
- Application Number
- CN202210789952.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-09-27
- Filing Date
- 2022-07-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-07-05
AI Technical Summary
In liquid chromatography, it is difficult for the prior art to separate peaks of multiple components overlapping each other on chromatography with high accuracy, resulting in inaccurate quantitative results.
By obtaining the spectral data of each component, the similarity of the wavelength region is calculated, the region with low similarity is set as the object range, and matrix decomposition is performed within this range to generate chromatographic data of each component.
The accuracy of chromatographic peak separation is improved and the accuracy of quantitative results is ensured.
Smart Images

Figure CN115856180B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and a system for processing three-dimensional chromatography data. Background Art
[0002] In a liquid chromatograph (LC) using a multi-channel detector such as a photo diode array (PDA) detector, an absorption spectrum of a sample eluted from an analytical column is continuously acquired, whereby three-dimensional chromatography data having three dimensions of time, wavelength, and signal intensity (absorbance) can be obtained.
[0003] When quantifying a target component in a sample using a liquid chromatograph, generally, a chromatogram is generated using the wavelength at which the absorbance of the target component is maximum, and the area value of the peak of the target component is obtained on the chromatogram for quantification. However, sometimes impurities other than the target component are contained in the sample, and the peak of the impurity sometimes overlaps with the peak of the target component on the chromatogram. In such a case, in a state where multiple peaks overlap, the peak area value of the target component or the impurity cannot be obtained, and thus a quantification result cannot be obtained. Therefore, it is necessary to separate multiple components whose peaks overlap on the chromatogram from each other.
[0004] As a method for separating the peaks of multiple overlapping components, there is a method of estimating the chromatogram of each component by applying a model function (peak model) such as an Exponential Modified Gaussian (EMG) function to the waveform of an actual chromatogram (see Patent Document 1); and a method of estimating the chromatogram of each component mathematically by performing matrix decomposition on the original three-dimensional chromatography data.
[0005] [Prior Art Documents]
[0006] [Patent Documents]
[0007] [Patent Document 1] International Publication No. 2016 / 035167 Summary of the Invention
[0008] [Problems to be Solved by the Invention]
[0009] Since the separation of peaks by matrix decomposition only separates the original three-dimensional chromatography data mathematically into a specified number, the shape of each separated peak may be completely different from the actual peak shape. On the other hand, the waveform of the chromatogram of each separated component does not depend on the peak model, and thus has a high degree of freedom, and it is also possible to obtain a higher separation accuracy than the method of applying the peak model.
[0010] An object of the present invention is to perform high-precision separation of peaks of multiple components overlapping each other on a chromatogram by matrix factorization.
[0011] [Technical means for solving the problem]
[0012] The inventors of the present invention considered that, before separating the peaks of multiple components overlapping on a chromatogram using matrix factorization, spectral data of each of the multiple components was obtained by a certain method and used as the basis for matrix factorization. However, when the overall waveforms of the spectra of the multiple components to be separated are similar to each other, it is difficult to perform high-precision peak separation if only these spectral data are simply used for matrix factorization. Here, the inventors of the present invention obtained the following insight: comprehensively evaluating the similarity between wavelength regions corresponding to the spectral data of the multiple components to be separated, and performing matrix factorization using the spectral data in wavelength regions with low mutual similarity, whereby the accuracy of peak separation achieved by matrix factorization can be improved. The present invention is based on this insight.
[0013] The data processing method of the present invention includes: a data preparation step of preparing actual data of a three-dimensional chromatogram including a chromatogram and a spectrum obtained by chromatographic analysis of a sample, and spectral data of each of multiple components in the sample where peaks overlap each other on the chromatogram of the actual data; a similarity calculation step of calculating the similarity between wavelength regions corresponding to the spectral data of each of the multiple components prepared in the data preparation step while changing the wavelength regions comprehensively for each wavelength region; an object range setting step of searching for a wavelength region where the similarity is lower than the overall similarity between the spectral data of each of the multiple components based on the calculation result in the similarity calculation step, and setting an object range; and a peak separation step of performing matrix factorization of the actual data in the object range set in the object range setting step using the spectral data of each of the multiple components, thereby generating chromatographic data of each of the multiple components.
[0014] The data processing system of the present invention includes: a data storage unit that stores actual data of a three-dimensional chromatogram including a chromatogram and a spectrum obtained by chromatographic analysis of a sample, and spectral data of each of a plurality of components in the sample in which peaks overlap each other on the chromatogram of the actual data; and a data processing unit configured to perform a peak separation process of the plurality of components in the sample using the actual data and the spectral data stored in the data storage unit. Moreover, the data processing unit is configured to execute: a similarity calculation step of calculating, for each wavelength region, the similarity between wavelength regions corresponding to the spectral data of each of the plurality of components stored in the data storage unit while changing the wavelength regions comprehensively; an object range setting step of searching, based on the calculation result in the similarity calculation step, for a wavelength region in which the similarity is lower than the overall similarity between the spectral data of each of the plurality of components, and setting an object range; and a peak separation step of performing matrix decomposition of the actual data in the object range set in the object range setting step using the spectral data of each of the plurality of components, thereby generating chromatographic data of each of the plurality of components.
[0015] [Advantages of the Invention]
[0016] According to the data processing method and data processing system of the present invention, actual data of a three-dimensional chromatogram of a sample and spectral data of each of a plurality of components in which peaks overlap each other on the chromatogram are prepared, the similarity between wavelength regions corresponding to the spectral data of the plurality of components is calculated while changing the wavelength regions comprehensively, a wavelength region with a low similarity is searched based on the calculation result, an object range is set based on the search result, and matrix decomposition of the actual data using the spectral data is performed within the set object range. Therefore, the peaks of the plurality of components can be separated with high precision. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flowchart schematically showing an embodiment of the data processing method.
[0018] Figure 2 is a block diagram schematically showing an embodiment of the data processing system that executes the data processing method.
[0019] Figure 3 is a flowchart showing an example of the data processing performed using the data processing system.
[0020] Figure 4 of (A), Figure 4(B) is a diagram showing an example of peak separation by applying a peak model, wherein (A) shows a chromatogram at a certain wavelength of actual data, and (B) shows a state where the peak model is applied to the chromatogram.
[0021] Figure 5 This is a diagram showing an example of a heat map of similarity calculation results in the data processing method. DETAILED DESCRIPTION
[0022] Hereinafter, embodiments of the chromatographic data processing method and data processing system of the present invention will be described with reference to the accompanying drawings.
[0023] Figure 1 An embodiment of a data processing method is schematically shown in FIG.
[0024] The data processing method of the embodiment described above is a method for separating peaks of a plurality of components overlapping each other on a chromatogram using three-dimensional chromatogram data including a spectrum and a chromatogram obtained by chromatographically analyzing a sample.
[0025] In the method, first, actual data on a three-dimensional chromatogram of a sample is prepared, and spectral data of each of a plurality of components whose peaks overlap with each other on the chromatogram of the actual data is prepared (step 101). The spectral data of each of the plurality of components to be separated may be data acquired by any method. When the plurality of components are known, the components may be analyzed separately to acquire their respective spectral data. When the plurality of components to be separated are unknown, peak separation by applying a peak model may be performed on the actual data of the three-dimensional chromatogram of the sample to generate spectroscopic inference data, and the inference data may be used as the spectral data of each component. As an algorithm for peak separation by applying a peak model, for example, the algorithm disclosed in Patent Document 1 (International Publication No. 2016 / 035167) may be cited.
[0026] Next, regarding the similarity between wavelength regions corresponding to each other among the spectral data of the plurality of components to be separated, the calculation is performed for each wavelength region while changing the wavelength regions comprehensively (step 102). After that, based on the calculation result of the similarity, a wavelength region where the similarity is lower than the mutual similarity of the spectral data of each component as a whole (for example, the wavelength region with the lowest similarity among the wavelength regions where the similarity is calculated) is searched for, and the wavelength region is set as the object range for matrix decomposition (step 103). Then, matrix decomposition of the three-dimensional chromatography based on the spectral data of each component is performed within the set object range to generate chromatographic data of each component (step 104). As the matrix decomposition, non-negative matrix factorization (NMF) or the like can be used. The matrix decomposition can be repeatedly performed until the synthesis result of the chromatographic data of each component generated has a certain degree of approximation to the actual data.
[0027] An embodiment of a data processing system for performing the data processing method is shown in Figure 2 .
[0028] The data processing system 1 includes a data storage unit 2 and a data processing unit 4. Analytical data obtained by using an analytical device 100 is input into the data processing system 1. The analytical device 100 is configured to perform liquid chromatography analysis on a sample and obtain an absorbance spectrum at regular intervals. That is, three-dimensional chromatography data including a chromatogram and a spectrum is input from the analytical device 100 into the data processing system 1.
[0029] The data storage unit 2 is a storage area for storing the actual data of the three-dimensional chromatography input from the analytical device 100 and the spectral data of each of the plurality of components to be separated. The data storage unit 2 can be implemented by a non-volatile memory, a hard disk drive, or the like.
[0030] The data processing unit 4 has: a first function of generating estimated data of the chromatogram and estimated data of the spectrum of each of the plurality of components where the peaks overlap each other on the chromatogram of the actual data by using a peak separation algorithm through applying a peak model; and a second function of adjusting the estimated data of the chromatogram and the estimated data of the spectrum of each component generated by the first function by using a matrix decomposition algorithm. The data processing unit 4 is a functional unit implemented by executing a program in a computer circuit including a central processing unit (CPU).
[0031] The data storage unit 2 can store the estimated data of the spectra of the respective components generated by the first function of the data processing unit 4 as the spectral data of the respective components. In the second function of the data processing unit 4, the estimated data of the spectra of the respective components stored in the data storage unit 2 can be used to limit the object range of matrix decomposition, and matrix decomposition is performed within the limited object range.
[0032] Use Figure 3 The flowchart of is used to illustrate an example of the peak separation process performed in the data processing system 1.
[0033] When the peak separation process starts, the data processing unit 4 searches for a peak model required to approximate the waveform of the chromatogram of the actual data in a pre-prepared database, and applies the peak model to the matching chromatogram (step 201). Then, by estimating the peak shapes of the respective components based on the peak model applied to the chromatogram, the peaks of the multiple components are separated from each other (step 202). For example, when the waveform of the chromatogram at a certain wavelength in the actual data is the waveform shown in (A) of Figure 4 , as shown in (B) of Figure 4 , three peak models are applied to approximate the waveform. As a result, it is estimated that the waveform of the chromatogram is formed by the overlap of three peaks P1 to P3, and is separated into three peaks P1 to P3.
[0034] The data processing unit 4 generates estimated data of the chromatograms and spectra of the respective components based on the peak separation results obtained by applying the peak model (step 203). Steps 201 to 203 up to this point are executed by the first function of the data processing unit 4.
[0035] Next, the data processing unit 4 calculates the similarity between the wavelength regions corresponding to the estimated data of the spectra of the respective components generated in step 203 for each wavelength region while changing the wavelength regions comprehensively (step 204). As the similarity, cosine (cos) similarity can be used. The method of comprehensively changing the wavelength regions in step 204 is not particularly limited. As an example, a method of changing the minimum wavelength and width (the range for calculating the similarity) of the wavelength regions respectively can be cited. Figure 5This is an example of a heat map generated by calculating the similarity while varying the minimum wavelength and width of the wavelength region. The horizontal axis of the heat map is the minimum wavelength in the wavelength region, and the vertical axis is the width of the wavelength region. The hatched triangular region in the upper right corner of the heat map is the region where the similarity cannot be calculated because the width of the wavelength region exceeds the actual data region. For example, when the minimum wavelength is 250 nm and the width of the wavelength region is 70 nm, the maximum wavelength of the wavelength region is 320 nm, which exceeds 300 nm, so the calculation cannot be performed. The data processing unit 4 may also have a function of generating such a heat map and displaying it on a display (not shown). In addition, it is not necessarily required to generate a heat map.
[0036] After the end of the step 204, the data processing unit 4 determines the wavelength region with the lowest similarity based on the calculation result in the step 204, and sets the wavelength region as the object range for matrix decomposition (step 205). In addition, since the similarity cannot be correctly evaluated when the width of the wavelength region is extremely narrow, it is desirable to set a wavelength region with a width of a certain value (for example, 25 nm) or more as the object range. In addition, in the step 204, the width of the wavelength region for which the similarity is to be calculated may also be limited to a certain value (for example, 25 nm) or more. In addition, here, the description is made in such a way that the data processing unit 4 automatically sets the object range for matrix decomposition, but the present invention is not limited thereto. The data processing unit 4 may also calculate the similarity for each wavelength region in the step 204, display the calculation result as Figure 5 shown, and let the user set the object range.
[0037] After setting the object range, the data processing unit 4 synthesizes the estimated data of the chromatogram and spectrum of each component generated in the step 203 to generate pseudo data of a three-dimensional chromatogram (step 206), and calculates the similarity of the pseudo data with respect to the actual data (step 207). The "similarity" here only needs to numerically represent how similar the pseudo data is to the actual data. Therefore, the calculation method of the similarity is not particularly limited. For example, the sum of the squares of the differences between the values of the pseudo data and the actual data at each point of the three-dimensional chromatogram may be set as the similarity.
[0038] The data processing unit 4 adjusts the parameters of the estimated data of each component using matrix factorization so that the similarity calculated in step 207 is good, that is, the pseudo data is closer to the actual data (step 209). After that, the data processing unit 4 generates pseudo data of the three-dimensional chromatogram based on the adjusted estimated data (step 206), and evaluates the similarity of the generated pseudo data with respect to the actual data (steps 207 and 208). In this way, steps 206 to 209 are repeated, and when the similarity of the pseudo data with respect to the actual data satisfies a specified condition, the adjustment of the estimated data is ended (step 208: Yes). As the specified condition, examples include: the similarity is lower (or higher) than a preset threshold value, or the similarity of the pseudo data after adjusting the estimated data converges to a certain value with respect to the actual data. Steps 204 to 209 are executed by the second function of the data processing unit 4.
[0039] Using steps 204 to 209 executed by the second function of the data processing unit 4, for the estimated data of the chromatogram and spectrum of each component generated by the first function, the data in the wavelength region where the similarity between the respective spectral data is low is used to perform adjustment without being restricted by the shape caused by the peak model. When the spectral data of the respective components to be separated are overall similar to each other, even if these spectral data are used to perform matrix factorization, it is difficult to accurately estimate the chromatogram of each component. However, by narrowing the analysis target range to the region where the similarity between the spectral data of the respective components is lower than the similarity evaluated in the overall spectral data, that is, the region showing the difference in the absorbance characteristics of each component, and performing matrix factorization, the estimation accuracy of the chromatogram of each component can be improved.
[0040] The embodiments described above are merely illustrative of the embodiments of the data processing method and data processing system of the present invention. The embodiments of the data processing method and data processing system of the present invention are as follows.
[0041] In an embodiment of the data processing method of the present invention, it includes:
[0042] A data preparation step of preparing actual data of a three-dimensional chromatogram including a chromatogram and a spectrum obtained by chromatographic analysis of a sample, and spectral data of each of a plurality of components in the sample in which peaks overlap each other on the chromatogram of the actual data;
[0043] A similarity calculation step of calculating the similarity between the wavelength regions corresponding to the spectral data of each of the plurality of components prepared in the data preparation step for each wavelength region while changing the wavelength regions comprehensively;
[0044] Object range setting step, based on the calculation result in the similarity calculation step, setting the wavelength region with the lowest similarity as the object range; and
[0045] Peak separation step, using the spectral data of each of the plurality of components, performing matrix decomposition of the actual data in the object range set in the object range setting step, thereby generating chromatographic data of each of the plurality of components.
[0046] In the first aspect of the one embodiment of the data processing method, in the data preparation step, by applying a pre-prepared peak model to approximate the waveform of the chromatogram of the actual data, using the peak model applied to the chromatogram to generate speculative data of the spectra and chromatograms of each of the plurality of components. Moreover, the spectral data used in the similarity calculation step and the peak separation step is the speculative data of the spectra generated in the data preparation step, and the chromatographic data generated by the peak separation step is based on the speculative data of the chromatograms generated in the data preparation step. According to this form, a matrix decomposition algorithm not restricted by the peak model can be used to adjust the speculative data of the chromatograms and spectra of each component generated by using a peak separation algorithm by applying a peak model, thereby obtaining high peak separation accuracy.
[0047] In the second aspect of the one embodiment of the data processing method, non-negative matrix factorization is used as the matrix decomposition. The second aspect can be combined with the first aspect.
[0048] In the third aspect of the one embodiment of the data processing method, in the similarity calculation step, the similarity is calculated while changing the minimum wavelength and wavelength width of the wavelength region. The third aspect can be combined with the first aspect and / or the second aspect.
[0049] In one embodiment of the data processing system of the present invention, it includes:
[0050] Data storage unit (2), storing actual data of a three-dimensional chromatogram including a chromatogram and a spectrum obtained by chromatographic analysis of a sample, and spectral data of each of a plurality of components in the sample where peaks overlap each other on the chromatogram of the actual data; and
[0051] Data processing unit (4), configured to use the actual data and the spectral data stored in the data storage unit (2) to perform separation processing of the peaks of the plurality of components in the sample,
[0052] The data processing unit (4) is configured to execute:
[0053] Similarity calculation step: For the similarity between wavelength regions corresponding to the spectral data of each of the multiple components stored in the data storage unit (2), the calculation is performed for each wavelength region while changing the wavelength regions comprehensively.
[0054] Object range setting step: Based on the calculation result in the similarity calculation step, the wavelength region with the lowest similarity is set as the object range; and
[0055] Peak separation step: Using the spectral data of each of the multiple components, matrix decomposition of the actual data is performed in the object range set in the object range setting step, thereby generating chromatographic data for each of the multiple components.
[0056] In the first aspect of the one embodiment of the data processing system, the data processing unit (4) is configured to: Before the similarity calculation step, execute a data preparation step. The data preparation step approximates the waveform of the chromatogram of the actual data by applying a pre-prepared peak model, and uses the peak model applied to the chromatogram to generate speculative data of the spectrum and chromatogram for each of the multiple components. In the similarity calculation step and the peak separation step, the speculative data of the spectrum generated in the data preparation step is used as the spectral data, and in the peak separation step, the chromatographic data is generated based on the speculative data of the chromatogram generated in the data preparation step. According to this form, a matrix decomposition algorithm not restricted by the peak model can be used to adjust the speculative data of the chromatogram and spectrum of each component generated by using a peak separation algorithm through applying a peak model, thereby obtaining high peak separation accuracy.
[0057] In the second aspect of the one embodiment of the data processing system, non-negative matrix factorization is used as the matrix decomposition. The second aspect can be combined with the first aspect.
[0058] In the third aspect of the one embodiment of the data processing system, the data processing unit (4) is configured to: In the similarity calculation step, calculate the similarity while changing the minimum wavelength and wavelength width of the wavelength region. The third aspect can be combined with the first aspect and / or the second aspect.
[0059] [Explanation of symbols]
[0060] 1: Data processing system
[0061] 2: Actual data storage unit
[0062] 4: Data processing unit
[0063] 100: Analyzer.
Claims
1. A data processing method, characterized in that, Comprising: A data preparation step of preparing actual data of a three-dimensional chromatogram including a chromatogram and a spectrum obtained by chromatographic analysis of a sample containing a plurality of components, and spectral data of each of the plurality of components in the sample where peaks overlap each other on the chromatogram of the actual data; A similarity calculation step of calculating, for each wavelength region, the similarity between wavelength regions corresponding to the spectral data of each of the plurality of components prepared in the data preparation step while changing the wavelength regions comprehensively; An object range setting step of searching for a wavelength region where the similarity is lower than the overall similarity between the spectral data of each of the plurality of components based on the calculation result in the similarity calculation step, and setting an object range; And A peak separation step of performing matrix decomposition of the actual data in the object range set in the object range setting step using the spectral data of each of the plurality of components, thereby generating chromatographic data of each of the plurality of components.
2. The data processing method according to claim 1, wherein in the data preparation step, the waveform of the chromatogram of the actual data is approximated by applying a pre-prepared peak model, and speculative data of the spectrum and speculative data of the chromatogram of each of the plurality of components are generated using the peak model applied to the chromatogram. The spectral data used in the similarity calculation step and the peak separation step is the speculative data of the spectrum generated in the data preparation step. The chromatographic data generated by the peak separation step is based on the speculative data of the chromatogram generated in the data preparation step.
3. The data processing method according to claim 1 or 2, wherein the matrix decomposition is non-negative matrix factorization.
4. The data processing method according to claim 1 or 2, wherein in the similarity calculation step, the similarity is calculated while changing the minimum wavelength and the wavelength width of the wavelength region.
5. A data processing system, characterized in that, Comprising: A data storage unit that stores actual data of a three-dimensional chromatogram including a chromatogram and a spectrum obtained by chromatographic analysis of a sample, and spectral data of each of a plurality of components in the sample where peaks overlap each other on the chromatogram of the actual data; and A data processing unit configured to perform a peak separation process of the plurality of components in the sample using the actual data and the spectral data stored in the data storage unit. The data processing unit is configured to execute: A similarity calculation step of calculating, for each wavelength region, the similarity between wavelength regions corresponding to the spectral data of each of the plurality of components stored in the data storage unit while changing the wavelength regions comprehensively; An object range setting step of searching for a wavelength region where the similarity is lower than the overall similarity between the spectral data of each of the plurality of components based on the calculation result in the similarity calculation step, and setting an object range; And A peak separation step, which uses the spectral data of each of the plurality of components to perform matrix decomposition of the actual data within the object range set in the object range setting step, thereby generating chromatographic data for each of the plurality of components.
6. The data processing system according to claim 5, wherein the data processing unit is configured to perform a data preparation step before the similarity calculation step. The data preparation step approximates the waveform of the chromatogram of the actual data by applying a pre-prepared peak model, and uses the peak model applied to the chromatogram to generate speculative data of the spectra and speculative data of the chromatogram for each of the plurality of components. In the similarity calculation step and the peak separation step, the speculative data of the spectra generated in the data preparation step is used as the spectral data, and in the peak separation step, the chromatographic data is generated based on the speculative data of the chromatogram generated in the data preparation step.
7. The data processing system according to claim 5 or 6, wherein the matrix decomposition is non-negative matrix factorization.
8. The data processing system according to claim 5 or 6, wherein the data processing unit is configured to calculate the similarity while changing the minimum wavelength and the wavelength width of the wavelength region in the similarity calculation step.
Citation Information
Patent Citations
Chromatogram data processing method and device
WO2016035167A1
Chromatogram data processing device and processing method
CN105026926A
Chromatogram data processing method and device
CN107076712A
Three-dimensional spectral data processing device and processing method
US20170356889A1