Data Processing System
The data processing system addresses the challenge of overlapping peaks in liquid chromatography by setting a maximum number of estimated components, ensuring accurate quantification of major and minor components while ignoring impurities.
Patent Information
- Application Number
- JP2022063252
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-04-06
- Publication Date
- 2026-01-15
- Estimated Expiration
- 2042-04-06
AI Technical Summary
Existing algorithms for quantifying components in liquid chromatography struggle with accurately determining peak components due to overlapping peaks, especially when impurities are present, and fail to allow selective quantification of major and minor components while ignoring impurities.
A data processing system that estimates peak components by repeatedly fitting a peak model function and sets a maximum number of components to be estimated, terminating the process when the set number is reached, regardless of data approximation, thereby preventing unnecessary peak components from being estimated.
The system effectively prevents the estimation of unnecessary peak components, allowing selective quantification of major and minor components by ignoring impurities, while ensuring accurate approximation to the original data.
Smart Images

Figure 0007799243000001 
Figure 0007799243000002 
Figure 0007799243000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a data processing system for processing three-dimensional chromatogram data. [Background technology]
[0002] In liquid chromatography (LC) using a multichannel detector such as a photodiode array (PDA) detector, three-dimensional chromatogram data having three dimensions, time, wavelength, and signal intensity (absorbance), can be obtained by continuously acquiring the absorption spectrum of a sample eluting from an analytical column.
[0003] When quantifying a target component in a sample using a liquid chromatograph, it is common to create a chromatogram using the wavelength at which the target component has the greatest absorbance, and then determine the area value of the target component's peak on that chromatogram to perform quantification. However, a sample may contain impurities other than the target component, and the peaks of these impurities may overlap with the target component's peak to form a single peak waveform. In such cases, it is not possible to determine the peak area values of the target component and impurities from the single peak waveform formed by the overlapping of multiple peaks, so it is necessary to estimate what peaks overlap to form that peak waveform.
[0004] As an algorithm for automatically estimating multiple peaks included in a peak waveform portion, an algorithm is known that fits a peak model function such as an EMG (Exponential Modified Gaussian) function to a chromatogram waveform while adjusting the parameters of the peak model function (see Patent Document 1). The algorithm disclosed in Patent Document 1 estimates a three-dimensional chromatogram for each peak component included in the peak waveform portion, and has an automatic component number estimation function that automatically estimates the number of peak components included in the target peak waveform portion by repeating the process of adding one peak if the loss of the estimation result (a value representing the degree of approximation of the three-dimensional chromatogram of synthesized data obtained by synthesizing the chromatograms and spectra of the estimated peak components to the original data; the smaller this value, the better the approximation to the original data). [Prior art documents] [Patent documents]
[0005] [Patent Document 1] International Publication No. 2016 / 035167 Summary of the Invention [Problem to be solved by the invention]
[0006] Although analysis using the algorithm with the automatic component number estimation function described above can be performed for any analysis range (wavelength range and retention time range), even a slight change in the analysis range can change the number of estimated peak components. This phenomenon is mainly due to the magnitude of the noise components contained in the data within the analysis range, and is difficult to resolve by modifying the algorithm, etc.
[0007] Furthermore, by using the above algorithm, it is possible to quantify not only the peaks of the major and minor components contained in the sample, but also the peaks of impurities that are at much lower concentrations than the major and minor components. However, this does not allow for cases where one wishes to quantify only the major and minor components while ignoring the impurities.
[0008] The present invention has been made in view of the above problems, and has as its object to prevent the presence of unnecessary peak components from being estimated while allowing the component number automatic estimation function to function effectively. [Means for solving the problem]
[0009] The data processing system according to the present invention comprises: an original data storage unit that stores original data of a three-dimensional chromatogram consisting of chromatogram data and spectra obtained by chromatography analysis; an arithmetic processing unit that is configured to execute a peak estimation process that estimates a peak included in one peak waveform portion of the original data stored in the original data storage unit by repeating a component estimation step that estimates a three-dimensional chromatogram of one peak component included in one peak waveform portion of the original data stored in the original data storage unit, until synthetic data obtained by synthesizing three-dimensional chromatograms of all estimated peak components whose three-dimensional chromatograms have been estimated in the component estimation step approximates the original data; and a maximum number storage unit that stores the maximum number of the estimated peak components, wherein the arithmetic processing unit is configured to terminate the peak estimation process when the number of the estimated peak components reaches the maximum number, regardless of the approximation status of the synthetic data to the original data.
[0010] That is, the data processing system according to the present invention is a system that executes a peak estimation process that estimates the number of peak components contained in one peak waveform portion and the three-dimensional chromatogram of each peak component. In the peak estimation process, a component estimation step that estimates a three-dimensional chromatogram of one peak component contained in a peak waveform portion is repeatedly executed, in principle, until the synthesized data obtained by synthesizing the respective three-dimensional chromatograms of all estimated peak components whose three-dimensional chromatograms have been estimated in the component estimation step approximates the original data. Meanwhile, a maximum number of estimated peak components is set, and when the number of estimated peak components reaches the set maximum number, the peak estimation process is terminated regardless of the state of approximation of the synthesized data of the estimated peak components to the original data.
[0011] For example, when analyzing a data range (peak waveform portion) where three peak components are estimated using an existing algorithm with an automatic component number estimation function, if the maximum number of estimated peak components is set to two, the peak estimation process will terminate when the number of estimated peak components reaches two without executing the next component estimation step. In this case, the combined data of the 3D chromatogram of the two estimated peak components may contain one or more missing peak components compared to the original data for the same range, but such missing components will be ignored. This is particularly useful when a sample is known to contain three components (major, minor, and impurity) but you want to quantify only the major and minor components while ignoring the presence of impurities. Conversely, when analyzing a peak waveform portion where two peak components are estimated using an existing algorithm, even if the maximum number of estimated peak components is set to three, the existing algorithm will not estimate the presence of three peak components, but will instead effectively estimate the presence of two peak components, as with the existing algorithm. This is clearly different from forcing the system to estimate the number of peak components specified by the user in the specified peak waveform portion.
[0012] Here, a peak waveform portion refers to a portion where one or more peaks are combined to form a single peak waveform. Furthermore, the composite data approximating the original data means that the difference between the composite data and the original data, calculated by a least squares method or the like, satisfies a predetermined condition, and thus the composite data can be evaluated as approximating the original data. One example of the predetermined condition is that the difference between the composite data and the original data is equal to or less than a predetermined threshold value. [Effects of the Invention]
[0013] As described above, the data processing system according to the present invention is equipped with an automatic component number estimation function that automatically estimates the number of peak components contained in a specified peak waveform portion by repeating the component estimation step, and is configured to terminate the peak estimation process when the component estimation step has been executed a set maximum number of times, regardless of the state of approximation of the original data by the synthesized data of the estimated peak components.Therefore, it is possible to prevent the presence of unnecessary peak components from being estimated from the data in the analysis range while allowing the automatic component number estimation function to function effectively. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 is a block diagram illustrating an example of a data processing system. [Figure 2] 10 is a flowchart illustrating a series of steps related to peak estimation processing. [Figure 3] 10 is a flowchart showing an example of an operation during peak estimation processing in the same embodiment. [Figure 4] 10A and 10B are diagrams for comparing the estimation results of peak estimation processing, where (A) is the case where the maximum number of estimated peak components is not set (comparison example), and (B) is the case where the maximum number of estimated peak components is set (example). DETAILED DESCRIPTION OF THE INVENTION
[0015] An embodiment of a data processing system according to the present invention will now be described with reference to the accompanying drawings.
[0016] FIG. 1 shows an embodiment of a data processing system.
[0017] The data processing system 1 includes an original data storage unit 2, an arithmetic processing unit 4, and a maximum number storage unit 6. Analysis data acquired by an analysis device 100 is input into the data processing system 1. The analysis device 100 is configured to perform liquid chromatography analysis on a sample and acquire an absorbance spectrum at regular intervals. In other words, the analysis data input into the data processing system 1 from the analysis device 100 is three-dimensional chromatogram data consisting of a chromatogram and a spectrum.
[0018] The original data storage unit 2 is a storage area that stores three-dimensional chromatogram data (hereinafter referred to as original data) acquired from the analysis device 100. The original data storage unit 2 can be realized by a nonvolatile flash memory, a hard disk drive, or the like.
[0019] The calculation processing unit 4 is configured to perform analysis processing of the original data of the three-dimensional chromatogram stored in the original data storage unit 2. The analysis processing of the original data by the calculation processing unit 4 includes a quantification processing that quantifies the concentrations of components contained in the sample from the area values of the peaks on the chromatogram of the original data, as well as a peak estimation processing that estimates the number of peak components contained in the peak waveform portion in the specified analysis range and the three-dimensional chromatogram of each peak component. The calculation processing unit 4 is a function realized by executing a program in a computer circuit equipped with a CPU (Central Processing Unit).
[0020] The maximum number storage unit 6 is a storage area for storing a set value of the maximum number of peak components (estimated peak components) estimated to be included in a specified peak waveform portion in the peak estimation process. The maximum number of estimated peak components can be set arbitrarily by the user.
[0021] A series of steps relating to the peak estimation process will be described with reference to the flowchart of FIG. 2 together with FIG.
[0022] First, when the user specifies the original data to be analyzed, the arithmetic processing unit 4 reads the specified original data (step 101). The arithmetic processing unit 4 displays the read original data of the three-dimensional chromatogram on a display (not shown) communicatively connected to the data processing system, and prompts the user to specify the analysis range (retention time range and wavelength range to be analyzed) (step 102), and further prompts the user to set the maximum number of estimated peak components (step 103). The maximum number of estimated peak components set by the user is stored in the maximum number storage unit 6. The setting of the maximum number of estimated peak components may be performed before the specification of the analysis range. Thereafter, when the user inputs an instruction to perform peak estimation processing, the arithmetic processing unit 4 executes peak estimation processing (step 104).
[0023] An example of the operation during the peak estimation process will be described with reference to the flowchart of FIG.
[0024] When the peak estimation process starts, the number N of estimated peak components is 0 (step 201). When the peak estimation process starts, the calculation processing unit 4 executes component estimation steps 202-204 using a peak model function prepared in advance to identify the position and magnitude of a peak estimated to be included in the peak waveform portion appearing in the chromatogram of the analysis target range. In the component estimation steps 202-204, the calculation processing unit 4 first fits the peak model function to the target peak waveform portion while adjusting parameters such as the height and width of the peak model function (step 202). The calculation processing unit 4 estimates the position and magnitude of the peak model function fitted to the peak waveform portion as a peak included in the peak waveform portion, and then calculates and estimates a three-dimensional chromatogram consisting of the chromatogram and spectrum of the peak component (step 203). This increases the number N of peak components (estimated peak components) for which a three-dimensional chromatogram has been estimated by 1 (step 204).
[0025] After adding one estimated peak component in the above component estimation steps 202-204, the calculation processing unit 4 determines whether the total number N of estimated peak components has reached a preset maximum number (step 205). If the number N of estimated peak components has not reached the set maximum number (step 205: No), the calculation processing unit 4 synthesizes the three-dimensional chromatograms of all estimated peak components to create synthetic data (step 206). The calculation processing unit 4 calculates the loss of the created synthetic data relative to the original data using the least squares method or the like (step 207), and determines whether the calculated loss is equal to or less than a predetermined value (step 208). If the loss is equal to or less than the predetermined value (step 208: Yes), it is determined that the synthetic data approximates the original data, and the peak estimation process is terminated.
[0026] On the other hand, if the loss of the synthesized data relative to the original data exceeds the predetermined value (step 208: No), the calculation processing unit 4 executes the component estimation steps 202-204 again to add one more estimated peak component. Thereafter, the calculation processing unit 4 determines whether the number N of estimated peak components has reached the set maximum number (step 205), and if the number N of estimated peak components has reached the set maximum number (step 205: Yes), the calculation processing unit 4 ends the peak estimation process without executing steps 206 and 207.
[0027] FIG. 4(A) shows the estimation result (comparison example) when the peak estimation process is performed without setting the maximum number of estimated peak components, and FIG. 4(B) shows the estimation result when the peak estimation process is performed with the maximum number of estimated peak components set.
[0028] If peak estimation processing is performed for a peak waveform portion of a certain data range without setting a maximum number of estimated peak components, the component estimation step is repeated until the loss of the estimated peak components from the combined data relative to the original data falls below a predetermined value, resulting in an estimated result that the peak waveform portion contains three peaks: main component A, subcomponent B, and impurity C, as shown in (A). If peak estimation processing is performed for the same peak waveform portion of the data range with the maximum number of estimated peak components set to 2, the number of estimated peak components reaches the set maximum of 2 before the loss of the estimated peak components from the combined data relative to the original data falls below the predetermined value, and only two peaks, main component A and subcomponent B, are estimated within the peak waveform portion. In other words, the presence of the peak of impurity C is not estimated and is ignored. Furthermore, if the data range to be analyzed is slightly changed and peak estimation processing is performed, the number of estimated peaks may be 2 or 3 if the maximum number of estimated peak components is not set. However, if the maximum number of estimated peak components is set to 2, the number of estimated peaks will remain constant.
[0029] The above-described embodiment merely exemplifies the embodiment of the data processing system according to the present invention. The embodiment of the data processing system according to the present invention is as follows.
[0030] In one embodiment of the data processing system according to the present invention, the system includes: an original data storage unit that stores original data of a three-dimensional chromatogram consisting of chromatogram data and spectra acquired by chromatography analysis; an arithmetic processing unit that is configured to execute a peak estimation process that estimates a peak included in one peak waveform portion of the original data stored in the original data storage unit, by repeating a component estimation step that estimates a three-dimensional chromatogram of one peak component included in one peak waveform portion of the original data stored in the original data storage unit, until synthetic data obtained by synthesizing three-dimensional chromatograms of all estimated peak components whose three-dimensional chromatograms have been estimated in the component estimation step approximates the original data; and a maximum number storage unit that stores a maximum number of the estimated peak components, wherein the arithmetic processing unit is configured to terminate the peak estimation process when the number of the estimated peak components reaches the maximum number, regardless of the approximation status of the synthetic data to the original data.
[0031] In a first aspect of the above embodiment, the calculation processing unit is configured to evaluate the loss of the composite data of the three-dimensional chromatogram of all the estimated peak components relative to the original data each time the component estimation step is executed until the number of the estimated peak components reaches the maximum number, terminate the peak estimation process when the loss satisfies a predetermined condition, and terminate the peak estimation process regardless of the loss when the number of the estimated peak components reaches the maximum number.
[0032] In a second aspect of the above embodiment, the maximum number can be set arbitrarily by the user.
[0033] In a third aspect of the above embodiment, the peak waveform portion is included in an analysis range specified by the user, allowing the user to arbitrarily select a peak waveform portion for which peak estimation processing is desired to be performed.
[0034] In the third aspect, the peak waveform portion may be a portion where multiple peaks overlap to form a single peak waveform. This allows the user to specify a portion that is considered to be a shape where multiple peaks overlap as the analysis range, and perform peak estimation processing on that peak waveform portion.
[0035] In a fourth aspect of the above embodiment, the calculation processing unit is configured to estimate the position and magnitude of one peak included in the peak waveform portion of the chromatogram by applying a peak model function prepared in advance to the peak waveform portion of the chromatogram while adjusting parameters in the component estimation step. [Explanation of symbols]
[0036] 1. Data Processing System 2. Original data storage section 4. Processing unit 6 Maximum number storage 100 Analyzer
Claims
1. an original data storage unit that stores original data of a three-dimensional chromatogram consisting of chromatogram data and spectra obtained by a chromatographic analysis; a calculation processing unit configured to execute a peak estimation process for estimating a peak included in one peak waveform portion of the original data stored in the original data storage unit, by repeating a component estimation step of estimating a three-dimensional chromatogram of one peak component included in one peak waveform portion of the original data stored in the original data storage unit until synthesized data obtained by synthesizing three-dimensional chromatograms of all estimated peak components whose three-dimensional chromatograms have been estimated in the component estimation step approximates the original data; a maximum number storage unit that stores the maximum number of the estimated peak components, The calculation processing unit is configured to terminate the peak estimation process when the number of estimated peak components reaches the maximum number, regardless of the approximation status of the synthesized data to the original data.
2. 2. The data processing system according to claim 1, wherein the calculation processing unit is configured to evaluate a loss of the composite data of the three-dimensional chromatogram of all the estimated peak components relative to the original data each time the component estimation step is executed until the number of the estimated peak components reaches the maximum number, terminate the peak estimation process when the loss satisfies a predetermined condition, and terminate the peak estimation process regardless of the loss when the number of the estimated peak components reaches the maximum number.
3. 2. The data processing system according to claim 1, wherein said maximum number can be arbitrarily set by a user.
4. 2. The data processing system according to claim 1, wherein the peak waveform portion is included in an analysis range designated by a user.
5. 5. The data processing system according to claim 4, wherein the peak waveform portion is a portion where a plurality of peaks overlap to form one peak waveform.
6. 6. The data processing system according to claim 1, wherein the calculation processing unit is configured to estimate the position and magnitude of one peak included in the peak waveform portion of the chromatogram by fitting a pre-prepared peak model function to the peak waveform portion of the chromatogram while adjusting parameters in the component estimation step.
Citation Information
Patent Citations
Comprehensive two-dimensional gas chromatography modulation peak multi-level limitation classification method based on characteristic value confirmation
CN111272925A
Chromatogram data processing device and processing method
WO2013035639A1
Peak detection method
WO2015033478A1
Chromatogram data processing method and device
WO2016035167A1
Adaptive asymmetrical signal detection and synthesis methods and systems
WO2020092955A1