Analysis device and analysis method

By constructing a learning model and using a U-Net neural network to analyze chromatograms or spectral waveforms, and calculating and displaying the confidence level of peaks, the problem of the inability to assess the confidence level of peak sampling results in existing technologies is solved, thereby improving the accuracy and efficiency of peak detection.

CN116519861BActive Publication Date: 2026-07-28SHIMADZU SEISAKUSHO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHIMADZU SEISAKUSHO LTD
Filing Date
2023-01-05
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing peak sampling methods using semantic segmentation techniques cannot output the confidence level of the peak sampling results, making it impossible to accurately assess the reliability of the peak.

Method used

By constructing a learned model, semantic segmentation technology is used to analyze the object waveform of a chromatogram or spectrum, segment it into multiple waveform parts, and use neural network models such as U-Net to make judgments, calculate the confidence level of the peak part, display the judgment result and confidence level, and provide a correction interface.

Benefits of technology

It enables the calculation and display of peak confidence during peak sampling using semantic segmentation technology, simplifying the user's confirmation and correction process and improving the accuracy and efficiency of peak detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116519861B_ABST
    Figure CN116519861B_ABST
Patent Text Reader

Abstract

The present application can calculate the confidence of peak picking when peak picking using a semantic segmentation technique is performed. The analysis device (1) divides an object waveform into a plurality of partial waveforms (S12), determines a peak waveform that is a peak portion among the plurality of divided partial waveforms using a learned model (S14), and calculates the confidence of the determination result of the peak waveform using data output from the learned model when the peak portion of the object waveform is determined using the learned model (S17).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to an analytical apparatus and method for analyzing the waveforms of chromatograms and spectra. Background Technology

[0002] Traditionally, chromatographs are used to identify or quantify the components contained in a sample. In a chromatograph, a column is used to separate the components in the sample, and the components eluting from the column are detected sequentially. A chromatogram is then created with time on the horizontal axis and detection intensity on the vertical axis.

[0003] To determine the peak height and area from a chromatogram, it is necessary to identify the starting and ending points of the peak rising from the baseline. This process of identifying the peak starting and ending points is called peak picking. By determining these points, the peak height and area can be determined. Based on the peak height and area, the concentration of the corresponding compound can be calculated.

[0004] In recent years, efforts have been made to automate peak sampling using deep learning. Methods utilizing deep learning for peak sampling include those employing object detection techniques and those employing semantic segmentation techniques.

[0005] International Publication No. 2020 / 225864 discloses a method that displays the confidence level of peak-picking results using a Single Shot Multibox Detector (SSD) by formalizing the peak-picking problem as object detection in the field of image recognition. The SSD outputs the peak-picking results along with the confidence level of the results. In contrast, Kanazawa S, 10 others, Fake metabolomics chromatogram generation for facilitating deep learning of peak-picking neural networks. J Biosci Bioeng. 2021 Feb; 131(2):207-212. doi:10.1016 / j.jbiosc.2020.09.013. Epub 2020 Oct10.PMID:33051155. discloses a method that performs peak-picking using U-Net by formalizing peak-picking as a semantic segmentation problem. Summary of the Invention

[0006] However, there is currently no method for calculating confidence level in peak sampling using semantic segmentation techniques. Therefore, existing peak sampling methods using semantic segmentation techniques output the sampling result but not the confidence level of the output result.

[0007] The purpose of this disclosure is to enable the calculation of confidence level in peak sampling when using semantic segmentation techniques.

[0008] The analytical apparatus of one aspect of this disclosure analyzes an object waveform as a chromatogram or spectrum, comprising: a processor; and a memory storing a learned model created through machine learning, the machine learning using multiple sets of groups including multiple partial waveforms created by segmenting a reference waveform whose peak positions are known, the processor segmenting the object waveform into multiple partial waveforms, using the learned model to determine the peaks of the object waveform, classifying the object waveform into peak regions with continuous peaks and non-peak regions outside the peak regions based on the determination results of the peaks of the object waveform, and calculating the confidence level of the determination results of the peaks using data output from the learned model when determining the peaks of the object waveform using the learned model.

[0009] One aspect of the analytical method disclosed herein is for analyzing an object waveform as a chromatogram or spectrum, comprising: a step of creating a learned model, the learned model determining the peak portions contained in the input waveform through machine learning, the machine learning using multiple sets including multiple partial waveforms created by segmenting a reference waveform whose peak positions are known; a step of segmenting the object waveform into multiple partial waveforms; a step of using the learned model to determine the peak portions of the object waveform; a step of classifying the object waveform into peak regions with continuous peak portions and non-peak regions outside the peak regions based on the determination results of the peak portions of the object waveform; and a step of calculating the confidence level of the determination results using data output from the learned model when determining the peak portions of the object waveform using the learned model.

[0010] The described and other objects, features, aspects and advantages of the invention will be shown by the following detailed description in relation to the invention, which is understood in conjunction with the accompanying drawings. Attached Figure Description

[0011] Figure 1 It is a block diagram representing the overall structure of the analytical device.

[0012] Figure 2 This is a diagram showing an example of a chromatogram.

[0013] Figure 3 It is a flowchart used to explain the sequence of making a learning model.

[0014] Figure 4 It is a flowchart used to explain the sequence of creating a learning model.

[0015] Figure 5 It is a flowchart used to explain the order of using learned models to determine chromatogram data.

[0016] Figure 6 This is a diagram representing an example of the judgment result of the learning model.

[0017] Figure 7 This is an example of a graph representing labeling processing based on the judgment results.

[0018] Figure 8 This is an example of an image that displays the confidence level along with the decision result.

[0019] Figure 9 This is an example of an image representing an operation that corrects the judgment result.

[0020] Figure 10 This is a graph showing the relationship between the confidence level of the peak and the correct solution rate.

[0021] Figure 11 These are figures representing various variations of the method for calculating the confidence level of a peak, examples 1 to 7.

[0022] Figure 12 This is a diagram used to illustrate variation 7. Detailed Implementation

[0023] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Furthermore, the same or equivalent symbols used to mark the same parts in the drawings will not be described again.

[0024] Figure 1 This is a block diagram showing the overall structure of the analysis device 1. The analysis device 1 includes a processor 10 that functions as a control unit, a memory 20 that functions as a storage unit, and an input / output port 30. A mouse 40, a keyboard 50, and a display device 60 are connected to the input / output port 30. A mass spectrometer or similar device can also be connected to the input / output port 30. One or more terminal devices can also be connected to the input / output port 30 via the Internet or an internal network.

[0025] The parsing device 1 is configured, for example, based on a personal computer. The parsing device 1 may include a server that can be accessed from one or more terminal devices via a network such as the Internet.

[0026] The input / output port 30 receives the measurement data (chromatogram data) of the target object and the learning data used for machine learning. Alternatively, the measurement data of the target object can be input via a mass spectrometer connected to the input / output port 30. A liquid chromatography-mass spectrometry system can also be constructed using a mass spectrometer, a liquid chromatograph connected to the mass spectrometer, and the analysis device 1.

[0027] The memory 20 stores at least the learning data 210 input to the input / output port 30, the measurement data 213 input to the input / output port 30, the estimation model 300 used for machine learning, and the analytical program 200 used to perform analytical processing and machine learning.

[0028] The learning data 210 is categorized into training data 211 and validation data 212. Training data 211 and validation data 212 are waveform data of chromatograms obtained by measuring samples containing various components using a chromatographic mass spectrometry apparatus. For example, a chromatogram may be a total ion chromatogram, which uses a mass spectrometer to perform mass spectrometry (MS) scans on components separated by liquid chromatography and represents the time change of the total intensity of ions with all detected mass-charge ratios. Chromatograms may also be mass chromatograms measured using selected ion monitoring (SIM) or multiple reaction monitoring (MRM) and representing the time change of the intensity of ions with specific mass-charge ratios.

[0029] The training data 211 and validation data 212 include data on the positions of peaks determined in advance by peak sampling. These waveform data are pre-standardized to a specified range (e.g., ±1.0) for intensity values. Standardization unifies multiple chromatograms with different intensity scales into a common intensity scale, thereby improving the accuracy of the learned model. Here, chromatograms obtained from measurements of actual samples are used as training data 211 and validation data 212, but chromatograms generated through simulation can also be used.

[0030] The chromatogram waveform is divided into a predetermined number of segments along the time axis. This predetermined number is, for example, 512 or 1024, and the width (length along the time axis) of each segment is set to be at least less than the peak width. This predetermined number is determined, for example, based on the peak width and the number of data points required to constitute a peak.

[0031] Each waveform data segment corresponds to information related to the characteristics of that segment (characteristic information). The characteristic information corresponding to a segment of the waveform includes at least information indicating whether the segment belongs to a peak region or a non-peak region.

[0032] The analysis program 200 comprises a segmentation unit 201, a model making unit 202, a judgment unit 203, a calculation unit 204, an image processing unit 205, and an output unit 206.

[0033] The segmentation unit 201 segments the waveform of the chromatogram into a predetermined number of partial waveforms. The model creation unit 202 uses the learning data 210 to perform machine learning on the estimation model 300, and creates a learned estimation model 300. The determination unit 203 uses the learned estimation model 300 to sample the peaks of the chromatogram. Hereinafter, the learned estimation model 300 will sometimes be referred to as the "learned model".

[0034] The calculation unit 204 calculates the confidence level of the determination result of the determination unit 203. The image processing unit 205 generates image data including the determination result and confidence level. The output unit 206 outputs a display signal including the image data from the input / output port 30 to the display device 60. Alternatively, the analysis device 1 may also include the display device 60.

[0035] Figure 2 This is an example of a chromatogram. Here, the names of the parts determined from the chromatogram are simply explained. A chromatogram can be classified into a baseline portion and a peak region. The rising portion from the baseline is called the peak start point and peak end point. The region between the peak start point and peak end point is called the peak region. The part of the peak region with the strongest detection intensity (the strongest part) is called the peak apex.

[0036] like Figure 2 As illustrated, a peak region includes individual peaks. In cases where unseparated peaks appear in the chromatograph waveform, the peak region includes both individual peaks and unseparated peaks. For example, the portion of the detection intensity of the valley between two connected mountain-shaped waveforms, where the peak apex is not lower than the intensity corresponding to the baseline, is called an unseparated peak.

[0037] Next, the sequence of creating the learning model will be explained with reference to the flowchart. Figure 3 It is a flowchart used to illustrate the sequence of creating a learning model. For example... Figure 3 As shown, the model creation unit 202 of the analysis device 1 functions as a learning device. The model creation unit 202 learns the estimation model 300 based on the input learning data 210. The estimation model 300 performs deep learning using a neural network. The estimation model 300 includes parameters such as weighting coefficients used in the calculations by the neural network.

[0038] To enable the inference model 300 to learn, a supervised learning algorithm is used, for example. The model creation unit 202 enables the inference model 300 to learn by using supervised learning of the learning data 210.

[0039] The learning of the estimation model 300 utilizes semantic segmentation. Semantic segmentation is typically used to parse images containing pixel data with a two-dimensional distribution. In this embodiment, semantic segmentation is applied to the parsing of waveforms from a chromatogram containing data arranged one-dimensionally along the time axis. Learning models capable of performing semantic segmentation can include, for example, U-Net, SeGNet, PSPNet, etc. In this embodiment, U-Net is used.

[0040] The model creation unit 202 receives a portion of the waveform from the chromatogram and its corresponding forward solution data. The forward solution data may be, for example, the results of pre-determined peak sampling. Peak sampling results may also include the peak apex.

[0041] The model making unit 202 determines the peak sampling result based on the input learning data 210 and the estimation model 300, and then enables the estimation model 300 to learn based on the determination result and the correct solution data. Specifically, the model making unit 202 enables the estimation model 300 to learn by adjusting the parameters within the estimation model 300 in a way that the result obtained using the estimation model 300 is close to the correct solution data.

[0042] Figure 4 This is a flowchart used to explain the sequence of creating a learning model. The processing of this flowchart is achieved by executing a part of the analysis program 200 through the processor 10 of the analysis device 1.

[0043] First, the processor 10 detects the operation of starting the learning of the inference model 300 (step S1). For example, the operation is detected in step S1 when the user uses the mouse 40 and keyboard 50 to start the learning of the inference model 300.

[0044] Next, the processor 10 reads the learning data 210 (training data 211 and validation data 212) from the memory 20 (step S2). Then, the processor 10 inputs the training data 211 into the estimation model 300 (step S3). Next, in the estimation model 300, learning processing using deep learning is performed (step S4). In this embodiment, the weights of the neural network used for learning the estimation model 300 are adjusted to obtain correct characteristic information based on partial waveforms.

[0045] More specifically, based on a portion of the waveforms in the training data 211 and the corresponding characteristic information of those portions, the parameters of the inference model 300 are adjusted. During the parameter adjustment process, the inference of individual peaks, unseparated peaks, peak start points, peak end points, and baselines are processed, as well as the comparison between the inference results and the correct solution data is performed.

[0046] Next, the processor 10 stores the inference model 300, which is generated based on the learning processing results of step S4, in the memory 20 (step S5). Next, the processor 10 confirms the correct answer rate of the characteristic information assigned to the partial waveform of the verification data 212 by the inference model 300 (step S6).

[0047] Next, the processor 10 determines whether a predetermined termination condition is met (step S7). For example, if the number of times the learning process is repeated using the training data 211 reaches a predetermined number, the processor 10 determines that the termination condition is met. If the termination condition is not met, the processor 10 repeats steps S3 to S6 until the termination condition is met.

[0048] If the termination condition is met, the processor 10 selects a suitable presumed model 300 from the plurality of presumed models 300 stored in the memory 20, and saves the selected presumed model 300 as a learned model in the memory 20 (step S8).

[0049] Thus, processor 10 ends. Figure 4 A series of processing steps are involved. The learned model is selected based on criteria such as the highest correct answer rate for the validation data 212, or the absence of overlearning. Furthermore, an example is shown here where the inference model 300 is stored in memory 20 after each learning iteration. However, it is also possible to repeatedly update the same inference model 300 before the predetermined number of learning iterations is reached, and then store the inference model 300 in memory 20 when the predetermined number of learning iterations is reached.

[0050] Next, the flowchart will be used to illustrate the sequence of analyzing the waveforms of unanalyzed chromatograms. Figure 5 This is a flowchart illustrating the sequence of determining chromatogram data using a learned model (learned presumption model 300). The processing described in this flowchart is implemented by the processor 10 of the analysis device 1 executing a portion of the analysis program 200.

[0051] First, the processor 10 obtains chromatogram data (measurement data) (step S11). The chromatogram data is input into the analysis device 1 via a measuring instrument such as a mass spectrometer connected to the input / output port 30, or a terminal device connected to the input / output port 30.

[0052] Next, the processor 10 divides the obtained chromatogram waveform into a predetermined number of partial waveforms (step S12). The number of chromatogram waveforms can be the same as the number of training data 211 and validation data 212, or it can be a different number.

[0053] The number of segments is determined based on the length of the waveform (the length of the chromatography-mass spectrometry execution time), ensuring that the width of each waveform segment (the length along the time axis) is at least less than the width of the predicted peak contained in the chromatogram. For example, consider setting the number of segments to 512 or 1024.

[0054] Next, the processor 10 inputs a portion of the waveform into the learned estimation model 300 (learned model) (step S13). Then, the learned model is used to determine whether the portion of the waveform belongs to a peak region, and labeling processing is performed (step S14). More specifically, the start and end points of peaks, baselines, single peaks, unseparated peaks, peak apexes, etc., are determined based on the portion of the waveform. Furthermore, the weight of each determination result is calculated. In step S14, characteristic information (information on whether it belongs to a peak region) is labeled for each portion of the waveform.

[0055] Next, the processor 10 calculates the confidence level of the peak (step S17). The confidence level of the peak is calculated based on the average of the weights corresponding to the peak start point determined by the learned model and the weights corresponding to the peak end point determined by the learned model.

[0056] Next, the processor 10 generates a graph representing the determination result and confidence level (step S18). In this embodiment, the processor 10 generates various graphs. The processor 10 outputs a display signal to the display device 60 to display the generated graph (step S19). The display device 60 then displays the determination result and confidence level. For example, the peak start point, peak end point, and confidence level are displayed on the waveform of the chromatogram on the screen of the display device 60.

[0057] Next, the processor 10 determines whether a correction instruction for the peak start point and end point has been detected (step S20). In this embodiment, the user can perform the operation of correcting the peak start point and end point on the screen of the display device 60. If no correction instruction is detected, the processor 10 proceeds to step S22.

[0058] When the user performs the operation of correcting the peak start point and end point using the mouse 40 and keyboard 50, the processor 10 corrects the data on the screen according to the correction instruction (step S21). As described above, the processor 10 accepts the user's correction instruction and corrects the peak start point and end point.

[0059] After correcting the data, the processor 10 determines whether an operation to determine the data has been detected (step S22). If no operation to determine the data is detected, the processor 10 returns to step S20. If an operation to determine the data is detected, the processor 10 stores the determination result (or the corrected determination result if the data has been corrected) in the memory 20 (step S23), and ends the processing based on this flowchart.

[0060] Figure 6 This is a diagram representing an example of the judgment result of the learning model. Figure 6 The curve above represents the waveform W0 of the input chromatogram. Figure 6 The graph below represents the learned model's judgment of the input chromatogram. The horizontal axis (index) of both graphs corresponds to the time axis. Figure 6 The vertical axis of the curve above represents intensity. Figure 6 The vertical axis of the graph below represents the weights output by the learned model. The weights are normalized to the range of 0 to 1.

[0061] The waveforms W1 to W5, as the determination results of the learning model, correspond to the baseline, a single peak, an unseparated peak, a peak start point, and a peak end point, respectively. By comparing the waveform W0 of the chromatogram with waveforms W1 to W5, it can be seen that, for example, in waveform W0 of the chromatogram, the peak start point corresponding to the position of the index Is has the highest weight. Similarly, it can be seen that in waveform W0 of the chromatogram, the peak end point corresponding to the position of the index Ie has the highest weight. In this case, for example, the analysis device 1 determines the position of the index Is in waveform W0 of the chromatogram as the peak start point and the position of the index Ie as the peak end point.

[0062] Here, the peak start point, peak end point, single peak, unseparated peak, and baseline are listed as examples of the judgment objects, but other elements such as peak tops can also be added to the judgment objects.

[0063] like Figure 6 As shown, the processor 10 determines the confidence level of a peak by calculating the average of the weight Ws corresponding to the peak start point Is determined by the learning model and the weight We corresponding to the peak end point Ie determined by the learning model.

[0064] Figure 7 This is an example of a graph representing labeling processing based on the judgment results. Figure 7 The curve above and Figure 6 The curve shown below is the same. Figure 7 The graph below is based on waveforms W1 to W5 against waveform W0 of the input chromatogram (refer to...). Figure 6 The curve is generated by labeling. Labels 0 to 4 correspond to the baseline, single peak, unseparated peak, peak start point, and peak end point, respectively.

[0065] For example, the labeling process can be performed in the following order: Among waveforms W1 to W5, select the waveform with the highest weight at a certain index Ix, and label the value of the index Ix on the selected waveform. Then, change x from its initial value to its final value and repeat the same process, thereby ending the labeling process. For example, Figure 7 The graph shows the range from index 0 to Is, with labels (label = 0) as the baseline.

[0066] Figure 8 This is an example of an image 61 that displays the confidence level along with the determination result. Image 61 is displayed by a display device 60. Image 61 shows the waveform of the chromatogram of the measured object and the peak start point Is and peak end point Ie corresponding to the determination result. Furthermore, image 61 displays the confidence level with respect to the determined peak start point Is and peak end point Ie. By observing image 61, the user can identify the accuracy of the determination result.

[0067] In addition to image 61, processor 10 is also capable of including Figure 6 The image shows two curves of the shape, including Figure 7 The images of the two curves shown, and the .... Figure 6 and Figure 7 The three graphs, arranged vertically, are selectively displayed on the display device 60. Each graph contains... Figure 8 The displayed form together indicates the level of confidence. The user can use the mouse 40 and keyboard 50 to input instructions into the analysis device 1 indicating whether to display a certain image.

[0068] Figure 9 This is an example of image 62 representing an operation to correct the judgment result. Image 62 is displayed by display device 60. Image 62 includes, in addition to... Figure 8 In addition to the content shown, icons 65 and 66 are also displayed to correct the positions of the peak start point Is and the peak end point Ie.

[0069] Icon 65 corresponds to the peak start point Is. When the user manipulates icon 65 using mouse 40 and keyboard 50, the position of the peak start point Is changes. When the user manipulates icon 66 using mouse 40 and keyboard 50, the position of the peak end point Ie changes. In conjunction with these changes in the positions of the peak start point Is and the peak end point Ie, the position of the index and the confidence level displayed below the graph also change.

[0070] After the user corrects the positions of the peak start point Is and the peak end point Ie to appropriate positions, the user performs a data confirmation operation. If the processor 10 detects the data confirmation operation, the corrected result is stored in the memory 20.

[0071] In addition, this is shown here with Figure 8 The example shown is based on image 61, used to display icons 65 and 66. However, it can also be used for other purposes, including... Figure 6 The image shows two curves of the shape, including Figure 7 The images of the two curves shown, and the .... Figure 6 and Figure 7 The image, composed of three graphs arranged vertically, displays icons 65 and 66 used to correct the judgment results.

[0072] As described above, in this embodiment, the judgment result and confidence level of the learned model are displayed on the display device 60. Therefore, the user can visually identify accurate peak information and peak information with lower reliability. As a result, it becomes simpler for the user to visually confirm or correct the instructions, thus reducing the user's burden in this operation. Moreover, when analyzing waveforms with multiple observed peaks, the number of peaks that the user should confirm is reduced, thereby preventing errors or omissions in the confirmation process.

[0073] Next, an example of creating a learning model using actual chromatographic data and performing waveform analysis of the chromatograms will be explained. When creating the learning model, 30 sets of chromatograms of primary metabolites were prepared. Each set included 475 chromatograms. Peaks were manually sampled from each prepared chromatogram. The waveforms of the chromatograms were then categorized into five types: baseline, peak start point, peak end point, single peak, and unseparated peak, and each was labeled. This created the learning data. Cross-validation was performed using the prepared learning data. In the cross-validation evaluation, 30 cross-validation tests were conducted, using one set from each of the 30 sets as validation data.

[0074] The average weight of the peak starting point output from the learned model is calculated by adding the weight of the peak starting point output from the learned model to the average weight of the peak starting point and dividing by 2. This average weight is then set as the confidence level of the peak. The relationship between the confidence level and the correct solution rate is then verified. Figure 10 The verification results are shown in the figure.

[0075] Figure 10 This is a graph showing the relationship between peak confidence and the positive resolution rate. In Figure 10 In this context, TP represents the number of positive solutions, and FP represents the number of negative solutions. For example... Figure 10As shown, the higher the confidence level, the higher the correct answer rate. Therefore, the confidence level calculation method disclosed in this embodiment is effective.

[0076] Next, refer to Figure 11 A variation of the method for calculating the confidence level of a peak is explained. Figure 11 These are figures representing variations 1 to 7 of the method for calculating the confidence level of the peak. Furthermore, waveforms W1 to W5 used in the following descriptions of the variations are shown in... Figure 6 and Figure 7 middle.

[0077] like Figure 11 As shown, the confidence level of a peak can be calculated using any one of the following: baseline (Modification 1), individual peak (Modification 2), peak start point (Modification 3), peak end point (Modification 4), and peak apex (Modification 5).

[0078] Variation Example 1 is an example of using a baseline to calculate the confidence level of a peak. For example... Figure 11 As shown, the confidence level can be calculated as "1 - (the average weight of the exponential portion belonging to the peak region in the baseline waveform W1)". Here, the exponential portion belonging to the peak region, for example, means... Figure 6 The range of the index Is to the index Ie.

[0079] Variation 2 is an example of calculating peak confidence using a single peak. For example... Figure 11 As shown, the confidence level can be calculated by "the average weight of the exponential portion belonging to the peak region in the waveform W2 of a single peak".

[0080] Variation 3 is an example of calculating peak confidence using the peak start point. For example... Figure 11 As shown, the confidence level can be calculated using the "average weight of the exponential portion corresponding to the waveform W4 at the peak start point". For example, in Figure 6 In this process, the weights corresponding to waveform W4 are determined by taking each index within the range from the initial value to the terminal value of the index as the object, and the average value of all determined weights is calculated, thereby deriving the confidence level.

[0081] Variation 4 is an example of using the peak termination point to calculate the confidence level of a peak. For example... Figure 11 As shown, the confidence level can be calculated by "the average weight of the exponential portion corresponding to the waveform W5 at the peak end point".

[0082] Variation 5 is an example of using the peak apex to calculate the confidence level of a peak. For example... Figure 11 As shown, the confidence level can be calculated by "the average weight of the index portion corresponding to the peak".

[0083] Variation 6 is an example of calculating peak confidence by combining individual peaks, unseparated peaks, and baselines. For example... Figure 11 As shown, the confidence level can be calculated using "(B+C) / (A+B+C)". Here, A, B, and C are as described below.

[0084] A: The sum of the weights of the exponential portion belonging to the peak region in the baseline waveform W1.

[0085] B: The sum of the weights of the exponential portion belonging to the peak region in the waveform W2 of a single peak.

[0086] C: The sum of the weights of the exponential portion belonging to the peak region in the waveform W3 without peak separation.

[0087] Variation 7 is an example of calculating peak confidence by combining the baseline, unseparated peak, peak start point, and peak end point. For example... Figure 11 As shown, the confidence level can be calculated using "X / (X+Y)". Here, X and Y are as described below.

[0088] X: The index corresponding to any of the labels 2 to 4 in the peak region.

[0089] Y: The index corresponding to label 0 in the peak region.

[0090] Reference Figure 12 Further detailed explanation of variation 7. Figure 12 This is a diagram used to illustrate variation 7. Figure 12 This involves assigning a graphical representation of regions Xa, Xb, and Ya to the curve obtained by labeling the judgment results, used to illustrate variation example 7. Figure 12 In the graph shown, part of the peak region includes the baseline. Sometimes, the relationship between the learned model and the measured object can be used to create a graph. Figure 12 The judgment result of the curve shown.

[0091] In the formula for calculating the confidence level related to Variation 7, X is the exponent corresponding to any of the labels 2 to 4 in the peak region. It is equivalent to adding the exponent of region Xa to the exponent of region Xb.

[0092] In the confidence calculation formula related to Variation 7, Y is the exponent corresponding to label 0 in the peak region. It is equivalent to the exponent of region Ya.

[0093] As explained above, the parsing apparatus 1 according to this embodiment is capable of calculating the confidence level of the determination result. In particular, the parsing apparatus 1 according to this embodiment is characterized in that it performs peak sampling using semantic segmentation techniques and calculates the confidence level of the determination result.

[0094] In peak sampling using deep learning, methods applying object detection techniques from the field of image recognition and methods applying semantic segmentation techniques are known. Non-patent literature 1 describes how standardizing the peak sampling problem using semantic segmentation improves performance compared to standardizing it using object detection. However, conventional peak sampling methods using semantic segmentation lack methods for calculating confidence levels.

[0095] The parsing device 1 according to this embodiment can perform peak sampling using semantic segmentation technology, calculate the confidence level of the determination result, and display the determination result and confidence level on the display device 60. Furthermore, the parsing device 1 provides an interface for the user to correct the determination result. Thus, the user can easily and efficiently confirm peak information such as the start and end points of the peaks detected by peak sampling, and make corrections as needed. As a result, according to this embodiment, a parsing device 1 capable of outputting highly accurate peak detection results can be provided.

[0096] These embodiments are all examples and can be modified appropriately according to the spirit of this disclosure. Here, the example of processing the waveform of a chromatogram obtained by chromatography-mass spectrometry is described. However, chromatograms obtained by chromatographs including detectors other than mass spectrometers (spectrometers) and gas chromatographs can also be analyzed by the analysis device 1. Furthermore, the object of analysis is not limited to chromatograms. For example, the spectrophotometer (a waveform representing the change of detection intensity relative to wavelength or wavenumber axis) obtained by spectrophotometer measurement can also be used as the object of analysis. Alternatively, any waveform obtained from liquid chromatography (LC), gas chromatography (GC), liquid chromatography with photodiode array detector (LC-PDA), liquid chromatography-mass spectrometry (LC / MS), gas chromatography-mass spectrometry (GC / MS), liquid chromatography-tandem mass spectrometry (LC / MS / MS), gas chromatography-tandem mass spectrometry (GC / MS / MS), or liquid chromatography-mass spectrometry-ion trap-time offlight (LC / MS-IT-TOF) can be used as the analytical object.

[0097] [form]

[0098] Those skilled in the art will understand that the embodiments and their variations are specific examples of the following forms.

[0099] (First item) A type of analytical apparatus is an analytical apparatus for analyzing an object waveform as a chromatogram or spectrum, comprising: a processor; and a memory storing a learned model created by machine learning, wherein the machine learning uses multiple sets of multiple partial waveforms created by dividing a reference waveform whose peak position is known, the processor divides the object waveform into multiple partial waveforms, uses the learned model to determine the peak waveform among the multiple divided partial waveforms, and uses the data output from the learned model when determining the peak waveform of the object waveform using the learned model to calculate the confidence level of the determination result of the peak waveform.

[0100] According to the parsing device described in the first item, when performing peak sampling using semantic segmentation technology, it is possible to calculate the confidence level of the peak sampling.

[0101] (Second item) In the parsing apparatus described in the first item, the processor calculates the confidence level using the value determined based on the data output from the learned model, or the data after performing labeling processing to label the data output from the learned model.

[0102] According to the analytical apparatus described in the second item, the confidence level can be appropriately calculated using the value determined from the data output from the learned model, or the data after performing labeling processing to label the data output from the learned model.

[0103] (Third item) In the analytical apparatus described in the first item, the processor labels the peak waveform and calculates the confidence level.

[0104] According to the analytical apparatus described in the third item, the peak waveform can be labeled and the confidence level can be calculated.

[0105] (Fourth item) In the analytical apparatus described in the second or third item, the label includes at least one of a single peak, an unseparated peak, a peak start point, a peak end point, a peak apex, and a baseline.

[0106] According to the analytical apparatus described in the fourth item, at least one of the following labels can be used: a single peak, an unseparated peak, a peak start point, a peak end point, a peak apex, and a baseline.

[0107] (Fifth item) In the parsing device described in the first item, the processor calculates the average value of the weight corresponding to the peak start point of the object waveform and the weight corresponding to the peak end point of the object waveform as the confidence level.

[0108] According to the analytical apparatus described in item 5, the degree of certainty can be calculated using a relatively simple formula by averaging the weight values ​​corresponding to the start point of the peak of the object waveform and the weight values ​​corresponding to the end point of the peak of the object waveform.

[0109] (Sixth item) The analytical apparatus described in any one of the first to fifth items further includes an output port, which outputs a display signal for displaying the determination result and the degree of certainty.

[0110] According to the analysis device described in item six, by inputting the display signal into the display device, the user can identify the relationship between the judgment result and the degree of certainty.

[0111] (Seventh item) The analysis device described in the sixth item also includes a display device, which displays the determination result and confidence level based on the display signal, and the processor accepts the operation of correcting the determination result when displaying the determination result and confidence level on the display device.

[0112] According to the analysis device described in item seven, the user can consider the degree of certainty and revise the judgment result to one that is considered more appropriate.

[0113] (Item 8) Other forms of analytical methods are analytical methods for analyzing object waveforms as chromatograms or spectra, including: the step of creating a learning model, wherein the learning model determines the peak portion contained in the input waveform through machine learning, wherein the machine learning uses multiple sets of multiple partial waveforms created by segmenting a reference waveform whose peak position is known; the step of segmenting the object waveform into multiple partial waveforms; the step of using the learning model to determine the peak waveform that is the peak portion among the segmented multiple partial waveforms; and the step of using the data output from the learning model when determining the peak portion of the object waveform using the learning model to calculate the confidence level of the determination result of the peak waveform.

[0114] According to the parsing method described in item 8, when performing peak sampling using semantic segmentation technology, the confidence level of peak sampling can be calculated.

[0115] In addition, the processor can also calculate the confidence level by calculating (second sum + third sum) / (first sum + second sum + third sum) when the sum of the weights of the parts belonging to the peak region in the baseline estimation results is set as the first sum, the sum of the weights of the parts belonging to the peak region in the estimation results of individual peaks is set as the second sum, and the sum of the weights of the parts belonging to the peak region in the estimation results of unseparated peaks is set as the third sum (variant example 6).

[0116] Furthermore, the processor can perform labeling processing on the data output from the learned model. When the total number of labels corresponding to any of the unseparated peaks, peak start points, and peak end points in the labels belonging to the peak region is set as the first total number, and the total number of labels corresponding to the baseline in the labels belonging to the peak region is set as the second total number, the processor can calculate the confidence level by calculating (first total number) / (first total number + second total number) (Variation Example 7).

[0117] Embodiments of the present invention have been described, but it should be considered that the embodiments disclosed herein are illustrative in all respects and not limiting. The scope of the invention is defined by the claims, and is intended to include all modifications within the meaning and scope equivalent to the claims.

Claims

1. A resolving apparatus for resolving object waveforms as chromatograms or spectra, characterized in that, And includes: Processor; and The memory stores learned models created through machine learning, which uses multiple sets of partial waveforms created by segmenting a reference waveform whose peak positions are known. The processor The waveform of the object is divided into multiple partial waveforms along the time axis. Using the learned model, the peak waveform of the segmented waveform portion is determined. The confidence level of the peak waveform determination result is calculated using the data output from the learning model when determining the peak portion of the object waveform using the learning model.

2. The analytical apparatus according to claim 1, wherein The processor calculates the confidence level using values ​​determined from data output from the learned model, or data after performing labeling processing to label the data output from the learned model.

3. The analytical apparatus according to claim 1 or 2, wherein The processor labels the peak waveform and calculates the confidence level.

4. The analytical apparatus according to claim 2, wherein The label includes at least one of the following: a single peak, an unseparated peak, a peak start point, a peak end point, a peak apex, and a baseline.

5. The analytical apparatus according to claim 1, wherein The processor calculates the average of the weight value corresponding to the peak start point of the object waveform and the weight value corresponding to the peak end point of the object waveform as the confidence level.

6. The analytical apparatus according to claim 1 or 2, wherein, Also includes: The output port outputs a display signal used to show the determination result and the degree of confidence.

7. The analytical apparatus according to claim 6, wherein, Also includes: The display device displays the determination result and the degree of certainty based on the display signal. When the processor displays the determination result and the confidence level on the display device, it accepts an operation to correct the determination result.

8. A method for analyzing object waveforms as chromatograms or spectra, characterized in that, And includes: The steps for creating a learning model are as follows: the learning model determines the peak portion of the input waveform through machine learning, and the machine learning uses multiple sets of groups including multiple partial waveforms created by segmenting a reference waveform whose peak position is known. The step of dividing the object waveform into multiple partial waveforms along the time axis; The step of using the learned model to determine the peak waveform of the peak portion among the multiple segmented waveforms; and The step of calculating the confidence level of the determination result of the peak waveform using the data output from the learning model when determining the peak portion of the object waveform using the learning model.