Learning data production method and learning data production device
Patent Information
- Application Number
- CN202180094903.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-19
- Filing Date
- 2021-10-05
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2041-10-05
AI Technical Summary
[0003]但是,在食品试样中含有多种多样的物质,仅对指标物质的浓度进行解析的话有时无法充分地捕捉品质的变化
[0031]例如在参照试样为食品试样的情况下上述属性是指新鲜程度或产地,例如在参照试样为生物体试样的情况下上述属性是指特定疾病的有无。
Smart Images

Figure CN117083523B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a technique for creating learning data based on measurement data obtained by measuring a sample using an analytical device such as a chromatography-mass spectrometry system. Background Technology
[0002] High-level quality management is required throughout the production, processing, and distribution of food, and this requirement has been further strengthened in recent years. Previously, the evaluation of quality deterioration at the food manufacturing and processing sites was generally based on subjective assessments of color, odor, and taste. To address this, attempts have been made to use analytical devices for quality evaluation to achieve a more objective and efficient assessment. For example, Patent Document 1 describes a method for evaluating the freshness of fish by quantitatively analyzing the non-volatile biogenic amines produced during the decay of fish.
[0003] However, food samples contain a wide variety of substances, and analyzing only the concentration of indicator substances is sometimes insufficient to fully capture changes in quality. Furthermore, in most cases, even foods of the same type and freshness can have varying levels of substance content, especially when food deterioration is low, which can easily lead to misjudgments or oversights.
[0004] Therefore, the following method was attempted: Measurement data obtained by measuring multiple reference samples (food samples, etc.) with known label information (freshness, etc.) representing the attributes of the object using a specific analytical device was used as training data. A learning model was constructed using machine learning with this training data. In this method, a learning model obtained by performing machine learning on the measurement data of multiple reference samples to converge the parameters of the learning model to a level that can identify the freshness of the target sample with an accuracy above a predetermined benchmark was used as the recognizer. The recognizer thus created identifies the freshness of the target sample by judging the correlation of common feature quantities (e.g., the position or intensity of peaks derived from various substances that assign features to reference samples with the same label information (of the same freshness) in the training data. Therefore, compared to analysis based solely on a single or specific type of indicator substance, it is less prone to misjudgment or omission.
[0005] Existing technical documents
[0006] Patent documents
[0007] Patent Document 1: Japanese Patent Application Publication No. 2017-122677
[0008] Patent Document 2: Japanese Patent Application Publication No. 2020-165847 Summary of the Invention
[0009] The problem the invention aims to solve
[0010] To create a highly accurate freshness identifier for a target sample, a learning model needs to be built using machine learning with multiple training data obtained by accurately extracting the feature quantities contained in the measurement data of a reference sample. For example, when using measurement data obtained by gas chromatography-mass spectrometry (mass chromatography), the training data is created by extracting the feature quantities (retention time, mass-to-charge ratio, and measurement intensity) corresponding to all peaks of each substance contained in the reference sample.
[0011] When extracting peaks from mass chromatography (peak selection), users choose a method from a range of pre-prepared methods (combined with peak extraction algorithms and parameters such as peak height-related thresholds) that best suits the shape of the chromatographic baseline. Even reference samples with identical label information may contain varying amounts of substances. Furthermore, sometimes multiple reference samples are measured repeatedly or using multiple instruments. In such cases, even if the amount of substance in each reference sample is the same, the shape of the baseline and the peak height of the measured data will differ slightly for each sample. Therefore, even using the same method and parameters for peak selection, it may sometimes be impossible to extract small peaks contained in the measured data of some reference samples, resulting in learning data lacking the characteristic quantities corresponding to the unextracted peaks. Using learning data with missing characteristic quantities makes it difficult for the learning model's parameters to converge. Furthermore, even if the learning model's parameters converge, the recognition accuracy may decrease if such a learning model is used as a recognizer. While it is possible to exclude some of the measurement data that was not extracted from the peaks and create training data based solely on the measurement data from which all peaks were extracted, the amount of training data becomes less than the amount of measurement data that was excluded. Therefore, there is still a possibility that the parameters of the learning model may not converge or the accuracy of the recognizer may decrease.
[0012] The problem to be solved by the present invention is to provide a technology that can produce learning data that suppresses the loss of characteristic quantities when extracting peaks from measurement data obtained by analyzing a reference sample whose properties are known to the sample.
[0013] Solution for solving the problem
[0014] The present invention, developed to address the aforementioned problems, is a method for creating learning data. This learning data is used to create a recognizer for multiple target samples with distinct identification attributes. In this learning data creation method...
[0015] For each of a plurality of reference samples with known properties, measurement data containing peaks of multiple measured intensities for specified parameters are obtained.
[0016] Using a pre-prepared first method, peak information related to the multiple peaks is extracted from each measurement data point of the plurality of reference samples.
[0017] The peak information extracted by the first method is compared with the peak information of a pre-prepared reference sample with the same properties to determine whether any peaks are missing.
[0018] A second method, using a different algorithm and / or parameters for peak extraction compared to the first method, is used to extract peak information related to multiple peaks from the measurement data of a reference sample determined to have missing peaks.
[0019] The peak information extracted by the second method is compared with the peak information of the reference sample to determine whether any peaks are missing.
[0020] Learning data is created by extracting feature quantities corresponding to each of the multiple peaks from measurement data that are determined to have no missing peaks extracted by the first method or the second method.
[0021] In addition, another aspect of the present invention, completed to solve the above-mentioned problems, is a learning data production apparatus for producing learning data for creating an identifier for recognizing multiple target samples with different identification attributes. The learning data production apparatus comprises:
[0022] The measurement data acquisition unit acquires mass chromatographic data for each of multiple reference samples with known properties by using a chromatography-mass spectrometry instrument.
[0023] The reference sample information storage unit stores peak information contained in the mass chromatogram of a reference sample with the same properties as the reference sample.
[0024] The method information storage unit stores information about a first method and a second method, wherein the first method is used to extract peaks from the mass chromatographic data, and the second method uses a different algorithm and / or parameters for peak extraction compared to the first method;
[0025] A first peak extraction unit extracts peak information related to the multiple peaks for each mass chromatographic data in the mass chromatographic data of the multiple reference samples using the first method.
[0026] The first determination unit determines whether there is a missing peak by comparing the peak information extracted by the first peak extraction unit with the peak information of the reference sample with the same properties.
[0027] The second peak extraction unit uses the second method to extract peak information related to the multiple peaks from the mass chromatographic data of the reference sample that was determined by the first determination unit to have peak missing.
[0028] The second determination unit compares the peak information extracted by the second method with the peak information of the reference sample to determine whether any peaks are missing; and
[0029] The learning data production unit obtains characteristic quantities corresponding to the multiple peaks from the mass chromatographic data that are determined by the first determination unit or the second determination unit to have no missing peaks, and produces learning data accordingly.
[0030] The effects of the invention
[0031] For example, when the reference sample is a food sample, the above attribute refers to freshness or place of origin; when the reference sample is a biological sample, the above attribute refers to the presence or absence of a specific disease.
[0032] In the learning data production method of the present invention, firstly, measurement data containing peaks with multiple measurement intensities for specified parameters is acquired for each of a plurality of samples with known properties. Secondly, in the learning data production apparatus of the present invention, mass chromatographic data is acquired for each of a plurality of reference samples with known properties using a chromatography-mass spectrometry (GC-MS) instrument, and used as measurement data. The acquisition of measurement data can be performed either by actually measuring the samples with an analytical apparatus or by reading measurement data obtained in a prior measurement.
[0033] Next, a pre-prepared first method is used to extract the peak intensity of each measurement data point from multiple samples. Then, by comparing the extracted peaks with the peak information of a pre-prepared reference sample with the same properties, it is determined whether all peaks have been extracted (whether any peaks are missing). For the peak information of the reference sample, peak information registered in a library can be used, for example. Alternatively, cumulative measurement data from multiple reference samples with the same properties can be accumulated to create cumulative measurement data, and the peak information extracted from this cumulative measurement data using the first method can be used as the peak information of the reference sample. Both the first and second methods are methods that include a combination of peak extraction algorithms and parameters such as thresholds related to peak height. For the peak extraction algorithm, either conventionally used algorithms or the characteristic algorithms described in the embodiments described later can be used.
[0034] For measurement data of samples identified as having missing peaks, a second method, using a different algorithm or parameters than the first method, is used to extract peaks of measured intensity. Then, the peaks extracted using the second method are compared with peak information from pre-prepared reference samples of the same properties to determine the presence or absence of missing peaks. Finally, from the measurement data where peaks extracted using either the first or second method are determined to be intact, feature quantities corresponding to multiple peaks in that measurement data are obtained to create training data.
[0035] In the learning data production method and apparatus of the present invention, since the learning data is produced by obtaining the characteristic quantity corresponding to the peak of the measurement intensity of the measurement data from the measurement data that is determined to have no missing peaks, it is possible to produce learning data that suppresses the loss of characteristic quantities. Furthermore, even if it is impossible to extract all peaks from the measurement data using the first method, the learning data is produced by obtaining the characteristic quantity from the measurement data from which all peaks can be extracted using the second method, thus suppressing the reduction in the amount of learning data. Attached Figure Description
[0036] Figure 1 This is a structural diagram of the main parts of a sample evaluation system, including an embodiment of the learning data production apparatus involved in this invention.
[0037] Figure 2 This is a flowchart illustrating an embodiment of the learning data production method involved in this invention.
[0038] Figure 3 This is a schematic diagram of the measurement data, i.e., the mass chromatographic data, obtained in this embodiment.
[0039] Figure 4 This is a schematic diagram related to previous peak extraction methods.
[0040] Figure 5 This is an example of the results obtained by extracting peaks from mass chromatography using conventional peak extraction methods.
[0041] Figure 6 This is another example of a result obtained by extracting peaks from mass chromatography using conventional peak extraction methods.
[0042] Figure 7 This is another example of a result obtained by extracting peaks from mass chromatography using conventional peak extraction methods.
[0043] Figure 8 This is a schematic diagram related to the baseline delineation based on the peak extraction method in this embodiment.
[0044] Figure 9 This is an example of defining the baseline of mass chromatography using the peak extraction method of this embodiment.
[0045] Figure 10 This is an example of the result obtained by using the peak extraction method of this embodiment to extract peaks from mass chromatography. Detailed Implementation
[0046] Hereinafter, an embodiment of the learning data production method and apparatus of the present invention will be described with reference to the accompanying drawings.
[0047] The sample evaluation system 1 of this embodiment is used to estimate the properties of a target sample. Specifically, it is used to estimate, for example, the freshness and origin of a target sample as a food product, or to determine whether a subject has a specific disease based on a biological sample.
[0048] exist Figure 1 The diagram shows the main structural components of the sample evaluation system 1, which includes the learning data production apparatus of this embodiment. The sample evaluation system 1 of this embodiment generally consists of a gas chromatography-mass spectrometry (GC-MS) instrument 2 and a control and processing unit 3.
[0049] The control processing unit 3 includes a storage unit 31. The storage unit 31 contains a method information storage unit 311 and a reference sample information storage unit 312. The method information storage unit 311 stores information about methods for extracting peaks from mass chromatographic data obtained by measuring samples using a gas chromatography-mass spectrometry (GC-MS) instrument 2. The method information includes information about peak extraction algorithms and combinations of parameters such as thresholds related to peak height. Peak extraction algorithms can include conventionally used methods such as the tie-point method, the level method, the new baseline method, and the method characteristic of this embodiment using a very small measurement point (described later).
[0050] The reference sample information storage unit 312 stores mass chromatographic data of multiple reference samples with different properties for various samples, as well as peak information (retention time, mass-to-charge ratio, and measurement intensity) related to multiple peaks appearing in the mass chromatogram. In addition, the storage unit 312 also stores the measurement conditions when measuring various samples using the gas chromatograph-mass spectrometer 2.
[0051] The control processing unit 3 also includes a measurement data acquisition unit 32, a reference sample data production unit 33, a first peak extraction unit 34, a second peak extraction unit 35, a judgment unit 36, a measurement data display unit 37, a learning data production unit 38, a learning model construction unit 39, and a recognizer production unit 40 as functional blocks. The control processing unit 3 is essentially a personal computer or a higher-performance computer called a workstation. It implements the aforementioned functional blocks by executing a sample evaluation system program pre-installed on the computer using the computer's processor. Furthermore, the control processing unit 3 is connected to an input unit 4 such as a keyboard and mouse, and a display unit 5 such as an LCD screen.
[0052] The method information storage unit 311, the reference sample information storage unit 312, the measurement data acquisition unit 32, the reference sample data production unit 33, the first peak extraction unit 34, the second peak extraction unit 35, the judgment unit 36, the measurement data display unit 36, and the learning data production unit 38, which are components of the control processing unit 3, constitute the learning data production apparatus 10 of this embodiment. In this embodiment, the learning data production apparatus 10 is assembled as part of the control processing unit 3, but it is also possible to configure the learning data production apparatus 10 as an independent device from the control processing unit 3.
[0053] Next, refer to Figure 2 The flowchart below illustrates the operation of the sample evaluation system 1 in this embodiment. In this example, an identifier is created to estimate the properties of the target sample.
[0054] Before creating the identifier, the user pre-sets a reference sample in the autosampler (illustration omitted) connected to the gas chromatography-mass spectrometry system 2. The reference sample is a sample of the same type as the target sample, and its properties are known. In addition, multiple reference samples are pre-set for each property.
[0055] When the user instructs the start of identifier fabrication, the measurement data acquisition unit 32 introduces the reference samples placed in the autosampler into the gas chromatography-mass spectrometry (GC-MS) instrument 2 in a predetermined order. In the GC-MS instrument 2, after the reference samples are separated into each component within the gas chromatograph column, they are introduced into the mass spectrometer, where each component is ionized using an electron ionization source or a plasma ionization source. After mass separation, they are detected using an ion detector. The output signal from the ion detector is sequentially sent to the control processing unit 3, and the measurement data (mass chromatographic data) for each reference sample is stored in the storage unit 31. Thus, mass chromatographic data is acquired for all reference samples (step 1). For example, in... Figure 3 As schematically illustrated, mass chromatographic data represent the measured intensity of ions for two parameters: time and mass-to-charge ratio. The peak with the mass-to-charge ratio of the ions generated by the component is the time (retention time centered at t1) during which the component contained in the sample elutes from the column of the gas chromatograph.
[0056] When all reference samples have been measured and the mass chromatographic data have been saved, the reference sample data generation unit 33 groups the reference samples according to each attribute. Then, the mass chromatographic data of reference samples with the same attribute are accumulated to generate cumulative mass chromatographic data (step 2). The reference sample data generation unit 33 then activates the first peak extraction unit 34 to generate an extraction ion current chromatogram (chromatogram of the intensity of ions with a specific mass-to-charge ratio) based on the cumulative mass chromatographic data, and extracts the peak using the algorithm and parameters stored in the method information storage unit 311 (step 3).
[0057] Here, the process of extracting peaks from the extraction ion current chromatography by the first peak extraction unit 34 will be described. Conventional methods for extracting peaks from extraction ion current chromatography, specifically total ion current chromatography representing the intensity of total ions, include methods such as the junction point method and the level method. Figure 4 ).exist Figure 4 In this method, shading lines are added to the peaks extracted by various methods. In the tie-point method, the point where the slope of the chromatographic waveform exceeds a predetermined value is designated as the peak start point S, and the point where the slope of the waveform is lower than a predetermined value is designated as the peak end point E. Then, the peak start point S and the peak end point E are connected to define the baseline and the peak is extracted. In the horizontal method, after defining the peak start point S and the peak end point E in the same way as above, a horizontal line is drawn through the point with the lower intensity of the two points. The intersection of this horizontal line and the vertical line drawn from another point is defined as the baseline, and the peak is extracted. Furthermore, when separating peaks with multiple superimposed peaks, for example, the new baseline method is used. In the new baseline method, the peak start point S and the peak end point E are defined in the same way as above, and the minimum point located between two peaks is designated as the peak separation point M to separate the peaks.
[0058] For measurement data with minimal baseline variation, the accuracy of peak extraction does not differ significantly regardless of whether the tie-point method or the level method is used. However, in the case of measurement data obtained using gas chromatography, due to a phenomenon known as column bleed, the baseline may sometimes increase over time in the latter part of the chromatogram. Previously, pre-set algorithms such as the tie-point method, level method, and new baseline method were used to automatically extract peaks from chromatographic data; however, if the algorithm or parameters are not suitable, peak extraction accuracy may occur. Figure 5 As shown, the baseline rise caused by the drift in the latter half of the chromatography is extracted as a peak.
[0059] To address such drift, it is sometimes possible to set parameters that take into account the baseline rise caused by the drift. Figure 6 The graph shows the parameter set to a value like 50, but even with this value, the rise in the baseline caused by drift is still extracted as a peak. On the other hand, Figure 7 The image shows the parameter set to a value of 100. Figure 7 In the first half of the chromatography, the drift was not extracted as a peak, which is appropriate in this respect. However, in the first half of the chromatography, the baseline slope was too steep, making it difficult to accurately obtain the peak height and area.
[0060] Therefore, in this embodiment, the minimum points of the chromatographic waveform are connected to define the baseline and extract the peaks. For example, in Figure 8 As schematically shown above, a chromatogram is composed of multiple measurement points connected together. In the algorithm of this embodiment, measurement points whose measurement intensity is smaller than that of any of the two adjacent measurement points are extracted as minimum measurement points. Then, by performing linear interpolation on the extracted minimum measurement points, a range is defined as follows: Figure 8 The baseline is as shown in the lower part. Furthermore, the method for defining the baseline based on the minimum measurement points is not limited to linear interpolation; it can also be an approximate curve connecting these minimum measurement points, etc. Then, the portion of the baseline whose height exceeds a predetermined threshold is extracted as a peak.
[0061] In extractive ion current chromatography, peaks will not overlap as long as the retention time and mass-to-charge ratio are not shared. Therefore, a baseline can be determined by linear interpolation. However, in chromatography with only one parameter for intensity measurement (such as total ion current chromatography), overlapping peaks may sometimes occur. In this case, a minimum point may appear between the two peaks. If linear interpolation is performed on the minimum measurement point, including the minimum point between peaks, in such cases, the baseline cannot be accurately determined. In such cases, it is preferable to obtain an approximate curve for all minimum measurement points, for example, instead of performing linear interpolation on the minimum measurement point, thereby reducing the influence of outliers such as minimum points within overlapping peaks and accurately determining the baseline. Alternatively, the analyst can remove outliers such as minimum points between peaks and perform linear interpolation on the remaining minimum measurement points to determine the baseline.
[0062] Figure 9 It shows that it is aimed at and Figures 5-7 The graph shown is a result of baseline determination of mass chromatographic data, which also exhibits large drift in the latter half. Figure 10 This graph shows the results obtained by extracting peaks (circles in the figure indicating peak apexes) that are higher than a predetermined threshold relative to the baseline. The results demonstrate that even in cases of large drift in the latter part of the chromatography, peaks can be accurately extracted using the method of this embodiment.
[0063] Even if a peak that is difficult to extract is present in one of the mass chromatographic data of multiple reference samples with the same properties, it can be supplemented by the measured intensity of other reference samples when they are accumulated, thus making peak extraction easier. Therefore, if accumulated mass chromatographic data is used, all peaks can be easily extracted. Therefore, in this embodiment, the position information of multiple peaks extracted from the accumulated mass chromatographic data is used as the peak information of the reference sample. The reference sample data generation unit 33 stores the peak information of the reference sample in the reference sample information storage unit 312 according to each attribute.
[0064] Next, the first peak extraction unit 34 extracts peaks from the mass chromatographic data of each reference sample using the algorithm and parameters stored in the method information storage unit 311, just as described above (step 5).
[0065] When a peak is extracted by the first peak extraction unit 34, the determination unit 36 reads the peak information of a reference sample with the same properties as the reference sample from the reference sample information storage unit 312. Then, the peak information of the reference sample is compared with the peak information of the reference sample (step 6).
[0066] When the determination unit 36 determines that the peak information of the reference sample is consistent with the peak information of the reference sample (i.e., there is no missing peak) ("No" in step 7), the characteristic quantities corresponding to the peaks extracted from the mass chromatographic data of the reference sample are acquired to create learning data (step 11). The extracted characteristic quantities include a combination of retention time and mass-to-charge ratio corresponding to the peak apex of each peak. In addition, the characteristic quantities may also include the height or area value of the peak. That is, data equivalent to a peak list is created as learning data.
[0067] When the determination unit 36 determines that the peak information of the reference sample is inconsistent with that of the reference sample (some of the multiple peaks that should be extracted were not extracted) ("Yes" in step 7), a screen is displayed to notify the reference sample of the missing peak. Additionally, the measurement data display unit 37 generates an extraction ion current chromatography (EIC) corresponding to the mass-to-charge ratio of the missing peak based on the mass chromatographic data of the reference sample, and displays it on the screen of the display unit 5. In the mass chromatogram displayed on the screen, marks are superimposed at the positions where peaks that should have been extracted but were not extracted by the first peak extraction unit 34 (i.e., the positions of peaks whose retention time and mass-to-charge ratio are included in the peak information of the reference sample but not in the peak information of the reference sample). The user can check the waveform of the EIC at the marked positions and determine whether the second peak extraction unit 35 can extract the peak by changing the algorithm or parameters.
[0068] Furthermore, the determination unit 36 determines whether the number of times the peak extraction process was performed on the reference sample has reached a predetermined number (specified number). This number is set to, for example, 5 times. At this stage, only the first peak extraction unit 34 has extracted the peak (the peak extraction process is performed once), so it is determined that the predetermined number has not been reached ("No" in step 8).
[0069] When the result is "No" in step 8, the second peak extraction unit 35 changes the algorithm and / or parameters used in the previous peak extraction (step 9). In this example, the algorithm is not changed, but the threshold (baseline height) that is identified as a peak is lowered. Specifically, for example, the threshold (parameter) used in peak extraction is changed to 90% of the threshold used when the first peak extraction unit 34 extracts the peak (the threshold is lowered by 10%).
[0070] Next, the second peak extraction unit 35 uses the modified threshold to extract peaks again from the mass chromatographic data of the reference sample (step 5). Then, it compares the peak information with that of the reference sample again (step 6). If it is determined that there are no missing peaks ("No" in step 7), it acquires the characteristic quantities corresponding to the extracted peaks to create learning data (step 11). When creating learning data, it is determined whether the processing of the mass chromatographic data of all reference samples is complete (step 12). Then, if there is unprocessed data, the processing after step 5 is performed on the mass chromatographic data of the next reference sample according to the same procedure as described above.
[0071] On the other hand, when it is determined again that a peak is missing ("Yes" in step 7), it is determined whether the number of peak extraction processes has reached the predetermined number (step 8). If the predetermined number has not been reached, the algorithm and / or parameters are changed again (step 9). In this example, the threshold for peak extraction processing performed by the first peak extraction unit 34 is set to 100, and the threshold is decreased by 10 each time. Alternatively, the threshold can be decreased by 10% based on the previous peak extraction processing.
[0072] If the missing peaks are not eliminated even after a specified number of peak extractions ("Yes" in step 8), the processing related to the mass chromatographic data of the reference sample ends, and the processing after step 5 is performed in the same manner as described above for the mass chromatographic data of the next reference sample.
[0073] When the processing of the mass chromatographic data of all reference samples is completed (Yes in step 12), the learning model construction unit 39 constructs a learning model by performing machine learning using the learning data prepared in the above process (step 13). As a machine learning method, methods such as supervised learning can be used. Specifically, in addition to representative machine learning methods such as support vector machines, neural networks, and random forests, multivariate analytical methods such as logistic regression, orthogonal least squares, and k-nearest neighbors can also be used. When the learning model is constructed using machine learning, the recognizer creation unit 40 uses the learning model to create a recognizer and saves it in the storage unit 31 (step 14).
[0074] Previously, when creating training data, if peaks extracted from the mass chromatographic data of a reference sample with known properties were missing, the training data was created by excluding those missing peaks and using only mass chromatographic data from which all peaks were extracted. As a result, the number of training data points decreased compared to the number of excluded mass chromatographic data points, potentially leading to difficulties in parameter convergence of the learning model or a reduction in the accuracy of the recognizer. Alternatively, in Patent Document 2, information on peaks that did not appear in all training data points (those missing in one training data point) was deleted to create the training data. In this case, training data was created by removing peak information that might be useful for determining peak properties, thus still raising the possibility of difficulty in parameter convergence of the learning model or a reduction in the accuracy of the recognizer.
[0075] In this embodiment, even if peaks are missing during peak extraction by the first peak extraction unit 34, the second peak extraction unit 35 attempts to extract peaks a predetermined number of times while changing the algorithm or parameters. Therefore, the reduction in the amount of training data can be suppressed. Furthermore, since the training data is created based on confirmation that all peaks have been extracted by comparing with the peak information of a reference sample, training data without missing feature quantities can be created.
[0076] The above embodiment is an example and can be appropriately modified in accordance with the spirit of the present invention. The values listed in the above embodiment are always just examples and can be appropriately modified according to the characteristics of the target sample, the measurement data, etc.
[0077] Alternatively, in the above embodiment, the user can confirm the mass chromatogram that is determined to have missing peaks on the display unit 5 screen, and only if the peak can be identified at the position of the missing peak, the user instructs the second peak extraction unit 35 to extract the peak. If configured in this way, time can be saved from processing mass chromatogram data where peak extraction is impossible (or extremely difficult), and learning data can be produced efficiently.
[0078] In the above embodiment, the structure is set as follows: only an algorithm that defines the baseline and extracts the peak based on the minimum measurement point is used. When extracting the peak by the second peak extraction unit 35, only the threshold is changed. However, the algorithm for extracting the peak can also be changed when extracting the peak by the second peak extraction unit 35.
[0079] In the above embodiments, the mass chromatographic data obtained by measuring the reference sample using a gas chromatograph-mass spectrometer 2 was processed. However, the same structure as in this embodiment can be applied to various measurement data containing peaks with multiple measurement intensities for specified parameters. Furthermore, the peaks mentioned here can also include downward peaks that appear in the absorption spectrum, etc.
[0080] [Way]
[0081] Those skilled in the art will understand that the above-described exemplary embodiments are specific examples of the following approaches.
[0082] (First item)
[0083] One aspect of the present invention is a method for creating learning data, used to create learning data for constructing a recognizer that identifies multiple target samples with distinct identification attributes. In this method,
[0084] For each of a plurality of reference samples with known properties, measurement data containing peaks of multiple measured intensities for specified parameters are obtained.
[0085] Using a pre-prepared first method, peak information related to the multiple peaks is extracted for each measurement data point of the plurality of reference samples.
[0086] The peak information extracted by the first method is compared with the peak information of a pre-prepared reference sample with the same properties to determine whether any peaks are missing.
[0087] A second method, using a different algorithm and / or parameters for peak extraction compared to the first method, is used to extract peak information related to multiple peaks from the measurement data of a reference sample determined to have missing peaks.
[0088] The peak information extracted by the second method is compared with the peak information of the reference sample to determine whether any peaks are missing.
[0089] Learning data is created by extracting feature quantities corresponding to each of the multiple peaks from measurement data that are determined to have no missing peaks extracted by the first method or the second method.
[0090] (Item 8)
[0091] Another aspect of the present invention is a learning data production apparatus for producing a learning model used in machine learning, wherein the machine learning is used to produce a learning model constituting a recognizer of multiple target samples with different recognition attributes, the learning data production apparatus comprising:
[0092] The measurement data acquisition unit acquires mass chromatographic data for each of multiple reference samples with known properties by using a chromatography-mass spectrometry instrument.
[0093] The reference sample information storage unit stores peak information contained in the mass chromatogram of a reference sample with the same properties as the reference sample.
[0094] The method information storage unit stores information about a first method and a second method, wherein the first method is used to extract peaks from the mass chromatographic data, and the second method uses a different algorithm and / or parameters for peak extraction compared to the first method;
[0095] A first peak extraction unit extracts peak information related to the multiple peaks for each mass chromatographic data in the mass chromatographic data of the multiple reference samples using the first method.
[0096] The first determination unit determines whether there is a missing peak by comparing the peak information extracted by the first peak extraction unit with the peak information of the reference sample with the same properties.
[0097] The second peak extraction unit uses the second method to extract peak information related to the multiple peaks from the mass chromatographic data of the reference sample that was determined by the first determination unit to have missing peaks.
[0098] The second determination unit compares the peak information extracted by the second method with the peak information of the reference sample to determine whether any peaks are missing; and
[0099] The learning data production unit obtains characteristic quantities corresponding to the multiple peaks from the mass chromatographic data that are determined by the first determination unit or the second determination unit to have no missing peaks, and produces learning data accordingly.
[0100] In the first method for preparing learning data, firstly, measurement data containing peaks with multiple measured intensities for specified parameters is acquired for each of multiple reference samples with known properties. Secondly, in the eighth apparatus for preparing learning data, mass chromatographic data is acquired for each of the multiple reference samples with known properties using a chromatography-mass spectrometry (GC-MS) instrument, and used as measurement data. The acquisition of measurement data can be performed either by actually measuring the sample with the analytical apparatus or by reading measurement data obtained in a prior measurement.
[0101] Next, for each measurement data point from multiple samples, a pre-prepared first method is used to extract the peak of the measured intensity. Then, by comparing the extracted peaks with the peak information of a pre-prepared reference sample with the same properties, it is determined whether all peaks have been extracted (whether any peaks are missing). For the peak information of the reference sample, peak information registered in a library can be used, for example.
[0102] For measurement data of samples identified as having missing peaks, a second method, using a different algorithm or parameters than the first method, is used to extract peaks of measured intensity. Then, the peaks extracted using the second method are compared with peak information from pre-prepared reference samples of the same properties to determine the presence or absence of missing peaks. Finally, from the measurement data where peaks extracted using either the first or second method are determined to be intact, feature quantities corresponding to multiple peaks in that measurement data are obtained to create training data.
[0103] In the learning data production method of the first item and the learning data production apparatus of the eighth item, since the learning data is produced by obtaining the characteristic quantity corresponding to the peak of the measurement intensity of the measurement data from the measurement data that is determined to have no missing peaks, it is possible to produce learning data that suppresses the loss of characteristic quantities. Furthermore, even if it is impossible to extract all peaks from the measurement data using the first method, the characteristic quantity is obtained from the measurement data using the second method to extract all peaks to produce the learning data, thus suppressing the reduction in the amount of learning data.
[0104] (Second item)
[0105] In the method for creating learning data described in the first item, wherein,
[0106] Measurement data that are determined to have missing peaks in the peak information extracted by the first method or the second method will be displayed on the screen.
[0107] (Third item)
[0108] In the second method for creating learning data, wherein,
[0109] For the measurement data specified by the user in the measurement data that has been displayed on the screen, peak information is extracted using the second method.
[0110] In the second method for creating learning data, the user can confirm the existence of a peak with extractable intensity by checking the measurement data identified as having missing peaks on the screen. Furthermore, in the third method for creating learning data, the user only specifies the measurement data containing peaks with extractable intensity, thereby reducing the processing load involved in extracting peak information using the second method.
[0111] (Item 4)
[0112] According to any one of the first to third items, the method for creating learning data, wherein,
[0113] The measurement data are obtained for each of multiple reference samples with the same properties.
[0114] Cumulative measurement data is generated by accumulating the measurement data of multiple reference samples with the same properties.
[0115] The peak information of the reference sample is obtained by extracting peak information related to multiple peaks contained in the cumulative measurement data using the first method.
[0116] In the fourth section on data production methods, peak information from reference samples can be used for samples that are not included in databases, such as samples that have not been measured in sufficient quantities in the past or unknown samples.
[0117] (Item 5)
[0118] According to any one of the first to fourth items, the method for creating learning data, wherein,
[0119] The process of extracting peak information related to multiple peaks using algorithms and / or methods with different parameters and determining whether any peak is missing will be repeated a specified number of times until it is determined that no peak is missing.
[0120] In the fifth method for creating learning data, since the automatic creation of learning data is repeated a predetermined number of times, users do not need to determine whether peak information needs to be collected each time, thus reducing the user's burden.
[0121] (Item 6)
[0122] According to any one of the methods for creating learning data described in items one through five, wherein,
[0123] The first method and / or the second method extracts, from among the multiple measurement points constituting the measurement data, the measurement point whose measurement intensity is lower than either of the measurement intensities of the adjacent measurement points on the left and right sides, as the minimum measurement point.
[0124] The baseline of the plurality of measurement points is determined using the minimum measurement point.
[0125] For each of the plurality of measurement points, a peak is extracted based on the fact that the value obtained by subtracting the baseline from the measurement intensity at that measurement point exceeds a predetermined threshold.
[0126] (Seventh item)
[0127] According to the method for creating learning data described in item six, wherein,
[0128] Both the first method and the second method use the minimum measurement point to determine the baseline and extract the peak, but the threshold values are different in the first method and the second method.
[0129] In mass chromatographic data obtained by gas chromatography-mass spectrometry analysis, the baseline size also changes over time. In the learning data preparation method of the sixth item, since a very small measurement point is used to determine the baseline, it can be well used to extract peaks from measurement data where both the measurement intensity and the baseline change with respect to the parameter. Furthermore, when extracting peaks from such measurement data, as in the learning data preparation method described in the seventh item, a method that only changes the threshold used to determine whether something is a peak compared to the first method can be used as a second method.
[0130] Explanation of reference numerals in the attached figures
[0131] 1: Sample evaluation system; 10: Data production device for learning; 2: Gas chromatography-mass spectrometry; 3: Control and processing unit; 31: Storage unit; 311: Method information storage unit; 312: Reference sample information storage unit; 32: Measurement data acquisition unit; 33: Reference sample data production unit; 34: First peak extraction unit; 35: Second peak extraction unit; 36: Judgment unit; 37: Measurement data display unit; 38: Data production unit for learning; 39: Learning model construction unit; 39: Recognizer production unit; 4: Input unit; 5: Display unit.
Claims
1. A method for creating learning data, used to create learning data for creating a recognizer that identifies multiple target samples with different identification attributes, wherein in the method for creating learning data, For each of a plurality of reference samples with known properties, measurement data containing peaks of multiple measured intensities for specified parameters are obtained. Using a pre-prepared first method, peak information related to the multiple peaks is extracted for each measurement data point of the plurality of reference samples. The peak information extracted by the first method is compared with the peak information of a pre-prepared reference sample with the same properties to determine whether any peaks are missing. A second method, using a different algorithm and / or parameters for peak extraction compared to the first method, is used to extract peak information related to multiple peaks from the measurement data of a reference sample determined to have missing peaks. The peak information extracted by the second method is compared with the peak information of the reference sample to determine whether any peaks are missing. Learning data is created by extracting feature quantities corresponding to each of the multiple peaks from measurement data that are determined to have no missing peaks extracted by the first method or the second method.
2. The method for generating learning data according to claim 1, wherein, Measurement data that are determined to have missing peaks in the peak information extracted by the first method or the second method will be displayed on the screen.
3. The method for creating learning data according to claim 2, wherein, For the measurement data specified by the user in the measurement data that has been displayed on the screen, peak information is extracted using the second method.
4. The method for generating learning data according to claim 1, wherein, The measurement data are obtained for each of multiple reference samples with the same properties. Cumulative measurement data is generated by accumulating the measurement data of multiple reference samples with the same properties. The peak information of the reference sample is obtained by using the first method to extract peak information related to multiple peaks contained in the cumulative measurement data.
5. The method for generating learning data according to claim 1, wherein, The process of extracting peak information related to multiple peaks using algorithms and / or methods with different parameters and determining whether any peak is missing will be repeated a specified number of times until it is determined that no peak is missing.
6. The method for generating learning data according to claim 1, wherein, The first method and / or the second method extracts, from among the multiple measurement points constituting the measurement data, the measurement point whose measurement intensity is lower than either of the measurement intensities of the adjacent measurement points on the left and right sides, as the minimum measurement point. The baseline of the plurality of measurement points is determined using the minimum measurement point. For each of the plurality of measurement points, a peak is extracted based on the fact that the value obtained by subtracting the baseline from the measurement intensity at that measurement point exceeds a predetermined threshold.
7. The method for generating learning data according to claim 6, wherein, Both the first method and the second method use the minimum measurement point to determine the baseline and extract the peak, but the threshold values are different in the first method and the second method.
8. A learning data production apparatus for producing learning data, the learning data being used to produce an identifier for recognizing multiple target samples with different identification attributes, the learning data production apparatus comprising: The measurement data acquisition unit acquires mass chromatographic data for each of multiple reference samples with known properties by using a chromatography-mass spectrometry instrument. The reference sample information storage unit stores peak information contained in the mass chromatogram of a reference sample with the same properties as the reference sample. The method information storage unit stores information about the first method and the second method, wherein... The first method is used to extract peaks from the mass chromatographic data, and the second method uses a different algorithm and / or parameters for peak extraction compared to the first method; A first peak extraction unit extracts peak information related to the multiple peaks for each mass chromatographic data in the mass chromatographic data of the multiple reference samples using the first method. The first determination unit determines whether there is a missing peak by comparing the peak information extracted by the first peak extraction unit with the peak information of the reference sample with the same properties. The second peak extraction unit uses the second method to extract peak information related to the multiple peaks from the mass chromatographic data of the reference sample that was determined by the first determination unit to have missing peaks. The second determination unit compares the peak information extracted by the second method with the peak information of the reference sample to determine whether any peaks are missing; and The learning data production unit obtains characteristic quantities corresponding to the multiple peaks from the mass chromatographic data that are determined by the first determination unit or the second determination unit to have no missing peaks, and produces learning data accordingly.
Citation Information
Patent Citations
Ion sensitive film, ion selective electrode, and ion sensor
JP2017122677A
Food quality determination method and food quality determination device
JP2020165847A
Peak detection method and data processing device
CN109073613A
Method for determining food-product quality and food-product quality determination device
US20200309746A1