Method and program for processing chromatograph mass spectrometry data

A computer-executed method for processing chromatographic mass spectrometry data using machine learning addresses the labor-intensive nature of feature identification by predicting and ranking feature items that affect physical properties, enhancing data analysis efficiency and accuracy.

JP2026013046APending Publication Date: 2026-01-28SHIMADZU SEISAKUSHO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024113196
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2026-01-28

AI Technical Summary

Technical Problem

Existing methods for identifying feature quantities in chromatographic mass spectrometry data are labor-intensive and lack assistance for operators when dealing with complex relationships between analysis results and physical properties.

Method used

A computer-executed method for processing chromatographic mass spectrometry data using machine learning to extract and predict feature items that affect specific physical properties, involving data preprocessing, alignment, feature selection, and model training to identify the importance of each feature item.

Benefits of technology

Assists operators in efficiently identifying feature quantities that affect specific physical properties, reducing manual labor and improving the accuracy of data analysis by using machine learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026013046000001_ABST
    Figure 2026013046000001_ABST
Patent Text Reader

Abstract

To provide a technique for assisting a worker who searches for a feature amount that affects a specific physical property by using chromatograph mass spectrometry data.SOLUTION: Values of a plurality of feature items are extracted from chromatograph mass spectrometry data of a sample. Each of the plurality of feature items corresponds to a combination of a retention time range and a mass-to-charge ratio. A machine learning model for predicting the material property information from values of at least some of the feature items is generated. For each of one or more feature items of the at least some feature items, an importance level in the machine learning model is identified and output.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for processing data, and in particular to a method for processing chromatographic mass spectrometry data. [Background technology]

[0002] In order to identify factors that affect specific physical properties of a material, feature quantities have been searched for in the analysis results of the material. In one example of feature quantity search, an operator identifies peaks in the chromatographic mass spectrometry data of each of multiple samples of the material, compiles the identified peaks together with the values ​​of the specific physical properties in a table, and then extracts peaks that affect the specific physical properties as markers using statistical processing and / or machine learning.

[0003] Chromatography mass spectrometry data contains unresolved or weak peaks and complex peaks. Therefore, the above-mentioned methods involving peak identification require an enormous amount of labor from the operator. In light of this background, various studies have been conducted on automating the extraction of markers from chromatography mass spectrometry data (for example, S. Abou-el-karam a, J. Ratel a, N. Kondjoyan a, C. Truan a, b, E. Engel, "Marker discovery in volatolomics based on systematic alignment of GC-MS signals: Application to food authentication," Analytica Chimica Acta, Elsevier, Netherlands, October 23, 2017, pp. 58-67 (see Non-Patent Document 1)). [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] S. Abou-el-karam a, J. Ratel a, N. Kondjoyan a, C. Truan a, b, E. Engel, "Marker discovery in volatolomics based on systematic alignment of GC-MS signals: Application to food authentication", Analytica Chimica Acta, Elsevier, Netherlands, October 23, 2017, pp. 58-67. Summary of the Invention [Problem to be solved by the invention]

[0005] On the other hand, there are cases where the operator still wants to identify the feature quantities himself, such as when a complex relationship is estimated between the analysis results and the physical properties. However, no technology has been considered to assist the operator in identifying the feature quantities in such cases.

[0006] The present invention has been devised in view of the above-described circumstances, and its purpose is to provide a technique for assisting an operator who uses chromatographic mass spectrometry data to search for features that affect specific physical properties. [Means for solving the problem]

[0007] A method for processing chromatographic mass spectrometry data according to an aspect of the present disclosure is a computer-executed method for processing chromatographic mass spectrometry data, comprising: a step of extracting values ​​of a plurality of feature items from the chromatographic mass spectrometry data of a sample, each of the plurality of feature items corresponding to a combination of a retention time range and a mass-to-charge ratio, the plurality of combinations corresponding to each of the plurality of feature items being different from one another; the chromatographic mass spectrometry data being associated with physical property information representing a given physical property; a step of generating training data for the plurality of samples, the training data including at least some of the plurality of feature items and physical property information for the chromatographic mass spectrometry data of each of the samples; a step of generating a machine learning model using the training data to predict the physical property information from the values ​​of at least some of the feature items; a step of identifying the importance of each of one or more of the at least some feature items in the machine learning model; and a step of outputting the importance of each of the one or more feature items.

[0008] A program according to an aspect of the present disclosure, when executed by one or more processors of a computer, causes the computer to perform the above-described method for processing chromatographic mass spectrometry data. [Effects of the Invention]

[0009] According to one aspect of the present disclosure, a technique is provided for assisting an operator who uses chromatography mass spectrometry data to search for feature quantities that affect specific physical properties. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating an overall configuration of an analysis system. [Figure 2] FIG. 10 is a diagram showing a specific example of chromatographic mass spectrometry data. [Figure 3] FIG. 2 is a diagram showing the flow of data processing in the present embodiment. [Figure 4]FIG. 1 shows peak information for two chromatograms. [Figure 5] FIG. 10 is a diagram for explaining the possibility of correspondence between a reference peak and a target peak. [Figure 6] FIG. 10 illustrates an example of exception matching. [Figure 7] FIG. 10 is a diagram showing specific examples of a plurality of feature items. [Figure 8] FIG. 10 is a diagram illustrating an example of a screen displaying importance levels. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the drawings, the same or corresponding parts are designated by the same reference numerals, and description thereof will not be repeated.

[0012] [Hardware configuration] In this embodiment, the computer that performs the processing method is configured to be able to communicate with a gas chromatograph mass spectrometer (hereinafter also referred to as "GC / MS") in the analysis system. Note that the computer only needs to have the function of processing chromatograph mass analysis data, and does not necessarily have to be configured to communicate with the GC / MS.

[0013] 1 is a diagram showing the overall configuration of an analysis system. The analysis system 100 includes a GC / MS 1 and a data processing device 3. The GC / MS 1 is configured to be able to communicate with the data processing device 3. In this embodiment, the data processing device 3 realizes a computer that performs the above processing method.

[0014] The GC / MS 1 includes a gas chromatograph 10 and a mass spectrometer 20. The gas chromatograph 10 includes an injector 11 that introduces a sample, and a column 12 that separates the components of the sample introduced by the injector 11. One or more components contained in the sample are separated while passing through the column 12. Each component is introduced sequentially into the mass spectrometer 20.

[0015] The mass spectrometer 20 includes a vacuum chamber 23 that is evacuated by a vacuum pump (not shown), and an ion source 21, a lens electrode 22, a quadrupole mass filter 24, and an ion detector 25 that are arranged inside the vacuum chamber 23.

[0016] Each component that passes through the column 12 is sequentially introduced into the ion source 21 of the mass spectrometer 20 and ionized. The ionized components are focused by the lens electrode 22, separated by the quadrupole mass filter 24 according to their mass-to-charge ratio (m / z), and then detected by the ion detector 25.

[0017] The mass spectrometer 20 is capable of performing a scan measurement. In a scan measurement, the mass spectrometer 20 scans the mass-to-charge ratios of ions passing through the quadrupole mass filter 24 over a predetermined range, and detects ions for each mass-to-charge ratio using the ion detector 25. The scan measurement is repeated at predetermined time intervals. The detection results obtained by the ion detector 25 are sequentially sent to the data processing device 3. Mass spectrum data is obtained at predetermined time intervals, thereby obtaining time-series data of the mass spectrum (chromatographic mass analysis data). Note that in the present disclosure, the chromatographic mass analysis data is not limited to data obtained as a result of measurement using a GC / MS, and may also be data obtained as a result of measurement using a liquid chromatograph mass analyzer.

[0018] FIG. 2 is a diagram showing a specific example of chromatographic mass spectrometry data. In the following description, chromatographic mass spectrometry data is also referred to as "GCMS data." As shown in graphs G1 and G2 in FIG. 2, GCMS data has three axes: intensity (intensity of detected ions), retention time, and m / z. As shown in graph G1, the chromatographic mass spectrometry data includes multiple mass spectrum data M corresponding to different retention times. As shown in graph G2, the chromatographic mass spectrometry data can also be interpreted as multiple chromatogram data C corresponding to different m / z. In other words, multiple chromatographic data can be identified from the GCMS data.

[0019] 1 again, the data processing device 3 includes a control device 30, a display device 31, and an input device 33. In addition to the function of processing the chromatographic mass spectrometry data, the data processing device 3 may also have the function of controlling each part of the GC / MS 1. Note that the function of controlling each part of the GC / MS 1 may be provided by a device different from the data processing device 3. The data processing device 3 may obtain the detection results (chromatographic mass spectrometry data) obtained by the ion detector 25 from a control device having the function of controlling each part of the GC / MS 1.

[0020] The control device 30 includes a processor 32 and a memory 34. The processor 32 is, for example, a CPU (Central Processing Unit) and is an example of a processing circuitry that executes predetermined arithmetic processing described in a program. The processor 32 reads out the programs and data stored in the memory 34 and executes various processes.

[0021] The memory 34 includes a non-volatile memory or a volatile memory such as a read-only memory (ROM) or a random access memory (RAM), and / or a large-capacity storage device such as a hard disk drive (HDD) or a solid state drive (SSD). The memory 34 non-temporarily stores a program 341 for executing various processes executed by the processor 32, and various data 342. The data 342 includes the detection results obtained by the ion detector 25.

[0022] The display device 31 and the input device 33 are connected to the control device 30. The display device 31 is a device, such as a liquid crystal display (LCD) or an organic electroluminescence (EL) display, that displays the results of calculations by the processor 32. The input device 33 is a device, such as a keyboard, a mouse, a pointing device, or a touch panel, that accepts information input by a user's operation.

[0023] [Example of data processing usage] The GCMS data has values ​​for a plurality of characteristic items. In data processing, the data processing device 3 outputs information suggesting a characteristic item that is expected to affect a certain physical property among the plurality of characteristic items.

[0024] More specifically, in the GCMS data of each of the plurality of samples, the values ​​of each of the plurality of characteristic items are identified as the characteristic quantities.

[0025] Then, for each GCMS data, a set of feature quantities is tagged with information about the physical properties (values ​​about the physical properties or classifications about the physical properties). In the following description, the "information about the physical properties" is also referred to as "physical property information."

[0026] Then, a set of feature sets tagged with the above information is generated as training data for a plurality of samples (a plurality of GCMS data).

[0027] Then, by using the training data, a machine learning model for predicting physical property information from values ​​of at least some of the plurality of feature items is generated. The machine learning model may be a regression model or a classification model.

[0028] Then, for each of one or more feature items among at least some of the feature items, the importance in the machine learning model is identified.

[0029] Then, the importance of each of the one or more feature items is output as an example of the above-mentioned "information suggesting feature items that are expected to affect physical properties."

[0030] In the above description, an example of a feature item is the combination of a retention time range of "9.0 to 9.5 min" and an m / z value (range) of "98." For such a feature item, a value representing the intensity characteristics in the range specified by the combination is specified as a feature amount from the GCMS data. One example of a feature amount is the integrated value of the intensity in the range specified by the feature item. Another example is the maximum intensity value in the range specified by the feature item.

[0031] [Processing flow] Fig. 3 is a diagram showing the flow of data processing in this embodiment. In one implementation example, in the data processing device 3, the processing of Fig. 3 is realized by the processor 32 (CPU) executing a given program. In one implementation example, the data processing device 3 starts the processing of Fig. 3 in response to input of an instruction to start processing of chromatographic mass spectrometry data. The contents of this processing will be described with reference to Fig. 3.

[0032] In step S10, the data processing device 3 reads out the GCMS data of the plurality of samples and adjusts the retention times among the GCMS data of the plurality of samples.

[0033] The adjustment of retention time is also generally referred to as “alignment.” An example of alignment is described in, for example, the literature (Noda, Akira, Chromatogram Retention Time Alignment Algorithm, Shimadzu Review, Special Edition, Shimadzu Corporation, Japan, 2012 (Heisei 24), Vol. 69, No. 3-4, pp. 265-269).

[0034] More specifically, one example of alignment utilizes an improved method of coarse-to-fine dynamic programming (DP), which is a type of DP commonly used in speech recognition and image processing.

[0035] "DP" is a technique that assumes that both retention time and peak intensity fluctuate, and investigates what kind of retention time fluctuations are necessary to best match the two. The DP procedure will be explained below with reference to Figures 4 and 5. Figure 4 shows peak information for two chromatograms. Figure 5 is a diagram illustrating the possibility of correspondence between a reference peak and a target peak.

[0036] The DP procedure includes the following steps a) to d). a) As shown in Figure 4, chromatogram A and chromatogram B are assumed. As shown in Figure 4, by taking into consideration the realistic range of retention time fluctuations (the "range of possible retention time fluctuations" in Figure 4), peaks B(1) to B(3) on chromatogram B are selected as peaks that may correspond to peak A(1) on chromatogram A, or it can be determined that there is no peak on chromatogram B that corresponds to peak A(1) on chromatogram A.

[0037] b) For each of the four candidates for A(1) obtained in a) above (no corresponding point, B(1), B(2), and B(3)), a realistic range of retention time variation is taken into consideration in the same way as in a) above, and peak candidates on chromatogram B are identified.

[0038] As a result, as shown in Figure 5, for A(2), B(1) to B(3) are obtained as candidates for "no corresponding point", B(2) to B(4) for B(1), B(3) to B(5) for B(2), and B(4) to B(6) for B(3) are obtained as candidates.

[0039] The candidates found are combined with "no corresponding point" and A(2) is linked to each of the four candidates for A(1).

[0040] c) By repeating the above steps a) and b), about 4 n candidates for the n peaks on chromatogram A are obtained.

[0041] d) Each candidate is scored based on intensity match and smoothness of time variation, and the candidate with the best score is selected for each peak on chromatogram A.

[0042] In the above example, there are four possible candidates each time from the range of possible retention time fluctuations, and as a result, 4 n candidates are obtained for the n peaks on chromatogram A. In reality, however, this may vary depending on the peak appearance time pattern. If m candidates are obtained on average, the order of the search range is m, where n is the number of peaks. n Since the search range increases exponentially with the number of peaks, DP that searches all of them is not realistic.

[0043] Next, we will explain the "coarse-dense DP". Typically, in DP, a beam search is performed to narrow the peak search range to a realistic range. In a beam search, only the top 100 candidates are always considered, and combinations that fall outside the top 100 are not searched. This has the advantage that the search can be completed in a constant calculation time proportional to the number of peaks n. However, the effect of the beam search is that the range in which a perfect search is possible is narrowed, and local optima in a narrow time range accumulate. This has the side effect of causing only a few incorrect matches in a row at an early time, and no matches at all in the latter half.

[0044] To address this side effect, in addition to beam search, DP processing divided into two stages, coarse and fine (coarse-fine DP), is used.

[0045] In the coarse stage, peaks are thinned out to reduce the value of n by the order of the search range, so that accurate alignment can be achieved for coarse retention time variations within a limited beam width.

[0046] Peaks with large intensities are considered to be reliable peaks, and only the large peaks are used in the coarse stage DP.

[0047] The retention time variation between multiple chromatograms consists of large slow variations and small variations for each peak. Therefore, after the coarse-stage DP, only small variations remain between multiple chromatograms. These small variations are then processed in the fine-stage DP.

[0048] In dense-stage DP, the variation between multiple chromatograms is small, so the search range m can be set small, allowing a large number of peaks n to be accurately processed even with a small beam range.

[0049] By performing the two-stage process as described above, it is possible to narrow the search range in both stages, thereby reducing the side effect of not being able to find a match due to the beam search.

[0050] In the implementation of coarse-to-fine DP, the results of the coarse stage are always assumed to be correct in the fine stage in order to efficiently reduce the search range. However, due to the characteristics of the algorithm, DP that performs beam search can only evaluate the degree of match on a one-way time axis from the past to the future. This can occasionally result in unnatural matches overall.

[0051] In order to remove such exceptional matches (unnatural matches), a filtering process may be performed. Fig. 6 is a diagram showing an example of exceptional matches. In determining exceptional matches, as shown in Fig. 6, the presence or absence of extreme deviations is detected using the local standard deviation of the amount of fluctuation in intensity per unit time ("time fluctuation" in Fig. 6).

[0052] Returning to FIG. 3 , after step S10, in step S12, the data processing device 3 performs baseline removal on the chromatogram data for each m / z. The chromatogram data is identified from each of the GCMS data of the multiple samples read in step S10. A percentile filter may be used for baseline removal. Furthermore, if blank measurements are periodically performed in the device that acquired the GCMS data, baseline removal may be achieved by subtracting the results of the blank measurements.

[0053] In step S14, the data processing device 3 adjusts the intensity of each of the GCMS data of the multiple samples read out in step S10. In adjusting the intensity, for each GCMS data, the data processing device 3 multiplies the intensity of all signal points by a uniform coefficient so that the integrated value of the intensities of all signal points in the two-dimensional intensity matrix becomes 1 (or some common value).

[0054] In step S16, the data processing device 3 extracts feature quantities for a plurality of feature items from each GCMS data.

[0055] In one implementation, multiple features have different ranges, each specified by a retention time interval of 60 seconds and an m / z interval of 1. For example, if the retention time range is 6,000 seconds and the m / z range is 500, 100 features are specified for the retention time and 500 features are specified for the m / z range, resulting in a total of 50,000 features. To mitigate the risk of peaks being separated midway or the risk of the same peak in different files belonging to different intervals due to retention time differences, adjacent features may overlap at both ends of their retention times by, for example, about 6 seconds. That is, a feature adjacent to other features at both ends may have a retention time range of 72 seconds (6 + 60 + 6). A feature adjacent to only one end may have a retention time range of 66 seconds (6 + 60 or 60 + 6).

[0056] Fig. 7 is a diagram showing specific examples of multiple characteristic items. In Fig. 7, four types of characteristic items are shown in columns F01 to F04. Column F01 shows a combination of retention time 9.0 to 9.5 minutes and m / z = 98 as a characteristic item. Column F02 shows a combination of retention time 10.0 to 10.5 minutes and m / z = 101 as a characteristic item. Column F03 shows a combination of retention time 18.0 to 18.5 minutes and m / z = 95 as a characteristic item. Column F04 shows a combination of retention time 16.0 to 16.5 minutes and m / z = 102 as a characteristic item.

[0057] Each of columns F01 to F04 shows chromatogram data corresponding to a feature item extracted from certain GCMS data. In each chromatogram data, the vertical axis represents intensity and the horizontal axis represents retention time. The feature amount of each feature item can be derived using the chromatogram data.

[0058] In step S18, the data processing device 3 performs a process of reducing the number of feature items used in generating a machine learning model, which will be described later, from the "plurality of feature items" in step S16. In one implementation example, feature items that satisfy a given condition are excluded from generating the machine learning model. In one example of the reduction process, the value (feature amount) of each feature item is used. More specifically, for each feature item, the maximum value of the feature amounts of the multiple GCMS data is identified. Then, feature items whose maximum value is equal to or less than a given threshold are excluded from the target of the machine learning model. Note that the feature amount of the excluded feature item may be assumed to be noise.

[0059] In another implementation example, a group of feature items that satisfy a given condition is treated as a single feature item, thereby reducing the number of feature items used to generate a machine learning model. More specifically, the data processing device 3 performs clustering on feature amounts of feature items that have a common retention time, using a correlation coefficient as a distance function. This generates a cluster of feature items. The data processing device 3 treats the cluster as a new feature item. The feature amount of the new feature item is identified by merging (adding or accumulating) the feature amounts of the original feature items.

[0060] In step S20, the data processing device 3 selects feature items to be used in generating a machine learning model, which will be described later. In the processing of FIG. 3, step S18 may be omitted. If step S18 is omitted, in step S20, feature items to be used in generating a machine learning model are selected from the "plurality of feature items" in step S16. If step S18 is performed, in step S20, feature items to be used in generating a machine learning model are selected from the feature items identified to be used in generating a machine learning model after the reduction in step S18 (i.e., at least some of the plurality of feature items).

[0061] In step S20, the data processing device 3 selects features by removing features that do not affect the target physical property using an existing machine learning-based feature selection method, such as Random Forest-based Boruta, PLS-based Boruta, or Recursive feature elimination (RFE).

[0062] In step S22, the data processing device 3 generates training data for the machine learning model. The training data includes data on a plurality of samples. The training data is generated by tagging a set of values ​​of feature items with physical property values ​​for each of the GCMS data of the plurality of samples. The physical property values ​​may be numerical values ​​(viscosity, glass transition temperature, etc.) or classifications (e.g., whether a patient corresponding to the sample has cancer or does not have cancer).

[0063] In step S24, the data processing device 3 generates a machine learning model using the training data generated in step S22. An existing machine learning method (Partial Least Squares (PLS) regression, Random Forest, etc.) can be used to generate the machine learning model. The generated machine learning model derives prediction results for physical property values ​​by inputting values ​​of feature items that make up the training data.

[0064] In step S26, the data processing device 3 derives the importance of each of one or more feature items in the generated machine learning model. The importance of each feature item may be the VIP of PLS ​​or the feature importance of RandomForest.

[0065] For example, assume that the following equation (1) is specified as the machine learning model. y=a·x1+b·x2+c·x3 …(1) In equation (1), y represents the explanatory function (prediction result). x1, x2, and x3 represent the values ​​of the feature items. a, b, and c represent the coefficients of x1, x2, and x3, respectively.

[0066] In the case of equation (1), the importance of each of x1, x2, and x3 is specified as a respective coefficient (i.e., a, b, and c).

[0067] In step S28, the data processing device 3 derives a performance index for the generated machine learning model. In one implementation example, the data processing device 3 uses the generated machine learning model to obtain a prediction result for the value of a physical property by applying values ​​of one or more feature items not included in the training data to the machine learning model, and derives a performance index using the prediction result and the actual value of the physical property. If the machine learning model is a regression model, the performance index may be a coefficient of determination. If the machine learning model is a classification model, the performance index may be an F-measure.

[0068] In step S30, the data processing device 3 outputs the importance derived in step S26. The output may be displayed on the display device 31 or transmitted to an external device. Thereafter, the data processing device 3 ends the processing in Fig. 3. Note that in step S30, the performance index derived in step S28 may also be output.

[0069] In the embodiment described above, a machine learning model is generated, the importance of each of one or more feature items in the machine learning model is identified, and each importance is output. Note that a machine learning model may be generated for each of two or more materials. Training data generated from each of multiple samples of a certain material may be used to generate a machine learning model for that material.

[0070] [Example of importance output] 8 is a diagram showing an example of a screen for displaying the importance level. In one implementation example, the data processing device 3 displays the importance level output in the above-mentioned step S30 on a screen 500 of FIG.

[0071] Screen 500 includes display fields 501 and 502. Display field 501 includes information about the generated machine learning model. In the example of Fig. 8, display field 501 displays a regression equation representing the machine learning model and an explanation of the variables in the regression equation.

[0072] In the example of FIG. 8, the display field 501 targets the same machine learning model as that explained above as equation (1). That is, y represents the explanatory function (prediction result). x1, x2, and x3 represent the values ​​of the feature items. a, b, and c represent the coefficients of x1, x2, and x3, respectively.

[0073] The information shown as the description ("physical property") of the variable y in the display field 501 corresponds to the prediction result of the machine learning model. For example, if a value of "viscosity" is derived as the prediction result, "viscosity" is displayed as the "physical property." If a classification of "has cancer" or "does not have cancer" is derived as the prediction result, "has cancer" is displayed as the "physical property."

[0074] The information displayed as explanations of variables x1, x2, and x3 in display field 501 displays the contents of the feature items (combinations of retention time ranges and m / z values) corresponding to each of variables x1, x2, and x3 in the machine learning model. That is, on the actual screen, specific combinations of retention time ranges and m / z values ​​are displayed as "feature item (1)," "feature item (2)," and "feature item (3)" in Fig. 8, respectively.

[0075] In a display field 502, the importance of each feature item in the display field 501 is displayed. In the example of Figure 8, the coefficient of each feature item in the machine learning model is displayed as the importance of each feature item. For example, the following formula (2) is generated as the machine learning model for material B.

[0076] y=5.43 x1+0.32 x2+5.43 x3 …(2) In formula (2), the coefficient for feature item (1) represented by x1 is "5.43." Therefore, in display field 502, "5.43" is displayed as the importance of feature item (1). Furthermore, the coefficient for feature item (2) represented by x2 is "0.32," so "0.32" is displayed as the importance. For feature item (3) represented by x3, the coefficient is "5.43," so "5.43" is displayed as the importance.

[0077] In the present embodiment described above, the importance of characteristic items that are expected to affect the physical properties of each material is displayed, allowing the operator to easily check whether it is worth spending money to search for characteristic items that affect the target physical properties of each material. The operator may also refer to the performance index of the machine learning model when making this check.

[0078] In this embodiment, by using a machine learning model, feature items that are expected to affect physical properties are derived without the need for detailed (in some sense, difficult) parameter adjustment by an operator. Furthermore, the derivation of feature items can be realized objectively by a machine learning model that operates robustly.

[0079] The operator may use the feature items and their importance (as well as the performance index of the machine learning model) output in this embodiment as information for primary screening. After confirming the likelihood of incurring costs, the operator performs costly secondary screening processes (such as improving separation by changing the method, painstakingly careful peak detection and peak identification) on the obtained candidate feature items, ultimately identifying feature items that affect physical properties.

[0080] In this embodiment, in generating a machine learning model, data correction as described in steps S10 to S14 is performed, thereby mitigating the effects of fluctuations in retention time and intensity due to long-term measurements in data analysis.

[0081] In this embodiment, with regard to the extraction of feature quantities in step S16, the feature items have a range of retention times rather than a single retention time, which enables robust analysis even when the retention time difference between GCMS data is not completely corrected.

[0082] In this embodiment, in the reduction of feature items in step S18, the values ​​of feature items in physicochemically meaningful groups (= mass spectra derived from the same component) are merged and used to generate a machine learning model, thereby reducing the multicollinearity of explanatory variables.

[0083] [Aspect] It will be appreciated by those skilled in the art that the exemplary embodiments described above are examples of the following aspects.

[0084] (Item 1) A method for processing chromatographic mass spectrometry data according to one aspect is a computer-executed method for processing chromatographic mass spectrometry data, the method comprising: extracting values ​​of a plurality of feature items from chromatographic mass spectrometry data of a sample, wherein each of the plurality of feature items corresponds to a combination of a range of retention time and a mass-to-charge ratio, and the plurality of combinations corresponding to each of the plurality of feature items are different from one another; the chromatographic mass spectrometry data is associated with physical property information representing a given physical property; generating training data for a plurality of the samples, wherein the training data includes at least some of the plurality of feature items and the physical property information for the chromatographic mass spectrometry data of each of the samples; generating a machine learning model using the training data to predict the physical property information from values ​​of the at least some of the feature items; identifying an importance level in the machine learning model for each of one or more of the at least some feature items; and outputting the importance level for each of the one or more feature items.

[0085] According to the method for processing chromatography mass spectrometry data described in paragraph 1, a technique is provided to assist an operator who uses chromatography mass spectrometry data to search for features that affect specific physical properties.

[0086] (Item 2) In the method for processing chromatographic mass spectrometry data according to item 1, the plurality of combinations may have overlapping retention time ranges.

[0087] According to the method for processing chromatographic mass spectrometry data described in paragraph 2, it is possible to avoid a situation in which a characteristic change such as a peak is separated by adjacent characteristic items, resulting in the characteristic change not being detected within the range defined by a single characteristic item.

[0088] (Clause 3) The method for processing chromatographic mass spectrometry data described in clause 1 or 2 may further include a step of identifying the at least some of the feature items by removing, from the plurality of feature items, feature items having values ​​lower than a given threshold, wherein the at least some of the feature items are fewer than the plurality of feature items.

[0089] According to the method for processing chromatographic mass spectrometry data described in Section 3, the number of feature items used to generate a machine learning model can be appropriately reduced.

[0090] (Item 4) The method for processing chromatographic mass spectrometry data according to any one of Items 1 to 3 may further include a step of identifying the at least some of the characteristic items by combining two or more characteristic items into one characteristic item, wherein the at least some characteristic items are fewer than the plurality of characteristic items.

[0091] According to the method for processing chromatographic mass spectrometry data described in Section 4, the number of feature items used to generate a machine learning model can be appropriately reduced.

[0092] (Item 5) The method for processing chromatographic mass spectrometry data according to any one of items 1 to 4 may further comprise a step of selecting at least some of the characteristic items from the plurality of characteristic items by removing characteristic items that do not affect the given physical property.

[0093] According to the method for processing chromatographic mass spectrometry data described in Section 5, the number of feature items used to generate a machine learning model can be appropriately reduced.

[0094] (Item 6) In the method for processing chromatographic mass spectrometry data described in any one of Items 1 to 5, the machine learning model may be expressed as a linear expression, and the importance may be determined by specifying a coefficient for each of the one or more feature items in the linear expression.

[0095] According to the method for processing chromatographic mass spectrometry data described in Section 6, the basis for identifying the importance becomes clear.

[0096] (Item 7) A program according to one aspect may be executed by one or more processors of a computer to cause the computer to carry out the method for processing chromatographic mass spectrometry data according to any one of items 1 to 6.

[0097] According to the program described in paragraph 7, a technique is provided to assist an operator who searches for features that affect specific physical properties using chromatography mass spectrometry data.

[0098] The embodiments disclosed herein should be considered to be illustrative in all respects and not restrictive. The scope of the present disclosure is defined by the claims, not by the description of the above embodiments, and is intended to include all modifications within the meaning and scope of the claims. Furthermore, it is intended that each technique in the embodiments can be implemented alone or, if necessary, in combination with other techniques in the embodiments to the extent possible. [Explanation of symbols]

[0099] 1 Gas chromatograph mass spectrometer, 3 Data processing device, 10 Gas chromatograph, 11 Injector, 12 Column, 20 Mass spectrometer, 21 Ion source, 22 Lens electrode, 23 Vacuum chamber, 24 Mass filter, 25 Ion detector, 30 Control device, 31 Display device, 32 Processor, 33 Input device, 34 Memory, 100 Analysis system, 500, 510 Screen, 501, 502, 520 Display column, C Chromatogram data, M Mass spectrum data.

Claims

1. 1. A computer-implemented method for processing chromatographic mass spectrometry data, comprising: extracting values ​​of a plurality of features from the chromatographic mass spectrometry data of the sample; each of the plurality of feature items corresponds to a combination of a retention time range and a mass-to-charge ratio; the plurality of combinations corresponding to the plurality of characteristic items are different from one another; the chromatographic mass spectrometry data is associated with physical property information representing a given physical property; generating training data for a plurality of said samples; the training data includes at least some of the plurality of feature items and the physical property information for the chromatographic mass spectrometry data of each of the samples; generating a machine learning model for predicting the physical property information from values ​​of at least some of the feature items using the training data; Identifying the importance of each of one or more feature items among the at least some feature items in the machine learning model; and outputting the importance of each of the one or more feature items.

2. The method for processing chromatographic mass spectrometry data according to claim 1 , wherein the plurality of combinations have overlapping retention time ranges.

3. the at least some characteristic items are fewer than the plurality of characteristic items; 3. The method for processing chromatographic mass spectrometry data according to claim 1, further comprising the step of identifying the at least some of the feature items by removing feature items having values ​​lower than a given threshold from the plurality of feature items.

4. the at least some characteristic items are fewer than the plurality of characteristic items; The method for processing chromatographic mass spectrometry data according to claim 1 or claim 2, further comprising a step of identifying at least some of the plurality of feature items by integrating two or more feature items into one feature item.

5. 3. The method for processing chromatographic mass spectrometry data according to claim 1, further comprising the step of selecting the at least some of the feature items by removing feature items that do not affect the given physical property from the plurality of feature items.

6. The machine learning model is expressed as a linear equation, 3. The method for processing chromatographic mass spectrometry data according to claim 1, wherein the coefficient of each of the one or more feature items in the linear expression is specified as the importance.

7. A program that, when executed by one or more processors of a computer, causes the computer to carry out the method for processing chromatography mass spectrometry data according to claim 1 or 2.