Information processing system, information processing method, and program
Patent Information
- Application Number
- PCT/JP2026/012674
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2026-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure JP2026012674_01102026_PF_FP_ABST
Abstract
Description
Information processing system, information processing method, and program
[0001] The present invention relates to an information processing system, an information processing method, and a program for generating information on a partial structure of a compound from mass spectrum information.
[0002] Conventionally, in order to estimate the molecular structure of an unknown compound, or to determine or estimate the chemical properties of an unknown compound, it has been desired to identify the partial structure (the partial chemical structure in the molecule of the unknown compound) of the unknown compound.
[0003] As methods for analyzing an unknown compound, there are methods including analysis using a nuclear magnetic resonance apparatus after fractional purification, and analysis using a mass spectrometer. From the perspective of throughput, mass spectra obtained by mass spectrometers are widely used.
[0004] In mass spectrum analysis, the molecular formula and partial structure can be analyzed based on information of peaks and peak intervals. In addition, the overall structure can be estimated based on a combination of such information.
[0005] As such a mass spectrum analysis method, the following Patent Document 1 discloses a partial structure estimation apparatus including a partial structure estimator that generates a first explanatory variable by performing composition estimation for each peak based on a peak group included in a mass spectrum obtained from a sample, generates a second explanatory variable by performing composition estimation for each peak interval based on the peak group, and estimates a partial structure of the sample as an objective variable based on the first explanatory variable and the second explanatory variable.
[0006] Japanese Unexamined Patent Publication No. 2023-124547
[0007] However, in the technique of the above-mentioned Patent Document 1, it is necessary to generate the first explanatory variable and the second explanatory variable by performing composition estimation through mathematical processing while exhaustively changing the combination of elements for each peak and each peak interval. Therefore, the number of dimensions of each explanatory variable becomes enormous, which increases the time and cost for estimating the partial structure.
[0008] The object of the present invention is to provide an information processing system, information processing method, and program that can easily and quickly predict the presence or absence of substructures of compounds contained in an unknown sample from the mass spectral information of the sample.
[0009] An information processing system according to one embodiment of the present invention comprises a control unit. The control unit acquires the peak intervals in a group of peaks included in a mass spectrum obtained from a sample, uses the peak intervals as explanatory variables, and uses the presence or absence of a substructure of the compound in the sample as the target variable, to generate a machine learning model that predicts the target variable from the explanatory variables.
[0010] An information processing system according to another embodiment of the present invention comprises a control unit. The control unit inputs the peak interval obtained from a mass spectrum obtained from an unknown sample and selection information for selecting at least one substructure to a machine learning model that has been trained to predict the presence or absence of a substructure of a compound in the sample, using the peak interval in a group of peaks included in a mass spectrum obtained from a sample as an explanatory variable and the presence or absence of a substructure of a compound in the sample as an objective variable. The control unit then generates structural information regarding the structure of the unknown sample based on the presence or absence of the selected substructure output from the machine learning model and transmits it to the user's terminal.
[0011] Another embodiment of the present invention involves obtaining the peak intervals in a group of peaks included in a mass spectrum obtained from a sample, using the peak intervals as explanatory variables, and using the presence or absence of a substructure of the compound in the sample as the target variable, to generate a machine learning model that predicts the target variable from the explanatory variables.
[0012] An information processing method according to yet another embodiment of the present invention includes inputting the peak interval obtained from a mass spectrum obtained from an unknown sample and selection information for selecting at least one substructure to a machine learning model that has been trained to predict the presence or absence of a substructure of a compound in the sample, with the peak interval in a group of peaks included in a mass spectrum obtained from a sample as an explanatory variable and the presence or absence of a substructure of a compound in the sample as an objective variable, and generating structural information regarding the structure of the unknown sample based on the presence or absence of the selected substructure output from the machine learning model and transmitting it to the user's terminal.
[0013] A program according to another embodiment of the present invention causes an information processing device to perform the steps of: obtaining the interval between peaks in a group of peaks included in a mass spectrum obtained from a sample; and generating a machine learning model that predicts the target variable from the explanatory variable, with the interval between peaks as an explanatory variable and the presence or absence of a substructure of the compound in the sample as the target variable.
[0014] A program according to yet another embodiment of the present invention causes an information processing device to perform the following steps: input the peak interval obtained from a mass spectrum obtained from an unknown sample and selection information for selecting at least one substructure to a machine learning model that has been trained to predict the presence or absence of a substructure of a compound in the sample, with the peak interval in a group of peaks included in a mass spectrum obtained from a sample as an explanatory variable and the presence or absence of a substructure of a compound in the sample as an objective variable; and generate structural information regarding the structure of the unknown sample based on the presence or absence of the selected substructure output from the machine learning model and transmit it to the user's terminal.
[0015] According to an information processing system, information processing method, and program in one embodiment of the present invention, it is possible to easily and quickly predict the presence or absence of substructures of compounds contained in an unknown sample from the mass spectral information of the sample. However, this effect is not limited to the present invention.
[0016] This figure shows the configuration of a sample structure analysis system according to one embodiment of the present invention. This figure shows the hardware configuration of a sample structure analysis server according to one embodiment of the present invention. This figure shows the configuration of a database owned by a sample structure analysis server according to one embodiment of the present invention. This is a flowchart showing the process of generating a machine learning model by a sample structure analysis server according to one embodiment of the present invention. This figure describes an adduct of a nucleophile and a sensitizing substance to be analyzed in one embodiment of the present invention. This figure shows the mass spectrum information obtained by analyzing the adduct in Figure 5. This figure shows the process of generating a machine learning model by a sample structure analysis server according to one embodiment of the present invention. This is a flowchart showing the process of generating sample structure information by a sample structure analysis server according to one embodiment of the present invention. This figure shows an example of structure information generated by a sample structure analysis server according to one embodiment of the present invention and displayed on a user terminal. This figure shows another example of structure information generated by a sample structure analysis server according to one embodiment of the present invention and displayed on a user terminal. This figure describes a substructure analyzed based on peak intervals by a sample structure analysis server according to one embodiment of the present invention.
[0017] Embodiments of the present invention will be described below with reference to the drawings.
[0018] [System Configuration] The system of this embodiment is a sample structure analysis system for human safety tests such as skin sensitization tests, and as shown in Figure 1, it includes a sample structure analysis server 100 on the Internet 5, a plurality of user terminals 200, an analytical laboratory terminal 300, and a mass spectrometer 400.
[0019] For products that come into contact with people, such as pharmaceuticals and cosmetics, it is crucial that they do not cause allergic reactions under actual usage conditions. Therefore, in product development, it is necessary to evaluate the skin sensitization risk of the ingredients used. Skin sensitization has been evaluated using animal experiments such as the Guinea Pig Maximization Test, the Buehler Test, and the Local Lymph Node Assay.
[0020] In recent years, due to legal regulations on animal testing and increased consumer awareness of animal welfare, alternative testing methods have been developed to evaluate the skin sensitization potential of chemicals in vitro, in chemico, or in silico without using animals. Several methods have been developed to replicate the events that occur in the skin as alternatives to skin sensitization, such as the Direct peptide reactivity assay (DPRA), Amino acid derivative assay (ADRA), KeratinoSens, and h-CLAT.
[0021] In particular, DPRA and ADRA are in chemico alternatives that reproduce the reaction between sensitizing substances and skin proteins. However, these tests rely on the rate of decrease of the peptide nucleophile when it reacts with the sensitizing substance. Therefore, it is difficult to identify the structure of skin sensitizing substances contained in unknown samples.
[0022] Therefore, the inventors have developed a technique for easily and quickly predicting the presence or absence of a partial structure of a skin-sensitizing substance contained in a sample.
[0023] The sample structure analysis server 100 is a server (information processing device) that performs a service to provide users with structural information regarding the partial structures of compounds contained in sample materials, etc., based on the analysis results of sample samples provided by users. The sample structure analysis server 100 is connected to multiple user terminals 200 via the Internet 5.
[0024] User terminals 200 (200A, 200B, 200C, etc.) are terminals used by users, such as notebook PCs (Personal Computers), smartphones, mobile phones, tablet PCs, desktop PCs, etc. Users are employees responsible for product safety evaluation in industries that handle raw materials and products that come into contact with human skin, such as raw material manufacturers, consumer goods manufacturers, medical device manufacturers, and toy manufacturers.
[0025] The analytical laboratory terminal 300 is a terminal used by the person in charge of analyzing the sample provided by the user at the analytical laboratory, and can be, for example, a notebook PC, smartphone, mobile phone, tablet PC, or desktop PC. A mass spectrometer 400 is used for the analysis, and the analytical laboratory terminal 300 transmits the analysis results (mass spectrum information) to the sample structure analysis server 100.
[0026] The mass spectrometer 400 is used to perform an alternative method for skin sensitization testing on the sample provided by the user. Specifically, the mass spectrometer 400 performs mass spectrometry on the sample (skin sensitizing substance) and the compound (adduct) obtained by reacting the sample with a nucleophile in order to detect trace amounts of chemical substances such as skin sensitizing substances from the sample, thereby generating a mass spectrum of the adduct. The mass spectrometer 400 is, for example, an LS-MS / MS (liquid chromatography-tandem mass spectrometry) instrument using liquid chromatography.
[0027] In addition to liquid chromatography, other types of chromatography can also be used, such as supercritical fluid chromatography, ion chromatography, gas chromatography, and thin-layer chromatography.
[0028] Examples of liquid chromatography include normal-phase chromatography, reverse-phase chromatography, size exclusion chromatography, and ion exchange chromatography. From the viewpoint of excellent detection performance, reverse-phase chromatography or ion exchange chromatography is preferred, and reverse-phase chromatography is more preferred.
[0029] Mass spectrometry is not particularly limited as long as it is an analytical method that utilizes the analysis of mass. In addition to the tandem mass spectrometry (MS / MS) method described above, conventional mass spectrum acquisition methods can also be used.
[0030] Methods for obtaining mass in MS / MS include, for example, the Selected Reaction Monitoring method, the Precursor Ion Scan method, and the Constant Neutral Loss Scan method. Fragmentation methods in MS / MS include collision-induced dissociation, EAD (electron-induced dissociation), UVPD (ultraviolet photodissociation), HAD (hydrogen attachment / extraction dissociation), and OAD (oxygen attachment dissociation). In addition to these, any analytical methods included in MS / MS that are developed in the future can also be adopted.
[0031] The sample structure analysis server 100 uses a machine learning model (classifier) 10 to generate structural information based on the above-mentioned substructure. The machine learning model 10 is trained to predict the presence or absence of a specific substructure of the compound in the adduct, using the peak position (measured (precise) mass-to-charge ratio) and / or the peak interval (corresponding to the measured (precise) mass-to-charge ratio difference) in the group of peaks (multiple peaks) included in the mass spectrum obtained from the adduct as explanatory variables. A substructure is a functional group or a unit containing a functional group, and multiple machine learning models 10 are prepared for each different substructure. The machine learning model 10 may be generated by a learning process via the sample structure analysis server 100, or it may be generated by another information processing device.
[0032] When the sample structure analysis server 100 receives a request for sample structure analysis from the user terminal 200, it inputs extracted information from the mass spectral information obtained from the analysis laboratory terminal 300 into the machine learning model 10. Based on the substructure (presence or absence) information output by the machine learning model, it generates structural information regarding the structure of the user's sample and sends it to the user terminal 200.
[0033] [Hardware configuration of the sample structure analysis server] As shown in Figure 2, the sample structure analysis server 100 includes a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, an input / output interface 15, and a bus 14 that connects these components to each other.
[0034] The CPU 11 accesses RAM 13 and other memory as needed, performing various calculations and comprehensively controlling each block of the sample structure analysis server 100. Multiple CPUs 11 may be provided depending on the processing. ROM 12 is a non-volatile memory in which the OS, programs, and firmware such as various parameters to be executed by the CPU 11 are permanently stored. RAM 13 is used as a working area for the CPU 11 and temporarily holds the OS, various running applications, and various data being processed.
[0035] The input / output interface 15 is connected to a display unit 16, an operation reception unit 17, a storage unit 18, a communication unit 19, and the like.
[0036] The display unit 16 is a display device that uses, for example, an LCD (Liquid Crystal Display), an OLED (Organic ElectroLuminescence Display), or a CRT (Cathode Ray Tube).
[0037] The operation reception unit 17 is, for example, a pointing device such as a mouse, a keyboard, a touch panel, or other input device. If the operation reception unit 17 is a touch panel, the touch panel may be integrated with the display unit 16.
[0038] The storage unit 18 is a non-volatile memory such as an HDD (Hard Disk Drive), flash memory (SSD; Solid State Drive), or other solid-state memory. The OS, various applications, and various data are stored in this storage unit 18.
[0039] As will be described later, particularly in the present embodiment, the storage unit 18 includes a compound information database, a mass spectrum information database, a user information database, and an analysis result information database, in addition to programs such as applications necessary for the generation processing of the machine learning model 10 described later and the sample structure analysis processing using the same.
[0040] The communication unit 19 is, for example, a network interface card (NIC) for Ethernet or various modules for wireless communication such as wireless LAN, and is responsible for communication processing with the user terminal 200 and the analysis institution terminal 300.
[0041] Although not shown in the figure, the basic hardware configuration of the user terminal 200 and the analysis institution terminal 300 is also substantially the same as the hardware configuration of the sample structure analysis server 100 described above.
[0042] [Database Configuration of Sample Structure Analysis Server] As shown in FIG. 3, the sample structure analysis server 100 includes, in the storage unit 18, a compound information database 31, a mass spectrum information database 32, a user information database 33, and an analysis result information database 34. Note that each of these databases may be stored not in the storage unit 18 but in a storage device or server externally connected to the sample structure analysis server 100.
[0043] The compound information database 31 stores information on a plurality of known compounds and information on partial structures constituting the compounds (for example, names, structural formulas, values of mass-to-charge ratio (m / z), etc.). The structural information is described, for example, in the SMILES (Simplified Molecular Input Line Entry System) format.
[0044] The mass spectrum information database 32 stores a plurality of pieces of mass spectrum information obtained by analyzing the above compounds, information on peak positions (product ion values) and peak intervals (neutral loss values) calculated from the mass spectrum information, and presence / absence information of a plurality of partial structures of each compound corresponding to the peak positions and peak intervals, etc.
[0045] The mass spectral information includes not only information that has already been publicly available, but also information that has been originally analyzed by the mass spectrometer 400 described above. The mass spectral information includes, for example, information such as compound name, measuring instrument name, SMILES information, collision energy, adduct, m / z of precursor ions, charge, polarity, and product ion spectrum.
[0046] Examples of partial structures to be analyzed in the present embodiment include, but are not limited to, -COOH (carboxyl group), -CO-NH2 (amide group), -COO- (ester bond), -C2H4O- (oxyethylene group), -CH3 (methyl group), -OSO3H (sulfate ester group), -OH (hydroxy group), -NH2 (amino group), -C3H6O- (oxypropylene group), and -CHO (aldehyde group).
[0047] The user information database 33 stores, for each user, attribute information of users of the sample structure information providing service provided by the sample structure analysis server 100. The user attribute information includes information such as full name, company name, department name, user ID, e-mail address, and past structural analysis history information.
[0048] The analysis result information database 34 stores structural information (such as presence / absence information of partial structures) of a sample provided by a user, which is obtained by inputting information of the aforementioned peak positions and peak intervals calculated from mass spectral information acquired from the aforementioned mass spectrometer 400 into the machine learning model 10, in association with the aforementioned user ID.
[0049] Each of these databases is mutually referenced and used as needed in the processing for generating the machine learning model 10 by the sample structure analysis server 100 described later and the sample structure analysis processing using the machine learning model 10.
[0050] [Operation of the Sample Structure Analysis Server] Next, the operation of the sample structure analysis server 100 configured as described above will be explained. This operation is performed through the cooperation of the hardware of the sample structure analysis server 100, such as the CPU 11 and communication unit 19, and the software stored in the storage unit 18. For convenience, in the following explanation, the CPU 11 will be considered the main operator.
[0051] (Machine Learning Model Generation Process) Figure 4 is a flowchart showing the process of generating the machine learning model 10 by the sample structure analysis server 100. Figure 7 is a diagram illustrating the general process of generating the machine learning model 10.
[0052] As shown in Figure 4, first the CPU 11 inputs mass spectral information for the target compound from the mass spectral information database 32 (step 41).
[0053] As described above, the mass spectral information includes the compound name, measuring instrument name, SMILES, collision energy, adduct, precursor ion m / z, charge, polarity, and product ion spectrum information. The product ion spectrum information can be cut at any percentage relative to the maximum intensity. In this embodiment, data cut at 1% was used. The CPU 11 then exports the SMILES and product ion information from this data, for example, in JSON format.
[0054] As mentioned above, the mass spectral information is obtained by performing mass spectrometry on the compound (adduct) formed by reacting a sample provided by the user with a nucleophile using the mass spectrometer 400. The nucleophile mainly consists of an organic compound containing a group selected from, for example, a thiol group or an amino group, such as N-acetylcysteine or N α -Acetyllysine and the like are used.
[0055] By using the mass spectral information of the compound (adduct) reacted with the nucleophile, even sensitizing substances present in trace amounts in the raw materials can be selectively and sensitively detected, enabling structural analysis.
[0056] The nucleophile may consist solely of the above-mentioned organic compound, or it may contain one or more additives such as pH adjusters, stabilizers, and surfactants in addition to the organic compound that is the main component of the measurement. The nucleophile may also be prepared by dissolving the above-mentioned organic compound and, if necessary, the additives in water, aqueous buffer solution, organic solvent, or a mixture thereof. Furthermore, the nucleophile may be in any form, such as a solution, liquid, or solid (powder, granules, freeze-dried product, tablet, etc.).
[0057] If the nucleophile is in solution form, the nucleophile solution is prepared in a buffer solution in accordance with existing alternative methods for skin sensitization testing, and the skin sensitizing substance (test substance) is dissolved in acetonitrile or water, acetone, or acetonitrile containing 5% DMSO. The test substance may also be dissolved in the solvent in which the raw materials are dissolved or in other solvents.
[0058] Furthermore, instead of nucleophiles, other derivatizing reagents that react with the skin-sensitizing substance (test substance) by mechanisms other than nucleophilic addition may be used. For example, compound products (adducts) obtained by reacting with derivatizing reagents that perform electrophilic reactions, radical reactions, or cycloaddition reactions such as the Diels-Alder reaction or the Huesgen cycloaddition reaction may be used. Examples of groups that perform electrophilic reactions include alkyl halides, carbonyl groups, nitrile groups, and Michael acceptors. Examples of groups that perform radical reactions include azo groups.
[0059] Furthermore, when the test substance is a resin, an extraction procedure is necessary. Components in the resin are extracted using polar / non-polar solvents as shown in Examples 1) and 2) below, and the extract is then tested. Example 1) Solvents listed in ISO-10993-10: methanol and acetone, olive oil, DMSO, hexane, cyclohexane, isopropyl alcohol. Example 2) Solvents listed in the Food Sanitation Law: 4% acetic acid (simulating acidic foods with a pH of 5 or less), water (simulating foods with a pH of 5 or more), 20% ethanol (simulating alcoholic beverages), heptane (simulating oils and fatty products).
[0060] Furthermore, the extract can be concentrated, and the dried sample can be redissolved in a polar or nonpolar solvent for further testing.
[0061] In this embodiment, as shown in Figure 5, N-acetylcysteine was used as the nucleophile and propyl gallate was used as the sensitizing agent, and the propanol unit derived from propyl gallate was analyzed from the adduct of N-acetylcysteine and propyl gallate.
[0062] Next, the CPU 11 calculates the peak position (product ion value) from the input mass spectrum information (step 42). Specifically, the CPU 11 extracts product ions with m / z values in the range of 70 to 1000 for the compound from the product ions exported from the mass spectrum information, as integer values. If there are multiple ions with the same integer part but different decimal parts, the one with the highest intensity is selected. The CPU 11 then stores this result in a data frame df1 containing SMILES and product ions. Alternatively, the measured precise mass data may be used directly to calculate the peak position instead of integer values.
[0063] Next, the CPU 11 calculates the peak interval (neutral loss) from the mass spectrum information (step 43). That is, the CPU 11 selects the top 5 m / z intensities from the product ion spectrum and calculates the neutral loss, which is the absolute value of the difference between them. The CPU 11 then stores this result in a data frame df2, which includes SMILES and the neutral loss.
[0064] Next, the CPU 11 receives information from the compound information database 31 indicating a specific substructure to be determined by the machine learning model 10, and determines whether the SMILES information in df1 and df2 has that specific substructure. If it does, it labels it 1, and if it does not, it labels it 0 (step 44). The substructure is entered, for example, in SMARTS (SMiles ARbitrary Target Specification) notation.
[0065] Furthermore, since the product ion spectrum values of the adduct will differ depending on the nucleophile (derivative reagent) used, a machine learning model will be generated according to the nucleophile used. Therefore, if multiple types of nucleophiles are used in the adduct, the CPU 11 inputs information indicating the type of nucleophile (derivative reagent) used in the adduct, in addition to the information indicating the substructure, in step 44.
[0066] Figure 6 shows the mass spectrum obtained by analyzing the adduct of N-acetylcysteine and propyl gallate. As shown in the figure, the difference (peak interval) between the m / z value of the compound shown in Figure (A) and the m / z value of the compound shown in Figure (B) indicates the mass-to-charge ratio (60 m / z) of the propanol unit derived from propyl gallate as a substructure included in the compound information database 31. In this case, the CPU 11 assigns label 1 indicating the presence of the propanol unit.
[0067] Next, the CPU 11 determines whether the number of compounds possessing the target substructure is a predetermined number (step 45). This is because it is considered difficult to create the machine learning model 10 if the number of compounds possessing the target substructure is too small. The predetermined number is, for example, 25 (50 or more including the compounds possessing the target substructure), but is not limited to this.
[0068] If the CPU determines that the number of compounds having the target substructure is not equal to or less than a predetermined number (No. in step 45), the CPU 11 adds more compounds to the compound information database 31 (step 46), returns to step 41, and repeats the subsequent processing.
[0069] If the CPU determines that the number of compounds possessing the target substructure is greater than or equal to a predetermined number (Yes in step 45), the CPU 11 adjusts the unbalanced data (step 47). That is, the CPU 11 corrects the class imbalance using methods such as undersampling to equalize the number of data points for labels 0 and 1. Other methods for adjusting unbalanced data may include underbagging and oversampling.
[0070] Next, the CPU 11 constructs a classification model using the above-mentioned product ion value (peak position) and neutral loss value (peak interval) as explanatory variables, and the presence or absence of a substructure in the target compound as the objective variable (step 48).
[0071] Examples of classification models that can be used include logistic regression, KNN (K Nearest Neighbor), SVM (Support-Vector Machine), decision trees, random forests, XGBoost, GBoost, LightGBM, CatBoost, neural networks, LDA, QDA, ridge classification, perceptrons, SGD, and AdaBoost. The CPU 11 may create these classification models using a predetermined library and compare them.
[0072] Next, the CPU 11 calculates evaluation metrics to assess the performance of the classification model constructed above (step 49). Possible evaluation metrics include Accuracy, AUC, Recall, Precision, F1 Score, Kappa, and MCC.
[0073] Next, the CPU 11 determines whether the evaluation index is equal to or greater than a predetermined value (step 50). When the F1 Score is used as the evaluation index, the predetermined value is, for example, 0.8, but is not limited to this.
[0074] If the CPU determines that the above evaluation indicator is equal to or greater than a predetermined value (Yes in step 50), the CPU 11 determines the classification model having that evaluation indicator as the machine learning model 10 (step 51).
[0075] On the other hand, if the CPU determines that the evaluation index is less than a predetermined value (No. in step 50), the CPU 11 returns to step 48 and repeats the subsequent processing until the evaluation index becomes equal to or greater than the predetermined value.
[0076] In this embodiment, by inputting the peak positions and peak intervals of the mass spectral information shown in Figure 6 into the machine learning model 10, it was possible to predict the presence of propanol units with a high score.
[0077] (Sample structure information generation process) Next, the sample structure information generation process using the machine learning model 10 generated by the above process will be explained. Figure 5 is a flowchart showing the flow of the sample structure information generation process by the sample structure analysis server 100.
[0078] As shown in the figure, first the CPU 11 inputs mass spectral information, which is an actual measured value obtained by mass analysis performed by the mass spectrometer 400 on the sample provided by the user terminal 200 along with the request for structural analysis of the sample, from the mass spectral information database 32 (step 61).
[0079] Next, the CPU 11 calculates the peak position (product ion value) from the input mass spectrum information, similar to the case in Figure 4 (step 62).
[0080] Next, the CPU 11 calculates the peak interval (neutral loss) from the mass spectral information, similar to the case in Figure 4 (step 63).
[0081] Next, the CPU 11 inputs substructure information (step 64). This substructure information is received from the user terminal 200 as substructure selection information along with the sample structure analysis request. The substructure selection information may select one substructure or multiple substructures. As mentioned above, along with this substructure information, information indicating the type of nucleophile used in the adduct is also input.
[0082] Next, the CPU 11 inputs the calculated product ion value and neutral loss value, as well as the input substructure information, into the machine learning model 10 (step 65).
[0083] Next, the CPU 11 receives from the machine learning model 10 a label information (0 or 1) indicating the presence or absence of a substructure and a score information indicating the reliability (probability) of the prediction, as a result of the machine learning model 10's prediction of the presence or absence of a substructure in the sample (step 66). From this prediction result, the CPU 11 generates structural information regarding the structure of the sample (step 67).
[0084] The CPU 11 then transmits the structural information to the user terminal 200 for display (step 68).
[0085] Figure 9 shows an example of a user interface that includes structural information displayed on the user terminal 200.
[0086] As shown in the figure, the user of the user terminal 200 selects, for example, in the library selection field 71 on the left side of the interface, the type of database of substructures from which the user's sample will be analyzed (for example, a publicly available database or a proprietary database of the sample structure analysis server 100), selects the substructures to be analyzed in the substructure selection field 72 below, and then presses the search button 73 below, thereby sending the sample structure analysis request along with the substructure selection information to the sample structure analysis server 100.
[0087] The sample structure analysis server 100 then performs a prediction using the machine learning model 10, and the result is displayed in the results display area 74 on the right side of the interface. Specifically, the CPU 11 displays a list of prediction results for each substructure selected by the user, including a label (0 or 1), which is a binary value indicating the presence or absence of a substructure, and a score indicating the reliability of the prediction (the likelihood of the substructure existing) on a value between 0 and 1. For example, a score of 0.85 indicates that there is an 85% chance that the substructure exists.
[0088] The CPU 11 may also generate and display a network diagram 75 showing the connections between substructures based on the similarity in the mass spectral information, for the substructures predicted to exist by the machine learning model 10.
[0089] Alternatively, or in addition to the above, the CPU 11 may, for the target sample, refer to the compound information database 31 based on the presence or absence of multiple substructures to generate the predicted overall structural formula of the compound and display it as structural information.
[0090] In this case, the CPU 11 may also use the precise mass of the target component obtained by HR-MS (high-resolution mass spectrometry), etc. That is, the CPU 11 calculates the composition of the compound contained in the target component from the precise mass and estimates its molecular formula, and based on the molecular formula and the substructure included in the prediction results by the machine learning model 10, it checks whether the overall structural formula of the compound (the estimated molecular formula) can be constructed from the substructure, and if it can be constructed, it may display the substructure and the overall structure.
[0091] Furthermore, the CPU 11 may store safety information or regulatory information corresponding to each of the multiple substructures, generate safety information or regulatory information corresponding to the substructures predicted by the machine learning model 10 (of which the substructures selected by the user) and display it in the additional information section 76 in the lower right of the figure.
[0092] Figure 10 shows an example of displaying other structural information. As shown in the figure, when analyzing the adduct of the nucleophile and sensitizing substance, as well as the raw material (main raw material) containing the sensitizing substance that constitutes the adduct, the analysis results of the adduct and the analysis results of the raw material can be displayed in combination.
[0093] This allows for more accurate detection of sensitizing substances in raw materials by tracking the decrease in components in the raw materials and the increase in components after the reaction through comparison of both analytical results. Furthermore, related substances to the sensitizing substance can also be identified. For example, in the figure, compound aa, which is observed in both analytical results at a point prior to the point in time when the reaction between the nucleophile and the sensitizing substance can be confirmed (rightmost point of the mass chromatogram in the figure), can be identified as the sensitizing substance.
[0094] As described above, according to this embodiment, the sample structure analysis server 100 can easily and quickly predict the presence or absence of substructures of compounds contained in a sample from the mass spectral information of an unknown sample.
[0095] [Modifications] Although embodiments of the present invention have been described above, the present invention is not limited to the embodiments described above, and various modifications can be made without departing from the spirit of the present invention.
[0096] In the above-described embodiment, the sample structure analysis server 100 generated a machine learning model 10 using both peak position and peak interval as explanatory variables. However, the machine learning model 10 may be generated using either the peak position or peak interval as an explanatory variable.
[0097] In particular, for substances with repeating units, such as surfactants, using the peak interval rather than the peak value as the explanatory variable allows for highly accurate determination of the presence or absence of substructures based on the unique characteristics that appear in the peak interval.
[0098] In other words, as shown in Figure 11, by using the peak interval as an explanatory variable, it is possible to analyze ions that are easily removed by fragmentation in MS / MS, i.e., substructures. Examples of substructures in which information appears in the peak interval include EO (ethylene oxide) / PO (propylene oxide) chains and alkyl groups.
[0099] In this case, as mentioned above, instead of using the peak intervals between all peak positions, noise can be removed during training by selecting major peak intensities (m / z) (e.g., the top 5) and using only their peak intervals. Furthermore, during this training process, it is possible to weight each peak value according to its fragment ion intensity.
[0100] In the above-described embodiment, the sample structure analysis server 100 generated structural information based on the prediction results from the machine learning model 10 regarding the presence or absence of a substructure selected by the user, and displayed it on the user terminal 200. In addition to this, if the sample structure analysis server 100 predicts that the selected substructure is absent, it may not only display the result but also send suggestion information to the user terminal 200 that proposes predictions for the presence or absence of other substructures, and generate structural information by predicting the presence or absence of those other substructures in response to an additional analysis request from the user terminal 200 regarding the suggestion information.
[0101] Alternatively, when the sample structure analysis server 100 receives a sample structure analysis request from the user terminal 200, it may predict the presence or absence of the substructure selected by the sample structure analysis server 100 without receiving the substructure selection information, and send the result as structural information to the user terminal 200.
[0102] In the above-described embodiment, only one sample structure analysis server 100 is shown, but the processing performed by the sample structure analysis server 100 may be distributed and executed across multiple servers. For example, the machine learning model 10 generation process and the process of predicting the presence or absence of substructures and generating structural information using the machine learning model 10 may be executed on separate servers.
[0103] In the embodiments described above, an example of the present invention being used in a skin sensitization test was shown, but the present invention can also be used in other human safety evaluation tests (for example, single-dose toxicity tests, repeated-dose toxicity tests, reproductive and developmental toxicity tests, etc.).
[0104] Of the inventions described in the claims of this application, the invention described as "information processing method" is one in which each step is performed automatically by at least one device such as a computer through information processing by software, and not by a human using a computer or other device. In other words, the "information processing method" is an information processing method using computer software, and not a method in which a human operates a computer as a calculating tool.
Claims
1. An information processing system comprising a control unit that obtains the peak intervals in a group of peaks contained in a mass spectrum obtained from a sample, uses the peak intervals as explanatory variables, and uses the presence or absence of a substructure of a compound in the sample as the objective variable, to generate a machine learning model that predicts the objective variable from the explanatory variables.
2. An information processing system comprising a control unit that inputs the peak interval obtained from a mass spectrum obtained from an unknown sample and selection information for selecting at least one substructure into a machine learning model trained to predict the presence or absence of a substructure of a compound in the sample, using the peak interval in a group of peaks included in a mass spectrum obtained from a sample as an explanatory variable, and the presence or absence of the selected substructure output from the machine learning model, and transmits structural information regarding the structure of the unknown sample to the user's terminal.
3. The information processing system according to claim 2, wherein the control unit receives the selection information from the user terminal and transmits label information indicating the presence or absence of the selected substructure or a score information indicating the predictive reliability of presence or absence as structural information.
4. The information processing system according to claim 2 or 3, wherein the control unit inputs the peak interval obtained from the mass spectrum obtained from the compound of the unknown sample and the derivatization reagent, which is the unknown sample, to the machine learning model.
5. The information processing system according to claim 2 or 3, wherein the control unit generates an overall structural formula predicted from the presence or absence of the predicted substructure as structural information.
6. The information processing system according to claim 2 or 3, wherein the control unit stores safety information or regulatory information corresponding to each of a plurality of substructures, and generates safety information or regulatory information corresponding to the substructures that are predicted to exist as structural information.
7. The information processing system according to claim 2 or 3, wherein the control unit inputs the selection information for a plurality of substructures to the trained model and generates a list of the presence or absence of the plurality of substructures as structural information.
8. The information processing system according to claim 3, wherein the control unit, when it predicts that there are no substructures selected by the user, transmits suggestion information to the user terminal that suggests the presence or absence of other substructures.
9. The information processing system according to claim 1, wherein the control unit further obtains the peak positions in the group of peaks included in the mass spectrum obtained from the sample, and generates a machine learning model that predicts the target variable from the explanatory variables, with the peak interval and the peak positions as explanatory variables and the presence or absence of a substructure of the compound in the sample as the target variable.
10. An information processing method comprising: obtaining the peak intervals in a group of peaks contained in a mass spectrum obtained from a sample; using the peak intervals as explanatory variables and the presence or absence of a substructure of a compound in the sample as the objective variable; and generating a machine learning model that predicts the objective variable from the explanatory variables.
11. An information processing method comprising: inputting the peak interval obtained from a mass spectrum obtained from an unknown sample and selection information for selecting at least one substructure into a machine learning model trained to predict the presence or absence of a substructure of a compound in the sample, using the peak interval in a group of peaks included in a mass spectrum obtained from a sample as an explanatory variable; generating structural information regarding the structure of the unknown sample based on the presence or absence of the selected substructure output from the machine learning model; and transmitting this information to the user's terminal.
12. A program that causes an information processing device to perform the following steps: acquire the peak intervals in a group of peaks contained in a mass spectrum obtained from a sample; and generate a machine learning model that predicts the target variable from the explanatory variable, with the peak intervals as explanatory variables and the presence or absence of a substructure of the compound in the sample as the target variable.
13. A program that causes an information processing device to execute the following steps: input the peak interval obtained from a mass spectrum obtained from an unknown sample and selection information for selecting at least one substructure to a machine learning model that has been trained to predict the presence or absence of a substructure of a compound in the sample, using the peak interval in a group of peaks included in a mass spectrum obtained from a sample as an explanatory variable and the presence or absence of a substructure of a compound in the sample as an objective variable; and generate structural information regarding the structure of the unknown sample based on the presence or absence of the selected substructure output from the machine learning model and transmit it to the user's terminal.