Method for analyzing unknown compound, method for analyzing unknown object, method for generating graph structure related to unknown object, system for analyzing unknown compound, and system for analyzing unknown object

The method and system create a graph structure from fragment spectrum similarities to efficiently predict compound properties and sample composition, addressing inefficiencies in existing mass spectrometry data analysis by utilizing graph analysis techniques.

WO2025206069A1PCT designated stage Publication Date: 2025-10-02OKINAWA INST OF SCI & TECH SCHOOL

Patent Information

Application Number
PCT/JP2025/012268
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-29
Filing Date
2025-03-26
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Current mass spectrometry data analysis methods struggle with inefficiencies in identifying and interpreting unknown compounds, lack comprehensiveness in standard sample spectral libraries, and inaccuracies in in silico fragmentation, while failing to consider interactions among multiple compounds in complex samples.

Method used

A method and system that generate a graph structure based on fragment spectrum similarities, allowing for efficient prediction of compound properties by organizing sample and reference nodes, performing similarity calculations, and utilizing graph analysis techniques like subgraph extraction and node vectorization.

Benefits of technology

Enables rapid, accurate prediction of compound properties and sample composition by leveraging graph structures to analyze unknown compounds and objects, facilitating comprehensive understanding beyond individual compound analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025012268_02102025_PF_FP_ABST
    Figure JP2025012268_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a method for analyzing an unknown compound, which enables efficient and accurate prediction of various properties of an unknown compound. In this method for analyzing an unknown compound: a fragment spectrum of the unknown compound is deemed a sample node; fragment spectra of one or more types of known compounds are deemed reference nodes; the degree of similarity between the fragment spectrum of the sample node and the fragment spectra of the reference nodes, and the degree of similarity between the fragment spectra of the reference nodes, if two or more types of known compounds are referenced, are deemed edges; a graph structure is generated from the sample node, the reference nodes, and the edges; and attribute information for the unknown compound is predicted on the basis of this graph structure.
Need to check novelty before this filing date? Find Prior Art

Description

Method for analyzing unknown compounds, method for analyzing unknown objects, method for generating graph structures relating to unknown objects, system for analyzing unknown compounds, and system for analyzing unknown objects

[0001] The present invention relates to a method for analyzing an unknown compound, a method for analyzing an unknown object, a method for generating a graph structure for an unknown object, a system for analyzing an unknown compound, and a system for analyzing an unknown object.

[0002] In recent years, mass spectrometry has played a central role in research and applications in many scientific and technological fields. Mass spectrometry is an analytical method with high sensitivity and throughput, capable of detecting thousands of compounds as mass peaks from an analytical sample in a short period of time. However, compounds detected by mass spectrometry are usually treated as unknown compounds, with unknown structures. For example, more than a thousand compound ions are detected from a serum sample by mass spectrometry. However, the majority of these are said to be unknown compounds with unknown structural information. Furthermore, analysis of data obtained by mass spectrometry is complex, and extracting useful information from the large amount of data is often difficult.

[0003] Conventional mass spectrometry data analysis requires analysts to manually identify and interpret specific peaks and patterns, which is time-consuming and can lead to misinterpretation or oversight. Against this background, there is a need for the development of new technologies that enable efficient and accurate analysis of mass spectrometry data.

[0004] Furthermore, with recent advances in research and changing demands in practical fields, there is an increasing need to identify unknown and unexpected compounds. To address this need, a new approach known as non-targeted analysis has been gaining attention. Non-targeted analysis aims to simultaneously analyze all components in a sample without targeting specific compounds.

[0005] Possible approaches for non-targeted analysis include methods that compare fragment spectra with those obtained from standard samples, methods based on in silico fragmentation that generate fragment spectra on a computer based on chemical structures, and methods that search for structural analogs based on the similarity of fragment spectra.

[0006] For example, Non-Patent Document 1 discloses a method for analyzing unknown compounds by generating a fragment spectrum in silico from a chemical structure and comparing it with an actually measured fragment spectrum.

[0007] Furthermore, for example, in Non-Patent Document 2, the distribution of structural analogs of small molecular weight natural products is expressed in a graph structure based on calculations of the similarity of fragment spectra. In addition to matching with the fragment spectra of standard samples, Non-Patent Document 2 also predicts the presence of unknown compounds that are structurally related to the standard samples.

[0008] Furthermore, for example, in Non-Patent Document 3, a graph structure is generated based on calculation of the similarity of fragment spectra of low molecular weight compounds, and the graph structure is used to assist in interpretation of the results of other analytical methods (in silico fragmentation).

[0009] Wolf S, Schmidt S, Muller-Hannemann M, Neumann S. In silico fragmentation for computer assisted identification of metabolite mass spectra. BMC Bioinformatics. 2010 Mar 22;11:148. doi: 10.1186 / 1471-2105-11-148. PMID: 20307295; PMCID: PMC2853470.Yang JY, Sanchez LM, Rath CM, Liu X, Boudreau PD, Bruns N, Glukhov E, Wodtke A, de Felicio R, Fenner A, Wong WR, Linington RG, Zhang L, Debonsi HM, Gerwick WH, Dorrestein PC. Molecular networking as a dereplication strategy. J Nat Prod. 2013 Sep 27;76(9):1686-99. doi: 10.1021 / np400413s.da Silva RR, Wang M, Nothias LF, van der Hooft JJJ, Caraballo-Rodriguez AM, et al. (2018), Propagating annotations of molecular networks using in silico fragmentation. PLOS Computational Biology 14(4): e1006089. https: / / doi.org / 10.1371 / journal.pcbi.1006089

[0010] However, the above-mentioned methods have several problems. First, the method of comparing fragment spectra with those obtained from standard samples has a problem of lacking comprehensiveness because the standard sample spectral library is limited in its repertoire. Furthermore, the method based on in silico fragmentation, such as that described in Non-Patent Document 1, has a problem of inaccuracy because the generated fragment spectra often differ significantly from the spectra obtained in actual analysis. Furthermore, the methods described in Non-Patent Documents 2 and 3 only use the similarity of fragment spectra to infer the relationship of "structural relatedness," and therefore cannot obtain specific chemical structures or other useful information.

[0011] Currently, a bottleneck in mass spectrometry-based analysis is the lack of technology to rapidly obtain the structure and related information of compounds. If high-throughput structural information analysis were possible, it would be a major breakthrough.

[0012] Furthermore, current mass spectrometry data analysis is primarily focused on individual compounds. This approach can be considered effective for detailed analysis of the properties and behavior of specific compounds. However, current mass spectrometry data analysis is biased toward compound-by-compound analysis and does not address the evaluation and analysis of the properties and behavior of the sample as a whole. For example, complex samples such as food, plants, and environmental samples contain numerous compounds, so it is essential to consider their interactions and overall effects. However, with current methods, it is difficult to comprehensively grasp information on the entire analytical sample, which is also a problem.

[0013] Therefore, an object of the present invention is to provide a method for analyzing an unknown compound that can efficiently and accurately predict the properties of the unknown compound. Another object of the present invention is to provide a method for analyzing an unknown object that can efficiently predict the properties of a sample that may contain multiple compounds. A further object of the present invention is to provide a method for generating a graph structure related to an unknown object that is suitable for use in the above-mentioned method for analyzing an unknown object. A further object of the present invention is to provide a system for analyzing an unknown compound that can efficiently and accurately predict the properties of the unknown compound. A further object of the present invention is to provide a system for analyzing an unknown object that can efficiently predict the properties of a sample that may contain multiple compounds.

[0014] The gist and configuration of the present invention are as follows.

[0015] [1] A method for analyzing an unknown compound, comprising: setting a fragment spectrum of the unknown compound as a sample node; setting fragment spectra of one or more known compounds as reference nodes; setting the similarity between the fragment spectrum of the sample node and the fragment spectrum of the reference node, and, when there are two or more known compounds, the similarity between the fragment spectra of the reference nodes as edges; generating a graph structure from these sample nodes, reference nodes, and edges; and predicting attribute information of the unknown compound based on the graph structure.

[0016] [2] The method according to [1], in the step of predicting attribute information of the unknown compound based on the graph structure, a local graph structure including the sample node and a nearby reference node connected to the sample node by an edge is extracted from the graph structure as a subgraph, attribute information of known compounds corresponding to the reference nodes in the subgraph is acquired, and the attribute information is aggregated and summarized, and the attribute information of the unknown compound is predicted from the results of the aggregation and summarization.

[0017] [3] The method according to [1], wherein in the step of predicting attribute information of the unknown compound based on the graph structure, the graph structure is vectorized, spatial information obtained by the vectorization is used to perform clustering, and the attribute information of the unknown compound is predicted from a result of the clustering.

[0018] [4] The method according to [1], wherein in the step of predicting attribute information of the unknown compound based on the graph structure, the graph structure is vectorized, and the attribute information of the unknown compound is predicted using a classifier trained using the vectors of the reference nodes obtained by the vectorization and attribute information of known compounds corresponding to the reference nodes as training data.

[0019] [5] A method for analyzing an unknown object, comprising: organizing fragment spectra of each compound contained in the unknown object into a spectrum group, and using the spectrum group as a sample node; organizing fragment spectra of each compound contained in one or more known objects into a spectrum group, and using the spectrum group as a reference node; using edges as the similarity between the spectrum group of the sample node and the spectrum group of the reference node, and, when there are two or more known objects, the similarity between the spectrum groups of the reference node; generating a graph structure from these sample nodes, reference nodes, and edges; and predicting attribute information of the unknown object based on the graph structure.

[0020] [6] The method according to [5], in the step of predicting attribute information of the unknown object based on the graph structure, a local graph structure including the sample node and a nearby reference node connected to the sample node by an edge is extracted from the graph structure as a subgraph, attribute information of known objects corresponding to the reference nodes in the subgraph is acquired, and the attribute information of the unknown object is aggregated and summarized, and the attribute information of the unknown object is predicted from the results of the aggregation and summarization.

[0021] [7] The method according to [5], in the step of predicting attribute information of the unknown object based on the graph structure, vectorization is performed on the graph structure, and the attribute information of the unknown object is predicted using a classifier trained using the vector of the reference node obtained by the vectorization and the attribute information of the known object corresponding to the reference node as training data.

[0022] [8] A method for generating a graph structure for an unknown object, comprising: organizing fragment spectra of each compound contained in the unknown object into a spectrum group, and using the spectrum group as a sample node; organizing fragment spectra of each compound contained in one or more known objects into a spectrum group, and using the spectrum group as a reference node; and using edges representing the similarity between the spectrum group of the sample node and the spectrum group of the reference node, and the similarity between the spectrum groups of the reference node themselves. The method is characterized by generating a graph structure from these sample nodes, reference nodes, and edges.

[0023] [9] A system for analyzing unknown compounds, the system being capable of accessing a database storing fragment spectra of two or more known compounds and the similarities between the fragment spectra, the system comprising: fragment spectrum acquisition means capable of acquiring the fragment spectrum of the unknown compound; graph structure generation means for generating a graph structure from the sample nodes, reference nodes, and edges, with the fragment spectrum of the unknown compound as a sample node, the fragment spectrum stored in the database as a reference node, and the similarity between the fragment spectrum of the sample node and the fragment spectrum of the reference node, and the similarity between the fragment spectra of the reference nodes as edges; prediction means for predicting attribute information of the unknown compound based on the graph structure; and output means for outputting the prediction result by the prediction means.

[0024]

[10] A system for analyzing an unknown object, the system being capable of accessing a database storing, for two or more types of known objects, spectral groups each formed by grouping together fragment spectra of each compound contained in each known object and similarities between the spectral groups, the system comprising: fragment spectrum acquisition means capable of acquiring fragment spectra of each compound contained in the unknown object; graph structure generation means for generating a graph structure from the sample nodes, reference nodes, and edges, with the spectral groups each formed by grouping together fragment spectra of each compound contained in the unknown object as sample nodes and the spectral groups stored in the database as reference nodes, and with the similarities between the spectral groups of the sample nodes and the spectral groups of the reference nodes, and the similarities between the spectral groups of the reference nodes as edges; prediction means for predicting attribute information of the unknown object based on the graph structure; and output means for outputting the prediction result by the prediction means.

[0025] According to the present invention, it is possible to provide a method for analyzing an unknown compound, which is capable of efficiently and accurately predicting the properties of the unknown compound. Also, according to the present invention, it is possible to provide a method for analyzing an unknown object, which is capable of efficiently predicting the properties of a sample that may contain a plurality of compounds. Furthermore, according to the present invention, it is possible to provide a method for generating a graph structure related to an unknown object, which is suitable for use in the above-mentioned method for analyzing an unknown object. Furthermore, according to the present invention, it is possible to provide a system for analyzing an unknown compound, which is capable of efficiently and accurately predicting the properties of the unknown compound. Furthermore, according to the present invention, it is possible to provide a system for analyzing an unknown object, which is capable of efficiently predicting the properties of a sample that may contain a plurality of compounds.

[0026] 1 is a schematic diagram showing how fragment spectra and related information are handled in a method for analyzing an unknown compound according to one embodiment of the present invention. It is a schematic flow diagram showing how the degree of similarity between fragment spectra of each compound is calculated and how information is generated based on the calculation when three types of compounds are targeted. It is a schematic flow diagram showing how a graph structure for analysis is generated from the fragment spectrum of an unknown compound and the fragment spectrum of a known compound. It is a schematic explanatory diagram showing how a subgraph including a sample node and its surrounding reference nodes is extracted from the graph structure. It is a schematic explanatory diagram showing how an unknown compound is analyzed using attribute information of reference nodes in the subgraph. It is a schematic flow diagram showing how vectorization is performed by node embedding in a graph structure in which the fragment spectra of compounds are nodes. It is a schematic diagram showing how spectrum groups and related information are handled in a method for analyzing an unknown object according to one embodiment of the present invention. It is a schematic flow diagram showing how the degree of similarity between spectrum groups of each object is calculated. It is a schematic flow diagram showing how a graph structure for analysis is generated from the spectrum group of an unknown object and the spectrum group of a known compound. It is a schematic explanatory diagram showing how a subgraph including a sample node and its surrounding reference nodes is extracted from the graph structure. The present invention relates to a method for analyzing an unknown object using attribute information of a reference node in a subgraph, and a method for vectorization by node embedding in a graph structure in which spectral groups of the object are nodes.

[0027] The present invention will be described below based on embodiments. However, such description is for the purpose of illustrating the present invention and does not limit the present invention in any way.

[0028] The present invention relates to a data analysis technique that uses spectral data that can be acquired by a data acquisition means such as a mass spectrometer to reveal various properties, such as the structure of an unknown compound to be analyzed, and the chemical composition and origin of a sample to be analyzed (unknown object) that may contain a mixture of multiple compounds.

[0029] The present invention constructs a graph structure based on the degree of spectral similarity, etc., for fragment spectra acquired by a data acquisition means such as a mass spectrometer, and performs various processes such as database creation and network calculations, which can be used to analyze unknown compounds or samples (target entities).

[0030] In the present invention, when analyzing an unknown compound contained in a sample (subject) to be analyzed, spectral data can be generated in advance using a standard substance or in silico, similarity calculations can be performed between the generated spectral data to generate a graph structure, and the generated graph structure can be stored in a database. Similarity calculations can be performed for the fragment spectrum of the unknown compound with the spectrum of a standard substance or in silico, and the similarity can be arranged on a graph structure already stored in a database. Nearby spectral data can be extracted based on their positional relationships on the network, and attribute information such as chemical structure, the type of sample (subject) from which the compound originated, and physiological activity can be extracted from the compound information related to the spectrum. Based on this attribute information, the properties of the unknown compound to be analyzed can be estimated.

[0031] The above-described approach according to the present invention has the following advantages over approaches that simply perform similarity searches on ordinary spectral libraries or spectral data. First, the graph structure allows for the simultaneous acquisition of a group of spectra that are highly similar, either directly or indirectly, to the spectrum being analyzed, and for this information to be used efficiently. Second, advanced analysis can be performed using graph analysis techniques such as subgraph extraction and node vectorization based on the network structure, rather than simply the similarity between spectra. Third, by comprehensively calculating similarities between spectra derived from standard samples or reference specimens in advance and creating a network structure, analysis can be performed significantly faster than by calculating similarities each time.

[0032] Furthermore, the technology of the present invention enables the identification, origin, and composition of a sample (object) containing a mixture of multiple compounds by grouping spectra according to their origin and sample information. Mass spectral data systematically acquired from various samples is grouped based on their origin and sample information. A brute-force similarity calculation is performed for each spectrum in the spectral bundle between these groups, generating a network based on the sample (object) itself. For samples (objects) with unknown origins, the obtained fragment spectra are calculated for each group with spectral groups derived from other objects to determine their position on the network, thereby revealing the sample's identity, origin, and composition. Furthermore, the fact that samples with similar chemical properties are more densely concentrated in a network of similar samples can be utilized to identify and predict the properties of the samples. For example, in a graph of samples derived from the same plant, spoiled samples form denser subnetworks, making it possible to determine the condition of the sample being analyzed from their position on the network.

[0033] In other words, while sharing the basic technique, the present invention can be used for various purposes, such as clarifying unknown compounds on a spectral basis, and clarifying unknown samples at a level beyond the compound level by treating spectra as a group.

[0034] (Fragment Spectrum) In this specification, the term "fragment spectrum" refers to a spectrum that indicates the mass and relative intensity of child ions generated from a parent ion when an ionized compound (ion) is fragmented (fragmentation) in mass analysis of the compound. This fragment spectrum can include information on the mass-to-charge ratio and a list of signal intensities (peak list) for each peak. The fragmented ions are called product ions.

[0035] The ionization method is not particularly limited, and examples thereof include electrospray ionization (ESI), electron impact ionization (EI), and the like.

[0036] The fragmentation method is not particularly limited, and examples thereof include high-energy collisions and photon absorption.

[0037] The fragment spectrum may be, for example, one obtained by a mass spectrometer in tandem mode, which refers to the use of two or more mass spectrometers in conjunction with one another, specifically, a mode in which a selected specific ion is fragmented by one mass spectrometer, and the mass of the fragmented ion is measured by a second mass spectrometer.

[0038] However, as fragment spectra, in addition to those acquired in tandem mode, when all spectra acquired relate to product ions (for example, in the case of data-independent acquisition, EI spectra in gas chromatography mass spectrometry, or all ion fragmentation), spectra generated by deconvoluting the data of all spectra for each compound can also be used.

[0039] Furthermore, the fragment spectrum may include spectral data derived from a mass spectrum database; spectral data acquired by a mass spectrometer; spectral data predicted from structural information using rule-based or machine learning methods; and fragment spectral data generated by quantum chemical processing.

[0040] Examples of the rule-based or machine learning methods include those described in (Djoumbou-Feunang Y, Pon A, Karu N, Zheng J, Li C, Arndt D, Gautam M, Allen F, and Wishart DS. (2019) Significantly Improved ESI-MS / MS Prediction and Compound Identification. Metabolites. 9(4):72.), and those described in (Wang F, Liigand J, Tian S, Arndt D, Greiner R, and Wishart D. (2021) CFM-ID 4.0: More Accurate ESI MS / MS Spectral Prediction and Compound Identification. Anal Chem. 93(34):11692-11700.). Examples of the quantum chemical treatment include those described in (Wang S, Kind T, Tantillo DJ, Fiehn O. Predicting in silico electron ionization mass spectra using quantum chemistry. J Cheminform. 2020 Oct 20;12(1):63. doi: 10.1186 / s13321-020-00470-3. PMID: 33372633; PMCID: PMC7576811.).

[0041] (Method for Analyzing Unknown Compounds) Next, a method for analyzing an unknown compound of the present invention will be described. The method for analyzing an unknown compound according to one embodiment of the present invention is characterized in that a fragment spectrum of the unknown compound is used as a sample node, fragment spectra of one or more known compounds are used as reference nodes, the similarity between the fragment spectrum of the sample node and the fragment spectrum of the reference node, and, when there are two or more known compounds, the similarity between the fragment spectra of the reference nodes are used as edges, a graph structure is generated from these sample nodes, reference nodes, and edges, and attribute information of the unknown compound is predicted based on the graph structure.

[0042] In this specification, an "unknown compound" as an analysis target refers to a compound for which at least some attribute information (for example, but not limited to, information such as chemical structure, compound class, physiological activity, or origin; also referred to as metadata) is unknown, and is not limited to compounds other than compounds for which the name or structure is known. In particular, an "unknown compound" also includes compounds for which the chemical structure is known but other attribute information is unknown, and the method of this embodiment can also be used to analyze such compounds. Specifically, the method of this embodiment can be suitably used to predict some attribute information other than chemical structure for compounds for which the chemical structure is known.

[0043] Known compounds include, but are not limited to, standards.

[0044] FIG. 1 is a schematic diagram showing how fragment spectra and related information are handled in this embodiment. First, in this embodiment, a fragment spectrum is acquired for an unknown compound to be analyzed by a data acquisition means such as mass spectrometry (P001). Meanwhile, a fragment spectrum is acquired for one or more known compounds by a data acquisition means such as mass spectrometry (P002). Note that only one known compound may be used, or two or more known compounds may be used in combination. Furthermore, when two or more known compounds are used, a fragment spectrum is acquired for each known compound. Furthermore, fragment spectra acquired in advance may be used as the fragment spectra of the known compounds.

[0045] In this specification, a fragment spectrum of an unknown compound may be referred to as a "sample spectrum," and a fragment spectrum of a known compound may be referred to as a "reference spectrum."

[0046] Furthermore, as shown in Figure 1, attribute information (e.g., but not limited to, information on chemical structure, compound class, physiological activity, origin, etc.) can be attached to the fragment spectrum (reference spectrum) of a known compound. The attribute information can be obtained from any information source, regardless of type, as long as it is structured so that compound identifiers are linked to information on the compound, such as a compound information database (e.g., Pubchem) or a toxicity information database (e.g., T3DB (https: / / www.t3db.ca / ) (Wishart D, Arndt D, Pon A, Sajed T, Guo AC, Djoumbou Y, Knox C, Wilson M, Liang Y, Grant J, Liu Y, Goldansaz SA, Rappaport SM. T3DB: the toxic exposome database. Nucleic Acids Res. 2015 Jan;43(Database issue):D928-34)). Furthermore, the attribute information to be attached to the reference spectrum is not limited to the databases described above. Element information extracted from articles (e.g., Wikipedia) that contain textual information about each compound can also be used. To cite a specific example, biological role information (e.g., sleep, thermoregulation) and / or history and etymology information (e.g., discoverer) can be collected from an article about "serotonin" and used as attribute information.

[0047] Next, in this embodiment, a graph structure for analysis is generated based on the sample spectrum and reference spectrum obtained above. In this regard, when two or more known compounds are used, the similarity between the respective reference spectra is calculated in advance. As an example of similarity calculation, FIG. 2 shows a general flow diagram for calculating the similarity between the fragment spectra of three compounds and generating information based on the calculation. As shown in the upper part of FIG. 2, the similarity can be calculated for each fragment spectrum of the target compound in a brute-force manner (P021).

[0048] Furthermore, when calculating the similarity between reference spectra in advance, as shown in the middle part of Figure 2, each fragment spectrum (reference spectrum) can be defined as a node, and the similarity between fragment spectra (reference spectra) can be defined as an edge, and a graph structure related to the reference spectra can be generated from these nodes and edges (P022). The generated graph structure and associated attribute information can be stored and saved in a table format as shown in the lower part of Figure 2, or in a graph information database (P023). In this way, by generating a graph structure related to the reference spectra and storing it in a database, the graph structure can be efficiently reused and managed. Furthermore, attribute information associated with the reference spectra can be stored as attribute information of the node or as another node derived from the reference spectrum. This allows a knowledge graph to be generated, and the knowledge graph can be stored, saved, and managed in a graph information database.

[0049] Methods for calculating the similarity between fragment spectra include, for example, (Frank AM, Bandeira N, Shen Z, Tanner S, Briggs SP, Smith RD, Pevzner PA. Clustering millions of tandem mass spectra. J Proteome Res. 2008 Jan;7(1):113-22. doi: 10.1021 / pr070361e. Epub 2007 Dec 8. PMID: 18067247; PMCID: PMC2533155.) (Huber, F., van der Burg, S., van der Hooft, JJJ et al. MS2DeepScore: a novel deep learning similarity measure to compare tandem mass spectra. J Cheminform 13, 84 (2021)). https: / / doi.org / 10.1186 / s13321-021-00558-4) However, in this embodiment, the present invention is not limited to a specific calculation method, and the above-mentioned spectral similarity calculation method or any method that can quantify the similarity between fragment spectra can be used.

[0050] FIG. 3 is a general flow diagram for generating a graph structure for analysis from fragment spectra of an unknown compound and fragment spectra of known compounds. First, referring also to FIG. 1 , the fragment spectrum acquired for the unknown compound to be analyzed is defined as a sample node (S1), and the fragment spectrum acquired for the known compound is defined as a reference node (R1). If two or more known compounds are used, the fragment spectra acquired for each known compound are defined as reference nodes (Rn) and listed. Then, the similarity between the fragment spectrum of the sample node (i.e., the sample spectrum) and the fragment spectrum of the reference node (i.e., the reference spectrum) is calculated (P041). The calculated similarity is defined as an edge (E1). If two or more known compounds are used, the similarity between each reference spectrum and the sample spectrum is calculated in a round-robin manner, and these similarities are defined as edges (En) and listed. Furthermore, if two or more known compounds are used, the similarity calculated between each reference spectrum is defined as an edge (RE1). When three or more known compounds are used, the similarities calculated in a round-robin fashion are defined as edges (REn) and listed. A graph structure is then generated from these sample nodes (S1), reference nodes (Rn), and edges (En, REn) (P042).

[0051] When two or more known compounds are used, for example, a first pre-graph structure is generated from the sample node (S1), the reference node (Rn), and the edge (En). On the other hand, when calculating the similarity (edge ​​(REn)) between the reference spectra, a graph structure (second pre-graph structure) related to the reference spectra is generated (corresponding to P022 in FIG. 2 ), and the second pre-graph structure is joined to the first pre-graph structure, thereby generating a graph structure. In this case, a graph structure reflecting all the similarities between the sample spectra and the reference spectra can be efficiently generated.

[0052] Next, the unknown compound can be analyzed based on the generated graph structure. In this embodiment, the above-described graph structure is used, so that the properties (attribute information) of the unknown compound can be predicted efficiently and accurately. Below, an example of a specific method for analyzing an unknown compound based on the graph structure will be described in detail.

[0053] A first method for analyzing unknown compounds based on a graph structure is to extract a local graph structure containing a sample node and a nearby reference node connected to the sample node by an edge from the graph structure as a subgraph, and use the subgraph. Figure 4 is an explanatory diagram outlining a method for extracting a subgraph containing a sample node (S1) and its surrounding reference nodes (Rn) from the graph structure. In Figure 4, the circled area corresponds to the subgraph.

[0054] For example, as shown in FIG. 4A, an egocentric network can be extracted as a subgraph by defining a sample node (S1) as a vertex and specifying the range of adjacent reference nodes (Rn) by the number of edges. Furthermore, as shown in FIG. 4B, community detection, which is frequently used in graph analysis, can be used to detect the community to which the sample node (S1) to be analyzed belongs and extract it as a subgraph. Furthermore, as shown in FIG. 4C, a clique in the graph structure, i.e., a set in which all nodes are directly connected to each other by edges, can be obtained and extracted as a subgraph. Furthermore, in this embodiment, any combination of the above-described subgraph extraction methods can be used.

[0055] The nodes in the subgraph extracted as described above have similar fragment spectra, and therefore it can be inferred that they are also similar in terms of attribute information (for example, but not limited to, information on chemical structure, compound class, physiological activity, origin, etc.) Based on this, after extracting a subgraph (local graph structure), attribute information of known compounds corresponding to reference nodes in the subgraph is obtained and tabulated and summarized, and the attribute information of the unknown compound can be predicted from the results of this tabulation and summarization.

[0056] Specifically, for example, referring to Figure 5, for known compounds (origins) corresponding to each reference node (Rn) in the subgraph, chemical structure information associated with the reference spectra is collected and the common portion of the chemical structures is calculated, thereby predicting the possible chemical structure of the unknown compound corresponding to the sample node (S1). Also, for example, referring to Figure 5, for known compounds (origins) corresponding to each reference node (Rn) in the subgraph, physiological activity information associated with the reference spectra is collected and statistics are taken, thereby predicting the possible physiological activity of the unknown compound corresponding to the sample node (S1). Also, for example, referring to Figure 5, for known compounds (origins) corresponding to each reference node (Rn) in the subgraph, origin information (i.e., information on the subject from which the compound originated) associated with the reference spectra is collected and statistics are taken, thereby predicting the type of subject in which the unknown compound corresponding to the sample node (S1) may exist. Furthermore, for the knowledge graph of the reference node stored in the graph information database, a large-scale language model (LLM) can be utilized to apply a graph embedding technique or a method typified by, but not limited to, GraphRAG (Graph Retrieval Augmented Generation) (Edge, Darren, David Gopstein, and Christopher Meek. 2024. “From Local to Global: A Graph RAG Approach to Query-Focused Summarization.” arXiv preprint, arXiv:2404.16130. https: / / doi.org / 10.48550 / arXiv.2404.16130.). This makes it possible to comprehensively summarize and / or infer the properties of unknown compounds by utilizing a variety of attribute information. Note that, in this embodiment, the type of attribute information predicted for unknown compounds is not particularly limited as long as it is attribute information that can be obtained from reference spectra.

[0057] A second method for analyzing unknown compounds based on graph structure involves vectorizing the graph structure and then performing clustering. Figure 6 shows an overview flow diagram of one such method. Using a vectorization technique such as node embedding, each spectrum can be represented as a low-dimensional vector in a feature matrix. This feature matrix is ​​then used to visualize and cluster the data.

[0058] The distance between nodes vectorized as described above can be interpreted as representing the relationship between the nodes (fragment spectra of compounds). Based on this, by performing clustering using the spatial information obtained by vectorization, spectra of nodes close to each other in the vector space can be interpreted as compounds that are highly related structurally or functionally. Based on this interpretation, attribute information of unknown compounds can be predicted from the clustering results.

[0059] Examples of vectorization techniques such as node embedding include node2vec (Aditya Grover and Jure Leskovec, "node2vec: Scalable Feature Learning for Networks," Proceedings of the 22nd ACM SIGKDD international conference on Knowledge discovery and data mining, 2016) and GraphSAGE (William L. Hamilton, Rex Ying, and Jure Leskovec, "Inductive Representation Learning on Large Graphs," In Proceedings of the Advances in Neural Information Processing Systems (NIPS), 2017). However, the present invention is not limited to these techniques, and any informatics technology that converts graph information into vectors and interprets them can be used.

[0060] Furthermore, the clustering method is not particularly limited, but examples that can be used include K-means (MacQueen, J. (1967). "Some Methods for Classification and Analysis of Multivariate Observations." In Proceedings of the 5th Berkeley Symposium on Mathematical Statistics and Probability. University of California Press.) and DBSCAN (Ester, M., Kriegel, H.-P., Sander, J., & Xu, X. (1996). "A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise." In Proceedings of the Second International Conference on Knowledge Discovery and Data Mining (KDD-96).). However, the clustering method is not limited to these, and any informatics technology that converts graph information into vectors and interprets them can be applied.

[0061] A third approach to analyzing unknown compounds based on graph structures involves vectorizing the graph structure and then using a trained classifier.

[0062] In this regard, graph-structured data may contain a huge number of reference spectrum nodes corresponding to known compounds whose attribute information, such as structure, is known. Therefore, a classifier is trained using the vectors of the reference nodes obtained by vectorization and the attribute information of the known compounds corresponding to the reference nodes as training data (machine learning in Figure 6). Using such a classifier, it is possible to predict the attribute information of unknown compounds.

[0063] A specific example of vectorization is the same as that described above.

[0064] The classifier (machine learning) is not particularly limited, and examples thereof include logistic regression, support vector machine (SVM), random forest, and neural network.

[0065] (Method for Analyzing an Unknown Object) In addition to the methods for analyzing a single compound as described above, the present disclosure also provides a method for analyzing an unknown object that may contain multiple compounds. A method for analyzing an unknown object according to one embodiment of the present invention includes: organizing fragment spectra of each compound contained in the unknown object into a spectrum group, using the spectrum group as a sample node; organizing fragment spectra of each compound contained in one or more known objects into a spectrum group, using the spectrum group as a reference node; using edges representing the similarity between the spectrum group of the sample node and the spectrum group of the reference node, and, when there are two or more known objects, the similarity between the spectrum groups of the reference node; generating a graph structure from these sample nodes, reference nodes, and edges; and predicting attribute information of the unknown object based on the graph structure.

[0066] As used herein, the term "object" refers to any tangible or intangible object in the universe that can be analyzed by extracting endogenous compounds and using a data acquisition method such as a mass spectrometer, including, for example, plants, animals, bacteria, tissues, food, environmental samples, and pharmaceuticals. As used herein, an "unknown object" as an analysis object refers to an object for which at least some attribute information (e.g., but not limited to, the object's name, taxonomy, category, origin, collection date and time, collection location, toxicity, allergy information, industry / business information used, health damage reports, physiological activity, summary text, etc.; also referred to as metadata) is unknown, and is not limited to objects other than those for which the name is known. Furthermore, an "unknown object" also includes an object for which it is unknown whether it contains multiple compounds.

[0067] Known objects are not particularly limited, but include known plant species such as "potato" and "tomato"; all biological species and their tissues, including invertebrates such as grasshoppers and vertebrates such as fish; samples obtained from the human body and living organisms, such as tears, urine, and feces; bacteria such as E. coli and yeast; environmental samples such as soil and river water from various locations; foods such as beverages, processed foods, and seasoning dishes; pharmaceuticals; and samples from industrial processes, all of which contain chemical substances (animal species, breeds, drug names, etc.), but are not limited to these.

[0068] FIG. 7 is a schematic diagram illustrating the handling of spectrum groups and related information in this embodiment. First, in this embodiment, a group of compounds inherent in an unknown target object to be analyzed is extracted (P201). Next, fragment spectra are acquired for each of the extracted compound groups using data acquisition means such as mass spectrometry (P202), and these are organized into a spectrum group. Meanwhile, a group of compounds inherent in one or more known targets is extracted (P203). Next, fragment spectra are acquired for each of the extracted compound groups using data acquisition means such as mass spectrometry (P204), and these are organized into a spectrum group. Note that only one type of known target object may be used, or two or more types may be used in combination. Furthermore, when two or more types of known targets are used, a group of compounds inherent in each known target object is extracted, and fragment spectra are acquired for each known target object. Furthermore, fragment spectra acquired in advance may be used as the fragment spectra of compounds inherent in the known targets.

[0069] That is, the fragment spectra of compounds contained in the target object are grouped together for each target object to which they belong, regardless of whether the compound is known or unknown. Fragment spectra obtained from compounds extracted from actual known targets, for example, fragment spectra obtained from a tomato extract, are grouped together as spectra belonging to the target object "tomato." Furthermore, even if a compound has not actually been extracted from a target object and data acquisition has not been performed, as long as the target object to which the compound originates (belongs) is known, the fragment spectra of the compound can be grouped together as spectra belonging to the target object. For example, for compounds inherent in a target object such as "potato," fragment spectra of standard compounds (known compounds) can be obtained from a mass spectral library based on a database showing the relationship between plant species or objects and the compounds inherent therein (e.g., KNApSAcK (Farit Mochamad Afendi, Taketo Okada, Mami Yamazaki, Aki Hirai-Morita, Yukiko Nakamura, Kensuke Nakamura, Shun Ikeda, Hiroki Takahashi, Md. Altaf-Ul-Amin, Latifah K. Darusman, Kazuki Saito, Shigehiko Kanaya, KNApSAcK Family Databases: Integrated Metabolite-Plant Species Databases for Multifaceted Plant Research, Plant and Cell Physiology, Volume 53, Issue 2, February 2012, Page e1, https: / / doi.org / 10.1093 / pcp / pcr165)). These can then be organized into groups.Furthermore, even if fragment spectra of the inherent compound do not exist, in silico fragment spectrum generation algorithms (e.g., Wang F, Liigand J, Tian S, Arndt D, Greiner R, and Wishart D. (2021) CFM-ID 4.0: More Accurate ESI MS / MS Spectral Prediction and Compound Identification. Anal Chem. 93(34):11692-11700.) can be used to generate fragment spectra of compounds that are hypothetically assigned to the target compound and group them together.

[0070] As shown in FIG. 7 , the spectrum group of a known object can be associated with attribute information (e.g., but not limited to, the object's name, taxonomy, category, origin, collection date and time, collection location, toxicity, health damage reports, physiological activity, summary text, etc.). The attribute information can be any type of structured information source related to the object, such as a structured database (e.g., FooDB (https: / / foodb.ca / )). Furthermore, entities extracted from text articles (e.g., Wikipedia) containing information about the object can also be used as attribute information, along with relevance.

[0071] Next, in this embodiment, the spectrum groups obtained above are compared and the similarity is quantified. Fig. 8 is a flow chart showing an example of this process.

[0072] 8 , first, assume a spectrum group X derived from a certain object. Spectrum group X contains a plurality of fragment spectra, and the mth spectrum (index m) among them is denoted as Xm. For example, the 100th spectrum (index 100) is expressed as X100. Also, assume a spectrum group Y derived from another certain object, and using the same concept as spectrum group X, the nth spectrum (index n) in spectrum group Y is denoted as Yn.

[0073] When quantifying the relationship between spectrum group X and spectrum group Y, the spectrum Yn that has the highest spectral similarity to Xm is found from group Y, and the similarity is expressed as S XmYn Furthermore, the signal intensity I normalized to 1 within each group is xm and I yn The similarity of the signal intensity is quantified between 0 and 1 using the following formula: xmyn = e-k|log(Ixm)-log(Iyn)|. In this calculation formula, the parameter k affects how strongly the difference in intensity is reflected, and can be adjusted based on the results of networking, which will be described later. The score indicating the similarity between spectrum group X and spectrum group Y is calculated by I xmyn ×S XmYn It should be noted that the above calculation method is just an example, and any calculation method that reflects the spectral similarity and the signal intensity of the standardized peak can be used instead.

[0074] This score can be calculated for all or a specific number of spectra in a spectrum group, and the average, median, or other calculated value can be used as the similarity (score) of spectrum group X to spectrum group Y (P221). This method is a process of grouping spectra together and networking information that scores the similarities of spectrum groups between groups. However, as long as the purpose can be achieved, the exact calculation method is not limited to the process described here.

[0075] In this embodiment, when two or more types of known objects are used, it is preferable to calculate the similarity between the spectrum groups of each known object in a round-robin manner in advance and generate a graph structure related to the spectrum groups of the known objects using this similarity. The generated graph structure and associated attribute information can be stored in table format or in a graph information database. In this way, by generating a graph structure related to the spectrum groups of known objects and storing it in a database, the graph structure can be efficiently reused and managed.

[0076] Next, in this embodiment, a graph structure for analysis is generated. FIG. 9 is a general flow diagram for generating a graph structure for analysis from a spectrum group of an unknown object and a spectrum group of a known compound. First, referring also to FIG. 7 , the spectrum group acquired for the unknown object to be analyzed is defined as a sample node (SO1), and the spectrum group acquired for the known object is defined as a reference node (RO1). If two or more known objects are used, the spectrum groups acquired for each known object are defined as reference nodes (ROn) and listed. Then, the similarity between the spectrum group of the sample node and the spectrum group of the reference node is calculated (P241). The calculated similarity is defined as an edge (EO1). If two or more known objects are used, the similarity between each spectrum group and the spectrum group of the sample node is calculated in a round-robin manner, and these similarities are defined as edges (EOn) and listed. Furthermore, when two or more types of known objects are used, the similarity calculated between each spectrum group is defined as an edge (REO1). When three or more types of known objects are used, the similarity calculated in a round-robin manner is defined as an edge (REOn) and listed. Then, a graph structure is generated from these sample nodes (SO1), reference nodes (ROOn), and edges (EOn, REOn) (P242).

[0077] In addition, when two or more types of known objects are used, for example, a first pre-graph structure is generated from the sample node (SO1), reference node (RO1), and edge (EOn), and when calculating the similarity (edge ​​(REOn)) between the spectrum groups of the known objects, a graph structure (second pre-graph structure) for the spectrum group of the known object can be generated in advance and the second pre-graph structure can be combined with the first pre-graph structure to generate a graph structure. In this case, a graph structure that reflects all the similarities between the spectrum group of the unknown object and the spectrum group of the known object can be efficiently generated.

[0078] Next, the unknown object can be analyzed based on the generated graph structure. In this embodiment, the above-described graph structure is used, so that the properties (attribute information) of a sample that may contain multiple compounds can be efficiently predicted. Below, an example of a specific method for analyzing an unknown object based on a graph structure will be described in detail.

[0079] A first method for analyzing an unknown object based on a graph structure is to extract a local graph structure including a sample node and a nearby reference node connected to the sample node by an edge from the graph structure as a subgraph, and use the subgraph. Figure 10 is an explanatory diagram outlining a method for extracting a subgraph including a sample node (SO1) and its surrounding reference nodes (ROn) from the graph structure. In Figure 10, the circled area corresponds to the subgraph.

[0080] For example, as shown schematically in FIG. 10A, an egocentric network can be extracted as a subgraph by defining a sample node (SO1) as a vertex and specifying the range of adjacent reference nodes (ROn) by the number of edges. Furthermore, as shown schematically in FIG. 10B, community detection, which is frequently used in graph analysis, can be used to detect the community to which the sample node (SO1) to be analyzed belongs and extract it as a subgraph. Furthermore, as shown schematically in FIG. 10C, a clique in the graph structure, i.e., a set in which all nodes are directly connected to each other by edges, can be obtained and extracted as a subgraph. Furthermore, in this embodiment, any combination of the above-described subgraph extraction methods can be used.

[0081] The nodes in the subgraph extracted as described above are likely to have similar spectral groups derived from compounds inherent in the target object, and therefore likely to contain highly related attribute information (e.g., but not limited to, chemical composition and physiological activity derived therefrom). Based on this, after extracting a subgraph (local graph structure), attribute information of known targets corresponding to reference nodes in the subgraph can be acquired and aggregated and summarized, and the attribute information of the unknown target can be predicted from the results of such aggregation and summarization. Furthermore, the attribute information associated with the nodes in the subgraph can be used as a knowledge graph, and techniques such as, but not limited to, large-scale language models (LLMs), graph embedding techniques, or GraphRAG (similar to those described above) can be applied. This makes it possible to comprehensively summarize and / or infer the properties of unknown targets by utilizing a variety of attribute information.

[0082] Specifically, referring to FIG. 11 , for known objects corresponding to (or originating from) each reference node (Rn) in the subgraph, attribute information such as the type (name), origin, and taxonomy (classification system) associated with the spectral group of the known object can be collected to predict the identity and origin of the unknown object. For example, if a food product contains a mixture of multiple plant extracts, the extract can be analyzed to create a sample node. By analyzing the connection with the reference node, the plant itself or a closely related plant contained in the food product can be predicted from the information in the reference node. For example, even if the plant species contained in the unknown object being analyzed is not identical to that of the reference node, the connection with closely related species nodes containing many structurally similar metabolites can reveal the presence of closely related species. In these respects, the present technology has advantages over existing technologies that attempt identification solely by comparing mass information itself, such as fingerprinting methods using mass spectrometry data. Furthermore, this technology makes it possible to predict various characteristics by utilizing graph structure and attribute information, such as predicting the possible physiological activity of the unknown subject being analyzed by collecting physiological activity information of known subjects corresponding to nearby reference nodes in the subgraph and taking statistics.

[0083] Furthermore, even when multiple objects share similar spectra of endogenous compounds, slight differences in the compound repertoire and differences in the signal intensity of normalized compound peaks are reflected in the scores (as weights) of the edges between nodes. By utilizing this property, even when multiple nodes share similar basic characteristics and chemical compositions, those with closer chemical compositions form a denser community-like structure. For example, even when multiple nodes originate from the same food sample, a local graph structure is obtained that reflects changes in signal intensity related to the presence or absence of contaminants or changes in compound concentration. This property enables rapid anomaly detection and identification of food samples, for example. While typical targeted analysis involves comparing signals of known compounds to detect anomalies, this technology utilizes the vast number of unknown compound signals obtained through non-targeted analysis to detect anomalies that reflect potential variations in chemical composition.

[0084] A second method for analyzing unknown objects based on graph structures involves vectorizing the graph structure and then performing clustering. Figure 12 shows a general flow diagram of one such method. In Figure 12, a vectorization technique such as node embedding can be used to represent each spectrum as a low-dimensional vector in a feature matrix. This feature matrix is ​​then used to visualize and cluster the data.

[0085] This vectorization allows for efficient visualization and clustering of the network of objects, reflecting the similarity of the underlying chemical groups. The distance between object nodes in the resulting vector space (including, but not limited to, Euclidean distance and cosine distance) can be used as an index of the chemical compositional relationships and similarities between objects. Based on this, clustering can be performed using the spatial information obtained through vectorization to identify groups of objects that are highly related based on chemical, physiological, and biological properties. Furthermore, this distribution can be used to predict the identity of unknown objects being analyzed.

[0086] Furthermore, by clustering, it is possible to group the unknown objects to be analyzed into groups of similar objects to which they belong, using clustering algorithms such as, but not limited to, k-means, hierarchical clustering, and DBSCAN.

[0087] The specific examples of vectorization and clustering are the same as those described above.

[0088] A third method for analyzing an unknown object based on a graph structure involves vectorizing the graph structure and using a trained classifier.

[0089] In this regard, graph-structured data may contain many spectral group nodes corresponding to known objects whose identities are known. Therefore, a classifier is trained using the vectors of reference nodes obtained by vectorization and attribute information of the known objects corresponding to the reference nodes (e.g., but not limited to, information on taxonomy, origin, and physiological activity) as training data (machine learning in FIG. 12). Using such a classifier, the identity of an unknown object can be clarified.

[0090] Specific examples of vectorization and classifiers (machine learning) are the same as those described above.

[0091] (Method for generating a graph structure for an unknown object) Next, a method for generating a graph structure for an unknown object of the present invention will be described. The method for generating a graph structure for an unknown object according to one embodiment of the present invention is characterized in that it comprises: organizing fragment spectra of each compound contained in the unknown object into a spectrum group, using the spectrum group as a sample node; organizing fragment spectra of each compound contained in one or more known objects into a spectrum group, using the spectrum group as a reference node; using edges as similarities between the spectrum group of the sample node and the spectrum group of the reference node, and as similarities between the spectrum groups of the reference node; and generating a graph structure from these sample nodes, reference nodes, and edges.

[0092] The graph structure generated by the method of this embodiment can be suitably used in the method for analyzing the unknown object described above.

[0093] The details of the method for generating a graph structure in this embodiment are the same as those already described for the method for analyzing an unknown object, and therefore, the details already described for the method for analyzing an unknown object will be used.

[0094] (System for Analyzing Unknown Compounds) Next, a system for analyzing unknown compounds will be described. A system for analyzing unknown compounds according to one embodiment of the present invention is configured to be able to access a database in which, for two or more known compounds, fragment spectra of each known compound and similarities between the fragment spectra of the known compounds are stored. The system is characterized by comprising: fragment spectrum acquisition means capable of acquiring fragment spectra of unknown compounds; graph structure generation means for generating a graph structure from the sample nodes, reference nodes, and edges, with the fragment spectra of the unknown compounds as sample nodes and the fragment spectra (fragment spectra of each known compound) stored in the database as reference nodes, and with the similarities between the fragment spectra of the sample nodes and the fragment spectra of the reference nodes, and the similarities between the fragment spectra of the reference nodes as edges; prediction means for predicting attribute information of the unknown compound based on the graph structure; and output means for outputting the prediction result by the prediction means.

[0095] This system allows the above-described method of analyzing unknown compounds to be carried out systematically. Furthermore, this system allows the analyzed unknown compounds to be treated as known compounds and the data to be stored and accumulated in a database. This allows the system to be used continuously and in an expansive manner.

[0096] The database stores fragment spectra of two or more known compounds and the similarities between the fragment spectra of the known compounds. Such a database may store both the fragment spectra of the known compounds and the similarities between the fragment spectra of the known compounds. Alternatively, the database may be composed of a first database storing fragment spectra of two or more known compounds and a second database storing the similarities between the fragment spectra of the known compounds.

[0097] Such a database may be installed in the system of this embodiment, or may be installed in a server computer connected to the system via a communication network, and such a server computer may be located either domestically or overseas.

[0098] Such a database may be a privately created database or may be a database that is open to the public.

[0099] The fragment spectrum acquired by the fragment spectrum acquisition means is not particularly limited and may be in various formats. The fragment spectrum acquisition means may also acquire the fragment spectrum via a communication network. For example, the fragment spectrum acquisition means may acquire the fragment spectrum via a communication network from a measurement device that measures the fragment spectrum, such as a mass spectrometer. The fragment spectrum acquisition means may also acquire the fragment spectrum via a computer-readable medium, etc.

[0100] The graph structure generation means may be, for example, a computer terminal equipped with a CPU, a GPU, and / or a memory. The graph structure generation means may also be a terminal operated by a user. Furthermore, if the database is installed in a server computer connected via a communication network, the graph structure generation means may also be the server computer. Furthermore, the graph structure generation means may also be a computer program itself for causing a computer or the like to execute a graph structure generation process.

[0101] The graph structure generated by the graph structure generating means can be stored in the database (the database storing data on known compounds) or can be stored in a separate database.

[0102] The prediction means may be, for example, a computer terminal equipped with a CPU, a GPU, and / or a memory. The prediction means may also be a terminal operated by a user. Furthermore, if the database is installed in a server computer connected via a communication network, the prediction means may also be the server computer. Furthermore, the prediction means may also be a computer program itself for causing a computer or the like to execute a prediction process.

[0103] The graph structure generating means and the predicting means may reside on the same terminal or on separate terminals.

[0104] The output means is not particularly limited, and may be a CRT, a printer, or the like, as long as it displays information including compound attribute information in a format that can be recognized by the user.

[0105] (System for Analyzing Unknown Objects) Next, a system for analyzing unknown objects will be described. A system for analyzing unknown objects according to one embodiment of the present invention is configured to be able to access a database that stores, for two or more known objects, spectrum groups each comprising a collection of fragment spectra of each compound contained in each known object, and similarities between the spectrum groups. The system is characterized by comprising: fragment spectrum acquisition means capable of acquiring fragment spectra of each compound contained in the unknown object; graph structure generation means that generates a graph structure from the sample nodes, reference nodes, and edges, using the spectrum groups each comprising the collection of fragment spectra of each compound contained in the unknown object as sample nodes and the spectrum groups (spectrum groups each comprising the fragment spectra of each compound contained in each known object) stored in the database as reference nodes, and using the similarities between the spectrum groups of the sample nodes and the spectrum groups of the reference nodes, and the similarities between the spectrum groups of the reference nodes as edges; prediction means that predicts attribute information of the unknown object based on the graph structure; and output means that outputs the prediction result by the prediction means.

[0106] This system allows the method of analyzing unknown objects described above to be carried out systematically. Furthermore, this system allows the analyzed unknown objects to be treated as known objects and the data to be stored and accumulated in a database. This allows the system to be used continuously and in an expansive manner.

[0107] The database stores spectral groups (hereinafter referred to as "spectrum groups related to known targets") each consisting of a group of fragment spectra of each compound contained in each known target, and the similarities between the spectral groups. Such a database may store both the spectral groups related to known targets and the similarities between the spectral groups. Alternatively, the database may be composed of a first database storing the spectral groups related to known targets and a second database storing the similarities between the spectral groups.

[0108] Such a database may be installed in the system of this embodiment, or may be installed in a server computer connected to the system via a communication network, and such a server computer may be located either domestically or overseas.

[0109] Such a database may be a privately created database or may be a database that is open to the public.

[0110] The fragment spectrum acquired by the fragment spectrum acquisition means is not particularly limited and may be in various formats. The fragment spectrum acquisition means may also acquire the fragment spectrum via a communication network. For example, the fragment spectrum acquisition means may acquire the fragment spectrum via a communication network from a measurement device that measures the fragment spectrum, such as a mass spectrometer. The fragment spectrum acquisition means may also acquire the fragment spectrum via a computer-readable medium, etc.

[0111] The graph structure generation means may be, for example, a computer terminal equipped with a CPU, a GPU, and / or a memory. The graph structure generation means may also be a terminal operated by a user. Furthermore, if the database is installed in a server computer connected via a communication network, the graph structure generation means may also be the server computer. Furthermore, the graph structure generation means may also be a computer program itself for causing a computer or the like to execute a graph structure generation process.

[0112] The graph structure generated by the graph structure generating means can be stored in the database (the database in which data relating to known objects is stored) or can be stored in a separate database.

[0113] The prediction means may be, for example, a computer terminal equipped with a CPU, a GPU, and / or a memory. The prediction means may also be a terminal operated by a user. Furthermore, if the database is installed in a server computer connected via a communication network, the prediction means may also be the server computer. Furthermore, the prediction means may also be a computer program itself for causing a computer or the like to execute a prediction process.

[0114] The graph structure generating means and the predicting means may reside on the same terminal or on separate terminals.

[0115] The output means is not particularly limited, and may be a CRT, a printer, or the like, as long as it displays information including attribute information of the target object in a format that can be recognized by the user.

[0116] According to the present invention, it is possible to provide a method for analyzing an unknown compound, which is capable of efficiently and accurately predicting the properties of the unknown compound. Also, according to the present invention, it is possible to provide a method for analyzing an unknown object, which is capable of efficiently predicting the properties of a sample that may contain a plurality of compounds. Furthermore, according to the present invention, it is possible to provide a method for generating a graph structure related to an unknown object, which is suitable for use in the above-mentioned method for analyzing an unknown object. Furthermore, according to the present invention, it is possible to provide a system for analyzing an unknown compound, which is capable of efficiently and accurately predicting the properties of the unknown compound. Furthermore, according to the present invention, it is possible to provide a system for analyzing an unknown object, which is capable of efficiently predicting the properties of a sample that may contain a plurality of compounds.

Claims

1. A method for analyzing an unknown compound, comprising: setting a fragment spectrum of the unknown compound as a sample node; setting fragment spectra of one or more known compounds as reference nodes; setting the similarity between the fragment spectrum of the sample node and the fragment spectrum of the reference node, and, when there are two or more known compounds, the similarity between the fragment spectra of the reference nodes as edges; generating a graph structure from these sample nodes, reference nodes, and edges; and predicting attribute information of the unknown compound based on the graph structure.

2. The method according to claim 1, wherein the step of predicting attribute information of the unknown compound based on the graph structure comprises extracting, from the graph structure, a local graph structure including the sample node and nearby reference nodes connected to the sample node by edges, as a subgraph, obtaining attribute information of known compounds corresponding to the reference nodes in the subgraph, aggregating and summarizing the information, and predicting the attribute information of the unknown compound from the results of the aggregation and summarization.

3. The method according to claim 1, wherein in the step of predicting attribute information of the unknown compound based on the graph structure, the graph structure is vectorized, spatial information obtained by the vectorization is used to perform clustering, and attribute information of the unknown compound is predicted from the results of the clustering.

4. The method according to claim 1, wherein in the step of predicting attribute information of the unknown compound based on the graph structure, the graph structure is vectorized, and the attribute information of the unknown compound is predicted using a classifier trained using the vectors of the reference nodes obtained by the vectorization and attribute information of known compounds corresponding to the reference nodes as training data.

5. A method for analyzing an unknown object, comprising: organizing the fragment spectra of each compound contained in the unknown object into a spectrum group, and using the spectrum group as a sample node; organizing the fragment spectra of each compound contained in one or more known objects into a spectrum group, and using the spectrum group as a reference node; using edges as the similarity between the spectrum group of the sample node and the spectrum group of the reference node, and, when there are two or more known objects, the similarity between the spectrum groups of the reference node; generating a graph structure from these sample nodes, reference nodes, and edges; and predicting attribute information of the unknown object based on the graph structure.

6. The method according to claim 5, wherein in the step of predicting attribute information of the unknown object based on the graph structure, a local graph structure including the sample node and a nearby reference node connected to the sample node by an edge is extracted as a subgraph from the graph structure, attribute information of known objects corresponding to the reference nodes in the subgraph is obtained and aggregated and summarized, and the attribute information of the unknown object is predicted from the results of such aggregation and summarization.

7. The method according to claim 5, wherein in the step of predicting attribute information of the unknown object based on the graph structure, the graph structure is vectorized, and the attribute information of the unknown object is predicted using a classifier trained using the vector of the reference node obtained by the vectorization and the attribute information of the known object corresponding to the reference node as training data.

8. A method for generating a graph structure for an unknown object, comprising: organizing fragment spectra of each compound contained in the unknown object into a spectrum group, and using the spectrum group as a sample node; organizing fragment spectra of each compound contained in one or more known objects into a spectrum group, and using the spectrum group as a reference node; and using edges representing the similarity between the spectrum group of the sample node and the spectrum group of the reference node, and the similarity between the spectrum groups of the reference nodes themselves. The method is characterized by generating a graph structure from these sample nodes, reference nodes, and edges.

9. A system for analyzing unknown compounds, the system being capable of accessing a database storing fragment spectra of two or more known compounds and the similarities between the fragment spectra, the system comprising: a fragment spectrum acquisition means capable of acquiring the fragment spectrum of the unknown compound; a graph structure generation means for generating a graph structure from the sample nodes, reference nodes, and edges, with the fragment spectra of the unknown compounds as sample nodes and the fragment spectra stored in the database as reference nodes, and with the similarities between the fragment spectra of the sample nodes and the fragment spectra of the reference nodes, and the similarities between the fragment spectra of the reference nodes as edges; a prediction means for predicting attribute information of the unknown compound based on the graph structure; and an output means for outputting the prediction results by the prediction means.

10. A system for analyzing an unknown object, the system being capable of accessing a database storing, for two or more types of known objects, spectral groups each consisting of a collection of fragment spectra of each compound contained in each known object and similarities between the spectral groups, the system comprising: fragment spectrum acquisition means capable of acquiring the fragment spectrum of each compound contained in the unknown object; graph structure generation means for generating a graph structure from the sample nodes, reference nodes, and edges, with the spectral groups each consisting of a collection of fragment spectra of each compound contained in the unknown object as sample nodes and the spectral groups stored in the database as reference nodes, and with the similarities between the spectral groups of the sample nodes and the spectral groups of the reference nodes, and the similarities between the spectral groups of the reference nodes as edges; prediction means for predicting attribute information of the unknown object based on the graph structure; and output means for outputting the prediction results by the prediction means.

Citation Information

Patent Citations

  • A method for identifying unknown substances in particular using mass spectrometry.

    JP2012515902A

Cited By

  • Circuit design method, and a system and program for said method.

    JP7886505B1