Method and apparatus for determining chemical structure of compound, and terminal device
By exploring the structure of compounds in the chemical space and determining the chemical structure of compounds using mass spectrometry and molecular formula information, the problem of unincluded compounds in the prior art is solved, and a wider analysis coverage and more comprehensive identification scheme are achieved.
Patent Information
- Application Number
- PCT/CN2024/135679
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-07
- Filing Date
- 2024-11-29
- Publication Date
- 2025-06-12
AI Technical Summary
The prior art cannot implement structural analysis of compounds not included in the database, resulting in a small analysis coverage.
By determining its molecular formula and adjacent mass spectrum based on the mass spectrum of the compound to be determined, the candidate compound is further determined based on the molecular formula, and the chemical structure of the compound is determined in the chemical space based on the spatial distance between the candidate compound and the adjacent compound.
The coverage of compound analysis has been broadened, and the problem that the prior art cannot perform structural analysis when dealing with unknown compounds is solved, providing a more comprehensive solution for the identification of unknown compounds.
Smart Images

Figure CN2024135679_12062025_PF_FP_ABST
Abstract
Description
Method, device and terminal equipment for determining chemical structure of compound
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 7, 2023, with application number 202311690470.0 and invention name “Method, device and terminal equipment for determining the chemical structure of a compound”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the technical field of compound analysis, and in particular to a method, apparatus, and terminal device for determining the chemical structure of a compound. Background Art
[0003] Mass spectrometry plays a crucial role in analyzing the diverse content of small molecules in complex systems, especially in biological complex systems such as metabolomics. In the qualitative analysis of small molecule compounds, traditional qualitative methods primarily match the mass spectrum of the compound to be analyzed with mass spectra in a standard mass spectrometry database to obtain relevant data and determine the structural characteristics of the compound to be analyzed. However, this method cannot achieve structural elucidation of compounds not included in the database, resulting in a limited analytical coverage issue that needs to be addressed urgently. Technical issues
[0004] The purpose of the embodiments of the present application is to provide a method, apparatus and terminal device for determining the chemical structure of a compound, aiming to solve the problem in the prior art that it is impossible to perform structural analysis on compounds not included in the database and that the analysis coverage is small. Technical Solutions
[0005] The purpose of this application is to provide a method, apparatus and terminal device for determining the chemical structure of a compound, aiming to solve the problem in the prior art that it is impossible to perform structural analysis on compounds not included in the database and that the analysis coverage is small.
[0006] A first aspect of the embodiments of the present application provides a method for determining the chemical structure of a compound, comprising:
[0007] Determining a molecular formula of the compound to be determined and at least one adjacent mass spectrum according to the mass spectrum of the compound to be determined, wherein each adjacent mass spectrum corresponds to a adjacent compound;
[0008] determining at least one candidate compound according to the molecular formula;
[0009] The chemical structure of the compound to be determined is determined according to the spatial distance between each candidate compound and each adjacent compound in the chemical space.
[0010] In some embodiments, before determining the molecular formula of the compound to be determined and at least one adjacent mass spectrum according to the mass spectrum of the compound to be determined, the method further comprises:
[0011] Convert each mass spectrum in the mass spectrum database to be matched into corresponding natural language;
[0012] Inputting the natural language corresponding to each mass spectrum into a data conversion model, and outputting the mass spectrum numerical vector corresponding to each mass spectrum from the data conversion model;
[0013] The mass spectrum numerical vector corresponding to each mass spectrum graph is saved in a numerical matrix.
[0014] In some embodiments, determining at least one adjacent mass spectrum of the compound to be identified based on the mass spectrum of the compound to be identified comprises:
[0015] Converting the mass spectrum of the compound to be determined into natural language and inputting it into the data conversion model, so that the data conversion model outputs a mass spectrum numerical vector corresponding to the compound to be determined;
[0016] In the numerical matrix, screening mass spectrum numerical vectors whose vector distance with the mass spectrum numerical vector corresponding to the compound to be determined meets a preset threshold, and setting at least one mass spectrum numerical vector obtained by screening as an adjacent mass spectrum numerical vector;
[0017] Determine the adjacent mass spectrum corresponding to each adjacent mass spectrum numerical vector; each adjacent mass spectrum corresponds to a adjacent compound.
[0018] In some embodiments, determining the molecular formula of the compound to be determined based on the mass spectrum of the compound to be determined comprises:
[0019] Determining the relative molecular mass of the compound to be determined according to the mass spectrum of the compound to be determined;
[0020] The molecular formula of the compound to be determined is determined according to the relative molecular mass of the compound to be determined.
[0021] In some embodiments, determining at least one candidate compound according to the molecular formula comprises:
[0022] According to the molecular formula, screening a database of structures of compounds to be matched to obtain at least one compound to be matched that has the same molecular formula as the molecular formula;
[0023] At least one of the compounds to be matched is determined as a candidate compound.
[0024] In some embodiments, before determining the chemical structure of the compound to be determined based on the spatial distance between each candidate compound and each adjacent compound in the chemical space, the method further includes:
[0025] Obtaining the chemical structure of each candidate compound and each adjacent compound;
[0026] Converting the chemical structures of each candidate compound and each adjacent compound into one-to-one corresponding structure numerical vectors;
[0027] The plurality of structure numerical vectors are projected into the chemical space respectively, and each structure numerical vector corresponds to a point in the chemical space.
[0028] In some embodiments, determining the chemical structure of the compound to be determined based on the spatial distance between each candidate compound and each adjacent compound in the chemical space comprises:
[0029] Calculating the average distance between the point corresponding to the candidate compound and each point corresponding to the adjacent compound;
[0030] Repeat the above calculation process to obtain the average distance between the points corresponding to all the candidate compounds and each point corresponding to the adjacent compounds;
[0031] All calculated average distances are sorted in order from small to large, and the candidate compound corresponding to the average distance at the first position is selected as the compound to be determined, and the chemical structure corresponding to the candidate compound is the chemical structure of the compound to be determined.
[0032] A second aspect of the embodiments of the present application provides a device for determining the chemical structure of a compound, comprising:
[0033] A first determination module is configured to determine a molecular formula of the compound to be determined and at least one adjacent mass spectrum according to the mass spectrum of the compound to be determined, wherein each adjacent mass spectrum corresponds to a nearby compound;
[0034] a second determination module, configured to determine at least one candidate compound according to the molecular formula;
[0035] The third determination module is configured to determine the chemical structure of the compound to be determined based on the spatial distance between each candidate compound and each adjacent compound in the chemical space.
[0036] A third aspect of an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the computer program.
[0037] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0038] A fifth aspect of the embodiments of the present application provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute the method described in the first aspect. Beneficial effects
[0039] In the technical solution adopted in the present application, in addition to determining the molecular formula of the compound to be determined and the adjacent compounds based on the mass spectrum of the compound to be determined, at least one candidate compound is further determined based on the molecular formula of the compound to be determined, and the chemical structure of the compound to be determined is determined by the spatial distance between each candidate compound and the adjacent compounds in the chemical space. Compared with the existing technology, the breakthrough of the method of the present application is that it not only relies on a pre-established mass spectrometry database, but also explores possible structures more comprehensively in the chemical space. This innovation not only broadens the coverage of compound analysis, but also solves the problem that the existing technology cannot perform structural analysis when dealing with unknown compounds, and provides a more comprehensive solution for the identification of unknown compounds. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0041] FIG1 is a flow chart of a method for determining the chemical structure of a compound provided in an embodiment of the present application;
[0042] FIG2 is a second flow chart of a method for determining the chemical structure of a compound provided in an embodiment of the present application;
[0043] FIG3 is a third flow chart of a method for determining the chemical structure of a compound provided in an embodiment of the present application;
[0044] FIG4 is a fourth flow chart of a method for determining the chemical structure of a compound provided in an embodiment of the present application;
[0045] FIG5 is a fifth flow chart of a method for determining the chemical structure of a compound provided in an embodiment of the present application;
[0046] FIG6 is a comparison chart of data results provided in the examples of the present application;
[0047] FIG7 is a schematic diagram of a device for determining the chemical structure of a compound provided in an embodiment of the present application;
[0048] FIG8 is a schematic structural diagram of a terminal device provided in an embodiment of the present application. Modes for Carrying Out the Invention
[0049] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0050] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0051] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0052] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0053] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0054] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0055] It should be understood that the size of the serial numbers of each step in this embodiment does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of this application.
[0056] Mass spectrometry plays a crucial role in analyzing the diverse content of small molecules in complex systems, especially in biological complex systems such as metabolomics. In the qualitative analysis of small molecule compounds, traditional qualitative methods primarily match the mass spectrum of the compound to be analyzed with mass spectra in a standard mass spectrometry database to obtain relevant data and determine the structural characteristics of the compound to be analyzed. However, this method cannot achieve structural elucidation of compounds not included in the database, resulting in a limited analytical coverage issue that needs to be addressed urgently.
[0057] In this regard, the present application provides a method for determining the chemical structure of a compound, which determines the molecular formula of the compound to be determined and at least one adjacent mass spectrum based on the mass spectrum of the compound to be determined, wherein each adjacent mass spectrum corresponds to a neighboring compound, and then determines at least one candidate compound based on the molecular formula. Thereafter, the chemical structure of the compound to be determined is determined based on the spatial distance between each candidate compound and each of the neighboring compounds in the chemical space.
[0058] In the technical solution adopted in the present application, in addition to determining the molecular formula of the compound to be determined and the adjacent compounds based on the mass spectrum of the compound to be determined, at least one candidate compound is further determined based on the molecular formula of the compound to be determined, and the chemical structure of the compound to be determined is determined by the spatial distance between each candidate compound and the adjacent compounds in the chemical space. This breaks through the limitation of the existing technology of only matching the compound to be determined in the mass spectrum database, broadens the analysis coverage of the compound to be determined, and solves the problem that the existing technology cannot complete structural analysis for the compound to be determined that is not in the database.
[0059] It should be noted that the mass spectrum is obtained through mass spectrometry analysis, which is an analytical method for measuring the mass-to-charge ratio (mass-to-charge ratio) of ions. Its basic principle is to ionize the components in the sample in the ion source to generate charged ions with different mass-to-charge ratios. After being accelerated by the electric field, an ion beam is formed and enters the mass analyzer. In the mass analyzer, the electric field and magnetic field are used to cause opposite velocity dispersion, and they are focused separately to obtain a mass spectrum, thereby determining their relative molecular mass. Mass spectrometry can also provide rich compound structure information in a single analysis.
[0060] In order to illustrate the technical solution of the present application, specific embodiments are provided below.
[0061] FIG1 is a flow chart of a method for determining the chemical structure of a compound provided in an embodiment of the present application.
[0062] 1 , a method for determining the chemical structure of a compound provided in an embodiment of the present application includes the following steps:
[0063] Step S101: determining the molecular formula of the compound to be determined and at least one adjacent mass spectrum according to the mass spectrum of the compound to be determined, wherein each adjacent mass spectrum corresponds to a adjacent compound;
[0064] Step S102: determining at least one candidate compound according to the molecular formula;
[0065] Step S103: determining the chemical structure of the compound to be determined based on the spatial distance between each candidate compound and each adjacent compound in the chemical space.
[0066] The above-mentioned steps of determining the compound by mass spectrum are not limited to matching known compounds in the mass spectrum database, but further determine the candidate compound by molecular formula information obtained by mass spectrum analysis. This new method takes into account the relative position of the candidate compound and the adjacent compound in the chemical space, and calculates the spatial distance between them to obtain the potential structure of the compound to be determined. Compared with the prior art, the breakthrough of the present application method is that it does not just rely on the pre-established mass spectrum database, but explores possible structures more comprehensively in the chemical space. This innovation not only broadens the coverage of compound analysis, but also solves the problem that the prior art cannot perform structural analysis when dealing with unknown compounds, and provides a more comprehensive solution for the identification of unknown compounds. Therefore, this method not only improves the analytical ability of unknown compounds, but also provides a new way for the accurate identification and analysis of compounds of unknown structure.
[0067] It should be noted that the unidentified compound in this application refers to a compound that is not included in the mass spectrometry database and may be an unknown compound, that is, a compound whose relevant chemical characteristics have not been identified, including a compound with an unknown chemical structure.
[0068] In order to efficiently and accurately match the neighboring compounds of the compound to be determined, referring to FIG. 2 , before performing step S101 , the chemical structure determination method in the embodiment of the present application further includes the following steps:
[0069] Step A1: converting each mass spectrum in the mass spectrum database to be matched into corresponding natural language;
[0070] Step A2: inputting the natural language corresponding to each mass spectrum into a data conversion model, and outputting the mass spectrum numerical vector corresponding to each mass spectrum from the data conversion model;
[0071] Step A3: Save the mass spectrum numerical vector corresponding to each mass spectrum graph into a numerical matrix.
[0072] In some embodiments, each mass spectrum peak in each mass spectrum in the mass spectrum database to be matched is represented as a word containing its position. In addition to all peaks in the mass spectrum, a certain range of neutral losses is also included, and the calculation of neutral losses is defined as the parent ion m / z minus the peak m / z. After the above processing, each mass spectrum can be represented as a natural language composed of several words, which can then be input into a data conversion model, which converts the natural language into a corresponding mass spectrum numerical vector, and then saves the converted mass spectrum numerical vector corresponding to each mass spectrum into a numerical matrix.
[0073] The mass spectrum numerical vector is a numerical vector converted from the natural language corresponding to the mass spectrum of the compound.
[0074] Exemplarily, for each mass spectrum in the database to be matched, each mass spectrum peak is represented as a word containing its position, accurate to a specified number of decimal places, and expressed as "peak@xxx.xx". In all the results presented, a grouping method including two decimal places is adopted, for example, the peak at m / z 200.445 is expressed as "peak@200.45". In addition to all the peaks in the mass spectrum, neutral losses ranging from 5.0Da to 200.0Da are also included, expressed as loss@xxx.xx. After processing, each mass spectrum can be represented as a natural language consisting of several words, which can then be converted into a numerical vector through a data conversion model and saved in a numerical matrix.
[0075] The above method converts each mass spectrum in the mass spectrum database to be matched into a mass spectrum numerical vector and stores it in a numerical matrix to complete the data preprocessing, which facilitates the subsequent calculation of the vector distance between the mass spectrum numerical vectors corresponding to the compound to be determined, so as to determine the adjacent mass spectra and obtain the adjacent compounds corresponding to the compound to be determined, thereby improving the analysis speed and accuracy.
[0076] It should be noted that, in the embodiments that can be implemented in this application, the mass spectrum database to be matched may include one or more databases, and specifically the following databases may be used:
[0077] NIST Mass Spectral Library: provides mass spectra and related information for a large number of compounds;
[0078] MassBank: a mass spectrometry database containing mass spectrometry data obtained from different laboratories;
[0079] METLIN: contains mass spectrometric data of metabolites in biological fluids;
[0080] GNPS (Global Natural Products Social Molecular Networking): provides mass spectrometry data, mainly about natural products;
[0081] MassIVE: stores large-scale mass spectrometry data, including raw data and analysis results;
[0082] Human Metabolome Database (HMDB): Contains detailed information on compounds related to human metabolism, including mass spectrometry data.
[0083] It should be noted that the above databases are only examples and are not intended to be limiting. Those skilled in the art can select other mass spectrum databases according to actual needs.
[0084] The data conversion model in the above embodiments of the present application can convert text words into numerical vectors, and can be specifically trained using the Word2vec model, GloVe model or Doc2Vec model.
[0085] The Word2Vec model is a two-layer neural network for processing text. Its input is a text corpus, and its output is a set of numeric vectors representing the numeric vectors of the words in the corpus. The purpose of the Word2Vec model is to group word vectors by similarity within a vector space and thereby identify mathematical similarities.
[0086] The input to the Word2Vec model is a sequence of words from a text corpus, each represented by a one-hot encoding. One-hot encoding is a vector representation method that assigns each word an N-dimensional vector, where N is the number of words in the vocabulary. Each word's vector has only one 1 and all other 0s, indicating that the word is unique.
[0087] The output of the Word2Vec model is a vocabulary in which each word has a corresponding vector. These vectors are obtained by training the input text corpus. They can be used to represent the contextual information, semantic information, and grammatical information of the word. The dimension of the output layer vector is usually N-dimensional, where N is the number of words in the vocabulary.
[0088] In one embodiment of the present application, the input data and output data are used to form training data to train the Word2Vec model. During the training process, the model continuously adjusts parameters through the back propagation algorithm to minimize the prediction error, and finally obtains the data conversion model of the present application.
[0089] The following is a detailed explanation of step S101.
[0090] 3 , to determine the neighboring compounds of the compound to be determined, step S101 may specifically include:
[0091] Step B1: converting the mass spectrum of the compound to be identified into natural language and inputting it into the data conversion model, and the data conversion model outputting a mass spectrum numerical vector corresponding to the compound to be identified;
[0092] Step B2: in the numerical matrix, screening mass spectrum numerical vectors whose vector distance with the mass spectrum numerical vector corresponding to the compound to be determined meets a preset threshold, and setting at least one mass spectrum numerical vector screened as a neighboring mass spectrum numerical vector;
[0093] Step B3: Determine the adjacent mass spectrum corresponding to each adjacent mass spectrum numerical vector; each adjacent mass spectrum corresponds to a nearby compound.
[0094] In some embodiments, certain data preprocessing has been performed in the aforementioned embodiments to convert each mass spectrum in the database of mass spectra to be matched into a mass spectrum numerical vector and store it in a numerical matrix. To facilitate subsequent calculations, the same process is performed on the mass spectrum of the compound to be identified. Specifically, the mass spectrum of the compound to be identified is converted into natural language and input into the data conversion model, which then outputs the mass spectrum numerical vector corresponding to the compound to be identified.
[0095] To find at least one mass spectrum vector in the numerical matrix whose vector distance with the mass spectrum vector corresponding to the compound to be identified meets a preset threshold, thereby screening for at least one corresponding adjacent compound, a vector distance calculation is first performed. Vector distances include, but are not limited to, Cosine distance, Mahalanobis distance, and Manhattan distance. After the calculation is completed, the results are sorted from smallest to largest, and at least one mass spectrum vector whose vector distance is within the preset threshold is selected and set as the adjacent mass spectrum vector.
[0096] It should be noted that each row in the numerical matrix corresponds to a mass spectrum in the mass spectrum database to be matched and also corresponds to a compound structure. After determining at least one adjacent mass spectrum numerical vector, the adjacent mass spectra corresponding to each adjacent mass spectrum numerical vector can be determined, and the chemical structure of the adjacent compound corresponding to each adjacent mass spectrum can be obtained. The chemical structure of each adjacent compound is considered to be similar to the chemical structure of the compound to be determined and can serve as a basis for subsequent compound analysis.
[0097] The above-mentioned vector distance calculation can screen out the neighboring compounds and corresponding chemical structures of the compound to be determined in the numerical matrix, which can improve the accuracy of the analysis and provide relevant basis for subsequent analysis.
[0098] 4 , to determine the molecular formula of the compound to be determined, step S101 may specifically include:
[0099] Step C1: determining the relative molecular mass of the compound to be determined based on the mass spectrum of the compound to be determined;
[0100] Step C2: Determine the molecular formula of the compound to be determined according to the relative molecular mass of the compound to be determined.
[0101] In some embodiments, relative molecular mass refers to the sum of the relative atomic masses (Ar) of the atoms in a chemical formula. To determine the relative molecular mass from a mass spectrum, it is necessary to locate the molecular ion peak (M+) in the mass spectrum of the compound to be determined. The molecular ion peak corresponds to a positively charged ion in the molecule, and its mass-to-charge ratio (m / z) is close to the relative molecular mass of the molecule. By measuring the mass-to-charge ratio of the molecular ion peak, the relative molecular mass of the compound to be determined is obtained.
[0102] To further determine the molecular formula of the compound to be determined, it is necessary to calculate the mass fraction of each element based on the relative atomic mass and atomic number of each element in the molecule, and then obtain the molecular formula based on the mass fraction and relative molecular mass. The specific method is: multiply the relative molecular mass by the mass fraction of an element, and then divide it by the relative atomic mass of the element to obtain the number of atoms of the element. Then, use the same method to calculate the number of atoms of other elements, and finally write the molecular formula.
[0103] The molecular formula of the compound to be determined obtained by the above method can be used to subsequently obtain candidate compounds to broaden the analysis coverage of the compound, so that the analysis scope of the compound to be determined is not limited to the compounds in each mass spectrum database, further improving the accuracy of the analysis.
[0104] The following is a detailed explanation of step S102.
[0105] To determine the candidate compound, step S102 may specifically include:
[0106] According to the molecular formula, at least one compound to be matched that has the same molecular formula as the molecular formula is screened in a database of compound structures to be matched; and the at least one compound to be matched is determined as a candidate compound.
[0107] The database of compound structures to be matched is a resource used to store and retrieve information about the structure, properties, and activity of compounds. Within the database, users can find detailed structural information for various compounds, including molecular formula, molecular weight, chemical bonds, stereo configuration, and electron distribution. Furthermore, users can use keywords such as molecular formula to search and filter the chemical structure information of compounds with the same molecular formula within the database.
[0108] In an achievable embodiment of the present application, the database of compound structures to be matched may include one or more databases, and specifically the following databases may be used:
[0109] HMDB (Human Metabolome Database): A database dedicated to human metabolites, providing detailed information on metabolites under physiological and pathological conditions;
[0110] KEGG (Kyoto Encyclopedia of Genes and Genomes): includes information on biology, biochemistry, and systems biology, and provides databases on biomolecules, metabolic pathways, and drugs.
[0111] ChEBI (Chemical Entities of Biological Interest): contains biologically important compounds;
[0112] PubChem: a public compound database containing a large amount of information on biologically active small molecules;
[0113] ChemSpider: a compound database containing a large number of organic, inorganic and biologically active compounds;
[0114] DrugBank: A database for drugs and drug targets, providing information on drug structure, mechanism of action, metabolic pathways, etc.
[0115] PDB (Protein Data Bank): contains a database of protein and nucleic acid structures, as well as structural information of small molecules and drugs that bind to them;
[0116] METLIN: Metabolomics database, containing mass spectrometry data of metabolites;
[0117] SwissBioisostere: provides a database of bioequivalents to help find compounds with similar biological activities;
[0118] ZINC (ZINC Is Not Commercial): A compound structure database mainly used for virtual screening and drug design.
[0119] It should be noted that the above databases are only examples and are not intended to be limiting. Those skilled in the art may select other compound structure databases according to actual needs.
[0120] In order to concretely reflect the spatial distance distribution between each candidate compound and each adjacent compound, and to judge the structural similarity between the candidate compound and the adjacent compounds, the present application can project each candidate compound and each adjacent compound into the chemical space for subsequent analysis.
[0121] In some embodiments, the chemical space of a compound can be described by a molecular fingerprint. A molecular fingerprint is a coding method for representing molecular structures, usually in binary form, characterized by a numerical vector that can describe the chemical structure of a molecule, and is used to compare and analyze the structural similarity between molecules in chemistry and drug discovery. Molecular fingerprints can capture the structural information in a molecule and map it to a binary numerical vector of a fixed length. From a mathematical point of view, when a molecule is represented as a numerical vector, it can be regarded as a point in a high-dimensional space, and the distance function between vectors can be used to describe the distance between these points. This representation method helps to process and compare molecular structures in computers, and is widely used in, for example, virtual screening, drug design, and chemical information retrieval. Molecular fingerprints include but are not limited to Morgan fingerprints, MACCS fingerprints, PubChem fingerprints, etc.
[0122] Correspondingly, in an embodiment that can be implemented in the present application, the chemical structure of each candidate compound and each adjacent compound is first obtained, and the chemical structure of each candidate compound and each adjacent compound is converted into a one-to-one corresponding structure numerical vector through molecular fingerprint calculation, and then the multiple structure numerical vectors are projected into the chemical space respectively, and each structure numerical vector corresponds to a point in the chemical space.
[0123] The structure numerical vector is a numerical vector converted according to the chemical structure of the compound.
[0124] The above conversion of the chemical structure of each candidate compound and each adjacent compound into a structural numerical vector and using it as a corresponding point in the chemical space helps to analyze the structural similarity between the candidate compound and the adjacent compounds and reduces the complexity of the analysis.
[0125] The following is a detailed explanation of step S103.
[0126] After converting each candidate compound and each adjacent compound into a structural numerical vector, it is necessary to further calculate the spatial distance between each candidate compound and each adjacent compound in the chemical space. Correspondingly, referring to FIG5 , step S103 may specifically include:
[0127] Step D1: Calculating the average distance between the point corresponding to the candidate compound and each point corresponding to the adjacent compound;
[0128] Step D2: Repeat the above calculation process to obtain the average distance between the points corresponding to all candidate compounds and each point corresponding to the adjacent compounds;
[0129] Step D3: sorting all the calculated average distances in ascending order, selecting the candidate compound corresponding to the first average distance as the compound to be determined, and the chemical structure corresponding to the candidate compound as the chemical structure of the compound to be determined.
[0130] In some embodiments, for a candidate compound, the vector distance between the structural numerical vector of the candidate compound and the structural numerical vector of each adjacent compound is calculated, and the vector distance includes but is not limited to Cosine distance, Mahalanobis distance, Manhattan distance, etc., and the average distance is further calculated; the above calculation process is repeated to obtain the average distance corresponding to all the candidate compounds; based on the calculated average distance, the compounds are sorted in order from small to large, and the higher the ranking, the more likely it is the true structure of the compound to be determined. The candidate compound corresponding to the first average distance is selected as the compound to be determined, and the chemical structure corresponding to the candidate compound is the chemical structure of the compound to be determined.
[0131] The above method obtains the candidate compound with the smallest average distance to each adjacent compound through calculation as the compound to be determined, and thereby determines the chemical structure of the compound to be determined. It concretely provides how to accurately determine the chemical structure of the compound to be determined after broadening the analysis coverage, and reduces the need for manual intervention and interpretation, thereby improving the speed and accuracy of compound analysis.
[0132] The following describes the implementation methods of this application in detail with reference to specific implementation data.
[0133] 600,289 different positive ion mode mass spectra from 43,653 different compounds from data sources such as GNPS were converted into corresponding natural language. The natural language was input into the Word2Vec model to train a data conversion model. This model was then used to convert all mass spectra into 256-length numeric vectors, resulting in a 600,289*256 matrix that was saved. The mass spectrum of the unidentified compound, N-Acetyl-5-aminosalicylic acid (its structure is shown in Figure 6A), was treated as an unknown mass spectrum and similarly converted into a 256-length numeric vector using the data conversion model. The top 10 mass spectra with the greatest cosine similarity to this vector were retrieved from the saved matrix. These are called neighboring mass spectra, and the corresponding compounds are called neighboring compounds (Top 10 Neigbors). It should be noted that the unidentified compound, N-Acetyl-5-aminosalicylic acid, is not included in the data included in the mass spectrum database, so this search is unlikely to find a compound and mass spectrum that are exactly the same.
[0134] Based on the compound's molecular formula, 34 candidate compounds were retrieved from compound structure databases such as HMDB and KEGG, each corresponding to a chemical structure. Molecular fingerprints were calculated for the top 10 neighboring compounds and all candidate compounds retrieved, projected into chemical space, and then projected into two-dimensional space using the Unified Mapping (UMAP) dimensionality reduction method for visualization, as shown in Figure 6B. The 34 potential candidate compounds were ranked based on the average distance between each candidate compound and each neighboring compound, with the candidate with the closest average distance ranked first. This candidate compound is the candidate to be determined. As can be seen in Figure 6B, the candidate compound represented by the dot has the closest distance to each neighboring compound, and the candidate compound corresponding to this dot (True Annotation) is determined as the candidate to be determined. Figures 6C and 6D, respectively, show the structures and corresponding adjacent mass spectra of two of these neighboring compounds, demonstrating their similar chemical structure to N-acetyl-5-aminosalicylic acid. The two adjacent mass spectra show significant correlation, providing a basis for compound analysis. This demonstrates that the chemical structure determination method of the present application has high accuracy.
[0135] FIG7 is a schematic structural diagram of a device for determining the chemical structure of a compound provided in an embodiment of the present application. For ease of explanation, only the portion related to the embodiment of the present application is shown.
[0136] The chemical structure determination device of a compound may specifically include the following modules:
[0137] A first determination module M1 is configured to determine a molecular formula of a compound to be determined and at least one adjacent mass spectrum according to a mass spectrum of the compound to be determined, wherein each adjacent mass spectrum corresponds to a nearby compound;
[0138] A second determination module M2, configured to determine at least one candidate compound according to the molecular formula;
[0139] The third determination module M3 is configured to determine the chemical structure of the compound to be determined based on the spatial distance between each candidate compound and each adjacent compound in the chemical space.
[0140] The first determination module M1 determines the adjacent compounds and molecular formula of the compound to be determined based on the mass spectrum of the compound to be determined, and then the second determination module M2 determines at least one candidate compound based on the molecular formula of the compound to be determined. Thereafter, the third determination module calculates the spatial distance between each candidate compound and each adjacent compound in the chemical space to determine the chemical structure of the compound to be determined. The above-mentioned operation adopted in this application not only relies on the pre-established mass spectrum database, but also explores possible structures more comprehensively in the chemical space. This innovation not only broadens the coverage of compound analysis, but also solves the problem that the existing technology cannot perform structural analysis when dealing with unknown compounds, and provides a more comprehensive solution for the identification of unknown compounds.
[0141] FIG8 is a schematic diagram of the structure of a terminal device provided in an embodiment of the present application. The terminal device E1 includes: at least one processor E2 (only one is shown in FIG8 ), a memory E3, and a computer program E4 stored in the memory E3 and executable on the at least one processor E2. When the processor E2 executes the computer program E4, it implements the steps of the above-mentioned method for determining the chemical structure of a compound.
[0142] The terminal device E1 can be a computing device such as a desktop computer, laptop, PDA, or cloud server. The terminal device may include, but is not limited to, a processor E2 and a memory E3. Those skilled in the art will appreciate that FIG8 is merely an example of the terminal device E1 and does not limit the terminal device E1. The terminal device E1 may include more or fewer components than shown, or may combine certain components or different components. For example, it may also include input / output devices, network access devices, etc.
[0143] The processor E2 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.
[0144] In some embodiments, the memory E3 may be an internal storage unit of the terminal device E1, such as a hard drive or memory of the terminal device E1. In other embodiments, the memory E3 may also be an external storage device of the terminal device E1, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the terminal device E1. Furthermore, the memory E3 may include both an internal storage unit of the terminal device E1 and an external storage device. The memory E3 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory E3 may also be used to temporarily store data that has been output or is about to be output.
[0145] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0146] In the above embodiments, the description of each embodiment has different emphases. If a rated part is not described or recorded in detail in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0147] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0148] An embodiment of the present application provides a computer program product. When the computer program product is run on a mobile terminal, the mobile terminal can implement the steps in the above-mentioned various method embodiments when executing the computer program product.
[0149] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0150] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0151] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0152] In addition, the functional units in the various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above units may be implemented in the form of hardware or software.
[0153] If the integrated module / unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium can be appropriately increased or decreased based on the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0154] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A method for determining the chemical structure of a compound, characterized in that: The chemical structure determination method comprises: Determine the molecular formula of the compound to be determined and at least one adjacent mass spectrum according to the mass spectrum of the compound to be determined, each adjacent mass spectrum corresponding to an adjacent compound; Determine at least one candidate compound according to the molecular formula; The chemical structure of the compound to be determined is determined according to the spatial distance between each of the candidate compounds and each of the adjacent compounds in the chemical space.
2. The method for determining a chemical structure according to claim 1, characterized in that: Before determining the molecular formula of the compound to be determined and at least one adjacent mass spectrum according to the mass spectrum of the compound to be determined, the chemical structure determination method further comprises: Convert each mass spectrum in the mass spectrum database to be matched into a corresponding natural language; Inputting the natural language corresponding to each of the mass spectra into a data conversion model, and outputting the mass spectrum numerical vector corresponding to each of the mass spectra by the data conversion model; The mass spectrum numerical vector corresponding to each mass spectrum graph is saved in a numerical matrix.
3. The method for determining a chemical structure according to claim 2, characterized in that: Determining at least one adjacent mass spectrum of the compound to be determined according to the mass spectrum of the compound to be determined comprises: Converting the mass spectrum of the compound to be determined into natural language and inputting it into the data conversion model, so that the data conversion model outputs a mass spectrum numerical vector corresponding to the compound to be determined; In the numerical matrix, a mass spectrum numerical vector whose vector distance with the mass spectrum numerical vector corresponding to the compound to be determined meets a preset threshold is screened, and at least one mass spectrum numerical vector screened is set as an adjacent mass spectrum numerical vector; Determine the adjacent mass spectrum corresponding to each adjacent mass spectrum numerical vector; each adjacent mass spectrum corresponds to a adjacent compound.
4. The method for determining a chemical structure according to claim 1, characterized in that: Determining the molecular formula of the compound to be determined according to the mass spectrum of the compound to be determined comprises: Determining the relative molecular mass of the compound to be determined according to the mass spectrum of the compound to be determined; The molecular formula of the compound to be determined is determined according to the relative molecular mass of the compound to be determined.
5. The method for determining a chemical structure according to claim 1, characterized in that: Determining at least one candidate compound according to the molecular formula comprises: According to the molecular formula, screening in a database of structures of compounds to be matched to obtain at least one compound to be matched that has the same molecular formula as the molecular formula; At least one of the compounds to be matched is determined as a candidate compound.
6. The method for determining a chemical structure according to claim 1, characterized in that: Before determining the chemical structure of the compound to be determined according to the spatial distance between each candidate compound and each adjacent compound in the chemical space, the chemical structure determination method further includes: Obtaining the chemical structure of each of the candidate compounds and each of the adjacent compounds; Convert the chemical structures of each candidate compound and each adjacent compound into one-to-one corresponding structure numerical vectors; The plurality of structure numerical vectors are projected into the chemical space respectively, and each structure numerical vector corresponds to a point in the chemical space.
7. The method for determining a chemical structure according to claim 1, characterized in that: Determining the chemical structure of the compound to be determined according to the spatial distance between each candidate compound and each adjacent compound in the chemical space includes: Calculating the average distance between the point corresponding to the candidate compound and each point corresponding to the adjacent compound; Repeat the above calculation process to obtain the average distance between the points corresponding to all the candidate compounds and each point corresponding to the adjacent compounds; All calculated average distances are sorted in order from small to large, and the candidate compound corresponding to the average distance at the first position is selected as the compound to be determined, and the chemical structure corresponding to the candidate compound is the chemical structure of the compound to be determined.
8. A device for determining the chemical structure of a compound, characterized in that: The chemical structure determination device comprises: A first determination module is used to determine the molecular formula of the compound to be determined and at least one adjacent mass spectrum according to the mass spectrum of the compound to be determined, each adjacent mass spectrum corresponding to an adjacent compound; A second determination module, used to determine at least one candidate compound according to the molecular formula; The third determination module is used to determine the chemical structure of the compound to be determined according to the spatial distance between each candidate compound and each adjacent compound in the chemical space.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Analysis method for identifying structure of compound outside NIST spectrum library and application
CN115248282A
Method and system for determining structure of organic compound by using spectrum data
CN115966262A
Vocs volatile organic compound component analysis method and device and storage medium
CN116525019A
Compound molecular fingerprint prediction algorithm based on learning structural relationship
CN117059174A
Chemical structure determination method and device of compound and terminal equipment
CN117854618A