Method, device and terminal equipment for determining chemical structure of compound

The chemical structure of the compound is determined through mass spectrometry analysis and chemical spatial distance, and the structural analysis problem of unincluded compounds in the prior art is solved, and a more comprehensive compound analysis is achieved.

CN117854618BActive Publication Date: 2025-08-19AGRICULTURAL GENOMICS INSTITUTE AT SHENZHEN CHINESE ACADEMY OF AGRICULTURAL SCIENCES (SHENZHEN BRANCH GUANGDONG LABORATORY FOR LINGNAN MODERN AGRICULTURE)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311690470.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2025-08-19
Estimated Expiration
2043-12-07

AI Technical Summary

Technical Problem

The prior art cannot effectively analyze the chemical structure of compounds not included in the mass spectrometry database, resulting in a small analysis coverage.

Method used

The molecular formula of the compound and the adjacent mass spectrum are determined by mass spectrometry analysis, and the chemical structure of the compound is determined by combining the spatial distance in the chemical space.

Benefits of technology

The coverage of compound analysis has been broadened, the structural analysis problem of unknown compounds has been solved, and the accuracy and coverage of the analysis have been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117854618B_ABST
    Figure CN117854618B_ABST
Patent Text Reader

Abstract

This application is applicable to the field of compound analysis technology, and provides a method, apparatus, and terminal device for determining the chemical structure of a compound. The chemical structure determination method includes: determining the molecular formula of the compound to be determined and at least one adjacent mass spectrum based on the mass spectrum of the compound to be determined, each adjacent mass spectrum corresponding to a neighboring compound; determining at least one candidate compound based on the molecular formula; and determining the chemical structure of the compound to be determined based on the spatial distance between each candidate compound and each neighboring compound in the chemical space. Through the above scheme, the problem of the inability to perform structural analysis on compounds not included in the database in the prior art and the small analysis coverage is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of compound analysis technology, and in particular relates to a method, apparatus and terminal device for determining the chemical structure of a compound. Background Art

[0002] Mass spectrometry plays a crucial role in analyzing the diverse content of small molecules in complex systems, especially in biological complex systems such as metabolomics. In the qualitative analysis of small molecule compounds, traditional qualitative methods primarily match the mass spectrum of the compound to be analyzed with mass spectra in a standard mass spectrometry database to obtain relevant data and determine the structural characteristics of the compound to be analyzed. However, this method cannot achieve structural elucidation of compounds not included in the database, resulting in a limited analytical coverage issue that needs to be addressed urgently. Summary of the Invention

[0003] The purpose of this application is to provide a method, apparatus and terminal device for determining the chemical structure of a compound, aiming to solve the problem in the prior art that it is impossible to perform structural analysis on compounds not included in the database and that the analysis coverage is small.

[0004] A first aspect of an embodiment of the present application provides a method for determining the chemical structure of a compound, characterized in that the method comprises:

[0005] Determining a molecular formula of the compound to be determined and at least one adjacent mass spectrum according to the mass spectrum of the compound to be determined, each adjacent mass spectrum corresponding to a adjacent compound;

[0006] determining at least one candidate compound according to the molecular formula;

[0007] The chemical structure of the compound to be determined is determined according to the spatial distance between each candidate compound and each adjacent compound in the chemical space.

[0008] Preferably, before determining the molecular formula and adjacent mass spectra of the compound to be determined based on the mass spectra of the compound to be determined, the chemical structure determination method further comprises:

[0009] Convert each mass spectrum in the mass spectrum database to be matched into corresponding natural language;

[0010] Inputting the natural language of each mass spectrum into a data conversion model, and outputting the mass spectrum numerical vector corresponding to each mass spectrum by the data conversion model;

[0011] The mass spectrum numerical vector corresponding to each mass spectrum graph is saved in a numerical matrix.

[0012] Preferably, determining at least one adjacent mass spectrum of the compound to be determined based on the mass spectrum of the compound to be determined comprises:

[0013] Converting the mass spectrum of the compound to be determined into natural language and inputting it into the data conversion model, so that the data conversion model outputs a mass spectrum numerical vector corresponding to the compound to be determined;

[0014] In the numerical matrix, screening mass spectrum numerical vectors whose vector distance with the mass spectrum numerical vector corresponding to the compound to be determined meets a preset threshold, and setting at least one mass spectrum numerical vector obtained by screening as an adjacent mass spectrum numerical vector;

[0015] Determine the adjacent mass spectrum corresponding to each adjacent mass spectrum numerical vector; each adjacent mass spectrum corresponds to a adjacent compound.

[0016] Preferably, determining the molecular formula of the compound to be determined based on the mass spectrum of the compound to be determined comprises:

[0017] Determining the relative molecular mass of the compound to be determined based on the mass spectrum of the compound to be determined;

[0018] The molecular formula of the compound to be determined is determined according to the relative molecular mass of the compound to be determined.

[0019] Preferably, determining at least one candidate compound according to the molecular formula comprises:

[0020] According to the molecular formula, screening a database of structures of compounds to be matched to obtain at least one compound to be matched that has the same molecular formula as the molecular formula;

[0021] The at least one compound to be matched is determined as a candidate compound.

[0022] Preferably, before determining the chemical structure of the compound to be determined based on the spatial distance between each candidate compound and each adjacent compound in the chemical space, the chemical structure determination method further comprises:

[0023] Obtaining chemical structure data of each candidate compound and the adjacent compounds;

[0024] Converting the chemical structure data of each candidate compound and the adjacent compound into one-to-one corresponding structure numerical vectors;

[0025] The multiple structural numerical vectors are projected into the chemical space respectively, and each structural numerical vector corresponds to a point in the chemical space.

[0026] Preferably, determining the chemical structure of the compound to be determined based on the spatial distance between each candidate compound and each adjacent compound in the chemical space comprises:

[0027] Calculating the average distance between the point corresponding to the candidate compound and each point corresponding to the adjacent compound;

[0028] Repeat the above calculation process to obtain the average distance between the points corresponding to the remaining candidate compounds and each point corresponding to the adjacent compounds;

[0029] All calculated average distances are sorted in order from small to large, and the candidate compound corresponding to the first average distance is selected as the compound to be determined, and the structural data corresponding to the candidate compound is the structural data of the compound to be determined.

[0030] A second aspect of an embodiment of the present application provides a device for determining the chemical structure of a compound, the device comprising:

[0031] A first determination module is configured to determine a molecular formula of the compound to be determined and at least one adjacent mass spectrum according to the mass spectrum of the compound to be determined, wherein each adjacent mass spectrum corresponds to a nearby compound;

[0032] a second determination module, configured to determine at least one candidate compound according to the molecular formula;

[0033] The third determination module is configured to determine the chemical structure of the compound to be determined based on the spatial distance between each candidate compound and each adjacent compound in the chemical space.

[0034] A third aspect of an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the computer program.

[0035] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.

[0036] A fifth aspect of the embodiments of the present application provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute the method described in the first aspect.

[0037] Beneficial effects of this application

[0038] In the technical solution adopted in the present application, in addition to determining the molecular formula of the compound to be determined and the adjacent compounds based on the mass spectrum of the compound to be determined, at least one candidate compound is further determined based on the molecular formula of the compound to be determined, and the chemical structure of the compound to be determined is determined by the spatial distance between each candidate compound and the adjacent compounds in the chemical space. Compared with the existing technology, the breakthrough of the method of the present application is that it not only relies on a pre-established mass spectrometry database, but also explores possible structures more comprehensively in the chemical space. This innovation not only broadens the coverage of compound analysis, but also solves the problem that the existing technology cannot perform structural analysis when dealing with unknown compounds, and provides a more comprehensive solution for the identification of unknown compounds. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0040] Figure 1 This is a flow chart of a method for determining the chemical structure of a compound provided in an embodiment of the present application;

[0041] Figure 2 This is a second flow chart of a method for determining the chemical structure of a compound provided in an embodiment of the present application;

[0042] Figure 3 This is a flowchart of a method for determining the chemical structure of a compound provided in an embodiment of the present application;

[0043] Figure 4 This is a fourth flow chart of a method for determining the chemical structure of a compound provided in an embodiment of the present application;

[0044] Figure 5 Flowchart 5 of a method for determining the chemical structure of a compound provided in an embodiment of the present application;

[0045] Figure 6 A comparison chart of data results provided in the examples of this application;

[0046] Figure 7 A schematic diagram illustrating the chemical structure of a compound provided in an embodiment of the present application;

[0047] Figure 8 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0048] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0049] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0050] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0051] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0052] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0053] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0054] It should be understood that the size of the serial numbers of each step in this embodiment does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of this application.

[0055] Mass spectrometry plays a crucial role in analyzing the diverse content of small molecules in complex systems, especially in biological complex systems such as metabolomics. In the qualitative analysis of small molecule compounds, traditional qualitative methods primarily match the mass spectrum of the compound to be analyzed with mass spectra in a standard mass spectrometry database to obtain relevant data and determine the structural characteristics of the compound to be analyzed. However, this method cannot achieve structural elucidation of compounds not included in the database, resulting in a limited analytical coverage issue that needs to be addressed urgently.

[0056] In this regard, the present application provides a method for determining the chemical structure of a compound, which determines the molecular formula of the compound to be determined and at least one adjacent mass spectrum based on the mass spectrum of the compound to be determined, wherein each adjacent mass spectrum corresponds to a neighboring compound, and then determines at least one candidate compound based on the molecular formula. Thereafter, the chemical structure of the compound to be determined is determined based on the spatial distance between each candidate compound and each of the neighboring compounds in the chemical space.

[0057] In the technical solution adopted in the present application, in addition to determining the molecular formula of the compound to be determined and the adjacent compounds based on the mass spectrum of the compound to be determined, at least one candidate compound is further determined based on the molecular formula of the compound to be determined, and the chemical structure of the compound to be determined is determined by the spatial distance between each candidate compound and the adjacent compounds in the chemical space. This breaks through the limitation of the existing technology of only matching the compound to be determined in the mass spectrum database, broadens the analysis coverage of the compound to be determined, and solves the problem that the existing technology cannot complete structural analysis for the compound to be determined that is not in the database.

[0058] It should be noted that the mass spectrum is obtained through mass spectrometry analysis, which is an analytical method for measuring the mass-to-charge ratio (mass-to-charge ratio) of ions. Its basic principle is to ionize the components in the sample in an ion source to generate charged ions with different mass-to-charge ratios. After being accelerated by an electric field, an ion beam is formed and enters a mass analyzer. In the mass analyzer, the electric and magnetic fields are used to cause opposite velocity dispersion, and they are focused separately to obtain a mass spectrum, thereby determining their relative molecular mass. Mass spectrometry can also provide rich compound structure information in a single analysis.

[0059] In order to illustrate the technical solution of the present application, specific embodiments are provided below.

[0060] Figure 1A flow chart of a method for determining the chemical structure of a compound provided in an embodiment of the present application is shown.

[0061] Reference Figure 1 The present invention provides a method for determining the chemical structure of a compound, which comprises the following steps:

[0062] Step S101: determining the molecular formula of the compound to be determined and at least one adjacent mass spectrum according to the mass spectrum of the compound to be determined, where each adjacent mass spectrum corresponds to a nearby compound;

[0063] Step S102: determining at least one candidate compound according to the molecular formula;

[0064] Step S103: determining the chemical structure of the compound to be determined based on the spatial distance between each candidate compound and each adjacent compound in the chemical space.

[0065] The above-mentioned steps of determining the compound by mass spectrum are not limited to matching known compounds in the mass spectrum database, but further determine the candidate compound by molecular formula information obtained by mass spectrum analysis. This new method takes into account the relative position of the candidate compound and the adjacent compound in the chemical space, and calculates the spatial distance between them to obtain the potential structure of the compound to be determined. Compared with the prior art, the breakthrough of the present application method is that it does not just rely on the pre-established mass spectrum database, but explores possible structures more comprehensively in the chemical space. This innovation not only broadens the coverage of compound analysis, but also solves the problem that the prior art cannot perform structural analysis when dealing with unknown compounds, and provides a more comprehensive solution for the identification of unknown compounds. Therefore, this method not only improves the analytical ability of unknown compounds, but also provides a new way for the accurate identification and analysis of compounds of unknown structure.

[0066] It should be noted that the unidentified compounds in this application refer to compounds that are not included in the mass spectrometry database and may be unknown compounds, i.e. compounds whose relevant chemical characteristics have not been identified, including compounds with unknown chemical structures;

[0067] In order to efficiently and accurately match the neighboring compounds of the compound to be determined, refer to Figure 2 Before performing step S101, the chemical structure determination method in the embodiment of the present application further includes the following steps:

[0068] Step A1: converting each mass spectrum in the mass spectrum database to be matched into corresponding natural language;

[0069] Step A2: inputting the natural language of each mass spectrum into a data conversion model, and outputting a mass spectrum numerical vector corresponding to each mass spectrum from the data conversion model;

[0070] Step A3: Save the mass spectrum numerical vector corresponding to each mass spectrum graph into a numerical matrix.

[0071] Specifically, each mass spectrum peak in each mass spectrum in the mass spectrum database to be matched is represented as a word containing its position. In addition to all peaks in the spectrum, a certain range of neutral loss is also included. The calculation of neutral loss is defined as the parent ion m / z minus the peak m / z. After the above processing, each mass spectrum can be represented as a natural language composed of several words, which can then be input into a data conversion model. The data conversion model converts the natural language into a corresponding mass spectrum numerical vector, and then the mass spectrum numerical vector corresponding to each mass spectrum obtained by conversion is saved in a numerical matrix.

[0072] The mass spectrum numerical vector is a numerical vector converted from the natural language corresponding to the mass spectrum of the compound.

[0073] Exemplarily, for each mass spectrum in the database to be matched, each mass spectrum peak is represented as a word containing its position, accurate to a specified number of decimal places, represented as "peak@xxx.xx". In all presented results, a grouping method including two decimal places is adopted, for example, the peak at m / z 200.445 is represented as "peak@200.45". In addition to all peaks in the spectrum, neutral losses ranging from 5.0Da to 200.0Da are also included, represented as "loss@xxx.xx"; after processing, each mass spectrum can be represented as a natural language consisting of several words, which can then be converted into a numerical vector through a data conversion model and saved in a numerical matrix.

[0074] The above method converts each mass spectrum in the mass spectrum database to be matched into a mass spectrum numerical vector and stores it in a numerical matrix to complete the data preprocessing, which facilitates the subsequent calculation of the vector distance between the mass spectrum numerical vectors corresponding to the compound to be determined, so as to determine the adjacent mass spectra and obtain the adjacent compounds corresponding to the compound to be determined, thereby improving the analysis speed and accuracy.

[0075] It should be noted that, in the embodiments that can be implemented in this application, the mass spectrum database to be matched may include one or more databases, and specifically the following databases may be used:

[0076] NIST Mass Spectral Library: provides mass spectra and related information for a large number of compounds;

[0077] MassBank: a mass spectrometry database containing mass spectrometry data obtained from different laboratories;

[0078] METLIN: contains mass spectrometric data of metabolites in biological fluids;

[0079] GNPS (Global Natural Products Social Molecular Networking): provides mass spectrometry data, mainly about natural products;

[0080] MassIVE: stores large-scale mass spectrometry data, including raw data and analysis results;

[0081] Human Metabolome Database (HMDB): Contains detailed information on compounds related to human metabolism, including mass spectrometry data.

[0082] It should be noted that the above databases are only examples and are not intended to be limiting. Those skilled in the art can select other mass spectrum data according to actual needs.

[0083] The data conversion model in the above embodiments of the present application can convert text words into numerical vectors, and can be specifically trained using the Word2vec model, GloVe model or Doc2Vec model.

[0084] The Word2Vec model is a two-layer neural network for processing text. Its input is a text corpus, and its output is a set of numerical vectors, that is, the numerical vectors of the words in the corpus. The purpose of the Word2Vec model is to group word vectors by similarity in the vector space and thereby identify mathematical similarities.

[0085] The input to the Word2Vec model is a sequence of words from a text corpus, each represented by a one-hot encoding. One-hot encoding is a vector representation method that assigns each word an N-dimensional vector, where N is the number of words in the vocabulary. Each word's vector has only one 1 and all other 0s, indicating that the word is unique.

[0086] The output of the Word2Vec model is a vocabulary in which each word has a corresponding vector. These vectors are obtained by training the input text corpus. They can be used to represent the contextual information, semantic information, and grammatical information of the word. The dimension of the output layer vector is usually N-dimensional, where N is the number of words in the vocabulary.

[0087] In one embodiment of the present application, input data and output data are used to form training data to train the Word2Vec model. During the training process, the model continuously adjusts parameters through the back propagation algorithm to minimize the prediction error, and finally obtains the data conversion model in the present application.

[0088] The following is a detailed explanation of step S101.

[0089] Reference Figure 3 To determine the neighboring compounds of the compound to be determined, step S101 may specifically include:

[0090] Step B1: converting the mass spectrum of the compound to be identified into natural language and inputting it into the data conversion model, and the data conversion model outputting a mass spectrum numerical vector corresponding to the compound to be identified;

[0091] Step B2: in the numerical matrix, screening mass spectrum numerical vectors whose vector distance with the mass spectrum numerical vector corresponding to the compound to be determined meets a preset threshold, and setting at least one mass spectrum numerical vector screened as a neighboring mass spectrum numerical vector;

[0092] Step B3: Determine the adjacent mass spectrum corresponding to each adjacent mass spectrum numerical vector; each adjacent mass spectrum corresponds to a nearby compound.

[0093] Specifically, in the aforementioned embodiments, certain data preprocessing has been performed, and each mass spectrum in the mass spectrum database to be matched is converted into a mass spectrum numerical vector and stored in a numerical matrix. To facilitate subsequent calculations, the same operation needs to be performed on the mass spectrum of the compound to be determined, that is, the mass spectrum of the compound to be determined is converted into natural language and input into the data conversion model, and the data conversion model outputs the mass spectrum numerical vector corresponding to the compound to be determined.

[0094] To find at least one mass spectrum vector in the numerical matrix whose vector distance with the mass spectrum vector corresponding to the compound to be identified meets a preset threshold, thereby screening for at least one corresponding adjacent compound, a vector distance calculation is first performed. Vector distances include, but are not limited to, Cosine distance, Mahalanobis distance, and Manhattan distance. After the calculation is completed, the results are sorted from smallest to largest, and at least one mass spectrum vector whose vector distance is within the preset threshold is selected and set as the adjacent mass spectrum vector.

[0095] It should be noted that each row in the above numerical matrix corresponds to a mass spectrum in the mass spectrum database to be matched, and at the same time corresponds to a compound structure. After determining at least one adjacent mass spectrum numerical vector, the adjacent mass spectra corresponding to each adjacent mass spectrum numerical vector can be determined, and the chemical structure of the adjacent compound corresponding to each adjacent mass spectrum can be obtained. The chemical structure of each adjacent compound is considered to be similar to the chemical structure of the compound to be determined, and can be used as a basis for subsequent analysis of the compound.

[0096] The above-mentioned vector distance calculation can screen out the neighboring compounds and corresponding chemical structures of the compound to be determined in the numerical matrix, which can improve the accuracy of the analysis and provide relevant basis for subsequent analysis.

[0097] Reference Figure 4 To determine the molecular formula of the compound to be determined, step S101 may specifically include:

[0098] Step C1: determining the relative molecular mass of the compound to be determined based on the mass spectrum of the compound to be determined;

[0099] Step C2: Determine the molecular formula of the compound to be determined according to the relative molecular mass of the compound to be determined.

[0100] Specifically, relative molecular mass refers to the sum of the relative atomic masses (Ar) of each atom in the chemical formula. To determine the relative molecular mass based on the mass spectrum, it is necessary to find the molecular ion peak (M+) in the mass spectrum of the compound to be determined. The molecular ion peak corresponds to an ion with a positive charge in the molecule, and its mass-to-charge ratio (m / z) is close to the relative molecular mass of the molecule. By measuring the mass-to-charge ratio of the molecular ion peak, the relative molecular mass of the compound to be determined can be obtained.

[0101] To further determine the molecular formula of the compound to be determined, it is necessary to calculate the mass fraction of each element based on the relative atomic mass and atomic number of each element in the molecule, and then obtain the molecular formula based on the mass fraction and relative molecular mass. The specific method is: multiply the relative molecular mass by the mass fraction of an element, and then divide it by the relative atomic mass of the element to obtain the number of atoms of the element. Then, use the same method to calculate the number of atoms of other elements, and finally write the molecular formula.

[0102] The molecular formula of the compound to be determined obtained by the above method can be used to subsequently obtain candidate compounds to broaden the analysis coverage of the compound, so that the analysis scope of the compound to be determined is not limited to the compounds in each mass spectrum database, further improving the accuracy of the analysis.

[0103] The following is a detailed explanation of step S102.

[0104] To determine the candidate compound, step S102 may specifically include:

[0105] According to the molecular formula, at least one compound to be matched that has the same molecular formula as the molecular formula is screened in a database of compound structures to be matched; and the at least one compound to be matched is determined as a candidate compound.

[0106] The database of compound structures to be matched is a resource used to store and retrieve information about the structure, properties, and activity of compounds. Within the database, users can find detailed structural information for various compounds, including molecular formula, molecular weight, chemical bonds, stereo configuration, and electron distribution. Furthermore, users can use keywords such as molecular formula to search and filter the chemical structure information of compounds with the same molecular formula within the database.

[0107] Specifically, in an embodiment that can be implemented in the present application, the database of compound structures to be matched may include one or more databases, and specifically the following databases may be used:

[0108] HMDB (Human Metabolome Database): A database dedicated to human metabolites, providing detailed information on metabolites under physiological and pathological conditions;

[0109] KEGG (Kyoto Encyclopedia of Genes and Genomes): includes information on biology, biochemistry, and systems biology, and provides databases on biomolecules, metabolic pathways, and drugs.

[0110] ChEBI (Chemical Entities of Biological Interest): contains biologically important compounds;

[0111] PubChem: a public compound database containing a large amount of information on biologically active small molecules;

[0112] ChemSpider: a compound database containing a large number of organic, inorganic and biologically active compounds;

[0113] DrugBank: A database for drugs and drug targets, providing information on drug structure, mechanism of action, metabolic pathways, etc.

[0114] PDB (Protein Data Bank): contains a database of protein and nucleic acid structures, as well as structural information of small molecules and drugs that bind to them;

[0115] METLIN: Metabolomics database, containing mass spectrometry data of metabolites;

[0116] SwissBioisostere: provides a database of bioequivalents to help find compounds with similar biological activities;

[0117] ZINC (ZINC Is Not Commercial): A compound structure database, mainly used for virtual screening and drug design.

[0118] It should be noted that the above databases are only examples and are not intended to be limiting. Those skilled in the art can select other mass spectrum data according to actual needs.

[0119] In order to concretely reflect the spatial distance distribution between each candidate compound and each adjacent compound, and to judge the structural similarity between the candidate compound and the adjacent compounds, the present application can project each candidate compound and each adjacent compound into the chemical space for subsequent analysis.

[0120] Specifically, the chemical space of a compound can be described using molecular fingerprints. Molecular fingerprints are a coding method for representing molecular structures, typically in binary form. They are characterized by numerical vectors that can describe the chemical structure of a molecule. They are used to compare and analyze the structural similarities between molecules in chemistry and drug discovery. Molecular fingerprints capture structural information in a molecule and map it to a fixed-length binary numerical vector. From a mathematical perspective, when a molecule is represented as a numerical vector, it can be viewed as a point in a high-dimensional space. Distance functions between vectors can be used to describe the distances between these points. This representation method facilitates computer processing and comparison of molecular structures and is widely used in fields such as virtual screening, drug design, and chemical information retrieval. Molecular fingerprints include, but are not limited to, Morgan fingerprints, MACCS fingerprints, and PubChem fingerprints.

[0121] Correspondingly, in the embodiments that can be implemented in the present application, the chemical structure data of each candidate compound and each adjacent compound are first obtained, and the chemical structure data of each candidate compound and each adjacent compound are converted into one-to-one corresponding structure numerical vectors through molecular fingerprint calculation, and then the multiple structure numerical vectors are projected into the chemical space respectively, and each structure numerical vector corresponds to a point in the chemical space.

[0122] The structure numerical vector is a numerical vector converted according to the chemical structure data of the compound.

[0123] The above conversion of the chemical structure data of each candidate compound and each adjacent compound into a structural numerical vector and using it as a corresponding point in the chemical space helps to analyze the structural similarity between the candidate compound and the adjacent compounds and reduces the complexity of the analysis.

[0124] The following is a detailed explanation of step S103.

[0125] After converting each candidate compound and each adjacent compound into a structural numerical vector, it is necessary to further calculate the spatial distance between each candidate compound and each adjacent compound in the chemical space. Figure 5 Step S103 may specifically include:

[0126] Step D1: Calculating the average distance between the point corresponding to the candidate compound and each point corresponding to the adjacent compound;

[0127] Step D2: Repeat the above calculation process to obtain the average distance between the points corresponding to the remaining candidate compounds and each point corresponding to the adjacent compounds;

[0128] Step D3: sorting all the calculated average distances in ascending order, selecting the candidate compound corresponding to the first average distance as the compound to be determined, and the structural data corresponding to the candidate compound is the structural data of the compound to be determined.

[0129] Specifically, for a candidate compound, the vector distance between the structural numerical vector of the candidate compound and the structural numerical vector of each adjacent compound is calculated. The vector distance includes but is not limited to Cosine distance, Mahalanobis distance, Manhattan distance, etc., and the average distance is further calculated. The above calculation process is repeated to obtain the average distance corresponding to the remaining candidate compounds; based on the calculated average distance, the compounds are sorted in order from small to large. The higher the ranking, the more likely it is the true structure of the compound to be determined. The candidate compound corresponding to the first average distance is selected as the compound to be determined, and the structural data corresponding to the candidate compound is the structural data of the compound to be determined.

[0130] The above method obtains the candidate compound with the smallest average distance to each adjacent compound through calculation as the compound to be determined, and thereby determines the chemical structure of the compound to be determined. It concretely provides how to accurately determine the chemical structure of the compound to be determined after broadening the analysis coverage, and reduces the need for manual intervention and interpretation, thereby improving the speed and accuracy of compound analysis.

[0131] The following describes the implementation methods of this application in detail with reference to specific implementation data.

[0132] 600,289 different mass spectra in positive ion mode belonging to 43,653 different compounds from data sources such as GNPS were converted into corresponding natural language. The natural language was input into the Word2Vec model to train a data conversion model. The model was used to convert all mass spectra into numerical vectors of length 256, and a 600,289*256 matrix was obtained and saved. The mass spectrum of the compound to be determined, N-Acetyl-5-aminosalicylicacid, was regarded as an unknown mass spectrum. Its compound structure is as follows Figure 6 As shown in Figure A, the data is converted into a 256-length numeric vector using the same data conversion model. The top 10 mass spectra with the greatest cosine similarity to the vector are retrieved from the stored matrix. These are called neighboring mass spectra, and the corresponding compounds are called neighboring compounds (Top 10 Neigbors). It should be noted that the compound to be identified, N-Acetyl-5-aminosalicylicacid, is not included in the data included in the mass spectrum database, so this search is unlikely to find a compound and mass spectrum that is exactly the same.

[0133] At the same time, according to the molecular formula of the compound, 34 candidate compounds (Candidates) were retrieved from compound structure databases such as HMDB and KEGG. Each candidate compound corresponds to a chemical structure. For the top 10 adjacent compounds retrieved and all candidate compounds, the molecular fingerprints were calculated and projected into the chemical space. The UMAP dimensionality reduction method was used to project the fingerprints into a two-dimensional space and visualize them, as shown in the following example: Figure 6 As shown in B. According to the average distance between each candidate compound and each adjacent compound, the 34 potential candidate compounds are sorted, and the candidate compound with the shortest average distance is ranked first. This candidate compound is the compound to be determined. Figure 6 As can be seen in B, the candidate compound represented by the dot is closest to each adjacent compound, and the candidate compound corresponding to this dot (True Annotation) is determined as the compound to be determined; Figure 6 C and Figure 6 D shows the structures and corresponding adjacent mass spectra of two adjacent compounds, showing that they have similar chemical structures to N-Acetyl-5-aminosalicylic acid. The two adjacent mass spectra show a significant correlation, which can serve as a basis for compound analysis. This indicates that the method for determining the chemical structure of the compound of this application has a high degree of accuracy.

[0134] Figure 7 This is a schematic diagram of the structure of a device for determining the chemical structure of a compound provided in an embodiment of the present application. For the sake of convenience, only the parts related to the embodiment of the present application are shown.

[0135] The chemical structure determination device of a compound may specifically include the following modules:

[0136] A first determination module M1 is configured to determine a molecular formula of a compound to be determined and at least one adjacent mass spectrum according to the mass spectrum of the compound to be determined, wherein each adjacent mass spectrum corresponds to a nearby compound;

[0137] A second determination module M2, configured to determine at least one candidate compound according to the molecular formula;

[0138] The third determination module M3 is configured to determine the chemical structure of the compound to be determined based on the spatial distance between each candidate compound and each adjacent compound in the chemical space.

[0139] The first determination module M1 determines the adjacent compounds and molecular formula of the compound to be determined based on the mass spectrum of the compound to be determined, and then the second determination module M2 determines at least one candidate compound based on the molecular formula of the compound to be determined. Thereafter, the third determination module calculates the spatial distance between each candidate compound and each adjacent compound in the chemical space to determine the chemical structure of the compound to be determined. The above operations adopted in this application not only rely on the pre-established mass spectrometry database, but also explore possible structures more comprehensively in the chemical space. This innovation not only broadens the coverage of compound analysis, but also solves the problem that the existing technology cannot perform structural analysis when dealing with unknown compounds, and provides a more comprehensive solution for the identification of unknown compounds.

[0140] Figure 8 This is a schematic diagram of the structure of a terminal device provided in an embodiment of the present application. The terminal device E1 includes: at least one processor E2 ( Figure 7 Only one is shown in the figure) a processor, a memory E3, and a computer program E4 stored in the memory E3 and executable on the at least one processor E2, wherein the processor E2 implements the steps in the above-mentioned embodiment of the automated testing method when executing the computer program E4.

[0141] The terminal device E1 can be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The terminal device may include, but is not limited to, a processor E2 and a memory E3. Those skilled in the art will understand that Figure 7 This is merely an example of the terminal device E1 and does not constitute a limitation on the terminal device E1. The terminal device E1 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the terminal device E1 may also include input and output devices, network access devices, etc.

[0142] The processor E2 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.

[0143] In some embodiments, the memory E3 may be an internal storage unit of the terminal device E1, such as a hard disk or memory of the terminal device E1. In other embodiments, the memory E3 may also be an external storage device of the terminal device E1, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device E1. Furthermore, the memory E3 may also include both an internal storage unit of the terminal device E1 and an external storage device. The memory E3 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory E3 may also be used to temporarily store data that has been output or is to be output.

[0144] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0145] In the above embodiments, the description of each embodiment has different emphases. If a rated part is not described or recorded in detail in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0146] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0147] An embodiment of the present application provides a computer program product. When the computer program product is run on a mobile terminal, the mobile terminal can implement the steps in the above-mentioned various method embodiments when executing the computer program product.

[0148] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0149] In the embodiments provided in this application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0150] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0151] In addition, the functional units in the various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above units may be implemented in the form of hardware or software.

[0152] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0153] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for determining the chemical structure of a compound, characterized in that: The chemical structure determination method comprises: Convert each mass spectrum in the mass spectrum database to be matched into the corresponding natural language, and encode the mass spectrum peak in the format of "peak@xxx.xx"; Inputting the natural language of each mass spectrum into a data conversion model, and outputting the mass spectrum numerical vector corresponding to each mass spectrum by the data conversion model; Saving the mass spectrum numerical vector corresponding to each mass spectrum into a numerical matrix; determining the molecular formula of the compound to be determined based on the mass spectrum of the compound to be determined; wherein the compound to be determined refers to an unknown compound not included in the mass spectrum database; Determining at least one adjacent mass spectrum of the compound to be determined based on the mass spectrum of the compound to be determined, including: determining an adjacent mass spectrum corresponding to each adjacent mass spectrum numerical vector, each adjacent mass spectrum corresponding to an adjacent compound; Determining at least one candidate compound according to the molecular formula includes: screening a database of structures of compounds to be matched to obtain at least one compound to be matched that has the same molecular formula according to the molecular formula; and determining the at least one compound to be matched as a candidate compound; Determining the chemical structure of the compound to be determined based on the spatial distance between each candidate compound and each adjacent compound in the chemical space, including: calculating the average distance between the point corresponding to the candidate compound and each point corresponding to the adjacent compound; repeating the above calculation process to obtain the average distance between the points corresponding to the remaining candidate compounds and each point corresponding to the adjacent compound; sorting all the calculated average distances in ascending order, selecting the candidate compound corresponding to the first average distance as the compound to be determined, and the structural data corresponding to the candidate compound is the structural data of the compound to be determined.

2. The method for determining the chemical structure according to claim 1, wherein The step of determining at least one adjacent mass spectrum of the compound to be determined based on the mass spectrum of the compound to be determined further comprises: Converting the mass spectrum of the compound to be determined into natural language and inputting it into the data conversion model, so that the data conversion model outputs a mass spectrum numerical vector corresponding to the compound to be determined; In the numerical matrix, mass spectrum numerical vectors whose vector distance with the mass spectrum numerical vector corresponding to the compound to be determined meets a preset threshold are screened, and at least one mass spectrum numerical vector obtained by screening is set as an adjacent mass spectrum numerical vector.

3. The method for determining the chemical structure according to claim 1, wherein Determining the molecular formula of the compound to be determined according to the mass spectrum of the compound to be determined includes: Determining the relative molecular mass of the compound to be determined based on the mass spectrum of the compound to be determined; The molecular formula of the compound to be determined is determined according to the relative molecular mass of the compound to be determined.

4. The method for determining the chemical structure according to claim 1, wherein Before determining the chemical structure of the compound to be determined based on the spatial distance between each candidate compound and each adjacent compound in the chemical space, the chemical structure determination method further includes: Obtaining chemical structure data of each candidate compound and each adjacent compound; Converting the chemical structure data of each candidate compound and each adjacent compound into one-to-one corresponding structure numerical vectors; The multiple structural numerical vectors are projected into the chemical space respectively, and each structural numerical vector corresponds to a point in the chemical space.

5. A device for determining the chemical structure of a compound, characterized in that: The chemical structure determination device comprises: a first determination module for determining a molecular formula of a compound to be determined based on a mass spectrum of the compound to be determined; wherein the compound to be determined refers to an unknown compound not included in a mass spectrum database; determining at least one adjacent mass spectrum of the compound to be determined based on the mass spectrum of the compound to be determined, including: determining an adjacent mass spectrum corresponding to each adjacent mass spectrum numerical vector, each adjacent mass spectrum corresponding to an adjacent compound; a second determination module for determining at least one candidate compound based on the molecular formula, including: screening, based on the molecular formula, at least one compound to be matched that has the same molecular formula as the compound to be matched in a structure database of compounds to be matched; and determining the at least one compound to be matched as a candidate compound; A third determination module is configured to determine the chemical structure of the compound to be determined based on the spatial distance between each candidate compound and each adjacent compound in the chemical space, comprising: calculating the average distance between a point corresponding to the candidate compound and each point corresponding to the adjacent compound; repeating the above calculation process to obtain the average distance between the points corresponding to the remaining candidate compounds and each point corresponding to the adjacent compound; sorting all the calculated average distances in ascending order, selecting the candidate compound corresponding to the first average distance as the compound to be determined, and the structural data corresponding to the candidate compound is the structural data of the compound to be determined; The chemical structure determination device is further used for: Convert each mass spectrum in the mass spectrum database to be matched into the corresponding natural language, and encode the mass spectrum peak in the format of "peak@xxx.xx"; Inputting the natural language of each mass spectrum into a data conversion model, and outputting the mass spectrum numerical vector corresponding to each mass spectrum by the data conversion model; The mass spectrum numerical vector corresponding to each mass spectrum graph is saved in a numerical matrix.

6. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 4 is implemented.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Molecular design scheme determination method and device, equipment and storage medium

    CN114300065A