A method, system, storage medium and terminal for automatically determining a structure of a compound based on a nuclear magnetic resonance spectrum

By constructing a nuclear magnetic resonance chemical shift database and calculating the cosine similarity of compound feature vectors, the problem of time-consuming compound structure analysis in existing technologies is solved, and fast and accurate automatic prediction of compound structures is achieved.

CN116773580BActive Publication Date: 2026-03-17YAORONGYUN DIGITAL TECH (CHENGDU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310741364.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-21
Publication Date
2026-03-17
Estimated Expiration
2043-06-21

AI Technical Summary

Technical Problem

Existing technologies are time-consuming and rely heavily on experience when automatically determining the structure of compounds, making it difficult to quickly and accurately resolve molecular structures from NMR spectra.

Method used

A nuclear magnetic resonance chemical shift database was constructed. By converting 13C NMR data into feature vectors of length 400, the cosine similarity between the compound to be predicted and the feature vectors in the database was calculated. Combined with known element information, the structure of the compound was automatically predicted.

Benefits of technology

It enables rapid and automated compound structure prediction based on NMR spectrum data, reducing data requirements and improving analysis efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116773580B_ABST
    Figure CN116773580B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on nuclear magnetic spectrum automatic determination compound structure method, system, storage medium and terminal, belong to biochemical field, comprising: construct nuclear magnetic resonance chemical shift database;13C nuclear magnetic data in database is converted into long 400 feature vectors;The compound to be predicted is converted into long 400 feature vectors, the cosine similarity of the feature vector of the compound to be predicted and all feature vectors in database is calculated;The structure of the compound to be predicted is predicted according to the calculation result of cosine similarity.The present application is based on a large number of nuclear magnetic spectrum data, realizes that only according to nuclear magnetic spectrum data to molecular structure fast automatic prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biochemistry, and in particular to a method, system, storage medium, and terminal for automatically determining the structure of compounds based on nuclear magnetic resonance spectroscopy. Background Technology

[0002] Computer-aided compound structure determination is a classic problem at the intersection of informatics, chemistry, and mathematics. Nuclear magnetic resonance (NMR) is the most widely used technique for characterizing organic molecular structures. It characterizes the local environment of the atoms that make up a molecule, providing the molecule's "fingerprint," which chemists can use to infer the structure of compounds. However, even relatively small molecules may have a large number of hydrogen NMR peaks with complex splitting patterns. Therefore, manual NMR data analysis is often time-consuming, error-prone, and requires extensive experience and accumulated chemical knowledge from chemical researchers. Due to the vast spatial dimensions of molecular structures, automatically determining molecular structures from their NMR spectra is very challenging. Therefore, automatically determining compound structures based on NMR spectra can help researchers accelerate chemical discovery.

[0003] Currently, auxiliary programs for structure resolution have been developed to resolve the chemical structures of small organic molecules. These typically require input of elemental composition, NMR data, and mass spectrometry data. These conditions generate a large number of candidate structures. To further narrow down the search, information on certain specific substructures is usually required, necessitating the combination of multiple data sets and resulting in lengthy processing times. Therefore, it is necessary to develop an automated method for rapidly resolving molecular structures that requires fewer types of data. Summary of the Invention

[0004] The purpose of this invention is to overcome the problems existing in the analysis of compound structures and to provide a method, system, storage medium and terminal for automatically determining the structure of compounds based on nuclear magnetic resonance spectroscopy.

[0005] The objective of this invention is achieved through the following technical solution:

[0006] In a first aspect, a method for automatically determining the structure of a compound based on nuclear magnetic resonance (NMR) spectroscopy is provided, the method comprising the following steps:

[0007] S1. Construct a nuclear magnetic resonance chemical shift database;

[0008] S2. Convert the 13C NMR data in the database into a feature vector of length 400;

[0009] S3. Convert the compound to be predicted into a feature vector of length 400 using the method in step S2, and calculate the cosine similarity between the feature vector of the compound to be predicted and all feature vectors in the database.

[0010] S4. Predict the structure of the compound to be predicted based on the calculation results of cosine similarity.

[0011] As a preferred embodiment, a method for automatically determining the structure of compounds based on nuclear magnetic resonance (NMR) spectra, wherein the construction of an NMR chemical shift database includes:

[0012] Collect nuclear magnetic resonance chemical shift data, including the structural codes of compounds and the 13C NMR data of compounds.

[0013] As a preferred embodiment, a method for automatically determining the structure of a compound based on NMR spectra, wherein converting 13C NMR data from a database into a 400-length feature vector includes:

[0014] S21. Remove duplicate data from the 13C NMR data, keeping only unique and non-repeating data;

[0015] S22. Round the 13C NMR data to the nearest integer.

[0016] S23. Add 100 to the rounded NMR data;

[0017] S24. Generate a vector of length 400, with indices starting from 0. Set the index of the NMR data to 1 and the rest to 0.

[0018] As a preferred embodiment, a method for automatically determining the structure of a compound based on NMR spectra, wherein converting 13C NMR data from a database into a 400-length feature vector, further includes:

[0019] S25. Spread the 1s at the location of the NMR data to both sides; if there are overlapping locations, add the values ​​together.

[0020] As a preferred embodiment, a method for automatically determining the structure of a compound based on nuclear magnetic resonance (NMR) spectra, wherein calculating the cosine similarity between the feature vector of the compound to be predicted and all feature vectors in a database includes:

[0021] The calculation formula is as follows:

[0022]

[0023] Where A and B are the feature vectors of the compound to be predicted and any feature vector from the database, respectively.

[0024] As a preferred embodiment, a method for automatically determining the structure of a compound based on nuclear magnetic resonance (NMR) spectra, wherein predicting the structure of the compound to be predicted based on the calculation results of cosine similarity includes:

[0025] All calculation results were sorted in descending order, and the 20 compounds corresponding to the most similar characteristic spectra were selected as candidate compounds for the NMR data to be tested.

[0026] If there is an NMR spectrum with a similarity of 1 among the first 20 characteristic spectra, then the corresponding compound is considered to be the compound structure of the NMR data to be tested.

[0027] If none of the top 20 most similar feature spectra have a similarity of 1, the structures of the compounds corresponding to the top 20 feature spectra are split to obtain functional group fragments, and then the fragments are recombined. The chemical shift of the recombined compounds is predicted, and the similarity is calculated based on the predicted NMR data. The compound structure corresponding to the maximum similarity is taken as the compound structure of the NMR data to be tested.

[0028] As a preferred option, a method for automatically determining the structure of a compound based on nuclear magnetic resonance spectroscopy involves screening based on the elements to be predicted, removing compounds that are impossible to exist, if the possible and impossible elements in the compound are known.

[0029] Secondly, a system for automatically determining the structure of compounds based on nuclear magnetic resonance (NMR) spectra is provided, the system comprising:

[0030] Nuclear magnetic resonance chemical shift database;

[0031] The feature vector calculation module is configured to convert 13C NMR data in the NMR chemical shift database into a feature vector of length 400.

[0032] The similarity calculation module is configured to calculate the cosine similarity between the feature vector of the compound to be predicted and all feature vectors in the database;

[0033] The compound structure prediction module is configured to predict the structure of the compound to be predicted based on the calculation results of cosine similarity.

[0034] Thirdly, a computer storage medium is provided, on which computer instructions are stored, wherein the computer instructions, when executed, perform any of the relevant contents of the method for automatically determining the structure of a compound based on nuclear magnetic resonance spectroscopy.

[0035] Fourthly, a terminal is provided, including a memory and a processor. The memory stores computer instructions that can be executed on the processor. When the processor executes the computer instructions, it performs any of the relevant contents of the method for automatically determining the structure of a compound based on nuclear magnetic resonance spectroscopy.

[0036] It should be further noted that the technical features corresponding to the above options can be combined or substituted to form new technical solutions if there is no conflict.

[0037] Compared with the prior art, the beneficial effects of the present invention are:

[0038] (1) Based on a large amount of NMR spectrum data, this invention constructs a NMR chemical shift database and automatically predicts the structure of the compound by calculating the cosine similarity between the feature vector of the compound to be predicted and all feature vectors in the database. It requires little data and realizes rapid automatic prediction of molecular structure based solely on NMR spectrum data.

[0039] (2) In one example, if the possible and impossible elements in a compound are known, the possible elements can be screened to remove compounds that cannot exist, further narrowing the scope and speeding up the structure prediction. Attached Figure Description

[0040] Figure 1 This is a flowchart illustrating a method for automatically determining the structure of a compound based on nuclear magnetic resonance (NMR) spectra, as shown in an embodiment of the present invention.

[0041] Figure 2 This is a specific process for determining the structure of a compound based on similarity, as illustrated in an embodiment of the present invention.

[0042] Figure 3 This is a schematic diagram of the mol code of xylene as shown in an embodiment of the present invention;

[0043] Figure 4 The present invention provides 13C NMR data for a compound to be predicted, as illustrated in an embodiment of the invention.

[0044] Figure 5 and Figure 6 These are the compounds corresponding to the 20 most similar characteristic spectra shown in the embodiments of the present invention;

[0045] Figure 7 This is a schematic diagram illustrating how, in an embodiment of the present invention, compounds containing F or S are screened out based on the fact that known compounds do not contain F and S elements.

[0046] Figure 8 This is a schematic diagram illustrating the breakdown of functional groups into basic functional groups and arranging these basic functional groups in descending order of their frequency, as shown in an embodiment of the present invention. Detailed Implementation

[0047] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0048] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0049] In one exemplary embodiment, reference is made to Figure 1 This paper provides a method for automatically determining the structure of compounds based on nuclear magnetic resonance (NMR) spectra, the method comprising the following steps:

[0050] S1. Construct a nuclear magnetic resonance chemical shift database;

[0051] S2. Convert the 13C NMR data in the database into a feature vector of length 400;

[0052] S3. Convert the compound to be predicted into a feature vector of length 400 using the method in step S2, and calculate the cosine similarity between the feature vector of the compound to be predicted and all feature vectors in the database.

[0053] S4. Predict the structure of the compound to be predicted based on the calculation results of cosine similarity.

[0054] The construction of the nuclear magnetic resonance chemical shift database includes:

[0055] a) Collect nuclear magnetic resonance chemical shift data, which mainly includes the structure codes (mol codes and Smiles codes) of compounds and the 13C NMR data of compounds;

[0056] b) The data is mainly stored in the database in the following form, and the following is a data point generated using xylene as an example;

[0057]

[0058] A single data entry includes the compound's mol code, SMILE code, 13C NMR chemical shift, corresponding atom index, solvent, and frequency of the NMR data. For example, the mol code for xylene is shown below. Figure 3 As shown.

[0059] c) The data is in units of molecules;

[0060] d) Store all collected nuclear magnetic resonance chemical shift data into the database in the data format described above, thus completing the construction of the database.

[0061] Further, data preprocessing is performed, including converting the 13C NMR data in the database into a feature vector of length 400, which includes:

[0062] S21. Remove duplicates from the 13C NMR data, keeping only the unique, non-repeating data. Taking xylene as an example, the deduplicated data is [136.42, 129.63, 125.85, 19.66].

[0063] S22. Round the 13C NMR data to the nearest integer. Taking xylene as an example, the rounded data is [136, 130, 126, 20].

[0064] S23. Add 100 to the rounded NMR data to get [236, 230, 226, 120];

[0065] S24. Generate a vector of length 400, starting with index 0. Set the index of the NMR data to 1 and the rest to 0, i.e.:

[0066]

[0067] S25. Considering the error in NMR data testing, the 1 at the position of the NMR data needs to be spread outwards, that is, replace 1 with 0.5, and fill in the numbers on both sides of 0.5 by decreasing by 0.1, i.e., 0.1, 0.2, 0.3, 0.4, 0.5, 0.4, 0.3, 0.2, 0.1. The vector after considering the testing error is as follows:

[0068]

[0069] S26. If there are overlapping positions, the values ​​are added together, and the following feature vector is obtained after addition:

[0070] .

[0071] S27. Store the feature vectors in the database and associate them with the compounds.

[0072] Further, the calculation of the cosine similarity between the feature vector of the compound to be predicted and all feature vectors in the database includes:

[0073] a) Input the 13C NMR spectrum data of the compound;

[0074] b) Extract peak positions from the input NMR spectrum data to obtain a list of chemical shifts;

[0075] c) Preprocess the obtained chemical shift data list according to the data preprocessing module described above to obtain a feature vector of length 400;

[0076] d) Calculate the cosine similarity between this feature vector and all feature vectors in the database. The calculation formula is as follows:

[0077]

[0078] Where A and B are the feature vectors of the compound to be predicted and any feature vector from the database, respectively. |A| and |B| are the magnitudes of the two vectors, respectively. The numerator is the dot product of the two vectors, and the denominator is the product of the magnitudes of the two vectors. Its value range is [-1, 1]. When the two vectors are in the same direction, the value is 1, which means that the two vectors are exactly the same. When the two vectors are in opposite directions, the value is -1, which means that the two vectors are completely different.

[0079] Furthermore, referring to Figure 2 The method of predicting the structure of the compound to be predicted based on the calculation results of cosine similarity includes:

[0080] After completing the cosine similarity calculation, all the calculation results are sorted in descending order, and the compounds corresponding to the top 20 most similar feature spectra are selected as candidate compounds for the NMR data to be tested.

[0081] If there is an NMR spectrum with a similarity of 1 among the first 20 characteristic spectra, then the corresponding compound is considered to be the compound structure of the NMR data to be tested.

[0082] If none of the top 20 most similar feature spectra have NMR data with a similarity of 1, then the structures of the compounds corresponding to the top 20 feature spectra need to be broken down to obtain functional group fragments. These fragments are then recombined, and the chemical shifts of the recombined compounds are predicted. Based on the predicted NMR data, the similarity is calculated as described above, and then sorted according to similarity. The compound structure corresponding to the maximum similarity is taken as the compound structure of the NMR data to be tested.

[0083] Furthermore, if the possible and impossible elements in the compound to be predicted are known, then the possible elements are screened to eliminate compounds that cannot exist, thus further narrowing down the range.

[0084] In one example, with Figure 4 Using the 13C NMR data of the compound shown as an example, we predict its structure:

[0085] 1. Extract peak positions from the input NMR spectrum to obtain a list of chemical shifts, namely: [166.32,156.68,135.29,129.08,126.51,124.2,98.39,35.39,12.9];

[0086] 2. Round the NMR data to the nearest integer. The rounded data is [166,157,135,129,127,124,98,35,13].

[0087] 3. Add 100 to the rounded NMR data, and you get [266,257,235,229,227,224,198,135,113];

[0088] 4. Convert the processed NMR data into a vector of length 400, which is:

[0089] 5. Perform cosine similarity calculations on this feature vector and all feature vectors in the database. Sort all the results in descending order and select the compounds corresponding to the top 20 most similar feature spectra as candidate compounds for the NMR data to be tested. Figure 5 and Figure 6 As shown, the captions in the images indicate the compound numbers and their similarity to the NMR spectra of the target compound.

[0090] 6. Furthermore, if it is clearly known that the compounds do not contain F and S elements, then compounds containing F or S elements are eliminated, further narrowing down the range of compounds to be identified. For example... Figure 7 As shown.

[0091] 7. For example Figure 7 It is known that no compound with a similarity score of 1 is matched. Therefore, these compounds need to be broken down into basic functional groups, and these basic functional groups should be arranged in descending order of frequency, such as... Figure 8 As shown.

[0092] 8. Then, these functional groups are recombine into new compounds. The chemical shifts of the recombine compounds are predicted. Based on the predicted NMR chemical shift data, steps 2-5 are repeated. After sorting by similarity, the structure most similar to the compound to be tested is found, thus obtaining the structure predicted based on the NMR data.

[0093] In another exemplary embodiment, a system for automatically determining the structure of a compound based on an NMR spectrum is provided, the system comprising:

[0094] Nuclear magnetic resonance chemical shift database;

[0095] The feature vector calculation module is configured to convert 13C NMR data in the NMR chemical shift database into a feature vector of length 400.

[0096] The similarity calculation module is configured to calculate the cosine similarity between the feature vector of the compound to be predicted and all feature vectors in the database;

[0097] The compound structure prediction module is configured to predict the structure of the compound to be predicted based on the calculation results of cosine similarity.

[0098] In another exemplary embodiment, the present invention provides a computer storage medium storing computer instructions thereon, which, when executed, perform relevant content in the method for automatically determining the structure of a compound based on an NMR spectrum.

[0099] Based on this understanding, the technical solution of this embodiment, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0100] In another exemplary embodiment, the present invention provides a terminal including a memory and a processor, wherein the memory stores computer instructions that can be executed on the processor, and the processor executes the relevant content of the method for automatically determining the structure of a compound based on an NMR spectrum when executing the computer instructions.

[0101] The processor may be a single-core or multi-core central processing unit or a specific integrated circuit, or one or more integrated circuits configured to implement the present invention.

[0102] The embodiments of the subject matter and functional operation described in this specification can be implemented in: tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by a data processing device or for controlling the operation of a data processing device. Alternatively or additionally, the program instructions may be encoded on artificially generated propagation signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information and transmit it to a suitable receiving device for execution by the data processing device.

[0103] The processing and logic flow described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform corresponding functions by operating on input data and generating output. The processing and logic flow can also be executed by dedicated logic circuitry—such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the device can also be implemented as dedicated logic circuitry.

[0104] Suitable processors for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as disks, magneto-optical disks, or optical disks, or the computer will be operatively coupled to such mass storage devices to receive data from or transfer data to them, or both. However, a computer is not required to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.

[0105] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.

[0106] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0107] The above detailed embodiments are a description of the present invention. It should not be considered that the specific embodiments of the present invention are limited to these descriptions. For those skilled in the art, several simple deductions and substitutions can be made without departing from the concept of the present invention, and all of these should be considered to fall within the protection scope of the present invention.

Claims

1. A method for automatically determining the structure of a compound based on a nuclear magnetic resonance spectrum, characterized by, The method comprises the following steps: S1, constructing a nuclear magnetic resonance chemical shift database; S2, converting 13C nuclear magnetic data in the database into a feature vector of 400; S3, converting the to-be-predicted compound into a feature vector of 400 using the method in step S2, and calculating the cosine similarity between the feature vector of the to-be-predicted compound and all feature vectors in the database; S4, predicting the structure of the to-be-predicted compound according to the calculation result of the cosine similarity; the prediction of the structure of the to-be-predicted compound according to the calculation result of the cosine similarity comprises: sorting all calculation results in descending order, and selecting the top 20 most similar feature spectra as candidate compounds corresponding to the to-be-tested nuclear magnetic data; if there is a nuclear magnetic spectrum with a similarity of 1 in the top 20 feature spectra, the compound corresponding to the nuclear magnetic spectrum is considered to be the compound structure of the to-be-tested nuclear magnetic data; if there is no nuclear magnetic spectrum with a similarity of 1 in the top 20 most similar feature spectra, the structures of the compounds corresponding to the top 20 feature spectra are split to obtain functional group fragments, and then the fragments are recombined; the chemical shift of the recombined compound is predicted, and the compound structure corresponding to the maximum similarity is taken as the compound structure of the to-be-tested nuclear magnetic data according to the predicted nuclear magnetic data.

2. The method for automatically determining the structure of a compound based on a nuclear magnetic resonance spectrum according to claim 1, wherein, The construction of the nuclear magnetic resonance chemical shift database comprises: collecting nuclear magnetic resonance chemical shift data, including compound structure encoding and 13C nuclear magnetic data of the compound.

3. The method of claim 1, wherein the method further comprises: determining a structure of the compound based on the NMR spectrum. The conversion of the 13C nuclear magnetic data in the database into a feature vector of 400 comprises: S21, removing duplicate 13C nuclear magnetic data and retaining only unique non-repeating data; S22, rounding off the 13C nuclear magnetic data; S23, adding 100 to the rounded-off nuclear magnetic data; S24, generating a vector of 400, with an index starting from 0, and setting the index position of the nuclear magnetic data to 1 and the remaining positions to 0.

4. The method of claim 3, wherein the method further comprises: determining the structure of the compound based on the NMR spectrum. The conversion of the 13C nuclear magnetic data in the database into a feature vector of 400 further comprises: S25, diffusing the 1 in the position of the nuclear magnetic data to both sides; if there are overlapping positions, the values are added.

5. The method of claim 1, wherein the method further comprises: The calculation of the cosine similarity between the feature vector of the to-be-predicted compound and all feature vectors in the database comprises: ​ The calculation formula is as follows: cos If the possible elements and the impossible elements in the to-be-predicted compound are known, the elements are screened to remove the compounds that are impossible to exist. = |A|·|B|(A·B) wherein A , B are the characteristic vector of the compound to be predicted and any one of the characteristic vectors in the database, respectively.

6. The method of claim 1, wherein the method further comprises: determining a structure of the compound based on the NMR spectrum. The system comprises:

7. A system for automatically determining the structure of a compound based on a nuclear magnetic spectrum, characterized by: a nuclear magnetic resonance chemical shift database; a feature vector calculation module configured to convert 13C nuclear magnetic data in the nuclear magnetic resonance chemical shift database into a feature vector of 400; a similarity calculation module configured to calculate the cosine similarity between the feature vector of the to-be-predicted compound and all feature vectors in the database; a compound structure prediction module configured to predict the structure of the to-be-predicted compound according to the calculation result of the cosine similarity; the prediction of the structure of the to-be-predicted compound according to the calculation result of the cosine similarity comprises: ​ All the calculation results are sorted in descending order, and the top 20 most similar feature spectra corresponding to the compounds are selected as the candidate compounds of the to-be-tested nuclear magnetic data; If there is a nuclear magnetic spectrum with a similarity of 1 in the top 20 feature spectra, the compound corresponding to the nuclear magnetic spectrum is considered to be the compound structure of the to-be-tested nuclear magnetic data; If there is no nuclear magnetic spectrum with a similarity of 1 in the top 20 most similar feature spectra, the structure of the compound corresponding to the top 20 feature spectra is split to obtain a functional group fragment, and then the fragment is recombined; the chemical shift of the recombined compound is predicted, and the similarity calculation is performed according to the predicted nuclear magnetic data, and the compound structure corresponding to the maximum similarity value is taken as the compound structure of the to-be-tested nuclear magnetic data.

8. A computer storage medium having stored thereon computer instructions, wherein the computer instructions comprise the steps of: The computer instructions perform the method of automatically determining the compound structure based on the nuclear magnetic spectrum in any one of claims 1-6.

9. A terminal comprising a memory and a processor, the memory having stored thereon computer instructions executable on the processor, wherein, The processor performs the method of automatically determining the compound structure based on the nuclear magnetic spectrum in any one of claims 1-6.

Citation Information

Patent Citations

  • Method and system for confirming structure of organic compound by carbon-13 nuclear magnetic resonance data

    CN103728330A

  • Method and system for determining structure of organic compound by using spectrum data

    CN115966262A