A method and system for determining the structure of an organic compound using spectral data

By mapping the spectral data and structure of organic compounds to a unified vector space using a deep learning model, the problem of time-consuming, labor-intensive, and inaccurate determination of organic compound structures in existing technologies is solved, achieving efficient and accurate determination of compound structures and large-scale database search.

CN115966262BActive Publication Date: 2026-04-17HANGZHOU CARBON SILICON SMART TECH DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU CARBON SILICON SMART TECH DEV CO LTD
Filing Date
2021-10-08
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, determining the structure of organic compounds using carbon NMR data is time-consuming, laborious, and not very accurate. Furthermore, it is limited by the number of reference organic compounds and the scientific nature of chemical shift matching, making it difficult to determine the structure of unknown compounds.

Method used

A training set is constructed using a deep learning model to establish the correspondence between organic compounds and spectral data. Through spectral data encoding models and organic compound encoding models, spectral data and compound structures are mapped to a unified multimodal vector space, and structure matching is performed using vector distance.

Benefits of technology

It improves the accuracy and efficiency of determining the structure of organic compounds, can handle any compound without being limited by reference compounds, is suitable for searching large-scale databases, and can handle compounds with different molecular complexities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115966262B_ABST
    Figure CN115966262B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for determining the structure of organic compounds using spectral data. The method includes: constructing a training set reflecting the correspondence between organic compound expressions and spectral data; initializing a spectral data encoding model and an organic compound encoding model; training the spectral data encoding model and the organic compound encoding model using the training set with a set loss function as the objective, and using a supervisory signal to select the correspondence between spectral data and the encoded organic compound expression; for target spectral data, inputting it into the trained spectral data encoding model to output a first vector, and inputting the organic compound expression into the trained organic compound expression encoding model to output a second vector; using the distance between the first and second vectors as a similarity measure between the spectral data and the organic compound molecular structure to determine the organic compound structure. This invention improves the accuracy and efficiency of determining the structure of organic compounds and has strong universality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of compound synthesis technology, and more specifically, to a method and system for determining the structure of organic compounds using spectral data. Background Technology

[0002] Spectroscopic data can be used to analyze the structure or state of organic compounds. For example, carbon-13 nuclear magnetic resonance (13C-NMR) data has advantages such as high accuracy, wide distribution range, low overlap, and easy identification, and has been widely used in compound structure analysis. However, the determination of the structure of unknown compounds using carbon-13 NMR data has traditionally been done manually, relying on personal experience or combining manual review and comparison of NMR data from literature. This manual approach is time-consuming, labor-intensive, and not very accurate.

[0003] In the prior art, patent application CN103728330A proposes a method and system for determining the structure of organic compounds using nuclear magnetic resonance carbon (NMR) spectroscopy data. The technical solution involves: pre-storing the structure of a reference organic compound, the corresponding NMR spectroscopy data for each reference organic compound, and the solvent used to obtain the NMR spectroscopy data for each reference organic compound; acquiring the NMR spectroscopy data, tolerance RC, and solvent used to obtain the NMR spectroscopy data for the organic compound to be tested; determining the number m of chemical shifts in the NMR spectroscopy data of the organic compound to be tested; selecting NMR spectroscopy data with the number of chemical shifts equal to m from the NMR spectroscopy data corresponding to the reference organic compounds to obtain first-stage selected NMR spectroscopy data; and then processing each of the first-stage selected NMR spectroscopy data... The chemical shifts in the carbon NMR spectrum data and the chemical shifts in the carbon NMR spectrum data of the organic compound to be tested are sorted according to the same sorting rules. The chemical shifts in each sorted primary screening carbon NMR spectrum data are compared one-to-one with the chemical shifts in the sorted carbon NMR spectrum data of the organic compound to be tested, obtaining the absolute values ​​of multiple chemical shift differences. The number S of absolute values ​​of chemical shift differences below RC is divided by m to obtain the matching rate between the reference organic compound and the organic compound to be tested for each primary screening carbon NMR spectrum data. The structure of the reference organic compound with a matching rate of 100% and using the same solvent as the solvent used to obtain the carbon NMR spectrum data of the organic compound to be tested is determined as the structure of the organic compound to be tested.

[0004] In the above technical solution, matching and screening are performed by comparing the carbon-13 NMR spectra of the organic compound to be tested with those of a reference organic compound. Therefore, the reference organic compound must be one whose corresponding carbon-13 NMR spectra have been obtained. However, obtaining carbon-13 NMR spectra requires calculation and instrumental measurement, which is not easy, and the proportion of organic compounds with obtainable carbon-13 NMR spectra among all known organic compounds is not large. This approach limits the number of reference organic compounds, and the organic compound to be tested is very likely not within the range of reference organic compounds, thus limiting the accuracy of determining the structure of the organic compound to be tested. Furthermore, organic compounds whose number of chemical shifts in the carbon-13 NMR spectra of the reference organic compound equals the number of chemical shifts in the carbon-13 NMR spectra of the organic compound to be tested are selected as primary screening compounds and then matched with the organic compound to be tested. This screening is not scientific in terms of "organic similarity" matching. If the reference organic compound does not include the organic compound to be tested, the primary screening products may all be compounds with the same number of samples as the organic compound to be tested, but whose structures are unrelated, resulting in no correct results. Customers' expectations for such systems may also include the desire for similar results if no correct results are found in the baseline library, to aid in the judgment. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and system for determining the structure of organic compounds using spectral data.

[0006] According to a first aspect of the present invention, a method for determining the structure of an organic compound using spectral data is provided. The method includes the following steps:

[0007] Step S1: Construct a training set that reflects the correspondence between the expressions of organic compounds and spectral data;

[0008] Step S2: Initialize the spectral data encoding model and the organic compound encoding model;

[0009] Step S3: Using the set loss function as the objective, train the spectral data encoding model and the organic compound encoding model using the training set, and supervise the correspondence between the selected signal spectral data and the organic compound expression encoding.

[0010] Step S4: For the target spectral data, input it into the trained spectral data encoding model and output a first vector. Input the organic compound expression into the trained organic compound expression encoding model and output a second vector. Use the distance between the first vector and the second vector as a similarity measure between the spectral data and the molecular structure of the organic compound to determine the structure of the organic compound.

[0011] According to a second aspect of the present invention, a system for determining the structure of an organic compound using spectral data is provided. The system comprises:

[0012] Data acquisition unit: used to construct a training set that reflects the correspondence between the expressions of organic compounds and spectral data;

[0013] Model building unit: used to initialize the spectral data encoding model and the organic compound encoding model;

[0014] Model training unit: used to train the spectral data encoding model and the organic compound encoding model using the training set with a set loss function as the objective, and to supervise the correspondence between the selected spectral data and the organic compound expression encoding;

[0015] Compound structure determination unit: For target spectral data, it inputs the data into a trained spectral data encoding model and outputs a first vector. For organic compound expressions, it inputs the data into a trained organic compound expression encoding model and outputs a second vector. The distance between the first and second vectors is used as a similarity measure between the spectral data and the molecular structure of the organic compound to determine the structure of the organic compound.

[0016] Compared with existing technologies, the advantages of this invention are that by using a deep learning model, the spectra and structural expressions of any organic compound can be mapped to a unified multimodal vector space and encoded into vectors for subsequent matching searches or obtaining synthetic products of multiple compounds, thereby improving the accuracy and efficiency of determining the structure of organic compounds and having strong universality.

[0017] Other features and advantages of the invention will become clear from the following detailed description of exemplary embodiments of the invention with reference to the accompanying drawings. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments of the invention and, together with their description, serve to explain the principles of the invention.

[0019] Figure 1 This is a schematic diagram of the overall process of determining the structure of an organic compound using nuclear magnetic resonance spectroscopy data according to an embodiment of the present invention;

[0020] Figure 2 This is a flowchart of a method for determining the structure of an organic compound using nuclear magnetic resonance spectroscopy data according to an embodiment of the present invention;

[0021] Figure 3 This is a schematic diagram of a deep learning model training process according to an embodiment of the present invention;

[0022] Figure 4 This is a schematic diagram of the structure of the C13-NMR encoding model and the SMILES expression encoding model according to an embodiment of the present invention;

[0023] Figure 5 This is a schematic diagram of a system for determining the structure of organic compounds using nuclear magnetic resonance spectroscopy data according to an embodiment of the present invention.

[0024] In the attached diagram: Feature matching; Reference SMILES Feature Library; input; SMILES string; Input Embedding; Positional Encoding; Multi-Head Attention; Feedforward; output vector. Detailed Implementation

[0025] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of the invention.

[0026] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use.

[0027] Techniques, methods, and devices known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be considered part of the specification.

[0028] In all the examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.

[0029] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0030] See Figure 1As shown, this invention utilizes a deep learning model to establish a multimodal joint representation of the nuclear magnetic resonance (NMR) spectral data and structural expressions of organic compounds. The NMR spectral data involved includes carbon NMR, proton NMR, mass spectrometry, or any combination thereof. The structural expressions of organic compounds include, but are not limited to, SMILES expressions (SMILES is a specification that explicitly describes the molecular structure using ASCII strings). Using this invention, the structural expression of organic compounds corresponding to spectral data can be accurately obtained without relying on the reference organic compound's carbon NMR spectral data.

[0031] In short, the method for determining the structure of organic compounds using spectral data includes: constructing a training set that reflects the correspondence (or pairwise relationship) between organic compound expressions and spectral data; initializing a spectral data encoding model and an organic compound encoding model, both of which can be deep learning models; training the spectral data encoding model and the organic compound encoding model using the training set with a set loss function as the objective, and using a supervisory signal to select the correspondence between spectral data and the encoded organic compound expression; for the target spectral data, inputting it into the trained spectral data encoding model to output a first vector, and inputting the organic compound expression into the trained organic compound expression encoding model to output a second vector; using the distance between the first vector and the second vector as a similarity measure between the spectral data and the molecular structure of the organic compound to determine the structure of the organic compound.

[0032] For clarity, the following explanation will primarily use carbon NMR spectroscopy data and SMILES expressions as examples.

[0033] Specifically, in combination Figure 2 and Figure 3 As shown, the method for determining the structure of organic compounds using carbon nuclear magnetic resonance (NMR) data includes the following steps:

[0034] Step S110: Construct a nuclear magnetic resonance carbon spectrum data encoding model to convert nuclear magnetic resonance carbon spectrum data into corresponding vectors.

[0035] In this step, a C13-NMR encoder model is established to encode NMR carbon spectrum data into vectors. For example, this model has a residual structure, consisting of stacked layers such as CNN (Convolutional Neural Network) and MLP (Multilayer Perceptron), with an input of (batch size, 4000) vectors and an output of (batch size, 768) vectors. Utilizing the residual structure, direct connections or shortcuts can be added to the model, thereby dynamically adjusting the model's complexity and enhancing gradient propagation. By leveraging the multilayer structure and nonlinear activation function of the multilayer perceptron, non-linearly separable data can be identified. The similarity between two vectors can be characterized using Euclidean distance or cosine similarity, etc.

[0036] After establishing the encoding model, the subsequent training process can first use randomly simulated carbon NMR spectrum data and data augmentation data to pre-train the model using infinity loss until a model that meets the set loss target is obtained (i.e., the model's weights and biases are determined).

[0037] Step S120: Construct an organic compound expression encoding model to convert organic compound expressions into corresponding vectors.

[0038] In one embodiment, a stacked multi-layer (e.g., 6-layer, 7-layer) transformer encoder structure is used as the encoder for SMILES expressions, encoding the SMILES expression strings into corresponding vectors. For example, the input is batch-size SMILES expression strings, and the output is a vector of (batch-size, 768). Similarly, cosine similarity can be used to characterize the similarity between two vectors. Subsequent model training can be performed using, for example, 100 million organic compound SMILES expressions and augmented data, employing infinity loss.

[0039] Step S130: The structure of the organic compound corresponding to the target spectral data is determined using the trained nuclear magnetic resonance carbon spectrum data encoding model and the organic compound expression encoding model.

[0040] For example, loading the C13-NMR encoder and SMILES expression encoder model, such as Figure 3 As shown, paired data of NMR carbon spectrum data and SMILES expression are fed into a deep learning model, and the model is trained using infinity loss, thereby forming a matching relationship between NMR carbon spectrum data and SMILES expression.

[0041] The specific process formula is as follows:

[0042] *****************************************

[0043] S_f=l2_normalize(smiles_encoder(S))#[n,d_s]

[0044] C_f=l2_normalize(nmr_encoder(C))#[n,d_n]

[0045] logits=np.exp(t)*np.dot(S_f,C_f.T)

[0046] label = np.arrange(n)

[0047] loss_s=cross_entropy_loss(logits,label,axis=0)

[0048] loss_c=cross_entropy_loss(logits,label,axis=1)

[0049] loss = (loss_s + loss_c) / 2

[0050] *********************************************

[0051] During model training, the supervisory signal is selected based on the correspondence between spectral data and the organic compound expression encoding. If there is a correspondence, it is considered as 1; otherwise, it is considered as 0 (or -1). Furthermore, preferably, the training process can be divided into two steps: first, pre-training is performed using large-scale pairwise data obtained through computation to obtain the aforementioned spectral data encoding model and organic compound expression encoding model; then, pairwise data collected from real experiments is used for model fine-tuning to obtain the final model.

[0052] Furthermore, after model training is complete, it is possible to map any organic compound's carbon spectrum and its SMILES structural expression to a unified multimodal vector space, encoding them into vectors for matching searches. For example, inputting carbon spectrum data into the trained spectral data encoding model results in a vector output; inputting an organic compound expression into the trained organic compound expression encoding model also results in a vector output; the distance between these two vectors is calculated and used as a measure of similarity between the carbon spectrum data and the organic compound's molecular structure.

[0053] It should be understood that the models mentioned above can employ various types of network structures. The first layer of the spectral data encoding model is a convolutional structure (Kernel is K, Stride is S, Pad is P), followed by N layers of fully connected structures with residual connections. For example... Figure 4 (a) is the C13-NMR encoder structure, which consists of an input layer, a convolutional layer, a flattened layer, a fully connected layer (FC), and an output layer. Figure 4 (b) is the expression encoder structure, whose input is the SMILES string, which is processed by the input embedding layer, the multi-head attention layer and the add&norm layer, and outputs a vector form.

[0054] Accordingly, the present invention also provides a system for determining the structure of organic compounds using spectral data, for implementing one or more aspects of the above-described method. For example, the system includes a front-end and a back-end. The front-end provides a user interface for selecting the form of the organic compound to be tested, and the back-end determines the structure of the organic compound based on the user's selection. In practical applications, the front-end and back-end can be implemented on a single computer or server, or the client can implement the front-end functionality while the server implements the back-end functionality; the present invention does not impose any limitations on this.

[0055] To achieve matching retrieval using a trained model, a search library can be pre-built. For example, multiple organic compound expressions can be pre-input into an organic compound expression encoding model to obtain multiple vectors as the search library. Spectral data with determined molecular structures can be fed into a spectral encoding model to obtain a feature vector for initiating a query. Finally, the feature vector for initiating the query can be compared with the multiple vectors in the search library to calculate similarity. The top N most similar results are then sorted by similarity measure and output, and these N results are considered to be the compounds most likely to correspond to the spectral data.

[0056] Specifically, in combination Figure 5 As shown, in one embodiment, the provided system performs the following process.

[0057] Step S210: Establish a reference organic compound library. Encode all compounds in the library into vectors using a SMILE expression encoder to create a SMILE expression library. Similarly, encode the NMR carbon spectrum data of reference organic compounds with NMR carbon spectrum data into vectors using a C13-NMR encoder to create an NMR carbon spectrum library. Then, input this data into the Milvus vector search engine.

[0058] Step S220. In the system, select whether to input the organic compound to be tested in the form of a SMILE expression or NMR carbon spectrum data, and complete the input. The deep learning model completes the data to vector conversion. Select whether to search in the SMILE expression library or the NMR spectrum library. Then, using cosine similarity as a reference, select the molecules whose ranking is in the top N, and click query to find the results.

[0059] Step S230: In the system, you can also choose to perform similarity calculation between vectors. Input the contents of the two data to be compared and select the data type (SMILES expression or NMR carbon spectrum data), and then click Calculate to perform similarity calculation.

[0060] Step S240: End the program.

[0061] In summary, compared with the prior art, the present invention has at least the following technical effects:

[0062] 1) By using a trained deep learning model, the carbon spectrum and structural SMILES expression of any organic compound can be mapped to a unified multimodal vector space, encoded into vectors, and used for matching searches. In this way, an organic compound can become a reference organic compound without needing a known carbon spectrum, so in principle, reference organic compounds can encompass all known compounds in the world, without being restricted by other conditions.

[0063] 2) This invention can not only search from carbon spectrum of organic compounds to SMILES expression of organic compound structure, but also search from carbon spectrum of organic compounds to carbon spectrum of organic compounds, from SMILES expression of organic compound structure to SMILES expression of organic compound structure, and from SMILES expression of organic compound structure to carbon spectrum of organic compounds. Furthermore, it can perform similarity detection of any number of organic compounds based on the features of any carbon spectrum or SMILES expression of a compound.

[0064] 3) Traditional methods for determining the structure of organic compounds vary in computation time depending on the molecular complexity or the number of chemical shifts in the carbon NMR spectrum. However, the deep learning feature extraction method used in this invention requires the same processing time for different molecules.

[0065] 4) This invention can be used to build an open-source search engine, such as a search platform using Milvus, which can achieve a search volume of hundreds of millions of records and a response time of milliseconds.

[0066] 5) This invention defaults to matching and searching among all reference organic compounds, but can also perform certain screening according to customer requirements before matching and searching.

[0067] 6) This invention utilizes deep learning to enable the model to learn features that are difficult to describe, such as the high similarity between different molecules with similar skeletal structures, thus avoiding the higher similarity of identical chemical shifts in the carbon NMR spectra of organic compounds compared to other traditional methods. It can also effectively identify the presence of small amounts of impurities.

[0068] 7) This invention can perform certain organic compound generation reasoning. For example, by merging the carbon spectra of two compounds and inputting them, there is a certain probability of obtaining the synthetic product of the two compounds.

[0069] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.

[0070] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0071] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0072] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, Python, etc., and conventional procedural programming languages ​​such as "C" or similar languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.

[0073] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0074] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0075] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0076] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. It will be known to those skilled in the art that implementation in hardware, implementation in software, and implementation using a combination of software and hardware are equivalent.

[0077] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the embodiments in the market, or to enable other those skilled in the art to understand the embodiments disclosed herein. The scope of the invention is defined by the appended claims.

Claims

1. A method for determining the structure of an organic compound using spectral data, comprising the following steps: Step S1: Construct a training set that reflects the correspondence between the expressions of organic compounds and spectral data; Step S2: Initialize the spectral data encoding model and the organic compound encoding model; Step S3: Using the set loss function as the objective, train the spectral data encoding model and the organic compound encoding model using the training set, and supervise the correspondence between the selected signal spectral data and the organic compound expression encoding. Step S4: For the target spectral data, input it into the trained spectral data encoding model and output a first vector. Input the organic compound expression into the trained organic compound expression encoding model and output a second vector. Use the distance between the first vector and the second vector as a similarity measure between the spectral data and the molecular structure of the organic compound to determine the structure of the organic compound. The method further includes: Multiple organic compound expressions are pre-input into a trained organic compound expression encoding model to obtain multiple vectors as the search base library; The spectral data that determines the molecular structure is fed into a trained spectral data encoding model to obtain the feature vector for initiating the query; The similarity measurement is calculated between the feature vector that initiates the query and multiple vectors in the base library to be searched. The results are then sorted by similarity measurement and the top N most similar results are output. These N results are considered to be the compounds that most likely correspond to the spectral data.

2. The method according to claim 1, characterized in that, The loss function used is InfoNCE loss.

3. The method according to claim 1, characterized in that, Step S3 includes the following sub-steps: The constructed spectral data encoding model and organic compound expression encoding model are pre-trained using paired data of large-scale organic compound expressions and spectral data obtained through calculation; The pre-trained model was fine-tuned using paired data collected from real experiments to obtain the final spectral data encoding model and organic compound encoding model.

4. The method according to claim 1, characterized in that, In step S4, the distance function is either the vector inner product normalized to the modulus or the cosine distance.

5. The method according to claim 1, characterized in that, The spectral data is one or more of the following: carbon spectrum, proton spectrum, and mass spectrum.

6. The method according to claim 1, characterized in that, The organic compound expression is the SMILES expression.

7. The method according to claim 1, characterized in that, The spectral data encoding model is a deep learning model, with a convolutional structure in the first layer followed by multiple fully connected structures with residual connections.

8. A system for determining the structure of an organic compound using spectral data, comprising: Data acquisition unit: used to construct a training set that reflects the correspondence between the expressions of organic compounds and spectral data; Model building unit: used to initialize the spectral data encoding model and the organic compound encoding model; Model training unit: used to train the spectral data encoding model and the organic compound encoding model using the training set with a set loss function as the target, and to supervise the correspondence between the selected spectral data and the organic compound expression encoding; Compound structure determination unit: For target spectral data, it inputs it into a trained spectral data encoding model and outputs a first vector. For organic compound expression, it inputs it into a trained organic compound expression encoding model and outputs a second vector. The distance between the first vector and the second vector is used as a similarity measure between the spectral data and the molecular structure of the organic compound to determine the structure of the organic compound. The system is also used for: Multiple organic compound expressions are pre-input into a trained organic compound expression encoding model to obtain multiple vectors as the search base library; The spectral data that determines the molecular structure is fed into a trained spectral data encoding model to obtain the feature vector for initiating the query; The similarity measurement is calculated between the feature vector that initiates the query and multiple vectors in the base library to be searched. The results are then sorted by similarity measurement and the top N most similar results are output. These N results are considered to be the compounds that most likely correspond to the spectral data.

9. A computer-readable storage medium having a computer program stored thereon, wherein, When executed by a processor, the program implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and system for confirming structure of organic compound by carbon-13 nuclear magnetic resonance data

    CN103728330A

  • Gene expression profile feature learning method based on auto-encoder

    CN111276187A

  • Spectral decomposition-based deep learning magnetic resonance spectrum reconstruction method

    CN113143243A