Compound structure analysis method and system based on nuclear magnetic resonance data

By preprocessing nuclear magnetic resonance data and using automated solvent peak identification, the problem of poor accuracy in compound structure analysis was solved, and efficient and accurate compound structure analysis was achieved.

CN121994856APending Publication Date: 2026-05-08赣江中药创新中心
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
赣江中药创新中心
Filing Date
2026-01-21
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing NMR data processing methods suffer from poor accuracy in compound structure analysis, particularly due to insufficient solvent correction precision and heavy reliance on manual comparison of multidimensional spectra, leading to errors in chemical shift extraction and long processing times.

Method used

By receiving NMR data, preprocessing it to convert it into frequency domain form, constructing solvent peak templates, matching and removing solvent peaks, calculating chemical shifts, and performing automated structure analysis using a chemical shift list, automatic phase correction and baseline correction algorithms are used to remove noise interference, and solvent peak templates are generated using the Lorentz function and matched with solvent peaks using cosine similarity.

Benefits of technology

It improves the accuracy of solvent correction, reduces the phenomenon of mislabeling and omission of solvent peaks, and enhances the accuracy and efficiency of compound structure analysis, replacing the traditional method of manually comparing multidimensional spectra.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121994856A_ABST
    Figure CN121994856A_ABST
Patent Text Reader

Abstract

The invention discloses a compound structure analysis method and system based on nuclear magnetic resonance data, and the method comprises the steps: receiving original nuclear magnetic resonance data, collected by a nuclear magnetic instrument, of a to-be-analyzed compound, preprocessing the original nuclear magnetic resonance data so as to convert the nuclear magnetic resonance data in a time domain form into nuclear magnetic resonance data in a frequency domain form, suppress noise interference and output a peak position list; constructing a solvent peak template matched with an actual spectrogram peak shape in the nuclear magnetic resonance data, delimiting candidate areas in the peak position list according to a configured search interval, and matching a signal of each candidate area with a signal of the solvent peak template to obtain a solvent peak; removing the solvent peak, calculating the offset between the actual chemical shift and the theoretical value, applying the offset to global chemical shift correction to obtain a chemical shift list, and performing structural analysis on the to-be-analyzed compound based on the chemical shift list. According to the invention, the problem of poor accuracy during compound structure analysis in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of nuclear magnetic resonance data processing technology, and in particular to a method and system for analyzing the structure of compounds based on nuclear magnetic resonance data. Background Technology

[0002] Current NMR data processing primarily relies on manual expert operation or semi-automated software tools. In terms of structural analysis, traditional techniques depend on expert experience comparing spectral databases (such as MicroSpectrum and PubChem) or rule-based systems (such as ACD / Labs), resulting in inefficiency and significant susceptibility to subjective factors. While recent studies have attempted to use machine learning for spectral classification, these are mostly limited to single spectral types and lack cross-modal (e.g., 13 (Mixed C NMR and HSQC spectra) matching capability.

[0003] Existing technologies suffer from the following main drawbacks: 1) Insufficient solvent correction accuracy: Current automated methods are ineffective at correcting compounds in certain solvents (such as deuterated methanol), leading to errors in chemical shift extraction. For example, when using the DP5 algorithm to process samples from *Sophora japonica*, mislabeling or omission of solvent peaks occurs. 2) High dependence on manual intervention for structural analysis: Experts need to manually compare multidimensional spectra, which is time-consuming and has poor repeatability. Therefore, current methods for compound structural analysis suffer from poor accuracy. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method and system for determining the structure of compounds based on nuclear magnetic resonance data, which aims to solve the problem of poor accuracy in existing compound structure determination methods.

[0005] This invention proposes a method for determining the structure of compounds based on nuclear magnetic resonance data, the method comprising: Receive raw NMR data of the compound to be analyzed from the NMR instrument, and preprocess the raw NMR data to convert the time-domain NMR data into the frequency-domain data, suppress noise interference, and output a list of peak positions; A solvent peak template matching the actual peak shape in the NMR data is constructed. Candidate regions are defined in the peak position list according to the configured search interval. The signal of each candidate region is matched with the signal of the solvent peak template to obtain the solvent peak. The solvent peaks are removed and the offset between the actual chemical shift and the theoretical value is calculated. The offset is applied to the global chemical shift correction to obtain a chemical shift list, and the structure of the compound to be analyzed is performed based on the chemical shift list.

[0006] Furthermore, in the above-mentioned compound structure analysis method based on nuclear magnetic resonance (NMR) data, the steps of preprocessing the raw NMR data to convert the time-domain NMR data into the frequency-domain data, suppressing noise interference, and outputting a peak position list include: The original nuclear magnetic resonance data was converted into a frequency domain spectrum by Fourier transform, and then the baseline drift was removed by automatic phase correction algorithm and baseline correction algorithm. Peaks in the frequency domain spectrum are identified by combining second derivative mutation point detection with local maximum search, and a list of peak positions is output.

[0007] Furthermore, in the above-mentioned compound structure determination method based on NMR data, the step of constructing a solvent peak template that matches the actual peak shape in the NMR data includes: The actual spectrum data is extracted from the nuclear magnetic resonance data within the target chemical shift range, and the global half width at half maximum (FWHM) is calculated as the peak shape parameter. Based on the theoretical coupling constant value corresponding to the solvent, a solvent peak template is generated by combining it with the Lorentz function; or Based on the statistical average template of historical measured solvent spectra, the corresponding solvent peak templates are obtained.

[0008] Furthermore, in the above-mentioned compound structure analysis method based on NMR data, the expression for generating solvent peak templates by combining the theoretical coupling constant value corresponding to the solvent with the Lorentz function is as follows: ; ; in, This represents the ratio of peak intensity. L This represents the half-width and height calculated in the previous step. x 0 is the center position of the current peak. J It is the coupling constant value corresponding to the solvent peak.

[0009] Furthermore, in the above-mentioned compound structure analysis method based on nuclear magnetic resonance data, the step of matching the signal of each candidate region with the signal of the solvent peak template to obtain the solvent peak includes: The cosine similarity between the signal of each candidate region and the signal of the solvent peak template is calculated, and the solvent peak is obtained based on the similarity. The formula for calculating cosine similarity is: ; in, Information for each point in the solvent peak template. This represents the information of points selected in the candidate region for each experiment; For methanol-based solvents, the top three candidate peaks with similarity values ​​higher than the threshold are selected, and the presence of quintet splitting characteristics and integral intensity constraints are further verified to determine the solvent peaks.

[0010] Furthermore, in the above-mentioned compound structure determination method based on nuclear magnetic resonance data, the step of performing structure determination of the compound to be determined based on the chemical shift list includes: In the list of calculated chemical shifts 13 The C peak list and the compounds in the constructed spectral shift list library 13 C-peak list difference index; The target compound with the smallest difference is identified based on the difference index, and structural analysis is performed based on the target compound; or Based on a chemical shift list, a pre-trained structure generation model is used to perform structural analysis on the compound to be analyzed. The structure generation model is trained using a contrastive learning framework or a graph neural network; or... The chemical shift list is input into the structure generation model and the spectral shift list library retrieval algorithm, respectively. The results are output separately and then comprehensively scored based on preset weights to obtain the final structure analysis result.

[0011] Furthermore, in the above-mentioned compound structure analysis method based on nuclear magnetic resonance data, the formula for calculating the difference index is as follows: ; in, and These are compounds from the chemical shift list and the spectral shift list library, respectively. 13 Number of C peaks Indicates in The number of peak values ​​matched within the threshold. Indicates the experimentally measured number of... i indivual 13 C chemical shift value, This indicates the predicted number of compounds in the spectral shift list library. j indivual 13 C chemical shift value.

[0012] Another object of the present invention is to provide a compound structure determination system based on nuclear magnetic resonance data, the system comprising: The acquisition module is used to receive the raw NMR data of the compound to be analyzed from the NMR instrument, and to preprocess the raw NMR data to convert the time-domain NMR data into the frequency-domain data, suppress noise interference, and output a list of peak positions. The construction module is used to construct a solvent peak template that matches the actual peak shape in the NMR data. Based on the configured search interval, candidate regions are defined in the peak position list, and the signal of each candidate region is matched with the signal of the solvent peak template to obtain the solvent peak. The analysis module is used to remove solvent peaks and calculate the offset between the actual chemical shift and the theoretical value. The offset is applied to global chemical shift correction to obtain a chemical shift list, and the structure of the compound to be analyzed is performed based on the chemical shift list.

[0013] Another object of the present invention is to provide a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.

[0014] Another object of the present invention is to provide an electronic device including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the program to implement the steps of the method described above.

[0015] This invention receives raw NMR data of a compound to be analyzed from an NMR instrument, preprocesses the raw NMR data to convert the time-domain NMR data into the frequency-domain data, suppresses noise interference, and outputs a peak position list. It constructs a solvent peak template that matches the actual peak shape in the NMR data, delineates candidate regions in the peak position list according to a configured search interval, and matches the signal of each candidate region with the signal of the solvent peak template to obtain solvent peaks. Solvent peaks are removed, and the offset between the actual chemical shift and the theoretical value is calculated. This offset is applied to global chemical shift correction to obtain a chemical shift list, and the structure of the compound to be analyzed is performed based on this list. The invention effectively avoids mislabeling and omission of solvent peaks by constructing a solvent peak template that matches the actual peak shape and combining it with precise matching of candidate regions, and then completes global chemical shift correction by calculating the offset, thus improving the accuracy of solvent correction. Through an automated process of spectrum preprocessing, solvent peak identification, and chemical shift correction, an accurate chemical shift list is output for structural analysis, replacing the traditional method of manually comparing multidimensional spectra. This solves the problem of poor accuracy in existing compound structure analysis methods. Attached Figure Description

[0016] Figure 1 This is a flowchart of the compound structure analysis method based on nuclear magnetic resonance data in the first embodiment of the present invention; Figure 2 This is a structural block diagram of a compound structure analysis system based on nuclear magnetic resonance data in the third embodiment of the present invention.

[0017] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation

[0018] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.

[0019] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0021] Example 1 Please see Figure 1 The figure shows a compound structure determination method based on nuclear magnetic resonance data in the first embodiment of the present invention, the method including steps S10 to S12.

[0022] Step S10: Receive the raw NMR data of the compound to be analyzed acquired by the NMR instrument, and preprocess the raw NMR data to convert the time-domain NMR data into the frequency-domain data, suppress noise interference, and output a list of peak positions.

[0023] The process involves receiving raw NMR data about the compound to be analyzed from the NMR instrument. This raw data is usually stored in the form of digital time-domain signals and contains core information such as the resonance frequency and intensity of the compound's atomic nuclei. However, it inevitably contains useless signals such as instrument noise, baseline drift, and solvent interference.

[0024] The raw NMR data is preprocessed. The core purpose of this preprocessing is to transform the time-domain NMR data into a frequency-domain format that is easy to analyze, while suppressing various noise interferences. The final output is a list of peak positions that can accurately reflect the characteristics of the compound. The list of peak positions must clearly record key parameters such as the chemical shift coordinates, peak intensity, and full width at half maximum (FWHM) of each peak, so as to provide high-quality basic data for subsequent solvent peak identification and structural analysis. Specifically, the raw NMR data is converted into a frequency domain spectrum using a Fast Fourier Transform (FFT). Compared to the traditional Discrete Fourier Transform (DFT), the FFT significantly reduces computational complexity; when processing 1024-point data, the FFT's computational efficiency is more than 100 times that of the DFT. Time-domain signals directly reflect changes in signal intensity over time and cannot be directly used for peak identification and chemical shift extraction. The Fourier Transform, however, can convert this time-domain signal into a frequency-domain signal, allowing the characteristic peaks of compounds to be presented in the form of intensities corresponding to specific chemical shifts, facilitating subsequent analysis.

[0025] After conversion to the frequency domain spectrum, automatic phase correction and baseline correction algorithms are used to remove baseline drift. Phase correction employs a zero-order / first-order joint correction algorithm based on peak symmetry, which can effectively handle the distortion problems of broad peaks and overlapping peaks in complex spectra. Baseline correction uses wavelet transform combined with morphological filtering technology, eliminating the need for manual parameter adjustment and exhibiting stronger generalization ability. During NMR data acquisition, due to factors such as instrument stability and sample shimming effects, phase shifts and baseline drift may occur in the spectrum. Phase shifts lead to peak shape distortion, while baseline drift affects the accuracy of peak intensity judgment and peak identification. The automatic phase correction algorithm can automatically adjust the phase of the spectrum to restore peak symmetry, while the baseline correction algorithm can eliminate baseline drift, keeping the spectrum baseline stable.

[0026] Subsequently, peak identification is performed on the frequency domain spectrum using second-derivative abrupt change point detection combined with local maximum search. Second-derivative abrupt change point detection can accurately capture locations in the spectrum where the signal intensity changes drastically; these locations are usually the edges of peaks. Combined with local maximum search, the vertex position of each peak can be accurately located, thereby determining the specific chemical shift coordinates of the peak. For overlapping peaks, this method achieves separation and identification by detecting multiple extrema of the second derivative, ultimately outputting a list of peak positions containing parameters such as peak chemical shift, peak intensity, and full width at half maximum (FWHM).

[0027] Through step-by-step preprocessing, various interference factors in the original data were gradually eliminated, transforming the time-domain signal into a clear and accurate frequency-domain spectrum and extracting comprehensive peak position information. The entire preprocessing process was automated, requiring no manual intervention, thus avoiding subsequent analytical errors caused by improper preprocessing parameter settings. This improved the reliability and repeatability of the entire analysis process, laying a solid foundation for the accurate identification of solvent peaks and the accurate extraction of chemical shifts. Step S11: Construct a solvent peak template that matches the actual peak shape in the NMR data. Delineate candidate regions in the peak position list according to the configured search interval. Match the signal of each candidate region with the signal of the solvent peak template to obtain the solvent peak.

[0028] The construction of solvent peak templates that match the actual peak shapes in the NMR data is crucial for accurate solvent peak identification. This is because different solvents exhibit significant differences in peak shape and chemical shift range, and the peak shape in actual experiments can be affected by factors such as instrument field strength, test temperature, and sample concentration. Candidate regions are defined in the previously output peak position list based on a pre-configured search interval. This search interval is set based on the characteristic chemical shift range of commonly used deuterated solvents, which narrows the subsequent matching range, reduces computational load, and improves processing efficiency. Then, the signal from each candidate region is meticulously matched with the signal from the solvent peak template. A specific matching algorithm is used to filter out signals that match the characteristics of the solvent peak, thus obtaining the solvent peak.

[0029] In practice, as one method of solvent peak template construction, the actual spectral data is first extracted from the NMR data within the target chemical shift range. The target chemical shift range is predetermined based on the known chemical shift characteristics of different solvents. Extracting data within this range can focus on the region where solvent peaks may exist, thus improving the targeting of template construction.

[0030] Then, the global half-width at half-maximum (HWHM) is calculated as a peak shape parameter. HWHM refers to the width of a peak when its intensity is half its peak value. The global HWHM is obtained by averaging the HWHMs of all significant peaks in the spectrum, reflecting the overall width and narrowness characteristics of the peak under experimental conditions, and is an important parameter for describing peak shape. Next, based on the theoretical coupling constant value corresponding to the solvent, a solvent peak template is generated by combining it with the Lorentz function. Different solvents have specific theoretical coupling constants, and the Lorentz function can accurately describe the linear characteristics of solvent peaks. The combination of the two ensures that the generated template can closely match the peak shape of the actual solvent peak.

[0031] For example, the expression for generating a solvent peak template based on the theoretical coupling constant value corresponding to the solvent and the Lorentz function is as follows: ; ; in, This represents the ratio of peak intensity. L This represents the half-width and height calculated in the previous step. x 0 is the center position of the current peak. J It is the coupling constant value corresponding to the solvent peak.

[0032] Solvent peaks typically exhibit specific splitting characteristics due to HD coupling. For example, the solvent peaks of deuterated methanol and deuterated dimethyl sulfoxide are quintets, while the solvent peak of deuterated chloroform is a triplet. Therefore, a combination of two formulas is needed to generate a complete solvent peak template. The first formula describes the line shape characteristics of a single peak. This formula calculates the intensity values ​​of a single peak at different positions, forming the Lorentz line shape of the single peak. This line shape accurately simulates the natural peak shape of the NMR signal. The second formula then stitches together the line shapes of the single peaks to form a template that conforms to the solvent peak splitting characteristics. This formula, based on the splitting rules of solvent peaks, superimposes the line shapes of the single peaks at different chemical shift positions according to a specific intensity ratio. For example, the intensity ratio of a quintet is 1:2:3:2:1. Through the calculation using this formula, a template signal that perfectly matches the actual solvent peak splitting characteristics is finally generated.

[0033] The generation process of solvent peak templates was quantified using a clear mathematical formula. Each parameter in the formula has a defined physical meaning and specific value range, allowing for flexible adjustment based on actual experimental data and solvent characteristics. The generated solvent peak templates not only match the peak shape profile but also accurately reproduce the splitting characteristics and intensity ratios, exhibiting a high degree of consistency with the actual solvent peak shape and splitting features. This solves the misidentification problem caused by traditional templates that only match peak positions, providing a precise reference for subsequent accurate solvent peak matching and significantly improving the accuracy of solvent peak identification. As another implementation method for constructing solvent peak templates, the corresponding solvent peak template can also be obtained based on the statistical average template of historical measured solvent spectra. For example, at least 50 sets of measured spectrum data of the same solvent under the same experimental conditions can be collected. After excluding outlier data, these spectrum data are normalized in intensity and aligned in coordinates, and then statistically averaged to eliminate random errors that may exist in a single experiment. The resulting statistically averaged template can better adapt to various variations in actual experiments, improving the robustness of the template. For hygroscopic solvents such as deuterated dimethyl sulfoxide and deuterated acetone, the peak shape information of characteristic water peaks can be added to the template to further improve the identification accuracy.

[0034] In this invention, two flexible methods for constructing solvent peak templates are provided. The first method, based on theoretical calculations and actual spectral parameters, can quickly generate highly targeted templates suitable for new solvents or special experimental conditions. The second method, based on a large amount of historical experimental data, has stronger practicality and anti-interference capabilities. The two methods can be selected according to the actual experimental scenario. By accurately matching the peak shape characteristics and chemical shift range of the actual spectrum, the matching degree between the solvent peak template and the actual spectrum peak shape is effectively improved, solving the problem of traditional solvent peak identification being easily affected by experimental conditions, and providing a guarantee for the accurate identification of subsequent solvent peaks.

[0035] Finally, the cosine similarity between the signal of each candidate region and the signal of the solvent peak template is calculated. Cosine similarity is an important indicator for measuring the similarity between two signal vectors, and its calculation formula is as follows: ; in, Information for each point in the solvent peak template. This represents the information of points selected in the candidate region for each experiment.

[0036] The similarity value between the candidate region signal and the template signal is obtained by calculating using this formula. The closer the similarity value is to 1, the higher the degree of similarity between the two, and the more likely it is to be a solvent peak.

[0037] For methanol-based solvents, the solvent peak exhibits a typical quintet splitting characteristic, and similarity screening alone may lead to misjudgment; therefore, further verification is required. The top three candidate peaks with similarity values ​​exceeding a preset threshold are selected. Then, it is verified whether these three candidate peaks exhibit the quintet splitting characteristic, i.e., whether the peak intensity ratio conforms to a 1:2:3:2:1 pattern. This is determined by calculating the intensity ratio of adjacent peaks, allowing an error range of ±10%. Simultaneously, integrated intensity constraint verification is performed. Integrated intensity refers to the peak area, calculated by integrating the peak signal intensity. An integrated intensity threshold is set at 5 times the integrated intensity of the strongest peak of the target compound in the sample. Only when the integrated intensity of a candidate peak exceeds this threshold is it confirmed as a solvent peak. These two layers of verification further eliminate interfering signals, ensuring the accuracy of solvent peak identification.

[0038] Step S12: Remove the solvent peak and calculate the offset between the actual chemical shift and the theoretical value. Apply the offset to the global chemical shift correction to obtain a chemical shift list, and perform structural analysis on the compound to be analyzed based on the chemical shift list.

[0039] Specifically, the identified solvent peaks are removed from the spectrum to avoid interference with the structural information of the target compound. Then, the offset between the actual and theoretical chemical shifts is calculated, and this offset is applied to global chemical shift correction to obtain an accurate chemical shift list. Chemical shifts are a core parameter for compound structure analysis, and their accuracy directly affects the reliability of subsequent functional group identification and chemical bond connection determination. Based on the corrected chemical shift list, appropriate structural analysis algorithms are used to analyze the structure of the compound, ultimately obtaining information such as the molecular structure and functional group composition.

[0040] In summary, the compound structure determination method based on NMR data in the above embodiments of the present invention receives raw NMR data of the compound to be determined from an NMR instrument, preprocesses the raw NMR data to convert the time-domain NMR data into the frequency-domain data, suppresses noise interference, and outputs a peak position list; constructs a solvent peak template that matches the actual peak shape in the NMR data, delineates candidate regions in the peak position list according to the configured search interval, matches the signal of each candidate region with the signal of the solvent peak template to obtain the solvent peak; removes the solvent peak and calculates the actual chemical shift. The offset from the theoretical value is applied to global chemical shift correction to obtain a chemical shift list, which is then used for structural analysis of the compound to be analyzed. Solvent peaks are identified by constructing a template matching the actual spectrum peak shape and combining it with precise candidate region matching. Global chemical shift correction is then performed by calculating the offset, effectively avoiding mislabeling and omission of solvent peaks and improving the accuracy of solvent correction. Through an automated process of spectrum preprocessing, solvent peak identification, and chemical shift correction, an accurate chemical shift list is output for structural analysis, replacing the traditional method of manually comparing multidimensional spectra. This solves the problem of poor accuracy in existing compound structure analysis methods.

[0041] Example 2 This embodiment also proposes a compound structure analysis method based on nuclear magnetic resonance (NMR) data. The difference between the compound structure analysis method based on NMR data in this embodiment and the compound structure analysis method based on NMR data in Embodiment 1 is as follows: The step of matching the signal of each candidate region with the signal of the solvent peak template to obtain the solvent peak includes: The cosine similarity between the signal of each candidate region and the signal of the solvent peak template is calculated, and the solvent peak is obtained based on the similarity. The formula for calculating cosine similarity is: ; in, Information for each point in the solvent peak template. This represents the information of points selected in the candidate region for each experiment; For methanol-based solvents, the top three candidate peaks with similarity values ​​higher than the threshold are selected, and the presence of quintet splitting characteristics and integral intensity constraints are further verified to determine the solvent peaks.

[0042] Furthermore, the step of performing structural analysis on the compound to be analyzed based on the chemical shift list includes: In the list of calculated chemical shifts 13 The C peak list and the compounds in the constructed spectral shift list library13 C-peak list difference index; The target compound with the smallest difference is identified based on the difference index, and structural analysis is performed based on the target compound; or Based on a chemical shift list, a pre-trained structure generation model is used to perform structural analysis on the compound to be analyzed. The structure generation model is trained using a contrastive learning framework or a graph neural network; or... The chemical shift list is input into the structure generation model and the spectral shift list library retrieval algorithm, respectively. The results are output separately and then comprehensively scored based on preset weights to obtain the final structure analysis result.

[0043] The first implementation method is a basic retrieval method, which calculates the chemical shift list. 13 The C peak list and the compounds in the constructed spectral shift list library 13 The C-peak list difference index, a spectral shift list library, is a pre-built database containing chemical shift data for known compounds, derived from experimental measurements and theoretical calculations. The difference index quantifies experimentally measured... 13 C peak list and compounds in the library 13 The degree of difference in the C-peak list. Based on the magnitude of the difference index, the target compound with the smallest difference is identified. The smaller the difference index, the closer the experimental data is to the data of compounds in the library, and the higher the structural similarity between the target compound and the compound to be resolved. Structural analysis is then performed based on the known structure of the target compound to infer the structure of the compound to be resolved.

[0044] Specifically, the formula for calculating the difference index is: ; in, and These are compounds from the chemical shift list and the spectral shift list library, respectively. 13 Number of C peaks Indicates in The number of peak values ​​matched within the threshold. Indicates the experimentally measured number of... i indivual 13 C chemical shift value, This indicates the predicted number of compounds in the spectral shift list library. j indivual 13 C chemical shift value.

[0045] The second implementation method is an intelligent matching algorithm. Based on a chemical shift list, a pre-trained structure generation model is used to analyze the structure of the compound to be analyzed. This model is pre-trained using a large amount of training data, including the chemical shift list of the compound and its corresponding real structural information. The structure generation model can be trained using a contrastive learning framework. This framework encodes the chemical shift data and compound structure into 128-dimensional vectors by constructing a spectral encoder and a structure encoder, respectively, and calculates cosine similarity in the latent space to achieve structure prediction. Alternatively, it can be trained using a graph neural network. The graph neural network uses a structure of 3 convolutional layers and 2 fully connected layers, with ReLU as the activation function. This allows it to capture the topological relationships of the compound's molecular structure and directly learn structural features from the chemical shift data to achieve structure generation.

[0046] The third implementation involves inputting the chemical shift list into both the structure generation model and the spectral shift list retrieval algorithm. Each method outputs its own structure analysis results and confidence scores. Then, the two results are combined and scored based on preset weights. These weights are determined by the performance of the two algorithms on the validation set; for example, the weight of the spectral shift list retrieval algorithm is set to 0.4, and the weight of the structure generation model is set to 0.6. A weighted average score is calculated for each candidate structure, and the highest score is selected as the final structure analysis result. When the results from the two algorithms differ significantly, the top three candidate structures are output for further validation. Alternatively, a decision tree method can be constructed, and the results from different models are fed into the decision tree model for further training and inference.

[0047] This invention provides several flexible structural analysis methods. The first method is simple and efficient, suitable for the rapid identification of known compounds. The second method has strong generalization ability and can handle the structural analysis of unknown compounds, making it particularly suitable for the research and development of novel compounds. The third method combines the advantages of two algorithms, improving the robustness and accuracy of structural analysis through comprehensive scoring, thus overcoming the limitations of a single algorithm in complex structural analysis. These three methods can be flexibly selected according to actual application scenarios, expanding the applicability of the method while reducing reliance on professional analytical personnel.

[0048] In summary, the compound structure determination method based on NMR data in the above embodiments of the present invention receives raw NMR data of the compound to be determined from an NMR instrument, preprocesses the raw NMR data to convert the time-domain NMR data into the frequency-domain data, suppresses noise interference, and outputs a peak position list; constructs a solvent peak template that matches the actual peak shape in the NMR data, delineates candidate regions in the peak position list according to the configured search interval, matches the signal of each candidate region with the signal of the solvent peak template to obtain the solvent peak; removes the solvent peak and calculates the actual chemical shift. The offset from the theoretical value is applied to global chemical shift correction to obtain a chemical shift list, which is then used for structural analysis of the compound to be analyzed. Solvent peaks are identified by constructing a template matching the actual spectrum peak shape and combining it with precise candidate region matching. Global chemical shift correction is then performed by calculating the offset, effectively avoiding mislabeling and omission of solvent peaks and improving the accuracy of solvent correction. Through an automated process of spectrum preprocessing, solvent peak identification, and chemical shift correction, an accurate chemical shift list is output for structural analysis, replacing the traditional method of manually comparing multidimensional spectra. This solves the problem of poor accuracy in existing compound structure analysis methods.

[0049] Example 3 Please see Figure 2 The figure shows a compound structure determination system based on nuclear magnetic resonance data proposed in the third embodiment of the present invention. The system includes: The acquisition module 100 is used to acquire multiple frames of images of the target area under consistent laser lighting and shooting conditions, and to generate a reference background image based on the multiple frames of images using a preset algorithm. The fusion module 200 is used to determine the corresponding difference image based on the reference background image, generate a guide image based on the image gradient magnitude, and fuse the normalized guide image into the difference image to obtain the target difference image. The processing module 300 is used to perform a two-dimensional Fourier transform on the target difference image to obtain a spectrogram, construct a two-dimensional mask, apply the two-dimensional mask to the spectrogram, and then perform an inverse Fourier transform to obtain the target image. Specifically, a data matrix is ​​constructed based on multiple frames of images, PCA is used to reduce the dimensionality of the data matrix, a preset number of principal components are retained to represent the stable background response, the principal components are used to reconstruct the object, and the reconstructed object is restored to an image to obtain the reference background image.

[0050] The functions or operation steps implemented by the above modules are largely the same as those in the above method embodiments, and will not be repeated here.

[0051] Example 4 In another aspect, the present invention provides a readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the steps of the method described in any one of Embodiments 1 to 2 above.

[0052] Example 5 In another aspect, the present invention provides an electronic device, the electronic device including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the program to implement the steps of any one of the methods described in Embodiments 1 to 2 above.

[0053] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0054] Those skilled in the art will understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequential list of executable instructions for implementing logical functions, and can be embodied in any computer-readable storage medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable storage medium" can mean any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0055] More specific examples (a non-exhaustive list) of computer-readable storage media include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable storage media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0056] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0057] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0058] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A method for determining the structure of compounds based on nuclear magnetic resonance data, characterized in that, The method includes: Receive raw NMR data of the compound to be analyzed from the NMR instrument, and preprocess the raw NMR data to convert the time-domain NMR data into the frequency-domain data, suppress noise interference, and output a list of peak positions; A solvent peak template matching the actual peak shape in the NMR data is constructed. Candidate regions are defined in the peak position list according to the configured search interval. The signal of each candidate region is matched with the signal of the solvent peak template to obtain the solvent peak. The solvent peaks are removed and the offset between the actual chemical shift and the theoretical value is calculated. The offset is applied to the global chemical shift correction to obtain a chemical shift list, and the structure of the compound to be analyzed is performed based on the chemical shift list.

2. The method for compound structure analysis based on nuclear magnetic resonance data according to claim 1, characterized in that, The steps of preprocessing the raw NMR data to convert the time-domain NMR data into the frequency-domain data, suppress noise interference, and output a list of peak positions include: The original nuclear magnetic resonance data was converted into a frequency domain spectrum by Fourier transform, and then the baseline drift was removed by automatic phase correction algorithm and baseline correction algorithm. Peaks in the frequency domain spectrum are identified by combining second derivative mutation point detection with local maximum search, and a list of peak positions is output.

3. The method for compound structure analysis based on nuclear magnetic resonance data according to claim 1, characterized in that, The steps for constructing a solvent peak template that matches the peak shape of the actual spectrum in the NMR data include: The actual spectrum data is extracted from the nuclear magnetic resonance data within the target chemical shift range, and the global half width at half maximum (FWHM) is calculated as the peak shape parameter. Based on the theoretical coupling constant value corresponding to the solvent, a solvent peak template is generated by combining it with the Lorentz function; or Based on the statistical average template of historical measured solvent spectra, the corresponding solvent peak templates are obtained.

4. The method for compound structure determination based on nuclear magnetic resonance data according to claim 3, characterized in that, The expression for generating the solvent peak template based on the theoretical coupling constant value corresponding to the solvent and the Lorentz function is as follows: ; ; in, This represents the ratio of peak intensity. L This represents the half-width and height calculated in the previous step. x 0 is the center position of the current peak. J It is the coupling constant value corresponding to the solvent peak.

5. The method for determining the structure of compounds based on nuclear magnetic resonance data according to claim 4, characterized in that, The step of matching the signal of each candidate region with the signal of the solvent peak template to obtain the solvent peak includes: The cosine similarity between the signal of each candidate region and the signal of the solvent peak template is calculated, and the solvent peak is obtained based on the similarity. The formula for calculating cosine similarity is: ; in, Information for each point in the solvent peak template. This represents the information of points selected in the candidate region for each experiment; For methanol-based solvents, the top three candidate peaks with similarity values ​​higher than the threshold are selected, and the presence of quintet splitting characteristics and integral intensity constraints are further verified to determine the solvent peaks.

6. The method for determining the structure of compounds based on nuclear magnetic resonance data according to claim 1, characterized in that, The steps for structural analysis of the compound to be analyzed based on the chemical shift list include: In the list of calculated chemical shifts 13 The C peak list and the compounds in the constructed spectral shift list library 13 C-peak list difference index; The target compound with the smallest difference is identified based on the difference index, and structural analysis is performed based on the target compound; or Based on a chemical shift list, a pre-trained structure generation model is used to perform structural analysis on the compound to be analyzed. The structure generation model is trained using a contrastive learning framework or a graph neural network; or... The chemical shift list is input into the structure generation model and the spectral shift list library retrieval algorithm respectively. The results are output separately and then comprehensively scored based on preset weights to obtain the final structure analysis result.

7. The method for compound structure determination based on nuclear magnetic resonance data according to claim 1, characterized in that, The formula for calculating the difference index is: ; in, and These are compounds from the chemical shift list and the spectral shift list library, respectively. 13 Number of C peaks Indicates in The number of peak values ​​matched within the threshold. Indicates the experimentally measured number of... i indivual 13 C chemical shift value, This indicates the predicted number of compounds in the spectral shift list library. j indivual 13 C chemical shift value.

8. A compound structure determination system based on nuclear magnetic resonance data, characterized in that, The system includes: The acquisition module is used to receive the raw NMR data of the compound to be analyzed from the NMR instrument, and to preprocess the raw NMR data to convert the time-domain NMR data into the frequency-domain data, suppress noise interference, and output a list of peak positions. The construction module is used to construct a solvent peak template that matches the peak shape of the actual spectrum in the NMR data. Based on the configured search interval, candidate regions are defined in the peak position list, and the signal of each candidate region is matched with the signal of the solvent peak template to obtain the solvent peak. The analysis module is used to remove solvent peaks and calculate the offset between the actual chemical shift and the theoretical value. The offset is applied to global chemical shift correction to obtain a chemical shift list, and the structure of the compound to be analyzed is performed based on the chemical shift list.

9. A readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method as described in any one of claims 1 to 7.

10. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor, when executing the program, implements the steps of the method as described in any one of claims 1 to 7.