A method and system for analyzing NMR spectra of mixtures under visualization conditions

By sorting and storing NMR spectral information on a computer, and automatically classifying and discriminating spectra using machine learning methods, the problem of low spectral analysis efficiency in the existing technology is solved, and rapid one-click spectral analysis is achieved.

CN114897031BActive Publication Date: 2025-06-20SICHUAN ACADEMY OF MEDICAL SCI SICHUAN PROVINCIAL PEOPLES HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210669361.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-14
Publication Date
2025-06-20
Estimated Expiration
2042-06-14

AI Technical Summary

Technical Problem

The existing spectral analysis technology has low analytical efficiency and cannot quickly and accurately analyze and identify organic drug structures.

Method used

A method for analyzing the mixture NMR spectrum under visual conditions is provided. NMR spectrum information is stored through computer sub-directory classification, Fourier transform and baseline adjustment are performed, and prediction models are established using supervised learning machine learning methods, automatically classify and distinguish the spectrum, record the peak position, match metabolites and display them.

Benefits of technology

It realizes rapid one-key processing of the spectrum, improves the spectrum analysis efficiency, solves the problem of low analytical efficiency in the prior art, and can quickly and accurately perform automatic analysis of the mixture NMR spectrum.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114897031B_ABST
    Figure CN114897031B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for analyzing NMR spectra of mixtures under visualization conditions, relating to nuclear magnetic resonance spectroscopy analysis technology, and solving the problem that there has been no automatic analysis of NMR spectra of mixtures for a long time. The main points of its technical solution are: classifying and storing the NMR spectrum information to be analyzed; converting the original attenuation oscillation file into a spectrum file after Fourier transform and performing baseline adjustment; converting the spectrum files of multiple samples into a matrix composed of multiple vectors of equal length, and constructing a sample classification information matrix after converting the classification information of the samples into dummy variables; establishing a prediction model; obtaining the weights of the contributions of the spectra of each segment of the samples to the sample classification information; sorting the weights of the spectra; matching to obtain the matching metabolites; calculating the Jaccard distance score between the peak positions and the matching metabolites; and displaying on the original NMR spectrum. The present invention achieves the purpose of quickly processing the spectrum with one key.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to nuclear magnetic resonance spectroscopy analysis technology, and more specifically, it relates to a method and system for analyzing NMR spectra of mixtures under visualization conditions. Background Art

[0002] Spectrum analysis is a class of analytical methods for determining the properties, contents, and structures of substances based on the electromagnetic radiation emitted by substances or the radiation signals generated or signal changes occurring after the interaction between substances and radiation.

[0003] Spectrum analysis has become one of the main methods for modern molecular structure analysis and identification of substances. With the development of technology, technological innovation, and computer applications, spectrum analysis has also developed rapidly. The spectrum analysis method has outstanding advantages and wide applications, and is an indispensable tool in many scientific research and production fields. With the continuous development of technology and the increasing requirements for analysis, scientific researchers are constantly innovating in spectrum analysis methods.

[0004] However, the existing spectrum analysis technology generally requires staff to operate a computer to complete spectrum analysis, which has low analysis efficiency, is time-consuming and laborious, and cannot provide timely and accurate data support for the research on the structure analysis and identification of organic drugs. Therefore, how to research and design a method and system for analyzing NMR spectra of mixtures under visualization conditions is an urgent problem for us to solve currently. Summary of the Invention

[0005] To solve the problems of low working efficiency and time-consuming and laborious of the existing spectrum analysis methods, the purpose of the present invention is to provide a method and system for analyzing NMR spectra of mixtures under visualization conditions, which can achieve the purpose of quickly processing spectra with one key.

[0006] The above technical purpose of the present invention is achieved through the following technical solutions:

[0007] In a first aspect, a method for analyzing NMR spectra of mixtures under visualization conditions is provided, including the following steps:

[0008] Obtain the NMR spectrum information to be analyzed, and classify and store the NMR spectrum information to be analyzed in units of computer subdirectories;

[0009] Convert the NMR spectrum information to be analyzed from the original attenuation oscillation file into a spectrum file after Fourier transform, and perform baseline adjustment on the converted NMR spectrum information to be analyzed;

[0010] The spectrum file is converted into a vector by piecewise integration, and the spectrum files of multiple samples are converted into a matrix consisting of multiple equal-length vectors, and the classification information of the samples is converted into dummy variables to construct a sample classification information matrix;

[0011] Use supervised learning machine learning methods to establish a prediction model for automatic classification and discrimination of spectra;

[0012] Compare the differences between the classified samples, and use the prediction model to obtain the weight of the spectrum of each segment of the sample to the sample classification information;

[0013] After sorting the weights of the spectrum, the positions of multiple peaks in the top contribution columns are recorded;

[0014] Matching metabolites are obtained from the database based on the peak positions;

[0015] The matching metabolites were ranked after calculating the Jaccard distance scores between the peak positions and the corresponding matching metabolites;

[0016] The peaks of multiple matching metabolites ranked at the top of the Jaccard score are selected and displayed in the original NMR spectrum.

[0017] Furthermore, the specific process of the baseline adjustment is:

[0018] Use Shannon entropy to determine whether the signal in the spectrum file is meaningful;

[0019] The signal area where the Shannon entropy is lower than the preset value is selected for analysis to complete the baseline adjustment.

[0020] Furthermore, the matching process of the matching metabolites is specifically as follows:

[0021] If there are NMR spectral data of n small molecule organic compounds in the known data list, all n metabolites are selected as candidate metabolites;

[0022] The peaks of n small molecule organic compounds are converted into 1-0 sequences, and the Jaccard distance scores corresponding to the n candidate metabolites are matched with the mixture samples obtained in the experiment, and the candidate metabolites with the highest scores are selected as the true metabolites.

[0023] Furthermore, the specific process of displaying each peak of the plurality of matching metabolites in the original NMR spectrum is as follows:

[0024] When selecting a peak image to be analyzed in a two-dimensional NMR spectrum image, the spectrum to be analyzed is converted into a semi-transparent state;

[0025] At the same time, the spectrum of the candidate metabolite is superimposed on the spectrum to be analyzed, and the name of the molecule is displayed, and the peak height of the candidate metabolite is adjusted.

[0026] Further, the machine learning method of supervised learning is any one of partial least squares method, PLS, support vector machine SVM, and shaply index.

[0027] Further, the database is any one of BMRB database, HMDB database, and Swidish NMR database.

[0028] Further, the number of selected peak positions is 30 - 50.

[0029] In a second aspect, a system for analyzing the NMR spectrum of a mixture under visualization conditions is provided, including:

[0030] A classification storage module, configured to obtain NMR spectrum information to be analyzed, and classify and store the NMR spectrum information to be analyzed in units of computer subdirectories;

[0031] A baseline adjustment module, configured to convert the NMR spectrum information to be analyzed from the original attenuation oscillation file into a spectrum file after Fourier transform, and perform baseline adjustment on the converted NMR spectrum information to be analyzed;

[0032] A matrix construction module, configured to convert the spectrum file into a vector in a segmented integration manner, and convert the spectrum files of multiple samples into a matrix composed of multiple equal-length vectors, and construct a sample classification information matrix after converting the classification information of the samples into dummy variables;

[0033] A model construction module, configured to establish a prediction model for automatically classifying and discriminating the spectrum by using a machine learning method of supervised learning;

[0034] A weight analysis module, configured to compare the differences between classified samples, and obtain the weights of the contributions of the spectrum of each segment of the sample to the sample classification information by using the prediction model;

[0035] A peak recording module, configured to record multiple peak positions ranked in the forefront of the contribution ranking after sorting the weights of the spectrum;

[0036] A metabolite matching module, configured to match and obtain matching metabolites from the database according to the peak positions;

[0037] A scoring analysis module, configured to calculate the Jaccard distance score between the peak positions and the corresponding matching metabolites, and sort the matching metabolites;

[0038] A peak display module, configured to select each peak of multiple matching metabolites ranked in the forefront of the Jaccard score and display it in the original NMR spectrum.

[0039] In a third aspect, a computer terminal is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements an analytical method for a mixture NMR spectrum under visualization conditions as described in any one of the first aspects.

[0040] In a fourth aspect, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it can implement an analytical method for a mixture NMR spectrum under visualization conditions as described in any one of the first aspects.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] The analytical method for a mixture NMR spectrum under visualization conditions proposed by the present invention completes data import, format conversion, fast Fourier transform, peak recognition integration, and finally performs machine learning on the spectrum through one-key data processing after inputting the spectrum, achieving the purpose of quickly processing the spectrum with one key and solving the problem that there has been no automatic analysis of mixture NMR spectra for a long time. Description of the Drawings

[0043] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not limit the embodiments of the present invention. In the drawings:

[0044] Figure 1 is a flowchart in an embodiment of the present invention;

[0045] Figure 2 is a schematic diagram of the analysis result in an embodiment of the present invention;

[0046] Figure 3 is a system block diagram in an embodiment of the present invention. Detailed Embodiments

[0047] To make the purpose, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and the drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and do not limit the present invention.

[0048] Embodiment 1: An analytical method for a mixture NMR spectrum under visualization conditions, as Figure 1 shown, includes the following steps:

[0049] S1: Obtain the NMR spectrum information to be analyzed and classify and store the NMR spectrum information to be analyzed in units of computer subdirectories;

[0050] S2: Convert the NMR spectral information to be parsed from the original attenuation oscillation file into a spectrogram file after Fourier transform, and perform baseline adjustment on the converted NMR spectral information to be parsed;

[0051] S3: Convert the spectrogram file into a vector in a segmented integration manner, and convert the spectrogram files of multiple samples into a matrix composed of multiple vectors of equal length. In addition, convert the classification information of the samples into dummy variables and then construct a sample classification information matrix;

[0052] S4: Use a machine learning method of supervised learning to establish a prediction model for automatically classifying and discriminating spectra;

[0053] S5: Compare the differences between each classified sample, and use the prediction model to obtain the weights of the spectral contributions of each segment of the sample to the sample classification information;

[0054] S6: After sorting the weights of the spectra, record multiple peak positions in the front row of the contributions;

[0055] S7: Match the matching metabolites from the database according to the peak positions;

[0056] S8: Calculate the Jaccard distance score between the peak positions and the corresponding matching metabolites, and then sort the matching metabolites;

[0057] S9: Select the peaks of multiple matching metabolites in the front row of the Jaccard score ranking and display them in the original NMR spectrum.

[0058] Essentially, NMR spectrum is an electromagnetic signal excited by calculating the co-rotating of seeds in a strong magnetic field to the chaotic rotation direction of neutrons (relaxation) in a magnetic field-free environment. This signal detected by the sensor is the original signal of the NMR spectrum. Essentially, this original signal is manifested as attenuation oscillation. Such attenuation oscillation comes from all hydrogen atom groups of the object to be detected. It also comes from the operation interference of various instruments and equipment. Therefore, we need to use spectral features to automatically select meaningful signals.

[0059] For this reason, the present invention determines whether it is meaningful according to the Shannon entropy of a segment of signal, and its calculation method is as follows: The equation of spectral entropy comes from the power spectrum of the signal and the probability distribution equation.

[0060] For the signal x(n), its power spectrum can be expressed as: S(m) = |X(m)| 2 , where X(m) is the discrete Fourier transform of x(n). And the probability distribution P(m) is equal to:

[0061]

[0062] Among them, S(m) is the power spectrum, and S(i) is the power spectrum of the i-th sample.

[0063] Sum over frequency points from m = 1 to the Nth frequency point and take the logarithm of the probability distribution. At this time, the calculated Shannon entropy H is as follows:

[0064]

[0065] After normalization, it is:

[0066]

[0067] Where N is the total number of frequency points, and the denominator log2N represents the maximum spectral entropy of white noise, which is uniformly distributed in the frequency domain. According to different application scenarios, we can select H n Less than a certain specific value, for example, less than 0.5, is a meaningful area that needs to be analyzed.

[0068] In addition, the obtained meaningful original attenuation oscillation signal is transformed from the time-domain signal to the frequency-domain signal by the classical Fourier transform algorithm.

[0069] For a vector x with n uniformly sampled points, the definition of the Fourier transform is:

[0070]

[0071] Where, is one of the complex roots of the nth unit roots, where i is the imaginary unit. For x and y, the ranges of the exponents j and k are from 0 to n - 1. Using the Fourier transform, the signal can be transformed from the time domain to the frequency domain.

[0072] After the meaningful spectrum is automatically selected, it is more meaningful to find out which small molecule organic compounds exist in the mixture NMR spectrum. The present invention estimates by calculating the Jaccard distance between the selected specific molecular spectrum and the specific molecules in the database. The specific method is as follows:

[0073] If there are two specific sequences A and B, where sequence B comes from the known metabolite spectrum and sequence A comes from the unknown mixture spectrum, we can calculate by:

[0074]

[0075] to obtain the Jaccard distance between the two sequences. In the above formula, d represents the Jaccard distance function of sequences A and B. T represents the proportion of the common peaks of sequences A and B in the total peaks of A and B.

[0076] If A and B are two binary vectors, representing the presence (1) or absence (0) of specific peaks at specific positions respectively, then the Jaccard distance can be calculated by the following formula:

[0077]

[0078] Among them, N AB represents the number of peaks appearing at the positions shared by the two sequences, N A represents the number of A peaks, N B represents the number of B peaks.

[0079] Assume that the NMR spectral data of n small molecule organic compounds are known in a certain data list, then all these n metabolites can be selected as possible metabolites. The peaks of these small molecule organic compounds known in the database are converted into a 1-0 sequence, where 1 indicates the presence of a peak and 0 indicates the absence of a peak at this site. Match with the mixture sample obtained from the experiment, calculate its Jaccard distance score according to the requirements described in S3. n possible metabolites can obtain n scores. Sort them according to the score size, and select the candidate metabolites with higher scores as the true metabolites. Mark the names of the metabolites on the spectrogram.

[0080] In the obtained two-dimensional NMR spectral image, when the mouse moves to the peak image that needs to be analyzed, the system converts the original spectrum to be analyzed into a semi-transparent state, and at the same time superimposes the candidate metabolite spectrum described in S5 on the spectrum to be analyzed. And display the name of the molecule. Rotating the mouse wheel can increase / decrease the peak height of the candidate metabolite.

[0081] While showing the peak height of the above-mentioned candidate metabolites, the peaks of the candidate metabolites are labeled with the colors mapped by the shaply index: Generally, long-wavelength colors (red, orange) are used to represent high shaply index values, and short-wavelength colors (green, blue) are used to represent low shaply index values.

[0082] In this embodiment, the machine learning methods of supervised learning include but are not limited to partial least squares method, PLS, support vector machine SVM, and shaply index. The shaply index is an index for judging the importance degree of each peak after the implementation of the above machine learning methods.

[0083] In this embodiment, the database includes but is not limited to the BMRB database, HMDB database, and Swidish NMR database.

[0084] Among them, the number of selected peak positions is 30 - 50.

[0085] In this embodiment, the number of selected peak positions is 50. The experimental results are as Figure 1As shown in the figure, the present invention provides a method for analyzing the NMR spectrum of a mixture under visual conditions. After inputting the spectrum, through one-key data processing, data import, format conversion, fast Fourier transform, peak recognition and integration are completed. Finally, machine learning is performed on the spectrum to achieve the purpose of quickly processing the spectrum with one key, and solve the problem that there has been no automatic analysis of the NMR spectrum of mixtures for a long time.

[0086] Example 2: An analytical system for the NMR spectrum of a mixture under visual conditions, as Figure 3 shown, includes a classification storage module, a baseline adjustment module, a matrix construction module, a model construction module, a weight analysis module, a peak recording module, a metabolite matching module, a scoring analysis module, and a peak display module.

[0087] Among them, the classification storage module is used to obtain the NMR spectrum information to be analyzed and classify and store the NMR spectrum information to be analyzed in units of computer subdirectories.

[0088] The baseline adjustment module is used to convert the NMR spectrum information to be analyzed from the original attenuation oscillation file into a spectrum file after Fourier transform, and perform baseline adjustment on the converted NMR spectrum information to be analyzed.

[0089] The matrix construction module is used to convert the spectrum file into a vector in a segmented integration manner, and the spectrum files of multiple samples are converted into a matrix composed of multiple equal-length vectors, and the classification information of the samples is converted into dummy variables to construct a sample classification information matrix.

[0090] The model construction module is used to establish a prediction model for automatically classifying and discriminating the spectrum by using a supervised learning machine learning method.

[0091] The weight analysis module is used to compare the differences between the classified samples and obtain the weights of the contributions of the spectra of each segment of the samples to the sample classification information by using the prediction model.

[0092] The peak recording module is used to record the positions of multiple peaks ranked in the front in terms of contribution after sorting the weights of the spectrum.

[0093] The metabolite matching module is used to match the matching metabolites from the database according to the peak positions.

[0094] The scoring analysis module is used to calculate the Jaccard distance score between the peak positions and the corresponding matching metabolites and then sort the matching metabolites.

[0095] The peak display module is used to select the peaks of multiple matching metabolites ranked in the front in terms of Jaccard score and display them in the original NMR spectrum.

[0096] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for analyzing the NMR spectrum of a mixture under visualization conditions, characterized in that, It includes the following steps: Obtain the NMR spectrum information to be analyzed, and classify and store the NMR spectrum information to be analyzed in units of computer subdirectories; Convert the NMR spectrum information to be analyzed from the original attenuation oscillation file into a spectrum file after Fourier transform, and perform baseline adjustment on the converted NMR spectrum information to be analyzed; Convert the spectrum file into a vector in a segmented integration manner, and convert the spectrum files of multiple samples into a matrix composed of multiple equal-length vectors, and convert the classification information of the samples into dummy variables and then construct a sample classification information matrix; Use a supervised learning machine learning method to establish a prediction model for automatically classifying and discriminating spectra; Compare the differences of each classified sample, and use the prediction model to obtain the weights of the contribution of the spectra of each segment of the sample to the sample classification information; After sorting the weights of the spectra, record multiple peak positions in the front row of the contribution ranking; Match the matching metabolites from the database according to the peak positions; Sort the matching metabolites after calculating the Jaccard distance score between the peak positions and the corresponding matching metabolites; Select the peaks of multiple matching metabolites ranked in the front row of the Jaccard score and display them in the original NMR spectrum; The specific process of the matching of the matching metabolites is as follows: If there are NMR spectrum data of n small molecule organic compounds in the known data list, then select all n metabolites as candidate metabolites; Convert the peaks of the n small molecule organic compounds into a 1-0 sequence, and match them with the mixture sample obtained by experiment according to the Jaccard distance scores corresponding to the n candidate metabolites, and select the candidate metabolites with higher scores as the true metabolites.

2. The method for analyzing the NMR spectrum of a mixture under visualization conditions according to claim 1, characterized in that, The specific process of the baseline adjustment is as follows: Judge whether the signals in the spectrum file are meaningful through Shannon entropy; Select the signal regions with Shannon entropy lower than the preset value for analysis to complete the baseline adjustment.

3. The method for analyzing the NMR spectrum of a mixture under visualization conditions according to claim 1, characterized in that, The specific process of displaying the peaks of multiple matching metabolites in the original NMR spectrum is as follows: When selecting the peak image to be analyzed in the two-dimensional NMR spectrum image, convert the spectrum to be analyzed into a semi-transparent state; At the same time, superimpose the candidate metabolite spectrum on the spectrum to be analyzed, and display the name of the molecule corresponding to the candidate metabolite, and adjust the peak height of the candidate metabolite.

4. The method for analyzing the NMR spectrum of a mixture under visualization conditions according to claim 1, characterized in that, The supervised learning machine learning method is any one of partial least squares method, PLS, support vector machine SVM, and shaply index.

5. The method for analyzing the NMR spectrum of a mixture under visualization conditions according to claim 1, characterized in that, The database is any one of the BMRB database, HMDB database, and Swidish NMR database.

6. The method for analyzing the NMR spectrum of a mixture under visualization conditions according to claim 1, characterized in that, The number of selected peak positions is 30-50.

7. A system for analyzing the NMR spectrum of a mixture under visualization conditions, characterized in that, It includes: A classification storage module for obtaining the NMR spectrum information to be analyzed and classifying and storing the NMR spectrum information to be analyzed in units of computer subdirectories; A baseline adjustment module for converting the NMR spectrum information to be analyzed from the original attenuation oscillation file into a spectrum file after Fourier transform and performing baseline adjustment on the converted NMR spectrum information to be analyzed; A matrix construction module, which is used to convert a spectral file into a vector in a segmented integration manner, and convert the spectral files of multiple samples into a matrix composed of multiple vectors of equal length, and construct a sample classification information matrix after converting the classification information of the samples into dummy variables; A model construction module, which is used to establish a prediction model for automatically classifying and discriminating spectra by using a supervised learning machine learning method; A weight analysis module, which is used to compare the differences of each classified sample and obtain the weights of the spectral contributions of each segment of the sample to the sample classification information by using the prediction model; A peak recording module, which is used to record the positions of multiple peaks ranked top in terms of contribution after sorting the weights of the spectra; A metabolite matching module, which is used to match the matching metabolites from the database according to the peak positions; A scoring analysis module, which is used to calculate the Jaccard distance score between the peak positions and the corresponding matching metabolites and then sort the matching metabolites; A peak display module, which is used to select the peaks of multiple matching metabolites ranked top in terms of Jaccard score and display them in the original NMR spectrum; The matching process of the matching metabolites is specifically as follows: If there are NMR spectral data of n small molecule organic compounds in the known data list, then all n metabolites are selected as candidate metabolites; Convert the peaks of the n small molecule organic compounds into a 1-0 sequence, and match the Jaccard distance scores corresponding to the n candidate metabolites with the experimentally obtained mixture samples, and select the candidate metabolites with higher scores as the true metabolites.

8. A computer terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements an analytical method for NMR spectra of mixtures under visualization conditions as described in any one of claims 1-6.

9. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it can implement an analytical method for NMR spectra of mixtures under visualization conditions as described in any one of claims 1-6.

Citation Information

Patent Citations

  • A method to measure tissue texture using NMR spectroscopy to identify chemical species of component textural elements in a targeted region of tissue

    CN111093473A

  • Polymer physical property estimation device and learning method

    JP2021107813A