Nuclear magnetic resonance spectrogram library construction and analysis system based on data standardization

Through the data-standardized NMR spectrum library construction and analysis system, the NMR data is automatically processed, and the structured database is constructed, which solves the problems of traditional low resolution efficiency and incomplete spectrum library, and achieves efficient spectrum analysis and compound structure reflection.

CN120277057APending Publication Date: 2025-07-08NANJING QINGSHI TESTING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510493309.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional nuclear magnetic resonance spectrogram analysis is inefficient and relies on manual processing of massive data. The spectrogram library construction lacks systematicity and integrity, and cannot fully reflect the structural information of the compound. Manual analysis makes it difficult to quickly identify potential components and structural information in complex spectrograms.

Method used

The NMR spectrum library is used to construct and analyze the system, including data acquisition, format conversion, standardization and visual analysis modules. By automatically processing the NMR data, a structured NMR database is constructed, and the key peak characteristics are integrated and visual analysis is performed.

Benefits of technology

The degree of automation of NMR spectrum analysis is improved, and the manual processing time is reduced. The built NMR spectrum database fully reflects the compound structure information, improving the comparability and analysis accuracy of the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277057A_ABST
    Figure CN120277057A_ABST
Patent Text Reader

Abstract

The invention discloses a nuclear magnetic resonance spectrogram library construction and analysis system based on data standardization, and relates to the technical field of data analysis, the system comprises a data acquisition module, a data format conversion module, a standardization module, a spectrogram library construction module and a visual analysis module; the data acquisition module is used for acquiring original data through a nuclear magnetic instrument and marking the original data; the data format conversion module is used for preprocessing original data and converting the preprocessed data format into a data point table form data file containing chemical shift and peak intensity; the standardization module is used for respectively carrying out operation processing and cleaning processing on X-axis data representing chemical shift and Y-axis data representing peak intensity; the spectrum library construction module is used for constructing a structured nuclear magnetic spectrum database; and the visual analysis module is used for visual presentation and analysis of the spectrogram data. According to the system and the method, spectrogram library construction and data analysis of the nuclear magnetic resonance data are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and specifically to a nuclear magnetic resonance spectrum library construction and analysis system based on data standardization. Background Art

[0002] With the continuous improvement of the accuracy requirements for material structure analysis in the scientific research field, nuclear magnetic resonance (NMR) technology has been increasingly deeply applied in fields such as component analysis, materials science, and drug research and development. As a key technology for revealing the molecular structure of substances, the accuracy of nuclear magnetic resonance spectrum analysis and the scientific nature of spectrum library construction affect the reliability of scientific research results and the innovation and development process of related fields.

[0003] However, traditional nuclear magnetic resonance spectrum analysis and spectrum library construction methods often face the following problems when dealing with the complex requirements of modern scientific research: First, the analysis efficiency is low, the spectrum data is complex, and the manual processing burden is heavy. A large amount of nuclear magnetic resonance spectrum data will be generated in work such as component analysis, materials science, and drug research and development. Traditional spectrum analysis highly depends on the knowledge and experience of professionals, and it is time-consuming and laborious to manually process this massive amount of data; Second, the construction of the spectrum library lacks systematicness and integrity, and the key spectral peak characteristics are not fully integrated. When currently constructing a nuclear magnetic resonance spectrum database, only the chemical shift values of the reference compounds are simply summarized to form a set of isolated data points, ignoring important spectral peak characteristics such as complex peak shape changes and peak intensities. This construction method results in the spectrum library being unable to comprehensively reflect the structural information of compounds; In addition, traditional spectrum analysis usually manually analyzes features such as chemical shifts and peak intensities in the spectrum one by one, mainly relying on personal experience judgment. This method is difficult to quickly identify potential components and structural information in complex spectra. Summary of the Invention

[0004] The purpose of the present invention is to provide a nuclear magnetic resonance spectrum library construction and analysis system based on data standardization to solve the problems raised in the prior art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A nuclear magnetic resonance spectrum library construction and analysis system based on data standardization, the system includes: a data acquisition module, a data format conversion module, a standardization module, a spectrum library construction module, and a visualization analysis module; the data acquisition module is used to obtain the original data of the nuclear magnetic resonance instrument and mark the original data; the data format conversion module is used to correct the original data and convert the corrected data format into a data file in the form of a data point table containing chemical shifts and peak intensities; the standardization module is used to perform arithmetic processing and cleaning processing on the X-axis data representing chemical shifts and the Y-axis data representing peak intensities respectively; the spectrum library construction module is used to construct a structured nuclear magnetic resonance spectrum database; the visualization analysis module is used for the visualization presentation and analysis of spectrum data.

[0006] The data acquisition module includes a data acquisition unit and a labeling unit; the data acquisition unit is used to acquire the original data of the nuclear magnetic instrument; the labeling unit is used to perform initial marking on the acquired original data, where the marking content includes the sample number and the acquisition time. The output end of the data acquisition module is connected to the input end of the data format conversion module. Among them, the data acquisition unit has the function of calling the data acquisition software supporting the nuclear magnetic instrument, and the labeling unit has the function of calling the data processing software supporting the nuclear magnetic instrument.

[0007] The data format conversion module includes a data correction unit and a format conversion unit; the data correction unit is used to perform correction operations on the acquired original data, including baseline correction, displacement correction, and phase correction;

[0008] The format conversion unit is used to verify the corrected data. The verification content includes the integrity of the data, that is, to check whether there are missing values in the data. For the data points with missing values and the data is continuous, interpolation method is used for filling. For the missing values that do not affect the statistical characteristics of the overall data, the data points containing the missing values are deleted; after the verification passes, the data format that passes the verification is converted into a data file in the form of a data point table containing X-axis data and Y-axis data. The output end of the data correction unit is connected to the input end of the format conversion unit, and the output end of the format conversion unit is connected to the input end of the standardization module.

[0009] The standardization module includes an X-axis data processing unit and a Y-axis data processing unit. The standardization module is used to perform arithmetic processing and cleaning processing on the converted data point table data file for the X-axis and Y-axis data respectively. For the X-axis data processing, a mapping processing method is adopted, and for the Y-axis data processing, a maximum normalization method is adopted to obtain the reference chemical substance spectrum data and the detection spectrum data of the chemical substance to be analyzed. By performing arithmetic processing on the X-axis and Y-axis data respectively, the data differences brought by different instruments and experimental conditions are eliminated, the comparability and analysis accuracy of the data are improved, and the chemical shift and peak intensity ranges are unified. The output end of the standardization module is connected to the input end of the spectral library construction module. Among them, the cleaning processing deletes the data outside the selected range of the X-axis data. The selected range of the X-axis data refers to the chemical shift range [δ min , δ max , where δ min represents the minimum value of the selected chemical shift data; δ max represents the maximum value of the selected chemical shift data.

[0010] The X-axis data processing unit processes the data through a mapping processing method, specifically as follows:

[0011] Select the X-axis data range in the converted data point table data file, that is, the chemical shift range [δ min , δ max , and perform mapping processing. The chemical shift range is set according to the type of nuclear magnetic data. The chemical shift range of proton nuclear magnetic resonance is usually [-4, 20], preferably [-1, 12]. This range covers the chemical shifts of hydrogen atoms in most organic compounds. Further preferably, [-0.2, 9.5] is determined based on the chemical shift intervals that frequently appear in the structures of common organic compounds and experiments, which helps to improve the data processing efficiency and pertinence; for the chemical shift range of carbon nuclear magnetic resonance, it is usually [-20, 220], preferably [-10, 210], and further preferably [-5, 190];

[0012] Select the target mapping range [v min , v max . The target mapping range is selected from the delimited ranges, including the chemical shift range and the second range. Among them, the second range includes [0, 90], [-20, 2000], [400, 4000], [3200, 10000], [50, 650], and [50, 6000];

[0013] After selecting the target mapping range, perform processing through the linear mapping algorithm. Among them, the linear mapping algorithm is defined as follows: v = v min +(δ - δ min ) * (v max - v min ) / (δ max - δ min ); where δ represents the chemical shift value; δ min represents the minimum value of the selected chemical shift data; δ max represents the maximum value of the selected chemical shift data; v represents the value obtained after linear mapping, v min represents the minimum value of the target mapping range; v max represents the maximum value of the target mapping range.

[0014] The Y-axis data processing unit processes the data through the maximum value normalization method, specifically as follows:

[0015] In the converted data point table data file, use the MAX function to find the maximum intensity value I in the Y-axis sequence max , divide all intensity values by the maximum intensity value to obtain the relative intensity value R I = k * I / I max , and process the Y-axis data through the maximum value normalization method. Among them, k represents the scaling factor, k ≠ 0; R I represents the relative intensity value; I maxUse the MAX function to find the maximum intensity value in the Y-axis sequence; I represents the peak intensity value in the Y-axis sequence of the data point table data file.

[0016] The spectral library construction module includes a data classification and storage unit and a data update management unit. The data classification and storage unit is used to classify and store the reference chemical substance spectral data according to the spectral type (such as hydrogen spectrum, carbon spectrum, etc.) and the categories, sources, and application fields of chemical substances, and construct a structured nuclear magnetic resonance spectral database. The output end of the spectral library construction module is connected to the input end of the visualization analysis module.

[0017] The data update management unit is used to execute the review process for adding data to the spectral library, specifically as follows:

[0018] After the data acquisition module obtains the original data and marks it, the data format conversion module corrects and converts it, and the normalization module performs operations and cleaning processing on the chemical shift and peak intensity data to obtain the reference chemical substance spectral data. Then, an operation for applying to enter the database is performed on the reference chemical substance spectral data, and the database is updated;

[0019] Judge whether the reference chemical substance spectral data meets the storage standards. Among them, the storage standards are specifically as follows:

[0020] In terms of data quality: The signal-to-noise ratio of the spectrum reaches the set signal-to-noise ratio threshold to ensure that the spectrum signal is clearly distinguishable; the baseline fluctuation range is within the set baseline fluctuation range; the measurement error of the chemical shift is controlled within the set measurement error range of the chemical shift;

[0021] In terms of data integrity: It includes complete spectral information, including but not limited to spectral type (hydrogen spectrum, carbon spectrum, etc.), basic information of chemical substances (name, structural formula, etc.), experimental conditions (temperature, solvent, etc.);

[0022] When the reference chemical substance spectral data meets the storage standards, the reference chemical substance spectral data that meets the storage standards is added to the structured nuclear magnetic resonance spectral database, and the index information of the structured nuclear magnetic resonance spectral database is updated; when the reference chemical substance spectral data does not meet the storage standards, the reference chemical substance spectral data that does not meet the storage standards is stored in a temporary data area, and a feedback report is generated. The feedback report includes the reasons why the data does not meet the storage standards.

[0023] The visualization analysis module is used to present the detection spectrum data of the chemical substance to be analyzed through a spectral curve, display the characteristics of chemical shift and peak intensity through the spectral curve, support users to perform data editing on the spectral curve, support automatic recognition of spectral features, database retrieval, and weighted fitting of multiple spectrograms. Among them, data editing includes a first function and a second function. The first function is used to select a section of the spectral curve to generate a baseline or a straight line; the second function is used to perform secondary normalization processing on the Y-axis data of the processed new spectral curve; the automatic recognition of spectral features is based on a threshold-based peak detection algorithm to identify spectral peaks and establish a temporary spectral peak data table for correlation retrieval based on the chemical shift of the peaks; there are two correlation retrieval methods for the database retrieval: correlation retrieval based on the obtained temporary spectral peak data table and correlation retrieval based on spectral data points. The correlation retrieval uses a spectral similarity algorithm between the compound to be measured and the standard compound for correlation retrieval; the weighted fitting of multiple spectrograms is used to assign corresponding weights to each standard NMR spectrogram data point according to the theoretical proportion of each component in the formula. On this basis, multiple standard NMR spectrogram data points are subjected to weighted summation operations to generate a fitting data point table. The generated fitting data point table is used to simulate the spectral peaks of different ratios of the formula to simulate or restore the formula;

[0024] The function of generating a baseline or a straight line of the first function is mainly used to remove local peaks in the NMR spectrogram. Taking the solvent peak as an example, when the signal intensity of the deuterated water solvent peak in the spectrogram is higher than other surrounding peaks and interferes with the spectrogram analysis (here, when the signal intensity exceeds [N] times the average intensity of the surrounding peaks, this multiple can be set according to the actual situation), where the parameter N represents a multiple standard for judging the signal intensity, the following operations can be adopted:

[0025] Select two endpoints (Xa, Ya) and (Xb, Yb) within the chemical shift range of the solvent peak whose vertical distance from the baseline is within a certain threshold (this threshold can be set to, for example, 0.05 times the peak height, that is, when the ratio of the peak height at the endpoint to the highest peak height of the solvent peak is less than 0.05, it is considered close to the baseline). Among them, Xa and Xb represent the positions on the chemical shift axis of the solvent peak, and these two values determine the interval of the chemical shift of the solvent peak of interest; Ya and Yb represent the vertical distance from the spectrogram curve to the baseline at the corresponding chemical shift values Xa and Xb;

[0026] For the points between the selected chemical shifts Xa and Xb, there are two processing methods: filling with the average endpoint value and filling with a straight-line equation;

[0027] Filling with the average endpoint value: These points are refilled as (X, 0.5Ya + 0.5Yb), where X represents the chemical shift value corresponding to the point being filled within the chemical shift interval [Xa, Xb] of the solvent peak when filling with the average endpoint value;

[0028] Straight line equation filling: First, obtain the straight line equation based on these two endpoints, then substitute the X-axis values into the equation to obtain the corresponding Y-axis values, and further refill the points between the selected chemical shifts Xa and Xb with the corresponding points on the straight line;

[0029] The second function is used to perform secondary normalization processing on the Y-axis data of the processed new spectral curve. Specifically:

[0030] When a local peak in the nuclear magnetic resonance spectrogram is generated as a baseline peak and this local peak is the maximum peak, it is necessary to re-normalize the spectrum to meet the requirements of the standardized spectrogram.

[0031] Compared with the prior art, the beneficial effects of the present invention are:

[0032] 1. By acquiring and marking the original data, the data format conversion module performs phase correction and format conversion, and the standardization module performs efficient arithmetic processing on the data. The entire process has a high degree of automation, reducing the time and energy consumption of manual processing. Facing the generated spectral data, the parsing work is completed, improving the problem of low efficiency of manual parsing;

[0033] 2. The data classification and storage unit classifies and stores the reference chemical substance spectral data in multiple dimensions such as spectrogram type, chemical substance category, source, and application field. During the data processing process, key spectral peak features such as complex peak shape changes and peak intensities are considered, and this information is integrated into the spectral library, making the constructed nuclear magnetic resonance spectral database reflect the structural information of the compound. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a schematic flowchart of a nuclear magnetic resonance spectral library construction and analysis system based on data standardization of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0036] In Embodiment 1: As Figure 1As shown in the figure, the present invention provides a technical solution, a nuclear magnetic resonance spectrum library construction and analysis system based on data standardization. The system includes: a data acquisition module, a data format conversion module, a standardization module, a spectrum library construction module, and a visualization analysis module; the data acquisition module is used to obtain the original data of the nuclear magnetic resonance instrument and mark the original data; the data format conversion module is used to correct the original data and convert the corrected data format into a data file in the form of a data point table containing chemical shift and peak intensity; the standardization module is used to perform arithmetic processing and cleaning processing on the X-axis data representing chemical shift and the Y-axis data representing peak intensity respectively; the spectrum library construction module is used to construct a structured nuclear magnetic resonance spectrum database; the visualization analysis module is used for the visualization presentation and analysis of spectrum data.

[0037] The data acquisition module includes a data acquisition unit and a marking unit; the data acquisition unit is used to obtain the original data of the nuclear magnetic resonance instrument; the marking unit is used to perform initial marking on the obtained original data, where the marking content includes the sample number and the acquisition time. The output end of the data acquisition module is connected to the input end of the data format conversion module. Among them, the data acquisition unit has the function of calling the data acquisition software supporting the nuclear magnetic resonance instrument, and the marking unit has the function of calling the data processing software supporting the nuclear magnetic resonance instrument.

[0038] Specifically, researchers use a nuclear magnetic resonance instrument to detect 10 new compound samples. The signal data acquisition unit obtains the original data, and the marking unit marks each sample with a unique number (Sample001 - Sample010) and the acquisition time (2024 - 10 - 01 - 10:00:00).

[0039] The data format conversion module includes a data correction unit and a format conversion unit; the data correction unit is used to perform correction operations on the obtained original data, including baseline correction, displacement correction, and phase correction;

[0040] The format conversion unit is used to verify the corrected data. The verification content includes the integrity of the data, that is, to check whether there are missing values in the data. For data points with missing values and the data is continuous, interpolation method is used to fill them. For missing values that do not affect the statistical characteristics of the overall data, the data points containing missing values are deleted; after passing the verification, the data format that passes the verification is converted into a data file in the form of a data point table containing X-axis data and Y-axis data. The output end of the data correction unit is connected to the input end of the format conversion unit, and the output end of the format conversion unit is connected to the input end of the standardization module.

[0041] Specifically, the data correction unit uses the Fourier transform algorithm to perform phase correction on the original data. For the original data of Sample001, there is a phase deviation in the time domain. After being transformed to the frequency domain through Fourier transform, the phase is adjusted;

[0042] The format conversion unit verifies the corrected data. When checking the data of Sample003, it is found that there are missing values in 3 consecutive data points, and the interpolation method is used for filling; for an isolated missing value in Sample007 that does not affect the overall data statistical characteristics, the data point containing the missing value is directly deleted. After the verification passes, all the data is converted into a data file in the form of a data point table with the chemical shift on the X-axis and the peak intensity on the Y-axis.

[0043] The normalization module includes an X-axis data processing unit and a Y-axis data processing unit. The normalization module is used to perform arithmetic processing and cleaning processing on the data point table data file obtained by conversion for the X-axis and Y-axis data respectively. For the X-axis data processing, the mapping processing method is adopted, and for the Y-axis data processing, the maximum value normalization method is adopted to obtain the reference chemical substance spectrum data and the detection spectrum data of the chemical substance to be analyzed. By performing arithmetic processing on the X-axis and Y-axis data respectively, the data differences brought by different instruments and experimental conditions are eliminated, the comparability and analysis accuracy of the data are improved, and the chemical shift and peak intensity ranges are unified. The output end of the normalization module is connected to the input end of the spectral library construction module. Among them, the cleaning processing deletes the data outside the selected range of the X-axis data. The selected range of the X-axis data refers to the chemical shift range [δ min , δ max , where δ min represents the minimum value of the selected chemical shift data; δ max represents the maximum value of the selected chemical shift data.

[0044] The X-axis data processing unit processes the data through the mapping processing method, specifically as follows:

[0045] Select the X-axis data range in the data point table data file obtained by conversion, that is, the chemical shift range [δ min , δ max , and perform mapping processing. The chemical shift range is set according to the type of nuclear magnetic data. The chemical shift range of nuclear magnetic hydrogen spectrum is usually [-4, 20], preferably [-1, 12]. This range covers the chemical shifts of hydrogen atoms in most organic compounds. Further preferably, [-0.2, 9.5] is determined based on the chemical shift intervals that frequently appear in the structures of common organic compounds and experiments, which helps to improve the data processing efficiency and pertinence; for the chemical shift range of nuclear magnetic carbon spectrum, it is usually [-20, 220], preferably [-10, 210], and further preferably [-5, 190];

[0046] Select the target mapping range [v min , v max, the target mapping range is selected from the delimited range, including the chemical shift range and the second range, where the second range includes [0, 90], [-20, 2000], [400, 4000], [3200, 10000], [50, 650] and [50, 6000];

[0047] After the target mapping range is selected, it is processed by the linear mapping algorithm, where the linear mapping algorithm is defined as follows: v = v min +(δ - δ min ) * (v max - v min ) / (δ max - δ min ); where δ represents the chemical shift value; δ min represents the minimum value of the selected chemical shift data; δ max represents the maximum value of the selected chemical shift data; v represents the value obtained after linear mapping, v min represents the minimum value of the target mapping range; v max represents the maximum value of the target mapping range.

[0048] Specifically, taking the hydrogen spectrum data of Sample005 as an example, its original chemical shift range is [-4, 20]. According to the setting, the selected chemical shift range [δ min , δ max is [-0.2, 9.5], and the target mapping range [v min , v max is [0, 90]. Using the linear mapping algorithm v = v min +(δ - δ min ) * (v max - v min ) / (δ max - δ min ), the original chemical shift value is processed. The original chemical shift value is 5, and after calculation, the new chemical shift value v = 49.5.

[0049] The Y-axis data processing unit processes the data by the maximum value normalization method, specifically as follows:

[0050] In the data point table data file obtained by conversion, use the MAX function to find the maximum intensity value I max in the Y-axis sequence, divide all intensity values by the maximum intensity value to obtain the relative intensity value R I = k * I / I max , and process the Y-axis data by the maximum value normalization method, where k represents the scaling factor, k ≠ 0; R I represents the relative intensity value; I maxUse the MAX function to find the maximum intensity value in the Y-axis sequence; I represents the peak intensity value in the Y-axis sequence of the data point table data file.

[0051] Specifically, in the data point table data file of Sample002, use the MAX function to find that the maximum intensity value Imax in the Y-axis sequence is 100. For one of the original peak intensity values I of 30, set the scaling factor k = 1, and calculate the relative intensity value RI = k * I / Imax = 1 * 30 / 100 = 0.3.

[0052] The spectral library construction module includes a data classification and storage unit and a data update and management unit. The data classification and storage unit is used to classify and store the reference chemical substance spectral data according to the spectral type (such as hydrogen spectrum, carbon spectrum, etc.) and the category, source, and application field of the chemical substance, and construct a structured nuclear magnetic resonance spectral database. The output end of the spectral library construction module is connected to the input end of the visualization analysis module.

[0053] Specifically, the spectral data of 10 new compounds after processing are classified and stored in the structured nuclear magnetic resonance spectral database according to the spectral type (hydrogen spectrum), chemical substance category (drug molecule), source (internal R & D and synthesis), and application field (anti-tumor drug R & D).

[0054] The data update and management unit is used to execute the audit process for adding data to the spectral library, specifically as follows:

[0055] After the data acquisition module obtains the original data and marks it, the data format conversion module corrects and converts it, and the standardization module performs operations and cleaning processing on the chemical shift and peak intensity data to obtain the reference chemical substance spectral data, an operation for applying to the database for the reference chemical substance spectral data is performed to update the database;

[0056] Judge whether the reference chemical substance spectral data meets the storage standards. Among them, the storage standards are specifically as follows:

[0057] In terms of data quality: the signal-to-noise ratio of the spectrum reaches the set signal-to-noise ratio threshold to ensure that the spectrum signal is clearly distinguishable; the baseline fluctuation range is within the set baseline fluctuation range; the measurement error of the chemical shift is controlled within the set measurement error range of the chemical shift;

[0058] In terms of data integrity: it includes complete spectral information, including but not limited to spectral type (hydrogen spectrum, carbon spectrum, etc.), basic information of chemical substances (name, structural formula, etc.), experimental conditions (temperature, solvent, etc.);

[0059] When the reference chemical substance spectrum data meets the storage criteria, the reference chemical substance spectrum data that meets the storage criteria is added to the structured nuclear magnetic resonance spectrum database, and the index information of the structured nuclear magnetic resonance spectrum database is updated; when the reference chemical substance spectrum data does not meet the storage criteria, the reference chemical substance spectrum data that does not meet the storage criteria is stored in a temporary data area, and a feedback report is generated. The feedback report includes the reasons why the data does not meet the storage criteria.

[0060] Specifically, a new compound Sample011 was synthesized later. After obtaining its nuclear magnetic resonance data, data quality and integrity checks were carried out. It was required that the signal-to-noise ratio of the spectrum reached more than 30 dB (the actual signal-to-noise ratio of Sample011 was 35 dB), the baseline fluctuation range was within ±0.05 (the actual value was ±0.03), and the chemical shift measurement error was controlled within ±0.01 ppm (the actual value was ±0.008 ppm); in terms of data integrity, it included information such as spectrum type, chemical substance name, structural formula, experimental temperature (25 °C), solvent (deuterated chloroform), etc. Since it met the storage criteria, it was added to the nuclear magnetic resonance spectrum database, and the index information was updated.

[0061] The visualization analysis module is used to present the detection spectrum data of the chemical substance to be analyzed through the spectral curve, display the characteristics of chemical shift and peak intensity through the spectral curve, support users to perform data editing on the spectral curve, support automatic recognition of spectral features, database retrieval, and weighted fitting of multiple spectra. Among them, data editing includes a first function and a second function. The first function is used to select a section of the spectral curve to generate a baseline or a straight line; the second function is used to perform secondary standardization processing on the Y-axis data of the processed new spectral curve; the automatic recognition of spectral features is based on a threshold-based peak detection algorithm to identify spectral peaks and establish a temporary spectral peak data table for correlation retrieval based on the chemical shift of the peaks; there are two correlation retrieval methods for database retrieval: correlation retrieval based on the obtained temporary spectral peak data table and correlation retrieval based on spectral data points. The correlation retrieval uses the spectral similarity algorithm between the compound to be measured and the standard compound for correlation retrieval; the weighted fitting of multiple spectra is used to assign corresponding weights to each standard nuclear magnetic resonance spectrum data point according to the theoretical proportion of each component in the formula. On this basis, multiple standard nuclear magnetic resonance spectrum data points are subjected to weighted summation operations to generate a fitting data point table. The generated fitting data point table is used to simulate the spectral peaks of different ratios of the formula to simulate or restore the formula.

[0062] The first function, generating the baseline or straight line function, is mainly used to remove local peaks in the NMR spectrogram. Taking the solvent peak as an example, when the signal intensity of the deuterium water solvent peak in the spectrogram is higher than that of other surrounding peaks and interferes with the spectrogram analysis (here, when the signal intensity exceeds [N] times the average intensity of the surrounding peaks, this multiple can be set according to the actual situation), where the parameter N represents a multiple standard for judging the signal intensity, the following operations can be adopted:

[0063] Select two endpoints (Xa, Ya) and (Xb, Yb) within the chemical shift range of the solvent peak whose perpendicular distances from the baseline are within a certain threshold (the threshold can be set to, for example, 0.05 times the peak height, that is, when the ratio of the peak height at the endpoint to the highest peak height of the solvent peak is less than 0.05, it is considered close to the baseline). Here, Xa and Xb represent the positions on the chemical shift coordinate axis of the solvent peak, and these two values determine the chemical shift interval of the solvent peak of interest; Ya and Yb represent the perpendicular distances from the spectrogram curve to the baseline at the corresponding chemical shift values Xa and Xb.

[0064] For the points between the selected chemical shifts Xa and Xb, there are two processing methods: filling with the average endpoint value and filling with the straight line equation.

[0065] Filling with the average endpoint value: These points are refilled as (X, 0.5Ya + 0.5Yb), where X represents the chemical shift value corresponding to the filled point when filling with the average endpoint value within the chemical shift interval [Xa, Xb] of the solvent peak.

[0066] Filling with the straight line equation: First, obtain the straight line equation based on these two endpoints, then substitute the X-axis values into the equation to obtain the corresponding Y-axis values, and then refill the points between the selected chemical shifts Xa and Xb as the corresponding points on the straight line.

[0067] The second function is used for the secondary standardization processing of the Y-axis data of the processed new spectrogram curve. Specifically: When the local peak in the NMR spectrogram is generated as the baseline peak and this local peak is the maximum peak, it is necessary to re-normalize the spectrogram to meet the requirements of the standardized spectrogram.

[0068] Specifically, the researchers viewed the spectrogram curve of Sample006 through the visualization analysis module, observed the characteristics of chemical shift and peak intensity, found that the peak shape in a certain area was complex and suspected to have impurity peaks, and removed this local spectral peak by directly operating on the spectrogram curve.

[0069] In Example 2: A nuclear magnetic resonance instrument was used to detect Competitor 1 and Competitor 2. The data acquisition unit obtained the original data. The annotation unit marked each sample with a unique number (SampleJ1, SampleJ2) and the acquisition time (2024-11-15-14:00:00). At the same time, information such as the source of the sample (competitor) and the sample type (surfactant or adhesive raw material) was recorded;

[0070] The data correction unit performed phase correction on the original data, including baseline correction, displacement correction, and phase correction. Taking SampleJ1 as an example, there was a phase deviation in the original data in the time domain. After being transformed to the frequency domain through Fourier transform, the phase was adjusted. The format conversion unit verified the corrected data to check if there were missing values. There were continuous missing values in the SampleJ2 data, and the interpolation method was used for filling; for isolated missing values that did not affect the overall data statistical characteristics, the data points containing the missing value were directly deleted. After passing the verification, all the data was converted into a data file in the form of a data point table with the chemical shift on the X-axis and the peak intensity on the Y-axis;

[0071] For chemical shift processing, the hydrogen spectrum data was processed, and the chemical shift range [δ min ,δ max for mapping processing was [-0.2, 9.5], and the target mapping range [v min ,v max was [0, 90]. Taking a data point with an original chemical shift value of 6 in SampleJ1 as an example, the new chemical shift value v = 54 was calculated using the linear mapping algorithm. The Y-axis data processing unit used the MAX function to find the maximum intensity value I max in the Y-axis sequence. The I max in the SampleJ2 data point table was 120. For a point with an original peak intensity value I of 40, the scaling factor k = 1 was set, and the relative intensity value R I = k*I / I max = 0.33;

[0072] The automatic identification of spectral features was based on a threshold-based peak detection algorithm to identify the spectral peaks of SampleJ1 and SampleJ2, and a temporary spectral peak data table was established. Using two methods, namely, correlation retrieval based on the temporary spectral peak data table and correlation retrieval based on the spectral data points, the similarity algorithm between the compound to be measured and the standard compound spectral diagram was used for retrieval. Standard compounds A and B with a similarity greater than the set threshold to SampleJ1 and standard compounds C and D with a similarity greater than the set threshold to SampleJ2 were retrieved in the spectral library;

[0073] Given the spectral data of standard compounds A, B, C, and D in the spectral library, as well as their chemical structure information, based on experience and preliminary analysis, the researchers determined that competitor 1 is composed of compounds A and B in a certain proportion, and competitor 2 is composed of compounds C and D in a certain proportion. Through inverse weighted fitting, an attempt was made to restore the formulations of the competitors. The theoretical proportion of compound A in competitor 1 was set as x1, and the theoretical proportion of compound B in competitor 1 was 1 - x1; the theoretical proportion of compound C in competitor 2 was set as x2, and the theoretical proportion of compound D in competitor 2 was 1 - x2. Corresponding weights were assigned to each standard NMR spectral data point, and multiple standard NMR spectral data points were subjected to weighted summation operations to generate a fitting data point table. The values of x1 and x2 were continuously adjusted to make the spectral peaks simulated by the fitting data point table similar to the actual spectral peaks of SampleJ1 and SampleJ2. Among them, making the spectral peaks simulated by the fitting data point table similar to the actual spectral peaks of SampleJ1 and SampleJ2 means that the Pearson correlation coefficient between the fitting spectral peak data and the actual spectral peak data of SampleJ1 and SampleJ2 is greater than 0.96. Through iterative calculations, the proportions of compounds A and B in competitor 1, and the proportions of compounds C and D in competitor 2 were determined.

[0074] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.

Claims

1. A nuclear magnetic resonance spectrum library construction and analysis system based on data standardization, characterized in that: The system includes: a data acquisition module, a data format conversion module, a normalization module, a spectral library construction module, and a visualization analysis module; the data acquisition module is used to obtain the original data of the nuclear magnetic resonance instrument and mark the original data; the data format conversion module is used to correct the original data and convert the corrected data format into a data file in the form of a data point table containing chemical shift and peak intensity; the normalization module is used to perform arithmetic processing and cleaning processing on the X-axis data representing chemical shift and the Y-axis data representing peak intensity respectively; the spectral library construction module is used to construct a structured nuclear magnetic resonance spectral database; the visualization analysis module is used for the visualization presentation and analysis of spectral data.

2. The nuclear magnetic resonance spectral library construction and analysis system based on data standardization according to claim 1, characterized in that: The data acquisition module includes a data acquisition unit and a marking unit; the data acquisition unit is used to obtain the original data of the nuclear magnetic resonance instrument; the marking unit is used to perform initial marking on the obtained original data, where the marking content includes the sample number and the acquisition time, and the output end of the data acquisition module is connected to the input end of the data format conversion module.

3. The nuclear magnetic resonance spectrum library construction and analysis system based on data standardization according to claim 2, wherein: The data format conversion module includes a data correction unit and a format conversion unit; the data correction unit is used to perform a correction operation on the obtained original data. The format conversion unit is used to verify the corrected data, and the verification content includes the integrity of the data, that is, to check whether there are missing values in the data. For the data points with missing values and continuous data, interpolation method is used for filling. For the missing values that do not affect the statistical characteristics of the overall data, the data points containing the missing values are deleted; after the verification passes, the data format that passes the verification is converted into a data file in the form of a data point table containing chemical shift and peak intensity. The output end of the data correction unit is connected to the input end of the format conversion unit, and the output end of the format conversion unit is connected to the input end of the normalization module.

4. The nuclear magnetic resonance spectrum library construction and analysis system based on data standardization according to claim 3, characterized in that: The normalization module includes an X-axis data processing unit and a Y-axis data processing unit. The normalization module is used to perform arithmetic processing and cleaning processing on the converted data point table data file for the X-axis and Y-axis data respectively. The mapping processing method is used for the X-axis data processing, and the maximum normalization method is used for the Y-axis data processing to obtain the reference chemical substance spectral data and the detection spectral data of the chemical substance to be analyzed. The output end of the normalization module is connected to the input end of the spectral library construction module. Among them, for the cleaning processing, the data outside the selected range of the X-axis data is deleted. The selected range of the X-axis data refers to the chemical shift range [δmin, δmax], where δmin represents the minimum value of the selected chemical shift data; δmax represents the maximum value of the selected chemical shift data.

5. The nuclear magnetic resonance spectrum library construction and analysis system based on data standardization according to claim 4, characterized in that: The X-axis data processing unit processes the data through the mapping processing method as follows: Select the X-axis data range in the converted data point table data file, that is, the chemical shift range [δ min , δ max , and perform mapping processing. The chemical shift range is set according to the type of nuclear magnetic data; Selected target mapping range [v min , v max ,], where the target mapping range is selected from the defined range, including a chemical shift range and a second range; After selecting the target mapping range, it is processed through a linear mapping algorithm, where the linear mapping algorithm is defined as follows: v = v min +(δ - δ min ) * (v max - v min ) / (δ max - δ min ); where δ represents the chemical shift value; δ min represents the minimum value of the selected chemical shift data; δ max represents the maximum value of the selected chemical shift data; v represents the value obtained after linear mapping, v min represents the minimum value of the target mapping range; v max represents the maximum value of the target mapping range.

6. The nuclear magnetic resonance spectrum library construction and analysis system based on data standardization according to claim 4, characterized in that: The Y-axis data processing unit processes the data through the maximum normalization method as follows: In the converted data point table data file, use the MAX function to find the maximum intensity value I in the Y-axis sequence max , divide all intensity values by the maximum intensity value to obtain the relative intensity value R I = k * I / I max , process the Y-axis data by the maximum value normalization method, where k represents the scaling factor, k ≠ 0; R I represents the relative intensity value; I max Use the MAX function to find the maximum intensity value in the Y-axis sequence; I represents the peak intensity value in the Y-axis sequence of the data point table data file.

7. A nuclear magnetic resonance spectral library construction and analysis system based on data standardization according to claim 6, characterized in that: The spectral library construction module includes a data classification and storage unit and a data update management unit. The data classification and storage unit is used to classify and store the reference chemical substance spectral data according to the spectral type, the category, source, and application field of the chemical substance, and construct a structured nuclear magnetic resonance spectral database. The output end of the spectral library construction module is connected to the input end of the visualization analysis module.

8. A nuclear magnetic resonance spectrum library construction and analysis system based on data standardization according to claim 7, characterized in that: The data update management unit is used to execute the review process for adding data to the spectral library, specifically as follows: After the data acquisition module obtains and marks the original data, the data format conversion module corrects and converts it, and the normalization module performs operations and cleaning processing on the chemical shift and peak intensity data to obtain the reference chemical substance spectral data. Then, an operation for applying to enter the database is performed on the reference chemical substance spectral data to update the database. Judge whether the reference chemical substance spectral data meets the storage criteria. When the reference chemical substance spectral data meets the storage criteria, the reference chemical substance spectral data that meets the storage criteria is added to the structured nuclear magnetic resonance spectral database, and the index information of the structured nuclear magnetic resonance spectral database is updated. When the reference chemical substance spectral data does not meet the storage criteria, the reference chemical substance spectral data that does not meet the storage criteria is stored in a temporary data area, and a feedback report is generated. The feedback report includes the reasons why the data does not meet the storage criteria.

9. The nuclear magnetic resonance spectral library construction and analysis system based on data standardization according to claim 8, characterized in that: The visualization analysis module is used to present the detection spectral data of the chemical substance to be analyzed through a spectral curve, display the characteristics of the chemical shift and peak intensity through the spectral curve, support the user to perform data editing on the spectral curve, support automatic recognition of spectral features, database retrieval, and weighted fitting of multiple spectra.