A systematic method for molecular structure-assisted analysis of heavy organic macromolecules based on exhaustive algorithm
By constructing a molecular fragment database and using exhaustive algorithms to perform molecular fragment combination and nuclear magnetic data comparison, the problem of inefficient molecular structure analysis of heavy organic mixtures is solved, and rapid and accurate molecular structure modeling and analysis are achieved.
Patent Information
- Application Number
- CN202411601878.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-11-11
AI Technical Summary
In the prior art, the molecular structure analysis of organic mixtures of heavy components is inefficient and the lack of a mass spectrometry database makes it impossible to quickly analyze the average molecular structure or the molecular structure of important components of the heavy components.
A molecular fragment database is constructed, a molecular fragment combination is performed based on an exhaustive algorithm, and a nuclear magnetic data comparison is used to achieve rapid analysis of the molecular structure of heavy organic matter.
It realizes rapid molecular structure modeling and analysis of heavy organic matter, saves manpower, and provides the basis for crude oil extraction and chemical kettle resource conversion.
Smart Images

Figure CN119560053B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of organic chemistry, and specifically relates to a system method for auxiliary analysis of the molecular structure of organic heavy macromolecules based on an exhaustive algorithm. Background Art
[0002] Organic molecular structure analysis techniques are widely used in fields such as basic chemical science, materials science, and biopharmaceutical research and development. Molecular structure analysis can help researchers gain a deeper understanding of the properties and reaction mechanisms of substances, providing important insights for chemical synthesis and separation, as well as the development and application of new materials.
[0003] Currently, gas chromatography-mass spectrometry (GC-MS) technology can more accurately determine the structural formula of small molecules in a mixture. This is due to the establishment of a small molecule organic mass spectrum database. By matching standard mass spectrum fingerprints, the composition of small molecule organic mixtures can be quickly analyzed. However, the efficiency of molecular structure analysis of heavy macromolecular organic compounds is relatively low. With the deepening implementation of the concepts of green environmental protection, energy conservation and carbon reduction in China, the development and utilization of unconventional oil and gas resources such as shale oil and tight oil with high heavy component content and high viscosity have been promoted. At the same time, the importance of resource conversion of large amounts of heavy macromolecular organic kettle residual hazardous waste in the chemical production process has also been emphasized, reducing the proportion of hazardous waste incineration. The physical properties of unconventional crude oil in different regions vary, and the types of fracturing fluids and demulsifiers required in the extraction and pretreatment stages vary. This matching process requires a large number of experiments and debugging, which puts tremendous pressure on the extraction department. At the same time, resource utilization of chemical reactor waste, which contains a high concentration of heavy components and has uncertain composition, requires comprehensive structural characterization and analysis to explore its physical, chemical, and toxicological properties to achieve chemical transformation and application. Rapidly predicting the physicochemical properties of large, heavy organic matter, such as heavy unconventional crude oil and heavy chemical reactor waste, is a topic of in-depth research.
[0004] Thanks to the increased computing power of modern computer technology, molecular dynamics simulations and quantum chemical computational simulation methods can be used to predict and analyze molecular structures based on basic physical properties such as boiling point, solubility, viscosity, and binary interaction parameters, as well as chemical reaction mechanisms. If relatively accurate modeling of the macromolecular structures of unconventional crude oil and heavy still residues can be achieved, molecular dynamics can be used to predict the matching results of different types of unconventional crude oil with different fracturing fluids and demulsifiers, reducing the experimental workload. Furthermore, quantum chemical computational simulation methods can be used to analyze the reactive sites of heavy macromolecules in chemical still residues to predict the chemical reactions they can undergo. Alternatively, molecular dynamics simulations can be used to analyze the physical characteristics of chemical still residues and propose chemical conversion or physical application plans for the still residues, ultimately realizing the resource conversion of hazardous still residues.
[0005] However, molecular structure analysis of organic mixtures with a large number of heavy components and complex composition currently relies on multiple characterizations, requiring the combined analysis of multiple methods such as nuclear magnetic resonance (NMR), mass spectrometry, and spectroscopy. This results in low efficiency and a high workload. Furthermore, due to the large variety of complex heavy macromolecules such as crude oil and chemical still residues, and the near-economic value of structural analysis, molecular structure-mass spectrometry databases for these materials are incomplete or nonexistent, resulting in a lack of applicable databases for liquid chromatography-mass spectrometry (LC-MS)-based comparative retrieval methods. Summary of the Invention
[0006] The purpose of this application is to provide a system method for auxiliary analysis of the molecular structure of organic heavy macromolecules based on an exhaustive algorithm, so as to solve the technical problems in the prior art that the molecular structure analysis of heavy component organic mixtures is based on multiple characterizations, the analysis efficiency is low and the workload is large, there is a lack of a mass spectrum database of organic heavy complex macromolecules, and it is impossible to quickly analyze the average molecular structure of the heavy mixture or the molecular structure of several important components.
[0007] In order to achieve the above objectives, the first aspect of the present application provides an auxiliary analysis method for the molecular structure of organic heavy macromolecules based on an exhaustive algorithm, comprising:
[0008] Constructing a molecular fragment database, wherein the molecular fragment database includes molecular fragments with different numbers of carbon atoms, and the types of the molecular fragments include cyclic structure functional groups, linear structure functional groups, terminal heteroatom functional groups, multi-directional linking functional groups, embedded functional groups, double bond-containing functional groups, and triple bond-containing functional groups;
[0009] Extracting molecular fragments from the molecular fragment database based on the number of atoms of each element in the target molecular formula to obtain a plurality of molecular fragment groups that meet the requirements, wherein the sum of the number of atoms of any element in each molecular fragment included in each molecular fragment group is consistent with the number of atoms of the element in the target molecular formula, and the target molecular formula is the average molecular formula of heavy organic matter or the molecular formula of a component of heavy organic matter;
[0010] Traversing the molecular fragment group, combining and splicing each molecular fragment in the currently traversed molecular fragment group, and obtaining corresponding molecular structure models and collecting them into a structure set;
[0011] Deduplication of molecular structure models in the structure set;
[0012] The simulated nuclear magnetic resonance data of the molecular structure model in the structure set and the real nuclear magnetic resonance data of the heavy organic matter are compared, and the analysis results are output.
[0013] In one or more embodiments, the step of extracting molecular fragments from the molecular fragment database based on the number of atoms of each element in the target molecular formula to obtain a plurality of molecular fragment groups that meet the requirements includes:
[0014] Based on the preset number of molecular fragment types and the target molecular formula, constructing a set of equations, wherein the set of equations is used to describe the functional relationship of the number of atoms of any element of each molecular fragment in the molecular fragment group;
[0015] Solving the system of equations to obtain a plurality of integer solution sets, wherein the integer solution sets include the number of atoms of any element of each molecular fragment in the molecular fragment group and the selected number of each molecular fragment;
[0016] The integer solution set is traversed to obtain the molecular formula of each molecular fragment in the molecular fragment group corresponding to the currently traversed integer solution set, and molecular fragments satisfying the molecular formula are extracted from the molecular formula database to obtain a plurality of molecular fragment groups.
[0017] In one or more embodiments, the target molecular formula is C M H N O P N Q S R In the step of constructing a set of equations based on the preset number of molecular fragment types and the target molecular formula, the set of equations is as follows:
[0018]
[0019] Where a is the number of preset molecular fragment types, x i is the number of selected molecular fragments of type i, i=1, 2, 3…a, is the molecular formula of the i-th molecular fragment, .
[0020] In one or more embodiments, each molecular fragment in the molecular fragment database is marked with a splicing site, and any two molecular fragments in the molecular fragment database have different structures or different splicing sites, and the splicing site is arranged at a carbon atom of a carbon-hydrogen bond or a heteroatom bonded to hydrogen of the molecular fragment of the type of cyclic structure functional group, and at one or more terminal carbon atoms of carbon-hydrogen bonds or heteroatoms bonded to hydrogen of other types of molecular fragments.
[0021] In one or more embodiments, the step of splicing and combining the molecular fragments in the currently traversed molecular fragment group includes:
[0022] Traversing the molecular fragments in the molecular fragment group in order based on priority, wherein the molecular fragments are prioritized according to type, and molecular fragments of the same type having a larger number of carbon atoms have a higher priority;
[0023] Each time a target molecular fragment is traversed, all carbon-hydrogen bonds and heteroatom positions that form bonds with hydrogen in the current matrix are scanned as connection sites, wherein the matrix is a molecular structure composed of spliced molecular fragments;
[0024] Traversing the connection sites of the current parent, and each time a target connection site is traversed, splicing the splicing sites of the target molecular fragments with the target connection site one by one, and obtaining a plurality of new parents after the traversal is completed;
[0025] The molecular fragments of the next priority in the molecular fragment group are continuously traversed, and after the traversal is completed, a plurality of molecular structure models corresponding to the molecular fragment group are obtained.
[0026] In one or more embodiments, the priority of the types of the molecular fragments is ranked from large to small as follows: cyclic structure functional group, linear structure functional group, terminal heteroatom functional group, multi-directional link functional group, embedded functional group, double bond-containing functional group, triple bond-containing functional group; and / or,
[0027] If the current parent has no carbon-hydrogen bonds and no heteroatoms bonded to hydrogen, the current parent is deleted.
[0028] In one or more embodiments, the step of removing duplicate molecular structure models in the structure set includes:
[0029] Based on the simulated nuclear magnetic resonance data of the molecular structure models in the structure set, determining whether there are repeated molecular structure models in the structure set;
[0030] If so, delete the duplicate molecular structure models.
[0031] In one or more embodiments, the step of determining whether there are repeated molecular structure models in the structure set based on the simulated NMR data of the molecular structure models in the structure set includes:
[0032] The peak point coordinates of the simulated NMR data of the molecular structure models in the structure set are compared. If there are multiple molecular structure models with exactly the same peak point coordinates, the multiple molecular structure models are repeated.
[0033] In one or more embodiments, the simulated nuclear magnetic resonance data is simulated based on a low-precision nuclear magnetic resonance simulation algorithm.
[0034] In one or more embodiments, the target molecular formula is the average molecular formula of the heavy organic matter;
[0035] The step of comparing the simulated NMR data of the molecular structure model in the structure set with the real NMR data of the heavy organic matter and outputting the analysis result comprises:
[0036] Extracting data from the simulated nuclear magnetic resonance data of the molecular structure model in the structure set and the real nuclear magnetic resonance data of the heavy organic matter to obtain comparison information, wherein the comparison information includes the number of peak points, the abscissa of the peak points, and the peak area;
[0037] Traversing the molecular structure models in the structure set, and determining whether the number of peak points of the currently traversed molecular structure model is the same as that of the heavy organic matter;
[0038] If they are the same, calculating a first root mean square error (RMS) between the horizontal coordinates of the peak points of the currently traversed molecular structure model and the corresponding peak of the heavy organic matter, and determining whether the first RMS error is less than a first threshold;
[0039] If so, calculating a second root mean square error between the peak area of the currently traversed molecular structure model and the corresponding peak of the heavy organic matter, and determining whether the second root mean square error is less than a second threshold;
[0040] If yes, retain the currently traversed molecular structure model and record the first root mean square error and the second root mean square error;
[0041] After the traversal is completed, the molecular structure model with the smallest first root mean square error and second root mean square error is selected as the analysis result.
[0042] In one or more embodiments, in the step of comparing the simulated NMR data of the molecular structure model in the structure set with the real NMR data of the heavy organic matter, the simulated NMR data is simulated based on a high-precision NMR simulation algorithm.
[0043] In one or more embodiments, the target molecular formula is the molecular formula of a component of the heavy organic matter;
[0044] The step of comparing the simulated NMR data of the molecular structure model in the structure set with the real NMR data of the heavy organic matter and outputting the analysis result comprises:
[0045] Extracting data from the simulated nuclear magnetic resonance data of the molecular structure model in the structure set and the real nuclear magnetic resonance data of the heavy organic matter to obtain comparison information, wherein the comparison information includes the number of peak points, the abscissa of the peak points, and the peak area of the corresponding peak;
[0046] Traversing the molecular structure models in the structure set to determine whether the number of peak points of the currently traversed molecular structure model is less than or equal to the number of peak points of the heavy organic matter;
[0047] If so, calculating a first root mean square error of the horizontal coordinates of the peak points of the currently traversed peaks corresponding to the molecular structure model and the heavy organic matter, and determining whether the first root mean square error is less than a third threshold;
[0048] If so, the molecular structure model currently traversed is collected into the candidate structure set of the target molecular formula. After the traversal is completed, the candidate structure set of the target molecular formula is obtained;
[0049] Obtaining candidate structure sets of all components of the heavy organic matter, and exhaustively combining the candidate structure sets to generate a candidate structure combination set of the heavy organic matter, wherein the candidate structure combination set includes a plurality of molecular structure combinations, each of which includes a molecular structure model of each component of the heavy organic matter;
[0050] Traversing the molecular structure combinations in the set of candidate structure combinations, and calculating the peak area of the peak corresponding to the peak point of the currently traversed molecular structure combination based on the comparison information of each molecular structure model in the currently traversed molecular structure combination and the content information of each component of the heavy organic matter;
[0051] Calculating a second root mean square error (RMS) between the peak areas of the currently traversed molecular structure combination and the corresponding peak of the heavy organic matter, and determining whether the second RMS error is less than a fourth threshold;
[0052] If so, the currently traversed molecular structure combination is collected into an output list. After the traversal is completed, the molecular structure of each component in the molecular structure combination included in the output list is used as the analysis result.
[0053] In order to achieve the above-mentioned purpose, the second aspect of the present application provides an auxiliary analysis system for the molecular structure of organic heavy macromolecules based on an exhaustive algorithm, comprising:
[0054] A database construction module is used to construct a molecular fragment database, wherein the molecular fragment database includes molecular fragments with different numbers of carbon atoms, and the types of molecular fragments include cyclic structure functional groups, linear structure functional groups, terminal heteroatom functional groups, multi-directional linking functional groups, embedded functional groups, double bond-containing functional groups, and triple bond-containing functional groups;
[0055] a combination selection confirmation module, configured to extract molecular fragments from the molecular fragment database based on the number of atoms of each element in the target molecular formula, to obtain a plurality of molecular fragment groups that meet the requirements, wherein the sum of the number of atoms of any element in each molecular fragment included in each molecular fragment group is consistent with the number of atoms of the element in the target molecular formula, and the target molecular formula is an average molecular formula of heavy organic matter or a molecular formula of a component of heavy organic matter;
[0056] A splicing modeling module is used to traverse the molecular fragment group, splice and combine the molecular fragments in the currently traversed molecular fragment group, obtain the corresponding molecular structure model and collect it into the structure set;
[0057] a deduplication module, used for deduplicating the molecular structure models in the structure set;
[0058] The comparison and analysis module is used to compare the simulated nuclear magnetic resonance data of the molecular structure model in the structure set with the real nuclear magnetic resonance data of the heavy organic matter, and output the analysis results.
[0059] In order to achieve the above-mentioned object, the third aspect of the present application provides an electronic device, including:
[0060] at least one processor; and
[0061] A memory storing instructions, which, when executed by the at least one processor, enables the at least one processor to perform the auxiliary analysis method for the molecular structure of organic heavy macromolecules based on an exhaustive algorithm as described in any of the above embodiments.
[0062] In order to achieve the above-mentioned objectives, the fourth aspect of the present application provides a machine-readable storage medium storing executable instructions, which, when executed, enable the machine to perform the auxiliary analysis method for the molecular structure of organic heavy macromolecules based on the exhaustive algorithm as described in any of the above embodiments.
[0063] Different from the prior art, the beneficial effects of this application are:
[0064] This application constructs a molecular fragment database. Based on the average molecular formula of heavy organic matter or the molecular formula of each component obtained through characterization, molecular fragments are extracted from the database, and fragment splicing and combination are performed through an exhaustive algorithm to obtain all feasible molecular structure models. After deduplication, the fragments are compared with the real NMR data. This can achieve rapid analysis of the average molecular structure of heavy organic matter or the molecular structure of each component; it can achieve rapid and comprehensive modeling of the molecular structures of a large number of complex organic compounds, and use computer-automated analysis and comparison methods to screen numerous structures, saving manpower to the greatest extent;
[0065] The method of the present application can quickly determine the physicochemical properties of heavy organic macromolecular organic mixtures such as crude oil or chemical still residues at the molecular scale, and can provide an analytical basis for the mechanism of action or transformation principle for crude oil extraction and resource-based physicochemical transformation of chemical still residues, and provide a modeling basis for molecular simulation and quantum chemical simulation of heavy components, and has pioneering significance for the structural analysis of heavy macromolecular organic matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0067] Figure 1 This is a schematic flow chart of an embodiment of the auxiliary analysis method for the molecular structure of organic heavy macromolecules based on an exhaustive algorithm of the present application;
[0068] Figure 2 yes Figure 1 A schematic flow chart of an embodiment corresponding to S200;
[0069] Figure 3 yes Figure 1 A schematic flow chart of an implementation method corresponding to S300;
[0070] Figure 4 yes Figure 1 A schematic flow chart of an implementation method corresponding to S400;
[0071] Figure 5 yes Figure 1 A schematic flow chart of an implementation method corresponding to S500;
[0072] Figure 6 yes Figure 1 A schematic flow chart of another embodiment corresponding to S500;
[0073] Figure 7 This is a schematic structural diagram of an embodiment of the auxiliary analysis system for the molecular structure of organic heavy macromolecules based on an exhaustive algorithm of the present application;
[0074] Figure 8 It is a structural diagram of an embodiment of the electronic device of the present application. DETAILED DESCRIPTION
[0075] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0076] Unconventional crude oils from different regions have varying physical properties, necessitating different types of fracturing fluids and demulsifiers for extraction and pretreatment. This matching process requires extensive experimentation and debugging, placing significant pressure on the extraction department. Furthermore, resource utilization of chemical reactor waste, which contains a high concentration of heavy components and has uncertain composition, requires comprehensive structural characterization and analysis to explore its physical, chemical, and toxicological properties for chemical conversion and application. Rapidly predicting the physical and chemical properties of large, heavy organic matter, such as heavy unconventional crude oil and heavy chemical reactor waste, is a topic of intensive research.
[0077] Currently, molecular structure analysis of heavy macromolecules primarily relies on LC-MS databases. However, these databases are primarily targeted at high-value macromolecules such as traditional Chinese medicine. For more complex heavy macromolecules like heavy chemical still residues or crude oil asphalt, these lack analytical value and, therefore, lack corresponding mass spectrometry databases for comparison and retrieval. Therefore, for organic macromolecules like heavy chemical still residues and heavy crude oil, experimental characterization methods such as proton nuclear magnetic resonance spectroscopy and gel permeation chromatography are currently commonly used to determine the average molecular structure model of heavy mixed organic matter using the Brown-Ladner (BL) method. However, this method is very slow, time-consuming, and inaccurate when used for manual modeling, making it difficult to determine the average molecular structure of heavy macromolecules.
[0078] In the absence of a mass spectrometry database for heavy organic complex macromolecules, it is currently impossible to quickly analyze the average molecular structure of heavy mixtures or the molecular structures of several important components. It is also necessary to rely on multiple characterizations such as nuclear magnetic resonance, mass spectrometry, and spectroscopy for joint analysis, which has low analysis efficiency and a large workload.
[0079] In order to solve the above problems, the applicant has developed an automated analysis method and system that combines actual characterization with molecular simulation. The system can analyze the average molecular structure of heavy organic mixtures, as well as the molecular structures of several key components in heavy organic mixtures. It can accurately give the molecular structure based on NMR characterization data, save human resources, and accelerate the analysis of complex organic macromolecular structures.
[0080] The following details the analysis method of this application. Figure 1 , Figure 1 It is a flow chart of an embodiment of the auxiliary analysis method of the molecular structure of organic heavy macromolecules based on the exhaustive algorithm of the present application.
[0081] like Figure 1 As shown, the auxiliary analysis method includes:
[0082] S100: Build a molecular fragment database.
[0083] The analysis method of the present application constructs and searches for a molecular structure model that meets the requirements by combining molecular fragments. In order to achieve the subsequent determination of the molecular fragment combination, it is first necessary to complete the construction of the molecular fragment database.
[0084] Among them, the molecular fragment database includes molecular fragments with different numbers of carbon atoms, and the types of molecular fragments include cyclic structure functional groups, linear structure functional groups, terminal heteroatom functional groups, multi-directional linking functional groups, embedded functional groups, double bond-containing functional groups and triple bond-containing functional groups.
[0085] It can be understood that the molecular fragments included in the molecular fragment database are intended to be able to be spliced to generate any common biological heavy macromolecules. The maximum number of carbon atoms in the molecular fragments can be selected based on actual needs. The specific elements and functional groups included in each molecular fragment in the molecular fragment database can be adjusted based on the actual application scenario. For example, the functional groups included in the analysis target can be determined by infrared characterization, and on this basis, a molecular fragment database with the functional groups can be constructed.
[0086] Since the difference in splicing positions during the splicing process of molecular fragments will generate completely different molecular structure models, each molecular fragment in the molecular fragment database is also marked with a splicing site, and molecular fragments with the same structure but different splicing sites are identified as different molecular fragments.
[0087] Specifically, any two molecular fragments in the molecular fragment database have different structures or different splicing sites, and the splicing sites are arranged at a carbon atom of a carbon-hydrogen bond or a heteroatom bonded to hydrogen in a molecular fragment of the type of a cyclic structure functional group, and at one or more terminal carbon atoms of a carbon-hydrogen bond or a heteroatom bonded to hydrogen in other types of molecular fragments.
[0088] Here, heteroatoms are atoms other than carbon and hydrogen, including nitrogen, oxygen, sulfur, phosphorus, boron, chlorine, bromine, iodine, etc.
[0089] It should be noted that in this embodiment, the number of splicing sites of different types of molecular fragments is limited. Among them, there is only one splicing site for the cyclic structure functional group, which can be arranged at the carbon atom of the carbon-hydrogen bond or at the heteroatom bonded to hydrogen; other types of molecular fragments can have one or more splicing sites, and their splicing sites can be arranged at the carbon atom of the terminal carbon-hydrogen bond or at the heteroatom bonded to hydrogen at any position.
[0090] For example, the splicing site of a linear functional group without heteroatoms can be arranged at the carbon atom of one carbon-hydrogen bond at either end, or at the carbon atoms of the carbon-hydrogen bonds at both ends, and there can be a maximum of two splicing sites; other functional groups select one or more splicing sites according to their properties.
[0091] Specifically, in one embodiment, the construction of the molecular fragment database can adopt an SMD format file similar to that in the ChemDraw software, which is an ASCII text file. The SMD format file contains the type of elements, the order of arrangement of different atoms, the number of bonds between various atoms and hydrogen atoms, and the order of links between different atoms. The purpose of modifying the structural model of the molecular fragment can be achieved by modifying the order of different atoms and the linking rules between different atoms through the program. The module can add possible molecular fragments to the database as needed, and mark the splicing sites that can be used for splicing in the added molecular fragment structure file, constructing a one-to-one correspondence between "serial number-name-structure file-molecular formula" to form a molecular fragment database.
[0092] S200 , extracting molecular fragments from a molecular fragment database based on the number of atoms of each element in the target molecular formula to obtain a plurality of molecular fragment groups that meet the requirements.
[0093] The sum of the number of atoms of any element in each molecular fragment included in each molecular fragment group is consistent with the number of atoms of the element in the target molecular formula.
[0094] After the molecular fragment database is constructed, molecular fragments can be extracted from the molecular fragment database based on the known target molecular formula, as long as the sum of the atoms of each element in each batch of extracted molecular fragments is consistent with the target molecular formula. Each batch of molecular fragments is taken as a group to obtain several molecular fragment groups.
[0095] The target molecular formula is the average molecular formula of the heavy organic matter or the molecular formula of a component of the heavy organic matter.
[0096] Specifically, the heavy organic matter in this application refers to a mixture of organic macromolecules or a pure organic macromolecule. In one embodiment, when the target substance for analysis in this application is a mixture of organic macromolecules, the target molecular formula can be the average molecular formula of the organic macromolecule mixture. When the target substance for analysis is a pure organic macromolecule, the target molecular formula can also be the molecular formula of the pure organic macromolecule. In this case, the analytical method of this application is applied to analyze the average molecular structure model of the heavy organic matter.
[0097] The average molecular formula of the organic macromolecular mixture and the molecular formula of the pure organic macromolecular substance can be obtained by elemental analysis and average molecular weight analysis, which will not be described in detail here.
[0098] In another embodiment, when the target substance for analysis of the present application is an organic macromolecular mixture, the target molecular formula may also be the molecular formula of a component of the organic macromolecular mixture, i.e., the molecular formula of a component of heavy organic matter. In this case, the analysis method of the present application is applied to analyze the molecular structure models of each component of the heavy organic matter.
[0099] The molecular formula of each component of the heavy organic mixture can be obtained by liquid chromatography-mass spectrometry (LC-MS), which will not be described in detail here.
[0100] The following details how to extract molecular fragments based on the target molecular formula. Figure 2 , Figure 2 yes Figure 1 A flow chart of an implementation method corresponding to S200.
[0101] like Figure 2 As shown, the extraction method of molecular fragments includes:
[0102] S201: Constructing a set of equations based on a preset number of molecular fragment types and a target molecular formula.
[0103] The equation group is used to describe the functional relationship of the number of atoms of any element in each molecular fragment in the molecular fragment group.
[0104] For example, when the target molecular formula is C M H N O P N Q S R When the number of molecular fragment types can be preset as a, the number of selected molecular fragments of type i is expressed as x. i , i=1, 2, 3…a; the molecular formula of the i-th molecular fragment is expressed as The constructed equations can be as follows:
[0105]
[0106] Where, .
[0107] It can be understood that the above set of equations is used to ensure that the sum of the number of atoms of the target element of each group of molecular fragments extracted is equal to the number of atoms of the target element in the target molecular formula; among them, since hydrogen atoms will be replaced during the splicing process of molecular fragments, the number of atoms of hydrogen elements in each group of molecular fragments needs to be converted later.
[0108] S202. Solve the system of equations to obtain several integer solution sets.
[0109] By solving the above equations, a plurality of integer solution sets can be obtained, each of which includes the number of atoms of any element of each molecular fragment in the molecular fragment group and the selected number of each molecular fragment.
[0110] S203 , traversing the integer solution set, obtaining the molecular formula of each molecular fragment in the molecular fragment group corresponding to the currently traversed integer solution set, and extracting molecular fragments that meet the molecular formula from the molecular formula database to obtain a plurality of molecular fragment groups.
[0111] It can be understood that each integer solution set includes the number of atoms of the elements of each molecular fragment in each molecular fragment group. Therefore, a corresponding set of molecular formulas can be obtained based on the integer solution set.
[0112] Based on the molecular formula, molecular fragments that meet the molecular formula can be extracted from the molecular fragment database, thereby realizing the extraction of several groups of molecular fragments.
[0113] It is worth noting that since the molecular fragments in the molecular fragment database of this application are marked with splicing sites, there may be different molecular fragments with the same molecular formula. Therefore, when molecular fragments are extracted based on the molecular formula of a set of molecular fragments that meet the requirements, multiple molecular fragment groups may be obtained, that is, each integer solution set may correspond to one or more molecular fragment groups.
[0114] S300 , traversing the molecular fragment group, combining and splicing each molecular fragment in the currently traversed molecular fragment group, obtaining corresponding molecular structure models and collecting them into a structure set.
[0115] Based on the above S200 , several molecular fragment groups are obtained. By traversing all molecular fragment groups and combining the molecular fragments included in each molecular fragment group, multiple possible molecular structure models of the target molecular formula can be obtained.
[0116] Specifically, in this application, an exhaustive algorithm is used to perform splicing and calling of molecular fragments based on priorities, thereby generating a molecular structure model of all combinations of molecular fragments that conform to the molecular formula.
[0117] See also Figure 3 , Figure 3 yes Figure 1 A flow chart of an implementation method corresponding to S300.
[0118] like Figure 3 As shown, the method for splicing and combining the molecular fragments in the traversed molecular fragment group includes:
[0119] S301 , traverse the molecular fragments in the molecular fragment group in order based on priority.
[0120] S302. Every time a target molecular fragment is traversed, all carbon-hydrogen bonds of the current parent and heteroatom positions that form bonds with hydrogen are scanned as connection sites.
[0121] S303, traversing the connection sites of the current parent, and each time a target connection site is traversed, splicing the splicing sites of the target molecular fragments with the target connection site one by one. After the traversal is completed, several new parents are obtained.
[0122] S304 , continue traversing the molecular fragments of the next priority in the molecular fragment group, and after the traversal is completed, obtain several corresponding molecular structure models in the molecular fragment group.
[0123] In order to improve the splicing efficiency, the present application prioritizes the molecular fragments according to their types, and the molecular fragments of the same type with a larger number of carbon atoms have a higher priority.
[0124] Exemplarily, in one embodiment, the priority of the types of molecular fragments is arranged from large to small as follows: cyclic structure functional group, linear structure functional group, terminal heteroatom functional group, multi-directional linking functional group, embedded functional group, double bond-containing functional group, triple bond-containing functional group.
[0125] At this time, the cyclic structure functional group with the largest number of carbon atoms is first traversed as the parent group, and then the molecular fragments of the next priority are traversed.
[0126] By scanning the matrix, the positions of all its carbon-hydrogen bonds and the positions of heteroatoms bonded to hydrogen can be obtained, and each carbon-hydrogen bond and heteroatom is used as a connection site of the matrix; the splicing sites of the next priority molecular fragments are spliced one by one with the connection sites of the matrix to obtain a new matrix.
[0127] It should be noted that during the splicing process, the parent has multiple connection sites, and the molecular fragments to be spliced may also have multiple splicing sites. For example, the parent has n connection sites and the molecular fragments to be spliced have m splicing sites. At this time, splicing the two can obtain n*m new parents; then continue to traverse the next molecular fragment, splice the next molecular fragment and the n*m new parents respectively, and then continue to traverse until the traversal is completed, and multiple possible molecular structure models can be obtained and collected into the structure set.
[0128] It is understandable that if the current matrix has no carbon-hydrogen bonds and heteroatoms bonded to hydrogen during the traversal process, it means that no site meets the splicing requirements of the current functional group. At this time, there are no splicing conditions and the current matrix can be deleted.
[0129] S400: De-duplicate the molecular structure models in the structure set.
[0130] Based on the above S300, a structure set including multiple molecular structure models is obtained. Due to molecular symmetry, completely identical molecular structure models may be generated during the splicing process of S300, so further deduplication operation is required.
[0131] In one embodiment, whether the molecular structure model is repeated can be determined based on the simulated NMR data of the molecular structure model. Figure 4 , Figure 4 yes Figure 1 A flow chart of an implementation method corresponding to S400.
[0132] like Figure 4 As shown in the figure, the deduplication methods of molecular structure models include:
[0133] S401. Based on the simulated NMR data of the molecular structure models in the structure set, determine whether there are repeated molecular structure models in the structure set.
[0134] Specifically, the method for determining whether there is repetition based on simulated nuclear magnetic resonance data can be:
[0135] The peak point coordinates of the simulated NMR data of the molecular structure models in the structure set are compared. If there are multiple molecular structure models with exactly the same peak point coordinates, the multiple molecular structure models are repeated.
[0136] In one embodiment, in order to improve efficiency, the simulated NMR data used for deduplication can be obtained by simulation based on a low-precision NMR simulation algorithm. For example, the Shoolery empirical formula can be used to predict the NMR data of the molecular structure model. Specifically, the calculation can be performed by programming in an editing language such as C+ or Python, or by automated scripting using known software (such as ChemDraw).
[0137] If there are duplicates, also include:
[0138] S402. Delete duplicate molecular structure models.
[0139] S500: Compare the simulated NMR data of the molecular structure model in the structure set with the real NMR data of the heavy organic matter, and output the analysis results.
[0140] After the deduplication operation based on S400, the simulated NMR data of the molecular structure models in the structure set and the real NMR data of heavy organic matter can be further used to score the molecular structure models, and the final molecular structure model can be selected as the analysis result for output.
[0141] The following describes in detail that, in one embodiment, the analysis method of the present application is used to analyze the average molecular structure model of heavy organic matter, and the target molecular formula is the average molecular formula of the heavy organic matter.
[0142] At this time, the dimensions of the simulated NMR data of the molecular structure model and the real NMR data of the heavy organic matter are the same, so they can be directly compared and judged. For details, please refer to Figure 5 , Figure 5 yes Figure 1 A flow chart of an implementation method corresponding to S500.
[0143] like Figure 5 As shown, the methods for outputting analysis results include:
[0144] S501a. Extract data from the simulated NMR data of the molecular structure model in the structure set and the real NMR data of the heavy organic matter to obtain comparison information.
[0145] The comparison information includes the number of peak points, the horizontal coordinates of the peak points and the peak area of the NMR data.
[0146] S502a: traverse the molecular structure models in the structure set, and determine whether the number of peak points of the currently traversed molecular structure model is the same as that of the heavy organic matter.
[0147] First, determine whether the number of peak points of the simulated NMR data and the real NMR data is the same.
[0148] If they are different, the currently traversed molecular structure model is deleted and the traversal continues to the next molecular structure model.
[0149] If the same, also include:
[0150] S503a: Calculate a first root mean square error (RMS) between the horizontal coordinates of the peak points of the currently traversed molecular structure model and the corresponding peak of the heavy organic matter, and determine whether the first RMS error is less than or equal to a first threshold.
[0151] The first threshold value may be preset based on actual needs.
[0152] For example, the abscissa of the peak point of the simulated NMR data of the currently traversed molecular structure model can be expressed as: ; The horizontal coordinate of the peak point of the real NMR data can be expressed as: ;
[0153] The first root mean square error of the horizontal coordinate of the peak point of the corresponding peak can be expressed as:
[0154] .
[0155] Then determine whether the first root mean square error is less than or equal to the first threshold, that is, determine Is it true?
[0156] If not, delete the currently traversed molecular structure model and continue to traverse the next molecular structure model.
[0157] If established, it also includes:
[0158] S504a: Calculate a second root mean square error between the peak area of the currently traversed molecular structure model and the corresponding peak of the heavy organic matter, and determine whether the second root mean square error is less than a second threshold.
[0159] The second threshold can be preset based on actual needs.
[0160] For example, the peak area corresponding to each peak point of the simulated NMR data of the currently traversed molecular structure model can be expressed as: ; The peak area corresponding to each peak point of the real NMR data can be expressed as: ;
[0161] The second root mean square error of the peak area of the corresponding peak can be expressed as:
[0162] .
[0163] Then determine whether the second root mean square error is less than or equal to the second threshold, that is, determine Is it true?
[0164] If not, delete the currently traversed molecular structure model and continue to traverse the next molecular structure model.
[0165] If established, it also includes:
[0166] S505a: retain the currently traversed molecular structure model, and record the first root mean square error and the second root mean square error.
[0167] S506a: After the traversal is completed, the molecular structure model with the smallest first root mean square error and second root mean square error is selected as the analysis result.
[0168] Based on the above steps, the molecular structure model with the highest similarity between the simulated NMR data and the real NMR data can be selected from the structure set as the average molecular structure analysis result of the heavy organic matter, thereby realizing the analysis and prediction of the average molecular structure of the heavy organic matter.
[0169] In another embodiment, the analysis method of the present application can be applied to analyze the molecular structure models of all components of heavy organic matter based on the analysis test results of LC-MS, and the target molecular formula is the molecular formula of the components of heavy organic matter.
[0170] At this time, the dimensions of the simulated NMR data of the molecular structure model and the real NMR data of the heavy organic compound are different. It is also necessary to combine the simulated NMR data of all components of the heavy organic mixture obtained by LC-MS analysis, and then compare them with the real NMR data of the heavy organic compound for judgment. For details, please refer to Figure 6 , Figure 6 yes Figure 1 A flow chart of another implementation method corresponding to S500 in FIG.
[0171] like Figure 6 As shown, the methods for outputting analysis results include:
[0172] S501b. Extract data from the simulated NMR data of the molecular structure model in the structure set and the real NMR data of the heavy organic matter to obtain comparison information.
[0173] Similar to S501a, the comparison information includes the number of peak points, the horizontal coordinates of the peak points, and the peak areas of the corresponding peaks.
[0174] S502b: traverse the molecular structure models in the structure set, and determine whether the number of peak points of the currently traversed molecular structure model is less than or equal to the number of peak points of heavy organic matter.
[0175] Since the molecular structure model in this embodiment corresponds to a component of heavy organic matter, the number of peak points of the simulated NMR data of the currently traversed molecular structure model must be less than or equal to the number of peak points of the real NMR data.
[0176] If not, the currently traversed molecular structure model is deleted and the next molecular structure model is traversed again.
[0177] If yes, also include:
[0178] S503b: Calculate a first root mean square error (RMS) between the horizontal coordinates of the peak points of the currently traversed molecular structure model and the corresponding peak of the heavy organic matter, and determine whether the first RMS error is less than a third threshold.
[0179] The third threshold value may be preset based on actual needs.
[0180] Since the number of peak points in the real NMR data of heavy organic matter must be greater than or equal to the number of peak points in the currently traversed molecular structure model, when calculating the horizontal coordinate of the peak point of the corresponding peak, after confirming the target peak point of the simulated NMR data, it is necessary to select the peak point closest to the target peak point from the real NMR data and calculate the root mean square error.
[0181] For example, the abscissa of each peak point of the simulated NMR data can be expressed as: ; The horizontal coordinates of each peak point of the real NMR data can be expressed as: .
[0182] The calculation formula of the first root mean square error can be expressed as:
[0183] .
[0184] Then determine whether the first root mean square error is less than the third threshold, that is, determine .
[0185] If not, delete the currently traversed molecular structure model and continue to traverse the next molecular structure model.
[0186] If established, it also includes:
[0187] S504b, collecting the currently traversed molecular structure models into a candidate structure set of the target molecular formula. After the traversal is completed, the candidate structure set of the target molecular formula is obtained.
[0188] Based on the above steps, a candidate structure set of the target molecular formula is obtained, and the candidate structure set includes several molecular structure models of the target molecular formula that temporarily meet the requirements.
[0189] After that, further judgment needs to be made based on the molecular structure models of all components of the heavy organic matter. Therefore, it is necessary to wait until all candidate structure sets of all components of the heavy organic matter are obtained.
[0190] S505b: Obtain candidate structure sets of all components of the heavy organic matter, and exhaustively combine the candidate structure sets to generate a candidate structure combination set of the heavy organic matter.
[0191] Specifically, it is first determined whether the analysis of the candidate structure sets of all components of the heavy organic matter is completed. If not, the analysis process of the candidate structure sets of other target molecular formulas is continued until the analysis of the candidate structure sets of all components of the heavy organic matter is completed.
[0192] At this point, a candidate structure set for all components of the heavy organic matter is obtained for the next step of exhaustive combination. Exhaustive combination specifically involves selecting one molecular structure model from each candidate structure set for all components of the heavy organic matter and combining them together to obtain all possible combinations between different components of the heavy organic matter, thereby generating a candidate structure combination set for the heavy organic matter. The candidate structure combination set includes several molecular structure combinations, each of which includes a molecular structure model for each component of the heavy organic matter.
[0193] For example, when the heavy organic matter includes two components, the candidate structure sets of the two components each include two molecular structure models. In this case, the generated candidate structure combination set includes four molecular structure combinations.
[0194] S506b. Traverse the molecular structure combinations in the set of candidate structure combinations, and calculate the peak area of the peak corresponding to the peak point of the currently traversed molecular structure combination based on the comparison information of each molecular structure model in the currently traversed molecular structure combination and the content information of each component of the heavy organic matter.
[0195] When the molecular formulas of the components of heavy organic matter are characterized by liquid chromatography-mass spectrometry (LC-MS), the content of each component can also be obtained simultaneously. For each molecular structure combination that meets the requirements, the peak area corresponding to the peak point of the currently traversed molecular structure combination can be obtained based on the simulated NMR data of each molecular structure model and the content of each component.
[0196] For example, the molecular structure combination includes a molecular structure models, and the peak area corresponding to the peak point of the i-th molecular structure model can be expressed as: , the content of each molecular structure model is w i , where i=1, 2, 3....a.
[0197] The peak area of the peak point of the molecular structure combination can be expressed as: .
[0198] S507b: Calculate a second root mean square error between the peak area of the currently traversed molecular structure combination and the corresponding peak of the heavy organic matter, and determine whether the second root mean square error is less than a fourth threshold.
[0199] The fourth threshold can be set based on actual needs.
[0200] Referring to the example in S506b, the peak area corresponding to the peak point of the real NMR data can be expressed as: .
[0201] Furthermore, the calculation formula of the second root mean square error can be expressed as: .
[0202] Then determine whether the second root mean square error is less than or equal to the second threshold, that is, determine Is it true?
[0203] If not, delete the currently traversed molecular structure combination and continue to traverse the next molecular structure combination.
[0204] If established, it also includes:
[0205] S508b: Collect the currently traversed molecular structure combinations into an output list. After the traversal is completed, the molecular structure combinations included in the output list are used as analysis results.
[0206] After all molecular structure combinations are traversed, an output list including several molecular structure combinations is obtained. Each molecular structure combination includes a molecular structure model of each component of the heavy organic matter, completing the molecular structure analysis and prediction of each component of the heavy organic matter.
[0207] This application also provides an auxiliary analysis system for the molecular structure of organic heavy macromolecules based on an exhaustive algorithm. Figure 7 , Figure 7 It is a structural schematic diagram of an embodiment of the auxiliary analysis system of the molecular structure of organic heavy macromolecules based on the exhaustive algorithm of the present application.
[0208] like Figure 7 As shown, the auxiliary analysis system includes a database construction module 21, a combination selection confirmation module 22, a splicing modeling module 23, a deduplication module 24, and a comparison analysis module 25.
[0209] The database construction module 21 is used to construct a molecular fragment database, which includes molecular fragments with different numbers of carbon atoms. The types of molecular fragments include cyclic structure functional groups, linear structure functional groups, terminal heteroatom functional groups, multi-directional link functional groups, embedded functional groups, double bond functional groups, and triple bond functional groups.
[0210] The combination selection confirmation module 22 is used to extract molecular fragments from the molecular fragment database based on the number of atoms of each element in the target molecular formula to obtain a plurality of molecular fragment groups that meet the requirements, wherein the sum of the number of atoms of any element in each molecular fragment included in each molecular fragment group is consistent with the number of atoms of the element in the target molecular formula, and the target molecular formula is the average molecular formula of the heavy organic matter or the molecular formula of a component of the heavy organic matter;
[0211] The splicing modeling module 23 is used to traverse the molecular fragment group, splice and combine the molecular fragments in the currently traversed molecular fragment group, and obtain the corresponding molecular structure model and collect it into the structure set;
[0212] The deduplication module 24 is used to dedupe the molecular structure models in the structure set;
[0213] The comparison and analysis module 25 is used to compare the simulated NMR data of the molecular structure model in the structure set with the real NMR data of the heavy organic matter and output the analysis results.
[0214] In one embodiment, the auxiliary analysis system also includes a nuclear magnetic resonance data simulation module 26, which is used to simulate the nuclear magnetic resonance data of the molecular structure model based on a high-precision nuclear magnetic resonance simulation algorithm and / or a low-precision nuclear magnetic resonance simulation algorithm, and is called by the deduplication module 24 and the comparison analysis module 25.
[0215] As above Figures 1 to 6 , the auxiliary analysis method for the molecular structure of organic heavy macromolecules based on the exhaustive algorithm according to the embodiment of this specification is described. The details mentioned in the above description of the method embodiment are also applicable to the auxiliary analysis device for the molecular structure of organic heavy macromolecules based on the exhaustive algorithm according to the embodiment of this specification. The above auxiliary analysis device for the molecular structure of organic heavy macromolecules based on the exhaustive algorithm can be implemented by hardware, software, or a combination of hardware and software.
[0216] This application also provides an electronic device, see Figure 8 , Figure 8 This is a schematic diagram of the structure of an embodiment of the electronic device of the present application. Figure 8 As shown, the electronic device 30 may include at least one processor 31, a memory 32 (e.g., a non-volatile memory), a storage 33, and a communication interface 34, and the at least one processor 31, the storage 32, the storage 33, and the communication interface 34 are connected together via a bus 35. The at least one processor 31 executes at least one computer-readable instruction stored or encoded in the storage 32.
[0217] It should be understood that the computer executable instructions stored in the memory 32, when executed, cause at least one processor 31 to perform the above combined operations in various embodiments of this specification. Figure 1-Figure 5 Describes the various operations and functions.
[0218] In the embodiments of the present specification, the electronic device 30 may include but is not limited to: a personal computer, a server computer, a workstation, a desktop computer, a laptop computer, a notebook computer, a mobile electronic device, a smart phone, a tablet computer, a cellular phone, a personal digital assistant (PDA), a handheld device, a messaging device, a wearable electronic device, a consumer electronic device, and the like.
[0219] According to one embodiment, a program product such as a machine-readable medium is provided. The machine-readable medium may have instructions (i.e., the above-mentioned elements implemented in software form), which, when executed by a machine, causes the machine to perform the above-mentioned combined embodiments of the present specification. Figure 1-Figure 5 Specifically, a system or device equipped with a readable storage medium can be provided, on which software program codes for implementing the functions of any of the above-mentioned embodiments are stored, and a computer or processor of the system or device can be enabled to read and execute the instructions stored in the readable storage medium.
[0220] In this case, the program code itself read from the machine-readable medium can implement the functions of any one of the above embodiments, and thus the machine-readable code and the machine-readable storage medium storing the machine-readable code constitute part of this specification.
[0221] Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks, magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code may be downloaded from a server computer or a cloud via a communication network.
[0222] Those skilled in the art will appreciate that the various embodiments disclosed above may be modified and altered in various ways without departing from the essence of the invention. Therefore, the scope of protection of this specification shall be defined by the appended claims.
[0223] It should be noted that not all steps and units in the above processes and system structure diagrams are required, and certain steps or units can be omitted according to actual needs. The execution order of each step is not fixed and can be determined as needed. The device structure described in the above embodiments can be a physical structure or a logical structure, that is, some units may be implemented by the same physical client, or some units may be implemented by multiple physical clients, or may be implemented by certain components in multiple independent devices.
[0224] In the above embodiments, hardware unit or module can be realized by mechanical means or electrical means. For example, a hardware unit, module or processor can include permanent dedicated circuit or logic (such as special processor, FPGA or ASIC) to complete the corresponding operation. Hardware unit or processor can also include programmable logic or circuit (such as general purpose processor or other programmable processor), can be temporarily set up to complete the corresponding operation by software. Concrete implementation (mechanical means or dedicated permanent circuit or temporary circuit) can be determined based on cost and time consideration.
[0225] The specific embodiments described above in conjunction with the accompanying drawings describe exemplary embodiments, but do not represent all embodiments that can be implemented or fall within the scope of protection of the claims. The term "exemplary" used throughout this specification means "used as an example, instance or illustration" and does not mean "preferred" or "having advantages" over other embodiments. For the purpose of providing an understanding of the described technology, the specific embodiments include specific details. However, these technologies can be implemented without these specific details. In some instances, in order to avoid obscuring the concepts of the described embodiments, well-known structures and devices are shown in block diagram form.
[0226] The foregoing description of the present disclosure is provided to enable any person skilled in the art to implement or use the present disclosure. Various modifications to the present disclosure will be readily apparent to those skilled in the art, and the general principles herein may be applied to other variations without departing from the scope of the present disclosure. Therefore, the present disclosure is not limited to the examples and designs described herein, but is intended to be consistent with the widest range of principles and novel features disclosed herein.
Claims
1. An auxiliary analysis method for the molecular structure of organic heavy macromolecules based on an exhaustive algorithm, characterized in that: include: Constructing a molecular fragment database, wherein the molecular fragment database includes molecular fragments with different numbers of carbon atoms, and the types of the molecular fragments include cyclic structure functional groups, linear structure functional groups, terminal heteroatom functional groups, multi-directional linking functional groups, embedded functional groups, double bond-containing functional groups, and triple bond-containing functional groups; Extracting molecular fragments from the molecular fragment database based on the number of atoms of each element in the target molecular formula to obtain a plurality of molecular fragment groups that meet the requirements, wherein the sum of the number of atoms of any element in each molecular fragment included in each molecular fragment group is consistent with the number of atoms of the element in the target molecular formula, and the target molecular formula is the average molecular formula of heavy organic matter or the molecular formula of a component of heavy organic matter; Traversing the molecular fragment group, combining and splicing each molecular fragment in the currently traversed molecular fragment group, and obtaining corresponding molecular structure models and collecting them into a structure set; Deduplication of molecular structure models in the structure set; comparing the simulated NMR data of the molecular structure model in the structure set with the real NMR data of the heavy organic matter, and outputting an analysis result; The step of extracting molecular fragments from the molecular fragment database based on the number of atoms of each element in the target molecular formula to obtain a plurality of molecular fragment groups that meet the requirements includes: Based on the preset number of molecular fragment types and the target molecular formula, constructing a set of equations, wherein the set of equations is used to describe the functional relationship of the number of atoms of any element of each molecular fragment in the molecular fragment group; Solving the system of equations to obtain a plurality of integer solution sets, wherein the integer solution sets include the number of atoms of any element of each molecular fragment in the molecular fragment group and the selected number of each molecular fragment; Traversing the integer solution set, obtaining the molecular formula of each molecular fragment in the molecular fragment group corresponding to the currently traversed integer solution set, and extracting molecular fragments that satisfy the molecular formula from the molecular fragment database to obtain a plurality of molecular fragment groups; The step of splicing and combining the molecular fragments in the currently traversed molecular fragment group includes: Traversing the molecular fragments in the molecular fragment group in order based on priority, wherein the molecular fragments are prioritized according to type, and molecular fragments of the same type having a larger number of carbon atoms have a higher priority; Each time a target molecular fragment is traversed, all carbon-hydrogen bonds and heteroatom positions that form bonds with hydrogen in the current matrix are scanned as connection sites, wherein the matrix is a molecular structure composed of spliced molecular fragments; Traversing the connection sites of the current parent, and each time a target connection site is traversed, splicing the splicing sites of the target molecular fragments with the target connection site one by one, and obtaining a plurality of new parents after the traversal is completed; The molecular fragments of the next priority in the molecular fragment group are continuously traversed, and after the traversal is completed, a plurality of molecular structure models corresponding to the molecular fragment group are obtained.
2. The auxiliary analysis method according to claim 1, characterized in that: The target molecular formula is C M H N O P N Q S R In the step of constructing a set of equations based on the preset number of molecular fragment types and the target molecular formula, the set of equations is as follows: ; Where a is the number of preset molecular fragment types, x i is the number of selected molecular fragments of type i, i=1, 2, 3…a, is the molecular formula of the i-th molecular fragment, .
3. The auxiliary analysis method according to claim 1, characterized in that Each molecular fragment in the molecular fragment database is marked with a splicing site, and any two molecular fragments in the molecular fragment database have different structures or different splicing sites. The splicing site is arranged at a carbon atom of a carbon-hydrogen bond or a heteroatom bonded to hydrogen of the molecular fragment of the type of cyclic structure functional group, and at one or more terminal carbon atoms of a carbon-hydrogen bond or a heteroatom bonded to hydrogen of other types of molecular fragments.
4. The auxiliary analysis method according to claim 1, characterized in that The priority of the types of the molecular fragments is arranged from large to small as follows: cyclic structure functional group, linear structure functional group, terminal heteroatom functional group, multi-directional linking functional group, embedded functional group, double bond-containing functional group, triple bond-containing functional group; and / or, If the current parent has no carbon-hydrogen bonds and no heteroatoms bonded to hydrogen, the current parent is deleted.
5. The auxiliary analysis method according to claim 1, characterized in that: The step of removing duplicate molecular structure models in the structure set includes: Based on the simulated nuclear magnetic resonance data of the molecular structure models in the structure set, determining whether there are repeated molecular structure models in the structure set; If so, delete the duplicate molecular structure models.
6. The auxiliary analysis method according to claim 5, characterized in that: The step of determining whether there are repeated molecular structure models in the structure set based on the simulated NMR data of the molecular structure models in the structure set comprises: The peak point coordinates of the simulated NMR data of the molecular structure models in the structure set are compared. If there are multiple molecular structure models with exactly the same peak point coordinates, the multiple molecular structure models are repeated.
7. The auxiliary analysis method according to claim 5, characterized in that: The simulated nuclear magnetic resonance data are obtained by simulation based on a low-precision nuclear magnetic resonance simulation algorithm.
8. The auxiliary analysis method according to claim 1, characterized in that: The target molecular formula is the average molecular formula of the heavy organic matter; The step of comparing the simulated NMR data of the molecular structure model in the structure set with the real NMR data of the heavy organic matter and outputting the analysis result comprises: Extracting data from the simulated nuclear magnetic resonance data of the molecular structure model in the structure set and the real nuclear magnetic resonance data of the heavy organic matter to obtain comparison information, wherein the comparison information includes the number of peak points, the abscissa of the peak points, and the peak area; Traversing the molecular structure models in the structure set, and determining whether the number of peak points of the currently traversed molecular structure model is the same as that of the heavy organic matter; If they are the same, calculating a first root mean square error (RMS) between the horizontal coordinates of the peak points of the currently traversed molecular structure model and the corresponding peak of the heavy organic matter, and determining whether the first RMS error is less than a first threshold; If so, calculating a second root mean square error between the peak area of the currently traversed molecular structure model and the corresponding peak of the heavy organic matter, and determining whether the second root mean square error is less than a second threshold; If yes, retain the currently traversed molecular structure model and record the first root mean square error and the second root mean square error; After the traversal is completed, the molecular structure model with the smallest first root mean square error and second root mean square error is selected as the analysis result; and / or, In the step of comparing the simulated nuclear magnetic resonance data of the molecular structure model in the structure set with the real nuclear magnetic resonance data of the heavy organic matter, the simulated nuclear magnetic resonance data is simulated based on a high-precision nuclear magnetic resonance simulation algorithm.
9. The auxiliary analysis method according to claim 1, characterized in that: The target molecular formula is the molecular formula of the component of the heavy organic matter; The step of comparing the simulated NMR data of the molecular structure model in the structure set with the real NMR data of the heavy organic matter and outputting the analysis result comprises: Extracting data from the simulated nuclear magnetic resonance data of the molecular structure model in the structure set and the real nuclear magnetic resonance data of the heavy organic matter to obtain comparison information, wherein the comparison information includes the number of peak points, the abscissa of the peak points, and the peak area of the corresponding peak; Traversing the molecular structure models in the structure set, and determining whether the number of peak points of the currently traversed molecular structure model is less than or equal to the number of peak points of the heavy organic matter; If so, calculating a first root mean square error of the horizontal coordinates of the peak points of the currently traversed peaks corresponding to the molecular structure model and the heavy organic matter, and determining whether the first root mean square error is less than a third threshold; If so, the molecular structure model currently traversed is collected into the candidate structure set of the target molecular formula. After the traversal is completed, the candidate structure set of the target molecular formula is obtained; Obtaining candidate structure sets of all components of the heavy organic matter, and exhaustively combining the candidate structure sets to generate a candidate structure combination set of the heavy organic matter, wherein the candidate structure combination set includes a plurality of molecular structure combinations, each of which includes a molecular structure model of each component of the heavy organic matter; Traversing the molecular structure combinations in the set of candidate structure combinations, and calculating the peak area of the peak corresponding to the peak point of the currently traversed molecular structure combination based on the comparison information of each molecular structure model in the currently traversed molecular structure combination and the content information of each component of the heavy organic matter; Calculating a second root mean square error (RMS) between the peak areas of the currently traversed molecular structure combination and the corresponding peak of the heavy organic matter, and determining whether the second RMS error is less than a fourth threshold; If so, the molecular structure combination currently traversed is collected into an output list. After the traversal is completed, the molecular structure of each component in the molecular structure combination included in the output list is used as the analysis result.
10. An auxiliary analysis system for the molecular structure of organic heavy macromolecules based on an exhaustive algorithm, characterized in that: include: A database construction module is used to construct a molecular fragment database, wherein the molecular fragment database includes molecular fragments with different numbers of carbon atoms, and the types of molecular fragments include cyclic structure functional groups, linear structure functional groups, terminal heteroatom functional groups, multi-directional linking functional groups, embedded functional groups, double bond-containing functional groups, and triple bond-containing functional groups; a combination selection confirmation module, configured to extract molecular fragments from the molecular fragment database based on the number of atoms of each element in the target molecular formula, to obtain a plurality of molecular fragment groups that meet the requirements, wherein the sum of the number of atoms of any element in each molecular fragment included in each molecular fragment group is consistent with the number of atoms of the element in the target molecular formula, and the target molecular formula is an average molecular formula of heavy organic matter or a molecular formula of a component of heavy organic matter; A splicing modeling module is used to traverse the molecular fragment group, splice and combine the molecular fragments in the currently traversed molecular fragment group, obtain the corresponding molecular structure model and collect it into the structure set; a deduplication module, configured to dedupe the molecular structure models in the structure set; a comparison and analysis module, configured to compare the simulated NMR data of the molecular structure models in the structure set with the real NMR data of the heavy organic matter, and output an analysis result; The step of extracting molecular fragments from the molecular fragment database based on the number of atoms of each element in the target molecular formula to obtain a plurality of molecular fragment groups that meet the requirements includes: Based on the preset number of molecular fragment types and the target molecular formula, constructing a set of equations, wherein the set of equations is used to describe the functional relationship of the number of atoms of any element of each molecular fragment in the molecular fragment group; Solving the system of equations to obtain a plurality of integer solution sets, wherein the integer solution sets include the number of atoms of any element of each molecular fragment in the molecular fragment group and the selected number of each molecular fragment; Traversing the integer solution set, obtaining the molecular formula of each molecular fragment in the molecular fragment group corresponding to the currently traversed integer solution set, and extracting molecular fragments that satisfy the molecular formula from the molecular fragment database to obtain a plurality of molecular fragment groups; The step of splicing and combining the molecular fragments in the currently traversed molecular fragment group includes: Traversing the molecular fragments in the molecular fragment group in order based on priority, wherein the molecular fragments are prioritized according to type, and molecular fragments of the same type having a larger number of carbon atoms have a higher priority; Each time a target molecular fragment is traversed, all carbon-hydrogen bonds and heteroatom positions that form bonds with hydrogen in the current matrix are scanned as connection sites, wherein the matrix is a molecular structure composed of spliced molecular fragments; Traversing the connection sites of the current parent, and each time a target connection site is traversed, splicing the splicing sites of the target molecular fragments with the target connection site one by one, and obtaining a plurality of new parents after the traversal is completed; The molecular fragments of the next priority in the molecular fragment group are continuously traversed, and after the traversal is completed, a plurality of molecular structure models corresponding to the molecular fragment group are obtained.
11. An electronic device comprising: at least one processor; as well as A memory storing instructions, which, when executed by the at least one processor, causes the at least one processor to execute the auxiliary analysis method for the molecular structure of organic heavy macromolecules based on an exhaustive algorithm as described in any one of claims 1 to 9.
12. A machine-readable storage medium storing executable instructions, which, when executed, enable the machine to perform the auxiliary analysis method for the molecular structure of organic heavy macromolecules based on an exhaustive algorithm according to any one of claims 1 to 9.
Citation Information
Patent Citations
Organic molecule virtual screening library construction method, device, equipment and medium
CN118412066A
Method for generating a database of molecular fragments
US20020062307A1