A method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry
By using graph theory and quantum chemistry-based methods, the biomass molecule is abstracted into a directed graph model and combined with quantum chemical calculations, which solves the problem of the difficulty in analyzing the biomass pyrolysis reaction mechanism in existing technologies, and realizes accurate analysis of the biomass pyrolysis reaction network and efficient prediction of product distribution.
Patent Information
- Application Number
- CN202511108744.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-08
AI Technical Summary
Existing technologies are insufficient to fully reveal the microscopic reaction mechanisms of biomass pyrolysis processes, and cannot achieve targeted control of target products. Traditional experimental analysis methods cannot capture reactive intermediates such as transient free radicals, and quantum chemical calculations are difficult to apply to the overall reaction analysis of complex biomass systems.
Using a graph theory and quantum chemistry-based approach, biomass molecules are abstracted into directed graph models. By combining molecular mechanics optimization and quantum chemical calculations, potential transformation paths are marked, and a hash algorithm is used for deduplication. Basic reaction blocks are predefined, and energy-optimal paths are screened to achieve precise analysis of chemical bond breaking paths and intermediate product transformations.
It achieves precise analysis of the biomass pyrolysis reaction network, improves the accuracy of product distribution prediction, balances the computational efficiency and accuracy of complex systems, and clearly understands the conversion path from raw materials to intermediate products.
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomass pyrolysis reaction technology, and more specifically, to a method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry. Background Art
[0002] Biomass pyrolysis refers to the process under anaerobic or oxygen-deficient conditions where complex organic macromolecules composed of cellulose, hemicellulose, and lignin undergo thermally induced chemical bond breaking and recombination, ultimately producing syngas, bio-oil, and solid char. Experimental studies have shown that when the temperature is in the range of 300-800℃, biomass macromolecules undergo thousands of parallel and tandem reactions, involving elementary reaction steps such as C / C bond breaking, hydrogen atom migration, and cyclization rearrangement, forming a complex reaction network containing free radical intermediates, small molecule products, and polycyclic aromatic hydrocarbons.
[0003] Existing experimental methods are insufficient to fully reveal the microscopic reaction mechanisms of biomass pyrolysis, thus hindering the targeted regulation of target products. Traditional experimental analytical methods have inherent limitations in analyzing biomass pyrolysis reaction networks: while thermogravimetric analysis (TGA) can obtain mass change curves during pyrolysis, it cannot reveal specific chemical bond breaking locations and intermediate product transformation pathways; while gas chromatography-mass spectrometry (GC-MS) can perform qualitative and quantitative analysis of some stable products, it struggles to capture reactive intermediates such as transient free radicals and cannot elucidate the synergistic reaction mechanisms among multiple components.
[0004] With the development of computational chemistry, molecular model-based simulation methods have been increasingly applied to related research. Early simulation methods often employed simplified reaction force field models, reducing biomass macromolecules to combinations of structural units. This approach neglected the detailed features and spatial configurations of molecular structures, leading to inaccurate descriptions of complex molecular structures and significant deviations from experimental values when predicting key reaction parameters. Later quantum chemical computational methods, while capable of accurately calculating some microscopic reaction parameters, were difficult to apply to the overall reaction analysis of complex biomass systems. Another analytical method, which abstracts molecular structures into graph structures, struggles to accurately predict the competition of reaction pathways involving different functional groups when dealing with biomass molecules containing multiple functional groups, due to a lack of effective characterization of the reactive properties of these functional groups. Summary of the Invention
[0005] To overcome the shortcomings of existing technologies, this invention provides a method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry.
[0006] The technical solution of this invention is as follows:
[0007] A method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry, including...
[0008] Analyze the molecular structure of the starting reactants in the biomass pyrolysis reaction, identify the atomic connection mode and chemical bond characteristics of the starting reactants, abstract atoms as nodes, chemical bonds as edges, and store and generate molecular models in a directed graph data structure;
[0009] The molecular model is optimized so that the energy gradient of the molecule is less than a certain stable energy gradient threshold.
[0010] In the molecular model, potential transformation pathways are marked and virtual edges and virtual nodes with fixed attributes are set;
[0011] The information of each atom is concatenated into a string according to the atom numbering rules, and a unique hash value is assigned to each atom using a hash algorithm.
[0012] Extract common functional group structures from molecular models and encapsulate them into subgraph modules;
[0013] The fundamental parameters of chemical bonds, including the chemical bond dissociation energy, are obtained through quantum chemical calculations. The rate constants of the elementary reaction steps associated with the breaking of a specific chemical bond are calculated and then normalized to convert the rate constants of the elementary reaction steps into the probability weights of the breaking of the chemical bond in that specific reaction path.
[0014] Predefined basic reaction blocks, which include bond change rules and matrix transformation methods;
[0015] Traverse the molecular model. If the chemical bond dissociation energy is less than the threshold, the chemical bond is determined to be a candidate for breaking. Select the probability weight of breaking according to the breaking reaction corresponding to the chemical bond. Determine whether to perform breaking by sampling according to probability and mark it as a breakable bond. Call the basic reaction building block to reassemble the broken molecular model fragments to obtain potential reaction products. Calculate the hash value of the potential reaction products and save them after deduplication based on the hash value.
[0016] Assign weights to the edges, where the weights represent the activation energy of the chemical bond breaking reaction.
[0017] Define a distance array and construct a priority queue. Add the starting molecular state node and the initial distance to the priority queue. Loop: Select the node with the smallest cumulative activation energy from the priority queue (this node represents the current molecular configuration). Traverse all breakable adjacent edges (breakable chemical bonds) in the molecular structure corresponding to this node. Calculate the path activation energy of the new molecular configuration after the breakable edge is broken. The path activation energy is equal to the current cumulative activation energy plus the edge weight of the breakable adjacent edge. If the new path activation energy is less than the path activation energy of the node corresponding to the breakable adjacent edge, update the path activation energy of the node corresponding to the adjacent edge and add the node corresponding to the adjacent edge to the priority queue.
[0018] (Nodes initially represent molecular configurations and are uniquely identified by combinations of atomic hash values; edges represent chemical bonds, and their weights correspond to the activation energies of chemical bond breaking reactions.)
[0019] Product combinations with optimal energy or feasible kinetics are screened out by reaction activation energy.
[0020] The above-mentioned method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry uses molecular editing software to draw the Fisher projection structure of the starting reactants and clarify the spatial positions of each atom in the six-membered ring structure.
[0021] After constructing the initial conformation, molecular mechanics optimization is adopted: a certain molecular force field is selected, the energy gradient convergence criterion is set as the energy gradient threshold, and iterative calculation is performed until the energy distribution within the molecule is uniform, thus obtaining a stable optimized conformation.
[0022] The aforementioned method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry can employ a hierarchical modeling approach during the modeling process. The first-level decomposition breaks down the initial reactants into the smallest repeating units that constitute macromolecules. The second-level decomposition further subdivides the basic structural units into functional group structures and connecting bonds. Among these, the functional group structures contain various characteristic groups, while the connecting bonds are the chemical bonds that maintain the molecular structure. The third-level decomposition clarifies the atomic connection modes and bond parameters within the functional groups, including bond type, bond length, and bond energy.
[0023] The above-mentioned method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry includes node attributes such as element type, valence state, and charge distribution, and edge attributes such as bond type, relative bond energy, and reactivity level.
[0024] The above-mentioned method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry uses the following atom numbering rules: first, atoms are classified according to element type; then, atoms of the same element are numbered according to their connection order in the molecule using a specific search method; the atom number, element type, valence state, and charge distribution information are concatenated into a string; and a hash value that can uniquely identify the atom is generated using a specific hash algorithm for deduplication of reaction products.
[0025] The aforementioned method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry involves the following steps when determining whether chemical bonds are breakable: A dissociation energy threshold and a breakage probability threshold are set, and chemical bonds with dissociation energies less than the dissociation energy threshold are considered candidates. Based on the normalized breakage probabilities, chemical bonds meeting the breakage probability threshold are selected and marked. The marking process records the position, type, and corresponding breakage probability of the breakable bonds.
[0026] The aforementioned method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry employs a subgraph isomorphism algorithm for functional group identification. This algorithm matches the substructure molecular graphs in the molecular model with functional group pattern graphs in a predefined functional group structure pattern library to construct the library. This library contains graph structure representations of each functional group, with each functional group pattern defined as a labeled subgraph. Node labels represent atomic element types, and edge labels represent chemical bond types. Mapping functions are established: one from the substructure molecular graphs in the molecular model to the functional group pattern graphs in the library, and the other from the functional group pattern graphs in the library to the substructure molecular graphs in the molecular model.
[0027] The above-mentioned method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry obtains atomic coordinates and product coordinates from a molecular model. The coordinate data of atomic coordinates and product coordinates, along with the calculation settings, are input into quantum chemistry software, which then adjusts the positions of the atoms until the energy of the molecules is reduced to the minimum. The quantum chemistry software then outputs the energy values of the reactants in the initial reactants.
[0028] The transition state energy at which the molecular energy is highest during the process of adjusting the position of atoms using quantum chemistry software is obtained. The reaction activation energy is obtained by subtracting the reactant energy from the transition state energy, and the reaction heat is obtained by subtracting the transition state energy from the reactant energy.
[0029] Furthermore, during the process of adjusting atomic positions using quantum chemistry software, iterative optimization is performed until the molecular energy is reduced to the minimum, resulting in a stable molecular conformation; the iterative formula is: , where Δx is the coordinate update vector, H is the Hessian matrix, G is the gradient matrix, μ is the damping factor, and g is the energy gradient.
[0030] According to the above-described scheme, the beneficial effects of this invention are as follows: by analyzing the molecular structure of biomass starting reactants and abstracting it into a directed graph molecular model, potential transformation paths are marked after molecular mechanics optimization, unique identifiers are generated using atom numbering rules and hash algorithms, functional groups are extracted and encapsulated into subgraph modules, chemical bond parameters are obtained by combining quantum chemical calculations and converted into breaking probabilities, basic reaction building blocks are predefined to traverse and mark breakable bonds, fragments are recombined and deduplicated by hashing, and then the activation energy weights of edge chemical bond breaking reactions are assigned. The optimal energy path is selected by using a priority queue search algorithm, and finally, accurate analysis of chemical bond breaking paths and intermediate product transformations in the biomass pyrolysis reaction network is achieved, reaction paths and product distributions are predicted efficiently, and the computational efficiency and accuracy of complex systems are balanced.
[0031] 1. Capable of accurately analyzing chemical bond breaking paths and intermediate product transformations. By abstracting the structure of biomass molecules into a directed graph composed of atomic nodes and chemical bond edges, and combining parameters such as chemical bond dissociation energy obtained through quantum chemical calculations, the positions and breaking sequences of easily broken chemical bonds in biomass molecules can be clearly identified. At the same time, the evolution trajectory of free radical intermediates can be tracked, clearly grasping the complete transformation path from raw materials to intermediate products and then to final products in the pyrolysis reaction.
[0032] 2. Efficient prediction of reaction pathways and product distribution. By using a priority queue search algorithm to traverse all possible reaction pathways and combining the activation energy of the reaction pathways to screen for the energy-optimal conversion channels, the system can systematically predict the competition of thousands of parallel and series reactions during biomass pyrolysis. This solves the prediction bias problem caused by the neglect of molecular structural details in early simplified force field models, and significantly improves the accuracy of product distribution prediction.
[0033] 3. Balancing computational efficiency and accuracy in handling complex systems: By dividing the molecular model into functional group subgraph modules and performing hash deduplication, and combining a hierarchical computational strategy optimized by density functional theory and molecular mechanics, the computation time for full reaction network analysis of complex biomass systems is significantly shortened while ensuring computational accuracy. This effectively solves the problem of low computational efficiency in quantum chemistry and enables efficient simulation of complex molecular systems in biomass. Detailed Implementation
[0034] To make the technical problems, technical solutions, and beneficial effects of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0035] A method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry, including...
[0036] Analyze the molecular structure of the starting reactants in the biomass pyrolysis reaction, identify the atomic connection mode and chemical bond characteristics of the starting reactants, abstract atoms as nodes, chemical bonds as edges, and store and generate molecular models in a directed graph data structure;
[0037] The molecular model is optimized to ensure that the energy gradient of the molecule is less than a certain stable energy gradient threshold. Potential transformation paths are marked in the molecular model, and virtual edges and nodes with fixed attributes are set. Information for each atom is concatenated into a string using atom numbering rules, and a unique hash value is assigned to each atom using a hash algorithm. Common functional group structures in the molecular model are extracted and encapsulated as subgraph modules. Fundamental parameters of chemical bonds, including bond dissociation energies, are obtained through quantum chemical calculations. The rate constants of elementary reaction steps related to the breaking of specific chemical bonds are calculated, and after normalization, these rate constants are converted into probability weights for bond breaking in that specific reaction path. Basic reaction building blocks are predefined, containing bond change rules and matrix transformation methods. The molecular model is traversed; if the bond dissociation energy is less than the threshold, the chemical bond is considered a candidate for breaking. The probability weights for breaking are selected based on the breaking reaction corresponding to the chemical bond, and a probability sampling method is used to determine whether to perform breaking, marking the bond as breakable. The process involves breaking bonds by reassembling the fragmented molecular model using basic reaction blocks to obtain potential reaction products. The hash values of these products are calculated, and duplicates are removed before saving. Edges are assigned weights, representing the activation energy of the bond-breaking reaction. A distance array is defined, and a priority queue is constructed. The starting molecular state node and its initial distance are added to the priority queue. A loop iterates through the priority queue, selecting the node with the lowest cumulative activation energy (representing the current molecular configuration). All adjacent edges (connected chemical bonds) that can be broken are traversed within the molecular structure corresponding to this node. The path activation energy of the new molecular configuration after the broken edge is calculated. This path activation energy equals the current cumulative activation energy plus the edge weight of the adjacent edge. If the new path activation energy is less than the path activation energy of the node corresponding to the adjacent edge, the path activation energy of the node corresponding to the adjacent edge is updated, and the node corresponding to the adjacent edge is added to the priority queue. Finally, the optimal energy or kinetically feasible product combination is selected based on the reaction activation energy.
[0038] When analyzing the molecular structure of the starting reactants in biomass pyrolysis, specialized molecular modeling software such as Ghemical or Avogadro can be used. This software can visually display the spatial arrangement and connections of atoms in the molecule, allowing for the addition of atoms, the construction of chemical bonds, and the preliminary drawing of the molecular structure through a graphical interface. When identifying atomic connections and chemical bond characteristics, it is necessary to rely on basic chemical theory to clarify parameters such as the number of bonds, bond length, and bond angle of each atom. For example, carbon atoms typically form four covalent bonds, while oxygen atoms form two covalent bonds.
[0039] When abstracting atoms as nodes and chemical bonds as edges, the directional characteristics of reactions must be considered when using a directed graph data structure for storage. Nodes in the directed graph must contain basic atomic properties, such as element type (carbon, hydrogen, oxygen, etc.), valence state (determined by the number of bonds and atom type), and charge distribution (which can be initially obtained through quantum chemical calculations). The construction of edges must reflect the type of chemical bond (single, double, triple, etc.), the relative magnitude of bond energies (ranked through empirical data or preliminary calculations), and the level of reactivity (classified according to the chemical environment of the bond). In practical implementation, an adjacency list can be used to store the directed graph, with each node corresponding to a list recording its connections to other nodes and the attributes of the edges. This storage method facilitates subsequent traversal and modification of the molecular model.
[0040] In the modeling process, a hierarchical modeling method can be adopted. First-level decomposition breaks down the starting reactants into the smallest repeating units that constitute the macromolecule. Second-level decomposition further subdivides the basic structural units into functional group structures and connecting bonds. The functional group structures contain various characteristic groups, while the connecting bonds are the chemical bonds that maintain the molecular structure. Tertiary decomposition clarifies the atomic connection patterns and bond parameters within the functional groups, including bond type, bond length, and bond energy. Specifically, taking lignin and cellulose as starting reactants as examples, first-level decomposition breaks down lignin into phenylpropane structural units (such as guaiacolyl and syringyl), and cellulose into glucose units (connected by β-1,4 glycosidic bonds). Second-level decomposition further breaks down each structural unit into functional group structures (such as benzene rings, hydroxyl groups, methoxy groups, etc.) and connecting bonds (such as COC and CC bonds). Tertiary decomposition defines the atomic connection patterns and bond parameters within the functional groups, such as the conjugated π bonds of benzene rings and the OH bonds of hydroxyl groups.
[0041] Common molecular force fields such as CHARMM, AMBER, and UFF each have different applicable ranges and parameter characteristics. Taking the CHARMM force field as an example, it is optimized for biomolecules and describes the covalent and non-covalent interactions in the structures of proteins and polysaccharides with relatively high accuracy. It is especially suitable for biomass molecules containing six-membered ring structures, such as glucose.
[0042] Taking glucose molecule optimization as an example, when using the CHARMM force field, the bond stretching potential energy function is usually expressed in harmonic form: ,in, Let b be the bond force constant and b be the current bond length. To balance bond lengths. For C / C single bonds, Approximately 500-600 kcal / mol・Ų, equilibrium bond length 1.54 Å; CO single bond Approximately 700-800 kcal / mol・Ų It is 1.43 Å. The bond angle bending potential energy function takes a similar form: Balance value of CCC bond angle 112°, force constant It is approximately 100-150 kcal / mol・rad².
[0043] When labeling potential transformation pathways in molecular models, it is necessary to comprehensively consider the reactivity of chemical bonds, the position of functional groups, and common mechanisms of pyrolysis reactions. By analyzing parameters such as the dissociation energy and charge distribution of chemical bonds, bonds that are prone to breakage can be identified as the starting point of potential transformation pathways. For example, CO bonds near hydroxyl groups (-OH) usually have high reactivity due to the strong electronegativity of oxygen atoms and can be considered as candidates for potential transformation pathways.
[0044] Virtual edge nodes must be assigned fixed attributes, including reaction type (e.g., bond breaking, rearrangement, hydrogen transfer), activation energy (obtainable through preliminary quantum chemical calculations or empirical estimation), reaction probability (related to the bond breaking probability), and corresponding product structural characteristics. The role of virtual edge nodes is to connect starting reactants and potential products in the reaction network, constructing a complete reaction pathway. Their attributes must be consistent with the actual reaction mechanism; for example, virtual edge nodes for bond breaking reactions should clearly define the type of chemical bond being broken and the type of free radical formed after the breakage.
[0045] Atom numbering rules include: classifying atoms according to their element type, such as carbon, hydrogen, oxygen, etc.; for atoms of the same element, numbering them according to their connection order in the molecule, using depth-first search (DFS) or breadth-first search (BFS) to traverse the molecular structure and assign a unique number to each atom. When concatenating atomic information into a string and generating a hash value, key information such as the atom's number, element type, valence state, and charge distribution must be included. The information for each atom is concatenated into a string in the format "element-number-valence-charge". For example, the information for a carbon atom is "C-12-4-0.15", indicating that it is the 12th carbon atom with a valence of 4 and a charge of +0.15.
[0046] A hash algorithm is used to calculate the hash value of the concatenated string, which uniquely identifies each atom, avoiding collisions caused by similar atom information. The hash value plays an important role in the subsequent deduplication of reaction products; by comparing the hash values of atoms in the products, it is possible to quickly determine whether a product is duplicated.
[0047] When extracting common functional group structures from molecular models, a functional group structure pattern library must first be established, including the structural features of common functional groups such as hydroxyl (-OH), carboxyl (-COOH), aldehyde (-CHO), and ester (-COO-). Then, functional groups are identified based on the data in the functional group structure pattern library. Functional group identification can employ a subgraph isomorphism algorithm, matching substructures in the molecular model with functional group structures in the pattern library to determine the types and positions of functional groups contained in the molecule.
[0048] In this embodiment, when identifying hydroxyl groups in a glucose molecule, the algorithm first traverses each oxygen atom node in the molecular graph, checking whether it is connected to a hydrogen atom and a carbon atom, and whether all chemical bonds are single bonds. For each candidate oxygen atom, the algorithm constructs a subgraph centered on that oxygen atom and performs isomorphic matching with the hydroxyl structure pattern diagram. During the matching process, not only is the consistency of atom type and bond type verified, but also the valence state of the atoms is checked to see if it meets the requirements (e.g., the valence of the oxygen atom is 2, and the valence of the hydrogen atom is 1).
[0049] When encapsulating functional group structures into subgraph modules, each subgraph module should include the atomic composition of the functional group, the chemical bonding mode, characteristic properties (such as reactivity and common reaction types), and connection points with other molecular structures. For example, the hydroxyl subgraph module should clearly indicate the connection mode between oxygen and hydrogen atoms, and the connection point between oxygen and carbon atoms. It should also record the reactivity level of the hydroxyl group (the reactivity level is determined by manually setting criteria, often through importing laboratory data and manually classifying the levels). Common reactions include dehydration or substitution reactions. The encapsulated subgraph modules can serve as independent structural units, enabling rapid identification and application in subsequent reaction pathway analysis, thus improving the efficiency of reaction network construction.
[0050] When obtaining fundamental parameters of chemical bonds through quantum chemical calculations, density functional theory (DFT) methods can be used. Specific functionals and basis sets are selected, and calculations are performed in quantum chemical software. Specifically, the optimized molecular model structure is converted into an input format recognizable by the quantum chemical software. The calculation tasks are set as geometry optimization and single-point energy calculation. Optimization parameters include convergence criteria (setting energy gradient and displacement values below different thresholds; in this embodiment, the energy gradient is less than 0.00045 hartree / Å, and the maximum displacement is less than 0.001 Å) and the maximum number of iterations. When calculating the chemical bond dissociation energy, the target chemical bond must first be broken to obtain two fragment molecules. Structural optimization and single-point energy calculations are then performed on the reactant molecule and the fragment molecule, respectively. The dissociation energy is the difference between the total energy of the fragment molecule and the energy of the reactant molecule. The calculation of the reaction rate constant must consider the temperature factor. According to transition state theory, the rate constant is related to the activation energy and temperature; the higher the temperature and the lower the activation energy, the larger the rate constant. During normalization, the reaction rate constants of all chemical bonds are divided by the maximum rate constant to obtain the breaking probability between 0 and 1, making the rate constants of different orders of magnitude comparable and facilitating the subsequent labeling of breakable bonds.
[0051] When predefining basic reaction blocks, they should be categorized based on common mechanisms of pyrolysis reactions, mainly including bond breaking reactions, rearrangement reactions, hydrogen transfer reactions, and addition reactions. Each basic reaction block should include the following: bond change rules, specifying which chemical bonds break and which new chemical bonds form during the reaction; and matrix transformation methods, describing the changes in the adjacency matrix of the molecular model before and after the reaction to accurately generate reaction products. Taking the bond breaking reaction block as an example, it is necessary to specify the type of chemical bond broken (e.g., C-C bond, CO bond), the type of free radical formed after breakage (e.g., methyl free radical, hydroxyl free radical), and the possible subsequent reaction pathways of the free radical. The matrix transformation method is manifested by deleting the corresponding bond connections in the adjacency matrix and adding new free radical nodes or adjusting the valence state of atoms.
[0052] When traversing the molecular model to label breakable bonds, a specific search method can be used to examine each chemical bond in the molecule sequentially. To determine whether a chemical bond is breakable, two parameters must be considered: the bond dissociation energy and the breakage probability. A dissociation energy threshold is set, and bonds meeting the threshold are considered candidates. Then, the normalized breakage probability is used to select bonds that meet the probability requirements for labeling. During labeling, the position, type, and corresponding breakage probability of the breakable bonds must be recorded for subsequent use of basic reaction blocks. When calling the blocks, the broken molecular model fragments are reassembled according to bond change rules to generate potential reaction products. The hash values of the products are calculated for deduplication to ensure the uniqueness of the stored product structures.
[0053] When assigning weights (reaction activation energies) to the edges, the transition state energies need to be calculated using quantum chemistry software. The specific process is as follows: First, obtain the atomic coordinates of the initial reactants from the molecular model and input them into the quantum chemistry software for transition state search, which can be done using the synchronous transfer-secondary perturbation method (STQN) or the flexible potential surface scan method (FES). Then, optimize the transition state structure until the energy gradient converges and obtain the transition state energy. Finally, subtract the reactant energy from the transition state energy to obtain the reaction activation energy, and subtract the transition state energy from the reactant energy to obtain the heat of reaction.
[0054] When constructing the priority queue, a min-heap data structure is used, with activation energy as the priority criterion. Initially, the starting node (starting reactant) and the initial distance (activation energy of 0) are added to the priority queue. During the loop, the node with the lowest activation energy is taken from the queue each time, and all its adjacent edges (i.e., possible reaction paths) are traversed. The path activation energy of the node (product) corresponding to each adjacent edge is calculated (current node activation energy plus edge weight). If the new path activation energy is less than the activation energy already recorded for the adjacent node, the activation energy of that node is updated, and it is added back to the priority queue. In this way, it is ensured that reaction paths with lower activation energies are explored first, ultimately selecting the product combination with optimal energy or kinetic feasibility.
[0055] When obtaining atomic and product coordinates and inputting them into quantum chemistry software, the coordinate data must be converted to a format supported by the software. The coordinate data should include the element symbols and three-dimensional coordinates of all atoms; the product coordinates are the atomic coordinates of the molecules after the reaction. When inputting the calculation settings, the optimization method, basis set, functional, and convergence criterion must be clearly specified. This provides an accurate initial reactant configuration for activation energy calculations, ensuring that the molecules are in a stable state with the lowest energy (local minimum), thus eliminating the influence of molecular strain on energy calculations.
[0056] In the process of adjusting atomic positions, quantum chemistry software continuously calculates the molecular energy and gradient, iteratively optimizing until the molecular energy is reduced to the minimum, resulting in a stable molecular conformation. The iterative formula is: Where Δx is the coordinate update vector, H is the Hessian matrix, G is the gradient matrix, μ is the damping factor, and g is the energy gradient. In the glucose reaction of this embodiment, the initial damping factor is set to 0.1, and it automatically increases to 0.5 when energy fluctuates during iteration to ensure convergence stability.
[0057] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry, characterized in that, include Analyze the molecular structure of the starting reactants in the biomass pyrolysis reaction, identify the atomic connection mode and chemical bond characteristics of the starting reactants, abstract atoms as nodes, chemical bonds as edges, and store and generate molecular models in a directed graph data structure; The molecular model is optimized so that the energy gradient of the molecule is less than a certain stable energy gradient threshold. In the molecular model, potential transformation pathways are marked and virtual edges and virtual nodes with fixed attributes are set; The information of each atom is concatenated into a string according to the atom numbering rules, and a unique hash value is assigned to each atom using a hash algorithm. Extract common functional group structures from molecular models and encapsulate them into subgraph modules; The fundamental parameters of chemical bonds, including the chemical bond dissociation energy, are obtained through quantum chemical calculations. The rate constants of the elementary reaction steps associated with the breaking of a specific chemical bond are calculated and then normalized to convert the rate constants of the elementary reaction steps into the probability weights of the breaking of the chemical bond in a specific reaction path. Predefined basic reaction blocks, which include bond change rules and matrix transformation methods; Traverse the molecular model. If the chemical bond dissociation energy is less than the threshold, the chemical bond is determined to be a candidate for breaking. Select the probability weight of breaking according to the breaking reaction corresponding to the chemical bond. Determine whether to perform breaking by sampling according to probability and mark it as a breakable bond. Call the basic reaction building block to reassemble the broken molecular model fragments to obtain potential reaction products. Calculate the hash value of the potential reaction products and save them after deduplication based on the hash value. Assign weights to the edges, where the weights represent the activation energy of the chemical bond breaking reaction. Define a distance array and construct a priority queue. Add the starting molecular state node and the initial distance to the priority queue. Loop: Select the node with the smallest cumulative activation energy from the priority queue. Traverse all breakable adjacent edges in the molecular structure corresponding to the node. Calculate the path activation energy of the new molecular configuration after the breakable edge is broken. The path activation energy is equal to the current cumulative activation energy plus the edge weight of the breakable adjacent edge. If the new path activation energy is less than the path activation energy of the node corresponding to the breakable adjacent edge, update the path activation energy of the node corresponding to the adjacent edge and add the node corresponding to the adjacent edge to the priority queue. Product combinations with optimal energy or feasible kinetics are screened out by reaction activation energy.
2. The method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry as described in claim 1, characterized in that, The Fisher projection structure of the starting reactants was drawn using molecular editing software to clarify the spatial positions of each atom in the six-membered ring structure. After constructing the initial conformation, molecular mechanics optimization is adopted: a certain molecular force field is selected, the energy gradient convergence criterion is set as the energy gradient threshold, and iterative calculation is performed until the energy distribution within the molecule is uniform, thus obtaining a stable optimized conformation.
3. The method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry as described in claim 1, characterized in that, In the modeling process, a hierarchical modeling method can be adopted. The first-level decomposition breaks down the starting reactants into the smallest repeating units that make up the macromolecule. The second-level decomposition further subdivides the basic structural units into functional group structures and connecting bonds. Among them, the functional group structures contain various characteristic groups, while the connecting bonds are the chemical bonds that maintain the molecular structure. The third-level decomposition clarifies the atomic connection mode and bond parameters within the functional group, including bond type, bond length and bond energy.
4. The method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry as described in claim 1, characterized in that, The attributes of a node include element type, valence state, and charge distribution, while the attributes of an edge include bond type, relative bond energy, and reactivity level.
5. The method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry as described in claim 1, characterized in that, Atom numbering rules: First, atoms are classified according to element type. Then, atoms of the same element are numbered according to their connection order in the molecule using a specific search method. The atom number, element type, valence state, and charge distribution information are concatenated into a string. A specific hash algorithm is used to generate a hash value that can uniquely identify the atom, which is used for deduplication of reaction products.
6. The method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry as described in claim 1, characterized in that, When determining whether a chemical bond is breakable: a dissociation energy threshold and a breakage probability threshold are set, and chemical bonds with dissociation energies less than the dissociation energy threshold are considered as candidates; based on the normalized breakage probability, chemical bonds that meet the breakage probability threshold are selected for marking; the marking process records the position, type, and corresponding breakage probability of the breakable bond.
7. The method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry as described in claim 1, characterized in that, Functional group identification employs a subgraph isomorphism algorithm, which matches the substructure molecular graphs in the molecular model with functional group pattern graphs in a predefined functional group structure pattern library to construct the functional group structure pattern library. This library contains graph structure representations of each functional group, with each functional group pattern defined as a labeled subgraph. Node labels represent atomic element types, and edge labels represent chemical bond types. Mapping functions are established: one from the substructure molecular graphs in the molecular model to the functional group pattern graphs in the library, and the other from the functional group pattern graphs in the library to the substructure molecular graphs in the molecular model.
8. The method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry as described in claim 1, characterized in that, Atomic and product coordinates are obtained from the molecular model. The coordinate data of atomic and product coordinates, along with the calculation settings, are input into the quantum chemistry software. The quantum chemistry software then adjusts the positions of the atoms until the energy of the molecule is reduced to the minimum. The quantum chemistry software outputs the energy values of the reactants from the initial reactants. The transition state energy at which the molecular energy is highest during the process of adjusting the position of atoms using quantum chemistry software is obtained. The reaction activation energy is obtained by subtracting the reactant energy from the transition state energy, and the reaction heat is obtained by subtracting the transition state energy from the reactant energy.
9. The method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry as described in claim 8, characterized in that, In the process of adjusting atomic positions using quantum chemistry software, iterative optimization is performed until the molecular energy is reduced to the minimum, resulting in a stable molecular conformation; the iterative formula is: , where Δx is the coordinate update vector, H is the Hessian matrix, G is the gradient matrix, μ is the damping factor, and g is the energy gradient.
Citation Information
Patent Citations
Biosynthesis path prediction method based on machine learning and user platform
CN117409872A
Drug response prediction method based on gene relation network and drug substructure
CN120148905A