Tar absorbent performance prediction and high-throughput screening method, system, equipment and medium

By constructing a structure library and deep learning models, combined with molecular dynamics simulations and experimental data, the problem of screening tar absorption aids was solved, achieving efficient and stable tar absorption effects and reducing experimental costs.

CN122091008APending Publication Date: 2026-05-26SOUTH CHINA UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTH CHINA UNIV OF TECH
Filing Date
2026-02-04
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing technologies, the composition of tar is complex and variable, the washing effect of crude coal gas is not good, and there is insufficient research on existing tar absorption aids, making it difficult to efficiently screen for high-efficiency absorption aids.

Method used

By constructing a library of absorption aids and tar component structures, and using molecular dynamics simulations and deep learning models for transfer learning, highly efficient absorption aids are screened out. The model is then fine-tuned using experimental data to achieve high-throughput screening.

Benefits of technology

It can quickly screen out highly efficient additives, reduce experimental costs, and the model takes into account both theoretical laws and real physical processes, with good predictive stability. It is also applicable to the screening of other high-throughput materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122091008A_ABST
    Figure CN122091008A_ABST
Patent Text Reader

Abstract

The invention discloses a tar absorbent performance prediction and high-throughput screening method, system and equipment and a medium. The method comprises the following steps: constructing an absorption assistant structure library and a tar component structure library; respectively sampling the two structure libraries, establishing a water phase-absorption aid-oil phase model, and carrying out sufficient molecular dynamics simulation until a data set with a specified scale is formed; pre-training the deep learning network model; selecting an actual tar component, determining a tar component descriptor, and screening out an absorption assistant subset by using a pre-trained deep learning network model; the method comprises the following steps: determining absorption aids of a plurality of representative structures in an absorption aid subset, determining absorption rates through experiments, performing fine tuning training on a model to obtain a high-throughput screening model, and predicting all structures in an absorption aid structure library to obtain the absorption aid with the highest absorption rate. The method has the advantages of high-precision prediction, low experiment cost, high expandability and the like, and is suitable for rapid research, development and screening of the absorption additive material.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method, system, equipment, and medium for predicting the performance of tar absorbents and conducting high-throughput screening, belonging to the field of chemical separation technology. Background Technology

[0002] Currently, pressurized coal-water slurry gasification in the coal chemical industry typically employs a quench process. The crude coal gas generally undergoes a five-stage scrubbing process for humidification and dust removal: quench chamber water bath, gasifier outlet spray, Venturi scrubbing, scrubbing tower water bath, and scrubbing tower tray scrubbing. However, this gas scrubbing process often suffers from poor scrubbing efficiency and easy ash accumulation in the scrubbing tower, affecting the long-term stable operation of the gasification system. To ensure scrubbing effectiveness and long-term stable operation of the equipment, researchers have achieved certain results by optimizing the scrubbing process, designing new devices, and adding scrubbing aids.

[0003] Chinese utility model patents with patent application publication numbers CN207159165U and CN207391363U address problems such as severe blockage of the washing water discharge pipeline, pipeline wear, and leakage of the quench water pump mechanical seal. They propose optimized design schemes for the washing tower, improving gas purification efficiency by reducing the dust and other ineffective component content in the crude gas. Guo Panchun et al. (Process Optimization of Crude Gas Scrubbing Tower in Coal Gasification Technology) optimized the crude gas scrubbing tower process by installing a venturi inside the tower and technically modifying the venturi and distributor. This effectively solved the problem of easy blockage of conventional venturi nozzles, throats, and distributors, resulting in a simple and reasonable structure with low investment, achieving cost reduction and efficiency improvement. However, such optimizations all increase the complexity of the equipment, raise equipment costs, and do not significantly improve the tar washing effect.

[0004] Chinese patent application CN103666584A discloses a tar removal device and method capable of long-term operation. This method uses alcohol, mineral oil, or tar complexes as absorbents to ensure a low tar content in the fuel gas entering the primary tar adsorption unit. However, the device is a multi-stage unit, making it relatively complex, and alcohols and mineral oils cannot be directly added to the aqueous phase; otherwise, emulsification would be very significant.

[0005] Studies by Zhou Peng et al. on the trial use of novel ash water dispersants in the GE coal-water slurry pressurized gasification process showed that the addition of dispersants to the ash water system of the coal-water slurry gasification unit significantly reduced the scaling rate of the quench water pipeline. However, the addition of dispersants significantly inhibited the subsequent coagulation and sedimentation treatment of ash water. Compared with dispersants, the core function of absorbents is to directionally optimize gas-liquid mass transfer efficiency, directly improving the elution effect of target pollutants (such as soluble organic matter, suspended ash, tar, etc.) in crude coal gas, rather than being limited to dispersing particles.

[0006] Zhang Ni et al.'s CFD simulation exploration of tar absorption using imidazole ionic liquids ([Bmim][Tf2N]) employed an absorption tower CFD model to simulate the absorption of tar by ionic liquids as absorbents, investigating the effects of liquid-to-gas ratio and tower structural parameters on the gas-liquid distribution and mass transfer characteristics within the tower. However, ionic liquids are expensive and have high regeneration energy consumption. Furthermore, the absorption performance of ionic liquids is highly dependent on their molecular structure, making selective design for specific gases difficult. Moreover, their overall absorption effect on complex gas mixtures (containing various hydrocarbons, sulfides, dust, etc. in crude coal gas) is often inferior to that of composite solvents, and they are easily affected by impurities. Feng Haijun et al.'s "Estimation of solubility of acid gases in ionic liquids using different machine learning methods" employed 12 machine learning algorithms to predict the solubility of acid gases in ionic liquids, providing some assistance for the targeted development of ionic liquids. However, the dataset was collected from literature, and current research on tar absorption aids is still limited; the literature data is currently insufficient to support the training of machine learning models.

[0007] Molecular dynamics simulations can track the interaction between tar molecules (such as polycyclic aromatic hydrocarbons) and absorbents at the atomic scale, visually presenting their adsorption sites, binding energies, and interfacial diffusion behavior. This accurately reveals the selective adsorption mechanism of absorbents on complex tar components, providing microscopic theoretical support for the design of highly efficient absorbents. Xu Yunfei et al. investigated the microscopic mechanism of asphaltenes and gums aggregation on oil-water interface stability using molecular dynamics simulations to alter the number of asphaltenes and gums molecules. The results provide a basis for explaining the oil-water emulsion stabilization mechanism and developing demulsification technology. Furthermore, large-scale simulations can generate datasets for training machine learning models. However, the accuracy of dynamic simulations directly depends on the force field parameters. Existing force fields often only provide qualitative analysis of absorbed tar, with some bias in quantitative analysis.

[0008] In view of the complex and variable composition of tar and the poor washing effect of crude coal gas, as well as the fact that the existing technology does not involve high-throughput screening of efficient tar absorption aids, it is necessary to develop a method for screening efficient absorption aids. Summary of the Invention

[0009] In view of this, the present invention provides a method, system, computer equipment, and storage medium for predicting the performance of tar absorbents and conducting high-throughput screening. It addresses the problem of insufficient datasets in deep learning models through molecular dynamics simulations, and uses a small amount of experimental data to perform transfer learning on the pre-trained deep learning model to correct accuracy issues, thereby achieving efficient screening of absorbent aids. Adding a small amount of the screened absorbent aid during the washing of crude coal gas can significantly improve the absorption effect, and it also has guiding significance for other high-throughput screening materials.

[0010] The first objective of this invention is to provide a method for predicting the performance of tar absorbents and for high-throughput screening.

[0011] The second objective of this invention is to provide a tar absorbent performance prediction and high-throughput screening system.

[0012] A third objective of this invention is to provide a computer device.

[0013] A fourth objective of this invention is to provide a computer-readable storage medium.

[0014] The first objective of this invention can be achieved by adopting the following technical solution:

[0015] A method for predicting the performance of tar absorbents and for high-throughput screening, the method comprising:

[0016] Construct a library of absorption aid structures and convert each structure or combination of structures into a molecular fingerprint;

[0017] Construct a tar component structure library and convert each structure or combination of structures into a molecular fingerprint;

[0018] Sampling was performed on the absorption aid structure library and the tar component structure library respectively, and an aqueous phase-absorption aid-oil phase model was established. Sufficient molecular dynamics simulations were conducted until a dataset of a specified size was formed.

[0019] Based on the dataset, a deep learning network model is pre-trained, where the molecular fingerprints of the absorption aids and tar components are used as structural descriptors and the absorption rate is used as the output.

[0020] Select the actual tar components, determine the tar component descriptors, and use a pre-trained deep learning network model to screen out a subset of absorption aids whose absorption rate of the selected tar components reaches a preset percentage.

[0021] Clustering algorithms were used to identify multiple representative structures of absorption aids within a subset of absorption aids, and the absorption rate was determined experimentally.

[0022] Based on the experimentally determined absorption rate, the pre-trained deep learning network model was fine-tuned to obtain a high-throughput screening model.

[0023] A high-throughput screening model was used to predict all structures in the absorption aid structure library to obtain the absorption aid with the highest absorption rate.

[0024] Furthermore, the molecular dynamics simulations include:

[0025] The simulation system was subjected to energy minimization to eliminate high-energy conformations. The system was then subjected to kinetic relaxation in the NVT ensemble for 500–1500 ps until it reached a stable state. The equilibrium stability criteria were that the system temperature and energy fluctuations were within 3%–10%. The system was then placed in the NPT ensemble for relaxation for 500–1500 ps to ensure that the densities of each component reached reasonable values ​​until the system stabilized. The convergence criteria were that the system temperature, energy, density, and box size fluctuations were within 3%–10%. Production data was then collected for 1000–2000 ps under the NPT ensemble, and the number of tar molecules in the aqueous phase was counted. The absorption rate was calculated, which was equal to the number of tar molecules in the aqueous phase divided by the total number of tar molecules.

[0026] Furthermore, the simulation system is as follows: periodic boundary conditions are used in each direction, the cutoff radius is 12.5 Å, the long-range Coulomb electrostatic interaction between particles is calculated using the Ewald algorithm, the van der Waals interaction between particles is calculated using the Atom-based algorithm, the temperature of the dynamic relaxation process is set to 298 K, the simulation pressure is 0.1 MPa, and the simulation step size is 1 fs.

[0027] Furthermore, the combined structure in the absorption aid structure library refers to a composite hydrocarbon substance composed of two or more hydrocarbon molecules in a preset ratio, and the molecular fingerprint of the composite hydrocarbon substance is obtained by weighted average of the molecular fingerprint data of pure hydrocarbons.

[0028] The combined structures in the tar component structure library refer to composite hydrocarbon substances composed of two or more hydrocarbon molecules in a preset ratio. The molecular fingerprint of the composite hydrocarbon substance is obtained by weighted averaging of the molecular fingerprint data of pure hydrocarbons.

[0029] Furthermore, the sampling method used for sampling the absorption aid structure library and the tar component structure library is Latin hypercube sampling.

[0030] Furthermore, the deep learning model employs a dual-branch convolutional neural network to extract features from the molecular fingerprints of the absorption aid and tar components, respectively. Each branch includes a combination of Conv1D convolutional layers and MaxPooling1D pooling layers. After extracting local patterns, the features are expanded into a one-dimensional vector through a Flatten layer. The two features are fused in a Concatenate layer, then combined and nonlinearly mapped through a fully connected layer, and finally the output layer regresses to predict the target absorption rate.

[0031] Furthermore, the clustering algorithm is the K-means clustering algorithm.

[0032] The second objective of this invention can be achieved by adopting the following technical solution:

[0033] A tar absorbent performance prediction and high-throughput screening system, the system comprising:

[0034] The first building module is used to build an absorption aid structure library, converting each structure or combination of structures into a molecular fingerprint;

[0035] The second building module is used to build a tar component structure library, converting each structure or combination of structures into a molecular fingerprint.

[0036] The simulation module is used to sample the absorption aid structure library and the tar component structure library respectively, establish an aqueous phase-absorption aid-oil phase model, and perform sufficient molecular dynamics simulations until a dataset of a specified size is formed.

[0037] The pre-training module is used to pre-train the deep learning network model based on the dataset, wherein the molecular fingerprints of the absorption aids and tar components are used as structural descriptors and the absorption rate is used as the output.

[0038] The screening module is used to select the actual tar components, determine the tar component descriptors, and use a pre-trained deep learning network model to screen out a subset of absorption aids whose absorption rate of the selected tar components reaches a preset percentage.

[0039] The measurement module is used to identify multiple representative structures of absorption aids in the subset of absorption aids using clustering algorithms, and then experimentally measure the absorption rate.

[0040] The fine-tuning training module is used to fine-tune the pre-trained deep learning network model based on the experimentally determined absorption rate, so as to obtain a high-throughput screening model.

[0041] The prediction module is used to predict all structures in the absorption aid structure library using a high-throughput screening model to obtain the absorption aid with the highest absorption rate.

[0042] The third objective of this invention can be achieved by adopting the following technical solution:

[0043] A computer device includes a processor and a memory for storing a processor-executable program, wherein when the processor executes the program stored in the memory, it implements the above-described method for predicting the performance of tar absorbents and high-throughput screening.

[0044] The fourth objective of this invention can be achieved by adopting the following technical solution:

[0045] A computer-readable storage medium storing a program that, when executed by a processor, implements the above-described method for predicting and screening the performance of tar absorbents.

[0046] The present invention has the following advantages over the prior art:

[0047] This invention enables rapid screening of highly efficient additives for improving the washing effect of crude coal gas, fully utilizes simulation data to reduce experimental costs, solves the problem of insufficient experimental data through molecular dynamics simulation, and utilizes transfer learning to adjust the model from the "simulation domain" to the "real experimental domain" with only a small amount of experimental data, thus solving the problem of difficult-to-obtain experimental data. The model trained on simulation data covers a wide parameter space, and the model after transfer learning can take into account both theoretical laws and real physical processes, resulting in better prediction stability. Furthermore, it has guiding significance for the screening of other high-throughput materials. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0049] Figure 1 This is a simplified flowchart of the tar absorbent performance prediction and high-throughput screening method of Embodiment 1 of the present invention.

[0050] Figure 2 This is a detailed flowchart of the tar absorbent performance prediction and high-throughput screening method of Embodiment 1 of the present invention.

[0051] Figure 3 This is a schematic diagram of the initial setup of the dynamic simulation model in Embodiment 1 of the present invention.

[0052] Figure 4 This is a schematic diagram of the deep learning model structure in Embodiment 1 of the present invention.

[0053] Figure 5 This is a block diagram of the tar absorbent performance prediction and high-throughput screening system of Embodiment 5 of the present invention.

[0054] Figure 6 This is a structural block diagram of a computer device according to Embodiment 6 of the present invention. Detailed Implementation

[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0056] Example 1:

[0057] like Figure 1 and Figure 2 As shown in the figure, this embodiment provides a method for predicting the performance of tar absorbents and for high-throughput screening. The method includes the following steps:

[0058] S201. Construct a library of absorption aid structures and convert each structure or combination of structures into a molecular fingerprint.

[0059] In this embodiment, the absorption aid structure library covers isomeric tridecyl alcohol polyoxyethylene ether, aromatic polyamino acid block copolymers, polyethylene oxide-propylene oxide polyethylene oxide-propylene oxide block copolymers, etc.; the combined structure refers to a composite absorption aid composed of two or more absorption aids in a preset ratio, and the molecular fingerprint of the composite absorption aid is obtained by weighted average of the molecular fingerprint data of the pure absorption aids.

[0060] S202. Construct a tar component structure library and convert each structure or combination of structures into a molecular fingerprint.

[0061] In this embodiment, the tar component structure library covers polycyclic aromatic hydrocarbons, monocyclic aromatic hydrocarbons, phenolic compounds, etc.; the combined structure refers to a composite hydrocarbon substance composed of two or more hydrocarbon molecules in a preset ratio, and the molecular fingerprint of the composite hydrocarbon substance is obtained by weighted averaging of the molecular fingerprint data of pure hydrocarbons.

[0062] The molecular fingerprints in steps S201 and S202 above are count-type fingerprints that can reflect the number of times a feature appears.

[0063] S203. Sample the structure libraries of absorbent aids and tar components respectively, establish an aqueous phase-absorbent aid-oil phase model, and perform sufficient molecular dynamics simulations until a dataset of a specified size is formed.

[0064] In this embodiment, the sampling method is Latin hypercube sampling, and a single sample consists of an absorbent aid structure and a tar component structure. The simulation system is as follows: periodic boundary conditions are used in each direction, the cutoff radius is 12.5 Å, the calculation of long-range Coulomb electrostatic interactions between particles uses the Ewald algorithm, the calculation of van der Waals interactions between particles uses the Atom-based algorithm, the temperature of the kinetic relaxation process is set to 298 K, the simulation pressure is 0.1 MPa, and the simulation step size is 1 fs.

[0065] like Figure 3 As shown, the specific process of molecular dynamics simulation in this embodiment is as follows: The simulation system is minimized to eliminate high-energy conformations; the simulation system is subjected to kinetic relaxation in the NVT ensemble for 500-1500 ps until the system reaches a stable state, with the standard for equilibrium stability being a temperature and energy fluctuation range of 3%-10%; the system is then placed in the NPT ensemble for relaxation for 500-1500 ps to ensure that the densities of each component reach reasonable values ​​until the system stabilizes, with the convergence standard being a temperature, energy, density, and box size fluctuation range of 3%-10%; data is generated by running the NPT ensemble for 1000-2000 ps, ​​the number of tar molecules in the aqueous phase is counted, and the absorption rate is calculated, which is equal to the number of tar molecules in the aqueous phase divided by the total number of tar molecules; this process is repeated until a dataset of a specified size is formed, where the specified size is more than 1000 samples.

[0066] S204. Based on the dataset, pre-train the deep learning network model.

[0067] In this embodiment, the molecular fingerprints of the absorption aid and tar components are used as structural descriptors, the absorption rate is used as the output, and the structure of the deep learning network model is as follows: Figure 3 As shown, it employs a dual-branch convolutional neural network to extract features from the molecular fingerprints of the absorption aid and tar components, respectively. Each branch includes a combination of Conv1D convolutional layers and MaxPooling1D pooling layers. After extracting local patterns, the features are expanded into a one-dimensional vector through a Flatten layer. The two features are fused in a Concatenate layer, then combined and nonlinearly mapped through a fully connected layer. Finally, the output layer regresses and predicts the target absorption rate.

[0068] S205. Select the actual tar components, determine the tar component descriptors, and use a pre-trained deep learning network model to screen out a subset of absorption aids whose absorption rate of the selected tar components reaches a preset percentage.

[0069] In this embodiment, the preset percentage is 80%. After pre-training the deep learning network model, the pre-trained deep learning network model is used to screen out a subset of absorption aids with an absorption rate of 80% for the selected tar components.

[0070] S206. Use clustering algorithms to identify multiple representative structures of absorption aids in the subset of absorption aids, and determine the absorption rate through experiments.

[0071] In this embodiment, the clustering algorithm used is the K-means clustering algorithm. The multiple representative structures are the structures represented by the centroid of each cluster plus 2 to 10 random structures within that cluster, so that the total number of structures is between 20 and 100.

[0072] S207. Based on the experimentally determined absorption rate, fine-tune the pre-trained deep learning network model to obtain a high-throughput screening model.

[0073] In this embodiment, during training, the weight parameters of all neurons except the last layer of neurons are frozen.

[0074] S208. Use a high-throughput screening model to predict all structures in the absorption aid structure library to obtain the absorption aid with the highest absorption rate.

[0075] It should be noted that although the method operations of the above embodiments are described in a specific order, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the described steps may be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0076] Example 2:

[0077] Following the process of Example 1 above, this example extracts 1000 samples from the absorbent structure library and the tar molecular structure library using the Latin hypercube sampling method. Molecular dynamics simulations are then performed on these 1000 samples to construct a simulated dataset. The dataset is divided into a training set and a validation set in an 8:2 ratio. The deep learning model structure employs a dual-branch convolutional neural network to extract features from the ECFPs molecular fingerprints of absorbents and tar components, respectively. In the deep learning model, each branch includes a combination of a Conv1D convolutional layer (the first convolutional layer has 32 kernels and a kernel size of 50; the second convolutional layer has 64 kernels and a kernel size of 10) and a MaxPooling1D pooling layer (pooling window size of 2). After extracting local patterns, the features are expanded into a one-dimensional vector using a Flatten layer. The two features are concatenated... The layers are fused, then feature combination and nonlinear mapping are performed through a fully connected layer (128 neurons). Dropout (scale 0.3) is used to reduce the risk of overfitting. Finally, the output layer regresses and predicts the target uptake rate, where the activation function is ReLU, the optimizer is Adam, and the loss function is mean squared error (MSE). After training, a pre-trained deep learning model A1 is obtained, and the regression coefficients R on the training set are... 2 The regression coefficient R0 for the test set is 0.98. 2 It is 0.93.

[0078] In this embodiment, the simulated tar component was selected as a mixture of naphthalene and anthracene in a 1:1 ratio. The molecular fingerprint data of the mixture was obtained by weighted averaging of the ECFPs molecular fingerprints of both components. A pre-trained deep learning model A1 was used to predict the molecular fingerprints of all structures in the absorption structure library. Structures with an absorption rate greater than 80% were then selected to form a subset of high-efficiency absorption aids. This subset was clustered using the K-Means algorithm, and 100 absorption aids were extracted from all clusters.

[0079] The isomeric tridecyl alcohol polyoxyethylene ether series materials in this embodiment are commercial materials.

[0080] The synthesis process of the aromatic polyamino acid block copolymer in this embodiment is as follows: First, the corresponding amino acid-N-carboxy ketal monomer (such as BLG-NCA, Phe-NCA) and initiator are dissolved in anhydrous DMF. Under a nitrogen atmosphere, the initiator solution is transferred to a container containing an amino acid-NCA monomer solution to start the polymerization reaction. The reaction temperature is controlled and the reaction time is set according to the specific polymer type (1-2 days; for block copolymers such as PBLG-b-PPA, the first block such as PBLG needs to be synthesized first, and after reacting for 1 day, the DMF solution of the second amino acid-NCA monomer such as Phe-NCA is added through a funnel, and the reaction continues for a specific time, such as 2 days). After the reaction is completed, excess anhydrous diethyl ether is added to the system to precipitate the polymer. The crude product is redissolved in DMF and precipitated again for purification. Finally, the purified product is vacuum dried to obtain the target polymer.

[0081] The synthesis process of the poly(ethylene oxide)-propylene oxide block copolymer in this embodiment is as follows: propoxylation and ethoxylation reactions are carried out in a high-pressure stainless steel reactor. First, propoxylation of N,N-dimethylethanolamine (DMEA) is performed: DMEA and KOH are added to the reactor, a vacuum is drawn, and the temperature is gradually increased. PO is gradually introduced through a pressure pipette connected to a nitrogen cylinder, and nitrogen is used to pressurize the PO into the reactor while stirring. After the PO has completely reacted and the pressure has dropped to a minimum, heating is stopped, and the contents are cooled to room temperature by circulating cold water through the cooling coils in the reactor, yielding a propoxylated DMEA intermediate. Subsequently, ethoxylation and subsequent reactions are performed: propoxylated DMEA and KOH are added to the reactor, and EO is introduced for ethoxylation; then, PO and EO are added sequentially in a similar manner to ensure complete reaction, finally yielding the target product.

[0082] The experimental procedure for determining the absorption rate in this embodiment is as follows: 0.1 g of the absorption aid was accurately weighed and added to 100 mL of deionized water. The mixture was stirred until homogeneous to prepare a dilute solution of the absorption aid. 50 mL of this solution was placed in a 250 mL Erlenmeyer flask. Subsequently, approximately 0.1 g of a specific tar simulation component was added to the solution. The Erlenmeyer flask was sealed and placed in a constant-temperature shaker at 25°C and 150 rpm for 4 hours to ensure sufficient contact between the oil phase and the absorption aid. After the shaker was completed, the solution was filtered through filter paper, the filtrate was collected, and the concentration of the tar simulation component in the filtrate was measured using a UV-Vis spectrophotometer at the characteristic absorption wavelength of the tar simulation component. The absorption rate of the absorption aid on the oil phase was calculated based on the difference between the concentration of the tar simulation component in the filtrate and the initial concentration. The blank control group did not add the absorption aid but instead added 0.1 g of deionized water; all other procedures were identical.

[0083] The absorption rates of these 100 absorption aids were experimentally determined to form a small dataset. This dataset was then used to perform transfer learning on a pre-trained deep learning model A1, i.e., freezing all weight parameters except for the last layer of neurons, and training for 10 rounds to obtain a high-throughput screening model B1. The regression coefficients R on the training set were... 2 The regression coefficient R0.97 for the test set is 0.97. 2 The value was 0.96. Using the high-throughput screening model B1, all structures in the absorption aid structure library were predicted. The absorption aid with the highest absorption rate was found to be a poly(ethylene oxide)-propylene oxide tetrablock copolymer, with an EO:PO molar ratio of approximately 0.7, a hydrophilic-lipophilic balance (HLB) value of around 8, and an absorption rate of 90%. The absorption rate of the blank control group was 10%, thus the screening was completed.

[0084] Example 3:

[0085] According to the process of Example 1 above, this example extracts 3000 samples from the absorption aid structure library and the tar molecular structure library using the Latin hypercube sampling method, performs molecular dynamics simulation on these 3000 samples, and constructs a simulation dataset; the dataset is divided into training set and validation set in an 8:2 ratio. The deep learning model employs a dual-branch convolutional neural network to extract features from the molecular fingerprints of ECFPs (electrochemically active polymeric substances) of the absorption aids and tar components, respectively. Each branch of the deep learning model includes a combination of a Conv1D convolutional layer (32 kernels, 50 kernel size in the first layer; 64 kernels, 10 kernel size in the second layer) and a MaxPooling1D pooling layer (pooling window size of 2). After extracting local patterns, these are expanded into one-dimensional vectors using a Flatten layer. The two feature streams are fused in a Concatenate layer, then combined and nonlinearly mapped through a fully connected layer (128 neurons). Dropout (scale 0.3) is used to reduce overfitting risk. Finally, the output layer regresses and predicts the target absorption rate. The activation function is ReLU, the optimizer is Adam, and the loss function is mean squared error (MSE). After training, a pre-trained deep learning model A2 is obtained, with regression coefficients R on the training set. 2 The regression coefficient R0 for the test set is 0.98. 2 It is 0.93.

[0086] In this embodiment, the simulated tar components are selected as a mixture of naphthalene and xylene in a ratio of 1:10. The molecular fingerprint data of the mixture is obtained by weighted averaging of the ECFPs molecular fingerprints of the two components. The pre-trained deep learning model A2 is used to predict all structures in the absorption structure library. Then, structures with an absorption rate greater than 80% are selected to form a subset of high-efficiency absorption aids. This subset is clustered using the K-Means algorithm, and then 80 absorption aids are extracted from all clusters.

[0087] This embodiment experimentally determined the absorption rate of these 80 absorption aids (the experimental determination process is the same as in Embodiment 2), forming a small dataset. This dataset was then used to perform transfer learning on the pre-trained deep learning model A2, i.e., freezing all weight parameters except for the last layer of neurons, and training for 10 rounds to obtain the high-throughput screening model B2. The regression coefficients R of the training set are... 2 The regression coefficient R0 for the test set is 0.98. 2 The value was 0.95. Using the high-throughput screening model B2, all structures in the absorption aid structure library were predicted. The absorption aid with the highest absorption rate was found to be a poly(L-tyrosine)-block-poly(L-phenylalanine) copolymer, with an HLB value of approximately 9.1 and an absorption rate of 95%. The absorption rate of the blank control group was 15%, thus completing the screening.

[0088] Example 4:

[0089] In this embodiment, a mixture of naphthalene, naphthalene, and phenol was selected as the simulated tar components, with a composition ratio of 3:3:2. The molecular fingerprint data of the mixture was obtained by weighted averaging the ECFPs molecular fingerprints of the three components. A pre-trained deep learning model A2 was used to predict all structures in the absorption structure library, and then structures with an absorption rate greater than 80% were selected to form a subset of high-efficiency absorption aids. This subset was then clustered using the K-Means algorithm, and 60 absorption aids were extracted from all clusters.

[0090] This embodiment experimentally determined the absorption rate of these 60 absorption aids (the experimental determination process is the same as in Embodiment 2), forming a small dataset. This dataset was then used to perform transfer learning on the pre-trained deep learning model A2, i.e., freezing all weight parameters except for the last layer of neurons, and training for 10 rounds to obtain the high-throughput screening model B3. The regression coefficients R of the training set were... 2 The regression coefficient R0 for the test set is 0.98. 2 The value was 0.95. Using the high-throughput screening model B3, all structures in the absorption aid structure library were predicted. The absorption aid with the highest absorption rate was a mixture of polyethylene oxide-propylene oxide tetrablock copolymer and isomeric tridecyl alcohol polyoxyethylene ether 1307 in a ratio of 3:1, with an absorption rate of 96%. The absorption rate of the blank control group was 20%, and the screening was completed.

[0091] Comparative Example 1:

[0092] Following the procedure of Example 1 above, 200 samples were extracted from the absorption aid structure library and the tar molecular structure library using the Latin hypercube sampling method. Molecular dynamics simulations were performed on these 200 samples to construct a simulation dataset. The dataset was divided into training and validation sets in an 8:2 ratio. The deep learning model structure adopted a dual-branch convolutional neural network to extract features from the ECFPs molecular fingerprints of absorption aids and tar components, respectively. In the deep learning model, each branch contains a combination of Conv1D convolutional layers (the first convolutional layer has 32 kernels and a kernel size of 50; the second convolutional layer has 64 kernels and a kernel size of 10) and MaxPooling1D pooling layers (pooling window size of 2). After extracting local patterns, the features were expanded into a one-dimensional vector through a Flatten layer. The two features were fused in the Concatenate layer, and then combined and nonlinearly mapped through a fully connected layer (128 neurons). Dropout (ratio 0.3) was used to reduce the risk of overfitting. Finally, the output layer regressed to predict the target absorption rate. The activation function is ReLU, the optimizer is Adam, and the loss function is mean squared error (MSE). After training, a pre-trained deep learning model C is obtained, and the regression coefficients R of the training set are... 2 The regression coefficient R0 for the test set is 0.98. 2 The value is 0.54. The sample size is insufficient, and the deep learning model C is severely overfitted and cannot be used for prediction.

[0093] Comparative Example 2:

[0094] Following the procedure in Example 1 above, 3000 samples were extracted from the absorbent structure library and the tar molecular structure library using the Latin hypercube sampling method. Molecular dynamics simulations were performed on these 3000 samples to construct a simulation dataset. The dataset was divided into a training set and a validation set in an 8:2 ratio. A machine learning model D was trained using the random forest algorithm, and the regression coefficients R of the training set were... 2 The regression coefficient R0.67 for the test set is 0.67. 2 The value is 0.45. The regression coefficient is low because the random forest algorithm cannot effectively learn from high-dimensional molecular fingerprint data.

[0095] Comparative Example 3:

[0096] Following the procedure in Example 1 above, 3000 samples were extracted from the absorbent structure library and the tar molecular structure library using the Latin hypercube sampling method. Molecular dynamics simulations were performed on these 3000 samples to construct a simulation dataset. The dataset was divided into training and validation sets in an 8:2 ratio. The deep learning model structure employed a dual-branch convolutional neural network to extract features from the non-counting molecular fingerprints (MACCS) of absorbents and tar components. The structure of the deep learning model is as follows. Figure 2 As shown, each branch contains a combination of Conv1D convolutional layers (the first convolutional layer has 32 kernels and a kernel size of 50; the second convolutional layer has 64 kernels and a kernel size of 10) and MaxPooling1D pooling layers (pooling window size of 2). After extracting local patterns, these are expanded into one-dimensional vectors through Flatten layers. The two feature streams are fused in a Concatenate layer, then combined and non-linearly mapped through a fully connected layer (128 neurons). Dropout (scale 0.3) is used to reduce overfitting risk, and finally, the output layer regresses to predict the target absorption rate. The activation function is ReLU, the optimizer is Adam, and the loss function is mean squared error (MSE). After training, a pre-trained deep learning model E is obtained, and the regression coefficients R of the training set are... 2 The regression coefficient R0.70 for the test set is 0.70. 2 The value is 0.55. This model cannot be used for further prediction because MACCS fingerprints are not count fingerprints and cannot identify differences in the number of functional groups. For example, the number of PO / EO has a significant impact on hydrophilicity and hydrophobicity, so it cannot complete the learning of this "structure-property relationship".

[0097] Comparative Example 4:

[0098] Phenol was selected as the simulated tar component. The A2 deep learning model was used to predict the absorption structure in the library, and structures with an absorption rate greater than 80% were selected to form a subset of high-efficiency absorption aids. This subset was then clustered using the K-Means algorithm, and five absorption aids were extracted from all clusters.

[0099] The absorption rates of these five absorption aids were experimentally determined (the experimental procedures were the same as in Example 1), forming a small dataset. This dataset was then used to perform transfer learning on the deep learning model A2, i.e., freezing all weight parameters except for the last layer of neurons, and training for 10 rounds to obtain the high-throughput screening model B4. The regression coefficients R on the training set were... 2 The regression coefficient R0 for the test set is 0.99. 2The value was 0.90. Using the high-throughput screening model B4, all structures in the absorption aid structure library were predicted. The absorption aid with the highest absorption rate was identified as a polyethylene oxide-propylene oxide block copolymer with an EO:PO molar ratio of approximately 0.3, an HLB value of approximately 4, and an absorption rate of 65%. The absorption rate of the blank control group was 30%. This discrepancy was due to insufficient experimental data samples, which prevented effective transfer learning and thus failed to correct errors in the simulation data, leading to prediction bias.

[0100] Example 5:

[0101] like Figure 5 As shown, this embodiment provides a tar absorbent performance prediction and high-throughput screening system. The system includes a first construction module 501, a second construction module 502, a simulation module 503, a pre-training module 504, a screening module 505, a measurement module 506, a fine-tuning training module 507, and a prediction module 508. Specific descriptions of each module are as follows:

[0102] The first building module 501 is used to build an absorption aid structure library, converting each structure or combination of structures into a molecular fingerprint.

[0103] The second building module 502 is used to build a tar component structure library and convert each structure or combination of structures into a molecular fingerprint.

[0104] Simulation module 503 is used to sample the absorption aid structure library and the tar component structure library respectively, establish an aqueous phase-absorption aid-oil phase model, and perform sufficient molecular dynamics simulations until a dataset of a specified size is formed.

[0105] The pre-training module 504 is used to pre-train the deep learning network model based on the dataset, wherein the molecular fingerprints of the absorption aids and tar components are used as structural descriptors and the absorption rate is used as the output.

[0106] The screening module 505 is used to select the actual tar components, determine the tar component descriptors, and use a pre-trained deep learning network model to screen out a subset of absorption aids whose absorption rate of the selected tar components reaches a preset percentage.

[0107] The measurement module 506 is used to identify multiple representative structures of absorption aids in the subset of absorption aids using a clustering algorithm, and to measure the absorption rate experimentally.

[0108] The fine-tuning training module 507 is used to fine-tune the pre-trained deep learning network model based on the experimentally determined absorption rate to obtain a high-throughput screening model.

[0109] The prediction module 508 is used to predict all structures in the absorption aid structure library using a high-throughput screening model to obtain the absorption aid with the highest absorption rate.

[0110] It should be noted that the system provided in this embodiment is only an example of the above-described division of functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure can be divided into different functional modules to complete all or part of the functions described above.

[0111] Example 6:

[0112] This embodiment provides a computer device, such as... Figure 6 As shown, it includes a processor 602, a memory, an input device 603, a display device 604, and a network interface 605 connected via a system bus 601. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium 906 and internal memory 607. The non-volatile storage medium 606 stores an operating system, computer programs, and a database. The internal memory 607 provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. When the processor 602 executes the computer programs stored in the memory, it implements the tar absorbent performance prediction and high-throughput screening method of Embodiment 1 described above, as follows:

[0113] A library of absorbent structures is constructed, converting each structure or combination of structures into a molecular fingerprint. A library of tar component structures is also constructed, converting each structure or combination of structures into a molecular fingerprint. Samples are taken from both the absorbent structure library and the tar component structure library to establish an aqueous-absorbent-oil phase model, and thorough molecular dynamics simulations are performed until a dataset of a specified size is formed. Based on this dataset, a deep learning network model is pre-trained, where the molecular fingerprints of the absorbents and tar components serve as structural descriptors, and the absorption rate is used as the output. Actual tar components are selected, and tar component descriptors are determined. The pre-trained deep learning network model is used to screen a subset of absorbents whose absorption rates reach a preset percentage for the selected tar components. A clustering algorithm is used to identify absorbents with multiple representative structures within the subset, and their absorption rates are experimentally determined. Based on the experimentally determined absorption rates, the pre-trained deep learning network model is fine-tuned to obtain a high-throughput screening model. The high-throughput screening model is used to predict the absorption rates of all structures in the absorbent structure library, identifying the absorbent with the highest absorption rate.

[0114] Example 7:

[0115] This embodiment provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the tar absorbent performance prediction and high-throughput screening method of Embodiment 1 above, as follows:

[0116] A library of absorbent structures is constructed, converting each structure or combination of structures into a molecular fingerprint. A library of tar component structures is also constructed, converting each structure or combination of structures into a molecular fingerprint. Samples are taken from both the absorbent structure library and the tar component structure library to establish an aqueous-absorbent-oil phase model, and thorough molecular dynamics simulations are performed until a dataset of a specified size is formed. Based on this dataset, a deep learning network model is pre-trained, where the molecular fingerprints of the absorbents and tar components serve as structural descriptors, and the absorption rate is used as the output. Actual tar components are selected, and tar component descriptors are determined. The pre-trained deep learning network model is used to screen a subset of absorbents whose absorption rates reach a preset percentage for the selected tar components. A clustering algorithm is used to identify absorbents with multiple representative structures within the subset, and their absorption rates are experimentally determined. Based on the experimentally determined absorption rates, the pre-trained deep learning network model is fine-tuned to obtain a high-throughput screening model. The high-throughput screening model is used to predict the absorption rates of all structures in the absorbent structure library, identifying the absorbent with the highest absorption rate.

[0117] It should be noted that the computer-readable storage medium in this embodiment can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0118] In this embodiment, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this embodiment, the computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable program. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0119] The computer-readable storage medium described above can be used to write computer programs for executing this embodiment in one or more programming languages ​​or combinations thereof. These programming languages ​​include object-oriented programming languages—such as Java, Python, and C++—and conventional procedural programming languages—such as C or similar programming languages. The program can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0120] In summary, this invention can rapidly screen for highly efficient additives to improve the washing effect of crude coal gas, fully utilize simulation data to reduce experimental costs, solve the problem of insufficient experimental data through molecular dynamics simulation, and utilize transfer learning to adjust the model from the "simulation domain" to the "real experimental domain" with only a small amount of experimental data, thus solving the problem of difficult-to-obtain experimental data. The model trained on simulation data covers a wide parameter space, and the model after transfer learning can take into account both theoretical laws and real physical processes, resulting in better prediction stability. Furthermore, it has guiding significance for the screening of other high-throughput materials.

[0121] The above description is merely a preferred embodiment of the present invention. It will be apparent to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects. The scope of the present invention is defined by the appended claims rather than the foregoing description, and thus all variations falling within the meaning and scope of the equivalents of the claims are intended to be included within the present invention.

Claims

1. A method for predicting the performance of tar absorbents and conducting high-throughput screening, characterized in that, The method includes: Construct a library of absorption aid structures and convert each structure or combination of structures into a molecular fingerprint; Construct a tar component structure library and convert each structure or combination of structures into a molecular fingerprint; Sampling was performed on the absorption aid structure library and the tar component structure library respectively, and an aqueous phase-absorption aid-oil phase model was established. Sufficient molecular dynamics simulations were conducted until a dataset of a specified size was formed. Based on the dataset, a deep learning network model is pre-trained, where the molecular fingerprints of the absorption aids and tar components are used as structural descriptors and the absorption rate is used as the output. Select the actual tar components, determine the tar component descriptors, and use a pre-trained deep learning network model to screen out a subset of absorption aids whose absorption rate of the selected tar components reaches a preset percentage. Clustering algorithms were used to identify absorption aids with multiple representative structures in the subset of absorption aids, and the absorption rate was determined experimentally. Based on the experimentally determined absorption rate, the pre-trained deep learning network model was fine-tuned to obtain a high-throughput screening model. A high-throughput screening model was used to predict all structures in the absorption aid structure library to obtain the absorption aid with the highest absorption rate.

2. The method for predicting and high-throughput screening the performance of tar absorbents according to claim 1, characterized in that, The molecular dynamics simulations include: The simulation system was subjected to energy minimization to eliminate high-energy conformations. The system was then subjected to kinetic relaxation in the NVT ensemble for 500–1500 ps until it reached a stable state. The equilibrium stability criteria were that the system temperature and energy fluctuations were within 3%–10%. The system was then placed in the NPT ensemble for relaxation for 500–1500 ps to ensure that the densities of each component reached reasonable values ​​until the system stabilized. The convergence criteria were that the system temperature, energy, density, and box size fluctuations were within 3%–10%. Production data was then collected for 1000–2000 ps under the NPT ensemble, and the number of tar molecules in the aqueous phase was counted. The absorption rate was calculated, which was equal to the number of tar molecules in the aqueous phase divided by the total number of tar molecules.

3. The method for predicting and high-throughput screening the performance of tar absorbents according to claim 2, characterized in that, The simulation system is as follows: periodic boundary conditions are used in each direction, the cutoff radius is 12.5 Å, the long-range Coulomb electrostatic interaction between particles is calculated using the Ewald algorithm, the van der Waals interaction between particles is calculated using the Atom-based algorithm, the temperature of the dynamic relaxation process is set to 298 K, the simulation pressure is 0.1 MPa, and the simulation step size is 1 fs.

4. The method for predicting and high-throughput screening the performance of tar absorbents according to any one of claims 1-3, characterized in that, The combined structures in the absorption aid structure library refer to composite hydrocarbon substances composed of two or more hydrocarbon molecules in a preset ratio. The molecular fingerprint of the composite hydrocarbon substance is obtained by weighted average of the molecular fingerprint data of pure hydrocarbons. The combined structures in the tar component structure library refer to composite hydrocarbon substances composed of two or more hydrocarbon molecules in a preset ratio. The molecular fingerprint of the composite hydrocarbon substance is obtained by weighted averaging of the molecular fingerprint data of pure hydrocarbons.

5. The method for predicting and high-throughput screening the performance of tar absorbents according to any one of claims 1-3, characterized in that, The sampling method used for sampling the absorption aid structure library and the tar component structure library is Latin hypercube sampling.

6. The method for predicting and high-throughput screening the performance of tar absorbents according to any one of claims 1-3, characterized in that, The deep learning model employs a dual-branch convolutional neural network to extract features from the molecular fingerprints of the absorption aids and tar components, respectively. Each branch includes a combination of Conv1D convolutional layers and MaxPooling1D pooling layers. After extracting local patterns, the features are expanded into a one-dimensional vector through a Flatten layer. The two features are fused in a Concatenate layer, then combined and nonlinearly mapped through a fully connected layer. Finally, the output layer regresses and predicts the target absorption rate.

7. The method for predicting and high-throughput screening the performance of tar absorbents according to any one of claims 1-3, characterized in that, The clustering algorithm is the K-means clustering algorithm.

8. A tar absorbent performance prediction and high-throughput screening system, characterized in that, The system includes: The first building module is used to build an absorption aid structure library, converting each structure or combination of structures into a molecular fingerprint; The second building module is used to build a tar component structure library, converting each structure or combination of structures into a molecular fingerprint. The simulation module is used to sample the absorption aid structure library and the tar component structure library respectively, establish an aqueous phase-absorption aid-oil phase model, and perform sufficient molecular dynamics simulations until a dataset of a specified size is formed. The pre-training module is used to pre-train the deep learning network model based on the dataset, wherein the molecular fingerprints of the absorption aids and tar components are used as structural descriptors and the absorption rate is used as the output. The screening module is used to select the actual tar components, determine the tar component descriptors, and use a pre-trained deep learning network model to screen out a subset of absorption aids whose absorption rate of the selected tar components reaches a preset percentage. The measurement module is used to identify multiple representative structures of absorption aids in the subset of absorption aids using clustering algorithms, and then experimentally measure the absorption rate. The fine-tuning training module is used to fine-tune the pre-trained deep learning network model based on the experimentally determined absorption rate, so as to obtain a high-throughput screening model. The prediction module is used to predict all structures in the absorption aid structure library using a high-throughput screening model to obtain the absorption aid with the highest absorption rate.

9. A computer device comprising a processor and a memory for storing a processor-executable program, characterized in that, When the processor executes the program stored in the memory, it implements the tar absorbent performance prediction and high-throughput screening method according to any one of claims 1-7.

10. A computer-readable storage medium storing a program, characterized in that, When the program is executed by the processor, it implements the tar absorbent performance prediction and high-throughput screening method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Tar removing device capable of operating for long period and method

    CN103666584A

  • A washing mechanism for coarse coal gas

    CN207159165U

  • Coarse coal gas purification device

    CN207391363U