Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

54 results about "Molecular descriptor" patented technology

Molecular descriptors play a fundamental role in chemistry, pharmaceutical sciences, environmental protection policy, and health researches, as well as in quality control, being the way molecules, thought of as real bodies, are transformed into numbers, allowing some mathematical treatment of the chemical information contained in the molecule.

Micro-emulsion interfacial tension efficient prediction method and system based on active learning and molecular dynamics

The invention discloses a microemulsion interfacial tension efficient prediction method and system based on active learning and molecular dynamics. According to the method, 217 molecular descriptors corresponding to each molecular structure are calculated by adopting an RDKit software package, and the descriptors are used for representing molecular structure characteristics and serve as input variables of a machine learning model, so that key structure information including molecular branching degree, polarity and the like is transmitted. For an oil-water-surfactant ternary interface system, the oil-water interfacial tension in the presence of a surfactant is simulated and calculated through molecular dynamics, and an IFT value is set as a model prediction target. An active learning mechanism is introduced, and iterative sample labeling in the molecular dynamics simulation process is guided; and integrating the obtained IFT data with the molecular descriptor features, constructing a machine learning data set, and training a random forest model. According to the method, the problem of screening a high-performance surfactant layer by a middle-phase microemulsion system can be solved, and the ultra-low oil-water interfacial tension can be rapidly and efficiently screened.
Owner:SICHUAN UNIV

Water pollutant molecular property prediction method and system based on pre-training model and multi-source coding feature fusion

The invention discloses a pre-training model and multi-source coding feature fusion-based water pollutant molecular property prediction method and system, and relates to the technical field of environmental science and artificial intelligence crossing. Comprising the steps of collecting pollutant data; the collected pollutant data is preprocessed; taking SMILES as a molecular input basis, and extracting molecular features; according to the molecular features, feature vectors are obtained from an encoder and then input to a vector interaction module for feature fusion; and embedding the comprehensive molecules obtained after feature fusion into an expert mixed structure module to obtain comprehensive representation of an expert mechanism, inputting the comprehensive representation of the expert mechanism into a preset prediction head, and completing prediction of molecular properties through the prediction head. According to the method, the molecular language model features, the molecular graph neural network features and the molecular descriptor features are integrated, and feature fusion is carried out by utilizing an attention mechanism and residual connection, so that high-precision prediction of the molecular properties of pollutants is realized.
Owner:HUIZHOU WATER TECHNOLOGY CO LTD +1

Method for predicting liquid chromatogram retention time of active ingredients of traditional Chinese medicine

The invention discloses a method for predicting liquid chromatography retention time of traditional Chinese medicine active ingredients. The method comprises the following steps: detecting retention time of flavonoid compounds or anthraquinone compounds under different chromatographic conditions through a reverse high performance liquid chromatograph; performing structure optimization on the flavonoid compound or the anthraquinone compound; performing molecular descriptor calculation on the optimized compound structure to obtain a molecular descriptor data set; performing genetic algorithm screening on the molecular descriptor data set to obtain characteristic molecular descriptors; combining the characteristic molecule descriptors with the different chromatographic conditions into a complete data set; the complete data set serves as an input variable, retention time serves as an output variable, and a gradient elevator GB algorithm or a random forest method RF is adopted for modeling to obtain a QSRR model; and predicting the retention time of flavonoid or anthraquinone compounds under different chromatographic conditions by using the constructed QSRR model.
Owner:CHONGQING UNIV OF TRADITIONAL CHINESE MEDICINE

Anthocyanin stabilizer screening method and system based on fusion of quantum chemistry calculation and machine learning

The invention relates to the technical field of electrical data processing, and provides an anthocyanin stabilizer screening method and system based on fusion of quantum chemistry calculation and machine learning. The method comprises the following steps: constructing a molecular structure data set; performing geometric structure optimization on the anthocyanin molecules and the candidate polyphenol molecules to obtain a ground state optimization structure; performing multi-dimensional molecule descriptor calculation on the candidate polyphenol molecules to obtain a polyphenol molecule descriptor matrix; performing structure optimization, single-point energy calculation and free energy calculation on the anthocyanin polyphenol compound system to obtain a combined free energy value set; integrating the polyphenol molecule descriptor matrix and the combination free energy value set, constructing a training data set, and performing feature screening; performing model training through a machine learning algorithm to obtain a combined free energy prediction model; and predicting the binding free energy of the anthocyanin polyphenol compound through the binding free energy prediction model to obtain candidate molecules of the anthocyanin stabilizer. The screening efficiency of the anthocyanin stabilizer is improved.
Owner:PEKING UNIV INST OF ADVANCED AGRI SCI +1

In silico toxicity risk evaluation method using machine learning technology

To provide an insilico toxicity risk evaluation method in which high prediction accuracy is compatible with the explicitness of a prediction basis.SOLUTION: A machine learning model for predicting the intensity of the toxicity risk of a chemical substance by an information processing system is constructed using a molecular descriptor based on a chemical structure and inchemico or invitro test data, the prediction result of the toxicity risk intensity output by the information processing system by the machine learning model is used as an index of similarity specific to toxicity, and the toxicity risk is evaluated by a Reed-Acros method based on the similarity by the index.SELECTED DRAWING: Figure 1
Owner:SUNSTAR INC

Formulation graph convolution networks (f-GCN) for predicting performance of formulated products

A formulation graph convolution network (F-GCN) with multiple GCNs assembled in parallel and connected to filters and an external learning architecture is able to predict the effectiveness of a formulation. Input into the multiple GCNs are molecular structures of formulants, which are processed as molecular graphs and output as molecular descriptors. The molecular descriptors are filtered by normalized ratios or fractions of the ingredient molecules in a formulation, such as a battery electrolyte or solvent. A formulation descriptor combines the filtered molecular descriptors to arrive at a predicted performance for the formulation, such as the battery capacity for an electrolyte formulation, by an external learning architecture. F-GCN may use a pre-trained GCN with physico-chemical properties of known molecular structures.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Multi-label smell description prediction method

The invention discloses a multi-label smell description prediction method, and relates to the field of compound smell prediction, and the method comprises the steps: obtaining compound identification information, molecular structure descriptors and smell label data, and constructing a multi-label smell data set; generating a molecular structure feature vector through a molecular fingerprint coding technology, and extracting a multi-dimensional descriptor reflecting the physicochemical properties of molecules; compressing the molecular fingerprint features to a low-dimensional space through a dimension reduction algorithm; performing unbalanced data processing on the training set, fusing the dimension-reduced molecular fingerprints with the molecular descriptors to form a joint feature matrix, and configuring a class weight balance mechanism and overfitting suppression parameters by adopting a multi-label classification architecture; independently optimizing a probability threshold for each odor label based on the verification set; and outputting a multi-odor label combination prediction result according to the target molecule identification information. According to the scheme, the multi-odor characteristics of the compound can be accurately depicted, and the combined recognition accuracy of the compound odor is remarkably improved.
Owner:RES CENT FOR ECO ENVIRONMENTAL SCI THE CHINESE ACAD OF SCI

A high trophic level food chain biomagnification prediction model for organic chemicals

This invention discloses a high-trophic-level food chain biomagnification prediction model for organic chemicals. The model comprises the following steps: S1. Dividing the sample into six training samples and three validation samples; S2. Retaining a 1169-dimensional molecular descriptor matrix after screening; S3. Generating an extended pollutant parameter database by combining laboratory biomagnification factors; S4. Forming a position-time normalized sample dataset; S5. Obtaining a fused amplification probability result set; S6. Identifying a candidate set of pollutant amplification paths on the coupled directed graph of amplification probability-flux-stability; S7. Performing uncertainty propagation and confidence assessment on the candidate set of pollutant amplification paths, outputting a priority control list of pollutant amplification paths and their confidence information. This invention addresses the problems of insufficient model coverage and weak extrapolation ability caused by traditional methods relying solely on single molecular descriptors or experimental values.
Owner:NANJING INST OF ENVIRONMENTAL SCI MINIST OF ECOLOGY & ENVIRONMENT OF THE PEOPLES REPUBLIC OF CHINA

A method and system for screening anthocyanin stabilizers based on the integration of quantum chemistry calculation and machine learning

The present application relates to the technical field of electric data processing, and provides a screening method and system for anthocyanin stabilizer based on quantum chemistry calculation and machine learning fusion. The method comprises the following steps: constructing a molecular structure dataset; performing geometric structure optimization on anthocyanin molecules and candidate polyphenol molecules to obtain ground state optimized structures; performing multidimensional molecular descriptor calculation on the candidate polyphenol molecules to obtain a polyphenol molecule descriptor matrix; performing structure optimization, single-point energy calculation and free energy calculation on the anthocyanin polyphenol complex system to obtain a binding free energy numerical set; integrating the polyphenol molecule descriptor matrix and the binding free energy numerical set, constructing a training dataset and performing feature screening; performing model training through a machine learning algorithm to obtain a binding free energy prediction model; and predicting the binding free energy of the anthocyanin polyphenol complex through the binding free energy prediction model to obtain candidate molecules for anthocyanin stabilizer. The present application improves the screening efficiency of anthocyanin stabilizer.
Owner:PEKING UNIV INST OF ADVANCED AGRI SCI +1

A method for predicting active structures of interfering biological pathways

ActiveCN116052772Bimprove data qualityImprove forecast resultsBiostatisticsHybridisationChemical compoundBiological pathway
The application discloses a method for predicting active structures interfering with biological pathways, which comprises the following steps: constructing a compound biological pathway interference database containing clear labels such as cell lines, exposure time and exposure concentration; evaluating the consistency of the biological pathway cross degree and the regulation trend of the training set and the test set through cumulative hypergeometric distribution and cumulative Bernoulli distribution; identifying a batch of potential compounds in the training set; evaluating the occurrence frequency of the molecular descriptors in the potential compounds through cumulative distribution probability; and finally predicting the potential active structures driving the change of the biological pathway through inputting the biological pathway.
Owner:NANJING UNIV

Prediction method of molecular frontline orbital energy level and related device

The invention provides a prediction method of molecular frontline orbit energy level and a related device, and relates to the technical field of chemistry. The prediction method for the molecular frontline orbit energy level can comprise the following steps: acquiring target information of a target molecule; wherein the target information comprises molecular structure information and a molecular descriptor; the molecular descriptors comprise a first molecular descriptor for predicting the highest occupied molecular orbital energy level and a second molecular descriptor for predicting the lowest unoccupied molecular orbital energy level; and inputting the target information into a pre-trained graph neural network model, and predicting the highest occupied molecular orbital energy level and the lowest unoccupied molecular orbital energy level of the target molecule through the graph neural network model. According to the technical scheme provided by the invention, the problem of low prediction efficiency of a scheme for predicting HOMO and LUMO energy level values through a DFT method in the prior art can be solved.
Owner:GUANGYIN (JIANGSU) NEW ENERGY CO LTD

Methods and systems for predicting the performance of polyimide, methods and systems for structural design, electronic devices, computer-readable storage media, and computer products.

This invention discloses a method and system for predicting the performance of polyimides, a method and system for structural design, an electronic device, a computer-readable storage medium, and a computer product. The method for predicting the performance of polyimides includes the following steps: inputting a first molecular descriptor into an AFP model to obtain the physical property parameter to be tested; the first molecular descriptor is the molecular descriptor of the polyimide to be tested; the AFP model is trained using a training dataset, which includes a second molecular descriptor and the first physical property parameter; the second molecular descriptor and the first physical property parameter are respectively the molecular descriptor and physical property parameter of a known polyimide; both the first and second molecular descriptors include node features and edge features, and the type of the first molecular descriptor is the same as the type of the second molecular descriptor; the type of the physical property parameter to be tested is the same as the type of the first physical property parameter. This invention has high accuracy and improves the effectiveness and reliability of de novo design.
Owner:EAST CHINA UNIV OF SCI & TECH

System and method for adjusting products to accommodate different regions using QSAR data

A method for predicting properties of a product having alternative ingredients, the method comprising: receiving data of an ingredient profile of a product, a molecular profile associated with the ingredient profile, and a proportioning profile of the product based on the ingredient profile; determining a set of molecular descriptors for the product based on the molecular profile; determining a set of material descriptors for the product based on the set of molecular descriptors; determining a set of proportioning descriptors for the product based on the set of material descriptors and the proportioning profile; generating a measured set of functional properties of the product using a machine learning model based on the one or more proportioning descriptors; receiving one or more alternative ingredients of the product; and generating a predicted set of functional properties of the product by using the machine learning model based on the one or more alternative components.
Owner:MARS INC

AIE fluorescent probe-oriented multi-index weighted scoring screening method and system

The invention provides an AIE fluorescent probe-oriented multi-index weighted scoring screening method and system. The method comprises the following steps: acquiring data such as a probe molecular structure and an experimental environment, and preprocessing the data to obtain a feature data set; for thermal stability, light stability, Stokes shift and AIE enhancement factors, optimal molecular descriptors and dichotomy prediction models are selected respectively; inputting a to-be-evaluated probe into the target model to obtain a prediction result, and calculating a synthetic accessibility score; combining the prediction result with the standardized synthetic accessibility score to form index data; the weight of each index is determined through mixed data factor analysis, and a comprehensive score is calculated to realize probe sorting and screening. According to the method, efficient screening of the AIE probe can be realized, the trial and error cost is reduced, and the scoring stability and applicability are enhanced.
Owner:CENT SOUTH UNIV

A method and device for screening potential substitutes of bisphenol a and a storage medium

The application relates to a bisphenol A potential substitute screening method, device and storage medium, the method comprising the following steps: acquiring a global AR antagonistic data set and a local BPA analogue data set; acquiring molecular characterization of the global and local BPA analogue data sets; acquiring a machine learning algorithm model; optimizing parameters of the machine learning algorithm model; processing the global AR antagonistic data set and the local BPA analogue data set by using the machine learning algorithm model; and acquiring a potential BPA substitute predicted by an optimal mixed model. The application provides a mixed deep learning architecture, which combines molecular descriptors and molecular graphs to predict the antagonistic activity of a compound on AR; compared with previous models, the mixed model can extract a large amount of chemical information from different molecular features, so that the generalization ability of the model for predicting BPA substitutes is improved; and the prediction result also shows that lignin derivatives are safer than bisphenol analogues as BPA substitutes.
Owner:JIANGHAN UNIVERSITY

An aie fluorescent probe multi-index weighted scoring screening method and system

ActiveCN121983174BFluoProbesFluorescence
The application provides a multi-index weighted scoring screening method and system for AIE fluorescent probes, and the method comprises the following steps: obtaining probe molecular structure, experimental environment and other data, and preprocessing to obtain a feature data set; for thermal stability, light stability, Stokes shift and AIE enhancement factor, the best molecular descriptor and binary classification prediction model are respectively optimized; the probe to be evaluated is input into the target model to obtain a prediction result, and a synthetic accessibility score is calculated; the prediction result is combined with the standardized synthetic accessibility score to form index data; the weight of each index is determined through mixed data factor analysis, and a comprehensive score is calculated to realize probe sorting and screening. The application can realize efficient screening of AIE probes, reduce trial and error cost, and enhance the stability and applicability of the score.
Owner:CENT SOUTH UNIV

Molecular representation and drug screening method based on transformer multi-translation model

The present application relates to the technical field of machine learning, and more particularly to a kind of molecular representation and drug screening method based on Transformer multi-translation model, comprising: S1: downloading small molecule compound from PubChem database, and carrying out molecular pretreatment, obtain molecular descriptor code set;S2: build Transformer multi-translation model and training;S3: using ChEMBL database obtains known activity small molecule dataset;S4: build machine learning classifier and training;S5: the drug to be tested is sequentially input into the trained Transformer multi-translation model and final machine learning classifier and is handled, obtains the screening result of the drug to be tested.The present application can improve screening accuracy, and speed up drug research and development.
Owner:HAINAN UNIV +1

A two-goal solvent screening paradigm based on machine learning

PendingCN122337406ASolubilityEngineering
This invention discloses a dual-objective solvent screening paradigm based on machine learning, belonging to the fields of chemical engineering and machine learning technology. The paradigm includes the following steps: constructing a dual-objective solvent screening paradigm with solubility and flammability as screening objectives; collecting and organizing solubility data and flammability risk data; constructing differentiated feature engineering for solubility prediction and flammability prediction respectively; constructing a multi-dimensional hybrid feature set integrating molecular fingerprints, physicochemical descriptors, and temperature for solubility prediction, and using multi-dimensional molecular descriptors for flammability prediction; optimizing the model using different hyperparameter optimization methods for different prediction tasks; and employing a Pareto ranking mechanism to output a Pareto-optimal solvent set with the objectives of maximizing solubility and minimizing risk, thus achieving dual-objective collaborative screening. This method achieves accurate prediction of solubility and flammability, possesses strong adaptability and generalization ability, and provides an effective tool for high-throughput solvent screening in crystallization process design.
Owner:HEBEI UNIV OF TECH

Method for constructing PFASs risk prediction model based on Koc and BCF

The invention relates to the technical field of pollutant risk prediction, and discloses a method for constructing a PFASs risk prediction model based on Koc and BCF, and the method comprises the steps: collecting various PFASs data, and processing the data as training data for modeling; calculating molecular descriptors of a plurality of PFASs, preliminarily screening the molecular descriptors through a Pearson correlation analysis method, and further screening the preliminarily screened molecular descriptors through multi-task elastic network regression; and constructing a PFASs prediction model by using the screened molecular descriptors based on a multi-task multiple linear regression algorithm in combination with a multi-task joint forward step-by-step selection strategy. The invention aims to construct a prediction model capable of synchronously predicting an organic carbon-water partition coefficient log Koc and a biological enrichment factor log BCF.
Owner:KUNMING UNIV OF SCI & TECH

Chemical evolution prediction method based on spectroscopy inversion neural network algorithm

The invention discloses a chemical evolution prediction method based on a spectroscopy inversion neural network algorithm, and relates to the technical field of inversion prediction. According to the method, a chemical and material database is firstly constructed, so that data support is provided for subsequent prediction; the molecular descriptor fusing the spectral characteristics, the microstructure and the physical property associated information is generated, and the defects that a traditional molecular descriptor is single in information and insufficient in representativeness are overcome; the molecular descriptor is used as input, a spectrum structure-effect relationship prediction model with common feature extraction and multi-branch special prediction capabilities is constructed, and spectrum-structure, structure-effect and spectrum-effect associated synchronous precise learning is realized; the compatibility and reliability of the input data and the model are ensured by carrying out noise reduction and standardization preprocessing on the target spectroscopic data subsequently, and finally, chemical structure change parameters, molecular interaction rules and physical property evolution trend data are directly output through model inversion, so that the defects that a traditional method needs multi-step splitting and experimental verification is tedious are avoided.
Owner:北京机数小来智能科技有限公司

Machine learning screening method for near-infrared two-zone organic photothermal co-crystals

The application belongs to the technical field of artificial intelligence empowered material innovation, and discloses a machine learning screening method for near-infrared two-zone organic photothermal co-crystals, which comprises the following steps: step 1, collecting organic co-crystal molecular data samples to form a first database; collecting molecular co-crystal donor-acceptor molecular formula and co-crystal maximum absorption wavelength to form a second database; step 2, converting the molecular formula in the database into molecular descriptors; step 3, using an organic photothermal co-crystal recognition sub-model to determine the key molecular descriptors and whether the co-crystals can be formed, and through feature knowledge transmission, using a NIR-II screening sub-model to take the key molecular descriptors as input to determine whether the NIR-II organic photothermal co-crystals can be formed; and step 4, inputting the molecular formula of the acceptor and the donor to be tested into the trained and verified rapid screening model to determine whether the absorption wavelength is greater than 1000 nm. The application can simultaneously predict the co-crystal formation ability and NIR-II absorption performance, and realizes one-stop screening.
Owner:TIANJIN UNIV

A method for predicting acute intraperitoneal toxicity based on maccs keys and machine learning

This patent relates to an acute intraperitoneal toxicity prediction method based on MACCS bonds and a machine learning framework. This method collects and processes toxicity data of chemical compounds, calculates molecular descriptors using MACCS bonds, and constructs a gradient boosting decision tree model for toxicity prediction through feature selection and removal of highly correlated features. The model's hyperparameters are optimized using a grid search method, and the interpretability of the model is improved using SHAP value analysis. This achieves high-precision and highly interpretable toxicity prediction. The method achieves a coefficient of determination of 0.914, a mean squared error of 0.047, a root mean square error of 0.218, and a mean absolute error of 0.148 on the test set. This invention significantly improves the accuracy of toxicity prediction and has strong generalization ability, making it widely applicable to toxicity screening in new drug development, reducing reliance on traditional animal experiments, and lowering R&D costs and time.
Owner:QINGDAO UNIV OF SCI & TECH

A machine learning-based method for judging ad five-target inhibitors

This invention relates to a machine learning-based method for identifying five-target inhibitors in Alzheimer's disease (AD). It involves collecting known inhibitor and non-inhibitor molecular structure data for five key target proteins associated with Alzheimer's disease, forming an initial dataset. Two-dimensional molecular descriptors for each molecular structure in the initial dataset are calculated, and a systematic feature selection operation incorporating multiple attribute screening methods and search strategies is used to select a subset of features for modeling from all molecular descriptors. Various machine learning algorithms are used to construct classification models for the five key targets and non-inhibitors, respectively, and these models are integrated to form a multi-target prediction system. The molecular descriptors of the molecules to be predicted are input into the multi-target prediction system. The multi-target prediction system processes the input descriptors and outputs the inhibitor identification result. Compared with existing technologies, this invention has advantages such as high efficiency, strong generalization ability, and strong robustness.
Owner:SHANGHAI UNIV

Prediction method for protoporphyrinogen oxidase (PPO) inhibitor

The present invention provides a technology that objectively and accurately predicts a degree of a PPO inhibitor using LUMO distribution as a “molecular descriptor” and a technology for selecting a PPO inhibitor based on the predicted degree of the PPO inhibitor. Provided are to predict PPO inhibitory activity based on a correlation between LUMO distribution and PPO inhibitory activity for each compound; to predict a degree of PPO inhibitory activity based on a correlation between a variation difference in PPO inhibitory activity and the degree of the PPO inhibitory activity between different plant species; and a method for selecting a PPO inhibitor, using the predicted PPO inhibitory activity as an index.
Owner:PAN ADVANCED BUSINESS RESEARCH LLC

Machine learning screening method for near-infrared two-region organic photo-thermal eutectic crystals

The invention belongs to the technical field of artificial intelligence energizing material innovation, and discloses a near-infrared two-zone organic photo-thermal eutectic machine learning screening method which comprises the following steps: step 1, collecting organic eutectic molecular data samples to form a first database; collecting a molecular eutectic donor-acceptor molecular formula and a eutectic maximum absorption wavelength to form a second database; 2, converting molecular formulas in the database into molecular descriptors; 3, the organic photo-thermal eutectic recognition sub-model is used for judging whether the key molecule descriptors can form eutectic or not, and through feature knowledge transfer, the NIR-II screening sub-model takes the key molecule descriptors as input to judge whether NIR-II organic photo-thermal eutectic can be formed or not; and 4, inputting the molecular formulas of the acceptor and the donor to be tested into the rapid screening model which is trained and verified, and judging whether the absorption wavelength is greater than 1000 nm or not. According to the method, the eutectic forming ability and the NIR-II absorption performance can be predicted at the same time, and one-stop screening is achieved.
Owner:TIANJIN UNIV

Selection of chromatography parameters for the manufacture of therapeutic proteins

UndeterminedES3075763T3Therapeutic proteinEngineering
In a method for facilitating the selection of chromatographic parameters for the fabrication of a therapeutic protein, one or more process parameter values ​​associated with a hypothetical chromatographic process are received, as well as one or more molecular descriptors of the therapeutic protein. The method also includes predicting a performance indicator for the hypothetical chromatographic process, at least by analyzing the process parameters and molecular descriptors using a machine learning model. This model can be a regression tree, an extreme gradient boosting model, or an elastic network model. The method also includes presenting the predicted performance indicator, or an indication of whether the indicator meets one or more acceptability criteria, to a user via a user interface.
Owner:AMGEN INC

Chemical risk prediction model based on interpretable machine learning and construction method thereof

The invention provides a chemical risk prediction model based on interpretable machine learning and a construction method thereof, and the method comprises the steps: S1, obtaining chemical data with risk endpoints, S2, generating a high-dimensional feature set containing physicochemical properties and structural topology for a molecular structure through employing a molecular descriptor and a fingerprint calculation platform; s3, performing comparison modeling on each risk endpoint by independently adopting different types of machine learning models, re-training an optimal model on a training set, and reporting a final result in an independent prediction set; s4, the model application range of the optimal model is determined, and an application domain framework of the model is constructed based on conformal prediction, and S5, the importance of global and sample-level key descriptors is output; and finally, batch reasoning and result exporting: outputting a prediction label, a prediction probability, an application domain framework mark, a significance level and an explanatory result one by one for a chemical list corresponding to the target chemical data.
Owner:FUJIAN NORMAL UNIV

A method of screening optimal descriptors and predicting classification of chemical activity

The application discloses a method for screening optimal descriptors and predicting classification of chemical activity, comprising the following steps: (1) constructing a chemical data set; (2) calculating molecular descriptors; (3) preliminary screening of molecular descriptors; (4) dividing data samples; (5) screening molecular descriptors by using a combination of support vector machines and ant colony optimization algorithm; (6) counting the most frequent descriptors; (7) predicting the chemical activity by using support vector machines according to the most frequent descriptors; the application can make the classifier more robust by iteratively dividing the original training set and running the ant colony optimization algorithm; the initial setting of different path pheromones can make the ant colony optimization algorithm converge in a small number of iterations; the pheromone increment limit setting can avoid the algorithm falling into local optimum; and the statistical frequency can better quantify the importance of features and thus screen the optimal descriptor combination.
Owner:JIANGSU UNIV OF SCI & TECH

MTL and GCN coupled new pollutant risk prediction method and system

The invention discloses a new pollutant risk prediction method and system coupling MTL and GCN, and belongs to the crossing field of new pollutant risk identification and artificial intelligence. The method comprises the following steps: selecting SMILES of chemicals, and obtaining durability, mobility and toxicity attribute indexes of the chemicals; the SMILES is standardized, and the problem that attribute indexes are unbalanced is solved; calculating molecular descriptors according to SMILES, and removing redundant molecular descriptors; taking the screened molecular descriptors as independent variables, taking three attribute indexes of the chemicals as dependent variables, substituting the independent variables and the dependent variables into a LightGBM model, and analyzing and calculating the molecular descriptors with the first five importance ranks under each attribute through SHAP; extracting a molecular map of each chemical according to SMILES to obtain atomic characteristics and chemical bond characteristics of molecules; constructing a prediction model coupled with multi-task learning and a graph convolutional neural network by taking the molecular graph features and the molecular descriptors as independent variables and three attribute indexes of chemicals as dependent variables; and deploying the constructed model as an online prediction tool.
Owner:SICHUAN UNIV

Liquid crystal polyimide molecule design method based on molecule descriptor and machine learning

The invention discloses a liquid crystal polyimide molecule design method based on a molecule descriptor and machine learning, which comprises the following steps: automatically reading an input liquid crystal polyimide molecular structure file through a molecule descriptor extraction module, and calculating the molecule descriptor; calculating thermodynamic and molecular simulation data of the liquid crystal polyimide system by using a molecular dynamics or quantum chemistry method; constructing a machine learning prediction model according to the molecular descriptor and thermodynamics and molecular simulation data in combination with online sensor data; performing forward prediction on the performance of the liquid crystal polyimide by using a machine learning model, and performing reverse generation from the set target performance through a reverse generation model to generate a candidate molecular structure meeting the target performance; and performing synthesis route reasoning on the candidate molecular structure through the reaction template and the chemical knowledge graph, and outputting a feasible synthesis route and experiment conditions. The research and development period of the liquid crystal polyimide can be greatly shortened, and the molecular performance prediction and optimization efficiency can be improved.
Owner:烟台国工智能科技有限公司