Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

95 results about "Molecular descriptor" patented technology

Molecular descriptors play a fundamental role in chemistry, pharmaceutical sciences, environmental protection policy, and health researches, as well as in quality control, being the way molecules, thought of as real bodies, are transformed into numbers, allowing some mathematical treatment of the chemical information contained in the molecule.

Micro-emulsion interfacial tension efficient prediction method and system based on active learning and molecular dynamics

The invention discloses a microemulsion interfacial tension efficient prediction method and system based on active learning and molecular dynamics. According to the method, 217 molecular descriptors corresponding to each molecular structure are calculated by adopting an RDKit software package, and the descriptors are used for representing molecular structure characteristics and serve as input variables of a machine learning model, so that key structure information including molecular branching degree, polarity and the like is transmitted. For an oil-water-surfactant ternary interface system, the oil-water interfacial tension in the presence of a surfactant is simulated and calculated through molecular dynamics, and an IFT value is set as a model prediction target. An active learning mechanism is introduced, and iterative sample labeling in the molecular dynamics simulation process is guided; and integrating the obtained IFT data with the molecular descriptor features, constructing a machine learning data set, and training a random forest model. According to the method, the problem of screening a high-performance surfactant layer by a middle-phase microemulsion system can be solved, and the ultra-low oil-water interfacial tension can be rapidly and efficiently screened.
Owner:SICHUAN UNIV

Method for predicting toxicity of rare and endangered organisms based on machine learning algorithm and quantitative structure-function relationship

The invention discloses a rare and endangered organism toxicity prediction method based on a machine learning algorithm and a quantitative structure-activity relationship, which constructs a toxicity prediction model through the machine learning algorithm, can effectively assess the toxicity influence of environmental pollutants on rare and endangered organisms, and provides technical support for rare and endangered organism protection and ecological environment risk assessment. Comprising the following steps: 1, collecting data related to rare and endangered organisms from a public database, and establishing a rare and endangered organism toxicity prediction database based on machine learning; 2, generating molecular descriptors for the chemical substances; step 3, data preprocessing; 4, development of acute and chronic toxicity prediction models of rare and endangered organisms based on multiple machine learning is carried out, and performance evaluation is carried out; 5, performing internal and external verification on the rare and endangered biotoxicity prediction model; step 6, analyzing the importance of the features by using the valuable, rare and endangered biotoxicity prediction model, and finding out the most influential features; and 7, predicting the toxicity value of the pollutants in combination with the optimal machine learning model.
Owner:BEIHANG UNIV

Prediction method for protoporphyrinogen oxidase (PPO) inhibitors

The present invention provides a technology for predicting the level of PPO inhibition objectively and with high accuracy using LUMO distribution as a "molecular descriptor," and a technology for selecting a PPO inhibitor based on the predicted level of PPO inhibition. [Solution] The present invention provides a method for predicting PPO inhibitory activity based on the correlation between LUMO distribution and PPO inhibitory activity for each compound, predicting the level of PPO inhibitory activity based on the correlation between the difference in variation in PPO inhibitory activity between different plant species and the level of PPO inhibitory activity, and selecting PPO inhibitors using the predicted PPO inhibitory activity as an index.
Owner:PAN ADVANCED BUSINESS RESEARCH LLC

Water pollutant molecular property prediction method and system based on pre-training model and multi-source coding feature fusion

The invention discloses a pre-training model and multi-source coding feature fusion-based water pollutant molecular property prediction method and system, and relates to the technical field of environmental science and artificial intelligence crossing. Comprising the steps of collecting pollutant data; the collected pollutant data is preprocessed; taking SMILES as a molecular input basis, and extracting molecular features; according to the molecular features, feature vectors are obtained from an encoder and then input to a vector interaction module for feature fusion; and embedding the comprehensive molecules obtained after feature fusion into an expert mixed structure module to obtain comprehensive representation of an expert mechanism, inputting the comprehensive representation of the expert mechanism into a preset prediction head, and completing prediction of molecular properties through the prediction head. According to the method, the molecular language model features, the molecular graph neural network features and the molecular descriptor features are integrated, and feature fusion is carried out by utilizing an attention mechanism and residual connection, so that high-precision prediction of the molecular properties of pollutants is realized.
Owner:HUIZHOU WATER TECHNOLOGY CO LTD +1

Reverse molecular design method and system of organic framework material COFs based on machine learning and quantitative calculation technology

The invention discloses a reverse molecular design method and system of an organic framework material COFs based on a machine learning and quantitative calculation technology, and relates to the technical field of material science and computer simulation. The photoelectric property of the COFs material can be accurately predicted, and a new material can be directionally synthesized under guidance. A COFs structure database is established, Gaussian is used for representing the fragment structure and electronic characteristics of the COFs, a molecular descriptor and a machine learning algorithm are used for establishing a structure-performance mapping relation, quantitative calculation software is used for obtaining an important target quantity for describing the photoelectric characteristics of the COFs, SHAP is used for carrying out importance analysis, and finally a machine learning model for predicting the photoelectric properties of the COFs is obtained. According to the model, through theoretical prediction of a novel COFs structure, a mapping relation between the COFs structure and photoelectric specificity can be established, and a novel COFs hypothetical structure can be reset according to important molecular fragments or word structures of the important molecular fragments; theoretical guidance can be provided for application in the fields of photoelectric catalysts and the like.
Owner:HAINAN NORMAL UNIV

New pollutant multi-medium PNEC prediction analysis system and method based on machine learning

The invention relates to the technical field of new pollutant risk identification, in particular to a new pollutant multi-medium PNEC predictive analysis system and method based on machine learning, and the method comprises the steps: collecting a compound SMILES structure and fresh water PNEC, BCF and Koc data thereof; using tools such as RDKit, rcdk and the like to automatically calculate molecular descriptors; key features are screened through redundant feature processing and a random forest algorithm; an RF model, an XGBoost model, a LightGBM model and a CatBoost model are respectively constructed for PNEC prediction, and a BCF model and a Koc model are constructed by adopting the LightGBM; and a visual platform is built based on an R language Shiny framework, so that full-process automation of SMILES input, model prediction and result display is realized. The system provides a convenient tool for multi-medium PNEC prediction, has the advantages of high efficiency, accuracy, low cost and the like, is suitable for large-scale new pollutant environmental risk identification, and has wide application prospects and important environmental protection significance.
Owner:SOUTH CHINA NORMAL UNIV

Method for predicting liquid chromatogram retention time of active ingredients of traditional Chinese medicine

The invention discloses a method for predicting liquid chromatography retention time of traditional Chinese medicine active ingredients. The method comprises the following steps: detecting retention time of flavonoid compounds or anthraquinone compounds under different chromatographic conditions through a reverse high performance liquid chromatograph; performing structure optimization on the flavonoid compound or the anthraquinone compound; performing molecular descriptor calculation on the optimized compound structure to obtain a molecular descriptor data set; performing genetic algorithm screening on the molecular descriptor data set to obtain characteristic molecular descriptors; combining the characteristic molecule descriptors with the different chromatographic conditions into a complete data set; the complete data set serves as an input variable, retention time serves as an output variable, and a gradient elevator GB algorithm or a random forest method RF is adopted for modeling to obtain a QSRR model; and predicting the retention time of flavonoid or anthraquinone compounds under different chromatographic conditions by using the constructed QSRR model.
Owner:CHONGQING UNIV OF TRADITIONAL CHINESE MEDICINE

Method for predicting flammability upper limit volume percentage of pure compounds

The invention provides a method for predicting the flammability upper limit volume percentage of a pure compound. The method can be used for predicting a mathematical model of the flammability upper limit volume percentage of the pure compound which is composed of 12 or less elements of hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus, arsenic and the like and has the atom number of 25 or less (excluding hydrogen) with high accuracy. Wherein the model is obtained as a universal quantitative structure-property relationship model by determining an optimal model from a plurality of multiple linear regression models by means of a step-by-step selection method, and the model takes some of the various molecular descriptors as independent variables, takes flammability upper limit volume percent as a dependent variable, and takes the flammability upper limit volume percent as a dependent variable. The value of the molecular descriptor included in the model can be received and input in a short time, and the flammability upper limit volume percentage can be output, so that only the specific value of the molecular descriptor included in the model is known. And the flammability upper limit volume percentage of the compound purely formed by the molecule can be predicted for any molecule meeting the requirements of the invention.
Owner:BEIJING AISEN ZHONGKE TECHNOLOGY CO LTD

Method for predicting saturated liquid density of pure compound at 298.15 K

Provided is a mathematical model that can predict, with high accuracy, the saturated liquid density of a pure compound at 298.15 K, said pure compound comprising 12 or less elements such as hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic and having 25 or less atoms (excluding hydrogen). The model is a quantitative structure-property relation model and is obtained by solving an optimal model from a plurality of multiple linear regression models by a step-by-step selection method, and the model takes some molecule descriptors in various molecule descriptors as independent variables, takes saturated liquid density under 298.15 K as a dependent variable, and takes the saturated liquid density under 298.15 K as a dependent variable. The value of the molecule descriptor included in the input model can be received in a short time, and the saturated liquid density at 298.15 K can be output, so that the saturated liquid density at 298.15 K of a compound formed by the molecule singly can be predicted for any molecule meeting the requirements of the invention as long as the specific value of the molecule descriptor included in the model is known.
Owner:BEIJING AISEN ZHONGKE TECHNOLOGY CO LTD

Anthocyanin stabilizer screening method and system based on fusion of quantum chemistry calculation and machine learning

The invention relates to the technical field of electrical data processing, and provides an anthocyanin stabilizer screening method and system based on fusion of quantum chemistry calculation and machine learning. The method comprises the following steps: constructing a molecular structure data set; performing geometric structure optimization on the anthocyanin molecules and the candidate polyphenol molecules to obtain a ground state optimization structure; performing multi-dimensional molecule descriptor calculation on the candidate polyphenol molecules to obtain a polyphenol molecule descriptor matrix; performing structure optimization, single-point energy calculation and free energy calculation on the anthocyanin polyphenol compound system to obtain a combined free energy value set; integrating the polyphenol molecule descriptor matrix and the combination free energy value set, constructing a training data set, and performing feature screening; performing model training through a machine learning algorithm to obtain a combined free energy prediction model; and predicting the binding free energy of the anthocyanin polyphenol compound through the binding free energy prediction model to obtain candidate molecules of the anthocyanin stabilizer. The screening efficiency of the anthocyanin stabilizer is improved.
Owner:PEKING UNIV INST OF ADVANCED AGRI SCI +1

In silico toxicity risk evaluation method using machine learning technology

To provide an insilico toxicity risk evaluation method in which high prediction accuracy is compatible with the explicitness of a prediction basis.SOLUTION: A machine learning model for predicting the intensity of the toxicity risk of a chemical substance by an information processing system is constructed using a molecular descriptor based on a chemical structure and inchemico or invitro test data, the prediction result of the toxicity risk intensity output by the information processing system by the machine learning model is used as an index of similarity specific to toxicity, and the toxicity risk is evaluated by a Reed-Acros method based on the similarity by the index.SELECTED DRAWING: Figure 1
Owner:SUNSTAR INC

Formulation graph convolution networks (f-GCN) for predicting performance of formulated products

A formulation graph convolution network (F-GCN) with multiple GCNs assembled in parallel and connected to filters and an external learning architecture is able to predict the effectiveness of a formulation. Input into the multiple GCNs are molecular structures of formulants, which are processed as molecular graphs and output as molecular descriptors. The molecular descriptors are filtered by normalized ratios or fractions of the ingredient molecules in a formulation, such as a battery electrolyte or solvent. A formulation descriptor combines the filtered molecular descriptors to arrive at a predicted performance for the formulation, such as the battery capacity for an electrolyte formulation, by an external learning architecture. F-GCN may use a pre-trained GCN with physico-chemical properties of known molecular structures.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Multi-label smell description prediction method

The invention discloses a multi-label smell description prediction method, and relates to the field of compound smell prediction, and the method comprises the steps: obtaining compound identification information, molecular structure descriptors and smell label data, and constructing a multi-label smell data set; generating a molecular structure feature vector through a molecular fingerprint coding technology, and extracting a multi-dimensional descriptor reflecting the physicochemical properties of molecules; compressing the molecular fingerprint features to a low-dimensional space through a dimension reduction algorithm; performing unbalanced data processing on the training set, fusing the dimension-reduced molecular fingerprints with the molecular descriptors to form a joint feature matrix, and configuring a class weight balance mechanism and overfitting suppression parameters by adopting a multi-label classification architecture; independently optimizing a probability threshold for each odor label based on the verification set; and outputting a multi-odor label combination prediction result according to the target molecule identification information. According to the scheme, the multi-odor characteristics of the compound can be accurately depicted, and the combined recognition accuracy of the compound odor is remarkably improved.
Owner:RES CENT FOR ECO ENVIRONMENTAL SCI THE CHINESE ACAD OF SCI

New pollutant identification method, system and device and storage medium

The invention discloses a new pollutant identification method, system and device, and a storage medium. The new pollutant identification method comprises the following steps: acquiring a multi-modal feature of a to-be-identified pollutant; wherein the multi-modal feature comprises at least one of a first molecular descriptor feature, a first molecular structure image feature and a first text feature; inputting the multi-modal features into a trained modal detection and completion model for modal completion to obtain second molecular descriptor features, second molecular structure image features and second text features; fusing the second molecular descriptor feature, the second molecular structure image feature and the second text feature to obtain a fused feature; and inputting the fused features into the trained new pollutant identification model for new pollutant identification to obtain a classification result of the to-be-identified pollutants, and improving the accuracy of new pollutant type identification by fusing the structured data, the image data and the semantic text data.
Owner:XIANGJIANG LAB

Method for predicting flash point of pure compound by using multiple linear regression model

The invention provides a method for predicting a flash point of a pure compound by using a multiple linear regression model. According to the method, a mathematical model for predicting the flash point of the pure compound with high accuracy is established to predict a flash point value. Specifically, the model is used as a quantitative structure-property relationship model, after the optimal model is obtained from a plurality of multiple linear regression models by using a stepwise selection method for most compounds with known flash point experimental values, the molecular descriptors included in the model are received in a short time, and the flash point values are output. As long as the specific values of the molecule descriptors included in the model are known, the flash point of the compound composed of the molecule is predicted for any molecule. As described above, the present invention provides a method and model capable of predicting a reliable flash point value even for many compounds having unknown experimental values, thereby saving the cost and time for physical performance measurement experiments or chemical structure prediction, and enabling research and development activities of related industries to become easier.
Owner:BEIJING AISEN ZHONGKE TECHNOLOGY CO LTD

A high trophic level food chain biomagnification prediction model for organic chemicals

This invention discloses a high-trophic-level food chain biomagnification prediction model for organic chemicals. The model comprises the following steps: S1. Dividing the sample into six training samples and three validation samples; S2. Retaining a 1169-dimensional molecular descriptor matrix after screening; S3. Generating an extended pollutant parameter database by combining laboratory biomagnification factors; S4. Forming a position-time normalized sample dataset; S5. Obtaining a fused amplification probability result set; S6. Identifying a candidate set of pollutant amplification paths on the coupled directed graph of amplification probability-flux-stability; S7. Performing uncertainty propagation and confidence assessment on the candidate set of pollutant amplification paths, outputting a priority control list of pollutant amplification paths and their confidence information. This invention addresses the problems of insufficient model coverage and weak extrapolation ability caused by traditional methods relying solely on single molecular descriptors or experimental values.
Owner:NANJING INST OF ENVIRONMENTAL SCI MINIST OF ECOLOGY & ENVIRONMENT OF THE PEOPLES REPUBLIC OF CHINA

Method for predicting melting point of nitrate-based fused cast explosive

The invention discloses a melting point prediction method for a nitrate-based fused cast explosive. The method comprises the following steps: S1, constructing a data set matched with a nitrate-based energetic material; s2, constructing a user-defined molecular descriptor set of the melting points of the nitrate-based energetic materials; s3, based on the data set, optimizing the LightGBM, XGBoost and BPNN machine learning models by adopting an ablation experiment method and a training experiment, and obtaining an optimal model for predicting the melting point of the nitrate-based energetic material by comparing model evaluation indexes and errors; and S4, converting the newly designed nitrate-based energetic material I into a corresponding user-defined molecular descriptor set I, and taking the user-defined molecular descriptor set I as the input of the prediction machine learning model to obtain the output value of the melting point Tm. The method for predicting the melting point of the nitrate-based fused cast explosive, provided by the invention, has good prediction accuracy and interpretation capability, and can be used for predicting the performance of nitrate-based energetic materials.
Owner:NANJING UNIV OF SCI & TECH +1

A method and system for screening anthocyanin stabilizers based on the integration of quantum chemistry calculation and machine learning

The present application relates to the technical field of electric data processing, and provides a screening method and system for anthocyanin stabilizer based on quantum chemistry calculation and machine learning fusion. The method comprises the following steps: constructing a molecular structure dataset; performing geometric structure optimization on anthocyanin molecules and candidate polyphenol molecules to obtain ground state optimized structures; performing multidimensional molecular descriptor calculation on the candidate polyphenol molecules to obtain a polyphenol molecule descriptor matrix; performing structure optimization, single-point energy calculation and free energy calculation on the anthocyanin polyphenol complex system to obtain a binding free energy numerical set; integrating the polyphenol molecule descriptor matrix and the binding free energy numerical set, constructing a training dataset and performing feature screening; performing model training through a machine learning algorithm to obtain a binding free energy prediction model; and predicting the binding free energy of the anthocyanin polyphenol complex through the binding free energy prediction model to obtain candidate molecules for anthocyanin stabilizer. The present application improves the screening efficiency of anthocyanin stabilizer.
Owner:PEKING UNIV INST OF ADVANCED AGRI SCI +1

A method for predicting active structures of interfering biological pathways

ActiveCN116052772Bimprove data qualityImprove forecast resultsBiostatisticsHybridisationChemical compoundBiological pathway
The application discloses a method for predicting active structures interfering with biological pathways, which comprises the following steps: constructing a compound biological pathway interference database containing clear labels such as cell lines, exposure time and exposure concentration; evaluating the consistency of the biological pathway cross degree and the regulation trend of the training set and the test set through cumulative hypergeometric distribution and cumulative Bernoulli distribution; identifying a batch of potential compounds in the training set; evaluating the occurrence frequency of the molecular descriptors in the potential compounds through cumulative distribution probability; and finally predicting the potential active structures driving the change of the biological pathway through inputting the biological pathway.
Owner:NANJING UNIV

Method for predicting solubility parameter of pure compound by using multiple linear regression model

The invention discloses a method and a mathematical model for high-precision prediction of solubility parameters of a pure compound which is composed of 12 or less elements of hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus, arsenic and the like and has the atom number of 25 or less (excluding hydrogen). Wherein the model, as a universal quantitative structure-property relationship model, is obtained by determining an optimal model from a plurality of multiple linear regression models by means of a stepwise selection method, and the model takes some descriptors in various molecular descriptors as independent variables, takes solubility parameters as dependent variables, and takes the dissolvability parameters as dependent variables. The value of the molecule descriptor included in the model can be input in a short time and the solubility parameter can be output, so that the solubility parameter of a compound purely composed of the molecule can be predicted for any molecule as long as the specific value of the molecule descriptor included in the model is known.
Owner:BEIJING AISEN ZHONGKE TECHNOLOGY CO LTD

A method for qualitative prediction of the aroma compound enhancement effect

The application relates to a qualitative prediction method for the flavor-enhancing effect of aroma compounds, which comprises the following steps: constructing an aroma compound database; obtaining the 3D structure of umami receptor protein T1R1 / T1R3 from a protein molecule database to obtain a receptor protein dataset, obtaining the structure of aroma compounds and monosodium glutamate from a ZINC database as a ligand molecule dataset, then performing molecular docking, and saving the binding energy data record; taking the binding energy data corresponding to each aroma compound and the molecular descriptor as characteristic variables, and taking whether the aroma compound has the flavor-enhancing effect as a target variable to construct a logistic regression model; inputting aroma compound information into the constructed logistic regression model to obtain a qualitative output result of the prediction of the flavor-enhancing effect of the aroma compound. Compared with the prior art, the application predicts whether the aroma compound has the flavor-enhancing effect by constructing a logistic regression model by using the binding energy and the molecular descriptor and other properties, the method is simple and fast, the result is intuitive and reliable, and the method is widely applicable.
Owner:SHANGHAI INST OF TECH +2

A knowledge graph-based terpenoid compound and disease correlation prediction method and system

The present application relates to the technical field of artificial intelligence and pharmacology, and proposes a terpenoid compound and disease correlation prediction method and system based on a knowledge graph, comprising the following steps: collecting terpenoid compound and biological activity data, including terpenoid compound molecular structure, protein target, gene target, cell line and disease correlation information, and constructing a terpenoid compound biological activity knowledge graph; inputting the molecular descriptor of the terpenoid compound to be predicted into a convolutional neural network for feature extraction, to generate a first feature vector; performing feature extraction on the molecular structure of the terpenoid compound to be predicted through a graph convolution network, to generate a second feature vector; inputting the molecular embedding vector obtained by splicing the first feature vector and the second feature vector into a knowledge graph embedding model, to generate a terpenoid compound-disease prediction score matrix; and performing descending order sorting based on the terpenoid compound-disease prediction score, to generate a disease recommendation list related to the terpenoid compound.
Owner:SUN YAT SEN UNIV

Method for predicting standard boiling point of pure compound by using multiple linear regression model

Provided is a mathematical model with which it is possible to accurately predict the standard boiling point of a pure compound having 25 or less atoms (excluding hydrogen), said pure compound comprising 12 or less elements such as hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, iodine, silicon, phosphorus, and arsenic. Wherein the model, as a universal quantitative structure-property relationship model, is obtained by determining an optimal model from a plurality of multiple linear regression models by means of a stepwise selection method, and the model takes some of the various molecular descriptors as independent variables, takes a standard boiling point as a dependent variable, and takes the standard boiling point as a dependent variable; according to the present invention, the values of the molecule descriptors included in the model can be received and input in a short time and the standard boiling point can be output, so that the standard boiling point of the compound purely composed of the molecule can be predicted for any molecule satisfying the requirements of the present invention as long as the specific values of the molecule descriptors included in the model are known.
Owner:BEIJING AISEN ZHONGKE TECHNOLOGY CO LTD

LC-MS database of kadsura characteristic lignans and triterpenes as well as construction method and application of LC-MS database

The invention belongs to the field of biological medicine and analytical chemistry, and provides an LC-MS database of kadsura characteristic lignans and triterpenes as well as a construction method and application of the LC-MS database. The construction method comprises the following steps: collecting mass spectrum data of a sample with a reference substance, carrying out deconvolution, peak alignment and other treatment, matching and associating the molecular formula, molecular weight and other information of the reference substance with a secondary mass spectrum, and constructing a first database; the method comprises the following steps: converting a chemical structure of a schizandriaceae sample without a reference substance into a simplified molecular linear input specification format, calculating a molecular descriptor by utilizing an R language, analyzing and predicting retention time and a mass spectrum splitting decomposition mode through a preset prediction model, and constructing a second database. According to the application, actually measured data and a prediction model are integrated through double databases, so that the accuracy and comprehensiveness of kadsura ingredient identification are improved, the problem of isomeride identification and the problem of reference substance dependence are solved, a standardized analysis process is established, and key technical support is provided for quality control of medicinal materials, metabonomics research and modernization of traditional Chinese medicines.
Owner:HUNAN UNIV OF CHINESE MEDICINE

Prediction method of molecular frontline orbital energy level and related device

The invention provides a prediction method of molecular frontline orbit energy level and a related device, and relates to the technical field of chemistry. The prediction method for the molecular frontline orbit energy level can comprise the following steps: acquiring target information of a target molecule; wherein the target information comprises molecular structure information and a molecular descriptor; the molecular descriptors comprise a first molecular descriptor for predicting the highest occupied molecular orbital energy level and a second molecular descriptor for predicting the lowest unoccupied molecular orbital energy level; and inputting the target information into a pre-trained graph neural network model, and predicting the highest occupied molecular orbital energy level and the lowest unoccupied molecular orbital energy level of the target molecule through the graph neural network model. According to the technical scheme provided by the invention, the problem of low prediction efficiency of a scheme for predicting HOMO and LUMO energy level values through a DFT method in the prior art can be solved.
Owner:GUANGYIN (JIANGSU) NEW ENERGY CO LTD

A method and model for screening umami peptides

The present invention discloses a method and a screening model for umami peptides. The screening method comprises the following steps: S1: organizing existing umami peptide data and establishing a database; S2: constructing molecular fingerprint feature data of umami peptides based on the structural fragments of the existing umami peptides themselves; S3: constructing intermolecular interaction residue feature data based on the interaction mode between the umami peptides and the umami receptors T1R1 / T1R3 analyzed by molecular docking technology; S4: obtaining molecular descriptor feature data of the physicochemical properties of the umami peptides based on molecular descriptors; S5: using a machine learning algorithm to establish umami peptide screening sub-models for the data obtained in steps S2-S4; S6: integrating the three umami peptide screening sub-models using a support vector machine algorithm to establish an umami peptide screening model; and S7: screening umami peptides using the umami peptide screening model established in step S6. The screening method of the present invention can quickly and accurately screen umami peptides, and the screening method is reusable.
Owner:SHANGHAI JIAOTONG UNIV

Methods and systems for predicting the performance of polyimide, methods and systems for structural design, electronic devices, computer-readable storage media, and computer products.

This invention discloses a method and system for predicting the performance of polyimides, a method and system for structural design, an electronic device, a computer-readable storage medium, and a computer product. The method for predicting the performance of polyimides includes the following steps: inputting a first molecular descriptor into an AFP model to obtain the physical property parameter to be tested; the first molecular descriptor is the molecular descriptor of the polyimide to be tested; the AFP model is trained using a training dataset, which includes a second molecular descriptor and the first physical property parameter; the second molecular descriptor and the first physical property parameter are respectively the molecular descriptor and physical property parameter of a known polyimide; both the first and second molecular descriptors include node features and edge features, and the type of the first molecular descriptor is the same as the type of the second molecular descriptor; the type of the physical property parameter to be tested is the same as the type of the first physical property parameter. This invention has high accuracy and improves the effectiveness and reliability of de novo design.
Owner:EAST CHINA UNIV OF SCI & TECH

System and method for adjusting products to accommodate different regions using QSAR data

A method for predicting properties of a product having alternative ingredients, the method comprising: receiving data of an ingredient profile of a product, a molecular profile associated with the ingredient profile, and a proportioning profile of the product based on the ingredient profile; determining a set of molecular descriptors for the product based on the molecular profile; determining a set of material descriptors for the product based on the set of molecular descriptors; determining a set of proportioning descriptors for the product based on the set of material descriptors and the proportioning profile; generating a measured set of functional properties of the product using a machine learning model based on the one or more proportioning descriptors; receiving one or more alternative ingredients of the product; and generating a predicted set of functional properties of the product by using the machine learning model based on the one or more alternative components.
Owner:MARS INC

AIE fluorescent probe-oriented multi-index weighted scoring screening method and system

The invention provides an AIE fluorescent probe-oriented multi-index weighted scoring screening method and system. The method comprises the following steps: acquiring data such as a probe molecular structure and an experimental environment, and preprocessing the data to obtain a feature data set; for thermal stability, light stability, Stokes shift and AIE enhancement factors, optimal molecular descriptors and dichotomy prediction models are selected respectively; inputting a to-be-evaluated probe into the target model to obtain a prediction result, and calculating a synthetic accessibility score; combining the prediction result with the standardized synthetic accessibility score to form index data; the weight of each index is determined through mixed data factor analysis, and a comprehensive score is calculated to realize probe sorting and screening. According to the method, efficient screening of the AIE probe can be realized, the trial and error cost is reduced, and the scoring stability and applicability are enhanced.
Owner:CENT SOUTH UNIV

A method and device for screening potential substitutes of bisphenol a and a storage medium

The application relates to a bisphenol A potential substitute screening method, device and storage medium, the method comprising the following steps: acquiring a global AR antagonistic data set and a local BPA analogue data set; acquiring molecular characterization of the global and local BPA analogue data sets; acquiring a machine learning algorithm model; optimizing parameters of the machine learning algorithm model; processing the global AR antagonistic data set and the local BPA analogue data set by using the machine learning algorithm model; and acquiring a potential BPA substitute predicted by an optimal mixed model. The application provides a mixed deep learning architecture, which combines molecular descriptors and molecular graphs to predict the antagonistic activity of a compound on AR; compared with previous models, the mixed model can extract a large amount of chemical information from different molecular features, so that the generalization ability of the model for predicting BPA substitutes is improved; and the prediction result also shows that lignin derivatives are safer than bisphenol analogues as BPA substitutes.
Owner:JIANGHAN UNIVERSITY