Green efficient corrosion inhibitor confirmation method, device and equipment and readable storage medium
By constructing and fine-tuning the pre-trained chemical language model, the toxicity and corrosion inhibition efficiency of drug compounds are predicted, and green and efficient corrosion inhibitors are screened out, which solves the problems of environmental pollution and human harm in traditional corrosion inhibitors, and achieves efficient and environmentally friendly metal anti-corrosion effects.
Patent Information
- Application Number
- CN202510355923.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
AI Technical Summary
Existing corrosion inhibitors have problems of environmental pollution and human body harm in industrial applications, and plant extracts are difficult to be used on a large scale and cannot effectively solve the problem of metal corrosion.
By constructing a pre-trained chemical language model, using SMILES format drug compound data for fine-tuning, predicting the toxicity index LD50 value and corrosion inhibition efficiency, and screening compounds with low toxicity and high corrosion inhibition efficiency as green and high corrosion inhibition agents.
The rapid screening of high-efficiency green corrosion inhibitors has been achieved, which significantly improves the anti-corrosion effect of metals in industrial applications, while avoiding environmental pollution and human harm.
Smart Images

Figure CN120220893A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of chemoinformatics, and particularly to a method, device, equipment and readable storage medium for confirming a green and efficient corrosion inhibitor. Background Art
[0002] Metal corrosion is a long-standing problem in industrial production, seriously affecting the service life and safety of equipment. As an effective anti-corrosion means, corrosion inhibitors inhibit the occurrence of corrosion reactions by forming a protective film on the metal surface; however, most traditional corrosion inhibitors contain components harmful to the environment and human body, such as organophosphorus salts, chromates, etc., and these substances may cause eutrophication of water bodies or irreversible harm to organisms; with the improvement of environmental awareness, it has become an urgent task to develop green, efficient and non-toxic corrosion inhibitors. Although many studies have used plant extracts as corrosion inhibitors, plant extracts have disadvantages such as difficult extraction and being closely related to seasons, and cannot be used in large-scale industrial applications. Therefore, screening for highly efficient and low-toxicity corrosion inhibitors from discarded and expired drugs has become one of the research directions for green corrosion inhibitors. Summary of the Invention
[0003] To solve the technical problem of being unable to extract corrosion inhibitors on a large scale, the embodiments of this application provide a method, device, equipment and readable storage medium for confirming a green and efficient corrosion inhibitor.
[0004] In a first aspect, the embodiments of this application provide a method for confirming a green and efficient corrosion inhibitor, and the method includes: Obtain drug compound data in SMILES format, and construct a pre-trained chemical language model based on the drug compound data in SMILES format; Obtain new drug compound data in SMILES format and the corresponding toxicity index LD50 value, and perform a first preprocessing on it to obtain a first data set, and obtain a corrosion inhibitor data set, and perform a second preprocessing on the corrosion inhibitor data set to obtain a second data set; Fine-tune the pre-trained chemical language model according to the first data set to obtain a toxicity prediction model; Fine-tune the pre-trained chemical language model according to the second data set to obtain a corrosion inhibition efficiency prediction model; Predict the corresponding toxicity index LD50 value of the new drug compound data in SMILES format according to the toxicity prediction model, and screen the drug compound data with the toxicity index LD50 value greater than the toxicity threshold to obtain a candidate corrosion inhibitor molecule data set; Use the corrosion inhibition efficiency prediction model to predict the corrosion inhibition efficiency of the candidate corrosion inhibitor molecule data set, and screen for green and efficient corrosion inhibitors with a corrosion inhibition efficiency greater than the corrosion inhibition efficiency threshold according to the prediction result.
[0005] In one embodiment, constructing a pre-trained chemical language model based on the drug compound data in the SMILES format includes: Performing word segmentation on the drug compound data in the SMILES format according to a tokenizer to obtain a Token sequence; Mapping the Token sequence into a vector matrix through an embedding layer; Extracting a feature matrix of the vector matrix through a Transformer encoder; Randomly masking some Tokens of the Token sequence through a masked language model, and based on the feature matrix, outputting the predicted masked Tokens through an output layer to train and obtain the pre-trained chemical language model.
[0006] In one embodiment, fine-tuning the pre-trained chemical language model according to the first data set to obtain a toxicity prediction model includes: Dividing the first data set according to a preset ratio to obtain a first training set and a first validation set; Performing word segmentation on the first training set and the first validation set through a tokenizer, and extracting a feature matrix through an embedding layer and a Transformer encoder; Performing dimensionality reduction on the feature matrix through a fully connected layer to obtain a new feature matrix; Outputting the toxicity index LD50 value through a Dropout layer and a linear regression layer, and finally training and obtaining the toxicity prediction model.
[0007] In one embodiment, the corrosion inhibitor data set includes corrosion inhibitor molecules in the SMILES format, environmental data, and performance data. The environmental data includes the type of corrosion medium, the type of metal to be protected, the concentration of the corrosion medium, and the temperature of the corrosion inhibitor data set. The performance data includes the corrosion inhibition efficiency.
[0008] In one embodiment, fine-tuning the pre-trained chemical language model according to the second data set to obtain a corrosion inhibition efficiency prediction model includes: Dividing the second data set according to a preset ratio to obtain a second training set and a second validation set; Performing word segmentation on the second training set and the second validation set through a tokenizer, and extracting a feature matrix through an embedding layer and a Transformer encoder; An auxiliary feature matrix obtained by performing one-hot encoding on the type of corrosion medium and the type of metal to be protected; Concatenating the temperature of the corrosion inhibitor data set and the concentration of the corrosion medium with the feature matrix and the auxiliary feature matrix to obtain a fused feature matrix; The dimensionality reduction processing is performed on the fusion feature matrix through a fully connected layer, a Dropout layer, and a linear regression layer to output the corrosion inhibition efficiency, and finally the corrosion inhibition efficiency prediction model is trained.
[0009] In one embodiment, before screening the drug compound data with the LD50 value of the toxicity index greater than the toxicity threshold, the method further includes: Setting the classification rules for the LD50 toxicity index according to international standards; And setting the toxicity threshold according to the classification rules.
[0010] In one embodiment, obtaining the corrosion inhibitor data set and performing a second preprocessing on the corrosion inhibitor data set includes: Using the statistical method Z-score to obtain the outliers in the corrosion inhibitor data set and deleting the outliers; Deleting the duplicate values and missing values in the corrosion inhibitor data set; Performing normalization processing on the corrosion inhibitor data set after deleting outliers, duplicate values, and missing values.
[0011] In a second aspect, an embodiment of the present application provides a green and efficient corrosion inhibitor confirmation device, and the green and efficient corrosion inhibitor confirmation device includes: A construction module, configured to obtain drug compound data in SMILES format and construct a pre-trained chemical language model based on the drug compound data in SMILES format; An acquisition module, configured to obtain new drug compound data in SMILES format and the corresponding LD50 value of the toxicity index, perform a first preprocessing on them to obtain a first data set, and obtain a corrosion inhibitor data set and perform a second preprocessing on the corrosion inhibitor data set to obtain a second data set; A first fine-tuning module, configured to fine-tune the pre-trained chemical language model according to the first data set to obtain a toxicity prediction model; A second fine-tuning module, configured to fine-tune the pre-trained chemical language model according to the second data set to obtain a corrosion inhibition efficiency prediction model; A first screening module, configured to predict the corresponding LD50 value of the toxicity index for the new drug compound data in SMILES format according to the toxicity prediction model, and screen the drug compound data with the LD50 value of the toxicity index greater than the toxicity threshold to obtain a candidate corrosion inhibitor molecule data set; A second screening module, configured to use the corrosion inhibition efficiency prediction model to perform a corrosion inhibition efficiency prediction on the candidate corrosion inhibitor molecule data set, and screen green and efficient corrosion inhibitors with a corrosion inhibition efficiency greater than the corrosion inhibition efficiency threshold according to the prediction result.
[0012] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor. The memory is used to store a computer program, and the computer program executes the green and efficient corrosion inhibitor confirmation method provided in the first aspect when running on the processor.
[0013] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program executes the green and efficient corrosion inhibitor confirmation method provided in the first aspect when running on a processor.
[0014] For the green and efficient corrosion inhibitor confirmation method provided in the present application above, drug compound data in SMILES format is obtained, and a pre-trained chemical language model is constructed based on the drug compound data in SMILES format; new drug compound data in SMILES format and the corresponding toxicity index LD50 value are obtained, and a first preprocessing is performed on them to obtain a first data set, and a corrosion inhibitor data set is obtained, and a second preprocessing is performed on the corrosion inhibitor data set to obtain a second data set; according to the first data set, the pre-trained chemical language model is fine-tuned to obtain a toxicity prediction model; according to the second data set, the pre-trained chemical language model is fine-tuned to obtain a corrosion inhibition efficiency prediction model; according to the toxicity prediction model, the corresponding toxicity index LD50 value of new drug compound data in SMILES format is predicted, and the drug compound data with the toxicity index LD50 value greater than the toxicity threshold is screened to obtain a candidate corrosion inhibitor molecule data set; the corrosion inhibition efficiency prediction model is used to predict the corrosion inhibition efficiency of the candidate corrosion inhibitor molecule data set, and a green and efficient corrosion inhibitor with a corrosion inhibition efficiency greater than the corrosion inhibition efficiency threshold is screened according to the prediction result. The present application can significantly improve the screening efficiency of high-efficiency green corrosion inhibitors. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as limiting the protection scope of the present application. In each drawing, similar components are numbered similarly.
[0016] Figure 1 FIG. 1 shows a schematic flow chart of the green and efficient corrosion inhibitor confirmation method provided by an embodiment of the present application; Figure 2 FIG. 2 shows another schematic flow chart of the green and efficient corrosion inhibitor confirmation method provided by an embodiment of the present application; Figure 3 FIG. 3 shows yet another schematic flow chart of the green and efficient corrosion inhibitor confirmation method provided by an embodiment of the present application; Figure 4It shows still another schematic flowchart of the method for confirming a green and efficient corrosion inhibitor provided by the embodiments of the present application; Figure 5 It shows a schematic structural diagram of a green and efficient corrosion inhibitor confirmation device provided by the embodiments of the present application.
[0017] Icon: 500 - Green and efficient corrosion inhibitor confirmation device, 501 - Construction module, 502 - Acquisition module, 503 - First fine-tuning module, 504 - Second fine-tuning module, 505 - First screening module, 506 - Second screening module. Detailed implementation manners
[0018] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.
[0019] Generally, the components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application claimed, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0020] Hereinafter, the terms "including", "having" and their cognates that can be used in various embodiments of the present application are only intended to represent specific features, numbers, steps, operations, elements, components or combinations of the foregoing items, and should not be construed as first excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing items or increasing the possibility of one or more features, numbers, steps, operations, elements, components or combinations of the foregoing items.
[0021] In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0022] Unless otherwise limited, all terms (including technical terms and scientific terms) used here have the same meaning as those generally understood by those of ordinary skill in the art to which the various embodiments of the present application belong. The terms (such as those defined in a general-use dictionary) will be construed as having the same meaning as the contextual meaning in the relevant technical field and will not be construed as having an idealized meaning or being overly formal, unless clearly defined in the various embodiments of the present application.
[0023] Embodiment 1 The embodiments of the present application provide a method for confirming a green and efficient corrosion inhibitor.
[0024] See Figure 1 , the method for confirming a green and efficient corrosion inhibitor includes steps S101 - S106: S101: Obtain drug compound data in SMILES format, and construct a pre - trained chemical language model based on the drug compound data in SMILES format.
[0025] In this embodiment, the Simplified Molecular Input LineEntry System (SMILES) is a symbolic system for representing chemical molecular structures with strings. By using simple text strings to describe the atoms, bonds, and structural information of molecules, complex chemical structures can be represented in a single line of text. Among them, element symbols directly represent atoms, such as C represents carbon, O represents oxygen, and N represents nitrogen; for two - letter element symbols (such as chlorine Cl, bromine Br), they need to be enclosed in square brackets, such as [Cl], [Br]; hydrogen atoms can usually be omitted unless it is necessary to explicitly represent them. For the representation of bonds, single bonds are not shown by default and are directly represented by atom connections, such as CC represents ethane (two carbon atoms are connected by a single bond); double bonds are represented by =, such as C = C represents ethylene; triple bonds are represented by #, such as C#C represents acetylene; aromatic bonds (such as the bonds in a benzene ring) can be represented by using lowercase letters for atoms (such as c represents aromatic carbon), or by :. Specifically, for example, water in SMILES format: O; ethanol: CCO; benzene: c1ccccc1 or C1 = CC = CC = C1, etc.; acetic acid: CC(=O)O; glucose: C(C1C(C(C(C(O1)O)O)O)O)O.
[0026] Furthermore, learn the structural and property information of drug compound molecules from a large amount of SMILES data, extract the feature representations of the molecules, and use these feature representations for downstream tasks, such as predicting the corrosion inhibition performance of drug compounds, predicting drug toxicity, etc.
[0027] See Figure 2 , in one embodiment, step S101 includes steps S1011 - S1014: S1011: Perform word - segmentation processing on the drug compound data in SMILES format according to a word - segmenter to obtain a Token sequence.
[0028] In this embodiment, the main function of the tokenizer is to decompose the SMILES string into a series of Tokens (markers) so that the model can understand and process this chemical structure information. For example, the tokenizer defines patterns through regular expressions, and the patterns it defines are shown in the following formula: "(\[[^\]]+]|Br?|Cl?|N|O|S|P|F|I|b|c|n|o|H|s|p|\(|\)|\.|=|#||\+|\\\\|\ / |_|:|~|@|\?|>|\*|\$|\%[0-9]{2}|[0-9])", where \[[^\]]+]: matches the content inside the square brackets (such as [Na+], [Cl]); Br?|Cl?: matches bromine (Br) or chlorine (Cl); N|O|S|P|F|I: matches nitrogen (N), oxygen (O), sulfur (S), phosphorus (P), fluorine (F), iodine (I); b|c|n|o|H|s|p: matches atoms represented by lowercase letters (such as aromatic carbon c); \(|\)|\.|=|#|-|\+|\\\\|\ / |_|:|~|@|\?|>|\*|\$|\%[0-9]{2}: matches special symbols in SMILES (such as parentheses, bonds, ring markers, etc.); [0-9]: matches numbers (used to represent the closing points of rings). And a molecule start symbol (start) and a molecule end symbol (end) are respectively added at the beginning and end of the Token sequence to represent the start and end of the molecule. Optionally, for example, decompose the SMILES string (such as C1=CC=CC=C1) into a series of Tokens, ['start', 'C', '1', '=', 'C', 'C', '=', 'C', 'C', '=', 'C', '1', 'end'].
[0029] S1012: Map the Token sequence to a vector matrix through the embedding layer.
[0030] In this embodiment, the Token sequence of the SMILES string is mapped into a vector matrix through an embedding layer and used as the input of the model. It should be noted that Token (token / symbol), in natural language processing (NLP) and cheminformatics, refers to the smallest unit obtained by decomposing the input data (such as text or SMILES string). For example: for the English sentence "I love chemistry", the Token sequence may be ["I", "love", "chemistry"]. For the SMILES string "C1=CC=CC=C1", the Token sequence may be ["C", "1", "=", "C", "C", "=", "C", "C", "=", "C", "1"]. The Token sequence (such as ['start', 'C', '1', '=', 'C', 'end']), where each Token is mapped to a vector of a fixed length. Assume that the length of the Token sequence is L (such as L = 6), and the dimension of the embedding vector is D (such as D = 128), then the output of the embedding layer is a matrix of shape (L, D), where each row in the vector matrix corresponds to the vector representation of a Token.
[0031] S1013: Extract the feature matrix of the vector matrix through a Transformer encoder.
[0032] In this embodiment, the Transformer encoder extracts the feature representation of chemical molecules from the vector matrix output by the embedding layer through a multi-head self-attention mechanism and a feed-forward neural network. The newly extracted feature representation captures the context relationship between Tokens and provides a basis for the masked language model to predict the masked Tokens. For example, the input of the Transformer encoder is the vector matrix output by the embedding layer, with a shape of (L, D), where: L is the length of the Token sequence, D is the dimension of the embedding vector. Through the multi-head self-attention mechanism (Multi-Head Self-Attention), feed-forward neural network (Feed-Forward Network, FFN), and residual connection and layer normalization (Residual Connection & Layer Normalization) of the Transformer encoder, a feature matrix with a shape of (L, D') is obtained, where D' is the dimension of the output feature. S1014: Randomly mask some Tokens in the Token sequence through a masked language model, and based on the feature matrix, output the predicted masked Tokens through an output layer to train the pre-trained chemical language model.
[0033] In this embodiment, pre-training is performed by randomly masking some tokens with a masked language model (MLM) and predicting the masked tokens, and finally the predicted tokens are output through an output layer. The masked language model randomly selects 15% of the tokens for masking, where 80% of the probability is replaced with the special symbol [MASK], 10% of the probability is replaced with a random token, and 10% of the probability remains unchanged. The masked language model predicts the masked tokens according to the context.
[0034] Optionally, for example, for a SMILES string (such as C1=CC=CC=C1), the tokenizer decomposes the SMILES string into a token sequence (such as ['C', '1', '=', 'C', 'C', '=', 'C', 'C', '=', 'C', '1']), and 15% of the tokens are randomly selected for masking. For example: the original sequence: ['C', '1', '=', 'C', 'C', '=', 'C', 'C', '=', 'C', '1'], the masked sequence: ['C', '[MASK]', '=', 'C', 'C', '=', '[MASK]', 'C', '=', 'C', '1'], and the model predicts the masked tokens according to the context. For example: the predicted [MASK] positions are 1 and C.
[0035] Furthermore, the cross-entropy loss function is used to calculate the difference between the predicted tokens and the true tokens. For a single masked token, the cross-entropy loss function is expressed as formula (1): (1) Where is the total number of all tokens, is the one-hot encoding of the true token, that is, if the true token is the i-th token, then , otherwise . is the probability of the i-th token predicted by the model.
[0036] Where can be expressed by formula (2): (2) Where is the original output of the model for the i-th token, is for performing an exponential operation, Sum the exponential operation results of all Tokens for normalization.
[0037] Furthermore, for multiple masked Tokens in the entire sequence, the total loss Total Loss is the average of the cross-entropy losses of each masked Token, which is expressed as formula (3): (3) where N is the number of masked Tokens, is the true one-hot encoding of the j-th masked Token, is the probability of the i-th Token of the j-th masked Token predicted by the model.
[0038] S102: Obtain the new drug compound data in the SMILES format and the corresponding toxicity index LD50 value, and perform a first preprocessing on them to obtain a first data set, and obtain a corrosion inhibitor data set, and perform a second preprocessing on the corrosion inhibitor data set to obtain a second data set.
[0039] In this embodiment, obtain the drug compound data in the SMILES format and its corresponding toxicity index LD50 value from a public database (such as PubChem, ChEMBL, DrugBank, etc.). Among them, the LD50 value represents the dose that can cause half of the experimental animals (usually mice, rats or other experimental animals) to die within a specific time, and the unit is milligrams per kilogram (mg / kg). This dose can be administered orally, by injection or through other routes. The LD50 value is used to measure the acute toxicity of a chemical substance, that is, the degree of harm that may be caused to an organism when exposed to the substance in the short term. The higher the LD50 value, the lower the toxicity of the substance. Perform a first preprocessing on the corresponding data LD50 value with repeated, missing values and outliers, that is, delete the drug compound data in the corresponding SMILES format, and the outliers are identified by the statistical method Z-score.
[0040] Optionally, obtain the corrosion inhibitor data set from a literature and patent database, laboratory data or public database. The corrosion inhibitor data set includes the corrosion environment data and performance data of the corrosion inhibitor molecule. Delete the data in the corrosion inhibitor data set with repeated, missing values and outliers, and the outliers are also identified by the statistical method Z-score. Since there are differences in dimensions between the variables of the environmental data of the corrosion inhibitor data, in order to minimize the influence of dimensional differences on each input feature, first perform a normalization process on the samples to ensure the quality and consistency of the data. In one embodiment, the corrosion inhibitor dataset includes corrosion inhibitor molecules in SMILES format, environmental data, and performance data. The environmental data includes temperature, corrosion medium type, protected metal type, corrosion medium concentration, and corrosion inhibitor dataset temperature. The performance data includes corrosion inhibition efficiency.
[0041] In this embodiment, the environmental data includes corrosion inhibitor dataset temperature, specific corrosion medium, specific protected metal, and specific corrosion medium concentration. The performance data includes corrosion inhibition efficiency, which represents the impact of the organic corrosion inhibitor on the corrosion mitigation of the protected metal and is represented by a value from -100% to 100%. A negative value represents promoting corrosion, that is, the organic corrosion inhibitor does not have a corrosion inhibition effect. A positive value represents inhibiting corrosion. -100% represents complete promotion of corrosion, and 100% represents complete inhibition of corrosion. In one embodiment, to obtain the corrosion inhibitor dataset and perform a second preprocessing on the corrosion inhibitor dataset, the following steps are included: using the statistical method Z-score to obtain the outliers in the corrosion inhibitor dataset and deleting the outliers; deleting the duplicate values and missing values in the corrosion inhibitor dataset; and normalizing the corrosion inhibitor dataset after deleting the outliers, duplicate values, and missing values.
[0042] In this embodiment, after obtaining the corrosion inhibitor dataset, first check the corrosion inhibitor data, including temperature, corrosion medium type, protected metal type, corrosion medium concentration, corrosion inhibitor dataset temperature, and corrosion efficiency for verification. Delete any numerical values with duplicates, missing values, and outliers and their corresponding corrosion inhibitor molecule data in SMILES format. The outliers are identified using the statistical method Z-score. Since there are differences in dimensions between the variables of the environmental data of the corrosion inhibitor data, to minimize the impact of dimensional differences on each input feature, after performing the screening actions of statistics on duplicates, missing values, and outliers on the data, normalize the dataset samples to ensure the quality and consistency of the data. The operation is as shown in formula (4): (4) In formula (4), the data matrix is , which includes n sample data and j environmental features. represents the environmental data. is the environmental data after standardization processing. and are the mean and standard deviation of the j-th column.
[0043] S103: Fine-tune the pre-trained chemical language model according to the first dataset to obtain a toxicity prediction model.
[0044] In this embodiment, the pre-trained chemical language model is adapted from a general chemical structure representation task to a toxicity prediction task, and the output of the model is adjusted from Token prediction to continuous LD50 value prediction, thereby obtaining a toxicity prediction model.
[0045] See Figure 3 , in one embodiment, step S103 includes steps S1031 - S1034.
[0046] S1031: Divide the first data set according to a preset ratio to obtain a first training set and a first validation set.
[0047] In this embodiment, after obtaining the drug toxicity data set (the first data set) for model training, the drug toxicity data is divided into a first training set , a first validation set and a first test set according to the ratio of 80:10:10 respectively, where 80% of the data is used for model training, 10% of the data is used for model tuning and hyperparameter selection, and 10% of the data is used for final model performance evaluation.
[0048] S1032: Tokenize the first training set and the first validation set through a tokenizer, and extract a feature matrix through an embedding layer and a Transformer encoder.
[0049] In this embodiment, the steps of S1032 are the same as the technical steps of S1011 - S1013, and will not be elaborated here.
[0050] S1033: Perform dimensionality reduction processing on the feature matrix through a fully connected layer to obtain a new feature matrix.
[0051] In this embodiment, based on the parameters of the pre-trained chemical language model, step S1032 will obtain a feature matrix with richer features, and then perform a non-linear transformation and dimensionality reduction on the feature matrix output by the Transformer encoder through a fully connected layer to extract high-level feature representations.
[0052] S1034: Output the toxicity index LD50 value through a Dropout layer and a linear regression layer, and finally train to obtain the toxicity prediction model.
[0053] In this embodiment, according to the feature matrix, model fine-tuning is performed on the drug toxicity data. Model fine-tuning means adjusting the output layer of the pre-trained chemical language model to a linear regression layer so that it can output the toxicity index LD50 value. Then the fine-tuned model is expressed as , where Reg represents the linear regression function, It represents a new feature matrix obtained by performing dimensionality reduction and non-linear transformation processing on the fully connected layer after obtaining the feature representation based on the Transformer encoder in the pre-trained chemical language model.
[0054] Furthermore, the feature output obtained through the Dropout layer, where the proportion of Dropout is usually set between 0.2 and 0.5. To prevent the model from overfitting, the model loss function is the mean squared error function with a regularization term added, and the regularization coefficient λ is determined through cross-validation. To obtain the best model performance, the Adam optimizer is used, the initial learning rate is set between 1e-4 and 1e-5, and the learning rate decay strategy is used to ensure that the learning rate is reduced when the performance on the validation set no longer improves. After passing through the Dropout layer, the LD50 value of the toxicity index is output through the linear regression layer.
[0055] S104: According to the second data set and set environmental data, fine-tune the pre-trained chemical language model to obtain a corrosion inhibition efficiency prediction model.
[0056] In this embodiment, based on the pre-trained chemical language model, the model is fine-tuned for corrosion inhibitor data to obtain a corrosion inhibition efficiency prediction model. The model fine-tuning includes adjusting the output layer of the pre-trained chemical language model to a linear regression layer so that it can output the corrosion inhibition efficiency value; it also includes embedding feature data such as the temperature of the corrosion inhibitor data set, the corrosion medium, the metal to be protected, the concentration of the corrosion medium, and the concentration of the corrosion inhibitor into the model input to enhance the model's understanding of these features. That is, the input of the fine-tuned model is the fusion feature representation obtained by splicing the above features with the feature representation embedding obtained based on the Transformer encoder in the pre-trained chemical language model.
[0057] See Figure 4 In one embodiment, the step S104 includes steps S1041-S1045.
[0058] S1041: Divide the second data set according to a preset ratio to obtain a second training set and a second validation set.
[0059] In this embodiment, after obtaining the corrosion inhibitor data set for model training, the corrosion inhibitor data is also divided into a second training set and a second validation set and a second test set according to the ratio 80:10:10 in S1031.
[0060] S1042: Perform word segmentation processing on the second training set and the second validation set through a tokenizer, and extract the feature matrix through the embedding layer and the Transformer encoder.
[0061] In this embodiment, the same technical features as those in steps S1011 - S1013 are used. The SMILES string is decomposed into a Token sequence by a tokenizer, the Token sequence is mapped into a vector representation, and a pre-trained Transformer encoder is used to extract a feature matrix.
[0062] S1043: An auxiliary feature matrix obtained by performing one-hot encoding on the corrosion medium type and the protected metal type.
[0063] In this embodiment, the categorical features in the environmental data (such as the corrosion medium type and the protected metal type) are encoded by One-Hot, each category is mapped to a unique binary vector, and the binary vectors corresponding to each category are combined to obtain an auxiliary feature matrix.
[0064] S1044: Concatenate the temperature of the corrosion inhibitor dataset and the corrosion medium concentration with the feature matrix and the auxiliary feature matrix to obtain a fused feature matrix.
[0065] In this embodiment, the numerical features (such as the temperature of the corrosion inhibitor dataset and the corrosion medium concentration) are standardized to eliminate the dimension difference, and then the feature matrix extracted by the Transformer encoder, the auxiliary feature matrix encoded by One-Hot, and the numerical features are concatenated together to form a fused feature matrix.
[0066] S1045: Perform dimensionality reduction on the fused feature matrix through a fully connected layer, a Dropout layer, and a linear regression layer, output the corrosion inhibition efficiency, and finally train the corrosion inhibition efficiency prediction model.
[0067] In this embodiment, the fused feature matrix passes through one or more fully connected layers to obtain a further feature representation, and finally passes through the Dropout layer to obtain the feature input of the linear regression layer. Finally, the corrosion inhibition efficiency is output through the linear regression layer. The proportion of Dropout is usually set between 0.2 and 0.5. To prevent overfitting of the model, the model loss function is the mean squared error function with a regularization term added, and the regularization coefficient λ is determined by cross-validation. To obtain the best model performance, the Adam optimizer is used, the initial learning rate is set between 1e-4 and 1e-5, and the learning rate decay strategy is used to ensure that the learning rate is reduced when the performance on the validation set no longer improves.
[0068] S105: Predict the corresponding toxicity index LD50 value for the new drug compound data in SMILES format according to the toxicity prediction model, and screen the drug compound data with the toxicity index LD50 value greater than the toxicity threshold to obtain a candidate corrosion inhibitor molecular dataset.
[0069] In this embodiment, after obtaining the trained toxicity prediction model, using the collected drug molecule data in SMILES format as input, first, based on the trained toxicity prediction model, predict the LD50 value of the toxicity index of the drug molecule, and screen out low-toxicity drug molecules as the candidate corrosion inhibitor molecule dataset.
[0070] In one embodiment, before the step S105, the method further includes: setting the classification rules for the LD50 toxicity index according to international standards; and setting the toxicity threshold according to the classification rules.
[0071] In this embodiment, according to international standards, toxicity classification is performed based on the LD50 value. Compounds with an LD50 value greater than 2000 mg / kg are classified into the fifth category, compounds with an LD50 value greater than 300 and less than or equal to 2000 mg / kg are classified into the fourth category, compounds with an LD50 value greater than 50 and less than or equal to 300 mg / kg are classified into the third category, compounds with an LD50 value greater than 5 and less than or equal to 50 mg / kg are classified into the second category, and compounds with an LD50 value less than or equal to 5 mg / kg are classified into the first category. Therefore, screen the drugs with a predicted LD50 value greater than 2000 mg / kg, and finally use the screening result as the candidate corrosion inhibitor molecule dataset.
[0072] S106: Based on the environmental data and using the corrosion inhibition efficiency prediction model, predict the corrosion inhibition efficiency of the candidate corrosion inhibitor molecule dataset, and screen out green and highly efficient corrosion inhibitors with a corrosion inhibition efficiency greater than the corrosion inhibition efficiency threshold.
[0073] In this embodiment, based on the candidate corrosion inhibitor molecule dataset, simultaneously set specific environmental data (temperature, corrosive medium, protected metal, corrosive medium concentration, corrosion inhibitor concentration) as the model input, use the corrosion inhibition efficiency prediction model to predict the corrosion inhibition efficiency of the candidate corrosion inhibitor molecule dataset, and screen out molecules with high corrosion inhibition efficiency (corrosion inhibition efficiency greater than 90%) as the molecule dataset of the green and highly efficient corrosion inhibitors after joint screening. That is, low-toxicity and high-corrosion-inhibition-efficiency compounds, namely highly efficient green corrosion inhibitors, can be obtained from the molecule dataset of the green and highly efficient corrosion inhibitors.
[0074] The method for confirming a green and efficient corrosion inhibitor provided in this embodiment obtains drug compound data in SMILES format and constructs a pre-trained chemical language model based on the drug compound data in SMILES format; obtains the new drug compound data in SMILES format and the corresponding toxicity index LD50 value, and performs a first preprocessing on them to obtain a first data set, and obtains a corrosion inhibitor data set, and performs a second preprocessing on the corrosion inhibitor data set to obtain a second data set; fine-tunes the pre-trained chemical language model according to the first data set to obtain a toxicity prediction model; fine-tunes the pre-trained chemical language model according to the second data set to obtain a corrosion inhibition efficiency prediction model; predicts the corresponding toxicity index LD50 value of the new drug compound data in SMILES format according to the toxicity prediction model, and screens the drug compound data with the toxicity index LD50 value greater than the toxicity threshold to obtain a candidate corrosion inhibitor molecule data set; uses the corrosion inhibition efficiency prediction model to predict the corrosion inhibition efficiency of the candidate corrosion inhibitor molecule data set, and screens out green and efficient corrosion inhibitors with a corrosion inhibition efficiency greater than the corrosion inhibition efficiency threshold according to the prediction result. This application can significantly improve the screening efficiency of high-efficiency green corrosion inhibitors.
[0075] Example 2 In addition, an embodiment of this application provides a device for confirming a green and efficient corrosion inhibitor, which is applied to an electronic device.
[0076] As Figure 5 shown, the device 500 for confirming a green and efficient corrosion inhibitor includes: A construction module 501, configured to obtain drug compound data in SMILES format and construct a pre-trained chemical language model based on the drug compound data in SMILES format; An acquisition module 502, configured to obtain the new drug compound data in SMILES format and the corresponding toxicity index LD50 value, and perform a first preprocessing on them to obtain a first data set, and obtain a corrosion inhibitor data set, and perform a second preprocessing on the corrosion inhibitor data set to obtain a second data set; A first fine-tuning module 503, configured to fine-tune the pre-trained chemical language model according to the first data set to obtain a toxicity prediction model; A second fine-tuning module 504, configured to fine-tune the pre-trained chemical language model according to the second data set to obtain a corrosion inhibition efficiency prediction model; A first screening module 505, configured to predict the corresponding toxicity index LD50 value of the new drug compound data in SMILES format according to the toxicity prediction model, and screen the drug compound data with the toxicity index LD50 value greater than the toxicity threshold to obtain a candidate corrosion inhibitor molecule data set; The second screening module 506 is used to predict the corrosion inhibition efficiency of the candidate corrosion inhibitor molecule dataset by using the corrosion inhibition efficiency prediction model, and screen out green and highly efficient corrosion inhibitors with a corrosion inhibition efficiency greater than the corrosion inhibition efficiency threshold according to the prediction results.
[0077] Optionally, the construction module 501 is further configured to tokenize the drug compound data in the SMILES format according to a tokenizer to obtain a Token sequence; map the Token sequence into a vector matrix through an embedding layer; extract a feature matrix of the vector matrix through a Transformer encoder; randomly mask some Tokens of the Token sequence through a masked language model, and output the predicted masked Tokens based on the feature matrix through an output layer to train the pre-trained chemical language model.
[0078] Optionally, the acquisition module 502 is further configured to obtain outliers in the corrosion inhibitor dataset by using the statistical method Z-score and delete the outliers; delete duplicate values and missing values in the corrosion inhibitor dataset; and perform normalization processing on the corrosion inhibitor dataset after deleting outliers, duplicate values and missing values.
[0079] Optionally, the first fine-tuning module 503 is further configured to divide the first dataset according to a preset ratio to obtain a first training set and a first validation set; tokenize the first training set and the first validation set through a tokenizer, and extract a feature matrix through an embedding layer and a Transformer encoder; perform dimensionality reduction processing on the feature matrix through a fully connected layer to obtain a new feature matrix; output the toxicity index LD50 value through a Dropout layer and a linear regression layer, and finally train the toxicity prediction model.
[0080] Optionally, the second fine-tuning module 504 is further configured to divide the second dataset according to a preset ratio to obtain a second training set and a second validation set; tokenize the second training set and the second validation set through a tokenizer, and extract a feature matrix through an embedding layer and a Transformer encoder; obtain an auxiliary feature matrix by performing one-hot encoding on the corrosion medium type and the protected metal type; splice the temperature of the corrosion inhibitor dataset and the corrosion medium concentration with the feature matrix and the auxiliary feature matrix to obtain a fused feature matrix; perform dimensionality reduction processing on the fused feature matrix through a fully connected layer, a Dropout layer and a linear regression layer, and output the corrosion inhibition efficiency, and finally train the corrosion inhibition efficiency prediction model.
[0081] Optionally, the first screening module 505 is further configured to set the classification rules for the LD50 toxicity index according to international standards; and set the toxicity threshold according to the classification rules.
[0082] The green and efficient corrosion inhibitor confirmation device 500 provided in this embodiment can implement the green and efficient corrosion inhibitor confirmation method provided in Embodiment 1. To avoid repetition, it will not be elaborated here.
[0083] The green and efficient corrosion inhibitor confirmation device provided in this embodiment obtains drug compound data in SMILES format and constructs a pre-trained chemical language model based on the drug compound data in SMILES format; obtains new drug compound data in SMILES format and the corresponding toxicity index LD50 value, and performs a first preprocessing on it to obtain a first data set, and obtains a corrosion inhibitor data set, and performs a second preprocessing on the corrosion inhibitor data set to obtain a second data set; according to the first data set, fine-tune the pre-trained chemical language model to obtain a toxicity prediction model; according to the second data set, fine-tune the pre-trained chemical language model to obtain a corrosion inhibition efficiency prediction model; according to the toxicity prediction model, predict the corresponding toxicity index LD50 value for the new drug compound data in SMILES format, and screen the drug compound data with the toxicity index LD50 value greater than the toxicity threshold to obtain a candidate corrosion inhibitor molecule data set; use the corrosion inhibition efficiency prediction model to predict the corrosion inhibition efficiency of the candidate corrosion inhibitor molecule data set, and screen out green and efficient corrosion inhibitors with a corrosion inhibition efficiency greater than the corrosion inhibition efficiency threshold according to the prediction result. This application can significantly improve the screening efficiency of high-efficiency green corrosion inhibitors.
[0084] Embodiment 3 In addition, an embodiment of the present application provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the computer program runs on the processor, it executes the green and efficient corrosion inhibitor confirmation method provided in Embodiment 1.
[0085] The electronic device provided in the embodiment of the present invention can execute the steps of the green and efficient corrosion inhibitor confirmation method provided in the above Method Embodiment 1. To avoid repetition, it will not be elaborated here.
[0086] Embodiment 4 The present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the green and efficient corrosion inhibitor confirmation method provided in Embodiment 1.
[0087] In this embodiment, the computer-readable storage medium can be a read-only memory (ROM for short), a random access memory (RAM for short), a magnetic disk, or an optical disc, etc.
[0088] The computer-readable storage medium provided in this embodiment can implement the green and efficient corrosion inhibitor confirmation method provided in Embodiment 1. To avoid repetition, it will not be elaborated here.
[0089] It should be noted that in this article, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or terminal. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or terminal including the element.
[0090] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of this application.
[0091] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of this application, those of ordinary skill in the art can also make many forms without departing from the purpose of this application and the scope protected by the claims, and all of them belong to the protection scope of this application.
Claims
1. A method for confirming a green and efficient corrosion inhibitor, characterized in that: The method comprises: Acquire drug compound data in SMILES format, and build a pre-trained chemical language model based on the drug compound data in SMILES format; Acquire the new drug compound data in the SMILES format and the corresponding toxicity index LD50 value, and perform a first preprocessing on the data to obtain a first data set, and acquire a corrosion inhibitor data set, and perform a second preprocessing on the corrosion inhibitor data set to obtain a second data set; Fine-tuning the pre-trained chemical language model according to the first data set to obtain a toxicity prediction model; According to the second data set, fine-tuning the pre-trained chemical language model to obtain a corrosion inhibition efficiency prediction model; Predicting the corresponding toxicity index LD50 value for the drug compound data in the new SMILES format according to the toxicity prediction model, screening the drug compound data whose toxicity index LD50 value is greater than the toxicity threshold, and obtaining a candidate corrosion inhibitor molecule data set; The corrosion inhibition efficiency prediction model is used to predict the corrosion inhibition efficiency of the candidate corrosion inhibitor molecule data set, and green and efficient corrosion inhibitors with a corrosion inhibition efficiency greater than a corrosion inhibition efficiency threshold are screened according to the prediction results.
2. The method according to claim 1, characterized in that The method of constructing a pre-trained chemical language model based on the drug compound data in the SMILES format includes: The drug compound data in the SMILES format is segmented according to a tokenizer to obtain a Token sequence; Mapping the Token sequence into a vector matrix through an embedding layer; Extracting a feature matrix of the vector matrix through a Transformer encoder; Part of the tokens of the token sequence are randomly masked by a mask language model, and based on the feature matrix, the predicted masked tokens are output through an output layer to train and obtain the pre-trained chemical language model.
3. The method according to claim 1, characterized in that The method of fine-tuning the pre-trained chemical language model according to the first data set to obtain a toxicity prediction model includes: Dividing the first data set according to a preset ratio to obtain a first training set and a first validation set; Performing word segmentation on the first training set and the first validation set through a word segmenter, and extracting a feature matrix through an embedding layer and a Transformer encoder; Performing dimensionality reduction processing on the feature matrix through a fully connected layer to obtain a new feature matrix; The toxicity indicator LD50 value is output through the Dropout layer and the linear regression layer, and the toxicity prediction model is finally obtained through training.
4. The method according to claim 1, characterized in that: The corrosion inhibitor data set includes corrosion inhibitor molecules in SMILES format, environmental data and performance data, the environmental data includes the type of corrosive medium, the type of protected metal, the concentration of the corrosive medium, and the temperature of the corrosion inhibitor data set, and the performance data includes the corrosion inhibition efficiency.
5. The method according to claim 4, characterized in that The method of fine-tuning the pre-trained chemical language model according to the second data set to obtain a corrosion inhibition efficiency prediction model includes: Dividing the second data set according to a preset ratio to obtain a second training set and a second validation set; Performing word segmentation on the second training set and the second validation set through a word segmenter, and extracting a feature matrix through an embedding layer and a Transformer encoder; An auxiliary feature matrix obtained by one-hot encoding the type of the corrosive medium and the type of the protected metal; The temperature of the corrosion inhibitor data set and the concentration of the corrosive medium are spliced with the feature matrix and the auxiliary feature matrix to obtain a fused feature matrix; The fusion feature matrix is subjected to dimension reduction processing through a fully connected layer, a Dropout layer, and a linear regression layer, and the corrosion inhibition efficiency is output, and finally the corrosion inhibition efficiency prediction model is trained.
6. The method according to claim 1, characterized in that Before screening the drug compound data whose toxicity index LD50 value is greater than the toxicity threshold, the method further comprises: Set classification rules for LD50 toxicity index according to international standards; And the toxicity threshold is set according to the classification rule.
7. The method according to claim 1, characterized in that The step of obtaining a corrosion inhibitor data set and performing a second preprocessing on the corrosion inhibitor data set includes: Using the statistical method Z-score to obtain outliers in the corrosion inhibitor data set, and deleting the outliers; Deleting duplicate values and missing values in the corrosion inhibitor data set; The corrosion inhibitor dataset was normalized to remove outliers, duplicate values, and missing values.
8. A green and efficient corrosion inhibitor confirmation device, characterized in that: The device comprises: A construction module is used to obtain drug compound data in SMILES format and construct a pre-trained chemical language model based on the drug compound data in SMILES format; An acquisition module, used to acquire the new drug compound data in the SMILES format and the corresponding toxicity index LD50 value, and perform a first preprocessing on the data to obtain a first data set, and to acquire a corrosion inhibitor data set, and perform a second preprocessing on the corrosion inhibitor data set to obtain a second data set; A first fine-tuning module, configured to fine-tune the pre-trained chemical language model according to the first data set to obtain a toxicity prediction model; A second fine-tuning module, used for fine-tuning the pre-trained chemical language model according to the second data set to obtain a corrosion inhibition efficiency prediction model; The first screening module is used to predict the corresponding toxicity index LD50 value for the drug compound data in the new SMILES format according to the toxicity prediction model, screen the drug compound data whose toxicity index LD50 value is greater than the toxicity threshold, and obtain the candidate corrosion inhibitor molecule data set; The second screening module is used to use the corrosion inhibition efficiency prediction model to predict the corrosion inhibition efficiency of the candidate corrosion inhibitor molecular data set, and screen green and efficient corrosion inhibitors whose corrosion inhibition efficiency is greater than the corrosion inhibition efficiency threshold according to the prediction results.
9. An electronic device, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program, and the computer program executes the green and efficient corrosion inhibitor confirmation method according to any one of claims 1 to 7 when the processor is running.
10. A computer-readable storage medium, characterized in that: It stores a computer program, which, when running on a processor, executes the green and efficient corrosion inhibitor confirmation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Compound property prediction model training method and compound property prediction method
CN116612835A
Molecular property prediction method based on molecule Image and SMILES character string pre-training
CN117238394A
Chemically modified small nucleic acid drug zero sample prediction method and system based on large language model and multi-view deep learning
CN119152929A
Amorphous forming ability prediction method based on Transform and table data conversion
CN119153001A
Methods and systems for screening and design of corrosion inhibitors
US20230282315A1