Prediction Model and Construction Method, Prediction Method and Device, Electronic Device
By constructing a prediction model of multiple sub-models, the animal welfare and accuracy problems in the prediction of eye corrosion or irritation of compounds is solved, and high accuracy prediction without increasing resources and ensuring animal welfare is achieved.
Patent Information
- Application Number
- CN202210305813.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-03-25
AI Technical Summary
The prior art has problems with animal welfare and poor prediction accuracy when predicting the corrosion or irritation of the compound to the eyes, especially in the case of small amounts of data.
By constructing a prediction model of a multi-sub model, different sub-models are trained using different compound samples, and the final prediction model is determined based on multiple sub-models to improve the prediction accuracy of the safety probability of the compounds to be tested.
Without increasing human, material and financial resources and ensuring animal welfare, the accuracy of predicting compounds for eye corrosion or irritation is improved, and new compounds can be effectively predicted.
Smart Images

Figure CN114649064B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of predicting compound toxicity, and particularly to a prediction model and a construction method, a prediction method and a device, and an electronic device. Background Art
[0002] The corneal and conjunctival tissues are directly exposed to the air and are vulnerable to the influence of chemicals, and thus are prone to eye corrosion or irritation due to various substances, such as chemicals used in manufacturing, agriculture, cosmetics, and ophthalmic drugs. In daily life, many people are easily exposed to chemicals that cause eye corrosion or irritation. Therefore, eye corrosion and eye irritation are important research topics in human health and should be considered in chemical hazard and risk assessment.
[0003] Currently, the research on the possibility of eye corrosion or irritation of compounds is carried out through animal test methods, such as the rabbit eye test method (OECD Test Guideline No. 405), the bovine corneal opacity and permeability test method (OECD Test Guideline No. 437), and the ex vivo chicken eye test method (OECD Test Guideline No. 438). In recent years, in order to reduce the number and suffering of experimental animals, scientists have studied alternative methods for animal tests, such as the fluorescein leakage test method (OECD Test Guideline No. 460), the reconstructed human corneal epithelium test method (OECD Test Guideline No. 492), and the in vitro macromolecule test method (OECD Test Guideline No. 496). However, each of the above methods has its own limitations, and moreover, it has a great impact on animal welfare. Therefore, scientific research personnel recommend using a combination of multiple non-animal methods, such as the integrated testing and assessment method (OECD Series on Testing and Assessment No. 255). Among them, the integrated testing and assessment method can include different types of test methods (such as chemical, in vitro, and in vivo tests), non-test methods (such as computational modeling), or data integration methods (such as integrated testing strategies, sequential testing strategies, weight of evidence). However, the amount of data on eye corrosion and eye irritation is small, resulting in poor accuracy after computational modeling. If a large number of labeled experiments are carried out to increase the amount of data, there are still problems such as affecting animal welfare and wasting human, material, and financial resources. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a prediction model and a construction method, a prediction method and a device, and an electronic device, which can improve the prediction accuracy while ensuring animal welfare and without increasing human, material, and financial resources.
[0005] In a first aspect, an embodiment of the present application provides a method for constructing a prediction model, including:
[0006] Obtain a training sample set, where the training sample set includes N compound samples;
[0007] Train a first sub-model through the N compound samples, determine a fourth sub-model through the first sub-model, and train a second sub-model through M compound samples among the N compound samples;
[0008] Use the second sub-model to screen out target compound samples from Q compound samples among the N compound samples;
[0009] Train a third sub-model through the target compound samples;
[0010] Determine a prediction model based on the first sub-model, the second sub-model, the third sub-model, and the fourth sub-model.
[0011] In a possible implementation manner, the determining the fourth sub-model through the first sub-model includes:
[0012] According to a combination rule, combine two or more of the multiple first sub-models to obtain a combined sub-model;
[0013] Determine the fourth sub-model based on a preset algorithm and the combined sub-model.
[0014] In a possible implementation manner, the using the second sub-model to screen out target compound samples from Q compound samples among the N compound samples includes:
[0015] Input the Q compound samples among the N compound samples into the second sub-model in sequence;
[0016] Obtain the actual result corresponding to each compound sample;
[0017] Determine the target compound samples based on the actual result corresponding to each compound sample and a reference value.
[0018] In a possible implementation manner, the determining the target compound samples based on the actual result corresponding to each compound sample and a reference value includes:
[0019] For each compound sample, calculate the difference between its actual result and the reference value;
[0020] If the difference is less than a preset threshold, determine the compound sample as the target compound sample.
[0021] In a possible implementation, each of the first sub-model, the second sub-model, and the third sub-model can be one or more, and the numbers of the first sub-model, the second sub-model, and the third sub-model are the same.
[0022] In a possible implementation, determining the prediction model based on the first sub-model, the second sub-model, the third sub-model, and the fourth sub-model includes:
[0023] Determining a first accuracy rate of the first sub-model, a second accuracy rate of the second sub-model, a third accuracy rate of the third sub-model, and a fourth accuracy rate corresponding to the fourth sub-model;
[0024] Determining the prediction model based on the first accuracy rate, the second accuracy rate, the third accuracy rate, and the fourth accuracy rate.
[0025] In a second aspect, an embodiment of the present application further provides a prediction model constructed by any of the above construction methods.
[0026] In a third aspect, an embodiment of the present application further provides a prediction method, including:
[0027] Processing the compound to be measured by using the above prediction model to obtain the safety probability of the compound to be measured; wherein, the compound to be measured is the same as or different from the compound samples included in the training sample set.
[0028] In a fourth aspect, an embodiment of the present application further provides a prediction device, including:
[0029] The prediction model described in the second aspect;
[0030] A processing module configured to: process the compound to be measured by using the prediction model to obtain the safety probability of the compound to be measured; wherein, the compound to be measured is the same as or different from the N compound samples included in the training sample set.
[0031] In a fifth aspect, an embodiment of the present application further provides an electronic device, including: a processor and a memory, the memory stores machine-readable instructions executable by the processor, when the electronic device runs, communication between the processor and the memory is through a bus, and when the machine-readable instructions are executed by the processor, the following steps are performed:
[0032] Processing the compound to be measured by using the pre-trained prediction model to obtain the safety probability of the compound to be measured; wherein, the compound to be measured is the same as or different from the compound samples included in the training sample set.
[0033] In the embodiments of the present application, different sub-models are trained using different compound samples, and then a prediction model is determined based on multiple sub-models to calculate the safety probability of a compound to be measured through the prediction model. Without increasing human, material, and financial resources while ensuring animal welfare, the purpose of improving prediction accuracy is achieved. Moreover, the prediction model can effectively predict new compounds. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments described in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0035] Figure 1 The flowchart showing a method for constructing a prediction model provided by the present application is shown;
[0036] Figure 2 The flowchart showing the screening of target compound samples in a method for constructing a prediction model provided by the present application is shown;
[0037] Figure 3 The flowchart showing the determination of the prediction model in a method for constructing a prediction model provided by the present application is shown;
[0038] Figure 4 The structural schematic diagram of a prediction device provided by the present application is shown;
[0039] Figure 5 The structural schematic diagram of an electronic device provided by the present application is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] Reference is made herein to the various aspects and features of the present application with reference to the drawings.
[0041] It should be understood that various modifications can be made to the embodiments described herein. Therefore, the above description should not be considered restrictive, but merely as an example of the embodiments. Those skilled in the art will envision other modifications within the scope and spirit of the present application.
[0042] The drawings included in the specification and forming a part of the specification illustrate the embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, are used to explain the principles of the present application.
[0043] These and other features of the present application will become apparent from the following description of the preferred forms of the embodiments given by way of non-limiting example with reference to the drawings.
[0044] It should also be understood that, although the present application has been described with reference to some specific examples, those skilled in the art can surely implement many other equivalent forms of the present application, which have the features as described in the claims and thus are all within the protection scope defined hereby.
[0045] When combined with the accompanying drawings, the above and other aspects, features and advantages of the present application will become more obvious in view of the following detailed description.
[0046] Specific embodiments of the present application will be described hereinafter with reference to the accompanying drawings; however, it should be understood that the embodiments claimed are merely examples of the present application, which can be implemented in various ways. Well-known and / or repetitive functions and structures are not described in detail to avoid obscuring the present application with unnecessary or redundant details. Therefore, the specific structural and functional details claimed herein are not intended to be limiting, but are merely used as a basis and representative basis for the claims to teach those skilled in the art to use the present application in substantially any suitable detailed structure in a variety of ways.
[0047] This specification may use the phrases "in one embodiment", "in another embodiment", "in yet another embodiment" or "in other embodiments", which may each refer to one or more of the same or different embodiments according to the present application.
[0048] In practical applications, the execution subject of the processing method of the process nodes in the embodiments of the present application can be the processor or controller of the system, etc. For the sake of convenience of explanation, the processor will be described in detail hereinafter. As Figure 1 shown, it is a flowchart of the method for constructing a prediction model provided by the embodiments of the present application, wherein the specific steps include S101-S105.
[0049] S101, obtain a training sample set, where the training sample set includes N compound samples.
[0050] In specific implementation, a large amount of reaction data of each compound on the eyes, such as no reaction, presence of corrosion, presence of irritation, etc., is collected from the test records of each hospital, each literature or each laboratory, and each compound in the reaction data is used as the compound sample in the training sample set. It is worth noting that in order to ensure the accuracy of the reaction of each compound on the eyes, the compound samples are obtained by removing salts, organometallic compounds, repeated compounds, inorganic substances and mixtures.
[0051] Moreover, each compound sample carries a label, which characterizes whether the chemical sample to which it belongs has corrosion or irritation to the eyes. For example, a label of 1 indicates that the chemical sample to which it belongs has corrosion or irritation to the eyes, and a label of 0 indicates that the chemical sample to which it belongs has no reaction to the eyes.
[0052] Taking all the compound samples as the training sample set, those skilled in the art should know that N in the embodiments of this application is not a fixed value, and it changes with the update of compounds or the update of the reaction data corresponding to the compounds.
[0053] S102, training a first sub-model with N compound samples, determining a fourth sub-model through the first sub-model, and training a second sub-model with M compound samples among the N compound samples.
[0054] Further, all the compound samples included in the training sample set are divided. Among them, a first sub-model is trained with N compound samples. Optionally, 80% of the N compound samples are used to train the first sub-model to be trained. After obtaining the first sub-model, the remaining 20% of the N compound samples are used to test the first sub-model, and the parameters of the first sub-model are adjusted based on the test results, thereby optimizing the first sub-model.
[0055] After obtaining the first sub-model, according to the combination rule, two or more of the multiple first sub-models are combined to obtain a combined sub-model; a fourth sub-model is determined based on a preset algorithm and the combined sub-model.
[0056] A second sub-model is trained with M compound samples among the N compound samples, where M is less than N. Similarly, 80% of the M compound samples can also be used to train the second sub-model to be trained. After obtaining the second sub-model, the remaining 20% of the M compound samples are used to test the second sub-model, and the parameters of the second sub-model are adjusted based on the test results, thereby optimizing the second sub-model.
[0057] It should be noted that the ratio between the compound samples used for training and the compound samples used for testing can be adjusted according to actual needs, and is not limited to 8:2 as described above.
[0058] The specific training process is as follows: For each compound sample or each group of compound samples, convert it into an input vector, and input the input vector into the model to be trained to obtain an actual result, which is the probability value of each compound sample in the compound sample or the group of compound samples being corrosive or irritating to the eyes. Further, determine whether the actual result is the same as the theoretical result represented by the label. If they are the same, determine the difference between the actual result and the reference value. The reference value can be set to 0.5. When a compound sample is corrosive or irritating to the eyes, its corresponding probability value is greater than 0.5. Of course, the reference value can also be set to other values, and the embodiments of the present application do not make specific limitations in this regard. Since when the difference is small, that is, the probability value is close to the reference value, it indicates that the model's judgment of the compound sample is not accurate enough and there may be errors. Therefore, when the difference is less than the set value, adjust the parameters of the model to be trained until the difference is greater than or equal to the set value to obtain a trained model. The above first sub-model and second sub-model can both be trained according to the above training process.
[0059] S103. Use the second sub-model to screen out target compound samples from Q compound samples among N compound samples.
[0060] After training the second sub-model, select Q compound samples from N compound samples. Preferably, the Q compound samples do not have the same compound samples as the above M compound samples. Screen the target compound samples from the Q compound samples through the second sub-model.
[0061] As an example, refer to Figure 2 the method flow chart shown to screen out target compound samples, where the specific steps include S201 - S203.
[0062] S201. Input the Q compound samples among N compound samples into the second sub-model in sequence.
[0063] S202. Obtain the actual result corresponding to each compound sample.
[0064] S203. Determine the target compound samples based on the actual result corresponding to each compound sample and the reference value.
[0065] In a specific implementation, after determining the Q compound samples, input the Q compound samples into the second sub-model in sequence, so that the second sub-model calculates the Q compound samples to obtain the actual result corresponding to each compound sample among the Q compound samples, that is, the probability value corresponding to each compound sample.
[0066] Afterwards, based on the actual results and the reference values corresponding to each compound sample, target compound samples are determined. Optionally, for each compound sample, the difference between its actual result and the reference value is calculated, that is, the differences between the probability values corresponding to each compound sample and the reference value are calculated respectively. The compound samples with the differences between the probability values and the reference value less than a preset threshold are selected and used as the target compound samples.
[0067] S104. A third sub-model is trained with the target compound samples.
[0068] After determining the target compound samples, the third sub-model to be trained is trained with the target compound samples to obtain the third sub-model. The specific training process is similar to that of the first sub-model and the second sub-model, and will not be elaborated here.
[0069] In the embodiments of the present application, the first sub-model to be trained, the second sub-model to be trained, and the third sub-model to be trained can all be one or more. Accordingly, the first sub-model, the second sub-model, and the third sub-model can all be one or more, and the numbers of the first sub-model, the second sub-model, and the third sub-model in the embodiments of the present application are the same.
[0070] As an example, for each compound sample, four molecular fingerprints of the compound sample are calculated using PaDEL-Descriptor, including Fingerprint (FP), Extended fingerprint (Ext), MACCS fingerprint (Maccs), and PubChem fingerprint (Pub); meanwhile, 78 descriptors of the compound sample are also calculated, including 11 descriptors based on the detour matrix, 42 information content descriptors, 22 path count descriptors, and 2 descriptors based on molecular weight and fragment contribution-based topological polar surface area (TopoPSA). After that, the low variance filter node in KNIME is used to set the variance threshold to 0.001 to remove descriptors with zero variance. The linear correlation node and correlation filter node in KNIME are used to set the threshold of the Pearson correlation coefficient to 0.999 to remove highly correlated descriptors. The strategy for selecting descriptors is set to backward feature elimination, and the evaluation criterion for arranging descriptors is the total prediction accuracy (ACC). When adding more descriptors but the total prediction accuracy changes little, the number of descriptors is considered appropriate. Finally, descriptors suitable for constructing models (the first sub-model to be trained, the second sub-model to be trained, and the third sub-model to be trained) are obtained, including 15 descriptors for predicting the presence of corrosion: piPC2, VR2_Dt, TopoPSA, CIC3, ZMIC4, BIC4, SpMax_Dt, VE1_Dt, R_TpiPCTPC, SIC3, piPC7, BIC0, IC0, MPC8, piPC6, and 12 descriptors for predicting the presence of irritation: TIC4, SIC3, piPC5, MPC10, MIC2, IC1, TIC3, MIC0, AMW, piPC4, MPC9, MPC5.
[0071] Afterwards, using five molecular description methods of four molecular fingerprints (FP, Ext, Maccs, Pub) and a set of descriptor groups (descriptors for predicting the presence of corrosion), combined with four machine learning methods of random forest (RF), tree ensemble (TE), gradient boosted trees (GBT), and radial basis function classification (RBF), 20 first submodels to be trained, 20 second submodels to be trained, and 20 third submodels to be trained for predicting the presence of corrosion are obtained; similarly, using five molecular description methods of four molecular fingerprints (FP, Ext, Maccs, Pub) and a set of descriptor groups (descriptors for predicting the presence of irritation), combined with four machine learning methods of random forest (RF), tree ensemble (TE), gradient boosted trees (GBT), and radial basis function classification (RBF), 20 first submodels to be trained, 20 second submodels to be trained, and 20 third submodels to be trained for predicting the presence of irritation are obtained.
[0072] Furthermore, according to the ten-fold cross-validation actual results and theoretical results of the training sample set, the counts of true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN) are obtained, and the total prediction accuracy ACC = (TP + TN) / (TP + FP + TN + FN), sensitivity SE = TP / (TP + FN), specificity SP = TN / (TN + FP), positive predictive value PPV = TP / (TP + FP), and negative predictive value NPV = TN / (TN + FP) are calculated.
[0073] The embodiments of the present application also provide a part of the training data during the training process: the first sub-model established and trained using the Maccs fingerprint and the tree integration method (for predicting whether corrosion exists), in the training data, ACC = 0.962, SE = 0.959, SP = 0.964, PPV = 0.943, NPV = 0.974, in the test data, ACC = 0.967, SE = 0.984, SP = 0.957, PPV = 0.938, NPV = 0.989. The first sub-model established and trained using the extended fingerprint and the random forest method (for predicting whether irritation exists), in the training data, ACC = 0.941, SE = 0.971, SP = 0.853, PPV = 0.950, NPV = 0.912, in the test data, ACC = 0.958, SE = 0.974, SP = 0.909, PPV = 0.969, NPV = 0.923. The second sub-model established and trained using the Maccs fingerprint and the tree integration method (for predicting whether corrosion exists), in the training data, ACC = 0.961, SE = 0.950, SP = 0.968, PPV = 0.950, NPV = 0.968, in the test data, ACC = 0.952, SE = 0.951, SP = 0.953, PPV = 0.931, NPV = 0.967. The second sub-model established and trained using the PubChem fingerprint and the radial basis function classification method (for predicting whether irritation exists), in the training data, ACC = 0.942, SE = 0.977, SP = 0.827, PPV = 0.948, NPV = 0.919, in the test data, ACC = 0.938, SE = 0.976, SP = 0.825, PPV = 0.943, NPV = 0.919. The third sub-model established and trained using the Maccs fingerprint and the gradient boosting tree method (for predicting whether corrosion exists), in the training data, ACC = 0.909, SE = 0.903, SP = 0.914, PPV = 0.897, NPV = 0.919, in the test data, ACC = 0.972, SE = 0.989, SP = 0.960, PPV = 0.943, NPV = 0.993. The third sub-model established and trained using the PubChem fingerprint and the tree integration method (for predicting whether irritation exists), in the training data, ACC = 0.900, SE = 0.943, SP = 0.796, PPV = 0.918, NPV = 0.852, in the test data, ACC = 0.955, SE = 0.982, SP = 0.875, PPV = 0.959, NPV = 0.943.
[0074] In addition, after obtaining the first sub-models, according to the combination rule, two or more of the multiple first sub-models are combined to obtain a combined sub-model. The combination rule is to randomly combine according to the types of machine learning methods. As can be seen from the above, there are 4 types of machine learning methods. Therefore, the obtained combined sub-models are It should be noted that 55 fourth sub-models for predicting whether corrosion exists are obtained by combining 20 first sub-models for predicting whether corrosion exists, and 55 fourth sub-models for predicting whether stimulation exists are obtained by combining 20 first sub-models for predicting whether stimulation exists.
[0075] After that, a fourth sub-model is determined based on a preset algorithm and the combined sub-model. The preset algorithm refers to the following formula (1), and the specific formula (1) is as follows:
[0076]
[0077] where P(Class = Positive) consensus represents the total probability value output by the fourth sub-model for the existence of corrosion / stimulation of the compound, represents the sub-probability value output by the first sub-model i for the existence of corrosion / stimulation of the compound; ACC i represents the accuracy rate of the first sub-model i, and n represents the number of first sub-models.
[0078] Among them, for the fourth sub-model (for predicting whether corrosion exists) established and trained using Maccs fingerprints and the tree ensemble and gradient boosting tree methods, in the training data, ACC = 0.963, SE = 0.959, SP = 0.966, PPV = 0.945, NPV = 0.974; in the test data, ACC = 0.967, SE = 0.984, SP = 0.957, PPV = 0.938, NPV = 0.989; for the fourth sub-model (for predicting whether stimulation exists) established and trained using extended fingerprints and the random forest and tree ensemble methods, in the training data, ACC = 0.941, SE = 0.973, SP = 0.849, PPV = 0.949, NPV = 0.915; in the test data, ACC = 0.959, SE = 0.977, SP = 0.905, PPV = 0.968, NPV = 0.930.
[0079] S105. Determine a prediction model based on the first sub-model, the second sub-model, the third sub-model, and the fourth sub-model.
[0080] After obtaining the first sub-model, the second sub-model, the third sub-model, and the fourth sub-model, refer to Figure 3 the shown method flow chart to determine the prediction model, where the specific steps include S301 and S302.
[0081] S301, determine the first accuracy rate of the first sub-model, the second accuracy rate of the second sub-model, the third accuracy rate of the third sub-model, and the fourth accuracy rate of the fourth sub-model.
[0082] S302, determine a prediction model based on the first accuracy rate, the second accuracy rate, the third accuracy rate, and the fourth accuracy rate.
[0083] In a specific implementation, use the same test sample set to test the first sub-model, the second sub-model, the third sub-model, and the fourth sub-model to obtain the first accuracy rate, the second accuracy rate, the third accuracy rate, and the fourth accuracy rate.
[0084] Select the highest accuracy rate from the first accuracy rate, the second accuracy rate, the third accuracy rate, and the fourth accuracy rate, and determine the model corresponding to it as the prediction model.
[0085] In a specific implementation, compare all models for predicting whether there is corrosion with all models for predicting whether there is irritation, and determine that the accuracy rate of the models for predicting whether there is corrosion is relatively high.
[0086] For all models for predicting whether there is corrosion and all models for predicting whether there is irritation respectively, compare the first sub-model and the second sub-model, and determine that the first accuracy rate of the first sub-model is higher than the second accuracy rate of the second sub-model. From this, it can be determined that the quantity and quality of the data set will have a certain impact on the accuracy rate of the model. Therefore, when training the model, as much accurate and comprehensive data as possible can be obtained to improve the accuracy rate of the model. Compare the second sub-model and the third sub-model, and determine that the third accuracy rate of the third sub-model is higher than the second accuracy rate of the second sub-model. Therefore, when training the model, data with relatively high uncertainty is added to improve the accuracy rate of the model. Compare the third sub-model and the fourth sub-model, and determine that the third accuracy rate of the third sub-model is similar to the fourth accuracy rate of the fourth sub-model, that is, it indicates that data with relatively high uncertainty is the main factor affecting the accuracy rate of the model. In summary, to improve the model performance such as the accuracy rate under the premise of limited cost, an active learning strategy should be used to query data with relatively high uncertainty instead of blindly adding data. Among them, the important advantage of the active learning strategy is to use a small amount of data to construct an effective model such as the third sub-model.
[0087] The embodiments of the present application use different compound samples to train different sub-models, and then determine a prediction model based on multiple sub-models to calculate the safety probability of the compound to be tested through the prediction model, so as to achieve the purpose of improving the prediction accuracy without increasing manpower, material resources, and financial resources while ensuring animal welfare. Moreover, the prediction model can effectively predict new compounds.
[0088] In a second aspect, an embodiment of the present application further provides a prediction method, which is to process a compound to be measured by using the above-determined prediction model to obtain the safety probability of the compound to be measured; wherein, the compound to be measured may be the same as or different from the compound samples included in the training sample set.
[0089] Considering that the training sample set cannot cover all compounds, the prediction model cannot be used to predict all compounds. At this time, the application domain of the prediction model is determined. Optionally, use the Enalos Domain-Similarity node in the KNIME platform to calculate the Euclidean distance between the compounds in the training sample set and the compounds in the test sample set, and then determine the APD value of the application domain of the prediction model. The calculation formula of the APD value is (APD = 'd' + Zσ), where 'd' represents the average value of all distances, σ represents the standard deviation of all distances, and Z is an empirical cut-off value. In the embodiment of the present application, Z is set to 0.5. When it is determined that the minimum distance between the compound to be measured and the compound sample is less than the APD value, that is, the compound to be measured is within the application domain of the prediction model, it is determined at this time that the prediction result of the prediction model for the compound to be measured is relatively accurate; when it is determined that the minimum distance between the compound to be measured and the compound sample is greater than or equal to the APD value, it is determined that the prediction result of the prediction model for the compound to be measured is not necessarily accurate.
[0090] Based on the training results, it is determined that the APD values of the application domains of the first sub-model and the fourth sub-model for predicting whether there is corrosion are 5.890, the APD value of the application domain of the second sub-model is 5.828, and the APD value of the application domain of the third sub-model is 5.725; it is also determined that the APD values of the application domains of the first sub-model and the fourth sub-model for predicting whether there is irritation are 11.975, the APD value of the application domain of the second sub-model is 8.532, and the APD value of the application domain of the third sub-model is 8.722.
[0091] In summary, the prediction model of the embodiment of the present application can effectively predict new compounds, especially new compounds corresponding to distances within the application domain can be predicted relatively accurately, and it has high practical value.
[0092] Based on the same inventive concept, a third aspect of the present application further provides a prediction model, which is constructed by the construction method provided in the first aspect.
[0093] Based on the same inventive concept, a fourth aspect of the present application further provides a prediction device corresponding to the prediction method. Since the principle of solving problems by the device in the present application is similar to the above prediction method of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0094] See Figure 4 As shown, the prediction device includes:
[0095] The prediction model 401 described in the third aspect;
[0096] A processing module 402, configured to: process a compound to be tested by using the prediction model to obtain a safety probability of the compound to be tested; wherein, the compound to be tested is the same as or different from the compound samples included in the training sample set.
[0097] In the embodiment of the present application, different sub-models are trained by using different compound samples, and then, a prediction model is determined based on multiple sub-models to calculate the safety probability of a compound to be tested through the prediction model, so as to achieve the purpose of improving the prediction accuracy without increasing the manpower, material resources and financial resources while ensuring animal welfare, and moreover, the prediction model can effectively predict new compounds.
[0098] The fifth aspect of the present application further provides a storage medium, which is a computer-readable medium and stores a computer program. When the computer program is executed by a processor, the method provided in any embodiment of the present application is implemented, including the following steps:
[0099] S11, obtaining a training sample set, where the training sample set includes N compound samples;
[0100] S12, training a first sub-model by using the N compound samples, determining a fourth sub-model by using the first sub-model, and training a second sub-model by using M compound samples among the N compound samples;
[0101] S13, using the second sub-model to screen out target compound samples from Q compound samples among the N compound samples;
[0102] S14, training a third sub-model by using the target compound samples;
[0103] S15, determining a prediction model based on the first sub-model, the second sub-model, the third sub-model and the fourth sub-model.
[0104] When the computer program is executed by the processor to determine the fourth sub-model through the first sub-model, the processor is further specifically executed as follows: combining two or more of the multiple first sub-models according to a combination rule to obtain a combined sub-model; and determining the fourth sub-model based on a preset algorithm and the combined sub-model.
[0105] When the computer program is executed by the processor to screen out the target compound samples from Q compound samples among the N compound samples by using the second sub-model, the following steps are specifically executed by the processor: input the Q compound samples among the N compound samples into the second sub-model in sequence; obtain the actual result corresponding to each compound sample; determine the target compound samples based on the actual result corresponding to each compound sample and the reference value.
[0106] When the computer program is executed by the processor to determine the target compound samples based on the actual result corresponding to each compound sample and the reference value, the following steps are specifically executed by the processor: for each compound sample, calculate the difference between its actual result and the reference value; if the difference is less than the preset threshold, determine the compound sample as the target compound sample.
[0107] When the computer program is executed by the processor to determine the prediction model based on the first sub-model, the second sub-model, the third sub-model and the fourth sub-model, the following steps are specifically executed by the processor: determine the first accuracy rate of the first sub-model, and determine the second accuracy rate of the second sub-model, and determine the third accuracy rate of the third sub-model, and determine the fourth accuracy rate corresponding to the fourth sub-model; determine the prediction model based on the first accuracy rate, the second accuracy rate, the third accuracy rate and the fourth accuracy rate.
[0108] When the computer program is executed by the processor to execute the prediction method, the following steps are further executed by the processor: process the compound to be tested by using the prediction model to obtain the safety probability of the compound to be tested; wherein, the compound to be tested is the same as or different from the compound samples included in the training sample set.
[0109] In the embodiments of the present application, different sub-models are trained by using different compound samples, and then, the prediction model is determined based on multiple sub-models to calculate the safety probability of the compound to be tested through the prediction model, so as to achieve the purpose of improving the prediction accuracy without increasing the manpower, material resources and financial resources while ensuring the animal welfare, and moreover, the prediction model can effectively predict new compounds.
[0110] It should be noted that the above storage medium of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any storage medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0111] The sixth aspect of the present application also provides an electronic device, as Figure 5 shown. The electronic device at least includes a memory 501 and a processor 502. A computer program is stored on the memory 501, and when the processor 502 executes the computer program on the memory 501, it implements the method provided by any embodiment of the present application. Exemplarily, the method executed by the computer program of the electronic device is as follows:
[0112] S21, obtain a training sample set, where the training sample set includes N compound samples;
[0113] S22, train a first sub-model through the N compound samples, determine a fourth sub-model through the first sub-model, and train a second sub-model through M compound samples among the N compound samples;
[0114] S23, use the second sub-model to screen out target compound samples from Q compound samples among the N compound samples;
[0115] S24. Train a third sub-model using the target compound samples;
[0116] S25. Determine a prediction model based on the first sub-model, the second sub-model, the third sub-model, and the fourth sub-model.
[0117] When the processor executes the computer program stored in the memory to determine the fourth sub-model using the first sub-model, it also performs the following: Combine two or more of the multiple first sub-models according to a combination rule to obtain a combined sub-model; Determine the fourth sub-model based on a preset algorithm and the combined sub-model.
[0118] When the processor executes the computer program stored in the memory to screen out the target compound samples from Q compound samples among the N compound samples using the second sub-model, it also performs the following: Input the Q compound samples among the N compound samples into the second sub-model in sequence; Obtain the actual result corresponding to each compound sample; Determine the target compound samples based on the actual result corresponding to each compound sample and a reference value.
[0119] When the processor executes the computer program stored in the memory to determine the target compound samples based on the actual result corresponding to each compound sample and a reference value, it also performs the following: For each compound sample, calculate the difference between its actual result and the reference value; If the difference is less than a preset threshold, determine the compound sample as the target compound sample.
[0120] When the processor executes the computer program stored in the memory to determine a prediction model based on the first sub-model, the second sub-model, the third sub-model, and the fourth sub-model, it also performs the following: Determine the first accuracy rate of the first sub-model, the second accuracy rate of the second sub-model, the third accuracy rate of the third sub-model, and the fourth accuracy rate corresponding to the fourth sub-model; Determine the prediction model based on the first accuracy rate, the second accuracy rate, the third accuracy rate, and the fourth accuracy rate.
[0121] When the processor executes the prediction method stored in the memory, it also performs the following: Process the compound to be measured using the prediction model to obtain the safety probability of the compound to be measured; where the compound to be measured is the same as or different from the compound samples included in the training sample set.
[0122] In the embodiments of the present application, different sub-models are trained using different compound samples, and then, a prediction model is determined based on multiple sub-models to calculate the safety probability of a compound to be tested through the prediction model. The purpose of improving the prediction accuracy is achieved while ensuring animal welfare and without increasing human, material, and financial resources. Moreover, the prediction model can effectively predict new compounds.
[0123] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the block may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0124] The above description is only the preferred embodiments of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present application.
[0125] In addition, although the operations are depicted in a specific order, this should not be construed as requiring the operations to be performed in the specific order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present application. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0126] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.
[0127] The above has described multiple embodiments of the present application in detail, but the present application is not limited to these specific embodiments. Based on the concept of the present application, those skilled in the art can make various variations and modifications to the embodiments, and these variations and modifications should all fall within the scope claimed by the present application.
Claims
1. A method for constructing a prediction model, characterized in that, Comprising: Obtain a training sample set, where the training sample set includes N compound samples, and each compound sample carries a label, which characterizes whether the chemical sample it belongs to corrodes or irritates the eyes; Train a first sub-model through the N compound samples, determine a fourth sub-model through the first sub-model, and train a second sub-model through M compound samples among the N compound samples; Use the second sub-model to screen out target compound samples from Q compound samples among the N compound samples; Train a third sub-model through the target compound samples; Determine a prediction model based on the first sub-model, the second sub-model, the third sub-model, and the fourth sub-model, so as to calculate the safety probability of a compound to be measured through the prediction model; The determining the fourth sub-model through the first sub-model includes: According to a combination rule, combine two or more of the multiple first sub-models to obtain a combined sub-model; Determine the fourth sub-model based on a preset algorithm and the combined sub-model.
2. The construction method according to claim 1, characterized in that, The using the second sub-model to screen out target compound samples from Q compound samples among the N compound samples includes: Input the Q compound samples among the N compound samples into the second sub-model in sequence; Obtain the actual result corresponding to each compound sample; Determine the target compound samples based on the actual result corresponding to each compound sample and a reference value.
3. The construction method according to claim 2, characterized in that, The determining the target compound samples based on the actual result corresponding to each compound sample and a reference value includes: For each compound sample, calculate the difference between its actual result and the reference value; If the difference is less than a preset threshold, determine the compound sample as the target compound sample.
4. The construction method according to claim 1, characterized in that, The numbers of the first sub-model, the second sub-model, and the third sub-model are the same.
5. The construction method according to claim 1, characterized in that, The determining the prediction model based on the first sub-model, the second sub-model, the third sub-model, and the fourth sub-model includes: Determine the first accuracy rate of the first sub-model, determine the second accuracy rate of the second sub-model, determine the third accuracy rate of the third sub-model, and determine the fourth accuracy rate corresponding to the fourth sub-model; Determine the prediction model based on the first accuracy rate, the second accuracy rate, the third accuracy rate, and the fourth accuracy rate.
6. A prediction model, characterized in that, Constructed by the construction method according to any one of claims 1 to 5.
7. A prediction method, characterized in that, Comprising: Use the prediction model according to claim 6 to process a compound to be measured, and obtain the safety probability of the compound to be measured; wherein, the compound to be measured is the same as or different from the compound samples included in the training sample set.
8. A prediction device, characterized in that, Comprising: The prediction model according to claim 6; A processing module configured to: use the prediction model to process a compound to be measured, and obtain the safety probability of the compound to be measured; wherein, the compound to be measured is the same as or different from the compound samples included in the training sample set.
9. An electronic device, characterized in that, Comprising: A processor and a memory, the memory storing machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory via a bus. When the machine-readable instructions are executed by the processor, the following steps are performed: Processing a compound to be tested using the prediction model according to claim 6 that has been pre-trained to obtain a safety probability of the compound to be tested; wherein the compound to be tested is the same as or different from the compound samples included in the training sample set.
Citation Information
Patent Citations
Model training method, model prediction method, molecule screening method and device thereof
CN114187980A