Model training method and device, medium and electronic equipment

By screening the molecular information with high uncertainty of the photosensitizer candidate molecules as sample molecular information, the target model is trained, and the training efficiency is improved by combining the correction model, which solves the problem of low photosensitizer screening efficiency in the existing technology and realizes the accurate and efficient determination of the physical parameters of the photosensitizer.

CN120636581APending Publication Date: 2025-09-12TSINGHUA UNIVERSITY
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510716781.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The existing technology has the problems of strong subjectivity and low calculation efficiency when screening photosensitizers, making it difficult to accurately and efficiently determine the physical parameters of the photosensitizers.

Method used

By obtaining the molecular information of candidate photosensitizer molecules, selecting part of the molecular information as basic molecular information, using multiple target models to determine the uncertainty, screening out the molecular information with uncertainty not lower than the threshold as sample molecular information, and training the target model through the training set, combined with the correction model to improve the training efficiency.

Benefits of technology

The screening efficiency of photosensitizers and the training efficiency of models have been significantly improved, and the physical parameters of candidate photosensitizer molecules can be accurately and efficiently determined.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636581A_ABST
    Figure CN120636581A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method and device, a medium and sub-equipment, and the method comprises the steps: obtaining a first molecular set comprising the molecular information of a plurality of photosensitizer candidate molecules, selecting the molecular information of a part of photosensitizer candidate molecules from the first molecular set as the first basic molecular information, and then, obtaining the first basic molecular information; determining the uncertainty corresponding to each piece of first basic molecular information according to predicted physical parameters output by a plurality of to-be-trained target models for each piece of first basic molecular information, screening out molecular information of which the uncertainty is not lower than a preset threshold value from each piece of first basic molecular information, and adding the molecular information as sample molecular information into a training set, and training at least part of to-be-trained target models through the training set. Through the method, the target model focuses on learning sample molecule information with high value, and the training efficiency of the model is remarkably improved. Moreover, through the trained target model, the screening efficiency of the photosensitizer can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of computer technology and biochemistry, and in particular to a model training method, device, medium and electronic equipment. Background Art

[0002] Photosensitizers, as key functional molecules in modern energy and biomedical technologies, are widely used in fields such as solar photovoltaic conversion, photodynamic therapy, and biosensing. To identify photosensitizers that meet desired performance requirements, current methods rely on empirical screening to identify the desired photosensitizer from a pool of candidate photosensitizers. However, due to the high degree of subjectivity inherent in empirical screening, it is often difficult to guarantee that the selected photosensitizers will meet the desired performance requirements.

[0003] In order to improve the screening accuracy, methods such as Time Domain Density Functional Theory (TD-DFT) can also be used to calculate the physical parameters of candidate photosensitizers, so as to screen suitable photosensitizers based on the calculated physical parameters. Although this method can accurately determine the physical parameters of candidate photosensitizers, the number of candidate photosensitizers is too large (often more than 10 6 ), so this method requires a lot of computing time, which seriously reduces the screening efficiency of photosensitizers. Summary of the Invention

[0004] The embodiments of the present application provide a model training method, device, medium and electronic device to partially solve the above-mentioned problems existing in the prior art.

[0005] This application adopts the following technical solutions:

[0006] The present invention provides a model training method, including:

[0007] Acquire a first molecule set, wherein the first molecule set includes molecular information of each photosensitizer candidate molecule;

[0008] Selecting molecular information of some photosensitizer candidate molecules from the first molecule set as first basic molecular information;

[0009] Determining the uncertainty corresponding to each first basic molecular information according to the predicted physical parameters respectively output by several target models to be trained for the first basic molecular information, where different target models have different numbers of model parameters and / or network layers;

[0010] Selecting molecular information with an uncertainty not less than a preset threshold from each first basic molecular information and adding it to the training set as sample molecular information;

[0011] At least a portion of the target model to be trained is trained based on the determined real physical parameters and predicted physical parameters corresponding to the sample molecular information contained in the training set.

[0012] Optionally, obtaining the first molecule set specifically includes:

[0013] Get the initial molecule set;

[0014] Redundant molecular information in the initial molecular set is removed to obtain a first molecular set.

[0015] Optionally, determining the uncertainty corresponding to each first basic molecular information according to the predicted physical parameters output by several target models to be trained for each first basic molecular information specifically includes:

[0016] For each piece of first basic molecular information, determine the average predicted physical information corresponding to the first basic molecular information according to the predicted physical parameters respectively output by the target models to be trained for the first basic molecular information;

[0017] The uncertainty corresponding to the first basic molecular information is determined according to the difference between the average predicted physical information and the predicted physical parameters output by each target model to be trained for the first basic molecular information.

[0018] Optionally, the real physical parameters include: a first singlet excitation energy and a first triplet excitation energy;

[0019] Determining the actual physical parameters corresponding to the sample molecular information contained in the training set specifically includes:

[0020] For each sample molecule information included in the training set, a preset first method is used to determine a first singlet excitation energy and a first triplet excitation energy corresponding to the sample molecule information at a first accuracy;

[0021] determining, based on the first singlet excitation energy and the first triplet excitation energy corresponding to the sample molecule information at the first precision, a singlet-triplet energy level difference corresponding to the sample molecule information at the first precision;

[0022] The sample molecule information and the single-three-state energy level difference corresponding to the sample molecule information at the first precision are input into a pre-trained calibration model, so that the calibration model outputs the physical parameters corresponding to the sample molecule information at a second precision as the true physical parameters, and the second precision is higher than the first precision.

[0023] Optionally, training a calibration model specifically includes:

[0024] Acquire a second molecule set, wherein the second molecule set includes molecular information of each photosensitizer candidate molecule;

[0025] Selecting molecular information of some photosensitizer candidate molecules from the second molecular set as second basic molecular information;

[0026] Using the first method, determining the single-state and triple-state energy level differences corresponding to the second basic molecular information at the first precision;

[0027] Inputting the second basic molecular information and the single-three-state energy level difference corresponding to the second basic molecular information at the first precision into the calibration model to be trained, so that the calibration model to be trained outputs the predicted physical parameter corresponding to the second basic molecular information;

[0028] The correction model to be trained is trained based on the deviation between the predicted physical parameters corresponding to the second basic molecular information and the physical parameters corresponding to the second basic molecular information at the second precision. The physical parameters corresponding to the second basic molecular information at the second precision are determined using a preset second method. The cost of determining the physical parameters using the second method is higher than the cost of determining the physical parameters using the first method.

[0029] Optionally, training at least part of the target model to be trained is performed based on the determined real physical parameters and predicted physical parameters corresponding to the sample molecular information contained in the training set. Training at least part of the target model to be trained specifically includes:

[0030] Based on the real physical parameters and predicted physical parameters corresponding to the sample molecular information contained in the determined training set, at least part of the target model to be trained is trained, and through the trained target model, the sample molecular information is screened from the remaining molecular information in the first molecular set, and the screened sample molecular information is added to the training set, so as to iteratively train at least part of the target model through the updated training set until a preset iteration stop condition is met.

[0031] Optionally, the method further includes:

[0032] Determining the real physical parameters corresponding to the pre-stored molecular information of each photosensitizer candidate molecule using at least a portion of the trained target model;

[0033] According to the determined real physical parameters, the molecular information used as the photosensitizer molecule is screened from the pre-stored molecular information of each photosensitizer candidate molecule.

[0034] The present invention provides a model training device, comprising:

[0035] an acquisition module, configured to acquire a first molecule set, wherein the first molecule set includes molecular information of each photosensitizer candidate molecule;

[0036] A selection module, configured to select molecular information of some photosensitizer candidate molecules from the first molecule set as first basic molecular information;

[0037] A first determination module is configured to determine the uncertainty corresponding to each piece of first basic molecular information based on the predicted physical parameters respectively output by a plurality of target models to be trained for the first basic molecular information, wherein different target models have different numbers of model parameters and / or network layers;

[0038] A selection module is used to select molecular information with an uncertainty not less than a preset threshold from each first basic molecular information, and add it to the training set as sample molecular information;

[0039] The first training module is used to train at least part of the target model to be trained according to the determined real physical parameters and predicted physical parameters corresponding to the sample molecular information contained in the training set.

[0040] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned model training method is implemented.

[0041] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned model training method when executing the program.

[0042] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects:

[0043] The model training method, apparatus, medium, and electronic device provided in the embodiments of the present application obtain a first molecular set including molecular information of multiple photosensitizer candidate molecules, select molecular information of some photosensitizer candidate molecules from the first molecular set as each first basic molecular information, then determine the uncertainty corresponding to each first basic molecular information based on the predicted physical parameters output by several target models to be trained for each first basic molecular information, screen out molecular information with an uncertainty not less than a preset threshold from each first basic molecular information, add the information as sample molecular information to the training set, and use the training set to train at least some of the target models to be trained.

[0044] The above method demonstrates that by selecting molecular information of high-uncertainty photosensitizer candidate molecules from the molecular set as sample molecular information for training the target model, the target model can focus on learning high-value sample molecular information, significantly improving model training efficiency. Furthermore, the trained target model can accurately and efficiently determine the physical parameters corresponding to a large number of photosensitizer candidate molecules, significantly improving photosensitizer screening efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this application. The illustrative embodiments of this application and their descriptions are used to explain this specification and do not constitute an improper limitation on this application. In the drawings:

[0046] Figure 1 A flowchart of a model training method provided in an embodiment of the present application;

[0047] Figure 2 A schematic diagram of screening sample molecular information with training value based on multiple dimensions provided in an embodiment of the present application;

[0048] Figure 3 A schematic diagram of a model training device provided in an embodiment of the present application;

[0049] Figure 4 A method corresponding to the embodiment of the present application is provided Figure 1 Schematic diagram of the electronic device. DETAILED DESCRIPTION

[0050] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this specification will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0051] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.

[0052] Figure 1 A flow chart of a model training method provided in an embodiment of the present application includes the following steps:

[0053] S101: Acquire a first molecule set, where the first molecule set includes molecule information of each photosensitizer candidate molecule.

[0054] In the embodiments of the present application, the execution entity for executing model training can be a terminal device such as a desktop computer or laptop computer, a client installed in the terminal device, a server, or a dedicated training machine specifically used for training artificial intelligence models. For ease of description, the following describes a model training method provided in the embodiments of the present application using the server as the execution entity as an example.

[0055] In an embodiment of the present application, a target model that can accurately determine the real physical parameters corresponding to the photosensitizer candidate molecules needs to be obtained through training. Once the training of the target model is completed, the trained target model can be deployed in a preset device (such as a server) to batch determine the real physical parameters corresponding to each photosensitizer candidate molecule.

[0056] Before executing the model training task, the server can first obtain a first molecular set containing molecular information of many photosensitizer candidate molecules. Among them, the server can collect molecular information that can serve as photosensitizer candidate molecules from public molecular databases and public literature. In addition to representative photosensitizer molecules, the collected molecular information can also include molecular information of molecules such as classic photosensitizer skeletons (such as porphyrin, phthalocyanine, etc.) and molecular information of their derivatives.

[0057] Since the molecular information collected from public molecular databases or public documents may contain a large amount of redundant information, the server needs to filter the collected molecular information to obtain the above-mentioned first molecular set.

[0058] In this process, the server can first construct an initial molecular set through the collected molecular information, and then identify repeated molecular information as redundant molecular information from the molecular information, and then remove the repeated redundant molecular information to obtain the first molecular set.

[0059] To facilitate subsequent processing, in an embodiment of the present application, the server can further convert the molecular information of the collected photosensitizer candidate molecules. In this process, the server can convert the molecular information of any photosensitizer candidate molecule into a SMILES string and normalize the converted SMILES string to obtain the molecular graph data corresponding to the photosensitizer candidate molecule. The molecular graph data corresponding to the photosensitizer candidate molecule can represent the molecular structure of the photosensitizer candidate molecule in the form of a graph.

[0060] The reason for converting it into a SMILES string is to unify the expression of the molecule, because the molecular information of the photosensitizer candidate molecules obtained may come from different public databases or public information. Different data sources have different representations for the same molecule or functional group. If they are not unified, it may affect subsequent steps such as parameter determination.

[0061] For example, CCO and COC (different ways of writing ethanol) can be uniformly expressed as CCO by converting them into SMILES strings.

[0062] In an embodiment of the present application, the server can standardize the SMILES strings corresponding to the photosensitizer candidate molecules using a preset tool, such as RDKit. In the process of screening redundant information, the server can process the SMILES strings using tools such as RDKit to determine the molecular fingerprints corresponding to the photosensitizer candidate molecules. Molecular fingerprints are mainly used to reflect the local structural characteristics and physicochemical properties of molecules. Therefore, based on the determined molecular fingerprints of each photosensitizer candidate molecule, the server can screen out redundant information from the collected photosensitizer candidate molecules to obtain the above-mentioned first molecular set.

[0063] The molecular fingerprints mentioned above may include information such as MACCS fingerprints, Morgan fingerprints, chirality information, molecular radius, molecular length, etc.

[0064] S102: Selecting molecular information of some photosensitizer candidate molecules from the first molecule set as first basic molecular information.

[0065] Due to the huge amount of molecular information collected on photosensitizer candidate molecules, if all the collected photosensitizer candidate molecules are directly used as sample molecular information at the beginning of the subsequent model training process, it will not only consume too much cost to determine the actual physical parameters corresponding to each photosensitizer candidate molecule, but also greatly reduce the efficiency of model training.

[0066] Therefore, in the embodiment of the present application, the server can first select the molecular information of a portion of the photosensitizer candidate molecules from the above-mentioned first molecular set as each first basic molecular information, and then in the subsequent process, complete the model training based on each first basic molecular information. There are many ways to select the molecular information of a portion of the photosensitizer candidate molecules from the first molecular set. For example, each first basic molecular information can be selected by random selection, or each first basic molecular information can be selected based on the importance of the molecule (e.g., representative molecules have a relatively high importance, while derivatives have a relatively low importance).

[0067] S103: Determine the uncertainty corresponding to each first basic molecular information according to the predicted physical parameters respectively output by several target models to be trained for the first basic molecular information, where different target models have different numbers of model parameters and / or network layers.

[0068] In order to further improve the training efficiency of the model and enable the model to use sample molecular information that is more valuable for training as much as possible during the training process, multiple target models can be deployed in an embodiment of the present application. These target models are all used to determine the physical parameters of the molecules, but the number of model parameters of these target models is different, or the number of network layers of these target models is different, or the number of model parameters and the number of network layers of these target models are different.

[0069] Therefore, these target models can be regarded as different expert networks, which will give different output results for the same molecular information.

[0070] On this basis, the server can input each piece of first basic molecular information selected in the above manner into each target model to be trained. For each piece of first basic molecular information, these target models will respectively determine the physical parameters of the first basic molecular information as predicted physical parameters. The server can then determine the uncertainty corresponding to each piece of first basic molecular information based on the predicted physical parameters determined by these target models. In subsequent processes, the server can further use the determined uncertainty to further select molecular information with greater training value from these first basic molecular information as sample molecular information.

[0071] The aforementioned uncertainty actually reflects the degree of uncertainty in each target model's recognition of the same basic molecular information. That is, for any given piece of basic molecular information, if the uncertainty corresponding to that basic molecular information is high, it indicates that different target models have significantly different understandings of that first basic molecular information, ultimately leading to significant differences in the predicted physical parameters output by different target models for that first basic molecular information. On the other hand, if the uncertainty corresponding to that basic molecular information is low, it indicates that different target models have similar understandings of that first basic molecular information, ultimately resulting in less variation in the predicted physical parameters output by different target models for that first basic molecular information.

[0072] Obviously, the basic molecular information that is close to the level of knowledge of each target model has a lower value for training, while the basic molecular information that has a large difference in the level of knowledge of the target models has a higher value for training, so it is necessary to use it as sample molecular information in the subsequent process to perform model training on at least part of these target models.

[0073] In an embodiment of the present application, the server may determine the uncertainty corresponding to each piece of first basic molecular information in the following manner. Specifically, for each piece of first basic molecular information, the server may determine the average predicted physical parameter corresponding to the first basic molecular information based on the predicted physical parameters respectively output by several target models to be trained for the first basic molecular information. Thereafter, the server may determine the uncertainty corresponding to the first basic molecular information based on the difference between the average predicted physical parameter and the predicted physical parameter output by each target model to be trained for the first basic molecular information.

[0074] If the predicted physical parameters output by a target model for the first basic molecule information are significantly different from the above-mentioned average predicted physical parameters, it means that the target model's understanding of the first basic molecule is significantly different from that of other target models.

[0075] Therefore, if the sum of the differences between the average predicted physical information and the predicted physical parameters output by each target model to be trained for the first basic molecular information is large, it means that there are significant differences in the degree of understanding of the first basic molecular information by most target models, and it can be added to the training set as sample molecular information with training value.

[0076] In addition, as mentioned in the above step S102, the server can ultimately convert the molecular information of the photosensitizer candidate molecule into its corresponding molecular graph data. Therefore, the server can input the molecular graph data corresponding to the first basic molecular information into the target model to be trained to determine the predicted physical parameters corresponding to the first basic molecule.

[0077] In the embodiment of the present application, the physical parameters corresponding to the first basic molecular information may include: the first singlet excitation energy and the first triplet excitation energy corresponding to the photosensitizer candidate molecule. These physical parameters are selected because they provide good guidance for selecting suitable photosensitizer molecules.

[0078] It should also be noted that since the accuracy of the physical parameters output by the target model is low when the training is not completed, it is called predicted physical parameters. For the sample molecular information used as training samples, it corresponds to real physical parameters. The so-called real physical parameters are accurate physical parameters, which can be used as labels to train the target model to be trained.

[0079] S104: Selecting molecular information with an uncertainty not less than a preset threshold from each of the first basic molecular information, and adding the information as sample molecular information to the training set.

[0080] After the uncertainty of each first basic molecular information is determined in the above manner, molecular information with an uncertainty not less than a preset threshold can be selected from the first basic molecular information and added to the training set as sample molecular information.

[0081] S105: Training at least part of the target model to be trained according to the determined real physical parameters and predicted physical parameters corresponding to the sample molecular information contained in the training set.

[0082] As mentioned above, during the training of the target model, it is necessary to determine the true physical parameters corresponding to each sample molecule in the training set. However, using a more precise method (such as TD-DFT) to determine the true physical parameters corresponding to each sample molecule requires a significant amount of time and resources.

[0083] To this end, in an embodiment of the present application, for each sample molecule information in the training set, the server may first use a preset first method to determine the first singlet excitation energy and the first triplet excitation energy corresponding to the sample molecule information at a first precision, and then further determine the difference between the first singlet excitation energy and the first triplet excitation energy corresponding to the sample molecule information at the first precision, that is, the single- and triplet energy level difference corresponding to the sample molecule information. The server may input the molecular graph data corresponding to the sample molecule information and the determined single- and triplet energy level difference corresponding to the sample molecule information at the first precision into a pre-trained calibration model, so that the calibration model outputs the physical parameters corresponding to the sample molecule information at a second precision as the true physical parameters corresponding to the sample molecule information, wherein the second precision mentioned above is higher than the first precision.

[0084] The above process can be understood as the server first determines the physical parameters corresponding to the sample molecular information in a less accurate manner, and then uses a pre-trained correction model to correct the less accurate physical parameters to obtain accurate physical parameters.

[0085] Therefore, the server actually uses a pre-trained correction model to batch calibrate the low-precision physical parameters corresponding to the sample molecular information, thereby efficiently determining the high-precision physical parameters (i.e., the real physical parameters) corresponding to the molecular information of each sample in the training set.

[0086] For the training process of the correction model, the server can use a portion of the photosensitizer candidate molecule information as training samples, and calculate the physical parameters of these training samples using low-precision and high-precision methods respectively, and then use the physical parameters determined by high-precision method as annotations to train the correction model, so that it has the ability to correct low-precision physical parameters into high-precision physical parameters, so as to determine the real physical parameters corresponding to each molecular information in batches in the subsequent process.

[0087] Specifically, the server may first obtain a second molecular set containing molecular information of each photosensitizer candidate molecule, wherein the second molecular set may be the same molecular set as the first molecular set in step S101, or a different molecular set. The terms "first" and "second" are used primarily to distinguish between different molecular sets used in the training process of different models, i.e., the first molecular set is primarily used in the training process of the target model, while the second molecular set is primarily used in the training process of the calibration model. The terms "first" and "second" themselves have no special meaning.

[0088] Furthermore, the server may select molecular information of some photosensitizer candidate molecules from the above-mentioned second molecular set as the second basic molecular information. The selection of some photosensitizer candidate molecules here is mainly to train the calibration model by using a small amount of data, thereby determining the real physical parameters corresponding to a large number of photosensitizer candidate molecules.

[0089] The server can use the first method to determine the single- and triplet energy level differences corresponding to the second basic molecular information at the first precision. The specific process of determining the single- and triplet energy level differences here is the same as described above, that is, the first singlet excitation energy and the first triplet excitation energy corresponding to the second basic molecular information are determined by the first method, and then the difference between the two is calculated to determine the single- and triplet energy level differences corresponding to the second basic molecular information.

[0090] The server can calculate the single and triple state energy levels corresponding to the second basic molecular information using the following formula:

[0091] S1=E singlet -E ground

[0092] T1=E triplet -E ground

[0093] ΔE ST =S1-T1

[0094] Among them, S1 is used to represent the first singlet excitation energy corresponding to the second basic molecular information, T1 is used to represent the first triplet excitation energy corresponding to the second basic molecular information, ΔE ST It is used to represent the single- and triple-state energy level differences corresponding to the second basic molecular information.

[0095] The server can use the first method to first calculate the ground state energy E corresponding to the second basic molecule information ground , the first excited singlet state energy E singlet and the first excited triplet state energy E triplet Then, the above formula can be used to determine the first singlet excitation energy and the first triplet excitation energy, and then determine the singlet and triplet energy level difference.

[0096] At the same time, the server can use a second method to determine the physical parameters corresponding to the second basic molecular information at the second precision (i.e., the first singlet excitation energy and the first triplet excitation energy at the second precision), and use them as annotation information. The server can input the molecular graph data corresponding to the second basic molecular information and the singlet and triplet energy level differences corresponding to the second basic molecular information at the first precision into the calibration model to be trained, so that the calibration model to be trained outputs the predicted physical parameters corresponding to the second basic molecular information.

[0097] Afterwards, the server can train the correction model to be trained based on the deviation between the predicted physical parameters corresponding to the second basic molecular information and the physical parameters corresponding to the second basic molecular information at the second precision, that is, minimize the deviation between the predicted physical parameters corresponding to the second basic molecular information and the actual physical parameters determined by the second method, and train the correction model.

[0098] In the embodiment of the present application, the first method can be a semi-empirical method such as xtb-sTDA, and the second method can be a method such as TD-DFT. Obviously, the first method is lower in computational accuracy than the second method, but the cost of the first method is significantly lower than the second method.

[0099] Furthermore, although the training of the correction model also consumes costs, since the number of collected photosensitizer candidate molecules is too large, the cost of selecting a small portion of the data to train the correction model (including the cost of calculating the low-precision physical parameters of this part of the data using the first method + the cost of calculating the high-precision physical parameters of this part of the data using the second method + the cost of training the correction model) is much lower than the cost of calculating the high-precision physical parameters of most or even all of the collected photosensitizer candidate molecules using the second method.

[0100] Therefore, the trained correction model can improve the efficiency of determining whether molecular information corresponds to real physical parameters, and also reduce the overall cost of this process (including the training process of the correction model).

[0101] After training the above correction model, it can actually be deployed in multiple device nodes to calibrate the physical parameters of each photosensitizer candidate molecule in batches through parallel operation to determine the actual physical parameters corresponding to the photosensitizer candidate molecules.

[0102] In the process of training the target model, an iterative training method can be adopted, that is, the target model is trained through the training set. After completing one round of training, the sample molecular information is screened out from the molecular information of the remaining photosensitizer candidate molecules, and the screened sample molecular information is added to the training set, so that the target model can be trained again through the updated training set until the preset iteration stop condition is met.

[0103] In each round of training, the goal is to ensure that the predicted physical parameters output by the target model are as close as possible to the actual physical parameters of the sample molecular information. This means minimizing the deviation between the predicted and actual physical parameters of the sample molecular information. The method for selecting sample molecular information from the first molecular set during each round of iterative training is essentially the same: first, a portion of molecular information is selected from the first molecular set as the first basic molecular information. The uncertainty of this first basic molecular information is then determined using multiple target models. Based on this uncertainty, sample molecular information is further selected from the first basic molecular information and added to the training set, thereby updating the training set.

[0104] In the embodiments of the present application, there are various specific iteration stopping conditions. For example, if it is determined that a preset number of iteration rounds has been reached, then the iteration stopping condition is determined to be satisfied. For another example, if it is determined that the model parameters of the target model have converged to a preset range, then the iteration stopping condition is determined to be satisfied. For another example, if it is determined that the ratio between the number of sample molecular information included in the training set and the number of molecular information in the first molecular set meets a preset threshold, then the iteration stopping condition is determined to be satisfied. Other forms of iteration stopping conditions are not illustrated here one by one.

[0105] In addition, since the amount of molecular information of the collected photosensitizer candidate molecules is too huge, it is necessary to further screen out molecular information with more training value as sample molecular information. Therefore, in addition to combining the above-mentioned uncertainty to screen valuable sample molecular information, other dimensions can also be combined for screening, such as Figure 2 shown.

[0106] Figure 2 A schematic diagram of screening sample molecular information with training value based on multiple dimensions provided in an embodiment of the present application.

[0107] Figure 2The sample molecular information is mainly screened based on four dimensions. First, through the above method, the molecular information with the largest uncertainty is selected as the sample molecular information. This process is introduced in the above content and will not be repeated here.

[0108] Second, because physical parameters of molecules suitable for photosensitizers typically fall within a preset range, the server can select molecules whose physical parameters fall within this range from the first basic molecular information as sample molecular information. In other words, even if molecular information of molecules unsuitable for photosensitizers is used as sample molecular information to train the target model, in actual applications, this molecular information is often not used to select suitable photosensitizers.

[0109] As mentioned above, training the target model can be considered a multi-round iterative process. During each round of training, new sample molecular information is added to the training set. To ensure that the sample molecular information added to the training set has a positive impact on the training of the target model and prevent useless training, molecular information with high diversity in the molecular feature space is selected as the sample molecular information.

[0110] In other words, the sample molecular information newly added to the training set needs to ensure that it is not too similar to the sample molecular information already in the training set in the feature space. Therefore, each time the server screens the sample molecular information from the first basic molecular information, it needs to calculate the molecular feature space similarity between the first basic molecular information to be added to the training set and the sample molecular information in the training set. If the similarity is high, it will not be used as sample molecular information. On the contrary, if the similarity is low, it is determined that the diversity between the first basic molecular information to be added and the sample molecular information already in the training set is strong, and then it is added to the training set as sample molecular information.

[0111] In an embodiment of the present application, there are multiple ways to calculate the diversity between molecular information. For example, the Tanimoto similarity between two molecular information can be calculated to ensure that the newly added sample molecular information and the sample molecular information already added to the training set have strong diversity in the molecular feature space.

[0112] Fourth, in practical applications, a suitable photosensitizer must not only meet the required physical parameters but also be easy to synthesize and suitable for industrial production. Therefore, it is necessary to further screen the first basic molecular information to identify suitable molecules for synthesis and add them as sample molecules to the training set.

[0113] In an embodiment of the present application, the server can input the first basic molecular information into a preset machine learning model to determine the synthesis path of the first basic molecular information, and then select molecular information that is easy to synthesize as sample molecular information based on the determined synthesis path.

[0114] The ease of synthesis of a molecule can be assessed based on factors such as the number of synthesis steps and synthesis conditions (e.g., temperature, pressure, etc.) included in the synthesis pathway of the molecule information. For example, if the number of synthesis steps included in the synthesis pathway of a molecule information is relatively small, the molecule corresponding to the molecule information is determined to be easy to synthesize. For another example, if the synthesis conditions in the synthesis pathway corresponding to the molecule information are relatively stringent, the molecule corresponding to the molecule information is determined to be difficult to synthesize.

[0115] Therefore, the server can actually combine the above four dimensions to select molecular information that has high uncertainty, meets the target property requirements, has high diversity in the molecular feature space with the existing sample molecular information in the training set, and is easy to synthesize, as the sample molecular information that needs to be newly added to the training set.

[0116] The server can train at least part or all of the target model in the above manner, so that in a subsequent process, the server can determine the real physical parameters corresponding to the molecular information of each photosensitizer candidate molecule through the trained at least part of the target model, and then use the determined real physical parameters to screen out the molecular information used as the photosensitizer molecule from the pre-stored molecular information of each photosensitizer candidate molecule.

[0117] Among them, since each target model can be regarded as an expert network, the process of determining the real physical parameters through the trained target model can actually be regarded as multiple expert networks giving a comprehensive opinion on a photosensitizer candidate molecule. Therefore, after receiving the screening instruction in actual application, the server can determine the real physical parameters corresponding to the photosensitizer candidate molecule based on the output results of multiple target models for the same photosensitizer candidate molecule (such as taking the average of the output results of each target model). Among them, the screening instruction mentioned here can be received by the server in response to the screening operation performed by the user.

[0118] When screening photosensitizer molecules, the server can select molecular information of molecules that meet the property requirements based on the determined real physical parameters (such as the values ​​of the above-mentioned physical parameters fall within the preset numerical range), and then combine the synthesis difficulty of the molecules, etc., to finally screen out suitable photosensitizer molecules. In this process, the server can obtain conditional parameters for screening photosensitizer molecules (such as the above-mentioned preset numerical range, parameters for evaluating synthesis difficulty, etc.). The conditional parameters can be input by the user according to actual needs, or they can be read by the server from a preset database. Afterwards, the server can screen out suitable photosensitizer molecules according to the conditional parameters and the real physical parameters determined by the target model.

[0119] The above method demonstrates that by selecting molecular information of high-uncertainty photosensitizer candidate molecules from the molecular set as sample molecular information for training the target model, the target model can focus on learning the more valuable sample molecular information, significantly improving model training efficiency. Furthermore, the trained target model can accurately and efficiently determine the physical parameters corresponding to a large number of photosensitizer candidate molecules, significantly improving photosensitizer screening efficiency.

[0120] In addition, during the training process of the target model, the real physical parameters of the sample molecular information can be quickly and accurately determined in batches through the trained correction model. The training of the correction model only requires a small amount of data. The cost of this process is significantly lower than the cost of determining the physical parameters using a high-precision method, and the overall training efficiency of the model is further improved.

[0121] It should be emphasized that, judging from the output results, both the calibration model and the target model are used to determine the physical parameters corresponding to the molecular information. However, there is a clear difference in the input between the two. That is, the calibration model requires the input of the molecular graph data and the single- and triple-state energy level differences corresponding to the molecular information, while the target model only requires the input of the molecular graph data corresponding to the molecular information.

[0122] This also determines that there are obvious differences in the usage scenarios of the two models. That is, the correction model needs to be used in model training to assist the training of the target model. Once the target model is trained, it can be used in subsequent processes to determine the true physical parameters of the molecular information. Whether it is existing molecular information or new molecular information, the target model is applicable.

[0123] From another perspective, since the correction model requires the first method to calculate the low-precision physical parameters of the molecular information, it is obviously not suitable for use in the process of predicting the real physical parameters of the molecular information.

[0124] It should also be noted that after the target model is trained, it can be deployed in multiple node devices. In this way, the server can run the target models on these node devices in parallel, and batch determine the real physical parameters corresponding to each photosensitizer candidate molecule in parallel, thereby significantly improving the efficiency of screening suitable photosensitizers.

[0125] The above is a model training method provided by one or more embodiments of the present application. Based on the same idea, the present application also provides a corresponding model training device, such as Figure 3 shown.

[0126] Figure 3 A schematic diagram of a model training device provided in an embodiment of the present application, specifically comprising:

[0127] An acquisition module 301 is configured to acquire a first molecule set, wherein the first molecule set includes molecular information of each photosensitizer candidate molecule;

[0128] A selection module 302 is configured to select molecular information of some photosensitizer candidate molecules from the first molecule set as first basic molecular information;

[0129] A first determination module 303 is configured to determine the uncertainty corresponding to each piece of first basic molecular information based on the predicted physical parameters respectively output by a plurality of target models to be trained for the first basic molecular information, wherein different target models may have different numbers of model parameters and / or network layers;

[0130] A selection module 304 is configured to select molecular information with an uncertainty not less than a preset threshold from each of the first basic molecular information and add the molecular information as sample molecular information to the training set;

[0131] The first training module 305 is configured to train at least a portion of the target model to be trained according to the determined real physical parameters and predicted physical parameters corresponding to the sample molecular information contained in the training set.

[0132] Optionally, the acquisition module 301 is specifically configured to acquire an initial molecule set; and remove redundant molecule information from the initial molecule set to obtain a first molecule set.

[0133] Optionally, the determination module 303 is specifically used to determine, for each first basic molecular information, the average predicted physical information corresponding to the first basic molecular information based on the predicted physical parameters output by several target models to be trained for the first basic molecular information; and determine the uncertainty corresponding to the first basic molecular information based on the difference between the average predicted physical information and the predicted physical parameters output by each target model to be trained for the first basic molecular information.

[0134] Optionally, the real physical parameters include: a first singlet excitation energy and a first triplet excitation energy;

[0135] The device further comprises:

[0136] The second determination module 306 is configured to determine, for each sample molecule information included in the training set, a first singlet excitation energy and a first triplet excitation energy corresponding to the sample molecule information at a first precision using a preset first method; determine, based on the first singlet excitation energy and the first triplet excitation energy corresponding to the sample molecule information at the first precision, a singlet-state and triplet-state energy level difference corresponding to the sample molecule information at the first precision; and input the sample molecule information and the singlet-state and triplet-state energy level difference corresponding to the sample molecule information at the first precision into a pre-trained calibration model, so that the calibration model outputs, as true physical parameters, physical parameters corresponding to the sample molecule information at a second precision, where the second precision is higher than the first precision.

[0137] Optionally, the device further comprises:

[0138] The second training module 307 is used to obtain a second molecular set, which contains molecular information of each photosensitizer candidate molecule; select molecular information of some photosensitizer candidate molecules from the second molecular set as second basic molecular information; use the first method to determine the single-three-state energy level difference corresponding to the second basic molecular information at the first precision; input the second basic molecular information and the single-three-state energy level difference corresponding to the second basic molecular information at the first precision into the calibration model to be trained, so that the calibration model to be trained outputs the predicted physical parameters corresponding to the second basic molecular information; train the calibration model to be trained based on the deviation between the predicted physical parameters corresponding to the second basic molecular information and the physical parameters corresponding to the second basic molecular information at the second precision, where the physical parameters corresponding to the second basic molecular information at the second precision are determined using a preset second method, and the cost of determining the physical parameters using the second method is higher than the cost of determining the physical parameters using the first method.

[0139] Optionally, the first training module 305 is specifically used to train at least part of the target model to be trained based on the real physical parameters and predicted physical parameters corresponding to the sample molecular information contained in the determined training set, and to filter the sample molecular information from the remaining molecular information in the first molecular set through the trained target model, and to add the filtered sample molecular information to the training set, so as to iteratively train at least part of the target model through the updated training set until a preset iteration stop condition is met.

[0140] Optionally, the device further comprises:

[0141] The screening module 308 is configured to determine the real physical parameters corresponding to the pre-stored molecular information of each photosensitizer candidate molecule using at least a portion of the trained target model; and to screen the molecular information for use as the photosensitizer molecule from the pre-stored molecular information of each photosensitizer candidate molecule based on the determined real physical parameters.

[0142] The present application also provides a computer-readable storage medium that stores a computer program that can be used to execute the above Figure 1 A model training method provided.

[0143] This manual also provides Figure 4 The one shown corresponds to Figure 1 Schematic diagram of the electronic equipment. Figure 4 As shown, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 The model training method described.

[0144] Of course, in addition to software implementation, this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0145] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages ​​and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0146] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.

[0147] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0148] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0149] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0150] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0151] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0152] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0153] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0154] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0155] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0156] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0157] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0158] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.

[0159] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

[0160] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.

Claims

1. A model training method, characterized in that: include: Acquire a first molecule set, wherein the first molecule set includes molecular information of each photosensitizer candidate molecule; Selecting molecular information of some photosensitizer candidate molecules from the first molecule set as first basic molecular information; Determining the uncertainty corresponding to each first basic molecular information according to the predicted physical parameters respectively output by several target models to be trained for the first basic molecular information, where different target models have different numbers of model parameters and / or network layers; Selecting molecular information with an uncertainty not less than a preset threshold from each first basic molecular information and adding it to the training set as sample molecular information; At least a portion of the target model to be trained is trained based on the determined real physical parameters and predicted physical parameters corresponding to the sample molecular information contained in the training set.

2. The method according to claim 1, wherein Get the first molecule set, including: Get the initial molecule set; Redundant molecular information in the initial molecular set is removed to obtain a first molecular set.

3. The method according to claim 1, wherein Determining the uncertainty corresponding to each piece of first basic molecular information based on the predicted physical parameters output by the target models to be trained for each piece of first basic molecular information specifically includes: For each piece of first basic molecular information, determine the average predicted physical information corresponding to the first basic molecular information according to the predicted physical parameters respectively output by the target models to be trained for the first basic molecular information; The uncertainty corresponding to the first basic molecular information is determined according to the difference between the average predicted physical information and the predicted physical parameters output by each target model to be trained for the first basic molecular information.

4. The method according to claim 1, wherein The real physical parameters include: first singlet excitation energy and first triplet excitation energy; Determining the actual physical parameters corresponding to the sample molecular information contained in the training set specifically includes: For each sample molecule information included in the training set, a preset first method is used to determine a first singlet excitation energy and a first triplet excitation energy corresponding to the sample molecule information at a first accuracy; determining, based on the first singlet excitation energy and the first triplet excitation energy corresponding to the sample molecule information at the first precision, a singlet-triplet energy level difference corresponding to the sample molecule information at the first precision; The sample molecule information and the single-three-state energy level difference corresponding to the sample molecule information at the first precision are input into a pre-trained calibration model, so that the calibration model outputs the physical parameters corresponding to the sample molecule information at a second precision as the true physical parameters, and the second precision is higher than the first precision.

5. The method according to claim 4, wherein Training the calibration model, specifically including: Acquire a second molecule set, wherein the second molecule set includes molecular information of each photosensitizer candidate molecule; Selecting molecular information of some photosensitizer candidate molecules from the second molecular set as second basic molecular information; Using the first method, determining the single-state and triple-state energy level differences corresponding to the second basic molecular information at the first precision; Inputting the second basic molecular information and the single-three-state energy level difference corresponding to the second basic molecular information at the first precision into the calibration model to be trained, so that the calibration model to be trained outputs the predicted physical parameter corresponding to the second basic molecular information; The correction model to be trained is trained based on the deviation between the predicted physical parameters corresponding to the second basic molecular information and the physical parameters corresponding to the second basic molecular information at the second precision. The physical parameters corresponding to the second basic molecular information at the second precision are determined using a preset second method. The cost of determining the physical parameters using the second method is higher than the cost of determining the physical parameters using the first method.

6. The method according to claim 1, wherein Training at least part of the target model to be trained according to the determined real physical parameters and predicted physical parameters corresponding to the sample molecular information contained in the training set, and training at least part of the target model to be trained specifically includes: Based on the real physical parameters and predicted physical parameters corresponding to the sample molecular information contained in the determined training set, at least part of the target model to be trained is trained, and through the trained target model, the sample molecular information is screened from the remaining molecular information in the first molecular set, and the screened sample molecular information is added to the training set, so as to iteratively train at least part of the target model through the updated training set until a preset iteration stop condition is met.

7. The method according to claim 1, wherein The method further comprises: Determining the real physical parameters corresponding to the pre-stored molecular information of each photosensitizer candidate molecule using at least a portion of the trained target model; According to the determined real physical parameters, the molecular information used as the photosensitizer molecule is screened from the pre-stored molecular information of each photosensitizer candidate molecule.

8. A model training device, characterized in that: include: an acquisition module, configured to acquire a first molecule set, wherein the first molecule set includes molecular information of each photosensitizer candidate molecule; A selection module, configured to select molecular information of some photosensitizer candidate molecules from the first molecule set as first basic molecular information; A first determination module is configured to determine the uncertainty corresponding to each piece of first basic molecular information based on the predicted physical parameters respectively output by a plurality of target models to be trained for the first basic molecular information, wherein different target models have different numbers of model parameters and / or network layers; A selection module is used to select molecular information with an uncertainty not less than a preset threshold from each first basic molecular information, and add it to the training set as sample molecular information; The first training module is used to train at least part of the target model to be trained according to the determined real physical parameters and predicted physical parameters corresponding to the sample molecular information contained in the training set.

9. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Luciferase inhibitor screening model construction method and luciferase inhibitor screening method

    CN112053741A

  • Model training method, model prediction method, molecule screening method and device thereof

    CN114187980A

  • Cooperative training method and device for evidence neural network model

    CN117010448A

  • Synthetic material selection method, material manufacturing method, synthetic material selection data structure and manufacturing method

    CN118160042A

  • Insulating gas molecule screening method and system based on neural network

    CN118351968A