Data processing method and device, equipment and storage medium

By constructing training data sets and determining target models, efficient prediction of the correlation between molecules and targets is achieved, the problem of inefficient efficiency in the existing technology is solved, and prediction accuracy and data processing efficiency are improved.

CN119943203APending Publication Date: 2025-05-06DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411997686.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is not efficient when predicting the correlation between molecules and targets, especially when processing large-scale data, requiring a large amount of computing resources and time, and lacking a complete prediction process.

Method used

By building the training data set, determining the target model, and using the trained target model to predict the correlation between molecules and targets, a fully automated process of database construction, model training and model application is realized.

Benefits of technology

The efficiency of data processing and the prediction accuracy of the correlation between molecules and targets is improved, and the problem of inefficiency in the prior art is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943203A_ABST
    Figure CN119943203A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method and device, equipment and a storage medium. The method includes: constructing a training data set based on a first group of molecules in a first molecule set, the training data set indicating a first group of relevance between the first group of molecules and a first group of targets, the first group of relevance being determined based on a docking tool; determining a target model matched with the training data set from a group of candidate models; based on the training data set, training the target model, so that a test result of the target model meets a preset condition, and the test result is determined based on a second group of molecules in the first molecule set; and determining a second set of associations between the second set of molecules and the second set of targets using the trained target model. The data processing efficiency can be effectively improved, and the accuracy of predicting the relevance between the molecule and the target spot is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to data processing methods, apparatuses, devices, and computer-readable storage media. Background Art

[0002] Predicting the association between molecules and targets based on computer simulation technology plays a key role in various fields, such as drug discovery and material science. How to efficiently and accurately predict the association between molecules and targets is crucial. Summary of the invention

[0003] In the first aspect of the present disclosure, a data processing method is provided. The method includes: constructing a training data set based on a first group of molecules in a first molecular set, the training data set indicating a first group of associations between the first group of molecules and a first group of targets, the first group of associations being determined based on a docking tool; determining a target model that matches the training data set from a set of candidate models; training the target model based on the training data set so that a test result of the target model satisfies a preset condition, the test result being determined based on a second group of molecules in the first molecular set; and determining a second group of associations between a second molecular set and a second group of targets using the trained target model.

[0004] In a second aspect of the present disclosure, a device for data processing is provided. The device includes: a construction module, configured to construct a training data set based on a first group of molecules in a first molecular set, the training data set indicating a first group of associations between the first group of molecules and a first group of targets, the first group of associations being determined based on a docking tool; a first determination module, configured to determine a target model that matches the training data set from a set of candidate models; a training module, configured to train the target model based on the training data set so that a test result of the target model satisfies a preset condition, the test result being determined based on a second group of molecules in the first molecular set; and a second determination module, configured to determine a second group of associations between a second molecular set and a second group of targets using the trained target model.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory, the at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit. When the instructions are executed by the at least one processing unit, the device executes the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.

[0007] In a fifth aspect of the present disclosure, a computer program product is provided, which includes computer executable instructions, and when the instructions are executed by a processor, the method according to the first aspect of the present disclosure is implemented.

[0008] It should be understood that the contents described in this content section are not intended to limit the key features or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;

[0011] Figure 2 A flowchart showing a data processing process according to some embodiments of the present disclosure is shown;

[0012] Figure 3 An example flow chart showing a data processing process according to some embodiments of the present disclosure;

[0013] Figure 4 A schematic structural block diagram of a device for data processing according to some embodiments of the present disclosure is shown;

[0014] Figure 5 A block diagram of an electronic device capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0016] It should be noted that the titles of any sections / subsections provided herein are not restrictive. Various embodiments are described throughout this article, and any type of embodiment may be included under any section / subsection. In addition, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0017] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.

[0018] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects are subject to the corresponding laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user knows and confirms. Accordingly, when implementing each embodiment of the present disclosure, the type, scope of use, usage scenario, etc. of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations. The specific notification and / or authorization method can vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this respect.

[0019] In this specification and the embodiments, if personal information processing is involved, it will be processed on the premise of having a legal basis (such as obtaining the consent of the subject of personal information, or it is necessary to perform a contract, etc.), and will only be processed within the scope of regulations or agreements. If a user refuses to process personal information other than the necessary information for basic functions, it will not affect the user's use of basic functions.

[0020] Traditionally, related technologies mainly rely on techniques such as molecular docking and molecular dynamics simulation to predict the association between molecules and targets. However, these methods require a lot of computing resources and time when the amount of data processed is too large, and are not efficient.

[0021] In addition, with the advancement of artificial intelligence technology, methods based on deep learning to predict the association between molecules and targets have begun to attract attention, but a complete prediction process that can predict the association between targets and molecules has not yet been proposed.

[0022] The embodiment of the present disclosure proposes a data processing scheme. According to the scheme, a training data set is constructed based on a first group of molecules in a first molecular set, the training data set indicates a first group of associations between the first group of molecules and a first group of targets, the first group of associations being determined based on a docking tool; a target model matching the training data set is determined from a set of candidate models; based on the training data set, the target model is trained so that the test result of the target model satisfies a preset condition, the test result being determined based on a second group of molecules in the first molecular set; and a second group of associations between a second molecular set and a second group of targets is determined using the trained target model.

[0023] Based on this approach, the embodiments of the present disclosure construct a set of fully automated processes that can achieve database construction, model training, and model application, effectively improving the efficiency of data processing and the prediction accuracy of the association between molecules and targets.

[0024] Example Environment

[0025] Figure 1 1 is a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Figure 1 As shown, the target model 130 - 1 with the pre-trained parameter values ​​and the target model 130 - 2 with the trained parameter values ​​may be collectively or individually referred to as the target model 130 . The target model 130 may be implemented or included in the electronic device 140 and / or the electronic device 150 .

[0026] exist Figure 1 In the environment 100 of the invention, it is desirable to train and use such a machine learning model (i.e., a target model 130) that is configured for a variety of application environments. For example, the model 130 can be used to determine the association between a molecule and a predetermined target based on the molecule and the predetermined target.

[0027] like Figure 1 As shown, a model training system may exist in electronic device 140, and a model application system may exist in electronic device 150. Of course, the model training system and the model application system may also exist in the same electronic device (not shown in the figure). Figure 1The upper part shows the process of the model training phase, and the lower part shows the process of the model application phase. Before training, the parameter values ​​of the target model 130 may have initial values, or may have pre-trained parameter values ​​obtained through a pre-training process. The target model 130-1 may be trained via forward propagation and back propagation, and the parameter values ​​of the target model 130-1 may be updated and adjusted during the training process. The target model 130-2 may be obtained after the training is completed. At this point, the parameter values ​​of the target model 130-2 have been updated, and based on the updated parameter values, the target model 130-2 may be used to implement the prediction task of the association between molecules and targets in the model application phase.

[0028] In the model training phase, the target model 130 can be trained based on the training data set 110 including a plurality of training samples 112 and using a model training system. Specifically, the training process can be iteratively performed using a large number of training samples. After the training is completed, the target model 130 can include knowledge about the task to be processed. In the model application phase, the target model 130 (at this time, the target model 130 has trained parameter values) can be used to perform the corresponding task.

[0029] In some embodiments, the model training system may also be composed of a construction part (e.g., a sample construction subsystem) for constructing training samples (e.g., for constructing training samples) and a training part (e.g., a training subsystem) for training models, thereby achieving the separation of sample construction and model training links. For example, the model training system may be composed of a group of devices, a part of which is used as the above-mentioned construction part, and another part of which is used as the above-mentioned training part. In some embodiments, the model training system may also use the same device to construct training samples and train models based on training samples. The present disclosure is not intended to be limiting in this regard.

[0030] To facilitate understanding, the following examples are given in which the model training system uses the same equipment to construct training samples and trains the model based on the training samples.

[0031] exist Figure 1 In the embodiment, the electronic device 140 and the electronic device 150 may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. The terminal device may involve any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. The server includes but is not limited to a mainframe, an edge computing node, a computing device in a cloud environment, and the like.

[0032] It should be understood that Figure 1 The components and arrangements in the environment 100 shown are merely examples, and a computing system suitable for implementing the exemplary implementation described in the present disclosure may include one or more different components, other components, and / or different arrangements. The implementation of the present disclosure is not limited in this respect.

[0033] Example Process

[0034] Figure 2 FIG. 2 is a flowchart of a process 200 of data processing according to some embodiments of the present disclosure. The process 200 may be implemented at the electronic device 140 and / or the electronic device 150. Figure 1 Describe the process 200. For ease of description, the process 200 is implemented in the electronic device 140 as an example for illustration.

[0035] In block 210 , the electronic device 140 constructs a training data set based on a first group of molecules in the first set of molecules, where the training data set indicates a first group of associations between the first group of molecules and a first group of targets, where the first group of associations are determined based on a docking tool.

[0036] In some embodiments, the molecule may be any appropriate chemical molecule, which may be composed of two or more atoms bonded together by chemical bonds. The target may be any appropriate specific biological molecule that can interact with the chemical molecule, which may include but is not limited to proteins, enzymes, receptors or other biological macromolecules. The interaction between the chemical molecule and the target may represent the non-covalent bond between the chemical molecule and the target or the formation of a covalent bond between the chemical molecule and the target, etc.

[0037] In some embodiments, the electronic device 140 can cluster the first set of molecules to determine multiple cluster centers. Specifically, the electronic device 140 can cluster the individual molecules in the first set of molecules based on a predetermined clustering algorithm to obtain multiple groups (or multiple clusters). For each group, the molecules included in the group have certain similarities in chemical structure or properties, and the group corresponds to a cluster center, which is the center point or representative point corresponding to the group. In some embodiments, the electronic device 140 can add multiple candidate molecules corresponding to multiple cluster centers to the first group of molecules.

[0038] In other embodiments, the electronic device 140 may also randomly select some molecules from the first set of molecules. Further, the electronic device 140 may perform deduplication processing on these molecules and the multiple candidate molecules corresponding to the multiple clustering centers to obtain target group molecules. Further, the electronic device 140 may extract a predetermined number of molecules from the target group molecules and add them to the first group of molecules. As an example, the electronic device 140 may extract a first predetermined ratio of molecules from the target group molecules to determine them as the first group of molecules, and the first predetermined ratio may be set as required, such as fifty percent, etc.

[0039] In some embodiments, the electronic device 140 may determine a first set of associations between the first set of molecules and the first set of targets based on the docking tool. The first set of associations may characterize the interactions between the first set of molecules and the first set of targets. The number of interactions between a molecule and a target may be one or more, which will not be described in detail herein.

[0040] Furthermore, the electronic device 140 can construct a training data set based on the first group of molecules, the first group of targets, and the first group of associations between the first group of molecules and the first group of targets, wherein the set of associations can also be referred to as a label corresponding to the training data set, and the combination of the first group of targets and any one of the targets in the first group of molecules and any one of the molecules can be referred to as a corresponding training sample in the training data set.

[0041] In order to ensure the accuracy and effectiveness of model training, in some embodiments, the electronic device 140 can determine the target satisfaction rate corresponding to the training data set based on the predetermined screening conditions and the training data set, wherein the target satisfaction rate can be used to characterize the proportion of training samples in the training data set that meet the predetermined screening conditions. The predetermined screening conditions can indicate at least one association constraint, such as the predetermined screening conditions that there is a correlation 1 and an interaction 2 between the molecule and the target.

[0042] Further, in response to the target satisfaction rate being greater than a threshold, the electronic device 140 may train the model with the training data set. In response to the target satisfaction rate being less than or equal to the threshold, the electronic device 140 may randomly select part of the data from the first molecular set to supplement the training data set until the satisfaction rate corresponding to the supplemented training data set is greater than the threshold.

[0043] In block 220 , the electronic device 140 determines a target model that matches the training data set from a set of candidate models.

[0044] Since different model structures may have different performances when processing data of different scales and qualities, in order to improve the accuracy of model prediction and the recall rate of the model, in some embodiments, the electronic device 140 can determine a target model that matches the number of training samples in the training data set from a group of candidate models. In some embodiments, this group of candidate models corresponds to different model structures.

[0045] In some embodiments, a set of candidate models may be any appropriate machine learning model that can be used to predict the association between molecules and targets, such as a multilayer perceptron, a graphical model, a language model, and the like.

[0046] As an example, the electronic device 140 may determine that the target model matching the training data set is a graph model in response to determining that the training data set corresponds to 100,000 training samples. As another example, the electronic device 140 may determine that the target model matching the training data set is a language model in response to determining that the training data set corresponds to 10 million training samples.

[0047] In other embodiments, the electronic device 140 may also determine a target model to be trained from a group of candidate models based on the number of training samples in the training data set and the model size corresponding to the group of candidate models.

[0048] In some embodiments, the number of molecules in the first group may be related to the computing resources used to train the target model. The computing resources may be any appropriate hardware resources and / or software resources required for training the target model, including but not limited to processors, memory, and the like.

[0049] In box 230, the electronic device 140 trains the target model based on the training data set so that the test result of the target model meets the preset condition, and the test result is determined based on the second group of molecules in the first molecule set.

[0050] In some embodiments, the electronic device 140 may input the training samples in the training data set into the target model to obtain a set of predicted associations between the first group of molecules and the first group of targets output by the target model. Further, the electronic device 140 may train the target model based on a comparison of a set of predicted associations between the first group of molecules and the first group of targets and a first set of associations between the first group of molecules and the first group of targets.

[0051] In some embodiments, in order to avoid overfitting of the target model, the electronic device may divide the second group of molecules in the first set of molecules into a verification data set and a test data set. Further, the electronic device 140 may verify the target model based on the verification data set in the second group of molecules in the first set of molecules to verify the prediction ability and performance of the target model on unknown data. During the verification process, the electronic device 140 may adjust the model weight corresponding to the target model based on the test results to determine the optimal model weight corresponding to the target model (when the target model corresponds to the optimal model weight, it can be confirmed that the target model training is complete).

[0052] Specifically, the validation data set indicates a set of annotated associations between the molecules included in the validation data set and a set of targets, and the set of annotated associations is determined based on the docking tool. Further, the electronic device 140 can use the target model to obtain a set of predicted associations between the molecules included in the validation data set output by the target model and the set of targets based on the validation training set. Further, the electronic device 140 can adjust the model weight of the target model based on the set of annotated associations and the set of predicted associations.

[0053] As an example, in order to ensure that the test data set is representative, the electronic device 140 can extract a second predetermined ratio of molecules from the target group molecules to determine as a verification data set, wherein the training data set is different from the verification data set. The second predetermined ratio can be set according to demand, such as 25 percent, etc.

[0054] Furthermore, the electronic device 140 may also test the target model based on the test data set to determine whether the test result of the target model meets a preset condition. The preset condition may be that the accuracy rate, recall rate, etc. meet a predetermined threshold.

[0055] Specifically, the test data set may indicate a first set of annotated associations between molecules included in the test data set and a set of targets, where the set of annotated associations is determined based on the docking tool. Further, the electronic device 140 may use the target model to obtain a set of predicted associations between molecules included in the test data set output by the target model and the set of targets based on the test training set. Further, the electronic device 140 may test whether the test result of the target model meets the preset conditions based on the set of annotated associations and the set of predicted associations.

[0056] As an example, in order to ensure that the test data set is representative, the electronic device 140 can also extract a third predetermined ratio of molecules from the target group molecules to determine as a test data set, wherein the test data set is different from the training data set and the verification data set. The third predetermined ratio can be set according to demand, such as 25 percent, etc.

[0057] It should be noted that the target group molecules are determined based on the first molecule set, and the process of determining the target group molecules has been described in the above embodiments and will not be repeated here.

[0058] At block 240 , the electronic device 140 determines a second set of associations between a second set of molecules and a second set of targets using the trained target model.

[0059] In some embodiments, the second molecular set may be the same as or different from the first molecular set, which will not be described in detail herein.

[0060] In some embodiments, the electronic device 140 may obtain a screening condition, and the screening condition may indicate at least one association constraint. As an example, the screening condition may indicate that correlation 1, interaction 2, and interaction 3 may be generated. Further, the electronic device 140 may determine a plurality of molecules that meet the screening condition from the second set of molecules based on the second set of associations.

[0061] As an example, if the screening condition can indicate that relevant interaction 1 and interaction 2 can be generated, the electronic device 140 can screen out multiple molecules that can generate interaction 1 and interaction 2 with target A from the second molecule set, and determine these multiple molecules as active molecules.

[0062] Figure 3 An example flow chart of a data processing process according to some embodiments of the present disclosure is shown. Figure 3 Provide explanation.

[0063] In block 301 , the electronic device 140 constructs a training data set.

[0064] As an example, the electronic device 140 can construct each training sample using the first molecule set and the predetermined target. Further, the electronic device 140 can determine the interaction corresponding to each training sample based on the docking tool, and determine the interaction corresponding to each training sample as a label of each training sample.

[0065] In block 302 , the electronic device 140 adaptively selects a target model based on the number of training samples in a training data set.

[0066] As an example, the electronic device 140 may determine a model that matches the training data set from Model 1, Model 2, ... Model n as the model to be trained based on the number of training samples in the training sample set.

[0067] In block 303 , the electronic device 140 trains a model based on a training data set.

[0068] As an example, the electronic device 140 may use the training data set to train the model to be trained, so that the model to be trained can learn the relationship between the interaction between molecules and targets.

[0069] At block 304 , the electronic device 140 performs a model test.

[0070] As an example, the electronic device 140 can test the model based on the validation data set to determine the optimal model weight of the model, and when the model corresponds to the optimal model weight, it can be determined that the model training is completed. Further, the electronic device 140 can also test the model based on the test data set to determine whether the recall rate, accuracy rate, etc. of the model meet the predetermined conditions.

[0071] In block 305 , the electronic device 140 performs model inference.

[0072] As an example, the electronic device 140 may use a model to determine at least one interaction between each molecule in the second molecule set and a predetermined target.

[0073] In block 306 , the electronic device 140 screens molecules satisfying the screening conditions from the large-scale molecule library.

[0074] As an example, the electronic device 140 may screen out target molecules that meet predetermined interaction conditions from the second molecule set, such as screening out molecules that have interaction 1 and interaction 2 with a predetermined target.

[0075] Based on this approach, the embodiments of the present disclosure construct a set of fully automated processes that can achieve database construction, model training, and model application, effectively improving the efficiency of data processing and the prediction accuracy of the association between molecules and targets.

[0076] The present disclosure can determine whether a molecule is an active molecule based on whether the predicted association between the molecule and the target satisfies the screening conditions.

[0077] Example devices and equipment

[0078] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 4 4 is a schematic structural block diagram of an apparatus 400 for data processing according to some embodiments of the present disclosure. The apparatus 400 may be implemented as or included in the electronic device 110 discussed above. Each module / component in the apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0079] like Figure 4As shown, the device 400 includes a construction module 410, which is configured to construct a training data set based on a first group of molecules in a first molecular set, wherein the training data set indicates a first group of associations between the first group of molecules and a first group of targets, and the first group of associations is determined based on a docking tool; a first determination module 420, which is configured to determine a target model that matches the training data set from a group of candidate models; a training module 430, which is configured to train the target model based on the training data set so that a test result of the target model meets a preset condition, and the test result is determined based on a second group of molecules in the first molecular set; and a second determination module 440, which is configured to determine a second group of associations between a second molecular set and a second group of targets using the trained target model.

[0080] In some embodiments, the first determination module 420 is further configured to: determine a target model that matches the number from a group of candidate models based on the number of training samples in the training data set.

[0081] In some embodiments, the apparatus 400 further comprises a first acquisition module configured to: acquire a screening condition indicating at least one association constraint; and a third determination module configured to: determine a plurality of molecules satisfying the screening condition from the second molecule set based on the second set of associations.

[0082] In some embodiments, the apparatus 400 further comprises a clustering module configured to: cluster the first set of molecules to determine a plurality of cluster centers; and an adding module configured to: add a plurality of candidate molecules corresponding to the plurality of cluster centers to the first group of molecules.

[0083] In some embodiments, the number of the first group of molecules is related to the computing resources used to train the target model.

[0084] In some embodiments, a set of candidate models corresponds to different model structures.

[0085] The units included in the device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units in the device 400 can be implemented at least in part by one or more hardware logic components. As an example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0086] Figure 51 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that Figure 5 The electronic device 500 shown is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown can be used to implement Figure 1 A receiving device 120 is shown.

[0087] like Figure 5 As shown, the electronic device 500 is in the form of a general electronic device. The components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 may be an actual or virtual processor and is capable of performing various processes according to a program stored in the memory 520. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the electronic device 500.

[0088] The electronic device 500 typically includes a plurality of computer storage media. Such media may be any accessible media that is accessible to the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 may be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 may be a removable or non-removable medium, and may include a machine-readable medium, such as a flash drive, a disk, or any other medium, which may be capable of being used to store information and / or data (e.g., training data for training) and may be accessed within the electronic device 500.

[0089] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 5 As shown in , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to the bus (not shown) by one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules that are configured to perform various methods or actions of various embodiments of the present disclosure.

[0090] The communication unit 540 implements communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented in a single computing cluster or multiple computing machines that can communicate through a communication connection. Therefore, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0091] The input device 550 may be one or more input devices, such as a mouse, a keyboard, a tracking ball, etc. The output device 560 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 may also communicate with one or more external devices (not shown) through the communication unit 540 as needed, such as a storage device, a display device, etc., communicate with one or more devices that allow a user to interact with the electronic device 500, or communicate with any device that allows the electronic device 500 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0092] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0093] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices, equipment, and computer program products implemented according to the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0094] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0095] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0096] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple implementations of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of a module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some implementations as replacements, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0097] The above descriptions of various implementations of the present disclosure are exemplary, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The selection of terms used herein is intended to best explain the principles of the implementations, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the various implementations disclosed herein.

Claims

1. A method for data processing, comprising: Based on a first group of molecules in the first set of molecules, constructing a training data set, the training data set indicating a first group of associations between the first group of molecules and a first group of targets, the first group of associations being determined based on a docking tool; Determine a target model matching the training data set from a set of candidate models; Based on the training data set, training the target model so that a test result of the target model satisfies a preset condition, wherein the test result is determined based on a second group of molecules in the first set of molecules; as well as Using the trained target model, a second set of associations between a second set of molecules and a second set of targets is determined.

2. The method according to claim 1, wherein determining a target model matching the training data set from a set of candidate models comprises: Based on the number of training samples in the training data set, the target model matching the number is determined from the set of candidate models.

3. The method according to claim 1, further comprising: Obtaining a screening condition, wherein the screening condition indicates at least one associativity constraint; as well as Based on the second set of associations, a plurality of molecules satisfying the screening condition are determined from the second set of molecules.

4. The method according to claim 1, further comprising: Clustering the first set of molecules to determine a plurality of cluster centers; as well as A plurality of candidate molecules corresponding to the plurality of cluster centers are added to the first group of molecules.

5. The method of claim 1, wherein the number of the first group of molecules is related to computing resources used to train the target model. The method of claim 1 , wherein the set of candidate models corresponds to different model structures.

7. A device for data processing, comprising: A construction module is configured to construct a training data set based on a first group of molecules in a first set of molecules, wherein the training data set indicates a first group of associations between the first group of molecules and a first group of targets, wherein the first group of associations is determined based on a docking tool; A first determination module is configured to determine a target model matching the training data set from a group of candidate models; A training module, configured to train the target model based on the training data set so that a test result of the target model satisfies a preset condition, wherein the test result is determined based on a second group of molecules in the first set of molecules; as well as The second determination module is configured to determine a second set of associations between a second set of molecules and a second set of targets using the trained target model.

8. An electronic device, comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 6 when executed by the at least one processing unit.

9. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 6 when executed by a processor.

10. A computer program product comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 6.