A model processing method, apparatus, device, storage medium, and program product

CN117219152BActive Publication Date: 2026-09-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310150166.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-06
Publication Date
2026-09-25
Estimated Expiration
2043-02-06

AI Technical Summary

Technical Problem

[0002]苗头化合物是指与受体存在初步结合可能性的配体;为确定苗头化合物,常常通过神经网络模型预测受体和配体之间的结合活性,并将预测出的结合活性大于活性阈值的配体确定为苗头化合物;然而,上述确定苗头化合物的神经网络模型,是基于受体样本与配体样本之间的结合活性训练得到的,而相比于待预测的受体,受体样本与配体样本对应的训练数据中常常存在噪声,影响了神经网络模型的准确度

Benefits of technology

[0032]本申请实施例至少具有以下有益效果:通过先获取待预测受体和一部分第一待预测配体集之间真实的第一活性标签集,并基于该第一活性标签集组成标准数据,以从训练数据集中确定出与该标准数据冲突的噪声数据集;再基于该噪声数据集优化初始预测模型,以将初始预测模型中从该噪声数据集学习到的信息移除,如此,优化出的目标预测模型能够准确地针对待预测受体进行活性预测;因此,能够提升神经网络模型的准确度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117219152B_ABST
    Figure CN117219152B_ABST
Patent Text Reader

Abstract

The application provides a model processing method and device, equipment, storage medium and program product; is applied to intelligent medical treatment, cloud technology, artificial intelligence, wisdom traffic and various active prediction scenes such as vehicle carries; the method comprises the following steps: training a to-be-trained model based on a training data set, obtaining an initial prediction model, wherein the to-be-trained model is a to-be-trained neural network model for predicting the binding activity of ligand and receptor;Obtain the first activity label set of the combination of the to-be-predicted receptor and the first to-be-predicted ligand set;The to-be-predicted receptor, the first to-be-predicted ligand set and the first activity label set are combined into a standard data set;From the training data set, obtain the noise data set conflicting with the standard data set;Optimize the initial prediction model based on the noise data set, obtain the target prediction model, wherein the target prediction model is used to predict the activity set of the combination of the to-be-predicted receptor and the second to-be-predicted ligand set.Through the application, the accuracy of the neural network model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to activity prediction technology in the field of artificial intelligence, and more particularly to a model processing method, apparatus, device, storage medium, and program product. Background Technology

[0002] A lead compound is a ligand that has the potential to initially bind to a receptor. To identify lead compounds, neural network models are often used to predict the binding activity between the receptor and the ligand, and ligands with predicted binding activity greater than the activity threshold are identified as lead compounds. However, the neural network models used to identify lead compounds are trained based on the binding activity between receptor and ligand samples. Compared to the receptor to be predicted, the training data corresponding to the receptor and ligand samples often contain noise, which affects the accuracy of the neural network model. Summary of the Invention

[0003] This application provides a model processing method, apparatus, device, storage medium, and program product that can improve the accuracy of neural network models.

[0004] The technical solution of this application embodiment is implemented as follows:

[0005] This application provides a model processing method, the method comprising:

[0006] The training model is trained based on the training dataset to obtain an initial prediction model, wherein the training model is a neural network model to be trained for predicting ligand and receptor binding activity.

[0007] Obtain the first set of active tags that bind to the first set of predicted receptors and the first set of predicted ligands;

[0008] The receptor to be predicted, the first set of ligands to be predicted, and the first set of active tags are combined into a standard dataset;

[0009] Obtain a noisy dataset that conflicts with the standard dataset from the training dataset;

[0010] The initial prediction model is optimized based on the noise dataset to obtain a target prediction model, wherein the target prediction model is used to predict the active set of the receptor to be predicted binding with the second set of ligands to be predicted.

[0011] This application provides a model processing apparatus, including:

[0012] The model training module is used to train the model to be trained based on the training dataset to obtain an initial prediction model, wherein the model to be trained is a neural network model to be trained for predicting ligand and receptor binding activity.

[0013] The tag acquisition module is used to acquire the first set of active tags that bind to the first set of predicted receptors and the first set of predicted ligands.

[0014] The data acquisition module is used to combine the receptor to be predicted, the first set of ligands to be predicted, and the first set of active tags into a standard dataset;

[0015] A noise determination module is used to obtain a noise dataset that conflicts with the standard dataset from the training dataset;

[0016] The model optimization module is used to optimize the initial prediction model based on the noise dataset to obtain a target prediction model, wherein the target prediction model is used to predict the active set of the receptor to be predicted binding with the second set of ligands to be predicted.

[0017] In this embodiment, the training dataset includes a receptor sample set and a ligand sample set. The noise determination module is further configured to: continuously learn the initial prediction model based on the standard dataset to obtain an intermediate prediction model; obtain a first prediction loss set combining the receptor sample set and the ligand sample set based on the intermediate prediction model; determine the conflict value of each receptor sample in the receptor sample set based on the difference between the first prediction loss set and the second prediction loss set, wherein the second prediction loss set is the loss set obtained by the initial prediction model for activity prediction of the receptor sample set and the ligand sample set; determine the noisy receptor set whose conflict value is greater than a conflict threshold from the receptor sample set; and determine the information corresponding to the noisy receptor set in the training dataset as the noisy dataset that conflicts with the standard dataset.

[0018] In this embodiment, the noise determination module is further configured to: predict a first predicted activity set of the receptor to be predicted and the first set of ligands to be predicted based on the initial prediction model; train the initial prediction model based on the difference between the first predicted activity set and the first set of active tags to obtain current model parameters; obtain first smoothing information between the current model parameters and the initial model parameters, wherein the initial model parameters are the model parameters of the initial prediction model, and the first smoothing information is used to retain the initial model parameters during continuous learning; and determine the intermediate prediction model by combining the current model parameters and the first smoothing information.

[0019] In this embodiment of the application, the noise determination module is further configured to obtain a first parameter difference between the current model parameters and the initial model parameters; obtain a first loss gradient of the initial prediction model for the training dataset; obtain parameter weights positively correlated with the first loss gradient; and fuse the first parameter difference and the parameter weights to obtain the first smoothing information.

[0020] In this embodiment of the application, the noise determination module is further configured to obtain the second loss gradient of the intermediate prediction model on the standard dataset; and obtain the conflict value that is negatively correlated with both the first loss gradient and the second loss gradient and positively correlated with the parameter weights.

[0021] In this embodiment of the application, the model optimization module is further configured to maximize the loss value of the initial prediction model for the noisy dataset to obtain the model parameters to be optimized; obtain the second parameter difference between the model parameters to be optimized and the initial model parameters; fuse the parameter weights with the second parameter difference to obtain second smoothing information; and combine the model parameters to be optimized and the second smoothing information to optimize the target prediction model.

[0022] In this embodiment of the application, the model optimization module is further configured to delete the noisy dataset from the training dataset to obtain a target training dataset; train the model to be trained based on the target training dataset to obtain the target prediction model; or, train the model to be trained by combining the target training dataset and the standard dataset to obtain the target prediction model.

[0023] In this embodiment of the application, the tag acquisition module is further configured to perform activity prediction on the predictable ligand library of the predictable receptor based on the initial prediction model to obtain a second predicted activity set; based on the second predicted activity set, to reverse the order of the predictable ligands in the predictable ligand library to obtain a predictable ligand sequence; and to determine a specified number of the predictable ligands selected sequentially from the predictable ligand sequence as the first predictable ligand set, wherein the second predictable ligand set consists of the predictable ligands in the predictable ligand library other than the first predictable ligand set.

[0024] In this embodiment of the application, the model optimization module is further configured to: predict a third predicted activity set of the receptor to be predicted binding to the second predicted ligand set based on the target prediction model; select a target predicted ligand set from the second predicted ligand set, and determine a target predicted activity set corresponding to the target predicted ligand set from the third predicted activity set; obtain a second activity tag set of the receptor to be predicted binding to the target predicted ligand set; obtain the activity difference between the second activity tag set and the target predicted activity set; and when the activity difference is lower than a specified difference, determine the activity set of the receptor to be predicted binding to the predicted ligand library by combining the first activity tag set, the second activity tag set, and the third predicted activity set.

[0025] In this embodiment of the application, the model optimization module is further configured to combine the receptor to be predicted, the target predicted activity set, and the second activity label set into a next standard dataset when the activity difference is equal to or higher than a specified difference; and optimize the target prediction model based on the datasets in the training dataset that conflict with the next standard dataset.

[0026] In this embodiment of the application, the training dataset further includes a third activity tag set combining the receptor sample set and the ligand sample set. The model training module is further configured to perform activity prediction on the receptor sample set and the ligand sample set based on the model to be trained, to obtain a fourth predicted activity set; determine a third predicted loss set based on the difference between the fourth predicted activity set and the third activity tag set; and train the model to be trained based on the third predicted loss set to obtain the initial prediction model.

[0027] This application provides an electronic device for model processing, including:

[0028] Memory is used to store executable instructions or computer programs.

[0029] The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the model processing method provided in the embodiments of this application.

[0030] This application provides a computer-readable storage medium storing computer-executable instructions or computer programs, which, when executed by a processor, implement the model processing method provided in this application.

[0031] This application provides a computer program product, including computer-executable instructions or a computer program, which, when executed by a processor, implements the model processing method provided in this application.

[0032] The embodiments of this application have at least the following beneficial effects: by first obtaining the real first active tag set between the receptor to be predicted and a portion of the first ligand set to be predicted, and forming standard data based on the first active tag set, the noisy dataset that conflicts with the standard data is determined from the training dataset; then, the initial prediction model is optimized based on the noisy dataset to remove the information learned from the noisy dataset in the initial prediction model. In this way, the optimized target prediction model can accurately predict the activity of the receptor to be predicted; therefore, the accuracy of the neural network model can be improved. Attached Figure Description

[0033] Figure 1 This is a schematic diagram of the architecture of the model processing system provided in the embodiments of this application;

[0034] Figure 2 This is one of the embodiments provided in this application. Figure 1 A schematic diagram of the server structure in the diagram;

[0035] Figure 3 This is a flowchart illustrating the model processing method provided in the embodiments of this application. Figure 1 ;

[0036] Figure 4 This is a flowchart illustrating the model processing method provided in the embodiments of this application. Figure 2 ;

[0037] Figure 5 This is a flowchart illustrating the model processing method provided in the embodiments of this application. Figure 3 ;

[0038] Figure 6 This is a flowchart illustrating the model processing method provided in the embodiments of this application. Figure 4 ;

[0039] Figure 7 This is a flowchart illustrating the model processing method provided in the embodiments of this application. Figure 5 ;

[0040] Figure 8 This is a flowchart illustrating the model processing method provided in the embodiments of this application. Figure 6 ;

[0041] Figure 9 This is an exemplary pharmaceutical manufacturing process diagram provided in an embodiment of this application;

[0042] Figure 10 This is an exemplary model correction diagram provided in an embodiment of this application. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0044] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0045] In the following description, the terms "first, second, third, fourth" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third, fourth" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0046] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0047] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0048] 1) Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0049] 2) Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It is used to study how computers simulate or implement human learning behavior to acquire new knowledge or skills; and to reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence. Its applications are widespread across all areas of artificial intelligence. Machine learning typically includes techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning.

[0050] 3) Artificial neural networks are mathematical models that mimic the structure and function of biological neural networks. Exemplary structures of artificial neural networks in this application include Graph Convolutional Networks (GCNs, a type of neural network for processing graph-structured data), Deep Neural Networks (DNNs), Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Neural State Machines (NSMs), and Phase-Functioned Neural Networks (PFNNs). The training model, initial prediction model, intermediate prediction model, and target prediction model involved in this application are all models corresponding to artificial neural networks.

[0051] 4) Receptors and ligands. Receptors are any biological macromolecules that can bind to hormones, neurotransmitters, drugs, or intracellular signaling molecules and cause changes in cell function, such as target proteins; ligands are biological small molecules that bind to receptors.

[0052] 5) Activity refers to the likelihood of receptors and ligands binding; generally speaking, the higher the activity, the higher the likelihood of binding, while the lower the activity, the lower the likelihood of binding.

[0053] 6) Wet experiments are biological experiments used to determine the binding activity of receptors and ligands.

[0054] 7) Virtual screening is a process of screening ligands for a receptor based on the activity of receptor-ligand binding; for example, receptor structure-based virtual screening (Structure Based Drug Discovery / Structure Based Virtual Screening, SBDD / SBVS), molecular structure-based virtual screening (LBDD), etc.

[0055] It should be noted that in the process of identifying lead compounds, neural network models are often used to predict the binding activity between receptors and ligands, and ligands with predicted binding activity greater than the activity threshold are identified as lead compounds. However, the neural network models used to identify lead compounds are trained based on the binding activity between receptor and ligand samples. Compared to the receptor to be predicted, the training data corresponding to the receptor and ligand samples often contain noise, such as incorrect activity labels, or the similarity between the receptor sample and the sample to be predicted (due to property differences or activity cliffs) being lower than the similarity threshold, which affects the accuracy of the neural network model.

[0056] Based on this, embodiments of this application provide a model processing method, apparatus, device, computer-readable storage medium, and computer program product, which can improve the accuracy of neural network models. The following describes exemplary applications of the electronic device for model processing (hereinafter referred to as the model processing device) provided in embodiments of this application. The model processing device provided in embodiments of this application can be implemented as various types of terminals such as smartphones, smartwatches, laptops, tablets, desktop computers, smart home appliances, set-top boxes, smart in-vehicle devices, portable music players, personal digital assistants, dedicated messaging devices, intelligent voice interaction devices, portable gaming devices, and smart speakers, or it can be implemented as a server. The following will describe exemplary applications when the model processing device is implemented as a server.

[0057] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of the model processing system provided in the embodiments of this application; as shown Figure 1 As shown, to support a model processing application, in the model processing system 100, terminal 200 (terminals 200-1 and 200-2 are shown as examples) connects to server 400 via network 300. Network 300 can be a wide area network (WAN), a local area network (LAN), or a combination of both. Additionally, the model processing system 100 also includes a database 500 for providing data support to server 400; and... Figure 1 The example shown illustrates a scenario where the database 500 is independent of the server 400. However, the database 500 can also be integrated into the server 400, and this embodiment does not limit this to any particular case.

[0058] Terminal 200 is used to receive, via network 300, the activity set of the receptor to be predicted binding with the ligand library to be predicted sent by server 400, and to display the activity set of the receptor to be predicted binding with the ligand library to be predicted (exemplary graphical interfaces 210-1 and 210-2 are shown (activity set, target + small molecule compound 1:9, target + small molecule compound 2:7, target + small molecule compound 3:6, ...)).

[0059] Server 400 is used to train a model to be trained based on a training dataset to obtain an initial prediction model, wherein the model to be trained is a neural network model to be trained for predicting ligand and receptor binding activity; to obtain a first set of activity tags for the binding of the receptor to be predicted and a first set of ligands to be predicted; to combine the receptor to be predicted, the first set of ligands to be predicted, and the first set of activity tags into a standard dataset; to obtain a noisy dataset that conflicts with the standard dataset from the training dataset; and to optimize the initial prediction model based on the noisy dataset to obtain a target prediction model, wherein the target prediction model is used to predict the activity set of the receptor to be predicted binding with a second set of ligands to be predicted. Server 400 is also used to obtain the activity set of the receptor to be predicted binding with a library of ligands to be predicted based on the target prediction model, and to send the activity set of the receptor to be predicted binding with the library of ligands to be predicted to terminal 200 via network 300.

[0060] In some embodiments, server 400 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals and servers can be connected directly or indirectly via wired or wireless communication, which is not limited in this embodiment.

[0061] See Figure 2 , Figure 2 This is one of the embodiments provided in this application. Figure 1 A schematic diagram of the server structure in the diagram; such as Figure 2 As shown, server 400 includes at least one processor 410, memory 450, and at least one network interface 420, and in some embodiments, a user interface 430. The various components in server 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 2 The general labeled all buses as Bus System 440.

[0062] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0063] User interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0064] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.

[0065] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.

[0066] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0067] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;

[0068] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, Wi-Fi, and Universal Serial Bus (USB), etc.

[0069] Presentation module 453 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with user interface 430;

[0070] The input processing module 454 is used to detect and translate one or more user inputs or interactions from one or more input devices 432.

[0071] In some embodiments, the model processing apparatus provided in this application can be implemented in software. Figure 2 A model processing device 455 stored in memory 450 is shown. This device can be software in the form of programs and plug-ins, and includes the following software modules: a model training module 4551, a label acquisition module 4552, a data acquisition module 4553, a noise determination module 4554, and a model optimization module 4555. These modules are logically connected and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.

[0072] In some embodiments, the model processing apparatus provided in this application can be implemented in hardware. As an example, the model processing apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the model processing method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0073] In some embodiments, the terminal or server can implement the model processing method provided in this application by running a computer program. For example, the computer program can be a native program or software module in an operating system; it can be a native application (APP), that is, a program that needs to be installed in the operating system to run, such as an activity prediction APP; it can also be a mini-program, that is, a program that only needs to be downloaded to a browser environment to run; or it can be a mini-program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of application, module or plugin.

[0074] The model processing method provided in this application will be described below with reference to exemplary applications and implementations of the model processing device provided in the embodiments of this application. Furthermore, the model processing method provided in the embodiments of this application is applicable to various activity prediction scenarios such as smart healthcare, cloud technology, artificial intelligence, smart transportation, and vehicle-mounted systems.

[0075] See Figure 3 , Figure 3 This is a flowchart illustrating the model processing method provided in the embodiments of this application. Figure 1 ,in, Figure 3 The execution entity for each step in the process is the model processing device; the following will combine... Figure 3 The steps shown are explained.

[0076] Step 101: Train the model to be trained based on the training dataset to obtain the initial prediction model.

[0077] In this embodiment of the application, when the model processing device obtains the training dataset and the model to be trained, it can trigger the model training process to train the model to be trained based on the training dataset; when the training ends, the obtained training result is the initial prediction model.

[0078] It should be noted that the training dataset is the dataset used to train the model to be trained. Each training data point in the training dataset includes a receptor sample, a ligand sample, and a third activity label indicating the binding of the receptor and ligand samples. The receptor sample is the receptor used to train the model, the ligand sample is the ligand used to train the model, and the third activity label is the activity label indicating the binding of the receptor and ligand samples. The model to be trained is a neural network model to be trained to predict ligand and receptor binding activity. It can be the original neural network model, a pre-trained neural network model, etc., and this embodiment does not limit this. The initial prediction model is the model to be trained after completion of training.

[0079] Step 102: Obtain the first active tag set that combines the receptor to be predicted and the first set of ligands to be predicted, and combine the receptor to be predicted, the first set of ligands to be predicted, and the first set of active tags to be predicted into a standard dataset.

[0080] In this embodiment, after obtaining the initial prediction model, the model processing device optimizes the initial prediction model based on the receptor to be predicted before performing activity prediction on the receptor to be predicted. When optimizing the initial prediction model, the model processing device first obtains the corresponding standard dataset based on the receptor to be predicted, and determines the noise in the training dataset based on the standard dataset. Thus, for the receptor to be predicted, the model processing device obtains a first set of ligands to be predicted, including multiple ligands to be predicted, and obtains the actual activity tags of the receptor to be predicted and each ligand to be predicted, referred to here as the first activity tags; here, all first activity tags are determined as the first activity tag set. Next, the model processing device combines the standard data based on the receptor to be predicted, the first set of ligands to be predicted, and the first activity tag set, and determines all the combined standard data as the standard dataset; wherein each piece of standard data in the standard dataset includes the receptor to be predicted, a first ligand to be predicted from the first set of ligands to be predicted, and a corresponding first activity tag from the first activity tag set.

[0081] It should be noted that the first predicted ligand is the ligand to be bound to the predicted receptor, while the first set of predicted ligands can be a subset of all ligands to be bound to the predicted receptor; the first active tag is the actual activity of the predicted receptor binding to the first predicted ligand, which can be obtained through wet experiments or other methods. Furthermore, there is a one-to-one correspondence between the first set of predicted ligands and the first set of active tags, and the first active tag corresponds to one predicted ligand and one predicted receptor within the first set of predicted ligands.

[0082] Step 103: Obtain the noisy dataset that conflicts with the standard dataset from the training dataset.

[0083] In this embodiment of the application, the model processing device compares the training dataset and the standard dataset to identify at least one piece of training data that conflicts with the standard dataset, and identifies the at least one piece of training data as a noisy dataset.

[0084] It should be noted that the conflict includes at least one of the following: label error, different activity prediction modes, and receptor property differences greater than the difference threshold; wherein, label error refers to the third activity label in the training dataset being incorrect, and different activity prediction modes refer to the activity deviation of receptors with the same ligand being greater than the deviation threshold when the receptor structures are similar. Here, the model processing device can determine the target standard dataset with an activity deviation greater than the deviation threshold from the standard dataset based on the activity prediction result (obtained by the initial prediction model to predict the activity of the standard dataset) and the first activity label set, and determine the training data in the training dataset that is similar to the target standard dataset (e.g., similar in features) as the noise dataset; the model processing device can also determine the noise dataset based on receptor samples in the training dataset whose property differences from the receptor to be predicted are greater than the difference threshold; the model processing device can also continuously learn on the standard dataset based on the initial prediction model, and determine the noise dataset based on the loss difference (or loss gradient difference) of the model obtained by continuous learning on the training dataset and the standard dataset; etc., the embodiments of this application do not limit this.

[0085] Step 104: Optimize the initial prediction model based on the noise dataset to obtain the target prediction model.

[0086] In this embodiment of the application, after the model processing device obtains the noise dataset, it can remove the information learned from the noise dataset in the initial prediction model to optimize the initial prediction model, so that the optimized initial prediction model is obtained based on the training data adapted to the receptor to be predicted; thereby, the targeting of the activity prediction of the receptor to be predicted can be improved; here, the optimized initial prediction model is the target prediction model.

[0087] It should be noted that optimization refers to removing information learned from the noisy dataset from the initial prediction model. The target prediction model is used to predict the active set of the receptor to be predicted binding with the second set of ligands to be predicted, so as to screen prospective ligands from the first set of ligands to be predicted and the second set of ligands to be predicted based on the predicted active set of the receptor to be predicted binding with the second set of ligands to be predicted, and the second set of ligands to be predicted is different from the first set of ligands to be predicted.

[0088] It is understandable that by first obtaining the true first active tag set between the receptor to be predicted and a portion of the first ligand set to be predicted, and then forming standard data based on this first active tag set, the noisy dataset that conflicts with the standard data can be identified from the training dataset; then, the initial prediction model can be optimized based on this noisy dataset to remove the information learned from the noisy dataset in the initial prediction model. In this way, the optimized target prediction model can accurately predict the activity of the receptor to be predicted; thus, the accuracy of the neural network model can be improved.

[0089] In this embodiment of the application, the training dataset includes a receptor sample set and a ligand sample set, wherein the receptor sample set refers to all receptor samples in the training dataset, and the ligand sample set refers to all ligand samples in the training dataset; in this case, see Figure 4 , Figure 4 This is a flowchart illustrating the model processing method provided in the embodiments of this application. Figure 2 ,in, Figure 4 The execution entity for each step in the process is the model processing device; such as Figure 4 As shown, Figure 3 Step 103 can be implemented through steps 1031 to 1035. That is, the model processing device obtains the noisy dataset that conflicts with the standard dataset from the training dataset, including steps 1031 to 1035. Each step is explained below.

[0090] Step 1031: Based on the standard dataset, continuously learn the initial prediction model to obtain an intermediate prediction model.

[0091] It should be noted that, since the initial prediction model is a neural network model, and neural network models learn information from the current data first, the model processing device updates the initial prediction model based on a standard dataset to enable the initial prediction model to learn information from the standard dataset. Here, the model processing device continuously learns from the standard dataset to update the initial prediction model, and the initial prediction model after continuous learning is the intermediate prediction model. Continuous learning refers to the process of learning new information while retaining the already learned information. Furthermore, the continuous learning process can be a one-time event or iterative one; this embodiment does not limit this.

[0092] Step 1032: Based on the intermediate prediction model, obtain the first prediction loss set combining the receptor sample set and the ligand sample set.

[0093] In this embodiment of the application, the model processing device performs activity prediction on the receptor sample set and ligand sample set in the training dataset to predict the binding activity between each receptor sample and each ligand sample, and determines the loss between each receptor sample and each ligand sample based on the difference between the predicted binding activity and the third activity tag, which is referred to here as the first prediction loss, and the full first prediction loss is determined as the first prediction loss set.

[0094] Step 1033: Based on the difference between the first prediction loss set and the second prediction loss set, determine the conflict value of each receptor sample in the receptor sample set.

[0095] In this embodiment, the model processing device can also obtain a second prediction loss set, which is the loss set obtained by the initial prediction model for activity prediction of the receptor sample set and ligand sample set; while the first prediction loss set is the loss set obtained by the intermediate prediction model for activity prediction of the receptor sample set and ligand sample set, as well as the information learned by the intermediate prediction model from the standard dataset; therefore, the loss corresponding to the dataset in the training dataset that conflicts with the standard dataset will be higher in the first prediction loss set than in the second prediction loss set; therefore, the model processing device obtains the difference between the first prediction loss set and the second prediction loss set, which is to obtain the loss difference between each receptor sample and each ligand sample; then, it obtains the loss difference set corresponding to each receptor sample and the full set of ligand samples, and determines the conflict value of each receptor sample based on the loss difference set; for example, the average value of the loss difference set is determined as the conflict value of the receptor sample, the mode of the loss difference set is determined as the conflict value of the receptor sample, and so on.

[0096] It should be noted that the loss obtained based on activity and activity tag in the embodiments of this application is the loss function value.

[0097] Step 1034: From the receptor sample set, determine the set of noisy receptors whose conflict value is greater than the conflict threshold.

[0098] In this embodiment of the application, the model processing device is provided with a conflict threshold, which is the minimum conflict value for determining whether there is a conflict with the standard dataset; thereby, the model processing device compares the conflict value of each receptor sample with the conflict threshold, and determines the receptor samples in the receptor sample set whose conflict value is greater than the conflict threshold as the noise receptor set.

[0099] Step 1035: Determine the information corresponding to the noise receptor set in the training dataset as the noise dataset that conflicts with the standard dataset.

[0100] In this embodiment of the application, after the model processing device obtains the noise receptor set, it retrieves the full training data corresponding to each noise receptor in the noise receptor set from the training dataset, and combines the full training data corresponding to each noise receptor in the noise receptor set into a noise dataset that conflicts with the standard dataset.

[0101] Understandably, for the receptor to be predicted, by obtaining the corresponding standard dataset and using this standard dataset to achieve continuous learning of the initial prediction model, and by identifying the noisy dataset that conflicts with the training dataset based on the loss difference between the model before and after continuous learning on the training dataset, data support is provided for the optimization of the initial prediction model, thereby improving the accuracy of the model's prediction.

[0102] In this embodiment of the application, step 1031, in which the model processing device continuously learns the initial prediction model based on a standard dataset to obtain an intermediate prediction model, includes: the model processing device first predicts a first predicted activity set based on the initial prediction model to predict the binding of the receptor to be predicted and the first set of ligands to be predicted; then, based on the difference between the first predicted activity set and the first active tag set, the initial prediction model is trained to obtain the current model parameters; next, the first smoothing information between the current model parameters and the initial model parameters is obtained; and finally, the intermediate prediction model is determined by combining the current model parameters and the first smoothing information.

[0103] It should be noted that the first predicted activity set is the activity set predicted by the initial prediction model for the receptor and the first predicted ligand set; the initial model parameters are the model parameters of the initial prediction model, and the first smoothing information is used to retain the initial model parameters during continuous learning. Furthermore, the current model parameters are the model parameters obtained after one training iteration of the initial prediction model. Here, the first smoothing information is also determined based on the current model parameters and the initial model parameters to reduce the deviation between the current model parameters and the initial model parameters, thereby retaining the information learned from the training dataset in the initial prediction model. In addition, the model processing device determines the combined result of the current model parameters and the first smoothing information as the model parameters after continuous learning, and uses these continuously learned model parameters as the model parameters of the intermediate prediction model, or iteratively continues to learn the continuously learned model parameters to obtain the intermediate prediction model.

[0104] In this embodiment, the first smoothing information is a regularization term for continuous learning. The process by which the model processing device obtains the first smoothing information between the current model parameters and the initial model parameters may include: the model processing device obtaining the first parameter difference between the current model parameters and the initial model parameters; obtaining the first loss gradient of the initial prediction model for the training dataset; obtaining the parameter weights that are positively correlated with the first loss gradient; and finally fusing the first parameter difference and the parameter weights to obtain the first smoothing information.

[0105] It should be noted that the model processing device can fuse the difference between the current model parameters and the initial model parameters with the transpose of that difference to obtain the first parameter difference. Furthermore, since a larger loss gradient indicates a higher importance of the model parameter, the parameter weights obtained by the model processing device are positively correlated with the importance of the initial prediction model parameters. Thus, when the initial model parameters change significantly, if the initial model parameters are more important, the first smoothing information is greater, which reduces the change in the initial model parameters and preserves the learned information in the initial model parameters.

[0106] In this embodiment of the application, after the model processing device continuously learns the initial prediction model based on the standard dataset in step 1031 to obtain the intermediate prediction model, and before the model processing device determines the set of noisy receptors with conflict values ​​greater than the conflict threshold from the receptor sample set in step 1034, the model processing method further includes: the model processing device obtaining the second loss gradient of the intermediate prediction model on the standard dataset; and obtaining the conflict value that is negatively correlated with both the first and second loss gradients and positively correlated with the parameter weights. That is, the model processing device can also determine the conflict value based on the loss gradient.

[0107] See Figure 5 , Figure 5This is a flowchart illustrating the model processing method provided in the embodiments of this application. Figure 3 ,in, Figure 5 The execution entity for each step in the process is the model processing device; such as Figure 5 As shown in the embodiments of this application, Figure 3 Step 104 can be implemented through steps 1041 to 1044; that is, the model processing device optimizes the initial prediction model based on the noise dataset to obtain the target prediction model, including steps 1041 to 1044. Each step is explained below.

[0108] Step 1041: Maximize the loss of the initial prediction model for the noisy dataset to obtain the model parameters to be optimized.

[0109] It should be noted that the initial prediction model is obtained by minimizing the loss of the training dataset. To ensure that the optimized initial prediction model is obtained by minimizing the loss of the training dataset excluding the noisy dataset, the loss of the initial prediction model for the noisy dataset is maximized, thus removing the information learned from the noisy dataset from the initial prediction model. Here, the model processing device adjusts the model parameters in the initial prediction model by maximizing the loss of the initial prediction model for the noisy dataset; the adjusted model parameters are the parameters to be optimized.

[0110] Step 1042: Obtain the second parameter difference between the parameters of the model to be optimized and the initial model parameters.

[0111] It should be noted that, in order to reduce the update magnitude of the initial model parameters, after obtaining the parameters of the model to be optimized, the model processing device also obtains the difference between the parameters of the model to be optimized and the initial model parameters, and determines the second parameter difference based on the obtained difference. For example, the model processing device can fuse the difference between the current model parameters and the initial model parameters with the transpose of the difference to obtain the second parameter difference.

[0112] Step 1043: Fuse the parameter weights with the difference in the second parameter to obtain the second smoothing information.

[0113] It should be noted that the process by which the model processing device fuses the parameter weights with the difference between the second parameter to obtain the second smoothing information is similar to the process by which it fuses the parameter weights with the difference between the first parameter to obtain the first smoothing information.

[0114] Step 1044: Combine the parameters of the model to be optimized and the second smoothing information to optimize the target prediction model.

[0115] It should be noted that the model processing device determines the optimized model parameters by combining the model parameters to be optimized and the second smoothing information, and uses the optimized model parameters as the model parameters of the target prediction model, or iteratively optimizes the optimized model parameters to obtain the target prediction model.

[0116] Understandably, optimizing the initial prediction model using a noisy dataset to obtain the target prediction model can improve the accuracy and efficiency of model optimization.

[0117] See Figure 6 , Figure 6 This is a flowchart illustrating the model processing method provided in the embodiments of this application. Figure 4 ,in, Figure 6 The execution entity for each step in the process is the model processing device; such as Figure 6 As shown in the embodiments of this application, Figure 6 Step 103 is followed by steps 105 and 106; that is, after the model processing device obtains the noisy dataset that conflicts with the standard dataset from the training dataset, the model processing method further includes steps 105 and 106, which are explained below.

[0118] Step 105: Remove noisy data from the training dataset to obtain the target training dataset.

[0119] In this embodiment, the model processing device can obtain a target prediction model by optimizing an initial prediction model, or by retraining the model to be trained using the training dataset after deleting the noisy dataset, etc., and this embodiment does not limit the scope of the application. When the model processing device retrains the model to be trained, it deletes the noisy dataset from the training dataset to obtain the dataset for retraining the model to be trained. The training dataset after deleting the noisy dataset is the target training dataset; that is, the dataset other than the noisy dataset in the training dataset is called the target training dataset.

[0120] Step 106: Train the model to be trained based on the target training dataset to obtain the target prediction model.

[0121] It should be noted that after the model processing device obtains the target training dataset for retraining the model to be trained, it trains the model to be trained based on the target training dataset and identifies the trained model as the target prediction model. Here, the process of training the model to be trained based on the target training dataset can be an iterative training process. Furthermore, the process of the model processing device training the model to be trained based on the target training dataset is similar to the process of training the model to be trained based on the training dataset.

[0122] In the embodiments of this application, the dataset used by the model processing device to retrain the model to be trained can also be the target training dataset and the standard dataset; thus, the model processing device can also combine the target training dataset and the standard dataset to train the model to be trained to obtain the target prediction model.

[0123] It is understandable that, since the target training dataset is adapted to the receptor to be predicted, retraining the model to be trained using the denoised target training dataset to obtain the target prediction model can improve the accuracy of the target prediction model.

[0124] See Figure 7 , Figure 7 This is a flowchart illustrating the model processing method provided in the embodiments of this application. Figure 5 ,in, Figure 7 The execution entity for each step in the process is the model processing device; such as Figure 7 As shown in the embodiments of this application, Figure 7 The method further includes steps 107 to 109 before step 102; that is, before the model processing device obtains the first active tag set of the receptor to be predicted and the first set of ligands to be predicted, the model processing method further includes steps 107 to 109, and each step is described below.

[0125] Step 107: Based on the initial prediction model, perform activity prediction on the library of ligands to be predicted for the receptor to be predicted, and obtain the second predicted activity set.

[0126] In this embodiment, all ligands to be bound to the predicted receptor constitute the predicted ligand library. The model processing device can select a first set of predicted ligands from the predicted ligand library based on an initial prediction model, or it can randomly select the first set of predicted ligands from the predicted ligand library, or it can be the combination described above, etc. This embodiment does not limit the specific choice. Here, when the model processing device selects a first set of predicted ligands from the predicted ligand library based on the initial prediction model, it first performs activity prediction on the predicted receptor and the predicted ligand library based on the initial prediction model. The predicted activity of the predicted receptor binding to each predicted ligand in the predicted ligand library is the second predicted activity, thus obtaining the second predicted activity set of the predicted receptor binding to the predicted ligand library.

[0127] Step 108: Based on the second predicted activity set, the ligands to be predicted in the ligand library are arranged in reverse order to obtain the ligand sequence to be predicted.

[0128] In this embodiment of the application, the model processing device can arrange the ligands to be predicted in the ligand library based on the second predicted activity set to obtain a sequence that changes monotonically based on the second predicted activity; here, when the arrangement is in reverse order, the obtained sequence is a sequence that decreases monotonically based on the second predicted activity, which is referred to as the ligand sequence to be predicted.

[0129] Step 109: Select a specified number of ligands to be predicted sequentially from the ligand sequences to be predicted, and determine them as the first set of ligands to be predicted.

[0130] In the embodiments of this application, the model processing device may sequentially select a specified number of ligands to be predicted from the ligand sequence to be predicted, and determine the selected specified number of ligands to be predicted as the first set of ligands to be predicted; it may also divide the ligand sequence to be predicted into stages, and select a specified number of ligands to be predicted from each stage to serve as the first set of ligands to be predicted; etc., the embodiments of this application do not limit this.

[0131] It should be noted that the ligand library to be predicted includes a first set of ligands to be predicted and a second set of ligands to be predicted. The second set of ligands to be predicted consists of the ligands to be predicted in the ligand library other than the first set of ligands to be predicted.

[0132] See Figure 8 , Figure 8 This is a flowchart illustrating the model processing method provided in the embodiments of this application. Figure 6 ,in, Figure 8 The execution entity for each step in the process is the model processing device; such as Figure 8 As shown in the embodiments of this application, Figure 8 Step 104 is followed by steps 110 to 114; that is, after the model processing device optimizes the initial prediction model based on the noise dataset to obtain the target prediction model, the model processing method further includes steps 110 to 114, which are explained below.

[0133] Step 110: Based on the target prediction model, predict the third predicted activity set of the receptor to be predicted and the second set of ligands to be predicted.

[0134] It should be noted that the activity of the target receptor binding to each second ligand in the second set of target ligands, as predicted by the model processing device based on the target prediction model, is called the third predicted activity. Thus, the third predicted activity set corresponding to the target receptor and the second set of target ligands can be obtained.

[0135] Step 111: Select a target set of ligands to be predicted from the second set of ligands to be predicted, and determine the target set of predicted activities corresponding to the target set of ligands to be predicted from the third set of predicted activities.

[0136] It should be noted that the target ligand set to be predicted is a partial set of ligands in the second set of ligands to be predicted. Here, the process by which the model processing device selects the target ligand set from the second set of ligands to be predicted is similar to the process of selecting the first set of ligands to be predicted from the ligand library, and will not be described again in this embodiment. Here, each target predicted activity in the target predicted activity set obtained by the model processing device represents the activity of the predicted receptor binding to one predicted ligand in the target predicted ligand set.

[0137] Step 112: Obtain the second set of active tags that bind to the target receptor and the target set of ligands.

[0138] It should be noted that the process of the model processing device acquiring the second set of active tags is similar to the process of acquiring the first set of active tags, and will not be described again in the embodiments of this application.

[0139] Step 113: Obtain the activity difference between the second set of active tags and the target predicted active set.

[0140] It should be noted that the second active label set and the target predicted active set are in one-to-one correspondence. The model processing device obtains the difference between each second active label and each target predicted active for the second active label set and the target predicted active set, thus obtaining the activity difference corresponding to the second active label set and the target predicted active set.

[0141] Step 114: When the activity difference is lower than the specified difference, combine the first activity tag set, the second activity tag set, and the third predicted activity set to determine the activity set of the receptor to be predicted and the ligand library to be predicted.

[0142] In this embodiment, the model processing device is configured with a specified difference, or the model processing device can obtain the specified difference from other devices (e.g., storage devices such as databases). The specified difference represents the minimum difference required to determine whether to perform model optimization. Thus, when the model processing device determines that the activity difference is lower than the specified difference, it indicates that the prediction accuracy of the target prediction model is high and model optimization has been completed. The model processing device then deletes the activity set corresponding to the second activity tag set from the third predicted activity set, and determines the deleted third predicted activity set, the first activity tag set, and the second activity tag set as the activity set for the binding of the receptor to be predicted with the ligand library to be predicted. Furthermore, based on the activity set for the binding of the receptor to be predicted with the ligand library to be predicted, a lead receptor is screened from the ligand library to be predicted.

[0143] It should be noted that the model processing device can end the model optimization after screening out the lead ligands that meet the quantity conditions from the ligand library to be predicted based on the first and second active tag sets.

[0144] See also Figure 8 Step 113 is followed by steps 115 and 116; that is, after the model processing device obtains the activity difference between the second activity tag set and the target predicted activity set, the model processing method further includes steps 115 and 116, which are explained below.

[0145] Step 115: When the activity difference is equal to or higher than the specified difference, combine the receptor to be predicted, the target predicted activity set, and the second activity tag set into the next standard dataset.

[0146] It should be noted that when the model processing device determines that the activity difference is equal to or higher than the specified difference, it indicates that the prediction accuracy of the target prediction model is low, and the model processing device continues to optimize the target prediction model. Here, the process of the model processing device optimizing the target prediction model is similar to the process of optimizing the initial prediction model; therefore, the model processing device first obtains the next standard dataset for further determining the noise in the training dataset; wherein, the process of obtaining the next standard dataset is similar to the process of obtaining the standard dataset, and will not be described again in this embodiment.

[0147] Step 116: Optimize the target prediction model based on the datasets in the training dataset that conflict with the next standard dataset.

[0148] It should be noted that the process by which the model processing device determines the dataset in the training dataset that conflicts with the next standard dataset is similar to the process by which it determines the noisy dataset in the training dataset that conflicts with the standard dataset; and the process by which the model processing device optimizes the target prediction model based on the dataset in the training dataset that conflicts with the next standard dataset is similar to the process by which it optimizes the initial prediction model based on the noisy dataset; these will not be described again in the embodiments of this application.

[0149] Understandably, the model processing device determines whether to iteratively optimize the model based on the prediction accuracy of the target prediction model, thereby improving the model's generalization ability.

[0150] In this embodiment of the application, the training dataset also includes a third activity tag set combining the receptor sample set and the ligand sample set; in step 101, the model processing device trains the model to be trained based on the training dataset to obtain an initial prediction model, including: the model processing device first performs activity prediction on the receptor sample set and the ligand sample set based on the model to be trained to obtain a fourth predicted activity set; and determines a third predicted loss set based on the difference between the fourth predicted activity set and the third activity tag set; then trains the model to be trained based on the third predicted loss set to obtain an initial prediction model.

[0151] It should be noted that the binding activity of the receptor and ligand samples predicted by the model processing device based on the model to be trained is called the fourth predicted activity. Therefore, when the model processing device first predicts the activity of the receptor and ligand sample sets based on the model to be trained, it can obtain the fourth predicted activity set corresponding to the receptor and ligand sample sets. Here, the process of the model processing device training the model to be trained based on the third predicted loss set can be a single step or iterative steps; this embodiment does not limit this.

[0152] It should be noted that the neural network model in this application embodiment can be a virtual screening model based on the receptor structure; in addition, the neural network model in this application embodiment can be a graph neural network model, used to combine the graph structure information corresponding to the receptor to be predicted and the graph structure information corresponding to the receptor to be predicted to perform activity prediction.

[0153] The following describes an exemplary application of the embodiments of this application in a practical application scenario. This exemplary application describes how, after an initial prediction model is trained using training targets (referred to as receptor samples), when activity prediction is performed on a target target (referred to as the sample to be predicted), training targets that conflict with the target target are identified based on a dataset consisting of the actual binding activities of the target target and multiple small molecule compounds. Then, the initial prediction model is corrected based on the training targets that conflict with the target target to optimize the prediction accuracy of the initial prediction model.

[0154] See Figure 9 , Figure 9 This is an exemplary pharmaceutical manufacturing process diagram provided in an embodiment of this application; as shown... Figure 9 As shown, this exemplary pharmaceutical process describes a pharmaceutical process based on an AI-powered intelligent drug platform, including target identification (9-11), lead compound acquisition (9-12), pilot compound acquisition (9-13), candidate compound acquisition (9-14), and clinical trials (9-15). Target identification (9-11) is achieved through protein structure prediction (9-21); lead compound acquisition (9-12) can be achieved through virtual screening (including virtual screening based on target structure (9-221) and virtual screening based on molecular structure (9-222)); pilot compound acquisition (9-13) can be achieved through property prediction (properties such as absorption, distribution, metabolism, excretion, and toxicity of molecules in the organism, etc.) (9-23); and candidate compound acquisition (9-14) can be achieved through synthetic route planning (9-24). Additionally, molecule generation (9-25) is used for lead compound acquisition (9-12) and pilot compound acquisition (9-13). The model corrected in this application's embodiments can improve the efficiency and accuracy of lead compound acquisition (9-12). The model correction process is described below.

[0155] See Figure 10 , Figure 10This is an exemplary model correction diagram provided in an embodiment of this application; as shown... Figure 10 As shown, this exemplary model correction process includes model training 10-1, activity prediction 10-2, real activity acquisition 10-3, conflict set determination 10-4, and model information removal 10-5. Each stage is described below.

[0156] First, in the 10-1 stage of model training, based on the training dataset... Training Model (referred to as the model to be trained); where p i For the i-th target site (called the receptor sample), x j For the j-th small molecule compound (called the ligand sample), y ij For target p i and small molecule compound x j The active tag (referred to as the third active tag), model It can be a structure-based drug discovery (SBDD) model, taking target p and small molecule compound x as inputs and outputting activity. For the model The model parameters. Additionally, the model... It can be a graph neural network, such as an attention-based graph neural network model (AttentiveFP).

[0157] It should be noted that the loss function l is used to train the model. When the loss function l is the mean squared error (MSE), it is as shown in formula (1).

[0158] l(f(p i x j ;θ1), y ij )=||y ij -f(p i x j ;θ1)|| 2 (1);

[0159] Among them, l(f(p) i x j ;θ1), y ij ) represents the loss function value (called the third prediction loss), f(p) i x j ;θ1) is the predicted target point p i and small molecule compound x j The activity; θ1 is the model The model parameter variables during training are initially set to... And it changes by continuously learning from the information in the training dataset during the training process.

[0160] Based on formula (1), the model The training process can be achieved through formula (2), as shown in formula (2).

[0161]

[0162] Where, θ o For the model The model parameters, and the model Based on training dataset Training Model Obtained; Used to obtain training dataset The expected value of each loss function value. Used to obtain the loss function value in the training dataset The model parameters corresponding to the minimum value.

[0163] Secondly, in the 10⁻² activity prediction stage, targeting the target site... and small molecule compound collection (referred to as the ligand library to be predicted), using the model Predict target points and each small molecule compound activity (referred to as the second predicted activity); where N is the set of small molecule compounds. The number of small and medium-sized molecule compounds.

[0164] Furthermore, in the 10⁻³ stage of obtaining true activity, based on activity... Small molecule compound collection Select activity The highest C (called the specified number) of small molecule compounds yields the set of small molecule compounds. (This is referred to as the first set of ligands to be predicted); where C is, for example, any integer from 5 to 20, and can be determined based on the efficiency of obtaining and consuming the actual activity. Here, the set of small molecule compounds can be obtained through wet experiments. The true active set (referred to as the first active tag set), and the small molecule compounds are grouped together. and the true active set Composition of feedback dataset (This is called the standard dataset).

[0165] Then, in the conflict set determination 10⁻⁴ phase, the training dataset is determined. In and feedback dataset A dataset with conflicts is called a conflict set (or a noisy dataset); where conflict can refer to the training dataset. The error labels in the training dataset can also refer to the training dataset. The similarity between the target in the dataset and the target in the feedback dataset is lower than the similarity threshold, etc.

[0166] It should be noted that, due to the model As a deep learning model, it prioritizes learning information from the current data during the learning process, thereby utilizing the feedback dataset. Update model At that time, the model Prioritize learning feedback datasets Information in the model; if the model is in the learning process Information and models to be learned If there are conflicts in the information already learned, the model will be modified. The information already learned that conflicts with the information to be learned. Here, the update process can be implemented using the Elastic Weight Consolidation (EWC) algorithm, as shown in Equation (3).

[0167]

[0168] in, For the model The model parameters, and the model It employs a continuous learning algorithm and is based on a feedback dataset. Update model Obtained; Used to obtain feedback dataset The expected value of each loss function value. Used to obtain the loss function value in the feedback dataset Minimize the model parameters to make the model Learning feedback dataset The information in the middle. (referred to as the first smoothing information) is the regularization term, and θ2 is the model. The model parameter variables during the update process are initially set to θ. o And it changes by continuously learning information from the training dataset during the update process; F is the loss function value l(f(p) i x j ;θ0), y ij The Fisher Information Matrix (called parameter weights) of θ o The importance of the parameters is positively correlated, as shown in formula (4).

[0169]

[0170] in, This represents the gradient of the loss function value.

[0171] It should be noted that, due to the model In the model Based on the learning feedback dataset The information obtained from the model By learning from the training dataset The information was obtained from the training dataset, therefore, for the training dataset In and feedback dataset Conflict datasets, models loss function value (referred to as the first prediction loss) is higher than the model. The loss function value l(f(p) i x ij ;θ0), y ij (This is called the second prediction loss), which is shown in formula (5).

[0172]

[0173] Therefore, the target point p is calculated. i and small molecule compound x j The difference in loss function values ​​r ij As shown in formula (6).

[0174]

[0175] Based on r ij Calculate target point p i With training dataset The average difference r(p) in the loss function values ​​of all small molecule compounds i (referred to as the conflict value), as shown in formula (7).

[0176]

[0177] Here, based on r(p) i The training dataset in reverse order. In the process, select a specified percentage (e.g., 20%) of the target points sequentially; and then use the training dataset. The data corresponding to the selected target points are identified as the conflict set.

[0178] In addition, target p i and small molecule compound x j The difference in loss function values ​​r ijIt can also be obtained through the influence function, as shown in formula (8).

[0179]

[0180] in, This is the second loss gradient. This is the first loss gradient.

[0181] Finally, in the model information removal 10⁻⁵ stage, information is used to extract data from the model. Lieutenant General in Conflict The learned information is removed. Here, the continuous learning algorithm is used in reverse to remove the model information, as shown in formula (9).

[0182]

[0183] Among them, the conflict set was obtained Then, in order to make the corrected model This is achieved by minimizing the training set excluding the conflict set. The loss function value corresponding to the target training dataset is obtained, while the model... By minimizing Obtain, and minimize, the training set excluding the conflict set. The loss function value can be obtained through Obtain; therefore, by minimizing It is possible to obtain from the model Lieutenant General in Conflict Removing the learned information is equivalent to maximizing the conflict set. The corresponding loss function value is expressed as θ3 is the model The model parameter variables in the information removal process are initially θ. o And it changes by continuously learning information from the dataset during the removal process; This is the second smoothing information mentioned above.

[0184] Subsequently, the model was adopted. (Referred to as the target prediction model) Perform activity prediction 10-2 and real activity acquisition 10-3 again. If the activity effect of small molecule compounds is outstanding in the wet experiment (for example, a preset number of lead small molecule compounds have been obtained), the virtual screening process ends. Otherwise, perform activity prediction 10-2, real activity acquisition 10-3, conflict set determination 10-4 and model information removal 10-5 again.

[0185] Furthermore, in the 10⁻⁵ stage of model information removal, it is also possible to use training sets other than the conflict set. The model is then retrained as shown in Equation (10).

[0186]

[0187] Understandably, once an activity prediction model is obtained... Subsequently, based on the true activity correction model of the target site and small molecule compound, the prediction accuracy of the corrected model can be improved. For example, on the test dataset, the corrected model obtained by using the model processing method provided in the embodiments of this application can improve the accuracy of activity prediction by more than 3% compared with the model before correction. Therefore, it is shown that the corrected model obtained by using the model processing method provided in the embodiments of this application can improve the accuracy of activity prediction of small molecule and target site binding, thereby improving the screening efficiency of lead compounds.

[0188] The following description continues to illustrate the exemplary structure of the model processing device 455 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the model processing device 455 of the memory 450 may include:

[0189] The model training module 4551 is used to train the model to be trained based on the training dataset to obtain an initial prediction model, wherein the model to be trained is a neural network model to be trained for predicting ligand and receptor binding activity.

[0190] Tag acquisition module 4552 is used to acquire a first set of active tags that bind to the first set of predicted receptors and the first set of predicted ligands;

[0191] Data acquisition module 4553 is used to combine the receptor to be predicted, the first set of ligands to be predicted, and the first set of active tags into a standard dataset;

[0192] The noise determination module 4554 is used to obtain the noise dataset that conflicts with the standard dataset from the training dataset;

[0193] The model optimization module 4555 is used to optimize the initial prediction model based on the noise dataset to obtain a target prediction model, wherein the target prediction model is used to predict the active set of the receptor to be predicted binding with the second set of ligands to be predicted.

[0194] In this embodiment, the training dataset includes a receptor sample set and a ligand sample set. The noise determination module 4554 is further configured to: continuously learn the initial prediction model based on the standard dataset to obtain an intermediate prediction model; obtain a first prediction loss set combining the receptor sample set and the ligand sample set based on the intermediate prediction model; determine the conflict value of each receptor sample in the receptor sample set based on the difference between the first prediction loss set and the second prediction loss set, wherein the second prediction loss set is the loss set obtained by the initial prediction model for activity prediction of the receptor sample set and the ligand sample set; determine the noisy receptor set whose conflict value is greater than a conflict threshold from the receptor sample set; and determine the information corresponding to the noisy receptor set in the training dataset as the noisy dataset that conflicts with the standard dataset.

[0195] In this embodiment of the application, the noise determination module 4554 is further configured to: predict a first predicted activity set of the receptor to be predicted and the first set of ligands to be predicted based on the initial prediction model; train the initial prediction model based on the difference between the first predicted activity set and the first set of active tags to obtain current model parameters; obtain first smoothing information between the current model parameters and the initial model parameters, wherein the initial model parameters are the model parameters of the initial prediction model, and the first smoothing information is used to retain the initial model parameters during continuous learning; and determine the intermediate prediction model by combining the current model parameters and the first smoothing information.

[0196] In this embodiment of the application, the noise determination module 4554 is further configured to obtain a first parameter difference between the current model parameters and the initial model parameters; obtain a first loss gradient of the initial prediction model for the training dataset; obtain parameter weights positively correlated with the first loss gradient; and fuse the first parameter difference and the parameter weights to obtain the first smoothing information.

[0197] In this embodiment of the application, the noise determination module 4554 is further configured to obtain the second loss gradient of the intermediate prediction model on the standard dataset; and obtain the conflict value that is negatively correlated with both the first loss gradient and the second loss gradient and positively correlated with the parameter weights.

[0198] In this embodiment of the application, the model optimization module 4555 is further configured to maximize the loss value of the initial prediction model for the noisy dataset to obtain the model parameters to be optimized; obtain the second parameter difference between the model parameters to be optimized and the initial model parameters; fuse the parameter weights with the second parameter difference to obtain second smoothing information; and combine the model parameters to be optimized and the second smoothing information to optimize the target prediction model.

[0199] In this embodiment of the application, the model optimization module 4555 is further configured to delete the noisy dataset from the training dataset to obtain a target training dataset; train the model to be trained based on the target training dataset to obtain the target prediction model; or, train the model to be trained by combining the target training dataset and the standard dataset to obtain the target prediction model.

[0200] In this embodiment of the application, the tag acquisition module 4552 is further configured to perform activity prediction on the predictable ligand library of the predictable receptor based on the initial prediction model to obtain a second predicted activity set; based on the second predicted activity set, to reverse the order of the predictable ligands in the predictable ligand library to obtain a predictable ligand sequence; and to determine a specified number of the predictable ligands selected sequentially from the predictable ligand sequence as the first predictable ligand set, wherein the second predictable ligand set consists of the predictable ligands in the predictable ligand library other than the first predictable ligand set.

[0201] In this embodiment of the application, the model optimization module 4555 is further configured to: predict a third predicted activity set of the receptor to be predicted binding to the second predicted ligand set based on the target prediction model; select a target predicted ligand set from the second predicted ligand set, and determine a target predicted activity set corresponding to the target predicted ligand set from the third predicted activity set; obtain a second activity tag set of the receptor to be predicted binding to the target predicted ligand set; obtain the activity difference between the second activity tag set and the target predicted activity set; and when the activity difference is lower than a specified difference, determine the activity set of the receptor to be predicted binding to the predicted ligand library by combining the first activity tag set, the second activity tag set, and the third predicted activity set.

[0202] In this embodiment of the application, the model optimization module 4555 is further configured to combine the receptor to be predicted, the target predicted activity set, and the second activity label set into a next standard dataset when the activity difference is equal to or higher than a specified difference; and optimize the target prediction model based on the datasets in the training dataset that conflict with the next standard dataset.

[0203] In this embodiment of the application, the training dataset further includes a third activity tag set combining the receptor sample set and the ligand sample set. The model training module 4551 is further configured to perform activity prediction on the receptor sample set and the ligand sample set based on the model to be trained, to obtain a fourth predicted activity set; determine a third predicted loss set based on the difference between the fourth predicted activity set and the third activity tag set; and train the model to be trained based on the third predicted loss set to obtain the initial prediction model.

[0204] This application provides a computer program product, which includes computer-executable instructions or a computer program stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions or computer program from the computer-readable storage medium and executes the computer-executable instructions or computer program, causing the electronic device to perform the model processing method described above in this application.

[0205] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the model processing method provided in this application. For example, ... Figure 3 The model processing method is shown.

[0206] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0207] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0208] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0209] As an example, computer-executable instructions can be deployed to execute on a single electronic device (in which case, this single electronic device is the model processing device), or to execute on multiple electronic devices located at one location (in which case, the multiple electronic devices located at one location are the model processing devices), or to execute on multiple electronic devices distributed across multiple locations and interconnected via a communication network (in which case, the multiple electronic devices distributed across multiple locations and interconnected via a communication network are the model processing devices).

[0210] It is understood that, in the embodiments of this application, data related to receptors, ligands, and active tags are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0211] In summary, the embodiments of this application first obtain a true first active tag set between the receptor to be predicted and a portion of the first ligand set to be predicted, and then form standard data based on the first active tag set to identify noisy datasets that conflict with the standard data from the training dataset; then optimize the initial prediction model based on the noisy dataset to remove the information learned from the noisy dataset in the initial prediction model. In this way, the optimized target prediction model can accurately predict the activity of the receptor to be predicted; therefore, it can improve the accuracy of the neural network model.

[0212] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A model processing method, characterized in that, The method includes: The training model is trained based on the training dataset to obtain an initial prediction model, wherein the training model is a neural network model to be trained for predicting ligand and receptor binding activity. Obtain the first set of active tags that bind to the first set of predicted receptors and the first set of predicted ligands; The receptor to be predicted, the first set of ligands to be predicted, and the first set of active tags are combined into a standard dataset; Obtain a noisy dataset that conflicts with the standard dataset from the training dataset; Maximize the loss value of the initial prediction model for the noisy dataset to obtain the model parameters to be optimized; obtain the second parameter difference between the model parameters to be optimized and the initial model parameters; fuse the parameter weights with the second parameter difference to obtain the second smoothing information; combine the model parameters to be optimized and the second smoothing information to optimize the target prediction model, wherein the target prediction model is used to predict the active set of the receptor to be predicted binding to the second set of ligands to be predicted.

2. The method according to claim 1, characterized in that, The training dataset includes a receptor sample set and a ligand sample set. The step of obtaining a noisy dataset that conflicts with the standard dataset from the training dataset includes: Based on the standard dataset, the initial prediction model is continuously learned to obtain an intermediate prediction model; Based on the intermediate prediction model, a first prediction loss set combining the receptor sample set and the ligand sample set is obtained; Based on the difference between the first prediction loss set and the second prediction loss set, the conflict value of each receptor sample in the receptor sample set is determined, wherein the second prediction loss set is the loss set obtained by the initial prediction model for activity prediction of the receptor sample set and the ligand sample set; From the receptor sample set, determine the set of noisy receptors whose conflict values ​​are greater than the conflict threshold; The information corresponding to the noise receptor set in the training dataset is used to determine the noise dataset that conflicts with the standard dataset.

3. The method according to claim 2, characterized in that, The step of continuously learning the initial prediction model based on the standard dataset to obtain an intermediate prediction model includes: Based on the initial prediction model, a first predicted activity set is predicted to be formed by the binding of the receptor to be predicted and the first set of ligands to be predicted. Based on the difference between the first predicted activity set and the first activity label set, the initial prediction model is trained to obtain the current model parameters; Obtain first smoothing information between the current model parameters and the initial model parameters, wherein the initial model parameters are the model parameters of the initial prediction model, and the first smoothing information is used to retain the initial model parameters during the continuous learning process; The intermediate prediction model is determined by combining the current model parameters and the first smoothing information.

4. The method according to claim 3, characterized in that, The step of obtaining the first smoothing information between the current model parameters and the initial model parameters includes: Obtain the first parameter difference between the current model parameters and the initial model parameters; Obtain the first loss gradient of the initial prediction model for the training dataset; Obtain the parameter weights that are positively correlated with the first loss gradient; The first smoothing information is obtained by combining the first parameter difference and the parameter weight.

5. The method according to claim 2, characterized in that, After continuously learning the initial prediction model based on the standard dataset to obtain an intermediate prediction model, and before determining the set of noisy receptors with conflict values ​​greater than the conflict threshold from the receptor sample set, the method further includes: Obtain the second loss gradient of the intermediate prediction model on the standard dataset; Obtain the conflict value that is negatively correlated with both the first loss gradient and the second loss gradient, and positively correlated with the parameter weights.

6. The method according to any one of claims 1 to 5, characterized in that, After obtaining the noisy dataset that conflicts with the standard dataset from the training dataset, the method further includes: The noisy dataset is removed from the training dataset to obtain the target training dataset; The model to be trained is trained based on the target training dataset to obtain the target prediction model; Alternatively, the target prediction model can be obtained by training the model to be trained using the target training dataset and the standard dataset.

7. The method according to any one of claims 1 to 5, characterized in that, Before obtaining the first set of active tags that bind to the first set of predicted receptors and the first set of predicted ligands, the method further includes: Based on the initial prediction model, the activity of the ligand library of the receptor to be predicted is predicted to obtain a second predicted activity set; Based on the second predicted activity set, the ligands to be predicted in the ligand library to be predicted are arranged in reverse order to obtain the ligand sequence to be predicted. A specified number of ligands to be predicted are selected sequentially from the ligand sequence to be predicted, and this number is determined as the first set of ligands to be predicted. The second set of ligands to be predicted consists of ligands in the ligand library other than the first set of ligands to be predicted.

8. The method according to any one of claims 1 to 5, characterized in that, After optimizing the target prediction model, the method further includes: Based on the target prediction model, a third predicted activity set is predicted for the binding of the receptor to be predicted to the second set of ligands to be predicted. Select a target set of ligands to be predicted from the second set of ligands to be predicted, and determine a target set of predicted activities corresponding to the target set of ligands to be predicted from the third set of predicted activities; Obtain a second set of active tags that bind to the target receptor and the target set of target ligands; Obtain the activity difference between the second set of activity tags and the target predicted activity set; When the activity difference is lower than a specified difference, the activity set of the receptor to be predicted binding to the ligand library is determined by combining the first activity tag set, the second activity tag set, and the third predicted activity set.

9. The method according to claim 8, characterized in that, After obtaining the activity difference between the second activity tag set and the target predicted activity set, the method further includes: When the activity difference is equal to or higher than the specified difference, the receptor to be predicted, the target predicted activity set, and the second activity tag set are combined into the next standard dataset; The target prediction model is optimized based on the datasets in the training dataset that conflict with the next standard dataset.

10. The method according to any one of claims 1 to 5, characterized in that, The training dataset also includes a third activity tag set combining receptor and ligand sample sets. The process of training the model to be trained based on the training dataset to obtain an initial prediction model includes: Based on the model to be trained, the activity of the receptor sample set and the ligand sample set is predicted to obtain a fourth predicted activity set. Based on the difference between the fourth predicted activity set and the third activity tag set, a third predicted loss set is determined; The model to be trained is trained based on the third prediction loss set to obtain the initial prediction model.

11. A model processing device, characterized in that, The model processing device includes: The model training module is used to train the model to be trained based on the training dataset to obtain an initial prediction model, wherein the model to be trained is a neural network model to be trained for predicting ligand and receptor binding activity. The tag acquisition module is used to acquire the first set of active tags that bind to the first set of predicted receptors and the first set of predicted ligands. The data acquisition module is used to combine the receptor to be predicted, the first set of ligands to be predicted, and the first set of active tags into a standard dataset; A noise determination module is used to obtain a noise dataset that conflicts with the standard dataset from the training dataset; The model optimization module is used to maximize the loss value of the initial prediction model for the noisy dataset to obtain the model parameters to be optimized; obtain the second parameter difference between the model parameters to be optimized and the initial model parameters; fuse the parameter weights with the second parameter difference to obtain second smoothing information; and combine the model parameters to be optimized and the second smoothing information to optimize the target prediction model, wherein the target prediction model is used to predict the active set of the receptor to be predicted binding to the second set of ligands to be predicted.

12. The apparatus according to claim 11, characterized in that, The training dataset includes a receptor sample set and a ligand sample set. The noise determination module is further configured to: continuously learn the initial prediction model based on the standard dataset to obtain an intermediate prediction model; obtain a first prediction loss set combining the receptor sample set and the ligand sample set based on the intermediate prediction model; determine the conflict value of each receptor sample in the receptor sample set based on the difference between the first prediction loss set and the second prediction loss set, wherein the second prediction loss set is the loss set obtained by the initial prediction model for activity prediction of the receptor sample set and the ligand sample set; determine the noisy receptor set whose conflict value is greater than a conflict threshold from the receptor sample set; and determine the information corresponding to the noisy receptor set in the training dataset as the noisy dataset that conflicts with the standard dataset.

13. The apparatus according to claim 12, characterized in that, The noise determination module is further configured to: predict a first predicted activity set of the receptor to be predicted and the first set of ligands to be predicted based on the initial prediction model; train the initial prediction model based on the difference between the first predicted activity set and the first set of active tags to obtain current model parameters; obtain first smoothing information between the current model parameters and the initial model parameters, wherein the initial model parameters are the model parameters of the initial prediction model, and the first smoothing information is used to retain the initial model parameters during continuous learning; and determine the intermediate prediction model by combining the current model parameters and the first smoothing information.

14. The apparatus according to claim 13, characterized in that, The noise determination module is further configured to obtain a first parameter difference between the current model parameters and the initial model parameters; obtain a first loss gradient of the initial prediction model for the training dataset; and obtain parameter weights that are positively correlated with the first loss gradient. The first smoothing information is obtained by combining the first parameter difference and the parameter weight.

15. The apparatus according to claim 12, characterized in that, The noise determination module is further configured to obtain the second loss gradient of the intermediate prediction model on the standard dataset; and to obtain the conflict value that is negatively correlated with both the first loss gradient and the second loss gradient and positively correlated with the parameter weights.

16. The apparatus according to any one of claims 11 to 15, characterized in that, After obtaining the noisy dataset that conflicts with the standard dataset from the training dataset, The model optimization module is also used to remove the noisy dataset from the training dataset to obtain the target training dataset; The target prediction model is obtained by training the model to be trained based on the target training dataset; or, the target prediction model is obtained by training the model to be trained by combining the target training dataset and the standard dataset.

17. The apparatus according to any one of claims 11 to 15, characterized in that, The tag acquisition module is further configured to predict the activity of the ligand library of the receptor to be predicted based on the initial prediction model, to obtain a second predicted activity set; based on the second predicted activity set, to reverse the order of the ligands to be predicted in the ligand library, to obtain a sequence of ligands to be predicted; and to determine a specified number of the ligands to be predicted sequentially selected from the sequence of ligands to be predicted as the first set of ligands to be predicted, wherein the second set of ligands to be predicted consists of the ligands to be predicted in the ligand library other than the first set of ligands to be predicted.

18. The apparatus according to any one of claims 11 to 15, characterized in that, The model optimization module is further configured to predict a third predicted activity set based on the target prediction model, wherein the receptor to be predicted binds to the second set of ligands to be predicted; select a target set of ligands to be predicted from the second set of ligands to be predicted; and determine a target predicted activity set corresponding to the target set of ligands to be predicted from the third set of predicted activity. Obtain a second set of active tags for the binding of the receptor to be predicted and the target set of predicted ligands; obtain the activity difference between the second set of active tags and the target predicted activity set; when the activity difference is lower than a specified difference, combine the first set of active tags, the second set of active tags and the third predicted activity set to determine the activity set for the binding of the receptor to be predicted and the library of predicted ligands.

19. The apparatus according to claim 18, characterized in that, The model optimization module is further configured to combine the receptor to be predicted, the target predicted activity set, and the second activity tag set into a next standard dataset when the activity difference is equal to or higher than a specified difference. The target prediction model is optimized based on the datasets in the training dataset that conflict with the next standard dataset.

20. The apparatus according to any one of claims 11 to 15, characterized in that, The training dataset also includes a third activity tag set combining the receptor sample set and the ligand sample set. The model training module is further configured to perform activity prediction on the receptor sample set and the ligand sample set based on the model to be trained, to obtain a fourth predicted activity set; determine a third predicted loss set based on the difference between the fourth predicted activity set and the third activity tag set; and train the model to be trained based on the third predicted loss set to obtain the initial prediction model.

21. An electronic device for model processing, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the model processing method according to any one of claims 1 to 10.

22. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the model processing method according to any one of claims 1 to 10 is implemented.

23. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the model processing method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Small molecule activity prediction method and device, and computing device

    CN111445945A

  • Model training method, related device and storage medium

    CN115392405A