Methods, apparatus, computer equipment and storage media for generating complex information

By acquiring protein structure and mutation information, molecular docking and energy calculations were performed, solving the problem of low efficiency in complex information generation and improving the accuracy of drug resistance prediction models.

CN116189766BActive Publication Date: 2026-05-19TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2021-11-29
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

The inefficiency of obtaining complex information in existing technologies leads to a decrease in the accuracy of drug resistance prediction models.

Method used

By acquiring protein structure and mutation information, adjusting protein structure information, performing molecular docking, and calculating energy information to determine target complex information, the efficiency of complex information generation can be improved.

Benefits of technology

Rapidly generating target complex information improves the efficiency of obtaining complex structural information, thereby enhancing the accuracy of drug resistance prediction models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189766B_ABST
    Figure CN116189766B_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, computer device, storage medium, and computer program product for generating complex information. The method includes: acquiring protein structure information, corresponding mutation information, and corresponding drug structure information; adjusting the protein structure information according to the mutation information to obtain mutant protein structure information; performing molecular docking of the protein structure information and mutant protein structure information with drug structure information respectively to obtain wild-type complex structure information and mutant complex structure information; determining the corresponding target wild-type complex structure information from the wild-type complex structure information, and determining the corresponding target mutant complex structure information from the mutant complex structure information; and obtaining target complex information based on the target wild-type complex structure information and the target mutant complex structure information. This method can improve the efficiency of obtaining target complex information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, computer device, storage medium, and computer program product for generating complex information. Background Technology

[0002] With the development of artificial intelligence (AI) technology, techniques for predicting drug resistance using AI have emerged. Predicting drug resistance typically requires obtaining information about the protein complex before and after mutation. This complex information is then used to extract features for training the drug resistance prediction model. Currently, obtaining complex information usually involves experimentally measuring the complex before and after protein mutation, for example, using structural biology methods (such as X-ray diffraction and single-molecule cryo-electron microscopy) to analyze the complex information between the protein or protein mutant and the drug. However, these methods are inefficient at obtaining complex information, resulting in small sample sizes and reduced accuracy of the trained drug resistance prediction model. Summary of the Invention

[0003] Therefore, it is necessary to provide a method, apparatus, computer device, storage medium, and computer program product for generating complex information that can improve the efficiency of complex information generation, in order to address the above-mentioned technical problems.

[0004] In a first aspect, this application provides a method for generating complex information. The method includes:

[0005] Obtain protein structure information and the mutation information corresponding to the protein structure information, and obtain drug structure information corresponding to the protein structure information;

[0006] The protein structure information is adjusted according to the mutation information to obtain the mutant protein structure information;

[0007] The structure information of the protein was molecularly docked with the structure information of the drug to obtain the structure information of each wild-type complex corresponding to the structure information of the protein. The structure information of the mutant protein was molecularly docked with the structure information of the drug to obtain the structure information of each mutant complex corresponding to the structure information of the mutant protein.

[0008] Calculate the wild-type energy information corresponding to the structure information of each wild-type complex, and determine the target wild-type complex structure information corresponding to the protein structure information from the structure information of each wild-type complex based on the wild-type energy information;

[0009] Calculate the mutant energy information corresponding to the structural information of each mutant complex, and determine the target mutant complex structural information corresponding to the mutant protein structural information from the structural information of each mutant complex based on the mutant energy information;

[0010] The target complex information is obtained based on the structural information of the target wild-type complex and the structural information of the target mutant complex.

[0011] Secondly, this application also provides a complex information generation apparatus. The apparatus includes:

[0012] The information acquisition module is used to acquire protein structure information and mutation information corresponding to the protein structure information, and to acquire drug structure information corresponding to the protein structure information.

[0013] The mutation module is used to adjust the protein structure information according to the mutation information to obtain the mutant protein structure information;

[0014] The molecular docking module is used to perform molecular docking between protein structure information and drug structure information to obtain the structure information of each wild-type complex corresponding to the protein structure information, and to perform molecular docking between mutant protein structure information and drug structure information to obtain the structure information of each mutant complex corresponding to the mutant protein structure information.

[0015] The wild-type complex determination module is used to calculate the wild-type energy information corresponding to the structure information of each wild-type complex, and to determine the target wild-type complex structure information corresponding to the protein structure information based on the wild-type energy information from the structure information of each wild-type complex.

[0016] The mutant complex determination module is used to calculate the mutant energy information corresponding to the structural information of each mutant complex, and to determine the target mutant complex structure information corresponding to the mutant protein structure information from the structural information of each mutant complex based on the mutant energy information.

[0017] The target complex information acquisition module is used to obtain target complex information based on the target wild-type complex structure information and the target mutant complex structure information.

[0018] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0019] Obtain protein structure information and the mutation information corresponding to the protein structure information, and obtain drug structure information corresponding to the protein structure information;

[0020] The protein structure information is adjusted according to the mutation information to obtain the mutant protein structure information;

[0021] The structure information of the protein was molecularly docked with the structure information of the drug to obtain the structure information of each wild-type complex corresponding to the structure information of the protein. The structure information of the mutant protein was molecularly docked with the structure information of the drug to obtain the structure information of each mutant complex corresponding to the structure information of the mutant protein.

[0022] Calculate the wild-type energy information corresponding to the structure information of each wild-type complex, and determine the target wild-type complex structure information corresponding to the protein structure information from the structure information of each wild-type complex based on the wild-type energy information;

[0023] Calculate the mutant energy information corresponding to the structural information of each mutant complex, and determine the target mutant complex structural information corresponding to the mutant protein structural information from the structural information of each mutant complex based on the mutant energy information;

[0024] The target complex information is obtained based on the structural information of the target wild-type complex and the structural information of the target mutant complex.

[0025] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0026] Obtain protein structure information and the mutation information corresponding to the protein structure information, and obtain drug structure information corresponding to the protein structure information;

[0027] The protein structure information is adjusted according to the mutation information to obtain the mutant protein structure information;

[0028] The structure information of the protein was molecularly docked with the structure information of the drug to obtain the structure information of each wild-type complex corresponding to the structure information of the protein. The structure information of the mutant protein was molecularly docked with the structure information of the drug to obtain the structure information of each mutant complex corresponding to the structure information of the mutant protein.

[0029] Calculate the wild-type energy information corresponding to the structure information of each wild-type complex, and determine the target wild-type complex structure information corresponding to the protein structure information from the structure information of each wild-type complex based on the wild-type energy information;

[0030] Calculate the mutant energy information corresponding to the structural information of each mutant complex, and determine the target mutant complex structural information corresponding to the mutant protein structural information from the structural information of each mutant complex based on the mutant energy information;

[0031] The target complex information is obtained based on the structural information of the target wild-type complex and the structural information of the target mutant complex.

[0032] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0033] Obtain protein structure information and the mutation information corresponding to the protein structure information, and obtain drug structure information corresponding to the protein structure information;

[0034] The protein structure information is adjusted according to the mutation information to obtain the mutant protein structure information;

[0035] The structure information of the protein was molecularly docked with the structure information of the drug to obtain the structure information of each wild-type complex corresponding to the structure information of the protein. The structure information of the mutant protein was molecularly docked with the structure information of the drug to obtain the structure information of each mutant complex corresponding to the structure information of the mutant protein.

[0036] Calculate the wild-type energy information corresponding to the structure information of each wild-type complex, and determine the target wild-type complex structure information corresponding to the protein structure information from the structure information of each wild-type complex based on the wild-type energy information;

[0037] Calculate the mutant energy information corresponding to the structural information of each mutant complex, and determine the target mutant complex structural information corresponding to the mutant protein structural information from the structural information of each mutant complex based on the mutant energy information;

[0038] The target complex information is obtained based on the structural information of the target wild-type complex and the structural information of the target mutant complex.

[0039] The aforementioned method, apparatus, computer equipment, storage medium, and computer program product for generating complex information acquire protein structure information, mutation information, and drug structure information. Then, based on the mutation information, they obtain mutant protein structure information. Next, they use the protein structure information and mutant protein structure information to perform molecular docking with the drug structure information, respectively, and calculate energy information to select the corresponding complex structure. Finally, they obtain the target complex information, thereby rapidly obtaining the target complex information, improving the efficiency of obtaining complex structure information, and solving the problem of low efficiency in obtaining complex structure information. Attached Figure Description

[0040] Figure 1 This is an application environment diagram of the complex information generation method in one embodiment;

[0041] Figure 2 This is a flowchart illustrating a method for generating complex information in one embodiment;

[0042] Figure 3This is a flowchart illustrating the process of obtaining a target drug resistance prediction model in one embodiment;

[0043] Figure 4 This is a schematic diagram of the process for predicting drug resistance in one embodiment;

[0044] Figure 5 This is a schematic diagram of the process for obtaining target drug resistance characteristics in one embodiment;

[0045] Figure 6 This is a schematic diagram of the process for obtaining drug structure information in one embodiment;

[0046] Figure 7 This is a schematic diagram of the process for obtaining the structure information of the wild-type complex in one embodiment;

[0047] Figure 8 This is a flowchart illustrating a method for generating complex information in a specific embodiment;

[0048] Figure 9 This is a schematic diagram of the process for obtaining training feature data in a specific embodiment.

[0049] Figure 10 This is a schematic diagram of the process for obtaining training feature data in another specific embodiment;

[0050] Figure 11 This is a comparative schematic diagram illustrating the characteristic distribution of drug properties in a specific embodiment;

[0051] Figure 12 This is a comparative schematic diagram illustrating the characteristic distribution of a mutation environment in a specific embodiment;

[0052] Figure 13 This is a schematic diagram illustrating the results of a comparative test of a drug resistance prediction model in a specific embodiment.

[0053] Figure 14 This is a structural block diagram of a complex information generation device in one embodiment;

[0054] Figure 15 This is an internal structural diagram of a computer device in one embodiment;

[0055] Figure 16 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0057] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0058] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.

[0059] The solutions provided in this application involve technologies such as machine learning in artificial intelligence, and are specifically illustrated through the following embodiments:

[0060] The complex information generation method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on another network server. Terminal 102 sends a complex information generation command to server 104. Server 104 retrieves protein structure information and corresponding mutation information from the data storage system according to the command, and also retrieves drug structure information corresponding to the protein structure information. Server 104 adjusts the protein structure information according to the mutation information to obtain mutant protein structure information. Server 104 performs molecular docking between the protein structure information and the drug structure information to obtain the structure information of each wild-type complex corresponding to the protein structure information, and performs molecular docking between the mutant protein structure information and the drug structure information to obtain the structure information of each mutant complex corresponding to the mutant protein structure information. Server 104 calculates the wild-type energy information corresponding to each wild-type complex structure information, and determines the target wild-type complex structure information corresponding to the protein structure information based on the wild-type energy information. Server 104 calculates the mutant energy information corresponding to each mutant complex structure information, and determines the target mutant complex structure information corresponding to the mutant protein structure information based on the mutant energy information. Server 104 obtains the target complex information based on the target wild-type complex structure information and the target mutant complex structure information. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle systems. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0061] In one embodiment, such as Figure 2 As shown, a method for generating complex information is provided, which can be applied to... Figure 1 Taking a server as an example, it can be understood that this method can also be applied to a terminal, or to a system that includes both a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the following steps are included:

[0062] Step 202: Obtain protein structure information and the corresponding mutation information, and obtain the drug structure information corresponding to the protein structure information.

[0063] Protein structure information refers to the information about the three-dimensional structure of a protein. Mutation information is used to mutate proteins and can refer to the amino acid information that has been mutated in the protein. Drug structure information refers to the information about the three-dimensional structure of a protein.

[0064] Specifically, the server can retrieve protein structure information and corresponding mutation information from a database, as well as the corresponding drug structure information. Alternatively, the server can retrieve protein structure information and corresponding mutation information from a third-party server (which could be a data service provider). The server can also collect protein structure information and corresponding mutation information from the internet, and obtain the corresponding drug structure information.

[0065] Step 204: Adjust the protein structure information according to the mutation information to obtain the mutant protein structure information.

[0066] Among them, mutant protein structural information refers to the three-dimensional structural information of the mutated protein.

[0067] Specifically, the server adjusts the protein structure information based on the mutated amino acid information in the mutation information to obtain the mutant protein structure information. This can involve adjusting the corresponding amino acid information in the protein structure information to match the mutated amino acid information in the mutation information.

[0068] Step 206: Molecularly dock the protein structure information with the drug structure information to obtain the structure information of each wild-type complex corresponding to the protein structure information, and molecularly dock the mutant protein structure information with the drug structure information to obtain the structure information of each mutant complex corresponding to the mutant protein structure information.

[0069] Molecular docking refers to the process of placing molecules from a known three-dimensional structural database one by one at the active site of a target molecule. By continuously optimizing the position, conformation, dihedral angles of rotatable bonds within the receptor compound, and the amino acid residue side chains and backbone of the receptor, the optimal conformation for interaction between the small receptor molecule and the target macromolecule is sought. This process also involves predicting the binding mode, affinity, and selecting ligands with the best affinity to the receptor and closest to the native conformation using a scoring function. In other words, molecular docking refers to the interaction between a protein molecule corresponding to the protein structure information as a receptor and a drug molecule corresponding to the drug structure information as a ligand. Wild-type complex structural information is used to characterize the complex obtained by molecular docking the pre-mutant protein structure information with the drug structure information. Mutant complex structural information is used to characterize the complex obtained by molecular docking the mutated protein structure information with the drug structure information.

[0070] Specifically, the server uses molecular docking tools to perform molecular docking between the pre- and post-mutation protein structure information and the drug structure information. This yields the structure information of each wild-type complex obtained from the molecular docking of the pre-mutation protein structure information and the drug structure information. Simultaneously, it also obtains the structure information of the post-mutation protein, i.e., the structure information of the mutant complex obtained from the molecular docking of the mutant protein structure information and the drug structure information. Molecular docking can include rigid docking, semi-flexible docking, flexible docking, etc. Molecular docking tools refer to software programs used for molecular docking, which can be accessed and used via an interface.

[0071] Step 208: Calculate the wild-type energy information corresponding to the structure information of each wild-type complex, and determine the target wild-type complex structure information corresponding to the protein structure information from the structure information of each wild-type complex based on the wild-type energy information.

[0072] Here, wild-type energy information refers to the binding free energy corresponding to the wild-type complex structure information. Binding free energy refers to the interaction between the ligand and the receptor. Target wild-type complex structure information refers to the structure information of the wild-type complex with the lowest binding free energy.

[0073] Specifically, the server calculates the wild-type energy information corresponding to each wild-type complex structure information, and selects the wild-type complex structure information with the lowest wild-type energy information from among the various wild-type complex structure information to obtain the target wild-type complex structure information corresponding to that protein structure information.

[0074] Step 210: Calculate the mutant energy information corresponding to the structural information of each mutant complex, and determine the target mutant complex structural information corresponding to the mutant protein structural information from the structural information of each mutant complex based on the mutant energy information.

[0075] Here, mutant energy information refers to the binding free energy corresponding to the structure information of the mutant complex. Target mutant complex structure information refers to the structure information of the mutant complex with the lowest binding free energy.

[0076] Specifically, the server calculates the mutant energy information corresponding to each mutant complex structure information, and then selects the mutant complex structure information with the lowest mutant energy information from all mutant complex structure information to obtain the target mutant complex structure information.

[0077] Step 212: Obtain target complex information based on the target wild-type complex structure information and the target mutant complex structure information.

[0078] The target complex information refers to the complexes corresponding to the protein before and after the mutation.

[0079] Specifically, the server obtains the corresponding complexes before and after the protein mutation based on the structural information of the target wild-type complex and the structural information of the target mutant complex.

[0080] The aforementioned method, apparatus, computer equipment, storage medium, and computer program product for generating complex information acquire protein structure information, mutation information, and drug structure information. Then, based on the mutation information, they obtain mutant protein structure information. Next, they use the protein structure information and mutant protein structure information to perform molecular docking with the drug structure information, respectively, and calculate energy information to select the corresponding complex structure. Finally, they obtain the target complex information, thereby rapidly obtaining the target complex information, improving the efficiency of obtaining complex structure information, and solving the problem of low efficiency in obtaining complex structure information.

[0081] In one embodiment, such as Figure 3 As shown, after step 212, that is, after obtaining the target complex information based on the target wild-type complex structure information and the target mutant complex structure information, the method further includes:

[0082] Step 302: Obtain the drug resistance tag corresponding to the target complex information.

[0083] The drug resistance tag is used to characterize whether a protein mutation leads to drug resistance, including the generation of a drug resistance tag and the absence of a drug resistance tag. The generation of a drug resistance tag indicates that the mutated protein has developed resistance to the drug, meaning the change in the binding affinity between the protein and the corresponding drug before and after the mutation is significant. The absence of a drug resistance tag indicates that the mutated protein has not developed resistance to the drug, meaning the change in the binding affinity between the protein and the corresponding drug before and after the mutation is relatively low.

[0084] Specifically, the server quickly obtains a large number of target complexes. It can then retrieve drug resistance tags corresponding to the target complex information from a database, or obtain drug resistance information associated with protein results and mutation information from a third-party server. Based on this drug resistance information, the server determines the drug resistance tag corresponding to the target complex information. This drug resistance information can refer to the change in the binding affinity between the mutated protein and the drug. The server can also obtain drug resistance tags corresponding to the target complex information submitted by the user terminal.

[0085] Step 304: Extract drug resistance features based on target complex information, drug structure information, protein structure information, and mutant protein structure information to obtain target drug resistance features.

[0086] Among them, target drug resistance features refer to the features extracted that are related to drug resistance.

[0087] Specifically, the server performs target drug resistance feature extraction, which can be achieved by using target complex information, drug structure information, protein structure information, and mutant protein structure information to obtain target drug resistance features.

[0088] Step 306: Input the target drug resistance characteristics into the initial drug resistance prediction model to predict drug resistance and obtain the initial drug resistance prediction results.

[0089] The initial resistance prediction model refers to the resistance prediction model with initialized model parameters. This model is built using machine learning algorithms, such as extreme random regression trees, random forests, linear regression, or neural networks. The initial resistance prediction result refers to the resistance prediction obtained using the initialized model parameters. This result indicates whether resistance has developed, including both developed and undeveloped resistance. The resistance prediction result can also be represented using changes in binding affinity.

[0090] Specifically, the server uses the target drug resistance characteristics as input to the initial drug resistance prediction model to predict drug resistance, and obtains the initial drug resistance prediction result output by the initial drug resistance prediction model.

[0091] Step 308: Calculate the loss based on the initial drug resistance prediction results and drug resistance labels to obtain loss information, and update the initial drug resistance prediction model in reverse based on the loss information until the training completion condition is met, thus obtaining the target drug resistance prediction model.

[0092] The loss information is used to characterize the error between the initial drug resistance prediction result and the corresponding drug resistance label. The training completion condition refers to the conditions under which the initial drug resistance prediction model has completed training, which may include reaching the maximum number of training iterations, the loss information reaching a preset threshold, and the model parameters no longer changing, etc.

[0093] Specifically, the server calculates the error between the initial drug resistance prediction result and the drug resistance label to obtain loss information. Then, using this loss information, the server updates the model parameters in the initial drug resistance prediction model in reverse using the gradient descent algorithm to obtain the updated drug resistance prediction model. The updated drug resistance prediction model is then used as the initial drug resistance prediction model and iteratively trained until the training completion condition is met. Finally, the initial drug resistance prediction model that has met the training completion condition is used as the final target drug resistance prediction model obtained through training.

[0094] In the above embodiments, the drug resistance prediction model is trained by using a large amount of target complex information. When the training is completed, the target drug resistance prediction model is obtained. Due to the increased sample size, the accuracy of drug resistance prediction is improved by the trained target drug resistance prediction model.

[0095] In one embodiment, such as Figure 4 As shown, after step 308, that is, after calculating the loss based on the initial drug resistance prediction results and drug resistance labels to obtain loss information, and updating the initial drug resistance prediction model in reverse based on the loss information, until the training completion condition is met and the target drug resistance prediction model is obtained, the following steps are also included:

[0096] Step 402: Obtain the data to be predicted, which includes the structure information of the protein to be predicted, the structure information of the mutant protein to be predicted, the structure information of the drug to be predicted, and the information of the complex to be predicted.

[0097] The data to be predicted refers to the data used to predict whether drug resistance will develop. The structural information of the protein to be predicted refers to the structural information corresponding to the protein to be predicted. The structural information of the mutant protein to be predicted refers to the structural information corresponding to the protein after mutation. The structural information of the drug to be predicted refers to the structural information corresponding to the drug to be predicted. This drug can interact with the protein to be predicted. The information of the complex to be predicted refers to the complex information corresponding to the protein before and after mutation, including the complex formed by the protein before mutation and the drug after mutation.

[0098] Specifically, the server can obtain the data to be predicted uploaded by the user terminal. This data includes the structure information of the protein to be predicted, the structure information of the mutant protein to be predicted, the structure information of the drug to be predicted, and the information of the complex to be predicted. The server can also obtain the data to be predicted from the database, and it can also obtain the data to be predicted from the business server.

[0099] Step 404: Extract drug resistance features based on the information of the complex to be predicted, the structure of the protein to be predicted, the structure of the drug to be predicted, and the structure of the mutant protein to be predicted, to obtain the drug resistance features to be predicted.

[0100] Among them, the antidrug resistance features to be predicted refer to the features extracted using the information of the antidrug complex, the structure of the antidrug protein, the structure of the antidrug, and the structure of the antidrug mutant protein, which are used to predict drug resistance.

[0101] Specifically, the server uses the information of the complex to be predicted, the structure of the protein to be predicted, the structure of the drug to be predicted, and the structure of the mutant protein to be predicted to extract drug resistance features. That is, it extracts features that are of reference value for predicting changes in binding affinity (binding affinity can also be called binding free energy) after protein mutation, thereby obtaining the drug resistance features to be predicted.

[0102] Step 406: Input the drug resistance characteristics to be predicted into the target drug resistance prediction model to predict drug resistance and obtain the drug resistance prediction results.

[0103] Specifically, a pre-trained target drug resistance prediction model is pre-deployed on the server. When the target drug resistance prediction model is needed, the server calls the model, then uses the drug resistance feature to be predicted as input to the model, and obtains the drug resistance prediction result output by the model.

[0104] In the above embodiments, by acquiring the data to be predicted, performing feature extraction, and then using the target drug resistance prediction model to predict drug resistance, the obtained drug resistance prediction results are more accurate.

[0105] In one embodiment, such as Figure 5 As shown, step 304 involves extracting drug resistance features based on the target complex information, drug structure information, protein structure information, and mutant protein structure information to obtain the target drug resistance features, including:

[0106] Step 502: Extract physicochemical property features based on target complex information, drug structure information, protein structure information, and mutant protein structure information to obtain target physicochemical property features.

[0107] Physicochemical properties refer to the characteristics that represent the physicochemical properties of molecules. Target physicochemical property characteristics refer to the characteristics obtained by extracting physicochemical properties.

[0108] Specifically, the server can extract physicochemical properties using target complex information, drug structure information, protein structure information, and mutant protein structure information. Specifically, the server can calculate features describing ligand properties, such as molecular weight and polar surface area. The server can calculate features describing the mutation environment, such as the radial count of protein atoms near the mutation, the number of polar / polar / charged residues in the binding pocket, etc. The server can calculate the properties of each amino acid, such as changes in side chain volume, hydrophilicity, and the number of hydrogen bond donors, and use these properties to calculate features describing changes in the chemical properties of amino acid residues at the mutation site. The server can also calculate the solvent-accessible surface area features of the pharmacophore, etc. Finally, the server obtains the target physicochemical properties.

[0109] Step 504: Based on the target complex information, drug structure information, protein structure information, and mutant protein structure information, the interaction energy features are extracted to obtain the target energy features.

[0110] Interaction energy refers to the binding free energy of a protein molecule. Target energy characteristics are features obtained by extracting interaction energy.

[0111] Specifically, the server can also use target complex information, drug structure information, protein structure information, and mutant protein structure information to extract interaction energy features. Among these features, the server can extract relevant characteristics of changes in protein folding free energy during mutation. The server can also use PLIP (Protein-Ligand Interaction Profiler) to calculate features describing protein-drug interactions. The server can also calculate relevant features of molecular docking tools, such as those obtained using the AutoDock Vina (a tool for processing proteins and small molecules) program, and then calculate the relevant features of Vina, etc.

[0112] Step 506: Obtain the target drug resistance characteristics based on the target physicochemical properties and target energy characteristics.

[0113] Specifically, the server can directly use the extracted target physicochemical properties and target energy characteristics as target drug resistance characteristics. Alternatively, the server can obtain the target drug resistance characteristics by weighting the extracted target physicochemical properties and target energy characteristics.

[0114] In the above embodiments, by extracting the target physicochemical properties and target energy characteristics, the target drug resistance characteristics are obtained. This enables the extracted target drug resistance characteristics to better characterize the changes in binding affinity after protein mutation, thereby improving accuracy.

[0115] In one embodiment, such as Figure 6 As shown, step 202, which involves obtaining protein structure information and the corresponding mutation information, and obtaining the drug structure information corresponding to the protein structure information, includes:

[0116] Step 602: Obtain basic drug resistance information, which includes the storage address of protein structure information, mutation information, and the storage address of the initial drug structure information.

[0117] Among them, the basic drug resistance information describes the basic information about drug resistance between proteins and their corresponding drugs. The protein structure information storage address refers to the address where protein structure information is stored; this address allows you to find the protein structure information, which can be its name, number, etc. The drug initial structure information storage address refers to the address where drug initial structure information is stored; this address allows you to find the drug's initial structure information. Drug initial structure information refers to drug structure information that cannot be directly used for molecular docking and requires processing before molecular docking can be performed.

[0118] Specifically, the server can retrieve basic resistance information from the database, which is presented in the form of a data table. The server can also retrieve a CSV (Comma-Separated Values) file, which records basic resistance data. CSV, sometimes called character-separated values, stores tabular data (numbers and text) in plain text. Plain text means the file is a sequence of characters and does not contain data that must be interpreted like binary numbers.

[0119] In one specific embodiment, the CSV file obtained by the server is shown in Table 1.

[0120] Table 1 Basic Information on Drug Resistance

[0121] UNIPROT_ID PDB_ID DRUG MUTATION WT.EXP MT.EXP DDG.EXP P07949 2IVU CBT L370I 0.24 2.96 1.488 P07949 2IVU CBT R455I_M366K 0.1 0.29 0.045

[0122] The database contains several key information: UNIPROT_ID records the protein's identifier in the Uniprot (Universal Protein) database; PDB_ID records the protein's identifier in the PDB (Protein Data Bank) database; DRUG records the drug name; MUTATION records the location of the protein mutation (e.g., in L370I, the first 'L' represents the wild-type amino acid information, the last 'I' represents the mutant amino acid information, and the '370' in the middle indicates the mutation location); WT.EXP records the IC50 experimental value for the wild-type protein; MT.EXP records the IC50 (half-inhibitory concentration) experimental value for the mutant protein; and DDG.EXP records the experimental value of the change in protein-drug binding affinity before and after the mutation. The fold change before and after the protein mutation can be directly calculated based on the WT.EXP and MT.EXP information, representing the change in the protein's affinity for the molecule before and after the mutation, which is an indicator for assessing drug resistance. Fold change is calculated based on the binding affinity after mutation and the binding affinity before mutation. For example, it can be obtained by calculating the ratio of the binding affinity after mutation to the binding affinity before mutation. When the fold change exceeds a preset threshold, for example, a preset threshold of 2, the protein is considered to have developed drug resistance after mutation.

[0123] Step 604: Obtain protein structure information from the protein structure information storage address, and obtain drug initial structure information from the drug initial structure information storage address.

[0124] Specifically, the server uses the protein structure information storage address to obtain the corresponding protein structure information, and uses the drug initial structure information storage address to obtain the corresponding drug initial structure information.

[0125] In one specific embodiment, the server can download the SMILES (Simplified Molecular Linear Input Specification, a specification that explicitly describes molecular structure using ASCII strings) string information of the drug molecule based on the UNIPROT_ID. Then, the SMILES string information of the drug molecule is converted into an SDF (Standard Delayed Format File, an easy-to-use file-based spatial data format that can store multiple geographic features in a tabular format, including various geometric types (points, lines, polygons, and arcs) and associated attribute information) file, which is the chemical structure format. The three-dimensional structure information of the protein can then be obtained based on the PDB_ID; if the PDB_ID does not exist, the UNIPROT_ID can also be used to obtain the three-dimensional structure information of the protein.

[0126] Step 606: Perform structural optimization based on the initial drug structural information to obtain drug structural information.

[0127] Among them, structure optimization refers to preprocessing the initial structure information of a drug so that the obtained drug structure information can be directly used for molecular docking.

[0128] Specifically, the server can adjust the drug structure corresponding to the initial drug structure information, such as by adding or deleting atoms in the drug structure, to obtain the drug structure information.

[0129] In the above embodiments, by obtaining basic information on drug resistance, the required information is obtained, thereby improving the efficiency of obtaining information.

[0130] In one embodiment, step 604, obtaining protein structure information from the protein structure information storage address, includes the following steps:

[0131] Download the target file from the protein structure information storage address. The target file includes protein structure information and the reference drug structure information associated with the protein structure information. Separate the protein structure information and the reference drug structure information according to the annotation information in the target file to obtain the protein structure information.

[0132] The target file refers to the file obtained from the protein structure information storage address. This file includes protein structure information and associated reference drug structure information. The reference drug structure information refers to the structure information of the drug associated with the protein stored at the protein structure information storage address. The annotation information in the target file is information that distinguishes between the protein and the drug.

[0133] Specifically, the server downloads a target file from the protein structure information storage address. This target file includes not only protein structure information but also associated reference drug structure information. To obtain the protein structure information, the protein and drug need to be separated. This involves separating the protein and reference drug structure information according to the annotation information in the target file, obtaining the separated protein and reference drug structure information. The reference drug structure information can then be saved, while the protein structure information can be used later.

[0134] In one specific embodiment, the server can obtain the protein's PDB file through PDB_ID or UNIPROT_ID information. This file records the protein's three-dimensional structural information and may also record associated ligands, i.e., drug structure information. Therefore, the protein and its corresponding drug need to be separated. During protein structural separation, water molecules can be removed from the protein, ultimately saving only the single-stranded protein. If there are multiple reference drug structures associated with the protein structure information, the number of heavy atoms in the reference drug structures is counted, and the reference drug structure with the highest number of heavy atoms is saved.

[0135] In one embodiment, step 606, which involves structural optimization based on the initial drug structure information to obtain drug structure information, includes the following steps:

[0136] Hydrogen atom information is removed from the initial drug structure information to obtain preliminary optimized structure information; polar hydrogen atom information is obtained, and the polar hydrogen atom information is combined with the preliminary optimized structure information to obtain the drug structure information.

[0137] Here, hydrogen atom information refers to the hydrogen atoms in the initial drug structure information. Preliminary optimized structure information refers to the drug structure information obtained after hydrogen atom deletion processing. Polar hydrogen atom information refers to polar hydrogen atoms.

[0138] Specifically, the server removes hydrogen atom information from the initial drug structure information; for example, the server can delete all hydrogen atoms from the initial drug structure information. Then, it obtains the added polar hydrogen atom information and adds it to the preliminary optimized structure information to obtain the drug structure information.

[0139] In the above embodiments, by optimizing the initial protein structure information, drug structure information is obtained, which facilitates subsequent molecular docking and improves efficiency.

[0140] In one embodiment, protein structure information corresponds to at least two mutation information;

[0141] Step 204: Adjust the protein structure information according to the mutation information to obtain the mutant protein structure information, including:

[0142] The protein structure information is adjusted according to at least two mutation information to obtain the target mutant protein structure information.

[0143] Among them, the protein structure information corresponds to at least two mutation information to indicate multiple mutations that have occurred in the protein.

[0144] Specifically, the server obtains protein structure information corresponding to at least two mutations, and then adjusts the protein structure information according to each mutation. For example, it replaces the original information in the protein structure information with the mutated information. When all mutation information has been replaced, the target mutant protein structure information is obtained. Subsequent processing is then performed using the target mutant protein structure information.

[0145] In a specific embodiment, as shown in Table 1, R455I_M366K, this protein has undergone two point mutations. In this case, the wild-type amino acid information R at position 455 in the protein structure information is replaced with I, and the wild-type amino acid information M at position 366 in the protein structure information is replaced with K. This yields the target mutant protein structure information after the mutation.

[0146] In the above embodiments, by adjusting the protein structure information according to at least two mutation information, the target mutant protein structure information is obtained, thereby ensuring the accuracy of the target mutant protein structure information. Then, the target mutant protein structure information is used for molecular docking, which can improve the accuracy of subsequent processing.

[0147] In one embodiment, mutation information includes the mutation location, the amino acid before the mutation, and the amino acid after the mutation;

[0148] Step 204: Adjust the protein structure information according to the mutation information to obtain the mutant protein structure information, including the following steps:

[0149] The mutant protein structure information is obtained by replacing the pre-mutation amino acid corresponding to the mutation position in the protein structure information with the post-mutation amino acid.

[0150] Specifically, when constructing the mutated protein, the server can directly replace the pre-mutation amino acid corresponding to the mutation position in the protein structure information with the post-mutation amino acid to obtain the mutant protein structure information.

[0151] In one embodiment, the server can also generate conformations of the amino acid residues at the mutation site based on the mutation location, the amino acid before mutation, and the amino acid after mutation, obtain various conformations, calculate the binding free energy of each conformation, and select the conformation with the smallest change in binding free energy after mutation as the structural information of the mutant protein. That is, by sampling in the conformational space of the protein, the conformation with the lowest energy is selected as the structural information of the mutant protein.

[0152] In one embodiment, such as Figure 7As shown, in step 206, the protein structure information is molecularly docked with the drug structure information to obtain the structure information of each wild-type complex corresponding to the protein structure information, including:

[0153] Step 702: When there is a related reference drug structure information for the protein structure information, determine the molecular docking region information based on the reference drug structure information.

[0154] Molecular docking region information refers to determining the docking region based on the reference drug structure information. This docking region is the specific three-dimensional area within the protein where molecular docking occurs. It can be implemented using a docking box, which includes center coordinates, dimensions, etc. The size of the region in this molecular docking region information can completely accommodate the drug structure information.

[0155] Specifically, the server determines whether the protein structure information has associated reference drug structure information, and can check if such information is stored. When associated reference drug structure information exists, the server calculates the molecular docking region information based on the reference drug structure information, that is, it calculates the size, center coordinates, etc. of the docking box based on the reference drug structure information, and generates the docking box.

[0156] Step 704: Perform molecular docking between the protein structure information and the drug structure information according to the molecular docking region information to obtain the structure information of each wild-type complex corresponding to the protein structure information.

[0157] Specifically, the server uses molecular docking region information to perform molecular docking between protein structure information and drug structure information. That is, the drug structure information is placed in the docking region in the protein structure information, and the drug conformation is adjusted or searched to obtain all possible binding conformations. Then, each binding conformation is scored and evaluated, and a certain number of binding conformations are selected based on the evaluation scores to obtain the structure information of each wild-type complex. For example, the three binding conformations with the best evaluation scores can be selected to obtain the structure information of each wild-type complex.

[0158] In the above embodiments, when there is associated reference drug structure information, determining the molecular docking region information based on the reference drug structure information can improve the accuracy of the molecular docking region information, thereby improving the accuracy of molecular docking.

[0159] In one embodiment, such as Figure 7 As shown, in step 206, the protein structure information is molecularly docked with the drug structure information to obtain the structure information of each wild-type complex corresponding to the protein structure information, including:

[0160] Step 706: When there is no associated reference drug structure information for the protein structure information, determine the docking region information of the target molecule based on the drug structure information.

[0161] Among them, the target molecule docking region information refers to the docking region determined based on the drug structure information.

[0162] Specifically, when the server determines that there is no associated reference drug structure information for the protein structure information, it directly uses the drug structure information to calculate the docking region, determine the size and center coordinates of the target region, and obtain the docking region information of the target molecule.

[0163] Step 708: Perform molecular docking between the protein structure information and the drug structure information according to the target molecule docking region information to obtain the structure information of each wild-type complex corresponding to the protein structure information.

[0164] Specifically, the server uses the target molecule docking region information to perform molecular docking between the protein structure information and the drug structure information. This involves placing the drug structure information within the docking region of the protein structure information, adjusting or searching for drug conformations, obtaining all possible binding conformations, and then scoring and evaluating each binding conformation. Based on the evaluation scores, a certain number of binding conformations are selected to obtain the structure information of each wild-type complex. For example, the three binding conformations with the best evaluation scores can be selected to obtain the structure information of each wild-type complex. In a specific embodiment, Smina (software based on Autodock Vina) can be used for molecular docking between the protein and the drug.

[0165] In the above embodiments, when there is no associated reference drug structure information for the protein structure information, the target molecule docking region information is determined based on the drug structure information, and then molecular docking is performed according to the target molecule docking region information to obtain the structure information of each wild-type complex corresponding to the protein structure information, thereby improving the accuracy of molecular docking.

[0166] In one embodiment, step 206 involves molecularly docking the mutant protein structure information with the drug structure information to obtain the structure information of each mutant complex corresponding to the mutant protein structure information, including:

[0167] Based on the drug structure information, the docking region information of the mutant molecule is determined; according to the docking region information of the target molecule, the structure information of the mutant protein is molecularly docked with the structure information of the drug to obtain the structure information of each mutant complex corresponding to the structure information of the mutant protein.

[0168] Among them, the information on the docking region of mutant molecules refers to the docking region when the structural information of mutant proteins is docked with molecules. This docking region refers to a specific three-dimensional region in the mutant protein.

[0169] Specifically, when performing molecular docking on mutant protein structure information, the server first uses drug structure information to determine the docking region information of the mutant molecule. Then, the server uses the docking region information of the mutant molecule to perform molecular docking between the mutant protein structure information and the drug structure information. That is, the drug structure information is placed in the docking region in the mutant protein structure information. The server adjusts or searches for drug conformations to obtain all possible binding conformations. Then, each binding conformation is scored and evaluated. Based on the evaluation scores, a certain number of binding conformations are selected to obtain the structure information of each wild-type complex. For example, the three binding conformations with the best evaluation scores can be selected to obtain the structure information of each mutant complex.

[0170] In a specific embodiment, such as Figure 8 As shown, a method for generating complex information is provided, specifically including the following steps:

[0171] Step 802: Obtain basic drug resistance information, which includes the storage address of protein structure information, mutation information, and the storage address of initial drug structure information. Obtain protein structure information from the protein structure information storage address and initial drug structure information from the initial drug structure information storage address. Perform structure optimization based on the initial drug structure information to obtain drug structure information.

[0172] Step 804: Replace the pre-mutation amino acid corresponding to the mutation position in the protein structure information with the post-mutation amino acid to obtain the mutant protein structure information.

[0173] Step 806: When there is a related reference drug structure information for the protein structure information, determine the molecular docking region information based on the reference drug structure information, and perform molecular docking between the protein structure information and the drug structure information according to the molecular docking region information to obtain the structure information of each wild-type complex corresponding to the protein structure information.

[0174] Step 808: When there is no associated reference drug structure information for the protein structure information, the target molecule docking region information is determined based on the drug structure information. The protein structure information and the drug structure information are then molecularly docked according to the target molecule docking region information to obtain the structure information of each wild-type complex corresponding to the protein structure information.

[0175] Step 810: Determine the docking region information of the mutant molecule based on the drug structure information; perform molecular docking between the mutant protein structure information and the drug structure information according to the target molecule docking region information to obtain the structure information of each mutant complex corresponding to the mutant protein structure information.

[0176] Step 812: Calculate the wild-type energy information corresponding to the structure information of each wild-type complex, and determine the target wild-type complex structure information corresponding to the protein structure information from the structure information of each wild-type complex based on the wild-type energy information.

[0177] Step 814: Calculate the mutant energy information corresponding to the structural information of each mutant complex, and determine the target mutant complex structural information corresponding to the mutant protein structural information from the structural information of each mutant complex based on the mutant energy information.

[0178] Step 816: Obtain target complex information based on the target wild-type complex structure information and the target mutant complex structure information.

[0179] This application also provides an application scenario in which the above-described complex information generation method is applied. Specifically, this method is applied in a drug resistance prediction scenario, where training data is required when training a drug resistance prediction model is needed. In this case, the training data can be obtained using the above-described complex information generation method, such as... Figure 9 The diagram illustrates the process of obtaining the training feature dataset for the drug resistance prediction model. First, basic information about the drug resistance dataset is obtained, such as that shown in Table 1. Then, data preprocessing is performed to obtain protein structure information and corresponding mutation information. The downloaded PDF file contains ligand (reference drug) structure information; therefore, the protein and ligand in the PDF file need to be separated. Next, mutant proteins are constructed using the mutation information and protein structure information. Then, the initial structure information of the drug to be docked is obtained, and all hydrogen atoms are removed, while polar hydrogen atoms are added to obtain the drug structure information. Then, molecular docking of the protein and drug is performed using a molecular docking tool, such as Smina (a software based on Autodock Vina used for molecular docking and structure-based virtual screening). After docking, the three-dimensional structure information of the mutant complex and the wild-type complex is obtained. A large number of different mutant complex three-dimensional structure information can be obtained through a large amount of different mutation information, allowing for feature extraction and the acquisition of training data. Specifically, through feature calculation, the physicochemical properties and energy characteristics corresponding to each training sample are obtained. The training samples refer to the samples adapted during training, and resistance labels are obtained based on basic resistance information. Then, the training data is used to train an initial resistance prediction model. Once training is complete, a target resistance prediction model is obtained, which is then deployed. In a specific embodiment, such as... Figure 10The diagram shows the flowchart for constructing the drug resistance dataset. Basic drug resistance information is collected from different data sources. This information is then used for data preprocessing to obtain wild-type protein structure information, corresponding mutation information, and drug structure information. Next, the wild-type protein structure information is subjected to gene mutation, resulting in mutant protein structure information. Then, the wild-type and mutant protein structure information are molecularly docked with ligands (i.e., drug structure information) to obtain wild-type and mutant complex structure information, respectively. Finally, feature calculations are performed to extract relevant features for drug resistance training, and these extracted features are used to train the drug resistance prediction model.

[0180] In one specific embodiment, when obtaining the training dataset for the drug resistance prediction model, each training sample in the training dataset includes wild-type complex structure information, mutant complex structure information, drug structure information, protein structure information, and mutant protein result information. Then, a comparative test is performed to verify the reliability of the training dataset. The comparative test datasets obtained at this time include Platinum (a database that extensively collects drug resistance information) dataset B, TKI (containing only 131 data points in total) dataset A, and KinaseMD (kinase mutation and drug response database) dataset C. The data in the Platinum and TKI datasets are co-crystal structure information of the protein binding to the drug molecule before and after mutation, obtained experimentally. The KinaseMD dataset only provides mutation site information. The complex is generated using the basic information in the KinaseMD dataset through any of the above embodiments, and the feature set predicting the affinity change after protein mutation is calculated. Then, the feature distribution of the KinaseMD dataset is compared with the feature distributions of the Platinum and TKI datasets. Figure 11 The image shows a comparison of the feature distributions of drug property description features A and B across three datasets, representing some of the calculated LIG (Liquidity-Induced Generic) features. Figure 12 The diagram shows the feature distributions of the calculated Mutant Context (MUT) features, specifically Mutant Context Feature A and Mutant Context Feature B, across three datasets. The feature distribution of the KinaseMD Mutant dataset is close to that of the experimentally obtained TKI Mutant dataset, while the feature distribution of the Platinum Mutant dataset differs somewhat from these two datasets. This is consistent with theoretical conclusions; therefore, the training dataset obtained using the complex structure information generated in this application is reliable.

[0181] In a specific embodiment, such as Figure 13 The image shows a schematic diagram illustrating the results of a comparative test of the drug resistance prediction model. The model was trained using a feature set obtained from the Platinum dataset, and then its predictive ability for drug resistance was tested on the TKI dataset. Figure 13 As shown in (a), the final drug resistance prediction model has a mean squared error (RMSE) of 0.87, a Pearson correlation coefficient of 0.12, and an area under the precision-recall curve (AUPRC) of 0.20 on the TKI dataset. Then, the feature sets extracted from the Platinum and KinaseMD datasets are used as the training feature set, and the TKI dataset is used as the test set to verify the predictive performance of the drug resistance prediction model after data gain. Figure 13 As shown in (b), it can be clearly seen that after data gain, that is, after training with a large number of training samples, the drug resistance prediction model has a significant improvement in the prediction performance of drug resistance. The mean squared error (RMSE) is 0.78, the Pearson correlation coefficient is 0.34, and the area under the precision and recall curve (AUPRC) is 0.33.

[0182] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0183] Based on the same inventive concept, this application also provides a complex information generation apparatus for implementing the complex information generation method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the complex information generation apparatus provided below can be found in the limitations of the complex information generation method described above, and will not be repeated here.

[0184] In one embodiment, such as Figure 14As shown, a complex information generation device 1400 is provided. This device can be a software module, a hardware module, or a combination of both as part of a computer device. Specifically, the device includes: an information acquisition module 1402, a mutation module 1404, a molecular docking module 1406, a wild-type complex determination module 1408, a mutant complex determination module 1410, and a target complex information acquisition module 1412, wherein:

[0185] The information acquisition module 1402 is used to acquire protein structure information and mutation information corresponding to the protein structure information, and to acquire drug structure information corresponding to the protein structure information.

[0186] Mutation module 1404 is used to adjust the protein structure information according to the mutation information to obtain the mutant protein structure information;

[0187] The molecular docking module 1406 is used to perform molecular docking between protein structure information and drug structure information to obtain the structure information of each wild-type complex corresponding to the protein structure information, and to perform molecular docking between mutant protein structure information and drug structure information to obtain the structure information of each mutant complex corresponding to the mutant protein structure information.

[0188] Wild-type complex determination module 1408 is used to calculate the wild-type energy information corresponding to the structure information of each wild-type complex, and to determine the target wild-type complex structure information corresponding to the protein structure information from the structure information of each wild-type complex based on the wild-type energy information.

[0189] The mutant complex determination module 1410 is used to calculate the mutant energy information corresponding to the structural information of each mutant complex, and to determine the target mutant complex structural information corresponding to the mutant protein structural information from the structural information of each mutant complex based on the mutant energy information;

[0190] The target complex information acquisition module 1412 is used to obtain target complex information based on the target wild-type complex structure information and the target mutant complex structure information.

[0191] In one embodiment, the complex information generation apparatus 1400 further includes:

[0192] The model training module is used to obtain the drug resistance label corresponding to the target complex information; extract drug resistance features based on the target complex information, drug structure information, protein structure information, and mutant protein structure information to obtain the target drug resistance features; input the target drug resistance features into the initial drug resistance prediction model to make drug resistance predictions and obtain the initial drug resistance prediction results; calculate the loss based on the initial drug resistance prediction results and drug resistance labels to obtain loss information, and update the initial drug resistance prediction model in reverse based on the loss information until the training completion condition is met, thus obtaining the target drug resistance prediction model.

[0193] In one embodiment, the complex information generation apparatus 1400 further includes:

[0194] The model prediction module is used to acquire the data to be predicted, which includes the structure information of the protein to be predicted, the structure information of the mutant protein to be predicted, the structure information of the drug to be predicted, and the information of the complex to be predicted. Based on the information of the complex to be predicted, the structure information of the protein to be predicted, the structure information of the drug to be predicted, and the structure information of the mutant protein to be predicted, drug resistance features are extracted to obtain the drug resistance features to be predicted. The drug resistance features to be predicted are input into the target drug resistance prediction model to predict drug resistance and obtain the drug resistance prediction results.

[0195] In one embodiment, the model training module is also used to extract physicochemical property features based on target complex information, drug structure information, protein structure information, and mutant protein structure information to obtain target physicochemical property features; to extract interaction energy features based on target complex information, drug structure information, protein structure information, and mutant protein structure information to obtain target energy features; and to obtain target drug resistance features based on target physicochemical property features and target energy features.

[0196] In one embodiment, the information acquisition module 1402 is further configured to acquire basic drug resistance information, which includes the storage address of protein structure information, mutation information, and the storage address of initial drug structure information; acquire protein structure information from the protein structure information storage address and acquire initial drug structure information from the initial drug structure information storage address; and perform structural optimization based on the initial drug structure information to obtain drug structure information.

[0197] In one embodiment, the information acquisition module 1402 is further configured to download a target file from the protein structure information storage address, the target file including protein structure information and reference drug structure information associated with the protein structure information; and to split the protein structure information and reference drug structure information according to the annotation information in the target file to obtain the protein structure information.

[0198] In one embodiment, the information acquisition module 1402 is further configured to delete hydrogen atom information from the initial structure information of the drug to obtain preliminary optimized structure information; acquire polar hydrogen atom information, and combine the polar hydrogen atom information with the preliminary optimized structure information to obtain drug structure information.

[0199] In one embodiment, the protein structure information corresponds to at least two mutation information; the mutation module 1404 adjusts the protein structure information according to the at least two mutation information to obtain the target mutant protein structure information.

[0200] In one embodiment, the mutation information includes the mutation location, the amino acid before the mutation, and the amino acid after the mutation; the mutation module 1404 replaces the amino acid before the mutation corresponding to the mutation location in the protein structure information with the amino acid after the mutation to obtain the mutant protein structure information.

[0201] In one embodiment, the molecular docking module 1406 is further configured to determine molecular docking region information based on the reference drug structure information when the protein structure information has associated reference drug structure information; and to perform molecular docking between the protein structure information and the drug structure information according to the molecular docking region information to obtain the structure information of each wild-type complex corresponding to the protein structure information.

[0202] In one embodiment, the molecular docking module 1406 is further configured to determine the target molecular docking region information based on the drug structure information when there is no associated reference drug structure information for the protein structure information; and to perform molecular docking between the protein structure information and the drug structure information according to the target molecular docking region information to obtain the structure information of each wild-type complex corresponding to the protein structure information.

[0203] In one embodiment, the molecular docking module 1406 is further configured to determine the mutant molecular docking region information based on the drug structure information; and to perform molecular docking between the mutant protein structure information and the drug structure information according to the target molecular docking region information to obtain the structure information of each mutant complex corresponding to the mutant protein structure information.

[0204] Each module in the aforementioned complex information generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0205] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 15As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The database stores basic information about drug resistance. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a method for generating complex information.

[0206] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 16 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a complex information generation method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.

[0207] Those skilled in the art will understand that Figure 15 Alternatively, the structure shown in Figure 16 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0208] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0209] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0210] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the steps in the above method embodiments.

[0211] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0212] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0213] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0214] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0215] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for generating complex information, characterized in that, The method includes: Obtain basic information on drug resistance, which includes the storage address of protein structure information, mutation information, and the storage address of the initial structure information of the drug. The protein structure information is obtained from the protein structure information storage address, and the initial drug structure information is obtained from the initial drug structure information storage address. Based on the initial drug structure information, structural optimization is performed to obtain the drug structure information; The protein structure information is adjusted according to the mutation information to obtain the mutant protein structure information; The protein structure information is molecularly docked with the drug structure information to obtain the structure information of each wild-type complex corresponding to the protein structure information, and the mutant protein structure information is molecularly docked with the drug structure information to obtain the structure information of each mutant complex corresponding to the mutant protein structure information. Calculate the wild-type energy information corresponding to each wild-type complex structure information, and determine the target wild-type complex structure information corresponding to the protein structure information based on the wild-type energy information from each wild-type complex structure information; Calculate the mutant energy information corresponding to the structural information of each mutant complex, and determine the target mutant complex structural information corresponding to the mutant protein structural information from the structural information of each mutant complex based on the mutant energy information; The target complex information is obtained based on the structural information of the target wild-type complex and the structural information of the target mutant complex.

2. The method according to claim 1, characterized in that, After obtaining the target complex information based on the target wild-type complex structure information and the target mutant complex structure information, the method further includes: Obtain the drug resistance label corresponding to the target complex information; Based on the target complex information, the drug structure information, the protein structure information, and the mutant protein structure information, drug resistance features are extracted to obtain the target drug resistance features; The target drug resistance characteristics are input into the initial drug resistance prediction model to predict drug resistance and obtain the initial drug resistance prediction results. Based on the initial drug resistance prediction results and the drug resistance labels, loss calculation is performed to obtain loss information. The initial drug resistance prediction model is then updated in reverse based on the loss information until the training completion condition is met, at which point the target drug resistance prediction model is obtained.

3. The method according to claim 2, characterized in that, After calculating the loss based on the initial drug resistance prediction result and the drug resistance label to obtain loss information, and updating the initial drug resistance prediction model in reverse based on the loss information until the training completion condition is met to obtain the target drug resistance prediction model, the method further includes: Acquire the data to be predicted, which includes the structure information of the protein to be predicted, the structure information of the mutant protein to be predicted, the structure information of the drug to be predicted, and the information of the complex to be predicted. Based on the information of the complex to be predicted, the structure information of the protein to be predicted, the structure information of the drug to be predicted, and the structure information of the mutant protein to be predicted, drug resistance features are extracted to obtain the drug resistance features to be predicted. The drug resistance characteristics to be predicted are input into the target drug resistance prediction model to predict drug resistance and obtain the drug resistance prediction results.

4. The method according to claim 2, characterized in that, The process of extracting drug resistance features based on the target complex information, the drug structure information, the protein structure information, and the mutant protein structure information to obtain target drug resistance features includes: Physicochemical property features are extracted based on the target complex information, the drug structure information, the protein structure information, and the mutant protein structure information to obtain the target physicochemical property features; Interaction energy features are extracted based on the target complex information, the drug structure information, the protein structure information, and the mutant protein structure information to obtain the target energy features; The target drug resistance characteristics are obtained based on the target's physicochemical properties and energy characteristics.

5. The method according to claim 1, characterized in that, The step of obtaining protein structure information from the protein structure information storage address includes: Download the target file from the protein structure information storage address. The target file includes protein structure information and reference drug structure information associated with the protein structure information. The protein structure information and the reference drug structure information are separated according to the annotation information in the target file to obtain the protein structure information.

6. The method according to claim 1, characterized in that, Based on the initial drug structure information, structural optimization is performed to obtain drug structure information, including: The hydrogen atom information in the initial structure information of the drug is deleted to obtain preliminary optimized structure information; The polar hydrogen atom information is obtained, and the polar hydrogen atom information is combined with the preliminary optimized structure information to obtain the drug structure information.

7. The method according to claim 1, characterized in that, The protein structure information corresponds to at least two mutation information; The step of adjusting the protein structure information according to the mutation information to obtain mutant protein structure information includes: The protein structure information is adjusted according to the at least two mutation information to obtain the target mutant protein structure information.

8. The method according to claim 1, characterized in that, The mutation information includes the mutation location, the amino acid before the mutation, and the amino acid after the mutation; The step of adjusting the protein structure information according to the mutation information to obtain mutant protein structure information includes: The mutant protein structure information is obtained by replacing the pre-mutation amino acid corresponding to the mutation position in the protein structure information with the post-mutation amino acid.

9. The method according to claim 1, characterized in that, The step of performing molecular docking between the protein structure information and the drug structure information to obtain the structure information of each wild-type complex corresponding to the protein structure information includes: When the protein structure information is associated with reference drug structure information, the molecular docking region information is determined based on the reference drug structure information; Based on the molecular docking region information, the protein structure information and the drug structure information are molecularly docked to obtain the structure information of each wild-type complex corresponding to the protein structure information.

10. The method according to claim 1, characterized in that, The step of performing molecular docking between the protein structure information and the drug structure information to obtain the structure information of each wild-type complex corresponding to the protein structure information includes: When there is no associated reference drug structure information for the protein structure information, the docking region information of the target molecule is determined based on the drug structure information; According to the target molecule docking region information, the protein structure information and the drug structure information are molecularly docked to obtain the structure information of each wild-type complex corresponding to the protein structure information.

11. The method according to claim 1, characterized in that, The step of performing molecular docking between the mutant protein structure information and the drug structure information to obtain the structure information of each mutant complex corresponding to the mutant protein structure information includes: Based on the drug structure information, determine the docking region information of the mutant molecule; According to the information of the mutant molecule docking region, the structure information of the mutant protein is molecularly docked with the structure information of the drug to obtain the structure information of each mutant complex corresponding to the structure information of the mutant protein.

12. A complex information generation device, characterized in that, The device includes: An information acquisition module is used to acquire basic drug resistance information, which includes the storage address of protein structure information, mutation information, and the storage address of initial drug structure information; to acquire protein structure information from the protein structure information storage address and to acquire initial drug structure information from the initial drug structure information storage address; and to perform structure optimization based on the initial drug structure information to obtain drug structure information. The mutation module is used to adjust the protein structure information according to the mutation information to obtain mutant protein structure information; The molecular docking module is used to perform molecular docking between the protein structure information and the drug structure information to obtain the structure information of each wild-type complex corresponding to the protein structure information, and to perform molecular docking between the mutant protein structure information and the drug structure information to obtain the structure information of each mutant complex corresponding to the mutant protein structure information. The wild-type complex determination module is used to calculate the wild-type energy information corresponding to the structure information of each wild-type complex, and to determine the target wild-type complex structure information corresponding to the protein structure information based on the wild-type energy information from the structure information of each wild-type complex. The mutant complex determination module is used to calculate the mutant energy information corresponding to the structural information of each mutant complex, and to determine the target mutant complex structural information corresponding to the mutant protein structural information from the structural information of each mutant complex based on the mutant energy information; The target complex information acquisition module is used to obtain target complex information based on the target wild-type complex structure information and the target mutant complex structure information.

13. The complex information generation apparatus according to claim 12, characterized in that, The device further includes a model training module, which is used to obtain the drug resistance label corresponding to the target complex information; and to extract drug resistance features based on the target complex information, the drug structure information, the protein structure information and the mutant protein structure information to obtain the target drug resistance features. The target drug resistance characteristics are input into the initial drug resistance prediction model to predict drug resistance and obtain the initial drug resistance prediction results. Based on the initial drug resistance prediction results and the drug resistance labels, loss calculation is performed to obtain loss information. The initial drug resistance prediction model is then updated in reverse based on the loss information until the training completion condition is met, at which point the target drug resistance prediction model is obtained.

14. The complex information generation apparatus according to claim 13, characterized in that, The device further includes a model prediction module, which is used to acquire data to be predicted, including the structure information of the protein to be predicted, the structure information of the mutant protein to be predicted, the structure information of the drug to be predicted, and the information of the complex to be predicted. Based on the information of the complex to be predicted, the structure information of the protein to be predicted, the structure information of the drug to be predicted, and the structure information of the mutant protein to be predicted, drug resistance features are extracted to obtain the drug resistance features to be predicted; the drug resistance features to be predicted are input into the target drug resistance prediction model to predict drug resistance and obtain the drug resistance prediction result.

15. The complex information generation apparatus according to claim 13, characterized in that, The model training module is also used to extract physicochemical property features based on the target complex information, the drug structure information, the protein structure information, and the mutant protein structure information to obtain the target physicochemical property features; Interaction energy features are extracted based on the target complex information, the drug structure information, the protein structure information, and the mutant protein structure information to obtain the target energy features; The target drug resistance characteristics are obtained based on the target's physicochemical properties and energy characteristics.

16. The complex information generation apparatus according to claim 12, characterized in that, The information acquisition module is also used to download a target file from the protein structure information storage address, the target file including protein structure information and reference drug structure information associated with the protein structure information; The protein structure information and the reference drug structure information are separated according to the annotation information in the target file to obtain the protein structure information.

17. The complex information generation apparatus according to claim 12, characterized in that, The information acquisition module is also used to delete hydrogen atom information from the initial structure information of the drug to obtain preliminary optimized structure information; The polar hydrogen atom information is obtained, and the polar hydrogen atom information is combined with the preliminary optimized structure information to obtain the drug structure information.

18. The complex information generation apparatus according to claim 12, characterized in that, The protein structure information corresponds to at least two mutation information; the mutation module is also used to adjust the protein structure information according to the at least two mutation information to obtain the target mutant protein structure information.

19. The complex information generation apparatus according to claim 12, characterized in that, The mutation information includes the mutation location, the amino acid before the mutation, and the amino acid after the mutation; the mutation module is also used to replace the amino acid before the mutation corresponding to the mutation location in the protein structure information with the amino acid after the mutation to obtain the mutant protein structure information.

20. The complex information generation apparatus according to claim 12, characterized in that, The molecular docking module is also used to determine molecular docking region information based on the reference drug structure information when the protein structure information has associated reference drug structure information; and to perform molecular docking between the protein structure information and the drug structure information according to the molecular docking region information to obtain the structure information of each wild-type complex corresponding to the protein structure information.

21. The complex information generation apparatus according to claim 12, characterized in that, The molecular docking module is also used to determine the target molecular docking region information based on the drug structure information when there is no associated reference drug structure information for the protein structure information; and to perform molecular docking between the protein structure information and the drug structure information according to the target molecular docking region information to obtain the structure information of each wild-type complex corresponding to the protein structure information.

22. The complex information generation apparatus according to claim 12, characterized in that, The molecular docking module is also used to determine the mutant molecule docking region information based on the drug structure information; and to perform molecular docking between the mutant protein structure information and the drug structure information according to the mutant molecule docking region information to obtain the structure information of each mutant complex corresponding to the mutant protein structure information.

23. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 11.

24. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.

25. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.