Multimodal compound generation system, method, and customized compound generation system

By integrating gene data and protein structures through a multimodal compound generation system and using neural network models to generate compounds, the problem of insufficient binding affinity in traditional drug development has been solved, enabling more efficient drug design.

CN122290777APending Publication Date: 2026-06-26IND TECH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510622090.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-25
Filing Date
2025-05-15
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In traditional drug development, only single-modality data is used to design potential drug molecules, failing to fully consider the complex interactions between the drug and the target protein, resulting in a high failure rate.

Method used

By integrating gene data, protein secondary structure, and three-dimensional structure into a multimodal compound generation system, a neural network model is used for heterogeneous data fusion, feature extraction, and prediction to generate compounds that meet the requirements.

Benefits of technology

The generated compounds bind more strongly to the target protein, improving the success rate of drug development, and the quality of compound output is optimized by training a neural network model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290777A_ABST
    Figure CN122290777A_ABST
Patent Text Reader

Abstract

This disclosure provides a multimodal compound generation system and method that simultaneously considers different dimensions of the target protein's characteristics to generate compounds. These compounds exhibit stronger binding properties to the target protein compared to molecules produced by considering only a single dimension in the past. Furthermore, modified compounds can be generated from known compound molecules, and this can be used to train a neural network model to improve the quality of future compound generation. This disclosure also provides a multimodal custom compound generation system that allows users to input desired parameters to produce compounds that meet their expectations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a compound generation system, method, and customized compound generation system, and particularly to a multimodal compound generation system, method, and customized compound generation system. Background Technology

[0002] Traditional drug development typically uses single-modal data, such as genome sequences or protein structures, to design potential drug molecules. However, such methods may not adequately account for the complex interactions between the drug and the target protein, leading to a high failure rate during development. Therefore, incorporating drug-target protein interactions into drug development parameters is an important research topic. Summary of the Invention

[0003] In view of this, this disclosure provides a multimodal compound generation system, method, and multimodal customized compound generation system, which integrates gene data, protein secondary structure, and three-dimensional structure data to form a unified analytical framework, thereby designing compound molecules that meet specific requirements.

[0004] The disclosed multimodal compound generation system includes a memory and a processor. The memory stores multiple models. The processor is coupled to the memory and configured to: input multiple parameters to a neural network model among the multiple models, wherein a first correlation exists between a first parameter and a second parameter among the multiple parameters, and a second correlation exists between a second parameter and a third parameter among the multiple parameters; and obtain information about a first compound from the neural network model, wherein a first characteristic of the first compound is greater than a first threshold, and a second characteristic of the first compound is greater than a second threshold.

[0005] In one embodiment of this disclosure, the processor of the multimodal compound generation system described above is further configured to: execute a heterogeneous data fusion algorithm to preprocess multiple parameters to obtain multiple preprocessed parameters, wherein each of the multiple preprocessed parameters has the same format.

[0006] In one embodiment of this disclosure, the parameters of the multimodal compound generation system described above include one-dimensional structure, two-dimensional structure, and three-dimensional structure.

[0007] In one embodiment of this disclosure, the processor of the multimodal compound generation system described above is further configured to: execute a feature extraction model in a plurality of models to extract a plurality of features of the target compound, wherein the plurality of features are related to a plurality of parameters.

[0008] In one embodiment of this disclosure, the memory of the multimodal compound generation system further stores multiple modules, and the processor is further configured to: execute the prediction module among the multiple modules to predict the activity and stability of the first compound.

[0009] In one embodiment of this disclosure, the processor of the multimodal compound generation system is further configured to: execute a training module among multiple modules to train a neural network model based on information of a first compound to obtain a trained neural network model; and input multiple parameters into the trained neural network model to obtain information of a second compound.

[0010] In one embodiment of this disclosure, the processor of the multimodal compound generation system described above is further configured to: determine that the information of the first compound conforms to a first expectation.

[0011] In one embodiment of this disclosure, the first parameter of the multimodal compound generation system is a polypeptide sequence, the second parameter is the secondary structure of the protein, and the third parameter is the three-dimensional structure of the protein.

[0012] The multimodal compound generation method disclosed herein includes: inputting multiple parameters into a neural network model, wherein there is a first correlation between a first parameter and a second parameter among the multiple parameters, and a second correlation between a second parameter and a third parameter among the multiple parameters; and obtaining information about a first compound from the neural network model, wherein a first characteristic of the first compound is greater than a first threshold, and a second characteristic of the first compound is greater than a second threshold.

[0013] In one embodiment of this disclosure, the above-described multimodal compound generation method further includes: performing a heterogeneous data fusion algorithm to preprocess multiple parameters to obtain multiple preprocessed parameters, wherein each of the multiple preprocessed parameters has the same format.

[0014] In one embodiment of this disclosure, the parameters of the above-described multimodal compound generation method include one-dimensional structure, two-dimensional structure, and three-dimensional structure.

[0015] In one embodiment of this disclosure, the above-described multimodal compound generation method further includes: performing a feature extraction model to extract multiple features of the target compound, wherein the multiple features are related to multiple parameters.

[0016] In one embodiment of this disclosure, the above-described multimodal compound generation method further includes: executing a prediction module to predict the activity and stability of a first compound.

[0017] In one embodiment of this disclosure, the above-described multimodal compound generation method further includes: executing a training module to train a neural network model based on information of a first compound to obtain a trained neural network model; and inputting multiple parameters into the trained neural network model to obtain information of a second compound.

[0018] In one embodiment of this disclosure, the above-described multimodal compound generation method further includes: determining that the information of the first compound conforms to a first expectation.

[0019] In one embodiment of this disclosure, the first parameter of the above-described multimodal compound generation method is a polypeptide sequence, the second parameter is the secondary structure of the protein, and the third parameter is the three-dimensional structure of the protein.

[0020] The disclosed multimodal customized compound generation system includes a memory and a processor. The memory stores multiple models. The processor is coupled to the memory and configured to: acquire information about a first compound; acquire multiple basic parameters based on the first compound; and perform the aforementioned methods to acquire information about a second compound.

[0021] In one embodiment of this disclosure, the above-described multimodal customized compound generation system further includes a user interface coupled processor, and the processor is further configured to: obtain a first threshold and a second threshold through the user interface.

[0022] In one embodiment of this disclosure, the memory of the aforementioned multimodal customized compound generation system further stores multiple modules, wherein the processor is further configured to: execute a quality monitoring module among the multiple modules to determine whether the information of the second compound meets a quality threshold.

[0023] In one embodiment of this disclosure, the processor of the aforementioned multimodal customized compound generation system is further configured to: execute a performance monitoring module among multiple modules to monitor the training performance of a neural network model.

[0024] Based on the above, this disclosure provides a multimodal compound generation system and method that can simultaneously consider the characteristics of target proteins in different dimensions to generate compounds. These compounds exhibit stronger binding properties to the target protein compared to molecules produced by considering only a single dimension in the past. Furthermore, modified compounds can be generated from known compound molecules, and this can be used to train neural network models to improve the quality of future compound generation. This disclosure also provides a multimodal customized compound generation system for users to input desired parameters to produce compounds that meet their expectations. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the multimodal compound generation system disclosed herein.

[0026] Figure 2 This is a schematic diagram of the implementation process of the disclosed multimodal compound generation system.

[0027] Figure 3 This is a detailed flowchart of the disclosed process.

[0028] Figure 4 This is a schematic diagram of the implementation process of the disclosed multimodal custom compound generation system. Detailed Implementation

[0029] The exemplary embodiments of this disclosure will now be described in detail with reference to the accompanying drawings. The terms "first," "second," etc., used throughout this specification (including the scope of the claims) are used to name components or distinguish different embodiments or scopes, and are not intended to limit the upper or lower limit of the number of components, nor to limit the order of the components. Furthermore, wherever possible, components / members using the same reference numerals in the drawings and embodiments represent the same or similar parts.

[0030] Figure 1 This is a schematic diagram of a multimodal compound generation system provided in this disclosure. The multimodal compound generation system 100 may include a processor 110 and a memory 120.

[0031] In embodiments of this disclosure, the processor may be, for example, a central processing unit (CPU), or other programmable general-purpose or special-purpose microcontrollers (MCUs), microprocessors, digital signal processors (DSPs), programmable controllers, application-specific integrated circuits (ASICs), graphics processing units (GPUs), image signal processors (ISPs), image processing units (IPUs), arithmetic logic units (ALUs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other similar components or combinations thereof. In the multimodal compound generation system 100, the processor 110 may be coupled to memory 120, and the processor 110 may execute the modules and models stored in memory 120.

[0032] The memory 120 may be any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid state drive (SSD), or similar components or combinations thereof, and is used to store multiple modules, multiple models, or various applications that can be executed by the processor 110. In this embodiment, the memory 120 may store at least a neural network model 121.

[0033] Please refer to this simultaneously. Figure 2 , Figure 2This is a schematic diagram of the implementation process of the disclosed multimodal compound generation system, which can be implemented by processor 110. In process S210, processor 110 can input multiple parameters into a neural network model of multiple models, wherein there is a first correlation between the first parameter and the second parameter, and a second correlation between the second parameter and the third parameter. In process S220, processor 110 can obtain information about a first compound from the neural network model, wherein a first characteristic of the first compound is greater than a first threshold, and a second characteristic of the first compound is greater than a second threshold.

[0034] In detail, the composition of a molecule includes a one-dimensional atomic arrangement, a secondary structure resulting from hydrogen bonds, and a folded three-dimensional structure. In the past, drug development or compound synthesis only considered one dimension and ignored the influence of other dimensions, resulting in weak binding affinity between the designed compound molecules and the target protein (such as the mutant protein that causes lesions).

[0035] In the following embodiments, compounds (drugs) designed for amyotrophic lateral sclerosis (ALS) are described in detail. Abnormal aggregation of TDP-43 (TAR DNA-binding protein 43) is a common pathological feature of ALS, and this abnormal aggregation of the protein leads to degenerative changes in neurons.

[0036] In the embodiments of this disclosure, TDP-43 is first disassembled to obtain its one-dimensional amino acid sequence, secondary hydrogen bond structure, and three-dimensional folding structure. This may include obtaining the base pair sequence (gene sequence characteristics) regulating TDP-43 through gene library data, and then obtaining the transcribed and translated amino acid sequences. Furthermore, due to the determination of the amino acid sequence, hydrogen bonds in the amino acids can be formed accordingly, resulting in specific folding patterns, including α-helices and β-sheets. In addition, the three-dimensional structure of the TDP-43 protein can be obtained through X-ray crystallography or cryo-electron microscopy data. In this way, this disclosure can extract key features from the gene sequence characteristics, secondary hydrogen bond structure, and three-dimensional structure of the TDP-43 protein, or it can perform heterogeneous data processing on the aforementioned multi-dimensional parameters to avoid information loss due to compression of any dimension of the parameters, thereby generating corresponding compounds. In the embodiments of this disclosure, the one-dimensional amino acid sequence (peptide sequence), secondary hydrogen bond structure, and three-dimensional folding structure of TDP-43 all belong to one modality.

[0037] In embodiments of this disclosure, processor 110 can execute a heterogeneous data fusion algorithm to preprocess parameters of multiple different units to obtain multiple preprocessed parameters, ensuring that the units or formats of the preprocessed parameters are identical. For example, as described above, in order to integrate parameters including one-dimensional amino acid sequences, secondary hydrogen bond structures, and three-dimensional folded structures of different modalities, processor 110 can use a heterogeneous data fusion algorithm to integrate the aforementioned parameters of different units to convert each parameter into the same unit. The aforementioned one-dimensional amino acid sequence is a one-dimensional structure, the secondary hydrogen bond structure is a two-dimensional structure, and the three-dimensional folded structure is a three-dimensional structure.

[0038] Specifically, in the embodiments of this disclosure, since the one-dimensional amino acid sequence of a protein affects its secondary structure, and the secondary structure in turn affects its three-dimensional structure, there is at least a first correlation between the one-dimensional amino acid sequence and the secondary structure, and at least a second correlation between the secondary structure and the three-dimensional structure. Therefore, in the embodiments of this disclosure, at least the related features of the one-dimensional amino acid sequence and the secondary structure can be extracted and represented in the same format, and the related features of the secondary structure and the three-dimensional structure can be extracted and represented in the same format. Finally, after preprocessing, the format of each modality can be organized into the same representation.

[0039] Please see Figure 3 , Figure 3 This is a detailed flowchart of the disclosure. After inputting the one-dimensional sequence, two-dimensional structure, and three-dimensional geometry of the TDP-43 compound, processor 110 can preprocess each compound parameter to obtain a single format. Processor 110 can first encode the one-dimensional sequence and obtain features, molecular fingerprints, and atomic and molecular-level feature embeddings through the Simplified Molecular Input Line Entry Specification (SMILES) and SMILES-Transformer. Processor 110 performs patterning processing on the two-dimensional structure before encoding. Processor 110 also performs encoding and parsing on the three-dimensional geometry to obtain multiple pieces of information.

[0040] Please continue reading. Figure 3After the processor 110 encodes the compound parameters of the one-dimensional sequence, two-dimensional structure, and three-dimensional geometry respectively, the processor 110 processes the data using a neural network algorithm. It fuses the first encoding of the obtained one-dimensional atomic and molecular features with the second encoding of the two-dimensional structure to obtain a first related format. The processor 110 also fuses multiple pieces of information from the second encoding of the two-dimensional structure with the third encoding of the three-dimensional geometric structure to obtain a second related format. In embodiments of this disclosure, the processor 110 can process the above process using a Graph Neural Network (GNN). In other embodiments of this disclosure, the processor 110 can also perform the above process using other neural network algorithms.

[0041] Continuing from the previous section, the processor 110 further integrates the first related format and the second related format into a single format. Furthermore, this single format, when input into the neural network model, can be further decoded to generate compounds. In embodiments of this disclosure, the aforementioned neural network model may be, for example, a Transformer.

[0042] Furthermore, processor 110 can extract features using three chemical molecular fingerprinting methods. Processor 110 executes MACCS and extracts the structural and functional features of molecules (e.g., rings, chains, functional groups), representing each structure and functional feature with a 166-bit binary code. This is then combined with a classifier to evaluate compound activity. Processor 110 can also use PubChem to set corresponding binary positions to 1 or 0 based on the presence of specific substructures in the molecule, forming an 881-bit binary code, which is then used for rapid comparison and screening of compounds in a compound library. Processor 110 can also use ErG to identify key features of molecule-receptor interactions, including hydrogen bonds and hydrophobic regions. These key features are converted into radial distribution functions and integrated into a descriptor vector, which is then used for virtual screening of small molecules.

[0043] Following the previous section, the processor 110's encoding of three-dimensional geometric structures can include: Task 1: masking atomic features and reconstructing atoms; Task 2: masking three-dimensional coordinates and reconstructing coordinates; and Task 3: independently masking atoms and three-dimensional coordinates and reconstructing the masked parts. In Task 1, the processor 110 aims to enhance the model's ability to extract two-dimensional graphical information. Therefore, the processor 110 masks the chemical features of all atoms, uses only three-dimensional coordinates to predict atomic features, preserves the topological structure of the atomic structure, and calculates the loss using the cross-entropy loss function. In Task 2, the processor 110 aims to enhance the model's ability to generate three-dimensional coordinates and extract relevant three-dimensional information. Therefore, the processor 110 masks the three-dimensional coordinates of all atoms, uses only atomic features to predict the three-dimensional structure, and uses a loss function to consider rotation and permutation invariance. In Task 3, the processor 110 aims to improve the model's comprehensive information fusion capability. Therefore, the processor 110 masks atoms and reconstructs the structure, which combines the extreme cases of Task 1 and Task 2 by masking and reconstructing the atoms and structure of the molecular graph.

[0044] In the embodiments of this disclosure, after the processor 110 performs the aforementioned GNN processing, it can integrate one-dimensional sequences, secondary structures, and three-dimensional structures into graphical parameters and input them into a neural network model to produce compounds.

[0045] In embodiments of this disclosure, processor 110 can analyze target proteins to produce corresponding compound structures, and the compounds can bind to the target proteins. Processor 110 also executes a prediction module to predict the activity and stability of the compounds. Specifically, the molecular characteristics of the produced compounds include the number of atoms (num_atoms), molecular weight (mwt), oil-water partition coefficient (logP), number of hydrogen bond acceptors (HAcceptors), number of hydrogen bond donors (HDonors), topological polar surface area (TPSA), number of rotatable bonds (rotatable bonds), drug-likeness (QED), and syntheticizability score (SAscore). In embodiments of this disclosure, the molecular characteristics of the produced compounds must at least meet the thresholds of the aforementioned two characteristics; that is, the first characteristic must be greater than the first threshold, and the second characteristic must be greater than the second threshold for the obtained compound to meet the requirements. The aforementioned first and second thresholds may have different values ​​depending on the target protein.

[0046] In addition to the aforementioned molecular characteristics, since the method of generating compounds by analyzing target proteins alone may result in compounds with low activity or instability, for example, compounds generated by machine learning may contain only one-dimensional sequences. They may fold into three-dimensional structures that cannot target the protein due to instability after generation. In this case, the generated compound is not a stable protein.

[0047] As mentioned earlier, the structure of compounds generated through machine learning may also be unreasonable. For example, in order to simultaneously target the one-dimensional sequence, secondary structure, and three-dimensional structure of a target protein, a compound with a molecular weight much larger than the target protein might be generated, resulting in the compound being unable to effectively bind to the target protein. The structure of such a compound is thus unreasonable. Therefore, after processing the compound generated by the machine learning model, the processor 110 can execute an analysis module to determine whether the compound's information meets expectations. For example, the aforementioned expectation might be that the molecular weight of the compound is no more than one-tenth of the target protein's. Implementers of this disclosure can adjust the expected items and corresponding thresholds according to actual applications.

[0048] In another embodiment of this disclosure, processor 110 can modify a known compound to make the modified compound bind more strongly to the target protein than the originally known compound. Processor 110 can execute a feature extraction model to extract multiple features of the compound from the known compound. Since the goal of the compound is to bind to the target protein, the multiple features of the compound are correlated with the parameters of the protein's one-dimensional sequence, secondary structure, and three-dimensional structure. Processor 110 can use the multiple features of the known compound to train a neural network model to obtain a trained neural network model. In addition, after obtaining a recommended compound from the neural network model and determining that the recommended compound meets expectations, processor 110 can retrain the neural network model with the composition and feature information of the recommended compound to obtain a trained neural network model. Processor 110 also inputs multiple parameters of the target protein into the trained neural network model to obtain information about a second compound. Through the aforementioned training method, the neural network model can be optimized, and processor 110 can obtain optimized compounds from the optimized neural network model.

[0049] This disclosure also provides a method for generating multimodal compounds, which can be implemented by the processor 110 of a multimodal compound generation system. Specific embodiments have been described above and will not be detailed here.

[0050] This disclosure also provides a multimodal customized compound generation system, which includes a processor and a memory. In embodiments of this disclosure, the processor of the multimodal customized compound generation system can first obtain information about a known compound, and the processor obtains multiple basic parameters based on the known compound. Next, the processor can obtain a second compound by executing a multimodal compound generation method.

[0051] Please see Figure 4 , Figure 4This is a schematic diagram of the implementation flow of the disclosed multimodal custom compound generation system. In process 410, the processor generates filter condition settings. In process 420, the processor can accept SMILE molecules input by the user. In process 430, the processor generates parameter settings. In process 440, the processor generates molecules. In process 450, the processor generates a summary. In process 460, the processor generates a model fine-tuning mechanism. As shown in the figure, after process 460, the processor can perform another molecule generation to produce another molecule.

[0052] In detail, the purpose of a multimodal custom compound generation system is to provide customers with threshold values ​​for the properties of the desired compounds. For example, if a customer wants a compound with fewer than 200 atoms and a molecular weight below 5000, the customer can input the threshold values ​​for these properties through the user interface of the multimodal custom compound generation system. The user must input at least two property threshold values ​​to ensure that the compounds produced by the multimodal custom compound generation system better meet the user's needs.

[0053] Furthermore, in process 410, the processor can generate filtering condition settings. For example, the processor can select to filter molecules that are not ZINC15. The processor can also select to filter molecules that are not within the characteristic range. This initially filters out most of the molecules, increasing the speed at which the neural network model recommends compounds. Next, in process 420, multiple molecules can be input. That is, the user can input SMILE molecules according to their needs, so that the compounds produced by the neural network model include SMILE molecules. The generation parameter settings in process 430 include setting the number of input molecules and setting the number of samples. The user can input the number of molecules from 1 to the number of SMILE strings. For example, the number of SMILEs input is 50. The implementer of this disclosure can determine the input number based on the required number of molecules and the amount of computational resources available. The user can also set the number of samples produced by the neural network model, which can include at least one sample to a maximum of ten samples.

[0054] Following the previous section, in process 440, the method for generating compound molecules can employ a molecular generation model finely tuned for a specific molecule (based on the MegaMolBART architecture). The model executes generation via an API and provides multiple generation requirements. During the compound molecule generation process, the multimodal customized compound generation system allows users to view the structural diagram and molecular properties of the input SMILE molecule through a display interface (user interface). The display interface also shows the molecular properties of all generated molecules at each generation stage. Molecular properties include the number of atoms (num_atoms), molecular weight (mwt), oil-water partition coefficient (logP), number of hydrogen bond acceptors (HAcceptors), number of hydrogen bond donors (HDonors), topological polar surface area (TPSA), number of rotatable bonds (rotatableBonds), drug-likeness (QED), and syntheticizability score (SAscore). This allows users to determine if these molecular properties meet their requirements.

[0055] Continuing from the previous section, in process 450, the processor can summarize and display the structural diagrams and molecular properties of molecules generated in all stages through a display interface. In this way, users can use the expert annotation function on the user interface and, based on their own experience, molecular structural diagrams, and summarized molecular properties, provide molecular evaluations of the compounds recommended by the neural network model. It should be understood that this evaluation can serve as a training basis for further optimization of the neural network model. Therefore, in process 460, the fine-tuning mechanism of the neural network model generation model can collect all molecules with high synthetic performance evaluated by experts as molecular data for the next fine-tuning. Furthermore, the processor can also use the generation model of the previously generated compound molecules as a foundation, and using the aforementioned expert-evaluated molecular data, the processor can periodically fine-tune the generation model so that the neural network model can recommend more expected compound molecules in subsequent generation processes.

[0056] In process 460, it should be understood that users can directly issue commands (actively) to the multimodal customized compound generation system to fine-tune the parameters of its neural network model. The multimodal customized compound generation system can also automatically (passively) fine-tune its generation model periodically (e.g., once a month). The multimodal customized compound generation system can use the fine-tuned neural network model to recommend subsequent compound generation methods.

[0057] In this disclosure, the memory of the multimodal customized compound generation system also stores multiple modules, and the processor can execute the quality monitoring module among the multiple modules to determine whether the information of the second compound meets the quality threshold. That is, in process 460, the processor further highlights whether the molecular characteristics of the compound meet the quality threshold in a prominent manner in the summary portion of the compound.

[0058] In this disclosure, in order to reduce the time for generating compound recommendations and improve the quality of compounds recommended by the neural network model, the processor can also execute a performance monitoring module in multiple modules to monitor the training performance of the neural network model, thereby ensuring that the training results of the neural network model are geared towards quickly recommending high-quality compounds.

[0059] In summary, this disclosure, through a multimodal compound generation system and method, can simultaneously consider the characteristics of target proteins across different dimensions to generate compounds. These compounds exhibit stronger binding properties to the target protein compared to molecules produced by considering only a single dimension in the past. Furthermore, modified compounds can be generated from known compound molecules, and this can be used to train neural network models to improve the quality of future compound generation. This disclosure also provides a multimodal customized compound generation system for users to input desired parameters to produce compounds that meet their expectations.

[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit them. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this disclosure.

Claims

1. A multimodal compound generation system, characterized in that, include: The memory stores multiple models; as well as A processor, coupled to the memory, is configured to: Multiple parameters are input into a neural network model among the multiple models, wherein there is a first correlation between the first parameter and the second parameter among the multiple parameters, and there is a second correlation between the second parameter and the third parameter among the multiple parameters; as well as Information about a first compound is obtained from the neural network model, wherein a first characteristic of the first compound is greater than a first threshold, and a second characteristic of the first compound is greater than a second threshold.

2. The multimodal compound generation system according to claim 1, characterized in that, The processor is further configured to: A heterogeneous data fusion algorithm is executed to preprocess the plurality of parameters to obtain a plurality of preprocessed parameters, wherein each of the plurality of preprocessed parameters has the same format.

3. The multimodal compound generation system according to claim 2, characterized in that, The parameters include one-dimensional structure, two-dimensional structure, and three-dimensional structure.

4. The multimodal compound generation system according to claim 1, characterized in that, The processor is further configured to: The feature extraction model among the plurality of models is executed to extract a plurality of features of the target compound, wherein the plurality of features are related to the plurality of parameters.

5. The multimodal compound generation system according to claim 1, characterized in that, The memory also stores multiple modules, wherein the processor is further configured to: The prediction module among the plurality of modules is executed to predict the activity and stability of the first compound.

6. The multimodal compound generation system according to claim 5, characterized in that, The processor is further configured to: Execute the training module among the plurality of modules to train the neural network model based on information from the first compound to obtain a trained neural network model; and The multiple parameters are input into the trained neural network model to obtain information about the second compound.

7. The multimodal compound generation system according to claim 5, characterized in that, The processor is further configured to: The information of the first compound is determined to be consistent with the first expectation.

8. The multimodal compound generation system according to claim 1, characterized in that, The first parameter is the polypeptide sequence, the second parameter is the secondary structure of the protein, and the third parameter is the three-dimensional structure of the protein.

9. A method for generating a multimodal compound, characterized in that, include: Multiple parameters are input into a neural network model, wherein there is a first correlation between the first parameter and the second parameter among the multiple parameters, and there is a second correlation between the second parameter and the third parameter among the multiple parameters; as well as Information about a first compound is obtained from the neural network model, wherein a first characteristic of the first compound is greater than a first threshold, and a second characteristic of the first compound is greater than a second threshold.

10. The method for generating multimodal compounds according to claim 9, characterized in that, Also includes: A heterogeneous data fusion algorithm is executed to preprocess the plurality of parameters to obtain a plurality of preprocessed parameters, wherein each of the plurality of preprocessed parameters has the same format.

11. The method for generating multimodal compounds according to claim 10, characterized in that, The parameters include one-dimensional structure, two-dimensional structure, and three-dimensional structure.

12. The method for generating multimodal compounds according to claim 9, characterized in that, Also includes: A feature extraction model is executed to extract multiple features of the target compound, wherein the multiple features are related to the multiple parameters.

13. The method for generating multimodal compounds according to claim 9, characterized in that, Also includes: The prediction module is executed to predict the activity and stability of the first compound.

14. The method for generating multimodal compounds according to claim 13, characterized in that, Also includes: The training module is executed to train the neural network model based on the information of the first compound to obtain a trained neural network model; as well as The multiple parameters are input into the trained neural network model to obtain information about the second compound.

15. The method for generating multimodal compounds according to claim 13, characterized in that, Also includes: The information of the first compound is determined to be consistent with the first expectation.

16. The method for generating multimodal compounds according to claim 9, characterized in that, The first parameter is the polypeptide sequence, the second parameter is the secondary structure of the protein, and the third parameter is the three-dimensional structure of the protein.

17. A multimodal customized compound generation system, characterized in that, include: The memory stores multiple models; as well as A processor, coupled to the memory, is configured to: Obtain information about the first compound; Several basic parameters were obtained based on the first compound; and Perform the method according to any one of claims 9 to 16 to obtain information about the second compound.

18. The multimodal customized compound generation system according to claim 17, characterized in that, It also includes a user interface coupled to the processor, wherein the processor is further configured to: The first threshold and the second threshold are obtained through the user interface.

19. The multimodal customized compound generation system according to claim 17, characterized in that, The memory also stores multiple modules, and the processor is further configured to: The quality monitoring module among the multiple modules is executed to determine whether the information of the second compound meets the quality threshold.

20. The multimodal customized compound generation system according to claim 19, characterized in that, The processor is further configured to: The performance monitoring module among the plurality of modules is executed to monitor the training performance of the neural network model.