Biomolecular structure prediction method, device, equipment and medium
By judging the basic information of the biomolecular structure prediction request, and using a combination of generation algorithms and database queries with advanced tools, the problem of low efficiency in biomolecular structure prediction in existing technologies is solved, achieving efficient and accurate biomolecular structure prediction, adapting to different input conditions, and providing unified database management and tool selection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2026-03-24
AI Technical Summary
In the existing technology, there are many kinds of biomolecular structure prediction tools, making it difficult for users to choose the right tool. Public databases do not cover all structural information, and it is difficult to choose a suitable database for retrieval when the basic information is known, resulting in low efficiency of biomolecular structure prediction.
By determining whether a biomolecule structure prediction request contains basic information, the system generates biomolecules using a preset biomolecule sequence generation algorithm or queries structures from a preset database. It then combines tools such as AlphaFold and Rosetta for structure prediction, automatically selects the most suitable tool, integrates it into the platform, and provides a unified database construction and management system.
It improves the efficiency and accuracy of biomolecular structure prediction, saves computational resources, makes full use of existing data, adapts to different input conditions, and provides an integrated biomolecular structure prediction solution.
Smart Images

Figure CN121725868A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of biomolecules, and in particular to a biomolecule structure prediction method, device, equipment and medium. BACKGROUND
[0002] Synthetic biology is in a period of rapid development, and polypeptides, enzymes, nucleic acids and other biomolecules are important research objects. In related research, the structures of these biomolecules are necessary, and based on the structures of these biomolecules, more and richer research can be carried out. Enterprises that need to predict the structures of biomolecules may encounter the following scenarios and needs, including but not limited to:
[0003] 1) Business needs: there are many existing public databases, and it is difficult to choose which database to search when some basic information of biomolecules is known;
[0004] 2) Business needs: public databases do not cover all known biomolecular structure information, and searching cannot achieve the goal;
[0005] 3) Business needs: there are many existing biomolecular structure prediction tools, and there is no platform to provide a complete set of tools;
[0006] 4) Business needs: there are many existing biomolecular structure prediction tools, and it is difficult for general users to select appropriate methods according to actual conditions;
[0007] 5) Business needs: when natural biomolecules cannot meet actual needs, it is difficult to continue related work;
[0008] In summary, how to achieve simple and direct prediction of biomolecular structure according to biomolecular structure prediction request is a technical problem to be solved in the field. SUMMARY
[0009] Therefore, the purpose of the present application is to provide a biomolecule structure prediction method, device, equipment and medium, which can achieve simple and direct prediction of biomolecular structure according to biomolecular structure prediction request. The specific scheme is as follows:
[0010] In a first aspect, the present application discloses a biomolecule structure prediction method, comprising:
[0011] determining whether the received biomolecule structure prediction request contains biomolecule basic information;
[0012] If the biomolecule structure prediction request does not contain the biomolecule basic information, a preset biomolecule sequence generation algorithm is used to generate a corresponding biomolecule according to biomolecule generation information of the biomolecule structure prediction request, and then the biomolecule is input into a target biomolecule structure prediction tool to obtain a corresponding biomolecule structure prediction result.
[0013] If the biomolecule structure prediction request contains the biomolecule basic information, a corresponding biomolecule structure is queried from a preset database based on the biomolecule basic information to obtain a corresponding query result.
[0014] If the query result is that no corresponding biomolecule structure is queried, the biomolecule basic information is input into a target biomolecule structure prediction tool to obtain a corresponding biomolecule structure prediction result.
[0015] If the query result is that the corresponding biomolecule structure is queried, the corresponding biomolecule structure is directly output as a biomolecule structure prediction result.
[0016] Optionally, the biomolecule structure prediction method further comprises:
[0017] The biomolecule basic information and biomolecule structure information of each biomolecule after standardization processing are collected, wherein the biomolecule basic information includes different biomolecule types of protein structure, deoxyribonucleic acid structure and ribonucleic acid structure, and structure complexity; and the biomolecule structure information includes biomolecule three-dimensional structure coordinates, biomolecule chain information and biomolecule sequence information.
[0018] An initial biomolecule basic information table and an initial biomolecule structure information table are constructed.
[0019] The biomolecule basic information and corresponding biomolecule structure information of each biomolecule are written into the initial biomolecule basic information table and the initial biomolecule structure information table respectively to obtain a target biomolecule basic information table and a target biomolecule structure information table.
[0020] The target biomolecule basic information table and the target biomolecule structure information table are table-associated through biomolecule unique identifiers of each biomolecule, and the table-associated target biomolecule basic information table and target biomolecule structure information table are stored into a preset database.
[0021] Optionally, the table-associated target biomolecule basic information table and target biomolecule structure information table are stored into a preset database, which comprises:
[0022] According to the biomolecule type of the biomolecule, the corresponding biomolecule basic information in the target biomolecule basic information table and the corresponding biomolecule structure information in the target biomolecule structure information table are processed, and each partitioned information table after the table association is stored in a preset database.
[0023] Optionally, if the query result is that no corresponding biomolecule structure is found, the biomolecule basic information is input into a target biomolecule structure prediction tool to obtain a corresponding biomolecule structure prediction result, including:
[0024] If the query result is that no corresponding biomolecule structure is found, a target biomolecule structure prediction tool is calculated and selected according to the target biomolecule type, structure complexity, computing resources and corresponding weight coefficients in the biomolecule basic information, and the biomolecule basic information is input into the target biomolecule structure prediction tool to obtain a corresponding biomolecule structure prediction result.
[0025] Optionally, if the biomolecule structure prediction request does not contain the biomolecule basic information, a preset biomolecule sequence generation algorithm is used to generate a corresponding biomolecule according to the biomolecule generation information of the biomolecule structure prediction request, including:
[0026] If the biomolecule structure prediction request does not contain the biomolecule basic information, the length information of the biomolecule to be generated, the amino acid arrangement sequence of the biomolecule to be generated, and the combination information of the amino acid residues of the biomolecule to be generated are obtained from the biomolecule structure prediction request.
[0027] A trained variational autoencoder is used to generate a corresponding biomolecule according to the length information, the amino acid arrangement sequence, and the combination information of the amino acid residues.
[0028] Or, a trained generative adversarial network is used to generate a corresponding biomolecule according to the length information, the amino acid arrangement sequence, and the combination information of the amino acid residues.
[0029] Optionally, the biomolecule is input into a target biomolecule structure prediction tool to obtain a corresponding biomolecule structure prediction result, including:
[0030] The biomolecule is input into the target biomolecule structure prediction tool, so that the target biomolecule structure prediction tool performs structure prediction on the biomolecule to obtain a three-dimensional structure of the biomolecule.
[0031] Optionally, the biomolecule structure prediction method further includes:
[0032] free energy of a biological molecule structure prediction result is calculated so as to evaluate stability of the biological molecule according to a numerical result of the free energy;
[0033] when the biological molecule structure prediction result is an enzyme structure prediction result, binding of the enzyme structure prediction result with a target substrate is simulated so as to verify a function of the enzyme structure prediction result according to a binding ability.
[0034] In a second aspect, the present application discloses a biological molecule structure prediction device, comprising:
[0035] an information judging module, configured to judge whether the biological molecule basic information is contained in the received biological molecule structure prediction request;
[0036] a first prediction module, configured to, if the biological molecule structure prediction request does not contain the biological molecule basic information, generate a corresponding biological molecule by using a preset biological molecule sequence generation algorithm and according to biological molecule generation information of the biological molecule structure prediction request, and then input the biological molecule into a target biological molecule structure prediction tool to obtain a corresponding biological molecule structure prediction result;
[0037] a second prediction module, configured to, if the biological molecule structure prediction request contains the biological molecule basic information, query a corresponding biological molecule structure from a preset database based on the biological molecule basic information to obtain a corresponding query result;
[0038] a third prediction module, configured to, if the query result is that the corresponding biological molecule structure is not queried, input the biological molecule basic information into the target biological molecule structure prediction tool to obtain a corresponding biological molecule structure prediction result;
[0039] a fourth prediction module, configured to, if the query result is that the corresponding biological molecule structure is queried, directly output the corresponding biological molecule structure as a biological molecule structure prediction result.
[0040] In a third aspect, the present application discloses an electronic device, comprising:
[0041] a memory, configured to save a computer program;
[0042] a processor, configured to execute the computer program to realize steps of the biological molecule structure prediction method disclosed above.
[0043] In a fourth aspect, the present application discloses a computer readable storage medium, configured to save a computer program; wherein the computer program is executed by a processor to realize steps of the biological molecule structure prediction method disclosed above.
[0044] It can be seen that the application discloses a biomolecule structure prediction method, which comprises the following steps: judging whether the biomolecule basic information is contained in the received biomolecule structure prediction request; if the biomolecule structure prediction request does not contain the biomolecule basic information, generating the corresponding biomolecule by using a preset biomolecule sequence generation algorithm and according to the biomolecule generation information of the biomolecule structure prediction request, and then inputting the biomolecule into a target biomolecule structure prediction tool to obtain the corresponding biomolecule structure prediction result; if the biomolecule structure prediction request contains the biomolecule basic information, querying the corresponding biomolecule structure from a preset database based on the biomolecule basic information to obtain the corresponding query result; if the query result is that the corresponding biomolecule structure is not queried, inputting the biomolecule basic information into the target biomolecule structure prediction tool to obtain the corresponding biomolecule structure prediction result; and if the query result is that the corresponding biomolecule structure is queried, directly outputting the corresponding biomolecule structure as the biomolecule structure prediction result. It can be seen that, by judging whether the biomolecule basic information is contained in the request, the subsequent processing flow can be quickly determined. For the case that the basic information is already available, the database is directly queried, unnecessary calculation and generation steps are avoided, and time and calculation resources are saved. By using the preset biomolecule sequence generation algorithm and the target biomolecule structure prediction tool, the prediction can be performed based on scientific and reasonable methods and advanced technology, and the accuracy of the prediction result is improved. When the basic information is contained and the corresponding structure is in the database, the result is directly outputted, the existing data accumulation is fully utilized, the same biomolecule structure is avoided from being repeatedly predicted, and the reliability of the result is ensured. The method can adapt to different input conditions, and has corresponding processing modes whether the basic information is contained or not, so that the method can be widely applied to various biomolecule structure prediction scenes. BRIEF DESCRIPTION OF DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only belong to the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on the provided drawings.
[0046] Figure 1 A biomolecule structure prediction method flow chart disclosed by the present application;
[0047] Figure 2 A specific database construction method flow chart disclosed by the present application;
[0048] Figure 3 A biomolecule structure prediction platform and platform part function schematic diagram disclosed by the present application;
[0049] Figure 4 A schematic diagram of a biomolecule structure prediction device disclosed in the present application;
[0050] Figure 5 A structural diagram of an electronic device disclosed in the present application. DETAILED DESCRIPTION
[0051] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the protection scope of the present application.
[0052] Synthetic biology is in a period of rapid development, and polypeptides, enzymes, nucleic acids and other biomolecules are important research objects. In related research, the structures of these biomolecules are necessary, and more and more research can be carried out based on the structures of these biomolecules. Enterprises that need to predict the structures of biomolecules may encounter the following scenarios and needs, including but not limited to:
[0053] 1) Business needs: There are many existing public databases, and it is difficult to choose which database to search when some basic information of biomolecules is known.
[0054] 2) Business needs: The public database does not cover all known biomolecular structure information, and the target cannot be achieved through searching;
[0055] 3) Business needs: There are many existing biomolecular structure prediction tools, and there is no platform to provide a complete set of tools;
[0056] 4) Business needs: There are many existing biomolecular structure prediction tools, and it is difficult for general users to select appropriate methods according to actual conditions;
[0057] 5) Business needs: When natural biomolecules cannot meet actual needs, it is difficult to continue related work;
[0058] Therefore, the present application provides a biomolecule structure prediction scheme, which can realize simple and direct prediction of biomolecule structure according to a biomolecule structure prediction request.
[0059] Reference Figure 1 As shown in the figure, the embodiment of the present application discloses a biomolecule structure prediction method, comprising:
[0060] Step S11: determining whether the received biomolecule structure prediction request contains biomolecule basic information.
[0061] In this embodiment, the biomolecule structure prediction platform acquires the biomolecule structure prediction request initiated by the user through the human-computer interaction channel, so that the biomolecule structure prediction platform determines whether the biomolecule basic information is directly contained in the biomolecule structure prediction request, and then facilitates the biomolecule structure prediction platform to determine the next step process for the current biomolecule structure prediction request. Wherein, the biomolecule includes biological macromolecules and biological small molecules, and the biological macromolecules are taken as an example in this embodiment: the biological macromolecules can include but are not limited to: proteins, deoxyribonucleic acids, ribonucleic acids, etc., and correspondingly, the biomolecule basic information includes different biomolecule types of protein structure, deoxyribonucleic acid structure, ribonucleic acid structure, and structural complexity.
[0062] Step S12: If the biomolecule structure prediction request does not contain the biomolecule basic information, a preset biomolecule sequence generation algorithm is used to generate a corresponding biomolecule according to the biomolecule generation information of the biomolecule structure prediction request, and then the biomolecule is input into a target biomolecule structure prediction tool to obtain a corresponding biomolecule structure prediction result.
[0063] In this embodiment, if the biomolecule structure prediction request does not contain the biomolecule basic information, the length information of the biomolecule to be generated, the amino acid arrangement sequence of the biomolecule to be generated, and the combination information of the amino acid residues of the biomolecule to be generated are obtained from the biomolecule structure prediction request. A trained variational autoencoder is used to generate a corresponding biomolecule according to the length information, the amino acid arrangement sequence, and the combination information of the amino acid residues. Alternatively, a trained generative adversarial network is used to generate a corresponding biomolecule according to the length information, the amino acid arrangement sequence, and the combination information of the amino acid residues. It can be understood that if the biomolecule structure prediction request does not contain the biomolecule basic information, the corresponding biomolecule structure cannot be directly queried from the preset database, and the biomolecule sequence generation module of the biomolecule structure prediction platform needs to generate a corresponding biomolecule based on the length information of the biomolecule to be generated, the amino acid arrangement sequence of the biomolecule to be generated, and the combination information of the amino acid residues of the biomolecule to be generated in the biomolecule structure prediction request. The biomolecule sequence generation module is a deep learning-based generative model for generating biomolecule sequences, such as a variational autoencoder (VAE) or a generative adversarial network (GAN). Specifically, the following examples are used to generate a corresponding biomolecule: define user requirements and specifications: first, collect user requirements through a user interface. For example, the user may wish to generate a protein of a specific length or a protein containing a specific amino acid sequence. These requirements and specifications will be used to guide the design of the generation algorithm to ensure that the generated biomolecule meets the user's requirements. Design and implement the generation algorithm: select an appropriate generation algorithm. For example, use deep learning-based generative models such as variational autoencoders or generative adversarial networks, which can learn and generate new biomolecule sequences. Implement the generation algorithm to generate biomolecule sequences that meet the specifications based on user input requirements. The specific generation model training process is consistent with the general model training process and will not be described here.
[0064] In this embodiment, after obtaining the generated biomolecule, the biomolecule is input into the target biomolecule structure prediction tool, so that the target biomolecule structure prediction tool performs structure prediction on the biomolecule to obtain the three-dimensional structure of the biomolecule. It can be understood that after obtaining the biomolecule, the biomolecule structure prediction tool connected with the biomolecule sequence generation module in the biomolecule structure prediction platform is used to perform structure prediction on the biomolecule. Specifically, the generated biomolecule is subjected to structure prediction by using a structure prediction tool (such as AlphaFold or Rosetta), and these tools can predict the three-dimensional structure of the molecule according to the sequence. The three-dimensional structure is evaluated by using a calculation method, such as calculation of free energy, stability and functional simulation. These calculation methods can help to judge whether the generated molecule has the expected function and stability. It should be noted that before using the structure prediction tool to perform structure prediction, the existing biomolecule structure prediction tool is evaluated, including AlphaFold, Rosetta and the like. These tools are integrated into the platform and classified, for example, according to the biomolecule type or the prediction algorithm. A user interface is developed, so that the user can directly call and use these tools. It should be noted that the biomolecule structure prediction tool used in the present application can not only be AlphaFold, Rosetta and the like, but also other prediction tools and algorithms can be used for structure prediction and verification. For example, the alternative prediction tool may include SWISS-MODEL, I-TASSER and the like, which are not limited in detail.
[0065] The process of evaluating the existing biomolecule structure prediction tool is as follows: research and evaluate the current most advanced biomolecule structure prediction tool, such as AlphaFold, Rosetta and the like. Analyze the characteristics, application scope, performance and use limitations of each tool. Evaluate AlphaFold: a deep learning-based protein structure prediction tool with high precision, suitable for most proteins. Evaluate Rosetta: a protein structure prediction tool based on physical and statistical models, suitable for a wide range of biomolecule modeling tasks, including protein folding, design, etc. Make an evaluation report to record the advantages and disadvantages of each tool and the applicable scenarios. AlphaFold performs well in protein structure prediction, but may be limited in predicting complex multi-protein complexes; Rosetta is suitable for a variety of biomolecule modeling tasks, but requires more computing resources.
[0066] Step S13: If the biomolecule structure prediction request contains the biomolecule basic information, the corresponding biomolecule structure is queried from the preset database based on the biomolecule basic information to obtain the corresponding query result.
[0067] In this embodiment, if the biomolecule structure prediction platform detects that the biomolecule structure prediction request contains biomolecule basic information, the corresponding biomolecule structure is queried from the embedded preset database based on the biomolecule basic information to obtain the corresponding query result. It can be understood that, since the biomolecule structure request contains biomolecule basic information, the biomolecule structure corresponding to the biomolecule basic information can be directly queried from the preset database based on the biomolecule basic information, that is, the corresponding biomolecule structure can be obtained directly by database query, without the generation of biomolecules and the prediction of biomolecule structures.
[0068] Step S14: If the query result is that the corresponding biomolecule structure is not queried, the biomolecule basic information is input to the target biomolecule structure prediction tool to obtain the corresponding biomolecule structure prediction result.
[0069] In this embodiment, if the query result is that the corresponding biomolecule structure is not queried, the target biomolecule structure prediction tool is calculated and selected according to the target biomolecule type, structure complexity, computing resources and corresponding weight coefficient in the biomolecule basic information, and the biomolecule basic information is input to the target biomolecule structure prediction tool to obtain the corresponding biomolecule structure prediction result. It can be understood that, if the query result is that the preset database returns a null value or returns a return message of query failure, the AlphaFold tool or the Rosetta tool selected by the platform automatically: the generated biomolecule sequence is input to the selected structure prediction tool to predict its three-dimensional structure.
[0070] The purpose of intelligent tool selection is to automatically select the most suitable target biomolecule structure prediction tool according to the biomolecule basic information input by the user. For this purpose, an algorithm or model needs to be designed to consider factors such as biomolecule type, target structure complexity, computing resources, etc., and the automatic tool selection function is integrated into the biomolecule structure prediction platform. The following are the detailed implementation steps;
[0071] 1. Define selection criteria.
[0072] Determine the key factors affecting tool selection: biomolecule type, target structure complexity, computing resources, etc. Determine the weight of each factor and define the selection criteria.
[0073] Biomolecule type: Protein, Nucleic Acid, etc.
[0074] Target structure complexity: Simple, Moderate, Complex.
[0075] Computational resources: Low, Medium, High.
[0076] Weights and criteria: Biomolecule type (30%), complexity of target structure (40%), computational resources (30%).
[0077] 2. Design selection algorithm.
[0078] Applicability score: Define a scoring rule for each biomolecule structure prediction tool. For example, the AlphaFold tool is suitable for most protein structure prediction but may be limited for complex multi-protein complex prediction; the Rosetta tool is suitable for a wide range of biomolecule modeling tasks but requires more computational resources. The applicability score code is as follows:
[0079] def calculate_score(tool, bio_type, complexity, resources):
[0080] scores = {
[0081] 'AlphaFold': {'Protein': 1, 'NucleicAcid': 0.5, 'Simple': 1,'Moderate': 0.8, 'Complex': 0.5, 'Low': 1, 'Medium': 0.8, 'High': 0.5},
[0082] 'Rosetta': {'Protein': 0.8, 'NucleicAcid': 1, 'Simple': 0.8,'Moderate': 1, 'Complex': 1, 'Low': 0.5, 'Medium': 0.8, 'High': 1}
[0083] }
[0084] return (scores[tool][bio_type] * 0.3 +
[0085] scores[tool][complexity] * 0.4 +
[0086] scores[tool][resources] * 0.3)
[0087] def select_best_tool(bio_type, complexity, resources):
[0088] tools = ['AlphaFold', 'Rosetta']
[0089] best_tool = max(tools, key=lambda tool: calculate_score(tool,bio_type, complexity, resources))
[0090] return best_tool
[0091] 3. Implement the selection algorithm.
[0092] Implement the API (Application Programming Interface): Implement the selection algorithm in the backend of the biomolecular structure prediction platform and provide an API interface for the frontend to call.
[0093] from flask import Flask, request, jsonify
[0094] app = Flask(__name__)
[0095] @app.route(' / select_tool', methods=['POST'])
[0096] def select_tool():
[0097] data = request.get_json()
[0098] bio_type = data['bio_type']
[0099] complexity = data['complexity']
[0100] resources = data['resources']
[0101] best_tool = select_best_tool(bio_type, complexity, resources)
[0102] return jsonify({'best_tool': best_tool})
[0103] if __name__ == '__main__':
[0104] app.run()
[0105] 4. Integration into the platform.
[0106] Front-end interface: Adding an automatic selection function of structure prediction tools at the front end of the biomolecular structure prediction platform, enabling users to input information and automatically select the most suitable structure prediction tool.
[0107] Through the above steps, the intelligent tool selection function is successfully designed and implemented. According to the user input of basic information of biomolecules, the most suitable structure prediction tool is automatically selected and integrated into the biomolecular structure prediction platform, improving user experience and operation efficiency.
[0108] Step S15: If the query result is to query the corresponding biomolecular structure, output the corresponding biomolecular structure as the biomolecular structure prediction result directly.
[0109] In this embodiment, if the query result is to directly query the corresponding biomolecular structure in the preset database, it indicates that there is a biomolecular structure corresponding to the current biomolecular structure prediction request in the preset database, and there is no need to use the structure prediction tool for structure prediction, which shortens the reaction time.
[0110] In this embodiment, after obtaining the biomolecular structure prediction result, further to judge the prediction accuracy of the obtained biomolecular structure prediction result, the three-dimensional structure (biomolecular structure prediction result) needs to be verified. Specifically, the free energy of the biomolecular structure prediction result is calculated, so as to evaluate the stability of the biomolecule according to the numerical result of the free energy; when the biomolecular structure prediction result is an enzyme structure prediction result, the binding of the enzyme structure prediction result and the target substrate is simulated, so as to verify the function of the enzyme structure prediction result according to the binding ability. It can be understood that the free energy of the generated molecule is calculated using molecular dynamics simulation or energy calculation software (such as GROMACS), and the stability thereof is evaluated. The function of the generated molecule is evaluated by molecular docking simulation (such as AutoDock), and the binding ability thereof with the target molecule (such as the substrate) is judged.
[0111] Example description:
[0112] The user generates a new enzyme protein sequence. This sequence predicts its three-dimensional structure through the AlphaFold tool. Next, the GROMACS software is used to perform energy calculation on the predicted three-dimensional structure, and the result shows that the structure has a low free energy, indicating that it is stable. Finally, the AutoDock software is used to simulate the binding of the enzyme and its target substrate, and the result shows that the enzyme can effectively bind the substrate, verifying its function.
[0113] It is important to regularly evaluate the performance and accuracy of the forecasting tools currently in use (such as AlphaFold and Rosetta). Compare the forecast results of different tools and select the optimal tool or a combination of tools to improve forecast accuracy.
[0114] Example description:
[0115] The prediction results of AlphaFold and Rosetta are evaluated quarterly, using common biomolecular structures as test samples to compare their prediction accuracy and computational efficiency. Based on the evaluation results, a decision is made as to whether the tools need to be updated or replaced.
[0116] In addition, structural prediction tools need to be updated and maintained: Stay informed about newly released prediction tools and algorithms, and assess their application potential within the platform. Update the prediction tools on the platform to ensure the use of the latest algorithms and models.
[0117] Example description:
[0118] If new biomolecular structure prediction tools (such as new machine learning models) are discovered, they should be tested and evaluated. Once their superiority is confirmed, they should be integrated into the platform. Existing tools should be updated regularly to leverage their new features and improvements, thereby enhancing the platform's predictive capabilities.
[0119] Integrate generation and prediction tools: Design an interface to connect the generation algorithm with the structure prediction tool, enabling the generated biomolecular sequences to be directly input into the prediction tool for structure prediction. Ensure interface compatibility and data transmission accuracy to achieve a smooth workflow.
[0120] Example description:
[0121] The user inputs their generation requirements on the interface, and the generation algorithm generates a new biomolecule sequence. This sequence is automatically transmitted to AlphaFold for structure prediction. After the structure prediction is complete, the system transmits the prediction results to GROMACS for energy calculation. The entire process is seamless for the user; the user only needs to input their requirements, and the system can automatically complete the generation, prediction, and verification.
[0122] As can be seen, this application discloses a method for predicting biomolecular structures, comprising: determining whether a received biomolecular structure prediction request contains basic biomolecular information; if the biomolecular structure prediction request does not contain the basic biomolecular information, generating a corresponding biomolecular structure using a preset biomolecular sequence generation algorithm and based on the biomolecular generation information in the biomolecular structure prediction request, and then inputting the biomolecular structure into a target biomolecular structure prediction tool to obtain a corresponding biomolecular structure prediction result; if the biomolecular structure prediction request contains the basic biomolecular information, querying a corresponding biomolecular structure from a preset database based on the basic biomolecular information to obtain a corresponding query result; if the query result is that no corresponding biomolecular structure is found, inputting the basic biomolecular information into the target biomolecular structure prediction tool to obtain a corresponding biomolecular structure prediction result; and if the query result is that a corresponding biomolecular structure is found, directly outputting the corresponding biomolecular structure as the biomolecular structure prediction result. Therefore, by first determining whether the request contains basic biomolecular information, the subsequent processing flow can be quickly determined. For cases where basic information is already available, directly querying the database avoids unnecessary calculation and generation steps, saving time and computational resources. By utilizing a pre-defined biomolecular sequence generation algorithm and target biomolecular structure prediction tool, predictions can be made based on scientifically sound methods and advanced technology, improving the accuracy of the prediction results. When basic information is included and a corresponding structure exists in the database, the result is directly output, fully utilizing existing data accumulation, avoiding repeated predictions of the same biomolecular structure, and ensuring the reliability of the results. This method can adapt to different input conditions, whether basic information is included or not, and has corresponding processing methods, making it widely applicable to various biomolecular structure prediction scenarios.
[0123] Reference Figure 2 As shown, this embodiment of the invention discloses a specific database construction method. Compared with the previous embodiment, this embodiment further explains and optimizes the technical solution. Specifically:
[0124] Step S21: Collect the basic biomolecular information and biomolecular structural information of each biomolecule after standardization; wherein, the basic biomolecular information includes protein structure, deoxyribonucleic acid structure, different biomolecular types of ribonucleic acid structure, and structural complexity; the biomolecular structural information includes the three-dimensional structural coordinates of the biomolecule, biomolecular chain information, and biomolecular sequence information.
[0125] In this embodiment, the construction of a unified database first involved collecting biomolecular structural information from publicly available databases, such as protein structures and DNA / RNA structures. This data was then cleaned and standardized to ensure its integrity and consistency. Next, a data storage and management system was designed, selecting a suitable database management system and designing the database structure to ensure usability and scalability. Finally, a user interface and access control system were designed to provide users with convenient data access and operation functions while ensuring data security.
[0126] In this embodiment, in order to construct a unified database, the scope of the biomolecular structure database requirement is first determined;
[0127] Communication with Experts and Users: We first conducted face-to-face interviews and questionnaires with researchers, bioengineers, and medical experts in the biological field to understand their needs and expectations in the area of biomolecular structure. Experts and users expressed a need for information on various biomolecular structures, including structural data for different types of proteins, DNA, and RNA. They also expressed the importance of accessing high-quality, accurate, and timely data to support their research and project implementation. The target user group includes researchers, bioengineers, drug developers, and faculty and students in educational institutions.
[0128] In this embodiment, after determining the scope of requirements, a suitable database is selected: based on the determined scope of requirements, a publicly available database with widely recognized and high-quality data is selected, such as Protein Data Bank (PDB) and EMBL-EBI. Next, a data collection strategy is formulated, specifically designing the data collection strategy, clarifying the data types, time ranges, and quality standards to ensure that the acquired data meets the requirements. After selecting the publicly available database, the biomolecular structure data that meets the requirements is obtained by accessing the selected database's website or API, using search functions or query tools according to preset conditions and keywords. The acquired data is then screened and filtered to ensure it meets the preset requirements and quality standards. Data that meets the conditions is downloaded using the database's download function or API interface. The downloaded data is then organized and standardized to ensure consistency in data format and structure. The processed data is stored in a local or cloud database to support subsequent data analysis and applications.
[0129] In this embodiment, cleaning and standardization rules are defined for data stored locally or in the cloud. Specifically, cleaning and standardization rules and standards are formulated based on the scope of requirements and the characteristics of the data. These rules may include requirements for data format, field content, naming conventions, etc. For example, for protein structure data, naming conventions for proteins are set, such as protein names should conform to specific naming conventions, such as the UniProt identifier. After the rules are specified, the data cleaning steps are performed, as follows:
[0130] The collected data is traversed, and data cleaning is performed on each record. Duplicate data removal: Duplicate data entries are removed by comparing unique identifiers, such as IDs (Identity Documents). Error data handling: Errors in the data are identified and corrected, such as invalid characters or invalid structural coordinates. Missing data filling: For missing fields or data, appropriate handling is performed, and inferences or additions can be made based on contextual information.
[0131] The data standardization steps are as follows: Standardize the format and structure of the cleaned data to ensure consistency and comparability. Unify naming conventions: Implement unified naming conventions for field names and content to facilitate subsequent data management and analysis. Convert formats: Convert the data to a unified format, such as a standard data file format (e.g., PDB (Protein Data Bank) format) or a data table format, such as CSV (Comma-Separated Values) format.
[0132] The data validation steps are as follows: After cleaning and standardizing, perform data validation and checks to ensure the accuracy and integrity of the data. Automated tools or scripts can be used for data validation to detect potential errors or anomalies in the data.
[0133] Step S22: Construct the initial biomolecule basic information table and the initial biomolecule structure information table.
[0134] In this embodiment, the database structure is designed based on the scope of requirements and data characteristics. This includes determining the necessary tables and the relationships between them. For protein structure data, two main tables can be designed: one is a protein information table, containing basic protein information such as PDB ID, name, and description; the other is a structure information table, containing specific information about the protein structure, such as structural coordinates and chain information.
[0135] For example, design two initial tables: a Protein table and a Structure table. The Protein table stores basic information about proteins, such as PDB ID, name, and description. The Structure table stores detailed information about the protein structure, such as structural coordinates and chain information, and is associated with the Protein table through the PDB ID.
[0136] Based on the designed database structure, create the two initial tables using SQL statements or database management tools in the database management system, and then define the fields and attributes of each initial table.
[0137] The following is the code for creating the Protein and Structure tables in a database management system using SQL statements, and defining their fields and attributes:
[0138] CREATE TABLE Protein (
[0139] PDB_ID VARCHAR(10) PRIMARY KEY,
[0140] Name VARCHAR(100),
[0141] Description TEXT );
[0143] CREATE TABLE Structure (
[0144] Structure_ID INT PRIMARY KEY AUTO_INCREMENT,
[0145] PDB_ID VARCHAR(10),
[0146] Coordinate TEXT,
[0147] FOREIGN KEY (PDB_ID) REFERENCES Protein(PDB_ID)
[0148] ).
[0149] Step S23: Write the basic biomolecular information and the corresponding biomolecular structural information of each biomolecule into the initial biomolecular basic information table and the initial biomolecular structural information table, respectively, to obtain the target biomolecular basic information table and the target biomolecular structural information table.
[0150] In this embodiment, the cleaned protein structure data is imported into the Protein and Structure tables. Data is imported row by row into the corresponding tables using database management tools or scripts.
[0151] Step S24: Associate the target biomolecule basic information table and the target biomolecule structural information table using the unique molecular identifiers of each biomolecule, and store the associated target biomolecule basic information table and target biomolecule structural information table in a preset database.
[0152] In this embodiment, relationships between tables are established based on the relationships between the data. Foreign key constraints are used to ensure that the PDB IDs in the Structure table correspond to the PDB IDs in the Protein table, thus establishing the relationships between the tables. A suitable database management system is selected based on the data volume, performance requirements, and the team's technology stack. Common choices include MySQL (MyStructured Query Language, a relational database management system) and PostgreSQL (The World's Most Advanced Open Source Relational Database). This method uses PostgreSQL because it supports complex queries, is highly scalable, and offers good performance.
[0153] In this embodiment, the basic biomolecular information in the target biomolecular information table and the structural biomolecular information in the target biomolecular structure information table are partitioned based on the biomolecular type of the biomolecule, and the partitioned information tables are then stored in a preset database. It is understood that partitioning based on the protein's origin or type improves query efficiency. For example, partitioning can be based on protein classification (such as enzymes, receptors, etc.), as shown in the partitioning code below:
[0154] CREATE TABLE Protein (
[0155] PDB_ID VARCHAR(10) PRIMARY KEY,
[0156] Name VARCHAR(100),
[0157] Description TEXT,
[0158] Type VARCHAR(50)
[0159] PARTITION BY LIST (Type);
[0160] CREATE TABLE Enzyme_Partition PARTITION OF Protein FOR VALUES IN ('Enzyme');
[0161] CREATE TABLE Receptor_Partition PARTITION OF Protein FOR VALUES IN ('Receptor');
[0162] After data partitioning is complete, you can design regular full and incremental backup strategies to ensure data security. For example, perform incremental backups daily and full backups weekly.
[0163] # Full backup script
[0164] pg_dump -U username -F c -b -v -f / path / to / backup / full_backup.pgsqldbname.
[0165] # Incremental backup script (using rsync or other tools)
[0166] rsync -av --delete / path / to / data / path / to / backup / incremental.
[0167] Furthermore, it is necessary to continuously collect and integrate new biomolecular structural data, and maintain the update and integrity of the pre-set database. Prediction tools should be regularly evaluated and updated to ensure the accuracy and reliability of the prediction results provided by the platform. The platform's functionality and performance should be continuously improved and optimized to adapt to changes in user needs and technological developments.
[0168] Continuously collect and integrate new biomolecular structure data:
[0169] Data source and update mechanism:
[0170] Identify and select reliable biomolecular databases, such as Protein Data Bank (PDB) and EMBL-EBI. Regularly check these databases and collect newly released biomolecular structure data.
[0171] Data collection process:
[0172] Set up automated scripts or tools to periodically extract new data from these databases. Ensure data format consistency by converting new data into the format required by the platform.
[0173] Example description:
[0174] Set up a scheduled task to extract newly released biomolecular structure data from the PDB and EMBL-EBI databases weekly. Use a Python script to automatically download the data files, parse and convert them into the JSON format required by the platform, and store them in the database.
[0175] Data integration and validation:
[0176] Import new data into the platform database, ensuring its compatibility and relevance with existing data. Design a data validation mechanism to verify the integrity and accuracy of the new data.
[0177] Example description:
[0178] After the new data is downloaded, it is processed using data cleaning and standardization scripts to remove duplicate and erroneous data. Then, the processed data is imported into the database, and predefined validation rules are used to check its integrity and accuracy. For example, this includes checking whether protein sequences conform to standard formats and whether three-dimensional structure data is complete.
[0179] In addition, data access interfaces need to be designed. Specifically, the query interface should provide a REST API (Representational State Transfer API) or a GraphQL API (Graph Query Language Application Programming Interface) to allow users to query protein structure data. The update and delete interfaces should provide APIs to allow users to update and delete data while ensuring data integrity and security.
[0180] It's important to note that, to ensure data integrity and usability, a data verification mechanism and a user access control system are also implemented. Specifically, the data verification mechanism aims to ensure the integrity and accuracy of all new data. Data verification can be implemented at both the database and application layers.
[0181] At the database level: Use database constraints to verify data integrity. Define foreign keys, unique constraints, and check constraints.
[0182] Application level: Use triggers to implement complex logic checks.
[0183] The goal of a user access control system is to determine access permissions based on the different needs of users. It employs a role-based access control model.
[0184] Define roles and permissions: Roles: Administrator, Researcher, Student. Permissions: View, Edit, Delete, Add. Database level: Define user and role tables, and establish the mapping relationship between roles and permissions. Application level: Implement role-based permission checks at the application layer.
[0185] As can be seen, through the above steps, a protein structure data storage and management system based on PostgreSQL (preset database data storage and management) has been successfully designed, ensuring efficient data storage, secure backup and convenient access, and providing users with powerful data management and operation functions.
[0186] Reference Figure 3 As shown, a biological structure prediction platform and its various functional components are disclosed. The system architecture of the biological structure prediction platform is designed as follows: A plug-in architecture is designed to allow different tools to be seamlessly integrated as plug-ins. Each tool provides a standardized API interface for the biological structure prediction platform to call. API interfaces are developed for AlphaFold and Rosetta, encapsulating their core functionalities.
[0187] # Example: Calling the AlphaFold prediction interface
[0188] def run_alphafold(sequence):
[0189] response = requests.post('http: / / alphafold-api / predict', json={'sequence': sequence})
[0190] return response.json().
[0191] # Example: Calling the Rosetta prediction interface
[0192] def run_rosetta(sequence):
[0193] response = requests.post('http: / / rosetta-api / predict', json={'sequence': sequence})
[0194] return response.json()
[0195] Installation and configuration management:
[0196] Develop installation scripts and configuration files to ensure the tools function correctly on the platform.
[0197] # Example: AlphaFold installation script
[0198] git clone https: / / github.com / deepmind / alphafold.git
[0199] cd alphafold
[0200] pip install -r requirements.txt.
[0201] # Configuration file example
[0202] alphafold_config.yaml:
[0203] api_url: http: / / alphafold-api.
[0204] Classification tools:
[0205] Create a classification table: Create a tool classification table in the database to record the tool categories and related information.
[0206] Developing the user interface:
[0207] Design the user interface: Develop the user interface using React or other front-end frameworks. Provide a tool selection interface, an input parameter configuration interface, and a results display interface.
[0208] The system will integrate seamless biomolecule generation and structure prediction into the user interface. After the user inputs their generation requirements, the system will automatically generate the biomolecule sequence and perform structure prediction and energy calculations. The interface should be intuitive, user-friendly, and easy to operate.
[0209] Example description:
[0210] Users input the desired enzyme characteristics and functional requirements through the platform's user interface. The system first calls the generation algorithm to generate an enzyme sequence that meets the requirements, and then automatically inputs the sequence into AlphaFold for structure prediction. After the structure prediction is complete, the system inputs the predicted three-dimensional structure into GROMACS for energy calculation, and finally displays the predicted structure and energy calculation results. Users can view these results on the interface for further analysis and research.
[0211] To create an intuitive and user-friendly interface that allows users to easily access and use the platform's various functions, front-end and back-end development is required, along with ensuring data interaction and coordination between functional modules. The following are the detailed implementation steps:
[0212] The steps for designing a user interface are as follows:
[0213] User needs analysis: Communicate with users and domain experts to understand their needs for platform functionality and user interface. Collect user requirements, including required functions, use cases, common operations, and interface design preferences.
[0214] User interface prototype design:
[0215] Use prototyping tools such as Sketch, Figma, or Adobe XD to create user interface prototypes. Design intuitive navigation and layouts so users can easily find and use the functions they need. Ensure the interface is clean and aesthetically pleasing, avoiding excessive visual clutter.
[0216] Example description:
[0217] To design a user interface for a biomolecule generation and validation platform, a prototype diagram was first drawn. The main page includes three main modules: "Molecular Generation," "Structure Prediction," and "Energy Calculation." Each module has simple input boxes and buttons, allowing users to easily input parameters and view results. A navigation bar is placed at the top of the interface, allowing users to quickly switch between different functional modules.
[0218] Next, develop the platform's backend. First, select a suitable backend technology stack, such as Python's Django (Django Web Framework) or Flask (Flask Framework), or Node.js's Express (Express.js Framework). Then, choose an appropriate database management system, such as MySQL, PostgreSQL, or MongoDB (Mongo Database). Next, execute the backend functionality development steps: develop API endpoints to handle user requests and perform operations such as biomolecule generation, structure prediction, and energy calculation. Finally, ensure data security and integrity by implementing necessary verification and error handling mechanisms.
[0219] Example description:
[0220] Develop the platform's backend using the Django framework, creating API endpoints such as / generate_molecule, / predict_structure, and / calculate_energy. These endpoints accept user input parameters, invoke the corresponding algorithms and tools, and return the generated sequences, predicted structures, and calculated energy results.
[0221] Furthermore, the front-end development steps for the platform are as follows:
[0222] Choose a front-end technology stack:
[0223] Choose a suitable front-end technology stack for your platform, such as React, Vue.js, or Angular.
[0224] Front-end feature development:
[0225] Develop user interface components, including input fields, buttons, and result display areas. Implement interaction with the backend API, sending user requests and receiving and displaying the results returned by the backend.
[0226] Example description:
[0227] Using the React development platform, create front-end components such as MoleculeGenerator, StructurePredictor, and EnergyCalculator. Each component contains input fields and buttons. After the user inputs parameters, the front-end sends a request to the back-end API via Axios and displays the returned results on the interface.
[0228] To enable data interaction and coordination between functional modules:
[0229] Design data flow and interaction logic: Ensure smooth data interaction between the front-end and back-end by defining clear data flow and interaction logic. Use state management tools (such as Redux or Vuex) to manage the front-end data state.
[0230] Example description:
[0231] The frontend uses Redux to manage the application's global state, including generated molecular sequences, predicted structures, and calculated energies. When the user clicks the "Generate Molecule" button, the frontend saves the input parameters to the Redux state and sends a request to the backend API. After the backend returns the result, the frontend saves the result to the Redux state and updates the interface display.
[0232] It should be noted that, in addition to the frameworks mentioned above, a biomolecular structure prediction platform can also be built using the Java-based Spring framework or the .NET-based ASP.NET framework for the backend, and Angular or Vue.js for the frontend.
[0233] Conduct user testing and collect feedback:
[0234] Design a user testing plan:
[0235] Select representative users to conduct user tests and collect their feedback on platform usage. Design test tasks, observe the users' task completion process, and record any problems encountered and suggestions for improvement.
[0236] Analyzing user feedback:
[0237] Based on user feedback, identify the platform's strengths and weaknesses, and develop an improvement plan. Optimize the interface design and functionality, resolve user issues, and improve platform usability.
[0238] Example description:
[0239] Several biological researchers were selected as test users, and a series of test tasks were designed, such as "generating a new protein sequence and predicting its structure." The process of users completing the tasks was observed, the difficulties they encountered in operating the interface were recorded, and their suggestions for improvement were collected. Based on the feedback, the interface layout and interaction flow were adjusted to enhance the user experience.
[0240] Optimize and improve the platform's user experience and performance:
[0241] Iterative development and improvement:
[0242] We employ an iterative development approach, continuously making small improvements and releases to gradually optimize the platform. Regular performance testing and optimization are conducted to ensure the platform's stability and responsiveness under high loads.
[0243] User experience optimization:
[0244] We continuously collect user feedback and optimize the interface and functionality based on it. We add new features and improve existing features to meet the ever-changing needs of users.
[0245] Example description:
[0246] After the platform went live, a user satisfaction survey was conducted every few months to collect user feedback on new features and usage. Based on user feedback, the platform was updated regularly, adding new biomolecule generation algorithms and prediction tools, and optimizing existing functional modules. Performance testing was conducted to identify and resolve platform performance bottlenecks, ensuring that the platform could still respond quickly to user operations under high concurrency requests.
[0247] Based on user feedback and needs, we regularly add new features or optimize existing ones. We stay abreast of new technologies and trends in the field of biomolecular research and introduce new features in a timely manner.
[0248] Example description:
[0249] We collected user feedback and discovered that users wanted to add visualization capabilities for biomolecular structures. We developed and integrated a WebGL-based 3D visualization tool, enabling users to directly view and manipulate 3D structures on the platform. We regularly add new algorithms and tools to meet diverse user needs.
[0250] Performance optimization:
[0251] Perform performance monitoring and analysis to identify and resolve performance bottlenecks. Optimize database queries, API response times, and front-end loading speed to improve overall performance.
[0252] Example description:
[0253] Use performance monitoring tools (such as New Relic) to monitor various performance metrics of the platform and identify API endpoints with long response times. By optimizing database queries, adding caching mechanisms, and improving front-end code, the platform's response speed and user experience can be significantly improved.
[0254] User experience improvements:
[0255] Conduct regular user experience tests to collect user feedback on the platform. Continuously optimize the user interface and interaction design based on this feedback to improve usability and user satisfaction.
[0256] Example description:
[0257] We conduct user experience testing quarterly, inviting users from diverse backgrounds to participate and observe their behavior and feedback while using the platform. We identify overly complex aspects of the interface and simplify the processes and design. For example, we redesign the navigation menu to make it more intuitive and add tooltips and operation guides to help users get started faster.
[0258] By following these steps, we ensure the platform's data updates and integrity, the accuracy and reliability of predictive tools, and continuous optimization of functionality and performance. We continuously collect user feedback and technological advancements to adapt to changing user needs and technological advancements, providing efficient and reliable research tools.
[0259] Reference Figure 4 As shown, the present invention also discloses a biomolecular structure prediction device, comprising:
[0260] Information judgment module 11 is used to determine whether the received biomolecular structure prediction request contains basic biomolecular information;
[0261] The first prediction module 12 is used to generate a corresponding biomolecule using a preset biomolecule sequence generation algorithm and based on the biomolecule generation information of the biomolecule structure prediction request if the biomolecule structure prediction request does not contain the basic information of the biomolecule, and then input the biomolecule into the target biomolecule structure prediction tool to obtain the corresponding biomolecule structure prediction result.
[0262] The second prediction module 13 is used to query the corresponding biomolecular structure from a preset database based on the basic biomolecular information if the biomolecular structure prediction request contains the basic biomolecular information, so as to obtain the corresponding query result.
[0263] The third prediction module 14 is used to input the basic information of the biomolecule into the target biomolecule structure prediction tool if the query result is that no corresponding biomolecule structure is found, so as to obtain the corresponding biomolecule structure prediction result.
[0264] The fourth prediction module 15 is used to directly output the corresponding biomolecular structure as the biomolecular structure prediction result if the query result is that the corresponding biomolecular structure is found.
[0265] As can be seen, this application discloses a method for determining whether a received biomolecular structure prediction request contains basic biomolecular information. If the biomolecular structure prediction request does not contain the basic biomolecular information, a preset biomolecular sequence generation algorithm is used to generate the corresponding biomolecular structure based on the biomolecular generation information in the biomolecular structure prediction request. This biomolecular structure is then input into a target biomolecular structure prediction tool to obtain the corresponding biomolecular structure prediction result. If the biomolecular structure prediction request contains the basic biomolecular information, a corresponding biomolecular structure is queried from a preset database based on the basic biomolecular information to obtain the corresponding query result. If the query result indicates that no corresponding biomolecular structure was found, the basic biomolecular information is input into the target biomolecular structure prediction tool to obtain the corresponding biomolecular structure prediction result. If the query result indicates that a corresponding biomolecular structure was found, the corresponding biomolecular structure is directly output as the biomolecular structure prediction result. Therefore, by first determining whether the request contains basic biomolecular information, the subsequent processing flow can be quickly determined. For cases where basic information is already available, directly querying the database avoids unnecessary calculation and generation steps, saving time and computational resources. By utilizing a pre-defined biomolecular sequence generation algorithm and target biomolecular structure prediction tool, predictions can be made based on scientifically sound methods and advanced technology, improving the accuracy of the prediction results. When basic information is included and a corresponding structure exists in the database, the result is directly output, fully utilizing existing data accumulation, avoiding repeated predictions of the same biomolecular structure, and ensuring the reliability of the results. This method can adapt to different input conditions, whether basic information is included or not, and has corresponding processing methods, making it widely applicable to various biomolecular structure prediction scenarios.
[0266] Furthermore, embodiments of this application also disclose an electronic device, Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.
[0267] Figure 5This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the biomolecular structure prediction method disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0268] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0269] The processor 21 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 21 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0270] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0271] The operating system 221 manages and controls the various hardware devices and computer programs 222 on the electronic device 20 to enable the processor 21 to perform calculations and processing on the massive amounts of data 223 in the memory 22. The operating system 221 can be Windows Server, Netware, Unix, Linux, etc. The computer program 222, in addition to including a computer program capable of performing the biomolecular structure prediction method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, may further include computer programs capable of performing other specific tasks. The data 223 may include data received by the electronic device from external devices, as well as data collected by its own input / output interface 25.
[0272] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned disclosed method for predicting biomolecular structures. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0273] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0274] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly in hardware, software modules executed by a processor, or a combination of both. The software module may be located in random access memory (RAM), memory, read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs (Compact Disc-Read Only Memory), or any other form of storage medium known in the art.
[0275] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0276] The solution provided by the present invention has been described in detail above. Specific examples have been used to illustrate the principle and implementation of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core idea of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation and application scope based on the idea of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for predicting the structure of biomolecules, characterized in that, include: Determine whether the received biomolecular structure prediction request contains basic biomolecular information; If the biomolecule structure prediction request does not contain the basic information of the biomolecule, then a preset biomolecule sequence generation algorithm is used to generate the corresponding biomolecule according to the biomolecule generation information of the biomolecule structure prediction request. Then the biomolecule is input into the target biomolecule structure prediction tool to obtain the corresponding biomolecule structure prediction result. If the biomolecular structure prediction request includes the basic biomolecular information, then the corresponding biomolecular structure is queried from the preset database based on the basic biomolecular information to obtain the corresponding query results; If the query result is that no corresponding biomolecular structure is found, the basic biomolecular information is input into the target biomolecular structure prediction tool to obtain the corresponding biomolecular structure prediction result. If the query result is a corresponding biomolecular structure, then the corresponding biomolecular structure is directly output as the biomolecular structure prediction result.
2. The method for predicting biomolecular structures according to claim 1, characterized in that, Also includes: Collect standardized biomolecular basic information and biomolecular structural information for each biomolecule; wherein, the biomolecular basic information includes protein structure, deoxyribonucleic acid structure, different biomolecular types of ribonucleic acid structure, and structural complexity; the biomolecular structural information includes biomolecular three-dimensional structural coordinates, biomolecular chain information, and biomolecular sequence information; Construct an initial basic biomolecule information table and an initial biomolecule structure information table; The basic biomolecular information and the corresponding biomolecular structural information of each biomolecule are written into the initial biomolecular basic information table and the initial biomolecular structural information table, respectively, to obtain the target biomolecular basic information table and the target biomolecular structural information table. The basic information table and the structural information table of the target biomolecule are linked by the unique molecular identifier of each biomolecule, and the linked basic information table and structural information table of the target biomolecule are stored in a preset database.
3. The method for predicting biomolecular structures according to claim 2, characterized in that, The step of storing the target biomolecule basic information table and the target biomolecule structural information table, after table association, into a preset database includes: Based on the biomolecule type of the biomolecule, the corresponding basic biomolecule information in the target biomolecule basic information table and the corresponding biomolecule structural information in the target biomolecule structural information table are partitioned, and the partitioned information tables after table association are stored in a preset database.
4. The method for predicting biomolecular structures according to claim 1, characterized in that, If the query result indicates that no corresponding biomolecular structure was found, the basic biomolecular information is input into the target biomolecular structure prediction tool to obtain the corresponding biomolecular structure prediction result, including: If the query result is that no corresponding biomolecular structure is found, then the target biomolecular structure prediction tool is calculated and selected based on the target biomolecular type, structural complexity, computing resources and corresponding weight coefficients in the basic biomolecular information, and the basic biomolecular information is input into the target biomolecular structure prediction tool to obtain the corresponding biomolecular structure prediction result.
5. The method for predicting biomolecular structures according to claim 1, characterized in that, If the biomolecule structure prediction request does not contain the basic biomolecule information, then a corresponding biomolecule is generated using a preset biomolecule sequence generation algorithm and based on the biomolecule generation information from the biomolecule structure prediction request, including: If the biomolecule structure prediction request does not contain the basic information of the biomolecule, then the length information of the biomolecule to be generated, the amino acid sequence of the biomolecule to be generated, and the combination information of the amino acid residues of the biomolecule to be generated are obtained from the biomolecule structure prediction request. A trained variational autoencoder is used to generate corresponding biomolecules based on the length information, the amino acid sequence, and the combination information of the amino acid residues. Alternatively, an adversarial network can be generated after training, and corresponding biomolecules can be generated based on the length information, the amino acid sequence, and the combination information of the amino acid residues.
6. The method for predicting biomolecular structures according to claim 5, characterized in that, The step of inputting the biomolecule into a target biomolecule structure prediction tool to obtain the corresponding biomolecule structure prediction result includes: The biomolecule is input into the target biomolecule structure prediction tool so that the target biomolecule structure prediction tool can predict the structure of the biomolecule to obtain the three-dimensional structure of the biomolecule.
7. The method for predicting biomolecular structures according to any one of claims 1 to 6, characterized in that, Also includes: Calculate the free energy of the predicted biomolecular structure in order to assess the stability of the biomolecule based on the numerical results of the free energy. When the predicted biomolecular structure is the same as the predicted enzyme structure, the binding of the predicted enzyme structure to the target substrate is simulated to verify the function of the predicted enzyme structure based on the binding ability.
8. A biomolecular structure prediction device, characterized in that, include: The information judgment module is used to determine whether the received biomolecular structure prediction request contains basic biomolecular information; The first prediction module is used to generate a corresponding biomolecule using a preset biomolecule sequence generation algorithm and based on the biomolecule generation information in the biomolecule structure prediction request if the biomolecule structure prediction request does not contain the basic information of the biomolecule. Then, the biomolecule is input into the target biomolecule structure prediction tool to obtain the corresponding biomolecule structure prediction result. The second prediction module is used to query the corresponding biomolecular structure from a preset database based on the basic biomolecular information if the biomolecular structure prediction request contains the basic biomolecular information, so as to obtain the corresponding query results. The third prediction module is used to input the basic information of the biomolecule into the target biomolecule structure prediction tool if the query result is that no corresponding biomolecule structure is found, so as to obtain the corresponding biomolecule structure prediction result. The fourth prediction module is used to directly output the corresponding biomolecular structure as the biomolecular structure prediction result if the query result is a corresponding biomolecular structure found.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the biomolecular structure prediction method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the biomolecular structure prediction method as described in any one of claims 1 to 7.