Protein aided design and analysis system and method based on cloud computing platform
By developing a highly integrated protein-assisted design and analysis system on the cloud computing platform, a variety of design-assisted software and intelligent algorithms are integrated, which solves the problem of low efficiency in the traditional protein design process, and realizes efficient and fast protein design and analysis, which lowers the technical threshold.
Patent Information
- Application Number
- CN202311616641.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
The traditional protein design process is inefficient, and scientific researchers need to have a large number of software construction, use and program writing capabilities. The threshold is high and cannot meet the efficient and fast biological research and development needs.
Based on the cloud computing platform, a highly integrated protein-assisted design and analysis system is developed, integrating complete design assistance software and intelligent algorithms, providing user interface modules, application warehouse modules, data storage and encryption modules, and resource scheduling modules, which reduces the R&D technology threshold and improves work efficiency.
It realizes efficient and fast protein structure and function design, reduces the requirements for the use of computer software, supports a variety of algorithms and functions, ensures data security, and is suitable for biological researchers who lack a computer background.
Smart Images

Figure CN120072027A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of protein design and information analysis, and relates to a protein assisted design and analysis system and method based on a cloud computing platform. Background Art
[0002] Proteins are the main executors of life functions. They are composed of amino acid sequences, and these amino acids are arranged in a specific order and rule to form a three-dimensional structure with specific functions. Protein engineering refers to the artificial modification or design of specific amino acids in protein molecules to generate new protein structures or functions, and is commonly used in the research and development of new drugs and biological materials.
[0003] Protein design and optimization is a complex task. When designing and optimizing protein sequences, researchers usually start from aspects such as amino acid sequences, three-dimensional conformations, and interactions. First, the amino acid sequence is the basis for protein structure design, and researchers construct different protein structures by changing the types and sequences of amino acids. Second, the three-dimensional conformation determines the function and stability of proteins. Finally, interaction is also a key factor in protein structure design. Proteins usually interact with other molecules, such as enzymes, receptors, and signal transducers.
[0004] Traditional protein design is carried out manually, generally including steps such as goal setting, sequence analysis, structure prediction, structure optimization, and experimental verification. Each step uses some independent auxiliary software tools respectively, such as AlphaFold, Rosetta, AutoDock, PyMOL, etc. The whole process requires the adjustment, screening, comparison, and verification of protein sequences, and requires scientific research workers to have a large amount of software building, using, and programming capabilities, with low work efficiency, so it cannot meet the requirements of efficient and fast biological research and development. Summary of the Invention
[0005] The purpose of the present invention is to provide a protein assisted design and analysis system and method based on a cloud computing platform, which is a highly integrated protein structure and function design platform, integrating complete design auxiliary software and intelligent algorithms, improving the research and development efficiency while reducing the research and development technical threshold.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] The first aspect of the present invention provides a protein assisted design and analysis system based on a cloud computing platform. The system is deployed on the cloud computing platform, and the system includes a front end and a back end;
[0008] The front end includes:
[0009] A user interface module for interacting with users and displaying protein design and / or analysis results;
[0010] The backend described above includes:
[0011] An application repository module for providing software or algorithms for protein design and / or analysis;
[0012] A data storage and encryption module for encrypting and storing data during the protein design and / or analysis process;
[0013] A resource scheduling module for providing computing power scheduling for applications running on the system.
[0014] Preferably, the system is deployed on a private cloud using container technology, and the private cloud includes multiple servers.
[0015] Preferably, the user interface module uses a web page as the interaction interface, including:
[0016] A user management page for setting user basic information and modifying passwords;
[0017] A project management page for creating, viewing, modifying, and deleting projects, and each project contains at least one protein analysis task;
[0018] A task execution page for selecting applications, configuring application parameters, inputting design and / or analysis files, and monitoring task execution progress;
[0019] A report display page for visually displaying protein design and / or analysis results.
[0020] Preferably, the applications in the application repository module are integrated into the cloud computing platform in the form of microservices, and the application repository module includes at least one of the following applications:
[0021] A protein basic parameter calculation application for calculating parameters describing protein structure and / or properties;
[0022] A structure prediction application for predicting the three-dimensional conformation of proteins;
[0023] A stability analysis application for analyzing the ability of proteins to maintain their original structure and function under different conditions;
[0024] A function analysis application for analyzing the functions of proteins in cells;
[0025] An interaction analysis application for analyzing the interactions between proteins;
[0026] A sequence design and optimization application for searching and generating potential amino acid sequences meeting specific structural and functional requirements;
[0027] Domain query application for querying the domains of proteins.
[0028] Preferably, the protein basic parameters include one or more of the sequence length of the protein, amino acid ratio statistics, molecular formula, molecular weight, isoelectric point, extinction coefficient, stability index, hydrophobicity index, aromaticity index, and protein flexibility; the structure prediction application adopts a deep neural network method based on MSA (Multiple Sequence Alignment) sequence features and / or a deep neural network method based on a pre-trained protein large language model PLM (Pre-trained protein Language Model); the stability analysis application predicts the change in protein stability before and after mutation by calculating the change in protein free energy ddG; the function analysis application infers the function of the protein by querying sequences and domains with significant similarity in the sequence; the interaction analysis application predicts the interaction between proteins by matching the interaction patterns mined from the protein sequence and / or structure database, and uses the protein sequence and / or structure information to predict the interaction sites and spatial conformations between proteins; the sequence design and optimization application searches and evaluates the protein sequence according to the target structure and / or function requirements set by the user to generate possible amino acid sequences that meet the requirements; the domain query application queries the domains of proteins according to the protein name and / or PDB (Protein Data Bank) number.
[0029] Preferably, in the data storage and encryption module, data encryption includes data transmission encryption and data storage encryption. The data transmission encryption is based on HTTPS (Hypertext Transfer Protocol Secure) technology, and the data storage encryption uniformly manages and encrypts the data through an object storage service and an encryption algorithm.
[0030] Preferably, the object storage service is MinIO, and the encryption algorithm is AES-256.
[0031] Preferably, in the resource scheduling module, the scheduling policies include label policies and / or priority policies.
[0032] The second aspect of the present invention provides a protein assisted design and analysis method based on a cloud computing platform. Based on the above system, the method includes the following steps:
[0033] S1: Create a project or select an existing project through the user interface module;
[0034] S2: Create a task under the item in step S1;
[0035] S3: Select the required application from the application repository module and configure the parameters;
[0036] S4: Input the compliant design and / or analysis files through the user interface module;
[0037] S5: The resource scheduling module provides computing power scheduling and the task starts to execute;
[0038] S6: After the task execution is completed, a report is generated and the report is viewed through the user interface module;
[0039] In steps S4 to S6, the data storage and encryption module performs data encryption and storage.
[0040] Preferably, when creating a task or selecting an application, specify tags and / or priority attributes, and the system schedules according to the specified tags and / or priority attributes.
[0041] Preferably, during the data encryption and storage process, the user's data and task calculation results are both stored in encrypted form in the file storage service and decrypted when read.
[0042] The third aspect of the present invention provides an electronic device, including:
[0043] A processor; and
[0044] A memory, on which executable code is stored, and when the executable code is executed by the processor, it causes the processor to execute the protein assisted design and analysis method based on the cloud computing platform.
[0045] The fourth aspect of the present invention provides a computer-readable storage medium, on which executable code is stored, and when the executable code is executed by the processor of an electronic device, it causes the processor to execute the protein assisted design and analysis method based on the cloud computing platform.
[0046] The fifth aspect of the present invention provides an application of a protein assisted design and analysis system based on a cloud computing platform, or a protein assisted design and analysis method based on a cloud computing platform, or an electronic device, or a computer-readable storage medium, in in vitro cell-free protein synthesis. The in vitro cell-free protein synthesis described in the present invention includes the in vitro cell-free protein synthesis (in vitro cell-free protein synthesis) described or mentioned in the following patent publication or announcement documents: CN106978349A, CN108535489A, CN108690139A, CN108949801A, CN108642076A, CN109022478A, CN109423496A, CN109423497A, CN109837293A, CN109971783A, CN109988801A, CN110551700A, CN109971775A, CN110551745A, CN110551700A, CN111378706A, CN111378707A, CN111378708A, CN111718419A, CN111748569A, CN112342248A, CN112876536A, CN110819647A, CN110845622A, CN110938649A, CN110964736A, CN111118065A, CN113215005A, CN113403360A and their cited documents.
[0047] Compared with the prior art, the present invention has the following characteristics:
[0048] (1) The present invention provides a highly integrated online design and analysis system, which adopts a clear modular design, supports computing power expansion and custom combination applications, integrates functions such as viewing, editing, structure prediction, mutation prediction, and stability analysis of protein structures, reduces the requirements for R & D personnel to use computer software, enables R & D personnel to avoid integrating various design and analysis software locally, building design and analysis processes, and recording design and analysis results, and allows biological researchers without a computer background and only with bioinformatics knowledge to quickly achieve in-depth AI analysis and design of biological big data. At the same time, the online design and analysis method also solves the problem of software difference for multiple offline users.
[0049] (2) The online design and analysis system supports multiple algorithms, and the system groups the algorithms according to their functions. For example, in the protein structure prediction group, there are three algorithms: MSAFold (an algorithm for protein structure prediction using deep learning, which makes predictions based on MSA features), Fast-MSAFold (an algorithm for protein structure prediction using deep learning, which improves the prediction speed by optimizing the MSA processing method), and PLMFold (an algorithm for protein structure prediction using deep learning, which does not rely on MSA and uses a pre-trained protein large language model to achieve fast structure prediction). The accuracies of the algorithms are slightly different, but the analysis speeds of the algorithms increase significantly in sequence. Users can flexibly select according to their needs.
[0050] (3) The problem of the leakage of bioinformatics protein sequence data is a very serious problem. The present invention encrypts the processed sequence data throughout the process to ensure data security during online processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a main component diagram of the protein assisted design and analysis system based on the cloud computing platform in the present invention.
[0052] Figure 2 It is an operation flowchart for protein analysis using the protein assisted design and analysis system based on the cloud computing platform in the present invention.
[0053] Figure 3 It is a schematic diagram of the data storage and encryption process of the protein assisted design and analysis system based on the cloud computing platform in the present invention.
[0054] Figure 4 It is a schematic diagram of the resource scheduling process of the protein assisted design and analysis system based on the cloud computing platform in the present invention.
[0055] Figure 5 It is a schematic diagram of the "structure prediction" configuration interface applied in the embodiment.
[0056] Figure 6 It is a preview diagram of the report after prediction using the "structure prediction" in the embodiment.
[0057] Figure 7 It is a preview diagram of the report after applying the "structure alignment" in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] In the present invention, a protein assisted design and analysis system based on a cloud computing platform includes a user interface module, an application repository module, a data storage and encryption module, and a resource scheduling module. Among them, the user interface module is the front end, implemented based on the web for interaction; the remaining modules are the back end, used for task analysis, data storage, and resource scheduling. The interaction between the modules in the system is as Figure 1 shown. Protein sequences or structures, analysis parameters, etc. are submitted to the back end through the user interface and stored in the data storage and encryption module in the form of tasks; then, the back end application computing service performs resource scheduling on the tasks. After the resource scheduling is successful and the task calculation is completed, a result report is generated and stored in the data storage and encryption module; finally, the user views the result report through the user interface. The present invention preferably deploys the system on a private cloud using container technologies such as docker and k8s. The private cloud consists of several data storage servers, can be horizontally scaled, and has a disaster recovery function.
[0059] Among them, the user interface module: includes pages such as user management, project and task creation, uploading sequences, configuring application parameters, monitoring progress, and result preview. The result preview includes the display of the three-dimensional structure of proteins and the display of charts, animations, etc., to help users better understand and use the analysis results.
[0060] Specifically, the user interface included in the present invention is implemented based on web pages using technologies such as HTML, JavaScript, and Vue.js. It mainly includes pages such as user management, project management, task execution, and report display. Compared with traditional terminal design / analysis software, using a web page as the interaction interface has the following advantages: users can access it on any device such as a computer or mobile phone without installing terminal software; the interface can be uniformly updated and maintained in the background; it can be shared and collaborated more easily. The user management page is used to set user basic information and modify passwords. The present invention is a multi-user system, and user data is mutually isolated, including projects created by users, uploaded sequence files, etc. The project management page includes creation, viewing, modification, deletion, etc. Each project contains at least one protein design or analysis task. The task execution page is used to select applications, configure application parameters, input design / analysis files, etc. The report display page is used to visually display the analysis results. As Figure 2 shown, the protein analysis operation using the interaction interface of the present invention preferably includes the following steps:
[0061] S1: Create a project. The project attributes include project members, project name, and project description, and the project is shared by all members.
[0062] S2: Create a task under the project. The task attributes include task name, task number, task status, creation time, start execution time, and completion time; multiple tasks can be included under one project.
[0063] S3: Select an application. The application repository contains software for common protein analysis and design. Select the required application software from it and configure specific parameters.
[0064] S4: Upload or select a protein sequence file from the file list. After the user selects a file, the system checks the file format and content. If it does not meet the requirements, the user will be prompted to upload it again.
[0065] S5: Wait for the task to execute. After the task is successfully created, it is scheduled and executed by the background. The user can monitor the task execution progress through the interface.
[0066] S6: View the task result report. After the result report is generated, it is archived in a fixed format, including text, charts, protein structure information, etc. The user can preview the result on the web page or download it to view locally.
[0067] Application repository module: The system constructs an application repository composed of protein design and analysis software. These applications are integrated into the platform in the form of microservices, supporting horizontal expansion of computing power. Users can also freely combine applications and configure applications according to their needs.
[0068] Specifically, the application repository includes applications such as protein basic parameter calculation, structure prediction, stability analysis, function analysis, interaction analysis, sequence design and optimization, and domain query, covering protein design and analysis links such as target setting, sequence analysis, structure prediction, and structure optimization.
[0069] Protein basic parameters are a set of parameters used to describe the structure and properties of proteins. They are used to represent the physical and chemical characteristics of proteins, including sequence length, amino acid proportion statistics, molecular formula, molecular weight, isoelectric point, extinction coefficient, stability index, hydrophobicity index, aromaticity index, protein flexibility, etc. Except for basic parameters such as sequence length, amino acid proportion statistics, molecular formula, and molecular weight which are directly calculated, other parameters can be realized based on past biological research papers. For example: The stability index is calculated by weighting the properties of each amino acid in the protein sequence based on the research method of Guruprasad et al.; the hydrophobicity index is based on the Kyte model, and the hydrophobicity of the protein is estimated by counting the hydrophobicity indices of different amino acid residues; the aromaticity index is based on the research of Lobry, and is calculated by predicting the position and number of aromatic groups; the isoelectric point is estimated according to the acidity and alkalinity of the protein; the extinction coefficient determination is based on the Beer-Lambert law, measuring the absorbance of the protein at 280 nm; protein flexibility is generally determined by X-ray crystallography and is also usually estimated using some bioinformatics calculation methods.
[0070] Structure prediction refers to inferring the three-dimensional conformation of proteins through computational methods or experimental techniques. The present invention integrates multiple structure prediction applications, using different methods for prediction respectively. There is a deep neural network method based on the sequence features of MSA, and there is also a deep neural network method based on the pre-trained protein large language model PLM. The difference between the two methods is that the former requires MSA alignment, while the latter does not. Comparatively speaking, since the latter does not require MSA alignment, the prediction speed is greatly accelerated. Compared with traditional structure prediction methods such as Rosetta and Phyre2, the two prediction methods based on deep learning technology in the present invention have greatly improved the prediction accuracy.
[0071] Stability analysis is to study the ability of proteins to maintain their structure and function under different conditions. The present invention mainly predicts the change in protein stability before and after mutation by calculating the change in protein free energy ddG.
[0072] Function analysis is to study the function and role of proteins in cells. The present invention infers its function by querying sequences and domains with significant similarity in the sequence, and tools such as BLAST are used.
[0073] Interaction analysis refers to studying the interaction between proteins, that is, the process of binding between proteins. The present invention adopts a computational method, uses a deep learning algorithm to learn from a known protein interaction database, and realizes protein interaction prediction based on protein sequence and / or structure information. At the same time, for proteins that have interactions or are predicted to have interactions, as needed, the sites and spatial conformations during the interaction between proteins are further predicted.
[0074] Sequence design and optimization is to search the protein sequence space using a search algorithm according to the user-set protein target structure and / or function requirements, and calculate and evaluate the searched sequences according to the target structure and / or function requirements, so as to conduct large-scale high-throughput screening of amino acid sequences that may meet the requirements and provide reasonable candidate sequences for experimental verification.
[0075] Domain query refers to querying the domains of a protein according to the protein name or PDB serial number. The present invention queries multiple protein domain databases such as Pfam and NCBI, and then screens and summarizes the query results. Domain query is the basis for de novo design of protein structures. By editing and combining the queried domain sequences, and then combining protein scaffold design, etc., the calculated protein structure and sequence can be obtained quickly.
[0076] Data storage and encryption module: The system provides a distributed data storage system for storing information such as protein sequences and structures. In addition, the system performs data stream encryption processing on the entire process from the protein sequence submitted by the user to the query calculation result.
[0077] Specifically, the use of the system of the present invention involves a large amount of core user data such as protein sequences and structures. Therefore, the present invention provides a set of data storage and encryption mechanisms to protect and store this data. Data protection includes data transmission encryption protection and data storage encryption protection. Data transmission encryption mainly relies on HTTPS to achieve; for data storage encryption, an independent object storage service needs to be selected first, and then an encryption algorithm that can be shared among each module is agreed upon, whereby the data can be uniformly managed and encrypted. The present invention selects MinIO as the object storage service and AES-256 as the encryption algorithm.
[0078] For the data encryption and storage mechanism in the task analysis process, J1-J4 are described in detail. Please refer to Figure 3 , and specifically includes the following steps:
[0079] J1: When the user submits a task through the browser, the input data such as protein sequences is encrypted first and then stored in the storage server.
[0080] J2: Then, it is downloaded from the storage server, decrypted after input, and then the task is calculated.
[0081] J3: After the calculation is completed, the task result is encrypted and then uploaded to the storage server.
[0082] J4: When the user views the result report, the system first downloads the data from the storage server, decrypts it, and then provides it to the user.
[0083] Resource scheduling module: The system provides a set of high-performance and horizontally scalable resource scheduling frameworks, which are deployed on multiple servers to build a high-performance computing cluster. It meets the requirements of multi-task parallel processing and AI large model calculations.
[0084] Specifically, the present invention provides a computing power source scheduling strategy to provide computing power scheduling for the applications running in the system. All applications in the application warehouse of the present invention are deployed in the form of containers. Therefore, the scheduling strategy is also implemented through container orchestration strategies, including two methods: setting tags and application priorities. As Figure 4 shown, when the user creates an analysis task, tags and / or priorities can be specified.
[0085] The label strategy is to set labels for the node server and containers respectively. When a container is executed, it is matched with nodes according to the labels. If the match is successful, the container is scheduled to that node; otherwise, it tries to match the next node until the match is successful or the scheduling fails. In the present invention, two types of labels are set according to the application's demand for a graphics card: GPU and GPU-Hight. The label GPU indicates that the application execution requires a graphics card, and the label GPU-Hight indicates that the application execution requires a graphics card with very high performance (such as A100). Similarly, the present invention also sets similar labels according to the application's demands for CPU and memory respectively.
[0086] The priority strategy is to set priority labels for application containers. When multiple applications are queuing for scheduling at the same time, the container with a higher priority setting is scheduled first. The priority strategy can ensure that important applications or instant applications can obtain better resource allocation and scheduling. For example, by setting a higher priority for interactive instant applications, the instantaneity of the interaction can be guaranteed, otherwise it will be blocked by applications with long execution times; in addition, when a user creates a task, the priority can be specified. For example, an urgent task is specified with a high priority, and a normal task is specified with a low priority, so as to perform task scheduling reasonably.
[0087] Term Introduction
[0088] As used in the present invention, "MSA" refers to multiple sequence alignment, which is a sequence alignment of three or more biological sequences, such as protein sequences, DNA sequences, or RNA sequences. The homology of the sequences can be deduced from the results of MSA, and the phylogenetic relationship can also lead to the common evolutionary ancestor of these sequences.
[0089] As used in the present invention, "MSAFold" is an algorithm for protein structure prediction using deep learning, which makes predictions based on MSA features.
[0090] As used in the present invention, "Fast-MSAFold" is an algorithm for protein structure prediction using deep learning, which improves the prediction speed by improving the MSA processing method.
[0091] As used in the present invention, "PLM" refers to a pre-trained protein large language model, which is a language model pre-trained in a self-supervised manner on a large-scale protein sequence database.
[0092] As used in the present invention, "PLMFold" is an algorithm for protein structure prediction using deep learning, which does not rely on MSA and uses a pre-trained protein large language model to achieve fast structure prediction.
[0093] As used in the present invention, "protein free energy change ddG" refers to the maximum available energy released or absorbed by a biological reaction from the initial state to the final state.
[0094] As used in the present invention, "PDB" refers to the Protein Data Bank, which is a database containing three-dimensional structure data of biological macromolecules such as proteins and nucleic acids.
[0095] As used in the present invention, "Https" refers to the Hypertext Transfer Protocol Secure, which is an Http channel with security as the goal. Based on Http, it ensures the security of the transmission process through transmission encryption and identity authentication.
[0096] As used in the present invention, "MinIO" is an object storage service based on the Apache License v2.0 open source protocol and can be used as a cloud storage solution to store a large amount of pictures, videos, and documents.
[0097] As used in the present invention, "AES-256" is a symmetric key algorithm that uses a 256-bit key and encrypts and decrypts data in 128-bit data block groups.
[0098] As used in the present invention, "docker" is an open-source application container engine that allows developers to package their applications and dependencies in a unified manner into a portable container and then publish it to any server with the docker engine installed (including popular Linux machines and Windows machines), and can also achieve virtualization.
[0099] As used in the present invention, "k8s", whose full name is Kubernetes, is an open-source platform for managing containers. It allows users to more conveniently deploy, expand, and manage containerized applications, and realizes functions such as load balancing, service discovery, and automatic elastic scaling through automation.
[0100] As used in the present invention, the full name of "HTML" is Hypertext Markup Language, which is a markup language that includes a series of tags. Through these tags, the formats of documents on the network can be unified, and scattered Internet resources can be connected into a logical whole.
[0101] As used in the present invention, "JavaScript" is a lightweight, interpreted or just-in-time compiled programming language with function priority.
[0102] As used in the present invention, "Vue.js" is a JavaScript framework for building user interfaces. It is built based on standard HTML, CSS (i.e., Cascading Style Sheets, a computer language used to present the styles of HTML and other files), and JavaScript, and provides a declarative and component-based programming model to help developers efficiently develop user interfaces.
[0103] "Rosetta" described in the present invention is a protein structure prediction method based on energy functions. Based on the physical properties and statistical information of proteins, it predicts the three-dimensional structure of proteins through simulation and optimization.
[0104] "Phyre2" described in the present invention is an online tool that can predict and analyze protein structures, functions, and mutations. Phyre2 is an upgraded version of Phyre, mainly using the method of remote homology detection for 3D modeling, predicting ligand binding sites and the effects of amino acid mutations.
[0105] The full name of "BLAST" described in the present invention is Basic Local Alignment Search Tool, that is, "Search Tool Based on Local Alignment Algorithm". It can align the input nucleic acid or protein sequence with the known sequences in the database to obtain information such as sequence similarity, so as to judge the source or evolutionary relationship of the sequence.
[0106] The "Long Short-Term Memory (LSTM) model" described in the present invention is a recurrent neural network model, mainly to solve the problems of gradient disappearance and gradient explosion during the training process of long sequences, and can conveniently process time series data.
[0107] It should be understood that within the scope of the present invention, the above technical features of the present invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be repeated one by one here.
[0108] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are implemented on the premise of the technical solution of the present invention, and give detailed implementation methods and specific operation processes, but the protection scope of the present invention is not limited to the following embodiments.
[0109] Embodiment:
[0110] This embodiment uses structure prediction and comparison of prediction results to illustrate the process of the present invention for performing bioanalysis tasks.
[0111] First, prepare the test sequence and the experimental resolution structure file. Download the sequence file seq-1DMM.fasta with PDB ID 1DMM and the experimentally resolved structure file strc-1DMM.pdb from the PDB website https: / / www.rcsb.org to the local.
[0112] Then, the operation steps of using the system of the present invention are as follows:
[0113] (1) Predict the structure of protein 1DMM
[0114] S11: Log in to the system, create a project, and name it "Protein 1DMM Structure Prediction Test".
[0115] S12: Create a new task "Structure Prediction 1DMM" under the project.
[0116] S13: Select the "Structure Prediction" application and set the label "GPU".
[0117] S14: Click "Local File", select the prepared sequence file seq-1DMM.fasta for upload. As Figure 5 This is a schematic diagram of the configuration interface of the "Structure Prediction" application in this embodiment.
[0118] S15: Submit the task and wait for the task to complete.
[0119] S16: After the task is completed, download the result and save it locally as predi-1DMM.pdb. Figure 6 This is a preview diagram of the report after prediction using the "Structure Prediction" in this embodiment.
[0120] (2) Compare the predicted structures of protein 1DMM
[0121] S21: Select the project created in S11.
[0122] S22: Create a new task "Structure Alignment 1DMM".
[0123] S23: Select the "Structure Alignment" application.
[0124] S24: Click "Local File", select the files predi-1DMM.pdb and strc-1DMM.pdb for upload.
[0125] S25: Submit the task and wait for the task to complete.
[0126] S26: After the task ends, through the result preview, as Figure 7 It can be seen that the two pdb files basically overlap. This proves that the structure prediction is very accurate.
[0127] In summary, through the system and method of the present invention, biological researchers can quickly perform computational simulations using protein design and analysis software without having to be familiar with various computational software and use commands. The present invention provides a new fast and easy-to-use tool for biological protein designers.
[0128] The above description of the embodiments and examples is for the convenience of those of ordinary skill in the art to understand and use the invention. It is obvious that those skilled in the art can easily make various modifications to these embodiments and apply the general principles described herein to other embodiments without creative efforts. Therefore, the present invention is not limited to the above embodiments, and all improvements and modifications made by those skilled in the art without departing from the scope of the present invention according to the disclosure of the present invention should be within the protection scope of the present invention.
Claims
1. A protein assisted design and analysis system based on a cloud computing platform, characterized in that, the system is deployed on a cloud computing platform, and the system includes a front end and a back end; the front end includes: a user interface module, which is used to interact with users and display protein design and / or analysis results; the back end includes: an application repository module, which is used to provide software or algorithms for protein design and / or analysis; a data storage and encryption module, which is used to encrypt and store data during the protein design and / or analysis process; a resource scheduling module, which is used to provide computing power scheduling for the applications running on the system.
2. The protein assisted design and analysis system based on a cloud computing platform according to claim 1, characterized in that, the system is deployed on a private cloud using container technology, and the private cloud includes multiple servers.
3. The protein assisted design and analysis system based on a cloud computing platform according to claim 1, characterized in that, the user interface module uses a web page as an interaction interface, including: a user management page, which is used to set user basic information and modify passwords; a project management page, which is used to create, view, modify, and delete projects, and each project contains at least one protein analysis task; a task execution page, which is used to select an application, configure application parameters, input design and / or analysis files, and monitor the task execution progress; a report display page, which is used to visually display protein design and / or analysis results.
4. The protein assisted design and analysis system based on a cloud computing platform according to claim 1, characterized in that, the applications in the application repository module are integrated into the cloud computing platform in the form of microservices, and the application repository module includes at least one of the following applications: a protein basic parameter calculation application, which is used to calculate parameters describing the structure and / or properties of proteins; a structure prediction application, which is used to predict the three-dimensional conformation of proteins; a stability analysis application, which is used to analyze the ability of proteins to maintain their original structure and function under different conditions; a function analysis application, which is used to analyze the functions of proteins in cells; an interaction analysis application, which is used to analyze the interactions between proteins; a sequence design and optimization application, which is used to search for and generate potential amino acid sequences that meet specific structural and functional requirements; a domain query application, which is used to query the domains of proteins.
5. The protein assisted design and analysis system based on a cloud computing platform according to claim 4, characterized in that, The described protein basic parameters include one or more of the protein's sequence length, amino acid proportion statistics, molecular formula, molecular weight, isoelectric point, extinction coefficient, stability index, hydrophobicity index, aromaticity index, and protein flexibility; the described structure prediction application uses a deep neural network method based on MSA sequence features and / or a deep neural network method based on a pre-trained protein large language model (PLM); the described stability analysis application predicts the change in protein stability before and after mutation by calculating the protein free energy change ddG; the described function analysis application infers the function of the protein by querying sequences and domains with significant similarity in the sequence; the described interaction analysis application predicts the interaction between proteins by matching the interaction patterns mined from the protein sequence and / or structure database, and uses the protein sequence and / or structure information to predict the interaction sites and spatial conformations between proteins. The described sequence design and optimization application searches for and evaluates protein sequences according to the target structure and / or function requirements set by the user to generate possible amino acid sequences that meet the requirements; the described domain query application queries the domains of a protein according to the protein name and / or PDB ID.
6. A protein assisted design and analysis system based on a cloud computing platform according to claim 1, characterized in that in the data storage and encryption module, data encryption includes data transmission encryption and data storage encryption. The data transmission encryption is based on HTTPS technology, and the data storage encryption uniformly manages and encrypts the data through an object storage service and an encryption algorithm.
7. A protein assisted design and analysis system based on a cloud computing platform according to claim 6, characterized in that the object storage service is MinIO, and the encryption algorithm is AES-256.
8. A protein assisted design and analysis system based on a cloud computing platform according to claim 1, characterized in that in the resource scheduling module, the scheduling policies include a tagging policy and / or a priority policy.
9. A protein assisted design and analysis method based on a cloud computing platform, based on the system according to any one of claims 1 to 8, characterized in that the method includes the following steps: S1: Create a project or select an existing project through the user interface module; S2: Create a task under the project in step S1; S3: Select the required application from the application repository module and configure the parameters; S4: Input a design and / or analysis file that meets the requirements through the user interface module; S5: The resource scheduling module provides computing power scheduling, and the task starts to execute; S6: After the task execution is completed, a report is generated, and the report is viewed through the user interface module; In steps S4 to S6, the data storage and encryption module performs data encryption and storage.
10. A protein assisted design and analysis method based on a cloud computing platform according to claim 9, characterized in that When creating a task or selecting an application, specify the tag and / or priority attribute, and the system schedules according to the specified tag and / or priority attribute.
11. A protein assisted design and analysis method based on a cloud computing platform according to claim 9, characterized in that during data encryption and storage, both the user's data and the task calculation results are stored in encrypted form in the file storage service and decrypted when read.
12. An electronic device, characterized in that comprising: a processor; and a memory having executable code stored thereon, which when executed by the processor causes the processor to execute the protein assisted design and analysis method based on a cloud computing platform according to any one of claims 9-11.
13. A computer-readable storage medium having executable code stored thereon, which when executed by a processor of an electronic device causes the processor to execute the protein assisted design and analysis method based on a cloud computing platform according to any one of claims 9-11.
14. Use of a protein assisted design and analysis system based on a cloud computing platform according to any one of claims 1-8, or a protein assisted design and analysis method based on a cloud computing platform according to any one of claims 9-11, or an electronic device according to claim 12, or a computer-readable storage medium according to claim 13, in cell-free protein synthesis in vitro.
Citation Information
Patent Citations
Kit for in vitro synthesis of protein and preparation method
CN106978349A
Protein synthesis system for in-vitro protein synthesis, kit and preparation method for protein through in-vitro synthesis
CN108535489A
Synthesis system, preparation, kit and preparation method of in-vitro DNA-to-Protein (D2P)
CN108642076A
Novel fusion protein preparation method and application of novel fusion protein for increasing protein synthesis
CN108690139A
Method for regulating in-vitro biosynthetic activity by knocking out nuclease system
CN108949801A
Cited By
Macromolecule analysis data sharing management method and system
CN120636557A
Macromolecule analysis data sharing management method and system
CN120636557B