Protein processing method, device, computer device, and storage medium
By displaying protein sequences and mutation statements on a computer device, and simulating protein mutation and structure prediction, the problem of low protein processing efficiency is solved, and efficient protein mutation and structure prediction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2022-10-24
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies have low protein processing efficiency, cumbersome biochemical experimental procedures, and difficulty in efficiently performing site-directed mutagenesis and structural prediction.
By displaying protein sequences and mutation statements on computer devices, protein mutations can be simulated in the computer according to the mutation instructions, and the structure after mutation can be predicted, thus simplifying the biochemical experimental process.
It improves the efficiency of protein processing, enables efficient protein mutation and structure prediction in computer equipment, and simplifies the operation process.
Smart Images

Figure CN117253541B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a protein processing method, apparatus, computer device, and storage medium. Background Technology
[0002] Proteins are organic molecules composed of amino acids and are the material basis of life. Mutations in proteins can affect their function and drug resistance, therefore understanding and predicting the structure of mutant proteins is of great significance in biological research.
[0003] In related technologies, site-directed mutagenesis of proteins is performed through biochemical experiments to obtain mutated proteins, and then the protein structure is determined. However, this method requires experimental work and is relatively cumbersome, resulting in low protein processing efficiency. Summary of the Invention
[0004] This application provides a protein processing method, apparatus, computer equipment, and storage medium, which can improve the efficiency of protein processing. The technical solution is as follows:
[0005] On one hand, a protein processing method is provided, the method comprising:
[0006] The input first protein sequence and mutation statement are displayed. The mutation statement indicates that the first protein sequence should be mutated. The mutation statement includes at least one mutation instruction, which includes the mutation location and mutation method.
[0007] In response to the processing request for the first protein sequence and the mutation statement, for each mutation instruction in the mutation statement, the mutation position in the first protein sequence is mutated according to the mutation method in the mutation instruction to obtain the second protein sequence;
[0008] The second protein sequence is shown.
[0009] Optionally, in response to the processing request for the first protein sequence and the mutation statement, for each mutation instruction in the mutation statement, the mutation position in the first protein sequence is mutated according to the mutation method in the mutation instruction to obtain the second protein sequence, including:
[0010] The mutated statement is validated;
[0011] If the mutation statement passes the verification, for each mutation instruction in the mutation statement, the mutation position in the first protein sequence is mutated according to the mutation method in the mutation instruction to obtain the second protein sequence;
[0012] The method further includes:
[0013] If the mutation statement fails validation, a mutation failure message will be displayed.
[0014] Optionally, after displaying the aligned first and second protein structures, the method further includes:
[0015] In response to a target operation performed on either the first protein structure or the second protein structure, the same target operation is performed on the other protein structure, the target operation including at least one of rotation, translation, magnification, or reduction.
[0016] On the other hand, a protein processing apparatus is provided, the apparatus comprising:
[0017] The display module is used to display the input first protein sequence and mutation statement, wherein the mutation statement indicates that the first protein sequence is mutated, and the mutation statement includes at least one mutation instruction, wherein the mutation instruction includes the mutation location and mutation method;
[0018] A mutation module is configured to, in response to a processing request for the first protein sequence and the mutation statement, mutate the mutation position in the first protein sequence according to the mutation method in the mutation statement for each mutation instruction in the mutation statement, to obtain a second protein sequence;
[0019] The display module is used to display the second protein sequence.
[0020] Optionally, the mutation module is configured to perform any of the following:
[0021] The mutation method includes an original amino acid and a target amino acid, replacing the original amino acid at the mutation position in the first protein sequence with the target amino acid. The original amino acid refers to the amino acid before the mutation, and the target amino acid refers to the amino acid after the mutation.
[0022] The mutation method includes at least one of the original amino acids and a deletion marker, wherein at least one of the original amino acids at the mutation position in the first protein sequence is deleted, and the deletion marker indicates the deletion of the amino acid;
[0023] The mutation method includes at least one target amino acid and an insertion marker, wherein at least one target amino acid is inserted at the mutation position in the first protein sequence, and the insertion marker indicates the inserted amino acid;
[0024] The mutation instruction includes at least one of the original amino acids, at least one of the target amino acids, and a deletion / insertion marker, which deletes at least one of the original amino acids at the mutation position in the first protein sequence and inserts at least one of the target amino acids, wherein the deletion / insertion marker indicates the deletion and insertion of the amino acid.
[0025] Optionally, the device further includes:
[0026] The prediction module is used to predict the structure of the second protein corresponding to the second protein sequence in response to the structure prediction request.
[0027] The display module is also used to display the structure of the second protein.
[0028] Optionally, the prediction module includes:
[0029] The prediction unit is configured to, in response to the structure prediction request, predict the second protein structure corresponding to the second protein sequence and predict the first protein structure corresponding to the first protein sequence.
[0030] An alignment unit is used to align the first protein structure with the second protein structure so that the first protein structure and the second protein structure can overlap through a translation operation;
[0031] The display module is also used to display the aligned first protein structure and the second protein structure.
[0032] Optionally, the alignment unit is used for:
[0033] A spatial coordinate system is established based on the first protein structure and the second protein structure;
[0034] The pose of the first protein structure or the pose of the second protein structure is adjusted so that the difference between the first coordinate information and the second coordinate information is less than a first threshold. The first coordinate information is the coordinate information of a plurality of target atoms in the first protein structure in the spatial coordinate system, and the second coordinate information is the coordinate information of the plurality of target atoms in the second protein structure in the spatial coordinate system.
[0035] Optionally, the device further includes an atom determination module for:
[0036] The confidence levels of multiple amino acids in the second protein structure are determined, where the confidence level of the amino acid represents the accuracy of the predicted position of the amino acid.
[0037] The preset atoms in amino acids with a confidence level greater than the second threshold are identified as the target atoms.
[0038] Optionally, the display module is configured to perform at least one of the following:
[0039] Different display methods are used to show amino acids at mutated and non-mutated positions;
[0040] Different display methods are used to show the amino acids at the mutation sites in the first protein structure and the second protein structure;
[0041] The first protein structure and the second protein structure are superimposed and displayed, with the transparency of the first protein structure being higher than that of the second protein structure;
[0042] Based on the confidence levels of amino acids in the first protein structure and the second protein structure, colors are set for the amino acids in the first protein structure and the second protein structure, respectively. The confidence level of the amino acid indicates the accuracy of the predicted position of the amino acid.
[0043] Optionally, the device further includes:
[0044] A first processing module is configured to perform the same target operation on another protein structure in response to a target operation performed on either the first protein structure or the second protein structure, wherein the target operation includes at least one of a rotation operation, a translation operation, a magnification operation, or a reduction operation.
[0045] Optionally, the prediction module is used to:
[0046] In response to the structure prediction request, sequence alignment information and template protein structure corresponding to the second protein sequence are obtained. The sequence alignment information includes multiple homologous protein sequences corresponding to the second protein sequence and difference information between the second protein sequence and the multiple homologous protein sequences. The template protein structure is the protein structure corresponding to the multiple homologous protein sequences.
[0047] The structure prediction model is invoked to predict the structure of the second protein corresponding to the second protein sequence based on the sequence alignment information and the template protein structure.
[0048] Optionally, the mutation module includes:
[0049] A verification unit is used to verify the mutated statement;
[0050] A mutation unit is configured to, when the mutation statement passes verification, mutate the mutation position in the first protein sequence according to the mutation method in the mutation instruction for each mutation instruction in the mutation statement to obtain the second protein sequence.
[0051] The display module is also used to display mutation failure information if the mutation statement fails the verification.
[0052] Optionally, the display module includes:
[0053] The task creation unit is used to create a protein processing task and display the protein processing interface in response to a request to create a protein processing task.
[0054] A display unit is configured to, in response to an input operation, display the input first protein sequence and the mutation statement on the protein processing interface, and add the first protein sequence and the mutation statement to the protein processing task.
[0055] Optionally, the display module is further configured to display a task interface, which includes created protein processing tasks and the status of each created protein processing task. The status includes editing status, prediction status, prediction failure status, and prediction success status. The editing status refers to editing a protein sequence or mutation statement.
[0056] Optionally, the device further includes a second processing module for performing any of the following:
[0057] In response to a triggering operation of the editing option corresponding to the first protein processing task, the protein sequence and mutation statement in the first protein processing task are displayed on the protein processing interface, and the first protein processing task is in the editing state.
[0058] In response to the triggering operation of the termination option corresponding to the second protein processing task, the prediction of the protein structure corresponding to the protein sequence in the second protein processing task is stopped, and the second protein processing task is in the prediction state.
[0059] In response to a trigger operation on the log option corresponding to the third protein processing task, the processing log of the third protein processing task is displayed, the processing log includes the prediction failure reason, and the third protein processing task is in a prediction failure state.
[0060] In response to a trigger operation on the result option corresponding to the fourth protein processing task, the predicted protein structure in the fourth protein processing task is displayed, and the fourth protein processing task is in a prediction success state.
[0061] Optionally, the protein processing task further includes the second protein sequence, and the apparatus further includes:
[0062] A prediction module is used to add the protein processing task to a task pool in response to a structure prediction request. The task pool includes protein processing tasks that have not yet been executed.
[0063] The prediction module is further configured to predict the second protein structure corresponding to the second protein sequence in the protein processing task when the protein processing task meets the execution conditions.
[0064] The display module is used to display the structure of the second protein.
[0065] Optionally, the prediction module is used to:
[0066] In response to the structure prediction request, the task identifier corresponding to the protein processing task is added to the identifier queue of the task pool, the identifier queue including task identifiers corresponding to protein processing tasks that have not yet been executed;
[0067] The protein processing task is stored in the storage space of the target account in the task pool. The target account is the account that created the protein processing task. The task pool includes the storage space of each account that created the protein processing task.
[0068] Optionally, the prediction module is used to:
[0069] If the task identifier is the first task identifier in the identifier queue, the protein processing task indicated by the task identifier is retrieved from the storage space.
[0070] Predict the structure of the second protein corresponding to the second protein sequence in the protein processing task;
[0071] Remove the task identifier from the identifier queue.
[0072] Optionally, the protein processing task further includes the second protein sequence, and the apparatus further includes:
[0073] The prediction module is used to query the database for the number of task identifiers corresponding to the target account in response to the structure prediction request. The target account is the account that created the protein processing task. The database stores any account and the task identifiers of the corresponding executed protein processing tasks.
[0074] The prediction module is used to predict the second protein structure corresponding to the second protein sequence in the protein processing task when the number of task identifiers corresponding to the target account is less than a third threshold.
[0075] The display module is used to display the structure of the second protein.
[0076] Optionally, the protein processing task further includes the second protein sequence, and the apparatus further includes:
[0077] The prediction module is used to obtain input resource configuration information in response to a structure prediction request, the resource configuration information including configuration information of the device used to perform the protein processing task;
[0078] The prediction module is further configured to determine a computing device that matches the resource configuration information; send the protein processing task to the computing device, wherein the computing device is configured to predict the second protein structure corresponding to the second protein sequence in the protein processing task, and return the second protein processing structure;
[0079] The prediction module is also configured to receive the second protein processing structure returned by the computing device;
[0080] The display module is also used to display the structure of the second protein.
[0081] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to perform the operations performed by the protein processing method as described above.
[0082] On the other hand, a computer-readable storage medium is provided that stores at least one computer program, which is loaded and executed by a processor to perform the operations performed by the protein processing method as described above.
[0083] On the other hand, a computer program product is provided, including a computer program that is loaded and executed by a processor to perform the operations performed by the protein processing method as described above.
[0084] The solution provided in this application embodiment allows for the input of the protein sequence and corresponding mutation statement to mutate a protein. The computer device then mutates the first protein sequence based on the mutation location and mutation method included in the mutation statement, resulting in a second protein sequence. This simulates the protein mutation process within a computer device, obtaining the mutated protein without the need for biochemical experiments, thus simplifying the protein mutation process and improving the efficiency of protein mutation, further enhancing the efficiency of protein processing. Attached Figure Description
[0085] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0086] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application;
[0087] Figure 2 This is a flowchart of a protein processing method provided in an embodiment of this application;
[0088] Figure 3 This is a flowchart of another protein processing method provided in the embodiments of this application;
[0089] Figure 4 This is a schematic diagram of a protein processing interface provided in an embodiment of this application;
[0090] Figure 5 This is a schematic diagram of another protein processing interface provided in an embodiment of this application;
[0091] Figure 6 This is a flowchart of another protein processing method provided in the embodiments of this application;
[0092] Figure 7 This is a flowchart of another protein processing method provided in the embodiments of this application;
[0093] Figure 8 This is a flowchart of another protein processing method provided in the embodiments of this application;
[0094] Figure 9 This is a schematic diagram of a product introduction interface provided in an embodiment of this application;
[0095] Figure 10 This is a schematic diagram of another product introduction interface provided in an embodiment of this application;
[0096] Figure 11 This is a schematic diagram of another product introduction interface provided in an embodiment of this application;
[0097] Figure 12 This is a schematic diagram of another product introduction interface provided in an embodiment of this application;
[0098] Figure 13 This is a schematic diagram of a task interface provided in an embodiment of this application;
[0099] Figure 14 This is a schematic diagram of a protein processing interface provided in an embodiment of this application;
[0100] Figure 15 This is a schematic diagram of another protein processing interface provided in an embodiment of this application;
[0101] Figure 16 This is a schematic diagram of another protein processing interface provided in an embodiment of this application;
[0102] Figure 17 This is a schematic diagram of another protein processing interface provided in an embodiment of this application;
[0103] Figure 18 This is a schematic diagram of another protein processing interface provided in an embodiment of this application;
[0104] Figure 19 This is a schematic diagram of another task interface provided in an embodiment of this application;
[0105] Figure 20 This is a schematic diagram of another task interface provided in an embodiment of this application;
[0106] Figure 21 This is a flowchart of another protein processing method provided in the embodiments of this application;
[0107] Figure 22 This is a schematic diagram of a protein processing method provided in an embodiment of this application;
[0108] Figure 23 This is a schematic diagram of another protein processing method provided in the embodiments of this application;
[0109] Figure 24 This is a schematic diagram of another protein processing method provided in the embodiments of this application;
[0110] Figure 25 This is a schematic diagram of the structure of a protein processing device provided in an embodiment of this application;
[0111] Figure 26This is a schematic diagram of another protein processing device provided in an embodiment of this application;
[0112] Figure 27 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application;
[0113] Figure 28 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation
[0114] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0115] It is understood that the terms "first," "second," etc., used in this application may be used to describe various concepts herein, but unless otherwise stated, these concepts are not limited by these terms. These terms are used only to distinguish one concept from another. For example, without departing from the scope of this application, a first protein sequence may be referred to as a second protein sequence, and similarly, a second protein sequence may be referred to as a first protein sequence.
[0116] "At least one" refers to one or more amino acids. For example, at least one amino acid can be one, two, three, or any integer number of amino acids greater than or equal to one. "Multiple" refers to two or more amino acids. For example, multiple amino acids can be two, three, or any integer number of amino acids greater than or equal to two. "Each" refers to each of the at least one amino acids. For example, each amino acid refers to each of the multiple amino acids. If the multiple amino acids consist of three amino acids, then each amino acid refers to each of the three amino acids.
[0117] It is understood that the embodiments of this application involve data such as user information, protein sequences, and mutation statements. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0118] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0119] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and intelligent transportation.
[0120] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.
[0121] The protein processing method provided in the embodiments of this application will be described below based on artificial intelligence and machine learning technologies.
[0122] The drug molecule generation method provided in this application embodiment can be used in a computer device. Optionally, the computer device is a terminal or a server. Optionally, the terminal is a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. Optionally, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0123] In one possible implementation, the computer program involved in the embodiments of this application may be deployed and executed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network can form a blockchain system.
[0124] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application. See also... Figure 1 The implementation environment includes a terminal 101 and a server 102. The terminal 101 and server 102 are connected via a wireless or wired network. Optionally, the terminal 101 is used to mutate a first protein sequence based on a mutation statement to obtain a second protein sequence, and sends the second protein sequence to the server 102. The server 102 is used to predict the second protein structure corresponding to the second protein sequence and return the second protein structure to the terminal 101.
[0125] It should be noted that, Figure 1 The example described here is based solely on the server 102 predicting the second protein structure corresponding to the second protein sequence and sending it to the terminal 101. In another embodiment, the terminal 101 may also predict the second protein structure corresponding to the second protein sequence.
[0126] The protein processing method provided in this application can be applied to any scenario where protein mutation is required.
[0127] For example, this can be applied to studies on the impact of protein mutations on drug resistance. Drug resistance is closely related to protein mutations, therefore understanding and predicting the affinity and resistance of mutated proteins to drugs is crucial. Using the method provided in this application, a computer device mutates a protein sequence based on mutation statements to obtain the mutated protein sequence, and then predicts the protein structure corresponding to the mutated protein sequence. This protein structure is the structure of the protein represented by the mutated protein sequence. Based on the mutated protein structure, features related to drug affinity and resistance in the mutated protein can be calculated, such as the protein-drug complex structure, the physicochemical properties of protein amino acid residues and the drug, and protein-drug interaction characteristics. Based on these features related to drug affinity and resistance, the impact of protein sequence mutations on drug affinity and resistance can be predicted.
[0128] Figure 2 This is a flowchart illustrating a protein processing method provided in an embodiment of this application. This embodiment is executed by a computer device. See also... Figure 2 The method includes:
[0129] 201. The computer device displays the input first protein sequence and mutation statement, the mutation statement instructing the first protein sequence to be mutated, the mutation statement includes at least one mutation instruction, the mutation instruction includes the mutation location and mutation method.
[0130] Proteins are organic molecules composed of amino acids and are the material basis of life. Amino acids are the basic building blocks of proteins, which are composed of 20 amino acids combined in different proportions. A protein sequence refers to the sequence of multiple amino acids in a protein. In this application, mutation refers to an alteration in the amino acids within a protein, such as deletion, duplication, or insertion.
[0131] The first protein sequence and mutation statement are input by the user. The mutation statement includes at least one mutation instruction, which specifies the mutation method at the mutation location. A mutation instruction indicates that a mutation should be performed at a specific location in the protein sequence. The mutation method can include various methods, such as amino acid deletion, substitution, or insertion. The mutation location can be represented using a number corresponding to its position in the first protein sequence.
[0132] 202. In response to a processing request for a first protein sequence and a mutation statement, the computer device, for each mutation instruction in the mutation statement, mutates the mutation position in the first protein sequence according to the mutation method in the mutation instruction to obtain a second protein sequence.
[0133] A processing request for a first protein sequence and a mutation statement instructs the computer device to mutate the first protein sequence based on the mutation statement. In response to the processing request, the computer device mutates the first protein sequence based on each mutation instruction in the mutation statement.
[0134] Taking one mutation instruction as an example, the computer device determines the mutation location and mutation method included in the mutation instruction, and mutates the mutation location in the first protein sequence according to the mutation method. After the computer device mutates the first protein sequence based on each mutation instruction, it obtains a second protein sequence, which is the protein sequence obtained by mutating the first protein sequence.
[0135] 203. The computer device displays the second protein sequence.
[0136] After obtaining the second protein sequence, the computer device displays the second protein sequence so that the user can view the mutated protein sequence.
[0137] In related technologies, using biochemical experiments to mutate protein sequences suffers from high cost, long time, and low efficiency. The method provided in this application, if a protein is to be mutated, inputs the protein sequence and the corresponding mutation statement. A computer device then mutates the first protein sequence based on the mutation location and mutation method included in the mutation statement, obtaining a second protein sequence. This simulates the protein mutation process within a computer device, obtaining the mutated protein sequence without the need for biochemical experiments, simplifying the protein mutation process and thus improving the efficiency of protein mutation, further enhancing the efficiency of protein processing.
[0138] In the above Figure 2 Based on the illustrated embodiment, after obtaining the second protein sequence, the computer device can also predict the protein structure corresponding to the second protein sequence. The specific process is detailed below. Figure 3 The example shown. Figure 3 This is a flowchart of another protein processing method provided in this application embodiment. This application embodiment is executed by a computer device. See also... Figure 3 The method includes:
[0139] 301. A computer device displays an input first protein sequence and a mutation statement, the mutation statement instructing the first protein sequence to be mutated, the mutation statement including at least one mutation instruction, the mutation instruction including the mutation location and mutation method.
[0140] In one possible implementation, the computer device displays a protein processing interface including a first input area and a second input area. The first input area is used to input a protein sequence, and the second input area is used to input a mutation statement indicating the mutation location and method for mutating the protein sequence. If a user wants to mutate a protein, they input a first protein sequence in the first input area and a mutation statement in the second input area. The computer device displays the input first protein sequence and mutation statement in the first and second input areas, respectively.
[0141] The first protein sequence comprises multiple amino acids arranged in sequence.
[0142] 302. In response to the processing request for the first protein sequence and the mutation statement, the computer device mutates the mutation position in the first protein sequence according to the mutation method in the mutation statement for each mutation instruction in the mutation statement, thereby obtaining the second protein sequence.
[0143] In one possible implementation, the computer device mutates the mutation site in the first protein sequence according to the mutation method in the mutation instruction, including any of the following:
[0144] (1) The mutation method includes an original amino acid and a target amino acid. The original amino acid at the mutation position in the first protein sequence is replaced with the target amino acid. The original amino acid refers to the amino acid before the mutation, and the target amino acid refers to the amino acid after the mutation.
[0145] One type of mutation, which replaces the original amino acid at a single mutation position with the target amino acid, is called a single-point mutation. Optionally, the format of this mutation instruction is [aa][position][bb], where [aa] represents the identifier of the original amino acid, [position] represents the number of the mutation position, and [bb] represents the identifier of the target amino acid. This mutation instruction means that the original amino acid [aa] at [position] is replaced with the target amino acid [bb].
[0146] For example, if the amino acid at position 74 of the first protein sequence is alanine A, and the mutation position is 74, and you want to replace the amino acid at position 74 with valine V, the mutation instruction is "A74V", which means replacing alanine A at position 74 with valine V.
[0147] It should be noted that this explanation uses only a single mutation instruction as an example. If multiple mutation instructions in a mutation statement are single-point mutations, then the mutation statement constitutes a multi-point mutation, which is composed of multiple single-point mutations. For example, the mutation statement is "L11T+E56G". "L11T" and "E56G" are each a mutation instruction. This mutation statement indicates that the leucine L at position 11 is replaced with threonine T, and the glutamic acid E at position 56 is replaced with glycine G.
[0148] (2) The mutation method includes at least one original amino acid and a deletion marker, wherein at least one original amino acid at the mutation position in the first protein sequence is deleted, and the deletion marker indicates the deletion of the amino acid.
[0149] Optionally, a mutation that deletes the original amino acid at a single mutation position is called a single deletion mutation, which means that only one amino acid is deleted. Optionally, the mutation instruction corresponding to this single deletion mutation has the format [aa][position]del[aa], where [aa] represents the identifier of the original amino acid, [position] represents the mutation position number, and del is the deletion identifier. This mutation instruction indicates that the original amino acid [aa] at [position] will be deleted.
[0150] For example, if the amino acid at position 15 of the first protein sequence is lysine K, and the mutation position is 15, then if you want to delete lysine K at position 15, the mutation instruction is "K15delK", which means to delete lysine K at position 15.
[0151] Optionally, a mutation method that deletes original amino acids at multiple consecutive mutation positions is called a multiple deletion mutation, which means deleting multiple amino acids. Optionally, the mutation instruction corresponding to this multiple deletion mutation has the format [aa1][postition1]_[aa2][position2]del[aa_seq], where [postition1] and [position2] represent the mutation positions, [aa1] represents the identifier of the original amino acid at [postition1], [aa2] represents the identifier of the original amino acid at [position2], [aa_seq] represents the identifier of multiple original amino acids from [postition1] to [position2], and del is the deletion identifier. This mutation instruction indicates that multiple original amino acids [aa_seq] from [postition1] to [position2] will be deleted.
[0152] For example, the mutation statement is "R84_L86delRLL", which indicates the deletion of multiple amino acids RLL from arginine R at position 84 to leucine L at position 86. The amino acid RLL represents arginine R, leucine L, and leucine L.
[0153] (3) The mutation method includes at least one target amino acid and an insertion marker, wherein at least one target amino acid is inserted at the mutation position in the first protein sequence, and the insertion marker indicates the inserted amino acid.
[0154] Optionally, a mutation method that inserts at least one target amino acid at the mutation position is called an insertion mutation. The format of the mutation instruction corresponding to the insertion mutation is [aa1][position1]_[aa2][position2]ins[aa_seq], where [position1] and [position2] represent the mutation positions, [aa1] represents the identifier of the original amino acid at [position1], [aa2] represents the identifier of the original amino acid at [position2], [aa_seq] represents the identifier of the multiple target amino acids inserted between [position1] and [position2], and ins is the insertion identifier. This mutation instruction indicates the insertion of multiple target amino acids [aa_seq] between [position1] and [position2].
[0155] For example, the mutation statement "R84_L85insAA" indicates the insertion of multiple amino acids AA between arginine (R) at position 84 and leucine (L) at position 85. Amino acids AA represent alanine A and alanine A.
[0156] (4) The mutation instruction includes at least one original amino acid, at least one target amino acid and a deletion insertion marker, which deletes at least one original amino acid at the mutation position in the first protein sequence and inserts at least one target amino acid, the deletion insertion marker indicating the deletion and insertion of the amino acid.
[0157] Optionally, a mutation method that deletes one original amino acid at a mutation position and inserts multiple target amino acids is called a single deletion / insertion mutation. Optionally, the mutation instruction corresponding to this single deletion / insertion mutation has the format [aa][position]delins[aa_seq], where [aa] represents the identifier of the original amino acid, [aa_seq] represents the identifier of multiple target amino acids, [position] represents the mutation position number, and delins is the deletion / insertion identifier. This mutation instruction indicates that the original amino acid [aa] at [position] is deleted and multiple target amino acids [aa_seq] are inserted.
[0158] For example, the mutation statement is "V97delinsAWS", which means to delete valine V at the 97th position and insert multiple target amino acids AWS at that mutation position. The multiple target amino acids AWS represent alanine A, tryptophan W and serine S.
[0159] Optionally, a mutation method that deletes multiple original amino acids at a mutation position and inserts multiple target amino acids is called a multiple deletion insertion mutation. Optionally, the mutation instruction corresponding to this multiple deletion insertion mutation has the format [aa1][postion1]_[aa2][position2]delins[aa_seq], where [postition1] and [position2] represent the mutation positions, [aa1] represents the identifier of the original amino acid at [postition1], [aa2] represents the identifier of the original amino acid at [position2], [aa_seq] represents the identifier of the multiple target amino acids to be inserted between [postition1] and [position2], and delins is the deletion insertion identifier. This mutation instruction indicates that multiple original amino acids between [postition1] and [position2] will be deleted and multiple target amino acids aa_seq will be inserted.
[0160] For example, the mutation statement is "V97_Q99delinsAWS". This mutation statement means that multiple amino acids will be deleted from valine V at position 97 to glutamine Q at position 99, and multiple target amino acids AWS will be inserted. The multiple target amino acids AWS represent alanine A, tryptophan W and serine S.
[0161] The above examples illustrate a single mutation instruction. A mutation statement can contain multiple mutation instructions of different types, forming a multi-site mixed mutation. For instance, the mutation statement "L11T+E56G+R84_L86delRLL" indicates that leucine (L) at position 11 is replaced with threonine (T), glutamic acid (E) at position 56 is replaced with glycine (G), and multiple amino acids (RLL) from arginine (R) at position 84 to leucine (L) at position 86 are deleted. These multiple amino acids (RLL) represent arginine (R), leucine (L), and leucine (L).
[0162] In this embodiment, protein sequences are mutated by writing mutation statements, providing various types of mutations, such as single-point mutations, multi-point mutations, deletion mutations, single deletion / insertion mutations, multiple deletion / insertion mutations, and mixed mutations at multiple sites. Different types of mutations can be combined with each other, improving the ease and flexibility of mutating protein sequences.
[0163] In one possible implementation, the computer device verifies the mutation statement. If the mutation statement passes the verification, for each mutation instruction in the mutation statement, the mutation is performed at the mutation position in the first protein sequence according to the mutation method specified in the instruction, resulting in a second protein sequence. If the mutation statement fails the verification, a mutation failure message is displayed.
[0164] In other words, before performing mutation based on a mutation statement, it is necessary to first verify whether the mutation statement is correct. If it is correct, the verification passes, and mutation is performed based on the mutation statement. If the mutation statement is incorrect, the verification fails. If mutation is still performed based on the mutation statement, an error will occur, so mutation is no longer performed based on the mutation statement. Instead, a mutation failure message is displayed to indicate that there is an error in the mutation statement and mutation cannot be performed. Optionally, a mutation statement failing verification includes the following situations.
[0165] (1) If a mutation statement includes multiple mutation instructions, the verification will fail if different mutation instructions include the same mutation position. For example, the mutation statement is “L85G+R84_L86delRLL”, which includes the mutation instructions “L85G” and “R84_L86delRLL”. Both mutation instructions include the mutation position 85, so the verification will fail.
[0166] If different mutation instructions include the same mutation location, the same position in the protein sequence will be mutated multiple times. This may result in two contradictory mutation instructions, leading to mutation errors. Therefore, in this case, the verification fails to avoid errors in subsequent mutations, which helps to improve the accuracy of protein sequence mutation.
[0167] (2) The mutation instruction includes the original amino acid and the mutation position. If the amino acid at the mutation position in the first protein sequence is not the original amino acid, the verification will fail. For example, the amino acid at position 73 in the first protein sequence is leucine (L), but the mutation instruction is "A73F". The original amino acid in the mutation instruction is alanine (A), which means that the amino acid at position 73 in the first protein sequence is not alanine (A), so the verification will fail.
[0168] If the amino acid at the mutation site in the first protein sequence is not the original amino acid in the mutation instruction, it means that the written mutation instruction does not match the actual situation of the first protein sequence. The mutation instruction contains an error, so the verification fails in this case. This helps to avoid errors in subsequent mutations and improves the accuracy of protein sequence mutation.
[0169] (3) If the mutation instruction does not conform to the regular expression for mutation, the validation will fail. The regular expression for mutation is the rule for writing the mutation instruction. If the mutation instruction does not conform to the rule, the computer device may not be able to recognize the mutation instruction, which will result in the inability to perform mutation according to the mutation instruction. Therefore, the validation will fail in this case to ensure that the mutation can be performed successfully in the future.
[0170] (4) If the amino acid identifier included in the mutation instruction is not a pre-set amino acid identifier, the verification fails. For example, the pre-set amino acid identifier is a letter in the natural amino acid alphabet, which is ['V', 'I', 'L', 'E', 'Q', 'D', 'N', 'H', 'W', 'F', 'Y', 'R', 'K', 'S', 'T', 'M', 'A', 'G', 'P', 'C']. In this embodiment, an amino acid identifier corresponding to each amino acid is pre-set. If the amino acid identifier in the mutation instruction is not a pre-set amino acid identifier, the computer device cannot determine which amino acid the amino acid identifier indicates, thus preventing mutation from being performed according to the mutation instruction. Therefore, the verification fails in this case to ensure successful mutation in the future.
[0171] (5) If the original amino acid in the mutation instruction is the same as the target amino acid, the verification will fail. The original amino acid refers to the amino acid before the mutation, and the target amino acid refers to the amino acid after the mutation. For example, if the mutation instruction is "L85L", and both the original and target amino acids are leucine (L), the verification will fail. Another example is if the mutation instruction is "V97delinsV", and both the original and target amino acids are valine (V), the verification will fail.
[0172] If the amino acids before and after the mutation are the same, it is illogical and the mutation instruction may be incorrect. Therefore, the verification will fail in this case to avoid errors in subsequent mutations and improve the accuracy of protein sequence mutation.
[0173] (6) The mutation instruction includes two position numbers, representing the mutation position from the first position number to the second position number. If the first position number in the mutation instruction is greater than the second position number, the verification will fail. For example, if the mutation instruction is "R84_S83delRS", the first position number 84 is greater than the second position number 83, so the verification will fail. Another example is the mutation instruction "V97_L96delinsWWW", where the first position number 97 is greater than the second position number 96, so the verification will fail.
[0174] Since the two position numbers in the mutation instruction represent the mutation position from the first position number to the second position number, it is illogical if the first position number in the mutation instruction is greater than the second position number. Therefore, the mutation instruction may be erroneous, and the verification will fail in this case. This is to avoid errors in subsequent mutations and improve the accuracy of protein sequence mutation.
[0175] (7) For mutation instructions that indicate the deletion of amino acids, if the number of original amino acids does not match the number of mutation positions, the verification fails, and the original amino acids are the amino acids to be deleted. For example, the mutation instruction is “R84_L86delRL”, which indicates the deletion of amino acids from position 84 to position 86. The mutation positions include positions 84, 85 and 86. The original amino acids include arginine (R) and leucine (L). There are 3 mutation positions and 2 original amino acids, so the verification fails.
[0176] If the number of original amino acids does not match the number of mutation sites, it indicates a contradiction in the mutation instruction, suggesting an error. Therefore, the verification fails in this case to avoid errors during subsequent mutations and improve the accuracy of protein sequence mutations.
[0177] (8) For mutation instructions indicating the deletion of amino acids, the mutation instruction includes the start mutation position and the end mutation position, as well as the amino acids at the start mutation position and the end mutation position, and the amino acid sequence to be deleted. If the first amino acid in the amino acid sequence is different from the amino acid at the start mutation position, or the last amino acid in the amino acid sequence is different from the amino acid at the end mutation position, the verification fails. For example, the mutation instruction is "R84_L86delRLP", the amino acid at the start mutation position 84 is arginine R, the amino acid at the end mutation position 86 is leucine L, the amino acid sequence is arginine R, leucine L and proline P. The last amino acid in the amino acid sequence is different from the amino acid at the end mutation position, therefore the verification fails.
[0178] The amino acid sequence consists of multiple amino acids arranged sequentially from the start mutation position to the end mutation position. If the first amino acid in the sequence is different from the amino acid at the start mutation position, or the last amino acid in the sequence is different from the amino acid at the end mutation position, it indicates a contradiction in the mutation instruction. Therefore, the mutation instruction may be erroneous, and the verification will fail in this case. This helps to avoid errors during subsequent mutations and improves the accuracy of protein sequence mutation.
[0179] (9) For mutation instructions that indicate the insertion of an amino acid, the mutation instruction includes two position numbers, indicating that an amino acid is inserted between the mutation positions of the first and second position numbers. If the two position numbers in the mutation instruction are not consecutive, the verification will fail. For example, the mutation instruction is “R84_L86insAA”, the first position number is 84 and the second position number is 86. Since 84 and 86 are not consecutive, the verification will fail.
[0180] If the two position numbers in the mutation instruction are not consecutive, the computer device cannot determine which position between the positions indicated by the two position numbers to insert the amino acid, which will prevent the mutation from being performed according to the mutation instruction. Therefore, the verification will fail in this case to ensure that the subsequent mutation can be performed successfully.
[0181] In one possible implementation, based on the above-mentioned possible implementations, mutation of the protein sequence includes the following.
[0182] (1) Obtain the array corresponding to the protein sequence: The first protein sequence is represented by an array. The computer device obtains the array input by the user. This array includes multiple amino acids, which are the multiple amino acids that make up the first protein sequence. For example, the first protein sequence is represented as an amino acid type-number array s = ((A_1, 1),…(A_n,n)), where A_n is one of the 20 natural amino acids, n represents the position number, and n is a positive integer.
[0183] (2) Verify mutation statement: If the mutation statement passes the verification, mutation is allowed based on the mutation statement. If the mutation statement fails the verification, mutation is not allowed based on the mutation statement, and mutation failure information is displayed.
[0184] (3) Compile mutation statement: Compiling mutation statement means mutating the first protein sequence based on the mutation statement. According to different types of mutation methods, the mutation statement is compiled into 4 types. The 4 types include substitution mutation, insertion mutation, deletion mutation and deletion-insertion mutation. Each type includes the mutation position, the original amino acid before mutation, the target amino acid after mutation and the mutation function.
[0185] (4) Generate the second protein sequence: Based on the mutation position in the class obtained by compilation, the original amino acid before mutation, the target amino acid after mutation, and the mutation function, modify the first protein sequence to obtain the second protein sequence.
[0186] (5) Storing the second protein sequence: After obtaining the second protein sequence, the second protein sequence is stored in a preset storage space.
[0187] 303. The computer device displays the second protein sequence.
[0188] In one possible implementation, the computer device displays a protein processing interface in which the second protein sequence is displayed.
[0189] 304. In response to a structure prediction request, the computer device predicts the second protein structure corresponding to the second protein sequence and predicts the first protein structure corresponding to the first protein sequence.
[0190] The second protein sequence is a sequence composed of multiple amino acids. Considering that it is difficult to determine the characteristics of the mutated protein based solely on this second protein sequence, it is also necessary to predict the second protein structure corresponding to the second protein sequence, that is, the structure of the protein represented by the second protein sequence. In addition, in order to facilitate the comparison between the protein before and after the mutation and obtain more information, the embodiments of this application will also predict the first protein structure corresponding to the first protein sequence.
[0191] In one possible implementation, taking the prediction of the second protein structure corresponding to the second protein sequence as an example, the prediction process includes: a computer device responding to a structure prediction request, acquiring sequence alignment information corresponding to the second protein sequence and a template protein structure. The sequence alignment information includes multiple homologous protein sequences corresponding to the second protein sequence and difference information between the second protein sequence and the multiple homologous protein sequences. The template protein structure is the protein structure corresponding to the multiple homologous protein sequences. The computer device invokes a structure prediction model and, based on the sequence alignment information and the template protein structure, predicts the second protein structure corresponding to the second protein sequence.
[0192] Here, the homologous protein sequence corresponding to the second protein sequence refers to a protein sequence whose similarity to the second protein sequence is higher than the fourth threshold. The structure prediction model is used to predict protein structure. This structure prediction model can be a model trained using artificial intelligence and machine learning techniques. For example, the structure prediction model can be a model trained based on AlphaFold (a protein structure prediction algorithm).
[0193] In this embodiment, considering the high similarity between the second protein sequence and homologous protein sequences, the protein structure corresponding to the second protein sequence also has a high similarity to the protein sequence corresponding to the homologous protein sequence. For example, if a part of the second protein sequence is the same as a part of the homologous protein sequence, then the protein structure corresponding to that part of the second protein sequence should also be the same as the protein structure corresponding to that part of the homologous protein sequence. Therefore, when predicting the protein structure corresponding to the second protein sequence, multiple protein structures corresponding to homologous protein sequences can be used as template protein structures. Using these template protein structures, the second protein structure corresponding to the second protein sequence can be predicted. Since the template protein structure contains common structures identical to those of the protein structure corresponding to the second protein sequence, using the template protein structure for prediction helps ensure the accuracy of the predicted protein structure. Furthermore, this method, by using the protein structures corresponding to homologous protein sequences for prediction, does not require in-depth analysis of the second protein sequence itself, thus the processing method is relatively simple and helps improve the efficiency of protein structure prediction.
[0194] The above explanation uses the prediction of the second protein structure as an example. The process of predicting the first protein structure corresponding to the first protein sequence is the same as the process of predicting the second protein structure corresponding to the second protein sequence, and will not be repeated here.
[0195] 305. The computer device aligns the first protein structure with the second protein structure so that the first protein structure and the second protein structure can overlap through a translation operation.
[0196] The first and second protein structures are three-dimensional. After obtaining the first and second protein structures, the computer aligns them. The aligned first and second protein structures can be superimposed through translation without rotation. If the first and second protein structures are projected onto the same plane after alignment, the projected first and second protein structures will have the same shape.
[0197] In one possible implementation, a computer device establishes a spatial coordinate system based on the first protein structure and the second protein structure, and adjusts at least one of the poses of the first protein structure or the second protein structure so that the difference between the first coordinate information and the second coordinate information is less than a first threshold. The first coordinate information is the coordinate information of multiple target atoms in the first protein structure in the spatial coordinate system, and the second coordinate information is the coordinate information of the multiple target atoms in the second protein structure in the spatial coordinate system.
[0198] The first protein structure includes multiple target atoms, and the second protein structure includes multiple target atoms. The positions of the target atoms in the first protein structure are the same as the positions of the target atoms in the second protein structure. Therefore, it can be considered that when the multiple target atoms in the first protein structure are aligned with the multiple target atoms in the second protein structure, the first protein structure and the second protein structure are also aligned. Thus, by adjusting at least one of the poses of the first protein structure or the second protein structure, the positions of the multiple target atoms in the first protein structure and the multiple target atoms in the second protein structure can be made closer and closer in the spatial coordinate system, thereby making the first protein structure and the second protein structure more and more aligned.
[0199] In this embodiment, considering that multiple target atoms in the first protein structure and the second protein structure are in the same position, the first protein structure and the second protein structure are aligned by means of multiple target atoms. By aligning multiple target atoms, the first protein structure and the second protein structure are aligned, which is simple to operate and improves the efficiency of aligning the first protein structure and the second protein structure.
[0200] Optionally, before adjusting the pose of the first protein structure or the pose of the second protein structure, the computer device needs to determine the plurality of target atoms. The process of the computer device determining the target atoms includes: the computer device determining the confidence level of a plurality of amino acids in the second protein structure, and determining a preset atom in the amino acid with a confidence level greater than a second threshold as the target atom.
[0201] The confidence level of an amino acid represents the accuracy of its predicted location. For example, this confidence level could be pLDDT (Predicted Local Distance Difference Test). A higher confidence level indicates a higher accuracy in the amino acid's location. Higher accuracy in the amino acid's location leads to higher accuracy in alignment based on the atoms within that amino acid. Therefore, the computer selects target atoms from amino acids with a confidence level greater than a second threshold. Since an amino acid contains multiple atoms, to improve processing efficiency, the computer only identifies preset atoms within that amino acid as target atoms. These preset atoms can be heavy atoms in the amino acid, including C (carbon), N (nitrogen), CA (C-alpha, alpha carbon), and O (oxygen).
[0202] In one possible implementation, based on the above-mentioned possible implementations, protein structure alignment includes the following.
[0203] (1) Load the first protein structure and the predicted candidate protein structures.
[0204] (2) Determine the confidence level of each candidate protein structure. The confidence level of a candidate protein structure is the average confidence level of multiple amino acids in that candidate protein structure. The candidate protein structure with a confidence level greater than the third threshold is determined as the second protein structure.
[0205] (3) In the first protein structure and the second protein structure, respectively, determine the amino acids with a confidence level greater than the second threshold, and denot them as set A_wt and set A_mt, respectively. Set A_wt and set A_mt are the amino acids used for alignment.
[0206] (4) Since the alignment of the protein structure as a whole needs to be considered, when aligning the selected sets A_wt and A_mt, only the heavy atoms in the amino acids, namely C, N, CA and O, are considered. The heavy atoms selected from sets A_wt and A_mt are respectively denoted as A_wt_backbone and A_mt_backbone.
[0207] (5) For the first protein structure and the second protein structure, use the select function to select the heavy atoms to be aligned, and use the align function to align the selected heavy atoms in the two protein structures.
[0208] (6) Store the aligned first and second protein structures.
[0209] 306. The computer device displays the aligned first and second protein structures.
[0210] After obtaining the aligned first and second protein structures, the computer device displays the aligned first and second protein structures, allowing users to compare the first protein structure before mutation and the second protein structure after mutation, thus improving the display effect of protein structures.
[0211] In this embodiment, the protein structure corresponding to the mutated protein sequence is predicted, and the three-dimensional structures of the protein before and after the mutation are generated and displayed. This is beneficial for studying the drug affinity and drug resistance of the protein based on the protein structure before and after the mutation, thereby providing clues for drug design and precision medicine.
[0212] In one possible implementation, the computer device displays a first protein sequence and a second protein sequence on a protein processing interface, and based on the protein processing interface, obtains a structure prediction request to predict the second protein structure corresponding to the second protein sequence and the first protein structure corresponding to the first protein sequence. The computer device then displays the aligned first and second protein structures on the protein processing interface.
[0213] In this embodiment, after the protein structure is predicted based on the protein processing interface, the predicted protein structure can be directly displayed on the protein processing interface. The prediction and display of the protein structure do not depend on third-party tools, thereby realizing end-to-end processing of the protein structure, simplifying the protein structure processing flow, and improving the processing efficiency of the protein structure.
[0214] In one possible implementation, a computer device displays the aligned first protein structure and the second protein structure, including at least one of the following.
[0215] (1) Computer equipment uses different display methods to display amino acids at mutation sites and amino acids at non-mutation sites.
[0216] The protein structure includes amino acids at mutated and non-mutated sites. Amino acids at mutated sites are more important because by examining them, the changes in the protein structure after mutation can be observed. Therefore, this application uses different display methods to display amino acids at mutated and non-mutated sites, thereby distinguishing them and making it easier for users to find them, further improving the display effect of the protein structure.
[0217] Optionally, the computer device highlights amino acids at mutated sites and displays amino acids at non-mutated sites in a non-highlighted manner, thereby emphasizing the amino acids at mutated sites. Optionally, the computer device uses a ball-and-stick model to display amino acids at mutated sites and a cartoon model to display amino acids at non-mutated sites. The ball-and-stick model can demonstrate the side chain structure of amino acids, highlighting the structural changes brought about by the mutation, and also making it easier for users to focus on the local structural changes caused by the mutation.
[0218] (2) The computer equipment uses different display methods to display the amino acids at the mutation sites in the first protein structure and the amino acids at the mutation sites in the second protein structure.
[0219] The first protein structure is the protein structure before the mutation, and the second protein structure is the protein structure after the mutation. The computer device displays the amino acids at the mutation sites in the first protein structure and the second protein structure using different display methods, so that users can easily compare the structures before and after the mutation.
[0220] For example, the computer device sets the amino acid at the mutation site in the first protein structure to a first color and sets the amino acid at the mutation site in the second protein structure to a second color, wherein the first color and the second color are different. Alternatively, the computer device sets the amino acid at the mutation site in the first protein structure to a semi-transparent state and sets the amino acid at the mutation site in the second protein structure to an opaque state.
[0221] (3) The computer device displays the first protein structure and the second protein structure superimposed, and the transparency of the first protein structure is higher than that of the second protein structure.
[0222] Because the first and second protein structures are aligned, when the computer displays them superimposed, the amino acids at non-mutated sites overlap, while those at mutated sites do not. This allows users to focus on the local structural changes caused by the mutation and to compare the structures before and after the mutation, identifying structural differences and thus improving the display effect of the protein structure. Furthermore, since the second protein structure is the mutated protein structure and is more important than the first, the computer sets the transparency of the first protein structure to be higher than that of the second. This makes it easier for users to view the second protein structure and avoids the problem of poor display quality caused by the first protein structure obscuring the second.
[0223] (4) The computer device sets the color of the amino acid in the first protein structure and the color of the amino acid in the second protein structure based on the confidence level of the amino acid in the first protein structure and the confidence level of the amino acid in the second protein structure, respectively. The confidence level of the amino acid indicates the accuracy of the predicted position of the amino acid.
[0224] The higher the confidence level of an amino acid, the more accurate its location. Optionally, the higher the confidence level of an amino acid, the brighter its dark color. Alternatively, the computer device can divide the confidence level of amino acids into several different intervals, with different colors for amino acids in different intervals. For example, amino acids with a confidence level of 80 to 100 are orange, those with a confidence level of 50 to 80 are green, and those with a confidence level of 0 to 50 are yellow.
[0225] In this embodiment, the color of an amino acid is set based on its confidence level, allowing users to intuitively understand the confidence level of an amino acid based on its color, thus increasing the amount of information displayed.
[0226] In one possible implementation, based on the above-mentioned possible implementations, the computer device displays the aligned first and second protein structures, including both superimposed display and independent display. Details are as follows.
[0227] Overlay display: Overlay display refers to displaying the first protein structure and the second protein structure in the same view. The first protein structure is set to a semi-transparent state, and the second protein structure is set to an opaque state, so as to intuitively perceive the structural differences between the two. Overlay display includes the following steps.
[0228] (1) Determine the first and second protein structures after alignment.
[0229] (2) Determine the amino acid at the mutation site and the amino acid at the adjacent site in the first protein structure, and determine the amino acid at the mutation site and the amino acid at the adjacent site in the second protein structure. Here, the adjacent site refers to the position adjacent to the mutation site.
[0230] (3) Render the first protein structure and the second protein structure in the same view and set their colors and styles. For the process of setting the colors and styles of the first protein structure and the second protein structure, see steps a-e below.
[0231] a. Set the first protein structure to a semi-transparent state.
[0232] b. The mutated site and adjacent amino acids are modeled as ball-and-stick amino acids, including the mutated site and adjacent amino acids in the first protein structure, and the mutated site and adjacent amino acids in the second protein structure. The ball-and-stick model can display the side chain structure of amino acids, highlighting the structural changes brought about by the mutation.
[0233] c. The amino acids at the mutation sites in the first protein structure and the second protein structure are rendered in different colors to enhance the contrast.
[0234] d. Set all amino acids except those at the mutation site and adjacent sites to a cartoon model. The cartoon model does not show the side chain structure of amino acids, but it will show the secondary structure of the protein.
[0235] e. Stain amino acids, excluding those at the mutation site and adjacent sites, according to their confidence levels. For example, divide the confidence levels into four categories: "greater than 90" (indicating very high confidence), "between 70 and 90" (indicating high confidence), "between 50 and 70" (indicating low confidence), and "less than 50" (indicating very low confidence), and render each category with a different color to enhance the contrast.
[0236] Figure 4 This is a schematic diagram of a protein processing interface provided in an embodiment of this application, as shown below. Figure 4 As shown, the computer device displays a protein processing interface. In the case of overlay display, the left side of the protein processing interface displays the first protein structure 401 and the second protein structure 402, and the right side of the protein processing interface displays a curve of the confidence level of the amino acids in the protein structure.
[0237] Independent display: Independent display refers to displaying the first protein structure and the second protein structure separately in different biological views. Independent display includes the following steps.
[0238] (1) Determine the first and second protein structures after alignment.
[0239] (2) Determine the amino acid at the mutation site and the amino acid at the adjacent site in the first protein structure, and determine the amino acid at the mutation site and the amino acid at the adjacent site in the second protein structure. Here, the adjacent site refers to the position adjacent to the mutation site.
[0240] (3) Render the first protein structure and the second protein structure in different views and set their colors and styles. For the process of setting the colors and styles of the first protein structure and the second protein structure, see steps a-c below.
[0241] a. The mutated site and its adjacent amino acids are modeled as a ball-and-stick model, including the mutated amino acid and its adjacent amino acid in the first protein structure, and the mutated amino acid and its adjacent amino acid in the second protein structure. The ball-and-stick model can display the side chain structure of amino acids, highlighting the structural changes brought about by the mutation.
[0242] b. The amino acids at the mutation sites in the first protein structure and the second protein structure are rendered in different colors to enhance the contrast.
[0243] c. Set all amino acids except those at the mutation site and adjacent sites to a cartoon model. The cartoon model does not show the side chain structure of amino acids, but it will show the secondary structure of the protein.
[0244] Figure 5This is a schematic diagram of another protein processing interface provided in an embodiment of this application, as shown below. Figure 5 As shown, the computer device displays a protein processing interface. When displayed independently, the left view of the protein processing interface shows a first protein structure 501, and the right view of the protein processing interface shows a second protein structure 502.
[0245] It should be noted that the embodiments in this application only illustrate the prediction of the first protein structure and the second protein structure. In another embodiment, the computer device, in response to a structure prediction request, predicts the second protein structure corresponding to the second protein sequence and displays the second protein structure. That is, only the second protein structure is predicted and displayed, without the need to predict and display the first protein structure.
[0246] 307. In response to a target operation performed on either the first protein structure or the second protein structure, the computer device performs the same target operation on another protein structure, the target operation including at least one of a rotation operation, a translation operation, a magnification operation, or a reduction operation.
[0247] In this embodiment, since the first protein structure and the second protein structure are already aligned, a view synchronization function is also provided for the first protein structure and the second protein structure. No matter what operation is performed on one of the protein structures, the same operation will be performed on the other protein structure at the same time, so that the first protein structure and the second protein structure always remain aligned, which makes it convenient for users to compare the structure of the first protein structure and the second protein structure.
[0248] In one possible implementation, the first protein structure and the second protein structure are displayed in different views. The computer device detects the interactive operations in each view. When a target operation on the protein structure is triggered in one view, the viewpoint information of that view is obtained and synchronized to the other view. Thus, the target operation on the protein structure is also triggered in the other view. Therefore, the user's target operation on one view can be synchronized to the other view in real time, which facilitates the user to compare the structures of the first and second protein structures.
[0249] In one possible implementation, the computer device provides a synchronization option for synchronously displaying a first protein structure and a second protein structure. When the synchronization option is selected, the computer device performs the same target operation on the other protein structure in response to a target operation performed on either the first or second protein structure. When the synchronization option is not selected, the computer device does not perform the same target operation on the other protein structure when it detects a target operation performed on either the first or second protein structure.
[0250] It should be noted that the embodiments in this application are only illustrated by simultaneously displaying the first protein structure and the second protein structure. In another embodiment, the first protein sequence and the second protein sequence may not be displayed simultaneously, that is, step 307 above may not be performed.
[0251] Figure 6 This is a flowchart of another protein processing method provided in the embodiments of this application, such as... Figure 6 As shown, the method includes the following steps.
[0252] 601. Enter the first protein sequence and mutation statement.
[0253] 602. Based on the mutation statement, mutate the first protein sequence to obtain the second protein sequence.
[0254] 603. Predicting Protein Structures. Computer equipment predicts the structure of a first protein sequence and the structure of a second protein sequence.
[0255] 604. Protein Structure Alignment. The computer equipment aligns the first and second protein structures.
[0256] 605. Show the alignment of the first and second protein structures.
[0257] The method provided in this application embodiment allows for the input of the protein sequence and corresponding mutation statement to mutate a protein. A computer device then mutates the first protein sequence based on the mutation location and mutation method included in the mutation statement, resulting in a second protein sequence. This simulates the protein mutation process within a computer device, obtaining the mutated protein without the need for biochemical experiments, thus simplifying the protein mutation process and improving the efficiency of protein mutation, further enhancing the efficiency of protein processing.
[0258] Based on the above embodiments, the computer device creates protein processing tasks to mutate protein sequences and predict protein structures, as detailed below. Figure 7 The example shown. Figure 7 This is a flowchart of another protein processing method provided in this application embodiment. This application embodiment is executed by a computer device. See also... Figure 7 The method includes:
[0259] 701. In response to a request to create a protein processing task, the computer device creates a protein processing task and displays the protein processing interface.
[0260] This creation request is used to request the creation of a protein processing task, and the protein processing interface is used to process proteins.
[0261] In one possible implementation, the computer device displays a task interface that includes previously created protein processing tasks and task creation options. When a user wants to create a protein processing task, a trigger operation is performed on the task creation option. Upon detecting the trigger operation, the computer device generates a request to create the protein processing task.
[0262] In some embodiments, the computer device displays a task interface that includes created protein processing tasks and the status of each created protein processing task, including editing status, prediction status, prediction failure status, and prediction success status.
[0263] Among them, the "editing" state means that the protein sequence or mutation statement is being edited, the "predicting" state means that the protein structure corresponding to the protein sequence is being predicted, the "prediction failure" state means that the protein structure corresponding to the protein sequence was not successfully predicted, and the "prediction success" state means that the protein structure corresponding to the protein sequence has been successfully predicted.
[0264] In one possible implementation, when the computer device displays a task interface, the method further includes any of the following.
[0265] (1) In response to the triggering operation of the editing option corresponding to the first protein processing task, the computer device displays the protein sequence and mutation statement in the first protein processing task on the protein processing interface, and the first protein processing task is in the editing state.
[0266] For the first protein processing task currently in editing mode, the task interface also includes editing options corresponding to that task. If the user wants to continue editing the first protein processing task, a trigger operation is executed on the editing option. In response to this trigger operation, the computer device jumps from the task interface to the protein processing interface, displaying the protein sequence and mutation statements from the first protein processing task, allowing the user to continue editing these elements.
[0267] (2) In response to the triggering operation of the termination option corresponding to the second protein processing task, the computer device stops predicting the protein structure of the protein sequence in the second protein processing task, and the second protein processing task is in the prediction state.
[0268] For a second protein processing task that is currently predicting protein structures, the task interface also includes a termination option for that task. If the user wants to terminate the existing prediction of the protein structure of the protein sequence in the second protein processing task, a trigger operation is performed on the termination option. In response to this trigger operation, the computer device stops predicting the protein structure of the protein sequence in the second protein processing task.
[0269] (3) In response to the triggering operation of the log option corresponding to the third protein processing task, the computer device displays the processing log of the third protein processing task, which includes the reason for prediction failure and the third protein processing task is in the prediction failure state.
[0270] For a third protein processing task that is in a prediction failure state, the task interface also includes a log option for that task. If the user wants to view the processing log for that third protein processing task, they can trigger an operation on this log option. In response to this trigger operation, the computer device displays the processing log for the third protein processing task. This log includes the reason for the prediction failure, allowing the user to view the reason for the failure in predicting the protein sequence structure and thus make adjustments and improvements based on the cause of the prediction failure.
[0271] (4) In response to the triggering operation of the result option corresponding to the fourth protein processing task, the computer device displays the protein structure predicted in the fourth protein processing task, and the fourth protein processing task is in the prediction success state.
[0272] For the fourth protein processing task that is in a successful prediction state, the task interface also includes a result option corresponding to the fourth protein processing task. If the user wants to view the predicted protein structure, a trigger operation is performed on the result option. In response to the trigger operation, the computer device displays the predicted protein structure in the fourth protein processing task, so that the user can trace the predicted protein structure.
[0273] In this embodiment, the task interface includes created protein processing tasks and the status of each created protein processing task. Different operations can be performed on protein processing tasks in different states, thereby realizing flexible management of protein processing tasks.
[0274] 702. In response to the input operation, the computer device displays the input first protein sequence and mutation statement on the protein processing interface, and adds the first protein sequence and the mutation statement to the protein processing task.
[0275] The user enters a first protein sequence and a mutation statement in the protein processing interface. In response to the input operation, the computer device displays the entered first protein sequence and mutation statement in the protein processing interface and adds the first protein sequence and the mutation statement to the protein processing task. The protein processing task includes tasks that perform processing based on the first protein sequence and the mutation statement.
[0276] 703. In response to a processing request for a first protein sequence and a mutation statement, the computer device mutates the mutation position in the first protein sequence according to the mutation method in the mutation statement for each mutation instruction in the mutation statement, thereby obtaining the second protein sequence.
[0277] The process of mutating the protein sequence in step 703 is the same as the process of mutating the protein sequence in step 302 above, and will not be described again here.
[0278] 704. The computer device displays the second protein sequence and adds the second protein sequence to the protein processing task.
[0279] After obtaining the second protein sequence, the computer device displays the second protein sequence on the protein processing interface and adds the second protein sequence to the protein processing task, which includes tasks that perform processing based on the second protein sequence.
[0280] 705. In response to a structure prediction request, the computer device predicts the second protein structure corresponding to the second protein sequence in the protein processing task and displays the second protein structure.
[0281] The second protein sequence is a sequence composed of multiple amino acids. Considering that it is difficult to determine the characteristics of the mutated protein based solely on the second protein sequence, the structure of the second protein corresponding to the second protein sequence is also predicted.
[0282] In one possible implementation, the computer device displays a second protein sequence on a protein processing interface, which also includes a structure prediction option. If the user wants to predict the protein structure corresponding to the second protein sequence, a trigger operation is executed for this structure prediction option. Upon detecting the trigger operation for the structure prediction option, the computer device generates a structure prediction request and, in response to the request, predicts the second protein structure corresponding to the second protein sequence in the protein processing task. After predicting the second protein structure, it is displayed on the protein processing interface.
[0283] In one possible implementation, the computer device, in response to a structure prediction request, adds the protein processing task to a task pool that includes protein processing tasks that have not yet been executed. If the protein processing task meets the execution conditions, the computer device predicts the second protein structure corresponding to the second protein sequence in the protein processing task and displays the second protein structure.
[0284] The task pool contains multiple unexecuted protein processing tasks. After a structure prediction request is submitted, the computer device first adds the current protein processing task to the task pool for waiting. Only when the execution conditions are met does it begin predicting the second protein structure corresponding to the second protein sequence in that protein processing task. This task pool can be viewed as a tool for batch computation of protein processing tasks. Batch protein processing tasks are added to the task pool, and the computer device intelligently allocates resources and manages task execution based on the execution conditions.
[0285] Optionally, if protein processing tasks in the task pool are executed in the order they were added to the task pool, the execution condition is that other protein processing tasks added earlier than the current protein processing task have completed their predictions. This "added time" refers to the time the task was added to the task pool. Alternatively, if each protein processing task has a corresponding execution time, which can be a user-defined time, the execution condition is that the current time reaches the execution time of the protein processing task.
[0286] Optionally, the task pool includes storage space corresponding to protein processing tasks and an identifier queue corresponding to task identifiers, such as... Figure 8 As shown, the process of predicting protein structure using computer equipment includes the following steps 7051-7055.
[0287] 7051. In response to the structure prediction request, the computer device adds the task identifier corresponding to the protein processing task to the identifier queue of the task pool, the identifier queue including task identifiers corresponding to protein processing tasks that have not yet been executed.
[0288] 7052. The computer device stores the protein processing task in the storage space of the target account in the task pool. The target account is the account that created the protein processing task. The task pool includes the storage space of each account that created the protein processing task.
[0289] The computer equipment provides each account with its own storage space. The protein processing task of the current target account is stored in the storage space of that target account. This isolates the protein processing tasks created by different accounts in terms of storage space, ensuring that the protein processing task of each account is not affected by other accounts. This guarantees both the sufficiency and availability of computing resources and the security of computing services.
[0290] 7053. When the task identifier is the first task identifier in the identifier queue, the computer device retrieves the protein processing task indicated by the task identifier from the storage space.
[0291] The identifier queue includes task identifiers corresponding to protein processing tasks that have not yet been executed. These task identifiers are ordered according to the time they were added to the queue, and the protein processing tasks in the task pool are also executed in this chronological order. Whenever a protein processing task completes its prediction, its corresponding task identifier is removed from the identifier queue. Therefore, when a task identifier is the first one in the identifier queue, it means that other protein processing tasks added earlier than the one indicated by that task identifier have completed their predictions, and it is time to execute that protein processing task. The computer then retrieves the protein processing task indicated by that task identifier from its storage space.
[0292] 7054. The computer device predicts the structure of the second protein corresponding to the second protein sequence in the protein processing task.
[0293] 7055. The computer device removes the task identifier from the identifier queue.
[0294] In one possible implementation, in response to a structure prediction request, the computer device queries a database for the number of task identifiers corresponding to a target account, which is the account that created the protein processing task. The database stores task identifiers for any account and corresponding executed protein processing tasks. If the number of task identifiers corresponding to the target account is less than a third threshold, the computer predicts the second protein structure corresponding to the second protein sequence in the protein processing task and displays the second protein structure.
[0295] In this embodiment of the application, the number of times the protein processing task created by each account can be executed is limited. The maximum number of executions for each account is a third threshold. If the number of task identifiers corresponding to the target account is less than the third threshold, it means that the number of times the target account requests to execute the protein processing task is less than the third threshold, that is, there is still a quota for executing the protein processing task. Then the computer device predicts the second protein structure corresponding to the second protein sequence in the protein processing task.
[0296] In another embodiment, if the number of task identifiers corresponding to the target account is not less than the third threshold, it means that the number of times the target account requests to execute the protein processing task is not less than the third threshold, that is, there is no quota for executing the protein processing task. Then the computer device displays a prompt message, which is used to indicate that the number of times the target account can execute has reached the limit.
[0297] In this embodiment of the application, only an account needs to be registered in the computer device to request the computer device to execute the protein processing task by creating a protein processing task, thereby realizing the prediction of protein structure. The operation process is simple and easy to manage the protein processing task.
[0298] In one possible implementation, in response to a structure prediction request, the computer device obtains input resource configuration information, including configuration information of the device used to perform the protein processing task. The computer device determines a computing device that matches the resource configuration information, sends the protein processing task to that computing device, which predicts the second protein structure corresponding to the second protein sequence in the protein processing task, and returns the second protein processed structure. The computer device receives the second protein processed structure returned by the computing device and displays the second protein structure.
[0299] In this embodiment, the computer device requests a computing device to execute a protein processing task, predicting the protein structure corresponding to the protein sequence in the protein processing task. The computer device determines the matching computing device based on resource configuration information input by the user, enabling the user to configure the computing device for executing the protein processing task. This computing device, which is also the computing resource for executing the protein processing task, ensures the easy availability and use of computing resources. Furthermore, it eliminates the need for manual maintenance of computing resources, greatly improving operational simplicity.
[0300] Optionally, the computing device is a cloud virtual machine (CVM). After obtaining the resource configuration information, the computer device creates a cloud server that matches the resource configuration information and sends the protein processing task to the cloud server. The cloud server then predicts the second protein structure corresponding to the second protein sequence in the protein processing task. The cloud server is a public cloud resource. This embodiment of the application utilizes public cloud resources to perform the protein processing task calculation, further improving the ease of operation.
[0301] Optionally, the resource configuration information includes at least one of computer type or model identifier, wherein the computer type is the model of the device used to perform the protein processing task, and the model identifier indicates the structure prediction model used to predict the protein structure, and different types of structure prediction models correspond to different model identifiers.
[0302] It should be noted that step 705 above is only used as an example to predict and display the second protein structure corresponding to the second protein sequence. In another embodiment, the computer device uses the same method as step 705 above to predict and display the first protein structure corresponding to the first protein sequence, which will not be described in detail here.
[0303] The method provided in this application embodiment allows for the input of the protein sequence and corresponding mutation statement to mutate a protein. A computer device then mutates the first protein sequence based on the mutation location and mutation method included in the mutation statement, resulting in a second protein sequence. This simulates the protein mutation process within a computer device, obtaining the mutated protein without the need for biochemical experiments, thus simplifying the protein mutation process and improving the efficiency of protein mutation, further enhancing the efficiency of protein processing.
[0304] In some embodiments, the protein processing method provided in this application is implemented based on a website. During the execution of this protein processing method, the interface displayed by the computer device is as follows: Figures 9-21 As shown.
[0305] Figure 9 This is a schematic diagram of a product introduction interface provided in an embodiment of this application, such as... Figure 9 As shown, the computer device displays a product introduction interface, which includes product information related to mutant protein structure prediction. This interface also includes a "Free Trial" option 901 and a "Contact Us" option 902. If a user wants to use the mutant protein structure prediction function, they trigger the "Free Trial" option 901. In response, the computer device displays a registration / login window 903 on the product introduction interface. Figure 10 As shown, the registration / login window 903 is used to enter an email address and password to log in. The registration / login window 903 also includes a "Register" option and a "Forgot Password" option. The "Register" option is used to register an account, and the "Forgot Password" option is used to reset the password. If the user has not yet registered an account, triggering the "Register" option will cause the computer device to respond by displaying the registration window 904 on the product introduction interface. Figure 11 As shown, registration window 904 is used to register an account. Registration window 904 includes input boxes for username, email, verification code, password, and confirm password. Registration window 904 also includes "Confirm" and "Cancel" options. After the user completes the information for registering the account, the "Confirm" option is triggered, and the computer device responds to this trigger by creating an account for the user. If the user forgets their password, the "Forgot Password" option is triggered, and the computer device responds by displaying the password setting window 905 on the product introduction interface. Figure 12 As shown, the password setting window 905 is used to reset the password. The password setting window 905 includes input boxes for email, verification code, new password, and confirm password. The registration window 904 also includes a "confirm" option and a "cancel" option. After the user has entered the information used to reset the password, the "confirm" option is triggered. The computer device responds to the trigger operation and determines the password reset by the user.
[0306] Figure 13 This is a schematic diagram of a task interface provided in an embodiment of this application, such as... Figure 13 As shown, after logging into the account, the computer device displays a task interface, which includes multiple protein processing tasks that have been created, along with their corresponding task status, processing time, and operation options. The task interface also includes a task creation option 1301. If the user wants to create a new protein processing task, they can trigger an operation on this task creation option 1301. In response to this trigger operation, the computer device will jump from the task interface to the protein processing interface.
[0307] Figure 14 This is a schematic diagram of a protein processing interface provided in an embodiment of this application, as shown below. Figure 14 As shown, the computer device displays a protein processing interface, which includes a first input region 1401 corresponding to the protein sequence and a second input region 1402 corresponding to the mutation statement. In addition, the protein processing interface includes "Upload," "View Example," "Compile," "Save as Draft," and "Structure Prediction" options. If the user wants to upload a protein sequence, a trigger operation is performed on the "Upload" option. In response to this trigger operation, the computer device displays a file upload window 1403 on the protein processing interface. Figure 15 As shown, the file upload window 1403 includes an upload area as well as "Upload" and "Cancel" options.
[0308] Figure 16 This is a schematic diagram of another protein processing interface provided in an embodiment of this application, as shown below. Figure 16 As shown, the computer device displays a protein processing interface. The first input area 1401 of the protein processing interface displays the protein sequence input by the user, and the second input area 1402 of the protein processing interface displays the mutation statement input by the user. If the mutation statement fails the verification, the protein processing interface also displays the mutation failure message "compilation error".
[0309] Figure 17 This is a schematic diagram of another protein processing interface provided in an embodiment of this application, as shown below. Figure 16 As shown, the computer device displays a protein processing interface. The first input area 1401 of the protein processing interface displays the protein sequence input by the user, and the second input area 1402 of the protein processing interface displays the mutation statement input by the user. If the mutation statement passes the verification, the protein processing interface also displays the mutation success message "Compilation successful". The display area 1404 of the protein processing interface displays the mutated protein sequence.
[0310] Once the user confirms that the mutated protein sequence is correct, if they wish to predict the protein structure corresponding to that sequence, they can trigger the "Structure Prediction" option in the protein processing interface. The computer device will respond to this trigger by displaying configuration window 1405 in the protein processing interface. For example... Figure 18 As shown, the configuration window 1405 is used to set resource configuration information, which is the configuration information of the equipment used to perform protein processing tasks. The configuration window 1405 includes input boxes for machine type configuration, database type, and the corresponding structural prediction model. The configuration window 1405 also includes "Complete" and "Cancel" options.
[0311] Figure 19 This is a schematic diagram of another task interface provided in an embodiment of this application, such as... Figure 19 As shown, the computer device displays a task interface, which includes multiple created protein processing tasks, their corresponding task status, processing time, and operation options. Task statuses include "Editing," "Predicting," "Predicted Successfully," and "Predicted Failed." The operation options for a protein processing task in the "Editing" state include "Edit" and "Delete." For a protein processing task in the "Predicting" state, the operation options include "Terminate" and "Delete." For a protein processing task in the "Predicted Successfully" state, the operation options include "View Results" and "Delete." For a protein processing task in the "Predicted Failed" state, the operation options include "View Log" and "Delete."
[0312] If a user wants to view the reason for the prediction failure of a protein processing task that is in a prediction failure state, they can trigger the "View Log" option corresponding to the protein processing task. The computer device responds to this trigger by displaying log window 2001 in the task interface. Figure 20 As shown, log window 2001 displays processing logs for predicted task execution failures.
[0313] The protein processing method provided in this application is executed by a computer device, which can be divided into a client and a server. The client provides page services, accessible via a URL, while the server provides processing and computing resources. When a user accesses the client, user actions on the client's page trigger interface calls in the server, thereby executing the protein processing task. For ease of understanding, this application provides a detailed description from both the client and server perspectives.
[0314] (1) Client side
[0315] The client includes a product introduction interface, a task interface, and a protein processing interface. After accessing the website, the client first displays the product introduction interface, then redirects to the task interface after logging in, and finally redirects to the protein processing interface after creating a protein processing task. In this embodiment, only the accounts registered in the client need to be maintained, and created protein processing tasks can be stored through the client. Each user has a limited number of protein processing tasks they can submit. For example, a newly registered account has 5 free submissions per week. Except for task failures due to scheduling errors, each successful and failed protein processing task deducts one free submission. When an account runs out of free submissions, tasks cannot be submitted, but created protein processing tasks can still be viewed until the following week when the 5 free submissions are restored.
[0316] Figure 21 This is a flowchart of a protein processing method provided in an embodiment of this application, such as... Figure 21 As shown, the method includes the following steps: 2101. Display the product introduction interface. 2102. Determine if an account is registered. If an account is registered, proceed directly to step 2103; otherwise, proceed to step 2103 after registering an account. 2103. Log in to the account. 2104. Display the task interface. 2105. Create a protein processing task. 2106. Display the protein processing interface. 2107. Execute the protein processing task. 2108. Determine if there are remaining execution attempts. If there are remaining attempts, proceed to step 2109; otherwise, return to step 2107. 2109. Determine if the prediction was successful. If the prediction was successful, proceed to step 2110; otherwise, return to step 2107. 2110. Display the predicted protein structure.
[0317] (2) Server-side aspects
[0318] The server-side consists of an interface server and a computation server. The interface server provides service-related interfaces to the client, while the computation server handles computation tasks. First, one account corresponds to one storage space. One account can create multiple protein processing tasks, and a single protein processing task can be submitted multiple times, corresponding to multiple computation records. The storage space corresponding to each account can be Cloud Object Storage (COS). Object storage is a distributed storage service without directory hierarchy or data format restrictions, capable of accommodating large amounts of data and supporting HTTP (Hypertext Transfer Protocol) / HTTPS (Hypertext Transfer Protocol over Secure Socket Layer) access. Users interact with the client through an interface displayed on their computer device. User operations are communicated from the client to the server via interfaces.
[0319] Figure 22 This is a schematic diagram of a protein processing method provided in an embodiment of this application, as shown below. Figure 22 As shown, the method includes the following steps: 2201. Log in to the account. 2202. Upload the protein sequence. The client uploads the protein sequence to the storage space corresponding to the account. 2203. The client sends a structure prediction request to the server, requesting prediction of the protein structure corresponding to the protein sequence. 2204. The server stores the task identifier and the account in the database. 2205. The server adds the task identifier to the identifier queue. 2206. The server creates a cloud host to batch compute protein processing tasks. 2207. The cloud host writes the prediction result to the storage space corresponding to the account; this prediction result is the predicted protein structure. 2208. The user requests the client to query the task, for example, requesting to view the prediction result. 2209. The server queries the prediction result in the storage space corresponding to the account. 2210. The server returns the prediction result to the client, which displays the prediction result.
[0320] Figure 23 This is a schematic diagram of another protein processing method provided in the embodiments of this application, as shown below. Figure 23As shown, the client connects to the interface server, which has access to the database, storage space, and identifier queue. The interface server writes the mapping between accounts and task identifiers in the database, protein sequences in the storage space, and task identifiers in the identifier queue. The computing server also accesses the storage space and identifier queue, reading protein sequences from the storage space and task identifiers from the identifier queue. The computing server connects to the cloud server, which creates cloud hosts used to predict the protein structures corresponding to the protein sequences.
[0321] Figure 24 This is a schematic diagram of another protein processing method provided in the embodiments of this application, as shown below. Figure 24 As shown, the method includes the following steps.
[0322] 2401. Log in to your account. 2402. Upload protein sequence. The user submits the protein sequence required for computation in the client, and the client uploads the protein sequence to the storage space corresponding to the account. 2403. Create a task. 2404. Store the account and task identifier in the database. 2405. Add task identifier to the identifier queue. The task identifier corresponding to the protein processing task is stored in the identifier queue. The server schedules the protein processing tasks indicated by the task identifiers in the identifier queue in a queue. 2406. When the execution condition is triggered, perform batch computation on the protein processing tasks indicated by the task identifiers in the identifier queue. After the protein processing task is scheduled, the server submits the batch computation task to the cloud host. During the computation process, a CVM is created based on the cloud host type and the container image corresponding to the computation type input by the user. The address of the storage space corresponding to the protein processing task is mounted to the host directory of the CVM. 2407. The cloud host retrieves the protein sequence from the storage space corresponding to the account and predicts the protein structure corresponding to the protein sequence. The predicted protein structure is stored in the storage space corresponding to the account. 2408. The client queries the server for the prediction results. 2409. The server retrieves the prediction results from the storage space. 2410. The server returns the prediction results to the client.
[0323] Figure 25 This is a schematic diagram of the structure of a protein processing device provided in an embodiment of this application. See also... Figure 25 The device includes:
[0324] Display module 2501 is used to display the input first protein sequence and mutation statement, the mutation statement indicating that the first protein sequence should be mutated, the mutation statement including at least one mutation instruction, the mutation instruction including mutation location and mutation method;
[0325] Mutation module 2502 is configured to respond to a processing request for the first protein sequence and the mutation statement, and for each mutation instruction in the mutation statement, mutate the mutation position in the first protein sequence according to the mutation method in the mutation instruction to obtain the second protein sequence.
[0326] The display module 2501 is used to display the second protein sequence.
[0327] The protein processing apparatus provided in this application embodiment allows for the input of the protein sequence and corresponding mutation statement to mutate a protein. The computer device then mutates the first protein sequence based on the mutation location and mutation method included in the mutation statement, resulting in a second protein sequence. This simulates the protein mutation process within the computer device, obtaining the mutated protein without the need for biochemical experiments, thus simplifying the protein mutation process and improving the efficiency of protein mutation, further enhancing the overall efficiency of protein processing.
[0328] Optionally, see Figure 26 The mutation module 2502 is used to perform any of the following:
[0329] The mutation method includes an original amino acid and a target amino acid, replacing the original amino acid at the mutation position in the first protein sequence with the target amino acid. The original amino acid refers to the amino acid before the mutation, and the target amino acid refers to the amino acid after the mutation.
[0330] The mutation method includes at least one original amino acid and a deletion marker, deleting at least one original amino acid at the mutation position in the first protein sequence, wherein the deletion marker indicates the deleted amino acid;
[0331] The mutation method includes at least one target amino acid and an insertion marker, inserting at least one target amino acid at the mutation position in the first protein sequence, the insertion marker indicating the inserted amino acid;
[0332] The mutation instruction includes at least one original amino acid, at least one target amino acid, and a deletion / insertion identifier, which deletes at least one original amino acid at the mutation position in the first protein sequence and inserts at least one target amino acid, wherein the deletion / insertion identifier indicates the deletion and insertion of the amino acid.
[0333] Optionally, see Figure 26 The device also includes:
[0334] Prediction module 2503 is used to predict the structure of the second protein corresponding to the second protein sequence in response to a structure prediction request.
[0335] The display module 2501 is also used to display the structure of the second protein.
[0336] Optionally, see Figure 26 The prediction module 2503 includes:
[0337] Prediction unit 2513 is configured to, in response to the structure prediction request, predict the second protein structure corresponding to the second protein sequence and predict the first protein structure corresponding to the first protein sequence.
[0338] Alignment unit 2523 is used to align the first protein structure with the second protein structure so that the first protein structure and the second protein structure can overlap through translation operation;
[0339] The display module 2501 is also used to display the aligned first protein structure and the second protein structure.
[0340] Optionally, see Figure 26 The alignment unit 2523 is used for:
[0341] Based on the first protein structure and the second protein structure, a spatial coordinate system is established;
[0342] The pose of the first protein structure or the pose of the second protein structure is adjusted so that the difference between the first coordinate information and the second coordinate information is less than a first threshold. The first coordinate information is the coordinate information of multiple target atoms in the first protein structure in the spatial coordinate system, and the second coordinate information is the coordinate information of the multiple target atoms in the second protein structure in the spatial coordinate system.
[0343] Optionally, see Figure 26 The device also includes an atom determination module 2504, used for:
[0344] The confidence level of multiple amino acids in the second protein structure is determined, whereby the confidence level of an amino acid represents the accuracy of the predicted position of that amino acid.
[0345] The preset atom in the amino acid whose confidence level is greater than the second threshold is identified as the target atom.
[0346] Optionally, see Figure 26 The display module 2501 is configured to perform at least one of the following:
[0347] Different display methods are used to show amino acids at mutated and non-mutated positions;
[0348] The amino acids at the mutation sites in the first protein structure and the second protein structure are displayed using different display methods;
[0349] The first protein structure and the second protein structure are superimposed and displayed, with the transparency of the first protein structure being higher than that of the second protein structure;
[0350] Based on the confidence levels of the amino acids in the first protein structure and the second protein structure, the colors of the amino acids in the first protein structure and the second protein structure are set respectively. The confidence level of the amino acid indicates the accuracy of the predicted position of the amino acid.
[0351] Optionally, see Figure 26 The device also includes:
[0352] The first processing module 2505 is configured to perform the same target operation on another protein structure in response to a target operation performed on either the first protein structure or the second protein structure, the target operation including at least one of rotation, translation, magnification or reduction.
[0353] Optionally, see Figure 26 The prediction module 2503 is used for:
[0354] In response to the structure prediction request, the sequence alignment information corresponding to the second protein sequence and the template protein structure are obtained. The sequence alignment information includes multiple homologous protein sequences corresponding to the second protein sequence and the difference information between the second protein sequence and the multiple homologous protein sequences. The template protein structure is the protein structure corresponding to the multiple homologous protein sequences.
[0355] The structure prediction model is invoked to predict the structure of the second protein corresponding to the second protein sequence based on the sequence alignment information and the template protein structure.
[0356] Optionally, see Figure 26 The mutation module 2502 includes:
[0357] Verification unit 2512 is used to verify the mutated statement;
[0358] Mutation unit 2522 is used to, when the mutation statement passes the verification, mutate the mutation position in the first protein sequence according to the mutation method in the mutation instruction for each mutation instruction in the mutation statement to obtain the second protein sequence.
[0359] The display module 2501 is also used to display mutation failure information if the mutation statement fails to pass the verification.
[0360] Optionally, see Figure 26 The display module 2501 includes:
[0361] The task creation unit 2511 is used to create a protein processing task and display the protein processing interface in response to a request to create a protein processing task.
[0362] Display unit 2521 is configured to, in response to an input operation, display the input first protein sequence and the mutation statement on the protein processing interface, and add the first protein sequence and the mutation statement to the protein processing task.
[0363] Optionally, see Figure 26 The display module 2501 is also used to display a task interface, which includes created protein processing tasks and the status of each created protein processing task. The status includes editing status, prediction status, prediction failure status, and prediction success status. The editing status means that the protein sequence or mutation statement is being edited.
[0364] Optionally, see Figure 26 The device also includes a second processing module 2506 for performing any of the following:
[0365] In response to the triggering operation of the editing option corresponding to the first protein processing task, the protein sequence and mutation statement in the first protein processing task are displayed in the protein processing interface, and the first protein processing task is in the editing state.
[0366] In response to the triggering operation of the termination option corresponding to the second protein processing task, the prediction of the protein structure corresponding to the protein sequence in the second protein processing task is stopped, and the second protein processing task is in the prediction state.
[0367] In response to the triggering operation of the log option corresponding to the third protein processing task, the processing log of the third protein processing task is displayed, which includes the reason for prediction failure and the third protein processing task is in a prediction failure state.
[0368] In response to the triggering operation of the result option corresponding to the fourth protein processing task, the predicted protein structure in the fourth protein processing task is displayed, and the fourth protein processing task is in a prediction success state.
[0369] Optionally, see Figure 26 The protein processing task also includes the second protein sequence, and the device further includes:
[0370] The prediction module 2503 is used to add the protein processing task to a task pool in response to a structure prediction request. The task pool includes protein processing tasks that have not yet been executed.
[0371] The prediction module 2503 is also used to predict the second protein structure corresponding to the second protein sequence in the protein processing task when the protein processing task meets the execution conditions.
[0372] The display module 2501 is used to display the structure of the second protein.
[0373] Optionally, see Figure 26 The prediction module 2503 is used for:
[0374] In response to the structure prediction request, the task identifier corresponding to the protein processing task is added to the identifier queue of the task pool, which includes task identifiers corresponding to protein processing tasks that have not yet been executed.
[0375] The protein processing task is stored in the storage space of the target account in the task pool. The target account is the account that created the protein processing task. The task pool includes the storage space of each account that created the protein processing task.
[0376] Optionally, see Figure 26 The prediction module 2503 is used for:
[0377] If the task identifier is the first task identifier in the identifier queue, the protein processing task indicated by the task identifier is retrieved from the storage space;
[0378] Predict the structure of the second protein corresponding to the second protein sequence in the protein processing task;
[0379] Remove the task identifier from the identifier queue.
[0380] Optionally, see Figure 26 The protein processing task also includes the second protein sequence, and the device further includes:
[0381] The prediction module 2503 is used to query the database for the number of task identifiers corresponding to the target account in response to the structure prediction request. The target account is the account that created the protein processing task. The database stores any account and the corresponding task identifiers of the executed protein processing tasks.
[0382] The prediction module 2503 is used to predict the second protein structure corresponding to the second protein sequence in the protein processing task when the number of task identifiers corresponding to the target account is less than a third threshold.
[0383] The display module 2501 is used to display the structure of the second protein.
[0384] Optionally, see Figure 26 The protein processing task also includes the second protein sequence, and the device further includes:
[0385] The prediction module 2503 is used to obtain input resource configuration information in response to a structure prediction request, the resource configuration information including configuration information of the device used to perform the protein processing task;
[0386] The prediction module 2503 is also used to determine the computing device that matches the resource configuration information; send the protein processing task to the computing device, and the computing device is used to predict the second protein structure corresponding to the second protein sequence in the protein processing task and return the second protein processing structure.
[0387] The prediction module 2503 is also used to receive the second protein processing structure returned by the computing device;
[0388] The display module 2501 is also used to display the structure of the second protein.
[0389] It should be noted that the protein processing apparatus provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the protein processing apparatus and the protein processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0390] This application also provides a computer device, which includes a processor and a memory. The memory stores at least one computer program, which is loaded and executed by the processor to perform the operations performed in the protein processing method of the above embodiments.
[0391] Optionally, the computer device is provided as a terminal. Figure 27 A schematic diagram of the structure of a terminal 2700 provided in an exemplary embodiment of this application is shown.
[0392] Terminal 2700 includes a processor 2701 and a memory 2702.
[0393] Processor 2701 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 2701 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). Processor 2701 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 2701 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 2701 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0394] Memory 2702 may include one or more computer-readable storage media, which may be non-transitory. Memory 2702 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in memory 2702 is used to store at least one computer program, which is used by processor 2701 to implement the protein processing method provided in the method embodiments of this application.
[0395] In some embodiments, the terminal 2700 may also optionally include a peripheral device interface 2703 and at least one peripheral device. The processor 2701, memory 2702, and peripheral device interface 2703 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 2703 via a bus, signal line, or circuit board. Optionally, the peripheral device includes at least one of the following: a radio frequency circuit 2704, a display screen 2705, a camera assembly 2706, an audio circuit 2707, or a power supply 2708.
[0396] Peripheral device interface 2703 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 2701 and memory 2702. In some embodiments, processor 2701, memory 2702 and peripheral device interface 2703 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 2701, memory 2702 and peripheral device interface 2703 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0397] The radio frequency (RF) circuit 2704 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 2704 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 2704 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 2704 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 2704 can communicate with other devices through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 2704 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0398] Display screen 2705 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 2705 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 2701 for processing. In this case, display screen 2705 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 2705, disposed on the front panel of terminal 2700; in other embodiments, there may be at least two display screens, disposed on different surfaces of terminal 2700 or in a folded design; in still other embodiments, display screen 2705 may be a flexible display screen, disposed on a curved or folded surface of terminal 2700. Furthermore, display screen 2705 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 2705 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0399] The camera assembly 2706 is used to acquire images or videos. Optionally, the camera assembly 2706 includes a front-facing camera and a rear-facing camera. The front-facing camera is disposed on the front panel of the terminal 2700, and the rear-facing camera is disposed on the back of the terminal 2700. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 2706 may also include a flash. The flash may be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cool light flash, which can be used for light compensation at different color temperatures.
[0400] The audio circuit 2707 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 2701 for processing, or to the radio frequency circuit 2704 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal 2700. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 2701 or the radio frequency circuit 2704 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 2707 may also include a headphone jack.
[0401] Power supply 2708 is used to power the various components in terminal 2700. Power supply 2708 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 2708 includes a rechargeable battery, the rechargeable battery can support wired or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0402] Those skilled in the art will understand that Figure 27 The structure shown does not constitute a limitation on terminal 2700 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0403] Optionally, the computer device is provided as a server. Figure 28 This is a schematic diagram of a server structure provided in an embodiment of this application. The server 2800 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 2801 and one or more memories 2802. The memories 2802 store at least one computer program, which is loaded and executed by the processor 2801 to implement the methods provided in the above-described method embodiments. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated upon here.
[0404] This application also provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to perform the operations of the protein processing method described above.
[0405] This application also provides a computer program product, including a computer program loaded and executed by a processor to perform operations as described in the protein processing method of the above embodiments. In some embodiments, the computer program involved in this application may be deployed and executed on a single computer device, or on multiple computer devices located in one location, or on multiple computer devices distributed across multiple locations and interconnected via a communication network. These multiple computer devices distributed across multiple locations and interconnected via a communication network can constitute a blockchain system.
[0406] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0407] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present application should be included within the protection scope of the present application.
Claims
1. A protein processing method, characterized in that, The method includes: The input first protein sequence and mutation statement are displayed. The mutation statement indicates that the first protein sequence should be mutated. The mutation statement includes at least one mutation instruction, which includes the mutation location and mutation method. In response to the processing request for the first protein sequence and the mutation statement, for each mutation instruction in the mutation statement, the mutation position in the first protein sequence is mutated according to the mutation method in the mutation instruction to obtain the second protein sequence; The second protein sequence is displayed; In response to a structure prediction request, predict the second protein structure corresponding to the second protein sequence and predict the first protein structure corresponding to the first protein sequence; Based on the first protein structure and the second protein structure, a spatial coordinate system is established; at least one of the poses of the first protein structure or the second protein structure is adjusted so that the difference between the first coordinate information and the second coordinate information is less than a first threshold, wherein the first coordinate information is the coordinate information of multiple target atoms in the first protein structure in the spatial coordinate system, and the second coordinate information is the coordinate information of the multiple target atoms in the second protein structure in the spatial coordinate system. The aligned first and second protein structures are shown.
2. The method according to claim 1, characterized in that, The step of mutating the mutation position in the first protein sequence according to the mutation method in the mutation instruction includes any one of the following: The mutation method includes an original amino acid and a target amino acid, replacing the original amino acid at the mutation position in the first protein sequence with the target amino acid. The original amino acid refers to the amino acid before the mutation, and the target amino acid refers to the amino acid after the mutation. The mutation method includes at least one of the original amino acids and a deletion marker, wherein at least one of the original amino acids at the mutation position in the first protein sequence is deleted, and the deletion marker indicates the deletion of the amino acid; The mutation method includes at least one target amino acid and an insertion marker, wherein at least one target amino acid is inserted at the mutation position in the first protein sequence, and the insertion marker indicates the inserted amino acid; The mutation instruction includes at least one of the original amino acids, at least one of the target amino acids, and a deletion / insertion marker, which deletes at least one of the original amino acids at the mutation position in the first protein sequence and inserts at least one of the target amino acids, wherein the deletion / insertion marker indicates the deletion and insertion of the amino acid.
3. The method according to claim 1, characterized in that, Before adjusting the pose of at least one of the first protein structure or the pose of the second protein structure, the method further includes: The confidence levels of multiple amino acids in the second protein structure are determined, where the confidence level of the amino acid represents the accuracy of the predicted position of the amino acid. The preset atoms in amino acids with a confidence level greater than the second threshold are identified as the target atoms.
4. The method according to claim 1, characterized in that, The alignment of the first and second protein structures is shown, including at least one of the following: Different display methods are used to show amino acids at mutated and non-mutated positions; Different display methods are used to show the amino acids at the mutation sites in the first protein structure and the second protein structure; The first protein structure and the second protein structure are superimposed and displayed, with the transparency of the first protein structure being higher than that of the second protein structure; Based on the confidence levels of amino acids in the first protein structure and the second protein structure, colors are set for the amino acids in the first protein structure and the second protein structure, respectively. The confidence level of the amino acid indicates the accuracy of the predicted position of the amino acid.
5. The method according to claim 1, characterized in that, The step of predicting the second protein structure corresponding to the second protein sequence in response to the structure prediction request includes: In response to the structure prediction request, sequence alignment information and template protein structure corresponding to the second protein sequence are obtained. The sequence alignment information includes multiple homologous protein sequences corresponding to the second protein sequence and difference information between the second protein sequence and the multiple homologous protein sequences. The template protein structure is the protein structure corresponding to the multiple homologous protein sequences. The structure prediction model is invoked to predict the structure of the second protein corresponding to the second protein sequence based on the sequence alignment information and the template protein structure.
6. The method according to any one of claims 1-5, characterized in that, The display of the input first protein sequence and mutation statement includes: In response to a request to create a protein processing task, a protein processing task is created and the protein processing interface is displayed. In response to the input operation, the input first protein sequence and the mutation statement are displayed on the protein processing interface, and the first protein sequence and the mutation statement are added to the protein processing task.
7. The method according to claim 6, characterized in that, The method further includes: The task interface displays the created protein processing tasks and the status of each created protein processing task. The status includes editing, prediction, prediction failure, and prediction success. The editing status means that the protein sequence or mutation statement is being edited.
8. The method according to claim 7, characterized in that, The method further includes any one of the following: In response to a triggering operation of the editing option corresponding to the first protein processing task, the protein sequence and mutation statement in the first protein processing task are displayed on the protein processing interface, and the first protein processing task is in the editing state. In response to the triggering operation of the termination option corresponding to the second protein processing task, the prediction of the protein structure corresponding to the protein sequence in the second protein processing task is stopped, and the second protein processing task is in the prediction state. In response to a trigger operation on the log option corresponding to the third protein processing task, the processing log of the third protein processing task is displayed, the processing log includes the prediction failure reason, and the third protein processing task is in a prediction failure state. In response to a trigger operation on the result option corresponding to the fourth protein processing task, the predicted protein structure in the fourth protein processing task is displayed, and the fourth protein processing task is in a prediction success state.
9. The method according to claim 6, characterized in that, The protein processing task further includes the second protein sequence, and the method further includes: In response to a structure prediction request, the protein processing task is added to a task pool, which includes protein processing tasks that have not yet been executed. When the protein processing task meets the execution conditions, predict the second protein structure corresponding to the second protein sequence in the protein processing task, and display the second protein structure.
10. The method according to claim 9, characterized in that, The step of adding the protein processing task to the task pool in response to the structure prediction request includes: In response to the structure prediction request, the task identifier corresponding to the protein processing task is added to the identifier queue of the task pool, the identifier queue including task identifiers corresponding to protein processing tasks that have not yet been executed; The protein processing task is stored in the storage space of the target account in the task pool. The target account is the account that created the protein processing task. The task pool includes the storage space of each account that created the protein processing task.
11. The method according to claim 10, characterized in that, The step of predicting the second protein structure corresponding to the second protein sequence in the protein processing task when the protein processing task meets the execution conditions includes: If the task identifier is the first task identifier in the identifier queue, the protein processing task indicated by the task identifier is retrieved from the storage space. Predict the structure of the second protein corresponding to the second protein sequence in the protein processing task; Remove the task identifier from the identifier queue.
12. The method according to claim 6, characterized in that, The protein processing task further includes the second protein sequence, and the method further includes: In response to the structure prediction request, the database is queried for the number of task identifiers corresponding to the target account, where the target account is the account that created the protein processing task, and the database stores any account and the corresponding task identifiers of the executed protein processing tasks. If the number of task identifiers corresponding to the target account is less than a third threshold, predict the second protein structure corresponding to the second protein sequence in the protein processing task and display the second protein structure.
13. The method according to claim 6, characterized in that, The protein processing task further includes the second protein sequence, and the method further includes: In response to a structure prediction request, input resource configuration information is obtained, including configuration information of the device used to perform the protein processing task; Identify the computing device that matches the resource configuration information; The computing device sends the protein processing task to the computing device, which predicts the second protein structure corresponding to the second protein sequence in the protein processing task and returns the second protein processed structure. The second protein processing structure returned by the computing device is received, and the second protein structure is displayed.
14. A protein processing apparatus, characterized in that, The device includes: The display module is used to display the input first protein sequence and mutation statement, wherein the mutation statement indicates that the first protein sequence is mutated, and the mutation statement includes at least one mutation instruction, wherein the mutation instruction includes the mutation location and mutation method; A mutation module is configured to, in response to a processing request for the first protein sequence and the mutation statement, mutate the mutation position in the first protein sequence according to the mutation method in the mutation statement for each mutation instruction in the mutation statement, to obtain a second protein sequence; The display module is used to display the second protein sequence; The prediction module is used to predict the second protein structure corresponding to the second protein sequence and the first protein structure corresponding to the first protein sequence in response to the structure prediction request. The prediction module is further configured to establish a spatial coordinate system based on the first protein structure and the second protein structure; and to adjust at least one of the pose of the first protein structure or the pose of the second protein structure so that the difference between the first coordinate information and the second coordinate information is less than a first threshold, wherein the first coordinate information is the coordinate information of multiple target atoms in the first protein structure in the spatial coordinate system, and the second coordinate information is the coordinate information of the multiple target atoms in the second protein structure in the spatial coordinate system. The display module is also used to display the aligned first protein structure and the second protein structure.
15. The apparatus according to claim 14, characterized in that, The mutation module is used to perform any of the following: The mutation method includes an original amino acid and a target amino acid, replacing the original amino acid at the mutation position in the first protein sequence with the target amino acid. The original amino acid refers to the amino acid before the mutation, and the target amino acid refers to the amino acid after the mutation. The mutation method includes at least one of the original amino acids and a deletion marker, wherein at least one of the original amino acids at the mutation position in the first protein sequence is deleted, and the deletion marker indicates the deletion of the amino acid; The mutation method includes at least one target amino acid and an insertion marker, wherein at least one target amino acid is inserted at the mutation position in the first protein sequence, and the insertion marker indicates the inserted amino acid; The mutation instruction includes at least one of the original amino acids, at least one of the target amino acids, and a deletion / insertion marker, which deletes at least one of the original amino acids at the mutation position in the first protein sequence and inserts at least one of the target amino acids, wherein the deletion / insertion marker indicates the deletion and insertion of the amino acid.
16. The apparatus according to claim 14, characterized in that, The device further includes an atom determination module for: The confidence levels of multiple amino acids in the second protein structure are determined, where the confidence level of the amino acid represents the accuracy of the predicted position of the amino acid. The preset atoms in amino acids with a confidence level greater than the second threshold are identified as the target atoms.
17. The apparatus according to claim 14, characterized in that, The display module is configured to perform at least one of the following: Different display methods are used to show amino acids at mutated and non-mutated positions; Different display methods are used to show the amino acids at the mutation sites in the first protein structure and the second protein structure; The first protein structure and the second protein structure are superimposed and displayed, with the transparency of the first protein structure being higher than that of the second protein structure; Based on the confidence levels of amino acids in the first protein structure and the second protein structure, colors are set for the amino acids in the first protein structure and the second protein structure, respectively. The confidence level of the amino acid indicates the accuracy of the predicted position of the amino acid.
18. The apparatus according to claim 14, characterized in that, The prediction module is used for: In response to the structure prediction request, sequence alignment information and template protein structure corresponding to the second protein sequence are obtained. The sequence alignment information includes multiple homologous protein sequences corresponding to the second protein sequence and difference information between the second protein sequence and the multiple homologous protein sequences. The template protein structure is the protein structure corresponding to the multiple homologous protein sequences. The structure prediction model is invoked to predict the structure of the second protein corresponding to the second protein sequence based on the sequence alignment information and the template protein structure.
19. The apparatus according to any one of claims 14-18, characterized in that, The display module includes: The task creation unit is used to create a protein processing task and display the protein processing interface in response to a request to create a protein processing task. A display unit is configured to, in response to an input operation, display the input first protein sequence and the mutation statement on the protein processing interface, and add the first protein sequence and the mutation statement to the protein processing task.
20. The apparatus according to claim 19, characterized in that, The display module is also used for: The task interface displays the created protein processing tasks and the status of each created protein processing task. The status includes editing, prediction, prediction failure, and prediction success. The editing status means that the protein sequence or mutation statement is being edited.
21. The apparatus according to claim 20, characterized in that, The device further includes a second processing module for performing any of the following: In response to a triggering operation of the editing option corresponding to the first protein processing task, the protein sequence and mutation statement in the first protein processing task are displayed on the protein processing interface, and the first protein processing task is in the editing state. In response to the triggering operation of the termination option corresponding to the second protein processing task, the prediction of the protein structure corresponding to the protein sequence in the second protein processing task is stopped, and the second protein processing task is in the prediction state. In response to a trigger operation on the log option corresponding to the third protein processing task, the processing log of the third protein processing task is displayed, the processing log includes the prediction failure reason, and the third protein processing task is in a prediction failure state. In response to a trigger operation on the result option corresponding to the fourth protein processing task, the predicted protein structure in the fourth protein processing task is displayed, and the fourth protein processing task is in a prediction success state.
22. The apparatus according to claim 19, characterized in that, The protein processing task further includes the second protein sequence, and the apparatus further includes: A prediction module is used to add the protein processing task to a task pool in response to a structure prediction request. The task pool includes protein processing tasks that have not yet been executed. The prediction module is further configured to predict the second protein structure corresponding to the second protein sequence in the protein processing task and display the second protein structure when the protein processing task meets the execution conditions.
23. The apparatus according to claim 22, characterized in that, The prediction module is used for: In response to the structure prediction request, the task identifier corresponding to the protein processing task is added to the identifier queue of the task pool, the identifier queue including task identifiers corresponding to protein processing tasks that have not yet been executed; The protein processing task is stored in the storage space of the target account in the task pool. The target account is the account that created the protein processing task. The task pool includes the storage space of each account that created the protein processing task.
24. The apparatus according to claim 23, characterized in that, The prediction module is used for: If the task identifier is the first task identifier in the identifier queue, the protein processing task indicated by the task identifier is retrieved from the storage space. Predict the structure of the second protein corresponding to the second protein sequence in the protein processing task; Remove the task identifier from the identifier queue.
25. The apparatus according to claim 19, characterized in that, The protein processing task further includes the second protein sequence, and the apparatus further includes: The prediction module is used to query the database for the number of task identifiers corresponding to the target account in response to the structure prediction request. The target account is the account that created the protein processing task. The database stores any account and the task identifiers of the corresponding executed protein processing tasks. The prediction module is further configured to predict the second protein structure corresponding to the second protein sequence in the protein processing task and display the second protein structure when the number of task identifiers corresponding to the target account is less than a third threshold.
26. The apparatus according to claim 19, characterized in that, The protein processing task further includes the second protein sequence, and the apparatus further includes: The prediction module is used to obtain input resource configuration information in response to a structure prediction request, the resource configuration information including configuration information of the device used to perform the protein processing task; The prediction module is also used to determine the computing device that matches the resource configuration information; The prediction module is further configured to send the protein processing task to the computing device, and the computing device is configured to predict the second protein structure corresponding to the second protein sequence in the protein processing task and return the second protein processing structure. The display module is further configured to receive the second protein processing structure returned by the computing device and display the second protein structure.
27. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one computer program, which is loaded and executed by the processor to perform the operations of the protein processing method as described in any one of claims 1 to 13.
28. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to perform the operations performed by the protein processing method as described in any one of claims 1 to 13.
29. A computer program product, comprising a computer program, characterized in that, The computer program is loaded and executed by a processor to perform the operations performed by the protein processing method as described in any one of claims 1 to 13.