Method and apparatus for recommending protein mutation, and computer device and storage medium

By specifying instructions and generating score results in protein mutation experiments, the tedious problem of multiple rounds of manual input in protein mutation experiments is solved, improving screening efficiency and convenience.

WO2026114306A1PCT designated stage Publication Date: 2026-06-04KANGMA (SHANGHAI) BIOTECH LTD +1

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
KANGMA (SHANGHAI) BIOTECH LTD
Filing Date
2025-11-27
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

In existing technologies, protein mutation experiments require multiple rounds of manual input of protein sequence information for experimental verification, which is cumbersome and has low screening efficiency.

Method used

By specifying the first mutation instruction, the second and third mutation instructions are generated to mutate the original protein sequence. The scoring results of the mutated protein sequence are displayed one by one with the instructions, forming a loop operation until the optimal sequence is found, reducing the number of manual input steps.

Benefits of technology

It improves the efficiency and convenience of protein mutation screening, and simplifies the user workflow through intuitive scoring result display and automated operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025138131_04062026_PF_FP_ABST
    Figure CN2025138131_04062026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a method and apparatus for recommending a protein mutation, and a computer device and a storage medium. The method comprises: receiving a mutation task request, wherein the request carries a specified first mutation instruction; in response to the request, generating third mutation instructions on the basis of the first mutation instruction and second mutation instructions; mutating an original protein sequence, so as to obtain mutated protein sequences corresponding to the third mutation instructions on a one-to-one basis; then, displaying scoring results and the third mutation instructions in a manner of corresponding to each other on a one-to-one basis; and determining whether a new mutation task request has been received, and if so, returning to continue to receive the new mutation task request. It can be seen that after the scoring results and the third mutation instructions are displayed in a manner of corresponding to each other on a one-to-one basis and a user performs experimental verification, a first mutation instruction is specified in a mutation task request again, thus forming a cyclic operation until the user finds an optimal protein sequence. Therefore, the user does not need to manually input a protein sequence, and thus the operation is convenient, and the efficiency of mutated-protein screening can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Recommended methods, apparatus, computer equipment, and storage media for protein mutation analysis. Technical Field

[0001] This invention belongs to the field of molecular biology experimental equipment technology, specifically relating to a method, apparatus, computer equipment, and storage medium for recommending protein mutations. Background Technology

[0002] Proteins are organic molecules composed of amino acids and are the material basis of life. Mutations in proteins can affect their function and drug resistance, so understanding and predicting the sequences of mutant proteins, and thus predicting their structures, is of great significance in biological research.

[0003] To obtain mutant proteins, various protein mutation models are currently available. Based on relevant information of the sequence to be mutated, mutation prediction models predict the mutation results, from which experimenters select for experiments. To obtain the desired mutant protein, multiple rounds of mutation need to be performed based on the experimental results using mutation prediction models. However, in the current method, each round of mutation requires experimenters to manually input the relevant information of the sequence that has been experimentally verified as optimal, and start a new round of mutation and mutant protein screening. The operation is cumbersome, resulting in low screening efficiency.

[0004] Therefore, there is an urgent need for a convenient method to recommend protein mutations in order to improve screening efficiency. Summary of the Invention

[0005] To address the aforementioned technical problems, this application provides a method, apparatus, computer device, and storage medium for recommending protein mutations. The method involves specifying a first mutation instruction in a mutation task, and generating a third mutation instruction based on the specified first mutation instruction and a newly generated second mutation instruction. This third mutation instruction is then used to mutate the original protein sequence. Finally, the scoring results of the mutated protein sequence are displayed in a one-to-one correspondence with the third mutation instruction for user reference and experimental verification. After the user performs experimental verification, they specify the first mutation instruction again in the mutation task request, thus forming a cyclical operation until the user finds the optimal protein sequence. Therefore, users do not need to manually input experimentally verified protein sequences; they only need to specify mutation instructions in the mutation task. This convenient operation improves the screening efficiency of mutated proteins.

[0006] This application provides a recommended method for protein mutation, including:

[0007] Step S1: Receive a mutation task request, which carries the original protein sequence, the location to be mutated, the mutation method, and the specified first mutation instruction.

[0008] Step S2: In response to the mutation task request, execute:

[0009] Step S21: Based on the mutation method and the mutation location, generate multiple second mutation instructions corresponding to the mutation location, generate multiple third mutation instructions based on the first mutation instructions and the second mutation instructions, mutate the original protein sequence, and obtain a mutant protein sequence that corresponds one-to-one with each of the third mutation instructions;

[0010] Step S22: Score at least each of the mutant protein sequences and obtain the score results. Display the score results corresponding to each mutant protein sequence in a one-to-one correspondence with the third mutation command.

[0011] Step S3: Determine whether a new mutation task request has been received. If so, return to step S1.

[0012] When the number of times the mutation task request is received is greater than 1, the first mutation instruction selects a specified good mutation instruction from the previous third mutation instruction.

[0013] In one specific embodiment, the corresponding display also corresponds to displaying the scoring result of the original protein sequence, which is obtained by scoring the original protein sequence.

[0014] When the number of times the mutation task request is received is greater than 1, the corresponding display will also display the corresponding score result of the mutant protein sequence corresponding to the first mutation instruction in correspondence with the first mutation instruction. The score result is obtained by scoring the mutant protein sequence corresponding to the specified first mutation instruction.

[0015] Receiving a mutation task request also includes the fourth mutation instruction, which is a specified bad mutation instruction selected from the previous third mutation instruction. When the number of times the mutation task request is received is greater than 1, the corresponding display will also display the corresponding score result of the mutant protein sequence corresponding to the fourth mutation instruction in correspondence with the fourth mutation instruction. The score result is obtained by scoring the mutant protein sequence corresponding to the specified fourth mutation instruction.

[0016] In one specific embodiment, the mutation task request information also carries protein structure data corresponding to the original protein sequence and at least one set of scoring positions selected on the original protein sequence, with a scoring fragment selected on each set of scoring positions.

[0017] In one specific embodiment, the corresponding display also includes a marker button for the user to mark the corresponding first mutation instruction.

[0018] In one specific embodiment, the scoring results are displayed in the form of scores; and / or, in the form of a sort, the sorting being determined based on the scores.

[0019] In one specific embodiment, the mutation method is substitution, and the mutation method includes one of the following mutation modes: single-point mutation, saturation mutation, and combined mutation, wherein the second mutation instruction is generated based on one of the single-point mutation, the saturation mutation, and the combined mutation;

[0020] The single-point mutation refers to replacing each of the mutated positions in each mutation task request with all other amino acids in turn to obtain a mutant protein sequence in which only a single mutated position is replaced.

[0021] The saturation mutation refers to replacing all the positions to be mutated in each mutation task request with all other amino acids in turn to obtain a mutant protein sequence in which all the positions to be mutated are replaced.

[0022] The combined mutation refers to the arbitrary combination of the first mutation instructions with different selected mutation positions in each mutation task request, and the sequential replacement of the corresponding positions in the original protein sequence.

[0023] The generation of multiple third mutation instructions based on the first mutation instruction and the second mutation instruction includes:

[0024] When the mutation method is the single-point mutation or the saturation mutation, each third mutation instruction is generated by a combination of any first mutation instruction and any second mutation instruction.

[0025] When the mutation method is the combined mutation, each of the third mutation instructions is generated by any combination of the first mutation instructions with different selected mutation positions.

[0026] In a specific embodiment:

[0027] When the number of times the mutation task request is received is equal to 1, if the position to be mutated is omitted, it means that all positions of the original protein sequence are mutated.

[0028] When the number of times the mutation task request is received is greater than 1, if the position to be mutated is omitted, it means that the position to be mutated is a mutation position other than the mutation position corresponding to the first mutation instruction specified in the previous task.

[0029] In one specific embodiment, when the number of times the mutation task request is received is equal to 1, the mutation task is input on a first input interface, which has multiple first input areas; wherein, input is performed by manual input, selection from a drop-down list, and input by clicking or dragging;

[0030] The plurality of first input regions are used to respectively input the mutation method, the protein structure data, the chain number corresponding to the original protein sequence in the three-dimensional structure, the position to be mutated, at least one scoring position group selected on the original protein sequence, and the original protein sequence;

[0031] When the number of times the mutation task request is received is greater than 1, the mutation task is created based on the previous mutation task on the second input interface, which has multiple second input areas:

[0032] The plurality of second input areas are used to respectively display the mutation method in the previous mutation task, the protein structure data, the chain number corresponding to the original protein sequence in the three-dimensional structure, the position to be mutated in the previous mutation task, the selection of at least one scoring position group on the original protein sequence, and the original protein sequence. Furthermore, any one or more of the mutation method, chain number, position to be mutated, and scoring position group in the previous mutation task can be changed in the second input interface.

[0033] The plurality of second input interfaces also include a second input region for displaying the mutant protein sequence corresponding to the first mutation command specified by the user;

[0034] Furthermore, the first input interface and / or the second input interface are provided with a task submission button, by clicking the button to send the mutation task request, and / or, each also includes a first input region or a second input region that displays the three-dimensional structure of the original protein sequence.

[0035] This application also provides an apparatus for recommending protein mutations, including: a receiving module, a response module, a mutation module, a scoring display module, and a judgment module;

[0036] The receiving module is used to receive a mutation task request, which carries the original protein sequence, the location to be mutated, the mutation method, and a specified first mutation instruction.

[0037] The response module is configured to, in response to the mutation task request, instruct the mutation module to execute:

[0038] Based on the mutation method and the location to be mutated, a plurality of second mutation instructions corresponding to the location to be mutated are generated; based on the first mutation instructions and the second mutation instructions, a plurality of third mutation instructions are generated; the original protein sequence is mutated to obtain a mutant protein sequence corresponding one-to-one with each of the third mutation instructions; and

[0039] Instruct the scoring display module to perform:

[0040] Each mutant protein sequence is scored at least once and a score result is obtained. The score result corresponding to each mutant protein sequence is then displayed in a one-to-one correspondence with the third mutation instruction.

[0041] The judgment module is used to determine whether a new mutation task request has been received. If so, it instructs the receiving module to receive the new mutation task request.

[0042] When the number of times the mutation task request is received is greater than 1, the first mutation instruction selects a specified good mutation instruction from the previous third mutation instruction.

[0043] This application also provides a computer device including a processor and a memory, wherein the memory stores at least one computer program, which is loaded and executed by the processor to perform the operations of the method for recommending protein mutations as described in any of the embodiments of this application.

[0044] This application also provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to perform the operations of the method for recommending protein mutations as described in any one of the embodiments of this application.

[0045] This application also provides a computer program product, including a computer program, characterized in that the computer program is loaded and executed by a processor to perform the operations of the method for recommending protein mutations as described in any of the embodiments of this application.

[0046] The beneficial effects of this application are as follows:

[0047] As can be seen from the above, by specifying the first mutation instruction in the mutation task, and generating a third mutation instruction based on the specified first mutation instruction and the newly generated second mutation instruction, the original protein sequence is mutated. Finally, the corresponding score results of the mutated protein sequence are displayed one-to-one with the third mutation instruction for users to refer to and verify experimentally. After the user performs experimental verification, the first mutation instruction is specified again in the mutation task request, thus forming a cyclical operation until the user finds the optimal protein sequence. Therefore, users do not need to manually input the experimentally verified protein sequence and related information; they only need to specify the mutation instruction in the mutation task. The operation is convenient and can improve the efficiency of mutated protein screening.

[0048] Furthermore, the scoring results of the mutant protein sequences are displayed one-to-one with the third mutation command. This correspondence display visualizes and intuitively presents the scoring results of each mutant protein sequence, making it convenient for users to quickly select mutant protein sequences for experimental verification. Moreover, this correspondence display also shows the scoring results of the original protein sequence. When the number of received mutation task requests is greater than 1, the correspondence display also shows the scoring results of the mutant protein sequences corresponding to the first mutation command and the first mutation command, providing users with a more intuitive reference when making selections.

[0049] Furthermore, the corresponding display also includes a button for users to mark the corresponding first mutation instruction. Users only need to click the mark button to mark the corresponding first mutation instruction. The operation is intelligent and does not require users to do much work, saving time and effort.

[0050] Furthermore, when the number of mutation task requests received is greater than 1, the mutation task is automatically created in the second input interface based on the previous mutation task. Only when the corresponding items need to be modified, the mutation task can be modified in the second input interface and then submitted. There is no need to manually re-enter all the data again, which is simple, time-saving and labor-saving, thereby improving the efficiency of protein sequence screening. Attached Figure Description

[0051] To more clearly illustrate the technical solution of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0052] Figure 1 is a flowchart of a method for recommending protein mutations provided in an embodiment of this application;

[0053] Figure 2 is a schematic diagram of the first input interface 200 provided in an embodiment of this application;

[0054] Figure 3 is a schematic diagram of the corresponding display 300 provided in the embodiment of this application;

[0055] Figure 4 is a schematic diagram of the second input interface 400 provided in an embodiment of this application;

[0056] Figure 5 is a schematic diagram of the corresponding display 500 provided in the embodiment of this application;

[0057] Figure 6 is a schematic diagram of the structure of the apparatus for recommending protein mutations provided in an embodiment of this application;

[0058] Figure 7 is an internal structural diagram of a computer device provided in an embodiment of this application;

[0059] Figure 8 is an internal structural diagram of a computer device provided in another embodiment of this application. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0061] It is understood that "at least one" as used in this application refers to one or more amino acids. For example, at least one amino acid can be any integer number of amino acids greater than or equal to one, such as one amino acid, two amino acids, three amino acids, etc. "More than one" refers to two or more amino acids. For example, more than one amino acid can be any integer number of amino acids greater than or equal to two amino acids, such as two amino acids, three amino acids, etc. "Each" refers to each of the at least one amino acids. For example, each amino acid refers to each of the more than one amino acids. If the more than one amino acids are three amino acids, then each amino acid refers to each of the three amino acids.

[0062] It is understood that the embodiments of this application involve data such as user information, protein sequences, and mutation instructions. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0063] This application provides a method for recommending protein mutations, which can be used in computer devices. Optionally, the computer device is a terminal or a server. Optionally, the terminal is a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. Optionally, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The method provided in this application can also be applied to systems including terminals and servers, and implemented through the interaction between the terminal and the server.

[0064] In one possible implementation, the computer program involved in the embodiments of this application may be deployed and executed on a computer device, or executed on multiple computer devices located in one location, or multiple computer devices distributed in multiple locations and interconnected through a communication network may form a blockchain system.

[0065] Figure 1 is a flowchart of a method for recommending protein mutations according to an embodiment of this application. Referring to Figure 1, the method includes:

[0066] Step S1: Receive a mutation task request, which carries the original protein sequence, the location to be mutated, the mutation method, and the specified first mutation instruction.

[0067] Step S2: In response to the mutation task request, execute:

[0068] Step S21: Based on the mutation method and the location to be mutated, generate multiple second mutation instructions corresponding to the location to be mutated, generate multiple third mutation instructions based on the first mutation instructions and the second mutation instructions, and mutate the original protein sequence based on the multiple third mutation instructions to obtain a mutant protein sequence that corresponds one-to-one with each third mutation instruction.

[0069] Step S22: At least score each mutant protein sequence and obtain the score results. Display the corresponding score results of each mutant protein sequence in a one-to-one correspondence with the third mutation instruction.

[0070] Step S3: Determine whether a new mutation task request has been received. If so, return to step S.

[0071] Specifically, when the number of received mutation task requests is greater than 1, the first mutation instruction is selected from the previous third mutation instructions. That is, the first mutation instruction in the current step S1 is selected as a good mutation instruction from the multiple third mutation instructions in the previous step S21. After the user experimentally verifies the mutated protein sequences corresponding to the multiple third mutation instructions, if the mutated protein sequence is found to be a good protein sequence, the mutation instruction corresponding to that protein sequence is designated as a good mutation instruction. Similarly, if the experimental verification reveals that the mutated protein is a bad protein sequence, the mutation instruction corresponding to that protein sequence is designated as a bad mutation instruction.

[0072] Specifically, proteins are organic molecules composed of amino acids and are the material basis of life. Amino acids are the basic building blocks of proteins. Generally, there are 20 common amino acids, including polar amino acids, nonpolar amino acids, charged amino acids, and special amino acids. Proteins can be composed of these 20 amino acids in different proportions. A protein sequence includes the amino acid sequence of the protein, from which information such as the types, numbers, and order of amino acids can be obtained. For example, the protein sequence of GB1 is a sequence of 56 amino acids as follows: “MTYKLILNGKTLKGETTTEAVDAATAEKVFKQYANDNGVDGEWTYDDATKTFTVTE”.

[0073] The original protein sequence in step S1 above can be obtained from a PDB (Protein Data Bank). The protein PDB file contains atomic coordinates, chemical bonds, and other relevant information related to the structure of the original protein sequence. Of course, it can also be obtained from other databases. The original protein sequence refers to a protein sequence that has not undergone any mutations, while the mutated protein sequence refers to a protein sequence that has undergone mutations based on the original protein sequence.

[0074] It should be noted that the first mutation instruction in step S1 above can be empty, that is, the user does not specify any first mutation instruction at this time, and the third mutation instruction is the same as the second mutation instruction.

[0075] For example, when performing a single-point mutation at position 40 of chain A of GB1, if the specified first mutation instruction is empty, the second mutation instruction generated based on the single-point mutation and the mutation at position 40 of chain A of GB1 is as follows: D[A40]T, D[A40]S, D[A40]E, D[A40]N, D[A40]H, D[A40]K, D[A40]A, D[A40]V, D[A40]Q, D[A40]I, D[A40]Y, D[A40]L, D[A40]F, D[A40]M, D[A40]R, D[A40]G, D[A40]W, D[A40]C, D[A40]P. Then, when generating multiple third mutation instructions based on the first and second mutation instructions, since the first mutation instruction is empty, the generated multiple third mutation instructions are the same as the second mutation instructions.

[0076] Of course, the first mutation instruction in step S1 above can also be non-empty, that is, the user specifies the first mutation instruction. For example, when performing a single-point mutation on the 40th position of the A chain of GB1 above, assuming the user specifies the first mutation instruction and the specified first mutation instruction is V

[0039] I, then according to the single-point mutation and the mutation on the 40th position of the A chain of GB1, the generated second mutation instruction is as follows: D[A40]T, D[A40]S, D[A40]E, D[A40]N, D[A40]H, D[A40]K, D[A40]A, D[A40]V, D[A40]Q, D[A40]I, D[A40]Y, D[A40]L, D[A40]F, D[A40]M, D[A40]R, D[A40]G, D[A40]W, D[A40]C, D[A40]P.

[0077] Then, based on the first mutation instruction V

[0039] I and the generated second mutation instruction, multiple third mutation instructions are generated as follows: V

[0039] ID[A40]T, V

[0039] ID[A40]S, V

[0039] ID[A40]E, V

[0039] ID[A40]N, V

[0039] ID[A40]H, V

[0039] ID[A40]K, V

[0039] ID[A40]A, V

[0039] ID[A40]V, V

[0039] ID[A40]Q, V

[0039] ID[A40]I, V

[0039] ID[A40]Y, V

[0039] ID[A40]Y, V

[0039] ID[A40]I ... ]ID[A40]L, V

[0039] ID[A40]F, V

[0039] ID[A40]M, V

[0039] ID[A40]R, V

[0039] ID[A40]G, V

[0039] ID[A40]W, V

[0039] ID[A40]C, V

[0039] ID[A40]P. It can be seen that at this time, the third mutation instruction is different from the second mutation instruction. Therefore, when the first mutation instruction is empty, that is, when the user does not specify the first mutation instruction, the third mutation instruction is the same as the second mutation instruction. However, when the user specifies the first mutation instruction, the third mutation instruction is different from the second mutation instruction.

[0078] In step S22 above, when scoring each mutant protein sequence and obtaining the scoring result, existing protein language models such as ESM (Embedding of Sequence Motifs), ProtGPT (Protein Generative Pre-Trained Transformer), and ProtSSN (Protein Spatial Shortcut Network) can be used to score each mutant protein sequence. For example, the basic neural network of the protein language model can be a diffusion model or a transformation network model. In one possible implementation, the protein language model can be trained using machine learning methods. For example, a sample protein sequence is input into a basic neural network, which outputs the mutation location in the sample protein sequence and the probability value of the mutation to a certain type of amino acid. Based on the difference between this prediction result and the annotation result of the sample protein sequence, the parameters of the basic neural network are iteratively adjusted to obtain the protein language model.

[0079] Specifically, the mutation mode in step S1 refers to the way in which amino acids in the original protein sequence are mutated, such as substitution mutations, deletion mutations, or insertion mutations in the original protein sequence.

[0080] in:

[0081] (1) Substitution mutation refers to replacing one or more original amino acids at the position to be mutated with the target amino acid. Optionally, the format of the mutation instruction is [xx][position][yy], where [xx] represents the identifier of the original amino acid, [position] represents the number of the position to be mutated, and [yy] represents the identifier of the target amino acid. The mutation instruction means replacing the original amino acid at [position] with the target amino acid.

[0082] For example, the amino acid at position 39 of the A chain of the GB1 protein sequence is valine V. If the mutation position is 39, and you want to replace valine V at position 39 with isoleucine I, the mutation instruction is "V[A39]I", which means replacing valine V at position 39 of the A chain with isoleucine I.

[0083] It should be noted that this explanation only uses a single mutation instruction as an example. If there are multiple mutation instructions, such as "V[A39]ID[A40]T", it means that the valine V at position 39 of chain A is replaced with isoleucine I, and the aspartic acid D at position 40 of chain A is replaced with threonine T.

[0084] (2) Deletion mutation refers to the deletion of one or more original amino acids at the position to be mutated, including at least one original amino acid and a deletion marker. For example, at least one original amino acid at the position to be mutated in the original protein sequence can be deleted, and the deletion marker indicates the deletion of the amino acid.

[0085] For example, the original amino acid at a position to be mutated can be deleted. The format of the mutation instruction corresponding to this single deletion mutation can be [xx][position]del[xx], where [xx] represents the identifier of the original amino acid, [position] represents the number of the position to be mutated, and del is the deletion identifier. This mutation instruction indicates that the original amino acid [xx] at [position] will be deleted.

[0086] For example, the fourth position of the A strand of the GB1 protein sequence is lysine K, and the position to be mutated is 4. If you want to delete lysine K at the fourth position, the mutation instruction is "K[A4]delK", which means to delete lysine K at the fourth position of the A strand of the GB1 protein sequence.

[0087] It should be noted that this explanation only uses the deletion of a single amino acid as an example. Of course, multiple amino acids can also be deleted. For example, the original amino acids at multiple consecutive mutation sites can be deleted. The format of the mutation command corresponding to this deletion mutation is [xx1][positon1]_[xx2][positon2]del[xx_seq]. [positon1] and [positon2] represent the mutation sites, [xx1] represents the identifier of the original amino acid at [positon1], [xx2] represents the identifier of the original amino acid at [positon2], [xx_seq] represents the identifier of multiple original amino acids from [positon1] to [positon2], and del is the deletion identifier. This mutation command indicates that multiple original amino acids [xx_seq] from [positon1] to [positon2] will be deleted.

[0088] For example, for the A chain of the GB1 protein sequence mentioned above, the mutation instruction is "V[A39]_G[A41]delVDG". This mutation instruction indicates the deletion of multiple amino acid VDGs from valine (V) at position 39 to glycine (G) at position 41. Here, amino acid VDG represents valine (V), aspartic acid (D), and glycine (G).

[0089] (3) Insertion mutation refers to inserting at least one target amino acid at the position to be mutated, including at least one target amino acid and an insertion marker. For example, at least one target amino acid can be inserted at the position to be mutated in the original protein sequence, and the insertion marker indicates the inserted amino acid.

[0090] For example, at least one target amino acid can be inserted at the position to be mutated. The format of the mutation instruction corresponding to this insertion mutation can be [xx1][positon1]_[xx2][positon2]ins[xx_seq], where [positon1] and [positon2] represent the positions to be mutated, [xx1] represents the identifier of the original amino acid at [positon1], [xx2] represents the identifier of the original amino acid at [positon2], [xx_seq] represents the identifier of multiple target amino acids to be inserted between [positon1] and [positon2], and ins is the insertion identifier. This mutation instruction indicates the insertion of multiple target amino acids [xx_seq] between [positon1] and [positon2].

[0091] For example, for the A chain of the GB1 protein sequence mentioned above, the mutation instruction is V[A39]_D[A40]insAA. This mutation instruction indicates that multiple amino acids AA will be inserted from valine (V) at position 39 to aspartic acid (D) at position 40. Here, amino acid AA represents alanine A and alanine A.

[0092] In some implementations, the mutation method can be a substitution mutation. In this case, the mutation method can include one of the following: single-point mutation, saturation mutation, and combination mutation. The second mutation instruction is generated based on one of the single-point mutation, saturation mutation, and combination mutation.

[0093] Single-point mutation refers to replacing each mutated position in each mutation request with all other amino acids in turn to obtain a mutant protein sequence in which only a single mutated position is replaced. For example, single-point mutation at position 40 of the GB1 protein sequence means replacing the aspartic acid D at position 40 of the GB1 protein sequence with the remaining 19 amino acids in turn.

[0094] Specifically, when a single-point mutation is performed on the original amino acid aspartic acid D at position 40 of the A chain of GB1 above, the resulting mutation instructions are as follows: D[A40]T, D[A40]S, D[A40]E, D[A40]N, D[A40]H, D[A40]K, D[A40]A, D[A40]V, D[A40]Q, D[A40]I, D[A40]Y, D[A40]L, D[A40]F, D[A40]M, D[A40]R, D[A40]G, D[A40]W, D[A40]C, D[A40]P, that is, the aspartic acid D at position 40 is replaced sequentially by all the other 19 amino acids.

[0095] For example, one mutation instruction, "D[A40]T," indicates replacing the aspartic acid (D) at position 40 of the A chain with a threonine (T). Mutating based on this instruction yields the mutant protein sequence MTYKLILNGKTLKGETTTEAVDAATAEKVFKQYANDNGVTGEWTYDDATKTFTVTE. Mutating based on the other instructions mentioned above can also yield corresponding mutant protein sequences, which will not be elaborated upon here.

[0096] Saturation mutation refers to replacing all the mutated positions in each mutation request with all other amino acids in sequence, resulting in a mutant protein sequence where all mutated positions are replaced. For example, saturation mutation targeting positions 39 and 40 of the GB1 protein sequence means replacing valine V at position 39 with the remaining 19 amino acids in sequence, and simultaneously replacing aspartic acid at position 40 with the remaining 19 amino acids in sequence.

[0097] Specifically, when saturation mutations are performed on the original amino acids valine (V) and aspartic acid (D) at positions 39 and 40 of the A chain of GB1, the generated partial mutation instructions are listed below: V

[0039] ID

[0040] T, V

[0039] ID

[0040] S, V

[0039] ID

[0040] N, V

[0039] ID

[0040] E, V

[0039] AD

[0040] E, V

[0039] LD

[0040] C, that is, at positions 39 and 40, valine (V) and aspartic acid (D) are replaced sequentially with all 19 other amino acids. For example, the mutation instruction V

[0039] ID

[0040] T means that valine (V) at position 39 of the A chain of GB1 is replaced with isoleucine (I), and aspartic acid (D) at position 40 is replaced with threonine (T). Other amino acid substitutions will not be listed one by one.

[0098] Combinatorial mutation refers to the arbitrary combination of first mutation instructions with different selected mutation positions in each mutation task request to replace corresponding positions in the original protein sequence. Here, "corresponding position" refers to the combination of mutation positions corresponding to the combined instructions obtained by arbitrarily combining first mutation instructions with different mutation positions. For example, for the A strand of the GB1 protein sequence mentioned above, assuming the selected first mutation instructions are V[A39]I, D[A40]T, K[A50]M, K[A50]N, then the arbitrary combination of these four first mutation instructions is as follows: V[A39]ID[A40]T (In this case, the "corresponding position" replacing the corresponding position in the original protein sequence is the 39th and 40th positions, which will not be elaborated further in the following examples).

[0099] V[A39]IK[A50]M

[0100] V[A39]IK[A50]N

[0101] D[A40]TK[A50]M

[0102] D[A40]TK[A50]N

[0103] V[A39]ID[A40]TK[A50]M

[0104] V[A39]ID[A40]TK[A50]N

[0105] It is important to note that "arbitrary combinations can be made between first mutation instructions with different selected mutation positions". That is, when the mutation method is a combination mutation, the combination mutation can only be arbitrarily combined with first mutation instructions with different mutation positions, but cannot be combined with first mutation instructions with the same mutation position. This is because an amino acid at the same position can only be replaced by another amino acid at a time. Therefore, for the above first mutation instructions K[A50]M and K[A50]N, a mutation combination cannot be formed, but they can be combined with other different mutation positions.

[0106] After the first mutation instructions at different mutation sites are combined, the corresponding positions in the original protein sequence are replaced according to the combined mutation instructions. For example, after sequentially replacing the original protein sequence of the A chain of GB1 with the combined mutation instructions V[A39]ID[A40]TK[A50]M, the resulting mutant protein sequence is as follows:

[0107] It should be noted that when the mutation method is a single-point mutation or a saturation mutation, each third mutation instruction can be generated by any combination of any first mutation instruction and any second mutation instruction; when the mutation method is a combined mutation, each third mutation instruction is generated by any combination of first mutation instructions with different selected mutation positions.

[0108] When the number of times the mutation task request is received is equal to 1, the mutation task received in step S1 above can be input in the first input interface 200 shown in Figure 2. Referring to Figure 2, the first input interface 200 can have multiple first input areas: first input area 11, first input area 12, first input area 13, first input area 14, first input area 15, first input area 16, first input area 17, first input area 18, etc.

[0109] The first input region 11 is used to input the mutation method, which can include one of the following: single-point mutation, saturation mutation, or combined mutation, as described above. For example, the mutation method input in the first input region 11 in Figure 2 is "new single-point mutation search", i.e., single-point mutation.

[0110] The first input area 12 is used to input the structural data of the original protein, that is, the three-dimensional structure of the protein to be optimized (the original protein) in the first input area 12 of Figure 2. The first input area 13 is used to input the chain number corresponding to the original protein in the three-dimensional structure. For example, when inputting the structural data of the original protein, i.e., the three-dimensional structure of the protein to be optimized, in the second input area of ​​Figure 2, you can input the structural data of the original protein by clicking "Select (*.pdb) file", or you can input the structural data of the original protein previously entered by clicking "Select (.pdb) file from resources". For example, the first input area 12 in Figure 2 is the structural data of the original protein sequence of GB1, and the first input area 13 is the A chain of GB1. After inputting the structural data of the original protein sequence of GB1, the protein structural data of the A chain of GB1 is also displayed in area 121.

[0111] Optionally, in one possible implementation, some buttons can also be set in region 121, for example, three buttons (not shown): button 1, button 2, and button 3. Button 1 indicates that the selected chain is set as the "chain number corresponding to the protein to be optimized (original protein) in the three-dimensional structure". When button 1 is clicked, the corresponding chain number input in the first input region 13 is updated synchronously. Button 2 indicates that the selected chain sequence is set as the "protein sequence to be optimized" and the chain sequences in the first input regions 14 to 16 below are updated synchronously. Button 3 indicates that the selected position is set as the position to be mutated and the position to be mutated in the first input region 14 below is updated synchronously.

[0112] The first input area 14 is used to input the mutation location. For example, in Figure 2, the mutation location input in the first input area 14 is "40". The mutation location can be input by positive or negative selection. The first input area 15 is used to select at least one scoring position group on the original protein sequence. The first input area 16 is used to input the original protein sequence, i.e., the protein sequence to be optimized. The first input area 17 is used to input the protein sequence corresponding to the specified good first mutation instruction. The first input area 18 is used to input the protein sequence corresponding to the specified bad first mutation instruction.

[0113] Specifically, you can enter multiple good first mutation instructions corresponding to mutant protein sequences by clicking the "+ Add a sequence" button in the lower box of the first input area 17; you can enter multiple bad first mutation instructions corresponding to mutant protein sequences by clicking the "+ Add a sequence" button in the lower box of the first input area 18.

[0114] The first input area 15 can be multiple, for example, there are two first input areas 15 in Figure 2, namely the first input area 15 with the upper left corner number "1" and the first input area 15 with the upper left corner number "2" in Figure 2. For each scoring position group, a scoring segment can be selected in each scoring position group. For example, the scoring segment selected in the first input area 15 with the upper left corner number "1" in Figure 2 is "AVDA" from position 20 to position 23 and "ANDNGVD" from position 34 to position 40. For example, the scoring segment selected in the first input area 15 with the upper left corner number "2" in Figure 2 is "L" at position 12, "DAATA" from position 22 to position 26, and "VDG" from position 39 to position 41.

[0115] Specifically, multiple first input areas 15 can be added by clicking the "+ Add a scoring location" button 151 in the lower square box of the first input area 15. After adding multiple first input areas 15, multiple scoring location groups can be added by selecting positively or negatively. Alternatively, the corresponding first input area 15 can be deleted by clicking the circular button 152 with "-" on the right side of the first input area 15, thereby deleting the corresponding scoring location group.

[0116] When using existing protein language models to score each mutant protein sequence, the scoring position group set in the first input region 15 and the selected scoring fragment on the scoring position group can be combined to score each mutant protein sequence, thereby increasing the accuracy of the scoring and improving the quality of protein recommendations.

[0117] It should be noted that, in the aforementioned first input interface 200, input to multiple first input areas refers to input through manual input, selection from a drop-down list, or input through dragging or clicking, such as using an input device like a mouse, touchpad, or touchscreen, or through gesture control, voice control, motion control, light signals, etc.

[0118] The first input interface 200 can also include a task "submit" button 19. Clicking this button sends a mutation task request. In response to the mutation task request, multiple second mutation instructions are generated based on the mutation method and the location to be mutated. Based on the specified first mutation instruction pair and the generated second mutation instructions, multiple third mutation instructions are generated to mutate the original protein sequence, resulting in a mutant protein sequence corresponding one-to-one with each third mutation instruction. Since no mutant protein sequence is specified in the first input area 17 (i.e., the first mutation instruction is empty), the third mutation instructions are the same as the second mutation instructions at this time.

[0119] Then, the existing protein language model is used to score each mutant protein sequence, output the score result corresponding to each mutant protein sequence, and display the third mutation instruction in a one-to-one correspondence with the corresponding score result of each mutant protein sequence.

[0120] Referring to Figure 3, Figure 3 shows the corresponding display 300 after entering the corresponding content based on multiple first input areas in the first input interface 200 and clicking the "Submit" button. As can be seen from Figure 3, display area 31 in Figure 3 is used to display each third mutation instruction, and display area 32 is used to display the scoring results corresponding to each mutated protein sequence. Since there are two first input areas 15 in the first input interface 200, that is, two scoring position groups are selected on the original protein sequence, display area 32 displays two scoring results for each first mutation instruction.

[0121] Optionally, the display 300 can also display the score result of the original protein sequence. This score result is obtained by scoring the original protein sequence using an existing protein language model. See Figure 3. The score result corresponding to "Native" in display area 31 is obtained by scoring the original protein sequence "MTYKLILNGKTLKGETTTEAVDAATAEKVFKQYANDNGVDGEWTYDDATKTFTVTE" using an existing protein language model, for user reference.

[0122] It should be noted that the scoring results corresponding to each mutant protein sequence displayed in display area 32 can be displayed in the form of scores, or in the form of sorting. The sorting is determined based on the scores. Alternatively, it can be displayed in a combination of scores and sorting. For example, display area 32 in Figure 3 displays scores and sorting in a combination of both.

[0123] In addition, in one possible implementation, the display area 31, besides displaying each third mutation instruction, can also be equipped with a "filter" button. By clicking the "filter" button, the third mutation instructions can be filtered to improve the efficiency of users in finding third mutation instructions when there are multiple third mutation instructions.

[0124] To enhance user convenience and improve protein recommendation efficiency, display area 300 also includes a labeling button for users to mark the corresponding third mutation command. For example, display area 33 in Figure 3 also displays a labeling button. After experimentally verifying the recommended mutant protein sequences, users can label the corresponding mutant protein sequences by clicking the labeling button.

[0125] After the user labels the corresponding mutant protein sequences, assuming the user experimentally verifies the mutant protein sequences corresponding to the mutation commands D[A40]T, D[A40]S, and D[A40]E in Figure 3, they find that the mutant protein sequence corresponding to the mutation command D[A40]T is “MTYKLILNGKTLKGETTTEAVDAATAEKVFKQYANDNGVTGEWTYDDATKTFTVTE”, and the mutant protein sequence corresponding to the mutation command D[A40]S is “MTYKLILNGKTLKGE”. The three mutant protein sequences “TTTEAVDAATAEKVFKQYANDNGVSGEWTYDDATKTFTVTE”, corresponding to the mutation instruction D[A40]E, are good mutant protein sequences. These sequences can be marked as good mutation instructions by clicking the corresponding marking buttons.

[0126] It should be noted that after users conduct experimental verification of mutant protein sequences, if they find any bad mutant protein sequences, they can also mark the bad mutant protein sequences. If users mark bad mutant protein sequences, after the next task submission, the existing protein language model will be used to score the bad mutant protein sequences and the score results will be displayed in the next corresponding display for users' reference.

[0127] After labeling, based on the above mutation task, for example, a mutation task can be created in the second input interface 400. See Figure 4. The second input interface 400 can have multiple second input areas: second input area 41, second input area 42, second input area 43, second input area 44, second input area 45, second input area 46, second input area 47, second input area 48, etc.

[0128] The second input area 41 is used to input the mutation method, which can include one of the following: single-point mutation, saturation mutation, or combined mutation. For example, the mutation method input in the second input area 41 in Figure 4 can still be "new single-point mutation search", i.e., single-point mutation, but it can also be other mutation methods.

[0129] The second input area 42 is used to display the structural data of the original protein obtained based on the mutation task input in the first input interface 200, i.e., the three-dimensional structure of the protein to be optimized (original protein) in the second input area 42 of Figure 4. The second input area 43 is used to display the chain number corresponding to the original protein obtained based on the mutation task in the three-dimensional structure. For example, the second input area 42 in Figure 4 displays the structural data of the original protein sequence of GB1 obtained based on the mutation task, the second input area 43 displays the A chain of GB1 obtained based on the mutation task, and the protein structural data of the A chain of GB1 obtained based on the mutation task is displayed in area 421.

[0130] Optionally, in one possible implementation, some buttons may also be provided in area 421. The way the buttons are set in area 421 and their functions are the same as those of the buttons set in area 121 as described above, and will not be repeated here.

[0131] It should be noted that since the mutation task in the second input interface 400 is created based on the mutation task input in the first input interface 200, the content in the corresponding second input area of ​​the second input interface 400 is the same as the corresponding content in the first input area when the first task was created. However, in the second input interface 400, the user can modify the corresponding content in multiple second input areas according to the specific application scenario. That is, in the second input interface 400, any one or more of the mutation method, chain number, mutation position, and scoring position group in the previous mutation task can be changed. Therefore, the content in multiple second input areas can be the same as or different from the content input in the first input area. In specific applications, the user can modify the content in the corresponding second input areas according to the specific application scenario.

[0132] The second input area 44 is used to display the mutation position obtained based on the above mutation task. At this time, the mutation position in the second input area 44 is the same as the mutation position in the first input area 14, which is position "40". However, the mutation position can be modified in the second input interface 400. For example, in Figure 4, the mutation position "40" displayed in the second input area 44 is modified to mutation position "41".

[0133] The second input area 45 is used to display two score position groups obtained based on the above mutation task. At this time, the score position groups in the second input area 45 are the same as the two score position groups in the first input area 15. That is, the score segments under one score position group are “AVDA” from position 20 to position 23 and “ANDNGVD” from position 34 to position 40. The other score segment is “L” at position 12, “DAATA” from position 22 to position 26 and “VDG” from position 39 to position 41. However, the two score position groups obtained based on the above mutation task can be modified in the second input interface 400. For example, in Figure 4, the other score position group is deleted, and only one score position group is kept. Specifically, the corresponding score position group can be deleted by clicking the circular button 452 with “-” on the right side of the second input area 45. Of course, depending on the specific application scenario, the rating segment under the rating position group can also be modified in the second input interface 400, and new rating position groups can also be added in the second input interface 400. These are all feasible.

[0134] It should be noted that if at least one second input region 45 is set in the second input interface 400, when using the existing protein language model to score each mutant protein sequence, the scoring position group set in the second input region 5 and the selected scoring fragment on the scoring position group can be combined to score each mutant protein sequence, so as to increase the accuracy of the scoring and thus improve the quality of protein recommendation.

[0135] Furthermore, when modifying the content of multiple second input areas in the aforementioned second input interface 400, it refers to modifying the input through manual input, selection of input from a drop-down list, or by dragging or clicking, such as using an input device like a mouse, touchpad, or touchscreen, or by using gesture control, voice control, motion control, light signals, etc. to achieve the click behavior.

[0136] The second input region 46 is used to input the original protein sequence, i.e. the protein sequence to be optimized; the second input region 47 is used to input known good mutation information; and the second input region 48 is used to input known bad mutation information.

[0137] It should be noted that the multiple second input interfaces also include a second input area 47 for displaying the mutant protein sequence corresponding to the mutation command marked by the user. Referring to Figure 4, in the second input area 47, the second input area 47 marked with "1" in the upper left corner displays the mutant protein sequence "MTYKLILNGKTLKGETTTEAVDAATAEKVFKQYANDNGVTGEWTYDDATKTFTVTE" corresponding to the mutation command D[A40]T; the second input area 47 marked with "2" in the upper left corner displays the mutant protein sequence "MTYKLILNGKTLKGETTTEAVDAATAEKVFKQYANDNGVSGEWTYDDATKTFTVTE" corresponding to the mutation command D[A40]S; and the second input area 47 marked with "3" in the upper left corner displays the mutant protein sequence "MTYKLILNGKTLKGETTTEAVDAATAEKVFKQYANDNGVEGEWTYDDATKTFTVTE" corresponding to the mutation command D[A40]E.

[0138] When it is necessary to delete the mutant protein sequences corresponding to multiple good mutation instructions specified in the second input area 47, the corresponding mutant protein sequences can be deleted by clicking the circular button 472 with "-" on the right side of the second input area 47. When it is necessary to add mutant protein sequences corresponding to specified good mutation instructions, the protein sequences corresponding to multiple good first mutation instructions can be entered by clicking the "+ Add a sequence" button 471 in the lower box of the second input area 47. Alternatively, the protein sequences corresponding to multiple bad mutation instructions can be entered by clicking the "+ Add a sequence" button in the lower box of the first input area 18. When multiple bad mutant protein sequences are specified, the deletion method is similar to that for deleting good mutant protein sequences, and will not be described in detail here.

[0139] The second input interface 400 can also be equipped with a task "submit" button 49. By clicking this button, a mutation task request is sent. Then, in response to the mutation task request, multiple second mutation instructions are generated according to the mutation method and mutation location. Multiple third mutation instructions are generated based on the specified first mutation instruction pair and the second mutation instructions. The original protein sequence is mutated to obtain a mutant protein sequence that corresponds one-to-one with each third mutation instruction.

[0140] It should be noted that, since the mutation instructions D[A40]T, D[A40]S, and D[A40]E are marked in Figure 3, the first mutation instruction is not empty, but rather D[A40]T, D[A40]S, and D[A40]E. Then, the good mutant protein sequences corresponding to these three mutation instructions are displayed in the second input area 47. After clicking the "Submit" button 49, based on the mutation method "single-point mutation" and the mutation location "4" entered in the second input interface 400, 1” generates 19 second mutation instructions, as follows: G[A41]A, G[A41]Y, G[A41]S, G[A41]N, G[A41]F, G[A41]T, G[A41]K, G[A41]L, G[A41]V, G[A41]D, G[A41]R, G[A41]H, G[A41]W, G[A41]I, G[A41]E, G[A41]C, G[A41]P, G[A41]M, G[A41]Q.

[0141] Then, based on the first mutation instructions D[A40]T, D[A40]S, and D[A40]E specified by the marker, and 19 second mutation instructions, 57 third mutation instructions are generated. In fact, the three first mutation instructions are combined with the 19 second mutation instructions respectively to generate each third mutation instruction. That is, each third mutation instruction is generated by combining any first mutation instruction and any second mutation instruction. Some of the generated third mutation instructions are as follows: D[A40]TG[A41]A, D[A40]TG[A41]Y, D[A40]TG[A41]S, D[A40]SG[A41]A, D[A40]SG[A41]Y, D[A40]SG[A41]S, D[A40]EG[A41]A, D[A40]EG[A41]Y, D[A40]EG[A41]S.

[0142] Based on the above 57 third mutation instructions, the original protein sequence is mutated to obtain a mutant protein sequence that corresponds one-to-one with each third mutation instruction. After the user performs experimental verification on the mutant protein sequence, the third mutation instructions can be marked again to initiate the next mutation task request. The marked third mutation instruction is the first mutation instruction specified in the next received mutation task request.

[0143] After mutating the original protein sequence based on the above 57 third mutation instructions, the existing protein language model is used to score each mutated protein sequence, output the score result corresponding to each mutated protein sequence, and display the correspondence between the third mutation instructions and the corresponding score result of each mutated protein sequence.

[0144] Referring to Figure 5, Figure 5 shows the corresponding display 500 after entering the corresponding content based on multiple second input areas in the second input interface 400 and clicking the "Submit" button. As can be seen from Figure 5, display area 51 in Figure 5 is used to display each third mutation instruction, and display area 52 is used to display the score result corresponding to each mutant protein sequence. Since at least one score position group has been deleted in the second input interface 400, only one score result is displayed in display area 52 for each third mutation instruction.

[0145] Referring to Figure 5, D[A40]T is a known good mutation instruction. Based on this mutant protein sequence, a single-point mutation is performed at position 41, resulting in the second mutation instruction G[A41]A. The specified first mutation instruction D[A40]T and the generated second mutation instruction G[A41]A are combined to generate the third mutation instruction D[A40]TG[A41]A.

[0146] Optionally, the corresponding display 500 can also display the scoring results of the mutant protein sequence specified by the user. The scoring results can be obtained by scoring the specified mutant protein sequence again using the existing protein language model, or they can be obtained from the scoring results obtained from the previous scoring of the mutant protein sequence. See Figure 5. The scoring results corresponding to "D[A40]T, D[A40]S, D[A40]S" in display area 51 are obtained by scoring the specified corresponding mutant protein sequence using the existing protein language model, for the user's reference.

[0147] Optionally, the display of 500 can also display the score result of the original protein sequence. This score result is obtained by scoring the original protein sequence using an existing protein language model. See Figure 5. The score result corresponding to "Native" in display area 51 is obtained by scoring the original protein sequence using an existing protein language model, for user reference.

[0148] Optionally, the mutation task request may also include a fourth mutation instruction. The fourth mutation instruction is a specified bad mutation instruction selected from the previous third mutation instruction. When the number of mutation task requests received is greater than 1, the corresponding score result of the mutant protein sequence corresponding to the fourth mutation instruction will also be displayed. The score result is obtained by scoring the mutant protein sequence corresponding to the specified fourth mutation instruction. In this way, users can be provided with an additional reference indicator for selecting mutations for subsequent experiments.

[0149] It should be noted that the scoring results corresponding to each mutant protein sequence displayed in display area 52 can be displayed in the form of scores, or in the form of sorting, with the sorting determined based on the scores. Alternatively, it can be displayed in a combination of scores and sorting, as shown in display area 52 in Figure 5, which is a combination of scores and sorting.

[0150] In addition, in one possible implementation, the display area 51 can be equipped with a "filter" button in addition to displaying each third mutation instruction. By clicking the "filter" button, the third mutation instructions can be filtered to improve the efficiency of users in finding third mutation instructions when there are multiple third mutation instructions.

[0151] To make the operation more convenient for users and improve the efficiency of protein recommendation, the corresponding display 500 can also display a marking button for users to mark the corresponding third mutation command. For example, in Figure 5, the display area 53 corresponding to display 500 also displays a marking button. After the user conducts experimental verification of the recommended third mutant protein sequence, they can mark the corresponding mutant protein sequence by clicking the marking button.

[0152] It should be noted that after users conduct experimental verification of mutant protein sequences, they can label good mutant protein sequences or bad mutant protein sequences.

[0153] After labeling, based on the above mutation tasks, mutation tasks can continue to be created in the new second input interface until the user finds the optimal mutant protein sequence.

[0154] As can be seen from the above, by specifying the first mutation instruction in the mutation task, and generating a third mutation instruction based on the specified first mutation instruction and the newly generated second mutation instruction, the original protein sequence is mutated. Finally, the corresponding score results of the mutated protein sequence are displayed one-to-one with the third mutation instruction for users to refer to and verify experimentally. After the user performs experimental verification, the first mutation instruction is specified again in the mutation task request, thus forming a cyclical operation until the user finds the optimal protein sequence. Therefore, users do not need to manually input the experimentally verified protein sequence; they only need to specify the mutation instruction in the mutation task. The operation is convenient and can improve the screening efficiency of protein mutations.

[0155] Furthermore, the scoring results of the mutant protein sequences are displayed one-to-one with the third mutation command. Through this correspondence display, such as the aforementioned correspondence display 300 and the aforementioned correspondence display 500, the scoring results of each mutant protein sequence are visualized and presented intuitively, which can facilitate users to quickly select mutant protein sequences for experimental verification. Furthermore, this correspondence display also displays the scoring results of the original protein sequence, such as the scoring results corresponding to "Native" in display area 31. When the number of received mutation task requests is greater than 1, this correspondence display will also display the scoring results of the mutant protein sequences corresponding to the first mutation command and the first mutation command, such as "D[A40]T, D[A40]S, D[A40]S" and their corresponding scoring results in display area 51, so as to provide users with a more intuitive reference when making subsequent experimental selections.

[0156] Furthermore, the corresponding display also shows a marking button for the user to mark the corresponding first mutation instruction, such as the marking button 33 and the marking button 53 mentioned above. The user only needs to click the marking button to mark the corresponding first mutation instruction. The operation is intelligent and does not require the user to do much work, saving time and effort.

[0157] Furthermore, when the number of mutation task requests received is greater than 1, the mutation task is automatically created based on the previous mutation task in the second input interface, such as the second input interface 400 mentioned above. When it is necessary to modify the corresponding items, the mutation task can be modified in the second input interface and then submitted. There is no need to manually re-enter all the data again. The operation is simple, time-saving and labor-saving, thereby improving the efficiency of protein sequence screening.

[0158] In one possible implementation, when the number of received mutation task requests is equal to 1, if the location to be mutated is omitted, it means that all positions of the original protein sequence are mutated. For example, in the first input area 14 of the first input interface 200 mentioned above, if the location to be mutated is omitted, it means that all positions of the protein sequence of chain A of GB1 are mutated.

[0159] When the number of received mutation task requests is greater than 1, if the mutation location is omitted, it means that the current mutation location is a location other than the mutation location corresponding to the first mutation instruction specified in the previous task. For example, in the second input area 44 of the second input interface 400, the current mutation location is the mutation location "40" of the task created in the first input interface 200 in the previous task. Since the mutation location "40" has already been mutated when the task was created in the first input interface 200, and since the mutation location corresponding to the first mutation instructions D[A40]T, D[A40]S, and D[A40]E specified in the previous task is also "40", in this case, the current mutation is performed on all locations other than "40", that is, all locations of the protein sequence of the A chain of GB1 except "40" are mutated.

[0160] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0161] Based on the same inventive concept, this application also provides an apparatus for recommending protein mutations to implement the aforementioned method for recommending protein mutations. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations of one or more embodiments of the apparatus for recommending protein mutations provided below can be found in the limitations of the method for recommending protein mutations described above, and will not be repeated here.

[0162] In one embodiment, as shown in FIG6, an apparatus 600 for recommending protein mutations is provided, including: a receiving module 601, a response module 602, a mutation module 603, a scoring display module 604, and a judgment module 605.

[0163] The receiving module 601 is used to receive a mutation task request, wherein the mutation task request carries the original protein sequence, the position to be mutated, the mutation method, and the specified first mutation instruction.

[0164] Response module 602, in response to the mutation task request, instructs mutation module 603 to execute:

[0165] Based on the mutation method and the location to be mutated, a plurality of second mutation instructions corresponding to the location to be mutated are generated; based on the first mutation instructions and the second mutation instructions, a plurality of third mutation instructions are generated; the original protein sequence is mutated to obtain a mutant protein sequence corresponding one-to-one with each of the third mutation instructions; and

[0166] Instruct the scoring display module 604 to execute:

[0167] Each mutant protein sequence is scored at least once and a score result is obtained. The score result corresponding to each mutant protein sequence is then displayed in a one-to-one correspondence with the third mutation instruction.

[0168] The judgment module 605 is used to determine whether a new mutation task request has been received. If so, it instructs the receiving module 601 to receive the new mutation task request.

[0169] When the number of times the mutation task request is received is greater than 1, the first mutation instruction selects a specified good mutation instruction from the previous third mutation instruction.

[0170] In one specific embodiment, the scoring display module is further configured to display the scoring result of the original protein sequence, which is obtained by scoring the original protein sequence.

[0171] When the number of times the mutation task request is received is greater than 1, it is also used to display the corresponding score result of the mutant protein sequence corresponding to the first mutation instruction in correspondence with the first mutation instruction. The score result is obtained by scoring the mutant protein sequence corresponding to the specified first mutation instruction.

[0172] Receiving a mutation task request also includes a fourth mutation instruction, which is a specified bad mutation instruction selected from the previous third mutation instruction. When the number of received mutation task requests is greater than 1, the corresponding display will also show the score result of the mutant protein sequence corresponding to the fourth mutation instruction. This score result is obtained by scoring the mutant protein sequence corresponding to the specified fourth mutation instruction.

[0173] In one specific embodiment, the mutation task request information also carries protein structure data corresponding to the original protein sequence and at least one set of scoring positions selected on the original protein sequence, and a scoring fragment is selected on each set of scoring positions.

[0174] In one specific embodiment, the scoring display module is further configured to display a marking button for the user to mark the corresponding first mutation instruction.

[0175] In one specific embodiment, the scoring result is displayed in the form of a score; and / or

[0176] The scores are displayed in a sorted manner, which is determined based on the scores of the ratings.

[0177] In a specific embodiment, the first mutation method is substitution, and the mutation method includes one of the following mutation modes: single-point mutation, saturation mutation, and combined mutation, and the second mutation instruction is generated according to one of the single-point mutation, the saturation mutation, and the combined mutation;

[0178] The single-point mutation refers to replacing each mutated position in each mutation task request with all other amino acids in turn, resulting in a mutant protein sequence in which only a single mutated position is replaced.

[0179] The saturation mutation refers to the process in which all positions to be mutated in each mutation task request are replaced by all other amino acids in sequence to obtain a mutant protein sequence in which all positions to be mutated are replaced.

[0180] The combined mutation refers to the arbitrary combination of the first mutation instructions with different selected mutation positions in each mutation task request, and the sequential replacement of the corresponding positions in the original protein sequence.

[0181] The mutation module is specifically used for:

[0182] When the mutation method is a single-point mutation or a saturation mutation, each third mutation instruction is generated by a combination of any first mutation instruction and any second mutation instruction.

[0183] When the mutation method is a combination mutation, each third mutation instruction is generated by any combination of the first mutation instructions with different selected mutation positions.

[0184] In a specific embodiment, when the number of mutation task requests received by the receiving module is equal to 1, if the position to be mutated is omitted, it means that all positions of the original protein sequence are mutated.

[0185] When the number of mutation task requests received by the receiving module is greater than 1, if the position to be mutated is omitted, it means that the position to be mutated is a mutation position other than the mutation position corresponding to the first mutation instruction specified in the previous task.

[0186] When the number of mutation task requests received by the receiving module is equal to 1, the mutation task is input in the first input interface, which has multiple first input areas; wherein, input can be made by manual input, selection from a drop-down list, or by clicking or dragging.

[0187] The plurality of first input regions are used to respectively input the mutation method, the protein structure data, the chain number corresponding to the original protein sequence in the three-dimensional structure, the position to be mutated, at least one scoring position group selected on the original protein sequence, and the original protein sequence;

[0188] When the receiving module receives more than one mutation task request, the mutation task is created on the second input interface based on the previous mutation task. The second input interface has multiple second input areas:

[0189] The plurality of second input areas are used to respectively display the mutation method in the previous mutation task, the protein structure data, the chain number corresponding to the original protein sequence in the three-dimensional structure, the position to be mutated in the previous mutation task, the selection of at least one scoring position group on the original protein sequence, and the original protein sequence. Furthermore, any one or more of the mutation method, chain number, position to be mutated, and scoring position group in the previous mutation task can be changed in the second input interface.

[0190] The plurality of second input interfaces also include a second input region for displaying the mutant protein sequence corresponding to the first mutation command specified by the user;

[0191] Furthermore, the first input interface and / or the second input interface are provided with a task submission button, by clicking the button to send the mutation task request, and / or, each also includes a first input region or a second input region that displays the three-dimensional structure of the original protein sequence.

[0192] The modules in the aforementioned recommended apparatus for protein mutation can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independent of the processor of a computer device, or stored in software in the memory of the computer device, so that the processor can invoke and execute the operations corresponding to each module.

[0193] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram is shown in Figure 7. The computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The memory stores at least one computer program, which is loaded and executed by the processor to perform the operations performed in the recommended method for protein mutation described in the above embodiment. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is connected to the system bus via the I / O interfaces. The processor of the computer device provides computational and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores raw protein sequences and corresponding protein structure data, etc. The I / O interfaces of the computer device are used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals via a network connection.

[0194] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure may be as shown in Figure 8. The computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a recommended method for protein mutation. The display unit of the computer device is used to form a visually visible image. It can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0195] Those skilled in the art will understand that the structure shown in Figure 8 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or may combine certain components, or may have different component arrangements.

[0196] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for receiving, storing, and displaying) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0197] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0198] This application also provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to perform the operations of the recommended method for protein mutation described above.

[0199] This application also provides a computer program product, including a computer program, characterized in that the computer program is loaded and executed by a processor to perform the operations of the recommended method for protein mutation described in the above embodiments. In some embodiments, the computer program involved in this application may be deployed and executed on a single computer device, or on multiple computer devices located in one location, or on multiple computer devices distributed across multiple locations and interconnected via a communication network. These multiple computer devices distributed across multiple locations and interconnected via a communication network can constitute a blockchain system.

[0200] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0201] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for recommending protein mutations, characterized in that, include: Step S1: Receive a mutation task request, which carries the original protein sequence, the location to be mutated, the mutation method, and the specified first mutation instruction. Step S2: In response to the mutation task request, execute: Step S21: Based on the mutation method and the mutation location, generate multiple second mutation instructions corresponding to the mutation location, generate multiple third mutation instructions based on the first mutation instructions and the second mutation instructions, mutate the original protein sequence, and obtain a mutant protein sequence that corresponds one-to-one with each of the third mutation instructions; Step S22: Score at least each of the mutant protein sequences and obtain the score results. Display the score results corresponding to each mutant protein sequence in a one-to-one correspondence with the third mutation command. Step S3: Determine whether a new mutation task request has been received. If so, return to step S1. When the number of times the mutation task request is received is greater than 1, the first mutation instruction selects a specified good mutation instruction from the previous third mutation instruction.

2. The method as described in claim 1, characterized in that, The corresponding display also displays the score result of the original protein sequence, which is obtained by scoring the original protein sequence. When the number of times the mutation task request is received is greater than 1, the corresponding display will also display the corresponding score result of the mutant protein sequence corresponding to the first mutation instruction in correspondence with the first mutation instruction. The score result is obtained by scoring the mutant protein sequence corresponding to the specified first mutation instruction. Receiving a mutation task request also includes the fourth mutation instruction, which is a specified bad mutation instruction selected from the previous third mutation instruction. When the number of times the mutation task request is received is greater than 1, the corresponding display will also display the corresponding score result of the mutant protein sequence corresponding to the fourth mutation instruction in correspondence with the fourth mutation instruction. The score result is obtained by scoring the mutant protein sequence corresponding to the specified fourth mutation instruction.

3. The method according to any one of claims 1-2, characterized in that, The mutation task request information also carries protein structure data corresponding to the original protein sequence and at least one set of scoring positions selected on the original protein sequence, with a scoring fragment selected on each set of scoring positions.

4. The method according to any one of claims 1-3, characterized in that, The corresponding display also includes a marker button for the user to mark the corresponding first mutation instruction.

5. The method according to any one of claims 1-4, characterized in that, The scoring results are displayed in the form of scores; and / or The scores are displayed in a sorted manner, which is determined based on the scores of the ratings.

6. The method according to any one of claims 1-5, characterized in that, The mutation method is substitution, and the mutation method includes one of the following mutation modes: single-point mutation, saturation mutation, and combined mutation. The second mutation instruction is generated based on one of the single-point mutation, the saturation mutation, and the combined mutation. The single-point mutation refers to replacing each of the mutated positions in each mutation task request with all other amino acids in turn to obtain a mutant protein sequence in which only a single mutated position is replaced. The saturation mutation refers to replacing all the mutated positions in each mutation task request with all other amino acids in turn to obtain a mutant protein sequence in which all the mutated positions are replaced. The combined mutation refers to the arbitrary combination of the first mutation instructions with different selected mutation positions in each mutation task request, and the sequential replacement of the corresponding positions in the original protein sequence. The generation of multiple third mutation instructions based on the first mutation instruction and the second mutation instruction includes: When the mutation method is the single-point mutation or the saturation mutation, each third mutation instruction is generated by a combination of any first mutation instruction and any second mutation instruction. When the mutation method is the combined mutation, each of the third mutation instructions is generated by any combination of the first mutation instructions with different selected mutation positions.

7. The method according to any one of claims 1-6, characterized in that, When the number of times the mutation task request is received is equal to 1, if the position to be mutated is omitted, it means that all positions of the original protein sequence are mutated. When the number of times the mutation task request is received is greater than 1, if the position to be mutated is omitted, it means that the position to be mutated is a mutation position other than the mutation position corresponding to the first mutation instruction specified in the previous task.

8. The method according to any one of claims 1-7, characterized in that, When the number of times the mutation task request is received is equal to 1, the mutation task is input in the first input interface, which has multiple first input areas; wherein, input is performed by manual input, selection from a drop-down list, or input by clicking or dragging; The plurality of first input regions are used to respectively input the mutation method, the protein structure data, the chain number corresponding to the original protein sequence in the three-dimensional structure, the position to be mutated, at least one scoring position group selected on the original protein sequence, and the original protein sequence; When the number of times the mutation task request is received is greater than 1, the mutation task is created based on the previous mutation task on the second input interface, which has multiple second input areas: The plurality of second input areas are used to respectively display the mutation method in the previous mutation task, the protein structure data, the chain number corresponding to the original protein sequence in the three-dimensional structure, the position to be mutated in the previous mutation task, the selection of at least one scoring position group on the original protein sequence, and the original protein sequence. Furthermore, any one or more of the mutation method, chain number, position to be mutated, and scoring position group in the previous mutation task can be changed in the second input interface. The plurality of second input interfaces also include a second input region for displaying the mutant protein sequence corresponding to the first mutation command specified by the user; Furthermore, the first input interface and / or the second input interface are provided with a task submission button, by clicking the button to send the mutation task request, and / or, each also includes a first input region or a second input region that displays the three-dimensional structure of the original protein sequence.

9. An apparatus for recommending protein mutations, comprising: The system includes a receiving module, a response module, a mutation module, a scoring display module, and a judgment module. The receiving module is used to receive a mutation task request, which carries the original protein sequence, the location to be mutated, the mutation method, and a specified first mutation instruction. The response module is configured to, in response to the mutation task request, instruct the mutation module to execute: Based on the mutation method and the mutation site to be mutated, a plurality of second mutation instructions corresponding to the mutation site are generated, and a plurality of third mutation instructions are generated based on the first mutation instructions and the second mutation instructions. The original protein sequence is mutated to obtain a mutant protein sequence that corresponds one-to-one with each of the third mutation instructions. as well as Instruct the scoring display module to perform: Each mutant protein sequence is scored at least once and a score result is obtained. The score result corresponding to each mutant protein sequence is then displayed in a one-to-one correspondence with the third mutation instruction. The judgment module is used to determine whether a new mutation task request has been received. If so, it instructs the receiving module to receive the new mutation task request. When the number of times the mutation task request is received is greater than 1, the first mutation instruction selects a specified good mutation instruction from the previous third mutation instruction.

10. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one computer program, which is loaded and executed by the processor to perform the operations of the recommended method for protein mutation as described in any one of claims 1 to 9.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to perform the operations of the recommended method for protein mutation as described in any one of claims 1 to 9.

12. A computer program product, comprising a computer program, characterized in that, The computer program is loaded and executed by a processor to perform the operations carried out by the recommended method for protein mutation as described in any one of claims 1 to 9.