Microfilament binding protein sequence design method and device

The microfilament-binding protein sequence design method, which utilizes a multi-task U-Net architecture and dynamic weight training, addresses the issues of large blind spots and low success rates in microfilament-binding protein design under small sample conditions, achieving efficient and precise residue-level design.

CN121034418APending Publication Date: 2025-11-28杭州市临安区青山湖未来产业科创中心
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510938576.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing technologies for designing microfilament-binding protein sequences, especially under small sample conditions, have shortcomings in functional site prediction and residue-level annotation, resulting in large design blind spots and low success rates.

Method used

We employ a multi-task U-Net architecture combined with dynamic weight training and homology annotation transfer technology. Through simulated annealing and Monte Carlo methods, combined with a scoring function, we optimize amino acid sequences to achieve efficient global search.

Benefits of technology

It enables precise residue-level design and high-success-rate generation of microfilament-binding proteins under small sample conditions, improving the accuracy and generalization performance of the model in microfilament-binding protein identification and design tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005488725760000061
    Figure BDA0005488725760000061
  • Figure BDA0005488725760000071
    Figure BDA0005488725760000071
  • Figure BDA0005488725760000081
    Figure BDA0005488725760000081
Patent Text Reader

Abstract

The invention provides a microfilament binding protein sequence design method and device. The method comprises the following steps: randomly generating an initial amino acid sequence; carrying out random mutation on single amino acid by starting from the initial amino acid sequence to obtain a mutated amino acid sequence, and scoring the amino acid sequences before and after mutation according to a scoring function; and selecting the amino acid sequence of which the score meets a preset condition to carry out next round of random mutation until a preset number of iterations is reached, thereby obtaining a design sequence. According to the method, the microfilament binding protein sequence is accurately designed under the small sample condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of protein design technology, specifically to a method and apparatus for designing microfilament-binding protein sequences. Background Technology

[0002] In modern bioinformatics research, functional annotation and structural prediction of biomolecule sequences have become core research directions, widely applied in disease mechanism research, new drug target discovery, synthetic biology design, and many other fields. Particularly in the study of interactions between proteins and cytoskeleton structures such as microfilaments, accurate identification of key binding residues is crucial for revealing physiological processes such as cell movement, division, and morphological changes. In recent years, deep learning methods have been widely applied in this field and have achieved significant progress. However, existing techniques still have significant limitations in functional site prediction, especially in residue-level annotation under small sample conditions. Summary of the Invention

[0003] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application provides a method and apparatus for designing microfilament-binding protein sequences. The method of this application, through a multi-task U-Net architecture combined with dynamic weight training and homology annotation transfer technology, can achieve precise residue-level design and high-success-rate generation of microfilament-binding proteins with only a small sample size.

[0004] Specifically, the technical solution of this application is as follows:

[0005] In a first aspect, this application proposes a method for designing microfilament-binding protein sequences. According to embodiments of this application, the method includes: randomly generating an initial amino acid sequence; randomly mutating individual amino acids starting from the initial amino acid sequence to obtain mutated amino acid sequences; scoring the amino acid sequences before and after mutation according to a scoring function; selecting amino acid sequences whose scores meet preset conditions for the next round of random mutation, until a preset number of iterations is reached to obtain the designed sequence.

[0006] Traditional protein design methods (such as ROSETTA and directed evolution) rely on known structural templates or experimental screening, resulting in high computational costs, large design blind spots, and low success rates. The method presented in this application utilizes simulated annealing and Monte Carlo methods, guided by a scoring function, to progressively optimize the amino acid sequence, making it more consistent with the characteristics of microfilament-binding proteins, thus achieving efficient global search.

[0007] In a second aspect, this application proposes a microfilament-binding protein sequence design device. According to an embodiment of this application, the device includes: an initialization module for randomly generating an initial amino acid sequence; a mutation module for randomly mutating individual amino acids starting from the initial amino acid sequence to obtain a mutated amino acid sequence, and scoring the amino acid sequences before and after mutation according to a scoring function; and an iteration module for selecting amino acid sequences whose scores meet preset conditions for the next round of random mutation, until a preset number of iterations is reached to obtain the designed sequence.

[0008] The description of the beneficial effects in the first aspect also applies to this aspect, and will not be repeated here.

[0009] In a third aspect, this application proposes an electronic device. According to an embodiment of this application, the device includes: a processor and a memory; the memory for storing a computer program; and the processor for executing the computer program to implement the method as described in the first aspect.

[0010] In a fourth aspect, this application provides a computer-readable storage medium. According to an embodiment of this application, the computer-readable storage medium stores computer instructions or programs that, when executed on a computer, cause the method described in the first aspect to be performed.

[0011] In a fifth aspect, this application provides a computer program product. According to an embodiment of this application, the computer program product includes computer instructions that, when some or all of the computer instructions are run on a computer, cause the method described in the first aspect to be performed.

[0012] The aforementioned electronic devices, computer-readable storage media, and computer program products achieve high efficiency and automation through the automatic execution of computer instructions for microfilament-binding protein sequence design. Furthermore, the instruction-based nature of these devices results in better stability in various environments.

[0013] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0014] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0015] Figure 1 This is a schematic flowchart of a microfilament-binding protein sequence design method provided in one embodiment of this application;

[0016] Figure 2A flowchart of protein generation provided in one embodiment of this application;

[0017] Figure 3 This is a schematic diagram of a data preprocessing process provided in one embodiment of this application;

[0018] Figure 4 A predictive model architecture diagram provided in one embodiment of this application;

[0019] Figure 5 A schematic diagram of a microfilament-binding protein sequence design device provided in one embodiment of this application;

[0020] Figure 6 A schematic diagram of an electronic device provided in one embodiment of this application;

[0021] Figure 7 A graph showing the predicted and generated microfilament-binding protein results provided for one embodiment of this application. Detailed Implementation

[0022] The embodiments of this application are described in detail below. The embodiments described below are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0023] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more.

[0024] The endpoints and any values ​​of the ranges disclosed herein are not limited to the precise ranges or values, and these ranges or values ​​should be understood to include values ​​close to these ranges or values. For numerical ranges, the endpoint values ​​of the various ranges, the endpoint values ​​of the various ranges and individual point values, and individual point values ​​can be combined with each other to obtain one or more new numerical ranges, which should be considered as specifically disclosed herein.

[0025] In this document, the terms “comprising” or “including” are open-ended expressions, meaning that they include the contents specified in this application but do not exclude other contents.

[0026] In this document, the terms “optionally,” “optionally,” or “optionally” generally refer to an event or condition that may, but may not, occur, and the description includes both cases in which the event or condition occurs and cases in which the event or condition does not occur.

[0027] In this document, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented, wholly or partially, using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of a larger module or unit that contains the functionality of that module or unit.

[0028] In this paper, the term "model" refers to a computational model or algorithm that can automatically perform tasks such as prediction, classification, recognition, or decision-making by learning and analyzing input data. The model's learning process is based on statistical principles and data pattern recognition, using a training dataset to tune model parameters and optimize the model to improve its predictive or inference capabilities. Models can employ various algorithms and techniques, such as neural networks, support vector machines, decision trees, random forests, and deep learning. These models can be trained and optimized through supervised learning, unsupervised learning, or reinforcement learning. In some examples of this application, machine learning is applied to the field of protein design technology, specifically for the design of microfilament-binding protein sequences.

[0029] In one aspect of this application, a method for designing microfilament-binding protein sequences is proposed, referring to... Figure 1 The method includes: S110 randomly generating an initial amino acid sequence;

[0030] According to some specific embodiments of this application, an initial amino acid sequence of fixed length is randomly generated.

[0031] S120 starts with the initial amino acid sequence and randomly mutates a single amino acid to obtain the mutated amino acid sequence. The amino acid sequences before and after the mutation are scored according to a scoring function.

[0032] S130 selects the amino acid sequence whose score meets the preset conditions for the next round of random mutation until the preset number of iterations is reached to obtain the designed sequence.

[0033] According to an embodiment of this application, the preset condition is: if the score corresponding to the randomly mutated amino acid sequence is greater than the score corresponding to the amino acid sequence before the random mutation, then the current randomly mutated amino acid sequence is accepted; if the score corresponding to the randomly mutated amino acid sequence is less than the score corresponding to the amino acid sequence before the random mutation, then the acceptance or rejection of the current randomly mutated amino acid sequence is determined according to the acceptance probability.

[0034] According to an embodiment of this application, the acceptance probability is: Where x is the amino acid sequence before mutation, x′ is the amino acid sequence after mutation, T is the temperature parameter, and E(x) is the scoring function.

[0035] For example, refer to Figure 2 For a sequence with an initial length of 100 amino acids, the process uses 170,000 Monte Carlo steps with an initial temperature parameter T = 8. The temperature parameter is halved every 10,000 cycles, and 170,000 optimization iterations are performed to finally make the obtained sequence converge within the optimal range of the scoring function.

[0036] According to an embodiment of this application, the scoring function consists of sequence scoring, semantic scoring, and binding score; the sequence scoring is calculated based on the KL divergence between the n-gram distribution of the designed sequence and the n-gram distribution of the natural protein library; the semantic scoring is calculated based on the ESMC language model; and the binding score is calculated based on a trained prediction model.

[0037] Sequence Score E ngram =∑ i∈(1,2,3) D KL (ngram i (x), ngram i,bg The KL divergence was used to measure the difference between the probability distribution of n-grams in the generated sequence and that in UniRef50. The amino acid sequence information in UniRef50 was statistically analyzed, with the frequencies of single amino acid occurrences, consecutive pairwise combinations of two amino acids, and consecutive combinations of three amino acids recorded as 1-gram, 2-gram, and 3-gram, respectively. The generated amino acid sequence must conform to the statistical regularity of amino acid distribution in UniRef50.

[0038] Semantic score E LM =-∑ i logp(x' i |x -i Using ESMC, the probability of an amino acid being the current amino acid is inferred from the context of each amino acid, and the joint probability of the entire sequence is calculated, which is the score of the entire sequence. This score is used to determine whether the generated sequence conforms to the rules of protein language.

[0039] Combination score E model =-Bottleneck_mlp(Encoder(x)) is based on the bottleneck layer prediction of the U-Net model and determines whether the protein binds to microfilaments.

[0040] The scoring function is E(x) = 2 × E LM (x)+E ngram (x)+E model .

[0041] According to embodiments of this application, the mutated amino acid resulting from the random mutation does not include cysteine. Therefore, excluding cysteine ​​from random mutations eliminates the risk of unintended disulfide bond formation, ensuring the structural stability of the generated sequence.

[0042] According to an embodiment of this application, the prediction model is trained as follows: an initial dataset is obtained, the initial dataset including protein sequences that bind to microfilaments; contact residue annotations are extracted from the protein sequences that bind to microfilaments; homologous protein sequences of the protein sequences that bind to microfilaments are obtained, and sequence alignment is performed to transfer the contact residue annotations to the homologous protein sequences; semantic embedding is performed on the protein sequences that bind to microfilaments and the homologous protein sequences to obtain a positive sample dataset; the prediction model is trained using the positive sample dataset to obtain a trained prediction model. Thus, by transferring annotations, the training dataset is effectively expanded, improving the model's ability to learn the features of protein sequences that bind to microfilaments, thereby enhancing the model's accuracy and generalization performance in the task of identifying and designing protein sequences that bind to microfilaments.

[0043] For example, refer to Figure 3 We collected protein sequences that might bind to microfilaments from previous studies, and used AlphaFold3 to predict their binding to complexes with six Actin proteins. The resulting binding interfaces were scored using pDockQ. For structures that passed the pDockQ scoring, the amino acid residues in contact with the Actin complex were extracted and annotated onto the protein's primary sequence. We used the MMSeqs2 tool to find homologous proteins of the input protein in the Swiss-prot database, and used sequence alignment to transfer the annotations of the binding residues to the new protein sequences. We used the ESMC-600M model to perform semantic embedding on these proteins, and then packaged the obtained protein sequence information, Actin segment information, and protein semantic embeddings into a pickle file as subsequent training data.

[0044] Furthermore, all other proteins in the human genome are considered negative samples, meaning that no amino acids bind to microfilaments and they are not microfilament-binding proteins. The same information packaging and ESMC semantic embedding are then performed on them.

[0045] According to an embodiment of this application, the prediction model is selected from U-Net.

[0046] According to embodiments of this application, the prediction model includes an encoder, a bottleneck layer, and a decoder; (Refer to...) Figure 4The encoder encodes the protein sequence to obtain encoded features; the bottleneck layer outputs the result of the protein sequence binding with microfilaments through a first multilayer perceptron; the decoder outputs the result of the amino acid binding with microfilaments; and the last transposed convolution of the decoder outputs the restored sequence through a second multilayer perceptron.

[0047] The ESMC-600M model uses a 1152-dimensional semantic embedding for proteins. This model first projects this embedding to 512 dimensions using a multilayer perceptron, then uses a one-dimensional U-net architecture as input. The output represents the regions on the sequence that may bind to microfilaments; regions with a high probability of binding have values ​​close to 1, while regions with a low probability of binding have values ​​close to 0. To simultaneously train for protein sequence classification, the bottleneck layer of U-net uses a multilayer perceptron projected to 1 dimension to determine whether a protein binds to microfilaments; high-probability binding values ​​approach 1, and low-probability binding values ​​approach 0. After the last transposed convolutional layer of U-net, a multilayer perceptron is added to project each row of the output of this layer into a 20-dimensional vector, representing the prediction of the amino acid of the original sequence corresponding to that row, as an auxiliary task.

[0048] According to some specific embodiments of this application, the original input positive sample sequence consists of 84 cases (see Table 1). Due to the limited training data for this model, severe overfitting is likely to occur. To prevent overfitting, the following regularization method is added to the model during training:

[0049] 1. The input sequence is reconstructed using the output of the last transposed convolution layer to ensure the preservation of protein semantics during training. This also generates additional gradient components, making the training task more complex and preventing overfitting.

[0050] 2. Use a larger Dropout rate at different levels, especially the feature extraction level, to randomly train different parameters, preventing the model from over-relying on a few features for judgments. Use a smaller Dropout rate in the decoder part to prevent impairing the model's expressive power. The Dropout rates for each level are as follows: Figure 4 As shown, the input encoder dropout rate is 0.4, the bottleneck layer dropout rate is 0.3, the bottleneck layer multilayer perceptron dropout rate is 0.2, the last transposed convolutional layer multilayer perceptron dropout rate is 0.2, and the decoder dropout rate is 0.2.

[0051] 3. During training, randomly select 1 to 40 tokens in the semantic embedding of ESMC and replace them with random noise masks. ESMC is a self-attention-based encoder, and partial masking does not lose much overall semantic information. Such replacement can greatly improve the diversity of training data and prevent the model from over-memorizing certain amino acid arrangement patterns with obvious features, thus preventing overfitting.

[0052] Table 1 Positive Sample Training Data

[0053]

[0054]

[0055]

[0056]

[0057]

[0058]

[0059]

[0060]

[0061]

[0062] According to an embodiment of this application, in the training of the prediction model, the bottleneck layer is trained using a first loss function to determine whether the protein sequence binds to microfilaments; the decoder is trained using a second loss function to segment amino acid residues in the protein sequence that bind to microfilaments; the last transposed convolutional layer of the decoder is trained using a third loss function to restore the sequence; the prediction model is trained by weighted fusion of the first loss function, the second loss function, and the third loss function.

[0063] The binary classification task, determining whether an input sequence is a microfilament-binding protein, utilizes the bottleneck layer semantics of U-net for binary classification. Its loss is the binary cross-entropy between the predicted probability of protein binding to microfilaments and the actual label indicating whether the protein binds to microfilaments: L label =BCE(pred,label) is the first loss function.

[0064] For the task of segmenting amino acid residues that bind to protein microfilaments, U-net is used to perform motif segmentation on the protein sequence. The loss function is the binary cross-entropy between the probability of each amino acid being predicted to bind to a microfilament and the actual binding probability: L motif =BCE(pred,motif) is the second loss function.

[0065] The task of restoring the original input sequence using the last transposed convolution layer has a loss equal to the cross-entropy between the predicted probabilities of the 20 amino acid values ​​for each bit of the input sequence and the actual amino acid values: L sequence =CrossEntropy(pred,aa) is the third loss function.

[0066] Due to the relatively small amount and high diversity of data, the model is prone to overfitting or training instability in the early stages of training. The inventors designed L... sequence This allows the final transposed convolutional layer to predict the input protein sequence, ensuring that the model's final output retains the semantics of the input. During training, L... sequence Compared to other loss functions, this one is more difficult to optimize. To aid training, dynamic weights are introduced, dynamically updating the weights based on the gradient of each loss function. The combined loss function is suitable for multi-task optimization problems, i.e.: L = ω1L motif +ω2L label +ω3L sequence Since the initial gradients generated by each loss function are of different magnitudes, and the optimization difficulty of the losses varies, in order to achieve the goal of simultaneous multi-task optimization, the gradient magnitude generated by each loss function in a single training iteration is used to constrain the coefficient values ​​of each loss function:

[0067] According to an embodiment of this application, the associativity score is calculated using the bottleneck layer of the trained prediction model.

[0068] In another aspect of this application, a device for designing microfilament-binding protein sequences is proposed. (Reference) Figure 5 The device 500 includes: an initialization module 510, a mutation module 520, and an iteration module 530, wherein...

[0069] The initial module 510 is used to randomly generate an initial amino acid sequence; the mutation module 520 is used to randomly mutate a single amino acid starting from the initial amino acid sequence to obtain a mutated amino acid sequence, and to score the amino acid sequences before and after mutation according to a scoring function; the iteration module 530 is used to select amino acid sequences whose scores meet preset conditions for the next round of random mutation, until a preset number of iterations is reached to obtain the designed sequence.

[0070] It should be understood that the device embodiments and method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, further details will not be provided here. Specifically, Figure 5 The device 500 shown can perform Figure 1The corresponding method embodiments, and the foregoing and other operations and / or functions of each module in device 500 are respectively implemented to achieve Figure 1 For the sake of brevity, the corresponding processes in each method are not described in detail here.

[0071] This application also proposes an electronic device. (Reference) Figure 6 The electronic device 600 can be, but is not limited to, the device that performs the above-described method. For example... Figure 6 As shown, the electronic device 600 may include:

[0072] The system includes a memory 610 and a processor 620. The memory 610 stores a computer program 630 and transfers the computer program 630 to the processor 620. In other words, the processor 620 can retrieve and run the computer program 630 from the memory 610 to implement the methods described in the embodiments of this application.

[0073] For example, the processor 620 can be used to execute the steps in the above method according to the instructions in the computer program 630.

[0074] In some embodiments of this application, the processor 620 may include, but is not limited to:

[0075] General-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0076] In some embodiments of this application, the memory 610 includes, but is not limited to:

[0077] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0078] In some embodiments of this application, the computer program 630 may be divided into one or more modules, which are stored in the memory 610 and executed by the processor 620 to perform the method provided in this application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 630 in the electronic device.

[0079] like Figure 6 As shown, the electronic device 600 may further include:

[0080] Transceiver 640, which can be connected to processor 620 or memory 610.

[0081] The processor 620 can control the transceiver 640 to communicate with other devices; specifically, it can send information or data to other devices or receive information or data sent by other devices. The transceiver 640 may include a transmitter and a receiver. The transceiver 640 may further include antennas, and the number of antennas may be one or more.

[0082] It should be understood that the various components in the electronic device 600 are connected through a bus system, which includes a data bus, a power bus, a control bus, and a status signal bus.

[0083] According to one aspect of this application, a computer-readable storage medium is provided that stores computer instructions or programs thereon, which, when executed by a computer, enable the computer to perform the methods of the above-described method embodiments. Alternatively, embodiments of this application also provide a computer program product containing instructions that, when executed by a computer, cause the computer to perform the methods of the above-described method embodiments.

[0084] According to another aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method described in the above-described method embodiments.

[0085] In other words, when implemented using software, it can be implemented wholly or partially in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0086] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0087] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0088] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to implement the solution of this embodiment according to actual needs. For example, the functional modules in the various embodiments of this application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0089] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0090] The following will explain the solution of this application with reference to embodiments. Those skilled in the art will understand that the following embodiments are for illustrative purposes only and should not be considered as limiting the scope of this application. Where specific techniques or conditions are not specified in the embodiments, they are performed according to the techniques or conditions described in the literature in the art or according to the product instructions. Reagents or instruments whose manufacturers are not specified are all conventional products that can be obtained commercially.

[0091] Example

[0092] In this embodiment, the prediction model trained using the method described above is used to predict the binding interface of known microfilament-binding proteins, and the results are as follows: Figure 7 As shown in (A), the prediction model trained in this application can accurately predict the binding interface of known microfilament-binding proteins, and the residue-level prediction results are in high agreement with the experimental results.

[0093] The microfilament-binding protein sequence was designed using the method described in this application, and its binding activity was verified by immunofluorescence co-localization. The results are as follows: Figure 7 As shown in (B), this method demonstrates its ability to identify novel human microfilament-binding proteins. Figure 7 (C) is the entry for this protein in the protein database.

[0094] Furthermore, using the method of this application, 22 peptides of 100 amino acids each were designed, of which 9 showed good co-localization with microfilaments, such as... Figure 7 As shown in (D).

[0095] Therefore, this application utilizes annotation migration, various regularization methods, and sequence alignment to obtain accurate microfilament-binding protein sequences using only a small number of positive samples.

[0096] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0097] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A method for designing microfilament-binding protein sequences, characterized in that, include: Randomly generate the initial amino acid sequence; Starting with the initial amino acid sequence, random mutations are performed on individual amino acids to obtain mutated amino acid sequences. The amino acid sequences before and after mutation are scored according to a scoring function. Select amino acid sequences that meet the preset criteria for the next round of random mutation until the preset number of iterations is reached to obtain the designed sequence.

2. The method according to claim 1, wherein the preset condition is: If the score corresponding to the randomly mutated amino acid sequence is greater than the score corresponding to the amino acid sequence before the random mutation, then the current randomly mutated amino acid sequence is accepted. If the score corresponding to the randomly mutated amino acid sequence is less than the score corresponding to the amino acid sequence before the random mutation, then the acceptance or rejection of the current randomly mutated amino acid sequence is determined based on the acceptance probability. Optionally, the acceptance probability is in, x is the amino acid sequence before the mutation, x ′ Here, T represents the mutated amino acid sequence, T is the temperature parameter, and E(x) is the scoring function. Optionally, the scoring function consists of sequence scoring, semantic scoring, and associativity scoring; the sequence scoring is calculated based on the KL divergence between the n-gram distribution of the designed sequence and the n-gram distribution of the natural protein library; the semantic scoring is calculated based on the ESMC language model; and the associativity scoring is calculated based on a trained prediction model. Optionally, the mutated amino acid of the random mutation does not include cysteine.

3. The method according to claim 2, characterized in that, The prediction model is trained in the following manner: Obtain an initial dataset, which includes protein sequences that bind to microfilaments; Contact residue annotations were extracted from the protein sequences that bound the microfilaments; Obtain the homologous protein sequence of the protein sequence that binds to the microfilament, and perform sequence alignment to migrate the contact residue annotation to the homologous protein sequence; The protein sequences that bind to the microfilaments and the homologous protein sequences are semantically embedded to obtain a positive sample dataset; The prediction model is trained using the positive sample dataset to obtain the trained prediction model. Optionally, the prediction model is selected from U-Net.

4. The method according to claim 3, characterized in that, The prediction model includes an encoder, a bottleneck layer, and a decoder; wherein, the encoder encodes the protein sequence to obtain encoded features; the bottleneck layer outputs the result of the protein sequence binding with microfilaments through a first multilayer perceptron; the decoder outputs the result of the amino acid binding with microfilaments; and the last transposed convolutional layer of the decoder outputs the restored sequence through a second multilayer perceptron.

5. The method according to claim 4, characterized in that, In the training of the prediction model, the bottleneck layer is trained using a first loss function to determine whether the protein sequence binds to the microfilament; the decoder is trained using a second loss function to segment the amino acid residues in the protein sequence that bind to the microfilament; and the last transposed convolutional layer of the decoder is trained using a third loss function to restore the sequence. The prediction model is trained by weighted fusion of the first loss function, the second loss function, and the third loss function.

6. The method according to any one of claims 2 to 5, characterized in that, The associativity score is calculated using the bottleneck layer of the trained prediction model.

7. A device for designing microfilament-binding protein sequences, characterized in that, include: The initial module is used to randomly generate the initial amino acid sequence; The mutation module is used to randomly mutate a single amino acid starting from the initial amino acid sequence to obtain a mutated amino acid sequence, and to score the amino acid sequences before and after mutation according to a scoring function. The iteration module is used to select amino acid sequences that meet the preset conditions for the next round of random mutation until the preset number of iterations is reached to obtain the designed sequence.

8. An electronic device, characterized in that, include: Processor and memory; The memory is used to store computer programs; The processor is configured to execute the computer program to implement the microfilament-binding protein sequence design method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions or programs that, when executed on a computer, cause the microfilament-binding protein sequence design method as described in any one of claims 1 to 6 to be performed.

10. A computer program product, characterized in that, The computer program product includes computer instructions that, when some or all of the computer instructions are run on a computer, cause the microfilament-binding protein sequence design method as described in any one of claims 1 to 6 to be executed.