A method for optimizing ALF antimicrobial peptides

Through the combination of feature engineering, machine learning and genetic algorithms, the antimicrobial peptide sequence is optimized, and the problem of poor transformation effect in the existing technology is solved, and efficient and accurate antimicrobial peptide transformation is achieved.

CN115472240BActive Publication Date: 2025-09-02BEIJING NORMAL UNIV AT ZHUHAI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211113727.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-14
Publication Date
2025-09-02
Estimated Expiration
2042-09-14

AI Technical Summary

Technical Problem

The prior art lacks effective methods for innovating the amino acid sequence of antibacterial peptides, resulting in poor transformation effect.

Method used

Using a combination of feature engineering, machine learning and genetic algorithms, the optimized antimicrobial peptide sequence is generated by generating position frequency matrix and positive charge distribution frequency, Pearson correlation coefficient calculation and fuzzy clustering analysis are performed, and multi-dimensional Euclidean distance and genetic algorithm optimization are combined to generate an optimized antimicrobial peptide sequence.

Benefits of technology

The antibacterial peptide sequence is effectively optimized, balanced the need for substitution minimization and antibacterial activity maximization, reduced experimental verification costs, and improved the success rate of transformation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115472240B_ABST
    Figure CN115472240B_ABST
Patent Text Reader

Abstract

The present invention provides a machine learning-based ALF antimicrobial peptide optimization method, which relates to the field of antimicrobial peptide modification technology and solves the technical problems of unguided modification and poor modification effects of antimicrobial peptide amino acid sequences in the prior art. The method includes: collecting published amino acid sequences and their minimum inhibitory concentrations (MICs) for Escherichia coli through a data collection module; constructing a site substitution fitness matrix 1 for the sequences to be optimized and preset sequences using partial least squares methods, etc. through a feature construction module; obtaining a positive charge interval distribution frequency matrix 2 for the preset sequence using probability distribution; and obtaining the correlation between the physicochemical properties of the preset sequence and the MIC using methods such as fuzzy clustering; integrating matrices 1 and 2 through a prediction model module to construct an amino acid site substitution fitness matrix 3; and finally evaluating the antimicrobial activity of the optimized sequence through a numerical scoring system; and completing the final model construction through an iterative verification module using an iterative method combining antimicrobial verification feedback and a genetic algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of antimicrobial peptide modification, and in particular to an ALF antimicrobial peptide optimization method. Background Art

[0002] Currently, due to the increasing resistance to antibiotics, there is an urgent need to find new alternative antimicrobial substances. Antimicrobial peptides (AMPs) are a class of polypeptides extracted from a variety of organisms, including bacteria, plants, invertebrates, and vertebrates. They are a large class of endogenous compounds widely distributed in nature and are polypeptides with broad-spectrum antimicrobial activity and certain antiviral properties. Anti-lipopolysaccharide factor (ALF) antimicrobial peptide is an antimicrobial peptide with broad-spectrum antimicrobial activity and certain antiviral properties. Currently, based on natural antimicrobial peptides, appropriate additions and deletions, and amino acid substitutions or modifications to their sequences have been shown to enhance the biological activity of natural antimicrobial peptides. However, there is currently no guidance for modifying the amino acid sequence of antimicrobial peptides. The existing technology has technical problems such as unguided modification of the amino acid sequence of antimicrobial peptides and poor modification effects. Summary of the Invention

[0003] The purpose of this application is to provide an ALF antimicrobial peptide optimization method to solve the technical problems of unguided modification and poor modification effect of antimicrobial peptide amino acid sequences in the prior art.

[0004] In a first aspect, the present invention provides an ALF antimicrobial peptide optimization method, which is applied to an ALF antimicrobial peptide optimization algorithm model; the method comprises:

[0005] Obtaining preset ALF antimicrobial peptide information and ALF antimicrobial peptide information to be optimized; wherein the preset ALF antimicrobial peptide information includes a preset LBD sequence of the preset ALF antimicrobial peptide, a preset antimicrobial index corresponding to the preset LBD sequence, and preset physicochemical property data;

[0006] Based on the preset LBD sequence, the preset antimicrobial index and the preset physicochemical property data, a position frequency matrix and a positive charge distribution frequency of the preset ALF antimicrobial peptide are generated; based on the ALF antimicrobial peptide information to be optimized, the preset ALF antimicrobial peptide information, the position frequency matrix and the positive charge distribution frequency, a first fitness matrix is ​​generated; and a first candidate LBD sequence set is obtained through the first fitness matrix; wherein the first candidate LBD sequence set includes each candidate LBD sequence with information on the number of times it is generated and candidate physicochemical property information;

[0007] Performing Pearson correlation coefficient calculation and fuzzy cluster analysis on the target physicochemical property data and target antimicrobial index corresponding to the target preset LBD sequence to obtain a range of physicochemical properties for the cluster of sequences with high antimicrobial activity, and generating a second set of candidate LBD sequences from the first set of candidate LBD sequences based on this range; wherein the target preset LBD sequence is an LBD sequence among the preset LBD sequences whose antimicrobial index meets the first preset standard;

[0008] calculating a multidimensional Euclidean distance value between the candidate physicochemical property of each candidate LBD sequence in the second candidate LBD sequence set and the target physicochemical property, and evaluating the antibacterial activity of each candidate LBD sequence in the second candidate LBD sequence set based on the multidimensional Euclidean distance value and the number of times each candidate LBD sequence is generated, and selecting the candidate LBD sequence whose evaluation result meets the second preset standard as the first optimized LBD sequence, to obtain a first optimized LBD sequence set;

[0009] The first fitness matrix is ​​updated based on the first optimized LBD sequence and the offspring sequence to obtain a second fitness matrix, and the first optimized LBD sequence set is optimized using the second fitness matrix to obtain a second optimized LBD sequence set; wherein the offspring sequence is generated by the first optimized LBD sequence through genetic algorithm calculation.

[0010] In one possible implementation, the ALF antimicrobial peptide information to be optimized includes the LBD sequence to be optimized of the ALF antimicrobial peptide to be optimized; generating the position frequency matrix and the positive charge distribution frequency of the preset ALF antimicrobial peptide based on the preset LBD sequence, the preset antimicrobial index, and the preset physicochemical property data, generating a first fitness matrix based on the ALF antimicrobial peptide information to be optimized, the preset ALF antimicrobial peptide information, the position frequency matrix, and the positive charge distribution frequency, and obtaining a first candidate LBD sequence set through the first fitness matrix, including:

[0011] Performing a multiple sequence alignment on the preset LBD sequence and the LBD sequence to be optimized, and generating a position frequency matrix based on the alignment results;

[0012] Performing multiple linear regression and regression tree construction on the ALF antimicrobial peptide information to be optimized, the position frequency matrix, and the antimicrobial index results using a partial least squares regression algorithm and a gradient boosting decision tree algorithm to determine the first contribution value of each sequence site to the antimicrobial index result, thereby obtaining an amino acid substitution matrix;

[0013] The interval distribution frequency of positively charged amino acids in the target preset LBD sequence is calculated to obtain an interval distribution frequency matrix;

[0014] Performing Monte Carlo simulation on the first contribution value in the amino acid substitution matrix and the lysine substitution contribution value in the interval distribution frequency matrix to determine the second contribution value of the lysine site to the antibacterial index result;

[0015] Replacing the first contribution value of the lysine site with the second contribution value to obtain a first fitness matrix;

[0016] A first candidate LBD sequence set is obtained through the first fitness matrix, including the number of times each candidate LBD sequence is generated and the candidate physicochemical property information; wherein, the first candidate LBD sequence set includes the number of times each candidate LBD sequence is generated and the candidate physicochemical property information, and the target preset LBD sequence is an LBD sequence in the preset LBD sequence whose antibacterial index meets the first preset standard.

[0017] In one possible implementation, the target physicochemical property data and target antibacterial index corresponding to the target preset LBD sequence are subjected to Pearson correlation coefficient calculation and fuzzy cluster analysis to obtain the physicochemical property range of the high antibacterial activity sequence cluster, and based on this range, a second candidate LBD sequence set is generated from the first candidate LBD sequence set, including:

[0018] Calculate the Pearson correlation coefficient between the target physicochemical property data and the target preset antibacterial index value;

[0019] If the Pearson correlation coefficient meets the third preset criterion, a two-dimensional and / or three-dimensional fuzzy cluster analysis is performed on the target physicochemical property data and the target preset antibacterial index value to obtain the optimal modified physicochemical property range of the LBD sequence to be optimized;

[0020] A candidate LBD sequence set is generated based on the optimal modified physicochemical property range.

[0021] In one possible implementation, the multidimensional Euclidean distance value between the candidate physicochemical property of each candidate LBD sequence in the second candidate LBD sequence set and the target physicochemical property is calculated, and the antibacterial activity of each candidate LBD sequence in the second candidate LBD sequence set is evaluated based on the multidimensional Euclidean distance value and the number of times each candidate LBD sequence is generated, and the candidate LBD sequence whose evaluation result meets the second preset standard is used as the first optimized LBD sequence to obtain the first optimized LBD sequence set, including:

[0022] Obtaining the physicochemical property data of any sequence in the candidate LBD sequence set, and calculating the multidimensional Euclidean distance between the above value and the target physicochemical property;

[0023] Get the number of times each candidate LBD sequence is generated;

[0024] determining the relative antibacterial activity of any one of the sequences according to the multidimensional Euclidean distance;

[0025] Normalizing the number of occurrences and the relative antibacterial activity, taking equal weights and summing them to obtain an evaluation result;

[0026] The candidate LBD sequence whose evaluation result meets the second preset standard is used as the first optimized LBD sequence to obtain a first optimized LBD sequence set.

[0027] In one possible implementation, updating the first fitness matrix based on the first optimized LBD sequence and the offspring sequence to obtain a second fitness matrix, and optimizing the first optimized LBD sequence set using the second fitness matrix to obtain a second optimized LBD sequence set includes:

[0028] performing antibacterial verification on a portion of the first optimized LBD sequence to obtain a verified first sub-optimized LBD sequence;

[0029] Crossing a target sub-optimized LBD sequence exhibiting high antibacterial activity with a second sub-optimized LBD sequence among the first sub-optimized LBD sequences using a genetic algorithm to obtain a crossover progeny sequence; wherein the second sub-optimized LBD sequence is a sequence among the first optimized LBD sequences that has not undergone antibacterial verification;

[0030] The first fitness matrix is ​​updated using the first optimized LBD sequence and the crossover offspring sequence to obtain a second fitness matrix;

[0031] Based on the second fitness matrix, a second optimization process is performed on the designated LBD sequence in the first optimized LBD sequence set to obtain a second optimized LBD sequence; wherein the designated LBD sequence is an LBD sequence whose antibacterial activity meets a fourth preset standard.

[0032] In one possible implementation, the preset ALF antimicrobial peptides are 31 crustacean ALF antimicrobial peptides; the preset antimicrobial index is the minimum inhibitory concentration value and / or the minimum bactericidal concentration value; and the preset physicochemical property data is the sequence net charge and hydrophobicity.

[0033] In a possible implementation, after taking the candidate LBD sequence whose evaluation result meets the second preset criterion as the first optimized LBD sequence to obtain the first optimized LBD sequence set, the method further includes:

[0034] The secondary structure of the first optimized LBD sequence was predicted using the PEP-FOLD3 method to obtain a polypeptide structure model.

[0035] In a second aspect, an embodiment of the present application provides an ALF antimicrobial peptide optimization device, which is applied to an ALF antimicrobial peptide optimization algorithm model. The device comprises:

[0036] An acquisition module is used to acquire preset ALF antimicrobial peptide information and ALF antimicrobial peptide information to be optimized; wherein the preset ALF antimicrobial peptide information includes a preset LBD sequence of the preset ALF antimicrobial peptide, a preset antimicrobial index corresponding to the preset LBD sequence, and preset physicochemical property data;

[0037] a generation module, configured to generate a position frequency matrix and a positive charge distribution frequency of the preset ALF antimicrobial peptide based on the preset LBD sequence, the preset antimicrobial index, and the preset physicochemical property data, and generate a first fitness matrix based on the ALF antimicrobial peptide information to be optimized, the preset ALF antimicrobial peptide information, the position frequency matrix, and the positive charge distribution frequency, and obtain a first candidate LBD sequence set through the first fitness matrix; wherein the first candidate LBD sequence set includes each candidate LBD sequence accompanied by generation count information and candidate physicochemical property information;

[0038] an analysis module, configured to calculate a Pearson correlation coefficient and perform fuzzy cluster analysis on the target physicochemical property data and target antimicrobial index corresponding to the target preset LBD sequence, obtain a range of physicochemical properties for clusters of sequences with high antimicrobial activity, and generate a second set of candidate LBD sequences from the first set of candidate LBD sequences based on this range; wherein the target preset LBD sequence is an LBD sequence among the preset LBD sequences whose antimicrobial index meets the first preset standard;

[0039] a calculation module, configured to calculate a multidimensional Euclidean distance value between the candidate physicochemical property of each candidate LBD sequence in the second candidate LBD sequence set and the target physicochemical property, and evaluate the antibacterial activity of each candidate LBD sequence in the second candidate LBD sequence set based on the multidimensional Euclidean distance value and the number of times each candidate LBD sequence is generated, and select the candidate LBD sequence whose evaluation result meets the second preset standard as the first optimized LBD sequence, to obtain a first optimized LBD sequence set;

[0040] An optimization module is used to update the first fitness matrix based on the first optimized LBD sequence and the offspring sequence to obtain a second fitness matrix, and optimize the first optimized LBD sequence set using the second fitness matrix to obtain a second optimized LBD sequence set; wherein the offspring sequence is generated by the first optimized LBD sequence through genetic algorithm calculation.

[0041] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the steps of the method described in the first aspect are implemented.

[0042] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to execute the steps of the method described in the first aspect above.

[0043] The embodiments of the present application bring the following beneficial effects:

[0044] The embodiment of the present application provides an ALF antimicrobial peptide optimization method, which first obtains preset ALF antimicrobial peptide information and ALF antimicrobial peptide information to be optimized, wherein the preset ALF antimicrobial peptide information includes a preset LBD sequence of the preset ALF antimicrobial peptide, a preset antimicrobial index corresponding to the preset LBD sequence, and preset physicochemical property data; the ALF antimicrobial peptide information to be optimized includes a to-be-optimized LBD sequence of the ALF antimicrobial peptide; a fuzzy clustering analysis is performed on the target preset physicochemical property data and the target preset antimicrobial index corresponding to the target preset LBD sequence; wherein the target preset LBD sequence is an LBD sequence in the preset LBD sequence whose antimicrobial index meets the preset standard; and then, based on the preset A A first fitness matrix is ​​created using the LF antimicrobial peptide information and the ALF antimicrobial peptide information to be optimized. The Euclidean distance between the candidate optimized LBD sequence set obtained from the optimization of the first fitness matrix and the target preset LBD sequence is calculated. An antimicrobial activity score evaluation system is constructed by equally weightedly adding the number of times the candidate optimized LBD sequence set is generated and the Euclidean distance value. This results in a first optimized LBD sequence set. The first fitness matrix is ​​then updated based on the first optimized LBD sequence and its genetic algorithm crossover offspring to obtain a second fitness matrix. The first optimized LBD sequence is then subjected to a second optimization process using the second fitness matrix and the score evaluation system to obtain a second optimized LBD sequence. In this approach, the antimicrobial peptide modification algorithm based on feature engineering, machine learning, and genetic algorithms is efficient, accurate, rapid, and convenient. It can simultaneously balance the modification requirements of minimizing substitutions and maximizing antimicrobial activity, making it suitable for gene editing of antimicrobial peptide sequences in hosts. By combining machine learning with genetic algorithms, the accuracy of the algorithm is guaranteed while reducing the experimental verification cost, solving the technical problems of unguided modification and poor modification effects of antimicrobial peptide amino acid sequences in existing technologies. Moreover, compared with blind synthesis after modification using machine learning alone and modification using genetic algorithms alone, the combination of the two methods can maximize the success rate of modification. In particular, the research and development ideas of this algorithm are applicable to the modification of any type of antimicrobial peptides and have broad application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0046] Figure 1 A schematic diagram of a process for optimizing an ALF antimicrobial peptide provided in an embodiment of the present application;

[0047] Figure 2A schematic structural diagram of an ALF antimicrobial peptide optimization model provided in an embodiment of the present application;

[0048] Figure 3 A schematic diagram of the fuzzy cluster analysis results of the net charge, hydrophobicity and MIC value of a sequence provided in an embodiment of the present application;

[0049] Figure 4 A schematic diagram of the Monte Carlo simulation results of amino acid position substitution and positively charged amino acid arrangement frequency provided in an embodiment of the present application;

[0050] Figure 5 A schematic diagram of the Pearson correlation coefficient between a physicochemical property and MIC value provided in an embodiment of the present application;

[0051] Figure 6 A schematic diagram of the structure of an ALF antimicrobial peptide optimization device provided in an embodiment of the present application;

[0052] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0054] The terms "including," "having," and any variations thereof, as used in the embodiments of this application, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.

[0055] Due to increasing antibiotic resistance, there is an urgent need to find new alternative antimicrobial agents. Antimicrobial peptides (AMPs) are a class of polypeptides extracted from a wide range of organisms, from bacteria to plants, invertebrates, and vertebrates. They are a large class of endogenous compounds widely distributed in nature and possess broad-spectrum antimicrobial activity and some antiviral properties. AMPs can specifically recognize carbohydrate-containing molecules, namely polysaccharides, glycolipids, glycoproteins, peptidoglycans, or lipopolysaccharides, from Gram-negative and Gram-positive bacteria, viruses, or fungi. Consequently, they possess broad-spectrum antimicrobial, antifungal, antiviral, antiprotozoal, and antiseptic properties. Furthermore, they can directly interact with host cells, modulating inflammatory processes and innate defenses. AMPs can be classified into cationic and anionic AMPs based on their charge. They can be divided into insect, mammalian, amphibian, fish, and microbial AMPs based on their origin. Based on their structure, they can be divided into α-helical AMPs, β-sheet AMPs, and other AMPs.

[0056] Anti-lipopolysaccharide factor (ALF) is an antimicrobial peptide with broad-spectrum antimicrobial activity and some antiviral properties. ALFs are a key member of the antimicrobial peptide family of shrimp and are a key component of host defense against bacteria, fungi, and viruses. ALFs are amphipathic peptides containing two highly conserved cysteine ​​residues that form a stable disulfide-bonded ring. They are generally composed of 114-124 amino acids, including a 16-26 amino acid signal peptide. The mature peptide has a molecular weight of approximately 11 kDa. ALFs bind to lipopolysaccharide (LPS) through their lipopolysaccharide-binding domain (LBD) and neutralize LPS. Because many naturally occurring antimicrobial peptides identified and isolated to date have suboptimal activity or exhibit toxicity to eukaryotic cells, the design and modification of natural antimicrobial peptides has proven to be a viable and effective approach in recent years. Currently, using natural antimicrobial peptides as a base, appropriate sequence additions, deletions, amino acid substitutions, or modifications have been shown to enhance their biological activity.

[0057] Although the physicochemical properties of AMPs can be changed by rationally replacing the amino acid sites of natural AMPs, thereby improving the antibacterial activity and hemolytic toxicity of antimicrobial peptides. In the prior art, in order to overcome the difficulties in improving the knowledge-based design of antimicrobial peptides, machine learning-based methods such as artificial neural network models combined with chemical informatics methods are used to capture more complex antimicrobial activity motifs, but they require relatively large data sets for accurate predictions, as well as careful adjustment of parameters based on the starting peptide library. The combination of genetic algorithms and experimental verification feedback has been successfully applied to the optimization and modification of small peptides. However, it has not been applied to sequences with more than 20 amino acids (e.g., 20 amino acids applied to 20-mer short peptides). 20The computational cost and expensive peptide synthesis verification are obviously impossible. Therefore, the existing technology has the technical problem of poor modification effect on the amino acid sequence of antimicrobial peptides.

[0058] Based on this, the present invention provides an ALF antimicrobial peptide optimization method, namely, an algorithm for modifying the amino acid sequence of the LBD binding domain of the ALF family of antimicrobial peptides. By collecting published LBD binding domain amino acid sequences and their minimum inhibitory concentration (MIC) against a certain bacterium (e.g., Escherichia coli), a fuzzy clustering algorithm is used to perform cluster analysis on the MIC value, hydrophobicity, net charge, and folding pattern of the LBD sequence. It is found that LBD sequences with low MIC values ​​(high antimicrobial activity) have obvious cluster populations on the three-dimensional plane. At the same time, partial least squares regression and gradient boosting decision trees are used to discover the motif that lays the foundation for the high antimicrobial activity of the LBD sequence at the amino acid arrangement positions of low and high MIC values. The frequency of the interval distribution of positive charge in the low MIC value sequence is calculated, thereby establishing an initial fitness matrix based on the collected LBD sequences. Finally, an iterative calculation method combining a genetic algorithm and verification feedback is used to modify the initial input LBD sequence in the direction of minimizing variation and maximizing antimicrobial activity, alleviating the technical problem of poor modification effect of antimicrobial peptide amino acid sequences in the prior art.

[0059] The embodiments of the present application are further described below with reference to the accompanying drawings.

[0060] Figure 1 A schematic flow chart of an ALF antimicrobial peptide optimization method provided in an embodiment of the present application.

[0061] Among them, this method can be applied to Figure 2 The ALF antimicrobial peptide optimization algorithm model shown in FIG, includes a data collection module, a feature construction module, an optimization operation module, and an iterative verification module. Figure 1 As shown, the method includes:

[0062] Step S110: obtaining preset ALF antimicrobial peptide information and ALF antimicrobial peptide information to be optimized.

[0063] The preset ALF antimicrobial peptide information includes a preset LBD sequence of the preset ALF antimicrobial peptide, a preset antimicrobial index corresponding to the preset LBD sequence, and preset physical and chemical property data.

[0064] For example, Figure 2As shown, the system can first obtain information on the 31 currently published crustacean ALF family antimicrobial peptides (preset ALF antimicrobial peptides) and the ALF antimicrobial peptides to be optimized through the data collection module. The antimicrobial peptide information includes LBD sequence information. For example, the LBD sequence information corresponding to the 31 published crustacean ALF antimicrobial peptides is LBD-1, LBD-2, LBD-3, etc., and the LBD sequence information of the ALF antimicrobial peptide to be optimized is LBD-0. The antimicrobial peptide information also includes the antimicrobial index (such as MIC value) and physicochemical properties (such as hydrophobicity, net charge, etc.) corresponding to each antimicrobial peptide.

[0065] Step S120: Generate a position frequency matrix and a positive charge distribution frequency of a preset ALF antimicrobial peptide based on a preset LBD sequence, a preset antimicrobial index, and preset physicochemical property data; generate a first fitness matrix based on the ALF antimicrobial peptide information to be optimized, the preset ALF antimicrobial peptide information, the position frequency matrix, and the positive charge distribution frequency; and obtain a first candidate LBD sequence set through the first fitness matrix.

[0066] The first candidate LBD sequence set includes each candidate LBD sequence and the accompanying generation number information and candidate physical and chemical property information.

[0067] For example, Figure 2 As shown, the system can construct a first fitness matrix through a feature construction module. For example, by comparing the differences between the ALF antimicrobial peptide LBD sequence to be optimized and the 31 published crustacean ALF antimicrobial peptide sequences (sequence site differences between the preset ALF antimicrobial peptide and the ALF antimicrobial peptide to be optimized, the frequency of the spacing of the positively charged amino acids in the preset ALF antimicrobial peptide, etc.), the contribution rate of each site in the sequence to the antimicrobial ability and the frequency of the spacing of the positively charged amino acids (e.g., K, R, H) in the sequence to the antimicrobial ability can be determined, that is, which amino acid is replaced at each site in the ALF antimicrobial peptide LBD sequence to improve the antimicrobial activity, and the first fitness matrix is ​​established based on the above information. Then, the first fitness matrix can be calculated 1 million times based on the law of large numbers to obtain the first candidate LBD sequence set, and each sequence in the set is accompanied by the number of times it was generated and the value of the physical and chemical properties.

[0068] In step S130, the target physicochemical property data and the target antibacterial index corresponding to the target preset LBD sequence are subjected to Pearson correlation coefficient calculation and fuzzy cluster analysis to obtain the physicochemical property range of the high antibacterial activity sequence cluster, and based on this range, a second candidate LBD sequence set is generated from the first candidate LBD sequence set.

[0069] The target preset LBD sequence is an LBD sequence whose antibacterial index meets the first preset standard among the preset LBD sequences.

[0070] For example, the preset standard can be specifically set according to the actual situation. In the embodiment of the present application, the sequences with MIC values ​​less than 40 among the 31 currently published crustacean ALF antimicrobial peptide LBD sequences are determined as target preset LBD sequences. By performing fuzzy cluster analysis on the physicochemical property data and the corresponding MIC values ​​of the target preset LBD sequence, the influence of the physicochemical property data of the LBD sequence on the MIC value can be better determined, and then it can be determined which specific site in the LBD sequence to be optimized and what kind of replacement to make, so as to achieve the improvement of the antimicrobial ability of the ALF antimicrobial peptide to be optimized and better achieve the transformation of the ALF antimicrobial peptide. Figure 3 As shown, after fuzzy cluster analysis, peptide sequences with MIC values ​​below 40 have hydrophobicity distributions in the range of 0.5-0.7 and net charge distributions in the range of 2-7. Therefore, based on the principle of minimizing substitutions and ensuring that the hydrophobicity of the replacement-generated sequences is likely to be in the range of 0.5-0.7 and the net charge is likely to be in the range of 2-7, only 3-6 amino acid sites can be selected for substitution in the sequence to be optimized, thereby generating the first candidate LBD sequence using the fitness matrix. Then, based on the above-mentioned physicochemical property ranges, the second candidate LBD sequence set is generated from the first candidate LBD sequence set.

[0071] Step S140, calculating the multidimensional Euclidean distance value between the candidate physicochemical properties and the target physicochemical properties of each candidate LBD sequence in the second candidate LBD sequence set, and evaluating the antibacterial activity of each candidate LBD sequence in the second candidate LBD sequence set based on the multidimensional Euclidean distance value and the number of times each candidate LBD sequence is generated, and the candidate LBD sequence whose evaluation result meets the second preset standard is used as the first optimized LBD sequence to obtain the first optimized LBD sequence set.

[0072] For example, the multidimensional Euclidean distance value between the physicochemical properties of the second candidate LBD sequence set generated by the first fitness matrix and the corresponding physicochemical properties of the target preset LBD sequence can be calculated, and the scoring system composed of this value and the number of times the second candidate LBD sequence is generated can be used to evaluate the antibacterial activity of the candidate LBD sequence, thereby obtaining the first optimized LBD sequence set.

[0073] Step S150 , updating the first fitness matrix based on the first optimized LBD sequence and the offspring sequence to obtain a second fitness matrix, and optimizing the first optimized LBD sequence set using the second fitness matrix to obtain a second optimized LBD sequence set.

[0074] The offspring sequence is generated by the first optimized LBD sequence through genetic algorithm calculation.

[0075] For example, Figure 2As shown, the specific optimization process can be as follows: The top three most frequently occurring sequences from the first optimized LBD sequence are extracted for peptide synthesis and antimicrobial validation, resulting in a peptide sequence with experimentally confirmed improved antimicrobial activity. An iterative run module then performs a genetic algorithm crossover between the peptide sequences with experimentally confirmed improved antimicrobial activity and 50 unverified sequences. The resulting offspring are then used to update the first fitness matrix using data features to obtain a second fitness matrix. The second fitness matrix is ​​then used to optimize the first set of optimized LBD sequences, and the sequence is reoptimized and transformed to obtain a second, further optimized LBD sequence. In practical applications, this process can be repeated multiple times to achieve even better transformation results.

[0076] It should be noted that the above “3 items with the highest frequency of occurrence” and “50 items that have not been verified” are for illustration only and can be set according to actual conditions in actual applications.

[0077] In the examples of this application, the antimicrobial peptide modification algorithm based on feature engineering, machine learning, and genetic algorithms is efficient, accurate, fast, and convenient. It can simultaneously balance the modification requirements of minimizing replacement and maximizing antimicrobial activity, making it applicable to gene editing of antimicrobial peptide sequences in hosts. By combining machine learning with genetic algorithms, the accuracy of the algorithm is guaranteed while reducing experimental verification costs. This solves the technical problems of unguided modification and poor modification effects of antimicrobial peptide amino acid sequences in the prior art. Moreover, compared with blind synthesis after modification using machine learning alone and genetic algorithms alone, the combination of the two methods can maximize the success rate of modification. In particular, the development strategy of this algorithm is applicable to the modification of any type of antimicrobial peptide and has broad application value.

[0078] The above steps are described in detail below.

[0079] In some embodiments, the ALF antimicrobial peptide information to be optimized includes the LBD sequence to be optimized of the ALF antimicrobial peptide to be optimized; the above step S120 may specifically include the following steps:

[0080] Step a) performs a multiple sequence alignment on the preset LBD sequence and the LBD sequence to be optimized, and generates a position frequency matrix based on the alignment results.

[0081] Step b) Multiple linear regression and regression tree construction are performed on the optimized ALF antimicrobial peptide information, position frequency matrix, and antimicrobial index results using the partial least squares regression algorithm and the gradient boosting decision tree algorithm to determine the first contribution value of each sequence site to the antimicrobial index result and obtain the amino acid substitution matrix.

[0082] Step c) calculating the interval distribution frequency of the positively charged amino acids in the target preset LBD sequence to obtain an interval distribution frequency matrix.

[0083] Step d) performing Monte Carlo simulation on the first contribution value in the amino acid substitution matrix and the lysine substitution contribution value in the interval distribution frequency matrix to determine the second contribution value of the lysine site to the antibacterial index result.

[0084] Step e) replacing the first contribution value of the lysine site with the second contribution value to obtain a first fitness matrix.

[0085] Step f) obtaining a first set of candidate LBD sequences through the first fitness matrix, including information on the number of times each candidate LBD sequence is generated and information on candidate physical and chemical properties.

[0086] For the above step f), the first candidate LBD sequence set includes each candidate LBD sequence with the number of times it is generated and the candidate physicochemical property information, and the target preset LBD sequence is the LBD sequence whose antibacterial index meets the first preset standard among the preset LBD sequences.

[0087] Exemplarily, the system can perform a multiple sequence alignment on the preset LBD sequence and the LBD sequence to be optimized, and establish a position frequency matrix (X) based on the comparison results; perform multiple linear regression and regression tree construction on the position frequency matrix (X) and MIC value (Y) of the ALF antimicrobial peptide to be optimized and the preset ALF antimicrobial peptide through partial least squares regression and gradient boosting decision tree to determine the first contribution value of each sequence site in the sequence site substitution matrix to the antibacterial index, as shown in Tables 1 and 2, where the smaller the value, the lower the MIC value can be achieved by replacing the site with the amino acid, thereby obtaining an amino acid substitution matrix.

[0088] Table 1 Position substitution matrix and MIC contribution value (a);

[0089] 1 2 3 4 5 6 7 8 9 10 G 0.55 -1.93 A V -0.17 -0.20 -0.65 L 0.55 -0.60 1.69 I -0.92 0.79 -1.39 -0.49 0.43 P F 0.29 -0.20 1.17 Y -0.20 -1.39 1.84 -0.38 W S 1.21 -0.39 0.25 -0.74 -0.21 T 0.36 1.44 0.11 0.24 C M 0.55 N 0.79 -1.40 0.55 Q -0.61 -0.05 -0.74 D 0.55 E 0.55 K -1.94 0.00 0.90 -0.38 0.86 0.00 -0.83 0.00 -1.33 -0.29

[0090] Table 2 Position substitution matrix and MIC contribution value (b);

[0091]

[0092]

[0093] By constructing a data feature module and creatively combining the contribution of positive charge arrangement elements in the amphipathic nature of antimicrobial peptides to the MIC value, the interval arrangement frequency of positively charged amino acids (K, R, H) in antimicrobial peptides with MIC values ​​lower than 40 was calculated (as shown in Table 3), and the interval distribution frequency matrix was obtained. The interval arrangement frequency and the contribution value of the amino acid substitution matrix were subjected to Monte Carlo simulation (the simulation results are shown in Table 3). Figure 4 ), determine the second contribution value of lysine site substitution to the antibacterial index, and replace the lysine site substitution contribution value in matrix 1, thereby creating a first fitness matrix. Based on the first fitness matrix and the sequence sites to be optimized, a first optimization process can be performed on the LBD sequence to be optimized. For example, using the law of large numbers, after repeating the calculation 1 million times, the obtained sequence results are relatively stable, and all the results are used as the first optimized LBD sequence.

[0094] Table 3 Frequency table of interval arrangement of positively charged amino acids;

[0095] Interval 0 Interval 1 Interval 2 Interval 3 Interval 4 Interval 5 Interval 6 Interval 7 frequency 0.16 0.41 0.07 0.18 0.05 0.05 0.03 0.03

[0096] In some embodiments, the above step S130 may specifically include the following steps:

[0097] Step g) calculating the Pearson correlation coefficient between the target physicochemical property data and the target preset antibacterial index value.

[0098] In step h), if the Pearson correlation coefficient meets the third preset criterion, a two-dimensional and / or three-dimensional fuzzy cluster analysis is performed on the target physicochemical property data and the target preset antibacterial index value to obtain the optimal modified physicochemical property range of the LBD sequence to be optimized.

[0099] Step i) generating a set of candidate LBD sequences based on the optimal modified physicochemical property range.

[0100] For example, Figure 5 As shown, the Pearson correlation coefficient between the hydrophobicity and net charge (preset physicochemical property data) and the corresponding MIC value of the target preset LBD sequence (a sequence with a MIC value of less than 40 among the 31 published crustacean ALF antimicrobial peptide LBD sequences) can be calculated, and the Pearson correlation coefficient between the hydrophobicity and net charge of the polypeptide and the MIC value is about 0.5. The preset threshold can be specifically set according to the actual situation. In the embodiment of the present application, the Pearson correlation coefficient of 0.5 does not meet the preset threshold. Therefore, it is not possible to simply use a linear model to describe the correspondence between physicochemical property data such as hydrophobicity and net charge and the MIC value. It is necessary to perform fuzzy cluster analysis on the physicochemical property data and MIC value of the LBD sequence to be optimized, so as to determine the sequence site to be optimized of the LBD sequence to be optimized and determine the range of the optimal modified physicochemical property.

[0101] By performing fuzzy cluster analysis on the physicochemical property data and MIC values ​​in the existing data, the optimal range of modified physicochemical properties can be obtained, which can better determine the impact of the physicochemical property data of the LBD sequence on the MIC value, and then determine which specific site in the LBD sequence to be optimized and what kind of replacement to make, so as to achieve the improvement of the antibacterial ability of the ALF antimicrobial peptide to be optimized and better realize the modification of the ALF antimicrobial peptide.

[0102] In some embodiments, the above step S140 may specifically include the following steps:

[0103] Step j) obtaining the physicochemical property data of any sequence in the candidate LBD sequence set, and calculating the multi-dimensional Euclidean distance between the above value and the target physicochemical property.

[0104] Step k), obtaining the number of times each candidate LBD sequence is generated.

[0105] Step 1) Determine the relative antibacterial activity of any sequence based on the multidimensional Euclidean distance.

[0106] In step m), the number of occurrences and the relative antibacterial activity are normalized and summed with equal weights to obtain the evaluation result.

[0107] Step n): taking the candidate LBD sequence whose evaluation result meets the second preset standard as the first optimized LBD sequence to obtain a first optimized LBD sequence set.

[0108] For example, the system can obtain the hydrophobicity and net charge values ​​of any sequence α in the first candidate LBD sequence set generated by the first fitness matrix, and calculate the Euclidean distance (α, β) between the above values ​​and the hydrophobicity and net charge values ​​of each sequence β in the target preset sequence. i ), select k distance (α, β k ) and calculate their sum. Then, the evaluation score system is used to obtain the number of times each sequence in the first candidate LBD sequence set is generated (seq_order); according to any sequence α in the first candidate LBD sequence set i Relative to any k β k Euclidean distance of physical and chemical properties, obtain α i The relative MIC value (relative_mic) of the peptide was obtained; after normalizing the two attributes seq_order and relative_mic, the sum of the two attributes with equal weights was taken as the final score of each short peptide, and then the candidate LBD sequence whose evaluation result met the second preset standard was taken as the first optimized LBD sequence to obtain the first optimized LBD sequence set.

[0109] In some embodiments, the above step S150 may specifically include the following steps:

[0110] Step o), performing antibacterial verification on a portion of the first optimized LBD sequence to obtain a verified first sub-optimized LBD sequence.

[0111] Step p) crosses the target sub-optimized LBD sequence exhibiting high antibacterial activity in the first sub-optimized LBD sequence and the second sub-optimized LBD sequence using a genetic algorithm to obtain a crossover progeny sequence.

[0112] Step q): updating the first fitness matrix by using the first optimized LBD sequence and the crossover offspring sequence to obtain a second fitness matrix.

[0113] Step r) Based on the second fitness matrix, a second optimization process is performed on the designated LBD sequence in the first optimized LBD sequence set to obtain a second optimized LBD sequence.

[0114] Regarding the above step p), the second sub-optimized LBD sequence is a sequence in the first optimized LBD sequence that has not been subjected to antibacterial verification.

[0115] In the above step r), the designated LBD sequence is an LBD sequence whose antibacterial activity meets the fourth preset standard.

[0116] For example, Figure 2 As shown, the top three sequences with the highest frequency of occurrence from the first optimized LBD sequence can be extracted for peptide synthesis and antibacterial validation. Table 4 shows a validation result. This yields a peptide sequence that has been experimentally verified to have improved antibacterial activity. An iterative run module then performs a genetic algorithm crossover between the peptide sequence with experimentally verified improved antibacterial activity and an unverified progeny sequence. The resulting progeny pair data features are used to update and optimize the first fitness matrix, resulting in a second fitness matrix. This second fitness matrix is ​​then used to further optimize the sequence, yielding a further optimized LBD sequence.

[0117] Table 4 Second optimized sequence information and antibacterial activity verification results;

[0118]

[0119] By cyclically updating and optimizing the fitness matrix based on the modified sequence, the modified sequence can be further optimized using the updated fitness matrix, making the results of the sequence modification more accurate and the antibacterial ability stronger.

[0120] In some embodiments, the preset ALF antimicrobial peptides are 31 crustacean ALF antimicrobial peptides; the preset antimicrobial index is the minimum inhibitory concentration value and / or the minimum bactericidal concentration value; and the preset physicochemical property data is the sequence net charge and hydrophobicity.

[0121] For example, the examples of this application use the LBD sequences of 31 currently published crustacean ALF antimicrobial peptides for illustration. The antimicrobial index used is the MIC value, and the physicochemical property data used are the sequence net charge and hydrophobicity. However, this scheme is also applicable to the modification of other types of antimicrobial peptides. Depending on actual needs, the minimum inhibitory concentration (MIC) value can be replaced with other antimicrobial indicators such as the minimum bactericidal concentration (MBC), and the physicochemical property data can be replaced with other physicochemical property data such as hydrophobic moment and folding mode. This example does not limit this.

[0122] In some embodiments, after the above step S140, the method may further include the following steps:

[0123] Step s), using the PEP-FOLD3 method to predict the secondary structure of the first optimized LBD sequence to obtain a polypeptide structure model.

[0124] For example, Figure 2 As shown, after obtaining the optimized first optimized LBD sequence, PEP-FOLD3 can be used to predict the secondary structure of the candidate polypeptide sequence, thereby generating a visual result to facilitate analysis and processing by experimenters.

[0125] Figure 6 This is a schematic diagram of the structure of an ALF antimicrobial peptide optimization device provided in the embodiment of this application. Figure 6 As shown, the ALF antimicrobial peptide optimization device 600 includes:

[0126] Acquisition module 601 is used to acquire preset ALF antimicrobial peptide information and ALF antimicrobial peptide information to be optimized; wherein the preset ALF antimicrobial peptide information includes a preset LBD sequence of the preset ALF antimicrobial peptide, a preset antimicrobial index corresponding to the preset LBD sequence, and preset physicochemical property data;

[0127] A generation module 602 is configured to generate a position frequency matrix and a positive charge distribution frequency of a preset ALF antimicrobial peptide based on a preset LBD sequence, a preset antimicrobial index, and preset physicochemical property data, generate a first fitness matrix based on the ALF antimicrobial peptide information to be optimized, the preset ALF antimicrobial peptide information, the position frequency matrix, and the positive charge distribution frequency, and obtain a first set of candidate LBD sequences using the first fitness matrix; wherein the first set of candidate LBD sequences includes information on the number of times each candidate LBD sequence is generated and candidate physicochemical property information.

[0128] Analysis module 603 is configured to calculate the Pearson correlation coefficient and perform fuzzy cluster analysis on the target physicochemical property data and target antimicrobial index corresponding to the target preset LBD sequence to obtain a range of physicochemical properties for the cluster of sequences with high antimicrobial activity, and generate a second set of candidate LBD sequences from the first set of candidate LBD sequences based on this range; wherein the target preset LBD sequence is an LBD sequence among the preset LBD sequences whose antimicrobial index meets the first preset standard;

[0129] a calculation module 604 for calculating a multidimensional Euclidean distance value between the candidate physicochemical property and the target physicochemical property of each candidate LBD sequence in the second candidate LBD sequence set, and evaluating the antibacterial activity of each candidate LBD sequence in the second candidate LBD sequence set based on the multidimensional Euclidean distance value and the number of times each candidate LBD sequence was generated, and selecting the candidate LBD sequence whose evaluation result meets the second preset standard as the first optimized LBD sequence to obtain the first optimized LBD sequence set;

[0130] The optimization module 605 is used to update the first fitness matrix based on the first optimized LBD sequence and the offspring sequence to obtain a second fitness matrix, and optimize the first optimized LBD sequence set using the second fitness matrix to obtain a second optimized LBD sequence set; wherein the offspring sequence is generated by the first optimized LBD sequence through genetic algorithm calculation.

[0131] In some embodiments, the ALF antimicrobial peptide information to be optimized includes the LBD sequence to be optimized of the ALF antimicrobial peptide to be optimized; the generating module 602 is specifically configured to:

[0132] Perform multiple sequence alignment on the preset LBD sequence and the LBD sequence to be optimized, and generate a position frequency matrix based on the comparison results;

[0133] The partial least squares regression algorithm and gradient boosting decision tree algorithm were used to construct multiple linear regression and regression trees for the optimized ALF antimicrobial peptide information, position frequency matrix, and antimicrobial index results, to determine the first contribution value of each sequence site to the antimicrobial index results and obtain the amino acid substitution matrix;

[0134] The interval distribution frequency of positively charged amino acids in the target preset LBD sequence is calculated to obtain an interval distribution frequency matrix;

[0135] Performing Monte Carlo simulation on the first contribution value in the amino acid substitution matrix and the lysine substitution contribution value in the interval distribution frequency matrix to determine the second contribution value of the lysine site to the antibacterial index result;

[0136] The first contribution value of the lysine site is replaced by the second contribution value to obtain the first fitness matrix;

[0137] A first candidate LBD sequence set is obtained through the first fitness matrix, including the number of times each candidate LBD sequence is generated and the candidate physicochemical property information; wherein, the first candidate LBD sequence set includes the number of times each candidate LBD sequence is generated and the candidate physicochemical property information, and the target preset LBD sequence is an LBD sequence in the preset LBD sequence whose antibacterial index meets the first preset standard.

[0138] In some embodiments, the analysis module 603 is specifically configured to:

[0139] Calculate the Pearson correlation coefficient between the target physicochemical property data and the target preset antibacterial index value;

[0140] If the Pearson correlation coefficient meets the third preset criterion, a two-dimensional and / or three-dimensional fuzzy cluster analysis is performed on the target physicochemical property data and the target preset antibacterial index value to obtain the optimal modified physicochemical property range of the LBD sequence to be optimized;

[0141] A set of candidate LBD sequences is generated based on the optimal modified physicochemical property range.

[0142] In some embodiments, the calculation module 604 is specifically configured to:

[0143] Obtain the physicochemical property data of any sequence in the candidate LBD sequence set and calculate the multidimensional Euclidean distance between the above value and the target physicochemical property;

[0144] Get the number of times each candidate LBD sequence is generated;

[0145] Determine the relative antibacterial activity of any sequence based on the multidimensional Euclidean distance;

[0146] The number of occurrences and relative antibacterial activity were normalized and summed with equal weights to obtain the evaluation results;

[0147] The candidate LBD sequence whose evaluation result meets the second preset standard is used as the first optimized LBD sequence to obtain a first optimized LBD sequence set.

[0148] In some embodiments, the optimization module 605 is specifically configured to:

[0149] Performing antibacterial verification on a portion of the first optimized LBD sequence to obtain a verified first sub-optimized LBD sequence;

[0150] A target sub-optimized LBD sequence exhibiting high antibacterial activity in the first sub-optimized LBD sequence is crossed with a second sub-optimized LBD sequence by a genetic algorithm to obtain a crossover progeny sequence; wherein the second sub-optimized LBD sequence is a sequence in the first optimized LBD sequence that has not been verified for antibacterial activity;

[0151] The first fitness matrix is ​​updated by the first optimized LBD sequence and the crossover offspring sequence to obtain a second fitness matrix;

[0152] Based on the second fitness matrix, a second optimization process is performed on the designated LBD sequence in the first optimized LBD sequence set to obtain a second optimized LBD sequence; wherein the designated LBD sequence is an LBD sequence whose antibacterial activity meets a fourth preset standard.

[0153] In some embodiments, the preset ALF antimicrobial peptides are 31 crustacean ALF antimicrobial peptides; the preset antimicrobial index is the minimum inhibitory concentration value and / or the minimum bactericidal concentration value; and the preset physicochemical property data is the sequence net charge and hydrophobicity.

[0154] In some embodiments, the apparatus may further comprise:

[0155] The prediction module uses the candidate LBD sequence whose evaluation results meet the second preset standard as the first optimized LBD sequence. After obtaining the first optimized LBD sequence set, the secondary structure of the first optimized LBD sequence is predicted using the PEP-FOLD3 method to obtain a polypeptide structure model.

[0156] The device provided in the embodiment of the present invention has the same implementation principle and technical effects as those in the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the system embodiment, reference can be made to the corresponding content in the aforementioned method embodiment.

[0157] An embodiment of the present invention provides an electronic device. Specifically, the electronic device includes a processor and a storage device. The storage device stores a computer program, and when the computer program is executed by the processor, it executes any one of the methods in the above embodiments.

[0158] Figure 7 This is a structural diagram of an electronic device provided in an embodiment of the present invention, which includes: a processor 701, a memory 702, a bus 703 and a communication interface 704, wherein the processor 701, the communication interface 704 and the memory 702 are connected via the bus 703; the processor 701 is used to execute an executable module stored in the memory 702, such as a computer program.

[0159] The memory 702 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage. The system network element communicates with at least one other network element via at least one communication interface 704 (which may be wired or wireless), and may utilize the Internet, a wide area network, a local area network, a metropolitan area network, or the like.

[0160] The bus 703 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 7 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0161] Among them, the memory 702 is used to store programs, and the processor 701 executes the program after receiving the execution instruction. The method executed by the device for flow process definition disclosed in any embodiment of the above-mentioned embodiment of the present invention can be applied to the processor 701 or implemented by the processor 701.

[0162] The processor 701 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor 701 or by software instructions. The above processor 701 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present invention can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 702, and processor 701 reads the information in memory 702 and performs the steps of the above method in conjunction with its hardware.

[0163] The computer program product of the readable storage medium provided in the embodiment of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method in the previous method embodiment. The specific implementation can be referred to the previous method embodiment and will not be repeated here.

[0164] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0165] Finally, it should be noted that the above embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for optimizing ALF antimicrobial peptide, characterized in that: Applied to the ALF antimicrobial peptide optimization algorithm model; the method comprises: Obtaining preset ALF antimicrobial peptide information and ALF antimicrobial peptide information to be optimized; wherein the preset ALF antimicrobial peptide information includes a preset LBD sequence of the preset ALF antimicrobial peptide, a preset antimicrobial index corresponding to the preset LBD sequence, and preset physicochemical property data; the ALF antimicrobial peptide information to be optimized includes the LBD sequence to be optimized of the ALF antimicrobial peptide to be optimized; A multiple sequence alignment is performed on the preset LBD sequence and the LBD sequence to be optimized, and a position frequency matrix is ​​generated based on the comparison results; multiple linear regression and regression tree construction are performed on the ALF antimicrobial peptide information to be optimized, the position frequency matrix and the antimicrobial index result by using the partial least squares regression algorithm and the gradient boosting decision tree algorithm to determine the first contribution value of each sequence site to the antimicrobial index result, and obtain an amino acid substitution matrix; the interval distribution frequency of positively charged amino acids in the target preset LBD sequence is calculated to obtain an interval distribution frequency matrix; the first contribution value in the amino acid substitution matrix and the lysine substitution contribution value in the interval distribution frequency matrix are subjected to Monte Carlo simulation to determine the second contribution value of the lysine site to the antimicrobial index result; the first contribution value of the lysine site is replaced by the second contribution value to obtain a first fitness matrix, and a first candidate LBD sequence set is obtained through the first fitness matrix; wherein, the first candidate LBD sequence set includes the number of times each candidate LBD sequence is accompanied by generation information and candidate physicochemical property information; Performing Pearson correlation coefficient calculation and fuzzy cluster analysis on the target physicochemical property data and target antimicrobial index corresponding to the target preset LBD sequence to obtain a range of physicochemical properties for the cluster of sequences with high antimicrobial activity, and generating a second set of candidate LBD sequences from the first set of candidate LBD sequences based on this range; wherein the target preset LBD sequence is an LBD sequence among the preset LBD sequences whose antimicrobial index meets the first preset standard; calculating a multidimensional Euclidean distance value between the candidate physicochemical property of each candidate LBD sequence in the second candidate LBD sequence set and the target physicochemical property, and evaluating the antibacterial activity of each candidate LBD sequence in the second candidate LBD sequence set based on the multidimensional Euclidean distance value and the number of times each candidate LBD sequence is generated, and selecting the candidate LBD sequence whose evaluation result meets the second preset standard as the first optimized LBD sequence, to obtain a first optimized LBD sequence set; The first fitness matrix is ​​updated based on the first optimized LBD sequence and the offspring sequence to obtain a second fitness matrix, and the first optimized LBD sequence set is optimized using the second fitness matrix to obtain a second optimized LBD sequence set; wherein the offspring sequence is generated by the first optimized LBD sequence through genetic algorithm calculation.

2. The method according to claim 1, characterized in that The target preset LBD sequence is an LBD sequence whose antibacterial index meets the first preset standard among the preset LBD sequences.

3. The method according to claim 1, characterized in that The target physicochemical property data and target antibacterial index corresponding to the target preset LBD sequence are calculated by Pearson correlation coefficient and fuzzy cluster analysis to obtain the physicochemical property range of the high antibacterial activity sequence cluster, and a second candidate LBD sequence set is generated from the first candidate LBD sequence set based on the range, including: Calculate the Pearson correlation coefficient between the target physicochemical property data and the target preset antibacterial index value; If the Pearson correlation coefficient meets the third preset criterion, a two-dimensional and / or three-dimensional fuzzy cluster analysis is performed on the target physicochemical property data and the target preset antibacterial index value to obtain the optimal modified physicochemical property range of the LBD sequence to be optimized; A candidate LBD sequence set is generated based on the optimal modified physicochemical property range.

4. The method according to claim 1, wherein Calculating a multidimensional Euclidean distance value between the candidate physicochemical property of each candidate LBD sequence in the second candidate LBD sequence set and the target physicochemical property, and evaluating the antibacterial activity of each candidate LBD sequence in the second candidate LBD sequence set based on the multidimensional Euclidean distance value and the number of times each candidate LBD sequence is generated, and taking the candidate LBD sequence whose evaluation result meets the second preset standard as the first optimized LBD sequence, to obtain a first optimized LBD sequence set, including: Obtaining the physicochemical property data of any sequence in the candidate LBD sequence set, and calculating the multidimensional Euclidean distance between the above value and the target physicochemical property; Get the number of times each candidate LBD sequence is generated; determining the relative antibacterial activity of any one of the sequences according to the multidimensional Euclidean distance; Normalizing the number of occurrences and the relative antibacterial activity, taking equal weights and summing them to obtain an evaluation result; The candidate LBD sequence whose evaluation result meets the second preset standard is used as the first optimized LBD sequence to obtain a first optimized LBD sequence set.

5. The method according to claim 1, wherein The updating of the first fitness matrix based on the first optimized LBD sequence and the offspring sequence to obtain a second fitness matrix, and optimizing the first optimized LBD sequence set by using the second fitness matrix to obtain a second optimized LBD sequence set, includes: performing antibacterial verification on a portion of the first optimized LBD sequence to obtain a verified first sub-optimized LBD sequence; Crossing a target sub-optimized LBD sequence exhibiting high antibacterial activity with a second sub-optimized LBD sequence among the first sub-optimized LBD sequences using a genetic algorithm to obtain a crossover progeny sequence; wherein the second sub-optimized LBD sequence is a sequence among the first optimized LBD sequences that has not undergone antibacterial verification; The first fitness matrix is ​​updated using the first optimized LBD sequence and the crossover offspring sequence to obtain a second fitness matrix; Based on the second fitness matrix, a second optimization process is performed on the designated LBD sequence in the first optimized LBD sequence set to obtain a second optimized LBD sequence; wherein the designated LBD sequence is an LBD sequence whose antibacterial activity meets a fourth preset standard.

6. The method according to claim 1, wherein The preset ALF antimicrobial peptides are 31 crustacean ALF antimicrobial peptides; the preset antimicrobial index is the minimum inhibitory concentration value and / or the minimum bactericidal concentration value; and the preset physicochemical property data is the sequence net charge and hydrophobicity.

7. The method according to claim 1, characterized in that After taking the candidate LBD sequence whose evaluation result meets the second preset standard as the first optimized LBD sequence to obtain the first optimized LBD sequence set, the method further includes: The secondary structure of the first optimized LBD sequence was predicted using the PEP-FOLD3 method to obtain a polypeptide structure model.

8. An ALF antimicrobial peptide optimization device, characterized in that: Applied to the ALF antimicrobial peptide optimization algorithm model; the device comprises: an acquisition module, configured to acquire preset ALF antimicrobial peptide information and ALF antimicrobial peptide information to be optimized; wherein the preset ALF antimicrobial peptide information includes a preset LBD sequence of the preset ALF antimicrobial peptide, a preset antimicrobial index corresponding to the preset LBD sequence, and preset physicochemical property data; and the ALF antimicrobial peptide information to be optimized includes the LBD sequence to be optimized of the ALF antimicrobial peptide to be optimized; A generation module is configured to perform a multiple sequence alignment on the preset LBD sequence and the LBD sequence to be optimized, and generate a position frequency matrix based on the comparison results; perform multiple linear regression construction and regression tree construction on the ALF antimicrobial peptide information to be optimized, the position frequency matrix, and the antimicrobial index result using a partial least squares regression algorithm and a gradient boosting decision tree algorithm to determine a first contribution value of each sequence site to the antimicrobial index result, and obtain an amino acid substitution matrix; calculate the interval distribution frequency of positively charged amino acids in the target preset LBD sequence to obtain an interval distribution frequency matrix; the target preset LBD sequence is an LBD sequence in the preset LBD sequence whose antimicrobial index meets a first preset standard; perform Monte Carlo simulation on the first contribution value in the amino acid substitution matrix and the lysine substitution contribution value in the interval distribution frequency matrix to determine a second contribution value of the lysine site to the antimicrobial index result; replace the first contribution value of the lysine site with the second contribution value to obtain a first fitness matrix, and obtain a first candidate LBD sequence set through the first fitness matrix; wherein the first candidate LBD sequence set includes each candidate LBD sequence with information on the number of times it is generated and candidate physicochemical property information; an analysis module, configured to calculate a Pearson correlation coefficient and perform fuzzy cluster analysis on the target physicochemical property data and target antimicrobial index corresponding to the target preset LBD sequence, obtain a range of physicochemical properties for clusters of sequences with high antimicrobial activity, and generate a second set of candidate LBD sequences from the first set of candidate LBD sequences based on this range; wherein the target preset LBD sequence is an LBD sequence among the preset LBD sequences whose antimicrobial index meets the first preset standard; a calculation module, configured to calculate a multidimensional Euclidean distance value between the candidate physicochemical property of each candidate LBD sequence in the second candidate LBD sequence set and the target physicochemical property, and evaluate the antibacterial activity of each candidate LBD sequence in the second candidate LBD sequence set based on the multidimensional Euclidean distance value and the number of times each candidate LBD sequence is generated, and select the candidate LBD sequence whose evaluation result meets the second preset standard as the first optimized LBD sequence, to obtain a first optimized LBD sequence set; An optimization module is used to update the first fitness matrix based on the first optimized LBD sequence and the offspring sequence to obtain a second fitness matrix, and optimize the first optimized LBD sequence set using the second fitness matrix to obtain a second optimized LBD sequence set; wherein the offspring sequence is generated by the first optimized LBD sequence through genetic algorithm calculation.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for forecasting antimicrobial activity of antimicrobial peptide and antimicrobial peptide

    CN104036155A

  • Multi-label learning based activity prediction method for antibacterial peptide

    CN104484580A