Structure-friendly cutting method for macromolecular protein and application of structure-friendly cutting method in molecular docking
By employing multi-source protein feature scoring and hard boundary constraints, large protein molecules are scientifically cut, solving the problems of structural integrity and computational stability in existing technologies, and enabling the efficient application of large proteins in structure prediction and molecular docking.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 郑雪
- Filing Date
- 2026-02-07
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies struggle to scientifically cleave large protein molecules, resulting in insufficient structural integrity, computational stability, and reproducibility. This affects the accuracy and efficiency of protein structure prediction and molecular docking. Furthermore, the lack of standardized cleavage methods across different research platforms limits the application of large protein molecules in drug development.
A method based on multi-source protein feature scoring is adopted. By acquiring multi-dimensional feature data, standardizing and comprehensively scoring it, and combining hard boundary constraints and fragment length constraints, a scientific cutting of large proteins is achieved, generating protein fragments suitable for structure prediction and molecular docking.
It improves the scientific rigor and stability of large protein cleavage, enhances the accuracy and feasibility of structure prediction and molecular docking, ensures the integrity of protein domains, and is applicable to various types of large protein sequence cleavage.
Smart Images

Figure CN122050482A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computational biology and intelligent drug development technology, specifically to a structure-friendly cleavage method for macromolecular proteins, and the application of this method in protein structure prediction, molecular docking, and drug development computation. Background Technology
[0002] Protein structure information is a core foundation for understanding biological functions, conducting drug development, and studying molecular mechanisms of action. With the development of computational biology and artificial intelligence technologies, protein structure-based prediction models and molecular docking algorithms have been widely applied in various stages such as target discovery, lead compound screening, and drug optimization, and have improved drug development efficiency to a certain extent.
[0003] Large proteins are widely distributed in humans and other organisms, playing a crucial role in signal transduction, transcriptional regulation, immune regulation, and tumor-related pathways. They often contain multiple domains, disordered regions, low-complexity regions, and transmembrane structures, exhibiting highly heterogeneous sequence and structural features. However, current mainstream protein structure prediction and molecular docking models are primarily designed for medium-length proteins with relatively simple structures. These characteristics of large proteins make them difficult to directly use as complete inputs in existing computational models, posing a key technological bottleneck in intelligent drug development.
[0004] To address the aforementioned issues, existing technologies typically employ methods such as preprocessing or cleaving large protein molecules to obtain protein fragments suitable for computation. However, current cleavage methods largely rely on human experience, single-domain annotation information, or simple length truncation rules, lacking a systematic and comprehensive consideration of the multidimensional structure and sequence characteristics of proteins. This makes it difficult to guarantee the stability and reproducibility of the cleavage results in subsequent calculations.
[0005] Specifically, cutting methods based solely on domain annotations often ignore factors such as disordered regions, low-complexity regions, and the continuity of secondary structures, which can easily lead to unreasonable breaks in key structural regions, resulting in protein conformational damage. On the other hand, cutting based solely on length thresholds cannot reflect the true structural organization of the protein and can easily lead to unstable computational results.
[0006] Furthermore, existing technologies typically employ fixed, singular cutting rules, failing to dynamically adjust for the varying structural complexity of different proteins. This makes it difficult to strike a balance between maintaining protein structural integrity and meeting the input requirements of computational models. Consequently, in practical applications, the cut protein units exhibit significant fluctuations in prediction success rate, result consistency, and computational efficiency.
[0007] Crucially, existing technologies lack a unified and standardized input system for macromolecular protein computation. Different researchers or platforms often adopt their own cutting strategies and processing procedures, resulting in a lack of comparability between computational results and limiting the promotion of such methods in large-scale applications and industrialization scenarios.
[0008] Therefore, existing technologies for processing large protein structures cannot simultaneously ensure structural integrity, computational stability, and engineering reproducibility, which has become a major fundamental problem restricting the further application of protein structure prediction and molecular docking technologies.
[0009] Based on the above situation, we will develop a cutting method for macromolecular proteins that has multidimensional feature constraints and standardized output capabilities. This method will improve the adaptability and stability of protein structure prediction and molecular docking calculations while ensuring the rationality of protein structure, thereby meeting the actual needs of intelligent drug development for complex targets. Summary of the Invention
[0010] The purpose of this invention is to provide an automated method for cutting large proteins based on multi-source protein sequence and structural feature data. By uniformly quantifying and fusion scoring the multi-dimensional features of proteins, potential cleavage sites within the protein are identified, enabling the scientific cutting of large proteins and providing standardized input for structure prediction, molecular docking, and drug target research.
[0011] To achieve the above objectives, the present invention provides the following technical solution: Technical Solution 1: A method for cutting large proteins based on multi-source protein feature scoring, the method comprising the following steps: (1) Acquisition of protein characteristic data; (2) Standardization processing of feature data; (3) Calculation of comprehensive score; (4) Setting hard boundary constraints; (5) Screening of candidate cleavage sites; (6) Fragment length constraints and segment generation.
[0012] Technical Solution 2: The large protein cleavage method based on multi-source protein feature scoring as described in Technical Solution 1, wherein the large protein is a protein sequence with a length greater than 1200 amino acids, preferably a protein sequence with more than 2000 amino acid residues.
[0013] Technical Solution 3: The large protein segmentation method based on multi-source protein feature scoring as described in Technical Solution 1 or 2, wherein in step (1), when a protein lacks any of the above-mentioned feature data, the corresponding feature data is not used in subsequent calculations.
[0014] Technical Solution 4: A large protein cleavage method based on multi-source protein feature scoring according to any one of Technical Solutions 1-3, wherein, in step (1), information on protein disordered regions, transmembrane structures, domain annotations, specific amino acid enriched regions, low-complexity regions, and secondary structures are preferentially acquired and used simultaneously.
[0015] Technical Solution 5: A large protein segmentation method based on multi-source protein feature scoring according to any one of Technical Solutions 1-4, wherein, in step (2), the standardization processing method for continuous feature data is: within a single protein, min–max normalization is performed based on its minimum and maximum values.
[0016] Technical Solution 6: A large protein cutting method based on multi-source protein feature scoring according to any one of Technical Solutions 1-5, wherein, in step (2), the standardization process of binary feature data is as follows: first, a sliding window smoothing process is performed along the protein sequence direction to obtain continuous values, and then min–max normalization is performed within the protein.
[0017] Technical Solution 7: The large protein cutting method based on multi-source protein feature scoring as described in Technical Solution 6, wherein in the sliding window smoothing process, the window size is adaptively set according to the protein length, preferably 5-50 amino acid residues per window.
[0018] Technical Solution 8: A large protein cleavage method based on multi-source protein feature scoring according to any one of Technical Solutions 1-7, wherein, in step (3), the final score reflects the overall suitability of the amino acid position as a protein cleavage site.
[0019] Technical Solution 9: A large protein cleavage method based on multi-source protein feature scoring according to any one of Technical Solutions 1-8, wherein, in step (4), the specific requirements of the hard boundary constraint are: the cleavage site must not be located inside the structural domain, and any cleavage result must maintain the integrity of the structural domain.
[0020] Technical Solution 10: A large protein cutting method based on multi-source protein feature scoring according to any one of technical solutions 1-9, wherein, in step (6), the protein fragment length constraint condition is: the length of the protein fragment formed between two adjacent cutting sites does not exceed a preset maximum length threshold.
[0021] Technical Solution 11: The large protein cleavage method based on multi-source protein feature scoring as described in Technical Solution 10, wherein the preset maximum length threshold can be customized, and is preferably no more than 1200 amino acid residues.
[0022] Technical Solution 12: According to the large protein cutting method based on multi-source protein feature scoring described in Technical Solution 10 or 11, in step (6), the prerequisite for determining the final cutting scheme is: the length of the protein fragment is less than the preset maximum length; under the condition of satisfying the length constraint, the larger continuous protein fragments are retained as much as possible; and the hard boundary constraints of the structural domain are not violated.
[0023] Technical Solution 13: The large protein fragmentation method based on multi-source protein feature scoring provided by this invention can be applied to fields such as protein structure prediction, small molecule docking, virtual screening, and computer-aided drug design. The protein fragments generated by this method can be directly used as input to structure prediction or molecular docking models such as AlphaFold and DiffDock, improving the stability and prediction success rate of related calculation processes.
[0024] The beneficial effects of this invention are: This invention acquires and standardizes multi-source protein feature data, then combines it with equal-weighted average fusion to obtain a comprehensive score for the suitability of amino acid position cleavage. Simultaneously, it sets dual constraints of hard domain boundaries and fragment length to achieve precise and scientific cleavage of large protein sequences. This method ensures the integrity of the protein domains after cleavage and allows the generated protein fragments to meet the input requirements of downstream computational models. It effectively solves the problem that large protein sequences cannot be directly predicted for structure or molecular docking, improving the accuracy and feasibility of downstream calculations. Furthermore, it can flexibly adapt to missing feature data and is applicable to the cleavage of various types of large protein sequences.
[0025] Amino acid sequence information: The protein amino acid sequence described in this invention is a publicly available protein sequence, which may be derived from a public protein database (such as Uniprot). This protein sequence is only used as the object of processing in the method of this invention, and its specific sequence content does not constitute a limitation of this invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and technical effects of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. The embodiments described below are some embodiments of the present invention, but not all embodiments. In conjunction with the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] In this application, the terms "first aspect," "second aspect," "third aspect," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or quantity, nor should they be construed as implicitly indicating the importance or quantity of the indicated technical features. Moreover, "first aspect," "second aspect," etc., serve only as a non-exhaustive enumeration and should be understood not to constitute a closed limitation on quantity.
[0028] In this document, terms such as "preferred," "better," and "good" are merely descriptions of implementation methods or embodiments that achieve better results, and should be understood as not constituting a limitation on the scope of protection of this application. If multiple "preferred" terms appear in a technical solution, unless otherwise specified and there are no contradictions or mutual constraints, each "preferred" term shall be independent.
[0029] In a first aspect, in one specific embodiment of this application, this application provides a method for obtaining protein cleavage sites, characterized by comprising the following steps: (1) Acquisition of protein characteristic data; (2) Standardization processing of feature data; (3) Calculation of comprehensive score; (4) Setting hard boundary constraints; (5) Screening of candidate cleavage sites; (6) Fragment length constraints and segment generation.
[0030] In one specific embodiment of this application, step (1) of obtaining protein feature data in technical solution 1 includes, but is not limited to: (1) Information on disordered regions of proteins, which is continuous numerical data reflecting the flexibility of amino acid positions and structures; (2) Low-complexity region information, which is binary data representing the complexity of amino acid sequences; (3) Information on specific amino acid enrichment regions, which is continuous numerical data reflecting the degree of enrichment of specific amino acids or amino acid combinations in the sequence; (4) Transmembrane structure information, which is binary data, where the amino acid position corresponding to the transmembrane structure region is marked as 1, and the other positions are marked as 0; (5) Domain annotation information (domain FT), which is binary data, where the amino acid position in the annotated domain is marked as 1, and the non-domain region is marked as 0; (6) Secondary structure information, which is binary data, where a position in a regular secondary structure such as an α-spiral or β-fold is marked as 1, and a position not in a regular secondary structure such as an α-spiral or β-fold is marked as 0.
[0031] In the specific implementation process, if a target protein lacks any of the above-mentioned feature data, then that feature data will not be processed or used in subsequent steps.
[0032] In some embodiments of this application, step (2) feature data standardization processing specifically includes: a. Continuous feature data: Within a single protein, the feature is normalized using the minimum and maximum values across the entire sequence using the min–max method; b. Binary feature data: First, smoothing is performed along the protein sequence using a sliding window method to convert the binary signal into a continuous value. Then, the smoothed data is normalized by min–max within the protein.
[0033] Through the above processing, multiple protein feature data from different sources and at different scales can be fused and calculated within the same numerical range.
[0034] In some embodiments of this application, step (3) of comprehensive score calculation specifically includes: For each amino acid position in the protein sequence, all available and normalized feature data corresponding to that position are summarized, and the feature data are fused using an equal-weighted average method to obtain a comprehensive score for that amino acid position, which is recorded as the final score.
[0035] The final score is used to characterize the overall suitability of the amino acid position as a potential protein cleavage site.
[0036] In some embodiments of this application, step (4) of setting hard boundary constraints specifically includes: When the target protein contains domain annotation information (domain FT), the start and end positions of the domains are treated as uncrossable hard boundary constraints, prohibiting the setting of cleavage sites inside the domains during subsequent cleavage processes, thereby ensuring the integrity of the known domains.
[0037] When the target protein does not contain domain annotation information, the above hard boundary constraints are not set.
[0038] In some embodiments of this application, step (5) candidate cleavage site screening specifically includes: Based on the results obtained in step (3), and under the premise of satisfying the hard boundary constraints in step (4), all amino acid positions in the protein sequence are sorted according to the comprehensive score, and several positions with higher scores are selected as candidate cleavage sites.
[0039] In some embodiments of this application, step (6) of segment length constraint and tangent point selection specifically includes: a. Length constraint: Based on the candidate cleavage sites, a protein fragment length constraint is introduced to limit the length of the protein fragment formed between any two adjacent cleavage sites to no more than a preset maximum length threshold. b. Cutting point selection: Under the premise of satisfying the length constraints and hard boundary constraints of the structural domains, the combination of cutting sites that can form larger continuous protein fragments is selected first to determine the final cutting scheme and generate a set of protein fragments suitable for downstream structure prediction or molecular docking calculation.
[0040] In some embodiments of this application, according to technical solution 2, the large protein is a protein sequence with a length greater than 1200 amino acid residues, preferably a protein sequence with a length greater than 2000 amino acid residues.
[0041] In a second aspect, the present invention also provides a large protein cleavage system based on multi-source protein feature scoring, the system being used to implement the large protein cleavage method as described in the first aspect, characterized in that it includes: (1) Feature data acquisition module: used to acquire any one or more feature data of the target protein at the amino acid sequence level; (2) Feature data standardization processing module: used to standardize or normalize the feature data according to its data type, so that feature data from different sources and with different dimensions can be uniformly integrated; (3) Comprehensive score calculation module: It is used to summarize all or part of the standardized feature data corresponding to each amino acid position in the protein sequence, and fuse them in an equal weighted average manner to obtain the comprehensive score of the amino acid position; (4) Hard boundary constraint module: When there is domain annotation information in the target protein sequence, the domain boundary is used as an insurmountable cutting hard constraint condition; (5) Candidate cleavage site screening module: under the premise of satisfying the hard boundary constraints, sorting the amino acid positions in the protein sequence according to the comprehensive score, and selecting the positions with higher scores as candidate cleavage sites; (6) Cutting generation module: Based on the candidate cutting sites, it introduces protein fragment length constraints and generates the final protein cutting scheme under the premise of meeting the preset constraints, thereby obtaining a set of protein fragments suitable for downstream structure prediction or molecular docking calculation.
[0042] In some implementations, the above modules can be implemented by software or by a combination of software and hardware. The system can automatically cut large protein sequences without relying on human experience rules.
[0043] Thirdly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to perform the large protein cleavage method based on multi-source protein feature scoring as described in the first aspect.
[0044] In some embodiments, the computer-readable storage medium includes, but is not limited to, a hard disk, a solid-state memory, a mobile storage device, or other media capable of storing program instructions, which are used to implement steps such as protein feature data acquisition, feature data standardization processing, comprehensive score calculation, hard boundary constraint setting, and cutting generation.
[0045] The beneficial effects of the present invention will be further illustrated below through specific embodiments.
[0046] Example 1 1. Construction of large protein datasets In this embodiment, human proteins annotated in the UniProt database are used as the research subjects.
[0047] All human protein sequence data were downloaded from the UniProt database. Proteins with more than 1200 amino acid residues were selected as large protein research objects, resulting in 5161 large protein sequences. All protein sequences were saved in FASTA format for subsequent feature calculation and cleavage processing.
[0048] 2. Acquisition of multi-source protein characteristic information For the 5161 large human protein sequences obtained in step 1, the following protein feature information was obtained or calculated: a. Low Complexity Region Information: Based on the annotation information of the corresponding protein in the UniProt database, low-complexity regions are extracted and annotated to form a binary sequence labeled by amino acid position, where the low-complexity region position is marked as 1 and the other positions are marked as 0.
[0049] b. Secondary region information: Based on the secondary structure annotation information provided by UniProt, the positions corresponding to α-helices and β-sheets in the protein are obtained and converted into binary sequences labeled by amino acid sites. Positions with secondary structures are marked as 1, and other positions are marked as 0.
[0050] c. Transmembrane region information: Based on the transmembrane segment annotation information in UniProt, a binary sequence of transmembrane structures is constructed, with the position of the transmembrane structure marked as 1 and the other positions marked as 0.
[0051] d. Domain annotation information (Domain FT): Based on UniProt's feature table (FT), protein domain segment information is extracted and a binary domain sequence is constructed, where positions inside the domain are marked as 1 and non-domain regions are marked as 0.
[0052] e. Information on specific amino acid enrichment regions (PG-rich): A sliding window scan of the protein sequence was performed using a Python program. The window size was set to 50 amino acid residues. The relative abundance of proline (P) and glycine (G) was calculated in each window, and the calculation results were mapped to the corresponding amino acid positions to form a continuous PG-rich score sequence.
[0053] f. Information about protein disorder regions: The IUPred2A tool was used to predict disordered regions for each protein sequence, resulting in a continuous disordered scoring sequence output by amino acid residues.
[0054] 3. Standardization processing of feature data The different types of feature data obtained in step 2 are processed as follows: a. For continuous feature data (including unordered region scores and PG-rich scores), standardization is performed within each protein using the min–max normalization method; b. For binary feature data (including low complexity, transmembrane structures, domain annotations and secondary structure information), first perform sliding window smoothing along the protein sequence direction to obtain continuous values, and then perform min–max normalization within a single protein. If a protein lacks any of the above-mentioned feature data, the corresponding feature will not be included in the comprehensive score calculation for that protein.
[0055] 4. Comprehensive scoring and cut-off point calculation For each amino acid position in the protein sequence, all available and normalized feature data for that position are summarized and fused using an equal-weighted average method to obtain a comprehensive score for that position, which is recorded as the final score.
[0056] During the calculation, if the protein has domain annotation information, the domain boundary is set as a hard constraint condition to prevent the generation of cleavage sites inside the domain, ensuring that the integrity of the domain is not compromised.
[0057] Based on the final score, the positions in the protein sequence that satisfy the hard boundary constraints are sorted, and the top 20 cleavage sites with the highest scores are selected as candidate cleavage sites.
[0058] 5. Protein cleavage and fragment generation Based on the candidate cleavage sites, a protein fragment length constraint is introduced, limiting the length of the protein fragment formed between adjacent cleavage sites to no more than 1200 amino acid residues.
[0059] Under the constraints of domain integrity and fragment length, the original protein FASTA sequence was cleaved into multiple contiguous protein fragments. After cleavage, the resulting protein fragments were saved as independent FASTA files for subsequent structure prediction or molecular docking calculations.
[0060] Example 2 To verify the effectiveness and rationality of the large protein cleavage method based on multi-source protein feature scoring described in this invention in real drug target scenarios, this embodiment selects a variety of human large proteins with experimentally verified binding sites as test objects. The predicted cleavage sites obtained by the method of this invention are compared and analyzed with the results of existing mainstream structure prediction or molecular docking models and the actual binding site range recorded in literature or databases.
[0061] 1. Test Objects and Methods Description The test proteins selected in this embodiment are all from the UniProt database, and the file format is FATSA. Their UniProt IDs include Q13315, P78527, Q15413, P21817, P00533, Q9UM73, P06213, and P08269. These proteins are all common, relatively long (>1200 amino acid residues) multi-domain proteins with very typical binding pocket characteristics. Directly inputting the entire protein into structure prediction or molecular docking models results in computational instability. The corresponding molecules have all been validated experimentally or in the literature.
[0062] For the above proteins, the key structural segments were identified using the following three methods.
[0063] a. AlphaFold3 prediction results Based on AlphaFold3's prediction of protein structure and molecular docking structure, the optimal three-dimensional coordinates of the corresponding molecules were obtained.
[0064] b. DiffDock Prediction Results First, AlphaFold3 is used to predict the structure of the protein, and then DiffDock is used to predict the molecular docking between the protein and small molecules, and the optimal three-dimensional coordinates of the molecule.
[0065] c. Calculate the small molecule binding segment The distances between all protein residues and small molecules were calculated using Python, with amino acid residues less than 5 Å being designated as small molecule binding sites.
[0066] d. Actual (control) locus data The actual sites are derived from protein functions or small molecule binding regions identified in existing experimental studies, literature reports, or databases, and are used as validation benchmarks.
[0067] The comparison between the predicted results of each protein and the actual site range is shown in Table 1.
[0068] The formula for calculating coverage is: Coverage = (AlphaFold3 prediction interval ∪ DiffDock prediction interval) ∩ Actual locus ÷ Actual locus length
[0069] 3. Results Analysis As shown in Table 1, the predicted cleavage segments involved in the method of this invention have a significant overlap with the actual known binding sites in most tested proteins, and the predicted segments as a whole cover the main functional regions of the actual sites. Although in some proteins, the predicted segments given by DiffDock show obvious deviations or fall into non-functionally related segments, the prediction results of AlphaFold3 remain accurate.
[0070] It is worth noting that while differences remain in the predicted sites regarding segment integrity and boundary rationality, all predicted protein sites contain over 90% of the critical residues. The prediction of critical residues is more important than the overall segment integrity because critical residues directly participate in direct contact between small molecules and ligands; hydrogen bonds, salt bridges, hydrophobic interactions; and conformational stability or catalytic reactions. This has already helped in finding more accurate drug targets for unknown proteins.
[0071] This invention introduces hard boundary constraints for structural domains and a multi-feature comprehensive scoring mechanism, which can more stably locate functionally related segments while ensuring the integrity of structural domains. This allows the protein fragments obtained from the cleavage to retain the actual functional site information to the maximum extent while maintaining controllable length.
[0072] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made to it in form and detail without departing from the scope of data collection and evaluation as defined in the claims of the present invention. Attached Figure Description
[0073] Figure 1 A flowchart illustrating a method according to an embodiment of the present invention is shown schematically.
Claims
1. A structure-friendly cleavage method for macromolecular proteins, characterized in that, Includes the following steps: (1) Obtain the amino acid sequence of the target macromolecular protein; (2) Obtain at least one protein structure-related feature data of the target macromolecular protein; (3) Standardize the protein structure-related feature data; (4) Based on the standardized protein structure-related feature data, the positions of each amino acid in the amino acid sequence are scored to obtain the suitability score of each amino acid position as a protein cleavage site. (5) Set hard constraints for cleavage based on the boundaries of protein domains and screen candidate cleavage sites that meet the hard constraints; (6) Under the premise of meeting the protein fragment length constraint, the target macromolecular protein is cut according to the candidate cleavage site to generate a protein fragment for downstream calculation.
2. The method according to claim 1, characterized in that: The macromolecular protein is a protein sequence with a length greater than 1200 amino acid residues.
3. The method according to claim 1, characterized in that: The protein structure-related feature data includes one or more of the following: protein disorder region information, domain annotation information, low-complexity region information, transmembrane structure information, secondary structure information, and specific amino acid enrichment region information.
4. The method according to claim 1, characterized in that: When the target macromolecular protein lacks structural feature data related to a certain type of protein, the corresponding feature data is not used in the scoring process.
5. The method according to claim 1, characterized in that: In step (3), the continuous feature data are standardized using a minimum-maximum normalization method within a single protein.
6. The method according to claim 1, characterized in that: In step (3), the binary feature data is first smoothed along the protein sequence direction and then normalized.
7. The method according to claim 1, characterized in that: In step (4), the suitability score is obtained by weighted or unweighted fusion of standardized protein structure-related feature data.
8. The method according to claim 1, characterized in that: In step (6), the length of the protein fragment formed between adjacent cleavage sites does not exceed the preset maximum length threshold.
9. The method according to claim 8, characterized in that: The maximum length threshold is no more than 1200 amino acid residues.
10. A protein cleavage system for implementing the method according to any one of claims 1 to 9, characterized in that, include: The module includes a feature data acquisition module, a feature data processing module, a scoring calculation module, a constraint control module, and a cutting generation module.