Multi-task polypeptide conformation optimization method based on hidden space alignment and reinforcement learning

CN122842684APending Publication Date: 2026-09-29XIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611170351.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-04
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

训练样本未做双向优质筛选,理化差异极大的工况样本分布偏移,网络泛化能力衰减,并且网络未绑定二面角立体化学约束,易输出非法构象,额外边界修复提升复杂度;且迁移概率固定不变,收敛慢的任务缺乏外源优质构象、早熟任务易被冗余迁移破坏稳定种群

Benefits of technology

(1)降低跨任务负迁移,降低多肽生产残次品率。传统方法直接复制不同环境的二面角构象,忽略温、pH、离子强度带来的折叠机理差异,易造成构象能量激增,导致实际生产中多肽松散聚集、活性丧失,产生大量残次产物。本发明通过VAE实现多环境构象隐空间特征对齐,仅迁移通用折叠规律而非机械复制数值,从根源规避负迁移;同时内置二面角物理约束,输出构象天然合规、结构稳定,贴合多肽天然活性构型,可有效筛选优质生产工况,提升成品纯度与有效产量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842684A_ABST
    Figure CN122842684A_ABST
Patent Text Reader

Abstract

This invention discloses a multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning, specifically including the following steps: Step 1, constructing a multi-task optimization mathematical model for peptide folding; Step 2, initializing the multi-task population, constructing an initial DE population, historical parameter storage matrix, elite archive set, reinforcement learning Q-table, and encoder cache space for each task; Step 3, constructing a variational autoencoder and using solution features to achieve alignment in the latent space; Step 4, for each peptide folding task under each physicochemical environment, adaptively updating the cross-task transfer probability RMP based on Q-learning, and periodically initiating cross-task knowledge transfer using the variational autoencoder; Step 5, setting an FE threshold to trigger adaptive population reduction, eliminating individuals and storing them in the archive set. This method can achieve multi-task peptide folding optimization that balances cross-environmental transfer adaptability, population conformation diversity, and stereochemical constraint compliance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of multi-task evolutionary optimization and computational biology interdisciplinary technology, and relates to a multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning. Background Technology

[0002] In three major industrial and scientific research scenarios—mass production of biopharmaceutical peptides, preparation of recombinant peptides in vitro, and prediction of target molecular structures—the simulation of 30-residue short peptide folding requires simultaneous control of three types of process environment parameters: temperature, pH, and ionic strength. Peptide chain folding patterns differ significantly under different physicochemical conditions. Low-temperature neutral solutions tend to form compact globular conformations, extreme acids and alkalis cause loose unfolding of peptide chains, and high ionic strength exacerbates the aggregation of hydrophobic residues. The optimal dihedral angle distribution and lowest free energy conformation corresponding to each environment are not interchangeable, making this a core problem in in vitro peptide preparation and peptide structure prediction.

[0003] Current mainstream methods employ dedicated force field optimization algorithms for single environments, searching for low free energy dihedral combinations only under fixed temperature or pH conditions. This approach has significant drawbacks: environmental parameters exhibit coupling effects, the optimal conformation for a single operating condition cannot be reused in other scenarios, and a complete re-iteration is required after environmental fine-tuning or changes in process buffer ratios, resulting in high computational costs. Furthermore, the model has weak robustness against molecular thermal noise and cannot adapt to dynamic experimental perturbations. Some studies have customized dedicated optimizers for single operating conditions, but their scenario adaptability is extremely poor, folding experience cannot be shared between different environments, and repeated debugging costs are high during process iteration and multi-condition batch simulations.

[0004] Multi-task differential evolution can simultaneously model multiple environmental folding subtasks, achieving conformational experience sharing. However, existing algorithms suffer from two major drawbacks: first, direct transfer without adaptation constraints easily leads to negative transfer, such as directly introducing a high-ionic-strength clustered conformation into an acidic, low-ionic environment, resulting in a significant increase in free energy due to electrostatic mismatch; second, the lack of diversity control methods leads to severe population homogenization in the later stages of iteration, making it prone to getting trapped in local optima and losing metastable natural folding configurations. Latent space transfer learning based on neural networks can alleviate negative transfer, but existing solutions still have shortcomings. The training samples are not subjected to bidirectional high-quality screening, resulting in a distribution shift of samples with vastly different physicochemical properties, weakening the network's generalization ability. Furthermore, the network is not bound to dihedral angles. Stereochemical constraints easily lead to the output of illegal conformations, and additional boundary repair increases complexity. Furthermore, the transfer probability remains constant, resulting in slow-converging tasks lacking high-quality external conformations, and premature tasks being susceptible to redundant transfers that disrupt the stable population. In addition, molecular thermal noise interferes with network feature extraction, causing the transferred conformations to deviate from the true optimal solution.

[0005] In summary, existing technologies struggle to simultaneously achieve parallel solution across multiple environments, secure negative transfer knowledge sharing, population diversity maintenance, and stereochemical compliance, further limiting the practical application of optimization algorithms in molecular simulation and in vitro peptide preparation scenarios. Therefore, designing a multi-task peptide folding optimization model that balances cross-environmental adaptability, population conformational diversity, and stereochemical compliance remains a critical challenge that urgently needs to be addressed in the field of biomolecular computing. Summary of the Invention

[0006] The purpose of this invention is to provide a multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning. The optimization model provided by this method can achieve multi-task peptide folding optimization that takes into account cross-environmental migration adaptability, population conformation diversity, and stereochemical constraint compliance in biopharmaceutical peptide mass production, in vitro recombinant peptide preparation, and target molecular structure prediction processes.

[0007] The technical solution adopted in this invention is a multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning, which specifically includes the following steps: Step 1: Construct a multi-task optimization mathematical model for peptide folding; Step 2, multi-task population initialization, constructing the initial DE population, historical parameter storage matrix, elite archive set, reinforcement learning Q table and encoder cache space for each task; Step 3: Construct a variational autoencoder and use the solution features to achieve alignment in the latent space; Step 4: For each peptide folding task under each physicochemical environment, the cross-task transfer probability RMP is adaptively updated based on Q-learning, and the variational autoencoder is periodically activated for cross-task knowledge transfer. Step 5: Set the FE threshold to trigger adaptive population reduction and eliminate individuals and store them in the archive set.

[0008] The invention is further characterized by: The specific process of step 1 is as follows: Step 1.1, define the total number of tasks T, and generate the following environmental triplet corresponding to the T groups through uniform sampling using Latin hypercube (LHS):

[0009] Temperature: pH: Ionic strength: ; Step 1.2, direct the 60-dimensional dihedral vector to a 30×3-dimensional vector. Transformation of atomic three-dimensional coordinate matrices, with decision variables as 60-dimensional vectors:

[0010] in, and These are the dihedral angles of the main chain for each residue; Step 1.3: Construct a fitness function to transform the folding stability of peptides under different environments into a problem of minimizing free energy.

[0011] in, This is the dihedral canonical term, which represents the environmental disturbance offset from the natural dihedral angle.

[0012] In step 1.3, the sum of the squared errors of the current dihedral angle and the ambient modulated target dihedral angle is calculated, specifically as follows: First, calculate the environmental modulation intensity:

[0013] Target reference dihedral angle after environmental disturbance:

[0014] Regular energy term:

[0015] in, The environmental disturbance modulation intensity coefficient for the mission; Euclidean distance between any two non-adjacent residues:

[0016] The specific rules are as follows:

[0017] Environmental penalties:

[0018] The thermal noise term is a Gaussian random perturbation that varies with temperature.

[0019] in, It is a standard normal Gaussian random number.

[0020] The specific process of step 2 is as follows: Step 2.1, define the DE individual structure, where each decision variable is a 60-dimensional variable storing 30 residues sequentially. The dihedral task populations in Random initialization, initial free energy Constraint violation value CV=inf; Step 2.2: Establish and initialize the auxiliary storage structure. Random sampling uses a fixed global random seed, and the archive set... An empty array stores high-quality historical conformation solutions that have been iteratively eliminated; a parameter memory pool. Stores the F and CR parameters of successful evolutions in the past, and sets the Q-learning parameters to be independent for each task. ; Step 2.3: Instantiate fully connected encoder and decoder networks independently for the current task, fix hyperparameters such as the number of network neurons, hidden layer dimensions, and activation functions, allocate network weights and bias memory space, and start the variational autoencoder network algebra. This allows migration to occur after the population has converged.

[0021] The specific process of step 3 is as follows: Step 3.1: Construct an independent encoder for each peptide folding optimization task t, with the dihedral angle as input and the mean value of the latent space distribution parameters as output. and logarithmic variance The latent space vector is obtained by reparameterization:

[0022] in, For Hadama multiplication element by element, It is three-dimensional standard normal noise. Given a covariance identity matrix, the mean can be obtained by multiplying each element. By superimposing random perturbations of the standard deviations of each corresponding dimension, VAE reparameter sampling is completed to obtain the latent vector. ; For each task t, an independent decoder is constructed, with H-dimensional latent variables as input and the reconstructed dihedral angle, the reconstructed dihedral angle vector, and the original input as output. With the same dimension and physical meaning, the network structure is symmetrical to the encoder; Step 3.2 For each task t, select the top K optimal solutions to form a high-quality set. For each high-quality solution x, evaluate it across all tasks:

[0023] Candidate solution conditions: ; Step 3.3: Design the loss function and perform latent space alignment; Step 3.4, Adaptive density peak clustering.

[0024] The specific process of step 3.3 is as follows: The VAE encoder maps solutions from different tasks to the same low-dimensional latent space. Assuming there are Q samples, a loss function is used to force the alignment of cross-task peptide folding knowledge in the latent space for "multi-task-friendly solutions". The total loss function is:

[0025] Reconstruction loss:

[0026] in, The dihedral angle reconstructed by the decoder; The original true dihedral angle of the encoder at the very beginning; KL divergence transforms the latent space into a standard normal distribution, defined as:

[0027] Comparative loss:

[0028] in, For positive sample pairs, These are negative sample pairs.

[0029] The specific process of step 3.4 is as follows: Calculate local density:

[0030] Where p represents the p nearest samples to j; Calculate relative distance:

[0031] Iterate through all samples j with a density higher than the current sample i, and take the Euclidean distance from the sample j closest to i as the distance to i. ; Final clustering decision value:

[0032] Before selection Each cluster is used as a cluster center, and smaller clusters with fewer than 4 samples are filtered out.

[0033] The specific process of step 4 is as follows: Step 4.1: For each task, based on the historical parameter pools MF and MCR, generate F using the Cauchy distribution and CR using the normal distribution, ensuring the parameters are within the effective range. Determine the mutation strategy based on the cross-task migration probability RMP(t). If migration is not triggered, use the classic current-to-pbest / 1 mutation strategy.

[0034] in, This is the optimal individual for this task. For random individuals within this task, i.e., only individuals from the population of this task are used to construct mutation vectors. If migration is triggered, source task c is randomly selected from other tasks, and elite individuals from the source task and random individuals are extracted to construct mutation vectors. Information from mature folded fragments that have converged under other environmental conditions is introduced into the current offspring generation:

[0035] Information on mature folding configurations under other physical and chemical environments is introduced into the current offspring generation to achieve cross-task knowledge injection, and finally binary crossover is performed:

[0036] Randomly determine whether to inherit the mutation vector components dimension by dimension to generate candidate offspring. And perform midpoint correction on decision variables that are outside the defined domain; Step 4.2: Adaptively update RMP through reinforcement learning; Step 4.3, Periodically triggered safe cross-task knowledge transfer module, which simultaneously meets the FE start threshold in a fixed period of each iteration.

[0037] The specific process of step 4.2 is as follows: Step 4.2.1: Calculate the migration success rate SR and the population diversity Div in sequence, and map the continuous (SR, Div) to 10 discrete states. Step 4.2.2: Design reinforcement learning actions, action set There are three actions: decreasing RMP, maintaining RMP, and increasing RMP. The action that maximizes the current Q value is selected during the phase to maximize the real-time cumulative reward.

[0038] The reward is set as follows:

[0039] The formula for calculating the Q value is:

[0040] Where s represents the state and a represents the action. For learning rate, For the next completely new state The maximum Q value among all available actions represents the optimal long-term return, and [] represents the time-series difference error.

[0041] The specific process of step 4.3 is as follows: Step 4.3.1: Randomly select two different elite original solutions from the currently validated valid clusters. The latent vector is obtained by feeding it into the encoder specific to the target task j. Then, latent space linear interpolation fusion is performed, and the interpolation coefficients are... Randomly select values ​​and linearly mix the two latent vectors:

[0042] Finally, the latent vector is input into the decoder of task j to obtain the original spatial reference solution base; Mutation generates new candidate individual vectors, and the new individual generation operator is:

[0043] in, This is the individual with the best fitness on the intrinsic target task j of this cluster. To randomly select two distinct individuals within a cluster The direction of the principal component of the cluster.

[0044] The beneficial effects of this invention are as follows: (1) Reduce negative migration across tasks and decrease the defect rate in peptide production. Traditional methods directly replicate dihedral conformations in different environments, ignoring the differences in folding mechanisms caused by temperature, pH, and ionic strength. This can easily lead to a surge in conformational energy, resulting in loose aggregation and loss of activity of peptides in actual production, producing a large number of defective products. This invention achieves alignment of latent space features of multi-environment conformations through VAE, migrating only the general folding rules rather than mechanically replicating values, thus avoiding negative migration at its source. At the same time, it incorporates dihedral physical constraints, outputting conformations that are naturally compliant and structurally stable, conforming to the natural active conformations of peptides. This can effectively screen for high-quality production conditions and improve the purity and effective yield of the finished product.

[0045] (2) Multi-task parallel optimization significantly shortens the drug process development cycle. Traditional single-environment simulation requires iterative processing, resulting in high computational redundancy and long development cycles for pharmaceutical companies when screening buffer formulations and optimizing production processes in batches. This invention supports parallel optimization under multiple operating conditions, improves search efficiency by relying on the reuse of latent space knowledge, and combines Q-learning to dynamically adjust the transition probability, adaptively accelerates convergence lag conditions, protects mature and high-quality conditions, and simplifies the computation by using a mid-to-late stage population reduction strategy. This allows for the rapid acquisition of multiple sets of optimal physicochemical process parameters, significantly reducing simulation costs and meeting the needs of high-throughput process screening in industry.

[0046] (3) Strong environmental adaptability, suitable for non-ideal industrial production conditions. In actual peptide preparation and mass production processes, temperature, buffer pH, and salt concentration fluctuate slightly. Traditional fixed-parameter algorithms are prone to conformational instability and process failure, requiring repeated debugging of experimental parameters. This invention relies on reinforcement learning to automatically adapt to the entire range of physicochemical environments without manual parameter tuning. It can quickly respond to minor environmental disturbances and maintain a stable low free energy conformation, improving the fault tolerance of the production process and adapting to dynamic experiments and continuous industrial production scenarios.

[0047] (4) High versatility, reducing the cost of secondary development of biomedical simulations. The framework of this invention is not limited to 30-residue peptides. It can be adapted to the simulation of folding of various biomolecules such as functional peptides, industrial enzymes, and targeted peptides through simple parameter adjustment. It is compatible with multi-task expansion and multi-optimization algorithm integration. It can be reused in multiple scenarios such as drug molecule modification and process iteration optimization without the need to repeatedly build simulation models. It has excellent engineering scalability and reusability.

[0048] (5) The visualization results are intuitive and interpretable, bridging the simulation and experimental production links. The present invention is equipped with a complete visualization analysis module, which can output three-dimensional conformation, hydrophobic interaction, structural compactness and convergence energy index. It can not only support the scientific research analysis of folding mechanism, but also quantitatively distinguish the working condition differences between high-quality and high-yield conformations and inactivated and defective conformations, accurately guide the optimization of buffer ratio and fermentation temperature control, and realize the direct application of simulation data to guide experimental and industrial production.

[0049] (6) Excellent resistance to thermal noise, and simulation results closely match real production effects. The actual folding process is subject to molecular thermal disturbances and interference from solution impurities. Traditional ideal simulation results deviate significantly from actual production, easily leading to problems such as optimal simulation but inactivated peptides and low yields in mass production. This invention optimizes the VAE noise loss function, effectively suppressing random environmental disturbances. It can still output stable and adapted conformations under harsh conditions of high temperature and strong disturbances, significantly narrowing the gap between simulation and actual experiments, and improving the reliability of algorithm engineering implementation. Attached Figure Description

[0050] Figure 1 This is a flowchart illustrating the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning of this invention.

[0051] Figure 2(a) shows the 3D backbone of the optimal peptide structure for Task 1 in the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning of this invention. Figure 2(b) shows the 3D skeleton of the optimal peptide structure for Task 2 in the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning of this invention. Figure 2(c) shows the 3D backbone of the optimal peptide structure for Task 3 in the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning of this invention. Figure 2(d) shows the 3D backbone of the optimal peptide structure for Task 4 in the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning of this invention. Figure 2(e) shows the 3D backbone of the optimal peptide structure for Task 5 in the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning of this invention. Figure 3(a) is a heatmap of contact intensity of residues in Task 1 in the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning of the present invention. Figure 3(b) is a heatmap of contact intensity of residues in Task 2 in the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning of this invention; Figure 3(c) is a heatmap of contact intensity of residues in Task 3 in the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning of this invention. Figure 3(d) is a heatmap of contact intensity of residues in Task 4 in the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning of this invention. Figure 3(e) is a heatmap of contact intensity of residue 5 in the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning of this invention. Figure 4(a) shows the centroid distance between structural domains in Task 1 of the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning in this invention. Figure 4(b) shows the centroid distance between structural domains in Task 2 of the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning in this invention. Figure 4(c) shows the centroid distance between domains in Task 3 of the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning in this invention. Figure 4(d) shows the centroid distance between domains in Task 4 of the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning in this invention. Figure 4(e) shows the centroid distance between structural domains in Task 5 of the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning in this invention. Figure 5 This is a bar chart of the minimum free energy of each task in the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning of this invention; Figure 6 This is a three-dimensional scatter plot of environmental parameters and free energy in the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning in this invention. Figure 7 This is a schematic diagram of the overall process of the peptide folding problem in the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning of this invention. Detailed Implementation

[0052] The following detailed description is provided in conjunction with specific implementation methods.

[0053] Example 1 This invention relates to a multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning, the process of which is as follows: Figure 1 As shown, this paper optimizes the multi-task optimization problem of peptide folding in multiple environments by combining variational autoencoder latent space alignment technology with Q-learning reinforcement learning to dynamically adjust cross-task transfer probabilities. The specific steps include: constructing the peptide folding multi-task optimization problem; initializing the multi-task population and shared parameters; multi-task differential evolution and cross-task knowledge transfer based on VAE latent space alignment; a clustering method based on adaptive density peaks; an adaptive population reduction strategy; and outputting and visualizing the optimal conformation scheme. The specific implementation is as follows: Example 2 Step 1: Construct a multi-task optimization mathematical model for peptide folding, defining task parameters, decision variables, three-dimensional coordinate transformation rules, and a fitness function that minimizes free energy.

[0054] Step 2: Multi-task population initialization, constructing the initial DE population, historical parameter storage matrix, elite archive set, reinforcement learning Q-table and encoder cache space for each task.

[0055] Step 3: Construct a variational autoencoder and use the solution features to achieve alignment in the latent space.

[0056] Step 4: Multi-task differential evolution iteration. For each peptide folding task under each physicochemical environment, the cross-task transfer probability RMP is adaptively updated based on Q-learning, and the variational autoencoder is periodically activated for cross-task knowledge transfer.

[0057] Step 5: Set the FE threshold to trigger adaptive population reduction, reduce the population size, and eliminate individuals and store them in the archive set.

[0058] Example 3 The specific process of step 1 is as follows: Step 1.1, Multi-task basic parameter definition: Total number of tasks T=50, corresponding to 50 groups of environmental triples generated by uniform sampling using Latin hypercube (LHS):

[0059] Among them, temperature pH: ; Ionic strength To ensure the comprehensiveness of the simulation scenario, the total number of amino acid residues is 30. The entire peptide chain is artificially divided into three independent structural domains: Domain 1 corresponds to residues 1-10, Domain 2 corresponds to residues 11-20, and Domain 3 corresponds to residues 21-30. Each amino acid residue contains... With two main chain dihedrals, the global decision variable dimension is 60, and the population size N for each task is 100.

[0060] Step 1.2: Construct a specific function to implement a 60-dimensional dihedral vector into a 30×3-dimensional vector. The transformation of the three-dimensional coordinate matrix of atoms, with the coordinate unit being angstroms.

[0061] The decision variables are 60-dimensional vectors:

[0062] in, and These represent the dihedral angles of the main chain for each residue. A dihedral angle is an angle that rotates around a chemical bond and is geometrically periodic. A complete periodic interval is selected. This allows for the complete capture of all independent and non-repeating spatial torsion states. The decision variables are then transformed into 30 amino acid residues. The three-dimensional spatial coordinates (X, Y, Z) of a carbon atom. The standard covalent bond length between atoms is 3.8 Å, and the standard bond angle of the peptide chain is... Initialize the global coordinate system: the first residue The second residue The iterative calculation rule is to start from the third residue, construct a local orthogonal coordinate system based on the first two residues, and utilize the torsion angle. The spatial deflection is calculated by iteratively solving for the three-dimensional coordinates of all residues point by point. Using this coordinate transformation, the hydrophobic contact energy of the residues can be calculated. This module further assists in the complete solution of the free energy objective function. It serves as a crucial bridge connecting abstract optimization variables with the actual spatial conformation of peptides and is also a necessary prerequisite module for calculating residue contact energies.

[0063] Step 1.3: Construct a fitness function to transform the folding stability of peptides under different environments into a problem of minimizing free energy.

[0064] in, The dihedral regularization term is the natural dihedral angle offset by environmental disturbance. Solve for the sum of squared errors between the current dihedral angle and the target dihedral angle modulated by the environment.

[0065] First, calculate the environmental modulation intensity:

[0066] Target reference dihedral angle after environmental disturbance:

[0067] Regular energy term:

[0068] in, This is the environmental perturbation modulation intensity coefficient for the task. This coefficient is a fundamental weighting factor used to uniformly define the upper limit of perturbation to the intrinsic baseline dihedral angle of the peptide under different physicochemical environments. Combined with normalized temperature and pH deviations, it jointly determines the deviation of the target's optimal dihedral angle relative to the global fixed baseline under the current environment. Higher temperatures and greater pH deviations from neutral 7 result in higher deviations. The larger the value, the greater the relative offset of the natural target dihedral angle to the reference under this environment, and the more the simulated physicochemical environment changes the inherent torsional preference of amino acids.

[0069] The energy is the hydrophobic contact energy of the residues. A residue spacing of less than 0.6 nm indicates that the hydrophobic residues are close to each other, forming a hydrophobic core, which improves the thermodynamic stability of the conformation and reduces the energy negatively, making the optimization target better. A residue spacing between 0.6 nm and 1 nm indicates that the hydrophobic interaction gradually decreases, the energy increases linearly, and the conformational stability gradually deteriorates. A spacing of 1 nm or more indicates that the residues are too far apart and there is no effective hydrophobic interaction, so this energy does not contribute.

[0070] Euclidean distance between any two non-adjacent residues:

[0071] The specific rules are as follows:

[0072] Environmental penalties:

[0073] The greater the pH deviation from neutral, the more severe the abnormal charging of amino acid side chains. Electrostatic repulsion disrupts stable folded structures, leading to a secondary amplification of energy penalties. Higher solution ionic strength alters intermolecular forces due to ion shielding effects, continuously applying a fixed proportion of positive energy penalties. Extreme acidity / alkalinity and high ionic strength environments naturally possess higher baseline total free energy.

[0074] The thermal noise term is a Gaussian random perturbation that varies with temperature.

[0075] in, It uses standard normal Gaussian random numbers to simulate random Brownian heat perturbations in molecules; temperature The higher the value, the greater the noise weighting coefficient, the more intense the molecular thermal motion at high temperatures, the stronger the random fluctuations in conformation, which is completely consistent with the physical laws of liquid-phase molecular dynamics.

[0076] Example 4 The specific process of step 2 is as follows: Step 2.1, Definition of Differential Evolution (DE) Individual Structure. Decision variables are all 60-dimensional, storing 30 residues sequentially. The dihedral task populations in Random initialization: initial free energy Objs=inf, constraint violation value CV=inf, initial differential evolution parameters F=0.5, CR=0.5.

[0077] Step 2.2: Establish and initialize the auxiliary storage structure. Random sampling uses a fixed global random seed to ensure reproducibility. Each of the 50 polypeptide folding tasks undergoes independent population initialization. The initial conformational solutions are uniformly distributed within the decision space and are distinct from each other, thus ensuring sufficient diversity of the polymorphic structure search starting points under each environmental condition. Archive Set An empty array stores high-quality historical conformation solutions that have been iteratively eliminated; a parameter memory pool. The length is 100, storing the F and CR parameters of successful evolutions in history, initially all set to 0.5; the Q-learning parameter is set to be independent for each task. This is used for cross-task transfer probability reinforcement learning updates. Ten state rows correspond to ten polypeptide folding evolution states discretely defined based on population success rate and conformational diversity. Three action columns correspond to three adjustment strategies for the cross-task transfer probability (RMP). The iterative process updates the Q-value based on the population optimization reward and adaptively adjusts the RMP in real time to control the frequency of cross-task knowledge transfer. The variational autoencoder training sample cache pool has a maximum cache size of 500 samples.

[0078] Step 2.3: Select the number of elites K=5 as the global hyperparameter of the algorithm, and adjust the population reduction timing coefficient. It also instantiates fully connected encoder and decoder networks independently for the current task, fixes hyperparameters such as the number of network neurons, hidden layer dimensions, and activation functions, allocates network weights and bias memory space, and sets the startup algebra for the variational autoencoder (VAE) network. The network is started later to avoid negative migration in the early stages, ensuring that migration occurs after the population converges. The network execution interval is 100, and the dimension of latent variables is 3.

[0079] Example 5 The specific process of step 3 is as follows: Step 3.1, Network Structure Initialization. An independent encoder is constructed for each peptide folding optimization task t, with a 60-dimensional dihedral input and the mean value of the latent space distribution parameters as the output. and logarithmic variance The network structure is as follows: Input layer (60) Fully connected layer (64) ReLU Fully connected layer (64) ReLU Output layer (2H). The latent space vector is obtained using reparameterization:

[0080] in, For Hadama multiplication element by element, It is three-dimensional standard normal noise. This is the covariance identity matrix. Multiplying them element-wise yields the mean. By superimposing random perturbations of the standard deviations of each corresponding dimension, VAE (Variational Autoencoder) reparameter sampling is completed to obtain the latent vector. .

[0081] For each task t, an independent decoder is constructed, with H-dimensional latent variables as input and the reconstructed 60-dimensional dihedral vector as output. The reconstructed dihedral vector and the original input... With the same dimension and physical meaning, the network structure is symmetrical to the encoder.

[0082] Step 3.2, Network Training Data Preparation. For each task t, select the top K=5 optimal solutions to form a high-quality set. These high-quality solutions represent the search directions for locally stable conformations of the peptide under these environmental conditions. For each high-quality solution x, evaluation is performed on all tasks: Candidate solution conditions: In other words, x is retained only if it is friendly to at least two tasks. These solutions are neither special fold states that are overfitted to a certain task, nor single environment-specific conformations that lack universality, thus providing high-quality cross-task shared samples for subsequent latent space alignment.

[0083] Step 3.3, Loss Function Design and Latent Space Alignment. A VAE encoder maps solutions from different tasks to the same low-dimensional latent space. Assuming a total of Q samples, the loss function forces "multi-task-friendly solutions" to be closer together in the latent space, achieving effective alignment of cross-task peptide folding knowledge. The total loss function is:

[0084] Reconstruction loss:

[0085] in, The 60-dimensional dihedral angle reconstructed by the decoder; The original true dihedral angles are initially input into the encoder. The reconstruction loss weight is 0.8, prioritizing reconstruction accuracy and ensuring that the conformational details within a single task are not lost due to dimensionality reduction.

[0086] The KL (Kullback-Leibler) divergence transforms the latent space into a standard normal distribution, defined as:

[0087] Each task's individual latent distribution is constrained to be near the same standard normal distribution, laying the distributional foundation for all subsequent tasks to share the same latent space. The weight of this latent distribution is 0.05, which is sufficient to ensure distribution stability.

[0088] Comparative loss:

[0089] in, For positive sample pairs, These are negative sample pairs. Cosine similarity is used to obtain latent vector similarity. Temperature is a hyperparameter used to scale the similarity of latent vectors. Finally, the gradients of all network parameters are adjusted by minimizing the total loss function (Loss) to achieve alignment.

[0090] Step 3.4, Adaptive density peak clustering. First, calculate the local density:

[0091] Where p represents the p nearest samples to j. The closer the distance, the better. The larger the value of the term, the more likely a mature folding pattern exists in the region and is frequently sampled. Dividing by p is used for mean normalization to eliminate density magnitude differences caused by the number of nearest neighbors.

[0092] Calculate relative distance:

[0093] Iterate through all samples j with a density higher than the current sample i, and take the Euclidean distance from the sample j closest to i as the distance to i. . The large size indicates that the sample is far from other, higher-density clusters. Even if there are many high-density samples around it, it does not stick to the center of the higher-density cluster, indicating that the conformation itself has the potential to become an independent cluster center.

[0094] Final clustering decision value:

[0095] Before selection Clusters with fewer than four samples are selected as cluster centers and then filtered out. Each remaining cluster corresponds to a mature, frequently occurring, stable polypeptide folding conformation pattern. These patterns reflect a set of metastable conformations that polypeptides repeatedly tend to under specific environmental conditions, and are also structural knowledge units that should be prioritized for transfer in subsequent cross-task transfer. Conformations within the same cluster are highly clustered in the latent space, indicating that they share similar folding backbone features. Even if they come from different environmental conditions, they are very likely to correspond to similar domain layouts or folding topologies.

[0096] Example 6 The specific process of step 4 is as follows: Step 4.1, DE parameter generation and mutation crossover. For each task, based on the historical parameter pool MF and MCR, F is generated using a Cauchy distribution, and CR is generated using a normal distribution, ensuring that the parameters are within the effective range. The mutation strategy is determined based on the cross-task migration probability RMP(t). If no migration is triggered, the classic current-to-pbest / 1 mutation strategy is used.

[0097] in, This is the optimal individual for this task. This refers to random individuals within this task. That is, only individuals from the current task's population are used to construct the mutation vector. If migration is triggered, source task c is randomly selected from the remaining 49 tasks. Elite individuals from the source task and random individuals are extracted to construct the mutation vector. Information from mature folded fragments that have converged under other environmental conditions is introduced into the current offspring generation.

[0098] Information on mature folding configurations under other physical and chemical environments is incorporated into the current offspring generation to achieve cross-task knowledge injection. Finally, binary crossover is performed.

[0099] Randomly determine whether to inherit the mutation vector components dimension by dimension to generate candidate offspring. And perform midpoint repair on decision variables that are outside the defined domain.

[0100] Step 4.2: Adaptively update the RMP through reinforcement learning. First, calculate the migration success rate (SR). A higher SR indicates a higher probability of generating effective and superior offspring through mutation and crossover in the current generation, indicating strong local folding search efficiency under the current environment and that the population has adapted well to the current physicochemical conditions. Next, calculate the population diversity (Div). A higher Div indicates greater differences in dihedral configurations among individuals within the population, a wider search space coverage, and no premature convergence. Finally, map the continuous (SR, Div) to 10 discrete states. State 1 represents low SR and low Div, indicating slow folding search speed and weak local sampling capability under this physicochemical environment, urgently requiring the introduction of mature folding configurations from other tasks to break the search stagnation; State 10 represents high SR and high Div, indicating that the population can stably generate low free energy offspring conformations, with rich conformational diversity and sufficient search range, and the current task does not require frequent external migrations.

[0101] Next, reinforcement learning actions are designed, and action sets are generated. Three actions are set up to represent decreasing RMP, maintaining RMP, and increasing RMP, respectively. The Q-learning strategy is an exploratory strategy. During the exploration phase, actions are randomly tested to avoid the controller becoming rigid due to long-term fixed single control actions, thus adapting to abnormal conditions such as fluctuations in peptide folding difficulty under different physicochemical environments and sudden convergence stagnation in the later stages of iteration. During the utilization phase, the action with the highest current Q value is selected, reusing the optimal RMP control experience accumulated from previous iterations to maximize the immediate cumulative reward. An exploration rate is set. .

[0102]

[0103] The reward is set as follows:

[0104] In the early stages of iteration, the difference between the population's average free energy and the global optimum free energy is large, individual configurations are dispersed, and population diversity is sufficient. There is still ample room for exploration before convergence, and high rewards will inversely incentivize Q-learning to choose actions. Actively increasing the RMP (Redirected Learning Process) enhances the frequency of cross-task knowledge transfer, and a large number of mature dihedral configurations from other tasks are introduced to broaden the search scope and quickly reduce the overall free energy. In the later stages of iteration, the reward is small, and the population-average free energy is close to the optimal individual free energy. The vast majority of individuals converge to a stable, low-free-energy natural conformation, indicating that optimization is essentially complete. Low rewards incentivize the controller to choose certain actions. Actively reduce RMP (Random Mating Probability), significantly reduce interference from the influx of unfamiliar external conformations, stabilize the optimal solution, and avoid the degradation and oscillation of the converged population.

[0105] The formula for calculating the Q value is:

[0106] Where s represents the state and a represents the action. The learning rate and step size are set to 0.05, resulting in a small single Q-value correction amplitude. The Q-table values ​​iterate smoothly without violent oscillations, which is suitable for the steady-state optimization requirements of multi-generation continuous iteration of DE. The discount factor is set to 0.9, emphasizing future rewards and demonstrating a long-term perspective. For the next completely new state The maximum Q-value among all available actions represents the optimal long-term return. [] indicates the time-series difference error, representing the deviation between the historical estimated value and the actual long-term return. An error greater than 0 indicates that the old Q-value was underestimated, and the actual total return from executing this action is higher than expected, requiring an upward adjustment. An error less than 0 indicates that the old Q value was overestimated, and the actual benefit was not as expected, so it needs to be adjusted downwards. The difference between the two is used to correct the Q-table, gradually aligning its empirical values ​​with the actual optimization patterns of peptide folding. The ε exploration rate is 0.1, with 10% random exploration and 90% utilization of the current optimal strategy.

[0107] Step 4.3, the periodically triggered safe cross-task knowledge transfer module, is activated only when the FE threshold is met at a fixed iteration period. First, a transfer admission check is performed; transfer is only allowed if the average free energy of the conformation within the target cluster is better than the median free energy of the target task population. This fundamentally prevents negative transfer from inferior conformations, i.e., avoids erroneous folding configurations in one physicochemical environment contaminating the population search direction in another environment. A baseline solution (base) is generated based on VAE common latent space interpolation. The base is a high-quality fusion baseline conformation within the cluster obtained from latent space interpolation fusion, serving as the starting point for the entire mutation vector. It inherently possesses a mature folding structure, ensuring the rationality of the initial configuration of the new individual.

[0108] First, randomly select two different elite original solutions from the currently validated valid clusters. The latent vector is obtained by feeding it into the encoder specific to the target task j. Then, latent space linear interpolation fusion is performed, and the interpolation coefficients are... Randomly select values ​​and linearly mix the two latent vectors:

[0109] Finally, the latent vector is input into the decoder of task j to obtain the original spatial reference solution base.

[0110] Mutation generates new candidate individual vectors, and the new individual generation operator is:

[0111] in, The individual with the best fitness on the intrinsic target task j of this cluster. Two distinct individuals are randomly selected within a cluster. The first principal component vector is selected by performing PCA (Principal Component Analysis) on all decision variables of the cluster, representing the overall evolutionary trend of the cluster in the folded configuration space. This overall cluster configuration evolution trend is incorporated into the differential evolution operator, ensuring the search direction aligns with the population's cluster evolutionary patterns. Finally, the worst-performing individual is replaced with a new one.

[0112] Example 7 The specific process of step 5 is as follows: Step 5.1, determining the timing of reduction. When the number of evolutionary generations... Reaching the total number of algebras of When the ratio is reached, the population reduction mechanism is activated. At this time, a large amount of polypeptide folding computational resources have been invested in locating low free energy regions under various environmental conditions. Reducing the population size can concentrate the computational budget on the most promising conformational candidate solutions, thereby improving the efficiency of subsequent local optimization.

[0113] Step 5.2, Elite Individual Selection. For each task, select the current population based on constraint violation degree. and fitness Perform non-dominated sorting, select the first The best individuals form a new generation of population. These elite conformations represent the structural state with the highest folding stability under the environmental conditions, ensuring that the population as a whole evolves towards lower free energy.

[0114] Step 5.3, Archive Set Management. Eliminated individuals are stored in an archive set. If the archive size exceeds the original population size, N individuals are randomly selected and retained to ensure efficient use of storage space. The algorithm terminates when the total number of evaluations reaches maxFE.

[0115] The algorithm's final output includes a 3D conformational diagram of the task, a residue contact heatmap, a visualization of inter-domain distance histograms, a histogram of the task's minimum free energy, and a 3D scatter plot of environmental parameters and free energy. Specifically, it includes the following parts: A 3D backbone comparison diagram of the optimal peptide structure under different environmental conditions was generated. Figure 1The skeletal diagrams illustrate the optimal three-dimensional conformations of the polypeptide chain under different physicochemical conditions. As shown in Figure 2(a), at a temperature of 41.7°C, a pH of 5.6, and an ionic strength of 0.33 M, a clear stretching trend is observed, with the three domains loosely separated, especially domain III, which extends outward most prominently. This conformation reflects the destructive effect of a high-temperature, weakly acidic environment on the native structure of the protein, consistent with the conclusion in the paper that high temperature leads to the loss of the hydrogen bond network. Its free energy is approximately 230. As shown in Figure 2(b), at a temperature of 27.7°C, a pH of 7.3, and an ionic strength of 0.18 M, the most compact spherical conformation is observed, with the three domains tightly bound together and the smallest overall size, corresponding to the native structure under physiological conditions. The stable state has a free energy of approximately 210. As shown in Figure 2(c), at a temperature of 28.8°C, a pH of 8.6, and an ionic strength of 0.50 M, an asymmetric configuration is observed where domains I and II are relatively close together, and domain III is deflected outward. This indicates that the alkaline high-salt environment leads to asymmetric rearrangement between the structural domains, with a free energy of approximately 198. As shown in Figure 2(d), at a temperature of 32.1°C, a pH of 6.7, and an ionic strength of 0.12 M, the conformation is between the compact state in Figure 2(b) and the extended state of Task 1 in Figure 2(a). The three structural domains maintain basic folding but exhibit some loosening, with a free energy of approximately 204. As shown in Figure 2(e), at a temperature of 25.4°C, a pH of 7.6, and an ionic strength of 0.35 M, the conformation is highly similar to that in Figure 2(b), but the angle between domains II and III is slightly increased. The high ionic strength slightly affects the relative orientation between the domains, with a free energy of approximately 200.

[0116] Heatmaps of contact intensity of residues in some tasks were generated. Figure 3(a) shows that the heatmap for Task 1 shows a blank band where the residues at both ends still have contact, but the contact in the middle is completely lost. This indicates that the high-temperature environment mainly disrupted the interaction network in the middle of the polypeptide chain, while the residues at both ends still retained some local folding. Figure 3(b) shows that the heatmap for Task 2 shows nearly uniform high-intensity contact from residues 1 to 26, with all regions being dark, corresponding to the dense internal interaction network of the compact globular conformation. The contact intensity is the highest among the five groups. Figure 3(c) shows that the heatmap for Task 3 shows two independent contact clusters (residues 1-6 and 16-26), but the cross-contact between the two clusters is significantly weakened, indicating that the contact between domains is selectively destroyed under alkaline conditions, while the folding units within the domains remain intact. Figure 3(d) shows that the heatmap for Task 4 shows medium-intensity contact in most regions, but the intensity in the region of residues 12-18 is significantly reduced, reflecting the mild destabilizing effect of the weakly acidic and low-salt environment on the overall polypeptide chain. As can be seen from Figure 3(e), the heat map of Task 5 is highly similar to that of Figure 3(b), with a uniform dark color overall and even stronger contact in some areas, indicating that moderate ion intensity actually enhances local interactions, which corresponds perfectly to the lowest free energy of Task 5.

[0117] Figures 4(a) to 4(e) show that the centroid distance between domains exhibits a clear environment-dependent characteristic. In Figure 4(a), the centroid distance of domains I-III in Task 1 increases the most significantly, reaching the highest level among all groups. Domains I-II also show a significant increase, exhibiting a typical extended conformation. This is consistent with the conclusion that the hydrogen bond network is disrupted and the domains are loosely separated under high temperature and weak acid conditions. In Figure 4(b), the distances between domains in Task 2 are all at a low level. The three domains are tightly packed to form a compact spherical conformation, corresponding to the natural stable state under physiological conditions. In Figure 4(c), Task 3 shows a significant asymmetric inter-domain rearrangement: the centroid distance of domains I-II is the smallest among all groups, and the two remain closely close. However, the centroid distances of domains I-III and II-III show a significant asymmetric rearrangement. The free energy of the system is significantly increased. The outward deflection of domain III leads to an increase in its distance from the other two domains. This asymmetric structural change induced by the alkaline high-salt environment can reduce the free energy of the system. Figure 4(d) shows that the overall extension of Task 4 is similar to that of Task 1, both belonging to the high level. The distance between domains I and III is comparable to that of Task 1, both in the highest range. Moreover, the separation of domains I and II is more significant than that of Task 1. The three domains are more loosely structured. Figure 4(e) shows that the conformation of Task 5 is highly similar to that of Task 2. The centroid distance between each domain remains at a low level. The protein still maintains a good native folded state, indicating good structural stability in a neutral to alkaline environment close to physiological conditions.

[0118] Depend on Figure 5 It can be seen that the optimal free energy of most tasks can converge to a small value. Some columns are obviously too high, which may be because the task environment is more extreme, such as high temperature and high pH, ​​or the algorithm is not optimized enough for this task.

[0119] Depend on Figure 6 As can be seen from the graph, the transition of the points' colors from blue to yellow and then red is consistent with the overall trend of increasing temperature, pH deviating from neutral, and increasing ionic strength. This indicates that the algorithm successfully identified the impact of different environmental stresses on peptide stability; that is, extreme environments generally lead to an increase in free energy, reflecting a decrease in peptide folding stability. Simultaneously, the uniform coverage of environmental parameters and the continuous color change indicate that the Latin hypercube sampling is reasonable, and the algorithm did not exhibit obvious outliers or performance collapse due to negative migration. This confirms that the multi-task optimization framework can find a low free energy conformation suitable for each environment, and that the migration strategy effectively improves convergence quality while maintaining diversity.

[0120] Figure 7This is a schematic diagram illustrating the overall process of the multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning, as presented in this invention. It fully presents the end-to-end solution logic, from multi-source data input, problem modeling, multi-task iterative optimization to the final conformation output. Combining VAE cross-task knowledge transfer and reinforcement learning adaptive RMP updates, through evolutionary iterative optimization, the optimal conformation, three-dimensional structure, and minimum free energy of the peptide are finally output.

Claims

1. A multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning, characterized in that: Specifically, the steps include the following: Step 1: Construct a multi-task optimization mathematical model for peptide folding; Step 2, multi-task population initialization, constructing the initial DE population, historical parameter storage matrix, elite archive set, reinforcement learning Q table and encoder cache space for each task; Step 3: Construct a variational autoencoder and use the solution features to achieve alignment in the latent space; Step 4: For each peptide folding task under each physicochemical environment, the cross-task transfer probability RMP is adaptively updated based on Q-learning, and the variational autoencoder is periodically activated for cross-task knowledge transfer. Step 5: Set the FE threshold to trigger adaptive population reduction and eliminate individuals and store them in the archive set.

2. The multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning according to claim 1, characterized in that: The specific process of step 1 is as follows: Step 1.1, define the total number of tasks T, and generate the following environmental triplet corresponding to the T groups through uniform sampling using Latin hypercube (LHS): Temperature: pH: Ionic strength: , Step 1.2, direct the 60-dimensional dihedral vector to a 30×3-dimensional vector. Transformation of atomic three-dimensional coordinate matrices, with decision variables as 60-dimensional vectors: in, and These are the dihedral angles of the main chain for each residue; Step 1.3: Construct a fitness function to transform the folding stability of peptides under different environments into a problem of minimizing free energy. in, This is the dihedral canonical term, which represents the environmental disturbance offset from the natural dihedral angle.

3. The multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning according to claim 2, characterized in that: In step 1.3, the sum of the squared errors of the current dihedral angle and the ambient modulated target dihedral angle is calculated as follows: First, calculate the environmental modulation intensity: Target reference dihedral angle after environmental disturbance: Regular energy term: in, The environmental disturbance modulation intensity coefficient for the mission; Euclidean distance between any two non-adjacent residues: The specific rules are as follows: Environmental penalties: The thermal noise term is a Gaussian random perturbation that varies with temperature. in, It is a standard normal Gaussian random number.

4. The multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning according to claim 3, characterized in that: The specific process of step 2 is as follows: Step 2.1, define the DE individual structure, where each decision variable is a 60-dimensional variable storing 30 residues sequentially. The dihedral task populations in Random initialization, initial free energy Constraint violation value CV=inf; Step 2.2: Establish and initialize the auxiliary storage structure. Random sampling uses a fixed global random seed, and the archive set... An empty array stores high-quality historical conformation solutions that have been iteratively eliminated; a parameter memory pool. Stores the F and CR parameters of successful evolutions in the past, and sets the Q-learning parameters to be independent for each task. ; Step 2.3: Instantiate fully connected encoder and decoder networks independently for the current task, fix hyperparameters such as the number of network neurons, hidden layer dimensions, and activation functions, allocate network weights and bias memory space, and start the variational autoencoder network algebra. This allows migration to occur after the population has converged.

5. The multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning according to claim 4, characterized in that: The specific process of step 3 is as follows: Step 3.1: Construct an independent encoder for each peptide folding optimization task t, with the dihedral angle as input and the mean value of the latent space distribution parameters as output. Sum of logarithmic variance The latent space vector is obtained by reparameterization: in, For Hadama multiplication element by element, It is three-dimensional standard normal noise. Given a covariance identity matrix, the mean can be obtained by multiplying each element. By superimposing random perturbations of the standard deviations of each corresponding dimension, VAE reparameter sampling is completed to obtain the latent vector. ; For each task t, an independent decoder is constructed, with H-dimensional latent variables as input and the reconstructed dihedral angle, the reconstructed dihedral angle vector, and the original input as output. With the same dimension and physical meaning, the network structure is symmetrical to the encoder; Step 3.2 For each task t, select the top K optimal solutions to form a high-quality set. For each high-quality solution x, evaluate it across all tasks: Candidate solution conditions: ; Step 3.3: Design the loss function and perform latent space alignment; Step 3.4, Adaptive density peak clustering.

6. The multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning according to claim 5, characterized in that: The specific process of step 3.3 is as follows: The VAE encoder maps solutions from different tasks to the same low-dimensional latent space. Assuming there are Q samples, a loss function is used to force the alignment of cross-task peptide folding knowledge in the latent space for "multi-task-friendly solutions". The total loss function is: Reconstruction loss: in, The dihedral angle reconstructed by the decoder; The original true dihedral angle of the encoder at the very beginning; KL divergence transforms the latent space into a standard normal distribution, defined as: Comparative loss: in, For positive sample pairs, These are negative sample pairs.

7. The multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning according to claim 6, characterized in that: The specific process of step 3.4 is as follows: Calculate local density: Where p represents the p nearest samples to j; Calculate relative distance: Iterate through all samples j with a density higher than the current sample i, and take the Euclidean distance from the sample j closest to i as the distance to i. ; Final clustering decision value: Before selection Each cluster is used as a cluster center, and smaller clusters with fewer than 4 samples are filtered out.

8. The multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning according to claim 7, characterized in that: The specific process of step 4 is as follows: Step 4.1: For each task, based on the historical parameter pools MF and MCR, generate F using the Cauchy distribution and CR using the normal distribution, ensuring the parameters are within the effective range. Determine the mutation strategy based on the cross-task migration probability RMP(t). If migration is not triggered, use the classic current-to-pbest / 1 mutation strategy. in, This is the optimal individual for this task. For random individuals within this task, i.e., only individuals from the population of this task are used to construct mutation vectors. If migration is triggered, source task c is randomly selected from other tasks, and elite individuals from the source task and random individuals are extracted to construct mutation vectors. Information from mature folded fragments that have converged under other environmental conditions is introduced into the current offspring generation: Information on mature folding configurations under other physical and chemical environments is introduced into the current offspring generation to achieve cross-task knowledge injection, and finally binary crossover is performed: Randomly determine whether to inherit the mutation vector components dimension by dimension to generate candidate offspring. And perform midpoint correction on decision variables that are outside the defined domain; Step 4.2: Adaptively update RMP through reinforcement learning; Step 4.3, Periodically triggered safe cross-task knowledge transfer module, which simultaneously meets the FE start threshold in a fixed period of each iteration.

9. The multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning according to claim 7, characterized in that: The specific process of step 4.2 is as follows: Step 4.2.1: Calculate the migration success rate SR and the population diversity Div in sequence, and map the continuous (SR, Div) to 10 discrete states. Step 4.2.2: Design reinforcement learning actions, action set There are three actions: decreasing RMP, maintaining RMP, and increasing RMP. The action that maximizes the current Q value is selected during the phase to maximize the real-time cumulative reward. The reward is set as follows: The formula for calculating the Q value is: Where s represents the state and a represents the action. For learning rate, For the next completely new state The maximum Q value among all available actions represents the optimal long-term return, and [] represents the time-series difference error.

10. The multi-task peptide conformation optimization method based on latent space alignment and reinforcement learning according to claim 9, characterized in that: The specific process of step 4.3 is as follows: Step 4.3.1: Randomly select two different elite original solutions from the currently validated valid clusters. The latent vector is obtained by feeding it into the encoder specific to the target task j. Then, latent space linear interpolation fusion is performed, and the interpolation coefficients are... Randomly select values ​​and linearly mix the two latent vectors: Finally, the latent vector is input into the decoder of task j to obtain the original spatial reference solution base; Mutation generates new candidate individual vectors, and the new individual generation operator is: in, This is the individual with the best fitness on the intrinsic target task j of this cluster. To randomly select two distinct individuals within a cluster The direction of the cluster principal component.