Protein language and flow matching-based antibody sequence light and heavy chain co-design method and evaluation system

The generation of antibody sequences through protein language and flow matching models solves the problem of insufficient semantic information in antibody drug development, realizes efficient and low-cost antibody design and evaluation, and improves the speed and quality of antibody generation.

CN120496620APending Publication Date: 2025-08-15LIANGZHU LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510479392.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

There are problems in the development of existing antibody drugs, such as shallow semantic information, low efficiency, high cost and imperfect evaluation system of computer-aided design methods, resulting in insufficient speed and quality of artificial antibodies development.

Method used

Using a method based on protein language and flow matching, light and heavy chain sequence embedding tensors are extracted through the protein language model, combined with the flow matching vector field prediction model to generate antibody sequences, and a multi-dimensional evaluation system is constructed, including evaluation units such as uniqueness, authenticity, diversity, and physical and chemical properties.

Benefits of technology

It has achieved high-quality, high-speed and low-cost antibody sequence generation co-designed by light and heavy chains, which has improved the speed of antibody generation by more than twice, and provided a comprehensive evaluation system to improve the reliability and diversity of antibody sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496620A_ABST
    Figure CN120496620A_ABST
Patent Text Reader

Abstract

The invention discloses a protein language and flow matching-based antibody sequence light and heavy chain co-design method and an evaluation system, and belongs to the crossing field of artificial intelligence and biotechnology. The method comprises the following steps: preprocessing antibody sequence data, inputting the preprocessed antibody sequence data into a protein language model coder, extracting light and heavy chain sequence embedded tensors, and inputting the tensors into a decoder for training after statistical analysis and normalization; a stream matching vector field prediction model is constructed based on the embedded tensor and the time step output by the decoder, light chain sequence embedding and heavy chain sequence embedding are generated through random noise sampling and ordinary differential equation solving, and then a final antibody sequence is output through decoding; and comprehensively evaluating the generated antibody sequence by using a matched evaluation system, and assisting in screening high-quality antibodies. According to the method, collaborative design and multi-dimensional evaluation of light and heavy chains are realized, the problems of shallow protein semantic information, high calculation cost, imperfect evaluation system and the like of a traditional method are solved, and a new way is provided for efficiently developing functional antibodies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the intersection of artificial intelligence and biotechnology, and in particular to a method and evaluation system for co-designing light and heavy chains of antibody sequences based on protein language and flow matching. Background Art

[0002] Antibodies (Abs) are large, Y-shaped proteins secreted by plasma cells (effector B cells) and used by the immune system to identify and neutralize foreign substances such as bacteria, viruses, and toxins. Antibodies specifically bind to these foreign substances (called antigens), helping the immune system eliminate these potential threats. Antibodies are typically composed of two heavy chains and two light chains, each of which is divided into a constant region and a variable region, connected by disulfide bonds. The complementarity determining regions (CDRs) are specific regions within the variable region of the antibody molecule that are responsible for binding to antigens (such as pathogens, toxins, etc.).

[0003] Antibodies not only play an important role in natural immune responses, but are also widely used in biology and medicine, for example in vaccine development, immune diagnosis and detection, immunotherapy, and the development of antibody drugs. However, the development and application of artificial antibody drugs are still subject to many limitations. For example, traditional antibody discovery requires time- and resource-intensive screening of large immune or synthetic libraries, which often requires a large investment of manpower and material resources and is inefficient. At the same time, most current artificial antibody development is limited to the design of CDRs and does not distinguish between light and heavy chains, ignoring considerations based on the rationality of the antibody as a whole. This can result in lead candidates with suboptimal binding and poor developability properties.

[0004] Furthermore, the current main AI-based antibody design methods lack a unified and comprehensive computer-assisted antibody evaluation standard or system. This results in computer-assisted antibody design being heavily dependent on the synthetic results of actual biochemical experiments, and thus has not effectively reduced the burden on biomedical researchers. Therefore, in order to improve the speed and quality of artificial antibody research and development, there is an urgent need to develop computer-assisted full-chain antibody design and evaluation methods.

[0005] In recent years, in the field of artificial intelligence, non-white regression deep generative models, represented by flow-matching methods, have demonstrated strong performance and competitiveness due to the similarity between their generation logic and the underlying logic of protein sequence structure. Flow-matching methods significantly reduce the number of generation steps and optimize the probability path by matching the optimal probability path between two data distributions, providing solid technical support for the efficient and high-quality design of protein sequences. At the same time, with the development of large language models in the field of protein science, some advanced protein language models, such as ESM-2, can re-extract and condense the semantic information of discrete amino acid sequences, forming a more difficult and efficient semantic space. However, current methods have yet to introduce this efficient protein language model semantic space.

[0006] In summary, there is an urgent need to develop a method for designing and evaluating the entire antibody sequence to achieve co-design of light and heavy chains, high-quality, high-speed, and low-cost artificial antibody sequence design, to solve the problem of shallow semantic information in current computer-aided protein design methods, and to provide high-quality candidate antibody sequences and antibody sequence evaluation methods for biological and medical antibody research. Summary of the Invention

[0007] The purpose of the present invention is to provide a method and evaluation system for the co-design of light and heavy chains of antibody sequences based on protein language and flow matching, so as to achieve co-design of light and heavy chains, high-quality, high-speed and low-cost artificial antibody sequence design, and solve the problem of shallow semantic information in current computer-aided protein design methods.

[0008] To achieve the above-mentioned purpose of the invention, the embodiment provides a method for co-designing light and heavy chains of antibody sequences based on protein language and flow matching, comprising the following steps:

[0009] Obtain antibody sequence data and perform preprocessing to generate sequence data of the light chain and heavy chain of the paired antibody;

[0010] The light chain and heavy chain sequence data of the paired antibody are input into the encoder of the protein language model, and the embedded tensors of the output light chain and heavy chain sequences are statistically analyzed and normalized to obtain the embedded data of the light chain and heavy chain sequences;

[0011] Input the embedded data and statistical information into the decoder of the protein language model for training to form a light chain decoder and a heavy chain decoder;

[0012] The embedding tensor and time step output from the encoder are used as input to train the flow matching model and form a flow matching vector field prediction model;

[0013] The random sampling noise and sampling time input stream are matched to the vector field prediction model to obtain the vector field corresponding to the sampling time. The antibody sequence embedding tensor corresponding to the time step is obtained from the vector field by solving the ordinary differential equation, and the final light chain sequence and heavy chain sequence are generated using the light chain decoder and heavy chain decoder.

[0014] In one embodiment, antibody sequence data is collected from the Observed Antibody Space database.

[0015] In one embodiment, the antibody sequence data includes paired antibody data and unpaired antibody data.

[0016] In one embodiment, the pre-processing comprises:

[0017] Pairing, used to filter out paired antibody data from antibody sequence data;

[0018] Clustering is used to cluster similar sequences of the screened paired antibody data and remove highly similar sequences;

[0019] Screening length is used to screen out antibody sequences that are too long or too short in the paired antibody data. The lower limit of the length threshold is 2 amino acids, the upper limit of the light chain length threshold is ≥148 amino acids, and the upper limit of the heavy chain length threshold is ≥149 amino acids;

[0020] Screening of amino acid species to eliminate uncommon amino acids in antibody sequences;

[0021] Length padding is used to pad the antibody sequences of uneven lengths in the paired antibody data to the same sequence length.

[0022] Furthermore, the similarity threshold for clustering is 95% and the coverage threshold is 80%.

[0023] Furthermore, the length is padded by using the ESM-2 built-in filler to fill all heavy chains to 149 positions and light chains to 148 positions.

[0024] In one embodiment, the protein language model is an ESM-2 protein language model with 8M parameters trained on the UniRef50 dataset.

[0025] In one embodiment, statistical information is obtained by calculating the average value and standard deviation of the overall embedding tensors of the output light chain and heavy chain sequences, and the embedding tensors of the output light chain and heavy chain sequences are normalized using the Z-score method based on the statistical information to obtain embedded data.

[0026] In one embodiment, the structure of the decoder of the protein language model includes two linear layers, an activation layer, and a detokenize layer;

[0027] The embedded data and statistical information are sequentially transformed through the first linear layer, activation layer, and second linear layer to transform the size of the embedded data. The token information is then extracted through the detokenize layer to obtain an embedding tensor for a pair of light chain and heavy chain sequences, and form a light chain decoder and a heavy chain decoder.

[0028] In one embodiment, the embedded tensor and time step output from the encoder are used as input, including: splicing the embedded tensors of the light chain and heavy chain sequences output from the encoder into a tensor z0, constructing a flow matching pair with the tensor z0 and the noise ε, and performing a second-order linear addition of the flow matching pair and the time step t to form a vector field u t , vector field u t The sum of the position information of tensor z0 is used as the input of the flow matching model.

[0029] In one embodiment, the input of the flow matching model further includes: the time step t is converted into an embedding tensor of the time step through a sinusoidal embedding layer, and then transformed through a linear layer to obtain the embedding tensor of the converted time step.

[0030] In one embodiment, the training flow matching model is performed using an optimal transmission condition flow algorithm.

[0031] In one embodiment, the method of solving an ordinary differential equation to obtain an embedding result corresponding to the sampling time from the vector field includes: using an ordinary differential equation solver to analyze the time step t′, and solving the embedding result corresponding to the time step t′ from the vector field.

[0032] In one embodiment, the method of using a light chain decoder and a heavy chain decoder to generate the final light chain sequence and heavy chain sequence includes: splitting the embedding results of the corresponding sampling time into an embedding tensor of the light chain sequence and an embedding tensor of the heavy chain sequence, inputting the results into the heavy chain decoder and the light chain decoder respectively, and using a length sampler to sample and truncate the length distribution of the light chain sequence and the heavy chain sequence to generate the final light chain sequence and heavy chain sequence.

[0033] The present invention also provides an antibody sequence light and heavy chain evaluation system based on protein language and flow matching, using the antibody sequence light and heavy chain co-design method based on protein language and flow matching, comprising:

[0034] A uniqueness calculation unit, used to calculate the percentage of non-repeated antibodies in the generated antibody sequence to the total number of antibodies, reflecting the novelty of the generated antibody sequence;

[0035] The authenticity assessment unit is used to calculate the similarity distance between the generated antibody sequence and the real antibody sequence. The real antibody sequence is the antibody sequence in the validation set, reflecting the authenticity of the generated antibody sequence;

[0036] A diversity evaluation unit is used to calculate the similarity distance between each antibody sequence and the remaining antibody sequences in the generated antibody sequences, reflecting the diversity of the generated antibody sequences;

[0037] The comprehensive physicochemical property evaluation unit is used to splice the generated antibody sequence with the real antibody sequence into one antibody sequence, calculate the physicochemical indicators of each spliced antibody sequence, convert the physicochemical indicators into a list format, and calculate the Wasserstein distance. The smaller the Wasserstein distance, the closer it is to the natural antibody.

[0038] Antibody sequence visualization unit, which predicts the antibody sequence logo based on the final light chain sequence and heavy chain sequence;

[0039] The human origin assessment unit adds a suffix to the generated antibody sequence, takes the suffixed antibody sequence as input, and outputs a human origin assessment file for evaluating the possibility that the generated antibody sequence is derived from humans;

[0040] The light and heavy chain quality assessment unit is used to take the final light chain sequence and heavy chain sequence as input, output the evaluation result files corresponding to the heavy chain and light chain, parse the chain type labels of the evaluation result files respectively, and calculate the consistency ratio between the predicted labels and the actual chain type;

[0041] The CDR quality assessment unit is used to generate diverse redesign samples of the specified CDR region using the final light chain and heavy chain sequences as input, and to calculate the matching degree of each CDR region after redesign with the original sequence;

[0042] a solubility evaluation unit, which is used to predict the solubility of the antibody sequence in a specific solution environment using the generated antibody sequence as input;

[0043] A structure prediction unit, which is used to predict the three-dimensional folding structure and connection mode of the antibody sequence using the generated antibody sequence as input;

[0044] The antibody structure designability assessment unit is used to calculate the TAP score indicator based on the three-dimensional folding structure of the antibody sequence output by the structure prediction unit to reflect the designability of the antibody sequence structure. The TAP score indicator includes CDR length, PSH, PPC and PNC parameters;

[0045] The antibody aggregation tendency assessment unit is used to calculate the SAP score index reflecting the aggregation tendency of the antibody sequence using the three-dimensional folding structure of the antibody sequence output by the structure prediction unit as input, where the SAP score index includes the total SAP value, the SAP value of each residue, and the number of residues;

[0046] The structure credibility prediction unit is used to predict the three-dimensional structure of the antibody sequence using the generated antibody sequence as input and calculate the pLDDT value and root mean square error, which reflect the credibility of the three-dimensional structure of the antibody sequence.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] (1) Distinguish the design of light and heavy chains, combine the protein language model with the antibody generation model, and design in the embedding space of the protein language model to cover the overall protein semantic information, effectively solving the problem of shallow semantic information in current computer-aided protein design methods, and can quickly, effectively and autonomously generate reliable, diverse, real and complete antibody pair sequences.

[0049] (2) The constructed flow matching vector field prediction model increases the antibody generation speed by more than 2 times, effectively solving the problems of high computational cost and slow speed of current computer-assisted methods.

[0050] (3) Based on the light and heavy chain co-design method, the present invention also provides an antibody sequence evaluation system, which includes a comprehensive evaluation of the antibody sequence foldability, novelty, diversity, authenticity, physicochemical properties, species origin, light and heavy chain accuracy, CDR design quality, solubility, aggregation tendency and other properties, providing a powerful means for screening high-quality generated antibody sequences. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for describing the embodiments or the prior art.

[0052] Figure 1 Schematic diagram of the process of the antibody sequence light and heavy chain co-design method based on protein language and flow matching provided in an embodiment of the present invention.

[0053] Figure 2 It is a schematic structural diagram of the skeleton model provided in an embodiment of the present invention.

[0054] Figure 3 Schematic diagram of the structure of the antibody sequence light and heavy chain evaluation system based on protein language and flow matching provided in an embodiment of the present invention.

[0055] Figure 4 It is a three-dimensional structure ribbon diagram, correlation length and pLDDT information of an antibody sequence generated as an example in an embodiment of the present invention. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.

[0057] In order to achieve the co-design of light and heavy chains, high quality, high speed and low cost of artificial antibody sequence design, the embodiment provides a method for co-designing light and heavy chains of antibody sequences based on protein language and flow matching, such as Figure 1 As shown, the following steps are included:

[0058] S1. Obtain antibody sequence data and perform preprocessing to generate sequence data of the light chain and heavy chain of the paired antibody.

[0059] In the examples, the antibody sequence data were collected from the antibody sequence data of the Observed Antibody Space (OAS) antibody database, including 2M (M, million) paired antibody data and 2.4B (B, billion) unpaired antibody data.

[0060] Next, the antibody sequence data is preprocessed, and the preprocessing is performed according to sequence pairing, clustering, length screening, amino acid type screening and length filling. Pairing refers to screening out 2M paired antibody data from the antibody sequence data; clustering refers to performing similar aggregation and removal of highly similar sequences on the paired antibody data through the clustering operation of the MMseqs2 tool. This step helps to remove highly similar or repeated sequences. Here, the similarity threshold (sequence identity) is selected as 95%, and the coverage threshold (coverage) is selected as 80%; screening length refers to removing antibody sequences that are too long or too short. Here, the lower limit of the length threshold is set to 2 amino acids, the upper limit of the light chain length threshold is ≥148 amino acids, and the upper limit of the heavy chain length threshold is ≥149 amino acids. The further upper limit is 149 amino acids for the heavy chain and 148 amino acids for the light chain; screening amino acid types refers to removing uncommon amino acids in the antibody sequence. Here, amino acids represented by single letters such as X, B, U, O, and J are considered uncommon amino acids; length filling refers to filling antibody sequences of varying lengths to the same sequence length to facilitate subsequent unified processing and use. Here, the filler is set to the ESM-2 built-in filler. <pad>, the heavy chain is filled to position 149 and the light chain is filled to position 148.

[0061] After preprocessing, we ultimately obtained 1.4M paired antibody data. We stored the antibody light and heavy chain sequences in two separate FASTA files, using the same identity for each paired chain. Preprocessing improves the overall performance and robustness of the model by enriching the dataset.

[0062] S2. Input the light chain and heavy chain sequence data of the paired antibody into the encoder of the protein language model, perform statistical information and normalization on the output embedding tensors of the light chain and heavy chain sequences, and obtain the embedding data of the light chain and heavy chain sequences.

[0063] In the embodiment, the protein language model is an ESM-2 protein language model with 8M parameters and trained with the UniRef50 dataset. Figure 2 As shown, the pre-processed heavy chain sequence x heavy The input to the ESM-28M protein language model is converted into a tensor of size (149,320) and the preprocessed light chain sequence x light The input to the ESM-2 8M protein language model was converted into a tensor of size (148,320) in the encoder. The overall mean and standard deviation of the two sets of tensor data sets were then calculated separately for subsequent data normalization.

[0064] In an embodiment, the preprocessed tensor data is Z-score normalized, wherein the values in the (0, 0.01) numerical interval in the normalized tensor are replaced with 0.01, and the values in the (-0.01, 0) numerical interval are replaced with -0.01. The normalized data is divided into a training set and a validation set in a ratio of 9:1. The normalization operation helps to unify the numerical range of the data set and helps to improve the success rate of flow matching training.

[0065] S3. Input the embedded data and statistical information into the decoder of the protein language model for training to form a light chain decoder and a heavy chain decoder.

[0066] In the embodiment, the normalized embedding data and statistical information are input into the decoder of the protein language model for training. The decoder consists of a linear layer of size (320, 320), a SiLU activation layer, a linear layer of size (320, 31), and an ESM-28M detokenize layer.

[0067] The normalized embedding data and statistical information were sequentially passed through the first linear layer, the activation layer, and the second linear layer to be converted into tensors of size (149, 31) and (148, 31). The token information was then extracted through the detokenize layer, and the 31-dimensional embeddings were converted into corresponding amino acids. Training was performed using the AdamW optimizer with a learning rate of 1e-4, a weight decay coefficient of 0.001, a beta value of (0.9, 0.98), and a training batch size of 64. The detokenize layer parameters were frozen, and the cross-entropy loss function was used to calculate the reconstruction loss between the decoder decoding the tensor-form antibody into the sequence-form antibody and the original antibody sequence.

[0068] S4. Using the embedding tensor and time step output from the encoder as input, the flow matching model is trained to form a flow matching vector field prediction model.

[0069] like Figure 2 As shown, after step 3 outputs the embedding tensors of the light chain and heavy chain sequences, we obtain z light and z heavy A pair of light and heavy chains are embedded in the tensor, and the heavy chain is embedded in the tensor z heavy and the light chain embedding tensor z light Use the concatenate operation to concatenate the first dimension into a tensor z0 of size (297, 320).

[0070] The embodiment adopts the optimal transport conditional flow algorithm, that is, the model learns the linear probability flow projection of the noise distribution and the real data distribution on the time interval [0, 1] dimension after optimal transport matching (Optimal Transport Mapping). In this embodiment, the standard normal distribution is used as the noise distribution, and the combined antibody sequence embedding data set is used as the real data distribution, that is, the standard normal distribution noise tensor ε of the size of (297, 320) of the same batch size is sampled. Subsequently, OT-CFM (Optimal Transport-Continuous Flow Matching, OT-CFM flow matching pair) is constructed for the tensor z0 and the standard normally distributed noise tensor ε, and the 2-Wasserstein distance between each tensor z0 and the standard normally distributed noise tensor ε is calculated, and then each tensor z0 is matched with the noise tensor closest to it.

[0071] Then, for each training example in the batch, a time step t is randomly sampled in the data interval [0, 1], and the input tensor z0 is linearly added with the standard normal distribution noise tensor ε according to t to form a vector field u t , vector field u t The actual integral flow field is then added to the position information of the tensor z0 and input into the flow matching model for training.

[0072] In the embodiment, the stream matching model uses BERT as the skeleton model, sets the number of hidden layers to 12, the hidden layer dimension to 320, the activation function to GeLU, the initializer range to 0.02, the number of dictionaries to 30522, the hidden layer dropout rate to 0.1, the attention heads to 16, the number of type dictionaries to 2, the maximum position embedding to 512, the Transformer feedforward network middle layer dimension to 3072, the layernorm epsilon to 1e-12, and the Transformer version to 4.6.0.dev0.

[0073] In addition, the stream matching model also includes a sinusoidal embedding layer for time step embedding and 12 linear layers for embedding time into the Transformer layer. The randomly sampled time step t is first converted into an embedding tensor of (297,320) through a sinusoidal embedding layer, and then transformed through a linear layer. The transformed tensor is added to the corresponding Transformer layer of BERT.

[0074] In the embodiment, the flow matching model learns the probability distribution π at different time steps t during training t , together constitute the optimal matching flow field v t The loss function uses the MSE loss function to calculate the actual integral flow field u t And the optimal matching flow field v learned by the model t The gap between.

[0075] Model training uses the constant_with_warmup learning rate scheduler to dynamically adjust the learning rate. The scheduler starts with a learning rate of 0 and linearly increases the learning rate to a constant of 1e-4 over 10,000 steps. The AdamW optimizer is used with a weight decay coefficient of 0.001, a beta value of (0.9, 0.98), and a training batch size of 64 to construct a flow matching vector field prediction model.

[0076] S5. Match the random sampling noise and sampling time input stream to the vector field prediction model to obtain the vector field corresponding to the sampling time. By solving the ordinary differential equation, the antibody sequence embedding tensor corresponding to the time step is obtained from the vector field. The light chain decoder and heavy chain decoder are used to generate the final light chain sequence and heavy chain sequence.

[0077] like Figure 2 As shown, after step 4, the optimal matching flow field v that can represent OT-CFM is obtained. t The model is constructed by combining the noise tensor ε′ with a sampling size of (297,320) in the standard normal distribution, the uniform sampling time step t′ in the data interval [0, 1] and the optimal matching flow field v t Input them into the ordinary differential equation analytical operator together, and output the new antibody sequence embedding tensor with a size of (297,320) corresponding to the sampling time. The new antibody sequence embedding tensor is split, that is, it is divided into two embedding tensors (149, 320) and (148, 320) along the first dimension, corresponding to the new light chain embedding tensor z′ respectively. light and the new heavy chain embedding tensor z′ heavy .

[0078] Then, the two embedded tensors are input into the light chain decoder and the heavy chain decoder respectively to generate the light chain sequence and heavy chain sequence data. Finally, a length sampler is used to sample the protein length, and the heavy chain and light chain sequences are truncated to obtain the heavy chain sequence x′ heavy and light chain sequence x′ light This embodiment uses the Dormand-Prince ordinary differential equation solver, and the analytical time step is a time step t′ with 25 steps evenly spaced from 1 to 0.

[0079] In order to obtain comprehensive evaluation information of generated antibody sequences so that biomedical experimenters can select the best antibodies for subsequent synthesis and immunoassays, the embodiment also provides an antibody sequence light and heavy chain evaluation system based on protein language and flow matching. Figure 3 Shown, including:

[0080] S301 Uniqueness Calculation Unit: Based on the antibody sequences generated by the light and heavy chain co-design method, the percentage of non-duplicate antibodies in the batch of sequences is calculated. This metric reflects the novelty of the generated antibodies to a certain extent.

[0081] S302 Authenticity Assessment Unit: By concatenating the generated antibody sequence and the validation set antibody sequence into a single sequence in the order of heavy chain and light chain, the edit distance and Hamming distance between each generated antibody sequence and each validation set antibody sequence are calculated. The minimum edit distance and minimum Hamming distance of each generated antibody sequence are retained and averaged to obtain the average minimum edit distance and average minimum Hamming distance. This unit reflects the authenticity of the generated sequence by evaluating the similarity distance between the generated sequence and the true sequence. The smaller the distance, the closer it is to the natural antibody.

[0082] S303 Diversity Assessment Unit: Based on the antibody sequences generated by the light and heavy chain co-design method, the edit distance between each antibody sequence in a batch of generated antibody sequences and other antibody sequences is calculated, and the minimum value is retained. The average of the minimum edit distances of all antibody sequences is taken as the average internal minimum edit distance IntDiv. This unit reflects the diversity of the generated antibody sequences by evaluating the similarity distance between the generated antibody sequences. The larger the IntDiv, the more diverse the antibody sequences.

[0083] S304 Comprehensive physicochemical property evaluation unit: The antibody sequence generated by the antibody sequence light and heavy chain co-design method and the validation set antibody sequence are spliced into a sequence in the order of heavy chain and light chain, and input into Biopython's ProteinAnalysis tool to calculate 15 types of physicochemical indicators for each antibody sequence, including: sequence length, molecular weight, aromaticity, instability index, isoelectric point, gravity, charge at pH 6, charge at pH 7, spiral fraction, turn structure fraction, sheet structure fraction, reduced molar extinction coefficient value, oxidized molar extinction coefficient value, average hydrophilicity and average surface accessibility. The index value is converted into a list format of length 15, and the Wasserstein distance (Wproperty) between the two is calculated. This unit uses the comprehensive physicochemical properties of real antibodies as the measurement object, quantifies the comprehensive physicochemical properties of the generated antibody sequence, and reflects the authenticity of the physicochemical properties of the generated antibody sequence. The smaller the Wproperty, the closer the physicochemical properties are to natural antibodies.

[0084] S305 Antibody Sequence Visualization Unit: Select representative antibody light and heavy chain sequences and input them into the WebLogo3 tool to predict sequence logo images.

[0085] S306 Human Origin Assessment Unit: The antibody sequences generated by the light and heavy chain co-design method of the antibody sequence are synthesized into a fasta file by adding the identity suffixes "HC" and "LC" to the heavy chain and light chain, respectively, and input into the BioPhi OASis tool. The OASis dataset of the BioPhi OASis tool uses OASis 9mers v 1, and the counting mode and CDR definition use kabat. The tool outputs a human origin assessment file in xlsx format, which comprehensively assesses the possibility that the light and heavy chains are derived from humans. It should be noted that the species origin of the dataset in this embodiment is mixed. Using this unit for evaluation requires adding a step to screen out human-derived antibody sequences in step 1 of the co-design method to train a model specifically for designing human-derived antibody sequences.

[0086] S307 Light and Heavy Chain Quality Assessment Unit: The antibody light and heavy chain sequences were input into the ANARCI tool, using IMGT for counting and CDR definition. The ANARCI tool outputs a samples_H.csv file for heavy chain evaluation and a samples_KL.csv file for light chain evaluation. The chain_type column of these two files was transposed and the percentage of predicted chain identity was calculated for each file.

[0087] S308CDR quality assessment unit: The antibody sequences generated by the antibody sequence light and heavy chain co-design method are merged into csV files with headers vh and v1 according to the heavy chain and light chain, and input into the IgLM model to predict the CDR sequence recovery rate of the antibody. The counting mode and CDR definition of the IgLM model adopt the Chothia model, and 1000 pairs of the antibody sequences are selected as seeds. The redesign amount of each pair is 10 groups. The redesigned regions are HCDR1, HCDR2, HCDR3, HCDR1HCDR2HCDR3, LCDR3, and the length design mode is fixed length. IgLM outputs the results to the iglm_fixed_samples_labeled.csv file. The device reads the sample_tag column and the seq_recovery column and calculates the sequence recovery rate of each redesigned region. This indicator reflects the generation confidence of each CDR region.

[0088] S309 Solubility Assessment Unit: The antibody sequences generated by the light and heavy chain co-design method and the validation set antibody sequences are concatenated into a single sequence in heavy chain and light chain order. This sequence is then input into the CamSol Intrinsic v2.2 tool to predict the antibody solubility. In this example, the pH is set to 7. The CamSol Intrinsic tool outputs the results as the files samples_pH7_CamSol_intrinsic.txt, and the device automatically reads the information in the Name and Instrinsic Solubility Profile columns.

[0089] S310 Structure Prediction Unit: This unit converts the antibody sequences generated by the light and heavy chain co-design method into a dictionary format with key values of "H" and "L" and inputs it into the IgFold v0.4.0 tool. The IgFold tool predicts the corresponding three-dimensional folding structure and connectivity of the antibody sequence in one go, outputting it as a pdb file. Figure 4 The generated three-dimensional structure ribbon diagram, correlation length and pLDDT information of the antibody sequence are displayed. The pLDDT information is used to measure the confidence of the local structure prediction for each amino acid residue in the model. It can be seen from the figure that the pLDDT is high, indicating that the predicted local conformation is highly credible.

[0090] S311 Antibody Structure Designability Evaluation Unit: The three-dimensional structure file predicted by S310 is input into the SAbPred tool to calculate the TAP score index, which includes four parameters: CDR length, PSH, PPC and PNC. The calculated TAP score index reflects the designability of the antibody sequence structure.

[0091] S312 Antibody Aggregation Propensity Assessment Unit: The three-dimensional structure file predicted by S310 is input into the SAP tool of Pyrosetta-4 to calculate the SAP score index. The tool outputs the total SAP value, the SAP value of each residue, and the number of residues. The calculated SAP score index reflects the aggregation tendency of the antibody sequence.

[0092] S313 structure credibility prediction unit: The antibody sequence generated by the light and heavy chain co-design method is input into the AlphaFold v2.3.O tool to generate a predicted antibody three-dimensional structure pdb file. The device calculates the pLDDT value and RMSD value based on the file, and evaluates the similarity between the model and the experimental results through the pLDDT value and RMSD value to measure the credibility of the three-dimensional structure of the antibody sequence.

[0093] Therefore, the present invention pre-processes the antibody sequence data and inputs it into the protein language model encoder, extracts the light / heavy chain sequence embedding tensor, and inputs it into the decoder for training after statistical analysis and normalization; based on the embedding tensor and time step output by the decoder, a flow matching vector field prediction model is constructed, and the antibody sequence embedding is generated by random noise sampling and solving ordinary differential equations, and then the light and heavy chain decoders are used to output the final sequence. The supporting evaluation system can comprehensively evaluate more than ten indicators such as the foldability, diversity, physicochemical properties, and CDR quality of the generated sequence to assist in the screening of high-quality antibodies. The global protein semantic information is integrated into the protein language model embedding space to improve the reliability of the generated sequence; the flow matching model increases the efficiency of antibody generation by more than 2 times; and the collaborative design and multi-dimensional evaluation of light and heavy chains are realized, which solves the problems of shallow semantic information, high computational cost, and imperfect evaluation system of traditional methods, providing a new way to efficiently develop functional antibodies.

[0094] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the scope of protection of the present invention.< / pad>

Claims

1. A method for co-designing light and heavy chains of antibody sequences based on protein language and flow matching, characterized in that: The following steps are involved: Obtain antibody sequence data and perform preprocessing to generate sequence data of the light chain and heavy chain of the paired antibody; The light chain and heavy chain sequence data of the paired antibody are input into the encoder of the protein language model, and the embedded tensors of the output light chain and heavy chain sequences are statistically analyzed and normalized to obtain the embedded data of the light chain and heavy chain sequences; Input the embedded data and statistical information into the decoder of the protein language model for training to form a light chain decoder and a heavy chain decoder; The embedding tensor and time step output from the encoder are used as input to train the flow matching model and form a flow matching vector field prediction model; The random sampling noise and sampling time input stream are matched to the vector field prediction model to obtain the vector field corresponding to the sampling time. The antibody sequence embedding tensor corresponding to the time step is obtained from the vector field by solving the ordinary differential equation, and the final light chain sequence and heavy chain sequence are generated using the light chain decoder and heavy chain decoder.

2. The method for co-designing light and heavy chains of antibody sequences according to claim 1, wherein: The antibody sequence data includes paired antibody data and unpaired antibody data.

3. The method for co-designing light and heavy chains of antibody sequences according to claim 1, wherein: The pretreatment includes: Pairing, used to filter out paired antibody data from antibody sequence data; Clustering is used to cluster similar sequences of the screened paired antibody data and remove highly similar sequences; Screening length is used to screen out antibody sequences that are too long or too short in the paired antibody data. The lower limit of the length threshold is 2 amino acids, the upper limit of the light chain length threshold is ≥148 amino acids, and the upper limit of the heavy chain length threshold is ≥149 amino acids; Screening of amino acid species to eliminate uncommon amino acids in antibody sequences; Length padding is used to pad the antibody sequences of uneven lengths in the paired antibody data to the same sequence length.

4. The method for co-designing light and heavy chains of antibody sequences according to claim 1, wherein: Statistical information is obtained by calculating the overall mean and standard deviation of the embedded tensors of the output light chain and heavy chain sequences. The embedded tensors of the output light chain and heavy chain sequences are normalized using the Z-score method based on the statistical information to obtain embedded data.

5. The method for co-designing light and heavy chains of antibody sequences according to claim 1, wherein: The decoder structure of the protein language model consists of two linear layers, an activation layer, and a detokenize layer; The embedded data and statistical information are sequentially transformed through the first linear layer, activation layer, and second linear layer to transform the size of the embedded data. The token information is then extracted through the detokenize layer to obtain an embedding tensor for a pair of light chain and heavy chain sequences, and form a light chain decoder and a heavy chain decoder.

6. The method for co-designing light and heavy chains of antibody sequences according to claim 5, characterized in that: The embedding tensor and time step output from the encoder are used as input, including: splicing the embedding tensors of the light chain and heavy chain sequences output from the encoder into tensor z0, constructing flow matching pairs with tensor z0 and noise ε, and performing second-order linear addition of the flow matching pairs and time step t to form a vector field u t , vector field u t The sum of the position information of tensor z0 is used as the input of the flow matching model.

7. The method for co-designing light and heavy chains of antibody sequences according to claim 6, characterized in that: The input of the flow matching model also includes: the time step t is converted into an embedding tensor of the time step through the sine embedding layer, and then transformed through the linear layer to obtain the embedding tensor of the converted time step.

8. The method for co-designing light and heavy chains of antibody sequences according to claim 1, wherein: The method of obtaining the antibody sequence embedding tensor corresponding to the time step from the vector field by solving the ordinary differential equation includes: using an ordinary differential equation solver to analyze the time step t', and solving the embedding result corresponding to the time step t' from the vector field.

9. The method for co-designing light and heavy chains of antibody sequences according to claim 8, wherein: The method of using a light chain decoder and a heavy chain decoder to generate the final light chain sequence and heavy chain sequence includes: splitting the embedding results of the corresponding sampling time into an embedding tensor of the light chain sequence and an embedding tensor of the heavy chain sequence, inputting the embedding tensors into the heavy chain decoder and the light chain decoder respectively, and using a length sampler to sample and truncate the length distribution of the light chain sequence and the heavy chain sequence to generate the final light chain sequence and heavy chain sequence.

10. A light and heavy chain evaluation system for antibody sequences based on protein language and flow matching, characterized in that: The antibody sequence light and heavy chain evaluation system uses the antibody sequence light and heavy chain co-design method based on protein language and flow matching according to any one of claims 1 to 9, comprising: A uniqueness calculation unit, used to calculate the percentage of non-repeated antibodies in the generated antibody sequence to the total number of antibodies, reflecting the novelty of the generated antibody sequence; The authenticity assessment unit is used to calculate the similarity distance between the generated antibody sequence and the real antibody sequence. The real antibody sequence is the antibody sequence in the validation set, reflecting the authenticity of the generated antibody sequence; A diversity evaluation unit is used to calculate the similarity distance between each antibody sequence and the remaining antibody sequences in the generated antibody sequences, reflecting the diversity of the generated antibody sequences; The comprehensive physicochemical property evaluation unit is used to splice the generated antibody sequence with the real antibody sequence into one antibody sequence, calculate the physicochemical indicators of each spliced antibody sequence, convert the physicochemical indicators into a list format, and calculate the Wasserstein distance. The smaller the Wasserstein distance, the closer it is to the natural antibody. Antibody sequence visualization unit, which predicts the antibody sequence logo based on the final light chain sequence and heavy chain sequence; The human origin assessment unit adds a suffix to the generated antibody sequence, takes the suffixed antibody sequence as input, and outputs a human origin assessment file for evaluating the possibility that the generated antibody sequence is derived from humans; The light and heavy chain quality assessment unit is used to take the final light chain sequence and heavy chain sequence as input, output the evaluation result files corresponding to the heavy chain and light chain, parse the chain type labels of the evaluation result files respectively, and calculate the consistency ratio between the predicted labels and the actual chain type; The CDR quality assessment unit is used to generate diverse redesign samples of the specified CDR region using the final light chain and heavy chain sequences as input, and to calculate the matching degree of each CDR region after redesign with the original sequence; a solubility evaluation unit, which is used to predict the solubility of the antibody sequence in a specific solution environment using the generated antibody sequence as input; A structure prediction unit, which is used to predict the three-dimensional folding structure and connection mode of the antibody sequence using the generated antibody sequence as input; The antibody structure designability assessment unit is used to calculate the TAP score indicator based on the three-dimensional folding structure of the antibody sequence output by the structure prediction unit to reflect the designability of the antibody sequence structure. The TAP score indicator includes CDR length, PSH, PPC and PNC parameters; The antibody aggregation tendency assessment unit is used to calculate the SAP score index reflecting the aggregation tendency of the antibody sequence using the three-dimensional folding structure of the antibody sequence output by the structure prediction unit as input, where the SAP score index includes the total SAP value, the SAP value of each residue, and the number of residues; The structure credibility prediction unit is used to predict the three-dimensional structure of the antibody sequence using the generated antibody sequence as input and calculate the pLDDT value and root mean square error, which reflect the credibility of the three-dimensional structure of the antibody sequence.