Prediction of multiple AAV properties by deep learning

Deep learning regression models predict AAV properties from capsid amino acid sequences, addressing the inefficiencies of traditional methods by enabling the rapid identification of AAV variants with improved productivity, infectivity, and neutralizing antibody escape.

WO2025104321A1PCT designated stage expired Publication Date: 2025-05-22GENETHON +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/082609
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2024-11-15
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Current methods for engineering Adeno-Associated Virus (AAV) capsids with improved properties are time-consuming and inefficient, as they require testing numerous sequence variants to identify those with enhanced productivity, transduction efficiency, and immunogenicity.

Method used

The development of deep learning regression models that accurately predict AAV properties such as productivity, infectivity, and neutralizing antibody escape directly from capsid amino acid sequences, using datasets with random mutations in multiple variable regions.

Benefits of technology

These models enable the rapid identification of AAV variants with improved multiple properties, facilitating the design of new AAV capsid variants with enhanced performance and allowing for the exploration of combinations of multiple variable regions that were previously unexplored.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024082609_22052025_PF_FP_ABST
    Figure EP2024082609_22052025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method implemented by computer means for the prediction of at least one property of an AAV vector.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] PREDICTION OF MULTIPLE AAV PROPERTIES BY DEEP LEARNING

[0002] FIELD OF THE INVENTION

[0003] The invention relates to a method implemented by computer means for the prediction of at least one property of an Adeno-Associated Virus (AAV) vector.

[0004] BACKGROUND OF THE INVENTION

[0005] Adeno-associated virus (AAV) vectors show great promise as gene delivery vectors, in part due to the ability of the AAV capsid to target various tissues for the treatment of a variety of human diseases.

[0006] AAV is a non-enveloped virus comprising three capsid proteins, virion protein 1 (VP1), VP2 and VP3, which assemble into an icosahedral 60-mer capsid. Usually, natural capsids are constituted by five VP1 proteins, five VP2 proteins and fifty VP3 proteins. The genome carries two genes, Rep and Cap, flanked by two palindromic regions named Inverted terminal Repeats (ITR) that serve as the viral origins of replication and the packaging signal. The Cap gene codes for the three structural proteins VP1, VP2 and VP3 that share the same C-terminal end which is all of VP3. The Rep gene encodes four proteins required for viral replication. In recombinant AAV vectors the Rep and Cap genes are replaced with a transgene expression cassette which is flanked by the ITRs and packaged into AAV capsid.

[0007] Structurally, VP3 monomer core contains highly conserved eight-stranded P-barrel motif (DiMattia et al., 2012). Nine surface-exposed variable regions, VR (VR1, VR2, VR3, VR4, VR5, VR6, VR7, VR8 and VR9) inserted between the P-strands result in local topological differences between serotypes and dictate virus-host interaction. Genetically modifying VRs can drastically change the AAV productivity, transduction efficiency, and immunogenicity (Li and Samulski, 2020; Ogden et al., 2019; Tseng and Agbandje-McKenna, 2014). For these reasons, VR regions are interesting targets to engineer AAV capsids with improved properties. A common approach for engineering AAV capsids with novel tropism is to select a random library of AAV capsids modified by peptide insertion into VR8 region through multiple rounds of selection to identify a few top-performing candidates.

[0008] However, Capsid production and selection represents a strong drawback because the quantity of sequence variants to test is such that the identification of one AAV capsids having appropriate productivity, and / or transduction efficiency, and / or immunogenicity takes years with conventional methods of selection.

[0009] Bryant et al. addressed this drawback by developing a machine learning models to identify capsid protein variants with improved productivity. However, said machine learning models allows to predict only one property of an AAV vector, based on the variation in only one VR (VR8).

[0010] Therefore, to improve AAV vectors used in gene therapy, there is a need for new machinelearning implemented methods able to predict AAV properties directly from capsid amino acid sequences.

[0011] SUMMARY OF THE INVENTION

[0012] The present invention provides different regression models with different architectures, that accurately predict AAV properties such as AAV productivity, AAV infectivity in target cells or tissue of interest or neutralizing antibody escape, directly from capsid amino acid sequences. The datasets were used for the training including random mutations (deletions, substitution, and insertion) simultaneously in at least two variable regions of AAV capsids, such as variable regions 4 and 8 (VR4 and VR8) of AAV capsid or variable regions 2 and 8 (VR2 and VR8) of AAV capsid. All models showed high accuracy on all tested AAV properties including AAV productivity, AAV infectivity in human differentiated myotubes (healthy control and Duchenne Muscular Dystrophy (DMD)-derived) or AAV neutralizing antibody escape. These models allow to engineer new AAV variants with improved properties, in particular improved multiple properties such as AAV productivity, AAV infectivity in skeletal muscles or other cells or tissue of interest, and AAV neutralizing antibody escape. Also, it helps to explore the combination of multiple VRs on AAV biology which has not been possible before. The invention relates to a method implemented by computer means for the prediction of at least one property of an AAV vector, said method comprising the steps of:

[0013] - obtaining the amino-acid sequences of at least two VR sequences of said AAV vector capsid protein,

[0014] - predicting a value of said property of the AAV vector, based on said at least two VR sequences, using a regression model.

[0015] In some embodiments, said regression model comprises an encoding part of a transformer model.

[0016] In some embodiments, said regression model comprises at least two input paths upstream of said encoding part, each input paths receiving as input one of said VR sequences, each input path comprising embedding layers and a convolution layer, said embedding layers comprising a token embedding layer and a positional embedding layer, said regression model comprising an aggregation layer aggregating the outputs of said input paths.

[0017] In some embodiments, said embedding layer comprises a mutational embedding layer.

[0018] In some embodiments, said regression model comprises an LSTM layer, a flatten layer and at least one fully connected layer downstream said encoding part.

[0019] In some embodiments, said at least one fully connected layer comprises two fully connected layers.

[0020] In some embodiments, a tensor representing a feature of f_aav to f_pls ratio or f_ma to f_aav ratio is aggregated to the input of the flatten layer and to the input of each fully connected layer; wherein f_aav is the frequency of AAV vector genome; f_pls is the frequency of ITR- containing plasmid used for AAV vector production; and f_ma is the frequency of transgene RNA level in AAV target cell transduced with AAV vector.

[0021] In some embodiments, at least a first and a second property are predicted, based on said at least two VR sequences, using a respectively a first and a second regression model. In some particular embodiments, the AAV vector productivity may be predicted using a first regression model, the AAV vector infectivity in target cells or tissue such as muscles may be predicted using a second prediction model and the neutralizing antibody escape of AAV vector may be predicted using a third prediction model.

[0022] Each regression model may have the same architecture. However, each regression model may be trained separately to predict the corresponding property.

[0023] In some embodiments, said at least two VR sequences are VR4 and VR8.

[0024] In some embodiments, said at least two VR sequences are VR2 and VR8.

[0025] Another aspect of the invention relates to a computer software, comprising instructions to implement at least a part of the method according to the present disclosure, when the software is executed by a processor.

[0026] Yet another aspect of the invention relates to a computer device comprising:

[0027] - an input interface to receive amino-acid sequences of at least two VR sequences,

[0028] - a memory for storing at least instructions of a computer program according to the present disclosure,

[0029] - a processor accessing to the memory for reading the aforesaid instructions and executing then the method according to any of the method according to the present disclosure,

[0030] - an output interface to provide an indication based on the predicted value of said property of the AAV vector.

[0031] The invention also relates to a computer-readable non-transient recording medium on which a computer software is registered to implement the method according to the present disclosure, when the computer software is executed by a processor.

[0032] DETAILED DESCRIPTION OF THE INVENTION

[0033] The invention provides a method implemented by computer means for the prediction of at least one property of an AAV vector. The invention encompasses a computer software, a computer device and a computer-readable non-transient recording medium to implement the method according to the present disclosure. The method according to the invention is useful in particular to engineer AAV vectors comprising new AAV capsid variants which have improved properties, in particular improved multiple properties such as productivity, infectivity in skeletal muscles or other target cells or tissue and neutralizing antibody escape.

[0034] In some embodiments, a method implemented by computer means for the prediction of at least one property of an AAV vector, comprises the steps of:

[0035] - obtaining the amino-acid sequences of at least two VR sequences of said AAV vector capsid protein,

[0036] - predicting a value of said property of the AAV vector, based on said at least two VR sequences, using a regression model.

[0037] AAV vector has the standard meaning in the art and relates to a recombinant AAV vector particle composed of an AAV capsid packaging a gene of interest. By “gene of interest”, it is meant a gene useful for a particular application, such as with no limitation, diagnosis, reporting, modifying, therapy and genome editing. Non limiting examples of genes of interest include: a gene of interest for therapy such as a transgene encoding a therapeutic protein or RNA; a gene of interest to assess AAV properties such as a reporter gene or the AAV Cap gene as disclosed in the present examples (Figure 2).

[0038] The AAV capsid may be of any AAV serotype. As used herein “AAV serotype” or “AAV capsid serotype” refers to an AAV capsid having distinct variable region (VR, also named hypervariable region or HVR) amino acid sequences compared to an AAV capsid of another serotype. Different AAV serotypes have amino acid variation in their VR sequences. The term AAV serotype encompasses any natural or artificial AAV capsid serotype including AAV capsid variants isolated from human or non-human species and AAV capsid variants engineered by various techniques known in the art such as for example rational design, directed evolution and in silica discovery. AAV serotype includes hybrid or chimeric AAV capsids. As used herein, the term AAV serotype refers to a functional AAV capsid which is able to form recombinant AAV viral particles which transduce a cell, tissue or organ, in particular a cell tissue or organ of interest (target cell, tissue or organ) and express a transgene in said cell, tissue or organ, in particular target cell tissue or organ.

[0039] The AAV capsid protein may be VP1, VP2 or VP3 protein. Using hybrid AAV9.rh74 capsid protein (SEQ ID NO: 1) as a reference sequence, VP1 corresponds to SEQ ID NO: 1; VP2 corresponds to the amino acid sequence from T138 to the end of SEQ ID NO: 1; VP3 corresponds to the amino acid sequence from M203 to the end of SEQ ID NO: 1.

[0040] An AAV vector may have various properties. Some examples of properties that could be associated with such AAV vectors may be:

[0041] - Productivity, also called Fitness: The property of productivity refers to the ability of the AAV vector to be produced by standard recombinant AAV production methods. Productivity may be measured as the number of AAV vector genomes (viral genome copy number or vg) generated per cell or the concentration or titer of AAV vector genomes of an AAV vector preparation.

[0042] - Infectivity: The property of infectivity refers to the ability of the AAV vector to efficiently enter and deliver its genetic material into target cells.

[0043] - Transduction Efficiency: Transduction efficiency represents how effectively the AAV vector can deliver and express the desired genetic material in target cells.

[0044] - Tropism: Tropism refers to the preference of the AAV vector for specific cell types or tissues. Different AAV vectors may exhibit distinct tropism profiles based on the VR sequences they contain.

[0045] - Immunogenicity: Immunogenicity refers to the potential of the AAV vector to elicit an immune response to the capsid or transgene when administered in vivo. Some AAV vectors may have lower immunogenicity, making them more suitable for gene therapy applications.

[0046] - Neutralizing antibody escape: neutralizing antibody escape relates to the ability of an AAV vector to avoid neutralization by antibodies present in human serum.

[0047] - Stability: The stability property relates to the resilience and integrity of the AAV vector under various conditions, such as storage, transportation, and delivery methods.

[0048] - Serotype Specificity: AAV vectors come in various serotypes, each with different surface properties conferred by the VR sequences. The property of serotype specificity relates to the ability of the AAV vector to interact with specific receptors of host cells and exhibit different transduction properties. These are just a few examples, and there could be other properties associated with AAV vectors depending on the specific context, research focus, or application of the vectors.

[0049] These properties may be assessed by standard assays that are well-known in the art.

[0050] AAV vectors are produced by standard co-transfection assays in appropriate cells for AAV production (see in particular Ayuso E. et al., Hum. Gene Ther. 2014, 25, 977-987). For example, HEK293T cells are transfected with 3 plasmids: i) a transgene plasmid (ITR- containing plasmid) containing AAV2 ITRs flanking a transgene expression cassette (corresponding to AAV vector genome) ii) the helper plasmid pXX6, containing adenoviral sequences necessary for AAV production, and iii) a plasmid containing AAV Rep and Cap genes, defining the serotype of AAV, as illustrated in Figure 2A (left panel). Alternatively, the Cap gene replaces the transgene in the transgene plasmid and the other plasmid codes only for the Rep gene (Figure 2A, right panel). Two days after transfection, the cells are lysed to release the AAV particles. The viral lysate is purified by affinity chromatography. Viral genomes are quantified by a TaqMan real-time PCR assay using primers and probes corresponding to the ITRs of the AAV vector genome (Rohr et al. J Virol Methods., 2002, 106,8 l-8.doi: 10.1016 / s0166-0934(02)00138-6). rAAV titers are expressed as viral genome copy number (vg).

[0051] AAV infectivity, transduction efficiency or tropism may be determined in vitro or in vivo by measuring the vector copy number (VCN) per cell or expression level of the gene of interest at the mRNA or protein level according to standard methods that are well-known in the art. The gene of interest may be a reporter gene such as luciferase. Luciferase gene expression level may be determined in vivo in mice by measuring bioluminescence level in various organs using in vivo imaging. It may also be determined in vitro or in vivo by measuring luciferase activity in cell or organ using luciferase assay.

[0052] In particular, AAV infectivity, transduction efficiency or tropism may be determined in a target or non-target cell, tissue or organ. The target refers to the cell, tissue or organ in which transduction with the AAV vector and expression of the gene of interest is desired. The nontarget refers to the cell, tissue or organ in which transduction with the AAV vector and / or expression of the gene of interest is avoided (detargeting). The target may be muscle, nervous system, liver cell or tissue or other cell or tissue. The non-target may be liver when the targeting of another cell or tissue such as muscle or nervous system is desired. As used herein, the term “muscle” refers to cardiac muscle (i.e. heart) and skeletal muscle. The term “muscle cells” refers to myocytes, myotubes, myoblasts, and / or satellite cells. In some embodiments, the target is muscle cell or tissue. The target cell or tissue may be from a healthy or diseased individual. As used herein, the term “individual” includes human and other mammalian subjects. Preferably, an individual according to the invention is a human. A diseased individual refers to an individual having a disease, in particular a genetic disease that can be treated by gene therapy with the AAV vector. In some embodiments, the genetic disease is muscular dystrophies. Muscular dystrophies include in particular Dystrophinopathies (DMD gene) and Limb-girdle muscular dystrophies (LGMDs) (CAPN3, DYSF, FKRP, ANO5, DNAJB6 genes and others such as SGCA, SGCB, SGCG).

[0053] AAV immunogenicity may be determined by measuring anti- AAV vector immune response induced after AAV vector administration into mice using standard methods that are well- known in the art. For example, the method may comprise measuring antibodies or T cells against AAV capsid or transgene.

[0054] Neutralizing antibody escape may be measured by determining the seroprevalence (human seroprevalence) of the AAV capsid which means the level of anti- AAV antibodies binding to an AAV capsid present in a human population and expressed as serie antibodies or immunoglobulins. The seroprevalence of an AAV capsid is measured using a cohort of human sera and standard assays that are well known in the art and disclosed for example in (Meliani et al., Hum Gene Ther Methods. 2015 Apr;26(2):45-53. doi: 10.1089 / hgtb.2015.037). The seroprevalence of an AAV capsid may be defined as the percentage of individuals having an ELISA titer of IgG specific for said capsid higher thanlO pg / mL. A seroprevalence of less than 30 % may be considered as low seroprevalence whereas a seroprevalence of more than 50 % may be considered as high seroprevalence.

[0055] Such property may be quantified by a numerical value as an output of the regression model.

[0056] A regression model is a statistical model used to estimate the relationship between one or more independent variables (often referred to as predictors, features, or input variables) and a dependent variable (also known as the response or output variable) that is a continuous or numerical value. The goal of a regression model is to predict or explain the value of the dependent variable based on the values of the independent variables.

[0057] Regression models assume that there is a functional relationship between the independent variables and the dependent variable. The model estimates the parameters of this relationship to make predictions or infer the impact of the independent variables on the dependent variable.

[0058] Each amino-acid sequence may be represented as an input tensor of dimension 1 x “a”, where “a” may be equal to 30. In this case, If the original sequence is shorter than 30 amino acids, said sequence may be padded to make it equal to 30.

[0059] In the context of AAV vectors, VR (Variable Region) sequences refer to specific regions within the viral capsid protein. These VR sequences play a critical role in determining the properties and behaviors of the AAV vector, including its productivity, tropism, transduction efficiency, interactions with target cells as well as its immunogenicity and neutralizing antibody escape. For example, using AAV9.rh74 capsid as reference sequence (VP1 of SEQ ID NO: 1), VR1 sequence comprises positions 262 to 274, VR2 sequence comprises positions 323 to 334, VR3 sequence comprises positions 382 to 392, VR4 sequence comprises positions 446 to 465, VR5 sequences comprises positions 489 to 503, VR6 sequence comprises positions 528 to 534, VR7 sequence comprises positions 545 to 558, VR8 sequence comprises positions 578 to 597, VR9 sequence comprises positions 704 to 714. VR1 to VR9 sequences are in the common C-terminal region corresponding to VP3 protein. A person skilled in the art can easily obtained the corresponding positions of the hypervariable regions in other AAV capsid serotypes by sequence alignment of any other AAV capsid sequence of any other serotype with SEQ ID NO: 1 using standard protein sequence alignment programs that are well-known in the art, such as for example BLAST, FASTA, CLUSTALW, MEGA and the like.

[0060] An amino acid is a building block of proteins, and in the context of a VR sequence of an AAV vector, an amino acid refers to the individual units or residues that make up the sequence. VR sequences are typically composed of a specific arrangement of amino acids that contribute to the structure and function of the viral capsid protein. For example, VR2 sequence of AAV9.rh74 consists of KEVTDNNGVKTI (SEQ ID NO: 8), VR4 sequence of AAV9.rh74 consists of YLSRTQSTGGTAGTQQLLFS (SEQ ID NO: 2) and VR8 sequence of AAV9.rh74 consists of YGVVADNLQQQNAAPIVGAV (SEQ ID NO: 3). Each amino acid in the sequence is represented by a specific three-letter or one-letter code according to the standard nomenclature. Alanine: Ala or A; Arginine: Arg or R; Asparagine: Asn or N; Aspartic acid: Asp or D; Cysteine: Cys or C; Glutamine: Gin or Q; Glutamic acid: Glu or E; Glycine: Gly or G; Histidine: His or H; Isoleucine: He or I; Leucine: Leu or L; Lysine: Lys or K; Methionine: Met or M; Phenylalanine: Phe or F; Proline: Pro or P; Serine: ser or S; Threonine: Thr or T; Tryptophane: Trp or W; Tyrosine: Tyr or Yand Valine: Vai or V. The specific amino acids present in the VR sequence can influence the properties of the AAV vector. Different amino acids have distinct physicochemical properties such as hydrophobicity, charge, size, and molecular interactions. These properties can impact the stability, binding affinity, receptor recognition, and other characteristics of the viral capsid and the AAV vector as a whole. By analyzing and understanding the amino acid composition and arrangement within VR sequences, researchers can gain insights into the structurefunction relationship of AAV vectors, as well as their behavior in various biological systems, which means their properties.

[0061] “a”, “an”, and “the” include plural referents, unless the context clearly indicates otherwise. As such, the term “a” (or “an”), “one or more” or “at least one” can be used interchangeably herein; unless specified otherwise, “or” means “and / or”.

[0062] The two or more VR sequences may be chosen among VR1, VR2, VR3, VR4, VR5, VR6, VR7, VR8 and VR9 of the AAV capsid (VP3 protein). In some embodiments, the two or more VR sequences are chosen from VR1, VR2, VR4, VR8 and VR9 sequences (of VP3 protein). In another embodiments, the two sequences are not VR4 and VR5 sequences. In some particular embodiments, the two VR sequences are VR4 and VR8 sequences. In some particular embodiments, the two VR sequences are VR2 and VR8 sequences. The AAV capsid is advantageously an AAV capsid variant having one or more mutations in at least one VR sequence, preferably both VR sequences compared to the VR sequences of the parent AAV capsid from which said AAV capsid is derived. The parent AAV capsid is a reference capsid whose properties are being improved using the prediction method according to the invention. As used herein a mutation may be an amino acid insertion, an amino acid deletion or an amino acid substitution. In some embodiments of the method according to the invention, said regression model comprises an encoding part of a transformer model.

[0063] The structure of a transformer model consists of an encoder-decoder architecture designed to handle sequence data and introduced in Vaswani, Ashish & Shazeer, Noam & Parmar, Niki & Uszkoreit, Jakob & Jones, Llion & Gomez, Aidan & Kaiser, Lukasz & Polosukhin, Illia, “Attention is all you need” , 2017.

[0064] Said encoding part may comprise a stack of N attention blocks where N may be comprised between 4 and 10, for example equal to 8. Said encoding part may comprise a stack of N attention blocks where N may be comprised between 2 and 10, for example equal to 2.

[0065] Each attention block may comprise two sub-layers. The first sub-layer may be a multi-head self-attention mechanism, and the second may be position- wise fully connected feed-forward network. A residual connection may be employed around each of the two sub-layers, followed by normalization layer. That is, the output of each sub-layer is LayerNorm(x + Sublayer(x)), where Sublayer(x) is the function implemented by the sub-layer itself.

[0066] In some embodiments of the method according to the invention, said regression model comprises at least two input paths upstream of said encoding part, each input paths receiving as input one of said VR sequences, each input path comprising embedding layers and a convolution layer, said embedding layers comprising a token embedding layer and a positional embedding layer, said regression model comprising an aggregation layer aggregating the outputs of said input paths.

[0067] In the context of machine learning model architecture, a "path" refers to the flow of information or computation within the model. It represents the sequence of operations or transformations that data undergoes as it passes through the model's layers.

[0068] A path typically starts at the input layer of the model and propagates through successive layers, with each layer performing a specific operation or transformation on the input data. The output of one layer becomes the input for the next layer. Each layer can be seen as a step along the path, where computations are applied.

[0069] The term "upstream" refers to the layers or operations that come before a specific layer or point in the model. In other words, it refers to the preceding stages of the data flow. On the other hand, "downstream" refers to the layers or operations that come after a specific layer or point in the model. It represents the subsequent stages of the data flow.

[0070] In the context of amino acid sequences, embedding layers may be used to represent discrete amino acids as continuous vector representations, capturing their characteristics and enabling neural networks to process and learn from the sequence data.

[0071] A token embedding layer may transform individual amino acids into dense vector (a vector where most of the elements are non-zero) representations. Each amino acid in the sequence is assigned a unique embedding vector. These embeddings capture the properties and characteristics of amino acids, such as hydrophobicity, charge, or structural features.

[0072] Token embeddings for amino acids can be initialized randomly or pretrained using techniques such as amino acid similarity matrices or domain- specific databases. Pretrained embeddings may leverage existing knowledge about amino acid properties and relationships, allowing the model to benefit from prior information.

[0073] For example, the amino acid sequence "ARNDC" would be transformed by the token embedding layer into a sequence of corresponding continuous vectors.

[0074] In addition, the order or position of amino acids in a sequence can be crucial for understanding their biological function or structure. To capture positional information, a positional embedding layer may be utilized.

[0075] The positional embedding layer may generate embeddings that encode the position or order of amino acids within the sequence. Each position in the sequence is assigned a unique vector, allowing the model to differentiate between amino acids based on their relative positions.

[0076] By adding positional embeddings to the token embeddings, the resulting combined representations capture both the meaning of the amino acids and their sequential relationships within the sequence.

[0077] For instance, in an amino acid sequence "ARNDC," the positional embedding layer would assign distinct vectors to the amino acids "A," "R," "N," "D," and "C" to denote their respective positions. By incorporating token embeddings and positional embeddings, neural network models can effectively capture both the characteristics of amino acids and the sequential dependencies within amino acid sequences. This facilitates the understanding and analysis of protein structures, functions.

[0078] A convolution layer is a fundamental component of convolutional neural network (CNN) architectures. Said convolution layer may operate on the embedded amino acid sequences to extract local patterns or features.

[0079] The convolutional layer applies a set of learnable filters or kernels to the input embeddings. These filters slide across the input tensor, performing element-wise multiplications and summations to produce a feature map. Each filter captures a specific pattern or feature at different positions in the sequence.

[0080] The learned filters act as feature detectors, extracting relevant information from the input tensors.

[0081] An aggregation layer refers to a component or operation within a model that combines or aggregates multiple inputs or features into a single representation. The aggregation layer can take various forms depending on the specific model architecture and task at hand.

[0082] Said aggregation layer may be a concatenation layer.

[0083] In a concatenation layer, multiple inputs or feature maps are concatenated along a specific dimension. This operation combines the individual representations into a single, larger representation. It allows the model to consider different aspects or perspectives of the data simultaneously.

[0084] Said convolutional layer may be a one-dimensional convolution layer (ConvlD).

[0085] The output of the embedding layers may be a 1 x a x b tensor, where a may be 30 and b may be equal to 20.

[0086] The output of the convolutional layer may be a 1 x a x c tensor, where c > b. For example, c may be equal to 256.

[0087] The output of the aggregation layer may be a 1 x 2a x c tensor. The output of the encoder part may also be a 1 x 2a x c tensor.

[0088] In some embodiments of the method according to the invention, said embedding layer comprising a mutational embedding layer.

[0089] A mutational embedding layer, in the context of machine learning for protein sequences, is a specialized layer that converts amino acid mutations into continuous vector representations or embeddings. It aims to capture the information about how mutations affect the structure or function of proteins and enable downstream machine learning models to learn from these representations.

[0090] When dealing with protein sequences, a mutational embedding layer takes as input the original amino acid sequence and the information about mutations or amino acid substitutions that have occurred. It then generates fixed- length vector representations that encode the effects of these mutations.

[0091] To incorporate mutations into the embeddings, the mutational embedding layer may perform operations such as element- wise addition or concatenation of the embeddings for the original amino acids and the mutated amino acids. This combination allows the resulting embeddings to capture the differences between the original sequence and the mutated sequence.

[0092] In some embodiments of the method according to the invention, said regression model comprise an LSTM layer, a flatten layer and at least one fully connected layer downstream said encoding part.

[0093] An LSTM (Long Short-Term Memory) layer is a type of recurrent neural network (RNN) layer commonly used for processing sequential data. LSTM networks are designed to address the vanishing gradient problem faced by traditional RNNs, enabling them to capture longterm dependencies in sequences.

[0094] A flatten layer is a common layer in neural network architectures that is used to convert multidimensional input data into a one-dimensional form. It reshapes the input tensor from a higher-dimensional representation into a flat vector, which can then be processed by subsequent layers, such as fully connected layers or output layers. A fully connected layer, also known as a dense layer, is a fundamental component in neural network architectures. In a fully connected layer, every neuron is connected to every neuron in the previous layer, forming a fully connected graph.

[0095] The final fully connected layer, known as the output layer, produces the network's final predictions or outputs. The parameters to be learned in a fully connected layer include the weights and biases associated with each neuron.

[0096] The output of the LSTM layer may be a 1 x 2a x d tensor, where d < c. For example, d may be equal to 16.

[0097] The output of the flatten layer may be a 1 x e tensor, where e = d x c. For example, e may be equal to 960.

[0098] The output of the fully connected layer may be a 1 x f tensor, where f > e. For example, f may be equal to 64.

[0099] In some embodiments of the method according to the invention, said at least one fully connected layer comprises two fully connected layers.

[0100] Using multiple fully connected layers instead of a single one in a neural network offers several advantages. Firstly, it increases the model's capacity to learn complex representations by introducing more parameters. This allows the network to capture intricate patterns and relationships in the data, potentially enhancing its performance. Secondly, the presence of multiple layers enables hierarchical feature learning, where each layer learns different levels of abstraction from the input data. Lower layers capture low-level features, while deeper layers capture higher-level features built upon them. This hierarchical representation learning can lead to more effective and expressive models. Additionally, each fully connected layer applies a non-linear activation function, allowing the network to learn and model complex non-linear transformations of the input data. Moreover, multiple layers aid in generalization and robustness by abstracting away noisy or irrelevant information and focusing on the most relevant features. Lastly, the stacked fully connected layers facilitate improved representation learning, as each layer learns a new representation of the data, refining and enriching the learned representations as information flows through the network. In some embodiments of the method according to the invention, a tensor representing a feature of f_aav to f_pls ratio or f_ma to f_aav ratio is aggregated to the input of the flatten layer and to the input of each fully connected layer, wherein f_aav is the frequency of AAV vector genome; f_pls is the frequency of ITR-containing plasmid (AAV genome plasmid or transgene plasmid) used for AAV vector production; and f_rna is the frequency of RNA level of the gene of interest (transgene) in AAV target cell transduced with AAV vector. AAV target cell is in particular wild-type (WT) and Duchenne Muscular Dystrophy (DMD) myotubes. AAV vector productivity may be defined as the ratio of f_aav to f_pls; AAV vector infectivity may be defined as the ratio of f_rna to f_aav. Said features may be determined in variants of a library of AAV vector variants having diversity in said at least two VR sequences of AAV vector capsid as disclosed in the examples. AAV vector productivity is determined in cells producing the library of AAV vector variants using the ITR-containing plasmids. AAV vector infectivity is determined in AAV target cell transduced with the library of AAV vector variants.

[0101] Such aggregation may be concatenation.

[0102] In some embodiments of the method according to the invention, at least one property of an AAV vector which is predicted, based on said at least two VR sequences, is neutralizing antibody escape, using a regression model.

[0103] In some embodiments of the method according to the invention, at least a first and a second property are predicted, based on said at least two VR sequences, using a respectively a first and a second regression model.

[0104] More particularly, the AAV productivity may be predicted using a first regression model, the AAV infectivity in muscles may be predicted using a second prediction model and the neutralizing antibody escape may be predicted using a third prediction model.

[0105] In some embodiments of the method according to the invention, said at least two VR sequences are VR4 and VR8.

[0106] In some embodiments of the method according to the invention, said at least two VR sequences are VR2 and VR8. Another aspect of the invention relates to a computer software, comprising instructions to implement at least a part of the method according to the present disclosure, when the software is executed by a processor.

[0107] Another aspect of the invention relates to a computer device comprising:

[0108] - an input interface to receive amino-acid sequences of at least two VR sequences,

[0109] - a memory for storing at least instructions of a computer program according to the preceding claim,

[0110] - a processor accessing to the memory for reading the aforesaid instructions and executing then the method,

[0111] - an output interface to provide an indication based on the predicted value of said property of the AAV vector.

[0112] A memory for storing at least instructions of a computer program is known as a program memory or instruction memory. This type of memory may store the instructions that the processor needs to execute a program, along with any associated data.

[0113] Several types of memory may be used, including Read-Only Memory (ROM), Flash Memory, Electrically Erasable Programmable Read-Only Memory (EEPROM), Random-Access Memory (RAM), Cache Memory and / or Virtual Memory.

[0114] Said output interface may comprise a display such as a monitor, a projector, a mobile device screen or a virtual reality headset screen.

[0115] Said indication may be the predicted numerical value.

[0116] Another aspect of the invention relates to a computer-readable non-transient recording medium on which a computer software is registered to implement the method according to the present disclosure, when the computer software is executed by a processor.

[0117] The term "non-transient" indicates that the data stored on the medium remains even after the medium is removed from the computer system or the power is turned off.

[0118] Examples of computer-readable non-transient recording media include hard disk drives, solid- state drives, optical disks (such as CDs or DVDs), USB drives, and memory cards. These storage devices are commonly used to store software programs, data files, and other digital content that can be accessed and processed by a computer system.

[0119] The present disclosure also relates to a method for evaluating cross-packaging level during AAV vector production as disclosed in the examples. This method uses an indicator plasmid that is added in the co-transfection step for AAV production; the indicator plasmid is a transgene plasmid containing an inactivated Cap gene comprising 3 stop codons to prevent expression of all 3 VP proteins (ITR_Cap9-STOP; Figure 3A). Since limited amount of Cap protein are produced from this plasmid, resulting AAV particles with ITR_Cap9-STOP genome indicate cross -packaging with other Cap proteins in the library. This plasmid is therefore used as indicator of cross -packaging level during AAV vector production.

[0120] The present disclosure also relates to a method of construction of AAV vector library as illustrated in Figure 2A (right panel), comprising transfecting cells with : i) a plasmid library of AAV vector genomes comprising an expression cassette for combinatorial variants of the capsid protein having diversity in at least two VR regions, wherein the expression cassette is flanked by AAV ITRs and further comprises a barcode for each VR introduced between the capsid coding sequence and the polyadenylation signal; (ii) an helper plasmid such as pXX6, containing adenoviral sequences necessary for AAV production, and (iii) a plasmid containing AAV Rep gene.

[0121] The present disclosure relates also to the method of construction of AAV vector library as illustrated in Figure 10.

[0122] Concatenation of barcodes, which correspond to each variant of each VR (Figure 1), allows high-throughput deep Illumina sequencing to quantify every variant in the library. Barcodes before polyA sequences also allow to examine the RNA level of transgene. Thank to barcode deep sequencing, multiple AAV properties can be studied in multiplex, including AAV productivity, AAV infection rate, gene expression after AAV infection, and AAV infection with the presence of pre-existing anti- AAV antibodies. It is important to note that since anti- AAV antibodies bind different regions of AAV capsid, single- VR mutation will not be able to help AAV to escape neutralization. On another hand, multi- VR library can be designed to mutate all VRs with known interaction with neutralizing antibodies. The data of each AAV property generated after screening can be used to train deep learning regression models according to the invention that predict the value of AAV properties from amino acid sequence of AAV capsid variants. AAV capsid amino acid sequences with desirable AAV properties are then selected using the prediction method according to the invention; for example, high AAV productivity, high infection rate in skeletal muscle, low infection rate in the liver, and be able to escape neutralizing antibodies.

[0123] The various embodiments of the present disclosure can be combined with each other and the present disclosure encompasses the various combinations of embodiments of the present disclosure.

[0124] The practice of the present invention will employ, unless otherwise indicated, conventional techniques, which are within the skill of the art. Such techniques are explained fully in the literature.

[0125] The invention will now be exemplified with the following examples, which are not limitative, with reference to the attached drawings in which:

[0126] FIGURE LEGENDS

[0127] Figure 1: The use of combinatory multi- VR AAV library and multiple AAV properties screening for AAV design

[0128] Figure 2: AAV library production. A. Modification of normal triple-transfection protocol for AAV library production. The Cap ORF in pRep-Cap plasmid was added 3 stop codons to remove expression of VP1, VP2, VP3 while unaffected the ORF of AAP, MAAP and X proteins. Cap proteins including VP1, VP2, VP3, however were produced from pITR_Transgene plasmid, which become AAV genome of final AAV product. Therefore, modification at protein level of Capsid protein can be determined by DNA sequence of the AAV genome by sequencing. B. AAV genome of the library with combinatory mutational variants from VR4 and VR8. Protein sequences of VR4 and VR8 can be determined by DNA sequence of the corresponding barcodes placed after the Capsid ORF. Barcodes before polyA sequences also allow to examine the RNA level of transgene.

[0129] Figure 3: Optimisation of trade-off between AAV production titer and Cross-Packaging level in AAV library production. A. Illustration of 4 pITR_Cap-ORF plasmids in the test library, which includes constructs encoding for Cap9, Cap9-3STOP - Cap9 ORF with 3 stop codon at the beginning of each VP protein, LICA1, and Cap9rh74. These 4 capsids variants are with known AAV productivity levels characterized in previous study (Cap9 > LICA1 > Cap9rh74 > Cap9-3STOP). B-D. Test plasmid library was subjected to AAV production with different amount of pITR plasmid during transfection step. B. AAV titers measured by ddPCR with ITR2 primers. C. AAV production level of individual capsid variants in the test library. AAV production level was defined as ratio of frequency of Cap-ORF AAV genome (f_aav) to frequency of pITR_Cap-ORF plasmid (f_pls), both measured by amplicon Illumina sequencing. The production level of Cap9-3STOP is considered as cross-packaging level during AAV library production. D. Comparison of AAV production titers and crosspackaging level of all diluted transfection conditions to the condition with 10 pls / c.

[0130] Figure 4: Dataset for different AAV properties. A-B. Comparison of log2 of reads per million (RPM) of every variants identified by deep amplicon sequencing between two replicates of libraries of plasmids (A) and purified AAV (B). R2: R score, PCC: Pearson correlation coefficients. C-E. Histograms of all filtered variants in AAVprod (C), ma_wt (D) and ma_dmd (E) datasets. Datasets were also divided in multiple subsets according to the distance of variants to Cap9rh74 wild-type sequence.

[0131] Figure 5: Model architecture of Example 1

[0132] Figure 6: Comparison of prediction accuracy measured by Spearman correlation coefficient between different inputs combinations, either aa_seq only or with f_input (f_pls in AAVprod measurement and f_aav in rna_wt or rna_dmd measurements).

[0133] Figure 7: Performance of trained models for each AAV property. Comparison of predicted values and measured values from NGS of different AAV variants in held-out test sets on AAV productivity (AAVprod), RNA level in WT myotubes (ma_wt) and RNA level in DMD myotubes (ma_dmd), respectively from left to right. Pearson correlation coefficient (PCC) and Spearman rank correlation coefficient (SRCC) were used to evaluate the model performance.

[0134] Figure 8: Sequences designed by deep learning models showed improvement in all 3 tested AAV properties. Histograms of 3 AAV properties; AAV productivity (A), RNA level in WT myotube (B), RNA level in DMD myotube (C), of library with random mutations (on VR4 and VR8) and library with designed mutations predicted to have both higher AAV productivity and RNA level in WT / DMD myotubes. The values of designed library were scaled by calibrating all common sequences found in both 2 libraries.

[0135] Figure 9: Computer device

[0136] Figure 10: Overview of multi- VR library construction. A. The final multi- VR capsid library with k VRs (kGZ, l<k<9) being diversified and barcoded separately (BCi). B. The scheme of oligo containing VR mutations and associated barcode. Two type-IIS enzymes are needed for the library construction. C. Step-by-step overview of library construction to obtain the final plasmid in A.

[0137] Figure 11: LIC library and datasets for simultaneously improving AAV production, transduction, and Nab-escape. A. Overview of LICA1 diversification where mutations are introduced in VR2 and VR8, while mutation in VR4 is preserved for muscle targeting. B-C. Reproducibility of the datasets across biological replicates. B. Comparison of AAV productivity across 2 biological replicates, with variants with more than one technical replicates - barcodes (left panel), or only one barcode per variant (right panel). C. Comparison of AAV RNA expression under different treatment of Nab+ serum. 50% and 95% neutralization conditions represent the serum concentration where 50% and 95% gene expression of LICA1-WT being reduced in a separate experiment.

[0138] Figure 12: Model architecture for predicting AAV viability from amino-acid sequence of VR2 and VR8 (Model of Example 2).

[0139] Figure 13. Multi- VR LICA library allow identification of Nab-escape variants. A. RNA level of 3 different sub-libraries (VR2-library (VR2 mutated+VR8 WT), VR8-library (VR2 WT+VR8 mutated), and VR2+VR8 library) (VR2 mutated+VR8 mutated) in different naturalization conditions (0%, 50%, and 95%, respectively). B. Sequence frequencies of all Nab-escaped variants with at least 2 descendants (variants with more mutations and derived from the same parent). The WT amino-acids are being excluded in each position. The sequences are grouped by the distance to the LICA1-WT.

[0140] EXAMPLES

[0141] Example 1: Predictive model of AAV9rh74 with random mutations in VR4 and VR8 for productivity and infectivity Material and methods

[0142] 1. AAV library production

[0143] All AAVs in the present study were produced in HEK293-T cells cultured in suspension using the modified triple transfection protocol, as further discussed in the result section. After 3 days post-transfection, cells were chemically lysed with Triton-XlOO (dilution 1 / 200, 2.5h, 37°C) and lysates were clarified by centrifugation (2000g, 15min, 4°C). AAV was precipitated from lysates by adding 40% PEG- 8000 at the final concentration of 8% and incubated overnight. The mixture was spun at 3500g for 30 minutes at 4°C, the pellet was collected and resuspended in 20ml TMS buffer. The material was then subjected to cesium chloride gradient purification, concentrated using an Amicon spin concentrator (Milipore), and stored at -80°C. The AAV titer was determined by qPCR after AAV DNA was extracted in triplicate using a MagNA Pure 96.

[0144] 2. Cell culture

[0145] Human immortalized myoblasts (originated from healthy control or DMD patients) were maintained in Skeletal Muscle Cell Growth Medium (PromoCell, C23060) and differentiated in Skeletal Muscle Differentiation Medium (PromoCell, C23061). In vitro AAV infection was performed by directly adding AAV into culture medium at the myotube stage (5 days after switching to differentiation medium) at the dose of 1E10 vg per 24-well plate well. After 48h post- infection, cells were washed, pelleted, and subjected to DNA and RNA analysis.

[0146] 3. RNA extraction and cDNA preparation

[0147] Total RNA was extracted from myotubes using an IDEAL32 extraction robot. DNA contamination was removed from the RNA samples with the TURBO DNA-free kit (Thermo Fisher Scientific). The cDNA was obtained by re verse- transcribing 1000 ng of total RNA with a mixture of random oligonucleotides and oligo-dT and the RevertAid H Minus First Strand cDNA Synthesis Kit (Thermo Fisher Scientific).

[0148] 4. Deep sequencing of barcodes

[0149] The barcodes from pITR-Cap plasmid, purified AAV DNA, and cDNA of myotubes were amplified by PCR using Q5 High-Fidelity 2X Master Mix (New England Biolas). For the amplification of plasmid and AAV DNA, 20 PCR cyclers were used to limit the PCR bias. For the amplification of cDNA samples, two sequential PCRs of 15 and 25 cyclers respectively were required. PCR products were subjected to deep amplicon sequencing at Genewiz (Germany), using the NovaSeq instrument (Illumina) following the Illumina protocol, resulting in approximately 350 million paired-end 2xl50bp reads per library. Custom python script was used to count the barcodes from fastq files. Only barcodes with 0 mismatch and Phred quality score of minimum 20 were selected for subsequent analysis. The sequencing depth was used to normalize the data using count per million (CPM) method.

[0150] Results

[0151] 1. AAV library production with minimized cross-packaging level

[0152] Library of pITR_Cap-ORF plasmids was subjected to AAV production using an adapted triple-transfection protocol for AAV library production (Figure 2A). In triple-transfection protocol, the Capsid protein comes from pRep-Cap plasmid while final AAV genome comes from pITR_Transgene plasmid, thus modification of Capsid protein sequence cannot be determined by AAV genome sequence. To enable multiplex correlation between AAV capsid protein sequence and AAV genome sequence using next-generation sequencing (NGS), capsid open reading frame (Cap-ORF) was used as transgene flanked between ITR sequences in pITR plasmid (Figure 2A-B). It is noteworthy that protein sequences of capsid protein can be determined by barcode sequences placed after Cap-ORF and before polyA sequence. These barcodes will allow short-read Illumina sequencing to determine the mutation in each VR used in library, regardless the distance between them. Besides, since barcodes were placed before polyA, it can be used to determine the RNA level of transgene (Capsid) expression upon infection. Capsid proteins generated from pRep-Cap plasmid was prevented by introducing 3 stop codons at the beginning of VP1, VP2, VP3 protein while keeping the ORFs of other essential proteins for AAV production, including AAP, MAAP, and X proteins. The proof- of-concept library presented in this study is the combination of mutation within VR4 and VR8 of hybrid Cap9rh74 capsid (Figure 2B).

[0153] Because pool of AAV variants was introduced in the same production, it risks multiple pITR_Cap-ORF plasmids enter the same producer cell, the resulting AAVs might have no correlation between genotype (DNA sequence) and phenotype (amino-acid sequence), usually referred as cross -packaging. To minimize cross-packaging during AAV library production, transfection at low copy of pITR_Cap-ORF plasmid was normally used to minimize the probability of multiple plasmid variants entering the same cell (Schmit et al., 2020). However, reduction of pITR_Cap-ORF plasmid significantly reduced AAV titer (Figure 3B). Therefore, the trade-off between cross-packaging level and resulting AAV titer need to be defined.

[0154] An assay to determine the cross-packaging level after AAV library production was then developed. The library of pITR_Cap-ORF plasmids include constructs of Cap9rh74, LICA1, and Cap9 with increasing order of the AAV productivity level, as previously characterized in the 400mL AAV production with normal transfection protocol (Figure 3A). This library also included similar construct of pITR_Cap9-STOP which contains Cap9 ORF with 3 stop mutations to prevent expression of all 3 VP proteins (Figure 3A). Since limited amount of Cap protein produced from this plasmid, resulting AAV particles with ITR_Cap9-STOP genome indicate cross-packaging with other Cap proteins in the library, will therefore be used as indicator of cross-packaging level during AAV production. Different quantity of pITR_Cap-ORF during AAV library production were tested and total AAV titers and crosspackaging levels were quantified. As expected, increased amount of pITR_Cap-ORF plasmids improved AAV production titers (Figure 2B), however increased cross -packaging levels as well (Figure 3C). Besides, differences in AAV production level between different capsid variants reduced with increasing level of cross-packaging (Figure 3C), yet can still be observed at condition of 500 pITR plasmids per cell (pls / c). However, while condition with 500 pls / c increased 28.2 times in AAV production titer compared to condition with 10 pls / c while increased cross-packaging level only 2.38-fold (Figure 3D). Therefore, AAV production can be significantly improved while remain low level of cross-packaging. In this study, the condition 10 pls / c was used to produce highest-quality data for training. Yet, for studies required much bigger AAV titers, conditions with 50-200 pls / c were used.

[0155] 2. Generation of training datasets for each AAV property

[0156] The study employed random mutations, including substitution, deletion, and insertion, to diversify the AAV library. The mutations were introduced simultaneously on VR4 and VR8 of Cap9rh74 wild-type amino-acid sequence. Each variant of the AAV capsid was encoded by a 32-nucleotide concatenated barcode, comprising a 16-Nu barcode for VR4 and another 16-Nu barcode for VR8 (Figure 2B). The study measured three properties of the AAV library, namely AAV productivity (AAVprod - Equation 7), RNA level of AAV transgene in healthy control human myotubes (rna_wt - Equation 2), and RNA level of AAV transgene in DMD human myotubes (ma_dmd - Equation 3). where f_pls and f_aav are the frequencies of each variant of ITR-containing plasmid and AAV genome in the population, and f_rna_wt and f_rna_dmd are the frequencies of transgene RNA levels of individual variant in WT and DMD myotubes.

[0157] The quantification of individual variants across replicates is reproducible, especially with high-count variants (Figure 4A-B). All three studied properties showed binary distribution with population of random mutations (Figure 4C-E), in agreement to previous study (Bryant et al., 2021). Next, the datasets were divided in multiple subsets according to number of mutations compared to wild-type sequence (Levenshtein distance). In both three datasets, the higher the distance to WT sequence, the higher proportion of variants with less AAV productivity and infectivity or lower percentage of variants with high AAV productivity and high transduction. It indicated a functional AAV library with meaningful biological outcomes.

[0158] After filtering low-quality and low-read variants, AAVprod dataset includes 690600 variants, ma_wt dataset includes 783006 variants, ma_dmd dataset includes 452780 variants. These datasets were split into training (used explicitly for training), validation (used for evaluating the model prediction during training) and held-out test sets (used for evaluating model prediction after training) in a ratio of 80%: 10%: 10%. The presented prediction was based on the test set, which was not involved in training the model.

[0159] 3. Model architecture for predicting AAV properties from capsid protein sequence

[0160] The study employed a deep learning model, the architecture of which is presented in Figure 5. The deep learning model accepted amino-acid sequences (aa_seq) of VR4 and VR8 as inputs. To prepare the aa_seq data for the model, each VR aa_seq was padded to the maximum length of 30 and projected to 2 or 3 embedding layers with a shape of (max_length, 20). These embedding layers included a token embedding layer to represent the 20 amino acids, a positional embedding layer to encode the raw position of each amino acid in the sequence, and an optional mutational embedding layer to represent the mutation of each amino acid compared to the wild-type sequence. The embedding layers were then summed to produce the final embedding of shape (max_length, 20). The embedding of each VR was then convolved by a ConvlD layer (filters=256, kemel_size=30) before being concatenated together and transformed by BatchNormalization. The resulting output was then passed through 8 transformer encoder-like blocks (Vaswani et al., 2017), which consist of self-attention (MultiHeadAttention layer hyperparameters: num_heads=8, key_dim=32, dropout=0.05) and 2 feed- forward layers. The output of the attention layer with a shape of (max_length_VR4+ max_length_VR8, 256) was then passed through a bidirectional LSTM, flattened, and a series of dense layers to obtain the final output (value of individual AAV property).

[0161] In addition, it should be noted that the frequencies of each variant's input (denominator) can be highly variable due to the nature of library preparation (f_pls in AAVprod measurement and f_aav in rna_wt or rna_dmd measurements). This variability may impact the final comparison between variants as some variants may be over- or under-represented within the library. To account for this, additional inputs were included in the model by concatenating f_pls for AAVprod and f_aav for ma level with the final dense layers (Figure 5). The inclusion of these values was found to significantly improve prediction accuracy compared to the model with aa_seq input only (Figure 6).

[0162] 4. Training configuration

[0163] Each regression model was trained using Tensorflow2 framework. The training was distributed to multiple GPUs (Nvidia A30) using MirroredStrategy (TensorFlow2). Adam optimiser with mean squared error loss were used for maximum 300 weight update steps. Callbacks for Learning Rate Schedule (ReduceLROnPlateau: from le-04 to le-06), EarlyStopping, and Modelcheckpoint were used during training.

[0164] 5. Model performance

[0165] The performance of the trained models was evaluated on the held-out test sets, consisting of 47,626 AAVprod, 44,432 ma_wt, and 32,627 rna_dmd variants. The models demonstrated high accuracy in predicting the properties of the new AAV variants (Figure 7). Specifically, the Pearson correlation coefficients (PCC) for the predicted values from the model predictions and experimental measured values were 0.87, 0.87, and 0.85 for the AAVprod, rna_wt, and ma_dmd models, respectively. Additionally, the Spearman rank correlation coefficients (SRCC) were 0.89, 0.90, and 0.91 for these three models.

[0166] 6. Using trained models to select sequences with combined improved AAV properties

[0167] The sets of 3 different trained models for each property (AAVprod, rna_wt, and ma_dmd) were used to evaluate new capsid sequences with mutations either in VR4 or VR8 as of in the initial library. The selected sequences for subsequent library (designed sequences) were predicted to be improved in AAVprod and RNA expression in either WT or DMD myotubes. Similar to the generation of training data, the designed capsid sequences were barcoded and subjected to AAV production, and subsequently infection in WT and DMD myotubes. Comparison of designed library with library of random mutations (initial training data) showed a significant increase in all three AAV properties (Figure 8). While in random library, only small percentage of sequences has high productivity or high RNA level in WT / DMD myotube, designed library show greater proportion of high AAV productivity in the binary distribution (Figure 8A). Similarly, majority of designed sequences showed shifted distribution towards higher RNA level in both WT and DMD myotubes, meaning that almost these sequences were able to infect WT / DMD myotubes (Figure 8B-C).

[0168] 7. Computer device

[0169] Figure 9 shows an example of a computer device 1 according to the invention. Said computer device 1 comprises:

[0170] - an input interface 2 to receive medical images,

[0171] - a memory 3 for storing at least instructions of a computer program,

[0172] - a processor 4 accessing to the memory 3 for reading the aforesaid instructions and executing the method according to the present document,

[0173] - an output interface 5 to provide an indication based on the above-mentioned score.

[0174] 8. Conclusions

[0175] The inventors have generated an AAV library which combined pools of mutations of different VRs. Therefore, each AAV variant from final library can carry mutations from multiple VRs. Concatenation of barcodes, which correspond to each variant of each VR (Figure 1), allows high-throughput deep Illumina sequencing to quantify every variant in the library. Thank to barcode deep sequencing, multiple AAV properties can be studied in multiplex, including AAV productivity, AAV infection rate, gene expression after AAV infection, and AAV infection with the presence of pre-existing anti- AAV antibodies. The data of each AAV property generated after screening was used to train deep learning regression models that predict the value of AAV properties from amino acid sequence of AAV capsid variants. Highly accurate deep learning models to predict AAV properties from capsid amino acid sequence were obtained. Using these models, amino acid sequences having desirable AAV properties were selected including high AAV productivity and high infection rate in skeletal muscle.

[0176] The use of homemade large multiple- VR AAV datasets containing a diverse range of mutations is conducive to the ability to learn and adapt with sequences of varying sizes by the models. Additionally, augmenting the datasets with information on plasmid frequencies used for AAV production or frequencies of AAV for infection significantly enhances the performance of the models. As a result, the models for each AAV property generated by this approach are capable of generalizing to unseen variants, thereby facilitating the design of novel sequences with accumulative desired properties.

[0177] Example 2: Predictive model of AAV LIC 1 with random mutations in VR2 and VR8 for productivity and Nab (Neutralizing antibody) escape

[0178] 1. Multi- VR library construction

[0179] The Inventors developed a versatile cloning technology that enables the construction of multidomain capsid libraries, allowing them to introduce mutations across multiple capsid regions without limitations imposed by the distance between domains or the number of domains targeted for mutation (Figure 10A). This approach utilizes two type Ils restriction enzymes, which facilitate the precise insertion of mutated regions while eliminating enzyme recognition sites from the final construct (Figures 10B-C, Enzyme 1 in black and Enzyme 2 in grey). The distinct orientations of these enzymes, marked by directional arrows, ensure accurate sequential assembly of mutations into the capsid. The library construction proceeds through two Golden Gate assembly steps for each mutated domain (Figure 10C). In the first step, an oligo pool containing the desired mutations and corresponding barcodes (restricted in length due to DNA synthesis constraints) is PCR- amplified and introduced into a recipient plasmid using the first Golden Gate reaction. In the second step, a PCR product bridging the two mutated domains (standardized across variants and without size constraints) is incorporated between the mutated region and its associated barcodes via a second Golden Gate reaction. This process is repeated sequentially for each VR domain, enabling the construction of k-domain libraries, where each additional domain is concatenated to complete the capsid open reading frame (ORF), followed by a series of barcodes representing each mutation.

[0180] The resulting barcode concatenation allows for high-throughput quantification of both mRNA (transgene expression) and DNA (viral genome) levels through next-generation sequencing. This technology allows for extensive diversification across multiple capsid domains, generating a highly complex library with a diversity potential that increases logarithmically relative to single-domain libraries (Diversity of k — VR library = (Diversity of VRil) X ■■■ X (Diversity of VRikf). This capability significantly expands the range of capsid phenotypes accessible for selection, providing a robust platform for optimizing AAV production, transduction efficiency, and Nab-escape properties.

[0181] 2. Combinatory library of VR2 / 8 derived from LIC 1 backbone

[0182] Using the technology, the Inventors developed a library of diversified LICA1 variants by introducing mutations in VR2 and VR8 - two regions critical for targeting and Nab evasion (Emmanuel, S.N. et al., 2022), while preserving VR4 for improved muscle targeting (Vu Hong, A. et al., 2024), in a LICA1 library (Figure 11A). This multi-variant library, libLICA (n=4.7 million), enabled a comprehensive assessment of AAV productivity, transduction efficiency, and Nab-escape potential.

[0183] To ensure data reliability, the Inventors confirmed reproducibility across biological replicates. AAV productivity, measured by the ratio of AAV genome and input plasmid DNA, showed high reproducibility when comparing two independent biological replicates (Figure 11B). Importantly, the reproducibility significantly improved with the inclusion of technical replicates within the library (different barcodes representing same capsid mutation). Similarly, reproducible trends were observed for AAV RNA expression under different neutralizing antibody (Nab) conditions, including serum concentrations that induced 50% and 95% neutralization of gene expression in LICA1-WT (Figure 11C). Of note, in the condition with high presence of Nabs, majority of variants was found absent or at low frequencies, indicating the strong neutralizing activity found in the serum.

[0184] The deep learning model used may have the architecture which is presented in Figure 12.

[0185] The multi- VR LICA library facilitated the identification of Nab-escape variants under various neutralization conditions. RNA expression levels of three distinct sub-libraries (VR2, VR8, and VR2+VR8 libraries) demonstrated differential Nab resistance profiles across conditions of 0%, 50%, and 95% neutralization (Figure 13A). Under strong neutralization conditions (50% and 95%), the majority of variants exhibited low expression levels, indicating susceptibility to Nab inhibition. However, a small subset within the double- VR library demonstrated high expression levels despite the presence of Nab, suggesting that these variants possess enhanced resistance to neutralization and maintain functionality under high Nab pressure. This highlights the potential of double- VR mutations to confer improved Nab escape properties compared to single- VR variants, which largely failed to express under similar conditions. Sequence analysis of Nab-escape variants showed an enrichment of specific mutations across VR2 and VR8 regions, with frequency distributions grouped by sequence distance from LICA1-WT, which further validated the Nab-escape properties of certain VR combinations (Figure 13B).

[0186] 3. Conclusion

[0187] In summary, the LICA library enabled a high-throughput analysis of AAV variants for optimizing production, transduction, and Nab evasion. By preserving muscle-targeting properties and strategically mutating key regions, the Inventors identified multiple LICA variants with promising therapeutic potential.

[0188] List of sequences disclosed in the application

[0189] SEQ ID NO : 1 Hybrid AAV9. rh74 capsid protein

[0190] MAADGYLPDWLEDNLSEGIREWWALKPGAPQPKANQQHQDNARGLVLPGYKYLGPGNGLD KGEPVNAADAAALEHDKAYDQQLKAGDNPYLKYNHADAEFQERLKEDTSFGGNLGRAVFQ AKKRLLEPLGLVEEAAKTAPGKKRPVEQSPQEPDSSAGIGKSGAQPAKKRLNFGQTGDTE SVPDPQP IGEPPAAPSGVGSLTMASGGGAPVADNNEGADGVGSSSGNWHCDSQWLGDRVI TTSTRTWALPTYNNHLYKQI SNSTSGGSSNDNAYFGYSTPWGYFDFNRFHCHFSPRDWQR LINNNWGFRPKRLNFKLFNIQVKEVTDNNGVKTIANNLTSTVQVFTDSDYQLPYVLGSAH EGCLPPFPADVFMIPQYGYLTLNDGSQAVGRSSFYCLEYFPSQMLRTGNNFQFSYEFENV

[0191] PFHSSYAHSQSLDRLMNPLIDQYLYYLSRTQSTGGTAGTQQLLFSQAGPNNMSAQAKNWL

[0192] PGPCYRQQRVSTTLSQNNNSNFAWTGATKYHLNGRDSLVNPGVAMATHKDDEERFFPSSG VLMFGKQGAGKDNVDYSSVMLTSEEEIKTTNPVATEQYGWADNLQQQNAAP IVGAVNSQ GALPGMVWQNRDVYLQGP IWAKIPHTDGNFHPSPLMGGFGMKHPPPQILIKNTPVPADPP TAFNKDKLNSFITQYSTGQVSVEIEWELQKENSKRWNPEIQYTSNYYKSNNVEFAVNTEG

[0193] VYSEPRP IGTRYLTRNL

[0194] SEQ ID NO : 2 VR4 of Hybrid AAV9. rh74 capsid protein

[0195] YLSRTQSTGGTAGTQQLLFS

[0196] SEQ ID NO : 3 VR8 of Hybrid AAV9. rh74 capsid protein

[0197] YGWADNLQQQNAAP I VGAV

[0198] SEQ ID NO : 8 VR2 of Hybrid AAV9. rh74 capsid protein

[0199] KEVTDNNGVKTI

[0200] SEQ ID NO : 9 of LICA1 capsid protein

[0201] MAADGYLPDWLEDNLSEGIREWWALKPGAPQPKANQQHQDNARGLVLPGYKYLGPGNGLDK GEPVNAADAAALEHDKAYDQQLKAGDNPYLKYNHADAEFQERLKEDTSFGGNLGRAVFQAK KRLLEPLGLVEEAAKTAPGKKRPVEQSPQEPDSSAGIGKSGAQPAKKRLNFGQTGDTESVP DPQP IGEPPAAPSGVGSLTMASGGGAPVADNNEGADGVGSSSGNWHCDSQWLGDRVITTST

[0202] RTWALPTYNNHLYKQI SNSTSGGSSNDNAYFGYSTPWGYFDFNRFHCHFSPRDWQRLINNN WGFRPKRLNFKLFNIQVKEVTDNNGVKTIANNLTSTVQVFTDSDYQLPYVLGSAHEGCLPP FPADVFMIPQYGYLTLNDGSQAVGRSSFYCLEYFPSQMLRTGNNFQFSYEFENVPFHSSYA HSQSLDRLMNPLIDQYLYYLSRTQSTDGRGDLGRLGPQQLLFSQAGPNNMSAQAKNWLPGP

[0203] CYRQQRVSTTLSQNNNSNFAWTGATKYHLNGRDSLVNPGVAMATHKDDEERFFPSSGVLMF GKQGAGKDNVDYSSVMLTSEEEIKTTNPVATEQYGWADNLQQQNAAP IVGAVNSQGALPG MVWQNRDVYLQGP IWAKIPHTDGNFHPSPLMGGFGMKHPPPQILIKNTPVPADPPTAFNKD KLNSFITQYSTGQVSVEIEWELQKENSKRWNPEIQYTSNYYKSNNVEFAVNTEGVYSEPRP IGTRYLTRNL

[0204] References

[0205] Bryant, D.H., Bashir, A., Sinai, S., Jain, N.K., Ogden, P.J., Riley, P.F., Church, G.M., Colwell, L.J., and Kelsic, E.D. (2021). Deep diversification of an AAV capsid protein by machine learning. Nature biotechnology 39, 691-696. DiMattia, M.A., Nam, H.J., Van Vliet, K., Mitchell, M., Bennett, A., Gurda, B.L., McKenna, R., Olson, N.H., Sinkovits, R.S., Potter, M., et al. (2012). Structural insight into the unique properties of adeno-associated virus serotype 9. Journal of virology 86, 6947-6958.

[0206] Emmanuel, S.N. et al. Structurally Mapping Antigenic Epitopes of Adeno-associated Virus 9: Development of Antibody Escape Variants. Journal of virology 96, e0125121 (2022).

[0207] Li, C., and Samulski, R.J. (2020). Engineering adeno-associated virus vectors for gene therapy. Nat Rev Genet 21, 255-272.

[0208] Ogden, P.J., Kelsic, E.D., Sinai, S., and Church, G.M. (2019). Comprehensive AAV capsid fitness landscape reveals a viral gene and enables machine-guided design. Science 366, 1139- 1143.

[0209] Schmit, P.F., Pacouret, S., Zinn, E., Telford, E., Nicolaou, F., Broucque, F., Andres-Mateos, E., Xiao, R., Penaud-Budloo, M., Bouzelha, M., et al. (2020). Cross -Packaging and Capsid Mosaic Formation in Multiplexed AAV Libraries. Molecular therapy. Methods & clinical development 17, 107-121.

[0210] Tseng, Y.S., and Agbandje-McKenna, M. (2014). Mapping the AAV Capsid Host Antibody Response toward the Development of Second Generation Gene Delivery Vectors. Front Immunol 5, 9.

[0211] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, E., and Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems 30.

[0212] Vu Hong, A. et al. An engineered AAV targeting integrin alpha V beta 6 presents improved myotropism across species. Nature communications 15, 7965 (2024).

[0213] Zhu, D., Brookes, D.H., Busia, A., Cameiro, A., Fannjiang, C., Popova, G., Shin, D., Donohue, K.C., Chang, E.F., Nowakowski, T.J., et al. (2022). Optimal trade-off control in machine learning-based library design, with application to adeno-associated virus (AAV) for gene therapy. bioRxiv.

Claims

CLAIMS1. A method implemented by computer means for the prediction of at least one property of an AAV vector, said method comprising the steps of:- obtaining the amino-acid sequences of at least two VR sequences of said AAV vector capsid protein,- predicting a value of said property of the AAV vector, based on said at least two VR sequences, using a regression model.

2. The method according to claim 1, wherein said regression model comprises an encoding part of a transformer model.

3. The method according to any of the preceding claims, wherein said regression model comprises at least two input paths upstream of said encoding part, each input paths receiving as input one of said VR sequences, each input path comprising embedding layers and a convolution layer, said embedding layers comprising a token embedding layer and a positional embedding layer, said regression model comprising an aggregation layer aggregating the outputs of said input paths.

4. The method according to any of the preceding claims, wherein said embedding layer comprises a mutational embedding layer.

5. The method according to claim 2 or claim 3, wherein said regression model comprises an LSTM layer, a flatten layer and at least one fully connected layer downstream said encoding part.

6. The method according to any of the preceding claims, wherein said at least one fully connected layer comprises two fully connected layers.

7. The method according to claim 5 or claim 6, wherein a tensor representing a feature of f_aav to f_pls ratio or f_rna to f_aav ratio is aggregated to the input of the flatten layer and to the input of each fully connected layer; wherein f_aav is the frequency of AAV vector genome; f_pls is the frequency of ITR-containing plasmid used for AAV vector production; and f_rna is the frequency of transgene RNA level in AAV target cell transduced with AAV vector.

8. The method according to any of the preceding claims, wherein said least one property of an AAV vector which is predicted is neutralizing antibody escape.

9. The method according to any of the preceding claims, wherein at least a first and a second property are predicted, based on said at least two VR sequences, using a respectively a first and a second regression model.

10. The method according to claim 9, wherein the AAV vector productivity may be predicted using a first regression model, the AAV vector infectivity in target cells or tissue may be predicted using a second prediction model and the neutralizing antibody escape of AAV vector may be predicted using a third prediction model.

11. The method according to claim 10, wherein the target is muscle cells or tissue.

12. The method according to any of the preceding claims, wherein said at least two VR sequences are either VR4 and VR8 or VR2 and VR8.

13. The method according to any one of claims 1 to 11, wherein said at least two VR sequences are not VR4 and VR5.

14. A computer software, comprising instructions to implement at least a part of the method according to any of the preceding claims, when the software is executed by a processor.

15. A computer device comprising:- an input interface to receive amino-acid sequences of at least two VR sequences,- a memory for storing at least instructions of a computer program according to the preceding claim,- a processor accessing to the memory for reading the aforesaid instructions and executing then the method according to any of the claims 1 to 12,- an output interface to provide an indication based on the predicted value of said property of the AAV vector.

16. A computer-readable non-transient recording medium on which a computer software is registered to implement the method according to any of the claims 1 to 13, when the computer software is executed by a processor.

Citation Information

Patent Citations

  • Novel adeno-associated virus capsid protein

    JP6994018B2

  • Engineered muscle and central nervous system compositions

    WO2023039476A1