Engineered envelope vectors and methods of use thereof

JP2025529142A5Pending Publication Date: 2026-09-07GIGAMUNE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025512668
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-31
Filing Date
2023-08-30
Publication Date
2026-09-07

AI Technical Summary

Technical Problem

Viral vectors face limitations such as low efficiency in delivering genes to specific target cells, particularly non-dividing or difficult-to-transduce cell types, induce immune responses, and lack precise targeting and selectivity, which hampers their clinical application.

Method used

Engineered envelope vectors comprising a viral envelope protein with limited mammalian cell binding and a non-viral membrane-associated protein for targeted gene delivery, optionally with a function modulator protein to enhance cell function modulation.

Benefits of technology

Enables targeted and efficient gene delivery to difficult-to-transduce cells, reduces immune response, and allows for precise screening and interaction testing with target cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure relates to novel engineered envelope vectors that can be used for gene delivery. The engineered envelope vectors comprise an engineered envelope that includes (a) a viral envelope protein and, optionally, (b) a non-viral membrane-associated protein. The present disclosure also provides methods for making and using engineered envelope vectors. The present disclosure is based on the discovery that a vector comprising a viral envelope protein (e.g., a "membrane fusogenic factor") with specific binding (e.g., tropism) to mammalian cells and overexpression of a second non-viral membrane-associated protein enables the vector to target and enter target cells, since the second non-viral membrane-associated protein functions for viral entry and gene delivery (e.g., a "target cell tropism protein").
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] 1. CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 402,936, filed August 31, 2022, which is hereby incorporated by reference in its entirety. 2. Sequence Listing

[0002] This application contains an electronically submitted Sequence Listing, which is hereby incorporated by reference in its entirety. The xml file, created on August 30, 2023, is named 53443 WO_sequencelisting.xml and is 22,452,493 bytes in size. [Background technology]

[0003] 3.Background Viral vectors, such as retroviruses, lentiviruses, adenoviruses, and adeno-associated viruses (AAVs), have revolutionized the field of gene delivery by providing gene transfer capabilities. Viral vectors have been used to develop gene therapies for the treatment of genetic disorders, cancer, and infectious diseases. However, they have limitations that prevent their more widespread clinical application.

[0004] One significant limitation is the low efficiency of gene delivery to specific target cells, particularly non-dividing or difficult-to-transduce cell types. This limits therapeutic potential and requires improved transduction efficiency. Furthermore, viral vectors often induce a host immune response, leading to complications and reduced efficacy. The immune response can result in vector elimination, vector particle neutralization, and inflammatory reactions, thereby limiting repeated administration and long-term gene expression. Precise targeting and selectivity of viral vectors to specific cell types or tissues is also required. Improving vector specificity minimizes off-target effects, enhances therapeutic efficacy, and reduces potential side effects. Furthermore, vectors with improved safety profiles are needed.

[0005] The development of vector systems without these improvements is crucial to accelerating the translation of gene therapy into effective clinical treatments. Summary of the Invention [Means for solving the problem]

[0006] 4. Abstract The present disclosure provides novel engineered envelope vectors that can be used for gene transfer. The engineered envelope vectors comprise an engineered envelope that contains (a) a viral envelope protein and, optionally, (b) a non-viral membrane-associated protein.

[0007] The present disclosure is based on the finding that a vector comprising a viral envelope protein (e.g., a "fusogen") with limited binding (e.g., tropism) to mammalian cells and overexpression of a second, non-viral membrane-associated protein allows the vector to target and enter target cells, since the second, non-viral membrane-associated protein functions for viral entry and gene delivery (e.g., a "target cell tropism protein"). The vector may further comprise a third membrane protein, i.e., a function modulator protein (e.g., a "function modulator"), that allows modulation of target cell function upon binding to the target cell.

[0008] These discoveries, as described herein, enable new and innovative methods for delivering nucleic acids to target cells in a target-specific manner. Furthermore, this vector can be used to screen cells (e.g., T cells) that are notoriously difficult to screen for specific antigens and functions, as well as to test interactions between surface proteins on the engineered envelope vector and target cells.

[0009] In one aspect, the disclosure provides an engineered envelope vector (e.g., an engineered lentivirus) comprising an engineered envelope, wherein the engineered envelope comprises (a) a viral envelope protein having at least 90%, 95%, 96%, 97%, or 99% sequence identity to a sequence selected from SEQ ID NOs: 1-167 or 8154-16497, and (b) optionally, a non-viral membrane-bound protein comprising: (i) an optional signal peptide (S); (ii) an extracellular targeting domain (ETD); and (iii) a membrane-binding domain (MBD). In some embodiments, the engineered envelope vector comprises: (i) a non-viral membrane-bound protein having the structure S-ETD-MBD, where S represents the signal sequence, ETD represents the extracellular targeting domain, and MBD represents the membrane-binding domain; and (ii) a mutant viral envelope protein comprising at least one mutation that reduces its native function. In some embodiments, the non-viral membrane-bound protein is encoded by a polynucleotide comprising the structure: S-ETD-MBD-IRES-R, where S encodes a signal sequence, ETD encodes an extracellular targeting domain, MBD encodes a membrane-bound domain, IRES encodes an internal ribosome entry site, and R encodes a reporter.

[0010] The vectors can include one or more mutant viral envelope proteins described in patent publications WO2020 / 236263A1 and US2020 / 0216502, and Nikolic et al., Nature Comm., 2018, 9:1029, the relevant disclosures of which are hereby incorporated by reference for this specific purpose.

[0011] Some aspects of the present disclosure provide compositions of engineered envelope vectors (e.g., engineered lentiviruses) comprising (i) a function modulator protein (e.g., a cell surface signal protein domain) and (ii) a non-mutated viral envelope protein, a viral envelope protein fragment, or a truncated viral envelope protein, or a mutant form (membrane fusogenic factor) containing at least one mutation that reduces its native function. The function modulator protein can be the extracellular domain of a cell surface protein that is naturally embedded in the cell membrane, or an endogenous secreted protein tethered to the engineered envelope vector with a fusion domain that embeds the protein in the cell membrane. In some embodiments, the function modulator protein binds to a receptor on a target cell.

[0012] Some aspects of the present disclosure provide methods for screening a population of cells, the methods comprising: (i) providing an engineered envelope vector (e.g., an engineered lentivirus) comprising a non-mutable "wild-type" viral envelope protein, a viral envelope protein fragment, or a truncated viral envelope protein, or a mutant (fusogenic factor) comprising at least one mutation that reduces its native function; a non-viral membrane-associated protein (target cell tropism protein) comprising a membrane-binding domain (MBD) and an extracellular targeting domain (ETD); a functional modulator protein (e.g., a protein ligand or receptor embedded in a lentivirus); and a nucleic acid encoding a reporter; (ii) allowing the engineered envelope vector to interact with a population of cells; and (iii) sorting the population of cells based on the presence or absence of the reporter. In some embodiments, the nucleic acid delivered by the engineered envelope vector encodes a nucleic acid barcode that is expressed by target cells upon delivery of the nucleic acid. In some embodiments, the target cells have been previously engineered to express distinct nucleic acid barcodes that identify a cell type or entity. In some embodiments, vector barcode and cell barcode are sequenced at single cell level to identify which vector is delivered to which cell target.In some embodiments, engineered envelope vector (for example, engineered lentivirus) comprises the nucleic acid of the coding sequence of non-viral membrane-bound protein, comprising the structure: S-ETD-MBD-IRES-R, wherein S codes for signal sequence, ETD codes for extracellular targeting domain, MBD codes for membrane-bound domain, IRES codes for internal ribosome entry site, and R codes for reporter.

[0013] In some embodiments, the engineered envelope vector comprises viral envelope proteins and / or other genomic sequences from a coronavirus, flavivirus, togavirus, arenavirus, bunyavirus, filovirus, orthomyxovirus, paramyxovirus, rhabdovirus, or any retrovirus other than a lentivirus.

[0014] In some embodiments, the target cell is a somatic cell (e.g., an antigen-specific cell, a T cell, or a B cell). In some embodiments, the cell is isolated from a subject (e.g., a human subject). In some embodiments, the cell is isolated from the subject's blood or tumor. In some embodiments, the cell is maintained in liquid culture prior to interaction with the engineered envelope vector. In some embodiments, the target cell is a primary human or mouse cell that has been immortalized through a cell engineering approach. In some embodiments, the target cell is an immortalized mammalian cell line that has been engineered to recombinantly express a target protein.

[0015] In some embodiments, the viral envelope protein comprises one or more of any of SEQ ID NOs: 1-167. In some embodiments, the viral envelope protein has at least 90%, 95%, 96%, 97%, or 99% sequence identity to a sequence selected from SEQ ID NOs: 1-167 or SEQ ID NOs: 8154-16497.

[0016] In some embodiments, the viral envelope protein is a chimera between two or more of any of SEQ ID NOs: 1-167 or 8154-16497. In some embodiments, the viral envelope protein is a sequence substantially comprising any fragment or domain of any of SEQ ID NOs: 1-167 or 8154-16497, e.g., a fragment of the transmembrane domain and the extracellular domain. In some embodiments, the viral envelope protein is a VSV-G envelope protein, a measles virus envelope protein, a Nipah virus envelope protein, or a coccus virus G protein. In some embodiments, the viral envelope protein is a VSV-G envelope protein having a mutation at any one or more of H8, K47, Y209, and / or R354. In some embodiments, the viral envelope protein is a measles virus envelope protein having a mutation at one or more of Y481, R533, S548, and / or F549. In some embodiments, the viral envelope protein is a Nipah virus envelope protein having a mutation at one or more of E501, W504, Q530, and / or E533. In some embodiments, the viral envelope protein is a coccal virus G protein having a mutation at K64 and / or R371. In some embodiments, the viral envelope protein comprises one or more mutations (e.g., deletions, insertions, substitutions, or chemical modifications) to any of SEQ ID NOs: 1-167 or 8154-16497 that confer beneficial properties, such as decreased or increased binding to a target cell of interest, or decreased or increased immunogenicity, or decreased or increased half-life in vivo or in vitro.

[0017] In some embodiments, the non-viral membrane-bound protein comprises a major histocompatibility complex (MHC) protein or a variant thereof. In some embodiments, the non-viral membrane-bound protein is a protein (e.g., interleukin-13), a peptide, or an antibody (e.g., an anti-CD19 antibody (SEQ ID NO: 8130), an anti-TCR antibody, an anti-MHC antibody, an anti-CD8 antibody (SEQ ID NO: 8137), an anti-CD5 antibody (SEQ ID NO: 16512), an anti-CD4 antibody (SEQ ID NO: 8135), an anti-CD7 antibody (SEQ ID NOs: 16513, 16514, 16515), an anti-CD3 antibody (SEQ ID NO: 8133), an anti-CD117 antibody (SEQ ID NO: 16511), or an anti-GPRC5D antibody (SEQ ID NOs: 16516-16519), or a variant thereof.

[0018] In some embodiments, the viral envelope protein is fused to one or more different molecules, in some embodiments, the one or more different molecules include a guide RNA and an exogenous endonuclease.

[0019] In some embodiments, the engineered envelope vector further comprises a nucleic acid construct enclosed in the engineered envelope.

[0020] In some embodiments, the nucleic acid construct comprises a coding sequence for a reporter protein. In some embodiments, the reporter is a fluorescent protein (e.g., green fluorescent protein, yellow fluorescent protein, red fluorescent protein) or an antibiotic resistance marker. In some embodiments, the nucleic acid construct comprises a coding sequence for a transgene, and optionally, the transgene is a therapeutic gene. In some embodiments, the nucleic acid construct comprises an inhibitory RNA, a catalytic RNA, or a sequence for CRISPR / Cas9 or other site-specific endonuclease-mediated mutagenesis. In some embodiments, the nucleic acid construct comprises a barcode sequence. In some embodiments, the nucleic acid construct comprises (i) a coding sequence for a reporter protein, (ii) a coding sequence for a therapeutic gene, (iii) an inhibitory RNA, (iv) a catalytic RNA, or (iii) a sequence for CRISPR / Cas9, CRISPR / Cas12a, or other site-specific endonuclease-mediated mutagenesis.

[0021] In some embodiments, the engineered envelope vector comprises a non-viral membrane-bound protein. In some embodiments, a linker is located between the membrane-binding domain and the extracellular targeting domain. The linker can be a rigid linker (e.g., a PDGFR stalk or a CD8a stalk), an Fc domain derived from IgG, a flexible linker (e.g., an amino acid sequence comprising GAPGAS, SEQ ID NO: 16548, or GGGGS, SEQ ID NO: 16549), or an oligomerization linker (e.g., an amino acid sequence capable of forming an IgG4 hinge or a tetrameric coiled-coil).

[0022] In some embodiments, the non-viral membrane-bound protein comprises any one or more of any of SEQ ID NOs: 168-5339, which are known or predicted to be membrane proteins. In some embodiments, these non-viral membrane-bound proteins are functional modulators. In some embodiments, these non-viral membrane-bound proteins are cell-tropic receptors. In some embodiments, the non-viral membrane-bound protein comprises any one or more of any of SEQ ID NOs: 168-5339. In some embodiments, the non-viral membrane-bound protein comprises a secreted protein, e.g., SEQ ID NOs: 5340-8121, which are known or predicted to be secreted proteins, which must be tethered to the surface of a vector, for example, by fusing the secreted protein to the transmembrane domain of another protein to create a recombinant chimeric protein. In some embodiments, the non-viral membrane-bound protein comprises one or more amino acid mutations to any of SEQ ID NOs: 168-8121 that confer beneficial properties, such as decreased or increased binding to a target cell of interest, or decreased or increased immunogenicity, or decreased or increased half-life in vivo or in vitro.

[0023] In some embodiments, a linker is positioned between the membrane-bound domain of a non-viral membrane-bound protein and its extracellular targeting domain. The linker may be a rigid linker (e.g., a PDGFR stalk or a CD8a stalk), an Fc domain from an IgG, a flexible linker (e.g., comprising an amino acid sequence comprising GAPGAS, SEQ ID NO: 16548, or GGGGS, SEQ ID NO: 16549), or an oligomerization linker (e.g., an amino acid sequence capable of forming an IgG4 hinge or a tetrameric coiled-coil).

[0024] In some embodiments, the non-viral membrane-bound protein comprises any one or more of SEQ ID NOs: 168-5339 that are known or predicted to be membrane proteins. In some embodiments, the non-viral membrane-bound protein comprises one or more chimeras selected from SEQ ID NOs: 168-5339. In some embodiments, the non-viral membrane-bound protein is a secreted protein, e.g., a known or predicted secreted protein selected from SEQ ID NOs: 5340-8121, which must be tethered to the vector surface by fusing the secreted protein to the transmembrane domain of another protein, i.e., generating a recombinant chimeric protein. In some embodiments, the non-viral membrane-bound protein comprises any amino acid mutation to any of SEQ ID NOs: 168-5339 that confers beneficial properties, such as reduced or increased binding to the target cell of interest, or reduced or increased immunogenicity, or reduced or increased half-life in vivo or in vitro. In some embodiments, a linker is positioned between the membrane-bound domain (MBD) of the non-viral membrane-bound protein and its extracellular targeting domain (ETD). The linker may be a rigid linker (e.g., a PDGFR stalk or a CD8a stalk), an Fc domain from an IgG, a flexible linker (e.g., comprising an amino acid sequence comprising GAPGAS, SEQ ID NO: 16548, or GGGGS, SEQ ID NO: 16549), or an oligomerization linker (e.g., an amino acid sequence capable of forming an IgG4 hinge or a tetrameric coiled-coil).

[0025] In some embodiments, the extracellular targeting domain (ETD) comprises a T cell receptor, an antibody, an MHC protein, or a modification thereof. In some embodiments, the extracellular targeting domain (ETD) comprises a targeting domain of a target cell-directed protein, and the membrane-binding domain (MBD) comprises a membrane domain of the target cell-directed protein. In some embodiments, the target cell-directed protein has a sequence selected from SEQ ID NOs: 168-8121, or a fragment thereof. In some embodiments, the extracellular targeting domain (ETD) comprises an antibody specific for CD3, CD4, CD5, CD7, CD8, CD19, CD20, or CD117.

[0026] In some embodiments, the non-viral membrane-bound protein further comprises an Fc domain and a linker located between the extracellular targeting domain (ETD) and the membrane-bound domain (MBD). In some embodiments, the viral envelope protein comprises at least one amino acid insertion, deletion, or substitution compared to a protein having a sequence selected from SEQ ID NOs: 1-167 or 8154-16497.

[0027] In another aspect, the disclosure provides a library of the engineered envelope vectors disclosed herein, in some embodiments, the library comprises 10, 100, 1,000, 10,000, 100,000, 1,000,000, or 10,000,000 unique clones of the engineered envelope vectors.

[0028] In some embodiments, each clone of the engineered envelope vector comprises a unique nucleic acid barcode. In some embodiments, each clone of the engineered envelope vector comprises a unique extracellular targeting domain (ETD). In some embodiments, each clone of the engineered envelope vector comprises a unique viral envelope protein.

[0029] In one aspect, the present disclosure provides a method for delivering a transgene or modifying a target cell using the engineered envelope vector or library disclosed herein. In another aspect, the present disclosure provides a method for gene therapy using the engineered envelope vector or library. In some embodiments, the engineered envelope vector is combined with a population of cells in vitro (ii) for 1 minute to 72 hours and at a temperature ranging from 4°C to 42°C. In some embodiments, the engineered envelope vector and the population of cells are combined in (ii) in the presence of: (a) cell culture medium, optionally RPMI or DMEM cell culture medium; (b) buffered saline, optionally phosphate-buffered saline or HEPES-buffered saline; and / or (c) a retroviral transduction enhancer, optionally heparin sulfate, polybrene, protamine sulfate, and / or dextran. In some embodiments, the extracellular targeting domain (ETD) of the engineered envelope vector is capable of binding to a cognate protein (e.g., a protein receptor) present on the cell surface of a subset of the cell population. In some embodiments, the population of cells is washed between (ii) and (iii) (e.g., with phosphate-buffered saline (PBS), e.g., to remove any remaining engineered envelope vector from the population of cells). In some embodiments, sorting of the cell population is performed using fluorescence-activated cell sorting, single-cell next-generation sequencing, or antibiotic selection.

[0030] In some embodiments, the target cell is an immune cell, optionally selected from a T cell and a Treg cell. In some embodiments, the target cell is an immune cell in a human subject.

[0031] In some embodiments, the nucleic acid construct comprises (i) a coding sequence for a reporter protein, (ii) a coding sequence for a therapeutic gene, (iii) an inhibitory RNA, (iv) a catalytic RNA, or (iii) a sequence for CRISPR / Cas9, CRISPR / Cas12a, or other site-specific endonuclease-mediated mutagenesis. In some embodiments, the nucleic acid construct comprises a coding sequence for a T cell receptor (TCR) or a chimeric antigen receptor (CAR). In some embodiments, the nucleic acid construct comprises a coding sequence for an endogenous gene, such as dystrophin.

[0032] In some embodiments, the engineered envelope vector is administered by intravenous or subcutaneous administration.

[0033] In some embodiments, the method further includes use of a second engineered envelope vector, wherein the second engineered envelope vector comprises a different extracellular targeting domain (ETD), and / or a different nucleic acid encoding a different reporter, compared to the first engineered envelope vector.

[0034] Another aspect of the present disclosure provides a method for delivering a nucleic acid (e.g., a gene of interest, e.g., a gene encoding a protein) to a target cell, the method comprising: (i) providing an engineered envelope vector comprising the nucleic acid, a viral envelope protein or a variant thereof comprising at least one mutation that reduces its native function (a fusogenic factor), and a non-viral membrane-bound protein comprising a membrane-binding domain and an extracellular targeting domain (a target cell tropism protein); and (ii) contacting the engineered envelope vector with a cell, thereby delivering the nucleic acid to the cell. In some embodiments, the engineered envelope vector enters the cell during (ii). In some embodiments, the method further comprises delivering a nucleic acid barcode, which can be used to trace back a transduction event to a particular engineered envelope vector among a library of two or more engineered envelope vectors. In some embodiments, gene delivery is performed in vivo by subcutaneous, intravenous, intramuscular, or intradermal injection. In some embodiments, gene delivery is performed in vitro or ex vivo using populations of primary or immortalized cells.

[0035] In some embodiments, engineered envelope vector is formulated to be encapsulated in lipid nanoparticles.In some embodiments, engineered envelope vector encapsulated in lipid nanoparticles is administered in vivo by subcutaneous, intravenous, intramuscular or intradermal injection.In some embodiments, encapsulation in lipid nanoparticles improves the half-life of engineered envelope vector, improves pharmacokinetics, and targets to therapeutically relevant cell types, thereby improving biodistribution.

[0036] In some embodiments, lipid nanoparticles comprising lipids tethered or linked to one or more proteins from any of SEQ ID NOS: 1-167 or 8154-15596 are used to deliver nucleic acids to target cells. In some embodiments, lipid nanoparticles comprising lipids tethered or linked to one or more proteins from any of SEQ ID NOS: 1-167 or 8154-15596 are used to deliver CRISPR / Cas proteins, zinc fingers, or other recombinant proteins for genome engineering.

[0037] Some aspects of the present disclosure provide a method for delivering a nucleic acid to a cell, the method comprising the steps of: (i) providing an engineered envelope vector comprising the nucleic acid, a viral envelope protein (membrane fusion factor), and a non-viral membrane-associated protein, e.g., SEQ ID NOs: 168-5339; and (ii) contacting the engineered envelope vector with a target cell, thereby delivering the nucleic acid to the cell.

[0038] Yet another aspect of the present disclosure provides a method for blocking interactions between an engineered envelope vector and a cell, the method comprising contacting a sample containing the engineered envelope vector and the cell with an antibody, wherein the engineered envelope vector comprises a membrane fusogenic factor and a non-viral membrane-associated protein, and the antibody binds to a cognate binder of the non-viral membrane-associated protein on the surface of the target cell. In some embodiments, the antibody blocks the interaction between the non-viral membrane-associated protein and its target cell protein target, thereby preventing transduction of the engineered envelope vector payload into the target cell. In some embodiments, the method is performed in parallel on tens, hundreds, thousands, or tens of thousands of non-viral membrane-associated protein targets, such that the antibodies can be used to identify interactions between complex mixtures of non-viral membrane-associated proteins and their cognate protein targets or target cell types.

[0039] Some aspects of the present disclosure provide a library of engineered envelope vectors comprising a plurality of unique engineered envelope vectors, each of which comprises a viral envelope protein (a fusogenic factor), a non-viral membrane-associated protein (e.g., a target cell tropism protein), and a nucleic acid encoding a reporter, and each of which comprises a different, unique extracellular targeting domain (ETD). Some aspects of the present disclosure provide a library of engineered envelope vectors comprising a plurality of unique vectors, each of which comprises a viral envelope protein (a fusogenic factor), a non-viral membrane-associated protein (e.g., a target cell tropism protein), a function modulator protein, and a nucleic acid encoding a reporter, and each of which comprises a different, unique extracellular targeting domain (ETD). In some embodiments, the library of engineered envelope vectors comprises a library of barcodes to be delivered as payloads to target cells. In some embodiments, the library of engineered envelope vectors comprises tens, hundreds, thousands, tens of thousands, hundreds of thousands, or millions of unique vector types in a single mixture.

[0040] In some embodiments, the library of engineered envelope vectors is derived from a library of packaging cells, and an envelope protein is engineered into multiple genomes of the packaging cells. In some embodiments, the multiple genomes of the packaging cells comprise a single envelope protein. In some embodiments, the library of packaging cells is used to generate the library of engineered envelope vectors by transfecting the library of packaging cells with a packaging plasmid that directs secretion of the engineered envelope vector from the packaging cells. In some embodiments, this library of engineered envelope vectors is used to transduce target cells. In some embodiments, the library of engineered envelope vectors comprises a library of lentiviral transgenes expressing viral envelopes. In some embodiments, the transduced target cells are subsequently sequenced to assess which viral envelope protein was associated with successful transduction. In some embodiments, the transduction is performed in vitro, for example, into an immortalized cell line, a library of different immortalized cell lines, or primary peripheral blood mononuclear cells. In some embodiments, the transduction is performed in vivo, ie, by injecting or injecting the library into a mouse, rat, or dog.

[0041] In some embodiments, the library of engineered envelope vectors is derived from a library of packaging cells, and a non-viral membrane-bound protein (e.g., an antibody or antibody fragment, or scFv) is engineered into the genomes of the packaging cells. In some embodiments, the genomes of the packaging cells comprise a single non-viral membrane-bound protein. In some embodiments, the library of packaging cells is used to generate the library of engineered envelope vectors by transfecting the library of packaging cells with a packaging plasmid that directs secretion of the engineered envelope vector from the packaging cells. In some embodiments, this library of engineered envelope vectors is used to transduce target cells. In some embodiments, the library of engineered envelope vectors comprises a library of lentiviral transgenes that express a non-viral membrane-bound protein (e.g., an antibody or antibody fragment, or scFv). In some embodiments, the transduced target cells are subsequently sequenced to assess which non-viral membrane-bound protein was associated with successful transduction. In some embodiments, the transduction is performed in vitro, e.g., on an immortalized cell line, a library of different immortalized cell lines, or primary peripheral blood mononuclear cells, hi some embodiments, the transduction is performed in vivo, i.e., by injecting or injecting the library into a mouse, rat, or dog.

[0042] In some embodiments, the library of engineered envelope vectors can be screened against a population of antigen-specific cells, optionally B cells or T cells. The library can be comprised of at least 10 2 Pieces, at least 10 3 Pieces, at least 10 4 Pieces, at least 10 5 Pieces, at least 10 6 Pieces, at least 10 7 Pieces, at least 10 8Pieces, at least 10 9 pieces, or at least 10 10 In some embodiments, the engineered envelope vectors may contain multiple unique vectors. In some embodiments, each different unique extracellular targeting domain (ETD) of the engineered envelope vectors is generated by site-directed mutagenesis. In some embodiments, a library or population of cells is recombinantly engineered to contain multiple unique cell types that express unique nucleic acid barcodes as RNA, contain unique nucleic acid barcodes in their genomes, or have unique RNA expression patterns that can be used to identify specific cells. In some embodiments, the population of cells is engineered into a multicellular organism, optionally as antigen-specific B cells or T cells transplanted into immunodeficient mice.

[0043] Some aspects of the present disclosure provide methods for identifying interactions between engineered envelope vectors and target cells using single-cell capture and high-throughput nucleic acid sequencing. In some embodiments, a single vector type is provided with a library of diverse cell types. In some embodiments, a single cell type is provided with a library of diverse vector types. In some embodiments, a diverse library of vectors is provided with a diverse library of cell types. Some aspects of the present invention include capturing single cells in emulsion microdroplets or microfluidic chambers such that multiple single cells are isolated with multiple single engineered envelope vectors. In some embodiments, vector transduction into target cells is measured by sequencing nucleic acid barcodes that are delivered to the cells and thereby expressed as RNA within the cells, and sequencing nucleic acid barcodes that are already expressed within the cells as RNA. In some embodiments, barcodes are not used; instead, sequencing is used to directly identify the delivered transgene. In some embodiments, transduction delivers a gene that confers selection resistance, and transduced cells are selected with a drug (e.g., an antibiotic). In some embodiments, reporters such as fluorescent proteins are delivered by transduction, and transduced cells are selected using FACS. In some embodiments, the selected transduced cells are then isolated in emulsion microdroplets or microfluidic chambers, and RNA barcodes are linked by overlap-extension RT-PCR. In some embodiments, overlap-extension RT-PCR molecules are sequenced en masse using high-throughput sequencing, thereby deconvoluting the interaction between the diverse library of engineered envelope vectors and target cells.In some embodiments, the RNA barcodes from the engineered envelope vectors are linked to the entire transcriptome of a cell, or a panel of 10, 100, or 1,000 transcripts, by overlap-extension RT-PCR in emulsion microdroplets or microfluidic chambers, such that the cellular entity or phenotype is linked to the RNA barcodes of the engineered envelope vectors via high-throughput sequencing of overlap-extension RT-PCR.

[0044] In some embodiments, the method includes: (a) contacting a plurality of single cells comprising a first nucleic acid barcode with an engineered envelope vector of the present disclosure, wherein the engineered envelope vector comprises a second nucleic acid barcode, such that the engineered envelope vector delivers the second nucleic acid barcode to said single cell; (b) isolating each of a plurality of said single cells into individual compartments, wherein each individual compartment is a microdroplet in an emulsion; (c) obtaining transcripts from each of the single cells; and (d) reacting the transcripts from each of the single cells with a first set of probes and a second set of probes. wherein each of the first probes is configured to bind to a first target polynucleotide comprising a first nucleic acid barcode and (ii) comprises the sequence of or a complementary sequence to a non-human exogenous sequence, and each of the second probes is configured to bind to a second target polynucleotide comprising a second nucleic acid barcode and (iv) comprises the sequence of or a complementary sequence to a non-human exogenous sequence; and (e) performing reverse transcription and subsequent PCR amplification using the first set of probes and the second set of probes, thereby generating a fusion complex comprising the first nucleic acid barcode and the second nucleic acid barcode.

[0045] In some embodiments, the method includes sequencing the fusion complexes. In some embodiments, each individual compartment has an average volume of 1 nanoliter (nL). In some embodiments, step (b) further includes introducing an mRNA capture agent and a cell lysis solution into the individual compartments. In some embodiments, the method includes obtaining sequences from at least 10,000 individual cells or at least 5,000 fusion complexes.

[0046] In some embodiments, step (e) comprises performing an overlap extension reverse transcriptase polymerase chain reaction to link the first nucleic acid barcode and the second nucleic acid barcode. In some embodiments, the first nucleic acid barcode and the second nucleic acid barcode are different.

[0047] In some embodiments, step (a) comprises contacting a plurality of single cells with a library of engineered envelope vectors disclosed herein. In some embodiments, the plurality of single cells comprises a library of 10, 100, 10,000, 100,000, 1,000,000 or more genetically distinct cells. In some embodiments, the plurality of single cells comprises a library of 10, 100, 10,000, 100,000 or 1,000,000 cells, each comprising a unique first nucleic acid barcode.

[0048] In another aspect, the present disclosure provides a method for identifying a pair of an engineered envelope vector and a target cell, the method comprising: (a) contacting a plurality of single cells with an engineered envelope vector disclosed herein, wherein the engineered envelope vector comprises an exogenous nucleic acid, such that the engineered envelope vector delivers the exogenous nucleic acid to the single cells; (b) isolating each of the plurality of single cells into individual compartments, each of the individual compartments being a microdroplet in an emulsion, each of the single cells comprising a first nucleic acid barcode; (c) obtaining transcripts from each of the single cells; and (d) performing reverse transcription followed by PCR amplification, thereby generating a fusion complex comprising the first nucleic acid barcode and the exogenous nucleic acid. In some embodiments, the exogenous nucleic acid encodes a reporter protein. In some embodiments, the exogenous nucleic acid comprises a second nucleic acid barcode. In some embodiments, the method further comprises sequencing the fusion complex. In some embodiments, the first barcode sequence is attached to a bead. In some embodiments, the exogenous nucleic acid comprises a coding sequence for a reporter protein, and the method further comprises sorting, selecting, or isolating cells that express the reporter protein.

[0049] In some embodiments, each individual compartment has an average volume of 1 nanoliter (nL). In some embodiments, step (b) further comprises introducing an mRNA capture agent and a cell lysis solution into the individual compartment.

[0050] In some embodiments, the method comprises obtaining sequences from at least 10,000 individual cells or at least 5,000 fusion complexes. In some embodiments, step (a) comprises contacting a plurality of single cells with a library of engineered envelope vectors disclosed herein. In some embodiments, the plurality of single cells comprises 10, 100, 10,000, 100,000, 1,000,000 or more unique cells.

[0051] Some aspects of the present disclosure provide a population of cells, a subset of which comprises a viral envelope protein (fusogenic factor), a non-viral membrane-associated protein, and an engineered envelope vector comprising a nucleic acid encoding a reporter. Some aspects of the present disclosure provide a population of cells, a subset of which comprises a viral envelope protein (fusogenic factor), a non-viral membrane-associated protein, a non-viral function modulator protein, and an engineered envelope vector comprising a nucleic acid encoding a reporter. In some embodiments, a subset of the population of cells (e.g., antigen-specific cells, e.g., B cells or T cells) comprises an engineered envelope vector described herein. In some embodiments, a subset of the cell population comprises an engineered envelope vector within each cell of the subset. The subset of the population comprising the engineered envelope vector can be isolated and / or sorted from cells of the population that do not comprise the vector. 5. Brief description of the drawings [Brief explanation of the drawings]

[0052] [Figure 1] Figure 1 shows an exemplary schematic diagram of an engineered envelope vector of the present invention having a functional modulator protein that interacts with its cognate ligand within a target cell, thereby generating a signal cascade in the target cell, a non-viral membrane-bound protein such as an scFv (e.g., a target cell tropism protein) that interacts with its cognate cell-type-specific cell surface target on the target cell, and a membrane-fusogenic pseudotype that triggers fusion of the vector to the target cell.

[0053] [Figure 2] Figure 2 shows an exemplary schematic diagram of an engineered envelope vector of the present invention having a non-viral membrane-bound protein (e.g., a target cell-tropic protein such as an scFv) that interacts with its cognate cell-type-specific cell surface target on a target cell, and a membrane-fusogenic pseudotype that triggers fusion of the vector to the target cell.

[0054] [Figure 3] Figure 3 shows an exemplary schematic diagram of an engineered envelope vector of the present invention having a functional modulator protein that interacts with its cognate ligand on a target cell, thereby generating a signal cascade in the target cell, and also serves as a non-viral membrane-associated protein (e.g., a target cell tropism protein) that interacts with its cognate cell-type-specific cell surface target on the target cell, and a membrane-fusogenic pseudotype that triggers fusion of the lentivirus or retrovirus to the target cell.

[0055] [Figure 4] Figure 4 shows the method used to identify novel fusogenic factors that can be used for the engineered envelope vectors of the present invention. A query sequence is used to search a genome sequencing database using a cloud-based sequence search algorithm. A tree of candidate sequences is generated, and the fusogenic factors are tested for in vitro function (e.g., cell type specificity, gene delivery efficiency, stability, etc.).

[0056] [Figure 5] Figure 5 shows a method used to test the function (e.g., cell-type specificity, gene delivery efficiency, stability, etc.) of novel membrane fusogenic factors in vitro. Each unique vector type has a unique barcode, and each unique cell or cell type has a unique barcode. After transducing a library of engineered envelope vectors into a library of cells, microfluidic technology is used to isolate single cells to ligate the barcodes of the engineered envelope vectors, thereby determining the functional properties of each unique vector in the library. Optionally, prior to microfluidic technology, transduced cells are purified or selected based on a reporter or cell selection gene.

[0057] [Figure 6]Figure 6 shows a method used to test the function (e.g., cell-type specificity, gene delivery efficiency, stability, etc.) of a novel fusogenic factor and its cognate antibody or scFv pair in vitro. Each unique vector type has a unique barcode, and each unique cell or cell type has a unique barcode. After transducing a library of engineered envelope vectors into a library of cells, microfluidic technology is used to isolate single cells and ligate the vector barcodes, thereby determining the functional properties of each unique vector in the library (specifically, the pairing of the scFv with the fusogenic factor). If necessary, prior to microfluidic technology, transduced cells are purified or selected based on a reporter or cell selection gene.

[0058] [Figure 7] Figure 7 illustrates the method used to test the function (e.g., cell-type specificity, gene delivery efficiency, stability, etc.) of novel membrane fusogenic and functional modulator proteins (i.e., a "surfaceome" library of over 1,000 endogenous cell surface proteins displayed on lentiviruses) in vitro. Each unique vector type has a unique barcode, and each unique cell or cell type has a unique barcode. After transducing a library of cells with a library of engineered envelope vectors, microfluidic technology is used to isolate single cells to ligate the vector barcodes, thereby determining the functional properties (specifically, surfaceome-to-cell type pairing) of each unique vector in the library. Optionally, prior to microfluidic technology, transduced cells are purified or selected based on reporter or cell-selection genes.

[0059] [Figure 8]Figure 8 illustrates the method used to test the function (e.g., cell-type specificity, gene delivery efficiency, stability, etc.) of novel membrane fusogenic and cognate pairs of functional modulator proteins (i.e., a "surfaceome" library of over 1,000 endogenous cell surface proteins displayed on lentiviruses) in vitro. Each unique vector type carries a unique barcode. After transducing target cells with the library of engineered envelope vectors, microfluidic technology is used to isolate single cells and ligate the vector barcode to the RNA transcript of interest in the target cell, thereby determining the functional properties of each unique vector in the library (specifically, any functional changes in the target cell predicted to be induced by the functional modulator protein). Optionally, prior to microfluidic technology, transduced cells are purified or selected based on a reporter or cell selection gene.

[0060] [Figure 9] Figure 9 illustrates a method used to identify cells transduced by a unique lentivector within a library of engineered envelope vectors. A population or library of cells is transduced with a library of vectors containing unique barcodes that identify the vector and its specific composition. Optionally, transduced cells are purified or selected based on a reporter or cell-selection gene. Microfluidic technology is used to isolate single cells using beads containing mRNA capture probes and nucleic acid barcodes. Subsequent cDNA synthesis and PCR are performed, in which the bead barcodes are ligated to cDNA generated from at least one RNA transcript from the single cell. The result is a library of cDNAs isolated from multiple single cells and barcoded at the single-cell level. The cDNA library is sequenced to identify RNA transcripts expressed by cells transduced with vectors bearing unique barcodes.

[0061] [Figure 10]Figure 10 illustrates a method used to identify cells transduced by unique vectors within a library of engineered envelope vectors. A population or library of barcoded cells is transduced with a library of engineered envelope vectors containing unique barcodes or other uniquely identifiable sequences that identify the vector and its specific composition. Optionally, transduced cells are purified or selected based on a reporter or cell selection gene. Microfluidic technology is used to isolate single cells using beads containing mRNA capture probes. Subsequent cDNA synthesis and PCR are performed to ligate cDNA generated from reporter or other RNA transcripts containing the vector barcode to cDNA generated from at least one barcoded RNA transcript ("cell barcode") from the single cell. The result is a library of fusion cDNAs isolated from multiple single cells, barcoded at the single-cell level. The cDNA library is sequenced to identify pairings between unique vector barcodes and unique cell barcodes, with the goal of mapping interactions between proteins expressed on the engineered envelope vectors and cell types.

[0062] [Figure 11] Figure 11 shows flow cytometry analysis demonstrating mRuby gene delivery to T cells by engineered envelope vectors containing (i) one of four different membrane fusogenic factor candidates (candidate A: SEQ ID NO: 84, candidate B: SEQ ID NO: 113, candidate C: SEQ ID NO: 121, and candidate D: SEQ ID NO: 2), and (ii) anti-CD19-targeting (encoded by DNA construct SEQ ID NO: 8130), anti-CD3-targeting (encoded by DNA construct SEQ ID NO: 8133), or non-targeting, non-viral membrane-bound proteins. A total of 12 engineered envelope vectors were evaluated.

[0063] [Figure 12]Figure 12 shows flow cytometry demonstrating mRuby gene delivery to specific PBMC populations using an engineered envelope vector containing fusogenic candidate C: SEQ ID NO: 121. The engineered envelope vector also contains non-viral membrane-bound proteins targeting CD19 (encoded by DNA construct SEQ ID NO: 8130), CD4 (encoded by DNA construct SEQ ID NO: 8135), and CD8 (encoded by DNA construct SEQ ID NO: 8137). Provided is vector-mediated delivery of mRuby (x-axis) directed toward anti-CD19, anti-CD4, or anti-CD8 scFv to CD4 and CD8 T cells (top, CD3+ parent gating) and CD19 B cells (bottom, CD20+ parent gating).

[0064] [Figure 13] Figure 13 shows how cell regulatory signals embedded in engineered envelope vectors can be used therapeutically to induce T cell activation while simultaneously delivering a gene that drives CAR expression. Activated CAR-expressing T cells are efficient at killing tumor cells.

[0065] [Figure 14] FIG. 14 shows a patient treatment regimen for AML using an engineered envelope vector that delivers anti-WT1 TCR to T cells in vivo.

[0066] [Figure 15] Figure 15 shows how anti-TFR2 CAR can be transduced into Tregs and transported towards the transplanted liver (because TFR2 is highly expressed on hepatocytes) to suppress host-versus-graft liver transplant killing by Teff (effector T cells).

[0067] [Figure 16]FIG. 16 shows a patient treatment regimen for liver transplantation using lentiviral or retroviral particles that deliver anti-TFR2 CARs to Tregs in vivo, thereby reducing or eliminating the need for patients to remain on traditional immunosuppressive drugs.

[0068] [Figure 17] Figure 17. An exemplary lentivector for delivery of Cas9 RNP to hematopoietic stem cells (HSCs) using anti-CD117 scFv and viral envelope protein for HSC targeting. The Cas9 RNP initially binds to the viral capsid protein but is proteolytically cleaved prior to transduction of target cells.

[0069] [Figure 18] Figure 18. Sequences including VSV-G and nine envelope proteins (SEQ ID NOs: 16522-16530) were subjected to multiple alignment using Clustalw. Clustalw was used to calculate the percent amino acid identity for each pair of sequence comparisons. The table shows these percent identities.

[0070] [Figure 19]Figure 19. Fifty lentivectors were generated containing VSV-G and nine envelope proteins (SEQ ID NOs: 16522-16530), delivering a GFP reporter transgene, with or without scFvs (non-viral membrane-bound proteins) directed against CD19, CD3, CD4, and CD8 for cell-type-specific tropism. The scFvs included the following sequences: CD19 (encoded by DNA construct SEQ ID NO: 8130), anti-human CD4 (SEQ ID NO: 8135), and anti-human CD8 scFv (encoded by DNA construct SEQ ID NO: 8137), as well as anti-CD3 antibody (encoded by DNA construct SEQ ID NO: 8133). Human PBMCs were transduced using the lentivectors and stained for the relevant cell type and GFP using flow cytometry. For each lentivector, sensitivity was calculated using the formula: sensitivity = true positives / (true positives + false negatives). The table shows the sensitivity of each lentivector.

[0071] [Figure 20] Figure 20. Fifty lentivectors were generated containing VSV-G and nine envelope proteins (SEQ ID NOs: 16522-16530), delivering a GFP reporter transgene, with or without scFvs (non-viral membrane-bound proteins) directed against CD19, CD3, CD4, and CD8 for cell-type-specific targeting. The scFvs included the following sequences: CD19 (encoded by DNA construct SEQ ID NO: 8130), anti-human CD4 (encoded by DNA construct SEQ ID NO: 8135), and anti-human CD8 scFv (encoded by DNA construct SEQ ID NO: 8137), as well as anti-CD3 antibody (encoded by DNA construct SEQ ID NO: 8133). Human PBMCs were transduced using the lentivectors and stained for the relevant cell type and GFP using flow cytometry. Specificity was calculated for each lentivector using the formula: specificity = true negatives / (true negatives + false positives). The table shows the specificity of each lentivector.

[0072] [Figure 21-1]Figure 21 is a sequence alignment of putative membrane fusion factors identified as described in Example 1. YP_008767242.1 is SEQ ID NO: 84, UBB42397.1 is SEQ ID NO: 2, AG00862.1 is SEQ ID NO: 36, YP_009505530.1 is SEQ ID NO: 76, ACB47442.1 is SEQ ID NO: 58, AAC02712.1 is SEQ ID NO: 9, AEI52254.1 is SEQ ID NO: 159, VSVG is SEQ ID NO: 16547, and AJR28459.1 is SEQ ID NO: 16548. Column number is 70, AJR28591.1 is sequence number 161, AFH89679.1 is sequence number 87, GI04017.1 is sequence number 8, AEG25354.1 is sequence number 60, BAA05163.1 is sequence number 121, ASK84898.1 is sequence number 69, CAH17547.1 is sequence number 155, and YP_009513006.1 is sequence number 142. [Figure 21-2] Same as above. [Figure 21-3] Same as above. DETAILED DESCRIPTION OF THE INVENTION

[0073] 6. Detailed Description of the Invention The present disclosure provides an engineered envelope vector comprising an engineered envelope, wherein the engineered envelope comprises: (a) a viral envelope protein having at least 90%, 95%, 96%, 97%, or 99% sequence identity to a sequence selected from SEQ ID NOs: 1-167; and (b) optionally, a non-viral membrane-bound protein (S-ETD-MBD) (i) an optional signal peptide (S); (ii) an extracellular targeting domain (ETD), and (iii) membrane-binding domain (MBD) Non-viral membrane-bound proteins, including The present invention provides an engineered envelope vector comprising:

[0074] The engineered envelope vectors can be used to deliver nucleic acids to target cells in a target-specific manner, and therefore this novel vector system can be used for gene therapy and experimental purposes.

[0075] Provided herein are novel and innovative methods for screening cells (e.g., T cells), which are notoriously difficult to screen for specific antigens and functions, for example. In some embodiments, described herein is a system that allows for repertoire-scale analysis of T cell receptor (TCR)-peptide-major histocompatibility complex (pMHC) specificities, which was previously an intractable bottleneck, for example, because previously described methods required significant effort to determine what a single T cell clone could recognize (e.g., as in a typical immune response).

[0076] In another aspect, the present invention describes a retrovirus or retrovirus-based system that reuses viral tropism as a method for selecting molecular interactions, for example, by encoding these protein variants on the corresponding transfer plasmid used to generate the virus, thereby ensuring that the resulting virus displays protein variants on its surface and packages corresponding gene sequences, and replaces the binding function of wild-type viral surface protein with the binding function of the target protein variant.Therefore, when the virus invades target cells (for example, has receptors that bind to the displayed extracellular targeting domain of protein variant), cell entry leads to the integration of the gene sequence of the displayed protein into the genome of target cells.

[0077] Previous approaches to studying T cell specificity required a combination of generated T cell lines, recombinant expression of T cell receptors, and / or individual validation of T cell binding or activity via candidate antigen-based approaches. Each of these elements provided inherent limitations in the throughput of T cells or antigens screened. For example, yeast display-based methods for deorphanizing T cell receptors alleviated the bottleneck in the number of antigens tested (10 8 While previous methods (with the ability to screen more than 10 ligands) were still severely limited by the need to recombinantly express the TCR, the current strategy of the invention described herein is 8 This represents a tremendous advance in the study of T cell specificity and screening of T cells by allowing the screening of individual ligands and eliminating the need for recombinant TCR expression. 6.1. Engineered envelope vectors

[0078] Described herein are engineered envelope vectors comprising a viral envelope protein (e.g., a "fusogenic factor"). In some embodiments, the engineered envelope vector further comprises a non-viral membrane-bound protein comprising a membrane-binding domain (MBD), an extracellular targeting domain (ETD), and, optionally, a signal peptide. In some embodiments, the engineered envelope vector further comprises a transgene nucleic acid (e.g., a reporter, a gene payload, or a nucleic acid barcode). The nucleic acid can be encapsulated in the engineered envelope comprising the viral envelope protein or a variant thereof and the non-viral membrane-bound protein.

[0079] Also described herein are engineered envelope vectors that include a viral envelope protein (membrane fusion factor), a non-viral membrane-bound protein (i.e., a target cell tropism protein) that includes a membrane-binding domain and an extracellular targeting domain, a function modulator protein (e.g., a non-viral membrane-bound protein involved in cell signaling), and a transgene nucleic acid (e.g., a reporter, gene payload, or nucleic acid barcode). In some embodiments, the engineered envelope vector includes a viral envelope protein that includes at least one mutation that reduces its native function, or a truncated or chimera of a viral envelope protein that reduces its native function. Examples of engineered vectors are shown in Figures 1-3.

[0080] The engineered envelope vector disclosed herein can comprise one or more elements (e.g., viral envelope proteins or genome sequences) derived from the retrovirus genome or lentivirus genome (natural or modified) of a suitable species. Retroviruses include seven families: alpharetroviruses (avian leukosis viruses), betaretroviruses (mouse mammary tumor viruses), gammaretroviruses (murine leukemia viruses), deltaretroviruses (bovine leukemia viruses), epsilonretroviruses (walleye cutaneous sarcoma viruses), lentiviruses (human immunodeficiency virus 1), and spumaviruses (human spumaviruses). Six additional examples of retroviruses are provided in U.S. Patent No. 7,901,671, which is hereby incorporated by reference.

[0081] In some embodiments, the retrovirus is a lentivirus.Lentivirus is a genus of retrovirus, and due to its ability to be integrated into the host genome, typically causes slowly progressing diseases.The modified lentivirus genome is useful as a viral vector for delivering nucleic acid to host cells.Host cells can be transfected with the lentivirus vector and, if necessary, an additional vector for expressing lentivirus packaging proteins (e.g., VSV-G, Rev, and Gag / Pol), to produce lentivirus particles in culture medium.When the engineered envelope vector comprises viral envelope proteins or genome sequences derived from retrovirus or lentivirus, the engineered envelope vector can be called retrovirus, lentivirus, retroviral vector, or lentiviral vector.

[0082] Retroviral and lentiviral constructs are well known in the art, and any suitable engineered envelope vector can be used to construct the engineered envelope vector (or a plurality of vectors or a library of vectors) described herein.Non-limiting examples of retroviral constructs include lentiviral vectors, human immunodeficiency virus (HIV) vectors, avian leukemia virus (ALV) vectors, murine leukemia virus (MLV) vectors, murine mammary tumor virus (MMTV) vectors, murine stem cell virus, and human T-cell leukemia virus (HTLV) vectors.These retroviral constructs contain proviral sequences derived from corresponding retroviruses.

[0083] The engineered envelope vectors described herein can contain viral elements, such as those described herein, derived from one or more suitable retroviruses, which are RNA viruses with single-stranded positive-sense RNA molecules. The engineered envelope vectors contain reverse transcriptase and integrase enzymes. Upon entering target cells, retroviruses and lentiviruses use their reverse transcriptase to transcribe their RNA molecules into DNA molecules. The DNA molecules are then integrated into the host cell genome using the integrase enzyme. Once integrated into the host cell genome, the sequence from the engineered envelope vector is referred to as a provirus (e.g., a proviral sequence or proviral sequence). The retroviral vectors described herein can further contain additional functional elements known in the art to address safety concerns and / or improve vector function, such as packaging efficiency and / or viral titer. Further information can be found in U.S. Patent Application Nos. 20150316511 and WO2015 / 117027, the relevant disclosures of each of which are hereby incorporated by reference for the purposes and subject matter referenced herein. Further information about lentiviruses and retroviruses can be found, for example, in WO2019 / 056015, the relevant disclosures of which are hereby incorporated by reference for this particular purpose.

[0084] In some embodiments, the engineered envelope vector is targeted to specific cells. In some embodiments, targeting is mediated by pMHC-TCR interactions or any other protein-protein or cell-cell interactions. In some embodiments, T cells with known relevant specificities can be enhanced (in the case of cancer or infectious diseases) or eliminated (in the case of autoimmunity) without affecting other T cells, dramatically limiting the risk of off-target effects. In some embodiments, the engineered envelope vector can include an extracellular domain for targeting any other surface-expressed molecule on the target cell. 6.1.1. Viral envelope proteins (membrane fusion factors)

[0085] In some embodiments, the viral envelope protein comprises one or more of any of SEQ ID NOs: 1-167 or 8154-16497. In some embodiments, the viral envelope protein has at least 90%, 95%, 96%, 97%, or 99% sequence identity to a sequence selected from SEQ ID NOs: 1-167. In some embodiments, the viral envelope protein is a chimera between two or more of any of SEQ ID NOs: 1-167 or 8154-16497. In some embodiments, the viral envelope protein is a sequence substantially comprising any fragment or domain of any of SEQ ID NOs: 1-167 or 8154-16497, e.g., fragments of the transmembrane domain and extracellular domain. In some embodiments, the viral envelope protein comprises at least one mutation or truncation that reduces its native function (i.e., the wild-type function of the non-mutated viral envelope protein). In some embodiments, the viral envelope protein is any viral envelope protein of any retrovirus (e.g., lentivirus).

[0086] In some embodiments, the viral envelope protein has no or very minimal natural cell tropism in mammals, humans, primates, and / or rodents. In some embodiments, the native function that is reduced by mutation of the viral envelope protein is viral tropism (e.g., the ability to infect a particular type of cell, the ability to bind to a cell, etc., via interaction with a cell surface target such as a protein). In some embodiments, the viral envelope protein contains one or more amino acid mutations (deletions, insertions, or substitutions) to any of SEQ ID NOS: 1-167 or 8154-16497 that confer beneficial properties, such as decreased or increased binding to a target cell of interest, decreased or increased immunogenicity, or decreased or increased half-life in vivo or in vitro. In some embodiments, the viral envelope protein has reduced viral tropism but retains the ability to fuse with the target cell membrane when combined with a second target cell tropism protein (a non-viral membrane-bound protein), such as an scFv, surface receptor, TCR, or pMHC.

[0087] In some embodiments, the viral envelope protein comprising at least one mutation or truncation that reduces its native function is a mutant VSV-G envelope protein or a mutated version of SEQ ID NOs: 1-167 or 8154-16497. In some embodiments, the viral envelope protein comprising at least one mutation or truncation that reduces its native function is a mutant measles virus envelope protein. In some embodiments, the viral envelope protein comprising at least one mutation or truncation that reduces its native function is a mutant Nipah virus envelope protein. In some embodiments, the viral envelope protein comprising at least one mutation or truncation that reduces its native function is a mutant coccoccus virus G protein.

[0088] In some embodiments, the viral envelope protein is mutated or truncated to reduce immunogenicity when administered to an organism in vivo. In some embodiments, HLA-presented peptides comprising amino acid sequences derived from the viral envelope protein are predicted using a computational method such as NetMHCPan. The predicted peptides are then mutated or protein domains comprising the predicted peptides are removed from the viral envelope protein. In some embodiments, the extramembrane domain of the viral envelope protein is substantially removed to reduce the epitopes available for antibody binding. In some embodiments, the intracellular domain of the viral envelope protein is replaced with the intracellular domain of another viral envelope protein so that engineered synthetic engineered envelope vectors can be produced more efficiently using chimeric proteins than using wild-type proteins.

[0089] In some embodiments, the mutant envelope protein is derived from any other enveloped virus, including, but not limited to, baculovirus, herpes simplex virus (HSV), cytomegalovirus (CMV), lymphocytic choriomeningitis virus (LCMV), Epstein-Barr virus (EBV), vaccinia virus, hepatitis A virus, hepatitis B virus, or hepatitis C virus, vaccinia virus, alphavirus, dengue virus, yellow fever virus, Zika virus, influenza virus, hantavirus, Ebola virus, rabies virus, human immunodeficiency virus (HIV), coronavirus, and other members of the Rhabdoviridae family.

[0090] In some embodiments, the viral envelope protein comprising at least one mutation comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more mutations. In some embodiments, the viral envelope protein comprising at least one mutation comprises a nucleotide sequence and / or amino acid sequence that is at least 50%, 60%, 70%, 80%, 90%, 95%, or 97% identical to the wild-type viral envelope protein. In some embodiments, the viral envelope protein comprising at least one mutation that reduces its native function retains less than 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, or 10% of the function of the wild-type viral envelope protein. In some embodiments, the viral envelope protein comprising at least one mutation lacks all of its native functions. In some embodiments, an engineered envelope vector comprising a viral envelope protein that contains at least one mutation that reduces its native function comprises less than 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20% or 10% of the cell infectivity of a retrovirus comprising a wild-type viral envelope protein.

[0091] In some embodiments, engineered envelope vectors comprising a viral envelope protein fused to a different protein or nucleotide (e.g., guide RNA, siRNA, ASO).

[0092] In some embodiments, the viral envelope protein is fused to a CRISPR / Cas enzyme, such that the guide RNA and the CRISPR / Cas enzyme are fused. The guide RNA and the CRISPR / Cas enzyme can be cleaved from the viral envelope protein after being delivered to the target cell. The CRISPR / Cas enzyme can include any type of nuclease-based gene editing technology, such as CRISPR / Cas9, CRISPR / Cas12a, or a base editor. An example of a construct for generating a fusion protein is a lentiviral packaging plasmid (SEQ ID NO: 16510) containing HIV-1 gag fused to SpCas9, with a 3xNES and a proteolytic cleavage site between them. See Figure 17.

[0093] The one or more DNA endonucleases may be Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas100, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cm r4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, or Cpf1 endonuclease; homologs thereof, recombinant naturally occurring molecules thereof, codon-optimized versions thereof, or modified versions thereof, and combinations thereof. In some embodiments, any of the DNA endonucleases disclosed in PCT / US2021 / 065554, filed December 29, 2021, and hereby incorporated by reference, is used. 6.1.2. Non-viral membrane-bound proteins

[0094] The engineered envelope vector described herein comprises a non-viral membrane-bound protein (target cell tropism protein).The non-viral membrane-bound protein may comprise a membrane-binding domain and an extracellular targeting domain.In some embodiments, the non-viral membrane-bound protein further comprises a signal peptide.

[0095] In some embodiments, the non-viral membrane-bound protein is a chimeric protein comprising sequences from at least two different proteins.In some embodiments, the non-viral membrane-bound protein (target cell-targeting protein) is a full-length or truncated protein comprising sequences from a single protein.In some embodiments, the target cell-targeting protein and the membrane fusogenic factor protein are fused to each other to form a chimeric protein.In some embodiments, the target cell-targeting protein is the same protein as the function modulator protein.

[0096] The membrane-bound domain of a non-viral membrane-bound protein is a protein or peptide having an amino acid sequence that allows the protein or peptide to be fully or partially embedded in or associated with the membrane (e.g., envelope) of an engineered envelope vector. In some embodiments, the membrane-bound domain allows the extracellular targeting domain to be presented and delivered to the extracellular environment. In some embodiments, the membrane-bound domain comprises an intracellular domain, a transmembrane domain, and / or an extracellular domain. In some embodiments, the membrane-bound domain comprises an intracellular domain and a transmembrane domain. In some embodiments, the membrane-bound domain comprises a major histocompatibility complex (MHC) protein or a fragment thereof. The MHC protein can be a class I or class II MHC protein, or a modified version thereof.

[0097] In some embodiments, the membrane binding domain comprises 10-50, 10-100, 25-100, 50-200, 50-150, 100-500, 100-250, 250-500, or any reasonable number of total amino acids.

[0098] In some embodiments, the engineered envelope vectors present in the library of engineered envelope vectors contain the same membrane-binding domain as some or all of the other vectors in the library. In some embodiments, each vector present in the library of engineered envelope vectors contains a different extracellular targeting domain (ETD) compared to some or all of the other vectors in the library. In some embodiments, the library of engineered envelope vectors includes a library of non-viral membrane-bound proteins, such as scFvs, TCRs, antibodies, or chimeric antigen receptors (CARs). In some embodiments, the library of engineered envelope vectors includes nucleic acid barcodes that can be used for high-throughput sequencing by identifying which vectors transduce which cells. In some embodiments, the library of antibodies or scFvs is a library of antigen-binding agents derived from immunized mice, human subjects, or yeast display libraries, and then rearranged as an engineered retroviral or lentiviral display library.

[0099] In some embodiments, the extracellular targeting domain (ETD) of a non-viral membrane-bound protein is any protein or peptide having an amino acid sequence that is a binding partner for a target molecule or ligand (e.g., a cognate protein) on the cell surface. The extracellular targeting domain can bind to a target cell when present in the extracellular environment beyond the interior of the engineered envelope vector. In some embodiments, the extracellular targeting domain binds to or targets a cognate protein or ligand (e.g., a protein receptor present on a target cell) present on the cell surface of a cell or a subset of a cell population. In some embodiments, the extracellular targeting domain binds to a cognate protein or ligand present on the cell surface of a single T cell or a subset of a T cell population. In some embodiments, the binding interaction between the extracellular targeting domain of the engineered envelope vector and the cognate protein or ligand of the cell allows the retrovirus to enter the cell (e.g., an antigen-specific cell, T cell). In some embodiments, the extracellular targeting domain comprises any of SEQ ID NOs: 168-8121 or any portion of any of SEQ ID NOs: 168-8121.

[0100] In some embodiments, the extracellular targeting domain comprises 10-50, 10-100, 25-100, 50-200, 50-150, 100-500, 100-250, 250-500, or any reasonable number of total amino acids, hi some embodiments, the extracellular targeting domain comprises at least 5, at least 10, at least 15, at least 20, or at least 50 amino acids.

[0101] In some embodiments, the extracellular targeting domain is a protein, antibody, or peptide. In some embodiments, the antibody is a full-length antibody, antibody fragment, nanobody, or single-chain antibody (scFv). In some embodiments, the extracellular targeting domain is an antibody that binds to a cognate protein on a target cell. In some embodiments, the extracellular targeting domain is an antibody that binds to a B cell antigen or a T cell antigen. In some embodiments, the extracellular targeting domain is an anti-CD19 antibody or antibody fragment (e.g., an antibody that binds to CD19). In some embodiments, the extracellular targeting domain is an antibody or antibody fragment that binds to a cell surface molecule. In some embodiments, the extracellular targeting domain is an antibody that binds to a lineage marker (e.g., CD3, CD4, CD8, CD20, integrin, or other receptor), a phenotypic marker (PD-1, CD25, CD45, or other). In some embodiments, the extracellular targeting domain is a protein or peptide that binds to a receptor (e.g., a receptor present on the surface of a target cell). In some embodiments, the extracellular targeting domain is a protein or peptide that binds to a cytokine receptor (e.g., interleukin-13 (IL-13) receptor). In some embodiments, the extracellular targeting domain is a cytokine (e.g., IL-2, IL-6, IL-12, IL-13). In some embodiments, the extracellular targeting domain is a chemokine ligand (e.g., CXCL9, CXCL10, CXCL11, etc.). In some embodiments, the extracellular targeting domain is a cellular receptor, including cytokine receptors (e.g., IL-13Rα1, IL-13Rα2, IL-2 receptor, common gamma chain), GPCRs (including chemokine receptors such as CSCR3, CXCR4), and integrins. In some embodiments, the extracellular targeting domain comprises stem cell factor (CSF), which binds to its receptor c-kit to direct targeting to CD34+ hematopoietic stem cells (HSCs). In some embodiments, the extracellular targeting domain is a peptide presented by an MHC protein.

[0102] In some embodiments, the non-viral membrane-bound protein comprises a membrane-bound domain (MBD) comprising an MHC protein or fragment and an extracellular targeting domain comprising a peptide presented by the MHC protein. In some embodiments, the extracellular targeting domain binds to a T cell receptor and / or a B cell receptor. T cell receptors are typically naturally expressed on the surface of T cells as α / β and γ / δ heterodimeric integral membrane proteins, with each subunit comprising a short intracellular segment, a single transmembrane α-helix, and two globular extracellular Ig superfamily domains. B cell receptors are transmembrane receptor proteins located on the outer surface of B cells. In some embodiments, the transmembrane domain of CD28 (e.g., SEQ ID NO: 8129) is used.

[0103] In some embodiments, the extracellular targeting domain is directed to a target cell or cell surface molecule. -9 ~10 -8 M, 10 -8 ~10 -7 M, 10 -7 ~10 -6 M, 10 -6 ~10 -5 M, 10 -5 ~10 -4 M, 10 -4 ~10 -3 M, or 10 -3 ~10 -2 In some embodiments, the extracellular targeting domain binds to a cognate protein or ligand on the target cell with a binding affinity of 10 -9 ~10 -8 M, 10 -8 ~10 -7 M, 10 -7 ~10 -6 M, 10 -6 ~10 -5 M, 10 -5 ~10 -4 M, 10 -4 ~10 -3 M, or 10 -3 ~10 -2M. In some embodiments, the binding affinity between the extracellular targeting domain and the cognate protein or ligand is in the picomolar to nanomolar range (e.g., about 10 -12 ~about 10 -9 In some embodiments, the binding affinity between the extracellular targeting domain and the cognate protein or ligand is in the nanomolar to micromolar range (e.g., about 10 -9 ~about 10 -6 In some embodiments, the binding affinity between the extracellular targeting domain and the cognate protein or ligand is in the micromolar to millimolar range (e.g., about 10 -6 ~about 10 -3 In some embodiments, the binding affinity between the extracellular targeting domain and the cognate protein or ligand is in the picomolar to micromolar range (e.g., about 10 -12 ~about 10 -6 In some embodiments, the binding affinity between the extracellular targeting domain and the cognate protein or ligand is in the nanomolar to millimolar range (e.g., about 10 -9 ~about 10 -3 M).

[0104] As used herein, the term antibody generally refers to a protein that includes at least one immunoglobulin variable domain or immunoglobulin variable domain sequence. For example, an antibody may include a heavy (H) chain variable region (herein referred to as a V H ) and / or light (L) chain variable region (abbreviated herein as V L In another example, the antibody may comprise two heavy (H) chain variable regions and / or two light (L) chain variable regions. The antibody may have structural characteristics of IgA, IgG, IgE, IgD, or IgM (and their subtypes). H Area and V L The regions can be further subdivided into regions of hypervariability called "complementarity determining regions" ("CDRs") interspersed with more conserved regions called "framework regions" ("FRs"). H and / or V LTypically, the V of an antibody is composed of three CDRs and four FRs, arranged in the following order from the amino terminus to the carboxy terminus: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4. H Chain or V L The chains can further comprise a heavy chain constant region or a light chain constant region, thereby forming an immunoglobulin heavy chain or an immunoglobulin light chain, respectively. In some embodiments, an antibody is a tetramer of two immunoglobulin heavy chains and two immunoglobulin light chains, where the immunoglobulin heavy chains and immunoglobulin light chains are interconnected, for example, by disulfide bonds. In IgG, the heavy chain constant region comprises three immunoglobulin domains: CH1, CH2, and CH3.

[0105] In some embodiments, an engineered envelope vector or library of vectors comprises the same extracellular targeting domain as some or all of the vectors in the library, hi some embodiments, each engineered envelope vector in a library of vectors comprises a different extracellular targeting domain compared to some or all of the other vectors in the library.

[0106] In some embodiments, the non-viral membrane-bound protein further comprises a signal sequence (also referred to as a localization sequence signal peptide). In some embodiments, the signal sequence is at the N-terminus or C-terminus of the non-viral membrane-bound protein. The signal sequence functions to translocate the non-viral membrane-bound protein to the retroviral membrane (or envelope). In some embodiments, the signal sequence is 5-10, 5-15, 10-20, 15-20, 15-30, 20-30, or 25-30 amino acids. In some embodiments, the signal sequence is an Ig kappa leader sequence (e.g., a mouse Ig kappa leader sequence comprising METDTLLLWVLLLWVPGSTG, SEQ ID NO: 16550) or a B2M signal peptide sequence (e.g., a B2M signal peptide sequence comprising MSRSVALAVLALLSLSGLEA, SEQ ID NO: 16551). In some embodiments, an engineered envelope vector present in a library of engineered envelope vectors comprises the same signal sequence as some or all of the other vectors in the library. In some embodiments, each vector present in a library of engineered envelope vectors contains a different signal sequence compared to some or all of the other engineered envelope vectors in the library.

[0107] In some embodiments, the nucleic acid encoding a non-viral membrane-bound protein further comprises an internal ribosome entry site (IRES). An IRES is an RNA sequence that allows for translation initiation during protein synthesis. In some embodiments, the IRES is located at or near the C-terminus. In some embodiments, the IRES is located C-terminal to the membrane-bound domain and the extracellular targeting domain (ETD). In some embodiments, the IRES is a viral IRES. In some embodiments, the IRES is a native IRES of a retrovirus. In some embodiments, the IRES is a sequence derived from encephalomyocarditis virus (EMCV). In some embodiments, vectors present in a library of engineered enveloped vectors contain the same IRES as some or all of the other vectors in the library. In some embodiments, each vector in a library of engineered enveloped viruses contains a different IRES compared to some or all of the other vectors in the library.

[0108] In some embodiments, the non-viral membrane-bound protein further comprises a linker located between the membrane-bound domain and the extracellular targeting domain. The linker can be an amino acid linker, such as a rigid linker, a flexible linker, or an oligomerization linker. A rigid linker is an amino acid sequence that lacks flexibility (e.g., it may contain at least one proline). In some embodiments, the rigid linker comprises a platelet-derived growth factor receptor (PDGFR) stalk, an Fc stalk, or a CD8α stalk. In some embodiments, the PDGFR stalk comprises amino acids comprising AVGQDTQEVIVVPHSLPFK (SEQ ID NO: 16552). In some embodiments, the PDGFR stalk comprises amino acids comprising ASAKPTTTPAPRPPTPAPTIASQPLSLRPEAARPAAGGAVHTRGLDFAK, SEQ ID NO: 16553.

[0109] A flexible linker is an amino acid sequence that has many degrees of freedom (e.g., it may contain multiple amino acids with small side chains, such as glycine or alanine). In some embodiments, a flexible linker comprises an amino acid sequence comprising GAPGAS (SEQ ID NO: 16548). In some embodiments, a flexible linker comprises an amino acid sequence consisting of GAPGSGGGGSGGGGSAS (SEQ ID NO: 16554). In some embodiments, a flexible linker comprises an amino acid sequence comprising GGGGS (SEQ ID NO: 16549). In some embodiments, a flexible linker comprises an amino acid sequence comprising (GAPGAS) N or (G4S) N where N is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more. An oligomerization linker is an amino acid that can oligomerize with another related amino acid. In some embodiments, an oligomerization linker is an amino acid sequence that can form a dimer, trimer, or tetramer. In some embodiments, an oligomerization linker comprises an IgG4 hinge domain (e.g., ASESKYGPPCPPCPAVGQDTQEVIVVPHSLPFK, SEQ ID NO: 16555). In some embodiments, an oligomerization linker comprises amino acids that can form a tetrameric coiled-coil (e.g., ASGGGGSGELAAIKQELAAIKKELAAIKWELAAIKQGAG, SEQ ID NO: 16556). In some embodiments, an oligomerization linker comprises an amino acid sequence that can form a dimeric coiled-coil (e.g., ASESKYGPPCPPCP, SEQ ID NO: 16557). 6.1.3. Function Modulator Proteins

[0110] In some embodiments, the engineered envelope vectors described herein further comprise a functional modulator protein. In some embodiments, the vector comprises a fusogenic factor and a functional modulator protein, but does not comprise a non-viral membrane-bound protein. In some embodiments, the retrovirus comprises a fusogenic factor and a non-viral membrane-bound protein for targeting and cell tropism, but does not comprise a functional modulator protein. In some embodiments, the vector comprises a fusogenic factor, a non-viral membrane-bound protein for targeting and cell tropism, and a functional modulator. In some embodiments, the functional modulator and the non-viral membrane-bound protein for targeting and cell tropism comprise the same protein. In some embodiments, the functional modulator and the non-viral membrane-bound protein for targeting and cell tropism are fused to each other to form a chimeric protein. In some embodiments, the functional modulator and the fusogenic factor are fused to each other to form a chimeric protein.

[0111] In some embodiments, the function modulator protein has activity on a target cell when the engineered envelope vector delivers a gene payload to the target cell. In some embodiments, the function modulator protein activates the target cell from a quiescent or low-activity state to a more active state. In some embodiments, the target cell is induced from a stem cell state to a non-stem cell state. In some embodiments, the function modulator protein induces a change in immune cell phenotype. In some embodiments, the functional modulator protein induces a change in the target cell that makes genome editing (e.g., modification using CRISPR / Cas9) more efficient. In some embodiments, the function modulator protein induces a change in the cell cycle state of the target cell. In some embodiments, the function modulator protein delivers a signal that induces apoptosis or programmed cell death. In some embodiments, the function modulator induces a change in chromatin structure, for example, from a heterochromatic state to a euchromatic state, or from a euchromatic state to a heterochromatic state. In some embodiments, the function modulator induces a change in the epigenetic state of the target cell. In some embodiments, the functional modulator binds to a cell surface protein on a target cell, thereby inducing an intracellular signaling cascade. In some embodiments, the functional modulator induces a kinase signaling cascade. In some embodiments, the functional modulator binds to a cell surface protein in a target cell, thereby directly or indirectly inducing the transcription, translation, or activation of a transcription factor. In some embodiments, the functional modulator binds to a cell surface protein in a target cell, thereby blocking the interaction between the cell surface protein and its endogenous ligand in the target cell. In some embodiments, the functional modulator protein blocks the interaction between the cell surface protein and its endogenous ligand in the target cell, thereby preventing the activation of an intracellular signal, thereby blocking a cellular function.In some embodiments, the function modulator protein provides a signal to lymphocytes to prevent the lentivirus from being degraded by lymphocytes. In some embodiments, the function modulator protein is a "don't eat me" signal, such as CD47, that signals lymphocytes not to degrade the lentivirus. In some embodiments, the function modulator protein induces a function in the same cells as the protein target of the target cell tropism protein. In some embodiments, the function modulator protein induces a function in cells other than the target cells of the target cell tropism protein. In some embodiments, the function modulator protein improves the serum half-life of the engineered envelope vector, thereby improving the pharmacological properties of the vector. In some embodiments, the function modulator contains a post-translational modification important for binding to a molecular target or stability.

[0112] In some embodiments, the function modulator protein comprises an extracellular signaling domain that is a binding partner for a cell surface target molecule or ligand (e.g., a cognate protein). When present in the extracellular environment beyond the interior of the viral envelope, the extracellular signaling domain is capable of binding to a target cell. In some embodiments, the extracellular signaling domain binds to or targets a cognate protein or ligand (e.g., a protein receptor present on a target cell) present on the cell surface of a cell or a subset of a cell population. In some embodiments, the extracellular signaling domain binds to a cognate protein or ligand present on the cell surface of a single T cell or a subset of a T cell population. In some embodiments, the extracellular targeting domain comprises any of SEQ ID NOs: 168-8121, or any portion of any of SEQ ID NOs: 168-8121. In some embodiments, the extracellular targeting domain has a sequence having at least 90%, 95%, 97%, 98%, or 99% identity to any one of SEQ ID NOs: 168-8121.

[0113] In some embodiments, the non-viral function modulator protein is a secreted protein, e.g., a known or predicted secreted protein having a sequence selected from SEQ ID NOs: 5340-8121, which can be tethered to the surface of a vector by fusing the secreted protein to the transmembrane domain of another protein, i.e., generating a recombinant chimeric protein. In some embodiments, the function modulator protein comprises one or more amino acid mutations (insertions, deletions, or substitutions) relative to any of SEQ ID NOs: 168-5339 that confer beneficial properties, such as decreased or increased binding to target cells of interest, or decreased or increased immunogenicity, or decreased or increased half-life in vivo or in vitro.

[0114] In some embodiments, the function modulator protein is a protein or peptide that binds to a receptor (e.g., a receptor present on the surface of a target cell). In some embodiments, the function modulator protein is a protein or peptide that binds to a cytokine receptor (e.g., an interleukin-13 (IL-13) receptor). In some embodiments, the function modulator protein is a cytokine (e.g., IL-2, IL-6, IL-12, IL-13). In some embodiments, the function modulator protein is a chemokine ligand (e.g., CXCL9, CXCL10, CXCL11, etc.). In some embodiments, the function modulator protein is a cellular receptor, including cytokine receptors (e.g., IL-13Rα1, IL-13Rα2, IL-2 receptor, common gamma chain), GPCRs (including chemokine receptors such as CSCR3, CXCR4), and integrins. In some embodiments, the function modulator protein is a peptide presented by an MHC protein. In some embodiments, the non-viral function modulator protein comprises an MHC protein or fragment and an extracellular targeting domain comprising a peptide presented by the MHC protein. In some embodiments, the function modulator protein binds to a T cell receptor and / or a B cell receptor. T cell receptors are typically naturally expressed on the surface of T cells as α / β and γ / δ heterodimeric integral membrane proteins, with each subunit containing a short intracellular segment, a single transmembrane α-helix, and two globular extracellular Ig superfamily domains. B cell receptors are transmembrane receptor proteins located on the outer surface of B cells.

[0115] In some embodiments, the function modulator protein binds to a target cell or cell surface molecule in a range of 10 -9 ~10 -8 M, 10 -8 ~10 -7 M, 10 -7 ~10 -6 M, 10 -6 ~10 -5 M, 10 -5 ~10 -4 M, 10-4 ~10 -3 M, or 10 -3 ~10 -2 In some embodiments, the function modulator protein binds to a cognate protein or ligand on the target cell with a binding affinity of 10 -9 ~10 -8 M, 10 -8 ~10 -7 M, 10 -7 ~10 -6 M, 10 -6 ~10 -5 M, 10 -5 ~10 -4 M, 10 -4 ~10 -3 M, or 10 -3 ~10 -2 In some embodiments, the binding affinity between the function modulator protein and the cognate protein or ligand is in the picomolar to nanomolar range (e.g., about 10 -12 ~about 10 -9 In some embodiments, the binding affinity between the functional modulator protein and the cognate protein or ligand is in the nanomolar to micromolar range (e.g., about 10 -9 ~about 10 -6 In some embodiments, the binding affinity between the function modulator protein and the cognate protein or ligand is in the micromolar to millimolar range (e.g., about 10 -6 ~about 10 -3 In some embodiments, the binding affinity between the function modulator protein and the cognate protein or ligand is in the picomolar to micromolar range (e.g., about 10 -12 ~about 10 -6 In some embodiments, the binding affinity between the function modulator protein and the cognate protein or ligand is in the nanomolar to millimolar range (e.g., about 10 -9 ~about 10 -3 M).

[0116] In some embodiments, the function modulator protein further comprises a signal sequence (also referred to as a signal peptide of a localization sequence). In some embodiments, the signal sequence is at the N-terminus or C-terminus of the function modulator protein. The signal sequence functions to translocate the function modulator protein to the membrane (or envelope) of the engineered envelope vector. In some embodiments, the signal sequence is 5-10, 5-15, 10-20, 15-20, 15-30, 20-30, or 25-30 amino acids. In some embodiments, the signal sequence is an Ig kappa leader sequence (e.g., a mouse Ig kappa leader sequence comprising METDTLLLWVLLLWVPGSTG, SEQ ID NO: 16550) or a B2M signal peptide sequence (e.g., a B2M signal peptide sequence comprising MSRSVALAVLALLSLSGLEA, SEQ ID NO: 16551). In some embodiments, an engineered envelope vector present in a library of vectors comprises the same signal sequence as some or all of the other vectors in the library. In some embodiments, each engineered envelope vector present in the library of vectors comprises a different signal sequence compared to some or all of the other engineered envelope vectors in the library.

[0117] In some embodiments, the nucleic acid encoding the function modulator protein further comprises an internal ribosome entry site (IRES). An IRES is an RNA sequence that allows for translation initiation during protein synthesis. In some embodiments, the IRES is located at or near the C-terminus. In some embodiments, the IRES is located C-terminal to the membrane-bound function modulator protein domain and the extracellular function modulator protein domain. In some embodiments, the IRES is a viral IRES. In some embodiments, the IRES is a native IRES of a retrovirus. In some embodiments, the IRES is a sequence derived from encephalomyocarditis virus (EMCV). In some embodiments, retroviruses present in a library of engineered envelope vectors contain the same IRES as some or all of the other vectors in the library. In some embodiments, each engineered envelope vector present in a library of vectors contains a different IRES compared to some or all of the other vectors in the library.

[0118] In some embodiments, the function modulator protein further comprises a linker located between the membrane-binding domain and the extracellular domain. The linker can be an amino acid linker, such as a rigid linker, a flexible linker, or an oligomerization linker. A rigid linker is an amino acid sequence that lacks flexibility (e.g., it may contain at least one proline). In some embodiments, the rigid linker comprises a platelet-derived growth factor receptor (PDGFR) stalk, an Fc stalk, or a CD8α stalk. In some embodiments, the PDGFR stalk comprises an amino acid sequence comprising AVGQDTQEVIVVPHSLPFK, SEQ ID NO: 16552. In some embodiments, the PDGFR stalk comprises amino acids comprising ASAKPTTTPAPRPPTPAPTIASQPLSLRPEAARPAAGGAVHTRGLDFAK, SEQ ID NO: 16553.

[0119] A flexible linker is an amino acid sequence that has many degrees of freedom (e.g., it may contain multiple amino acids with small side chains, such as glycine or alanine). In some embodiments, a flexible linker comprises an amino acid sequence comprising GAPGAS (SEQ ID NO: 16548). In some embodiments, a flexible linker comprises an amino acid sequence consisting of GAPGSGGGGSGGGGSAS (SEQ ID NO: 16554). In some embodiments, a flexible linker comprises an amino acid sequence comprising GGGGS (SEQ ID NO: 16549). In some embodiments, a flexible linker comprises an amino acid sequence comprising (GAPGAS) N or (G4S) N where N is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more. An oligomerization linker is an amino acid that can oligomerize with another related amino acid. In some embodiments, an oligomerization linker is an amino acid sequence that can form a dimer, trimer, or tetramer. In some embodiments, an oligomerization linker comprises an IgG4 hinge domain (e.g., ASESKYGPPCPPCPAVGQDTQEVIVVPHSLPFK, SEQ ID NO: 16555). In some embodiments, an oligomerization linker comprises amino acids that can form a tetrameric coiled-coil (e.g., ASGGGGSGELAAIKQELAAIKKELAAIKWELAAIKQGAG, SEQ ID NO: 16556). In some embodiments, an oligomerization linker comprises an amino acid sequence that can form a dimeric coiled-coil (e.g., ASESKYGPPCPPCP, SEQ ID NO: 16557).

[0120] In some embodiments, the function modulator protein tethered to the engineered envelope vector comprises at least one mutation, or comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more mutations compared to the corresponding wild-type sequence. In some embodiments, the function modulator protein comprising at least one mutation comprises a nucleotide sequence and / or amino acid sequence that is at least 50%, 60%, 70%, 80%, 90%, 95%, or 97% identical to the wild-type function modulator protein. In some embodiments, the function modulator protein comprising at least one mutation that reduces its native function retains less than 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, or 10% of the function of the wild-type function modulator protein. In some embodiments, the function modulator protein comprising at least one mutation lacks some or all of its native function. In some embodiments, the function modulator protein comprises at least one mutation that increases its native function by 1,000%, 500%, 100%, 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, or 10% over the function of the wild-type function modulator protein. In some embodiments, the increase or decrease in function of the function modulator protein is associated with an increase or decrease in binding affinity between the function modulator protein and a cognate ligand or receptor. In some embodiments, the increase or decrease in function of the function modulator protein is associated with abrogation of binding between the function modulator protein and a cognate ligand or receptor. 6.1.4. Transgene

[0121] The engineered envelope vector may comprise a nucleic acid construct enclosed in an engineered envelope. In some embodiments, the nucleic acid construct comprises a transgene.

[0122] The introduction of a specific foreign gene or native gene into a host or target cell is facilitated by incorporating the transgene sequence into a suitable gene delivery vector. The transferred nucleic acid, including structural and / or regulatory sequences, is then referred to as a transgene, and cells or organisms containing the transgene are called transgenic cells and transgenic organisms, respectively. Various methods have been developed to introduce such recombinant gene delivery vectors into desired host cells. In contrast to methods involving DNA transformation or transfection, the use of viral vectors can result in the rapid introduction of recombinant nucleic acids (and their RNA transcripts and / or translated protein products) into a wide variety of host cells. In particular, viral vectors have been used to increase the efficiency of introducing recombinant nucleic acid (gene) vectors into host or target cells.

[0123] Retroviral and lentiviral based vectors have been used as tools to achieve stable and integrative gene transfer of foreign genes into cells. Retroviruses that have been used as vectors for the introduction and expression of exogenous genes in cells include Moloney murine sarcoma virus (T. Curran et al., J. ViroL. 44, 674-682 (1982); A. Gazit et al., J. Virol. 60, 19-28 (1986)) and murine leukemia virus (MuLV; A.D. Miller, Curr. Tsp. Microbiol. Immunol. 158, 1-24 (1992)). Other viruses that have been used as vectors for the transduction and expression of exogenous genes in mammalian cells include SV40 virus (see, e.g., H. Okayama et al., Molec. Cell Biol. 5, 1136-1142 (1985)), bovine papilloma virus (see, e.g., D. DiMalo et al., Proc. Natl. Acad. Sci. USA 79, 4030-4034). (1982)), adenovirus (see, e.g., J.E. Morin et al., Proc. Natl. Acad. Sci. USA 84, 4626 (1987)), adeno-associated virus (AAV, see, e.g., N. Muzyczka et al., J. Clin. Inveit. 94, 1351 (1994)), herpes simplex virus (see, e.g., A.I. Geller, et al., Science 241, 1667 (1988)), and the like.

[0124] In some embodiments, the transgene is encoded in a DNA plasmid that is transfected into the packaging cell, whereby it is packaged in an engineered envelope vector and secreted. In some embodiments, the transgene is inserted into the genome of the packaging cell, whereby it is packaged in an engineered envelope vector and secreted. In some embodiments, the transgene inserted into the genome of the packaging cell is under the control of an inducible promoter, for example, a promoter that contains an element that is conditionally activated in the presence of a molecule such as tetracycline or mifepristone.

[0125] In some embodiments, the nucleic acid construct contains regulatory elements that affect the transcription or translation of the transgene. For example, the nucleic acid construct can contain a promoter of eukaryotic or prokaryotic origin sufficient to direct the transcription of a distally located sequence within the cell (e.g., a sequence linked to the 5' end of the promoter sequence). In some embodiments, the promoter region further contains control elements for enhancing or suppressing transcription. Suitable promoters include the cytomegalovirus immediate-early promoter (pCMV), the Rous sarcoma virus long terminal repeat promoter (pRSV), and the SP6, T3, or T7 promoters. An enhancer sequence upstream from the promoter or a terminator sequence downstream of the coding region can also be included in the gene delivery vector of the present invention to facilitate transgene expression. The nucleic acid construct of the present invention can also contain additional nucleic acid sequences (e.g., polyadenylation sequences, localization sequences, or signal sequences) sufficient to enable the cell to efficiently and effectively process the protein expressed by the nucleic acid of the engineered envelope vector. Examples of preferred polyadenylation sequences are the SV40 early region polyadenylation site (CV Hall et al., J Molec. App. Genet. 2, 101 (1983)) and the SV40 late region polyadenylation site (S. Carswell and JC Alwine, Mol. Cell Biol. 9, 4248 (1989)). Such additional sequences can be included in the gene delivery vector so that they are operably linked to a promoter sequence, if transcription is desired, or further operably linked to an initiation sequence and processing sequence, if translation and processing are desired. Alternatively, the inserted sequence can be located at any position in the vector.The term "operably linked" is used to describe the linkage between a gene sequence and a promoter or other regulatory or processing sequence such that transcription of the gene sequence is directed by the operably linked promoter sequence, translation of the gene sequence is directed by the operably linked translational control sequence, and post-translational processing of the gene sequence is directed by the operably linked processing sequence.

[0126] As will be understood by those skilled in the art, the nucleotide sequence of the inserted heterologous transgene sequence(s) can be any nucleotide sequence. For example, the inserted heterologous transgene sequence can be a reporter gene sequence or a selectable marker gene sequence. As used herein, a reporter transgene sequence is any gene sequence that, when expressed, results in the production of a protein whose presence or activity can be monitored. Examples of suitable reporter genes include green fluorescent protein (GFP), luciferase, and the like. Alternatively, a reporter gene sequence can be any gene sequence whose expression produces a gene product that affects the physiological function of a cell. A preferred reporter or selectable marker gene sequence is sufficient to enable the recognition or selection of the vector in normal cells. In one embodiment of the present invention, the reporter gene sequence encodes an enzyme or other protein not normally present in mammalian cells, and therefore, its presence can definitively indicate the presence of the vector in such cells. The heterologous gene sequence of the present invention can include one or more gene sequences that already possess one or more promoters, initiation sequences, or processing sequences.

[0127] The heterologous gene sequence may also comprise a coding sequence for a desired product, such as a suitable biologically active, immunogenic or antigenic, or therapeutically active protein or polypeptide. Alternatively, the heterologous gene sequence may comprise a sequence complementary to an RNA sequence, e.g., an antisense RNA sequence, which can be administered to an individual to inhibit expression of the complementary polynucleotide in the individual's cells.

[0128] Expression of the heterologous gene can result in an immunogenic or antigenic protein or polypeptide to achieve an antibody response, after which antibodies can be collected from the animal in body fluids such as blood, serum, or ascites. Expression of the heterologous gene can result in an immunogenic or antigenic protein or polypeptide to achieve a T cell response, after which T cell receptor sequences can be captured from T cells.

[0129] The transgenes of the present invention can also be applied to provide a means for controlling the expression of proteins and assessing their ability to modulate cellular events. Some functions of proteins, such as their role in differentiation, can be studied in tissue culture, while other functions will require reintroduction into in vivo systems at various times during development to monitor changes in relevant properties.

[0130] Transgenes also have substantial potential applications in understanding disease states and providing treatments. There are many genetic diseases for which the defective genes are known and have been cloned. In some cases, the functions of these cloned genes are known. Generally, these disease states fall into two classes: deficiency states, which are usually enzyme deficiency states and are generally inherited in a recessive manner, and imbalance states, which are inherited in a dominant manner and, at least sometimes, involve regulatory or structural proteins. For deficiency state diseases, transgenes delivered by engineered envelope vectors could be used to introduce normal genes into affected tissues for replacement therapy, as well as antisense mutations to create animal models of the disease. In the case of imbalanced disease states, transgenes delivered by engineered envelope vectors could be used to create disease states in model systems, which could then be used to combat the disease state. Thus, the methods of the present invention enable the treatment of genetic diseases. As used herein, a disease state is treated by partially or completely correcting the deficiency or imbalance that causes or makes the disease more severe. Site-specific integration of nucleic acid sequences to create mutations or correct defects can also be used, for example, using the CRISPR / Cas9 engineering system.

[0131] Hematopoietic stem cells, lymphocytes, vascular endothelial cells, respiratory epithelial cells, keratinocytes, skeletal and cardiac muscle cells, satellite muscle cells, neurons, and cancer cells have been proposed as targets for therapeutic gene transfer, either ex vivo or in vivo. See, e.g., AD Miller, Nature 357, 455-460 (1992); RC Mulligan, Science 260, 926-932 (1993). These and other cells are suitable target cells for the engineered retroviral or lentiviral vectors and methods of the present invention.

[0132] In the present disclosure and claims, the term "expression cassette" can refer to a transgene containing a regulatory sequence including an adjacent structural gene, or can refer to a transgene containing only a regulatory sequence. The transgene introduced by the lentiviral-based vector of the present invention can include at least one structural gene and / or regulatory gene, preferably of animal origin, more preferably of vertebrate or mammalian origin, or of human origin. Alternatively or additionally, the transgene can include a nucleotide sequence encoding an antisense RNA, a ribozyme, or an siRNA (inhibitory RNA).

[0133] An expression cassette can contain one or more, e.g., two, three or more, structural genes, preferably each flanked by its control sequences, e.g., a tissue-specific promoter, enhancer sequences, internal ribosome entry sites (IRES sequences) and additional selectable markers.

[0134] In one embodiment, the engineered envelope vector can be used to transduce cells of a subject without causing significant toxicity or immunogenicity in the subject, and the transgene is expressed after transduction. Once the cells of the subject are transduced, the therapeutic protein is expressed in a therapeutically acceptable amount. In some embodiments, the transduced cells are non-dividing cells, such as nerve cells, muscle cells, liver cells, skin cells, cardiac cells, lung cells, and bone marrow cells. In some embodiments, the transduced cells are dividing cells, such as T cells, B cells, hematopoietic stem cells (HSCs), macrophages, NK cells, dendritic cells, or monocytes. In some embodiments, the cells of the subject are hepatocytes.

[0135] In some embodiments, the expression of the transgene is under the control of a tissue-specific promoter and / or enhancer. For example, the promoter or other expression control sequence selectively enhances the expression of the transgene in hepatocytes. Examples of liver-specific promoters include, but are not limited to, the mouse thyretin promoter (mTTR), the endogenous human factor VIII promoter (F8), the human alpha-1-antitrypsin promoter (hAAT), the human albumin minimal promoter, and the mouse albumin promoter. In some embodiments, the mTTR promoter is used. The mTTR promoter is described in RH Costa et al., 1986, Mol. Cell. Biol. 6:4697. The F8 promoter is described in Figueiredo and Brownlee, 1995, J. Biol. Chem. 270:11828-11838. In some embodiments, the promoter is specific for neurons, muscle, liver, skin, heart, lung, bone marrow cells, stem cells, T cells, regulatory T cells, B cells, hematopoietic stem cells, macrophages, NK cells, dendritic cells, or monocytes.

[0136] To achieve therapeutic effect, one or more enhancers can be used to further enhance expression level.One or more enhancers can be provided alone or together with one or more promoter elements.Typically, expression control sequence comprises multiple enhancer elements and tissue-specific promoter.

[0137] In one embodiment of the invention, a nucleic acid construct is constructed to encode an antibody, a fragment of an antibody (scFv) or a variant thereof.

[0138] In one embodiment of the present invention, a nucleic acid construct is constructed to encode a T cell receptor (TCR). TCRs are receptors found on the surface of T cells and play a key role in recognizing antigens presented by major histocompatibility complex (MHC) molecules. Nucleic acid constructs encoding TCRs are designed to express TCR α and β chains, or γ and δ chains, depending on the specific T cell subset or therapeutic strategy. These transgenes incorporate regulatory elements, such as promoters and enhancers, to drive expression of the TCR gene in target cells. In some embodiments, expression of the TCR transgene is conditional on delivery to a specific cell type, such as a T cell. In some embodiments, TCRs are used to redirect T cells to tumor targets. In some embodiments, TCRs are used to redirect regulatory T cells to tissues or cell types, through which engineered regulatory T cells regulate tolerance of autoimmune diseases or organ transplants. In some embodiments, the TCR is directed against NY-ESO-1, MAGE-A3, MART-1, gp100, WT1, PRAME, LAGE-1, AFP, HER2, MUC1, survivin, p53, CEA, PSMA, HPV, PR1, hTERT, TRP2, or a neoantigen that arises during the progression of the patient's cancer.

[0139] In another embodiment, the nucleic acid construct comprises a chimeric antigen receptor (CAR). CAR is a synthetic receptor that combines an antigen-binding domain, typically derived from an antibody fragment, with an intracellular signaling domain derived from a T cell receptor signaling molecule. The nucleic acid construct encoding the CAR allows these chimeric receptors to be expressed on the surface of T cells, enhancing their antigen recognition and cytotoxic activity. The transgene contains the regulatory elements necessary to drive CAR expression, including promoters, enhancers, and signaling sequences. In some embodiments, the CAR is directed to CD123 (DNA construct encoded by SEQ ID NO: 16504), CD19 (DNA construct encoded by SEQ ID NO: 16506), CD20 (DNA construct encoded by SEQ ID NO: 16505), CD22, GPRC5D (DNA construct encoded by SEQ ID NOs: 16507, 16508, 16509), BCMA, CD30, CD33, EGFRvIII, HER2, GD2, mesothelin, PSMA, ROR1, CD38, CLL1, CD138, NKG2D ligand, MUC1, OX40, or CD171.

[0140] The nucleic acid constructs described herein can be customized to target specific antigens associated with specific diseases or conditions.In the case of TCR-based transgenes, antigen specificity is conferred by the α and β chains, or the γ and δ chains, which recognize specific antigen-MHC complexes.On the other hand, CAR-based transgenes can be designed to directly recognize cell surface antigens without the need for MHC presentation.

[0141] In some embodiments, the nucleic acid constructs described herein drive expression of antigen-MHC complexes in target cells. In some embodiments, the antigen-MHC complexes are expressed in T cells. In some embodiments, the antigen-MHC complexes expressed in T cells redirect the target T cells to kill other T cells expressing TCRs directed toward the anti-MHC complexes expressed by the transgene. In some embodiments, the T cells killed are T cell malignancies such as adult T-cell leukemia / lymphoma (ATLL), angioimmunoblastic T-cell lymphoma (AITL), or large granular lymphocytic leukemia (LGLL). In some embodiments, a TCR clone giving rise to adult T-cell leukemia / lymphoma (ATLL), angioimmunoblastic T-cell lymphoma (AITL), or large granular lymphocytic leukemia (LGLL) is identified in a patient, followed by identification of the cognate antigen-MHC complex.

[0142] In one embodiment, the extracellular binding domain of the CAR of the present invention, for example, the scFv portion, is encoded by a nucleic acid construct whose sequence is codon-optimized for expression in mammalian cells. In one embodiment, the entire CAR construct of the present invention is encoded by a transgene whose entire sequence is codon-optimized for expression in mammalian cells. Codon optimization refers to the discovery that the frequency of synonymous codons (i.e., codons that encode the same amino acid) in coding DNA is biased in different species. Such codon degeneracy allows the same polypeptide to be coded by various nucleotide sequences. Various codon optimization methods are known in the art, including, for example, at least, the methods disclosed in U.S. Patent Nos. 5,786,464 and 6,114,148.

[0143] The present invention encompasses lentiviral or retroviral-based gene therapy strategies for genetic disorders that aim to restore or modulate the function of mutated genes. Lentiviral or retroviral gene delivery vectors are utilized to deliver therapeutic transgenes to target cells, where the therapeutic transgenes integrate into the genome to produce functional proteins. The engineered lentiviral or retroviral-based gene therapy described herein offers innovative solutions for a wide range of genetic disorders and provides potential treatment options for affected individuals.

[0144] In one embodiment of the present invention, engineered envelope vectors are used to replace defective genes with functional copies. The engineered envelope vectors contain a transgene cassette encoding a functional version of a mutant gene under the control of specific regulatory elements, such as promoters and enhancers. The engineered envelope vectors can efficiently transduce target cells, including dividing and non-dividing cells, allowing the delivery and integration of therapeutic genes into the host genome.

[0145] In another embodiment, the engineered envelope vector corrects gene mutations by combining nuclease-based gene editing technology, such as CRISPR / Cas9, CRISPR / Cas12a, or base editors, with the engineered envelope vector. The engineered envelope vector can carry the components necessary for gene editing, including guide RNA sequences and nuclease coding sequences. Once transduced, the engineered envelope vector can induce the expression of gene editing mechanisms in target cells, allowing for precise modification of mutated genes. This approach allows for the correction of gene mutations at the genomic DNA level, potentially providing a cure for genetic disorders.

[0146] The one or more DNA endonucleases may be Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas100, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cm The DNA endonucleases may be Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, or Cpf1 endonucleases, homologs thereof, recombinant naturally occurring molecules thereof, codon-optimized versions thereof, or modified versions thereof, and combinations thereof. In some embodiments, any of the DNA endonucleases disclosed in PCT / US2021 / 065554, filed December 29, 2021, and hereby incorporated by reference, is used.

[0147] Furthermore, engineered envelope vectors can be used to modulate gene expression in genetic disorders.In certain cases, the underlying genetic defect may not involve the complete loss or mutation of a gene, but may involve dysregulated gene expression.Engineered envelope vectors can be designed to deliver transgenes containing regulatory elements such as microRNA or transcription factors to modulate the transcription or translation of specific genes.This approach aims to restore the balance of gene expression and alleviate the symptoms associated with genetic disorders.

[0148] The engineered envelope vectors described herein can be applied to a wide range of genetic disorders, including, but not limited to, cystic fibrosis, Duchenne muscular dystrophy, hemophilia, sickle cell disease, and lysosomal storage disorders. The selection of therapeutic genes and the design of engineered envelope vectors can be tailored to the specific genetic defects and requirements of the target cells.

[0149] In one embodiment of the present invention, the engineered envelope vector utilizes siRNA as an RNA inhibitor to silence the expression of disease-causing genes. The siRNA is designed to be complementary to the target RNA sequence, causing the degradation of the RNA molecule through the RNA interference (RNAi) pathway. The engineered envelope vector contains a transgene cassette encoding the siRNA sequence under the control of specific regulatory elements such as promoters and enhancers. When transduced into target cells, the engineered envelope vector delivers the siRNA, resulting in the specific degradation of the target RNA and the subsequent modulation of gene expression.

[0150] In some embodiments, engineered envelope vector-based gene therapy uses antisense oligonucleotides (ASOs) as RNA inhibitors to modulate gene expression by interfering with RNA processing, splicing or translation.ASOs are oligonucleotides designed to be complementary to specific RNA sequences, allowing them to hybridize with target RNA and modulate its function.In some embodiments, engineered envelope vectors carry transgene cassettes encoding ASO sequences, which can be designed to either promote RNA degradation or prevent the translation or processing of target RNA.Engineered envelope vectors can efficiently deliver ASOs to target cells, allowing gene expression modulation and therapeutic intervention.

[0151] Furthermore, engineered envelope vectors can deliver transgenes containing other RNA inhibitors, such as ribozymes or RNA aptamers, to achieve specific gene modulation or targeting. Ribozymes are catalytic RNA molecules that can selectively cleave target RNA molecules at specific sites, while RNA aptamers are short RNA sequences that can bind to specific target molecules with high affinity and specificity. Engineered envelope vectors carrying transgene cassettes encoding ribozymes or RNA aptamers enable precise and targeted modulation of gene expression or interactions with specific cellular components.

[0152] The engineered envelope vector utilizing the RNA inhibitor described herein can be applied to various genetic disorders, including but not limited to neurodegenerative diseases, muscular dystrophy, inherited metabolic disorders, and tumors.The selection of RNA inhibitor and the design of engineered envelope vector can be tailored to the requirements of specific target RNA molecules and target cells.

[0153] In one embodiment of the present invention, a nucleic acid construct contains a transgene including a promoter for driving the expression of a therapeutic gene in a target cell. A promoter is a DNA sequence that interacts with transcription factors and RNA polymerase to initiate transcription. The nucleic acid construct incorporates a transgene cassette containing a promoter region upstream of the therapeutic transgene of interest. The choice of promoter depends on the desired expression profile, tissue specificity, and regulatory requirements of the therapeutic gene. Examples of commonly used promoters include viral promoters (e.g., cytomegalovirus (CMV) promoter), cellular promoters (e.g., human elongation factor-1α (EF1α) promoter), and tissue-specific promoters (e.g., neuron-specific enolase (NSE) promoter). The choice of promoter can be adjusted to achieve optimal expression levels and tissue specificity in the target cells.

[0154] In another embodiment, the nucleic acid construct incorporates an enhancer to enhance promoter activity and enhance gene expression. Enhancers are DNA sequences that interact with specific transcription factors to increase the transcriptional activity of promoters. The engineered envelope vector contains a transgene containing an enhancer element within the transgene cassette, either together with or independently of the promoter. The selection and combination of enhancers can be customized to achieve the desired level and pattern of gene expression in target cells. Enhancers can be derived from viral or cellular sources, or they can be synthetic. Examples of commonly used enhancers include the cytomegalovirus immediate-early enhancer (CMV enhancer) and the SV40 enhancer. The incorporation of enhancers allows for fine-tuning of gene expression and precise control of therapeutic outcomes.

[0155] Engineered envelope vectors utilizing the promoters and enhancers described herein can be applied to a variety of diseases, including, but not limited to, genetic disorders, cancer, autoimmune diseases, neurodegenerative diseases, and cardiovascular diseases. The selection of an appropriate promoter and enhancer depends on the specific therapeutic gene, target cell, and desired expression profile. 6.1.5. Reporter Transgenes and Nucleic Acid Barcodes

[0156] In some embodiments, the nucleic acid constructs described herein comprise a reporter transgene (e.g., a gene encoding a reporter protein). In some embodiments, the nucleic acid construct encodes a reporter transgene (e.g., a reporter protein). As used herein, a reporter transgene is generally a protein or gene that can be detected when expressed in a target cell. In some embodiments, the presence or absence of the reporter in a target cell or a subset of target cells within a cell population enables the ability to sort the cells (e.g., using flow cytometry and / or fluorescence-activated cell sorting).

[0157] In some embodiments, the reporter is a fluorescent protein. The fluorescent protein may be green fluorescent protein (GFP), yellow fluorescent protein (YFP), or red fluorescent protein (RFP). The fluorescent protein may be as described in U.S. Patent No. 7,060,869, entitled "Fluorescent protein sensors for detection of analytes." In some embodiments, the fluorescent protein is used in conjunction with FACS to isolate or purify a library of cells transduced with a retroviral or lentiviral library.

[0158] In some embodiments, the reporter is an antibiotic resistance marker. In some embodiments, the antibiotic resistance marker is a protein or gene that confers a competitive advantage to target cells that contain the marker. In some embodiments, the antibiotic resistance marker comprises a hygromycin resistance protein or gene, a kanamycin resistance protein or gene, an ampicillin resistance protein or gene, a streptomycin resistance protein or gene, or a neomycin resistance protein or gene. In some embodiments, the antibiotic resistance marker is used to isolate or purify a library of cells transduced with an engineered retroviral or lentiviral library.

[0159] In some embodiments, the engineered envelope vectors described herein comprise a nucleic acid comprising a nucleic acid barcode. In some embodiments, the engineered envelope vectors described herein deliver a nucleic acid payload to a target cell, such that the target cell subsequently expresses an RNA transcript comprising the nucleic acid barcode. In some embodiments, the target cell further expresses an RNA nucleic acid barcode or comprises a DNA nucleic acid barcode, and in addition to isolating the transduced cells in a reaction vessel (e.g., a microfluidic chamber, a well plate, or an emulsion microdroplet), overlap extension RT-PCR, or overlap extension PCR and high-throughput sequencing, is used to match the barcode delivered by the engineered retrovirus or lentivirus and the cellular barcode in a high-throughput manner. 6.2.Cells

[0160] The cells described herein can be any bacterial, mammalian, or yeast cell. In some embodiments, the cells are human, mouse, rat, or non-human primate cells. In some embodiments, the cells are somatic or germ cells. In some embodiments, the cells are epithelial cells, neural cells, hormone-secreting cells, immune cells, secretory cells, blood cells, stromal cells, or germ cells. In some embodiments, the cells are antigen-specific cells (e.g., cells that bind to a specific antigen). In some embodiments, the antigen-specific cells are immune cells. In some embodiments, the antigen-specific cells are B cells or T cells. In some embodiments, the cells are target cells (e.g., containing a cognate protein or ligand that can be targeted by the engineered envelope vectors described herein).

[0161] The population of cells described herein can be any bacterial, mammalian, or yeast cell population. In some embodiments, the population of cells is a population of human, mouse, rat, or non-human primate cells. In some embodiments, the population of cells is a somatic cell population or a germ cell population. In some embodiments, the population of cells comprises epithelial cells, neural cells, hormone-secreting cells, immune cells, secretory cells, blood cells, stromal cells, and / or germ cells. In some embodiments, the population of cells comprises antigen-specific cells (e.g., cells that bind to a specific antigen). In some embodiments, the population of antigen-specific cells comprises immune cells. In some embodiments, the population of antigen-specific cells comprises B cells and / or T cells. In some embodiments, the population of cells comprises a homogeneous population of cells. In some embodiments, the population of cells comprises a heterogeneous population of cells.

[0162] In some embodiments, the population of cells is a population of cells isolated from a subject. The subject can be a human subject (e.g., a human subject suffering from a disease), a mouse subject, a rat subject, or a non-human primate subject. In some embodiments, the population of cells is isolated from the subject's blood or tumor. In some embodiments, the subject comprises the population of cells. In some embodiments, the population of cells remains within the subject. In some embodiments, the population of cells is removed from the subject.

[0163] In some embodiments, the population of cells has been previously frozen and thawed (e.g., 1, 2, 3, 4, 5 or more freeze / thaw cycles). In some embodiments, the population of cells is maintained in liquid medium. In some embodiments, the population of cells has been passaged 1, 2, 3, 4, 5 or more times using any known method. In some embodiments, the population of cells is maintained in liquid medium before combining with the engineered envelope vector. In some embodiments, the population of cells is maintained in liquid medium after combining with the engineered envelope vector or multiple engineered envelope vectors. In some embodiments, the population of cells is maintained in liquid medium before combining with the engineered envelope vector or multiple engineered envelope vectors.

[0164] In some embodiments, the population of cells comprises any of the engineered envelope vectors described herein (e.g., an engineered retrovirus or lentivirus). In some embodiments, a subset of the population of cells contains any of the engineered envelope vectors described herein (e.g., an engineered retrovirus or lentivirus). In some embodiments, a subset of the population of cells contains an engineered envelope vector within each cell of the subset (e.g., within the nucleus of each cell of the subset). In some embodiments, the population of cells or a subset thereof expresses a reporter (e.g., a fluorescent protein or an antibiotic resistance marker). In some embodiments, the population of cells or a subset thereof (e.g., containing an engineered envelope vector) is isolated and / or sorted based on the presence or absence of the reporter. In some embodiments, a subset of the population of cells containing an engineered envelope vector described herein is isolated and / or sorted from cells of the population that do not contain the engineered envelope vector based on the presence or absence of the reporter. In some embodiments, at least 50%, 60%, 70%, 80%, 90%, or 95% of the population of cells before cell sorting contain the engineered envelope vector, hi some embodiments, at least 70%, 80%, 90%, 95%, or 100% of the population of cells contain the engineered envelope vector after isolation and / or sorting based on the presence or absence of the reporter.

[0165] In some embodiments, the population of cells comprises packaging cells used to produce engineered envelope vectors (lentiviruses or other retroviruses). In some embodiments, the packaging cells are a population of cells that are substantially equivalent to one another, i.e., largely clonal. In some embodiments, the packaging cells comprise a library of cells that produce a library of distinct engineered envelope vectors (e.g., lentiviruses or other retroviruses). In some embodiments, the packaging cells comprise a library of cells encoding a library of nucleic acid barcodes. In some embodiments, the transgene comprises a nucleic acid barcode. In some embodiments, nucleic acid barcodes are not used; instead, envelope sequences or scFv sequences or other surface receptors are delivered as transgenes to target cells and directly sequenced. In some embodiments, the library of cells is engineered to express nucleic acid barcodes using a library of engineered envelope vectors (engineered lentiviruses or retroviruses). In some embodiments, the library of cells is engineered to express nucleic acid barcodes using a library of plasmids. In some embodiments, the nucleic acid barcodes in the library of cells are measured or characterized using high-throughput DNA sequencing. In some embodiments, the library of cells expresses a library of recombinant proteins. In some embodiments, the nucleic acid barcodes comprise the same RNA transcripts as those expressing the recombinant proteins. In some embodiments, the nucleic acid barcodes comprise transcripts that are different from those expressing the recombinant proteins. In some embodiments, the library of recombinant proteins expressed by the packaging cells comprises a library of non-viral membrane-bound proteins. In some embodiments, the library of recombinant proteins expressed by the packaging cells comprises a library of membrane fusogenic factors. In some embodiments, the library of recombinant proteins expressed by the packaging cells comprises a library of function modulator proteins.In some embodiments, the lentiviral transgene comprises the recombinant protein of the library.

[0166] In some embodiments, the population of cells comprises target cells of an engineered envelope vector (an engineered lentivirus or population of lentiviruses, or an engineered retrovirus or population of retroviruses). In some embodiments, the target cells are a population of cells that are substantially equivalent to one another, i.e., largely clonal. In some embodiments, the target cells comprise a library of cells that produce a library of distinct recombinant proteins. In some embodiments, the target cells comprise a library of cells that encode a library of nucleic acid barcodes. In some embodiments, the library of target cells is engineered to express nucleic acid barcodes using a lentivirus or retrovirus library. In some embodiments, the library of cells is engineered to express nucleic acid barcodes using a library of CRISPR / Cas9 guide RNAs. In some embodiments, the nucleic acid barcodes in the library of cells are measured or characterized using high-throughput DNA sequencing. In some embodiments, the library of cells comprises a library of recombinant proteins. In some embodiments, the library of cells comprises a library of immortalized primary cells. In some embodiments, the library of cells comprises a library of immortalized primary cells, wherein each cell type or clone in the library expresses or comprises a unique nucleic acid barcode. In some embodiments, the nucleic acid barcodes comprise the same RNA transcript as the RNA transcript expressing the recombinant protein. In some embodiments, the nucleic acid barcodes comprise a transcript that is different from the RNA transcript expressing the recombinant protein. In some embodiments, the library of target cells expresses a sequence comprising one or more of SEQ ID NOs: 168-5339 or 5340-8121, or a fragment or fusion protein of one or more of SEQ ID NOs: 168-5339 or 5340-8121. In some embodiments, the library of target cells comprises 10, 100, 1,000, 10,000, 100,000, 1,000,0000, 10,000,000, or more than 10,000,000 unique nucleic acid barcodes.In some embodiments, the library of cells comprises a mutation library of one or more proteins. In some embodiments, the library of cells comprises an alanine scanning mutagenesis library of one or more proteins. In some embodiments, the library of cells comprises a library of one or more proteins generated using error-prone PCR. 6.3. Polynucleotide Constructs Encoding Proteins of Engineered Envelope Vectors

[0167] In another aspect, the present disclosure provides a polynucleotide construct encoding one or more proteins of the engineered envelope vectors provided herein. In some embodiments, the polynucleotide construct encodes a viral envelope protein (membrane fusogenic factor). In some embodiments, the polynucleotide construct encodes a non-viral membrane-associated protein for targeting and tropism. In some embodiments, the polynucleotide construct encodes a function modulator protein. In some embodiments, the polynucleotide construct comprises a transgene sequence. In some embodiments, the transgene construct comprises one or more long terminal repeat (LTR) sequences. In some embodiments, the LTR sequence is wild-type, and in some embodiments, the LTR sequence is chimeric or mutant. In some embodiments, the polynucleotide construct is used to generate an engineered envelope vector.

[0168] In some embodiments, the polynucleotide construct further comprises a regulatory sequence. In some embodiments, the regulatory sequence is a tissue-specific or cell-type-specific regulatory element such as an enhancer or promoter. In some embodiments, the regulatory element is a synthetic sequence. In some embodiments, the regulatory element is an endogenous, naturally occurring sequence.

[0169] In some embodiments, the polynucleotide construct is a plasmid. In some embodiments, the polynucleotide construct is a non-plasmid vector. In some embodiments, the polynucleotide construct is a viral vector. In some embodiments, the polynucleotide construct is inserted into the genome of a cell.

[0170] In some embodiments, the plasmid encodes two or more components selected from a viral envelope protein (fusogenic factor), a non-viral membrane-bound protein for targeting, and a function modulator protein. When the plasmid comprises multiple components, the plasmid may contain one or more promoters. In some embodiments, the plasmid contains an IRES (internal ribosome entry site). In some embodiments, the plasmid encodes a viral envelope protein (fusogenic factor), a non-viral membrane-bound protein for targeting and targeting, or a function modulator protein. 6.4. Methods for generating engineered envelope vectors

[0171] Another aspect of the present disclosure provides a method for producing the engineered envelope vector disclosed herein.In some embodiments, the method comprises the steps of delivering one or more polynucleotide constructs (e.g., a polynucleotide construct encoding a viral envelope protein (membrane fusion factor), a non-viral membrane-associated protein, a functional modulator protein, and / or a transgene) disclosed above into packaging cells, culturing the packaging cells, and harvesting the engineered envelope vector.In some embodiments, the method further comprises the step of enriching the fully engineered envelope vector containing the nucleic acid construct.

[0172] In some embodiments, one or more of the polynucleotide constructs are delivered to the packaging cell as one or more plasmids. In some embodiments, one or more of the polynucleotide constructs are integrated into the genome of the packaging cell. The second-generation lentiviral packaging plasmid encodes the Gag, Pol, Pro, Rev, and Tat genes from a single plasmid. The plasmid psPAX2 (SEQ ID NO: 16498) is an example of a second-generation lentiviral packaging plasmid. The second-generation lentiviral transfer plasmid expresses viral RNA from the 5'LTR, which is Tat-dependent. The pLOC-TurboRFP is an example of a second-generation transfer plasmid (SEQ ID NO: 16499). Third-generation lentiviral packaging plasmids split the packaging system into two plasmids, one encoding Rev (e.g., pRSV-Rev, SEQ ID NO: 16500) and the second encoding Gag, Pro, and Pol (e.g., pMDLg / pRRE, SEQ ID NO: 16501, and pCgpV, SEQ ID NO: 16502). The Tat gene is removed from third-generation packaging systems. Therefore, an exogenous promoter, such as CMV or RSV, must be used to drive viral RNA expression from the transfer vector. pReceiver-EF1a-GFP (SEQ ID NO: 16503) is an example of a third-generation lentiviral transfer plasmid. Fourth-generation lentiviral packaging plasmids further split the packaging system, with one plasmid encoding Gag and Pro, a second plasmid encoding Pol, and a third plasmid encoding Tat and Rev. Fourth-generation lentiviral packaging systems typically use tetracycline (Tet)-Off expression plasmids to drive the expression of Gag-Pro and Tat-Rev expression plasmids containing Tat transactivator sequences. In this system, the expression of Gag, Pro, Tat, and Rev requires the expression of Tet-Off in tetracycline-free medium. Fourth-generation lentiviral packaging systems can package either second- or third-generation transfer plasmids.

[0173] In some embodiments, packaging cells are human and animal (eg, pig, cow, dog, horse, donkey, mouse, hamster, monkey) cells. 6.5. Libraries of engineered envelope vectors

[0174] Described herein is a library of engineered envelope vectors. In some embodiments, the library comprises a plurality of unique engineered envelope vectors, each unique vector comprising a nucleic acid encoding a viral fusogenic protein, a non-viral membrane-associated protein, and a reporter or nucleic acid barcode, and each unique vector comprising a different unique extracellular targeting domain. Also described herein is a library of engineered envelope vectors, each unique engineered envelope vector comprising a nucleic acid encoding a viral fusogenic protein, a non-viral membrane-associated protein, a functional modulator protein, and a reporter or nucleic acid barcode, and each unique vector comprising a different unique extracellular targeting domain. Also described herein is a library of engineered envelope vectors, each unique vector comprising a nucleic acid encoding a viral fusogenic protein, a functional modulator protein, and a reporter or nucleic acid barcode, and each unique vector comprising a different unique extracellular targeting domain. Also described herein is a library of cells comprising an engineered envelope vector, the library comprising a plurality of unique cells, each unique cell comprising a unique engineered envelope vector.

[0175] In some embodiments, the library of engineered envelope vectors comprises a library of transfer vectors that further comprises a library of nucleic acids encoding viral envelope proteins. In some embodiments, the library of engineered envelope vectors comprises a library of transfer vectors that further comprises a library of nucleic acids encoding non-viral membrane-bound proteins. In some embodiments, the library of engineered envelope vectors comprises a library of transfer vectors that further comprises a library of nucleic acids encoding genome editing proteins and / or nucleic acids required for genome editing, such as a library of CRISPR / Cas proteins and / or cognate guide RNAs for genome editing.

[0176] The one or more DNA endonucleases may be Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas100, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cm r4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4 or Cpf1 endonuclease, homologs thereof, recombinant forms of these naturally occurring molecules, codon-optimized versions thereof, or modified versions thereof, and combinations thereof.

[0177] In some embodiments, the library of engineered envelope vectors comprises a library of packaging cells further comprising a library of nucleic acids encoding viral envelope proteins. In some embodiments, the library of engineered envelope vectors comprises a library of packaging cells further comprising a library of nucleic acids encoding non-viral membrane-bound proteins. In some embodiments, the library of engineered envelope vectors comprises a library of packaging cells further comprising a library of nucleic acids encoding genome-editing proteins and / or nucleic acids required for genome editing, e.g., a library of CRISPR / Cas proteins and / or cognate guide RNAs for genome editing. In some embodiments, the non-viral membrane-bound proteins, viral envelope proteins, and / or genome-editing proteins are stably engineered into multiple genomes of the cells.

[0178] In some embodiments, the library comprises a pMHC-encoded (peptide / MHC-encoded) retroviral (e.g., lentiviral) library for use in screening T cell populations. In such libraries, pMHC displayed on the viral surface allows for T cell infection in a TCR-specific manner. Infected T cells can be collected and sequenced, allowing for the identification of pMHC ligands that can infect subsets of T cell populations of interest and the ability to simultaneously track TCR sequences and reactive pMHC ligands. In some embodiments, a pMHC retroviral library minimally comprises a randomized transfer vector containing randomized pMHC targeting elements. In some embodiments, randomly derived libraries are generated using degenerate oligonucleotide primers. In some embodiments, targeted libraries specific to a unique set of antigens (e.g., all possible viral or bacterial antigens for a particular target of interest, such as human immunodeficiency virus, tuberculosis (TB), or all possible neoantigens for a particular subject) are generated.

[0179] In some embodiments, the library may be screened against a population of antigen-specific cells (e.g., B cells or T cells). In some embodiments, the library contains at least 10 2 Pieces, at least 10 3 Pieces, at least 10 4 Pieces, at least 10 5 Pieces, at least 10 6 Pieces, at least 10 7 Pieces, at least 10 8 Pieces, at least 10 9 pieces, or at least 10 10 In some embodiments, the library of unique engineered envelope vectors comprises an extracellular targeting domain that is at least 5, at least 10, at least 15, at least 20, or at least 50 amino acids in length. In some embodiments, each different unique extracellular targeting domain is generated by site-directed mutagenesis.

[0180] Retroviral, lentiviral or cellular libraries can vary in size from hundreds to hundreds of thousands, millions or more of unique retroviruses, lentiviruses or unique cells. In some embodiments, the libraries of the present disclosure contain at least 500,000 unique engineered envelope vectors (retroviruses or lentiviruses) or unique cells. The libraries of the present invention include retroviral libraries and cellular libraries. A library is a synthetic (i.e., isolated, synthetically produced, free of components naturally found together in cells, purified before being placed in the library) collection of members that share a common element and at least one different element. Libraries contain 1000 or more members (e.g., at least 1,000, 2,000, 3,000, 4,000, 5,000, 10,000, 50,000, 100,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 2,000,000, 3,000,000, 4,000,000 or more). The upper limit of library size is defined by the combination of domains or modules that provide the difference or diversity between members. For example, the upper limit may be 4,000,000 members. Thus, in some embodiments, engineered libraries are highly diverse and contain at least 500,000 different members. Highly diverse libraries are 6 In some embodiments, the library of engineered envelope vectors is generated using site-directed mutagenesis of nucleic acid as described herein. In some embodiments, site-directed mutagenesis comprises the use of primers and low-fidelity RNA polymerase to allow randomized mutagenesis of common nucleic acid as described herein.

[0181] In some embodiments, the population of cells comprises packaging cells used to produce engineered envelope vectors. In some embodiments, the packaging cells are a population of cells that are substantially equivalent to one another, i.e., largely clonal. In some embodiments, the packaging cells comprise a library of cells that produce a library of distinct engineered envelope vectors (lentiviruses or other retroviruses). In some embodiments, the packaging cells comprise a library of cells encoding a library of nucleic acid barcodes. In some embodiments, the library of cells is engineered to express nucleic acid barcodes using a library of engineered envelope vectors (engineered envelope vectors). In some embodiments, the library of cells is engineered to express nucleic acid barcodes using a library of plasmids (e.g., transfer plasmids encoding lentiviral transgenes). In some embodiments, the nucleic acid barcodes in the library of cells are measured or characterized using high-throughput DNA sequencing. In some embodiments, no barcodes are used. In some embodiments, the library of cells expresses a library of recombinant proteins. In some embodiments, the nucleic acid barcodes comprise the same RNA transcripts as those expressing the recombinant proteins. In some embodiments, the nucleic acid barcode comprises a transcript that is distinct from the RNA transcript that expresses the recombinant protein. In some embodiments, the library of recombinant proteins expressed by the packaging cells comprises a library of non-viral membrane-bound proteins (e.g., target cell tropism proteins). In some embodiments, the library of recombinant proteins expressed by the packaging cells comprises a library of membrane fusogenic factors (e.g., viral envelope proteins). In some embodiments, the library of envelope proteins comprises a transgene for transduction of target cells.In some embodiments, transduction of a library of envelope proteins in vivo or in vitro, followed by sequencing of the transgenes delivered to target cells, is used to screen and identify useful envelope proteins. In some embodiments, the library of recombinant proteins expressed by packaging cells comprises a library of function modulator proteins. In some embodiments, the library of recombinant proteins expressed by packaging cells comprises TCR, CAR, or pMHC lentiviral transgenes. In some embodiments, in vivo or in vitro transduction with a library of TCR, CAR, or pMHC lentiviral transgenes, followed by sequencing and / or functional analysis (i.e., antigen activation, followed by flow cytometry of activation markers) is used to identify TCRs, CARs, or pMHCs with high therapeutic potential.

[0182] In some embodiments, the library of engineered envelope vectors comprises a library of antibodies, scFvs, or other antibody fragments. In some embodiments, the library of antibodies, scFvs, or other antibody fragments is generated by first immunizing at least one humanized or wild-type mouse, chicken, rat, monkey, or human subject with an immunogen (e.g., a protein, glycoprotein, peptide, cell, lysed cell, etc.), amplifying nucleic acids encoding antibodies, scFvs, or other antibody fragments from B cells, plasmablasts, plasma cells, or other antibody-producing cells from the subject, and then generating a library of engineered envelope vectors from the amplified nucleic acids. In some embodiments, the amplified nucleic acid library comprising antibodies, scFvs, or other antibody fragments from a mammalian subject is generated by a step comprising isolating a plurality of single cells from the mammalian subject into emulsion microdroplets to obtain pairs of linked heavy and light chain sequences at the single-cell level. In some embodiments, the amplified nucleic acid library comprising antibodies, scFvs, or other antibody fragments from a mammalian subject is generated by including yeast, mammalian, bacterial, or phage scFv, full-length antibody, or Fab display to identify binding agents of interest, and then rearranging the library of binding agents into a lentiviral or other retroviral library. In some embodiments, the yeast, mammalian, bacterial, or phage scFv, full-length antibody, or Fab display library is generated using single-cell pairing between intact heavy and light chain sequences; in other embodiments, the pairing between heavy and light chain sequences is random.

[0183] In some embodiments, the library of engineered envelope vectors (retroviral or lentiviral) comprises a library of TCRs. In some embodiments, the library of TCRs is generated by first immunizing at least one humanized or wild-type mouse, chicken, rat, monkey, or human subject with an immunogen (e.g., a protein, glycoprotein, peptide, cell, lysed cell, etc.), amplifying nucleic acids encoding TCRs or TCR fragments from T cells from the subject, and then generating an engineered retroviral or lentiviral library from the amplified nucleic acids. In some embodiments, the amplified nucleic acid library comprising TCRs from a mammalian subject is generated by a step comprising isolating a plurality of single cells from the mammalian subject into emulsion microdroplets. In some embodiments, the amplified nucleic acid library TCRs or TCR fragments from a mammalian subject is generated by including yeast, mammalian, bacterial, or phage-displayed TCRs to identify binders of interest, and then reassembling the library of binders into a lentiviral or other retroviral library. In some embodiments, the TCR library is generated using single-cell pairing between intact alpha and beta chain TCR sequences, while in other embodiments, the pairing between alpha and beta chain TCR sequences is random.

[0184] In some embodiments, the library of engineered envelope vectors is derived from a library of packaging cells, and an envelope protein is engineered into the genomes of the packaging cells. In some embodiments, the genomes of the packaging cells comprise a single envelope protein. In some embodiments, the library of packaging cells is used to generate the library of engineered envelope vectors by transfecting the library of packaging cells with a packaging plasmid that directs secretion of the engineered envelope vector from the packaging cells. In some embodiments, this library of engineered envelope vectors is used to transduce target cells. In some embodiments, the transduced target cells are then sequenced to assess which envelope protein was associated with successful transduction. In some embodiments, the transduction is performed in vitro. In some embodiments, the transduction is performed in vivo, i.e., by injecting or injecting the library into mice, rats, or dogs.

[0185] In some embodiments, the library of engineered envelope vectors is derived from a library of packaging cells, and a non-viral target cell tropism protein (e.g., an antibody or antibody fragment, or scFv) is engineered into the multiple genomes of the packaging cells. In some embodiments, the multiple genomes of the packaging cells contain a single non-viral target cell tropism protein. In some embodiments, the library of packaging cells is used to generate the library of engineered envelope vectors by transfecting the library of packaging cells with a packaging plasmid that directs secretion of the engineered envelope vector from the packaging cells. In some embodiments, this library of engineered envelope vectors is used to transduce target cells. In some embodiments, the transduced target cells are then sequenced to evaluate which non-viral target cell tropism protein was associated with successful transduction. In some embodiments, the transduction is performed in vitro. In some embodiments, the transduction is performed in vivo, i.e., by injecting or injecting the library into mice, rats, or dogs. 6.6. Screening Methods

[0186] The present specification describes a screening method using the engineered envelope vector disclosed herein.In some embodiments, the screening method includes: (i) providing an engineered envelope vector comprising a viral envelope fusogenic protein, a target cell tropism protein, and a nucleic acid encoding a reporter; (ii) combining the engineered envelope vector with a population of cells; and (iii) sorting the population of cells based on the presence or absence of the reporter.The method can further include identifying cells or cell types that can be targeted by the engineered envelope vector.In some embodiments, the method includes identifying target cell tropism proteins that can be used to target specific cells or cell types.

[0187] Also described herein is a method for screening a population of cells, the method comprising: (i) providing an engineered envelope vector comprising a viral envelope fusogenic protein, a target cell tropism protein, a function modulator protein, and a nucleic acid encoding a reporter; (ii) combining the retrovirus with a population of cells; and (iii) sorting the population of cells based on the presence or absence of the reporter. The method may further comprise identifying a cell or cell type that is responsive to the function modulator protein. In some embodiments, the method comprises identifying a function modulator protein that can modulate a specific cell or cell type. In some embodiments, the method further comprises identifying a cell or cell type that can be targeted by the engineered envelope vector. In some embodiments, the method comprises identifying a target cell tropism protein that can be used to target a specific cell or cell type.

[0188] Any of the engineered envelope vectors disclosed herein can be used for this method. In some embodiments, the engineered envelope vector (i) comprises a nucleic acid having the structure: S-ETD-MBD and a viral envelope fusogenic protein, where S encodes a signal sequence, ETD encodes an extracellular targeting domain, and MBD encodes a membrane-binding domain. In some embodiments, the engineered envelope vector (i) comprises a nucleic acid having the structure: S-ETD-MBD-IRES-R and a viral envelope fusogenic protein, where S encodes a signal sequence, ETD encodes an extracellular targeting domain, MBD encodes a membrane-binding domain, IRES encodes an internal ribosome entry site, and R encodes a reporter.

[0189] As used herein, the term "combining" (which in some embodiments is synonymous with the terms "providing" and "contacting") generally refers to the act of bringing an engineered envelope vector into intimate physical contact with a population of cells such that the extracellular targeting domain of the vector is capable of binding to a cognate ligand present on a subset of cells in the population. In some embodiments, the combination of the engineered envelope vector with the population of cells occurs when a solution comprising the engineered envelope vector is mixed with a solution comprising the population of cells. In some embodiments, the combination of the engineered envelope vector with the population of cells occurs when a lyophilized engineered envelope vector is mixed with a solution comprising the population of cells. In some embodiments, the combination of the engineered envelope vector with the population of cells occurs when a lyophilized engineered envelope vector is mixed with a lyophilized population of cells and reconstituted in solution. In some embodiments, the cells of the population are maintained in a monolayer of cells in cell culture medium and / or attached to a tissue culture plate or Petri dish.

[0190] Generally, the engineered envelope vector and the population of cells are combined (e.g., physically combined or contacted) for a defined period of time. In some embodiments, the period is measured in seconds, minutes, hours, or days. In some embodiments, the period is 0-30 seconds, 15-45 seconds, 30-60 seconds, 45-90 seconds, 60-90 seconds, or 60-120 seconds. In some embodiments, the retrovirus and the population of cells are combined and contacted for 0-30 seconds, 15-45 seconds, 30-60 seconds, 45-90 seconds, 60-90 seconds, or 60-120 seconds. In some embodiments, the time period is 1-2 minutes, 1-5 minutes, 1-10 minutes, 2-10 minutes, 5-10 minutes, 5-20 minutes, 10-20 minutes, 25-30 minutes, 25-60 minutes, 30-45 minutes, 30-40 minutes, 40-60 minutes, 50-70 minutes, or 60-120 minutes. In some embodiments, the engineered envelope vector and the population of cells are combined and contacted for 1-2 minutes, 1-5 minutes, 1-10 minutes, 2-10 minutes, 5-10 minutes, 5-20 minutes, 10-20 minutes, 25-30 minutes, 25-60 minutes, 30-45 minutes, 30-40 minutes, 40-60 minutes, 50-70 minutes, or 60-120 minutes. In some embodiments, the period is at least 1 minute, at least 2 minutes, at least 5 minutes, at least 10 minutes, at least 20 minutes, at least 30 minutes, at least 60 minutes, at least 45 minutes, at least 40 minutes, at least 70 minutes, or at least 120 minutes. In some embodiments, the engineered envelope vector and the population of cells are combined and contacted for 1-2 minutes, 1-5 minutes, 1-10 minutes, 2-10 minutes, 5-10 minutes, 5-20 minutes, 10-20 minutes, 25-30 minutes, 25-60 minutes, 30-45 minutes, 30-40 minutes, 40-60 minutes, 50-70 minutes, or 60-120 minutes.

[0191] In some embodiments, the period is 1 to 2 hours, 1 to 5 hours, 1 to 3 hours, 2 to 5 hours, 3 to 6 hours, 3 to 12 hours, 6 to 12 hours, 12 to 18 hours, 12 to 24 hours, 15 to 30 hours, 18 to 24 hours, 24 to 48 hours, 24 to 36 hours, or 36 to 50 hours. In some embodiments, the period is at least 1 hour, at least 2 hours, at least 5 hours, at least 3 hours, at least 6 hours, at least 12 hours, at least 18 hours, at least 24 hours, at least 15 hours, at least 30 hours, at least 48 hours, at least 36 hours, or at least 50 hours. In some embodiments, the engineered envelope vector and the population of cells are combined and contacted for 1-2 hours, 1-5 hours, 1-3 hours, 2-5 hours, 3-6 hours, 3-12 hours, 6-12 hours, 12-18 hours, 12-24 hours, 15-30 hours, 18-24 hours, 24-48 hours, 24-36 hours, or 36-50 hours. In some embodiments, the period is 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 8 days, 9 days, 10 days, or 5-15 days. In some embodiments, the engineered envelope vector and the population of cells are combined and contacted for 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 8 days, 9 days, 10 days, or 5-15 days.

[0192] In some embodiments, cell populations are selected based on the presence or absence of reporters.In some embodiments, a subset of cell populations that contain reporters (e.g., express reporters) is selected from the remaining subset of cell populations that do not contain reporters.In some embodiments, cell populations are selected using flow cytometry (e.g., fluorescence-activated cell sorting), next-generation DNA sequencing (e.g., single-cell next-generation sequencing), or antibiotic selection.

[0193] In some embodiments, the conditions of step (ii) that allow the engineered envelope vector to interact with a subset of the cell population include combining the engineered envelope vector with the cell population in the presence of a defined solution, composition, and at a specific temperature. In some embodiments, the engineered envelope vector and the cell population are combined in the presence of cell culture medium (e.g., RPMI or DMEM cell culture medium). In some embodiments, the lentivirus or retrovirus and the cell population are combined in the presence of buffered saline. In some embodiments, the buffered saline is phosphate-buffered saline or HEPES-buffered saline. In some embodiments, the buffered saline contains bovine serum albumin and / or EDTA. In some embodiments, the retrovirus and the cell population are combined in the presence of an enhancer of retroviral or lentiviral transduction (e.g., heparin sulfate, polybrene, protamine sulfate, or dextran). In some embodiments, the engineered envelope vector and the population of cells are combined in (ii) at a temperature ranging from 4°C to 42°C, 4°C to 8°C, 4°C to 10°C, 8°C to 15°C, 10°C to 20°C, 18°C ​​to 23°C, 20°C to 30°C, 25°C to 35°C, 30°C to 40°C, or 37°C to 42°C.

[0194] In some embodiments, the screening method described herein further comprises washing the cell population with a wash solution between step (ii) and step (iii). In some embodiments, the wash solution is any liquid solution that allows for the maintenance of healthy cells (e.g., a solution having a neutral pH and a low to moderate ionic strength). In some embodiments, washing the cell population removes excess and / or remaining engineered envelope vectors from the cell population. In some embodiments, the cell population is washed using cell culture medium (e.g., RPMI or DMEM cell culture medium). In some embodiments, the cell population is washed using buffered saline. In some embodiments, the buffered saline is phosphate-buffered saline or HEPES-buffered saline. In some embodiments, the buffered saline comprises bovine serum albumin and / or EDTA. In some embodiments, the population of cells is washed at a temperature ranging from 4°C to 42°C, 4°C to 8°C, 4°C to 10°C, 8°C to 15°C, 10°C to 20°C, 18°C ​​to 23°C, 20°C to 30°C, 25°C to 35°C, 30°C to 40°C, or 37°C to 42°C.

[0195] In some embodiments, the population of cells is maintained in liquid medium before being combined with the engineered envelope vector. In some embodiments, the population of cells is maintained in liquid culture after being combined with the engineered envelope vector. In some embodiments, the population of cells is maintained in liquid medium during the step of combining with the engineered envelope vector. In some embodiments, the population of cells is attached to a cell culture plate or Petri dish. In some embodiments, the population of cells is maintained as a monolayer, embryoid body, or any cell aggregate.

[0196] In some embodiments, the screening method comprises the use of a plurality of engineered envelope vectors. In certain embodiments, the plurality of engineered envelope vectors comprises at least 10 2 pieces, 10 3 pieces, 10 4 pieces, 10 5 pieces, 10 6pieces, 10 7 pieces, 10 8 pieces, 10 9 pieces, 10 10 pieces, 10 11 pieces or 10 12 In some embodiments, at least 10 of each unique retrovirus comprises at least 10 unique engineered envelope vectors. 2 pieces, 10 3 pieces, 10 4 pieces, 10 5 pieces, 10 6 pieces, 10 7 pieces, 10 8 pieces, 10 9 pieces, 10 10 pieces, 10 11 pieces or 10 12 This copy may be present in multiple retroviruses.

[0197] In some embodiments, the screening method comprises screening a cell population using at least two different, unique engineered envelope vectors. In some embodiments, the different, unique engineered envelope vectors comprise different extracellular targeting domains and / or different reporters. In some embodiments, the screening method comprises a first retrovirus (or lentivirus) and a second retrovirus (or lentivirus), wherein the first and second retrovirus (or lentivirus) comprise different extracellular targeting domains and / or different reporters. In some embodiments, the screening method comprises screening a cell population using 2, 3, 4, 5, 6, 7, 8, 9, 10, 50, 100 or more different engineered envelope vectors. In some embodiments, the screening method comprises screening a cell population using a library of engineered envelope vectors. In some embodiments, the library of engineered envelope vectors comprises at least 10 2 Pieces, at least 10 3 Pieces, at least 10 4 Pieces, at least 10 5 Pieces, at least 10 6Pieces, at least 10 7 Pieces, at least 10 8 Pieces, at least 10 9 pieces, or at least 10 10 Each vector contains a unique retrovirus or lentivirus. 6.7. Methods for Screening Libraries of Engineered Envelope Vectors and Cells

[0198] Specific quantitative genetic analyses of living tissues and organisms are best performed at the single-cell level. However, single cells contain only picograms of genetic material. Traditional methods (e.g., polymerase chain reaction (PCR), RNA sequencing (Mortazavi et al., 2008 Nature Methods 5:621-8), chromatin immunoprecipitation sequencing (Johnson et al., 2007 Science 316:1497-502), or whole-genome sequencing (Lander et al., 2001 Nature 409:860-921) require more genetic material than can be found in a single cell and are typically performed using thousands to millions of cells. While these methods provide useful genetic information at the cell population level, they have significant limitations for elucidating biological phenomena at the single-cell level.

[0199] Single cells are used as reaction compartments to perform various genetic analyses (Embleton et al., 1992 Nucleic Acids Research 20:3831-37; Hviid, 2002 Clinical Chemistry 48:2115-2123; U.S. Patent No. 5,830,663). Single cells are sorted into water-in-oil microdroplet emulsions, and molecular analyses are performed in the microdroplets (Johnston et al., 1996 Science 271:624-626; Brouzes et al., 2009 PNAS 106:14195-200; Kliss et al., 2008 Anal Chem 80:8975-81; Zeng et al., 2010 Anal Chem 82:3183-90). These single-cell assays are limited to single-cell PCR in emulsion or in situ PCR on fixed and permeabilized single cells. Furthermore, when analyzing large populations of cells, it is difficult to trace each gene product back to a single cell or subpopulation of cells.

[0200] Massively parallel methods for analyzing nucleic acids in single cells are disclosed in US45960010P and related patents by Johnson, which are hereby incorporated by reference in their entirety. These methods provide protocols for performing bulk sequencing reactions to generate sequence information for at least 100,000 fusion complexes from at least 10,000 cells in a population of cells, and this sequence information is sufficient to co-localize a first target nucleic acid sequence and a second target nucleic acid sequence in a single cell from a population of at least 10,000 cells. The screening method or any other method using the engineered envelope vector disclosed herein can be performed using a single-cell analysis system.

[0201] In some embodiments, target cells transduced with engineered envelope vectors of the present invention are subjected to single-cell genetic analysis (e.g., Figures 5-10). In one aspect, single cells are isolated in emulsion microdroplets. In another aspect, single cells are isolated in a reaction vessel. In some embodiments, single-cell genetic analysis is used to obtain single-cell coexpression or colocalization between nucleic acid barcodes delivered by retroviruses and barcodes comprising target cells prior to transduction. In some embodiments, single-cell genetic analysis is used to obtain single-cell coexpression or colocalization between nucleic acid barcodes delivered by engineered vectors and any nucleic acid comprising the target cell prior to transduction, including the entire transcriptome or 10, 100, 1,000, 10,000, 100,000, or 1,000,000 endogenous nucleic acid targets or RNA transcripts. In some embodiments, single-cell genetic analysis is used to obtain single-cell co-expression or co-localization between nucleic acids encoding reporters delivered by engineered envelope vectors and any nucleic acid, including the entire transcriptome or 10, 100, 1,000, 10,000, 100,000, or 1,000,000 endogenous nucleic acid targets or RNA transcripts, of target cells before transduction. In some embodiments, single-cell genetic analysis is used to obtain single-cell co-expression or co-localization between nucleic acids encoding protein reporters delivered by retroviruses (or lentiviruses) and barcodes, including target cells before transduction. In some embodiments, massively parallel single-cell analysis is performed to profile 100, 1,000, 10,000, 100,000, 1,000,000, or 10,000,000 or more single cells in parallel.In some embodiments, massively parallel single-cell analysis is performed to profile 100, 1,000, 10,000, 100,000, 1,000,000, or 10,000,000 or more single cells that have been transduced in parallel with 100, 1,000, 10,000, 100,000, 1,000,000, or 10,000,000 or more retrovirus (or lentivirus) type libraries. In some embodiments, multiple single nucleic acid barcodes are isolated with single cells, and the single nucleic acid barcodes are used to uniquely identify nucleic acids derived from single cells. In some embodiments, beads are used to deliver the nucleic acid barcodes to the isolated single cells.

[0202] In one embodiment, the amplifying step comprises performing a polymerase chain reaction, wherein the first and third probes are forward primers for the polymerase chain reaction and the second and fourth probes are reverse primers. In another embodiment, the amplifying step comprises performing a polymerase chain reaction, wherein the first and third amplification primers are forward primers for the polymerase chain reaction and the second and fourth amplification primers are reverse primers. In some embodiments, the amplifying step comprises performing a ligase chain reaction. The amplifying step may comprise a polymerase chain reaction, a reverse transcriptase polymerase chain reaction, a ligase chain reaction, or a ligase chain reaction followed by a polymerase chain reaction.

[0203] In some embodiments, the method for analyzing a single cell comprises a single cell contained in a population of at least 25,000 cells, at least 50,000 cells, at least 75,000 cells, or at least 100,000 cells. In some embodiments, the single cell is a unique cell relative to the remaining cells in the population. In other embodiments, the single cell is representative of a subpopulation of cells in the population. In some embodiments, the population can be considered to be the total number of cells analyzed in the method of the present invention. In one embodiment, performing a high-throughput sequencing reaction to generate sequence information is performed on at least 1,000,000 fusion complexes from at least 10,000 cells in the cell population.

[0204] In some embodiments, a method includes introducing a unique barcode sequence comprising at least six nucleotides into each of a plurality of single cells, wherein each barcode sequence is selected from a pool of barcode sequences having greater than 1,000-fold sequence diversity. For each of the plurality of single cells, the method includes providing at least one set of nucleic acid probes. The method includes analyzing at least two nucleic acid sequences in a single cell contained within a population of at least 10,000 cells, wherein the step includes isolating each of the plurality of single cells from the population of at least 10,000 cells in emulsion microdroplets or reaction vessels. The method includes introducing a unique barcode sequence comprising at least six nucleotides into each of the plurality of single cells, wherein each barcode sequence is selected from a pool of barcode sequences having greater than 1,000-fold sequence diversity.

[0205] In some embodiments, the barcode sequence is fixed to a bead or a solid surface. The bead or solid surface can be isolated in an emulsion microdroplet or a reaction vessel. In other aspects, the method includes introducing a unique barcode sequence, including fusing an emulsion microdroplet or a reaction vessel containing a barcode sequence fixed to a bead or a solid surface with an emulsion microdroplet or a reaction vessel containing a single cell. The second target nucleic acid sequence can be complementary to an RNA sequence. The second target nucleic acid sequence can be complementary to a DNA sequence. In certain embodiments, the amplification includes performing a polymerase chain reaction, a ligase chain reaction, or a ligase chain reaction followed by a polymerase chain reaction.

[0206] In one embodiment, the single cell is contained within a population of at least 25,000 cells. In other embodiments, the single cell is contained within a population of at least 50,000 cells. The single cell may be contained within a population of at least 75,000 cells or at least 100,000 cells. In certain embodiments, the method also includes a step of quantifying the fusion complex.

[0207] In some embodiments, a microfluidic device is used to generate single-cell emulsion droplets. The microfluidic device ejects single cells in an aqueous reaction buffer into a hydrophobic oil mixture. The device can generate thousands of emulsion droplets per minute. After the emulsion droplets are generated, the device ejects the emulsion mixture into a trough. The mixture can be pipetted into a standard reaction tube for thermal cycling or collected.

[0208] Custom microfluidic devices for single-cell analysis are routinely fabricated in academic and commercial laboratories (Kintses et al., 2010 Current Opinion in Chemical Biology 14:548-555). For example, chips can be fabricated from polydimethylsiloxane (PDMS), plastic, glass, or quartz. In some embodiments, fluids move through the chip by pressure or the action of syringe pumps. Single cells can even be manipulated on programmable microfluidic chips using custom dielectrophoresis devices (Hunt et al., 2008 Lab Chip 8:81-87). In one embodiment, a pressure-based PDMS chip composed of flow-focusing geometries fabricated using soft lithography techniques is used (Dolomite Microfluidics, Royston, UK) (Anna et al., 2003 Applied Physics Letters 82:364-366). The stock design can typically generate 10,000 water-in-oil microdroplets per second, ranging in size from 10 to 150 μm in diameter. In some embodiments, the hydrophobic phase consists of a fluorinated oil containing an ammonium salt of carboxy-perfluoropolyether, which ensures optimal conditions for molecular biology and reduces the probability of droplet coalescence (Johnston et al., 1996 Science 271:624-626). To measure the periodicity of cell and droplet streams, images are recorded at 50,000 frames per second using standard techniques such as a Phantom V7 camera or Fastec InLine (Abate et al., 2009 Lab Chip 9:2628-31).

[0209] Microfluidic systems can optimize droplet size, input cell density, chip design, and cell loading parameters so that >98% of droplets contain a single cell. There are three common methods for achieving this: (i) extreme dilution of the cell solution; (ii) fluorescent selection of droplets containing a single cell; and (iii) optimization of the cell input cycle. For each method, measures of success include: (i) encapsulation rate (i.e., the number of droplets containing exactly one cell); (ii) yield (i.e., the proportion of the original cell population that results in droplets containing exactly one cell); (iii) multi-hit rate (i.e., the proportion of droplets containing two or more cells); (iv) negative rate (i.e., the proportion of droplets containing no cells); and (v) encapsulation rate per second (i.e., the number of droplets containing a single cell formed per second).

[0210] In some embodiments, a simple microfluidic chip with a droplet-making junction is used, where a stream of water flows through a 10 μm square nozzle to dispense a water-in-oil emulsion mixture into a reservoir. The emulsion mixture can then be pipetted from the reservoir and subjected to thermal cycling in a standard reaction tube. This method appears to predictably yield high encapsulation rates and low multi-hit rates, but the encapsulation rate per second is low. Designs capable of achieving 1000 Hz packed droplet throughput have been shown to produce up to 10 encapsulations in under 17 minutes. 6 Individual cells can be selected. In some embodiments, the method of the present invention uses single cells in a reaction vessel instead of emulsion droplets. Examples of such reaction vessels include 96-well plates, 0.2 mL tubes, 0.5 mL tubes, 1.5 mL tubes, 384-well plates, 1536-well plates, etc.

[0211] PCR is used to amplify many types of sequences, including but not limited to SNPs, short tandem repeats (STRs), variable protein domains, methylated regions, and intergenic regions.Methods for overlap extension PCR are used to create fusion amplicon products of several independent genomic loci in a single tube reaction (Johnson et al., 2005 Genome Research 15:1315-24; U.S. Patent No. 7,749,697).

[0212] In some embodiments, at least two nucleic acid target sequences (e.g., first and second nucleic acid target sequences, or first and second loci) are selected in a cell and designated as target loci. A forward primer and a backward primer are designed for each of the two nucleic acid target sequences, and these primers are used to amplify the target sequences. A "minor" amplicon is generated by amplifying the two nucleic acid target sequences separately, and then fused by amplification to generate a fusion amplicon, also known as a "major" amplicon. In one embodiment, the "minor" amplicon is a nucleic acid sequence amplified from a target genome locus, and the "major" amplicon is a fusion complex generated from the sequences amplified between multiple genome loci.

[0213] PCR primers are designed for the target of interest, 20–50 nucleotides in length, using standard parameters, i.e., a melting temperature (Tm) of approximately 55–65°C. Primers are used with standard PCR conditions, e.g., 1 mM Tris-HCl pH 8.3, 5 mM potassium chloride, 0.15 mM magnesium chloride, 0.2–2 μM primers, 200 μM dNTPs, and a thermostable DNA polymerase. Many commercial kits are available for performing PCR, such as Platinum Taq (Life Technologies), Amplitaq Gold (Life Technologies), Titanium Taq (Clontech), Phusion polymerase (Finnzymes), and HotStartTaq Plus (Qiagen). Any standard thermostable DNA polymerase, e.g., Taq polymerase or Stoffel fragment, can be used for this step.

[0214] In one embodiment, a set of nucleic acid probes (or primers) is used to amplify a first target nucleic acid sequence and a second target nucleic acid sequence to form a fusion complex. The first probe comprises a sequence complementary to the first target nucleic acid sequence (e.g., the 5' end of the first target nucleic acid sequence). The second probe comprises a sequence complementary to the first target nucleic acid sequence (e.g., the 3' end of the first target nucleic acid sequence) and a second sequence complementary to an exogenous sequence. In some embodiments, the exogenous sequence is a non-human nucleic acid sequence and is not complementary to any of the target nucleic acid sequences. The first and second probes are forward and reverse primers for the first target nucleic acid sequence.

[0215] The third probe contains a sequence complementary to the portion of the second probe that is complementary to the exogenous sequence and a sequence complementary to the second target nucleic acid sequence (e.g., the 5' end of the second target nucleic acid sequence). The fourth probe contains a sequence complementary to the second target nucleic acid sequence (e.g., the 3' end of the second target nucleic acid sequence). The third and fourth probes are forward and reverse primers for the second target nucleic acid sequence.

[0216] The second and third probes are also called the "internal" primers of the reaction (i.e., the reverse primer for the first locus and the forward primer for the second locus) and are limited in concentration (e.g., 0.01 μM for the internal primers and 0.1 μM for all other primers). This allows amplification of the major amplicon to be driven preferentially over the minor amplicon. The first and fourth probes are called "external" primers.

[0217] In other embodiments, multiple barcodes are fused to RNA transcripts from a single cell by binding to bead-immobilized probes followed by first-strand cDNA synthesis and subsequent PCR.

[0218] The first and second nucleic acid sequences are independently amplified, with the first nucleic acid sequence being amplified using the first and second probes, and the second nucleic acid sequence being amplified using the third and fourth probes. The complementary sequence regions of the amplified first and second nucleic acid sequences are then hybridized, and the hybridized sequence is amplified using the first and fourth probes to generate a fusion complex. This is called overlap extension PCR amplification. In another embodiment, multiple barcodes are fused with RNA transcripts from a single cell by binding to probes immobilized on beads, followed by first-strand cDNA synthesis and subsequent PCR.

[0219] During overlap extension PCR amplification, the complementary sequence regions of the amplified first and second nucleic acid sequences act as primers for extension in each direction on both strands by DNA polymerase molecules. In subsequent PCR cycles, the outer primers prime the complete fusion sequence so that the fusion complex is replicated by DNA polymerase. This method generates multiple fusion complexes. In another embodiment, multiple barcodes are fused with RNA transcripts from a single cell by binding to probes immobilized on beads, followed by first-strand cDNA synthesis and subsequent PCR.

[0220] In some embodiments, multiple loci can be targeted in a single cell, multiple sets of probes can be multiplexed into a single analysis to analyze several loci, or even the entire transcriptome or genome. Multiplex PCR is a variation of PCR that uses multiple primer sets in a single PCR mixture to generate amplicons of various sizes specific to different DNA sequences. By targeting multiple genes at once, additional information can be obtained from a single test run that would otherwise require several times more reagents and more time. In one embodiment, 10-20 different transcripts are targeted in a single cell and linked to a second target nucleic acid (e.g., linked to a mutated gene sequence, barcode, or variable region such as an immune variable region). In some embodiments, multiple barcodes are fused to RNA transcripts from a single cell by binding to bead-immobilized probes followed by first-strand cDNA synthesis and subsequent PCR.

[0221] In one embodiment, single cells are encapsulated in picoliter water-in-oil microdroplets. The droplets allow for compartmentalization of reactions, allowing molecular biology procedures to be performed on millions of single cells in parallel. Monodisperse water-in-oil microdroplets can be generated on microfluidic devices in the size range of 10-150 μm in diameter. Alternatively, droplets can be generated by vortexing or a TissueLyser (Qiagen). Two embodiments of oil and aqueous solutions for generating PCR microdroplets are: (i) a PCR buffer containing 0.5 μg / μL bovine serum albumin (New England Biolabs) in combination with a mixture of fluorocarbon oil (3M), Krytox 157FSH surfactant (Dupont), and PicoSurf (Sphere Microfluidics); and (ii) a PCR buffer containing 0.1% Tween® 20 (Sigma) in combination with a mixture of light mineral oil (Sigma), EM90 (Evonik), and Triton® X-100 (Sigma). Several replicate assays quantifying one million amplicons by next-generation sequencing showed that both chemistries form monodisperse microdroplets with greater than 99.98% stability after 40 cycles of PCR. PCR can occur in standard thermocycling tubes, 96-well plates, or 384-well plates using a standard thermocycler (Life Technologies). PCR can also occur in a heated microfluidic chip or any other type of vessel capable of holding an emulsion and transferring heat.

[0222] After thermocycling and PCR, the amplified material must be recovered from the emulsion. In one embodiment, ether is used to break the emulsion, followed by evaporation of the ether from the aqueous / ether layer to recover the amplified DNA in solution. Other methods include adding a detergent to the emulsion, flash-freezing with liquid nitrogen, and centrifugation. Once the ligated and amplified products are recovered from the emulsion, there are many ways to prepare the products for bulk sequencing. In one embodiment, the major amplicon is isolated from the minor amplicon using gel electrophoresis. If the yield is insufficient, the major amplicon is amplified again using PCR and two external primers. This material can then be directly sequenced using bulk sequencing. In some embodiments, external primers are used to generate molecules that can be directly sequenced. In other embodiments, adapters must be added to the major amplicon before bulk sequencing. Once the sequencing library is synthesized, bulk sequencing can be performed using standard methods without major modifications.

[0223] Bulk sequencing requires disruption of cells or emulsion microdroplets so that all polynucleic acid analytes are pooled into a single reaction mixture. It is typically not possible to trace specific sequence targets back to specific cells from bulk sequencing data. However, many applications may require tracing sequences back to their single cell of origin. For example, a researcher may wish to analyze the single-cell expression patterns of a cell population for two RNA transcripts. Overlap extension reverse transcriptase PCR amplification of the two RNA transcript targets, followed by bulk sequencing, is not appropriate for such analyses because all of the transcripts are mixed together, making it impossible to distinguish transcripts from high-expressing cells from those from low-expressing cells. To address this issue, polynucleic acid barcodes are used. Each single-cell emulsion microdroplet or physical reaction vessel contains a single, unique, clonal polynucleic acid barcode. This barcode is then ligated to the target polynucleic acid (i.e., RNA transcript) and used to trace the major amplicon back to a single cell (see Figures 4-10). By tracing each sequence back to the single cell of origin, the genetic data for each single cell can be tabulated, which subsequently allows single-cell quantification (i.e., gene expression levels in a single cell).

[0224] In one embodiment, the linker barcode oligonucleotides are highly diluted so that emulsion microdroplets less than 1 picoliter carry two or more linker barcodes. This allows a single cell to be linked to a single barcode. The linker barcode oligonucleotides are amplified by PCR using universal primers in each droplet, so that each droplet contains millions of copies of only one linker barcode sequence, and that barcode is unique to that droplet. The dilution follows Poisson statistics, such that for P(k=1)≈0.99, the linker barcode needs to be diluted to λ≈0.01. The barcodes are then physically linked to the target molecule by overlap extension PCR or by binding to probes immobilized on beads followed by first-strand cDNA synthesis. Nucleic acid amplification using bead emulsions is described in U.S. Patent No. 7,842,457.

[0225] In another embodiment, a microfluidic device injects beads coated with clonal linker barcode oligonucleotides into single-cell emulsion microdroplets. Such a device allows visualization of single beads and single cells in each droplet, eliminating the need for highly diluted linker barcode oligonucleotides. In this embodiment, PCR is also used to amplify the linker barcode oligonucleotides so that each droplet contains millions of copies of the same barcode sequence, but each barcode is unique to a single microdroplet. The barcodes are then ligated to target nucleic acid sequences using overlap extension PCR. During overlap extension PCR amplification, complementary sequence regions of the amplified first and second nucleic acid sequences act as primers for extension in each direction on both strands by DNA polymerase molecules. In other embodiments, barcodes are ligated to endogenous RNA transcripts by binding to bead-immobilized probes, followed by first-strand cDNA synthesis and subsequent PCR. In some embodiments, in overlap extension PCR, the outer primers prime the entire fusion sequence to be replicated by DNA polymerase in subsequent PCR cycles. These methods generate multiple fusion complexes.

[0226] Many commercial methods exist for polynucleic acid sequencing. These technologies are often referred to as "high-throughput sequencing," "next-generation sequencing," "massively parallel sequencing," or "bulk sequencing." These terms are used interchangeably to describe any sequencing method capable of obtaining more than one million polynucleic acid sequences in a single run. Typically, these methods work by performing highly parallel measurements, i.e., parallel screening of millions of DNA clones on glass slides. Methods for linking multiple polynucleic acid targets in a single cell can be used in conjunction with any commercial bulk sequencing method. These methods include reversible terminator chemistry (Illumina), single-molecule sequencing (Pacific Biosciences), and others (IonTorrent, Element Biosciences, Ultima Genomics, etc.).

[0227] After the molecular ligation protocol is performed, it is useful to specifically amplify and purify the major amplicons before high-throughput sequencing to reduce the overall sequencing required to obtain useful data. Otherwise, many minor amplicons and other types of unwanted background sequences will be unnecessarily sequenced. This can be achieved by performing PCR using only external primers and nucleic acid analytes obtained from lysed cells, followed by size selection using methods such as gel agarose electrophoresis. Other methods, such as size exclusion columns, microfluidic electrophoresis, or micropore filters, can be used to select molecules of appropriate size.

[0228] In one embodiment, the method provides for performing a high-throughput sequencing reaction to generate sequence information for at least 100,000 fusion complexes from at least 10,000 cells in a cell population. In another embodiment, the high-throughput sequencing reaction generates sequence information for at least 75,000, 50,000, or 25,000, or 10,000 fusion complexes from at least 10,000 cells in a cell population. The fused complexes can then be used to quantify a specific biological or clinical phenomenon of interest.

[0229] In the case of functional T cell or B cell analysis, the CDR3 peptide sequence of the fusion complex is first determined, and then specific clonal types can be analyzed by tabulating the instances of the CDR3 peptide linked to a nucleic acid encoding a specific barcode or reporter protein.In this way, high-throughput sequencing quantifies which clonal types are transduced by the retrovirus of the present invention, or which clonal types are displayed on the surface of lentivirus or retrovirus to efficiently transduce target cells.In some embodiments, the present invention includes a single-cell method for determining a lentivirus or retrovirus library containing nucleic acid barcodes linked to scFv or other target cell-specific proteins displayed on the surface of an engineered envelope vector, followed by transduction of target cells by the library, and then identifying which scFv or other target cell-specific proteins are associated with successful transduction of target cells.

[0230] In the case of linkage between barcodes and transcript targets, high-throughput sequencing data is hierarchical by barcode and then tabulated by instances of specific barcodes linked to transcript targets. When primers targeting multiple transcripts are multiplexed in a single assay, barcodes are used to infer multiplexed gene expression patterns for a single cell that can be traced back to a single droplet. In the case of linkage between mutant or variable sequences and other mutant or variable sequences, bulk sequencing data is analyzed to determine the sequence at each locus for each molecule in the high-throughput sequencing library, and then instances of each sequence type are tabulated. 6.8. Methods for delivering nucleic acids into cells

[0231] The engineered envelope vectors and methods of the present invention are useful in in vitro expression systems, where the inserted heterologous gene comprising the engineered envelope vector encodes a protein or peptide that is desired to be produced in vitro.The engineered envelope vectors and methods of the present invention are useful in in vivo expression systems, where the inserted heterologous gene comprising the engineered envelope vector encodes a protein or peptide that is desired to be produced in vivo.

[0232] Described herein is a method for delivering a nucleic acid to a cell, the method comprising: (i) providing an engineered envelope vector of the present invention; and (ii) contacting the engineered envelope vector with a cell so that the engineered envelope vector enters or infects the cell. In some embodiments, the nucleic acid encodes an mRNA molecule, and optionally the mRNA is a gene of interest. In some embodiments, the nucleic acid encodes double-stranded RNA, antisense RNA, microRNA, or any other RNA molecule. In some embodiments, the gene of interest encodes a protein. In some embodiments, the nucleic acid encodes a CRISPR / Cas protein, a zinc finger, or other recombinant protein for genome engineering. In some embodiments, the gene of interest encodes a therapeutic protein (e.g., a protein for compensating for a disease state in a subject).

[0233] In some embodiments, the nucleic acid is delivered to the cell when the engineered envelope vector enters or infects the cell during step (ii). In some embodiments, the methods of delivering nucleic acids described herein do not require a transfection agent (e.g., a lipophilic transfection agent such as Lipofectin).

[0234] In some embodiments, the engineered encapsulation vector is further formulated to be encapsulated in a lipid nanoparticle.

[0235] In some embodiments, nucleic acids are delivered to target cells using lipid nanoparticles comprising lipids tethered or conjugated to one or more of any of SEQ ID NOs: 1-167. In some embodiments, lipid nanoparticles comprising lipids tethered or conjugated to one or more of any of SEQ ID NOs: 1-167 are used to deliver CRISPR / Cas proteins, zinc fingers, or other recombinant proteins for genome engineering. 6.9. Methods for vaccination and gene therapy

[0236] The pharmaceutical preparations, e.g., vaccines, of the present invention comprise an immunogenic amount of the engineered envelope vector disclosed herein in combination with a pharmaceutically acceptable carrier. An "immunogenic amount" is an amount of infectious virus particles sufficient to induce an immune response in a subject to which the pharmaceutical preparation is administered. Exemplary pharmaceutically acceptable carriers include, but are not limited to, sterile pyrogen-free water and sterile pyrogen-free saline. Subjects that can be administered an immunogenic amount of the infectious replication-defective virus particles of the present invention include, but are not limited to, humans and animals (e.g., pigs, cows, dogs, horses, donkeys, mice, hamsters, and monkeys).

[0237] Pharmaceutical formulations of the present invention include those suitable for parenteral (e.g., subcutaneous, intradermal, intramuscular, intravenous, and intraarticular), oral, or inhalation administration. Alternatively, pharmaceutical formulations of the present invention may be suitable for administration to a mucous membrane of a subject (e.g., intranasal administration). The formulations can be conveniently prepared in unit dosage form and can be prepared by any method well known in the art. In some embodiments, the formulation involves post-translational modification of the engineered envelope vector, e.g., the addition of a polyethylene glycol (PEG) moiety to the engineered envelope vector. In some embodiments, the PEG moiety is added to the purified engineered envelope vector prior to formulation, filling, and finishing in a storage buffer. In some embodiments, the formulation buffer for the engineered envelope vector comprises a formulation of 50 mM HEPES, 20 mM MgCl2, pH 7.5, containing a stabilizer, e.g., 10% sucrose. In some embodiments, the formulation buffer is PBS or HEPES.

[0238] The engineered envelope vectors, methods, and pharmaceutical preparations of the present invention are further useful in methods of administering a protein or peptide to a subject in need thereof as a method of treatment or otherwise. In some embodiments of the present invention, a heterologous gene comprising an engineered envelope vector of the present invention encodes a desired protein or peptide, and packaging cells or pharmaceutical preparations containing packaging cells of the present invention are administered to a subject in need of the desired protein or peptide. In this way, the protein or peptide can be produced in vivo in the subject. The subject may need the protein or peptide because the subject has a deficiency of the protein or peptide, or because producing the protein or peptide in the subject as a method of treatment or otherwise may provide some therapeutic effect, as further described below.

[0239] The gene transfer technology of the present invention has several research applications. In some embodiments, cloned DNA or genomic sequences for proteins can be introduced in vivo into multicellular organisms, such as patients, to study cell-specific differences in processing and cell fate. In some embodiments, by placing the coding sequence under the control of a strong promoter, significant quantities of the desired protein can be produced. In some embodiments, specific residues involved in protein processing, intracellular sorting, or biological activity are determined by mutational changes in distinct residues of the coding sequence.

[0240] The gene transfer technology of the present invention is also applied to provide a means of controlling protein expression and evaluating its ability to modulate cellular events. In some embodiments, some functions of proteins, such as their role in differentiation, are studied in tissue culture, while other functions require reintroduction into an in vivo system at different time points in development to monitor changes in related properties. In some embodiments, gene transfer provides a means for studying nucleic acid sequences and cellular factors that regulate the expression of specific genes. In some embodiments, the regulatory element to be studied is fused to a reporter gene, and the expression of this reporter gene is then assayed.

[0241] Gene transfer also has considerable utility in providing treatment for disease states. In some embodiments, the retrovirus contains a transgene. In some embodiments, the transgene is any nucleic acid of interest that is transcribed. In some embodiments, the transgene encodes a polypeptide. In some embodiments, the polypeptide has some therapeutic benefit. There are many genetic diseases for which defective genes are known and have been cloned, such as sickle cell anemia, muscular dystrophy, cystic fibrosis, thalassemia, phenylketonuria, color blindness, skeletal dysplasia, hemophilia, immunodeficiency, and thousands of other conditions. Generally, the disease states described above fall into two classes: deficiency states, usually enzymes, which are generally inherited in a recessive manner, and imbalance states, at least sometimes involving regulatory or structural proteins, which are inherited in a dominant manner. In some embodiments, for deficiency state diseases, retroviral gene transfer of the invention is used to bring normal genes into affected tissues for replacement therapy and to generate animal models for the disease using antisense mutations. In some embodiments, in the case of an imbalanced disease state, gene transfer using the engineered envelope vector of the present invention is used to create a disease state in a model system, which is then used to combat the disease state.Therefore, in some embodiments, the compositions and methods of the present invention enable the treatment of genetic diseases.As used herein, in some embodiments, a disease state is treated by partially or completely correcting the defect or imbalance that causes the disease or makes the disease more severe.In some embodiments, site-specific integration of nucleic acid sequences is used to cause mutations or correct defects.

[0242] In some embodiments, hematopoietic stem cells (HSCs), lymphocytes, vascular endothelial cells, respiratory epithelial cells, keratinocytes, skeletal and cardiac muscle cells, satellite cells (e.g., resident muscle stem cells), neurons, and cancer cells are targets for therapeutic gene transfer either ex vivo or in vivo. See, for example, AD Miller, Nature 357, 455-460 (1992); RC Mulligan, Science 260, 926-932 (1993). These and other cells are suitable target cells for the engineered envelope vectors and methods of the present invention. In some embodiments, the engineered envelope vectors of the present invention are used to deliver chimeric antigen receptors (CARs) or T cell receptors (TCRs) to T cells in vivo or ex vivo. In some embodiments, the CAR or TCR is preferentially expressed on the surface of tumor cells or is directed to a protein or peptide or other molecular marker expressed on the surface of tumor cells. In some embodiments, the CAR or TCR directs T cells, or in particular regulatory T cells, to alleviate the pathology of autoimmunity, graft-versus-host disease, or host-versus-graft disease. In some embodiments, the engineered envelope vector of the present invention further encodes a promoter specific to regulatory T cells, such as a FoxP3 promoter, and thus specifically expresses the gene cargo in regulatory T cells. In some embodiments, the engineered envelope vector delivers DNA encoding the FoxP3 gene to T cells, which drives the phenotype of T cells toward a regulatory T cell phenotype. In some embodiments, the FoxP3 gene is delivered to CD4+ cells in vivo, thereby differentiating the CD4+ cells into regulatory T cells through the expression of the FoxP3 protein. In some embodiments, the engineered envelope vector of the present invention is used to directly deliver a gene payload to tumor cells, thereby killing the tumor cells, for example, by the expression of a protein that induces programmed cell death.In some embodiments, the engineered envelope vector is specifically directed to skeletal muscle cells, satellite muscle cells, cardiomyocytes, muscle cells, pancreatic cells, hepatocytes, kidney cells, epithelial cells, stem cells, neurons, dendritic cells, macrophages, regulatory T cells, central memory T cells, CD4+ T cells, CD8+ T cells, memory B cells, plasmablasts, NK cells, osteoblasts, chondrocytes, adipocytes, eggs, sperm, melanocytes, keratinocytes, Merkel cells, Langerhans cells, neutrophils, eosinophils, basophils, lung cells, stomach cells, colon cells, small intestine cells, brain cells, skin cells, hematopoietic stem cells, CD34+ cells, cancer cells, or any other cell type or subtype with clinical utility. In some embodiments, the target cell-targeting molecule contained in the engineered envelope vector directs cell-type-specific gene delivery. In some embodiments, cell-type-specific gene delivery is important for clinical efficacy because delivering genes to certain cells poses safety risks, such as the risk of genotoxicity. In some embodiments, cell-type specific gene delivery is important for efficacy because more specific on-target binding reduces the required effective dose, and the dose of engineered envelope vectors is limited by toxicity, such as hepatotoxicity. In some embodiments, the cell-type specificity of gene delivery is further enhanced by retroviruses containing DNA gene payloads that further include tissue- or cell-type specific regulatory elements, such as enhancers or promoters. In some embodiments, the regulatory elements are synthetic sequences, while in some embodiments, the regulatory elements are endogenous naturally occurring sequences.

[0243] In some embodiments, it is desirable to use the lentivirus of the present invention to modulate the expression of gene regulatory molecules in cells.In this context, the term "modulate" refers to the suppression of gene expression when gene is overexpressed, or the enhancement of gene expression when gene is underexpressed.In some embodiments, when cell proliferation disorder is related to gene expression, the nucleic acid sequence that interferes with gene expression at the translation level is used.In some embodiments, this approach utilizes, for example, antisense nucleic acid, ribozyme or triple chain element, and blocks the transcription or translation of specific mRNA, either by masking this mRNA with antisense nucleic acid or triple chain element, or by cleaving it with ribozyme.

[0244] Antisense nucleic acids are DNA or RNA molecules complementary to at least a portion of a specific mRNA molecule (Weintraub, 1990, Sci. Am. 262:40). In cells, antisense nucleic acids hybridize to the corresponding mRNA to form double-stranded molecules. Because cells do not translate double-stranded mRNA, antisense nucleic acids interfere with mRNA translation. Antisense oligomers of about 15 nucleotides or more are preferred because they are easily synthesized and are less likely to cause problems than larger molecules when introduced into target cells. The use of antisense methods to inhibit in vitro gene translation is well known in the art (Marcus-Sakura, 1988, Anal. Biochem. 172:289). In some embodiments of the present invention, retroviruses deliver antisense nucleic acids to block the expression of mutant proteins or dominantly active gene products, such as amyloid precursor protein, which accumulates in Alzheimer's disease, or mutant dystrophin. In some embodiments, such compositions and methods are used for the treatment of Huntington's disease, hereditary Parkinsonism, and other diseases. In some embodiments, retroviruses are used to deliver antisense nucleic acids for the inhibition of expression of proteins associated with toxicity.

[0245] Ribozymes are RNA molecules that have the ability to specifically cleave other single-stranded RNA in a manner similar to DNA restriction endonucleases.Through the modification of the nucleotide sequence that codes for these RNAs, it is possible to engineer molecules that recognize and cleave specific nucleotide sequences in RNA molecules (Cech, 1988, J. Aer. Med Assn. 260:3030).The main advantage of this approach is that only mRNAs with specific sequences are inactivated.In some embodiments, the retrovirus of the present invention is used to deliver ribozymes to target cells for gene therapy in clinical applications.

[0246] In some embodiments, lentiviruses of the present invention are used in clinical gene therapy to deliver nucleic acids encoding biological response modifiers. This category includes immune enhancers, including nucleic acids encoding several cytokines classified as interleukins, such as interleukins 1-12. Also included in this category, although they do not necessarily act via the same mechanisms, are interferons, particularly gamma interferon (γ-IFN), tumor necrosis factor (TNF), and granulocyte-macrophage colony-stimulating factor (GM-CSF). In some embodiments, it is desirable to deliver such nucleic acids to bone marrow cells or macrophages to treat congenital enzyme deficiencies or immune deficiencies. In some embodiments, nucleic acids encoding growth factors, toxic peptides, ligands, receptors, or other physiologically important proteins are delivered to specific cells in vivo or ex vivo.

[0247] CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) evolved in bacteria as an adaptive immune system to defend against viral attack. Upon exposure to a virus, a short segment of viral DNA is integrated into the CRISPR locus. RNA is transcribed from the portion of the CRISPR locus containing the viral sequence. The RNA, which contains a sequence complementary to the viral genome, mediates targeting of the Cas9 protein to the target sequence in the viral genome. The Cas9 protein cleaves the viral target, thereby silencing it. Recently, the CRISPR / Cas system has been adapted for genome editing in eukaryotic cells. Introduction of a site-specific double-strand break (DSB) allows for target sequence modification by one of two endogenous DNA repair mechanisms: nonhomologous end joining (NHEJ) or homology-directed repair (HDR). The CRISPR / Cas system has also been used for gene regulation, including transcriptional repression and activation, without altering the target sequence. Target gene regulation based on the CRISPR / Cas system uses enzymatically inactive Cas9 (also known as catalytically inactive Cas9). In addition to the standard Cas9 nuclease, other nucleic acid-guided nucleases have been discovered, including CasX, Cas12a (including MAD7), Cas12b, Cas12c, and Cas13. In some embodiments of the present invention, retroviruses deliver Cas9, CasX, Cas12a (including MAD7), Cas12b, Cas12c, or Cas13 to any target cell that has therapeutic or research utility.

[0248] In preferred embodiments, the guide nucleic acid forms a complex with a compatible nucleic acid-guided nuclease. In some embodiments, the nucleic acid-guided nuclease is used together with a heterologous guide nucleic acid. In some embodiments, the guide nucleic acid and the heterologous guide nucleic acid are derived from two different species. In some embodiments, the guide nucleic acid and the heterologous guide nucleic acid are derived from the same species. In some embodiments, the guide nucleic acid and the heterologous guide nucleic acid are derived from the same species but do not naturally exist in the same cell.

[0249] The compatibility of the nucleic acid-guided nuclease and the guide nucleic acid can be determined by empirical testing. The heterologous guide nucleic acid may be derived from a different bacterial species, or may be non-naturally occurring, synthetic, or engineered. In some embodiments, the guide nucleic acid is DNA. In some embodiments, the guide nucleic acid is RNA. In some embodiments, the guide nucleic acid comprises both DNA and RNA. In some embodiments, the guide nucleic acid comprises non-naturally occurring nucleotides. When the guide nucleic acid comprises RNA, the RNA guide nucleic acid may be encoded by a DNA sequence.

[0250] In some embodiments, guide nucleic acid comprises one or more polynucleotides.In some embodiments, guide nucleic acid comprises a guide sequence that can hybridize with target sequence and a scaffold sequence that can interact with or form a complex with nucleic acid-guided nuclease.In some embodiments, guide sequence and scaffold sequence are in a single polynucleotide.In some embodiments, guide sequence and scaffold sequence are in two or more separate polynucleotides.

[0251] The guide nucleic acid may include a scaffold sequence. Generally, a "scaffold sequence" includes any sequence that has a sequence for promoting the formation of a ribonucleoprotein particle (RNP), and the RNP includes a nucleic acid-guided nuclease and a guide nucleic acid. In some embodiments, the scaffold sequence promotes the formation of an RNP by having two sequence regions within the scaffold sequence, for example, one or two sequence regions that are involved in the formation of a secondary structure, along their length. In some cases, the one or two sequence regions are on the same polynucleotide. In some cases, the one or two sequence regions are on separate polynucleotides. Optimal alignment can be determined by any suitable alignment algorithm, and can further take into account secondary structures such as self-complementarity within one or two sequence regions. In some embodiments, the degree of complementarity between one or two sequence regions along the length of the shorter of the two when optimally aligned is about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99% or more, hi some embodiments, at least one of the two sequence regions is about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50 or more nucleotides in length.

[0252] In some embodiments, the scaffold sequence of the guide nucleic acid comprises a secondary structure. The secondary structure can comprise a pseudoknot region. In some cases, the binding kinetics of the guide nucleic acid to the nucleic acid-guided nuclease is determined in part by the secondary structure within the scaffold sequence. In some cases, the binding kinetics of the guide nucleic acid to the nucleic acid-guided nuclease is determined in part by the nucleic acid sequence having the scaffold sequence.

[0253] In some embodiments, the guide nucleic acid comprises a guide sequence. The guide sequence is a polynucleotide sequence that has sufficient complementarity with the target polynucleotide sequence to hybridize with the target sequence and guide the sequence-specific binding of the complexed nucleic acid-guided nuclease to the target sequence. The degree of complementarity between the guide sequence and its corresponding target sequence can be about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more when optimally aligned using a suitable alignment algorithm. Optimal alignment can be determined using any suitable algorithm for aligning sequences. In some embodiments, the guide sequence is about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75 or more nucleotides in length. In some embodiments, the guide sequence is less than about 75, 50, 45, 40, 35, 30, 25, 20 nucleotides in length. In preferred embodiments, the guide sequence is 10-30 nucleotides in length. The guide sequence may be 15-20 nucleotides in length. The guide sequence may be 15 nucleotides in length. The guide sequence may be 16 nucleotides in length. The guide sequence may be 17 nucleotides in length. The guide sequence may be 18 nucleotides in length. The guide sequence may be 19 nucleotides in length. The guide sequence may be 20 nucleotides in length.

[0254] A guide nucleic acid can be engineered to target a desired target sequence by modifying the guide sequence so that it is complementary to the target sequence, thereby allowing hybridization between the guide sequence and the target sequence. A guide nucleic acid with an engineered guide sequence can be referred to as an engineered guide nucleic acid. Engineered guide nucleic acids often do not exist in nature.

[0255] In one aspect, the present disclosure provides a method for modifying a target region of a eukaryotic or prokaryotic genome using the gene editing system provided herein. The method may include (1) contacting a sample containing the target region with (i) a nucleic acid-guided nuclease and (ii) a guide nucleic acid complexed with the nucleic acid-guided nuclease, and (2) allowing the nucleic acid-guided nuclease to modify the target region. In some embodiments, the sample is further contacted with (iii) a homology template configured to bind to the target region. In some embodiments, the sample comprises a eukaryotic cell, a bacterial cell, a plant cell, a mammalian cell, or a human cell. In some embodiments, the sample comprises an immune cell. In some embodiments, the immune cell is a B cell or a T cell. In some embodiments, one or more vectors encoding one or more components of the gene editing system are introduced into a host cell. In some embodiments, the nucleic acid-guided nuclease and the guide nucleic acid are operably linked to separate regulatory elements on separate vectors. In some embodiments, two or more elements expressed from the same or different regulatory elements combined in a single vector are introduced. When several elements are combined in a single vector, the coding sequence of one element can be located on the same or opposite strand of the coding sequence of the second element, and can be oriented in the same or opposite direction. In some embodiments, a single promoter drives the expression of transcripts encoding the nucleic acid-guided nuclease and one or more guide nucleic acids. In some embodiments, the nucleic acid-guided nuclease and one or more guide nucleic acids are operably linked to and expressed from the same promoter. In other embodiments, one or more guide nucleic acids or polynucleotides encoding one or more guide nucleic acids are introduced into cells in an in vitro environment already containing the nucleic acid-guided nuclease or polynucleotide sequence encoding the nucleic acid-guided nuclease.

[0256] In some embodiments, the engineered envelope vector of the present invention is used to deliver gRNA to target cells. In some embodiments, the engineered envelope vector of the present invention is used to deliver nucleic acids encoding both gRNA and Cas9, CasX, Cas12a (including MAD7), Cas12b, Cas12c, or Cas13, or any related nuclease, to target cells. In some embodiments, the delivery of nucleic acids encoding gRNA and / or nuclease is used to modify the target cell genome for gene therapy using any of the therapeutic methods described herein. In some embodiments, the nuclease is anchored in the membrane of a lentivirus or retrovirus in protein form, tethered using a chimeric transmembrane domain derived from another protein, and then released in a functional form into the cytoplasm of the target cell.

[0257] In some embodiments, the packaging cell line for the engineered envelope vector is edited to remove HLA expression from the cell membrane. In some embodiments, when the engineered envelope vector is produced using such an HLA-deficient cell line, the engineered envelope vector does not present any peptide:MHC and is therefore less immunogenic when administered in vivo. In some embodiments, the engineered envelope vector of the present invention has additional proteins embedded in the membrane of the engineered envelope vector that extend its half-life in vivo, for example, by reducing immunogenicity or by a "don't eat me" signal such as CD47. 6.10. Nucleic acids

[0258] As used herein, the term "nucleic acid" generally refers to a molecule comprising multiple linked nucleotides (i.e., a sugar (e.g., ribose or deoxyribose) linked to an interchangeable organic base, which is either a pyrimidine (e.g., cytosine (C), thymidine (T), or uracil (U)) or a purine (e.g., adenine (A) or guanine (G))). Nucleic acids include DNA and RNA, such as D-DNA and L-DNA, and various modifications thereof. Modifications include base modifications, sugar modifications, and backbone modifications.

[0259] It should be understood that the nucleic acids used in the engineered envelope vectors and methods of the present invention can be of homogeneous or heterogeneous nature. For example, they can be entirely DNA in nature, or they can be composed of DNA and non-DNA (e.g., LNA) monomers or sequences. Thus, any combination of nucleic acid elements can be used. Modifications can make nucleic acids more stable under certain conditions and / or less susceptible to degradation. For example, in some cases, the nucleic acid is nuclease-resistant. Methods for synthesizing nucleic acids, including automated nucleic acid synthesis, are also known in the art.

[0260] Nucleic acids may contain modifications to their bases. Modified bases include modified cytosine (such as 5-substituted cytosine (e.g., 5-methyl-cytosine, 5-fluoro-cytosine, 5-chloro-cytosine, 5-bromo-cytosine, 5-iodo-cytosine, 5-hydroxy-cytosine, 5-hydroxymethyl-cytosine, 5-difluoromethyl-cytosine, and unsubstituted or substituted 5-alkynyl-cytosine)), 6-substituted cytosine, N4-substituted cytosine (e.g., N4-ethyl-cytosine), 5-aza-cytosine, 2-mercaptocytosine, 5-methyl-cytosine, 5-fluoro-cytosine, 5-chloro-cytosine, 5-bromo-cytosine, 5-iodo-cytosine, 5-hydroxy-cytosine, 5-hydroxymethyl-cytosine, 5-difluoromethyl-cytosine, and unsubstituted or substituted 5-alkynyl-cytosine), 6-substituted cytosine, N4-substituted cytosine (e.g., N4-ethyl-cytosine), 5-aza-cytosine, 2-mercaptocytosine, 5-methyl-cytosine, 5-bromo-cytosine, 5-iodo-cytosine, 5-hydroxymethyl-cytosine, 5-difluoromethyl-cytosine, and unsubstituted or substituted 5-alkynyl-cytosine). puto-cytosine, isocytosine, pseudo-isocytosine, cytosine analogs with fused ring systems (e.g., N,N'-propylenecytosine or phenoxazine), as well as uracil and its derivatives (e.g., 5-fluoro-uracil, 5-bromo-uracil, 5-bromovinyl-uracil, 4-thio-uracil, 5-hydroxy-uracil, 5-propynyl-uracil), modified guanines, e.g., 7-deazaguanine, 7-deaza-substituted guanines (7-deaza-7(C2 C6)alkynylguanine, etc.), 7-deaza-8-substituted guanines, hypoxanthine, N2-substituted guanines (e.g., N2-methyl-guanine), 5-amino-3-methyl-3H,6H-thiazolo[4,5-d]pyrimidine-2,7-dione, 2,6-diaminopurine, 2-aminopurine, purine, indole, adenine, substituted adenines (e.g., N6-methyl-adenine, 8-oxo-adenine), 8-substituted guanines (e.g., 8-hydroxyguanine and 8-bromoguanine), and 6-thioguanine. Nucleic acid can comprise universal base (for example, 3-nitropyrrole, P-base, 4-methyl-indole, 5-nitro-indole and K-base) and / or aromatic ring system (for example, fluorobenzene, difluorobenzene, benzimidazole or dichloro-benzimidazole, 1-methyl-1H-[1,2,4]triazole-3-carboxylic acid amide).The specific base pair that can be incorporated into the oligonucleotide of the present invention is the dZ and dP non-standard nucleobase pair reported by Yang et al. NAR, 2006, 34(21):6095-6101.The pyrimidine analog, dZ, is 6-amino-5-nitro-3-(1'-β-D-2'-deoxyribofuranosyl)-2(1H)-pyridone, and its Watson-Crick complement, dP, the purine analog, is 2-amino-8-(1'-β-D-1'-deoxyribofuranosyl)-imidazo[1,2-a]-1,3,5-triazin-4(8H)-one. 6.11. Amino Acid Substitutions

[0261] In some embodiments, the amino acid residue variation is a conservative amino acid residue substitution. As used herein, "conservative amino acid substitution" refers to an amino acid substitution that does not change the relative charge or size characteristics of the protein in which the amino acid substitution is made. Variants can be prepared according to methods for modifying polypeptide sequences known to those skilled in the art, such as those found in references that compile such methods, for example, Molecular Cloning: A Laboratory Manual, J. Sambrook, et al., eds., Second Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989, or Current Protocols in Molecular Biology, FM Ausubel, et al., eds., John Wiley & Sons, Inc., New York. Conservative amino acid substitutions include substitutions between amino acids within the following groups: (a) M, I, L, V; (b) F, Y, W; (c) K, R, H; (d) A, G; (e) S, T; (f) Q, N; and (g) E, D.

[0262] The "percent identity" of two amino acid sequences is determined using the algorithm of Karlin and Altschul Proc. Natl. Acad. Sci. USA 87:2264-68, 1990, modified as in Karlin and Altschul Proc. Natl. Acad. Sci. USA 90:5873-77, 1993. Such an algorithm has been incorporated into the NBLAST and XBLAST programs (version 2.0) of Altschul, et al. J. Mol. Biol. 215:403-10, 1990. BLAST protein searches can be performed with the XBLAST program, score=50, wordlength=3, to obtain amino acid sequences homologous to a protein molecule of interest. When gaps exist between the two sequences, Gapped BLAST can be utilized as described in Altschul et al., Nucleic Acids Res. 25(17):3389-3402, 1997. When utilizing BLAST and Gapped BLAST programs, the default parameters of the respective programs (e.g., XBLAST and NBLAST) can be used. 6.12. Other Embodiments

[0263] All of the features disclosed herein can be combined in any combination. Each feature disclosed herein may be replaced by an alternative feature serving the same, equivalent, or similar purpose. Thus, unless expressly stated otherwise, each feature disclosed is only an example of a generic series of equivalent or similar features.

[0264] From the above description, those skilled in the art can easily ascertain the essential features of the present invention, and can make various changes and modifications to the present invention to adapt it to various uses and conditions without departing from the spirit and scope thereof. Accordingly, other embodiments are within the scope of the following claims. [Example]

[0265] 7. Working Example Example 1 7.1. Example 1. Identification and sequence analysis of novel membrane fusogenic proteins from sequence databases Public databases of raw (i.e., unassembled) reads currently number in the tens of petabases (petabase = 10 bases 15 The total number of sequences in this corpus is exponentially increasing. Traditional search tools such as BLAST are impractical for searching this data corpus due to the enormous computational cost, which would likely exceed the cloud resources available to large pharmaceutical companies. Serratus (Edgar et al., 2022, doi: 10.1038 / s41586-021-04332-2) is a biological search engine highly optimized for the Amazon Web Services infrastructure, enabling alignment of large sets of query sequences (up to hundreds of Mb) against petabase-scale sequence databases in a practical timeframe (days). See Figure 4. Using Serratus, we searched SRA for novel fusion gene pseudotypes using SEQ ID NO: 1 as the query sequence. Serratus output SEQ ID NOs: 2-167, which are the putative membrane fusion factors used to generate the novel engineered retroviruses disclosed herein.

[0266] Mutant fusogenic factors were generated by identifying functional or otherwise important amino acids using a variety of means, including sequence alignment. One such output is shown in Figure 21.

[0267] This sequence alignment shows that certain amino acids are important for evolutionary conservation. For example, the transmembrane domain of VSVG has previously been annotated as IASFFIIGLIIGLFLVLRVGIHLC (SEQ ID NO: 16558). In the sequence alignment, there is considerable diversity between the listed sequences, but the underlined I and R amino acids indicate the common amino acids found in these widely diverse proteins: [ka] These amino acids are likely to be functional based on the sequence conservation. Other conserved or non-conserved residues using sequence alignment of any two or more sequences containing SEQ ID NOS: 1-167 strongly indicate functionally important residues in fusion factor molecules whose sequences are widely diverse.

[0268] In another Serratus sequence search, sequences including SEQ ID NOs: 15597-16497 were used as bulk search queries to identify SEQ ID NOs: 8154-15596 from SRA and other sequence databases. These query sequences and search results are membrane fusogenic or envelope proteins for engineering novel retroviruses. Example 2 7.2. Example 2. Construction of Engineered Envelope Vectors

[0269] Four candidate fusion gene pseudotypes (Candidate A: SEQ ID NO: 84, Candidate B: SEQ ID NO: 113, Candidate C: SEQ ID NO: 121, and Candidate D: SEQ ID NO: 2) were selected from the Serratus output and used to generate engineered envelope vectors delivering mRuby reporters with or without scFvs (non-viral membrane-bound proteins) directed against CD19, CD3, CD4, and CD8 for cell-type specific tropism.

[0270] Examples of plasmids containing fusogenic factors include: GM.PMD-6510 pTwist CMV BGlobin Chandipura VSV-G (SEQ ID NO: 8125) containing SEQ ID NO: 113, GM.PMD-6511 pTwist CMV-BGlobin Piry VSV-G (SEQ ID NO: 8124) containing SEQ ID NO: 121, GM.PMD-6509 pTwist CMV-BGlobin ABVV-G (SEQ ID NO: 8123) containing SEQ ID NO: 84, and GM.PMD-6512 pTwist CMV-BGlobin Rhinolophus VSV-G (SEQ ID NO: 8122) containing SEQ ID NO: 2. Similar constructs are made for any of SEQ ID NOs: 2-167. Wild-type and mutant (K47Q, R354A) VSV-G pseudotyped lentiviruses are also made using similar nucleic acid construct designs. Third generation lentiviruses are packaged in Lenti-Pac 293Ta cells (GeneCopoeia) using equimolar amounts of packaging plasmids: pRSV-Rev (Cell Biolabs), pCgpV (Cell Biolabs), pTwist-CMV-pseudotype (Twist Bio) and mRuby lentiviral expression plasmid (made in-house).

[0271] The anti-CD19, CD3, CD4 and CD8 scFvs are optionally further tethered to an extracellular Fc domain and a CD28 transmembrane domain (lacking the intracellular signaling domain). Examples of plasmids containing cytotropic scFvs include PMD-6521 pTwist-CMV-aCD19-PDGFR (SEQ ID NO: 8127), which contains an anti-CD19 scFv tethered to a PDGFR stalk; PMD-6522 pTwist-CMV-aCD19-Fc-CD28™ (SEQ ID NO: 8129), which contains an anti-CD19 scFv tethered to an Fc and CD28 domain; PMD-6540 pTwist-CMV-ammCD19-Fc-CD28™ (SEQ ID NO: 8130), which contains an anti-mouse CD19 scFv tethered to an Fc and CD28 transmembrane domain; PMD-6523 pTwist-CMV-aCD3-PDGFR (SEQ ID NO: 8131), which contains an anti-CD3 scFv tethered to a PDGFR domain; and PMD-6524 pTwist-CMV-aCD3-PDGFR (SEQ ID NO: 8132), which contains an anti-CD3 scFv tethered to an Fc and CD28 transmembrane domain. pTwist-CMV-aCD3-Fc-CD28™ (SEQ ID NO: 8132), PMD-6531, comprising anti-mouse CD3 tethered to the Fc and CD28 transmembrane domains; pTwist-CMV-ammCD3-Fc-CD28™ (SEQ ID NO: 8133), PMD-6536, comprising anti-CD4 scFv tethered to the CD28 transmembrane domain; pTwist-CMV-ahCD4.1-Fc-CD28™ (SEQ ID NO: 8134), PMD-6533, comprising anti-CD4 scFv tethered to the Fc and CD28 transmembrane domains; pTwist-CMV-amCD4.1-Fc-CD28™ (SEQ ID NO: 8135), PMD-6534, comprising anti-CD8 scFv tethered to the Fc and CD28 transmembrane domains; pTwist-CMV-amCD8.1-Fc-CD28™ (SEQ ID NO: 8136), PMD-6538 pTwist-CMV-ahCD8.1-Fc-CD28™ (SEQ ID NO: 8137) containing an anti-CD8 scFv tethered to the Fc and CD28 transmembrane domains. Similar constructs are made for antibodies and scFvs against any molecular target. Optionally, a function modulator protein is further introduced into the producer cells on a plasmid.Examples of plasmids include PMD-6520 pTwist-CMV-hsCD80 (SEQ ID NO: 8138), which contains CD80; PMD-6024 PSF-CMV-atezoH_1D11ScFv (SEQ ID NO: 8139), which contains an scFv derived from the atezolizumab sequence; and PMD-3805 pReceiver_EF1a-aA0201-PMEL.4 (SEQ ID NO: 8140), an anti-PMEL TCR. Similar constructs can be made for antibodies and scFvs against any molecular target, or TCRs against any molecular target, or any protein containing SEQ ID NOs: 168-8121. These plasmids are transfected into Lenti-Pac 293Ta cells using Lipofectamine 3000 (ThermoFisher), and lentiviral supernatants are harvested and pooled 24 and 48 hours after transfection. The lentivirus supernatant was centrifuged at 2000 × g for 10 minutes to remove cell debris and filtered through a 45 μm filter. The lentivirus was concentrated using a Lenti-X Concentrator (Takara), and the pelleted virus particles were resuspended in cell culture medium to a concentration of approximately 20×. The lentivirus titer was determined using the isolated viral RNA using a NucleoSpin RNA Virus Kit (Macherey-Nagel) and a Lenti-X qRT-PCR Titration Kit (Takara). Example 3 7.3. Example 3. Testing the Transduction Efficiency and Specificity of Engineered Envelope Vectors

[0272] To test whether the mutant (K47Q, R354A) VSV-G completely abrogates gene delivery, as reported elsewhere (Dodson et al., 2022 doi: 10.1038 / s41592-022-01436-z; Yu et al., 2021 doi: https: / / doi.org / 10.1101 / 2021.12.13.472464), 2.5 × 10 HEK-293 cells were cultured. 4 pieces / cm 2and incubated with a 1:4 dilution of unconcentrated lentiviral particles for 3 days. The mutant (K47Q, R354A) VSV-G-containing virus had approximately 5-fold lower transduction efficiency than wild-type VSV-G virus, as measured by mRuby expression (15-17% vs. 76-78%, respectively). Thus, mutant VSV-G paired with cell-type-specific scFvs appears to result in significant gene delivery to undesired cell types.

[0273] To assess the functionality of candidate fusion factor pseudotypes, 1 x 10 purified T cells were cultured in a 2000-well plate. 6 Cells were seeded at a final cell density of cells / mL and incubated with 1 mL of approximately 20× concentrated anti-CD3 engineered envelope vectors (pseudotyped lentiviral candidate A: SEQ ID NO: 84, candidate B: SEQ ID NO: 113, candidate C: SEQ ID NO: 121, and candidate D: SEQ ID NO: 2) for 2 days. Lentiviral transduction was measured by flow cytometry on a CytoFLEX LX (Beckman Coulter) for mRuby expression in live cells (FIG. 11). All four candidates efficiently delivered mRuby to target T cells, with candidate C being particularly efficient, achieving >95% positivity for mRuby with minimal background signal in anti-CD19 and non-targeting controls.

[0274] We also evaluated the cell-specific affinity of candidate C (SEQ ID NO: 121) using lentiviruses engineered with anti-human CD19 (SEQ ID NO: 8130), anti-human CD4 (SEQ ID NO: 8135), and anti-human CD8 scFv (SEQ ID NO: 8137). Primary peripheral blood mononuclear cells (PBMCs, not purified T cells only) were cultured at 1 x 10 6Cells were seeded at a final cell density of 1 / mL and incubated with 1 mL of approximately 20x concentrated LV for 2 days. Lentiviral transduction was measured by flow cytometry on a CytoFLEX LX (Beckman Coulter) for mRuby expression in live cells. Staining with antibodies against CD19 and CD8 was also performed (Figure 12). Cell type tropism was highly specific; for example, nearly all CD19+ cells became mRuby+ in the anti-CD19 scFv lentivirus experiment, and nearly all CD8+ cells became mRuby+ in the anti-CD8 scFv LV experiment. Example 4 7.4. Example 4. DNA Payloads for Delivery by Engineered Envelope Vectors

[0275] The DNA payload for the retrovirus can be delivered in the form of a plasmid, for example, PMD-6520 pTwist-CMV-hsCD80 (SEQ ID NO: 8138) containing CD80, PMD-6024 PSF-CMV-atezoH_1D11ScFv (SEQ ID NO: 8139) containing an scFv based on the antibody atezolizumab, PMD-3805 pReceiver_EF1a-aA0201-PMEL.4 (SEQ ID NO: 8140) containing an anti-PMEL TCR, PMD-6437_pET21b_GIG17-nuc-myc (SEQ ID NO: 8150) containing a Cas12a-like nuclease, PMD-6541 pReceiver-EF1a-hsCD19-28z-CAR-GFP (SEQ ID NO: 8151) containing an anti-CD19 CAR and a GFP reporter, PMD-6529 pReceiver-EF1a-hsCD19-28z-CAR-GFP (SEQ ID NO: 8152) containing an anti-CD19 CAR and a GFP reporter. pReceiver-EF1a-mmCD19-28z-CAR-GFP (SEQ ID NO: 8152), PMD-6530 pReceiver-EF1a-mmCD19-28z-CAR-Foxp3 (SEQ ID NO: 8153) containing the anti-CD19 CAR and FoxP3 gene. Make similar constructs for antibodies and scFvs against any molecular target, or TCRs against any molecular target, or any protein containing SEQ ID NOs: 168-8121, or any reporter or cell selection marker, or any nuclease for genome editing, or any gRNA for genome editing.

[0276] In one example, lentiviruses are engineered using the above-described method to express lentivirus-embedded OX40L and deliver anti-tumor CARs or anti-tumor TCRs. The lentiviruses are delivered to T cells in vivo or ex vivo. T cells are activated by OX40L via binding to OX40 on the T cell surface, and the activated T cells are more efficient at killing tumors via activated CARs or TCRs (Figure 13). In another example, a similar method is used to deliver anti-tumor TCRs (e.g., anti-WT1 TCRs) to patients with AML (Figure 14), which has the advantage of eliminating the need for lymphodepletion before administering lentiviruses, which is usually required in cell therapy. In another example, a similar method is used to deliver CARs directed to CD123 (SEQ ID NO: 16504), CD19 (SEQ ID NO: 16506), CD20 (SEQ ID NO: 16505), or GPRC5D (SEQ ID NOs: 16507, 16508, 16509) for cancer treatment.

[0277] In another example, lentiviruses are engineered to deliver anti-TFR2 CARs to Tregs using the methods described above, either through tropism directed to Tregs, or delivery of the CAR under the control of the FoxP3 promoter, or delivery of the FoxP3 gene to target cells (thereby inducing a Treg phenotype). Anti-TFR2 directs Tregs to the patient's, for example, liver in liver transplant patients, where they secrete signals such as cytokines to suppress anti-transplant effector T cells (Teff) (Figure 15). This mechanism can be used to reduce or eliminate immunosuppression after liver transplantation without the need for lymphodepletion or conventional cell therapy (Figure 16). Example 5 7.5. Example 5. Delivery of Genome-Engineering Enzymes and Related Nucleic Acids for Gene Therapy

[0278] β-Hemoglobinopathies are the most common single-gene disorders worldwide. These genetic disorders affect the normal production of adult hemoglobin due to mutations in the β-globin gene. The two most common diseases are i) β-thalassemia, which is characterized by reduced or absent β-globin production, and ii) sickle cell disease (SCD), in which a mutant form of β-globin is produced that results in red blood cells (RBCs) with a "sickle"-like shape rather than the normal disc shape.

[0279] Ex vivo gene therapy for β-thalassemia and SCD involves removing hematopoietic stem cells (HSCs) from patients, editing the HSCs with CRISPR / Cas ribonucleoprotein (RNP), and then infusing the edited cells back into the patient for the long-term production of healthy red blood cells. Vertex and CRISPR Therapeutics' Exa-cel knocks out the B-cell lymphoma protein 11A (BCL11A) gene, which normally suppresses the production of fetal hemoglobin (HbF). Over 12 consecutive months, Exa-cel achieved functional cure in 31 of 31 SCD patients. Unfortunately, ex vivo cell therapy is available only at a few cell therapy centers, limiting access and increasing the risk of infertility due to busulfan bone marrow ablation.

[0280] In vivo gene therapy is more accessible because it can be administered in most hospitals and does not require bone marrow ablation with busulfan. Adenoviruses have been used to knock in β-globin in preclinical models for in vivo gene therapy of SCD (Li et al., 2021, https: / / doi.org / 10.1016 / j.omtm.2021.12.003), but such methods are complicated by widespread human seropositivity for adenovirus. Lipid nanoparticles (LNPs) and adeno-associated viruses (AAVs) can be used to deliver CRISPR / Cas machinery for BCL11A knockout, but they are not cell-type specific and have not been able to further knock in β-globin due to payload limitations. Traditional lentiviruses are generally not immunogenic at the initial dose and deliver larger payloads than LNPs and AAVs. However, because traditional lentiviruses lack cell-type specificity, researchers have attempted to engineer HSC-tropic lentivectors containing membrane-bound human stem cell factor (hSCF), which binds to CD117 on HSCs (Froelich et al., 2009, DOI: 10.3109 / 08923970903420582).

[0281] An in vivo delivery vector using an engineered lentivirus containing a CRISPR / Cas9 RNP, further comprising a guide RNA, SEQ ID NO: 16520, is generated in HEK293 lentiviral packaging cells. This gRNA has been shown to be capable of knocking out BCL11A by repressing its enhancer element in vivo and in vitro. For a description of the gRNA, see U.S. Patent Application No. 20190201553, which is hereby incorporated by reference. The lentiviral particle contains a protein derivative generated from plasmid SEQ ID NO: 16510, which contains HIV-1 gag fused to SpCas9, with a 3xNES and proteolytic cleavage site between them. This lentiviral vector further comprises an envelope protein, such as any of SEQ ID NOs: 1-167 or 8154-16497, that specifically and efficiently delivers to HSCs. For further HSC targeting, a plasmid expressing anti-CD117 antibody (SEQ ID NO: 16511) and / or a plasmid expressing hSCF is added to the packaging cells. See Figure 17.

[0282] To test the HSC-targeting BCL11A-editing lentiviral vector, mobilized human peripheral blood CD34+ cells from human donors 1–3 were cultured for 2 days in serum-free StemSpan Medium containing CD34+ expansion supplement. 100,000 cells were washed and transduced with the engineered lentivector. Cells were allowed to recover for 2–3 days before switching to erythroid differentiation medium (IMDM+Glutamax supplemented with 5% human serum, 10 μg / ml insulin, 20 ng / ml SCF, 5 ng / ml IL-3, 3 U / ml EPO, 1 μM dexamethasone, 1 μM β-estradiol, 330 μg / ml holotransferrin, and 2 U / ml heparin). Subsequently, high-throughput sequencing (Illumina) was used to determine the percentage of insertions / deletions ("indels") in the transduced cells. These cells are differentiated in erythroid differentiation medium for 12 days after which RNA is collected and hemoglobin levels assessed by quantitative real-time PCR.

[0283] After one day, single erythroid precursors are generated using flow cytometry and cultured in erythroid differentiation medium to expand and grow as colonies. Each colony is split and collected for DNA and RNA analysis 12 days after sorting. Sister colonies are collected 15 days after sorting for hemoglobin protein analysis. Globin expression (γ / 18sRNA ratio or γ / α ratio) is determined by quantitative real-time PCR and compared for each of the edited erythroid colonies.

[0284] Lentiviral vectors containing CRISPR / Cas RNPs have also been used to engineer satellite muscle cells to correct genetic disorders such as DMD. An in vivo delivery vector (Xiang et al., 2021; https: / / doi.org / 10.1016 / j.omtn.2021.03.005; incorporated herein by reference in its entirety) using engineered lentiviruses containing CRISPR / Cas RNPs and dual guide RNAs designed to correct DMD was generated in HEK293 lentiviral packaging cells. The lentiviral particles contain a protein derivative generated from plasmid SEQ ID NO: 16510, which contains HIV-1 gag fused to SpCas9, with a 3xNES and proteolytic cleavage site between them. The lentiviral vector further contains an envelope protein, such as any of SEQ ID NOs: 1-167 or 8154-16497, for specific and efficient delivery to satellite muscle cells. For further HSC targeting, a plasmid expressing an anti-CD117 antibody (SEQ ID NO: 16511) is added to the packaging cells.

[0285] The one or more DNA endonucleases may be Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas100, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cm r4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4 or Cpf1 endonuclease, homologs thereof, recombinant forms of these naturally occurring molecules, codon-optimized versions thereof, or modified versions thereof, and combinations thereof. Example 6 7.6. Example 6. Testing Transduction Efficiency and Specificity of Engineered Envelope Vectors

[0286] Nine candidate fusion factor pseudotype / envelope proteins were selected from the Serratus output. Multiple sequence alignments of the nine proteins to VSV-G were obtained using Clustalw. Pairwise sequence similarities ranged from a low of 25% to a high of 85%, indicating a wide range of envelope protein sequence diversity (Figure 18). This set of envelope proteins was formatted and synthesized as packaging plasmids (SEQ ID NOs: 16522-16530) to generate third-generation lentiviruses delivering a GFP reporter transgene with or without scFvs (non-viral membrane-bound proteins) directed against CD19, CD3, CD4, and CD8 for cell-type-specific tropism. The scFvs included the following sequences: CD19 (SEQ ID NO: 8130), anti-human CD4 (SEQ ID NO: 8135), and anti-human CD8 scFv (SEQ ID NO: 8137), and anti-CD3 antibody (SEQ ID NO: 8133). A VSV-G (SEQ ID NO: 16521) packaging plasmid was also used. A total of 50 lentiviral vectors were generated (10 envelope proteins × 5 targeting modalities = 50 lentiviral vectors). These vectors were used to transduce human PBMC samples, and 72 hours later, GFP signals were measured by flow cytometry for B cells (CD19+CD20+) and T cells (markers: CD3, CD4, CD8, TCR). For each lentivector, sensitivity was calculated using the formula: sensitivity = true positive / (true positive + false negative) (Figure 19). For each lentivector, specificity was calculated using the formula: specificity = true negative / (true negative + false positive) (Figure 20). An ideal lentivector would achieve high sensitivity and specificity. The data show that various combinations of targeting scFvs and viral envelope proteins have varying sensitivities and specificities, with some sensitivities exceeding 80% and some specificities exceeding 99%. 8. Equivalents

[0287] While several embodiments of the present invention have been described and illustrated herein, those skilled in the art will readily envision various other means and / or structures for performing the functions and / or obtaining one or more of the results and / or advantages described herein, and each such variation and / or modification is deemed to be within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and that the actual parameters, dimensions, materials, and / or configurations will depend on the specific application or applications for which the teachings of the present invention are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. Accordingly, the foregoing embodiments are presented by way of example only, and it should be understood that, within the scope of the appended claims and their equivalents, embodiments of the invention may be practiced otherwise than as specifically described and claimed. The inventive embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits and / or methods is included within the inventive scope of the present disclosure, if such features, systems, articles, materials, kits and / or methods are not mutually inconsistent.

[0288] All definitions defined and used herein should be understood to provide dictionary definitions, definitions in documents incorporated by reference herein, and / or ordinary meanings of the defined terms.

[0289] All references, patents, and patent applications disclosed herein are incorporated by reference with respect to the subject matter for which each is cited, and in some cases may include the entire document.

[0290] The indefinite articles "a" and "an," as used in this specification and claims, unless expressly indicated otherwise, should be understood to mean "at least one."

[0291] The phrase "and / or," as used in the specification and claims, should be understood to mean "either or both" of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with "and / or" should be construed in the same manner, i.e., "one or more" of the elements so conjoined. Other elements other than the elements specifically identified by the "and / or" clause may optionally be present, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to "A and / or B," when used in combination with open-ended language such as "comprising," can refer in one embodiment to A only (optionally including elements other than B); in another embodiment to B only (optionally including elements other than A); in yet another embodiment to both A and B (optionally including other elements); etc.

[0292] As used herein and in the claims, "or" should be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" shall be interpreted as being inclusive, i.e., including at least one of several elements or a list of elements, but also including more than one, and optionally including additional unlisted items. Only terms clearly indicated to the contrary, such as "only one of" or "exactly one of," or, when used in the claims, "consisting of," shall refer to the inclusion of exactly one element of several elements or a list of elements. In general, the term "or" as used herein shall only be interpreted as indicating exclusive alternatives (i.e., "one or the other, but not both") when preceded by terms of exclusivity, such as "either," "one of," "only one of," or "exactly one of." "Consisting essentially of," when used in the claims, shall have its ordinary meaning as used in the field of patent law.

[0293] As used herein and in the claims, the phrase "at least one" in connection with a list of one or more elements should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed in the list of elements, or excluding any combination of elements in the list of elements. This definition also allows for elements other than those specifically identified in the list of elements to which the phrase "at least one" refers, whether related or unrelated to those specifically identified elements, may optionally be present. Thus, as a non-limiting example, "at least one of A and B" (or, equivalently, "at least one of A or B," or, equivalently, "at least one of A and / or B") can refer in one embodiment to at least one, and optionally two or more, A, and no B (and optionally including elements other than B); in another embodiment to at least one, and optionally two or more, B, and no A (and optionally including elements other than A); and in yet another embodiment to at least one, and optionally two or more, A, and at least one, and optionally two or more, B (and optionally including other elements).

[0294] It is also to be understood that, unless expressly stated otherwise, in any method claimed herein that includes more than one step or action, the order of the method steps or actions is not necessarily limited to the order in which the method steps or actions are recited.

Claims

1. An manipulated envelope vector including a manipulated envelope, wherein the manipulated envelope is (a) A viral envelope protein having at least 90%, 95%, 96%, 97%, or 99% sequence identity with a sequence selected from SEQ ID NOs: 1-167 or SEQ ID NOs: 8154-16497, and (b) If necessary, (i) Signal peptide (S) as needed, (ii) Extracellular targeting domain (ETD), and (iii) Membrane-bound domain (MBD) Nonviral membrane-bound proteins including An manipulated envelope vector, including [the specified element].

2. The manipulated envelope vector according to claim 1, wherein the viral envelope protein has a sequence selected from SEQ ID NOs: 1 to 167 or SEQ ID NOs: 8154 to 16497.

3. The manipulated envelope vector according to claim 1, wherein the viral envelope protein is fused to one or more different molecules.

4. The manipulated envelope vector according to claim 3, wherein the one or more different molecules comprise a guide RNA and an exogenous endonuclease.

5. The manipulated envelope vector according to claim 1, comprising the nonviral membrane-bound protein.

6. The manipulated envelope vector according to claim 1, further comprising a nucleic acid construct encapsulated in the manipulated envelope.

7. The manipulated envelope vector according to claim 1, wherein the manipulated envelope further comprises a functional modulator protein having at least 90%, 95%, 96%, 97%, or 99% sequence identity with respect to a sequence selected from SEQ ID NOs. 168 to 8121.

8. The manipulated envelope vector according to claim 7, wherein the functional modulator protein has a sequence selected from SEQ ID NOs: 168 to 8121.

9. The manipulated envelope vector according to claim 1, wherein the viral envelope protein is derived from a retrovirus, and optionally the retrovirus is a lentivirus.

10. The manipulated envelope vector according to claim 1, comprising the nonviral membrane-bound protein, wherein the extracellular targeting domain (ETD) comprises a T cell receptor, an antibody, an MHC protein, or a variant thereof, and / or the extracellular targeting domain (ETD) comprises a targeting domain of a target cell-targeting protein, and the membrane-bound domain (MBD) comprises a membrane domain of the target cell-targeting protein.

11. The manipulated envelope vector according to claim 10, wherein the target cell-targeting protein has a sequence or fragment thereof selected from SEQ ID NOs. 168 to 8121.

12. The manipulated envelope vector according to claim 10, wherein the extracellular targeting domain (ETD) contains an antibody specific to CD3, CD4, CD5, CD7, CD8, CD19, CD20, or CD117.

13. The manipulated envelope vector according to claim 1, comprising the nonviral membrane-bound protein, wherein the nonviral membrane-bound protein further comprises an Fc domain and a linker located between the extracellular targeting domain (ETD) and the membrane-bound domain (MBD).

14. The manipulated envelope vector according to claim 1, wherein the viral envelope protein comprises at least one amino acid insertion, deletion, or substitution compared to a protein having a sequence selected from SEQ ID NOs: 1 to 167 or SEQ ID NOs: 8154 to 8497.

15. The engineered envelope vector according to claim 6, wherein the nucleic acid construct comprises a barcode sequence, a reporter protein coding sequence, a transgene coding sequence, and / or repressive RNA, catalytic RNA, or a sequence for CRISPR / Cas9 or other site-specific endonuclease-mediated mutagenesis, and optionally the transgene is a therapeutic gene.

16. A library of manipulated envelope vectors comprising 10, 100, 1,000, 10,000, 100,000, 1,000,000 or 10,000,000 unique clones of the manipulated envelope vector described in any one of claims 1 to 15.

17. The library according to claim 16, wherein each clone of the manipulated envelope vector comprises a unique nucleic acid barcode, a unique extracellular targeting domain (ETD), and / or a unique viral envelope protein.

18. A composition comprising an engineered envelope vector according to any one of claims 1 to 15 for use in a method for modifying target cells, or a library of engineered envelope vectors comprising 10, 100, 1,000, 10,000, 100,000, 1,000,000 or 10,000,000 unique clones of the engineered envelope vector according to any one of claims 1 to 15, wherein the method comprises the step of contacting the engineered envelope vector or the library with the target cells.

19. The composition or library according to claim 18, wherein the target cells are immune cells selected from T cells and Treg cells as needed, and / or the target cells are immune cells in a human subject.

20. The composition or library according to claim 18, wherein the manipulated envelope vector comprises a nucleic acid construct, which optionally comprises (i) a coding sequence for a reporter protein, (ii) a coding sequence for a therapeutic gene, (iii) an inhibitory RNA, (iv) a catalytic RNA, or (iii) a sequence for CRISPR / Cas9, CRISPR / Cas12a, or other site-specific endonuclease-mediated mutagenesis.

21. A composition for gene therapy, comprising an engineered envelope vector according to any one of claims 1 to 15, wherein the engineered envelope vector comprises a nucleic acid construct for gene therapy.

22. The composition according to claim 21, wherein the nucleic acid construct comprises (i) a coding sequence for a reporter protein, (ii) a coding sequence for a therapeutic gene, (iii) an inhibitory RNA, (iv) a catalytic RNA, or (iii) a sequence for CRISPR / Cas9, CRISPR / Cas12a, or other site-specific endonuclease-mediated mutagenesis, and / or the nucleic acid construct comprises a coding sequence for a T cell receptor (TCR) or a chimeric antigen receptor (CAR), and / or the nucleic acid construct comprises a coding sequence for an endogenous gene such as dystrophin.

23. The composition or library according to claim 18, characterized in that the manipulated envelope vector is administered by intravenous or subcutaneous administration.

24. The composition according to claim 21, characterized in that the manipulated envelope vector is administered by intravenous or subcutaneous administration.

25. A composition for use in a method comprising an operated envelope vector according to any one of claims 1 to 15, wherein the method is (a) A step of bringing a plurality of single cells, each containing a first nucleic acid barcode, into contact with the manipulated envelope vector, wherein the manipulated envelope vector contains a second nucleic acid barcode, and as a result, the manipulated envelope vector delivers the second nucleic acid barcode to the single cells; (b) A step of isolating each of the plurality of single cells into separate compartments, wherein each separate compartment is a microdroplet in an emulsion; (c) The step of obtaining a transcript from each of the single cells; (d) A step of reacting the transcripts from each of the single cells with a set of first probes and a set of second probes, wherein each of the first probes is configured to (i) bind to a first target polynucleotide comprising the first nucleic acid barcode and (ii) comprises a sequence of a non-human exogenous sequence or a sequence complementary thereto, and each of the second probes is configured to (iii) bind to a second target polynucleotide comprising the second nucleic acid barcode and (iv) comprises a sequence of the non-human exogenous sequence or a sequence complementary thereto; and (e) The step of performing reverse transcription and subsequent PCR amplification using the first set of probes and the second set of probes to generate a fusion complex containing the first nucleic acid barcode and the second nucleic acid barcode. A composition containing the following:

26. The composition according to claim 25, further comprising the step of sequencing the fusion complex.

27. ​​The composition according to claim 25, wherein each individual compartment has an average volume of 1 nanoliter (nL).

28. The composition according to claim 25, wherein step (b) further comprises introducing an mRNA capture agent and a cell lysate solution into the individual compartments.

29. The composition according to claim 25, wherein the method comprises the step of obtaining a sequence from at least 10,000 individual cells or at least 5,000 fusion complexes.

30. Step (e) is A step of linking the first nucleic acid barcode and the second nucleic acid barcode by performing an overlap extension reverse transcriptase polymerase chain reaction. The composition according to claim 25, comprising:

31. The composition according to claim 25, wherein the first nucleic acid barcode and the second nucleic acid barcode are different.

32. The composition according to claim 25, wherein step (a) comprises contacting the plurality of single cells with a library of manipulated envelope vectors containing 10, 100, 1,000, 10,000, 100,000, 1,000,000 or 10,000,000 unique clones of the manipulated envelope vector.

33. The composition according to claim 25, wherein the plurality of single cells comprises a library of 10, 100, 10,000, 100,000, 1,000,000 or more genetically distinct cells.

34. The composition according to claim 33, wherein the plurality of single cells comprises a library of 10, 100, 10,000, 100,000, and 1,000,000 cells, each containing a unique first nucleic acid barcode.

35. A composition for use in a method comprising an operated envelope vector according to any one of claims 1 to 15, wherein the method is (a) A step of bringing a plurality of single cells into contact with the manipulated envelope vector, wherein the manipulated envelope vector contains an exogenous nucleic acid, and as a result, the manipulated envelope vector delivers the exogenous nucleic acid to the single cells; (b) A step of isolating each of the plurality of single cells into separate compartments, each of which is a microdroplet in an emulsion, and each of the single cells contains a first nucleic acid barcode; (c) the step of obtaining a transcript from each of the single cells; and (d) Reverse transcription and subsequent PCR amplification to generate a fusion complex containing the first nucleic acid barcode and the exogenous nucleic acid. A composition containing the following:

36. The composition according to claim 35, wherein the exogenous nucleic acid encodes a reporter protein.

37. The composition according to claim 35, wherein the exogenous nucleic acid comprises a second nucleic acid barcode.

38. The composition according to claim 35, further comprising the step of sequencing the fusion complex.

39. The composition according to claim 35, wherein the first barcode sequence is attached to a bead.

40. The composition according to claim 35, wherein the exogenous nucleic acid comprises a coding sequence for a reporter protein, and the method further comprises the step of sorting, selecting, or isolating cells expressing the reporter protein.

41. The composition according to claim 35, wherein each individual compartment has an average volume of 1 nanoliter (nL).

42. The composition according to claim 35, wherein step (b) further comprises introducing an mRNA capture agent and a cell lysate solution into the individual compartments.

43. The composition according to claim 35, wherein the method comprises obtaining a sequence from at least 10,000 individual cells or at least 5,000 fusion complexes.

44. The composition according to claim 35, wherein step (a) comprises contacting the plurality of single cells with a library of manipulated envelope vectors containing 10, 100, 1,000, 10,000, 100,000, 1,000,000 or 10,000,000 unique clones of the manipulated envelope vector.

45. The composition according to claim 44, wherein the plurality of single cells include 10, 100, 10,000, 100,000, 1,000,000 or more unique cells.