Methods and compositions for high-throughput protein delivery, screening, and detection.
Peptide barcodes enable high-throughput detection and quantification of therapeutic agents in complex systems by associating with cargo polypeptides for nucleic acid sequencing, addressing the limitations of current techniques.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- MANIFOLD BIOTECHNOLOGIES INC
- Filing Date
- 2024-05-02
- Publication Date
- 2026-06-02
AI Technical Summary
Current techniques for evaluating agents, particularly in complex systems like in vivo environments, are time-consuming, costly, and limited by the availability of fluorogenic substrates, and DNA barcoding methods lack stability and immunogenicity, making it difficult to accurately detect and quantify multiple agents.
The use of peptide barcodes associated with cargo polypeptides, which can be sequenced without covalent bonding, allowing for high-throughput detection and quantification of multiple agents in complex systems using nucleic acid sequencing, enabling simultaneous measurement of multiple characteristics such as concentration and localization.
This approach provides accurate and efficient evaluation of therapeutic agents in complex environments by enabling high-throughput detection and quantification of multiple agents, overcoming limitations of existing methods like mass spectrometry and DNA barcoding.
Smart Images

Figure 2026517806000049 
Figure 2026517806000050 
Figure 2026517806000051
Abstract
Description
Technical Field
[0001] Cross - reference to related applications This application claims the benefit of U.S. Provisional Application No. 63 / 463,844, filed May 3, 2023; U.S. Provisional Application No. 63 / 598,007, filed Nov. 10, 2023; and U.S. Provisional Application No. 63 / 633,381, filed Apr. 12, 2024, the entire contents of which are hereby incorporated by reference in their entirety.
[0002] Sequence listing This specification refers to a sequence listing (submitted electronically as an.xml file named "2013703 - 0026_ST26.xml" on May 2, 2024). This.xml file was created on Oct. 18, 2022 and is 12,323,567 bytes in size. The entire contents of the sequence listing are hereby incorporated by reference in this specification.
Background Art
[0003] The evaluation of agents (e.g., cargo components encoding cargo polypeptides, delivery particles, etc.) is central to much of molecular biology and pharmaceutical biology.
Summary of the Invention
[0004] This disclosure provides insights and techniques for achieving improved or desirable evaluation of agents (e.g., nucleic acids that contain and / or encode cargo agents (e.g., one or more nucleic acid sequence components, e.g., cargo polypeptide agents), therapeutic agents, and / or in some embodiments, delivery particles that contain and / or encode such nucleic acids that contain cargo agents and / or therapeutic agents).
[0005] In particular, the present disclosure understands that many current techniques for assessing and particularly identifying the presence and / or abundance of one or more target agents may typically rely on mass spectrometry and / or affinity (e.g., immunoaffinity) detection. The present disclosure understands that many available affinity detection techniques are time-consuming and / or costly to implement or perform, that many such techniques need to be performed one at a time, and that many are limited by, for example, the availability of fluorogenic substrates (e.g., it can be evaluated by related techniques such as optical microscopy).
[0006] The present disclosure further understands that certain other techniques that may be used to evaluate target agents, such as, for example, DNA barcoding techniques, may also have drawbacks. For example, DNA barcodes may lack stability and / or exhibit undesirable immunogenicity when used, for example, in vivo. The present disclosure understands that such techniques may therefore encounter problems, particularly when evaluating agents (e.g., cargo agents) in complex environments (e.g., in vivo).
[0007] In particular, the present disclosure encompasses an appreciation of the causes of certain problems with available techniques commonly used to evaluate target agents and particularly to evaluate target cargo agents and / or delivery particles. In particular, the present disclosure identifies the causes of certain problems (e.g., the inability to detect and / or measure quantities such as concentration, e.g., the inability to detect and / or quantify functional information) faced by such techniques for the evaluation of multiple agents, particularly when such agents are present in complex systems (e.g., in complex solutions and / or in vivo).
[0008] Furthermore, the present disclosure provides certain techniques that, in some embodiments, achieve such evaluations with surprisingly high accuracy. Those skilled in the art will understand that there are several situations where the detection and / or measurement (e.g., of an accurate amount) of multiple agents within a complex system is desirable. Furthermore, those skilled in the art will understand the advantages of high accuracy in many such situations.
[0009] In particular, the present disclosure provides techniques for detecting and / or measuring (e.g., with high accuracy and / or otherwise precise measurement) one or more, and in some embodiments, multiple drugs (e.g., nucleic acids containing and / or encoding cargo drugs, e.g., delivery particles containing and / or encoding cargo drugs) in complex systems (e.g., in vivo). In some embodiments, the detected drug(s) may be, or include, a cargo drug and / or its form (e.g., aggregated, complex, covalently, modified, e.g., by the formation of disulfide bonds, glycosylated, pegylated, phosphorylated, truncated, e.g., by proteolytic cleavage, etc.).
[0010] In some embodiments, the detected drug(s) may be delivered under various conditions (e.g., physiological conditions, e.g., target tissue) via delivery particles (e.g., target virus particles, virus-like particles, lipid-based particles, polymer-based particles, bead-based, metal-based, or polysaccharide-based particles, e.g., of the same and / or different types). In some embodiments, nucleic acids are placed within the delivery particles. In some embodiments, the nucleic acid sequence comprises (a) a cargo component encoding a cargo polypeptide, and (b) a barcode component, the cargo component being operably linked to the barcode component. In some embodiments, the cargo component includes other types of cargo as described herein. In some embodiments, detection of the cargo component associated with the barcode component may then be used to evaluate and / or quantify the phenotype (e.g., directivity, etc.) of the target delivery particle.
[0011] In some embodiments, the provided technology is particularly useful or effective for evaluating therapeutic agents. For example, in some embodiments, the provided technology may be particularly useful for evaluating one or more characteristics (e.g., properties (e.g., concentration, localization, persistence, affinity, etc.)) of a drug(s) of interest. In some such embodiments, the drug(s) in question may be characterized by one or more attributes that are appropriate or desirable for therapeutic use. For example, in some embodiments, the provided technology may be used to screen for promising therapeutic agents (e.g., polypeptide entities) for one or more characteristics (e.g., properties, attributes) that are suitable for therapeutic use. In some embodiments, the characteristics of a promising therapeutic agent(s) may be measured one at a time. In some embodiments, two or more characteristics of a promising therapeutic agent(s) may be measured simultaneously. For example, in some embodiments, one or more therapeutic agents may be screened for affinity to a target drug, but other desirable properties, such as molecular stability in a physiologically relevant environment, are still unknown. In some embodiments, for example, one or more therapeutic agents may be screened for affinity to a target drug and simultaneously for other desirable properties, such as molecular stability in a physiologically relevant environment.
[0012] This disclosure acknowledges that many current methods for polypeptide measurement, such as Western blotting or ELISA, rely on measuring the abundance of light at a specific wavelength, or the overall emission (Towbin 1979, Engvall 1972). Due to the limitations of visible light wavelengths, these methods measure only a small number of different polypeptides, often fewer than four, in a single reaction (Elshal, 2006). This disclosure acknowledges that many applications, including drug discovery, would benefit from (and in some cases require) dramatically higher throughput.
[0013] This disclosure further acknowledges the development of nucleic acid sequencing technologies (e.g., DNA sequencing technologies) that enable the analysis of billions of individual DNA molecules in a single experiment (Shendure, 2005). Various strategies have been developed to apply this high throughput achievable with nucleic acid sequencing technologies to protein detection and measurement, particularly by tagging proteins with accompanying DNA fragments (commonly referred to as "DNA barcodes"). These DNA fragments can then be sequenced to indirectly detect the protein (Trads, 2017) or one or more features of the protein.
[0014] This disclosure acknowledges the ability to apply high-throughput nucleic acid sequencing technology to the evaluation of other drugs, and in particular cargo drugs, but also identifies the root causes of certain problems associated with many approaches used to test polypeptides by DNA barcoding. For example, this disclosure acknowledges that protein modification by DNA barcoding can often alter its function (Trads, 2017), which may negate the purpose of using DNA barcoding to evaluate polyptides.
[0015] Known techniques for quantifying or screening multiple therapeutic components are described by WO2020097254 (Gordian Biotechnology). However, this disclosure identifies the root causes of problems associated with such approaches and further offers certain advantages compared to them, including the ability to directly evaluate and / or quantify cargo polypeptides. Approaches such as those described by Gordian Biotechnology fail to exhibit such characteristics and rely on cell-based analysis. Furthermore, as can be understood by those skilled in the art, reading this disclosure reveals that cell-based analysis is limited by experimental complexity, the number of outputs, requires additional enrichment steps, and is costly. In comparison, this disclosure is not limited by such disadvantages, as nucleic acid sequences encoding one or more polypeptide binders associated with one or more peptide barcodes can be sequenced and measured to quantify cargo components without requiring additional analysis and / or enrichment steps.
[0016] Known techniques for quantifying barcoded cargo polypeptides include those described by Egloff et al. (2019), which use mass spectrometry to determine the presence or absence of protein sequences in a mixture (Egloff 2019). However, this disclosure identifies the root causes of problems associated with such approaches and, further, offers certain advantages related thereto, for example, by using nucleic acids (e.g., DNA) for amplification of the original signal. Approaches such as those described by Egloff et al. cannot include (or benefit from) such features. Furthermore, as will be understood by those skilled in the art, reading this disclosure reveals that mass spectrometry only reads the mass-to-charge ratio of the associated sequences, and consequently, methods using mass spectrometry are limited in their total throughput because different sequences may have the same mass-to-charge ratio. In comparison, the present invention is not limited by such disadvantages because nucleic acid sequences that associate with one or more binders and then associate with each barcode are sequenced and measured to identify and quantify the barcoded cargo polypeptide.
[0017] Other techniques available in the art use antibodies presented on phages (Fab phages) to identify the presence of endogenous proteins expressed on the cell surface (Pollock, 2018). In such methods, one Fab phage is generated for each endogenous protein (i.e., the target protein to be evaluated), and no barcodes are used. In contrast, the present technique envisions the use of a generalizable, manipulated barcode sequence such that the barcode sequence is used to label any protein, whether endogenous or exogenous to the context in which it is applied, and that each barcode, and thus each barcoded cargo polypeptide, can subsequently be measured using one or more binders to which it specifically associates (i.e., the “barcode fingerprint” as described elsewhere in this disclosure). The complex association of one or more binders and the barcode is then measured, and the precise quantification of the associated protein is achieved, for example, using a composite algorithm (i.e., the “decoding” as described elsewhere in this disclosure).
[0018] This disclosure recognizes the ability of an antigen presented to a phage to determine the epitope of an antibody in the blood to which the phage can bind (Mohan, 2018). However, this method cannot determine the sequence of the antibody to which the antigen binds, and therefore provides only limited information about any antibody that specifically binds to the antigen presented to the phage. However, this disclosure provides systems, compositions, and methods that offer the advantage of using common barcodes with known affinity to one or more binders or conjugates, which can be used to tag any target(s) of interest in complex mixtures, including but not limited to blood, to identify and quantify such targets(s).
[0019] This disclosure provides, in particular, a technique that enables the evaluation (e.g., detection and / or quantification) of multiple drugs (e.g., multiple nucleic acids containing and / or encoding cargo drugs, multiple delivery particles containing nucleic acids encoding cargo drugs, and / or combinations thereof) within a pool of such drugs using DNA sequencing, without requiring (direct or indirect) covalent bonding between the DNA and the drug to be evaluated, or any other form of constraint on the drug to be evaluated.
[0020] This specification describes peptide barcodes (also known as "barcodes") and techniques for producing and / or using them. In some embodiments, barcodes are used to label cargo. In particular, such approaches can enable integrated measurement of cargo without modifying non-polypeptide identifiers. In some embodiments, the peptide barcode is an amino acid polypeptide sequence. In some embodiments, the peptide barcode is contained within the cargo (e.g., endogenous to the polypeptide to be measured (e.g., an antibody, e.g., a cell surface antigen)). In some embodiments, the peptide barcode is not contained with the cargo polypeptide to be measured (e.g., an antibody, e.g., a cell surface antigen) (e.g., exogenous to the cargo polypeptide to be measured). In some embodiments, the barcode is, for example, a sequence contained within the cargo (e.g., the cargo to be measured) (e.g., a designed sequence). In some embodiments, the cargo contains a cargo polypeptide. In some embodiments, the barcode associates with the N-terminus of the cargo polypeptide (e.g., the cargo polypeptide to be measured) (e.g., binds to it (e.g., covalently)). In some embodiments, the barcode associates with the C-terminus of the cargo polypeptide (e.g., the cargo polypeptide to be measured) (e.g., binds to it covalently). In some embodiments, the barcode associates with the proximal N-terminus of the cargo polypeptide (e.g., the cargo polypeptide to be measured) (e.g., binds to it covalently). In some embodiments, the barcode associates with the proximal C-terminus of the cargo polypeptide (e.g., the cargo polypeptide to be measured) (e.g., binds to it covalently).
[0021] The methods disclosed herein may use peptide barcodes designed to have varying lengths. In some embodiments, the peptide barcode may have amino acid lengths ranging from 1 to 100, 5 to 50, 8 to 25, 9 to 25, or 9 to 15. In some embodiments, the peptide barcode may have a length of at least 25 amino acids. In some embodiments, the peptide barcode may have a length of up to 8 amino acids. In some embodiments, the peptide barcode may have a length of 10 amino acids.
[0022] The barcode sequences described herein can be reused to quantify different drugs (e.g., nucleic acids containing and / or encoding the cargo of interest, e.g., delivery particles of interest, each containing a nucleic acid encoding the cargo of interest) or mixtures of drugs (a mixture of cargoes of interest to be measured and / or a mixture of delivery particles to be measured, e.g., via the detection of barcoded cargoes). In some embodiments, the barcodes are generated so that they can be easily reused in different experiments between several different drugs (e.g., nucleic acids containing and / or encoding the cargo of interest, e.g., delivery particles of interest, each containing a nucleic acid encoding the cargo of interest).
[0023] In particular, the barcodes described herein are designed to be distinct from each other (e.g., unique). In some embodiments, barcodes are designed to have different sequences (e.g., different sequences from another barcode). For example, each barcode is designed to be distinct from (e.g., unique to) any other barcode used in the experiment, such that each drug (e.g., nucleic acid containing the cargo to be measured, e.g., delivery particles containing nucleic acid containing the cargo to be measured) associates with at least one barcode, and each barcode (e.g., a barcode with a specific sequence) associates with only one cargo. As can be understood by those skilled in the art, the diversity of barcodes contained in the pool is limited only by the possible diversity of amino acid sequences for a given barcode length. For example, for a barcode length "N", 20 of length N NThere are different amino acid barcode sequences.
[0024] The methods described herein relate to the detection of one or more barcodes using a binder. In some embodiments, the barcodes are associated with or in contact with a binder containing a detectable nucleic acid. For example, in some embodiments, the binder may be or may contain a phage, ribosome, mRNA, DNA, etc. In some embodiments, the binder is a phage having a binding motif (e.g., a polypeptide binder described herein) on its surface. In some embodiments, the binder contains a detectable nucleic acid. In some embodiments, the binder expresses the detectable nucleic acid. In some embodiments, the binder expresses the detectable nucleic acid on (e.g., on the surface of) the binder (e.g., a binder). In some embodiments, the binder is a polypeptide. In some embodiments, the binder associates with the barcode (e.g., with known specificity and affinity). In some embodiments, the binder associates with one or more barcodes (e.g., with different known specificity and affinity). In some embodiments, the binder is an antibody (e.g., expressed on the surface of the binder). In some embodiments, for example, to detect the presence of a specific (e.g., different) barcode, the present disclosure assumes the association of different detectable nucleic acids (e.g., DNA sequences, RNA sequences, etc.) with a specific barcode. This is achieved through contact of a binder (e.g., on its surface) which may be expressed with a binder containing different detectable nucleic acids.
[0025] This specification describes binders. In some embodiments, the binder is a polypeptide. In some embodiments, for example, the binder is produced to have known specificity and affinity to a given barcode. In some embodiments, the binder is produced to have known specificity and affinity to one barcode. In some embodiments, the binder is produced to have known specificity and affinity to multiple (e.g., two or more, three or more, etc.) barcodes. In some embodiments, the binder is produced to have known specificity and affinity to at least one barcode. In some embodiments, the binder is expressed on the surface of a binder (e.g., a phage, a ribosome, etc.) using, for example, a method known to those skilled in the art.
[0026] In particular, the systems and methods described herein, for example, identify the advantages of nucleic acid sequencing technologies and apply them effectively to protein detection and measurement methods. For example, the methods described herein may use several binders having known specificity and affinity for different barcodes, which can be expressed on a binder and mixed in a single pool. After mixing with a pool of barcoded cargoes (i.e., cargo polypeptides, each associated with a barcode described herein), the binders expressed on the binder bind to any given barcode in the pool, which has known but different affinities. The spectrum of binder affinity to various barcodes is referred to herein as the “binder fingerprint.” Conversely, a barcode can bind to any given binder in the binder pool, which has known but different affinities. The spectrum of barcode affinity to various binders is referred to herein as the “barcode fingerprint.” Therefore, the presence of a specific barcoded cargo can be detected, for example, by extracting and sequencing the associated nucleic acid (e.g., detectable nucleic acid (e.g., DNA sequence, RNA sequence, etc.)) of a group of binders (e.g., phages) bound to the barcode associated with the cargo in a complex solution.
[0027] Other methods have been developed that use binders to identify polypeptide sequences. However, these methods face several challenges, including difficulties in generating and characterizing binders, and in effectively decoding their bindings to specifically identify polypeptides. Another limitation associated with binders developed in the past is their nonspecific binding, which negatively impacts detection accuracy due to a poor signal-to-noise ratio. In contrast, this technique generates a large number of binders rapidly (e.g., in about 1 week, 2 weeks, 3 weeks, 4 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, or 1 year). In some embodiments, for example, about 100 to 1000 binders can be rapidly generated. In some embodiments, about 10 to 1000 binders can be rapidly generated. In some embodiments, about 10 to 10,000 binders can be rapidly generated. In some embodiments, at least about 10,000 binders can be rapidly generated.
[0028] The binders described herein are robust. The binders can be bound to barcodes (e.g., with robust affinity to one or more barcodes) under a variety of conditions and / or environments as described herein. For example, the binders described herein can be bound to barcodes (e.g., with robust affinity to one or more barcodes) in a variety of composite environments (e.g., in blood, tissue, serum, plasma, etc.). Therefore, the binders of this disclosure can be used to detect targets (e.g., nucleic acids encoding a cargo of interest) under a variety of conditions (e.g., physiological conditions, e.g., target tissue of interest). Furthermore, the binders of this disclosure can also be used to detect delivery particles (e.g., virus particles of interest, virus-like particles, lipid-based particles, polymer-based particles, bead-based, or polysaccharide-based particles) under a variety of conditions (e.g., physiological conditions, e.g., target tissue of interest).
[0029] Similarly, barcodes can be generated quickly and robustly as described herein. In some embodiments, the barcodes described herein are specific to the binder described herein. In some embodiments, for example, about 100 to about 2000 barcodes can be generated quickly (e.g., in about 1 week, about 2 weeks, about 3 weeks, about 4 weeks, about 1 month, about 2 months, about 3 months, about 4 months, about 5 months, about 6 months, or about 1 year). In some embodiments, about 10 to about 1000 barcodes can be generated quickly. In some embodiments, about 10 to about 10,000 barcodes can be generated quickly. In some embodiments, at least about 10,000 barcodes can be generated quickly.
[0030] The barcodes described herein are robust. The barcodes can be bound to binders (e.g., with robust affinity to one or more binders) under a variety of conditions and / or environments, as described herein. For example, the barcodes described herein can be bound to binders (e.g., with robust affinity to one or more binders) in a variety of combined environments (e.g., in blood, tissue, serum, plasma, etc.). Therefore, the barcodes of this disclosure can be used to detect targets (e.g., drugs of interest) under a variety of conditions (e.g., physiological conditions). Accordingly, this disclosure corrects the disadvantages and shortcomings of existing methods (e.g., nonspecific binding, variable binding in different environments, etc.) by rapidly generating a large number of robust binders and barcodes that can be used in combination with the computational methods described herein (e.g., deconvolution methods) to enable specific and well-characterized binder-barcode binding / association and accurate detection methods.
[0031] The disclosure also envisions the ability to modify one or more peptide barcode sequences so that they are readily distinguishable from each other and / or from potential background protein sequences. Similarly, the disclosure also envisions the ability to modify one or more polypeptide binder sequences so that they are readily distinguishable from each other and / or from potential background protein sequences.
[0032] In particular, the present invention as described herein provides a method for testing "n" different candidate proteins, where n is 1 or more, in a single assay or animal model. In some embodiments, the candidate proteins are therapeutic candidate proteins. In some embodiments, multiple candidate proteins are designed, and each different candidate protein associates with its own unique peptide barcode as described herein. Such barcoding has many advantages, including, but not limited to, cost-effective and time-efficient injection of all candidate proteins in an assay and / or a single injection into an animal. Samples (e.g., tissue samples, serum samples, blood samples, extracellular samples, single-cell samples, etc.) are then obtained from the injected animals, and the barcodes can be extracted. In some embodiments, such extracted barcodes provide a measure of the relative abundance of the initially injected candidate protein. For example, one or more extracted barcodes may be identified by contacting them with a pool of binders (e.g., expressed on a binder) known to bind to the barcode initially bound to the candidate protein. After the barcode and binder are bound, the bound binder (e.g., phage) is selected, and their detectable nucleic acids (e.g., DNA sequences, RNA sequences, etc.) are extracted. In some embodiments, the extracted nucleic acids are subjected to sequencing (e.g., next-generation sequencing). The sequenced nucleic acids can then be used to identify one or more barcodes to which they are designed to bind. This, along with previously established information regarding binding affinities between various binder-barcode pairs, can be used to identify and determine the relative abundance of each protein initially injected.
[0033] Also described herein are methods used, for example, to convert nucleic acid counts obtained from sequencing experiments into relative or absolute protein quantifications. In some embodiments, nucleic acid sequences are counted and translated in silico into protein sequences. As described herein, nucleic acid sequences correspond to binder sequences whose affinity is established and characterized for each barcode given in the pool. In some embodiments, binder counts are compared to a database where the tendency to bind to a single barcode is known. In some embodiments, binder counts are compared to a database where the tendency to bind to multiple (e.g., two or more, three or more, etc.) barcodes is known. In some embodiments, for example, in a sequencing experiment, the relative proportion of binder counts is directly compared to identify the relative proportion of barcodes and / or proteins associated with the barcodes. In some embodiments, as may be known to those skilled in the art, sequences (e.g., control sequences or accessory sequences) of known abundance (e.g., counts, quantifications, concentrations, etc.) are used to determine the absolute abundance (e.g., counts, quantifications, concentrations, etc.) of a given binder(s) (e.g., added to a sequencing experiment). This can be used to estimate the absolute abundance (e.g., counts, quantifications, concentrations, etc.) of barcodes(s) and / or proteins(s) associated with the barcodes(s) using either direct counting or one of the linear models described herein.
[0034] In some embodiments, the nucleic acid comprises a cargo component encoding a cargo polypeptide. In some embodiments, the cargo polypeptide is or comprises a therapeutic polypeptide. In some embodiments, the cargo component further comprises one or more sequence elements. In some embodiments, the cargo component associates (e.g., is operably linked) with a nucleotide sequence encoding a barcode as described herein.
[0035] In particular, this disclosure provides methods for evaluating barcodes, binders (e.g., conjugates (e.g., surface-expressing binders)), and cargo (e.g., barcoded cargo (e.g., barcoded cargo polypeptides)) as described herein. In some embodiments, one method includes subjecting a population of barcoded cargo (e.g., barcoded cargo polypeptides) to evaluation, separating members of the population that satisfy the evaluation from those that do not, identifying either a positive population or a negative population, or both, contacting the positive population, the negative population, or each population separately from the other, with a set of binders containing at least one specific binder specific to each barcode in the population, and identifying the binders that bind to the separated members, thereby identifying the barcoded cargo (e.g., barcoded cargo polypeptides) present in the contacted population(s).
[0036] This disclosure provides a method comprising contacting a set of binders separately with either a first group, a second group, or each of the first and second groups of barcoded cargo (e.g., barcoded cargo polypeptides), and identifying the binders of the set that bind to members of the first group, the second group, or both, thereby identifying the barcoded cargo (e.g., barcoded cargo polypeptides) present in the contacted group(s). In some embodiments, each binder binds specifically to one or more barcodes (e.g., with known affinity). In some embodiments, the set of binders collectively comprises at least one binder specific to each barcode in the first and second groups. In some embodiments, the first and second groups are separated from each other based on performance in evaluation.
[0037] In some embodiments, the method further includes identifying differences between the first and second groups and identifying the functional effects of performance evaluation. In some embodiments, the method includes separating the binder bound to at least one cargo (e.g., barcoded cargo (e.g., barcoded cargo polypeptide)).
[0038] In some embodiments, the identification step includes quantifying the number of binders bound to the barcoded cargo (e.g., barcoded cargo polypeptide). In some embodiments, quantification may be performed by decoding the nucleotide sequence of each binder bound to the barcoded cargo (e.g., barcoded cargo polypeptide). In some embodiments, quantifying the number of binders bound to the cargo (e.g., barcoded cargo polypeptide) provides a measure of the cargo (e.g., protein) in the population.
[0039] In some embodiments, the identification step includes amplifying the nucleic acid of the bound phage particle. In some embodiments, the identification step includes identifying the nucleotide sequence of the amplified nucleic acid. In some embodiments, one or more of the identified nucleotide sequences correspond to the coding sequence of the binder. In some embodiments, the identification step includes detecting one or more cargoes (e.g., proteins) from a population of barcoded cargoes (e.g., barcoded cargo polypeptides) using the identified sequence(s) of the binder coding sequence. In some embodiments, the identification step includes identifying one or more barcoded cargoes (e.g., barcoded cargo polypeptides) as therapeutic agents or targets for treating a disease, disorder, or condition.
[0040] In some embodiments, the specified step includes performing one or more of amplification, amplification, and sequencing (e.g., amplification, amplification, and / or sequencing of nucleic acids (e.g., DNA, RNA)). In some embodiments, amplification may be performed using one or more of the following known techniques: polymerase chain reaction (PCR), loop-mediated isothermal amplification (LAMP), rolling circle amplification (RCA), or similar. In some embodiments, sequencing may be performed using one or more of the following known techniques: Illumina, next-generation sequencing (NGS), nanopore sequencing, Pac Bio long-read sequencing, or similar.
[0041] In some embodiments, the separation step includes purifying one or more barcoded cargoes (e.g., barcoded cargo polypeptides) from the sample. In some embodiments, the barcoded cargoes (e.g., barcoded cargo polypeptides) are purified from a composite sample. In some embodiments, the barcoded cargoes (e.g., barcoded cargo polypeptides) are purified from a composite mixture. In some embodiments, the barcoded cargoes (e.g., barcoded cargo polypeptides) are purified using affinity purification methods (e.g., FLAG IP, Protein G / A) or protein precipitation methods.
[0042] In some embodiments, the method further includes injecting a population of barcoded cargo into an animal. In some embodiments, the method further includes injecting a population of barcoded cargo (e.g., barcoded cargo polypeptide) into an animal. In some embodiments, each barcode is bound to a specific binder expressed on a phage. In some embodiments, the method further includes taking a sample from the animal and subjecting it to evaluation.
[0043] In some embodiments, the method described herein includes identifying the relative amount of each binder present in the sample, thereby identifying a subset of the injected population of barcoded cargo (e.g., barcoded cargo polypeptide) present in the sample. In some embodiments, the method described herein includes identifying the absolute amount of each binder present in the sample by comparing the relative amount to a standard of known concentration.
[0044] In some embodiments, the methods described herein optionally include repeating one or more of the steps of the methods described herein using a subset of identified cargo (e.g., proteins).
[0045] In some embodiments, the methods described herein include identifying one or more cargoes (e.g., cargo polypeptides) as therapeutic agents or targets for treating a disease, disorder, or condition.
[0046] In some embodiments, the methods described herein include identifying one or more delivery particles as a therapeutic agent or target for treating a disease, disorder, or condition.
[0047] In some embodiments, the methods described herein include removing any unassociated (e.g., unbound) binders. In some embodiments, removal may be carried out by washing.
[0048] In some embodiments, the barcoded cargo (e.g., barcoded cargo polypeptide) is present in a sample. In some embodiments, the barcoded cargo (e.g., barcoded cargo polypeptide) is present in a composite sample. In some embodiments, the barcoded cargo (e.g., barcoded cargo polypeptide) is present in a composite mixture. In some embodiments, the barcoded cargo (e.g., barcoded cargo polypeptide) is present in a purified sample. In some embodiments, the barcoded cargo (e.g., barcoded cargo polypeptide) is present in a mammal.
[0049] In some embodiments, the delivery particles (e.g., including nucleic acids described herein) are present in the sample. In some embodiments, the delivery particles (e.g., including nucleic acids described herein) are present in the composite sample. In some embodiments, the delivery particles (e.g., including nucleic acids described herein) are present in the composite mixture. In some embodiments, the barcoded cargo (e.g., barcoded cargo polypeptide) is present in the purified sample. In some embodiments, the barcoded cargo (e.g., barcoded cargo polypeptide) is present in the mammal.
[0050] In some embodiments, the sample is or includes one or more of serum, blood, tissue, or tumor. In some embodiments, the sample is a control (e.g., a positive control or a negative control). In some embodiments, the sample is or includes cells or a population of cells.
[0051] In some embodiments, the sample is a composite sample. In some embodiments, the composite sample is or contains tissue. In some embodiments, the composite sample is or contains blood. In some embodiments, the composite sample is a composite mixture. In some embodiments, the composite sample is or contains one or more of serum, blood, or tissue.
[0052] In some embodiments, the barcode is one or more amino acids. In some embodiments, the barcode is contained within the complementarity-determining region (CDR) of the cargo (e.g., protein). In some embodiments, the barcode is synthetic. In some embodiments, the barcode is 1-100, 5-50, 8-25, 9-25, or 9-15 amino acid long. In some embodiments, the barcode is 10 amino acid long. In some embodiments, the barcode has a relatively small or no effect on the function of the cargo (e.g., protein). In some embodiments, the barcode does not induce an immune response. In some embodiments, the barcodes are orthogonal to each other. In some embodiments, at least one barcode is linked to the polypeptide of interest (e.g., polypeptide binder, cargo).
[0053] In some embodiments, the barcode is attached to the cargo (e.g., cargo polypeptide). In some embodiments, the barcode is attached to an appropriate location on the cargo (e.g., cargo polypeptide). In some embodiments, the appropriate location is the N-terminus or the C-terminus.
[0054] In some embodiments, the binder is a binding site presented on the phage, or includes such a site. In some embodiments, each binder in a set of binders is expressed on the phage. In some embodiments, the binder is expressed on the surface of the phage particle.
[0055] In some embodiments, the phage is selected from the group consisting of M13, T4, T7, lambda, and filamentous phages. In some embodiments, the phage is M13.
[0056] This disclosure provides nucleic acids, in particular, whose nucleotide sequence is a sequence encoding a peptide barcode, or a nucleic acid comprising such a sequence. In some embodiments, the peptide barcode has an amino acid length in the range of 1 to 100, 5 to 50, 8 to 25, 9 to 25, or 9 to 15. In some embodiments, the peptide barcode has an amino acid length of 8 to 25. In some embodiments, the peptide barcode has an amino acid length of 10. In some embodiments, the peptide barcode has been confirmed to bind specifically to a particular group of polypeptide binders in a set of binders.
[0057] In some embodiments, the peptide barcode has an amino acid sequence selected from the group consisting of SEQ ID NOs. 5347 to 8398. In some embodiments, the coding sequence is selected from the group consisting of SEQ ID NOs. 1148 to 4199.
[0058] This disclosure provides a library comprising multiple nucleic acids. In some embodiments, the multiple nucleic acids together encode, among other things, a set of peptide barcodes. In some embodiments, each nucleic acid comprises, in 5' to 3' or 3' to 5' order, one or more of the following: a) a first invariant sequence (e.g., a linker sequence or cargo sequence), b) a variant sequence that is at least 9 nucleotides long, and c) a second invariant sequence (e.g., a linker sequence, a stop codon, or a cargo sequence).
[0059] In some embodiments, the variant sequence is at least 15, 24, 27, 45, 150, or 300 nucleotides long.
[0060] In some embodiments, the library further includes one or more of the following: d) sequences that encode one or more short helix motifs; e) sequences that encode one or more mutated motifs; and f) invariant sequences that concatenate sequences (e.g., barcode components) to cargo (e.g., cargo components).
[0061] In some embodiments, each peptide barcode in the collection specifically binds to a particular group of polypeptide binders within the set of binders. In some embodiments, each peptide barcode in the collection specifically binds to one or more polypeptide binders within the set of binders.
[0062] This disclosure provides nucleic acids whose nucleotide sequence is a sequence encoding a polypeptide binder moiety, or a nucleic acid containing such a sequence. In some embodiments, the polypeptide binder moiety has an amino acid length in the range of 10 to 400. In some embodiments, the polypeptide binder moiety has been shown to specifically bind to a particular group of peptide barcodes within a collection of barcodes.
[0063] In some embodiments, the polypeptide binder portion has an amino acid sequence selected from the group consisting of SEQ ID NOs: 4200 to 5346. In some embodiments, the coding sequence is selected from the group consisting of SEQ ID NOs: 1 to 1147.
[0064] This disclosure provides a library comprising multiple nucleic acids. In some embodiments, multiple nucleic acids together encode a set of polypeptide binder moieties. In some embodiments, each nucleic acid comprises, in 5' to 3' or 3' to 5' order, a) a first invariant sequence (e.g., an antibody germline sequence (e.g., IGHV / IGKV)), b) a first variant sequence at least 10 nucleotides long (e.g., a CDR (e.g., a CDR3) sequence), and c) a second invariant sequence (e.g., an antibody germline sequence (e.g., IDHJ / IGKJ)).
[0065] In some embodiments, each nucleic acid further comprises one or more of the following: d) a stop codon (e.g., after a second invariant sequence), e) a linker sequence, f) a third invariant sequence (e.g., an antibody germline sequence (e.g., IGHV / IGKV)), g) a second variant sequence at least 10 nucleotides long (e.g., a CDR (e.g., a CDR3) sequence), and h) a fourth invariant sequence (e.g., an antibody germline sequence (e.g., IDHJ / IGKJ)).
[0066] This disclosure provides, in particular, a library of phage particles, each of which comprises one or more nucleic acids as described herein.
[0067] In some embodiments, the phage is selected from the group consisting of M13, T4, T7, lambda, and filamentous phages. In some embodiments, the phage is M13.
[0068] This disclosure provides a set of barcodes and binders. In some embodiments, each barcode is a peptide having a length of 1 to 100, 5 to 50, 8 to 25, 9 to 25, or 9 to 15 amino acids that specifically binds to a particular group of binders in the set. In some embodiments, each binder is a polypeptide that specifically binds to at least one barcode in the set.
[0069] In some embodiments, specific binding is observed when the binder is expressed on the phage and comes into contact with the barcode. In some embodiments, each binder is expressed on the phage.
[0070] This disclosure provides a kit comprising a set of binders, each being a polypeptide that specifically binds to at least a specific peptide barcode in a collection of barcodes. In some embodiments, each binder is provided as a polypeptide, a nucleic acid encoding a polypeptide, or both. In some embodiments, one or more of the binders are provided as phage particles or an aggregate thereof, engineered to express the binder. In some embodiments, one or more of the binders are provided as nucleic acids in a phagemide vector, or as an insert suitable for cloning into a phage vector.
[0071] In some embodiments, the kit further includes information specifying a peptide barcode for each binder. In some embodiments, each binder has been confirmed to bind to at least one specific peptide barcode within a set of barcodes, each of which binds specifically to at least one binder in the set.
[0072] In some embodiments, the kit further includes a set of instructions for performing sequencing of one or more phage particles coupled to one or more barcodes. In some embodiments, the kit further includes a computer-readable program for decoding the sequencing data. In some embodiments, the kit further includes reagents for expressing a binder on the phage particles.
[0073] In some embodiments, the kit includes nucleic acids that encode one or more barcodes. In some embodiments, the kit includes nucleic acids that encode one or more binders.
[0074] This disclosure provides a method for pharmacokinetic screening. In some embodiments, the method includes injecting a set of barcoded therapeutic candidate cargo polypeptides into an animal. In some embodiments, each barcoded therapeutic candidate protein contains a specific peptide barcode. In some embodiments, the method includes taking a sample from an animal, purifying one or more barcoded therapeutic candidate cargo polypeptides from the sample, contacting the sample with a set of binders (e.g., binders expressing the binders) containing at least one specific binder specific to each barcode in the sample, and determining the relative amount of each binder present in the sample to determine the pharmacokinetic properties or biodistribution of each barcoded therapeutic candidate protein.
[0075] In some embodiments, the purified proteins may be a subset of barcoded therapeutic candidate cargo polypeptides administered to animals.
[0076] In some embodiments, multiple samples may be collected from an animal.
[0077] In some embodiments, the animal is a mammal. In some embodiments, the animal is a human. In some embodiments, the animal is genetically modified to express barcode cargo (e.g., barcode cargo polypeptide).
[0078] In some embodiments, the animal is a model of a disease, disorder, or condition. In some embodiments, the disease, disorder, or condition is a cancerous, autoimmune, neurodegenerative, or pathogenic (e.g., viral / bacterial) disease, disorder, or condition.
[0079] In some embodiments, the identifying step includes (i) sequencing nucleic acids from a binder expressing a binder, (ii) identifying the relative amount of each therapeutic candidate protein by decoding the relative amount of each barcode present, and / or (iii) performing one or more of FACS or MACS (magnetically activated cell sorting), affinity purification.
[0080] In some embodiments, the identification step includes quantifying the number of binders bound to the barcoded cargo (e.g., barcoded cargo polypeptide (e.g., barcoded therapeutic cargo polypeptide)). In some embodiments, quantification is performed by decoding the nucleotide sequence of each binder bound to the barcoded cargo (e.g., barcoded cargo polypeptide). In some embodiments, the identification step includes identifying one or more delivery particles via, for example, the barcoded cargo polypeptide.
[0081] In some embodiments, the number of nucleotide sequences provides a measure of cargo (e.g., target protein) within a population of barcoded cargo (e.g., barcoded cargo polypeptides).
[0082] In some embodiments, the administration step includes administering a barcoded cargo (e.g., a barcoded cargo polypeptide, a barcoded therapeutic candidate protein, a nucleic acid encoding the barcoded cargo polypeptide, a nucleic acid encoding the therapeutic candidate protein, etc.) that is placed within a delivery particle. In some embodiments, the administration step includes administering a barcoded cargo (e.g., a nucleic acid encoding the barcoded cargo polypeptide, a nucleic acid encoding the therapeutic candidate protein, etc.) that modifies the surface of the delivery particle.
[0083] This disclosure provides a method for characterizing a collection of peptide barcodes, the method comprising: (i) providing a library of phage particles, each phage particle designed to express a polypeptide binder, each binder binding to one or more peptide barcodes; and (ii) providing a collection of peptide barcodes, contacting each phage particle with each barcode to form bound phage-barcode particles, determining the amount of binding between each phage particle and barcode, and identifying phage-barcode pairs that bind specifically to each other between barcodes in the collection and phages in the library.
[0084] This disclosure provides a method for characterizing a collection of peptide barcodes, the method comprising (i) providing a set of binders, each binder being a polypeptide bound to one or more peptide barcodes; (ii) providing a collection of peptide barcodes; forming binder-barcode particles by contacting and binding each binder to each barcode; identifying the relative amount of binding between each polypeptide binder and the peptide barcode; and identifying binder-barcode pairs that specifically bind to each other between barcodes in the collection and binders in the set.
[0085] This disclosure provides a database of amino acid sequences or coding nucleic acid sequences for a collection of peptide barcodes, the database being embodied in a computer-readable format. In some embodiments, each barcode sequence has an amino acid length in the range of 1 to 100, 5 to 50, 8 to 25, 9 to 25, or 9 to 15. In some embodiments, each barcode sequence has been confirmed to specifically bind to one or more polypeptide binders in a set of binders, each specifically binding to one or more barcodes in the collection.
[0086] In some embodiments, the binding pattern of one or more polypeptide binders to a barcode is used to identify the peptide barcode.
[0087] This disclosure provides a database of amino acid sequences or coding nucleic acid sequences for a set of polypeptide binders. In some embodiments, the database is embodied in a computer-readable format. In some embodiments, each binder sequence has an amino acid length in the range of 10 to 400. In some embodiments, each binder sequence has been confirmed to specifically bind to one or more peptide barcodes in a set of barcodes, each of which specifically binds to one or more binders in the set.
[0088] This disclosure provides, among other things, a database of amino acid sequences or coding nucleic acid sequences for a set of barcode-binder associations, the database being embodied in a computer-readable format. In some embodiments, each barcode is a peptide with an amino acid length of 1 to 100, 5 to 50, 8 to 25, 9 to 25, or 9 to 15 amino acids. In some embodiments, each binder is a polypeptide that specifically binds to one or more barcodes in the set.
[0089] This disclosure provides a set of barcode-binder associations that are embodied in a computer-readable format. In some embodiments, each barcode is a peptide with a length of 1 to 100, 5 to 50, 8 to 25, 9 to 25, or 9 to 15 amino acids. In some embodiments, each binder is a polypeptide that specifically binds to one or more barcodes in the set.
[0090] In some embodiments, specific binding is observed when the binder is expressed on a phage particle and subsequently comes into contact with the barcode.
[0091] This disclosure provides therapeutic methods using the techniques described herein. In some embodiments, a method includes administering a therapeutic cargo polypeptide, or a characteristic portion thereof, that has been confirmed to satisfy evaluation. In some embodiments, satisfaction with evaluation may be achieved by a process comprising: a) subjecting a population of barcoded cargo polypeptides to evaluation; b) separating members of the evaluation-satisfying population from those that do not satisfy evaluation, and identifying either a positive population or a negative population, or both; c) contacting the positive population or the negative population, or each population separately from the other, with a set of binders containing at least one specific binder specific to each barcode in the population; d) identifying the binder that binds to the separated member, thereby identifying the barcoded cargo polypeptide present in the contacted population(s); and e) identifying a therapeutic cargo polypeptide from the barcoded cargo polypeptide identified to be present in the contacted population(s).
[0092] This disclosure provides a therapeutic method, which includes administering a therapeutic cargo polypeptide, or a characteristic portion thereof, that has been confirmed to satisfy evaluation by a process comprising: a) contacting a set of binders separately with a first group, a second group, or each of the first and second groups of barcoded cargo polypeptides; b) identifying the binders of the set that bind to members of the first group, the second group, or both, thereby identifying the barcoded cargo polypeptides present in the contacted group(s); and c) identifying a therapeutic cargo polypeptide from the barcoded cargo polypeptides identified as present in the contacted group(s). In some embodiments, each binder binds specifically to one or more barcodes compared to other barcodes. In some embodiments, the set of binders collectively comprises binders specific to each barcode in the first and second groups. In some embodiments, the first and second groups are separated from each other based on their performance in the evaluation.
[0093] In particular, this disclosure applies to nucleic acids as described herein, libraries of nucleic acids as described herein, a plurality of delivery particles as described herein, or cells containing delivery particles as described herein.
[0094] In particular, this disclosure covers nucleic acids as described herein, libraries of nucleic acids as described herein, a plurality of delivery particles as described herein, or populations of cells containing delivery particles as described herein.
[0095] In particular, this disclosure covers nucleic acids as described herein, libraries of nucleic acids as described herein, a plurality of delivery particles as described herein, or compositions (e.g., pharmaceutical compositions) comprising delivery particles as described herein.
[0096] In particular, the present disclosure provides a composition (e.g., a pharmaceutical composition) comprising one or more therapeutic polypeptides or one or more nucleic acids encoding a characteristic portion thereof, wherein the therapeutic polypeptide is identified from a population of barcoded cargo polypeptides by the method described herein.
[0097] In particular, the present disclosure provides a composition (e.g., a pharmaceutical composition) comprising one or more barcoded cargo polypeptides or one or more nucleic acids encoding a characteristic portion thereof, wherein the barcoded cargo polypeptides are produced by the methods described herein.
[0098] In particular, the present disclosure provides a composition (e.g., a pharmaceutical composition) comprising one or more therapeutic polypeptides or characteristic portions thereof, wherein the one or more therapeutic polypeptides are identified from a group of barcoded cargo polypeptides by the method described herein.
[0099] In particular, the present disclosure provides a composition (e.g., a pharmaceutical composition) comprising one or more barcoded cargo polypeptides or characteristic portions thereof, the one or more barcoded cargo polypeptides being produced by the methods described herein.
[0100] In particular, this disclosure provides a method for producing a composition (e.g., a pharmaceutical composition) comprising one or more therapeutic polypeptides or characteristic portions thereof, wherein the one or more therapeutic polypeptides are identified from a group of barcoded cargo polypeptides by the method described herein.
[0101] In particular, the present disclosure provides a method for producing a composition (e.g., a pharmaceutical composition) comprising one or more therapeutic polypeptides or one or more nucleic acids encoding a characteristic portion thereof, wherein the therapeutic polypeptides are identified from a population of barcoded cargo polypeptides by the method described herein.
[0102] In particular, this disclosure provides nucleic acids. In some embodiments, the nucleic acid comprises (a) a cargo component in which the nucleotide sequence encodes a cargo polypeptide, or a cargo component comprising the same, and (b) a barcode component in which the nucleotide sequence encodes a peptide barcode, or a barcode component comprising the same. In some embodiments, the barcode component may be characterized in that (i) the peptide barcode has an amino acid length in the range of 1 to 100, 5 to 50, 8 to 25, 9 to 25, or 9 to 15, and (ii) it has been confirmed to bind specifically to a particular group of polypeptide binders in a set of binders. In some embodiments, the cargo component is operably linked to the barcode component.
[0103] In some embodiments, the cargo component further comprises one or more sequence elements or their complements. In some embodiments, the cargo component further comprises one or more sequence elements or their complements selected from the group consisting of promoters, enhancers, silencers, insulators, transcription factors, translation factors, splice donors, splice acceptors, transcription terminators, translation start sites, translation termination sites, packaging signals, integration signals, and any combination thereof. In some embodiments, the cargo component further comprises one or more of the following: capping moieties, 5' untranslated regions (UTRs), 3'UTRs, polyadenylated (Poly-A) tails, or their complements, or any combination thereof. In some embodiments, the cargo component comprises an intra-sequence ribosome entry site (IRES). In some embodiments, the cargo component further encodes a cleavable moiety (e.g., a self-cleaving peptide (e.g., a 2A peptide)). In some embodiments, the cargo component, or a portion thereof, is codon-optimized.
[0104] In some embodiments, the cargo polypeptide further comprises a localization moiety. In some embodiments, the localization moiety is selected from the group consisting of secretory signaling moieties and intracellular localization moieties. In some embodiments, the cargo polypeptide further comprises an intermediate or procomponent. In some embodiments, the cargo polypeptide further comprises a tag moiety. In some embodiments, the cargo polypeptide further comprises a targeting moiety (e.g., a shuttle moiety). In some embodiments, the cargo polypeptide further comprises a ligand-binding moiety (e.g., a shuttle moiety). In some embodiments, the cargo polypeptide further comprises a stability-modifying moiety. In some embodiments, the cargo polypeptide further comprises a masking moiety. In some embodiments, the cargo polypeptide further comprises an allosteric regulatory moiety. In some embodiments, the localization moiety, tag moiety, targeting moiety, ligand-binding moiety, stability-modifying moiety, masking moiety, or allosteric regulatory moiety is cleavable.
[0105] In some embodiments, the cargo polypeptide is or includes a wild-type (e.g., spontaneously occurring) polypeptide. In some embodiments, the cargo polypeptide is or includes a variant polypeptide (e.g., a variant cargo polypeptide). In some embodiments, the variant polypeptide is a variant of a reference polypeptide, which is or includes a wild-type (e.g., spontaneously occurring) polypeptide. In some embodiments, the variant polypeptide is or includes at least one mutation of a reference polypeptide (e.g., a wild-type polypeptide).
[0106] In some embodiments, the variant cargo polypeptide associates with a barcode (e.g., is operably linked) as described herein (i.e., is a barcoded variant cargo polypeptide). In some embodiments, the variant cargo polypeptide has improved functionality (e.g., reduced toxicity, improved pharmacokinetic measures (e.g., dissociation constant (Kd), improved biophysical properties, improved scalability, improved expression, etc.) compared to a reference polypeptide (e.g., wild-type polypeptide).
[0107] In some embodiments, the cargo nucleic acid (e.g., cargo component) is or includes a wild-type (e.g., spontaneously occurring) nucleic acid. In some embodiments, the cargo nucleic acid (e.g., cargo component) is or includes a variant nucleic acid (e.g., variant cargo nucleic acid). In some embodiments, the variant nucleic acid is a variant of a reference nucleic acid, which is or includes a wild-type (e.g., spontaneously occurring) nucleic acid (e.g., a nucleic acid encoding a wild-type polypeptide). In some embodiments, the variant nucleic acid is or includes at least one mutation with respect to the reference nucleic acid (e.g., a wild-type nucleic acid (e.g., a nucleic acid encoding a wild-type polypeptide)).
[0108] In some embodiments, the variant cargo nucleic acid (e.g., the variant cargo component) associates with a barcode (e.g., is operably linked) as described herein (i.e., barcoded variant cargo nucleic acid). In some embodiments, the variant cargo nucleic acid has improved functionality (e.g., reduced toxicity, improved pharmacokinetic measures (e.g., dissociation constant (Kd), improved biophysical properties, improved scalability, improved expression, etc.) compared to the reference nucleic acid (e.g., the wild-type nucleic acid (e.g., the nucleic acid encoding the wild-type polypeptide)).
[0109] In particular, the present disclosure provides a composition (e.g., a pharmaceutical composition) comprising one or more variant polypeptides or characteristic portions thereof, wherein the one or more variant polypeptides are identified from a group of barcoded variant polypeptides by the method described herein.
[0110] In particular, the present disclosure provides a composition (e.g., a pharmaceutical composition) comprising one or more variant nucleic acids or characteristic portions thereof that encode one or more variant polypeptides, the one or more variant nucleic acids being identified from a population of barcoded variant nucleic acids by the method described herein.
[0111] In particular, the present disclosure provides a composition (e.g., a pharmaceutical composition) comprising one or more therapeutic polypeptides or one or more variant nucleic acids encoding a characteristic portion thereof, wherein the therapeutic polypeptide is identified from a population of barcoded variant cargo polypeptides by the method described herein.
[0112] In particular, the present disclosure provides compositions (e.g., pharmaceutical compositions) comprising one or more barcoded variant cargo polypeptides or one or more variant nucleic acids encoding a characteristic portion thereof, wherein the barcoded variant cargo polypeptides are produced by the methods described herein.
[0113] In particular, the present disclosure provides a composition (e.g., a pharmaceutical composition) comprising one or more therapeutic polypeptides or characteristic portions thereof, wherein the one or more therapeutic polypeptides are identified from a group of barcoded variant cargo polypeptides by the method described herein.
[0114] In particular, the present disclosure provides a composition (e.g., a pharmaceutical composition) comprising one or more barcoded variant cargo polypeptides or characteristic portions thereof, the one or more barcoded variant cargo polypeptides being produced by the methods described herein.
[0115] In some embodiments, the encoded peptide barcode has an amino acid sequence selected from the group consisting of SEQ ID NOs. 5347 to 8398. In some embodiments, the encoded peptide barcode is encoded by a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 1148 to 4199. In some embodiments, the encoded peptide barcode has a length of 8 to 25 amino acids. In some embodiments, the encoded peptide barcode has a length of 10 amino acids.
[0116] In some embodiments, the nucleotide sequence of the barcode component includes, in 5' to 3' or 3' to 5' order, one or more of the following: (a) a first invariant sequence (e.g., a linker sequence or payload sequence), (b) a variant sequence that is at least 9 nucleotides long, and (c) a second invariant sequence (e.g., a linker sequence, a stop codon, or a payload sequence). In some embodiments, the nucleotide sequence of the barcode component further includes one or more of the following: (d) a sequence encoding a short helix motif, (e) a sequence encoding a denatured motif, and (f) an invariant sequence that anneals the barcode component to the cargo component.
[0117] In some embodiments, the variant sequence is at least 15, 24, 27, 45, 150, or 300 nucleotides long.
[0118] In some embodiments, each polypeptide binder in the polypeptide binder group has an amino acid sequence selected from the group consisting of SEQ ID NOs: 4200-5346. In some embodiments, each polypeptide binder in the polypeptide binder group is encoded by a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 1-1147. In some embodiments, each polypeptide binder is expressed on a phage. In some embodiments, the phage is selected from the group consisting of M13, T4, T7, lambda, and filamentous phages. In some embodiments, the phage is M13.
[0119] In some embodiments, the nucleic acid encodes a barcoded cargo polypeptide. In some embodiments, the barcoded cargo polypeptide, or a characteristic portion thereof, is expressed on the surface of delivery particles (e.g., viral particles, lipid-based particles [e.g., cell-produced or non-cell-produced lipid nanoparticles (LNPs), liposomes, micelles, extracellular vesicles (e.g., exosomes, microparticles, etc.)], polymer-based particles (e.g., PGLA), polysaccharide-based particles, etc.).
[0120] In some embodiments, the nucleic acid is DNA or includes it. In some embodiments, the nucleic acid is RNA or includes it.
[0121] In some embodiments, the nucleic acid is disposed within the delivery particle. In some embodiments, the nucleic acid is disposed on the surface of the delivery particle.
[0122] This disclosure provides a library comprising multiple nucleic acids. In some embodiments, each nucleic acid is a nucleic acid described herein.
[0123] This disclosure provides a plurality of delivery particles. In some embodiments, one or more of the plurality of delivery particles contain nucleic acids as described herein. In some embodiments, the nucleic acids of each delivery particle of the plurality of delivery particles are the same. In some embodiments, the delivery particles contain at least two different nucleic acids. In some embodiments, the delivery particles containing at least two different nucleic acids contain different cargo components. In some embodiments, the delivery particles contain cargo components encoding at least two different cargo polypeptides. In some embodiments, the cargo polypeptide is a variant of a reference polypeptide, which is or contains a wild-type (e.g., spontaneously occurring) polypeptide. In some embodiments, the variant contains an amino acid sequence. In some embodiments, the variant contains an amino acid sequence that is at least 70% identical to one another (e.g., at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to one another).
[0124] In some embodiments, the delivery particle includes one or more associated (e.g., covalently or non-covalently) targeting moieties. In some embodiments, the one or more targeting moieties are of the same type. In some embodiments, the one or more targeting moieties are of different types.
[0125] In some embodiments, the multiple delivery particles are substantially the same type of delivery particle. In some embodiments, the multiple delivery particles comprise two or more types of delivery particles. In some embodiments, the multiple delivery particles are or comprise viral particles, lipid-based particles [e.g., lipid nanoparticles (LNPs), liposomes, micelles, extracellular vesicles (e.g., exosomes, microparticles, etc.), which are either cell-produced or not cell-produced], polymer-based particles (e.g., PGLA), polysaccharide-based particles, or a combination thereof.
[0126] In some embodiments, the multiple delivery particles are or include viral particles. In some embodiments, the multiple delivery particles are or include two or more types of viral particles. In some embodiments, the viral particles are or include one or more of AAV delivery particles, lentivirus delivery particles, adenovirus delivery particles, herpesvirus delivery particles, and anerovirus delivery particles. In some embodiments, the AAV delivery particles are or include two or more serotypes (e.g., AAV2, AAV5, AAV6, AAV8, AAV9, AAV.DJ, AAV.PHP, any variant thereof, or a combination thereof).
[0127] In some embodiments, two or more types of delivery particles are or comprise two or more types of lipid-based particles (e.g., LNPs) (e.g., having different formulations).
[0128] In particular, this disclosure provides delivery particles comprising nucleic acids as described herein.
[0129] In particular, this disclosure provides a group of delivery particles containing nucleic acids as described herein.
[0130] In particular, this disclosure provides nucleic acids as described herein, libraries as described herein, a plurality of delivery particles as described herein, or cells comprising delivery particles as described herein.
[0131] This disclosure provides a population of cells comprising nucleic acids as described herein, libraries as described herein, a plurality of delivery particles as described herein, delivery particles as described herein, or a population of delivery particles as described herein.
[0132] This disclosure provides compositions (e.g., pharmaceutical compositions). In some embodiments, the compositions include nucleic acids as described herein, libraries as described herein, a plurality of delivery particles as described herein, delivery particles as described herein, or a collection of delivery particles as described herein.
[0133] In particular, the Disclosure provides a kit comprising (a) a set of nucleic acids, and (b) a set of binders, each of which is a polypeptide or nucleic acid encoding a polypeptide, and which specifically binds to at least certain peptide barcodes in a set of barcodes. In some embodiments, each nucleic acid in the set is as described herein. In some embodiments, the kit comprises one or more binders provided as phage particles or a set thereof, which have been engineered to express the binders. In some embodiments, the kit comprises one or more binders provided as nucleic acids in a phagemid vector or as inserts suitable for cloning into a phage vector.
[0134] In some embodiments, the kit further includes information specifying a peptide barcode for each binder. In some embodiments, each binder has been confirmed to bind specifically to at least one specific peptide barcode in the set of barcodes. In some embodiments, each peptide barcode binds specifically to at least one binder in the set.
[0135] In some embodiments, the kit further includes a set of instructions for performing sequencing of one or more phage particles coupled to one or more barcodes. In some embodiments, the kit further includes a computer-readable program for decoding the sequencing data. In some embodiments, the kit further includes reagents for expressing a binder on the phage particles.
[0136] In some embodiments, the kit includes nucleic acids that encode one or more barcodes. In some embodiments, the kit includes nucleic acids that encode one or more binders.
[0137] In particular, this disclosure provides methods for identifying therapeutic polypeptides or target polypeptides for treating diseases, disorders, or conditions. In some embodiments, a method includes a) a step of subjecting a population of barcoded cargo polypeptides to evaluation; b) a step of separating members of a population that meet the evaluation criteria from those that do not, in order to identify a positive population or a negative population, or both; c) a step of contacting the positive population or the negative population, or each population separately from the other, with a set of binders containing at least one binder specific to each barcode in the population; and d) a step of identifying the binders that bind to the separated member, thereby identifying the barcoded cargo polypeptide present in the contacted population(s). In some embodiments, the barcoded cargo polypeptides are encoded by nucleic acids described herein. In some embodiments, a method further includes a) administering a population of nucleic acids encoding barcoded cargo polypeptides to an animal; and b) obtaining a sample from the animal for further evaluation.
[0138] In some embodiments, the separation step includes purifying one or more barcoded cargo polypeptides from the sample. In some embodiments, the barcoded cargo polypeptides are purified from a composite sample. In some embodiments, the composite sample is tissue. In some embodiments, the composite sample is blood. In some embodiments, the barcoded cargo polypeptides are purified using affinity purification methods (e.g., FLAG IP, Protein G / A) or protein precipitation methods.
[0139] In some embodiments, each binder in the set of binders is expressed on the phage.
[0140] In some embodiments, the identifying step includes a) amplifying the nucleic acid of a bound phage particle; b) identifying the nucleotide sequence of the amplified nucleic acid, wherein one or more of the identified nucleotide sequences correspond to the coding sequence of a binder; c) detecting one or more cargo polypeptides from a population of barcoded cargo polypeptides using the identified sequence(s) of the binder coding sequence; and f) identifying one or more barcoded cargo polypeptides as therapeutic agents or targets for treating a disease, disorder, or condition.
[0141] In particular, this disclosure provides a method for pharmacokinetic screening. In some embodiments, a method includes a) administering an animal a population of nucleic acids encoding a barcoded therapeutic candidate polypeptide or a set of characteristic portions thereof; b) collecting a sample from the animal; c) purifying one or more barcoded therapeutic candidate polypeptides from the sample; d) contacting the sample with a set of binders (e.g., binders expressing binders) containing at least one binder specific to each barcode in the sample; and e) identifying the relative amounts of each binder present in the sample (e.g., simultaneously) and identifying the pharmacokinetic properties, biodistribution, half-life, tissue-mediated drug pharmacokinetics (TMDD), epitope properties, affinity, thermal stability, pH sensitivity, or in vivo stability of each barcoded therapeutic candidate polypeptide. In some embodiments, each therapeutic candidate polypeptide contains a specific peptide barcode.
[0142] In some embodiments, multiple samples are taken from an animal. In some embodiments, the animal is a model of a disease, disorder, or condition.
[0143] In some embodiments, the animal is a mammal. In some embodiments, the animal is a human. In some embodiments, the animal is genetically modified to express a barcoded therapeutic candidate polypeptide.
[0144] In some embodiments, the disease, disorder, or condition is a cancerous, autoimmune, neurodegenerative, or pathogenic (e.g., viral / bacterial) disease, disorder, or condition.
[0145] In some embodiments, the purified therapeutic candidate polypeptide is a subset of the barcoded therapeutic candidate polypeptide administered to the animal.
[0146] In some embodiments, the sample is blood, tissue, or tumor. In some embodiments, the sample is a control.
[0147] In some embodiments, the identifying step includes (i) sequencing nucleic acids from a binder expressing a binder, (ii) identifying the relative amount of each therapeutic candidate polypeptide by decoding the relative amount of each barcode present, and / or (iii) performing one or more of FACS, MACS (magnetically activated cell sorting), or affinity purification.
[0148] In some embodiments, a method includes removing any unassociated (e.g., unbound) binders. In some embodiments, removal is carried out by washing.
[0149] In some embodiments, the specified step includes performing one or more of amplification, amplification, and sequencing (e.g., amplification, amplification, and / or sequencing of nucleic acids (e.g., DNA, RNA)). In some embodiments, amplification is performed using PCR, LAMP, or RCA. In some embodiments, sequencing is performed using Illumina, NGS, nanopore sequencing, or Pac Bio long-read sequencing.
[0150] In some embodiments, the identification step includes quantifying the number of binders that bind to the barcoded therapeutic candidate polypeptide. In some embodiments, quantification is performed by decoding the nucleotide sequence of each binder that binds to the barcoded therapeutic candidate polypeptide. In some embodiments, the number of nucleotide sequences provides a measure of the target polypeptide in a population of barcoded therapeutic candidate polypeptides.
[0151] In some embodiments, the administration step includes administering the barcoded therapeutic candidate polypeptide orally or intravenously.
[0152] In some embodiments, the barcoded therapeutic candidate polypeptide is delivered by a plurality of delivery particles described herein, a collection of delivery particles described herein, or a group of delivery particles described herein.
[0153] In particular, the present disclosure provides a therapeutic method, the method comprising administering a therapeutic polypeptide, or a nucleic acid encoding a therapeutic polypeptide or a characteristic portion thereof, that has been confirmed to satisfy an evaluation by a process comprising the steps of: a) subjecting a population of nucleic acids encoding a set of barcoded cargo polypeptides to evaluation; b) separating members of a population that satisfies the evaluation from those that do not satisfy the evaluation in order to identify a positive population or a negative population, or both; c) contacting a positive population or a negative population, or each population separately from the other, with a set of binders containing at least one binder specific to each barcode of the population; d) identifying a binder that binds to the isolated member, thereby identifying a barcoded cargo polypeptide present in the contacted population(s); and e) identifying a therapeutic cargo polypeptide from the barcoded cargo polypeptides identified as present in the contacted population(s).
[0154] In particular, the present disclosure provides a therapeutic method, which includes administering a therapeutic polypeptide, or a nucleic acid encoding a therapeutic polypeptide or a characteristic portion thereof, that has been confirmed to satisfy an evaluation by a process comprising: a) contacting a set of binders separately with a first group, a second group, or each of the first and second groups of barcoded cargo polypeptides; b) identifying the binders of the set that bind to members of the first group, the second group, or both, thereby identifying the barcoded cargo polypeptides present in the contacted group(s); and c) identifying a therapeutic polypeptide from the barcoded cargo polypeptides identified as present in the contacted group(s). In some embodiments, i) each binder binds specifically to one barcode compared to other barcodes; and ii) the set of binders collectively comprises binders specific to each barcode in the first and second groups. In some embodiments, the barcoded cargo polypeptides are encoded by the nucleic acids described herein. In some embodiments, the first and second groups are separated from each other based on their performance in the evaluation.
[0155] In particular, this disclosure provides a therapeutic method comprising administering a therapeutic polypeptide or a characteristic portion thereof. In some embodiments, the therapeutic polypeptide is identified from a population of barcoded cargo polypeptides by the method described herein.
[0156] This disclosure provides a therapeutic method comprising administering a therapeutic polypeptide or a nucleic acid encoding a characteristic portion thereof. In some embodiments, the therapeutic polypeptide is identified from a population of barcoded cargo polypeptides by the method described herein.
[0157] This disclosure provides compositions (e.g., pharmaceutical compositions) comprising one or more therapeutic polypeptides or characteristic portions thereof. In some embodiments, one or more therapeutic polypeptides are identified from a group of barcoded cargo polypeptides by the methods described herein.
[0158] This disclosure provides compositions (e.g., pharmaceutical compositions) comprising one or more barcoded cargo polypeptides or characteristic portions thereof. In some embodiments, one or more barcoded cargo polypeptides are produced by the methods described herein.
[0159] This disclosure provides compositions (e.g., pharmaceutical compositions) comprising one or more therapeutic polypeptides or one or more nucleic acids encoding characteristic portions thereof. In some embodiments, the therapeutic polypeptide is identified from a population of barcoded cargo polypeptides by the methods described herein.
[0160] This disclosure provides a method for producing a composition (e.g., a pharmaceutical composition) comprising one or more therapeutic polypeptides or characteristic portions thereof. In some embodiments, one or more therapeutic polypeptides are identified from a group of barcoded cargo polypeptides by the method described herein.
[0161] This disclosure provides a method for producing a composition (e.g., a pharmaceutical composition) comprising one or more therapeutic polypeptides or one or more nucleic acids encoding a characteristic portion thereof. In some embodiments, the therapeutic polypeptide is identified from a population of barcoded cargo polypeptides by the method described herein.
[0162] These and other aspects covered by this disclosure are described in more detail below and in the claims. [Brief explanation of the drawing]
[0163] [Figure 1A]This is a schematic diagram of a barcoded cargo as described herein, according to an exemplary embodiment. It shows the barcoded cargo and the DNA encoding the corresponding barcoded cargo. LN refers to the "linker N-terminus" and LC refers to the "linker C-terminus". In some embodiments, the LN and LC sequences are constant and encode amino acids that link the cargo to a barcode. In some embodiments, the LN and LC sequences are constant and are nucleic acid sequences used for modular cloning of barcodes using different cargoes. In some embodiments, an IIS-type restriction enzyme recognition site sequence is located alongside the LN and LC sequences. [Figure 1B] This is a schematic diagram of a barcoded cargo as described herein, according to an exemplary embodiment. It shows the nucleic acid sequence encoding the barcode and / or barcoded cargo. LN refers to the "linker N-terminus" and LC refers to the "linker C-terminus". In some embodiments, the LN and LC sequences are stationary and encode amino acids that link the cargo to the barcode. In some embodiments, the LN and LC sequences are stationary and are nucleic acid sequences used for modular cloning of barcodes using different cargoes. In some embodiments, an IIS-type restriction enzyme recognition site sequence is located alongside the LN and LC sequences. [Figure 2]This is a schematic diagram illustrating a method for detecting and / or quantifying and / or characterizing cargo (e.g., cargo polypeptides) in a pool using the barcodes and binders described herein, according to exemplary embodiments. A library of barcoded cargoes is brought into contact with a library of binders containing the DNA to be identified. A washing step is applied to remove binders that do not associate (e.g., link (e.g., form a strong bond)) with any of the barcoded cargoes, leaving only the binders that are associated with the barcodes. Following washing, a DNA sequencing process is applied to the associated binders. In some embodiments, sequencing may be performed using next-generation sequencing (NGS) (e.g., operated by an Illumina sequencer). The relative abundance of the DNA sequences is reported as a computer file (e.g., .fastq data). A computer algorithm is applied to the .fastq data, combined with prior biophysical characterization of the binders, to estimate the abundance of each barcoded cargo in the pool. [Figure 3A] This is a schematic diagram for capturing a barcode as described herein, according to an exemplary embodiment, so that the binder described herein can come into contact with it. This shows a capture scaffold to which the barcode may associate (e.g., immobilized on its surface), and the binder (e.g., a phage expressing a binder on its surface (e.g., a phage comprising binder DNA)) is brought into contact to characterize a biophysical interaction. In some embodiments, the biophysical characterization is a measurement of the dissociation constant (Kd) between the binder and the peptide barcode. [Figure 3B]This is a schematic diagram of an exemplary embodiment for capturing a barcode as described herein so that the binder described herein can come into contact with it. This is a schematic diagram of an exemplary embodiment of a barcode-binder platform as described herein. This schematic diagram shows a magnetic bead having a bead-binding domain conjugated to a universally tagged (e.g., HALO, chitin BD, Avitag (Strep), etc.) barcoded cargo. To detect the captured barcoded cargo, a binder having a known affinity for the barcode (e.g., a phage expressing a binder on its surface (e.g., a phage with binder DNA / lib)) is conjugated to the immobilized cargo. The DNA in the phage encoding the binder is then amplified and subjected to NGS for detection of the cargo. [Figure 3C] This is a schematic diagram of an exemplary embodiment for capturing a barcode as described herein so that the binder described herein can come into contact with it. This is a schematic diagram of an exemplary embodiment of a barcode-binder platform as described herein. This schematic diagram shows magnetic beads with Fc / protein A conjugate barcoded cargo. To detect the captured barcoded cargo, a binder having a known affinity for the barcode (e.g., a phage expressing the binder on its surface, e.g., a phage with the binder's DNA / lib) is bound to the immobilized cargo. The DNA in the phage encoding the binder is then amplified and subjected to NGS for detection of the cargo. [Figure 4]This is a schematic diagram of a barcode fingerprint learning method for a given barcode described herein, according to an exemplary embodiment. A peptide barcode presented on a capture scaffold is subjected to contact with a library of binders containing the DNA to be identified. A washing step is applied to remove binders that do not associate with any of the barcodes (e.g., ligate (e.g., form a strong bond)), leaving only the binders that are associated with the barcodes. After washing, the DNA sequencing process is applied to the associated binders. In some embodiments, sequencing may be performed using next-generation sequencing (NGS) (e.g., operated by an Illumina sequencer). The relative abundance of the DNA sequences is reported as a computer file (e.g., in .fastq format). A computer algorithm is applied to the .fastq data to calculate the barcode fingerprint, which is a vector of relative counts of members in the binder library. The barcode fingerprint learning method can be repeated for any barcode to identify a unique fingerprint. In some embodiments, steps 1–4 in Figure 4 may be repeated, each time starting with a focused binder library to improve the fingerprint of an existing or new barcode. In some embodiments, the focused binder library is prepared by oligonucleotide library synthesis. [Figure 5]This is a schematic diagram illustrating a method for determining the relative abundance of a mixture of barcodes using a fingerprint matrix of a set of barcodes, according to an exemplary embodiment. A set of barcodes, each with its own fingerprint identified, is mixed in known proportions and presented (e.g., on a scaffold) for subsequent contact with a binder library. The binder library is subjected to contact with the set of barcodes, and non-specific binders are washed away. Specific binders are quantified by NGS and reported as a mixture measurement computer file (e.g., .fastq format). This data is provided to a computer algorithm that uses the mixture measurements to learn the relative scaling of the readouts relative to the original fingerprints and assembles the scaled fingerprints into a scaled matrix. This scaled fingerprint matrix can then be used to quantify the relative abundance of the barcoded cargo. [Figure 6A] The results of the quantification of a composite mixture of barcodes are shown. Up to six barcodes were pooled and then measured using the decoding method described herein. The actual relative proportion of a given barcode (left panel) and the measured relative proportion of a given barcode (right panel) are shown. Rows represent individual experimental conditions, columns represent barcodes, and colors represent measured values (100% barcode = white, 0% barcode = black). [Figure 6B] The results of the quantification of barcode composite mixtures are shown. Up to six barcodes were pooled and then measured using the decoding method described herein. A plot of the measured barcode concentration against the actual barcode concentration for all experiments, compared across all barcodes, is shown. A Pearson of 0.95 was calculated between the measured proportion and the actual proportion across all experiments and mixtures. [Figure 6C]The results of the quantification of a composite mixture of barcodes are shown. Up to six barcodes were pooled and then measured using the decoding method described herein. Plots of single barcode measurements and NGS count values normalized to count per million for each mixture are shown and used to predict the relative abundance of each barcode in the mixture. Rows represent experiments, and therefore all values in a row are generated from a single .fastq file, and columns represent binders. Sequence IDs 8400-8413 are disclosed in order of appearance, respectively. [Figure 7] Figures A and B show schematic diagrams and obtained data of a method using a decoding method for cargo polypeptides with barcodes contained within the internal region of the polypeptide sequence (i.e., endogenous barcodes). The schematic diagrams show the results of a pooled synthetic barcode measurement assay. Figure A shows two barcoded cargoes (BC1 and BC2) combined at various known concentrations in different wells of a 96-well plate. Each mixture was subjected to contact with the same binder pool and decoded as described herein. Each mixture was quantified and then compared with known values for the barcoded cargoes. Figure B shows that the actual relative proportion of each barcode (X axis) correlates with the measured relative proportion (Y axis) using Pearson .96. [Figure 8A] A schematic diagram of a method for detecting cargo polypeptides in serum using the barcode-binder platform described herein, according to an exemplary embodiment, is shown. The cargo polypeptide has a barcode (i.e., an endogenous barcode) contained within the internal region of the polypeptide sequence. The diagram shows that the barcoded therapeutic antibody drug of interest (barcoded mAb) was mixed at a known concentration and then added to serum. The barcoded cargo was then purified, contacted with a binder, and subjected to decoding. [Figure 8B]A schematic diagram of a method for detecting cargo polypeptides in serum using the barcode-binder platform described herein, according to an exemplary embodiment, is shown. These cargo polypeptides have barcodes (i.e., endogenous barcodes) contained within the internal region of the polypeptide sequence. For three experimental conditions, the relative percentages of actual barcoded antibodies (left) and the relative percentages of measured antibodies (right) are shown in three replicates each. Rows correspond to experimental conditions, columns to barcodes, and the color of the heatmap cells is a measure of the percentage of barcoded antibodies present. [Figure 8C] A schematic diagram of a method for detecting cargo polypeptides in serum using the barcode-binder platform described herein, according to an exemplary embodiment, is shown. These cargo polypeptides have barcodes (i.e., endogenous barcodes) contained within the internal region of the polypeptide sequence. A scatter plot of all data for all experimental conditions for all barcodes is shown, with a Spearman correlation of 0.926 across all experimental measurements. [Figure 9A] A schematic diagram of the experiments provided in Examples 1, 2, and 8 is shown. Six unique barcodes (BC1, BC2, BC3, BC4, BC5, and BC6) were mixed in known proportions, contacted with a binder, and subjected to the decoding described herein. Two barcodes were experimentally excluded as negative controls, but since the prediction of these barcodes was possible, background predictions could be identified. [Figure 9B] This shows data on the accuracy of the decoding procedure across a 10x concentration range for six unique barcodes. It plots the measured data obtained after decoding for a mixture of real data (input) and known barcode concentrations. The input known concentrations (left bar) are shown next to the predicted / measured data (right bar) for each of the three replicated barcodes. [Figure 9C]This document presents data on the accuracy of the decoding procedure across a 10x concentration range for six unique barcodes. It plots the measured data obtained after decoding for five different mixtures (i.e., pools 1-5) of real data (input) and known barcode concentrations. The input known concentrations (left bar) are shown next to the predicted / measured data (right bar) for each of the three replicated barcodes. [Figure 10A] This specification provides a method and data for determining the absolute concentration of a single test barcode. A schematic diagram of the experiment is shown. A single test barcode was assayed at several concentrations, and a "spike-in" barcode (i.e., a reference barcode) was added to each assay mixture at known concentrations. Test barcodes of various concentrations were brought into contact with a binder and decoded as described herein. The absolute amount of test barcode to be measured was determined using the prediction of the "spike-in" barcode. [Figure 10B] This specification provides a method and data for determining the absolute concentration of a single test barcode. A plot comparing the measured absolute amount of the test barcode (right bar) with the input known concentration of the test barcode (left bar) is shown for each titration of the test barcode. The Y-axis is the logarithm of the test barcode concentration in nanograms per milliliter (ng / mL). [Figure 10C] This specification provides a method and data for determining the absolute concentration of a single test barcode. The results for determining the absolute concentration for six different barcodes are shown. The plots show the input known concentration (left bar) and the measured concentration (right bar) for the six different barcodes. [Figure 11]A method for determining the relative abundance of two polypeptides after in vivo injection using the binder-barcode system described herein, according to an exemplary embodiment, is shown. This figure shows a graphical representation of the experimental setup. In Group 1 (top), mice were injected with one barcoded cargo. In Group 2 (middle), mice were injected with two barcoded cargoes. In Group 3 (bottom), mice were injected without a barcoded cargo. For each of the three groups, serum samples were collected at 24 hours, and the barcoded cargo(s) were captured using the binder described herein and subjected to decoding. For each group, the measured barcoded cargo concentration (right bar) is shown in comparison to the input known concentration (left bar). [Figure 12A] This demonstrates the identification of 24 barcodes contained within a single mixture. A graphical representation of the experiment is shown. Of the 24 barcodes that the algorithm could predict, 10 were present in the mixture at equal concentrations. The remaining barcodes were excluded from this pool, although prediction was computationally possible. Three separate pools covering all possible barcodes were measured by replication. [Figure 12B] This shows the identification of 24 barcodes contained within a single mixture. It also shows the prediction for the first pool. The input concentration (left bar) and the measured concentration (right bar) are shown. [Figure 12C] This shows the identification of 24 barcodes contained within a single mixture. Predictions for all three pools are shown. Similar to B, the input concentration is the bar on the left, and the measured concentration is the bar on the right. [Figure 12D]This shows the identification of 24 barcodes contained within a single mixture. It also shows barcode fingerprints for the 24 barcodes used to computationally determine the relative abundance of the barcodes in three pools. Columns represent barcode fingerprints, and rows represent binder fingerprints. In order of appearance, they are sequence numbers 8414, 8415, 8414, 8416, 8414, 8413, 8414, 8417, 8414, 8418, 8414, 8419, 8414, 8420~8425, 8422, 8426, 8427, 8426, 8428~8431, 8430, 8432, 8433, 8432, 8434, 8432, 8435, 8432, 8430, 8432, 8436~8453, 8413, 8453, 8454, 8453, 8455~8475, 8474, 8476~8480, 8479, 8481~8484, 8483, 8484~ 8493, 8472, 8494, 8472, 8495, 8472, 8496, 8472, 8497, 8472, 8498~8502, 8501, 8503~8505, 8504, 8506, 8504, 8507, 8504, 8508~8516, 8515, 8517, 8518, We disclose 8417, 8519, 8520, 8519, 8521-8532, 8403, 8533-8542, 8541, 8543-8545, 8544, 8546, 8544, 8547, 8544, 8548-8552, 8551, 8553, 8551, and 8554-8562. [Figure 12E]This shows the identification of 24 barcodes contained within a single mixture. It also shows binder counts from three pools used to computationally determine the proportion of each pool. The rows are binder counts, the columns are pools, and the cells are binder counts within a particular pool. In order of appearance, they are sequence numbers 8414, 8415, 8414, 8416, 8414, 8413, 8414, 8417, 8414, 8418, 8414, 8419, 8414, 8420~8425, 8422, 8426, 8427, 8426, 8428~8431, 8430, 8432, 8433, 8432, 8434, 8432, 8435, 8432, 8430, 8432, 8436~8453, 8413, 8453, 8454, 8453, 8455~8475, 8474, 8476~8480, 8479, 8481~8484, 8483, 8484~ 8493, 8472, 8494, 8472, 8495, 8472, 8496, 8472, 8497, 8472, 8498~8502, 8501, 8503~8505, 8504, 8506, 8504, 8507, 8504, 8508~8516, 8515, 8517, 8518, We disclose 8417, 8519, 8520, 8519, 8521-8532, 8403, 8533-8542, 8541, 8543-8545, 8544, 8546, 8544, 8547, 8544, 8548-8552, 8551, 8553, 8551, and 8554-8562. [Figure 13A] This is a schematic diagram of a method for detecting and / or quantifying and / or characterizing 14 exemplary cargoes (e.g., cargo polypeptides) in a pool using the binder-barcode platform described herein. A library of barcoded cargoes was contacted with a library of a binder containing the DNA to be identified ("binder-barcode particles"). The binder-barcode particles were injected in vivo into wild-type (wt) BALB / c mice (n=3 per time point) as a pooled library. Blood was collected from individual mice at 30 minutes, 6 hours, 24 hours, and 48 hours (n=3 per time point), and serum was extracted. The binder-barcode particles were captured and subjected to the decoding procedure described herein. [Figure 13B]The following plots show the clearance of 14 exemplary binder-barcode particles injected in vivo into wild-type (wt) BALB / c mice (n=3 per time point). Data were collected at 30 minutes, 6 hours, 24 hours, and 48 hours by the decoding procedure described herein. The Y-axis is normalized to 100% of the injected volume for each exemplary binder-barcode particle. The plots shown were measured simultaneously. Each plot contains an exemplary binder-barcode particle characterized as having a certain measurable phenotype. The left plot shows the clearance (injection %) of a clinical control with known properties. The middle plot shows the clearance (injection %) of an exemplary binder-barcode particle characterized as having slow clearance properties. The right plot shows the clearance (injection %) of an exemplary binder-barcode particle characterized as having fast clearance properties. [Figure 14A] This is a schematic diagram of a method for detecting and / or quantifying and / or characterizing 36 cargoes (e.g., cargo polypeptides) in a pool using the binder-barcode platform described herein. A library of barcoded cargoes was contacted with a library of a binder containing the DNA to be identified ("binder-barcode particles"). The binder-barcode particles were injected as a pooled library into tumor-bearing NSG mice that had been pre-implanted in vivo with two tumor cell lines ("tumor 1", "tumor 2") (n=2-4 per time point). Blood and tumor tissue were collected from individual mice at 30 minutes, 6 hours, 24 hours, and 48 hours (n=3 per time point). The tissues were lysed using standard lysis buffer, and serum was separated from the blood. The binder-barcode particles were captured and subjected to the decoding procedure described herein. [Figure 14B]This is a heatmap of data collected from 36 exemplary binder-barcode particles using the decoding procedure described herein. Rows identify each exemplary binder-barcode particle tested in this embodiment. Columns show mouse data at each time point for serum, tumor 1, or tumor 2. Color intensity indicates the relative units of the drug measured via the decoding procedure described herein. Color intensity shows the normalized readout of relative concentrations measured via next-generation sequencing (NGS). [Figure 14C] Figure 14B shows a plot of binder-barcode particles using the decoding procedure described herein. Various properties were measured simultaneously. For example, binder-barcode particle P14_A5 was rapidly eliminated from the serum and accumulated only slightly in tumor 1 or tumor 2, while binder-barcode particle P17_A10 was eliminated more slowly and maintained over time in tumor 1. [Figure 15A] The plot shows the ELISA quantification of two groups of cargo (Group 1: cargo polypeptide without barcodes; Group 2: a pool of 8 binder-barcoded particles, each containing the same cargo polypeptide as used in Group 1, with each particle barcoded with a different barcode). [Figure 15B] A plot showing the quantification of the second group using the decoding procedure described herein is shown. [Figure 15C] The following shows a comparison of the half-life measurements of Group 1 and Group 2, respectively, quantified using ELISA and the decoding procedure described herein. [Figure 16A] This is a schematic diagram of a method for detecting and / or quantifying and / or characterizing 35 cargoes (e.g., cargo polypeptides) in different pools of different concentrations with different numbers of barcoded cargoes, using the binder-barcode platform described herein. [Figure 16B]The plot shows the measured barcode level (in arbitrary units) against the predicted barcode cargo level (ng) generated by aligning 96 different mixtures, each containing 10 to 35 barcode cargoes, with each barcode cargo having a known concentration of 1 pg to 1 μg. Each data point represents a comparison between the known concentration of binder-barcode particles from one of the 96 different mixtures and the concentration identified by the decoding procedure described herein. [Figure 17] A schematic diagram shows an exemplary method for providing high-throughput cargo delivery, production, screening, identification, and / or characterization as described herein. (1) A nucleotide sequence encoding a cargo polypeptide, or a cargo component containing such a sequence, and (2) a nucleotide sequence encoding a peptide barcode, or a barcode component containing such a sequence, are placed in one or more delivery particles and administered to an animal (e.g., a mammal). The functional cargo is expressed in the tissue of interest. A decoding method is used to identify cargo and / or delivery particles having the desired properties. [Figure 18] A schematic diagram shows an exemplary method according to embodiments of the present disclosure that provides tracking and / or evaluation and / or quantification of different nucleic acids encoding cargo components placed within different types of delivery particles. Two exemplary nucleic acid constructs were designed, namely (1) a first nucleic acid comprising (a) a cargo component encoding a cargo polypeptide containing a secretion signal peptide and (b) a barcode component, and (2) a second nucleic acid comprising (a) a cargo component encoding a cargo polypeptide without a secretion signal peptide and (b) a barcode component. Each nucleic acid design was placed within different delivery particles exhibiting different tissue directivity (e.g., AAV delivery particles, e.g., AAV2, AAV9, AAV.PHPB). The delivery particles were administered to mice and decoded according to the method described herein. [Figure 19]A–C are bar graphs illustrating high-throughput screening, identification, and / or quantification of two different cargo polypeptides (with or without secretory signaling peptides) delivered via different delivery particles (AAV2, AAV9, AAV.PHPB) across different tissue types (brain, liver, serum). [Figure 20] A schematic diagram illustrates how high-throughput screening provides simultaneous screening of multiple cargoes, formats, targets, and tissues across different models. [Figure 21] The following shows octet biolayer interferometry (BLI) data illustrating the dissociation of each cargo polypeptide relative to the transferrin receptor (TfR). [Figure 22] The following shows ELISA data illustrating the dissociation of each cargo polypeptide relative to the transferrin receptor (TfR). [Figure 23] Variant cargoes of cargoes previously detected, evaluated, and / or characterized (e.g., wild-type cargoes) may be generated and, for example, may be subject to further detection, evaluation, and / or characterization using the methods described herein. A schematic diagram of an exemplary method is shown. In some embodiments, such variant cargoes may have improved functionality (e.g., improved scalability, improved expression, improved affinity, etc.). [Figure 24] This specification shows plots illustrating high-throughput in vivo screening of brain shuttle candidates using the binder-barcode platform described herein. Panel (a) shows the VHH of anti-TfRs nominated for in vivo screening, possessing unique properties including epitope, affinity, thermal stability, and pH sensitivity. Panel (b) shows the VHH of 239 anti-TfRs simultaneously screened for in vivo abundance in brain, serum, and other tissues using the binder-barcode platform after 24 hours, in sets of 15–96, at doses ranging from 0.5–1 mg / kg, depending on batch size. [Figure 25]The plots shown herein illustrate PK analysis of selected screened TfR1 brain shuttle candidates across brain, cell-free fraction (parenchyma), serum, and muscle tissue, analyzed in multiplexed experiments using the binder-barcode platform described herein. [Modes for carrying out the invention]
[0164] definition When used herein in relation to a value, the term "about" refers to a similar value in the context of the value mentioned. Generally, a person skilled in the art familiar with the context will understand the reasonable degree of variation that "about" encompasses in that context. For example, in some embodiments, the term "about" may encompass a range of values within or less of 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% of the value mentioned.
[0165] Administering: As used herein, the terms “administering” or “dosing” usually refer to administering a composition to a subject or system in order to achieve delivery of a drug contained in the composition. Those skilled in the art will know the various routes that may be used for administration to a subject, e.g., a human, in appropriate circumstances. For example, in some embodiments, administration may be intraocular, oral, parenteral, topical, etc. In some specific embodiments, administration may be bronchial (e.g., by intrabronchial instillation), buccal mucosa, percutaneous (e.g., one or more of the following, or may include, topical, intradermal, interdermal, transdermal, etc.), enteral, intra-arterial, intradermal, gastric, intramedullary, intramuscular, intranasal, intraperitoneal, subarachnoid, intravenous, intraventricular, intraspecific organ (e.g., intrahepatic), mucosa, nasal, oral, rectal, subcutaneous, sublingual, topical, trachea (e.g., by intratracheal instillation), transvaginal, vitreous, etc. In some embodiments, administration may consist of only a single dose. In some embodiments, administration may include the application of a certain number of doses. In some embodiments, administration may include intermittent administration (e.g., multiple doses spaced apart in time) and / or cyclic administration (individual doses separated by a common set of time). In some embodiments, administration may include continuous administration (e.g., perfusion) for at least a selected set of time.
[0166] Affinity: As is known in the art, "affinity" is a measure of the closeness with which two or more binding partners associate with each other. Those skilled in the art will have knowledge of the various assays available for assessing affinity and will also be aware of appropriate controls for such assays. In some embodiments, affinity is assessed by quantitative assays. In some embodiments, affinity (e.g., of binding partners at one time) is assessed over multiple concentrations. In some embodiments, affinity is assessed in the presence of one or more potential competitors (e.g., relevant, e.g., those that may be present in a physiological context). In some embodiments, affinity is compared to a reference (e.g., a known affinity above a certain threshold [see “positive control”]) or a known affinity below a certain threshold [see “negative control”]. In some embodiments, affinity may be assessed in comparison to a coexisting reference, and in some embodiments, affinity may be assessed in comparison to a background reference. Typically, when affinity is assessed in comparison to a reference, it is assessed under equivalent conditions.
[0167] Agent: Generally, as used herein, the term “agent” is used to refer to an entity (e.g., lipids, metals, nucleic acids, polypeptides, polysaccharides, small molecules, etc., or complexes, combinations, mixtures, or systems thereof [e.g., cells, tissues, organisms]) or a phenomenon (e.g., heat, electric current or electric field, magnetic force or magnetic field, etc.). Under appropriate circumstances, as will be apparent to those skilled in the art, the term may be used to mean an entity that is or contains a cell or organism, or a fraction, extract, or component thereof. Or, or further, as will be apparent to the context, the term may be used to mean a natural product in the sense that it is found in nature and / or obtained from nature. In some cases, as will also be apparent to the context, the term may be used to mean one or more entities that are designed, engineered, and / or produced by human hands and / or are artificial in the sense that they are not found in nature. In some embodiments, the agent may be available in isolated or pure form, and in some embodiments, the agent may be available in crude form. In some embodiments, potential agents may be provided as aggregates or libraries, for example, that can be screened to identify or analyze the active agents contained therein. In some cases, the term “agent” may mean a compound or entity that is or contains a polymer. In some cases, the term may mean a compound or entity that contains one or more polymeric moieties. In some embodiments, the term “agent” may mean a compound or entity that is not a polymer, and / or a compound or entity that substantially does not contain any polymer, and / or a compound or entity that substantially does not contain one or more specific polymeric moieties. In some embodiments, the term may refer to a compound or entity that lacks or substantially does not contain any polymeric moieties.
[0168] Amino acid: In its broadest sense, as used herein, refers to any compound and / or substance that can be incorporated into a polypeptide chain, for example, through the formation of one or more peptide bonds. In some embodiments, an amino acid has the general structure H2N-C(H)(R)-COOH. In some embodiments, an amino acid is a naturally occurring amino acid. In some embodiments, an amino acid is a non-natural amino acid; in some embodiments, an amino acid is a D-amino acid; in some embodiments, an amino acid is an L-amino acid. "Standard amino acid" refers to any of the 20 standard L-amino acids commonly found in naturally occurring peptides. "Non-standard amino acid" refers to any amino acid other than a standard amino acid, whether synthetically prepared or obtained from a natural source. In some embodiments, amino acids, including carboxyl-terminated and / or amino-terminated amino acids in a polypeptide, may contain structural modifications compared to the general structure described above. For example, in some embodiments, amino acids may be modified compared to their general structure by methylation, amidation, acetylation, pegylation, glycosylation, phosphorylation, and / or substitution (e.g., of an amino group, a carboxylic acid group, one or more protons, and / or a hydroxyl group). In some embodiments, such modifications may alter, for example, the cyclic half-life of a polypeptide containing a modified amino acid compared to one containing the same amino acid except that which is unmodified. In some embodiments, such modifications do not significantly alter the relevant activity of a polypeptide containing a modified amino acid compared to one containing the same amino acid except that which is unmodified. As will be apparent from the context, in some embodiments, the term “amino acid” may be used to refer to a free amino acid. In some embodiments, the term may be used to refer to an amino acid residue of a polypeptide.
[0169] Animal: As used herein, refers to any member of the animal kingdom. In some embodiments, “animal” refers to a human of either sex and any developmental stage. In some embodiments, “animal” refers to a non-human animal of any developmental stage. In certain embodiments, a non-human animal is a mammal (e.g., rodents, mice, rats, rabbits, monkeys, dogs, cats, sheep, cattle, primates, and / or pigs). In some embodiments, animals include, but are not limited to, mammals, birds, reptiles, amphibians, fish, insects, and / or parasites. In some embodiments, an animal may be a genetically modified animal, a genetically modified animal, and / or a clone.
[0170] Antibody: As used herein, the term “antibody” refers to a polypeptide containing canonical immunoglobulin sequence elements sufficient to result in specific binding to a particular target antigen. As known in the art, intact antibodies, such as those produced in nature, are approximately 150 kD tetrameric chemicals composed of two identical heavy-chain polypeptides (each about 50 kD) and two identical light-chain polypeptides (each about 25 kD) that associate with each other, commonly referred to as a “Y-shaped” structure. Each heavy chain consists of at least four domains (each about 110 amino acids long): an amino-terminal variable (VH) domain (located at the tip of the Y structure) followed by three constant domains: CH1, CH2, and carboxy-terminal CH3 (located at the base of the Y stem). A short region known as the “switch” connects the heavy-chain variable and constant domains. The “hinge” links the CH2 and CH3 domains to the rest of the antibody. Two disulfide bonds in this hinge region connect two heavy-chain polypeptides in an intact antibody. Each light chain consists of two domains: an amino-terminal variable (VL) domain followed by a carboxy-terminal constant (CL) domain, which are separated from each other by another "switch." An intact antibody tetramer consists of two heavy-chain-light-chain dimers, in which the heavy and light chains are linked to each other by one disulfide bond; two other disulfide bonds connect the heavy-chain hinge regions, so that the dimers are linked together and a tetramer is formed. Also, naturally produced antibodies are usually glycosylated at the CH2 domain. Each domain in a natural antibody has a structure characterized by an "immunoglobulin fold" formed from two β-sheets (e.g., 3-chain, 4-chain, or 5-chain sheets) packed together in a compressed antiparallel β-barrel. Each mutable domain includes three hyper-mutability loops (CDR1, CDR2, and CDR3) known as "complementarity determination regions," and four somewhat immutable "framework" regions (FR1, FR2, FR3, and FR4).When a natural antibody folds, the FR region forms a beta sheet, creating a structural framework for the domain, and the CDR loop regions of both the heavy and light chains bind in three-dimensional space, resulting in a single hypervariable antigen-binding site located at the tip of the Y structure. The Fc region of naturally occurring antibodies binds to complement system elements and also to receptors on effector cells, such as effector cells that mediate cytotoxicity. As is known in the art, the affinity and / or other binding properties of the Fc region to Fc receptors can be regulated through glycosylation or other modifications. In some embodiments, antibodies produced and / or utilized according to the present invention include a glycosylated Fc domain, for example, an Fc domain that has undergone modification or engineering such as glycosylation. In the present invention, in certain embodiments, any polypeptide or polypeptide complex containing a sufficient immunoglobulin domain sequence, such as those found in natural antibodies, may be referred to and / or used as an “antibody,” regardless of whether such polypeptide is naturally occurring (e.g., produced by an organism that reacts to an antigen) or produced by recombinant operations, chemical synthesis, or other artificial systems or methodologies. In some embodiments, the antibody is polyclonal; in some embodiments, the antibody is monoclonal. In some embodiments, the antibody has a constant region sequence characteristic of mouse, rabbit, primate, or human antibodies. In some embodiments, the sequence elements of the antibody are humanized, primated, chimeric, etc., as known in the art. Furthermore, as used herein, the term “antibody” may, in appropriate embodiments (unless otherwise specified or evident from the context), refer to any construct or format known or developed in the art for utilizing the structural and functional characteristics of an antibody in an alternative presentation.For example, in embodiments, antibodies used in accordance with the present invention are not limited to, but include, intact IgA, IgG, IgE, or IgM antibodies, bispecific or multispecific antibodies (e.g., Zybodies®), antibody fragments (e.g., Fab fragments, Fab' fragments, F(ab')2 fragments, Fd' fragments, Fd fragments, and isolated CDRs or sets thereof), single-chain Fv, polypeptide-Fc fusions, single-domain antibodies (e.g., shark single-domain antibodies such as IgNAR or fragments thereof), camelid antibodies, mask antibodies (e.g., Probodies®), and small modular antibodies. The format is selected from ImmunoPharmaceuticals ("SMIPs" trademark), single-stranded diabody or tandem diabody (TandAb® trademark), VHH, Anticalin® trademark, Nanobody® minibody, BiTE® trademark, Ankyrin repeat protein or DARPINs® trademark, Avimer® trademark, DART, TCR-like antibody, Adnectin® trademark, Affilin® trademark, Trans-body® trademark, Affibody® trademark, TrimerX® trademark, MicroProteins, Fynomer® trademark, Centyrin® trademark, and KALBITOR® trademark. In some embodiments, the antibody may lack covalent modifications (e.g., glycan linkages) that it would have if naturally produced. In some embodiments, the antibody may include covalent modifications (e.g., glycan linkage), cargo [e.g., a detectable portion, a therapeutic portion, a catalytic portion, etc.], or other pendant groups [e.g., polyethylene glycol, etc.].
[0171] Antibody Agents: As used herein, the term “antibody agent” refers to a drug that specifically binds to a particular antigen. In some embodiments, the term encompasses any polypeptide or polypeptide complex containing sufficient immunoglobulin structural elements to confer specific binding. Exemplary antibody agents include, but are not limited to, monoclonal or polyclonal antibodies. In some embodiments, an antibody agent may comprise one or more constant region sequences characteristic of mouse, rabbit, primate, or human antibodies. In some embodiments, an antibody agent may comprise one or more sequence elements that are humanized, primated, chimeric, etc., as known in the art. In many embodiments, the term “antibody agent” is used to mean one or more constructs or formats known or developed in the art for utilizing the structural and functional characteristics of an antibody in an alternative presentation.For example, in embodiments, the antibody agents used according to the present invention are not limited to, but include, intact IgA, IgG, IgE, or IgM antibodies, bispecific or multispecific antibodies (e.g., Zybodies®), antibody fragments (e.g., Fab fragments, Fab' fragments, F(ab')2 fragments, Fd' fragments, Fd fragments, and isolated CDRs or sets thereof), single-chain Fv, polypeptide-Fc fusions, single-domain antibodies (e.g., shark single-domain antibodies such as IgNAR or fragments thereof), camelid antibodies, mask antibodies (e.g., Probodies®), and small modular antibodies. The format is selected from ImmunoPharmaceuticals ("SMIPs" trademark), single-stranded diabody or tandem diabody (TandAb® trademark), VHH, Anticalin® trademark, Nanobody® minibody, BiTE® trademark, Ankyrin repeat protein or DARPINs® trademark, Avimer® trademark, DART, TCR-like antibody, Adnectin® trademark, Affilin® trademark, Trans-body® trademark, Affibody® trademark, TrimerX® trademark, MicroProteins, Fynomer® trademark, Centyrin® trademark, and KALBITOR® trademark. In some embodiments, the antibody may lack covalent modifications (e.g., glycan linkages) that it would have if naturally produced. In some embodiments, the antibody may include covalent modifications (e.g., glycan linkage), cargo [e.g., a detectable portion, a therapeutic portion, a catalytic portion, etc.], or other pendant groups [e.g., polyethylene glycol, etc.]. In many embodiments, the antibody agent is a polypeptide or comprises such polypeptide, in which one or more structural elements recognized by those skilled in the art as complementarity-determining regions (CDRs) are contained in the amino acid sequence. In some embodiments, the antibody agent is a polypeptide or comprises such polypeptide, in which at least one CDR (e.g., at least one heavy-chain CDR and / or at least one light-chain CDR) is contained in the amino acid sequence that is substantially identical to that found in a reference antibody.In some embodiments, the included CDR is substantially identical to the reference CDR in that it is either sequence-identical or contains 1 to 5 amino acid substitutions when compared to the reference CDR. In some embodiments, the included CDR is substantially identical to the reference CDR in that it exhibits 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the reference CDR. In some embodiments, the included CDR is substantially identical to the reference CDR in that it exhibits 96%, 96%, 97%, 98%, 99%, or 100% sequence identity with the reference CDR. In some embodiments, the included CDR is substantially identical to the reference CDR in that at least one amino acid in the included CDR is deleted, added, or substituted compared to the reference CDR, but the included CDR has an amino acid sequence that is otherwise identical to that of the reference CDR. In some embodiments, the included CDR is substantially identical to the reference CDR in that at least one amino acid in the included CDR is substituted compared to the reference CDR, but the included CDR has an amino acid sequence that is otherwise identical to that of the reference CDR. In some embodiments, the included CDR is substantially identical to the reference CDR in that at least one amino acid in the included CDR is substituted compared to the reference CDR, but the included CDR has an amino acid sequence that is otherwise identical to that of the reference CDR. In some embodiments, the included CDR is substantially identical to the reference CDR in that at least one amino acid in the included CDR is substituted compared to the reference CDR, but the included CDR has an amino acid sequence that is otherwise identical to that of the reference CDR. In some embodiments, the antibody agent is a polypeptide containing, or comprises, a structural element in its amino acid sequence that is recognized by those skilled in the art as an immunoglobulin variable domain. In some embodiments, the antibody agent is a polypeptide protein having a binding domain homologous to an immunoglobulin binding domain or a binding domain that is largely homologous to an immunoglobulin binding domain.
[0172] Associated: The term “associated” is used herein when two events or entities are related to each other if the existence, level, degree, type and / or form of one is related to that of the other. For example, a particular entity (e.g., polypeptide, gene signature, metabolite, microorganism, etc.) is considered associated with a particular disease, disorder, or condition if its existence, level and / or form correlates with the incidence and / or susceptibility of the disease, disorder, or condition (e.g., across a relevant population). In some embodiments, two or more entities are physically “associated” with each other if they interact directly or indirectly, and as a result they are physically close to and / or remain close to each other. In some embodiments, two or more entities that are physically associated with each other are covalently bonded to each other, and in some embodiments, two or more entities that are physically associated with each other are not covalently bonded to each other but are non-covalently associated by, for example, hydrogen bonding, van der Waals interactions, hydrophobic interactions, magnetism, and combinations thereof.
[0173] Barcode, barcode component, or barcode peptide or peptide barcode: As used herein, the term “barcode” refers to a sequence (nucleic acid or amino acid) that associates (e.g., covalently or noncovalently) with a cargo as described herein. In some embodiments, the nucleic acid comprises a “barcode component” encoding a peptide barcode. In some embodiments, the barcode component is operably linked to a cargo component encoding a cargo polypeptide. In some embodiments, the peptide barcode is linked to a cargo polypeptide. As described herein, the barcode associates with a binder of known specificity and affinity. In some embodiments, the barcode binds to a specific antibody agent. In some embodiments, the barcode may be contained within a particular cargo of interest. In some embodiments, the barcode may be at the end of a particular cargo of interest. In some embodiments, the barcode may be synthetic. In some embodiments, the barcode may be designed. For example, a barcode sequence may be ordered as a DNA polynucleotide and cloned into a cargo of interest using molecular cloning methods known to those skilled in the art.
[0174] Binder: As used herein, the terms “binder” or “binder portion” refer to a polypeptide sequence that associates with a barcode of known specificity and affinity. In some embodiments, the binder is or comprises an antibody. In some embodiments, the binder is expressed on the surface of a binder. In some embodiments, the binder may bind to one or more barcodes.
[0175] Bonding: As used herein, the terms “bonding” or “joining” are generally understood to mean a non-covalent association between two or more entities. “Direct” bonding involves physical contact between entities or parts, while indirect bonding involves physical interaction through physical contact with one or more intervening entities. Bonding between two or more entities can typically be evaluated in any of a variety of contexts, including when considering the interacting entities or parts individually or in more complex systems (e.g., while covalently or otherwise bonded to a carrier entity, and / or within a biological system or cell).
[0176] Binding Agent: Generally, the term “binding agent” is used herein to refer to any entity that binds to a target of interest described herein (e.g., a barcode, a barcoded target, etc.). In many embodiments, the binding agent of interest is one that specifically binds to its target in such a way that it distinguishes that target from other potential binding partners in a particular interaction situation. Generally, a binding agent may be or include an entity of any chemical classification (e.g., polymers, nonpolymers, small molecules, polypeptides, carbohydrates, lipids, nucleic acids, etc.) or a biological classification (e.g., bacteria, phages, ribosomes, mRNA, DNA, etc.). In some embodiments, the binding agent is a single chemical substance. In some embodiments, the binding agent is a complex of two or more different chemical substances that associate with each other by non-covalent interactions under the relevant conditions. For example, it will be understood by those skilled in the art that in some embodiments, the binding agent may include a “general” binding moiety (e.g., one of biotin / avidin / streptavidin and / or class-specific antibodies) and a “specific” binding moiety (e.g., an antibody or aptamer having a specific molecular target) linked to a partner of the general binding moiety. In some embodiments, such an approach may enable the modular assembly of multiple binders through the linking of different specific binding sites with partners of the same common binding site. In some embodiments, the binder is or contains a phage. In some embodiments, the binder is or contains a polypeptide (e.g., an antibody or antibody fragment). In some embodiments, the binder is or contains a small molecule. In some embodiments, the binder is or contains a nucleic acid. In some embodiments, the binder is or contains an aptamer. In some embodiments, the binder is a polymer. In some embodiments, the binder is not a polymer. In some embodiments, the binder is nonpolymer in that they lack a polymer portion. In some embodiments, the binder is a carbohydrate. In some embodiments, the binder is a lectin.In some embodiments, the binder is or contains a peptide mimetic. In some embodiments, the binder is or contains a scaffold protein. In some embodiments, the binder is or contains a mimeotope. In some embodiments, the binder is or contains a staple peptide. In certain embodiments, the binder is or contains a nucleic acid, such as DNA or RNA.
[0177] Biological Sample: As used herein, the term “biological sample” typically refers to a sample obtained or derived from the biological source of interest as described herein (e.g., tissue or organism or cell culture). In some embodiments, the source of interest includes an animal or an organism such as a human. In some embodiments, the biological sample is or includes biological tissue or bodily fluid. In some embodiments, the biological sample may be or include bone marrow, blood, blood cells, ascites, tissue or fine-needle biopsy specimens, cell-containing bodily fluids, suspended nucleic acids, sputum, saliva, urine, cerebrospinal fluid, peritoneal fluid, pleural fluid, feces, lymph, gynecological bodily fluids, skin swabs, vaginal swabs, oral swabs, nasal swabs, lavage fluids or washes such as mammary duct lavage fluid or bronchoalveolar lavage fluid, aspirates, scrapes, bone marrow specimens, tissue biopsy specimens, surgical specimens, feces, other bodily fluids, secretions, and / or excretions, and / or cells derived therefrom. In some embodiments, the biological sample is or includes cells obtained from an individual. In some embodiments, the obtained cells are or include cells derived from the individual from which the sample is obtained. In some embodiments, the sample is a “primary sample” obtained directly from the source of interest by any suitable means. For example, in some embodiments, the primary biological sample is obtained by a method selected from the group consisting of biopsy (e.g., fine-needle aspiration or tissue biopsy), surgery, collection of bodily fluids (e.g., blood, lymph, feces, etc.). In some embodiments, as will be apparent from the context, the term “sample” refers to a preparation obtained by processing the primary sample (e.g., by removing one or more of its components and / or by adding one or more agents to it). For example, filtration using a semipermeable membrane. Such “processed sample” may include nucleic acids or proteins, for example, extracted from the sample or obtained by subjecting the primary sample to techniques such as mRNA amplification or reverse transcription, isolation and / or purification of certain components.
[0178] Cargo, Cargo Component, or Cargo Polypeptide: As used herein, the term “cargo” refers to a payload that can associate with a barcode (e.g., covalently or non-covalently). In some embodiments, the cargo comprises a nucleic acid (referred to herein as the “cargo component”) that encodes a cargo polypeptide. In some embodiments, the cargo is or comprises a cargo component (e.g., encoding a cargo polypeptide). In some embodiments, the cargo is or comprises a cargo polypeptide (e.g., encoded by the cargo component). In some embodiments, the cargo component is operably linked to a nucleic acid referred to herein as the “barcode component”. In some embodiments, the barcode component encodes a peptide barcode. In some embodiments, the peptide barcode is linked to a cargo polypeptide (e.g., covalently and / or non-covalently). In some embodiments, the cargo polypeptide is detected in a pool of polypeptides. In some embodiments, the cargo polypeptide is an unmodified polypeptide detected in a pool of polypeptides without association with a peptide barcode. In some embodiments, the cargo polypeptide is a modified polypeptide detected in a pool of polypeptides. In some embodiments, the cargo polypeptide may not be associated with a barcode (e.g., a peptide barcode). In some embodiments, the cargo comprises one or more sequences (nucleic acid sequences or amino acid sequences) that modify the expression of the cargo polypeptide. In some embodiments, such one or more sequences are associated (directly or indirectly) with the barcode throughout the period of evaluation of the cargo polypeptide described herein.
[0179] CDR: As used herein, “CDR” refers to the complementarity-determining region within the antibody variable region. There are three CDRs in each of the heavy chain and light chain variable regions, which are named CDR1, CDR2, and CDR3, respectively, for each variable region. A “set of CDRs” or “CDR set” refers to a group of three or six CDRs present in either a single variable region capable of binding to an antigen or a group of CDRs in a congeneral heavy chain and light chain variable region capable of binding to an antigen. Certain systems have been established in the art to define CDR boundaries (e.g., Kabat, Chothia, etc.), and those skilled in the art can understand the differences between these systems and the CDR boundaries to the extent necessary to understand and implement the claimed invention.
[0180] Characteristic parts: As used herein, the term “characteristic parts” refers, in its broadest sense, to parts of a substance whose presence (or absence) correlates with the presence (or absence) of a particular feature, attribute, or activity of that substance. In some embodiments, a characteristic part of a substance is a part found in a given substance and in related substances that share a particular feature, attribute, or activity, but not in related substances that do not share that particular feature, attribute, or activity. In some embodiments, a characteristic part shares at least one functional property with the intact substance. For example, in some embodiments, a “characteristic part” of a protein or polypeptide comprises a sequence of amino acid stretches, or a collection of sequence of amino acid stretches, that are both characteristic of the protein or polypeptide. In some embodiments, each such sequence of stretches typically contains at least 2, 5, 10, 15, 20, 50, or more amino acids. Generally, a characteristic part of a substance (e.g., a protein, an antibody, etc.) shares at least one functional property with the related intact substance, in addition to the sequence and / or structural identity specified above. In some embodiments, a characteristic part may be biologically active.
[0181] Characteristic sequences: As used herein, the term “characteristic sequences” refers to sequences found in all members of a family of polypeptides or nucleic acids, and can therefore be used by those skilled in the art to define members of a family.
[0182] Characteristic Sequence Elements: As used herein, the term “characteristic sequence elements” means sequence elements found in a polymer (e.g., a polypeptide or nucleic acid) that represent a characteristic portion of that polymer. In some embodiments, the presence of characteristic sequence elements correlates with the presence or level of a particular activity or property of the polymer. In some embodiments, the presence (or absence) of characteristic sequence elements defines a particular polymer as a member (or not a member) of a particular family or group of such polymers. Characteristic sequence elements typically comprise at least two monomers (e.g., amino acids or nucleotides). In some embodiments, characteristic sequence elements comprise at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, or more monomers (e.g., continuously linked monomers). In some embodiments, characteristic sequence elements comprise at least first and second stretches of continuous monomers separated by one or more spacer regions, which may or may not vary in length between polymers sharing the sequence elements.
[0183] Combination therapy: As used herein, the term “combination therapy” refers to a situation in which a subject is simultaneously exposed to two or more treatment regimens (e.g., two or more therapeutic agents). In some embodiments, two or more agents may be administered simultaneously. In some embodiments, two or more agents may be administered sequentially. In some embodiments, two or more agents may be administered in overlapping dosing regimens.
[0184] Comparable: As used herein, the term “comparable” means two or more drug, entity, situation, or condition settings that do not necessarily have to be identical to one another, but are similar enough to allow comparison between them, such that a person skilled in the art would understand that conclusions can be reasonably drawn based on observed differences or similarities. In some embodiments, comparable condition settings, situations, individuals, or groups are characterized by several substantially identical features and one or a few modified features. A person skilled in the art will understand, depending on the context, what degree of identity is required for two or more such drug, entity, situation, or condition settings to be considered comparable in any given situation. For example, a person skilled in the art will understand that a setting, individual, or group is equivalent to one another when it is characterized by a sufficient number and type of substantially identical features to ensure a reasonable conclusion that differences in results obtained or observed phenomena under or using different settings, individuals, or groups are caused by or exhibit variations in diverse features.
[0185] Comprising: A composition or method described herein as "comprising" one or more of the named elements or steps is open-ended, meaning that the named elements or steps are mandatory, but other elements or steps may be added to the scope of the composition or method. To avoid redundancy, any composition or method described as "comprising" (or "comprises") one or more of the named elements or steps also represents a more limited composition or method that "consisting essentially of" (or "consists essentially of") the same named elements or steps, meaning that the composition or method may include additional elements or steps that include the mandatory named elements or steps but do not substantially affect the basic and novel properties (or properties) of the composition or method. Furthermore, any composition or method described herein as "comprising" or "consisting essentially of" one or more named elements or steps also represents a corresponding, more limited, closed-end composition or method that "consistes of" (or "consists of") the named elements or steps, excluding any other elements or steps that are not named. In any composition or method disclosed herein, any known or disclosed equivalent of any of the named essential elements or steps may be used in place of that element or step.
[0186] Decoding: As used herein, the term “decoding” refers to a laboratory and / or bioinformatics process that identifies and quantifies a specific set of amino acids within a barcode. In some embodiments, such identification and quantification is achieved by using nucleic acid (e.g., DNA) counts from sequencing experiments and measuring the abundance of binder counts. In some embodiments, relationships between unknown barcode mixtures being decoded are identified by using previously measured fingerprints (e.g., binder fingerprints or barcode fingerprints) and comparing, for example, binder counts of known mixtures between binders with known affinity to certain barcodes in a pool and binders with varying affinity.
[0187] Designed / Made: As used herein, the term “designed / made” means (i) a drug whose structure is selected or chosen by human hands, (ii) a drug produced by a process that requires human hands, and / or (iii) a drug that is different from natural substances and other known drugs.
[0188] Determining / Identifying: Many methodologies described herein include a “determining / identifying” step. Those skilled in the art will understand, upon reading this specification, that such “determining” may be achieved by utilizing, or by using, any of the various techniques available to those skilled in the art, including, for example, specific techniques explicitly mentioned herein. In some embodiments, determining involves the manipulation of a physical sample. In some embodiments, determining involves the consideration and / or manipulation of data or information, for example, using a computer or other processing unit adapted to perform the relevant analysis. In some embodiments, determining involves receiving relevant information and / or materials from a source. In some embodiments, determining involves comparing one or more features of a sample or entity with a comparable reference.
[0189] Manipulated: Generally, the term “manipulated” refers to the aspect of being manipulated by human hands. For example, in some embodiments, a small molecule may be considered manipulated if its structure and / or production is designed and / or implemented by human hands. Similarly, in some embodiments, a polynucleotide may be considered “manipulated” if two or more sequences that are not linked to each other in nature in that order are manipulated by human hands to be directly linked to each other in the manipulated polynucleotide. For example, in some embodiments of the present invention, the manipulated polynucleotide includes a regulatory sequence which is found in nature to operably associate with a first sequence (e.g., a coding sequence) but not with a second sequence (e.g., a coding sequence), and is manipulated to operably associate with the second sequence. Similarly, a cell or organism is considered "manipulated" if its genetic information is altered (e.g., new genetic material that was not previously present is introduced by transformation, mating, somatic hybridization, transfection, transduction, or other mechanisms, or if previously present genetic material is altered or removed, for example, by substitution or deletion mutations, or by mating protocols). Although common and understood by those skilled in the art, the expression products of manipulated polynucleotides, and / or the offspring of manipulated polynucleotides or cells, are usually still referred to as "manipulated" even if the actual manipulation was performed on the original entity.
[0190] Expression: As used herein, “expression” of a nucleic acid sequence means one or more of the following events: (1) generation of an RNA template from a DNA sequence (e.g., by transcription), (2) processing of an RNA transcript (e.g., by splicing, editing, 5' cap formation, and / or 3' end formation), (3) translation of RNA into a polypeptide or protein, and / or (4) post-translational modification of a polypeptide or protein.
[0191] Fingerprint: As used herein, the term “fingerprint” refers to a count of one or more unknown agents to which a known agent may bind or associate. In some embodiments, the fingerprint may be to a known barcode or barcode mixture. In some embodiments, the fingerprint may be to a known binder or binder mixture. For example, in some embodiments, the fingerprint (e.g., barcode fingerprint) may refer to a count of one or more binders that specifically bind to a known barcode or barcode mixture (e.g., identified through sequencing analysis). That is, in some embodiments, the barcode fingerprint refers to a count of one or more binders, some of which may have a high affinity for the barcode, and some of which may have a low affinity for the barcode. In some embodiments, the fingerprint may be used in a decoding process to determine the relative or absolute abundance of a given barcode in a pool of barcodes. As will be understood by those skilled in the art, the fingerprint may be identified in relation to a known barcode or barcode mixture, or to a known binder or binder mixture. For example, in some embodiments, a fingerprint (e.g., a binder fingerprint) may refer to a count of one or more barcodes that specifically bind to a known binder or binder mixture (e.g., identified through sequencing analysis). That is, in some embodiments, a binder fingerprint refers to a count of one or more barcodes, some of which may have a high affinity for the binder, and some of which may have a low affinity for the binder. Thus, the fingerprint may also be used in the decoding process, and in some embodiments, the relative or absolute abundance of a given binder in a pool of binders is identified.
[0192] Fragments: The “fragments” of the materials or entities described herein have a structure that includes individual parts of the whole but lacks one or more parts found in the whole. In some embodiments, a fragment consists of such individual parts. In some embodiments, a fragment consists of or includes characteristic structural elements or parts found in the whole. In some embodiments, the polymer fragment comprises or consists of at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500 or more monomer units (e.g., residues) found throughout the polymer. In some embodiments, polymer fragments contain or consist of at least about 5%, 10%, 15%, 20%, 25%, 30%, 25%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more monomer units (e.g., residues) found throughout the polymer. The whole substance or entity may, in some embodiments, be referred to as the “parent” of the whole.
[0193] Human: In some embodiments, a human is an embryo, fetus, infant, child, teenager, adult, or elderly person.
[0194] To improve, increase, inhibit, or decrease: As used herein, the terms “improve,” “increase,” “inhibit,” “decrease,” or their grammatical equivalents, indicate values relative to a baseline or other reference measure. In some embodiments, a suitable reference measure may be, or include, a measurement in a particular system (e.g., in a single individual) under otherwise equivalent conditions (e.g., in the absence of a particular drug or treatment) or in the presence of a suitable equivalent reference drug. In some embodiments, a suitable reference measure may be, or include, a measurement in an equivalent system known or expected to respond in a particular way in the presence of a relevant drug or treatment. In some embodiments, “improve,” “increase,” “inhibit,” and “decrease” may be collectively referred to as “modify.”
[0195] Invariant Sequences: As used herein, the term “invariant sequence” refers to substantially identical sequences in a library of nucleic acids. In some embodiments, each nucleic acid includes, among other things, a barcode component. For example, in some embodiments, the barcode component may further include one or more of the following: (1) a nucleic acid sequence encoding a short helix motif, (2) a nucleic acid encoding a denatured motif, and (3) an invariant sequence annexing the barcode component to a cargo component. The nucleic acid sequence encoding the short helix motif and the nucleic acid encoding the denatured motif may differ from one nucleic acid library to another. In contrast, each invariant sequence in a pool of nucleic acids is substantially identical.
[0196] In vitro: As used herein, the term "in vitro" refers to events that occur in an artificial environment, such as in a test tube or reactor, or in a cell culture, rather than within a multicellular organism.
[0197] In vivo: As used herein, this term refers to events occurring within multicellular organisms such as humans and non-human animals. In the context of cell-based systems, this term may also be used to mean events occurring within living cells (e.g., not in vitro systems).
[0198] Library: As used herein, the term “library” refers to a mixture of one or more different molecules. In some embodiments, all elements of the library share one or more common components. In some embodiments, all elements of the library do not share any common components. In some embodiments, one or more elements of the library are distinguished by one or more unique components. In some embodiments, as may be evident from the context, the library may also refer to a mixture of binders. In some embodiments, the library may be a phage library. In some embodiments, for example, a phage library may consist of phages that have different binders presented (e.g., on their surface) and encapsulate DNA encoding this binder within the phages. In some embodiments, the library may refer to a mixture of barcoded cargo proteins. In some embodiments, the library may refer to a mixture of barcodes (e.g., peptide barcodes).
[0199] Linker: As used herein, this term is used to refer to a multi-component drug segment that connects different elements from one another. For example, it is understood by those skilled in the art that a polypeptide having two or more functional or organizational domains in its structure often contains stretches of amino acids between them that link such domains together. In some embodiments, a polypeptide containing a linker element has the general form S1-L-S2 overall structure, where S1 and S2 may be the same or different and represent two domains associated with each other by a linker. In some embodiments, the polypeptide linker has an amino acid length of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more. In some embodiments, the linker tends not to adopt a rigid three-dimensional structure, but rather is characterized by providing flexibility to the polypeptide. Various different linker elements that can be appropriately used when manipulating polypeptides (e.g., fusion polypeptides) are known in the art (see, for example, Holliger, P., et al. (1993) Proc. Natl. Acad. Sci. USA 90:6444-6448, Poljak, RJ, et al. (1994) Structure 2:1 121-1123).
[0200] Nucleic acid: As used herein, in its broadest sense, means any compound and / or substance that is incorporated into or can be incorporated into an oligonucleotide chain. In some embodiments, nucleic acid is a compound and / or substance that is incorporated into or can be incorporated into an oligonucleotide chain by a phosphodiester bond. As is evident from the context, in some embodiments, “nucleic acid” refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides), and in some embodiments, “nucleic acid” refers to an oligonucleotide chain comprising individual nucleic acid residues. In some embodiments, “nucleic acid” is or includes RNA. In some embodiments, “nucleic acid” is or includes DNA. In some embodiments, nucleic acid is one or more native nucleic acid residues, or includes or consists of them. In some embodiments, nucleic acid is one or more nucleic acid analogs, or includes or consists of them. In some embodiments, nucleic acid analogs differ from nucleic acids in that they do not utilize a phosphodiester backbone. For example, in some embodiments, the nucleic acid is, contains, or consists of one or more "peptide nucleic acids," which are known in the art, have peptide bonds in their backbone instead of phosphodiester bonds, and are considered to be within the scope of the present invention. Alternatively, or further, in some embodiments, the nucleic acid has one or more phosphorothioate and / or 5'-N-phosphoramidite bonds instead of phosphodiester bonds. In some embodiments, the nucleic acid is, contains, or consists of one or more natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine).In some embodiments, the nucleic acid is one or more nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, C-5 propynylcytidine, C-5 propynyluridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyluridine, C5-propynylcytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, 0(6)-methylguanine, 2-thiocytidine, methylated bases, intercalated bases, and combinations thereof), includes, or consists of. In some embodiments, the nucleic acid comprises one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) compared to those of the natural nucleic acid. In some embodiments, the nucleic acid has a nucleotide sequence encoding a functional gene product such as RNA or protein. In some embodiments, the nucleic acid contains one or more introns. In some embodiments, the nucleic acid is prepared by one or more of the following: isolation from a natural source, enzymatic synthesis by polymerization based on a complementary template (in vivo or in vitro), replication in recombinant cells or systems, and chemosynthesis. In some embodiments, the nucleic acid has a residue length of at least 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 20, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, or more. In some embodiments, the nucleic acid is partially or completely single-stranded, and in some embodiments, the nucleic acid is partially or completely double-stranded.In some embodiments, the nucleic acid has a nucleotide sequence comprising at least one element that encodes a polypeptide or is a complement to a polypeptide-encoding sequence. In some embodiments, the nucleic acid has enzymatic activity.
[0201] Operablely linked: As used herein, this refers to a juxtaposition that is in a relationship that enables the components described to function in the intended manner. A regulatory element “operably linked” to a functional element is associated with the functional element such that the expression and / or activity of the functional element is achieved under conditions that are compatible with the regulatory element. In some embodiments, the “operably linked” regulatory element is contiguous (e.g., covalently linked) to the coding element of interest, and in some embodiments, the regulatory element is either trans to the functional element of interest or otherwise acts away from it. In some embodiments, “operably linked” refers to a functional link between a regulatory element and a heterologous nucleic acid sequence, resulting in the expression of the latter. For example, the first nucleic acid sequence is operably linked to the second nucleic acid sequence if the first nucleic acid sequence is positioned in a functional relationship with the second nucleic acid sequence. In some embodiments, for example, the functional link may include transcriptional regulation. For example, a promoter is operably linked to a coding sequence if the promoter affects the transcription or expression of the coding sequence. Operablely linked DNA sequences can be adjacent to each other and, for example, are on the same reading frame if it is necessary to link two protein coding regions. In some embodiments, the cargo component is operably linked to the barcode component.
[0202] Peptides: As used herein, the term "peptide" typically refers to polypeptides that are relatively short, for example, having a length of less than about 100 amino acids, less than about 50 amino acids, less than about 40 amino acids, less than about 30 amino acids, less than about 25 amino acids, less than about 20 amino acids, less than about 15 amino acids, or less than 10 amino acids.
[0203] Pharmaceutical Composition: As used herein, the term “pharmaceutical composition” refers to a composition in which an active agent is formulated with one or more pharmaceutically acceptable carriers. In some embodiments, the active agent is present in a unit dose suitable for administration in a therapeutic regimen that exhibits a statistically significant probability of achieving a predetermined therapeutic effect when administered to an appropriate population. In some embodiments, the pharmaceutical composition may be specifically formulated for administration in solid or liquid form, including, for example, an administered-appropriate formulation such as an aqueous or non-aqueous solution or suspension, or an injectable formulation which is a droplet designed to be administered into the external auditory canal. In some embodiments, the pharmaceutical composition may be formulated for administration via infusion, either in a specific organ or compartment, such as direct administration into the ear, or systemically, such as intravenously. In some embodiments, the formulation may be or include oral medications (aqueous or non-aqueous solutions or suspensions), tablets, boluses, powders, granules, pastes, capsules, etc. In some embodiments, the active agent may be or include an isolated, purified, or pure compound.
[0204] Polypeptide: As used herein, refers to any polymer chain of residues (e.g., amino acids) that are typically linked by peptide bonds. In some embodiments, the polypeptide has a naturally occurring amino acid sequence. In some embodiments, the polypeptide has a non-natural amino acid sequence. In some embodiments, the polypeptide has a modified amino acid sequence in that it is artificially designed and / or manufactured. In some embodiments, the polypeptide may contain or consist of natural amino acids, non-natural amino acids, or both. In some embodiments, the polypeptide may contain or consist of only natural amino acids or only non-natural amino acids. In some embodiments, the polypeptide may contain D-amino acids, L-amino acids, or both. In some embodiments, the polypeptide may contain only D-amino acids. In some embodiments, the polypeptide may contain only L-amino acids. In some embodiments, the polypeptide may include one or more pendant groups or other modifications, e.g., modifications or attachments to one or more amino acid side chains, at the N-terminus of the polypeptide, the C-terminus of the polypeptide, or any combination thereof. In some embodiments, such pendant groups or modifications may be selected from the group consisting of acetylation, amidation, lipidation, methylation, pegylation, etc. (including combinations thereof). In some embodiments, the polypeptide may be cyclic and / or contain a cyclic portion. In some embodiments, the polypeptide may not be cyclic and / or contain a cyclic portion. In some embodiments, the polypeptide is linear. In some embodiments, the polypeptide may be or contain a staple polypeptide. In some embodiments, the term “polypeptide” may be predicated on the name of a reference polypeptide, activity, or structure; in such cases, it is used herein to refer to polypeptides that share a relevant activity or structure and can therefore be considered members of the same class or family of polypeptides.For each such class, this specification provides exemplary polypeptides within the class whose amino acid sequence and / or function is known, and / or which will be recognized by those skilled in the art; in some embodiments, such exemplary polypeptides are reference polypeptides of the class or family of polypeptides. In some embodiments, members of the class or family of polypeptides exhibit significant sequence homology or identity with the reference polypeptide of the class; and in some embodiments, with all polypeptides within the class; they share common sequence motifs (e.g., characteristic sequence elements) and / or common activity (in some embodiments, at equivalent levels or within a specified range). For example, in some embodiments, the member polypeptide comprises at least about 30–40% and often exhibits an overall degree of sequence homology or identity with the reference polypeptide exceeding about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more, and / or includes at least one region (for example, a conserved region that may be, or may contain, a characteristic sequence element) that often exhibits a very high degree of sequence identity exceeding 90%, or even 95%, 96%, 97%, 98%, or 99%. Such a conserved region typically comprises at least 3–4 amino acids, often up to 20 or more, and in some embodiments, the conserved region comprises at least one interval of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more consecutive amino acids. In some embodiments, the useful polypeptide may contain or consist of a fragment of the parent polypeptide.In some embodiments, a useful polypeptide may contain or consist of multiple fragments, each of which is found in the same parent polypeptide in a different spatial arrangement from that found in the polypeptide of interest (for example, a fragment directly linked to the parent may be spatially separated in the polypeptide of interest, or vice versa, and / or the fragments may be present in the polypeptide of interest in a different order than that of the parent), and thus the polypeptide of interest is a derivative of its parent polypeptide. In some embodiments, the polypeptide may be a protein.
[0205] Pro-component: As used herein, the term “pro-component” refers to an inactive component. In some embodiments, the pro-component may become active upon expression and, for example, have an intended effect. In some embodiments, the pro-component may be expressed in vitro. In some embodiments, the pro-component may be expressed in vivo. In some embodiments, the pro-component may be expressed in tissue (e.g., of an animal (e.g., a mammal)).
[0206] Protein: As used herein, the term “protein” refers to a polypeptide (i.e., a string of at least two amino acids linked together by peptide bonds). Proteins may include non-amino acid portions (e.g., glycoproteins, proteoglycans, etc.) and / or may be separately processed or modified. Those skilled in the art will understand that a “protein” may be a complete polypeptide chain (with or without a signal sequence) produced by a cell, or a characteristic portion thereof. Those skilled in the art will understand that a protein may include two or more polypeptide chains linked, for example, by one or more disulfide bonds, or linked by other means. A polypeptide may include L-amino acids, D-amino acids, or both, and may include any of the various amino acid modifications or analogs known in the art. Useful modifications include, for example, terminal acetylation, amidation, and methylation. In some embodiments, a protein may include native amino acids, non-native amino acids, synthetic amino acids, and combinations thereof. In some embodiments, a protein may be an antibody, an antibody fragment, its biologically active portion, and / or its characteristic portion.
[0207] Reference: As used herein, this refers to a standard substance or control on which the comparison is performed. For example, in some embodiments, a drug, animal, individual, population, sample, sequence, or value of interest is compared to a reference or control drug, animal, individual, population, sample, sequence, or value. In some embodiments, the reference or control is tested and / or measured substantially simultaneously with the test or measurement of interest. In some embodiments, the reference or control is optionally a historical reference or control recorded in a tangible medium of expression. As will be understood by those skilled in the art, the reference or control is usually measured or characterized under conditions or circumstances equivalent to those being evaluated. Those skilled in the art will understand when sufficient similarity exists to justify reliance on and / or comparison with a particular reference or control.
[0208] Regulatory Factors: As used herein, the terms “regulatory factor” or “regulatory sequence” refer to a non-coding region of DNA that modulates the expression of one or more specific genes in some way. In some embodiments, such genes are adjacent to or in the “neighborhood” of a given regulatory factor. In some embodiments, such genes are located very far from a given regulatory factor. In some embodiments, the regulatory factor attenuates or enhances the transcription of one or more genes. In some embodiments, the regulatory factor may be cis-positioned with respect to the gene being regulated. In some embodiments, the regulatory factor may be trans-positioned with respect to the gene being regulated. For example, in some embodiments, a regulatory sequence refers to a nucleic acid sequence that modulates the expression of a gene product operably linked to the regulatory sequence. In some such embodiments, this sequence may be an enhancer sequence and other regulatory factors that modulate the expression of a gene product.
[0209] Sample: As used herein, the term “sample” usually refers to an aliquot of a substance obtained or derived from a source of interest. In some embodiments, as will be understood from the context to those skilled in the art, the term “sample” may be used synonymously with terms such as “mixture” or “compound mixture” or “compound sample.” In some embodiments, the source of interest is a biological source or an environmental source. In some embodiments, the source of interest may be or include cells or organisms, e.g., microorganisms, plants, or animals (e.g., humans). In some embodiments, the source of interest may be or include biological tissue or biological fluids. In some embodiments, biological tissue or biological fluids may be or include cells, serum, extracellular matrix, CSF, and / or combinations or components thereof. In some embodiments, biological tissue or bodily fluid may be or include amniotic fluid, aqueous humor, ascites, bile, bone marrow, blood, breast milk, cerebrospinal fluid, earwax, chyle, porridge, ejaculated semen, endolymph, exudate, feces, gastric acid, gastric juice, lymph, mucus, pericardial fluid, perilymph, peritoneal fluid, pleural fluid, pus, mucosal secretions, saliva, sebum, semen, serum, smegma, sputum, synovial fluid, sweat, tears, urine, vaginal secretions, vitreous fluid, vomit, and / or combinations or components thereof. In some embodiments, bodily fluid may be or include intracellular fluid, extracellular fluid, intravascular fluid (plasma), interstitial fluid, lymph, and / or cell permeable fluid. In some embodiments, bodily fluid may be or include plant exudate. In some embodiments, biological tissue or biological specimens may be obtained, for example, by aspiration, biopsy (e.g., fine-needle biopsy or tissue biopsy), swab (e.g., oral swab, nasal swab, skin swab, or vaginal swab), scraping, surgery, washing or washing with a washing solution (e.g., bronchoalveolar lavage or washing solution, tubal lavage or washing solution, nasal lavage or washing solution, eye lavage or washing solution, oral lavage or washing solution, uterine lavage or washing solution, vaginal lavage or washing solution, or other washing or washing solution). In some embodiments, the biological specimen is or contains cells obtained from an individual.In some embodiments, the sample is a “primary sample” obtained directly from a source of interest by any suitable means. In some embodiments, as is evident from the context, the term “sample” refers to a preparation obtained by processing the primary sample (e.g., by removing one or more components of the primary sample and / or by adding one or more agents), for example, by filtration using a semipermeable membrane. Such a “processed sample” may include nucleic acids or proteins, obtained, for example, by extraction from the sample or by subjecting the primary sample to one or more techniques such as nucleic acid amplification or reverse transcription, isolation and / or purification of certain components.
[0210] Specific: When the term “specific” is used herein in reference to an active drug, those skilled in the art will understand that it means the drug distinguishes potential target entities or aspects from each other. For example, in some embodiments, a drug is said to bind “specifically” to a target if it preferentially binds to that target in the presence of one or more competing alternative targets. In many embodiments, specific interaction depends on the presence of specific structural properties of the target entity (e.g., epitopes, clefts, binding sites). It should be understood that specificity does not have to be absolute. In some embodiments, specificity may be evaluated in comparison to the specificity of the binding agent to one or more other potential target entities (e.g., competing substances). In some embodiments, specificity is evaluated in comparison to that of a reference specific binding agent. In some embodiments, specificity is evaluated in comparison to that of a reference nonspecific binding agent. In some embodiments, the agent or entity does not directly bind to competing alternative targets under conditions that it binds to its target entity. In some embodiments, the binder binds to the target entity with a higher on-rate, lower off-rate, increased affinity, decreased dissociation, and / or increased stability compared to competing alternative targets.
[0211] Subject: As used herein, the term “subject” refers to an organism, usually a mammal (e.g., a human, and in some embodiments, a prenatal human form). In some embodiments, the subject is suffering from the relevant disease, disorder, or condition. In some embodiments, the subject is susceptible to the disease, disorder, or condition. In some embodiments, the subject exhibits one or more symptoms or characteristics of the disease, disorder, or condition. In some embodiments, the subject exhibits no symptoms or characteristics of the disease, disorder, or condition. In some embodiments, the subject is a person having one or more characteristics that indicate susceptibility or risk to the disease, disorder, or condition. In some embodiments, the subject is a patient. In some embodiments, the subject is an individual being diagnosed and / or treated.
[0212] Substantial: As used herein, the term “substantial” refers to a qualitative state in which the characteristics or properties of an object of interest are exhibited to a degree or extent that is complete or nearly complete. Those skilled in the art of the biological field will understand that biological and chemical phenomena rarely, if ever, proceed toward completion and / or toward completion or achieve or avoid absolute results. Thus, the term “substantial” is used herein to capture the potential lack of completeness inherent in many biological and chemical phenomena.
[0213] Therapeutic Agent / Therapeutic Agent: As used herein, the term “therapeutic agent / therapeutic agent” generally refers to any agent that, when administered to an organism, induces a desired pharmacological effect. In some embodiments, an agent is considered a therapeutic agent if it exhibits a statistically significant effect across a suitable population. In some embodiments, the suitable population may be a population of model organisms. In some embodiments, the suitable population may be defined by a variety of criteria, such as a particular age group, sex, genetic background, or pre-existing clinical condition. In some embodiments, a therapeutic agent is a substance that can be used to alleviate, improve, reduce, inhibit, prevent, delay the onset, reduce the severity, and / or reduce the incidence of one or more symptoms or characteristics of a disease, disorder, and / or condition. In some embodiments, a “therapeutic agent” is an agent that has been or needs to be approved by a government agency before it can be marketed for administration to humans. In some embodiments, a “therapeutic agent” is an agent that requires a prescription for administration to humans. In some embodiments, a therapeutic agent is a therapeutic protein.
[0214] Variant: As used herein, the term “variant” refers to any, for example, a type of gene sequence that is different in some way from another type. To determine whether something is a variant, a reference type is usually selected, and the variant is different from that reference type. In some embodiments, the level of activity or functionality that a variant has may be the same as or different (e.g., higher or lower) from that of the wild-type sequence. For example, in some embodiments, a variant may have improved functionality compared to the wild-type sequence if it is mutated to confer, for example, reduced toxicity in cells. In another column, in some embodiments, a variant may have improved functionality compared to the wild-type sequence if it is mutated to confer, for example, improved protein production in cells. As another example, as used herein, “variant polypeptide” is a variant polypeptide that contains one or more mutations compared to a reference polypeptide.
[0215] I. Barcoded Cargo Methods and systems for generating and using barcodes and barcoded cargo are described herein. In some embodiments, cargo polypeptides are encoded by cargo components. In some embodiments, peptide barcodes are encoded by barcode components. In some embodiments, cargo components are operably linked to barcode components. Other exemplary cargoes are described throughout this disclosure.
[0216] In particular, this disclosure provides methods for detecting and / or characterizing cargo. In some embodiments, the methods disclosed herein are used to detect and / or characterize cargo polypeptides (e.g., therapeutic polypeptides) encoded by cargo components. In some embodiments, the methods disclosed herein are used to detect and / or characterize therapeutic or non-therapeutic polypeptides. In some embodiments, the methods disclosed herein are used to detect and / or characterize cargo by tagging the cargo with a barcode (e.g., barcoded cargo components). In some embodiments, the methods disclosed herein are used to detect and / or characterize cargo in vitro. In some embodiments, the methods disclosed herein are used to detect and / or characterize cargo in vivo. In some embodiments, the methods disclosed herein are used to detect and / or characterize cargo. In some embodiments, the methods disclosed herein are used to detect and / or characterize multiple (e.g., two or more, three or more, four or more, etc.) cargoes.
[0217] i. Barcode In some embodiments, the barcode is or contains an amino acid sequence. In some embodiments, the barcode is or contains a naturally occurring amino acid sequence. In some embodiments, the barcode is or contains an amino acid sequence that does not exist in nature. In some embodiments, the barcode is or contains a synthetic amino acid sequence. In some embodiments, the barcode contains naturally occurring amino acids. In some embodiments, the barcode contains non-naturally occurring amino acids (e.g., modified amino acids). In some embodiments, the barcode is or contains a peptide barcode.
[0218] The barcodes of this disclosure may be of various lengths. For example, in some embodiments, a barcode may have an amino acid length ranging from 1 to 100. In some embodiments, a barcode may have an amino acid length ranging from 5 to 50. In some embodiments, a barcode may have an amino acid length ranging from 8 to 25. In some embodiments, a barcode may have an amino acid length ranging from 9 to 25. In some embodiments, a barcode may have an amino acid length ranging from 9 to 15. In some embodiments, a barcode may have an amino acid length of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25. In some embodiments, a barcode may have at least 5 amino acid lengths. In some embodiments, a barcode may have up to 100 amino acid lengths.
[0219] The barcodes described herein may be available in different formats in the library. For example, in some embodiments, the barcodes described herein may be described as nucleic acid sequences. In other embodiments, the barcodes described herein may be described as amino acid sequences. Those skilled in the art will understand that a barcode described in one format may be converted to another format using basic biological principles. Thus, a barcode described as a nucleic acid sequence may be translated into a protein, which can then be used to detect the presence or absence of cargo (e.g., cargo polypeptide) in a mixture. Such a translated barcode is referred to herein as a peptide barcode.
[0220] Therefore, when described using nucleic acids, the barcodes of this disclosure may have lengths different from the lengths of the amino acid sequences disclosed in the above paragraphs. For example, in some embodiments, the barcode may have a nucleotide length of 3 to 300. In some embodiments, the barcode may have a nucleotide length of 15 to 150. In some embodiments, the barcode may have a nucleotide length of 24 to 75. In some embodiments, the barcode may have a nucleotide length of 27 to 75. In some embodiments, the barcode may have a nucleotide length of 27 to 45. In some embodiments, the barcode may have a length of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, or 75 nucleotides. In some embodiments, the barcode may have a length of at least 15 nucleotides. In some embodiments, the barcode may have a length of up to 300 nucleotides.
[0221] The barcodes of this disclosure may have one or more properties. In some embodiments, the barcodes may be spontaneously occurring. In some embodiments, the barcodes may not be spontaneously occurring (e.g., synthetic). In some embodiments, the barcodes may not significantly affect the function of the cargo. For example, in some embodiments, tagging cargo (e.g., cargo components) with a barcode as described herein does not significantly alter or change the function of the tagged cargo. In some embodiments, the barcodes may affect (e.g., positively or negatively) the function of the cargo. For example, in some embodiments, tagging cargo (e.g., cargo components) with a barcode as described herein may relatively alter or change the function of the tagged cargo (e.g., half-life (e.g., extension of half-life), increased targeting to specific tissues, etc.). In some embodiments, the barcodes may not induce an immune response (e.g., IgG response, complement response, etc.). In some embodiments, the barcodes are orthogonal to each other. In some embodiments, the barcodes are not orthogonal to each other.
[0222] The barcodes of this disclosure can be attached to various locations on the cargo. For example, in some embodiments, the barcode may assemble (e.g., covalently or non-covalently) at a suitable location on the cargo. In some embodiments, the barcode may assemble (e.g., covalently or non-covalently) at an unsuitable location on the cargo. In some embodiments, the barcode may assemble (e.g., covalently or non-covalently) at a suitable location on the cargo. For example, in some embodiments, the barcode may assemble (e.g., covalently or non-covalently) at the N-terminus of the cargo polypeptide. In some embodiments, the barcode may assemble (e.g., covalently or non-covalently) at the C-terminus of the cargo polypeptide. In some embodiments, the barcode may assemble (e.g., covalently or non-covalently) at a non-terminal location (e.g., a side chain) on the cargo polypeptide. In some embodiments, the barcode may assemble (e.g., covalently or non-covalently) at an unsuitable location on the cargo polypeptide.
[0223] In particular, further sequences (e.g., nucleic acid sequences, amino acid sequences, etc.) may be placed alongside the barcode of this disclosure (e.g., a peptide barcode or a barcode component encoding a peptide barcode). In some embodiments, further sequences may be placed at the 5' end of the barcode. In some embodiments, further sequences may be placed at the 3' end of the barcode. In some embodiments, further sequences may be placed at the 3' and 5' ends of the barcode. In some embodiments, the further sequences may be a primer binding site, a restriction endonuclease recognition sequence, a restriction enzyme site (e.g., a cleavage site), a sequence encoding an amino acid sequence, a sequence not encoding an amino acid sequence, an amino acid sequence, or a nucleic acid sequence. For example, in some embodiments, a nucleic acid sequence encoding an amino acid sequence may be placed alongside the barcode. In some embodiments, a nucleic acid sequence not encoding an amino acid sequence may be placed alongside the barcode. In some embodiments, an amino acid sequence may be placed alongside the barcode. In some embodiments, an amino acid sequence (e.g., glycine-serine (GS), e.g., another linker amino acid sequence) may be placed next to the peptide barcode. Similarly, in some embodiments, further sequences (e.g., nucleic acid sequences, amino acid sequences, etc.) may be placed next to the barcode component encoding the peptide barcode of this disclosure. In some embodiments, a nucleic acid sequence may be placed at the 5' end next to the barcode component encoding the peptide barcode. In some embodiments, a nucleic acid sequence may be placed at the 3' end next to the barcode component encoding the peptide barcode. In some embodiments, nucleic acid sequences may be placed at the 3' and 5' ends next to the barcode component encoding the peptide barcode. In some embodiments, a nucleic acid sequence encoding an amino acid sequence containing glycine-serine (GS) may be placed next to the barcode component encoding the peptide barcode. In some embodiments, a nucleic acid encoding an amino acid sequence containing the linker amino acid sequence described herein may be placed next to the barcode component encoding the peptide barcode.
[0224] In some embodiments, a restriction endonuclease recognition sequence may be located next to the barcode. In some embodiments, a restriction endonuclease recognition sequence may be located at the 5' end of the barcode. In some embodiments, a restriction endonuclease recognition sequence may be located at the 3' end of the barcode. In some embodiments, restriction endonuclease recognition sequences may be located at both the 3' and 5' ends of the barcode. In some embodiments, a restriction endonuclease recognition sequence may be located next to the nucleic acid encoding the peptide barcode. In some embodiments, a restriction endonuclease recognition sequence may be located at the 5' end of the nucleic acid encoding the peptide barcode. In some embodiments, a restriction endonuclease recognition sequence may be located at the 3' end of the nucleic acid encoding the peptide barcode. In some embodiments, restriction endonuclease recognition sequences may be located at both the 3' and 5' ends of the nucleic acid encoding the peptide barcode. In some embodiments, the restriction endonuclease recognition sequence may be recognized by one or more restriction enzymes (e.g., BsaI, BsmBI, BbsI, SapI, etc.). In some embodiments, the restriction endonuclease recognition sequence is a type I, type II, or type IIs restriction endonuclease recognition sequence. Such recognition sequences may be used, for example, to generate universal overhangs that can be used to clone peptide barcodes to different locations on various cargoes. Such flexibility makes it possible to use the barcodes to detect different cargo polypeptides in different experiments.
[0225] In some embodiments, a barcode component encoding a barcode (e.g., a peptide barcode) may associate (e.g., bind, ligate) with a second cargo component encoding a cargo polypeptide (e.g., a cargo polypeptide of interest). Such a nucleic acid sequence may be translated, for example, to form a barcoded cargo (e.g., a barcoded cargo polypeptide). In some embodiments, the barcode component encoding a barcode (e.g., a peptide barcode) is separated from the cargo component encoding the cargo polypeptide (e.g., a cargo polypeptide of interest). In such a nucleic acid sequence, for example, the barcode component encoding the peptide barcode may be translated separately from the cargo component sequence encoding the cargo polypeptide and then joined using one or more methods known in the art for ligating different amino acid sequences (e.g., using a linker).
[0226] The barcodes of this disclosure can be associated with a cargo (e.g., directly or indirectly attached) to form a barcoded cargo (or a barcoded cargo component as described herein). For example, in some embodiments, each barcode sequence (e.g., a peptide barcode sequence) may be associated with only one target cargo in the mixture (e.g., a target cargo polypeptide). In some embodiments, each barcode sequence may be associated with multiple target cargoes in the mixture (e.g., cargoes having different sequences). In some embodiments, multiple (e.g., two or more, three or more, four or more, etc.) barcode sequences may be associated with one target cargo in the mixture. For example, in some embodiments, one or more barcode sequences may be associated with various different positions on a given cargo, such configurations may be useful, for example, for testing and identifying the stability and / or cleavage of such barcoded cargo. In some embodiments, each cargo in the mixture has a unique sequence (e.g., each cargo has a different sequence from all the other cargoes in the mixture). In some embodiments, each cargo in the mixture is in a non-unique arrangement.
[0227] Various methods and parameters can be used to select an appropriate barcode for a given cargo. For example, the stability of the barcoded cargo is important in determining whether the cargo can be tagged by that barcode. In some embodiments, the barcode may be tagged on a particular cargo in different experiments. In some embodiments, the barcode may be tagged on different cargoes in different experiments. For example, in some embodiments, the barcode may be tagged on two or more, three or more, four or more, ten or more, one hundred or more, one thousand or more, or one hundred,000 or more different cargoes in different experiments.
[0228] In some embodiments, a barcode may be associated with only one cargo in a given experiment. In some embodiments, a barcode may be associated with multiple cargoes in a given experiment. For example, in some embodiments, one or more barcodes may be associated with multiple cargoes in a mixture such that each cargo is associated with a unique set of barcodes in the mixture (i.e., a barcode is tagged with multiple cargoes). That is, each cargo may be associated with a unique "pattern" of barcodes in the mixture. Similarly, in some embodiments, multiple cargoes may be associated with the same barcode.
[0229] In particular, the barcodes described herein are designed to have different (i.e., unique) sequences. In some embodiments, the barcodes are designed to have different sequences (e.g., sequences different from another barcode). For example, each barcode is designed to be different (e.g., unique) from any other barcode used in the experiment, such that each cargo (e.g., protein to be measured) is attached to at least one barcode, and each barcode (e.g., a barcode with a particular sequence) associates with only one cargo. As can be understood by those skilled in the art, the diversity of barcodes contained in the pool is limited only by the possible diversity of amino acid sequences for a given barcode length. For example, for a barcode length "N", 20 of length N NThere are different amino acid barcode sequences (only when unmodified / spontaneous amino acids are used). That is, if the barcode length is 15, the theoretical limit is 20. 15 , or 3.2768x10 19 That is the case.
[0230] In some embodiments, barcodes may be designed and / or developed through machine learning methods as described herein.
[0231] Examples of barcodes according to various embodiments of this disclosure are listed in the sequence listing submitted with this specification. In some embodiments, the barcode (e.g., a peptide barcode) is or comprises an amino acid sequence selected from SEQ ID NOs. 5347–8398. In some embodiments, the barcode (e.g., a peptide barcode) is encoded by a nucleic acid sequence selected from SEQ ID NOs. 1148–4199, or a sequence comprising such a sequence.
[0232] ii. Cargo polypeptides The methods and systems disclosed herein may be used to detect one or more cargoes described herein.
[0233] In one embodiment, the systems and methods disclosed herein may be used to detect cargo (e.g., cargo polypeptides) in a mixture. Specifically, the barcodes disclosed herein are used to tag cargo in a mixture (e.g., barcoded cargo components) and to detect such cargo in the mixture. In some embodiments, each cargo is distinct from all other cargo in the mixture. In some embodiments, each cargo in the mixture differs from all other cargo in the mixture by at least one amino acid. In some embodiments, each cargo in the mixture differs from all other cargo in the mixture by two or more amino acids. In some embodiments, cargo (e.g., in a mixture) may be tagged with a barcode. In some embodiments, each cargo (e.g., in a mixture) may be tagged with the same barcode. In some embodiments, each cargo (e.g., in a mixture) may be tagged with a different barcode. In some embodiments, cargo (e.g., in a mixture) may be tagged with a barcode that differs from all other barcodes in the mixture (e.g., associated with other cargoes) by at least one amino acid. In some embodiments, a cargo (e.g., in a mixture) may be tagged with a barcode that differs by two or more amino acids from all other barcodes in the mixture (e.g., associated with other cargoes).
[0234] As discussed elsewhere in this specification, cargoes may be tagged with different barcodes (e.g., different mixtures, different experiments, etc.). For example, as stated above, in some embodiments, each barcode sequence may be associated with only one cargo of interest in the mixture (e.g., covalently or non-covalently). In some embodiments, each barcode sequence may be associated with multiple cargoes of interest in the mixture (e.g., cargoes having different sequences) (e.g., covalently or non-covalently). In some embodiments, multiple (e.g., two or more, three or more, four or more, etc.) barcode sequences may be attached to one cargo of interest in the mixture. For example, in some embodiments, one or more barcode sequences may be attached to various different positions on a given cargo, and such configurations may be useful, for example, for testing and identifying the stability of such barcoded cargoes. In some embodiments, each cargo in the mixture has a unique sequence (e.g., each cargo has a different sequence from all other cargoes in the mixture). In some embodiments, each cargo in the mixture has a non-unique sequence.
[0235] In some embodiments, cargo may be tagged with specific barcodes in different experiments. In some embodiments, cargo may be tagged with different barcodes in different experiments. For example, in some embodiments, cargo may be tagged with two or more, three or more, four or more, ten or more, 100 or more, 1,000 or more, or 10,000 or more different barcodes in different experiments.
[0236] In some embodiments, a cargo may be associated with only one barcode in a given experiment. In some embodiments, a cargo may be associated with multiple barcodes in a given experiment. For example, in some embodiments, one or more barcodes (e.g., in a mixture) are associated with multiple cargoes in a mixture such that each cargo is associated with a unique set of barcodes in the mixture (i.e., the barcodes are tagged to multiple cargoes). That is, each cargo may be associated with a unique "pattern" of barcodes in the mixture. In some embodiments, multiple cargoes may be associated with the same barcode.
[0237] In particular, this disclosure provides nucleic acids comprising a cargo component encoding a cargo polypeptide of interest, for example. In some embodiments, the cargo polypeptide has therapeutic function. In some embodiments, the cargo polypeptide does not have therapeutic function (for example, it may assist another cargo that has therapeutic function). For example, possible cargo polypeptides may be desirable to screen as drugs, such as monoclonal antibodies, single-domain antibodies, enzymes, bispecific antibodies, or any other cargo polypeptide that may have therapeutic function.
[0238] In some embodiments, the cargo polypeptide further includes a targeting moiety. In some embodiments, the targeting moiety targets the cargo polypeptide to a site of interest (e.g., a cell of interest, a tissue of interest, or an organ of interest). In some embodiments, the targeting moiety targets the cargo polypeptide to a cell receptor drug of interest. In some embodiments, the targeting moiety is expressed on the surface of the delivery particle described herein. The targeting moiety is known in the art.
[0239] In some embodiments, the cargo polypeptide further includes a localization moiety. In some embodiments, the localization moiety is a secretory peptide signal. In some embodiments, the localization moiety is a nuclear localization signal. Other localization moieties are known in the art.
[0240] In some embodiments, the cargo polypeptide further comprises a pro-component. In some embodiments, the pro-component described herein refers to an inactive component that takes on an active form to exhibit the intended effect when expressed in the target tissue. For example, the pro-component includes moieties such as carboxylic acid, hydroxyl, amine, or phosphate / phosphonate groups. In some embodiments, the pro-component can be activated when exposed to environmental conditions such as pH, the presence (or absence) of a drug, and the like.
[0241] In some embodiments, the cargo polypeptide further comprises a tag moiety. In some embodiments, the tag moiety includes a detectable moiety. The tag moiety is known in the art.
[0242] In some embodiments, the cargo polypeptide further comprises a ligand-binding moiety. In some embodiments, the ligand-binding moiety targets the cargo polypeptide to the target tissue. In some embodiments, the ligand-binding moiety targets the cargo polypeptide to a target factor within a cell, tissue, or organ (e.g., in vivo). In some embodiments, the ligand-binding moiety targets the cargo polypeptide to a target factor on the surface of a cell, tissue, or organ (e.g., in vivo). In some embodiments, the cargo polypeptide further comprises a stability-modifying moiety. In some embodiments, the cargo polypeptide further comprises a masking moiety. In some embodiments, the cargo polypeptide further comprises an allosteric regulatory moiety.
[0243] In some embodiments, the targeting moiety may also be referred to as the shuttle moiety (or the "shuttle" described herein). In some embodiments, the ligand-binding moiety may also be referred to as the shuttle moiety (or the "shuttle" described herein). In some embodiments, the shuttle moiety is an antibody or comprises an antibody. In some embodiments, the shuttle moiety is or comprises a variant or fragment of an antibody. In some embodiments, the ligand-binding moiety is the targeting moiety or comprises the targeting moiety. In some embodiments, the targeting moiety is the ligand-binding moiety or comprises the ligand-binding moiety.
[0244] In some embodiments, the targeting moiety (e.g., the shuttle moiety) can be designed and / or developed through machine learning methods as described herein. In some embodiments, the ligand-binding moiety (e.g., the shuttle moiety) can be designed and / or developed through machine learning methods as described herein.
[0245] In some embodiments, the cargo is an antibody or includes an antibody. In some embodiments, the cargo is an antibody associated with a targeting moiety (e.g., a shuttle moiety) as described herein. In some embodiments, the cargo is an antibody associated with a ligand-binding moiety (e.g., a shuttle moiety) as described herein. In some embodiments, the cargo is an antibody-drug conjugate (ADC) or includes an ADC. In some embodiments, the cargo is an ADC associated with a targeting moiety (e.g., a shuttle moiety) as described herein. In some embodiments, the cargo is an ADC associated with a ligand-binding moiety (e.g., a shuttle moiety) as described herein. In some embodiments, the cargo is an antibody or includes an antibody associated with an oligonucleotide (e.g., covalently, e.g., noncovalently). In some embodiments, the cargo is an antibody associated with an oligonucleotide associated with a targeting moiety (e.g., a shuttle moiety) as described herein. In some embodiments, the cargo is an antibody associated with an oligonucleotide associated with a ligand-binding moiety (e.g., a shuttle moiety), as described herein, or comprises such an antibody.
[0246] In some embodiments, the oligonucleotide includes DNA. In some embodiments, the oligonucleotide includes RNA. In some embodiments, the oligonucleotide includes both DNA and RNA. In some embodiments, the oligonucleotide includes or is an RNA interference (RNAi) molecule. In some embodiments, the oligonucleotide includes or is a DNA interference (DNAi) molecule. In some embodiments, the oligonucleotide includes or is an antisense oligonucleotide (ASO). In some embodiments, the oligonucleotide includes or is an antisense oligonucleotide (ASO). In some embodiments, the oligonucleotide includes or is shRNA. In some embodiments, the oligonucleotide includes or is miRNA. In some embodiments, the oligonucleotide includes or is gRNA. In some embodiments, the oligonucleotide includes or is siRNA.
[0247] In some embodiments, the cargo polypeptide is or includes a wild-type (e.g., spontaneously occurring) polypeptide. In some embodiments, the cargo polypeptide is or includes a variant polypeptide (e.g., a variant cargo polypeptide). In some embodiments, the variant polypeptide is a variant of a reference polypeptide, which is or includes a wild-type (e.g., spontaneously occurring) polypeptide. In some embodiments, the variant polypeptide is or includes at least one mutation of a reference polypeptide (e.g., a wild-type polypeptide).
[0248] In some embodiments, the variant cargo polypeptide associates with a barcode (e.g., is operably linked) as described herein (i.e., is a barcoded variant cargo polypeptide). In some embodiments, the variant cargo polypeptide has improved functionality (e.g., reduced toxicity, improved pharmacokinetic measures (e.g., dissociation constant (Kd), improved biophysical properties, etc.) compared to a reference polypeptide (e.g., wild-type polypeptide).
[0249] In some embodiments, the cargo may be designed and / or developed through machine learning methods as described herein. In some embodiments, the cargo polypeptide may be designed and / or developed through machine learning methods as described herein. For example, in some embodiments, the cargo polypeptide (including, for example, a targeting portion or ligand-binding portion (e.g., a shuttle portion) as described herein) may be designed and / or developed through machine learning methods (e.g., refined through multiple iterations).
[0250] While pharmacokinetic (PK) characterization is a crucial criterion in nominating therapeutic leads, it is typically performed late in the drug discovery process and only for a limited number of candidates. The binder-barcode platform described herein enables the characterization of the PKs of many therapeutic candidates early in the drug discovery process.
[0251] In some embodiments, cargo designed and / or developed through machine learning methods has improved functionality (e.g., reduced toxicity, improved pharmacokinetic (pK) metrics (e.g., dissociation constant (Kd), improved biophysical properties, epitope properties, affinity, thermal stability, pH sensitivity, etc.)) compared to a reference cargo (e.g., wild-type cargo).
[0252] iii. Linker In particular, the systems and methods described herein may utilize linkers. In some embodiments, the cargo and barcode described herein are separated by linkers. In some embodiments, the linker (L) provides distance between the cargo (P) and the barcode (b). That is, structurally, the barcoded cargo may have an array of PLb in some embodiments. This may contribute, for example, to folding characteristics, cargo functionality, and / or cargo stability.
[0253] In some embodiments, the linker may be a nucleic acid. In some embodiments, the linker may be an amino acid. The linkers described herein may have a variety of lengths. For example, in some embodiments, the linker may have a length of at least 3 amino acids. In some embodiments, the linker may have a length of 1 to 50 amino acids (e.g., 1 to 30 amino acids). In some embodiments, for example, the linker may be the sequence GGGS or include it.
[0254] In some embodiments, the linker of the present invention may be cut during processing. For example, in some embodiments, the linker may include one or more motifs that may be cut during processing.
[0255] In some embodiments, the linker of the present invention may exhibit resistance to cleavage. In some embodiments, the linker of the present invention may exhibit resistance to cleavage in an assay. In some embodiments, the linker of the present invention may exhibit resistance to cleavage in vivo.
[0256] In one embodiment, linkers can be used to tag barcodes. In some embodiments, each linker sequence is associated with a different barcode sequence. For example, in some embodiments, a linker sequence can be used as a unique tag associated with a different barcode sequence (e.g., a nucleic acid sequence) in a mixture. That is, in some embodiments, such linkers can be used to amplify the associated barcode sequences. For example, in some embodiments, such linkers can be used as primers for amplifying the associated barcode sequences. Then, in some embodiments, the amplified linkers may be used to isolate the associated barcode sequences, allowing for the recovery of the barcode sequences (e.g., nucleic acid sequences) from a given linker-barcode pair. In some embodiments, the linker-barcode pairs may be subjected to DNA sequencing for identification of the barcode sequences.
[0257] In some embodiments, the nucleic acid sequence encoding the linker-barcode pair may be used to associate (e.g., ligate) the linker-barcode pair with a new cargo.
[0258] II. Binders and Binding Agents i. Binder In some embodiments, the binder (i.e., the binder portion) is or contains a nucleic acid sequence. In some embodiments, the binder is or contains a naturally occurring nucleic acid sequence. In some embodiments, the binder is or contains a nucleic acid sequence that does not exist naturally. In some embodiments, the binder is or contains a synthetic nucleic acid sequence. In some embodiments, the binder contains a naturally occurring nucleic acid. In some embodiments, the binder contains a non-naturally occurring nucleic acid (e.g., a modified nucleic acid).
[0259] In some embodiments, the nucleic acid sequence of the binder is or includes a sequence encoding a polypeptide sequence. For example, in some embodiments, the nucleic acid sequence of the binder may include a region encoding a polypeptide sequence that confers high affinity and / or specificity to a given barcode (e.g., a peptide barcode). In some embodiments, the nucleic acid sequence of the binder is or includes a sequence encoding an antibody. In some embodiments, the nucleic acid sequence of the binder is or includes a sequence encoding a fragment of an antibody. In some embodiments, the nucleic acid sequence of the binder is or includes a sequence encoding a single-chain variable fragment (scFv). As may be known to those skilled in the art, scFv is a heavy chain (V) of immunoglobulin. H ) and light chain (V L It is a fusion protein of the variable region of ). In some embodiments, V H Chain and V L The chains can be linked by short linker peptides (e.g., linkers approximately 5-50 amino acids long and 10-25 amino acids long).
[0260] In some embodiments, for example, the binder is generated to have known specificity and affinity for a given barcode. In some embodiments, the binder is generated to have known specificity and affinity for one barcode. In some embodiments, the binder is generated to have known specificity and affinity for multiple (e.g., two or more, three or more, etc.) barcodes. In some embodiments, the binder is generated to have known specificity and affinity for at least one barcode. In some embodiments, the binder is expressed on the surface of a binder (e.g., a phage, a ribosome, etc.) using, for example, a method known to those skilled in the art.
[0261] In some embodiments, the binder associates with the barcode (for example, with known specificity and affinity).
[0262] In some embodiments, the binder is or contains a naturally occurring polypeptide sequence. In some embodiments, the binder is or contains a polypeptide sequence that does not exist naturally. In some embodiments, the binder is or contains a synthetic polypeptide sequence. In some embodiments, the binder contains naturally occurring amino acids. In some embodiments, the binder contains non-naturally occurring amino acids (e.g., modified amino acids).
[0263] The binders of the present invention can be of various lengths. For example, in some embodiments, the binder can have an amino acid length ranging from 5 to 1000. In some embodiments, the binder can have an amino acid length ranging from 5 to 800. In some embodiments, the binder can have an amino acid length ranging from 6 to 500. In some embodiments, the binder can have an amino acid length ranging from 10 to 400. In some embodiments, the binder can have an amino acid length ranging from 5 to 500. In some embodiments, the binder can have an amino acid length ranging from 5 to 1000. In some embodiments, the binder can have a length of 10 amino acids. In some embodiments, the binder can have a length of at least 5 amino acids. In some embodiments, the binder can have a maximum length of 1000 amino acids.
[0264] The binders described herein can be available in different formats in a library. For example, in some embodiments, the binders described herein can be described as nucleic acid sequences. In other examples, the binders described herein can be described as amino acid sequences. Those skilled in the art will understand that a binder described in one format can be converted to another format using basic biological principles. Thus, a binder described as a nucleic acid sequence can be translated into a protein and used to detect the presence or absence of cargo (e.g., barcoded cargo (e.g., barcoded cargo polypeptide)) in a mixture. Such a translated binder is referred to herein as a polypeptide binder or a polypeptide binder moiety.
[0265] Therefore, when described using nucleic acids, the binders of this disclosure may have lengths different from the lengths of the amino acid sequences disclosed in the above paragraphs. For example, in some embodiments, the binder may have a nucleotide length ranging from 15 to 3000. In some embodiments, the binder may have a nucleotide length ranging from 15 to 2400. In some embodiments, the binder may have a nucleotide length ranging from 24 to 1500. In some embodiments, the binder may have a nucleotide length ranging from 30 to 1200. In some embodiments, the binder may have a nucleotide length of 30. In some embodiments, the binder may have a length of at least 15 nucleotides. In some embodiments, the binder may have a length of up to 3000 nucleotides.
[0266] The binders of this disclosure may have one or more specific properties. In some embodiments, the binder may be spontaneously occurring. In some embodiments, the binder may not be spontaneously occurring (e.g., synthetic). In some embodiments, the binder may not induce an immune response (e.g., an IgG response, a complement response, etc.).
[0267] In particular, similar to the barcodes described above, further sequences (e.g., nucleic acid sequences, amino acid sequences, etc.) may be placed alongside the binder of this disclosure (e.g., polypeptide binders, nucleic acids encoding the binder). In some embodiments, further sequences may be placed at the 5' end of the binder. In some embodiments, further sequences may be placed at the 3' end of the binder. In some embodiments, further sequences may be placed at the 3' and 5' ends of the binder. In some embodiments, the further sequences may be primer binding sites, restriction endonuclease recognition sequences, restriction enzyme sites (e.g., cleavage sites), sequences encoding amino acid sequences, sequences not encoding amino acid sequences, amino acid sequences, or nucleic acid sequences.
[0268] In one aspect of the present invention, a binder nucleic acid sequence can be associated (e.g., bound, ligated) with another nucleic acid sequence. For example, in some embodiments, a binder nucleic acid sequence can be associated with nucleic acid sequences encoding one or more genes. In some embodiments, a binder nucleic acid sequence can be associated with nucleic acid sequences encoding one or more genes of a phage (e.g., m13). In some embodiments, a binder nucleic acid sequence can be associated with a nucleic acid sequence encoding a polypeptide. In some embodiments, a binder nucleic acid sequence can be associated with a nucleic acid sequence encoding a polypeptide of a phage (e.g., gene 3 protein of m13). This binder-gene 3 protein fusion can be expressed and incorporated into an m13 phage.
[0269] In particular, the binders described herein are designed to have different (i.e., unique) sequences. In some embodiments, the binders are designed to have different sequences (e.g., different sequences from other binders). For example, each binder is designed to be different (i.e., unique) from all other binders used in the experiment.
[0270] In some embodiments, the binder may be designed and / or developed through machine learning methods as described herein.
[0271] In one aspect of the present invention, the binder binds to the barcode or barcoded cargo with high specificity and high affinity, as described herein. In some embodiments, the barcode or barcoded cargo (e.g., the barcoded cargo polypeptide to be measured) binds to one binder, and each binder (e.g., a binder having a specific sequence) binds to one barcode or barcoded cargo. In some embodiments, the barcode or barcoded cargo (e.g., the barcoded cargo polypeptide to be measured) binds to at least one binder. In some embodiments, each binder (e.g., a binder having a specific sequence) binds to at least one barcode or barcoded cargo. In some embodiments, multiple binders (e.g., having different sequences (e.g., polypeptide sequences)) bind to a single barcode. In some embodiments, multiple barcodes (e.g., having different sequences (e.g., peptide sequences)) bind to a single binder.
[0272] Examples of binders according to various embodiments of this disclosure are listed in the sequence listing submitted with this specification. In some embodiments, the binder (e.g., a polypeptide binder) is or comprises an amino acid sequence selected from SEQ ID NOs: 4200-5346. In some embodiments, the binder (e.g., a polypeptide binder) is encoded by a nucleic acid sequence selected from SEQ ID NOs: 1-1147, or a sequence comprising such a sequence.
[0273] ii. Binder The methods described herein relate to the detection of one or more barcodes using a binder. In some embodiments, the binder is associated with or comprises a detectable nucleic acid. In some embodiments, the binder expresses a detectable nucleic acid. In some embodiments, the binder expresses a detectable nucleic acid on its surface (e.g., a binder). In some embodiments, the binder expresses an antibody on its surface.
[0274] In some embodiments, for example, to detect the presence of a specific (e.g., different) barcode, the present invention assumes the association of different detectable nucleic acids (e.g., DNA sequences, RNA sequences, etc.) with a specific barcode. This is achieved by bringing the barcode into contact with a binder. In some embodiments, one or more barcodes may be brought into contact with the binder. In some embodiments, one or more binders may be brought into contact with the barcode.
[0275] In some embodiments, the binder may be or may include phages, ribosomes, mRNA, DNA, etc. In some embodiments, the binder is a phage. In some embodiments, the binder may be an M13 phage, a T4 phage, a T7 phage, a lambda phage, or a filamentous phage. In some embodiments, the binder may be an M13 phage.
[0276] The binders disclosed herein can be expressed on a binder using methods known in the art. For example, those skilled in the art may express nucleic acids encoding a polypeptide binder on a phage (e.g., on its surface) using techniques and methods available in the art.
[0277] III. Generation i. Barcode generation Disclosed herein are methods and systems for generating barcodes for use in the systems and methods of this disclosure. In some embodiments, the barcodes described herein can be generated rapidly (e.g., in about one week, two weeks, three weeks, four weeks, one month, two months, three months, four months, five months, six months, or one year). In some embodiments, for example, about 100 to about 1,000 barcodes can be generated rapidly. In some embodiments, about 10 to about 1,000 barcodes can be generated rapidly. In some embodiments, about 10 to about 10,000 barcodes can be generated rapidly. While a large number of barcodes can be generated rapidly according to this specification, such barcodes generated using the methods disclosed herein are also robust in that they can be bound to a known set of binders with specific and different affinities.
[0278] According to various embodiments, the barcodes described herein may be synthesized using nucleic acid (e.g., oligonucleotide) arrays. In some embodiments, the barcodes described herein may be synthesized using DNA arrays. In some embodiments, the nucleic acids (e.g., oligonucleotides) of the nucleic acid array are expressed in the barcodes. In some embodiments, the barcodes described herein may be synthesized using nucleic acid libraries. In some embodiments, the nucleic acid libraries are synthesized using nucleic acid arrays. In some embodiments, the nucleic acids (e.g., oligonucleotides) of the nucleic acid library are expressed in the barcodes.
[0279] In some embodiments, the nucleic acid library of barcodes includes about 1 or more, about 2 or more, about 3 or more, about 4 or more, about 5 or more, about 10 or more, about 50 or more, about 100 or more, about 200 or more, about 300 or more, about 400 or more, about 500 or more, about 600 or more, about 700 or more, about 800 or more, about 900 or more, about 1000 or more, about 2000 or more, about 3000 or more, about 4000 or more, or about 5000 or more potential barcodes. In some embodiments, the nucleic acid library includes one or more potential barcode sequences. Such potential barcode sequences can be screened for functionality as peptide barcodes using one or more methods described herein (i.e., after translation of the nucleic acid sequences of potential barcodes).
[0280] The barcodes of this disclosure may be screened for one or more specific properties. In some embodiments, the barcode may be screened for specific binding to a binder (e.g., specificity, binding affinity). In some embodiments, the barcode may be screened for specific binding to one or more binders. In some embodiments, the barcode may be screened for specific binding to at least one binder. In some embodiments, the barcode may be screened for specific binding to up to one binder. In some embodiments, the barcode may be screened for specific binding to multiple binders.
[0281] As will be understood by those skilled in the art, barcodes are designed to be distinguishable within a pool of barcodes (i.e., unique (e.g., having a unique sequence)). Such distinction can be achieved in some embodiments by altering one or more amino acids in the barcode. In some embodiments, a barcode differs from other barcodes in the pool of barcodes by one amino acid. In some embodiments, a barcode differs from other barcodes in the pool of barcodes by at least one amino acid. In some embodiments, a barcode differs from other barcodes in the pool of barcodes by up to one amino acid. In some embodiments, a barcode differs from other barcodes in the pool of barcodes by 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids. In some embodiments, a barcode differs from other barcodes in the pool of barcodes by at least two amino acids. In some embodiments, the barcode differs from other barcodes in the barcode pool by up to 50 amino acids.
[0282] ii. Generation of barcoded cargo The barcoded cargo according to the present invention can be generated in various ways. In some embodiments, the cargo-barcode nucleic acid sequence pair can be inserted into a plasmid to enable expression in different expression systems (e.g., protein expression systems). In some embodiments, at least one cargo-barcode nucleic acid sequence pair is inserted into the plasmid. In some embodiments, at least two cargo-barcode nucleic acid sequence pairs are inserted into the plasmid. In some embodiments, at least three cargo-barcode nucleic acid sequence pairs are inserted into the plasmid. In some embodiments, one or more cargo-barcode nucleic acid sequence pairs are inserted into the plasmid.
[0283] In some embodiments, the cargo-barcode nucleic acid sequence may include further sequences. In some embodiments, the cargo-barcode nucleic acid sequence may include further nucleic acid sequences. In some embodiments, the cargo-barcode nucleic acid sequence may include universal motif sequences. In some embodiments, the cargo-barcode nucleic acid sequence may include at least one universal motif sequence. In some embodiments, the cargo-barcode nucleic acid sequence may include at least two universal motif sequences. In some embodiments, the cargo-barcode nucleic acid sequence may include two or more universal motif sequences.
[0284] In some embodiments, at least one cargo-barcode nucleic acid sequence in the cargo-barcode nucleic acid sequence pool may include a universal motif sequence. In some embodiments, all cargo-barcode nucleic acid sequences in the cargo-barcode nucleic acid sequence pool may include a universal motif sequence.
[0285] The techniques described herein can be produced using different plasmids. In some embodiments, the plasmid is a DNA plasmid. In some embodiments, the plasmid is an RNA plasmid. In some embodiments, the plasmid is a fertility F plasmid. In some embodiments, the plasmid is a resistance plasmid. In some embodiments, the plasmid is a pathogenic plasmid. In some embodiments, the plasmid is a degradation plasmid. In some embodiments, the plasmid is a cholin-producing plasmid.
[0286] The techniques described herein can be produced using different hosts (e.g., host cells, host cell lines, etc.). In some embodiments, the host is a mammalian host. In some embodiments, the host is a non-mammalian host. In some embodiments, the host is an insect. In some embodiments, the host is a bacterium. In some embodiments, the host is E. coli.
[0287] In some embodiments, cargo-barcode pairs are expressed in vitro. In some embodiments, cargo-barcode pairs are expressed in vivo. In some embodiments, cargo-barcode pairs are expressed from RNA. In some embodiments, cargo-barcode pairs are expressed from recombinant RNA. In some embodiments, cargo-barcode pairs are expressed from DNA. In some embodiments, cargo-barcode pairs are expressed using protein components (e.g., those required for protein translation).
[0288] Following expression of the barcoded cargo construct, the construct can be purified from the pool. In some embodiments, purification may be performed using a universal motif. In some embodiments, purification may be performed using an HIS tag, FLAG tag, HALO tag, SNAP tag, Avitag, Twin strep tag, or any other tag-based protein purification method known in the art.
[0289] iii. Binder generation Disclosed herein are methods and systems for generating binders for use in the systems and methods of this disclosure. In some embodiments, the binders described herein can be generated rapidly (e.g., in about one week, two weeks, three weeks, four weeks, one month, two months, three months, four months, five months, six months, or one year). In some embodiments, for example, about 100 to about 1,000 binders can be generated rapidly. In some embodiments, about 10 to about 1,000 binders can be generated rapidly. In some embodiments, about 10 to about 10,000 binders can be generated rapidly. In some embodiments, about 10 to about 10,000 binders can be generated rapidly. While a large number of binders can be generated rapidly according to this specification, binders generated using the methods disclosed herein are also robust in that they can be bound to a known set of barcodes with specific and different affinities.
[0290] In some embodiments, the nucleic acid library of binders includes about 1 or more, about 2 or more, about 3 or more, about 4 or more, about 5 or more, about 10 or more, about 50 or more, about 100 or more, about 200 or more, about 300 or more, about 400 or more, about 500 or more, about 600 or more, about 700 or more, about 800 or more, about 900 or more, about 1000 or more, about 2000 or more, about 3000 or more, about 4000 or more, or about 5000 or more potential binders. In some embodiments, the nucleic acid library includes one or more potential binder sequences. Such potential binder sequences can be screened for functionality as a polypeptide binder using one or more methods described herein (i.e., post-translation of potential nucleic acid binder sequences).
[0291] The binder according to the present invention can be produced in various ways. In some embodiments, the nucleic acid sequence of the binder can be inserted into a plasmid to enable expression in different expression systems. In some embodiments, the nucleic acid sequence of at least one binder is inserted into the plasmid. In some embodiments, the nucleic acid sequences of at least two binders are inserted into the plasmid. In some embodiments, the nucleic acid sequences of at least three binders are inserted into the plasmid. In some embodiments, the nucleic acid sequences of one or more binders are inserted into the plasmid.
[0292] In some embodiments, the binder nucleic acid sequence is attached to one or more genes. In some embodiments, after the binder nucleic acid sequence is attached to one or more genes, it is inserted into a plasmid. In some embodiments, after the binder nucleic acid sequence is inserted into a plasmid, it is attached to one or more genes. In some embodiments, the binder nucleic acid sequence is attached to a bacteriophage gene. In some embodiments, the binder nucleic acid sequence is attached to the m13 bacteriophage gene. In some embodiments, the binder nucleic acid sequence is attached to gene 3 of the m13 bacteriophage (i.e., the one encoding the gene 3 protein).
[0293] In some embodiments, plasmids (e.g., those containing a binder sequence, or containing a binder and a bacteriophage sequence) can be transformed into a host. In some embodiments, plasmids can be transformed into and expressed in a host. In some embodiments, plasmids can be transformed into bacteria. In some embodiments, plasmids can be transformed into E. coli.
[0294] In some embodiments, plasmid expression results in phage production. In some embodiments, plasmid expression results in binder presentation on the phage surface. In some embodiments, plasmid expression results in the presentation of two binders on the phage surface. In some embodiments, plasmid expression results in the presentation of at least one binder on the phage surface. In some embodiments, plasmid expression results in the presentation of one or more binders on the phage surface. In some embodiments, plasmid expression results in the presentation of one or more binders on one or more phage surfaces. In some embodiments, plasmid expression results in the presentation of at least one binder on one or more phage surfaces.
[0295] Following phage production, the resulting pool can be purified to identify the presence of one or more polypeptide binders. In some embodiments, purification may be performed using a universal motif. In some embodiments, purification may be performed using HIS tags, FLAG tags, HALO tags, SNAP tags, Avitag, Twin strep tags, or any other tag-based protein purification method known in the art.
[0296] In some embodiments, the pool of purified binders may be highly diverse. In some embodiments, the pool of purified binders may not be highly diverse. In some embodiments, the pool of purified binders is subjected to a screening method for selecting the binder of interest.
[0297] The binders of this disclosure can be screened for one or more specific characteristics. In some embodiments, the binder can be screened for specific binding to barcodes. In some embodiments, the binder can be screened for specific binding to one or more barcodes. In some embodiments, the binder can be screened for specific binding to at least one barcode. In some embodiments, the binder can be screened for specific binding to up to one barcode. In some embodiments, the binder can be screened for specific binding to multiple barcodes.
[0298] As will be understood by those skilled in the art, binders are designed to be distinct within a pool of binders (i.e., unique (e.g., having a unique sequence)). Such distinction can be achieved in some embodiments by altering one or more amino acids in the binder. In some embodiments, a binder differs from other binders in the pool of binders by one amino acid. In some embodiments, a binder differs from other binders in the pool of binders by at least one amino acid. In some embodiments, a binder differs from other binders in the pool of binders by up to one amino acid. In some embodiments, a binder differs from other binders in the pool of binders by 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids. In some embodiments, a binder differs from other binders in the pool of binders by at least two amino acids. In some embodiments, the binder differs from other binders in the binder pool by up to 1000 amino acids.
[0299] IV. Characterization i. Sample As described elsewhere in this disclosure, the sample may be a biological sample. In some embodiments, the sample may comprise one or more barcoded cargoes. In some embodiments, the sample may comprise one or more barcoded cargo polypeptides.
[0300] In some embodiments, the sample is derived from a living organism. In some embodiments, the sample is derived from an animal. In some embodiments, the sample is derived from a disease animal model. In some embodiments, the sample is derived from a non-mammalian. In some embodiments, the sample is derived from a mammal (e.g., rodents, mice, rats, rabbits, monkeys, dogs, cats, sheep, cattle, primates, and / or pigs). In some embodiments, the sample is derived from a mouse. In some embodiments, the sample is derived from a human. In some embodiments, the sample is derived from cells (e.g., in vitro). In some embodiments, the sample is a human cell line.
[0301] In some embodiments, the sample may be purified. In some embodiments, the sample may not be purified.
[0302] In some embodiments, the sample is obtained from cells treated with barcoded cargo. In some embodiments, the sample is obtained from cells not treated with barcoded cargo. In some embodiments, the sample is obtained from animals treated with barcoded cargo. In some embodiments, the sample is obtained from animals not treated with barcoded cargo. For example, in some embodiments, the sample is obtained from humans treated with barcoded cargo polypeptide.
[0303] In some embodiments, the sample is obtained from genetically modified cells. In some embodiments, the sample is obtained from cells modified by gene therapy. In some embodiments, the sample is obtained from cells genetically modified to contain one or more barcoded cargoes. In some embodiments, the sample is obtained from cells genetically modified to express barcoded cargoes. In some embodiments, the sample is obtained from cells genetically modified to contain one or more barcodes. In some embodiments, the sample is obtained from cells genetically modified to express barcodes. In some embodiments, the sample is obtained from cells genetically modified to contain one or more binders. In some embodiments, the sample is obtained from cells genetically modified to express binders.
[0304] In some embodiments, the sample is obtained from a genetically modified animal. In some embodiments, the sample is obtained from an animal modified by gene therapy. In some embodiments, the sample is obtained from an animal genetically modified to contain one or more barcoded cargoes. In some embodiments, the sample is obtained from an animal genetically modified to express barcoded cargoes. In some embodiments, the sample is obtained from an animal genetically modified to contain one or more barcodes. In some embodiments, the sample is obtained from an animal genetically modified to express barcodes. In some embodiments, the sample is obtained from an animal genetically modified to contain one or more binders. In some embodiments, the sample is obtained from an animal genetically modified to express binders.
[0305] ii. Fingerprints In particular, the systems and methods described herein identify the advantages of nucleic acid sequencing technologies and apply them effectively to protein detection and measurement methods. For example, the methods described herein may use several binders having known binding specificity and affinity for different barcodes, which can be expressed on a binder and mixed in a single pool. After mixing with a pool of barcoded cargo polypeptides (i.e., proteins, each of which associates with a barcode described herein), each binder expressed on the binder binds to one or more barcodes in the pool that have known but different affinities. Such a spectrum of affinity of a given barcode to one or more binders can be identified through NGS, resulting in a specific distribution of binder counts with respect to a given barcode, which is referred herein to as a “barcode fingerprint”. In some embodiments, the collective barcode fingerprint for a set of barcodes is referred herein to as a “fingerprint matrix”. Similarly, the spectrum of affinity of binders to various (e.g., one or more) barcodes is referred herein to as a “binder fingerprint”. In some embodiments, the presence of barcoded cargo polypeptides(s) can be detected using the provided technique by, for example, extracting and sequencing the associated nucleic acids (e.g., detectable nucleic acids (e.g., DNA sequences, RNA sequences, etc.)) of a group of binders (e.g., phages) that bind to the barcoded cargo polypeptide(s) in a complex solution. That is, in some embodiments, for example, the presence of a protein in a complex solution is identified not by a single binder, but by a specific combination of multiple binders that bind to the barcode associated with the protein in fixed, known proportions.
[0306] The fingerprinting methods disclosed herein offer several advantages. In some embodiments, the fingerprinting approach to detection allows for noise reduction. For example, using multiple binders to detect barcodes in a composite solution makes the detection method redundant, thereby reducing the signal-to-noise ratio. Furthermore, another advantage of the "fingerprinting" approach is that partial nonspecificity of the binder (e.g., barcodes other than the target barcode to be detected) is tolerated and can be compensated for by a computational prediction method.
[0307] In some embodiments, the binder array may be modified to alter the fingerprint. In some embodiments, the binder array may be modified to improve the fingerprint.
[0308] A barcode fingerprint described herein for a given barcode may include affinity information for the given barcode to one or more binders. In some embodiments, the barcode fingerprint may include affinity information for a given barcode to one binder. In some embodiments, the barcode fingerprint may include affinity information for a given barcode to at least one binder. In some embodiments, the barcode fingerprint may include affinity information for 2, 3, 4, 5, 10, 20, 25, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10,000 or more binders. In some embodiments, the barcode fingerprint may include affinity information for a given barcode to up to 10,000 binders.
[0309] A binder fingerprint described herein for a given binder may include affinity information of the given binder to one or more barcodes. In some embodiments, the binder fingerprint may include affinity information of the given binder to one barcode. In some embodiments, the binder fingerprint may include affinity information of the given binder to at least one barcode. In some embodiments, the binder fingerprint may include affinity information of the given binder to 2, 3, 4, 5, 10, 20, 25, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10,000 or more barcodes. In some embodiments, the binder fingerprint may include affinity information of a given binder to up to 10,000 barcodes.
[0310] As discussed herein, in some embodiments, multiple barcode fingerprints for a set of barcodes may be grouped together and referred to herein as a “fingerprint matrix”. In some embodiments, a fingerprint matrix may include one barcode fingerprint. In some embodiments, a fingerprint matrix may include at least one barcode fingerprint. In some embodiments, a fingerprint matrix may include 2, 3, 4, 5, 10, 20, 25, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10,000 or more barcode fingerprints. In some embodiments, a fingerprint matrix may include up to 10,000 barcode fingerprints.
[0311] The techniques described herein enable the generation and characterization of unique fingerprints for each barcode. This makes available methods for detecting cargo (i.e., targets (e.g., proteins)) that may not require orthogonality between barcode-binder pairs. In some embodiments, the barcode-binder pairs may be orthogonal. In some embodiments, the barcode-binder pairs may not be orthogonal. As may be apparent to those skilled in the art, the barcode-binder pairs described herein offer the advantage of greater robustness in terms of the availability of unique fingerprints, thus reducing concerns about nonspecific binding, and thus providing a major advantage in complex environments (e.g., serum, blood, etc.).
[0312] iii. Decryption analysis A key component of the present invention is a method used to estimate the relative or absolute protein concentration from the DNA sequencing of binders. In this invention, a binder count table is obtained by translating the DNA sequence in silico into the amino acid sequence corresponding to each binder and compiling it into a table. The binder count table measured for any given barcode individually will hereafter be understood as the “fingerprint” of the barcode. When the present invention is applied to a mixture of unknown barcoded cargoes, the relative or absolute abundance of individual barcodes is determined by comparing the binder count table to a given fingerprint of the individual barcodes and applying the calculation-based prediction method described below. In some embodiments, the binder count table for a mixture of m unknown barcodes is assumed to be a linear combination of their respective fingerprints. The coefficients of the linear combination are estimated by the least squares method of the equation Ax=b. In the formula, A is an n × m matrix of fingerprints, b is a vector of length n of binder counts, and x is a vector of undetermined length m of the abundance of each barcode. In some embodiments, the abundance of each barcode is estimated using Bayesian methods, thereby assuming a suitable prior probability distribution over the abundance of barcodes, the likelihood ratio of the abundance of a given barcode in the observed count table is calculated from a model of the uncertainty of the experimental system, and the posterior probability distribution is estimated as the product of the prior and the likelihood ratio. In some embodiments, the posterior distribution is estimated using Monte Carlo sampling. In some embodiments, the maximum value of the posterior distribution is identified by a computational optimization procedure. In some embodiments, the binder count table is assumed to be a nonlinear function of the abundance of various barcodes to account for saturation of a particular barcode-binder interaction or competition between different barcodes or different binders.
[0313] In some embodiments, the relative proportion of binder counts is directly compared to determine the relative proportion of barcodes. In some embodiments, sequences of known abundance are mixed into the experiment and used to determine the absolute abundance of a given binder, which is then used to estimate the absolute density of barcodes.
[0314] V. Nucleic acids i. Cargo Nucleic Acids In particular, this disclosure provides nucleic acids that can be deployed, for example, within delivery particles as described herein. Nucleic acids according to this disclosure include all nucleic acids known in the art, including cargo polypeptides, cosmids incorporating nucleic acids encoding characteristic portions thereof, plasmids (e.g., naked or liposome-containing), and viral constructs (e.g., lentivirus, retrovirus, adenovirus, and adeno-associated virus constructs). Those skilled in the art will be able to select suitable constructs and cells for producing any of the polynucleotides described herein. In some embodiments, the nucleic acid is a plasmid (i.e., a circular DNA molecule capable of autonomous replication within a cell). In some embodiments, the nucleic acid may be a cosmid (e.g., the pWE or sCos series).
[0315] In some embodiments, the cargo nucleic acid (e.g., cargo component) is or includes a wild-type (e.g., spontaneously occurring) nucleic acid. In some embodiments, the cargo nucleic acid (e.g., cargo component) is or includes a variant nucleic acid (e.g., variant cargo nucleic acid). In some embodiments, the variant nucleic acid is a variant of a reference nucleic acid, which is or includes a wild-type (e.g., spontaneously occurring) nucleic acid (e.g., a nucleic acid encoding a wild-type polypeptide). In some embodiments, the variant nucleic acid is or includes at least one mutation with respect to the reference nucleic acid (e.g., a wild-type nucleic acid (e.g., a nucleic acid encoding a wild-type polypeptide)).
[0316] In some embodiments, the variant cargo nucleic acid (e.g., the variant cargo component) associates with a barcode (e.g., is operably linked) as described herein (i.e., barcoded variant cargo nucleic acid). In some embodiments, the variant cargo nucleic acid has improved functionality (e.g., reduced toxicity, improved pharmacokinetic measures (e.g., dissociation constant (Kd), improved biophysical properties, improved scalability, improved expression, etc.) compared to the reference nucleic acid (e.g., the wild-type nucleic acid (e.g., the nucleic acid encoding the wild-type polypeptide)).
[0317] 1. Viral nucleic acids In some embodiments, the nucleic acid is a viral construct. In some embodiments, the viral construct is a lentivirus, retrovirus, adenovirus, or adeno-associated virus construct. In some embodiments, the nucleic acid is an adeno-associated virus (AAV) construct (see, for example, Asokan et al., Mol. Ther. 20:699-7080, 2012, which is incorporated herein by reference as a whole). In some embodiments, the viral construct is an adenovirus construct. In some embodiments, the viral construct may also be based on or derived from an alphavirus. Alphaviruses include Sindbis (and VEEV) virus, Aura virus, Babanki virus, Burma forest virus, Beval virus, Cabasou virus, Chikungunya virus, Eastern equine encephalitis virus, Everglades virus, Fort Morgan virus, Geta virus, Highlands J virus, Kiziragati virus, Mayaro virus, Metri virus, Middleburg virus, Mosso das Pedras virus, Mukambo virus, Ndumu virus, Onyonnyon virus, Pixuna virus, Rio Negro virus, Ross River virus, Salmonid pancreatic disease virus, Semlik Forest virus, Southern elephant seal virus, Tonate virus, Trocara virus, Una virus, Venezuelan horse encephalitis virus, Western equine encephalitis virus, and Wataroa virus. Generally, the genomes of such viruses encode non-structural proteins (e.g., replicons) and structural proteins (e.g., capsids and envelopes) that can be translated within the cytoplasm of host cells. Ross River virus, Sindbis virus, Semlik Forest virus (SFV), and Venezuelan encephalitis virus (VEEV) have all been used in the development of viral constructs for coding sequence delivery. Pseudotyped viruses can be formed by combining alphaviral envelope glycoproteins with retroviral capsids. Examples of alphaviral constructs can be found in U.S. Publications 20150050243, 20090305344, and 20060177819.The structures and methods for constructing them are incorporated herein by reference to each of the respective publications.
[0318] In some embodiments, the nucleic acid is a viral construct and may have a total number of nucleotides up to 10 kb. In some embodiments, the virus construct is approximately 1kb to 2kb, 1kb to 3kb, 1kb to 4kb, 1kb to 5kb, 1kb to 6kb, 1kb to 7kb, 1kb to 8kb, 1kb to 9kb, 1kb to 10kb, 2kb to 3kb, 2kb to 4kb, 2kb to 5kb, 2kb to 6kb, 2kb to 7kb, 2kb to 8kb, 2kb to 9kb, 2kb to 10kb, 3kb to 4kb, 3kb to 5kb, 3kb to 6kb, 3kb to 7kb, 3kb to 8kb, 3kb to 9kb The total number of nucleotides may be in the range of kb, approximately 3kb to approximately 10kb, approximately 4kb to approximately 5kb, approximately 4kb to approximately 6kb, approximately 4kb to approximately 7kb, approximately 4kb to approximately 8kb, approximately 4kb to approximately 9kb, approximately 4kb to approximately 10kb, approximately 5kb to approximately 6kb, approximately 5kb to approximately 7kb, approximately 5kb to approximately 8kb, approximately 5kb to approximately 9kb, approximately 5kb to approximately 10kb, approximately 6kb to approximately 7kb, approximately 6kb to approximately 8kb, approximately 6kb to approximately 9kb, approximately 6kb to approximately 10kb, approximately 7kb to approximately 8kb, approximately 7kb to approximately 9kb, approximately 7kb to approximately 10kb, approximately 8kb to approximately 9kb, approximately 8kb to approximately 10kb, or approximately 9kb to approximately 10kb.
[0319] In some embodiments, the nucleic acid is a lentiviral construct and may have a total number of nucleotides up to 8kb. In some examples, the lentiviral construct has approximately 1kb to 2kb, approximately 1kb to 3kb, approximately 1kb to 4kb, approximately 1kb to 5kb, approximately 1kb to 6kb, approximately 1kb to 7kb, approximately 1kb to 8kb, approximately 2kb to 3kb, approximately 2kb to 4kb, approximately 2kb to 5kb, approximately 2kb to 6kb, approximately 2kb to 7kb, approximately 2kb to 8kb, and approximately 3kb to 4kb. b. May have a total number of nucleotides of approximately 3kb to 5kb, 3kb to 6kb, 3kb to 7kb, 3kb to 8kb, 4kb to 5kb, 4kb to 6kb, 4kb to 7kb, 4kb to 8kb, 5kb to 6kb, 5kb to 7kb, 5kb to 8kb, 6kb to 8kb, 6kb to 7kb, or approximately 7kb to 8kb.
[0320] In some embodiments, the nucleic acid is an adenovirus construct and may have a total number of nucleotides up to 8kb. In some embodiments, the adenovirus construct has approximately 1kb to approximately 2kb, approximately 1kb to approximately 3kb, approximately 1kb to approximately 4kb, approximately 1kb to approximately 5kb, approximately 1kb to approximately 6kb, approximately 1kb to approximately 7kb, approximately 1kb to approximately 8kb, approximately 2kb to approximately 3kb, approximately 2kb to approximately 4kb, approximately 2kb to approximately 5kb, approximately 2kb to approximately 6kb, approximately 2kb to approximately 7kb, approximately 2kb to approximately 8kb, and approximately 3kb to approximately 4kb b. May have a total number of nucleotides in the range of approximately 3kb to 5kb, approximately 3kb to 6kb, approximately 3kb to 7kb, approximately 3kb to 8kb, approximately 4kb to 5kb, approximately 4kb to 6kb, approximately 4kb to 7kb, approximately 4kb to 8kb, approximately 5kb to 6kb, approximately 5kb to 7kb, approximately 5kb to 8kb, approximately 6kb to 7kb, approximately 6kb to 8kb, or approximately 7kb to 8kb.
[0321] Any nucleic acid described herein may further include a control sequence, selected from the group of control sequences, such as transcription start sequences, transcription termination sequences, promoter sequences, enhancer sequences, RNA splicing sequences, polyadenylation (poly(A)) sequences, Kozak consensus sequences, and / or additional untranslated regions capable of accommodating pre-transcriptional or post-transcriptional regulatory and / or regulatory elements. In some embodiments, the promoter may be a native promoter, a constitutive promoter, an inducible promoter, and / or a tissue-specific promoter. Non-limiting examples of control sequences are described herein.
[0322] In some embodiments, the disclosure further provides cargo components comprising one or more sequence elements, or their complements, selected from the group consisting of promoters, enhancers, silencers, insulators, transcription factors, translation factors, splice donors, splice acceptors, transcription terminators, translation start sites, translation termination sites, packaging signals, integration signals, inverse end sequences (ITRs), and any combination thereof. Exemplary sequence elements are described herein.
[0323] 2. Plasmid In some embodiments, the nucleic acid (e.g., cargo nucleic acid) is or contains a plasmid. In some embodiments, the nucleic acid is a DNA plasmid. In some embodiments, the nucleic acid is an RNA plasmid. In some embodiments, the plasmid can replicate independently in cells. In some embodiments, the plasmid may contain an origin of replication sequence. In some embodiments, the plasmid is a nanoplasmid.
[0324] The nucleic acids provided herein may be of different sizes. In some embodiments, the nucleic acids are plasmids and may include total lengths of up to about 1 kb, up to about 2 kb, up to about 3 kb, up to about 4 kb, up to about 5 kb, up to about 5 kb, up to about 6 kb, up to about 7 kb, up to about 8 kb, up to about 9 kb, up to about 10 kb, up to about 11 kb, up to about 12 kb, up to about 13 kb, up to about 14 kb, or up to about 15 kb. In some embodiments, nucleic acids are plasmids and may have a total length in the range of approximately 1kb to 2kb, 1kb to 3kb, 1kb to 4kb, 1kb to 5kb, 1kb to 6kb, 1kb to 7kb, 1kb to 8kb, 1kb to 9kb, 1kb to 10kb, 1kb to 11kb, 1kb to 12kb, 1kb to 13kb, 1kb to 14kb, or 1kb to 15kb.
[0325] In some embodiments, the Disclosure further provides plasmids comprising a cargo component containing one or more sequence elements, or complements thereof, selected from the group consisting of promoters, enhancers, silencers, insulators, transcription factors, translation factors, splice donors, splice acceptors, transcription terminators, translation start sites, translation termination sites, packaging signals, integration signals, inverse end sequences (ITRs), and any combination thereof. Exemplary sequence elements are described herein.
[0326] 3. RNA In certain embodiments, the compositions of this disclosure include nucleic acids. In some embodiments, the nucleic acids are RNA. In some embodiments, the nucleic acids include modified nucleic acids. In some embodiments, the nucleic acids include modified RNA. In particular, this disclosure states that the selection and combination of nucleic acids described herein affect the properties of cargo nucleic acids, such as stability and ionizability.
[0327] A. Modified RNA In certain embodiments, the compositions and / or nucleic acids of the present disclosure include modified nucleic acids, including modified RNA.
[0328] Modified nucleosides or nucleotides may be present in RNA, such as mRNA. For example, mRNA containing one or more modified nucleosides or nucleotides is referred to as “modified” RNA to represent the presence of one or more non-spontaneous and / or spontaneously occurring components or structures used in place of or in addition to standard A, G, C, and U residues. In some embodiments, modified RNA is synthesized with non-standard nucleosides or nucleotides and is referred to herein as “modified.”
[0329] Modified nucleosides and nucleotides may include one or more of the following: (i) alterations, e.g., substitution of one or both unbound phosphate oxygens and / or one or more bound phosphate oxygens in a phosphodiester backbone bond (exemplary backbone modification); (ii) alterations, e.g., substitution of a component of a ribose sugar, e.g., a 2' hydroxyl on a ribose sugar (exemplary sugar modification); (iii) large-scale substitution of a phosphate moiety by a "dephosphorylation" linker (exemplary backbone modification); (iv) alteration or substitution of a spontaneously occurring nucleic acid base by a non-standard nucleic acid base, etc. (exemplary base modification); (v) substitution or modification of the ribose-phosphate backbone (exemplary backbone modification); (vi) alteration of the 3' or 5' end of an oligonucleotide, e.g., removal, modification, or substitution of a terminal phosphate group, or conjugation of a part, cap, or linker (such 3' or 5' cap alterations may include sugar and / or backbone modifications); and (vii) alteration or substitution of sugars (exemplary sugar modifications). Certain embodiments include 5' end modifications to mRNA or nucleic acids. Certain embodiments include 3' end modifications to mRNA or nucleic acids. Modified RNA may include modifications at both the 5' and 3' ends. Modified RNA may include one or more modified residues at non-terminal positions. In certain embodiments, mRNA includes at least one modified residue.
[0330] Unmodified nucleic acids may be susceptible to degradation by, for example, intracellular nucleases or nucleases found in serum. For example, nucleases can hydrolyze phosphodiester bonds in nucleic acids. Thus, in one embodiment, the RNA (e.g., mRNA) described herein may contain one or more modified nucleosides or nucleotides to introduce stability, for example, against intracellular or serum-based nucleases. The term “innate immune response” includes cellular responses to exogenous nucleic acids, including single-stranded nucleic acids, that are involved in cytokine expression and release, particularly interferons, and induction of cell death.
[0331] Accordingly, in some embodiments, the RNA or nucleic acids in the compositions, preparations, nanoparticles, and / or nanomaterials of the Disclosure include at least one modification that confers increased or improved stability to the nucleic acid, including, for example, improved resistance to nuclease digestion in vivo. As used herein, the terms “modified” and “modified,” when such terms relate to nucleic acids provided herein, preferably include at least one modification that improves stability and makes the RNA or nucleic acid more stable (e.g., resistant to nuclease digestion) than wild-type or naturally occurring RNA or nucleic acid. As used herein, the terms “stable” and “stability,” when such terms relate to nucleic acids of the Invention, and in particular to RNA, refer to increased or improved resistance to degradation by nucleases (i.e., endonucleases or exonucleases) that are normally capable of degrading such RNA. Increased stability may include, for example, extending or increasing the time such RNA is present in target cells, tissues, objects, and / or cytoplasm due to reduced sensitivity to hydrolysis or other disruption by endogenous enzymes (e.g., endonucleases or exonucleases) or conditions within target cells or tissues. The stabilized RNA molecules provided herein exhibit longer half-lives compared to their naturally occurring, unmodified counterparts (e.g., the wild type of the mRNA). When the terms “modified” and “modified” are used in relation to mRNA in the LNP compositions disclosed herein, such terms refer to changes that increase or improve the translation of mRNA nucleic acids, including, for example, the inclusion of sequences that function in the initiation of protein translation (e.g., Kozak consensus sequences). (Kozak, M., Nucleic Acids Res 15(20):8125-48 (1987), the contents of which are incorporated herein by reference in their entirety).
[0332] In some embodiments, the RNA or nucleic acids in the compositions, preparations, nanoparticles, and / or nanomaterials disclosed herein have been chemically or biologically modified to make them more stable. Examples of modifications to RNA include base depletion (e.g., by deletion or substitution of one nucleotide with another) or base modification, such as chemical modification of bases. As used herein, the term “chemical modification” includes covalent modifications that introduce chemical processes different from those found in naturally occurring RNA, such as the introduction of modified nucleotides (e.g., nucleotide analogs, or inclusion of pendant groups not naturally found in such RNA molecules).
[0333] In some embodiments of skeletal modification, the phosphate group of the modified residue may be modified by substituting one or more oxygen atoms with different substituents. Furthermore, the modified residue, for example, the modified residue present in a modified nucleic acid, may include extensive substitution of the unmodified phosphate moiety by the modified phosphate group described herein. In some embodiments, skeletal modification of the phosphate skeleton may include changes that result in either an uncharged linker or a charged linker with an asymmetric charge distribution. Examples of modified phosphate groups include phosphorothioates, phosphoroselenates, boranophosphates, boranophosphate esters, hydrogen phosphonates, phosphoramidates, alkyl or aryl phosphonates, and phosphotryesters. The phosphorus atom in the unmodified phosphate group is achiral. However, the phosphorus atom may become chiral if one of the non-bridged oxygen atoms is replaced with one of the atoms or groups of atoms described above. The phosphorus atom at the chiral center may have either an "R" configuration (Rp as herein) or an "S" configuration (Sp as herein). The skeleton can also be modified by replacing the crosslinking oxygen (i.e., the oxygen linking the phosphate to the nucleoside) with nitrogen (crosslinking phosphoramidate), sulfur (crosslinking phosphorothioate), and carbon (crosslinking methylenephosphonate). The substitution may occur with any or both of the linking oxygens. The phosphate group can be replaced with a phosphorus-free conjugate in certain skeleton modifications. In some embodiments, the charged phosphate group can be replaced with a neutral moiety. Examples of moies that can replace the phosphate group include, but are not limited to, methylphosphonates, hydroxylaminos, siloxanes, carbonates, carboxymethyls, carbamates, amides, thioethers, ethylene oxide linkers, sulfonates, sulfonamides, thioformacetals, formacetals, oximes, methyleneiminos, methylenemethyliminos, methylenehydrazos, methylenedimethylhydrazos, and methyleneoxymethyliminos.
[0334] 4. Others In certain embodiments, the cargo nucleic acid comprises other components. In some embodiments, the cargo nucleic acid comprises one or more components, such as a promoter, enhancer, untranslated region (UTR), intra-sequence ribosome entry site (IRES), splice site, polyadenylated sequence, destabilization domain, reporter sequence or element, and / or other further sequences. In particular, this disclosure states that the selection and combination of one or more components described herein affects the properties of the nucleic acid, such as stability, expression, localization, and directivity.
[0335] A. Promoter In some embodiments, the nucleic acid includes a promoter. The term “promoter” refers to a DNA sequence recognized by an enzyme / protein that can promote and / or initiate the transcription of an operably linked gene (e.g., a nucleic acid encoding a cargo polypeptide). For example, a promoter typically refers to a nucleotide sequence to which, for example, RNA polymerase and / or any related factor can bind and from which transcription can be initiated. Thus, in some embodiments, the nucleic acid (e.g., placed within a delivery particle) includes a promoter operably linked to one of the non-limiting examples of promoters described herein.
[0336] In some embodiments, the promoter is an inducible promoter, a constitutive promoter, a mammalian cell promoter, a viral promoter, a chimeric promoter, an engineered promoter, a tissue-specific promoter, or any other type of promoter known in the art. In some embodiments, the promoter is an RNA polymerase II promoter, such as a mammalian RNA polymerase II promoter. In some embodiments, the promoter is an RNA polymerase III promoter including, but not limited to, the HI promoter, the human U6 promoter, the mouse U6 promoter, or the porcine U6 promoter. A promoter is generally one that can promote transcription in a target cell, tissue, organ, organoid, or organism. In some embodiments, the promoter is a mammalian cell-specific promoter.
[0337] A variety of promoters are known in the art and can be used herein. Non-limiting examples of promoters that can be used herein include human EFlα, human cytomegalovirus (CMV) (U.S. Patent No. 5,168,062, which is hereby incorporated by reference in its entirety), human ubiquitin C (UBC), mouse phosphoglycerate kinase 1, polyoma adenovirus, simian virus 40 (SV40), β-globin, β-actin, α-fetoprotein, γ-globin, β-interferon, γ-glutamyltransferase, mouse mammary tumor virus (MMTV), Rous sarcoma virus, rat insulin, glyceraldehyde-3-phosphate dehydrogenase, metallothionein II (MT II), amylase, cathepsin, MI muscarinic receptor, retroviral LTR (e.g., human T cell leukemia virus HTLV), AAV ITR, interleukin-2, collagenase, platelet-derived growth factor, adenovirus 5 E2, stromelysin, mouse MX gene, glucose-regulated protein (GRP78 and GRP94), α-2-macroglobulin, vimentin, MHC class I gene H-2 KExamples of promoters include b, HSP70, proliferin, tumor necrosis factor, thyroid-stimulating hormone a gene, immunoglobulin light chain, T cell receptor, HLA DQa and DQ, interleukin-2 receptor, MHC class II, MHC class II HLA-DRa, muscle creatine kinase, prealbumin (transthyretin), elastase I, albumin gene, c-fos, c-HA-ras, neuronal cell adhesion molecule (NCAM), H2B (TH2B) histone, rat growth hormone, human serum amyloid (SAA), troponin I (TN I), Duchenne muscular dystrophy, human immunodeficiency virus, and gibbon leukemia virus (GALV) promoter. Further examples of promoters are known in the art. See, for example, Lodish, Molecular Cell Biology, Freeman and Company, New York 2007, each incorporated herein by reference as a whole. In some embodiments, the promoter is the CMV pre-initial promoter. In some embodiments, the promoter is a CAG promoter or a CAG / CBA promoter. The term “constitutive” promoter refers to a nucleotide sequence that, when operably linked to a nucleic acid encoding a cargo polypeptide, causes RNA to be transcribed from the nucleic acid within the cell under almost all physiological conditions.
[0338] Examples of constitutive promoters include, but are not limited to, the retroviral Roussarcoma virus (RSV) LTR promoter, the cytomegalovirus (CMV) promoter (see, for example, Boshart et al, Cell 41:521-530, 1985, which is incorporated herein in whole by reference), the SV40 promoter, the dihydrofolate reductase promoter, the beta-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EFl-alpha promoter (Invitrogen).
[0339] Inducible promoters allow for the regulation of gene expression, which can be modulated by exogenously supplied compounds, environmental factors such as temperature, or the presence of specific physiological conditions, such as the acute phase, a particular differentiation state of cells, or cell restriction during replication. Inducible promoters and induction systems are available from a variety of commercial sources, including, but not limited to, Invitrogen, Clontech, and Ariad. Further examples of inducible promoters are known in the art.
[0340] Examples of inducible promoters regulated by exogenously supplied compounds include the zinc-inducible sheep metallothionein (MT) promoter, the dexamethasone (Dex)-inducible mouse mammary cancer virus (MMTV) promoter, the T7 polymerase promoter system (WO98 / 10088 incorporated herein as a whole by reference), the ecdysone insect promoter (see, for example, No et al, Proc. Natl. Acad Sci. US.A. 93:3346-3351, 1996 incorporated herein as a whole by reference), the tetracycline repression system (see, for example, Gossen et al, Proc. Natl. Acad Sci. US.A. 89:5547-5551, 1992 incorporated herein as a whole by reference), and the tetracycline-inducible gene expression system (see, for example, Gossen et al, Science 268:1766-1769, 1995, and Harvey et al, each incorporated herein as a whole by reference by reference). Examples include al., Curr. Opin. Chem. Biol. 2:512-518, 1998), RU486-inducible gene expression systems (see, for example, Wang et al., Nat. Biotech. 15:239-243, 1997 and Wang et al., Gene Ther. 4:432-441, 1997, each incorporated herein by reference as a whole), and rapamycin-inducible gene expression systems (see, for example, Magari et al., J Clin. Invest. 100:2865-2872, 1997, incorporated herein by reference as a whole).
[0341] The term "tissue-specific" promoter refers to a promoter that is active only in specific cell types and / or tissues (for example, transcription of a particular gene occurs only in cells that express transcriptional regulatory and / or regulatory proteins that bind to a tissue-specific promoter).
[0342] In some embodiments, the regulatory and / or control sequences confer tissue-specific gene expression capabilities. In some cases, the tissue-specific regulatory and / or control sequences bind to tissue-specific transcription factors that induce transcription in a tissue-specific manner.
[0343] In some embodiments, the nucleic acid provided comprises a promoter sequence selected from the CAG, CBA, CMV, or CB7 promoters.
[0344] B. Enhancer In some cases, a construct may include an enhancer sequence. The term “enhancer” refers to a nucleotide sequence that can increase the transcription level of the nucleic acid encoding the protein of interest (e.g., cargo polypeptide). Enhancer sequences (typically 50–1500 bp in length) generally increase the transcription level by providing an additional binding site to transcription-related proteins (e.g., transcription factors). In some embodiments, the enhancer sequence resides within an intron sequence. Unlike promoter sequences, enhancer sequences can act at a much greater distance from the transcription start site (e.g., compared to a promoter). Non-limiting examples of enhancers include RSV enhancers, CMV enhancers, and / or SV40 enhancers.
[0345] C. Adjacent untranslated regions, 5'UTR and 3'UTR In some embodiments, any of the nucleic acids described herein may include an untranslated region (UTR), e.g., a 5'UTR or a 3'UTR. The UTR of a gene is transcribed but not translated. The 5'UTR begins at the transcription start site and continues to the start codon, but does not include the start codon. The 3'UTR begins immediately after the stop codon and continues to the transcription termination signal. The regulatory and / or controllable properties of the UTR can be incorporated into any of the constructs, compositions, kits, or methods described herein to enhance or otherwise modulate the expression of cargo polypeptides.
[0346] The natural 5'UTR contains sequences involved in translation initiation. In some embodiments, the 5'UTR may contain sequences such as the Kozak sequence, which is commonly known to be involved in the process by which ribosomes initiate translation of many genes. The Kozak sequence has a consensus sequence CCR(A / G)CCAUGG, where R is a purine (A or G) three bases upstream of the start codon (AUG), followed by another "G" after the start codon. The 5'UTR is also known to form secondary structures involved in elongation factor binding.
[0347] In some embodiments, the 5'UTR is included in any of the constructs described herein. Non-limiting examples of 5'UTRs, including those derived from the following genes, namely albumin, serum amyloid A, apolipoprotein A / B / E, transferrin, alpha-fetoprotein, erythropoietin, and factor VIII, can be used to enhance the expression of nucleic acid molecules, such as mRNA.
[0348] 3'UTRs are known to have stretches of adenosine and uridine (RNA form) or thymidine (DNA form) embedded within them. These AU-rich signatures are particularly widespread in genes with high turnover rates. Based on their sequence and functional characteristics, AU-rich elements (AREs) can be classified into three classes (see, for example, Chen et al., Mal. Cell. Biol. 15:5777-5788, 1995 and Chen et al., Mal. Cell. Biol. 15:2010-2018, 1995, each of which is incorporated herein by reference as a whole). Class I AREs contain several dispersed copies of the AUUUA motif within the U-rich region. For example, the mRNAs of c-Myc and MyoD contain Class I AREs. Class II AREs have two or more overlapping UUAUUUA(U / A)(U / A) nonomers. GM-CSF and TNF-alpha mRNA are examples of molecules containing class II AREs. Class III AREs are less clearly defined. These U-rich regions do not contain the AUUUA motif, and two well-studied examples of this class are c-Jun and myogenin mRNA.
[0349] Most proteins that bind to AREs are known to destabilize messengers, but members of the ELAV family, particularly HuR, have been reported to increase mRNA stability. HuR binds to all three classes of AREs. Manipulating the HuR-specific binding site within the 3'UTR of nucleic acid molecules leads to HuR binding and, consequently, in vivo message stabilization.
[0350] In some embodiments, the stability of mRNA encoding cargo polypeptides can be regulated by introducing, removing, or modifying the ARE in the 3'UTR. In other embodiments, the ARE can be removed or mutated to increase intracellular stability and, consequently, increase the translation and production of cargo polypeptides.
[0351] In other embodiments, non-ARE sequences may be incorporated into the 5' or 3' UTR. In some embodiments, an intron or a portion of an intron sequence may be incorporated into a polynucleotide adjacency region in any of the constructs, compositions, kits, and methods provided herein. Incorporation of an intron sequence may increase protein production and mRNA levels.
[0352] D. Intra-sequence ribosome entry sites (IRES) In some embodiments, the nucleic acid containing the cargo component may include an intrasequence ribosome entry site (IRES). The IRES forms a complex secondary structure that initiates translation from any position with mRNA immediately downstream of the IRES (see, for example, Pelletier and Sonenberg, Mal. Cell. Biol. 8(3):1103-1112, 1988).
[0353] For example, there are several IRES sequences known to those skilled in the art, including sequences derived from foot-and-mouth disease virus (FMDV), encephalomyocarditis virus (EMCV), human rhinovirus (HRV), cricket paralysis virus, human immunodeficiency virus (HIV), hepatitis A virus (HAV), hepatitis C virus (HCV), and poliovirus (PV) (see, for example, Alberts, Molecular Biology of the Cell, Garland Science, 2002, and Hellen et al., Genes Dev. 15(13):1593-612, 2001, each of which is incorporated herein by reference as a whole).
[0354] In some embodiments, the IRES sequence incorporated into a cargo polypeptide, or a construct encoding the C-terminal portion of a cargo polypeptide, is the foot-and-mouth disease virus (FMDV) 2A sequence. The FMDV 2A sequence is a small peptide (approximately 18 amino acids long) that has been shown to mediate polyprotein cleavage (see, for example, Ryan, MD et al., EMBO 4:928-933, 1994, Mattion et al., J Virology 70:8124-8127, 1996, Furler et al., Gene Therapy 8:864-873, 2001, and Halpin et al., Plant Journal 4:453-459, 1999, each of which is incorporated herein by reference as a whole). The cleavage activity of the 2A sequence has been previously demonstrated in artificial systems including plasmids and gene therapy constructs (AAV and retroviruses) (see, for example, Ryan et al., EMBO 4:928-933, 1994, Mattion et al., J Virology 70:8124-8127, 1996, Furler et al., Gene Therapy 8:864-873, 2001, and Halpin et al., Plant Journal 4:453-459, 1999, de Felipe et al., Gene Therapy 6:198-208, 1999, de Felipe et al., Human Gene Therapy II:1921-1931, 2000, and Klump et al., Gene Therapy 8:811-817, 2001, each of which is incorporated herein by reference as a whole).
[0355] IRESs may be used in the delivery particles described herein. In some embodiments, the nucleic acid encoding the C-terminal portion of the cargo polypeptide may include an intra-sequence ribosome entry site (IRES) of a polynucleotide. In some embodiments, the IRES may be part of a composition comprising multiple nucleic acids. In some embodiments, the IRES is used to produce multiple cargo polypeptides from a single gene transcript.
[0356] E. Splice site In some embodiments, any of the nucleic acids provided herein may include splice donor and / or splice acceptor sequences that function during RNA processing occurring during transcription. In some embodiments, the splice site is involved in trans-splicing.
[0357] F. Polyadenylated sequence In some embodiments, the constructs provided herein may include a polyadenylation (poly(A)) signal sequence. Most neonatal eukaryotic mRNAs have a poly(A) tail at their 3' end, which is added during a complex process involving cleavage of the primary transcript and a conjugated polyadenylation reaction driven by the poly(A) signal sequence (see, e.g., Proudfoot et al., Cell 108:501-512, 2002, incorporated herein by reference in whole). The poly(A) tail confers stability and transmissibility to the mRNA (see, e.g., Molecular Biology of the Cell, Third Edition by B. Alberts et al., Garland Publishing, 1994, incorporated herein by reference in whole). In some embodiments, the poly(A) signal sequence is positioned 3' relative to the coding sequence.
[0358] As used herein, “polyadenylation” refers to the covalent bonding of a polyadenylyl moiety, or a modified variant thereof, to a messenger RNA molecule. In eukaryotes, most messenger RNA (mRNA) molecules are polyadenylated at their 3' end. The 3' poly(A) tail is a long sequence of adenine nucleotides (e.g., 50, 60, 70, 100, 200, 500, 1000, 2000, 3000, 4000, or 5000) that is added to premRNA through the action of the enzyme polyadenylate polymerase. In some embodiments, the poly(A) tail is added to a transcript containing a specific sequence, e.g., a poly(A) signal. The poly(A) tail and associated proteins help protect mRNA from degradation by exonucleases. Polyadenylation also plays a role in transcription termination, mRNA export from the nucleus, and translation. Polyadenylation typically occurs in the nucleus immediately after the transcription of DNA to RNA, but can later occur in the cytoplasm. After transcription is complete, the mRNA strand is cleaved by the action of an endonuclease complex associated with RNA polymerase. The cleavage site is usually characterized by the presence of the base sequence AAUAAA near the cleavage site. After mRNA is cleaved, an adenosine residue is added to the free 3' end of the cleavage site.
[0359] As used herein, “poly(A) signal sequence” or “polyadenylation signal sequence” is a sequence that induces endonuclease cleavage of mRNA and the addition of a series of adenosines to the 3' end of the cleaved mRNA.
[0360] Bovine growth hormone (bGH) (see, for example, Woychik et al., Proc. Natl. Acad Sci. US. A. 81(13):3944-3948, 1984, and U.S. Patent No. 5,122,458, each incorporated herein by reference as a whole), mouse β-globin, mouse α-globin (see, for example, Orkin et al., EMBO J 4(2):453-456, 1985, and Thein et al., Blood 71(2):313-319, 1988, each incorporated herein by reference as a whole), human collagen, polyomavirus (see, for example, Batt et al., Mal. Cell Biol. 15(9):4783-4790, 1995, each incorporated herein by reference as a whole), herpes simplex virus thymidine kinase gene (HSV) There are several poly(A) signal sequences that can be used, including those derived from TK), the IgG heavy chain gene polyadenylation signal (US2006 / 0040354, incorporated herein by reference as a whole), human growth hormone (hGH) (see, for example, Szymanski et al., Mal. Therapy 15(7):1340-1347, 2007, incorporated herein by reference as a whole), and a group consisting of poly(A) sites of SV40, for example, late and early poly(A) sites of SV40 (see, for example, Schek et al., Mal. Cell Biol. 12(12):5386-5393, 1992, incorporated herein by reference as a whole).
[0361] The poly(A) signal sequence may be AATAAA. The AATAAA sequence may be replaced with other hexanucleotide sequences homologous to AATAAA and capable of signaling polyadenylation, including ATTAAA, AGTAAA, CATAAA, TATAAA, GATAAA, ACTAAA, AATATA, AAGAAA, AATAAT, AAAAAA, AATGA, AATCA, AACAAA, AATCA, AATAC, AATGA, AATTA, or AATAG (see, for example, WO06 / 12414, which is incorporated herein by reference in its entirety).
[0362] In some embodiments, the poly(A) signal sequence may be a synthetic polyadenylation site (see, for example, the Promega pCl-neo expression construct based on Levitt el al, Genes Dev. 3(7):1019-1025, 1989, which is incorporated herein by reference in whole). In some embodiments, the poly(A) signal sequence is a polyadenylation signal of soluble neuropilin-1 (sNRP) (AAATAAAATACGAAATG) (see, for example, WO05 / 073384, which is incorporated herein by reference in whole). In some embodiments, the poly(A) signal sequence includes or consists of the poly(A) site of SV40.
[0363] G. Destabilization Domain In some embodiments, any of the nucleic acids provided herein may optionally include a sequence encoding a destabilization domain for temporal control of protein expression ("destabilization sequence"). Non-limiting examples of destabilization sequences include the FK506 sequence, a sequence encoding a dihydrofolate reductase (DHFR) sequence, or other exemplary destabilization sequences.
[0364] In the absence of a stabilizing ligand, protein sequences operably linked to the destabilizing sequence are degraded by ubiquitination. In contrast, in the presence of a stabilizing ligand, protein degradation is inhibited, thereby leading to the active expression of protein sequences operably linked to the destabilizing sequence. As a positive regulation of protein expression stabilization, protein expression can be detected by conventional means including enzymes, radiation, colorimetric analysis, fluorescence, or other spectroscopic assays, fluorescence-activated cell classification (FACS) assays, and immunoassays (e.g., enzyme-linked immunosorbent assays (ELISA), radioimmunoassays (RIA), and immunohistochemistry).
[0365] Further examples of destabilizing sequences are known in the art. In some embodiments, the destabilizing sequence is the FK506 and rapamycin-binding protein (FKBP12) sequence, and the stabilizing ligand is Shield-1 (Shld1) (see, for example, Banaszynski et al. (2012) Cell 126(5):995-1004, which is incorporated herein by reference in whole). In some embodiments, the destabilizing sequence is the DHFR sequence, and the stabilizing ligand is trimethoprim (TMP) (see, for example, Iwamoto et al. (2010) Chem Biol 17:981-988, which is incorporated herein by reference in whole).
[0366] In some embodiments, the destabilizing sequence is an FKBP12 sequence, and the presence of nucleic acid containing the FKBP12 gene in the target cell (e.g., target cell (e.g., glial cells, liver cells, tumor cells, etc.)) is detected by Western blotting. In some embodiments, the destabilizing sequence can be used to verify the time-specific activity of the delivery particles described herein.
[0367] H. Reporter array or element In some embodiments, the nucleic acids provided herein may optionally include a sequence encoding a reporter polypeptide and / or protein ("reporter sequence"). Non-limiting examples of reporter sequences include DNA sequences encoding beta-lactamase, beta-galactosidase (LacZ), alkaline phosphatase, thymidine kinase, green fluorescent protein (GFP), red fluorescent protein, mCherry fluorescent protein, yellow fluorescent protein, chloramphenicol acetyltransferase (CAT), and luciferase. Further examples of reporter sequences are known in the Art. When associated with regulatory elements that drive their expression, reporter sequences may provide a signal detectable by conventional means, including enzyme, radiation, colorimetric analysis, fluorescence, or other spectroscopic assays, fluorescence-activated cell classification (FACS) assays, and immunoassays (e.g., enzyme-linked immunosorbent assay (ELISA), radioimmunoassay (RIA), and immunohistochemistry).
[0368] In some embodiments, the reporter sequence is the LacZ gene, and the presence of a construct containing the LacZ gene in mammalian cells (e.g., cells of interest (e.g., glial cells, liver cells, tumor cells, etc.)) is detected by an assay for beta-galactosidase activity. If the reporter is a fluorescent protein (e.g., green fluorescent protein) or luciferase, the presence of a construct containing the fluorescent protein or luciferase in mammalian cells (e.g., cells of interest (e.g., glial cells, liver cells, tumor cells, etc.)) can be measured by fluorescence techniques (e.g., fluorescence microscopy or FACS) or photogeneration using an illuminometer (e.g., spectrophotometer or IVIS imaging device). In some embodiments, the reporter sequence can be used to verify the tissue-specific targeting ability and tissue-specific promoter modulation and / or regulatory activity of any of the constructs described herein.
[0369] In some embodiments, the reporter sequence is a FLAG tag (e.g., a 3xFLAG tag), and the presence of a construct with the FLAG tag in mammalian cells (e.g., cells of interest (e.g., glial cells, liver cells, tumor cells, etc.)) is detected by protein binding or detection assays (e.g., Western blotting, immunohistochemistry, radioimmunoassay (RIA), mass spectrometry).
[0370] I. Further arrangements In some embodiments, the nucleic acids of the Disclosure may comprise T2A elements or sequences. In some embodiments, the nucleic acids of the Disclosure may comprise one or more cloning sites. In some such embodiments, the cloning sites may not be completely removed before production for administration to a subject. In some embodiments, the cloning sites may have a functional role, such as as a linker sequence or as a portion of a Kozak site. As will be understood by those skilled in the art, the cloning sites may undergo significant alteration of the primary sequence while retaining their desired function.
[0371] J. Inverted terminal sequence (ITR) In some embodiments, the delivery particles are AAV delivery particles. The AAV-derived nucleic acid sequence of the construct typically includes cis-acting 5' and 3' ITRs (see, for example, B.J. Carter, in “Handbook of Parvoviruses”, ed., P. Tijsser, CRC Press, pp. 155-168 (1990), which is incorporated herein by reference as a whole). Generally, the ITRs can form hairpins. The ability to form hairpins contributes to the self-priming ability of the ITRs, enabling primase-independent synthesis of the second DNA strand. The ITRs can also contribute to the efficient capsid formation of the AAV construct in the AAV delivery particles.
[0372] The rAAV delivery particles of this disclosure (e.g., AAV2 delivery particles) may contain nucleic acids comprising a cargo component encoding a cargo polypeptide and associated elements, with the ITR sequence of the AAV positioned at 5' and 3'. In some embodiments, the ITR is or comprises about 145 nucleic acids. In some embodiments, all or substantially all of the sequence encoding the ITR is used. The ITR sequence of the AAV may be obtained from any known AAV, including currently identified mammalian AAV types. In some embodiments, the ITR is the ITR of AAV2.
[0373] Examples of construct molecules used in this disclosure are “cis-acting” constructs containing a transgene, where the ITR sequences of AAV are positioned 5' or “left” and 3' or “right” next to the selected transgene sequence and associated regulatory factors. The 5' and “left” notations refer to the position of the ITR sequence relative to the entire construct, read from left to right in the sense direction. For example, in some embodiments, the 5' or “left” ITR is the ITR closest to the promoter (opposite the polyadenylation sequence) for a given construct, if the construct is depicted linearly in the sense direction. Simultaneously, the 3' and “right” notations refer to the position of the ITR sequence relative to the entire construct, read from left to right in the sense direction. For example, in some embodiments, the 3' or “right” ITR is the ITR closest to the polyadenylation sequence (opposite the promoter sequence) for a given construct, if the construct is depicted linearly in the sense direction. The ITRs provided herein are shown in 5' to 3' order according to the sense strand. Therefore, those skilled in the art will understand that when converting from sense orientation to antisense orientation, a 5' or "left" oriented ITR can also be depicted as a 3' or "right" ITR. Furthermore, converting a given sense ITR sequence (e.g., a 5' / left AAV ITR) to an antisense sequence (e.g., a 3' / right ITR sequence) is well within the scope of the skills possessed by those skilled in the art. Those skilled in the art will understand how to modify a given ITR sequence for use as either a 5' / left or 3' / right ITR, or its antisense type.
[0374] VI. Delivery particles In particular, this disclosure provides delivery particles. In some embodiments, the delivery particles are viral particles, lipid-based particles [(e.g., cell-produced or non-cell-produced), lipid nanoparticles (LNPs), liposomes, micelles, extracellular vesicles (e.g., exosomes, microparticles, etc.)], polymer-based particles (e.g., PGLA), polysaccharide-based particles, etc. In some embodiments, the delivery particles described herein include nucleic acids. In some embodiments, the nucleic acids described herein are disposed within the delivery particles. In some embodiments, the nucleic acids described herein associate with the surface of the delivery particles (e.g., covalently or non-covalently). In some embodiments, the nucleic acids include, among other things, cargo components encoding cargo polypeptides, which, if expressed, are expressed on the surface of the delivery particles.
[0375] i. Billion: In particular, this disclosure provides virions comprising nucleic acids and capsids as described herein. In some embodiments, the virion is a delivery particle comprising nucleic acids comprising a cargo polypeptide or a characteristic portion thereof as described herein, and a capsid as described herein. Exemplary delivery particles are AAV delivery particles. Exemplary delivery particles are lentiviral delivery particles. However, other delivery particles may be used.
[0376] In some embodiments, the delivery particle is an AAV delivery particle. The AAV delivery particle comprises a nucleic acid containing a cargo polypeptide or a characteristic portion thereof as described herein, and a capsid as described herein. In some embodiments, the AAV delivery particle may be described as having a serotype, which is a description of a construct strain and a capsid strain. For example, in some embodiments, the AAV delivery particle may be described as AAV2, in which case the particle has a construct comprising an AAV2 capsid and a characteristic AAV2 inverse terminal sequence (ITR). In some embodiments, the AAV delivery particle may be described as a pseudotype, in which case the capsid and construct are derived from different AAV strains, for example, AAV2 / 9 refers to an AAV delivery particle comprising a construct utilizing the AAV2 ITR and the AAV9 capsid.
[0377] 1. AAV structures This disclosure provides nucleic acids comprising a cargo polypeptide or a cargo component encoding a characteristic portion thereof. In some embodiments described herein, nucleic acids comprising a cargo polypeptide or a cargo component encoding a characteristic portion thereof may be placed within AAV delivery particles.
[0378] In some embodiments, the nucleic acid comprises one or more components derived from or modified from a naturally occurring AAV genome construct. In some embodiments, the sequence derived from the AAV construct is the AAV1 construct, AAV2 construct, AAV3 construct, AAV4 construct, AAV5 construct, AAV6 construct, AAV7 construct, AAV8 construct, AAV9 construct, AAV2.7m8 construct, AAV8BP2 construct, AAV293 construct, AAV.DJ construct, or AAV Anc80 construct. In some embodiments, the rAAV Anc80 capsid is the rAAV Anc80L65 capsid. Further exemplary AAV constructs that may be used herein are known in the art (see, for example, Kanaan et al., Mol.Ther.Nucleic Acids 8:184-197, 2017; Li et al., Mol.Ther. 16(7):1252-1260, 2008; Adachi et al., Nat.Commun. 5:3075, 2014; Isgrig et al., Nat.Commun. 10(1):427, 2019; and Gao et al., J.Virol. 78(12):6381-6388, 2004, each incorporated herein by reference as a whole).
[0379] In some embodiments, the nucleic acid provided comprises, for example, a cargo component encoding a cargo polypeptide, one or more regulatory and / or control sequences, and optionally 5' and 3' AAV-derived inverted end sequences (ITRs). In some embodiments in which 5' and 3' AAV-derived ITRs are used, the polynucleotide construct may be referred to as a recombinant AAV (rAAV) construct. In some embodiments, the rAAV construct provided is packaged in an AAV capsid to form an AAV delivery particle.
[0380] In some embodiments, the AAV-derived sequence (included in the polynucleotide construct) typically includes cis-active 5' and 3' ITR sequences (see, e.g., B.J. Carter, in “Handbook of Parvoviruses,” ed., P. Tijsser, CRC Press, pp. 155-168, 1990, each incorporated herein by reference as a whole). A typical AAV2-derived ITR sequence is approximately 145 nucleotides long. In some embodiments, at least 80% (e.g., at least 85%, at least 90%, or at least 95%) of a typical ITR sequence is incorporated into the constructs provided herein. The ability to modify these ITR sequences is within the scope of the skills possessed by those skilled in the art (see, e.g., Sambrook et al., “Molecular Cloning. A Laboratory Manual”, 2nd ed., Cold Spring Harbor Laboratory, New York, 1989, and K. Fisher et al., J Virol. 70:520-532, 1996, each incorporated herein by reference as a whole). In some embodiments, the ITR sequence of the AAV is positioned at 5' and 3' alongside any of the code sequences and / or constructs described herein. The ITR sequence of the AAV can be obtained from any known AAV, including currently identified AAV types.
[0381] In some embodiments, the nucleic acids described in accordance with this disclosure and in patterns known in the art (see, for example, Asokan et al., Mol. Ther. 20:699-7080, 2012, which are incorporated herein in whole by reference) typically consist of a coding sequence or a portion thereof, at least one and / or regulatory sequence, and optionally 5' and 3' inverted end sequences (ITRs) of AAV. In some embodiments, the construct provided may be packaged in a capsid to create AAV delivery particles. The AAV delivery particles may be delivered to selected target cells. In some embodiments, the construct provided includes further optional coding sequences that are heterogeneous nucleic acid sequences (e.g., repressive nucleic acid sequences) to the nucleic acid sequence encoding the polypeptide, protein, functional RNA molecule (e.g., miRNA, miRNA inhibitor), or other gene product of interest. In some embodiments, the nucleic acid coding sequence is operably ligated to and / or is a regulatory component in such a way that it enables transcription, translation, and / or expression of the coding sequence in cells of the target tissue.
[0382] In some embodiments, the nucleic acid is rAAV nucleic acid. In some embodiments, the rAAV nucleic acid may include at least 500 bp, at least 1 kb, at least 1.5 kb, at least 2 kb, at least 2.5 kb, at least 3 kb, at least 3.5 kb, at least 4 kb, or at least 4.5 kb. In some embodiments, the AAV construct may include up to 7.5 kb, up to 7 kb, up to 6.5 kb, up to 6 kb, up to 5.5 kb, up to 5 kb, up to 4.5 kb, up to 4 kb, up to 3.5 kb, up to 3 kb, or up to 2.5 kb. In some embodiments, the AAV structure may include approximately 1kb to approximately 2kb, approximately 1kb to approximately 3kb, approximately 1kb to approximately 4kb, approximately 1kb to approximately 5kb, approximately 2kb to approximately 3kb, approximately 2kb to approximately 4kb, approximately 2kb to approximately 5kb, approximately 3kb to approximately 4kb, approximately 3kb to approximately 5kb, or approximately 4kb to approximately 5kb.
[0383] Any nucleic acid described herein may further include regulatory and / or control sequences, such as transcription start sequences, transcription termination sequences, promoter sequences, enhancer sequences, RNA splicing sequences, polyadenylation (poly(A)) sequences, Kozak consensus sequences, and / or any combination thereof. In some embodiments, the promoter may be a native promoter, a constitutive promoter, an inducible promoter, and / or a tissue-specific promoter. Non-limiting examples of control sequences are described herein.
[0384] 2. AAV capsid This disclosure provides one or more nucleic acids to be coupled with an AAV capsid. In some embodiments, the AAV capsid is derived from, or derived from, an AAV capsid of AAV2, 3, 4, 5, 6, 7, 8, 9, 10, DJ, PHP-B, rh8, rh10, rh39, rh43, or Anc80 serotypes, or a hybrid thereof of one or more. In some embodiments, the AAV capsid is derived from an AAV ancestral serotype.
[0385] As provided herein, any combination of AAV capsid and AAV nucleic acid (e.g., including the ITR of AAV) may be used in the recombinant AAV (rAAV) particles of this disclosure. For example, the ITR of wild-type or variant AAV2 and an Anc80 capsid, the ITR of wild-type or variant AAV2 and an AAV6 capsid, etc. In some embodiments of this disclosure, the AAV delivery particles consist entirely of AAV2 components (e.g., the capsid and ITR are the AAV2 serotype). In some embodiments, the AAV delivery particles are AAV2 / 6, AAV2 / 8, or AAV2 / 9 particles (e.g., AAV6, AAV8, or AAV9 capsids with an AAV construct having the ITR of AAV2).
[0386] ii. Lipid-based delivery particles In particular, the present disclosure provides compositions, preparations, and / or delivery particles (e.g., lipid-based delivery particles) comprising lipids. In some embodiments, the lipid-based delivery particles are produced by cells. In some embodiments, the lipid-based delivery particles are not produced by cells. The present invention provides lipid-based delivery particles that can be of various types. In some embodiments, the lipid-based delivery particles may be lipid nanoparticles (LNPs). In some embodiments, the lipid-based delivery particles may be liposomes. In some embodiments, the lipid-based delivery particles may be micelles. In some embodiments, the lipid-based delivery particles may be extracellular vesicles (e.g., exosomes).
[0387] 1. Lipid nanoparticles In some embodiments, the Disclosure provides compositions, preparations, and / or delivery particles comprising lipid nanoparticles. In some embodiments, the lipid nanoparticles comprise one or more components. In some embodiments, the lipid nanoparticles comprise one or more components, such as compounds, ionic lipids, sterols, conjugate-linker lipids, and phospholipids. In particular, the Disclosure states that the selection and combination of one or more components described herein affects the properties of the lipid nanoparticles, such as diameter, pKa, stabilization, and ionizability.
[0388] In particular, this disclosure describes how the selection and combination of one or more components described herein affects the functional activity of lipid nanoparticles, such as directivity, stabilization, and drug delivery effectiveness. For example, this disclosure describes how combinations of components may be well suited to the delivery of nucleic acids containing cargo described herein (e.g., nucleic acids encoding cargo polypeptides). In some embodiments, the cargo includes RNA. In some embodiments, the cargo includes DNA.
[0389] In some embodiments, the lipid nanoparticles comprise one or more compounds described herein. In some embodiments, the lipid nanoparticles comprise one or more ionic lipids described herein. In some embodiments, the lipid nanoparticles comprise one or more sterols described herein. In some embodiments, the lipid nanoparticles comprise one or more conjugate-linker lipids described herein. In some embodiments, the lipid nanoparticles comprise one or more phospholipids described herein.
[0390] A. Ionic lipids In particular, this disclosure describes compositions, preparations, delivery particles, and / or methods comprising one or more ionic lipids de...
Claims
1. nucleic acids, (a) A nucleotide sequence that encodes a cargo polypeptide, or a cargo component that contains such a sequence, (b) A barcode component in which the nucleotide sequence is a sequence that encodes a peptide barcode, (i) The peptide barcode has an amino acid length in the range of 1 to 100, 5 to 50, 8 to 25, 9 to 25, or 9 to 15, and (ii) The barcode component is characterized by being confirmed to specifically bind to a particular group of polypeptide binders within a set of binders, The nucleic acid wherein the cargo component is operably linked to the barcode component.
2. The nucleic acid according to claim 1, wherein the cargo component further comprises one or more sequence elements, or complements thereof, selected from the group consisting of promoters, enhancers, silencers, insulators, transcription factors, translation factors, splice donors, splice acceptors, transcription terminators, translation start sites, translation termination sites, packaging signals, integration signals, and any combination thereof.
3. The nucleic acid according to any of the prior claims, wherein the cargo component includes an intra-sequence ribosome entry site (IRES).
4. The nucleic acid according to any of the prior claims, wherein the cargo component further encodes a cleavable portion (e.g., a self-cleaving peptide (e.g., a 2A peptide)).
5. The nucleic acid according to any of the prior claims, wherein the nucleic acid is DNA or comprises DNA.
6. The nucleic acid according to any of the prior claims, wherein the nucleic acid is RNA or comprises RNA.
7. The nucleic acid according to claim 6, wherein the cargo component further comprises one or more of the following: a capping portion, a 5' untranslated region (UTR), a 3'UTR, a polyadenylated (poly-A) tail, or complements thereof, or any combination thereof.
8. The nucleic acid according to any of the prior claims, wherein the cargo polypeptide further comprises a localization moiety.
9. The nucleic acid according to claim 8, wherein the localization portion is selected from the group consisting of a secretory signal and an intracellular localization portion.
10. The nucleic acid according to any of the prior claims, wherein the cargo polypeptide further comprises an intermediate or procomponent.
11. The nucleic acid according to any of the prior claims, wherein the cargo polypeptide further includes a tag portion.
12. The nucleic acid according to any of the prior claims, wherein the cargo polypeptide further comprises a ligand-binding portion (e.g., a shuttle portion).
13. The nucleic acid according to any of the prior claims, wherein the cargo polypeptide further comprises a stability modification portion.
14. The nucleic acid according to any of the prior claims, wherein the cargo polypeptide further comprises a masking portion.
15. The nucleic acid according to any of the prior claims, wherein the cargo polypeptide further comprises an allosteric regulatory moiety.
16. The nucleic acid according to any one of claims 8 to 15, wherein the localization portion, the tag portion, the ligand-binding portion, the stability modification portion, the masking portion, or the allosteric regulation portion is cleavable.
17. The nucleic acid according to any of the prior claims, wherein the cargo polypeptide is a wild-type (e.g., spontaneously occurring) polypeptide or comprises the same.
18. The nucleic acid according to any one of claims 1 to 16, wherein the cargo polypeptide is a variant polypeptide (e.g., a variant cargo polypeptide) or comprises the same.
19. The nucleic acid according to claim 18, wherein the variant polypeptide is a variant of a reference polypeptide, and the reference polypeptide is a wild-type (e.g., spontaneously occurring) polypeptide or comprises the same.
20. A nucleic acid according to any of the prior claims, which is disposed within a delivery particle.
21. A nucleic acid according to any of the prior claims, disposed on the surface of a delivery particle.
22. The nucleic acid according to any one of the prior claims, wherein the encoded peptide barcode has an amino acid sequence selected from the group consisting of SEQ ID NOs: 5347 to 8398.
23. The nucleic acid according to any one of the prior claims, wherein the encoded peptide barcode is encoded by a nucleic acid sequence selected from the group consisting of sequence numbers 1148 to 4199.
24. The nucleic acid according to any one of the prior claims, wherein the encoded peptide barcode has a length of 8 to 25 amino acids.
25. The nucleic acid according to any one of the prior claims, wherein the encoded peptide barcode has a length of 10 amino acids.
26. The nucleotide sequence of the barcode component is in the order of 5' to 3' or 3' to 5', (a) First invariant sequence (e.g., linker sequence or payload sequence), (b) A variant sequence having a length of at least 9 nucleotides, and (c) The nucleic acid according to any one of the prior claims, comprising one or more of the following: a second invariant sequence (e.g., a linker sequence, a stop codon, or a payload sequence).
27. The nucleic acid according to claim 26, wherein the variant sequence has a length of at least 15, 24, 27, 45, 150, or 300 nucleotides.
28. The nucleotide sequence of the aforementioned barcode component is further, (d) A sequence that codes for a short helix motif, (e) Sequences that code for a mutated motif, (f) The nucleic acid according to any one of the prior claims, comprising one or more of the following: an invariant sequence that links the cargo component to the barcode component.
29. The nucleic acid according to any one of the prior claims, wherein each polypeptide binder in the group of polypeptide binders has an amino acid sequence selected from the group consisting of SEQ ID NOs: 4200 to 5346.
30. The nucleic acid according to any one of the prior claims, wherein each polypeptide binder in the group of polypeptide binders is encoded by a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 1 to 1147.
31. The nucleic acid according to any one of the prior claims, wherein each polypeptide binder is expressed on a phage.
32. The nucleic acid according to claim 31, wherein the phage is selected from the group consisting of M13, T4, T7, lambda, and filamentous phages.
33. The nucleic acid according to claim 31, wherein the phage is M13.
34. A nucleic acid according to any one of the prior claims, encoding a barcoded cargo polypeptide, wherein the barcoded cargo polypeptide, or a characteristic portion thereof, is expressed on the surface of a delivery particle (e.g., a viral particle, a lipid-based particle [e.g., a cell-produced or non-cell-produced lipid nanoparticle (LNP), liposome, micelle, extracellular vesicle (e.g., exosome, microparticle, etc.)], a polymer-based particle (e.g., PGLA), a polysaccharide-based particle, etc.).
35. The nucleic acid according to any one of the prior claims, wherein the cargo component, or a portion thereof, is codon-optimized.
36. A library comprising multiple nucleic acids, wherein each nucleic acid is a nucleic acid described in any one of the prior claims.
37. A plurality of delivery particles, wherein one or more of the delivery particles contain the nucleic acid described in any one of claims 1 to 35.
38. A plurality of delivery particles according to claim 37, wherein the nucleic acid in each of the delivery particles is the same.
39. The plurality of delivery particles according to claim 37, wherein the delivery particles comprise at least two different nucleic acids.
40. The plurality of delivery particles according to claim 39, wherein at least two different nucleic acids comprise different cargo components.
41. The plurality of delivery particles according to claim 39 or 40, wherein the delivery particles comprise cargo components that encode at least two different cargo polypeptides.
42. A plurality of delivery particles according to any one of claims 39 to 41, wherein the cargo polypeptide is a variant of a reference polypeptide, and the reference polypeptide is a wild-type (e.g., spontaneously occurring) polypeptide or comprises the same.
43. The plurality of delivery particles according to claim 42, wherein the variants include amino acid sequences that are at least 70% identical to each other (for example, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, and at least 99% identical to each other).
44. The plurality of delivery particles according to any one of claims 37 to 43, wherein the delivery particle comprises one or more associated (e.g., covalently or non-covalently) targeting portions.
45. The plurality of delivery particles according to claim 44, wherein one or more of the targeting portions are of the same type.
46. The plurality of delivery particles according to claim 44, wherein one or more of the targeting portions are of different types.
47. A plurality of delivery particles according to any one of claims 37 to 46, which are substantially the same type of delivery particle.
48. A plurality of delivery particles according to any one of claims 37 to 46, comprising two or more types of delivery particles.
49. A plurality of delivery particles according to any one of claims 37 to 48, comprising viral particles, lipid-based particles [e.g., lipid nanoparticles (LNPs), liposomes, micelles, extracellular vesicles (e.g., exosomes, microparticles, etc.), which are produced by or not produced by cells], polymer-based particles (e.g., PGLA), polysaccharide-based particles, or a combination thereof.
50. A plurality of delivery particles according to any one of claims 37 to 49, which are or contain a virus particle.
51. A plurality of delivery particles according to any one of claims 37 to 50, which are two or more types of virus particles, or include them.
52. A plurality of delivery particles according to any one of claims 49 to 51, wherein the virus particle is one or more of the following: AAV delivery particles, lentivirus delivery particles, adenovirus delivery particles, herpesvirus delivery particles, and anerovirus delivery particles.
53. The plurality of delivery particles according to claim 51, wherein the AAV delivery particle is one or more serotypes (e.g., AAV2, AAV5, AAV6, AAV8, AAV9, AAV.DJ, AAV.PHP, any variant thereof, or a combination thereof).
54. The plurality of delivery particles according to claim 48 or 49, wherein the two or more types of delivery particles are two or more types of lipid-based particles (e.g., LNPs) (e.g., having different formulations), or comprise the same.
55. A delivery particle containing nucleic acid according to any one of claims 1 to 35.
56. A group of delivery particles containing nucleic acids according to any one of claims 1 to 35.
57. A cell comprising a nucleic acid according to any one of claims 1 to 35, a library according to claim 36, a plurality of delivery particles according to any one of claims 37 to 54, or a delivery particle according to claim 53.
58. A population of cells comprising a nucleic acid according to any one of claims 1 to 35, a library according to claim 36, a plurality of delivery particles according to any one of claims 37 to 54, a delivery particle according to claim 55, or a population of delivery particles according to claim 56.
59. A composition (for example, a pharmaceutical composition) comprising a nucleic acid according to any one of claims 1 to 35, a library according to claim 36, a plurality of delivery particles according to any one of claims 37 to 54, a delivery particle according to claim 55, or a collection of delivery particles according to claim 56.
60. (a) A set of nucleic acids, wherein each nucleic acid is a nucleic acid according to any one of claims 1 to 35, and (b) A kit comprising a set of binders, each of which is a polypeptide or a nucleic acid encoding a polypeptide, and which specifically binds to at least certain peptide barcodes in a collection of barcodes.
61. The kit according to claim 60, wherein one or more of the binders are provided as phage particles or an aggregate thereof that have been manipulated to express the binders.
62. The kit according to claim 60, wherein one or more of the binders are provided as nucleic acids in a phagemide vector or as inserts suitable for cloning into a phage vector.
63. Furthermore, the kit according to any one of claims 60 to 62 includes information specifying a peptide barcode for each binder, it is confirmed that each binder specifically binds to at least a specific peptide barcode in the set of barcodes, and each peptide barcode specifically binds to at least one of the binders in the set.
64. The kit according to any one of claims 60 to 63, further comprising a set of instructions for performing sequencing of one or more phage particles coupled to one or more barcodes.
65. Furthermore, the kit according to claim 64 includes a computer-readable program for decoding sequencing data.
66. Furthermore, the kit according to any one of claims 60 to 65 further comprises a reagent for expressing a binder on phage particles.
67. A kit according to any one of claims 60 to 66, comprising nucleic acids that encode one or more barcodes.
68. A kit according to any one of claims 60 to 67, comprising nucleic acids encoding one or more binders.
69. A method for identifying therapeutic polypeptides or targeted polypeptides for treating a disease, disorder, or condition, a) A step of evaluating a group of barcoded cargo polypeptides, wherein the barcoded cargo polypeptides are encoded by the nucleic acid described in any one of claims 1 to 35, b) In order to identify a positive group, a negative group, or both, the step of separating members of the group that meet the criteria from those that do not meet the criteria, c) The steps of bringing the positive group, or the negative group, or each group separately from the other, into contact with a set of binders containing at least one binder specific to each barcode of the group, and d) The method comprising the step of identifying a binder that binds to the separated member, thereby identifying a barcoded cargo polypeptide present in the contacted group(s).
70. moreover, a) Administering the population of nucleic acids encoding the barcoded cargo polypeptide to an animal, and b) The method according to claim 69, comprising obtaining a sample from the animal for further evaluation.
71. The method according to claim 70, wherein the separation step comprises purifying one or more barcoded cargo polypeptides from the sample.
72. The method according to claim 71, wherein the barcoded cargo polypeptide is purified from the composite sample.
73. The method according to claim 72, wherein the composite sample is a tissue.
74. The method according to claim 73, wherein the composite sample is blood.
75. The method according to any one of claims 71 to 74, wherein the barcoded cargo polypeptide is purified using an affinity purification method (e.g., FLAG IP, Protein G / A) or a protein precipitation method.
76. The method according to any one of claims 69 to 75, wherein each binder in the set of binders is expressed on a phage.
77. The aforementioned step of identification is a) Amplifying the nucleic acid of the bound phage particle, b) Identifying the nucleotide sequence of the amplified nucleic acid, wherein one or more of the identified nucleotide sequences correspond to the coding sequence of the binder. c) Using the identified sequence(s) of the code sequence of the binder, detect one or more cargo polypeptides from the group of barcoded cargo polypeptides, and f) The method according to claim 76, comprising identifying one or more barcoded cargo polypeptides as a therapeutic agent or target for treating a disease, disorder, or condition.
78. A method for pharmacokinetic screening, a) Administering to an animal a population of nucleic acids encoding a barcoded therapeutic candidate polypeptide or a set of characteristic portions thereof, wherein each therapeutic candidate polypeptide contains a specific peptide barcode, b) Collecting samples from the animals, c) Purify one or more barcoded therapeutic candidate polypeptides from the sample. d) Contacting the sample with a set of binders (e.g., binder-expressing binders) containing at least one binder specific to each barcode in the sample, and e) The method comprising (for example, simultaneously) identifying the relative amounts of each binder present in the sample, and identifying the pharmacokinetic properties, biodistribution, half-life, tissue-mediated drug pharmacokinetics (TMDD), epitope properties, affinity, thermal stability, pH sensitivity, or in vivo stability of each barcoded therapeutic candidate polypeptide.
79. The method according to claim 78, wherein multiple samples are collected from the animal.
80. The method according to claim 78, wherein the animal is a model of a disease, disorder, or condition.
81. The method according to claim 80, wherein the disease, disorder, or condition is a cancerous, autoimmune, neurodegenerative, or pathogenic (e.g., viral / bacterial) disease, disorder, or condition.
82. The method according to any one of claims 78 to 81, wherein the purified therapeutic candidate polypeptide is a subset of the barcoded therapeutic candidate polypeptide administered to the animal.
83. The method according to any one of claims 78 to 81, wherein the sample is blood, tissue, or tumor.
84. The method according to any one of claims 78 to 83, wherein the sample is a control.
85. The method according to any one of claims 78 to 84, wherein the identifying step comprises (i) sequencing nucleic acids from the binder expressing the binder; (ii) identifying the relative amount of each therapeutic candidate polypeptide by decoding the relative amount of each barcode present; and / or (iii) performing one or more of FACS, MACS (magnetically activated cell sorting), or affinity purification.
86. The method according to any one of claims 78 to 85, comprising removing any unassociated (e.g., unbound) binders.
87. The method according to claim 86, wherein the removal is performed by washing.
88. The method according to any one of claims 78 to 87, wherein the identifying step includes performing one or more of amplification, proliferation, and sequencing (e.g., amplification, proliferation, and / or sequencing of nucleic acids (e.g., DNA, RNA)).
89. The method according to claim 88, wherein the amplification is performed using PCR, LAMP, or RCA.
90. The method according to claim 88, wherein the sequencing is performed using Illumina, NGS, nanopore sequencing, or Pac Bio long-read sequencing.
91. The method according to any one of claims 78 to 90, wherein the identifying step includes quantifying the number of binders that bind to the barcoded therapeutic candidate polypeptide, and the quantification is performed by decoding the nucleotide sequence of each binder that binds to the barcoded therapeutic candidate polypeptide.
92. The method according to claim 91, wherein the number of nucleotide sequences provides a measure of the target polypeptide in the population of barcoded therapeutic candidate polypeptides.
93. The method according to any one of claims 78 to 92, wherein the administration step comprises administering the barcoded therapeutic candidate polypeptide orally or intravenously.
94. The method according to any one of claims 78 to 93, wherein the barcoded therapeutic candidate polypeptide is delivered by a plurality of delivery particles according to any one of claims 37 to 54, a delivery particle according to claim 55, or a group of delivery particles according to claim 56.
95. The method according to any one of claims 70 to 94, wherein the animal is a mammal.
96. The method according to any one of claims 70 to 95, wherein the animal is a human.
97. The method according to any one of claims 70 to 96, wherein the animal is genetically modified to express the barcoded therapeutic candidate polypeptide.
98. It is a treatment method, a) A step of evaluating a population of nucleic acids that encode a set of barcoded cargo polypeptides, b) In order to identify a positive group, a negative group, or both, the step of separating members of the group that meet the criteria from those that do not meet the criteria, c) The step of bringing the positive group, or the negative group, or each group separately from the other, into contact with a set of binders containing at least one binder specific to each barcode of the group. d) Identifying the binder that binds to the separated member, thereby identifying the barcoded cargo polypeptide present in the contacted group(s), and e) The method comprising administering a therapeutic polypeptide, or a nucleic acid encoding the therapeutic polypeptide or a characteristic portion thereof, which has been confirmed to satisfy the evaluation by a process comprising the step of identifying a therapeutic polypeptide from the barcoded cargo polypeptide identified to be present in the contacted population(s) in the population(s) in contact.
99. It is a treatment method, a) A step of bringing a set of binders into contact separately with either a first group, a second group, or each of the first and second groups, wherein the barcoded cargo polypeptides are encoded by the nucleic acid described in any one of claims 1 to 35. i) Each binder is uniquely bound to one barcode compared to other barcodes. ii) The set of binders collectively includes a binder specific to each of the barcodes in the first and second groups, The first and second groups are separated from each other based on their performance in the evaluation, the contact step, b) Identifying the binder of the set that is bound to a member of the first group, the second group, or both, thereby identifying the barcoded cargo polypeptide present in the contacted group(s), and c) The method comprising administering a therapeutic polypeptide, or a nucleic acid encoding the therapeutic polypeptide or a characteristic portion thereof, which has been confirmed to satisfy the evaluation by a process including the step of identifying a therapeutic cargo polypeptide from the barcoded cargo polypeptide identified to be present in the contacted population(s) said population(s).
100. A method of treatment comprising administering a therapeutic polypeptide or a characteristic portion thereof, wherein the therapeutic polypeptide is identified from a group of barcoded cargo polypeptides by the method described in any one of claims 69 to 97.
101. A therapeutic method comprising administering a therapeutic polypeptide or a nucleic acid encoding a characteristic portion thereof, wherein the therapeutic polypeptide is identified from a group of barcoded cargo polypeptides by the method described in any one of claims 69 to 97.
102. A composition (for example, a pharmaceutical composition) comprising one or more therapeutic polypeptides or characteristic portions thereof, wherein the one or more therapeutic polypeptides are identified from a group of barcoded cargo polypeptides by the method described in any one of claims 69 to 97.
103. A composition (for example, a pharmaceutical composition) comprising one or more barcoded cargo polypeptides or characteristic portions thereof, wherein the one or more barcoded cargo polypeptides are produced by the method according to any one of claims 69 to 97.
104. A composition (e.g., a pharmaceutical composition) comprising one or more therapeutic polypeptides or one or more nucleic acids encoding a characteristic portion thereof, wherein the therapeutic polypeptide is identified from a group of barcoded cargo polypeptides by the method described in any one of claims 69 to 97.
105. A method for producing one or more therapeutic polypeptides or a composition (e.g., a pharmaceutical composition) comprising a characteristic portion thereof, wherein the one or more therapeutic polypeptides are identified from a group of barcoded cargo polypeptides by the method described in any one of claims 69 to 97.
106. A method for producing a composition (e.g., a pharmaceutical composition) comprising one or more therapeutic polypeptides or one or more nucleic acids encoding a characteristic portion thereof, wherein the therapeutic polypeptide is identified from a group of barcoded cargo polypeptides by the method described in any one of claims 69 to 97.