Methods for assessing the clinical relevance of genetic variance
A high-throughput screening method via saturation mutagenesis and phosphorylation measurement identifies ERBB2 variants linked to cancer, offering accurate diagnostic and therapeutic insights.
Patent Information
- Application Number
- JP2025507517
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-12
- Filing Date
- 2023-08-09
- Publication Date
- 2025-08-20
AI Technical Summary
Current methods lack a systematic approach to identify ERBB2 polypeptide variants associated with cancer progression, particularly in regions 679 to 992 of the ERBB2 polypeptide, which are implicated in increased phosphorylation and drug resistance, necessitating a comprehensive screening method to determine residues that contribute to cancer progression.
A method involving saturation mutagenesis of amino acids 679 to 992 of ERBB2, generating a library of variants expressed in mammalian cells, followed by high-throughput screening to measure phosphorylation levels and identify variants with enhanced activity using barcoded plasmids and next-generation sequencing.
Enables robust identification of ERBB2 variants associated with cancer, providing a comprehensive profile for diagnosis and treatment guidance with high accuracy and reproducibility, surpassing existing low-throughput methods.
Smart Images

Figure 2025527321000001_ABST
Abstract
Description
[Technical Field]
[0001] (cross reference) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 370,837, filed August 9, 2022, and U.S. Provisional Patent Application No. 63 / 375,369, filed September 12, 2022, the entire contents of which are incorporated herein by reference. Summary of the Invention [Means for solving the problem]
[0002] (overview) Disclosed herein is a method for identifying Erb-B2 receptor tyrosine kinase 2 (ERBB2) polypeptide variants associated with cancer, the method comprising: (a) providing a plasmid library for expression of a library of ERBB2 polypeptide variants, each plasmid in the plasmid library comprising: (i) a promoter; (ii) a polynucleotide sequence operably linked to the promoter, encoding an ERBB2 polypeptide variant from the library of ERBB2 polypeptide variants, wherein each ERBB2 polypeptide variant in the library of ERBB2 polypeptide variants independently and substantially comprises a single amino acid substitution in a region of the ERBB2 polypeptide ranging from residue 679 to residue 992 of SEQ ID NO: 1; and the library of ERBB2 polypeptide variants collectively comprises amino acid substitutions of substantially all 20 amino acids at substantially every amino acid residue in the region of the ERBB2 polypeptide ranging from residue 679 to residue 992 of SEQ ID NO: 1. and (iii) barcodes, wherein each plasmid in the plasmid library independently has a different barcode associated with the polynucleotide sequence encoding the ERBB2 polypeptide variant; (b) contacting a plurality of mammalian cells with the plasmid library, wherein the contacting results in expression of a single ERBB2 polypeptide variant from the library of ERBB2 polypeptides at the surface of a single mammalian cell from the plurality of mammalian cells, thereby producing a plurality of mammalian cells expressing the library of ERBB2 polypeptide variants, wherein a subset of the plurality of mammalian cells expressing the library of ERBB2 polypeptide variants have an ERBB2 polypeptide variant that is phosphorylated to a greater extent than a wild-type ERBB2 polypeptide expressed on the surface of the mammalian cells, and the ERBB2 polypeptide variant that is phosphorylated to a greater extent than the wild-type ERBB2 polypeptide is(c) identifying a subset of mammalian cells expressing an ERBB2 polypeptide variant that is phosphorylated to a greater extent than the wild-type ERBB2 polypeptide; and (d) sequencing the barcode of a plasmid present in each mammalian cell of the subset of mammalian cells, thereby identifying the ERBB2 polypeptide variant associated with the cancer. In some embodiments, the cancer comprises carcinoma. In some embodiments, the cancer comprises ovarian cancer, gastric cancer, bladder cancer, salivary cancer, or lung cancer. In some embodiments, the method further comprises detecting the presence of the ERBB2 polypeptide variant identified in (d) in a sample obtained from the subject. In some embodiments, the method further comprises diagnosing the subject as having or at risk of developing the cancer. In some embodiments, the method further comprises administering an anti-cancer treatment to the subject, if necessary, based on the presence of the ERBB2 polypeptide variant identified in (d) in a sample obtained from the subject. In some embodiments, the anti-cancer treatment is effective against cancer cells expressing the ERBB2 polypeptide variant. In some embodiments, the identifying step in (c) of the method comprises contacting a plurality of mammalian cells expressing the library of ERBB2 polypeptide variants with an agent that binds to phosphorylated ERBB2. In some embodiments, the agent that binds to phosphorylated ERBB2 is an antibody against phosphorylated ERBB2. In some embodiments, the identifying step in (c) further comprises performing an immunoassay using the antibody against phosphorylated ERBB2. In some embodiments, the identifying step in (c) further comprises performing cell sorting based on the immunoassay. In some embodiments, the cell sorting is fluorescence-activated cell sorting. In some embodiments, the promoter is an inducible promoter. In some embodiments, the inducible promoter is a doxycycline-inducible promoter. In some embodiments,The plasmids in the library of plasmids are viral plasmids. In some embodiments, the viral plasmids are lentiviral or adeno-associated viral plasmids. In some embodiments, the method further comprises packaging the viral plasmid into a viral capsid prior to the contacting step (b), thereby generating virions containing the viral plasmid. In some embodiments, the contacting step (b) comprises contacting the virions with mammalian cells from the plurality of mammalian cells. In some embodiments, the plurality of mammalian cells comprises HEK293 cells. In some embodiments, the sequencing step comprises next-generation sequencing. In some embodiments, the method further comprises stably expressing a cancer-associated ERBB2 polypeptide variant identified in (d) in mammalian cells and measuring the activity of the ERBB2 polypeptide when stably expressed. In some embodiments, the method further comprises comparing the activity of the ERBB2 polypeptide when stably expressed with that of stably expressed wild-type ERBB2. In some embodiments, at least 80% of the variants identified in (d) exhibit increased activity when stably expressed compared to stably expressed wild-type ERBB2. In some embodiments, the method further comprises entering the cancer-associated ERBB2 polypeptide variants identified in (d) into a database.
[0003] Also disclosed herein are databases containing ERBB2 polypeptide variants identified by the methods described herein.
[0004] Also disclosed herein are compositions comprising ERBB2 polypeptide variants identified by the methods described herein.
[0005] The novel features of the exemplary embodiments are set forth with particularity in the appended claims. A better understanding of these features and advantages will be obtained by reference to the following detailed description that sets forth illustrative embodiments in which the principles of the disclosed systems and methods are utilized, and the accompanying drawings in which: [Brief explanation of the drawings]
[0006] [Figure 1-1] Figure 1A is a schematic diagram showing exemplary steps in the method described herein for pERBB2 activation. Recombinant cells are grown under toxic selection. Cell sorting is based on GFP receptor expression or immunostaining. gDNA is isolated, and a targeted ERBB2 amplicon library is prepared and sequenced by NGS.
[0007] [Figure 1-2] FIG. 1B is a schematic diagram showing the ERBB2 activation assay used in the methods described herein.
[0008] [Figure 2] Figure 2 is a table showing the parameters of various project designs.
[0009] [Figure 3] Figure 3 is a table showing the project statistics summary.
[0010] [Figure 4] Figure 4 is a table showing the variants analyzed using the methods described herein, including activity scores (GML scores) and confidence values (P values). WT = wild type; GOF = gain of function; LOF = loss of function.
[0011] [Figure 5] FIG. 5 is a table showing the validation statistics for the assays of the methods described herein.
[0012] [Figure 6] Figure 6 is a heat map showing the sequence reads for each of the altered amino acids at each position in Erbb2. The color gradient indicates the number of associated reads. White boxes are wild-type amino acids and gray boxes are null values. Darker shades are associated with higher read counts.
[0013] [Figure 7] Figure 7 is a heatmap showing the number of barcodes for each amino acid variant at each position in ERBB2. The color gradient indicates the number of associated barcodes. White boxes are wild-type amino acids, and gray boxes are null values. Darker shades are associated with higher read counts.
[0014] [Figure 8] Figure 8 is a pie chart showing the read summary for ERBB2 variants. Variant distribution across different levels of read depth was quantified.
[0015] [Figure 9] Figure 9 is a pie chart summarizing the number of barcode reads for ERBB2 variants. Variant distribution across different levels of barcoding depth was quantified.
[0016] [Figure 10] Figure 10 is a heat map showing the number of barcodes for each amino acid variant at each position in ERBB2. The color gradient indicates the relative activity of the variant. Red variants are RF variants, and green variants are GOF variants. White boxes have wild-type levels of activity. Black boxes are reference amino acids, and gray boxes are null values.
[0017] [Figure 11]Figure 11 is a heat map showing the number of barcodes for each amino acid variant at each position in ERBB2. Statistically relevant activities are plotted (p<0.05). Red indicates RF variants and green indicates GOF variants. White boxes have wild-type levels of activity, black boxes are reference (Ref) amino acids, and gray boxes are null values.
[0018] [Figure 12] Figure 12 is an overlaid heat map of both p-values and activity for variants showing the variant amino acid for each position in ERBB2. The color gradient indicates the relative activity of the variant. White boxes have wild-type levels of activity, black boxes are reference (Ref) amino acids, and gray boxes are null values. The activity levels of the variants are colored based on the gradient key. Red is RF variants and green is GOF variants. The number of small blue boxes indicates p-values (p<0.05 (no box); <0.05 (1 box); <0.001 (2 boxes); <0.0001 (3 boxes); <0.00001 (4 boxes)).
[0019] [Figure 13] Figure 13 is a pie chart showing the classification of variants. The number of statistically relevant variants (p<0.05) was quantified for each category and is shown.
[0020] [Figure 14] Figure 14 is a flow sorting profile showing individual clones in cell colonies that were tested through the pERBB2 assay before performing the GML experiment. The relative fluorescence above each control is shown.
[0021] [Figure 15]Figure 15 is a 3D structural rendering showing the effect of HER2 variants on function. All surface maps are on the wild-type ERBB2 structure (PDB: 3PPO), with one member of each pair rotated 180 degrees around its Y axis: A) amino acid positions on the erbb2 backbone; B) the activation loop of ERBB2; C) the ATP binding site; D) hyperactivation variants; E) hyperactivation variants derived from the methods described herein; F) the phosphorylation site of ERBB2; G) LOF variants; H) RF variants derived from the methods described herein; and I) the ubiquitination site of ERBB2. For all plots, at least one substitution at the indicated position had relevant activity.
[0022] [Figure 16-1] Figures 16A and 16B are flow sorting profiles showing individual clones in cell colonies that were tested through pERBB2 assays before performing GML experiments. Relative fluorescence above each control is shown. [Figure 16-2] Same as above. DETAILED DESCRIPTION OF THE INVENTION
[0023] (Detailed explanation) (definition) The following definitions supplement the definitions in the art, and are intended for this application, and are not transferred to any related or unrelated matter (for example, to any patent or application that is commonly owned). Although any method and material similar or equivalent to the methods and materials described herein can be used in the implementation of the present disclosure, preferred materials and methods are described herein.Therefore, the terminology used herein is intended only to describe specific embodiments, and is not intended to be limiting.
[0024] In this application, the use of the singular includes the plural unless specifically stated otherwise. It must be noted that as used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. In this application, the use of "or" means "and / or" unless stated otherwise.
[0025] References herein to "some embodiments," "an embodiment," "one embodiment," or "other embodiments" mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least some embodiments of the present disclosure, but not necessarily in all embodiments.
[0026] Conditional language such as "can," "could," "might," or "may" is generally intended to convey that a particular implementation includes or excludes particular features, elements, and / or steps, unless specifically stated otherwise or understood otherwise from the context in which it is used. Thus, such conditional language is generally not intended to imply that features, elements, and / or steps are in any way required for one or more implementations.
[0027] Conjunctive language such as the phrase "at least one of X, Y, and Z," unless specifically stated otherwise, is understood in a different manner to generally convey that, in the context in which it is used, an item, term, etc., can be either X, Y, or Z. Thus, such conjunctive language is generally not intended to imply that a particular implementation requires the presence of at least one X, at least one Y, and at least one Z.
[0028] As used herein, the term "about" or "approximately" means within an acceptable error range for a particular value, including a range of up to 10% of the given value. When a particular value is stated in this application and claims, unless otherwise stated, the term "about" meaning within an acceptable error range for the particular value should be assumed.
[0029] The term "substantially" refers to the qualitative state of exhibiting at least 70% of the full extent or degree of a characteristic or property of interest.
[0030] As used herein, the terms "ERBB2" and "Her2" are used interchangeably to refer to the same polypeptide or the same gene encoding the same polypeptide.
[0031] The term "operably linked" refers to a functional linkage between a regulatory sequence and a nucleic acid sequence that results in expression of the latter nucleic acid sequence.
[0032] Although the present disclosure has been described with respect to particular implementations and uses, other implementations and uses (including implementations and uses that do not provide all of the features and advantages set forth herein) are also within the scope of the present disclosure. Components, elements, features, acts, or steps may be arranged or performed differently from that described, and components, elements, features, acts, or steps may be combined, integrated, added, or omitted in various implementations. All possible combinations and subcombinations of the elements and components described herein are intended to be included in the present disclosure. No single element or group of elements is essential or indispensable.
[0033] Any portion of any of the steps, processes, structures, and / or devices disclosed in one implementation or example in this disclosure can be combined with or used with (or in place of) any other portion of any of the steps, processes, structures, and / or devices disclosed or illustrated in a different implementation, flowchart, or example. The implementations and examples described herein are not intended to be separate and distinct from one another. Combinations, variations, and implementations of the features of the present disclosure are within the scope of the present disclosure.
[0034] Although operations may be described herein in a particular order, such operations need not be performed in the particular order or sequential order described, or even all of the operations need not be performed, to achieve desirable results. Other operations not depicted or described may be incorporated into the example methods and processes. For example, one or more additional operations may be performed before, after, simultaneously with, or between any of the described operations. Furthermore, the operations may be rearranged or reordered in some implementations. Also, the separation of various components in the above implementations should not be understood as requiring such separation in all implementations, and it should be understood that the described components and systems may generally be incorporated together into a single item or packaged into multiple items. Furthermore, several implementations are within the scope of the present disclosure.
[0035] (overview) Disclosed herein is a method for identifying ERBB2 variants that are related to the progression of cancers associated with mutations in ERBB2.ERBB2 is a receptor tyrosine kinase with intrinsic tyrosine kinase activity.Currently, there is no known substrate for its receptor, and therefore, it is believed that the ERBB2 extracellular domain remains in an "open position" similar to other members of the mammalian EGFR family when not bound to its natural ligand.In the "open" state, ERBB2 can easily bind to other members of the mammalian EGFR family.
[0036] Amplification and / or overexpression of EGFR is thought to be involved in many cancers (e.g., ovarian cancer, gastric cancer, bladder cancer, salivary gland cancer, and lung cancer). Furthermore, increased phosphorylation of ERBB2 due to mutations that increase its intrinsic tyrosine kinase activity has been implicated in the progression of such cancers, which may lead to drug resistance in certain cancer cells that express mutant ERBB2 that results in increased phosphorylation. In some embodiments, the mutation is with respect to the ERBB2 polypeptide of SEQ ID NO: 1.
[0037] (SEQ ID NO: 1) [ka] [ka]
[0038] In some cases, these mutations are in the tyrosine kinase domain (residues 715-992 of SEQ ID NO:1), the junction membrane (JM) region (residues 679-714 of SEQ ID NO:1), or both. However, a systematic approach to determining which residues in ERBB2, when mutated, result in increased autophosphorylation associated with cancer progression is needed. The disclosed method thus provides this systematic approach by saturation mutagenesis of amino acids 679-992 of SEQ ID NO:1, thereby generating a library of ERBB2 variants that substantially encompasses the entire sequence space. The library is then screened directly in mammalian cells to recapitulate the native environment of the ERBB2 variants. Finally, the degree of phosphorylation for each ERBB2 variant present on the surface of the mammalian cells is directly measured, thereby enabling binning of ERBB2 variants based on their degree of phosphorylation. By comparing the degree of phosphorylation to that of wild-type or benchmark controls, a comprehensive profile of ERBB2 variants associated with cancer was constructed. This approach provides for the identification of novel ERBB2 variants associated with cancer progression. Using the cancer-associated ERBB2 variant profiles generated herein as a guide, novel methods for detecting and / or treating cancer are also provided, in which one or more ERBB2 variants identified using the methods described herein can be detected in a sample obtained from a subject, thereby diagnosing the subject as having or at risk of developing the cancer.
[0039] (Screening method) Disclosed herein is a method for identifying Erb-B2 receptor tyrosine kinase 2 (ERBB2) polypeptide variants associated with cancer. Figure 1A provides an overall overview of embodiments of the methods described herein. In some embodiments, the methods described herein may include one or more of steps 1-7 described in Figure 1A. In some embodiments, the methods may include one or more of the following: (i) Assay validation; (ii) generation of synthetic variant libraries; (iii) generation of lentiviral variant libraries; (iv) infecting cells with the lentiviral variant library; (v) cell sorting based on the readout of that assay; (vi) variant identification using targeted next-generation sequencing; and (vii) Biochemical / bioinformatic characterization of identified variants.
[0040] In some embodiments, the method comprises one or more of the following steps (a)-(e): (a) providing a plasmid library for expression of a library of ERBB2 polypeptide variants, wherein each plasmid in the plasmid library comprises one or more of the following (i) to (iii): (i) a promoter; (ii) a polynucleotide sequence operably linked to the promoter, encoding an ERBB2 polypeptide variant from the library of ERBB2 polypeptide variants, wherein each ERBB2 polypeptide variant in the library of ERBB2 polypeptide variants independently and substantially comprises a single amino acid substitution in a region of the ERBB2 polypeptide ranging from residue 679 to residue 992 of SEQ ID NO: 1; and the library of ERBB2 polypeptide variants comprises ERBB2 polypeptide variants collectively having amino acid substitutions of substantially all 20 amino acids at substantially every amino acid residue in the region of the ERBB2 polypeptide ranging from residue 679 to residue 992 of SEQ ID NO: 1; and (iii) barcodes, wherein each plasmid in the plasmid library independently comprises a different barcode associated with the polynucleotide sequence encoding the ERBB2 polypeptide variant; (b) contacting a plurality of mammalian cells with the plasmid library, wherein the contacting results in expression of a single ERBB2 polypeptide variant from the library of ERBB2 polypeptides at the surface of a single mammalian cell from the plurality of mammalian cells, thereby generating a plurality of mammalian cells expressing the library of ERBB2 polypeptide variants, wherein a subset of the plurality of mammalian cells expressing the library of ERBB2 polypeptide variants has an ERBB2 polypeptide variant that is phosphorylated to a greater extent than a wild-type ERBB2 polypeptide expressed on the surface of the mammalian cells, and the ERBB2 polypeptide variant that is phosphorylated to a greater extent than the wild-type ERBB2 polypeptide is the ERBB2 variant associated with the cancer; (c) contacting a plurality of mammalian cells expressing the library of ERBB2 polypeptide variants with an agent that binds to phosphorylated ERBB2; (d) identifying a subset of mammalian cells that express an ERBB2 polypeptide variant that is phosphorylated to a greater extent than the wild-type ERBB2 polypeptide based on binding of the agent to phosphorylated ERBB2 expressed on the surface of the subset of mammalian cells; and (e) sequencing the barcode of the plasmid present in each mammalian cell of the subset of mammalian cells, thereby identifying an ERBB2 polypeptide variant associated with the cancer.
[0041] In some embodiments, each plasmid has the same promoter. In some embodiments, each plasmid has a different promoter. In some embodiments, the promoter is a constitutive promoter (e.g., SV40, CMV, UBC, EF1A, PGK, or CAGG). In some embodiments, the promoter is an inducible promoter (e.g., doxycycline or tetracycline). The embodiments described herein utilize an assay for detecting ERBB2 variants associated with cancer. FIG. 1B provides a diagram of the ERBB2 signaling pathway. Without wishing to be bound by theory, autophosphorylated ERBB2 may be involved in the AKT signaling pathway and the ERK signaling pathway, and abnormally high phosphorylation of ERBB2 may lead to the progression of the cancers described herein. Therefore, the presence of phosphorylated ERBB2 using an agent that binds to phosphorylated ERBB2 can be used in an assay to screen for ERBB2 variants with abnormally high phosphorylation (which may therefore be associated with cancer progression).
[0042] Using the methods described herein, comprehensive profiles can be generated to model gene variants (e.g., phospho-Her2) based on, for example, an assay in mammalian cell culture (e.g., a method for phospho-Her2 Her2 tyrosine kinase domain activation) shown in Figure 1B. Plasmid libraries encoding these variants can be constructed, which are then analyzed using the methods described herein to generate comprehensive variant effects on gene activity, as well as variant activity profiles.
[0043] The high-throughput cellular molecular function assay method described herein has several advantages over existing methods.For example, performing saturation mutagenesis of the entire region of ERBB2 polypeptide allows all or almost all variations to be directly measured in biologically native environment, rather than inferring activity from the difference between pre-screening samples and screened samples.In addition, the method described herein can utilize plasmids that each encode individual ERBB2 polypeptide variants and each have individual barcodes associated with specific ERBB2 polypeptide variants.Because of the barcoding of individual molecules and high-throughput performance at the single-cell assay level, the method described herein allows signal averaging of multiple individual measurements, resulting in robust reproducibility, high accuracy, and statistical reliability of the activity of each variant.In addition, all ERBB2 variants are assayed under standardized conditions in the same cells with the same genetic background, which produces a consistent data set.
[0044] This disclosure illustrates this method as it relates to the production of ERBB2 (a protein produced from the Erbb2 gene) (containing a constitutively active YVMA indel variant that induces ERBB2 phosphorylation due to activation of the MAPK pathway). Detecting phosphorylated ERBB2 is a direct measurement of ERBB2 / Her2 activity. Alternatively, indirect measurement of ERBB2 / Her2 activity can be performed by measuring p-ERK activation. However, p-ERK activation can also result from other pathways and mitogen receptors, which require more stringent control and have a lower signal-to-noise ratio.
[0045] Utilizing the embodiments presented herein that directly detect ERBB2 / Her2 activity through the generation of phosphorylated ERBB2 / Her2 can be combined with downstream / global measurements of tumorigenicity through cell proliferation, utilizing known variants as controls to increase the pathway-specific accuracy of those assays, which can help, at least in part, reduce ambiguity in variant interpretation. In some cases, quantification of the tumorigenic potential of known or identified variants using low-throughput phospho-Her2 flow cytometry assays (in triplicate) can be combined with the methods described herein to provide standardization of the methods described herein.
[0046] The methods provided herein are modular, high-throughput, one-pot assay systems for measuring the molecular function of a large number of genetic variants (e.g., hundreds or thousands of variants) simultaneously with high accuracy. In some cases, the methods provided herein measure millions of genetic variants. In some applications, the high-throughput methods provided herein have an overall accuracy of more than 75%, for example, compared to activity measurements using low-throughput flow cytometry. In some embodiments, the accuracy of the methods described herein is about 80%, about 82%, about 84%, about 86%, about 88%, or about 90%, for example, compared to activity measurements using low-throughput flow cytometry. In some embodiments, the accuracy of the methods described herein for identifying variants with gain-of-function activity (i.e., higher activity than wild-type) is about 80%, about 82%, about 84%, about 86%, about 88%, or about 90%, for example, compared to activity measurements using low-throughput flow cytometry. In some embodiments, the accuracy of the methods described herein has perfect or near perfect agreement with other standardized assays (100% accuracy or near 100% accuracy).
[0047] (Barcoded variant library) Plasmids can be created by introducing compatible restriction enzyme sites (e.g., EcoRI, SalI, and AsiSI) and inserting desired clones of the ERBB2 / Her2 variant library. ERBB2 / Her2 variant encodings can be PCR amplified from templates with polymerase and cloned into the digested plasmid. In some embodiments, the plasmid is a retrovirus. In some embodiments, the plasmid is a lentivirus plasmid. In some embodiments, the plasmid is an AAV plasmid. In some embodiments, the viral vector corresponds to a virus of a specific serotype. In some instances, the serotype is selected from AAV1 serotype, AAV2 serotype, AAV3 serotype, AAV4 serotype, AAV5 serotype, AAV6 serotype, AAV7 serotype, AAV8 serotype, AAV9 serotype, AAV10 serotype, AAV11 serotype, AAV12 serotype, avian AAV, bovine AAV, canine AAV, equine AAV, or ovine AAV.
[0048] In some embodiments, the plasmid comprises DNA. In some embodiments, the plasmid comprises RNA. In some instances, the plasmid comprises circular double-stranded DNA. In some instances, the plasmid may be linear. Various selection markers (e.g., puromycin resistance gene or blasticidin S resistance gene) can be used to select for plasmid transduction. For example, the selection marker amplicon and the Her2 variant amplicon can be fused by inverse PCR using a polymerase. The fused amplicon is then cloned into a plasmid digested with a compatible restriction enzyme.
[0049] In some embodiments, a double-stranded (ds) DNA library containing Her2 cDNAs with the sequences of all possible single amino acid variants is synthesized. The dsDNA from each well can be pooled, and a single overlap PCR extension adds random-mer oligonucleotides to the 3' untranslated region. The synthesized dsDNA library has a 3' overhang sequence after the stop codon that overlaps with the 5' overhang sequence upstream of the random-mer oligonucleotide sequence. The pooled ds DNA library and its random oligomers can be mixed, denatured, and annealed in molar ratios of 1:2, 1:5, 1:10, 1:20, and 1:50. The hybridized DNA can be extended during one cycle of PCR with DNA polymerase. The PCR reaction mixture can then be treated with exonuclease and purified.
[0050] The purified DNA is digested with a compatible restriction enzyme and ligated into the digested plasmid with ligase. The ligation reactions can be pooled, purified, and dialyzed. The purified ligation reaction mixture can be electroporated into electrocompetent cells, plated, and incubated. Transformants can be scraped from the plates, and a plasmid library from the pooled cell suspension can be isolated.
[0051] A viral library (e.g., a lentiviral library) is generated in cells compatible with transformation (e.g., LentiX 293T cells). In some embodiments, tens of thousands to billions of cells are seeded in a Petri dish. In some embodiments, approximately 3 million LentiX 293T cells are seeded in a Petri dish and grown in complete medium (e.g., DMEM + 10% fetal bovine serum). In some embodiments, the plasmid contains one or more regulatory elements. In some embodiments, the plasmid contains packaging elements used in the transformation of the viral plasmid. For example, the plasmid library; a vector encoding the packaging element Gag and packaging element Pol; a vector encoding the packaging element Rev; and a vector encoding an envelope protein are combined. CaCl2 is added to the plasmid mixture. 2x HBS can be added to the transfection mixture with stirring. The transfection mixture is incubated and added to the cells in the Petri dish. The cells may be incubated in a CO2 incubator at 37°C with a 5% CO2 atmosphere. After transfection, the calcium phosphate-containing medium is replaced with complete medium (DMEM + 10% FBS) and incubated in a CO2 incubator at 37°C with a 5% CO2 atmosphere. Spent medium from confluent transfected cells (e.g., LentiX 293T) is filtered. Aliquots of the filtered spent medium containing the lentivirus may be frozen and stored.
[0052] A viral vector for a specific clone can be produced in cells (e.g., LentiX 293T). For example, the cells are seeded into the wells of a well plate. After 24 hours, the cells are co-transfected with the above-mentioned plasmid clone and one or more packaging elements or envelope proteins and transfection with a transfection reagent (e.g., Lipofectamine LTX (Invitrogen)). After incubation, the medium is replaced and the cells are cultured in complete medium. The cell supernatant is collected, filtered, frozen, and stored.
[0053] Virus can be titered by seeding cells into wells of a well plate and culturing in complete medium (e.g., DMEM + 10% FBS). After removing most of the spent medium from the wells, serial dilutions of virus can be added and incubated. Complete medium can be added and incubated. The spent medium is removed and replaced with complete medium containing a selection compound (e.g., puromycin) and incubated. The cells are examined for viability under a microscope, and colonies are counted to calculate the infectious units / ml.
[0054] (GFP reporter cell line) Cells are seeded into wells of a well plate and grown in complete medium (e.g., DMEM). For example, a GFP reporter plasmid carrying LTR-GFP and a resistance marker (e.g., the blasticidin S resistance (BSR) gene) is transfected into the cells (e.g., LentiX 293T) and incubated. The transfected cells are selected for marker resistance (e.g., blasticidin S), and the medium (e.g., DMEM) containing the marker is replaced every three days. The cells are trypsinized and serially diluted in the well plate. After incubation, single colonies are grown and screened.
[0055] To confirm viral integration, gDNA is isolated. Tat amplicons are subcloned and sequenced. Tat transcription activity is measured in subcultures of each clonal cell line. Cell cultures in well plates are transfected with wild-type Tat expression vectors and cultured. Transactivation-induced GFP expression is assessed by epifluorescence microscopy. The clonal reporter cell lines are propagated, frozen, and stored.
[0056] (Cell lines and libraries) Cells (e.g., LentiX 293T / LTR-GFP) can be transduced with the Her2 variant virus library at a multiplicity of infection (MOI) of 0.1. After infection, cells are cultured and maintained in complete medium supplemented with a selection agent (e.g., puromycin). Confluent cells are harvested, counted, washed once with 1x PBS, and then fixed and gDNA isolated for next generation sequencing of the Her2 amplicon.
[0057] Cells (e.g., Jurkat / LTR-GFP) are seeded and transduced with the viral library at 0.1 MOI. After transduction, the cells are selected for viral survival in medium (e.g., RPMI 1640 + 10% FBS) supplemented with a selection agent (e.g., puromycin). The cells are then counted, washed with 1x PBS, and fixed for flow sorting, after which gDNA is isolated.
[0058] To evaluate the performance of the system and method for high-throughput cellular molecular function assay, random variants of Her2, as well as empty vectors and wtHer2, are stably expressed in cells (e.g., LentiX 293T / LTR-GFP). The cells are seeded into wells of a well plate and incubated. The cells are transduced with viruses and selected and maintained in complete medium containing a selection agent (e.g., puromycin). The cells are harvested and analyzed by flow cytometry to evaluate LTR transactivation GFP expression. The same stable cell line can be created in Jurkat / LTR-GFP cells. The clones selected for empty vectors, wtHer2, and Her2 variants are frozen and stored.
[0059] In some embodiments, one-quarter of the LentiX293T / LTR-GFP cells and one-tenth of the Jurkat / LTR-GFP cells are harvested, and gDNA is isolated and sequenced to evaluate library display before flow sorting. The remaining cells are fixed in 2% paraformaldehyde / PBS, washed twice with 1x PBS, and resuspended in 1x PBS for analysis by flow sorting (e.g., Sony 800S cell sorter). Cells can be sorted into three bins of GFP signal intensity (low GFP, medium GFP, and high GFP), gated using a threshold determined for cells stably expressing wt-Tat for maximum transactivation of LTR-GFP, and a threshold determined for cells stably expressing Her2 variants or empty vectors for low background transactivation of LTR-GFP.
[0060] To perform deep sequencing, primers can be designed to flank the Her2 targeting region from gDNA and incorporate NGS sequencing adapters. gDNA is amplified by PCR. The NGS library for each sample category can use 10 NGS library forward primers and one NGS library reverse primer. The forward primers can be common to all sample categories, and the reverse primer is unique to each sample. The Her2 amplicons are pooled and purified (e.g., gel extracted). All samples are pooled and sequenced (e.g., on a Novaseq 6000 sequencing platform). Samples are sequenced (synthetic dsDNA Her2 variant library, plasmid library), cell libraries are selected (in duplicate), and low-, medium-, and high-GFP cells are flow-sorted (in duplicate) for each cell line.
[0061] (Bioinformatics) Provided herein is a database containing ERBB2 variants associated with cancer progression identified by the methods described herein. Sequencing data derived from the methods described herein are processed and stored in a database. For example, paired-end reads can be processed through a multi-stage bioinformatics pipeline (e.g., BaseSpace), and the resulting reads in a file (e.g., bcl) are converted into a processed file (e.g., FASTQ). Read quality can be assessed by an algorithm (e.g., FASTQC). Paired-end reads for all samples are merged together to construct complete Her2 contigs (e.g., FLASH®). The contigs are quality trimmed (e.g., Trimmomatic). Adapters are trimmed, barcodes are isolated (e.g., CutAdapt), and the barcodes are grouped (e.g., Starcode). The sequence reads are demultiplexed into a subset of read sequences for each cell clone based on unique barcodes using a custom script (e.g., Python) that processes the output of the grouped barcodes. The resulting reads are then aligned to the Her2 cDNA. A file containing the nucleotide variants is retrieved for each subset of Her2 contigs (cell clones) and output as a file. A custom Python script can be used to identify the amino acid substitutions for the output file, the number of reads for each barcode in each sample, and the barcode group for cells with the same amino acid substitution. This library can be used in a script that aggregates information about each variant from the output files. The read count and read depth for each barcode and each amino acid substitution in each sample are normalized to reads per million, and activity is measured by the percentage of GFP+ reads for each barcode and each variant.
[0062] Statistics can be calculated for each variation. In some cases, there are n cell lines (biological replicates), and each cell line has m technical replicates. For each barcode (group) in the sample, the percentage of reads in the GFP+ group relative to the total number of reads in both the GFP+ and GFP- groups can be calculated and expressed as the h ratio (h∈[0,1]). In some cases, a high h percentage indicates a wild type, while a low h percentage indicates a variant. Then, for each variant, the average h ratio for all barcodes assigned to the same variant is calculated and expressed as a variant-level summary score. In some cases, a one-sample t-test is used to evaluate 1) and 2) below: 1) whether the variant has a significantly different number of reads in the GFP+ group compared to the GFP- group within the technical replicate, and 2) whether the variant has a significantly different number of reads in the GFP+ group compared to the GFP- group between different cell lines based on biological replicates (null hypothesis H0: h=0.5).
[0063] In some cases, variants with high h percentages are classified as wild-type, and variants with low h percentages are classified as LOF variants.In order to estimate the first-order error of this classification, in some cases, a list of true variants with wild-type transcriptional activity and true LOF variants with low activity is compiled.Then, these h percentages are fitted with a beta distribution as the null distribution.Specifically, for wild-type detection, in some cases, the true variant is used as the null, and vice versa, for variant detection, the wild-type is used as the null.Moment estimators are used to estimate the parameters of the model.The p-values for various cell lines are combined into a global test p-value using Fisher's method.
[0064] Performance metrics such as accuracy, sensitivity, specificity, positive predictive value and negative predictive value may be based on standard formulas.
[0065] Figures can be created in PowerPoint, Excel, FlowJo, and Pymol. Bin, bar, and pie plots, as well as saturation mutagenesis heat maps, can be created in Excel. Values for saturation mutagenesis heat maps and 3D surface plots can be generated with custom Python scripts. 3D surface plots of amino acid tolerance at each position show the accuracy of physiochemical properties as a color gradient, with the highest accuracy indicated. Accuracy is a standard formula and is calculated for groups of amino acids with similar physiochemical properties. Solvent-accessible surface area (SASA) can be calculated for the Her2 structure. A residue is considered buried if less than 10% of its surface area is exposed to solvent.
[0066] For example, an MCC formula is calculated using the following data definitions for large hydrophobic amino acids at a position in Her2: if Phe, Tyr, or Trp have >50% activity, they are true positives; if other amino acids have <50% activity, they are true negatives. If any of Phe, Tyr, or Trp have <50% activity, they are false positives; if other amino acids have >50% activity, they are false negatives. Also, the wild-type amino acid is considered a true positive if it is in its physiochemical group, and a true negative if it is not in its physiochemical group. The MCC is a novel visual mining approach to capture the tolerance of amino acid types at each position and, when mapping the surface of the 3D structure, reveal the spatial relationships of amino acid tolerance and their association with other Her2 functions. [Example]
[0067] Example 1: ERBB2 Plasmid Library ERBB2 (HER2) is an oncogene implicated in cancer progression. Furthermore, activating missense variations in the tyrosine kinase domain and juxtamembrane regions of ERBB2 (HER2) may be drivers of variation in multiple tumor types. Historically, assays used by different groups to characterize these variants lacked uniformity, resulting, at least in part, in a lack of clarity regarding each variant's impact on protein function, tumorigenic potential, or both.
[0068] The parameters for this project are summarized in the table in Figure 2. The generated ERBB2 plasmid library contained 99% of the expected 5,966 variants, with the most frequently deleted variants coming from four positions. The ERBB2 library included the tyrosine kinase domain (residues 715-992) and the JM region (residues 679-714), which were not part of the original design. The p-Her2 GML results were filtered to remove variants with low read counts (<1 barcode) and variants not designed in the original experiment. The filtered set had measurements for 5,886 (99%) of the variants designed in the experiment (table in Figure 3).
[0069] Example 2: Results and Interpretation The plasmid library was sequenced using PacBio CCS long reads (n = 1,519,453), and variants were called using the Heligenics pipeline (Table in Figure 3). Next, a lentiviral library was generated and transduced into dox-inducible LentiX293T cells. Expression of variant proteins was induced with doxycycline for 72 hours before cell harvest. Harvested cells were fixed, permeabilized, and then immunostained with antibodies against Her2 and p-Her2 (Tyr1248). The immunostained cell library was then flow-sorted first for Her2 expression and then into four fluorescence-graded bins based on p-Her2 expression. After sorting, the GMLs were sequenced and filtered by next-generation sequencing (NGS), generating 29,603,876 short reads for calculating the unique molecular identifier (UMI)-barcode frequency in each bin. Of the 15,845,059 barcoded cells sorted, there were 268,325 unique barcodes, with an average of 46 barcodes / variant. Relative activity levels were calculated and variant activity was classified as wild-type (WT), loss-of-function (RF), or gain-of-function (GOF) by comparing each variant to WT using a two-sample t-test. The method described herein was optimized to primarily identify GOF variants and not RF variants.
[0070] The number of reads and barcodes associated with each variant is shown in heat maps (Figures 6-7), and the read and barcode counts for all variants are shown in pie charts (Figures 8-9). Figures 10-11 show the saturation mutagenesis heat maps of p-Her2 activity, statistical tests for each variant, and an overlay combining both data. These data indicate which variants are WT, GOF, and RF. The number of variants that met statistical significance for either GOF or RF activity was quantified (Figure 12). Using p<0.05 as the classification threshold, approximately 4.6% of the variants were GOF, 41.8% were RF, and 53.6% were WT (Figure 13). The list of these variants, the number of barcodes, their relative activity, classification, and p-value are reported in the table in Figure 4.
[0071] Seventeen benchmark control variants were tested; one additional variant (K676R) provided was outside the range of its test window. Stable cell lines expressing each of these variants were generated, and their activity was measured individually by flow cytometry (Figure 14). Three variants (Q679L, H878Y, and L726F) were identified that had some additional considerations. L726F was not a GOF variant as determined by flow cytometry profile, and both H878Y and Q679L had GOF activity in their flow profiles, although this was mild compared to most other GOF variants (Figure 14). Results for these variants were generally consistent with other methods. HPAFII cells expressing Q679L showed no increase in p-Her2, whereas its expression in HPNE cells promoted elevated p-Her2 levels, suggesting cell line specificity of this variant's autophosphorylation. Expression of H878Y in HEK293, BaF3, and BEAS-2B cells promoted only a slight increase in p-Her2 expression compared with cells expressing the WT protein. Finally, MCF10A cells expressing L726F exhibited decreased phosphorylation at multiple sites compared with MCF10A cells expressing the WT protein.
[0072] To compare these results with other high-throughput approaches, we performed a multiplexed assays of variant effect (MAYE) enrichment screen for oncogenic variants. In some implementations, only the oncogenic variants G776S and D769Y were significantly enriched in the MAYE assay, which in this case had an accuracy of 41% despite high reproducibility between replicates with an R2 of 0.94 for all variants tested.
[0073] Her2 phosphorylation is measured at a single site, and other pathways may contribute to GOF or RF activity, which can be tested in separate assays.
[0074] Given that the p-Her2 assay was designed to identify GOF variants with increased p-Her2, we observed exceptional assay performance when removing three benchmark variants lacking increased Her2 phosphorylation, similar to published results for HIV Tat [5]. When the GOF Her2 activity data was compared to the benchmark data, it showed high accuracy (Acc = 100%), with other performance metrics = 100% (Table in Figure 5). When the RF Her2 activity data was compared to the benchmark data, it showed high accuracy (Acc = 88%), with other performance metrics in Table 4.
[0075] When the three benchmark variants were included in the analysis, mildly reduced performance was observed. When the GOF Her2 activity data was compared to the benchmark data, it showed high accuracy (Acc=82%), along with the other performance metrics in Table 4. When the RF Her2 activity data was compared to the benchmark data, it also showed accuracy=82%, along with the other performance metrics in the table in Figure 5.
[0076] Although two RF variants (D845A, K753M) were correctly identified as RF in the assay, they did not meet statistical significance for their RF classification. However, it should be noted that this particular assay was not designed to detect LOF variants, and these variants fell just below the statistical significance threshold. These variants had p-values between 0.1 and 0.05, and were the only LOF variants that were not confirmed.
[0077] These variants were plotted on the surface of the crystal structure of the Her2 tyrosine kinase domain (PDBid: 3PP0). The ribbon diagram shows the rainbow from the N-terminus to the C-terminus (blue to red; Figure 15A). The kinase activation loop and ATP-binding site regions are shown for comparison (Figures 15B-C). The 271 GOF variants are shown in Figures 15D-E. PhosphoSite-derived phosphorylation sites are shown in Figure 15F. LOF and RF variants determined using the methods described herein are shown in Figures 15G-H. PhosphoSite-derived ubiquitination sites are shown in Figure 15I. Each position has 19 amino acid substitutions measured. For GOF, at least one substitution has GOF activity (p<0.05), and for RF, at least one substitution has RF activity (p<0.05). It is common to find positions with substitutions that have both RF and GOF activity (see the heatmap in Figure 6). For these plots, counting is performed only if there is a variation at that position, not which of the 19 amino acids are substituted.Only statistically significant variants derived from the method described herein are plotted (p<0.05).No structure of human Her2 has both the tyrosine kinase domain and the JM region.Therefore, for structural evaluation of variants in the JM region, the structure predicted by Alphafold V2.0 is used.The GOF variants and RF variants in the JM region are shown on the Her2 structure predicted by Alphafold V2.0 (Figure 15).
[0078] While exemplary embodiments have been shown and described herein, such embodiments are provided for illustrative purposes only. Numerous variations, modifications, and substitutions are within the scope of this disclosure. It should be understood that various alternatives to the embodiments described herein may be used. It is intended that the appended claims define the scope of the disclosure, and that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
1. 1. A method for identifying Erb-B2 receptor tyrosine kinase 2 (ERBB2) polypeptide variants associated with cancer, comprising: (a) providing a plasmid library for expression of a library of ERBB2 polypeptide variants, each plasmid in the plasmid library comprising: (i) a promoter; (ii) a polynucleotide sequence encoding an ERBB2 polypeptide variant from the library of ERBB2 polypeptide variants, the polynucleotide sequence being operably linked to the promoter, wherein each ERBB2 polypeptide variant in the library of ERBB2 polypeptide variants independently and substantially comprises a single amino acid substitution in a region of the ERBB2 polypeptide ranging from residue 679 to residue 992 of SEQ ID NO:1; and the library of ERBB2 polypeptide variants comprises ERBB2 polypeptide variants collectively having amino acid substitutions of substantially all 20 amino acids at substantially every amino acid residue in a region of the ERBB2 polypeptide ranging from residue 679 to residue 992 of SEQ ID NO:1; and (iii) a barcode, wherein each plasmid in the plasmid library independently has a different barcode associated with a polynucleotide sequence encoding the ERBB2 polypeptide variant in the plasmid. a process comprising: (b) contacting a plurality of mammalian cells with the plasmid library, wherein the contacting results in expression of a single ERBB2 polypeptide variant from the library of ERBB2 polypeptides at the surface of a single mammalian cell from the plurality of mammalian cells, thereby generating a plurality of mammalian cells expressing the library of ERBB2 polypeptide variants, wherein a subset of the plurality of mammalian cells expressing the library of ERBB2 polypeptide variants has an ERBB2 polypeptide variant that is phosphorylated to a higher degree than a wild-type ERBB2 polypeptide expressed on the surface of the mammalian cells, and the ERBB2 polypeptide variant that is phosphorylated to a higher degree than the wild-type ERBB2 polypeptide is the ERBB2 variant associated with the cancer; (c) identifying a subset of mammalian cells that express an ERBB2 polypeptide variant that is phosphorylated to a greater extent than the wild-type ERBB2 polypeptide; and (d) sequencing the barcode of the plasmid present in each mammalian cell of the subset of mammalian cells, thereby identifying an ERBB2 polypeptide variant associated with the cancer. A method comprising:
2. 10. The method of claim 1, wherein the cancer comprises a carcinoma.
3. 10. The method of claim 1, wherein the cancer comprises ovarian cancer, gastric cancer, bladder cancer, salivary gland cancer, or lung cancer.
4. detecting the presence of the ERBB2 polypeptide variant identified in (d) above in a sample obtained from the subject; The method of claim 1 further comprising:
5. diagnosing said subject as having or at risk of developing said cancer. The method of claim 4 further comprising:
6. optionally administering an anti-cancer treatment to the subject based on the presence of the ERBB2 polypeptide variant identified in (d) from a sample obtained from the subject. The method of claim 4 further comprising:
7. The method of claim 6 , wherein the anti-cancer treatment is effective against cancer cells that express the ERBB2 polypeptide variant.
8. 2. The method of claim 1, wherein the identifying step in (c) comprises contacting a plurality of mammalian cells expressing the library of ERBB2 polypeptide variants with an agent that binds to phosphorylated ERBB2.
9. The method of claim 8, wherein the agent that binds to phosphorylated ERBB2 is an antibody against phosphorylated ERBB2.
10. The method of claim 8, wherein the identifying step (c) further comprises performing an immunoassay using an antibody against the phosphorylated ERBB2.
11. The method of claim 10 , wherein the identifying step (c) further comprises performing cell sorting based on the immunoassay.
12. 12. The method of claim 11, wherein the cell sorting is fluorescence-activated cell sorting.
13. The method of claim 1 , wherein the promoter is an inducible promoter.
14. 14. The method of claim 13, wherein the inducible promoter is a doxycycline-inducible promoter.
15. 2. The method of claim 1, wherein the plasmids in the library of plasmids are viral plasmids.
16. 16. The method of claim 15, wherein the viral plasmid is a lentiviral plasmid or an adeno-associated viral plasmid.
17. packaging the viral plasmid into a viral capsid prior to the contacting step (b), thereby producing virions containing the viral plasmid.
16. The method of claim 15, further comprising:
18. 18. The method of claim 17, wherein the contacting step of (b) comprises contacting the virion with a mammalian cell from the plurality of mammalian cells.
19. 2. The method of claim 1, wherein the plurality of mammalian cells comprises HEK293 cells.
20. 10. The method of claim 1, wherein the sequencing step comprises next generation sequencing.
21. a step of stably expressing the cancer-associated ERBB2 polypeptide variant identified in (d) above in mammalian cells and measuring the activity of the ERBB2 polypeptide when stably expressed; The method of claim 1 further comprising:
22. comparing the activity of said ERBB2 polypeptide when stably expressed with that of stably expressed wild-type ERBB2.
22. The method of claim 21 further comprising:
23. 23. The method of claim 22, wherein at least 80% of the variants identified in (d) exhibit increased activity when stably expressed compared to stably expressed wild-type ERBB2.
24. Entering the cancer-associated ERBB2 polypeptide variants identified in (d) into a database. The method of claim 1 further comprising:
25. A database comprising ERBB2 polypeptide variants identified by the method of any one of claims 1 to 24.
26. A composition comprising an ERBB2 polypeptide variant identified by the method of any one of claims 1 to 24.