Kit and its application in constructing recombinant pig cells with tissue-specific expression of human type II 5α-reductase gene
By integrating the human type II 5α-reductase gene into pig cells and driving its expression using the hair follicle tissue-specific promoter KAP6.1, a pig model of hair loss was created. This solves the problem of inconsistencies in existing animal models in simulating human hair loss, provides an experimental tool that is closer to human physiology and pathology, reduces costs, and improves research efficiency.
Patent Information
- Application Number
- CN202210535262.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-17
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2042-05-17
AI Technical Summary
Existing animal models, such as mice and primates, have significant differences in body size, organ size, and physiological and pathological characteristics when simulating human hair loss diseases, especially androgenetic alopecia. They cannot realistically simulate human physiological and pathological states, and their cloning efficiency is low and the cost is high, making them unsuitable as effective experimental tools.
Specific DNA molecules were integrated into the genome of pig cells using the CRISPR/Cas9 system and homologous recombination technology. The expression of human type II 5α-reductase gene was driven by the hair follicle tissue-specific promoter KAP6.1 to prepare recombinant pig cells. A hair loss model pigs were then prepared by somatic cell nuclear transfer cloning.
This study provides a pig model of hair loss that is closer to the physiological and pathological state of humans, which can be used for drug screening, efficacy evaluation and hair loss mechanism research, reducing the difficulty and cost of model making and improving the effectiveness and reliability of experiments.
Smart Images

Figure BDA0003647737970000151 
Figure BDA0003647737970000171 
Figure BDA0003647737970000181
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology and relates to a kit and its application in constructing recombinant porcine cells that specifically express the human type II 5α-reductase gene in hair follicle tissue. The kit provided by this invention uses the CRISPR / Cas9 system and homologous recombination technology to prepare recombinant porcine cells with a specific DNA molecule integrated at a specific location in the genome. The expression of the human type II 5α-reductase gene in this specific DNA molecule is driven by the KAP6.1 promoter. These recombinant cells can be used as nuclear transfer donor cells to prepare cloned pigs, which can then be used as a hair loss model pig. Background Technology
[0002] There are many types of hair loss, among which androgenetic alopecia (AGA) is one of the most common. It is a progressive hair loss disorder characterized by the miniaturization of hair follicles that begins in adolescence or late adolescence. It can affect both men and women, but manifests in different patterns and prevalences. Current research indicates that androgens are the decisive factor in the pathogenesis of AGA, while other factors, including perifollicular inflammation, increased life stress, tension and anxiety, and unhealthy lifestyle and dietary habits, can exacerbate AGA symptoms.
[0003] In men, androgens primarily originate from testosterone secreted by the testes; in women, androgens are mainly synthesized by the adrenal cortex, with the ovaries also secreting small amounts. Androgens are primarily androstenediol, which can be metabolized into testosterone and dihydrotestosterone. Although androgens are a key factor in the pathogenesis of alopecia areata (AGA), almost all AGA patients maintain normal levels of circulating androgens in their blood. Studies have shown that increased expression of androgen receptor genes and / or type II 5α-reductase genes in hair follicles in the bald area leads to increased androgen effects on susceptible hair follicles. In AGA, the dermal cells of susceptible hair follicles contain specific type II 5α-reductases that convert circulating testosterone to dihydrotestosterone. This dihydrotestosterone binds to androgen receptors within the cells, triggering a series of reactions that lead to progressive miniaturization and hair loss, ultimately resulting in baldness.
[0004] Hair loss sufferers often experience psychological problems such as anxiety, depression, and even suicidal thoughts due to damage to their self-image, thus increasing public awareness of its treatment. Current treatments, due to significant side effects or high costs, are not widely accepted. Therefore, exploring new treatment methods is crucial for promoting public awareness and acceptance, which requires relevant animal disease models as experimental tools. Currently, the most common animal model is the mouse model; however, mice differ greatly from humans in body size, organ size, physiology, and pathology, and cannot realistically simulate normal human physiological and pathological states. Pigs, as large animals, have long been a primary source of meat for humans. Their size and physiological functions are similar to humans, making them easy to breed and raise on a large scale. Furthermore, they face lower ethical and animal protection requirements, making them ideal animal models for human diseases.
[0005] Gene editing is a biotechnology that has seen significant development in recent years. It encompasses everything from homologous recombination-based gene editing to nuclease-based technologies such as ZFN, TALEN, and CRISPR / Cas9, with CRISPR / Cas9 currently being the most advanced. Gene editing technology is increasingly being applied to the creation of animal models. Summary of the Invention
[0006] This invention relates to a kit and its application in constructing recombinant porcine cells that specifically express the human type II 5α-reductase gene in hair follicle tissue. The kit provides employs the CRISPR / Cas9 system and homologous recombination technology to prepare recombinant porcine cells with a specific DNA molecule integrated at a specific location in the genome. The expression of the human type II 5α-reductase gene in this specific DNA molecule is driven by the KAP6.1 promoter. These recombinant cells can be used as nuclear transfer donor cells to prepare cloned pigs, which can then be used as a hair loss model pig.
[0007] This invention provides a method for preparing recombinant pig cells, comprising the following steps: integrating a DNA molecule named DNA molecule A into the genomic DNA of a pig cell to obtain recombinant pig cells; wherein the DNA molecule A contains a KAP6.1 promoter and a human type II 5α-reductase gene, and the expression of the human type II 5α-reductase gene is driven by the KAP6.1 promoter.
[0008] Specifically, human type II 5α-reductase (hSRD5A2 protein) is shown in SEQ ID NO: 19.
[0009] The KAP6.1 promoter is a hair follicle-specific promoter that drives the specific expression of downstream genes in hair follicle tissue.
[0010] The KAP6.1 promoter is a sheep hair keratin-binding protein promoter.
[0011] Specifically, the KAP6.1 promoter is shown as nucleotides 1088-2139 in SEQ ID NO: 20.
[0012] Human SRD5A2 gene (hSRD5A2 gene) information: located on human chromosome 2; GeneID is 6716.
[0013] Specifically, the human type II 5α-reductase gene (hSRD5A2 gene) is shown as nucleotides 2140-2904 in SEQ ID NO: 20.
[0014] Specifically, the DNA molecule A contains the hSRD5A2 gene expression cassette.
[0015] In the hSRD5A2 gene expression cassette, the expression of the hSRD5A2 gene is driven by the KAP6.1 promoter.
[0016] In the hSRD5A2 gene expression cassette, there is Poly(A) downstream of the hSRD5A2 gene.
[0017] Specifically, Poly(A) is EF1αPoly(A).
[0018] Specifically, EF1αPoly(A) is shown as nucleotides 2905-3477 in SEQ ID NO: 20.
[0019] Specifically, the hSRD5A2 gene expression cassette is shown as nucleotides 1088-3477 in SEQ ID NO: 20.
[0020] Specifically, the DNA molecule A contains an expression cassette for an antibiotic resistance selection gene.
[0021] The resistance selection gene may be a gene encoding a resistance selection protein.
[0022] The specific resistance screening protein is Neomycin resistance protein.
[0023] Specifically, the resistance selection gene expression cassette is shown as nucleotides 3605-5247 in SEQ ID NO: 20.
[0024] The DNA molecule A also includes a LoxP sequence.
[0025] The DNA molecule A specifically includes two LoxP sequences, as shown in nucleotides 3507-3540 and 5304-5337 of SEQ ID NO: 20, respectively.
[0026] The DNA molecule A also includes insulators.
[0027] The DNA molecule A specifically includes two insulators, as shown in nucleotides 887-1087 and 5358-5559 of SEQ ID NO: 20, respectively.
[0028] The DNA molecule A comprises the following segments from upstream to downstream: KAP6.1 promoter, hSRD5A2 gene, EF1αPoly(A), LoxP sequence, pGK promoter, nucleotide encoding Neomycin resistance protein, bGH Poly(A), and LoxP sequence.
[0029] The DNA molecule A comprises the following segments from upstream to downstream: insulator 1, KAP6.1 promoter, hSRD5A2 gene, EF1αPoly(A), LoxP sequence, pGK promoter, nucleotide encoding Neomycin resistance protein, bGHPoly(A), LoxP sequence, and insulator 5.
[0030] Specifically, the DNA molecule A is shown as nucleotides 887-5559 in SEQ ID NO: 20.
[0031] Specifically, the DNA molecule A is shown as nucleotides 881-5559 in SEQ ID NO: 20.
[0032] The method for "integrating a DNA molecule named DNA molecule A into the genomic DNA of a pig cell" is as follows: a DNA molecule named DNA molecule B is introduced into a pig cell or a recombinant plasmid containing the DNA molecule B is introduced into a pig cell; the DNA molecule B contains the DNA molecule A and has an upstream homologous arm upstream of the DNA molecule A and a downstream homologous arm downstream of the DNA molecule A, the upstream homologous arm and the downstream homologous arm being used to integrate the DNA molecule A into the genomic DNA of the pig cell.
[0033] The homologous arms are homologous arms targeting the COL1A1 gene, with the upstream homologous arm being the left arm of SH4 and the downstream homologous arm being the right arm of SH4. The left arm of SH4 is shown as nucleotides 9-880 in SEQ ID NO: 20, and the right arm of SH4 is shown as nucleotides 5560-6286 in SEQ ID NO: 20.
[0034] The DNA molecule B comprises the following segments from upstream to downstream: upstream homologous arm, insulator 1, KAP6.1 promoter, hSRD5A2 gene, EF1αPoly(A), LoxP sequence, pGK promoter, nucleotide encoding Neomycin resistance protein, bGH Poly(A), LoxP sequence, insulator 5, and downstream homologous arm.
[0035] Specifically, the DNA molecule B is shown as nucleotides 9-6286 in SEQ ID NO: 20.
[0036] Specifically, the recombinant plasmid containing the DNA molecule B is shown in SEQ ID NO: 20.
[0037] Specifically, DNA molecule A integrates into the COL1A1 gene of the pig cell's genomic DNA.
[0038] Specifically, DNA molecule A integrates into the COL1A1 safe harbor insertion site of the genomic DNA in pig cells.
[0039] The integration of DNA molecule A into the COL1A1 gene of the genomic DNA of pig cells refers to the insertion of DNA molecule A between the left and right arms of SH4 in the genomic DNA; the left arm of SH4 is shown as nucleotides 9-880 in SEQ ID NO: 20, and the right arm of SH4 is shown as nucleotides 5560-6286 in SEQ ID NO: 20.
[0040] In the method, the recombinant plasmid containing the DNA molecule B is introduced into pig cells along with two gRNAs (COL1A1-gRNA1 and COL1A1-gRNA3) and the NCN protein.
[0041] The specific mass ratio of the recombinant plasmid containing the DNA molecule B, COL1A1-gRNA1, COL1A1-gRNA3, and NCN protein can be 3:1:1:4.
[0042] The specific ratio of porcine cells, recombinant plasmid containing the DNA molecule B, COL1A1-gRNA1, COL1A1-gRNA3, and NCN protein can be: 100,000 primary porcine fibroblasts: 3 μg recombinant plasmid: 1 μg COL1A1-gRNA1: 1 μg COL1A1-gRNA3: 4 μg NCN protein.
[0043] Specifically, the COL1A1 safe harbor insertion site and its surrounding region in the pig genome are shown in SEQ ID NO: 25.
[0044] The recombinant pig cells are homozygous recombinants (i.e., DNA molecules A are integrated into the same position on two homologous chromosomes).
[0045] The recombinant pig cells are heterozygous recombinations (i.e., DNA molecule A is integrated into a homologous chromosome).
[0046] This invention also protects a kit comprising any of the DNA molecules described above.
[0047] The present invention also protects a kit comprising a recombinant plasmid having any of the DNA molecules B described above.
[0048] The kit also includes two gRNAs (COL1A1-gRNA1 and COL1A1-gRNA3).
[0049] The kit also includes NCN protein.
[0050] The specific mass ratio of the recombinant plasmid containing the DNA molecule B, COL1A1-gRNA1, COL1A1-gRNA3, and NCN protein can be 3:1:1:4.
[0051] The kit also includes PRONCN protein.
[0052] The kit also includes specific plasmids.
[0053] The kit also includes porcine cells.
[0054] This invention also protects the use of any of the above-described DNA molecules B in the preparation kit.
[0055] The present invention also protects the use of recombinant plasmids having any of the above-described DNA molecules B in preparation kits.
[0056] The present invention also protects the use of recombinant plasmids having any of the above-described DNA molecules B, two gRNAs (COL1A1-gRNA1 and COL1A1-gRNA3), and NCN protein in the preparation kit.
[0057] The above-described kits are used for the following purposes (a) or (b): (a) to prepare recombinant porcine cells that specifically express the human type II 5α-reductase gene in hair follicle tissue; (b) to prepare a hair loss model pig.
[0058] The present invention also protects the application of any of the above-described DNA molecules B, recombinant plasmids having any of the above-described DNA molecules B, or any of the above-described kits, as follows (a) or (b): (a) preparing recombinant porcine cells that specifically express the human type II 5α-reductase gene in hair follicle tissue; (b) preparing a hair loss model pig.
[0059] gRNA is also known as sgRNA.
[0060] The target sequence binding region of COL1A1-gRNA1 is shown as nucleotides 3-22 in SEQ ID NO: 23. Specifically, COL1A1-gRNA1 is shown in SEQ ID NO: 23.
[0061] The target sequence binding region of COL1A1-gRNA3 is shown as nucleotides 3-22 in SEQ ID NO: 24. Specifically, COL1A1-gRNA3 is shown in SEQ ID NO: 24.
[0062] The target sequence binding region refers to the region in sgRNA that binds to the target sequence (which is located in the target region of the target gene).
[0063] The NCN protein is a Cas9 protein or a fusion protein containing a Cas9 protein.
[0064] This invention also protects recombinant porcine cells prepared by any of the methods described above.
[0065] Specifically, the recombinant porcine cells may be recombinant porcine cells as follows: compared with porcine cells, the only difference in the genomic DNA of the recombinant porcine cells is that the DNA molecule represented by nucleotides 881-5559 of SEQ ID NO: 20 is inserted between the left and right arms of SH4 in the cell genomic DNA.
[0066] The recombinant pig cells are homozygous recombinants (i.e., DNA molecules A are integrated into the same position on two homologous chromosomes).
[0067] The recombinant pig cells are heterozygous recombinations (i.e., DNA molecule A is integrated into a homologous chromosome).
[0068] The recombinant porcine cells were homozygous recombinants (i.e., the DNA molecule represented by nucleotides 881-5559 of SEQ ID NO: 20 was inserted between the left and right arms of SH4 on both homologous chromosomes).
[0069] The recombinant porcine cells are heterozygous recombinants (i.e., the DNA molecule represented by nucleotides 881-5559 of SEQ ID NO: 20 is inserted between the left and right arms of the SH4 chromosome of a homologous chromosome).
[0070] This invention also protects the use of the recombinant porcine cells in the preparation of a hair loss model pigs.
[0071] Using the recombinant porcine cells as nuclear transfer donor cells for somatic cell cloning, cloned pigs can be obtained, which are the alopecia model pigs. These model pigs can be used in further biomedical fields such as drug screening and efficacy evaluation, gene and cell therapy, and research on the mechanisms of hair loss. These model pigs can provide a powerful tool for studying the pathogenesis of human AGA or exploring related treatment methods.
[0072] Specifically, the NCN protein is shown in SEQ ID NO: 3.
[0073] The method for preparing the NCN protein includes the following steps:
[0074] (1) Plasmid pKG-GE4 was introduced into Escherichia coli BL21(DE3) to obtain recombinant bacteria;
[0075] (2) The recombinant bacteria were cultured in liquid culture medium at 30°C, then IPTG was added and induced at 25°C, and then the bacterial cells were collected.
[0076] (3) The collected bacterial cells were broken down to collect the crude protein solution;
[0077] (4) The His6-tagged fusion protein was purified from the crude protein solution by affinity chromatography;
[0078] (5) The His6-tagged fusion protein was digested with His6-tagged enterokinase, and then the His6-tagged protein was removed with Ni-NTA resin to obtain purified NCN protein.
[0079] The plasmid pKG-GE4 contains the fusion gene shown in nucleotides 5209-9852 of SEQ ID NO: 1.
[0080] The preparation method of the NCN protein specifically includes the following steps:
[0081] (1) Plasmid pKG-GE4 was introduced into Escherichia coli BL21(DE3) to obtain recombinant bacteria.
[0082] (2) Inoculate the recombinant bacteria obtained in step (1) into liquid LB medium containing ampicillin and culture with shaking;
[0083] (3) Inoculate the bacterial culture obtained in step (2) into liquid LB medium and culture at 30°C with shaking at 230 rpm until OD. 600nm The value was 1.0, then IPTG was added to make the concentration in the system 0.5mM, and then the cells were cultured at 25℃ and 230rpm for 12 hours with shaking, and then the cells were collected by centrifugation.
[0084] (4) Take the bacterial cells obtained in step (3) and wash them with PBS buffer;
[0085] (5) Take the bacterial cells obtained in step (4), add crude extraction buffer and suspend the bacterial cells, then break the bacterial cells, then centrifuge and collect the supernatant, filter with a 0.22 μm pore size filter membrane and collect the filtrate;
[0086] (6) The His6-tagged fusion protein (the fusion protein shown in SEQ ID NO: 2) was purified from the filtrate obtained in step (5) by affinity chromatography;
[0087] (7) Take the column-passed solution collected in step (6), concentrate it using an ultrafiltration tube, and then dilute it with 25 mM Tris-HCl (pH 8.0);
[0088] (8) Add the recombinant bovine enterokinase with the His6 tag to the solution obtained in step (7) and digest it with enzymes;
[0089] (9) Mix the solution from step (8) with Ni-NTA resin, incubate, and then centrifuge to collect the supernatant.
[0090] (10) Take the supernatant obtained in step (9), concentrate it using an ultrafiltration tube, and then add it to the enzyme storage solution to obtain the NCN protein solution.
[0091] The specific method for purifying the His6-tagged fusion protein from the filtrate obtained in step (5) using affinity chromatography is as follows:
[0092] First, equilibrate the Ni-NTA agarose column with 5 column volumes of equilibration buffer (flow rate: 1 ml / min); then load 50 ml of the filtrate obtained in step (5) (flow rate: 0.5-1 ml / min); then wash the column with 5 column volumes of equilibration buffer (flow rate: 1 ml / min); then wash the column with 5 column volumes of buffer (flow rate: 1 ml / min) to remove contaminating proteins; then elute with 10 column volumes of elution buffer at a flow rate of 0.5-1 ml / min, and collect the post-column solution (90-100 ml).
[0093] The PRONCN protein described above comprises the following components from upstream to downstream: signal peptide, molecular chaperone protein, protein tag, protease cleavage site, nuclear localization signal, Cas9 protein, and nuclear localization signal.
[0094] The function of the signal peptide is to promote the secretory expression of a protein. The signal peptide can be selected from the Escherichia coli alkaline phosphatase (phoA) signal peptide, Staphylococcus aureus protein A signal peptide, Escherichia coli outer membrane protein (ompa) signal peptide, or any other prokaryotic gene signal peptide, preferably the alkaline phosphatase signal peptide (phoA signal peptide). The alkaline phosphatase signal peptide is used to guide the secretory expression of the target protein into the bacterial periplasmic lumen, thereby separating it from the intracellular protein. The target protein secreted into the bacterial periplasmic lumen is soluble and can be cleaved by the signal peptidase in the bacterial periplasmic lumen.
[0095] The function of the molecular chaperone protein is to increase the solubility of the protein. The molecular chaperone can be any protein that helps form disulfide bonds, preferably a thioreduction protein (TrxA protein). A thioreduction protein, acting as a molecular chaperone, helps the co-expressed target protein (e.g., Cas9 protein) form disulfide bonds, improving protein stability, correct folding, and increasing the solubility and activity of the target protein.
[0096] The protein tag is used for protein purification. The tag can be a His tag (His-Tag, His6 protein tag), GST tag, Flag tag, HA tag, c-Myc tag, or any other protein tag, with a His tag being more preferred. The His tag can bind to a Ni column, enabling one-step Ni column affinity chromatography to purify the target protein, greatly simplifying the purification process.
[0097] The function of the protease cleavage site is to cleave the non-functional segment after purification to release the native form of Cas9 protein. The protease can be selected from enterokinase, factor Xa, thrombin, TEV protease, HRV 3C protease, WELQut protease, or any other endopeptide, with enterokinase being more preferred. EK is an enterokinase cleavage site, facilitating the cleavage of the fused TrxA-His segment using enterokinase to obtain the native form of Cas9 protein. In this application, after cleaving the fusion protein with a His-tagged commercial enterokinase, the TrxA-His segment and the His-tagged enterokinase can be removed by a single affinity chromatography to obtain the native form of Cas9 protein, avoiding the damage and loss of the target protein caused by multiple purification dialysis processes.
[0098] The nuclear localization signal can be any nuclear localization signal, preferably the SV40 nuclear localization signal and / or the nucleoplasmin nuclear localization signal. The NLS is the nuclear localization signal; an NLS site is designed at both the N-terminus and C-terminus of Cas9, enabling Cas9 to more effectively enter the cell nucleus for gene editing.
[0099] The Cas9 protein may be saCas9 or spCas9, preferably spCas9 protein.
[0100] The PRONCN protein is shown in SEQ ID NO: 2.
[0101] Each of the above-mentioned specific plasmids comprises the following elements from upstream to downstream: promoter, operon, ribosome binding site, gene encoding PRONCN protein, and terminator.
[0102] The promoter may specifically be the T7 promoter. The T7 promoter is a strong prokaryotic expression promoter that can efficiently drive the expression of exogenous genes.
[0103] The operon can specifically be the Lac operon. The Lac operon is a regulatory element for lactose-induced expression. After the bacteria have grown to a certain number, the expression of the target protein can be induced by IPTG at low temperature, which can avoid the impact of premature expression of the target protein on the growth of the host bacteria. Induction at low temperature also significantly improves the solubility of the expressed target protein.
[0104] The ribosome binding site is the ribosome binding site during protein translation, which is essential for protein translation.
[0105] The terminator can specifically be a T7 terminator. The T7 terminator can effectively terminate gene transcription at the end of the target gene, preventing other downstream sequences outside the target gene from being transcribed and translated.
[0106] For the codons of spCas9 protein, this application has optimized the codons to fully adapt to the codon preferences of the high-efficiency E. coli expression strain E. coli BL21(DE3) selected in this application, thereby improving the expression level of Cas9 protein.
[0107] The T7 promoter is shown as nucleotides 5121-5139 in SEQ ID NO: 1.
[0108] The Lac operon is shown as nucleotides 5140-5164 in SEQ ID NO: 1.
[0109] The ribosome binding site is shown as nucleotides 5178-5201 in SEQ ID NO: 1.
[0110] The coding sequence of the alkaline phosphatase signal peptide is shown as nucleotides 5209-5271 in SEQ ID NO: 1.
[0111] The coding sequence of the TrxA protein is shown as nucleotides 5272-5598 in SEQ ID NO: 1.
[0112] The coding sequence of His-Tag is shown as nucleotides 5620-5637 in SEQ ID NO: 1.
[0113] The coding sequence of the enterokinase cleavage site is shown as nucleotides 5638-5652 in SEQ ID NO: 1.
[0114] The coding sequence of the nuclear localization signal is shown as nucleotides 5656-5670 in SEQ ID NO: 1.
[0115] The coding sequence of the spCas9 protein is shown as nucleotides 5701-9801 in SEQ ID NO: 1.
[0116] The coding sequence of the nuclear localization signal is shown as nucleotides 9802-9849 in SEQ ID NO: 1.
[0117] The T7 terminator is nucleotides 9902-9949 in SEQ ID NO: 1.
[0118] Specifically, the specific plasmid is plasmid pKG-GE4.
[0119] The plasmid pKG-GE4 contains the DNA molecule represented by nucleotides 5121-9949 of SEQ ID NO: 1.
[0120] Specifically, any of the plasmids pKG-GE4 described above is shown in SEQ ID NO: 1.
[0121] The pig cells are derived from male pigs.
[0122] The pig cells are somatic cells derived from male pigs.
[0123] The porcine cells mentioned are primary porcine fibroblasts.
[0124] The porcine cells are primary fibroblasts derived from male pigs.
[0125] The pig can be any breed, but preferably, it can be a Congjiang Xiang pig.
[0126] The pigs mentioned can specifically refer to newborn pigs.
[0127] Compared with the prior art, the present invention has at least the following beneficial effects:
[0128] (1) The research object of this invention (pig) has better applicability than other animals (mice, mice, primates).
[0129] Rodents such as mice and rats differ greatly from humans in body size, organ size, physiology, and pathology, making it impossible to realistically simulate normal human physiological and pathological states. Studies have shown that over 95% of drugs proven effective in mice and rats are ineffective in human clinical trials. Among large animals, primates are the closest relatives to humans, but they are small, reach sexual maturity late (mating begins at 6-7 years old), and are single-birth animals, resulting in extremely slow population expansion and high rearing costs. Furthermore, primate cloning is inefficient, difficult, and costly.
[0130] Pigs, as model animals, do not have the aforementioned drawbacks. Pigs are the closest relatives to humans besides primates, sharing similar body size, weight, and organ size with humans. They are also remarkably similar to humans in anatomy, physiology, immunology, nutritional metabolism, and disease pathogenesis. Furthermore, pigs reach sexual maturity early (4-6 months), have high reproductive capacity, producing multiple offspring per litter, and can form a large herd within 2-3 years. In addition, pig cloning technology is highly advanced, and the costs of cloning and raising pigs are much lower than those for primates.
[0131] (2) This invention explored the expression of four safe harbor sites in the pig genome after gene knock-in, and selected the best safe harbor sites in the pig genome for foreign gene insertion, which can effectively improve the expression of the target gene after gene knock-in.
[0132] (3) Using the homozygous knock-in KAP6.1-hSRD5A2 single-cell clone obtained in this invention, somatic cell nuclear transfer animal cloning can directly produce KAP6.1-hSRD5A2 homozygous knock-in cloned pigs, and the homozygous inserted gene can be stably inherited. Furthermore, it can be used in the biomedical field for further drug screening and efficacy evaluation in AGA, gene and cell therapy, pathogenesis research, etc.
[0133] In mouse model creation, the common practice is to microinject gene-edited material into fertilized eggs followed by embryo transfer. However, the probability of directly obtaining homozygous mutant offspring is extremely low (less than 5%), necessitating crossbreeding and selection of the offspring. This approach is not well-suited for creating models of large animals (such as pigs) with long gestation periods. Therefore, this invention employs a technically challenging method of in vitro editing of primary cells and screening for positive edited single-cell clones. Subsequently, somatic cell nuclear transfer animal cloning technology is used to directly obtain the corresponding model pigs. This significantly shortens the model pig creation cycle and saves manpower, material resources, and financial resources.
[0134] (4) This invention is the first to integrate the human type II 5α-reductase gene into the pig genome and express it specifically in hair follicles, and a homozygous knock-in single-cell clone has been obtained after verification.
[0135] This invention utilizes gene editing technology to obtain recombinant porcine cells that specifically express human type II 5α-reductase in hair follicle tissue. These recombinant cells can be used as nuclear transfer cell donors to clone and produce alopecia model pigs, which will contribute to the study and elucidation of the pathogenesis of alopecia areata (AGA). Furthermore, they can be used for drug screening, efficacy testing, gene and cell therapy research, and other studies, providing effective experimental data for further clinical applications and thus offering powerful experimental tools for the prevention and treatment of human AGA. This invention has significant application value for the study of the pathogenesis of human AGA, drug development, and preclinical trials. Attached Figure Description
[0136] Figure 1 This is a schematic diagram of the structure of plasmid pET-32a.
[0137] Figure 2 This is a schematic diagram of the structure of plasmid pKG-GE4.
[0138] Figure 3 This is an electrophoresis diagram showing the optimized ratio of gRNA to NCN protein in Example 1.
[0139] Figure 4 This is an electrophoresis diagram comparing the gene editing efficiency of NCN protein and commercial Cas9 protein in Example 1.
[0140] Figure 5 This is a schematic diagram of the structure of plasmid PB-1G 2R 3-puro-ROSA26.
[0141] Figure 6 Diagram showing the regulation of GFP green fluorescence expression at different safe harbor sites.
[0142] Figure 7 Results of quantitative real-time PCR for regulating GFP gene transcription levels at different safe harbor sites.
[0143] Figure 8 FACS detection results for regulating GFP protein expression at different safe harbor sites.
[0144] Figure 9 This is a schematic diagram of the structure of plasmid KAP6.1-hSRD5A2.
[0145] Figure 10 The sequencing results show the linkage sequence between the COL1A1 gene sequence and insulator 1.
[0146] Figure 11The sequencing results show the linkage sequence between insulator 1 and the EF1α promoter.
[0147] Figure 12 The sequencing results show the linkage sequence between the EF1α promoter and the human SRD5A2 gene.
[0148] Figure 13 The sequencing results show the linkage sequence between the human SRD5A2 gene and EF1αPoly(A).
[0149] Figure 14 Sequencing results for the linkage sequence of EF1αPoly(A) and LoxP. Detailed Implementation
[0150] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.
[0151] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available. The recombinant plasmids constructed in the examples have all been sequenced and verified. Complete culture medium (% by volume): 15% fetal bovine serum (Gibco) + 83% DMEM medium (Gibco) + 1% Penicillin-Streptomycin (Gibco) + 1% HEPES (Solarbio). Cell culture conditions: 37°C, incubator with 5% CO2 and 5% O2.
[0152] The primary porcine fibroblasts used in Examples 1 and 2 were prepared from the ear tissue of newly hatched Jiangxiang pigs. The primary porcine fibroblasts used in Example 3 were prepared from the ear tissue of newly hatched male Jiangxiang pigs. Method for preparing primary porcine fibroblasts: ① Take 0.5g of porcine ear tissue, remove hair and bone tissue, then soak in 75% alcohol for 30-40s, wash 5 times with PBS buffer containing 5% (v / v) Penicillin-Streptomycin (Gibco), and then wash once with PBS buffer; ② Cut the tissue into small pieces with scissors, digest with 5mL of 0.1% collagenase solution (Sigma) at 37℃ for 1h, then centrifuge at 500g for 5min and discard the supernatant; ③ Resuspend the pellet in 1mL of complete culture medium, then plate it into a 10cm diameter cell culture dish containing 10mL of complete culture medium and sealed with 0.2% gelatin (VWR), and culture until the cells reach approximately 60% confluence with the bottom of the dish; ④ After completing step ③, digest with trypsin and collect the cells, then resuspend them in complete culture medium for subsequent electroporation experiments.
[0153] The plasmid pKG-GE3 is a circular plasmid, as shown in SEQ ID NO: 2 of patent application 202010084343.6. (SEQ ID NO: 2 in patent application 202010084343.6) In NO:2, nucleotides 395-680 form the CMV enhancer, nucleotides 682-890 form the EF1a promoter, nucleotides 986-1006 encode the nuclear localization signal (NLS), nucleotides 1016-1036 encode the nuclear localization signal (NLS), nucleotides 1037-5161 encode the Cas9 protein, nucleotides 5162-5209 encode the nuclear localization signal (NLS), nucleotides 5219-5266 encode the nuclear localization signal (NLS), and nucleotides 5276-5332 encode the self-cleaving polypeptide P2A (the amino acid sequence of the self-cleaving polypeptide P2A is "ATNFSLLKQAGDVEENPGP", and the self-cleavage occurs at the following location: Nucleotides 5333-6046 (between the first and second amino acid residues at the C-terminus) encode the EGFP protein, nucleotides 6056-6109 encode the self-cleaving polypeptide T2A (the amino acid sequence of the self-cleaving polypeptide T2A is “EGRGSLLTCGDVEENPGP”, and the self-cleavage occurs between the first and second amino acid residues at the C-terminus), nucleotides 6110-6703 encode the Puromycin protein (abbreviated as Puro protein), nucleotides 6722-7310 form the WPRE sequence element, nucleotides 7382-7615 form the 3'LTR sequence element, and nucleotides 7647-7871 form the bGH poly(A)signal sequence element. In SEQ ID NO: 2 of patent application 202010084343.6, nucleotides 911-6706 form a fusion gene to express the fusion protein. Due to the presence of the self-cleaving peptide P2A and the self-cleaving peptide T2A, the fusion protein spontaneously forms the following three proteins: a protein with Cas9 protein, a protein with EGFP protein, and a protein with Puro protein.
[0154] The pKG-U6gRNA vector, or plasmid pKG-U6gRNA, is a circular plasmid, as shown in SEQ ID NO: 3 of patent application 202010084343.6. In SEQ ID NO: 3 of patent application 202010084343.6, nucleotides 2280-2539 form the hU6 promoter, and nucleotides 2558-2637 are used for transcription to form the gRNA backbone. In use, a DNA molecule of approximately 20 bp (the target sequence binding region for gRNA transcription) is inserted into the plasmid pKG-U6gRNA to form a recombinant plasmid. The recombinant plasmid is then transcribed into gRNA in cells.
[0155] Example 1: Preparation, purification, and performance of NCN protein
[0156] I. Construction of a high-efficiency prokaryotic Cas9 expression vector
[0157] A schematic diagram of the structure of plasmid pET-32a is shown below. Figure 1 .
[0158] Plasmid pKG-GE4 was obtained by modifying plasmid pET-32a. Plasmid pET32a-T7lac-phoA:SP-TrxA-His-EK-NLS-spCas9-NLS-T7ter (abbreviated as plasmid pKG-GE4), as shown in SEQ ID NO: 1, is a circular plasmid; its structural diagram is shown below. Figure 2 .
[0159] In SEQ ID NO: 1, nucleotides 5121-5139 form the T7 promoter, nucleotides 5140-5164 encode the Lac operator, nucleotides 5178-5201 form the ribosome binding site (RBS), nucleotides 5209-5271 encode the alkaline phosphatase signal peptide (phoA signal peptide), nucleotides 5272-5598 encode the TrxA protein, nucleotides 5620-5637 encode the His-Tag (also known as the His6 tag), nucleotides 5638-5652 encode the enterokinase cleavage site (EK cleavage site), nucleotides 5656-5670 encode the nuclear localization signal, nucleotides 5701-9801 encode the spCas9 protein, nucleotides 9802-9849 encode the nuclear localization signal, and nucleotides 9902-9949 form the T7 terminator. The nucleotides encoding the spCas9 protein have been codon-optimized for Escherichia coli BL21(DE3) strain.
[0160] The main modifications to plasmid pKG-GE4 are as follows: ① The coding region of the TrxA protein was retained. The TrxA protein can help the expressed target protein form disulfide bonds, increasing the solubility and activity of the target protein. An alkaline phosphatase signal peptide coding sequence was added before the TrxA protein coding region. The alkaline phosphatase signal peptide can guide the expressed target protein to be secreted into the bacterial periplasmic lumen and can be cleaved by prokaryotic periplasmic signal peptidase. ② A His-Tag coding sequence was added after the TrxA protein coding sequence. The His-Tag can be used for... Enrichment of the target protein; ③ Add the coding sequence of the enterokinase cleavage site DDDDK (Asp-Asp-Asp-Asp-Lys) downstream of the His-Tag coding sequence. The purified protein will remove His-Tag and the upstream fused TrxA protein under the action of enterokinase; ④ Insert the Cas9 gene of suitable Escherichia coli BL21(DE3) strain with optimized codons, and add nuclear localization signal coding sequences upstream and downstream of this gene to increase the nuclear localization ability of the purified Cas9 protein in the later stage.
[0161] The fusion gene in plasmid pKG-GE4, as shown in nucleotides 5209-9852 of SEQ ID NO: 1, encodes the fusion protein shown in SEQ ID NO: 2 (fusion protein TrxA-His-EK-NLS-spCas9-NLS, abbreviated as PRONCN protein). Due to the presence of alkaline phosphatase signal peptide and enterokinase cleavage site, the fusion protein is cleaved by enterokinase to form the protein shown in SEQ ID NO: 3. The protein shown in SEQ ID NO: 3 is named NCN protein.
[0162] II. Induced Expression
[0163] 1. Plasmid pKG-GE4 was introduced into Escherichia coli BL21(DE3) to obtain recombinant bacteria.
[0164] 2. Inoculate the recombinant bacteria obtained in step 1 into liquid LB medium containing 100 μg / ml ampicillin and culture overnight at 37°C with shaking at 200 rpm.
[0165] 3. Inoculate the bacterial culture obtained in step 2 into liquid LB medium and incubate at 30°C with shaking at 230 rpm until OD reaches 100%. 600nm The concentration was set to 1.0, then isopropyl thiogalactoside (IPTG) was added to a concentration of 0.5 mM in the system. The mixture was then cultured at 25°C and 230 rpm for 12 hours with shaking. Finally, the cells were collected by centrifugation at 4°C and 10,000 g for 15 minutes.
[0166] 4. Take the bacterial cells obtained in step 3 and wash them with PBS buffer.
[0167] III. Purification of the fusion protein TrxA-His-EK-NLS-spCas9-NLS
[0168] 1. Take the bacterial cells obtained in step 2, add crude extraction buffer and suspend the bacterial cells, then homogenize the bacterial cells using a homogenizer (3 cycles at 1000 par), then centrifuge at 4℃ and 15000g for 30 min, collect the supernatant, filter the supernatant through a 0.22μm pore size filter membrane, and collect the filtrate. In this step, 10 ml of crude extraction buffer is prepared for every gram of wet bacterial cells.
[0169] Crude extraction buffer: containing 20mM Tris-HCl (pH 8.0), 0.5M NaCl, 5mM Imidazole, 1mM PMSF, with the balance being ddH2O.
[0170] 2. Affinity chromatography was used to purify the fusion protein.
[0171] First, equilibrate the Ni-NTA agarose column with 5 column volumes of equilibration buffer (flow rate: 1 ml / min); then load 50 ml of the filtrate obtained in step 1 (flow rate: 0.5-1 ml / min); then wash the column with 5 column volumes of equilibration buffer (flow rate: 1 ml / min); then wash the column with 5 column volumes of buffer (flow rate: 1 ml / min) to remove contaminating proteins; finally, elute with 10 column volumes of elution buffer at a flow rate of 0.5-1 ml / min, and collect the post-column solution (90-100 ml).
[0172] Ni-NTA agarose column: GenScript, L00250 / L00250-C, 10ml packing material.
[0173] Equilibrium solution: contains 20 mM Tris-HCl (pH 8.0), 0.5 M NaCl, 5 mM Imidazole, and the balance is ddH2O.
[0174] Buffer solution: containing 20 mM Tris-HCl (pH 8.0), 0.5 M NaCl, 50 mM Imidazole, with the balance being ddH2O.
[0175] Eluent: Contains 20 mM Tris-HCl (pH 8.0), 0.5 M NaCl, 500 mM Imidazole, and the balance is ddH2O.
[0176] IV. Enzymatic digestion of the fusion protein TrxA-His-EK-NLS-spCas9-NLS and purification of the NCN protein
[0177] 1. Take 15 ml of the post-column solution collected in step 3, concentrate it to 200 μl using an Amicon ultrafiltration tube (Sigma, UFC9100, 15 ml capacity), and then dilute it to 1 ml with 25 mM Tris-HCl (pH 8.0). Use 6 ultrafiltration tubes to obtain a total of 6 ml.
[0178] 2. Add the commercially available His6-tagged recombinant bovine enterokinase (Sangon Biotech, C620031, Recombinant Bovine Enterokinase Light Chain, His6-tagged) to the solution obtained in step 1 (approximately 6 ml), and digest at 25°C for 16 hours. Add 2 units of enterokinase per 50 μg of protein.
[0179] 3. Take the solution from step 2 (about 6 ml), mix it with 480 μl of Ni-NTA resin (GenScript, L00250 / L00250-C), mix by rotation at room temperature for 15 min, then centrifuge at 7000 g for 3 min, and collect the supernatant (4-5.5 ml).
[0180] 4. Take the supernatant obtained in step 3 and concentrate it to 200 μl using an Amicon ultrafiltration tube (Sigma, UFC9100, capacity 15 ml). Then add it to the enzyme storage solution and adjust the protein concentration to 5 mg / ml to obtain the NCN protein solution.
[0181] Sequencing revealed that the N-terminal 15 amino acid residues in the NCN protein solution are as shown in positions 1 to 15 of SEQ ID NO: 3, which is the NCN protein.
[0182] The NCN protein used in subsequent steps and examples is provided by an NCN protein solution.
[0183] Enzyme stock solution (pH 7.4): contains 10 mM Tris, 300 mM NaCl, 0.1 mM EDTA, 1 mM DTT, 50% (v / v) glycerol, with the balance being ddH2O.
[0184] V. Performance of NCN Protein
[0185] The following two gRNA targets targeting the TTN gene were selected:
[0186] TTN-gRNA1: AGAGCACAGTCAGCCTGGCG;
[0187] TTN-gRNA2: CTTCCAGAATTGGATCTCCG.
[0188] The primers used to identify target fragments containing gRNA from the TTN gene are as follows:
[0189] TTN-F55: TACGGAATTGGGGAGCCAGCGGA;
[0190] TTN-R560: CAAAGTTAACTCTCTGTGTCT.
[0191] 1. Preparation of gRNA
[0192] (1) Preparation of TTN-T7-gRNA1 transcription template and TTN-T7-gRNA2 transcription template
[0193] The TTN-T7-gRNA1 transcription template is a double-stranded DNA molecule, as shown in SEQ ID NO: 4.
[0194] The TTN-T7-gRNA2 transcription template is a double-stranded DNA molecule, as shown in SEQ ID NO: 5.
[0195] (2) Obtain gRNA by in vitro transcription
[0196] Using TTN-T7-gRNA1 as a transcription template, in vitro transcription was performed using the Transcript Aid T7 High Yield Transcription Kit (Fermentas, K0441), followed by MEGA clearing. TM The TTN-gRNA1 was recovered and purified using a Transcription Clean-Up Kit (Thermo, AM1908). TTN-gRNA1 is a single-stranded RNA, as shown in SEQ ID NO: 6.
[0197] Using TTN-T7-gRNA2 as a transcription template, in vitro transcription was performed using the Transcript Aid T7 High Yield Transcription Kit (Fermentas, K0441), followed by MEGA clearing. TM The TTN-gRNA2 was recovered and purified using a Transcription Clean-Up Kit (Thermo, AM1908). TTN-gRNA2 is a single-stranded RNA, as shown in SEQ ID NO: 7.
[0198] 2. Optimization of the ratio of gRNA to NCN protein dosage
[0199] (1) Co-transfection of porcine primary fibroblasts
[0200] Group 1: TTN-gRNA1, TTN-gRNA2, and NCN protein were co-transfected into porcine primary fibroblasts. The ratio was approximately 100,000 porcine primary fibroblasts: 0.5 μg TTN-gRNA1 : 0.5 μg TTN-gRNA2 : 4 μg NCN protein.
[0201] Group 2: TTN-gRNA1, TTN-gRNA2, and NCN protein were co-transfected into porcine primary fibroblasts. The ratio was approximately 100,000 porcine primary fibroblasts: 0.75 μg TTN-gRNA1 : 0.75 μg TTN-gRNA2 : 4 μg NCN protein.
[0202] Group 3: TTN-gRNA1, TTN-gRNA2, and NCN protein were co-transfected into porcine primary fibroblasts. The ratio was approximately 100,000 porcine primary fibroblasts: 1 μg TTN-gRNA1 : 1 μg TTN-gRNA2 : 4 μg NCN protein.
[0203] Group 4: TTN-gRNA1, TTN-gRNA2, and NCN protein were co-transfected into porcine primary fibroblasts. The ratio was approximately 100,000 porcine primary fibroblasts: 1.25 μg TTN-gRNA1 : 1.25 μg TTN-gRNA2 : 4 μg NCN protein.
[0204] Group 5: TTN-gRNA1 and TTN-gRNA2 were co-transfected into porcine primary fibroblasts. Ratio: approximately 100,000 porcine primary fibroblasts: 1 μg TTN-gRNA1: 1 μg TTN-gRNA2.
[0205] Co-transfection was performed using electroporation with a mammalian nuclear transfection kit (Neon kit, Thermofisher) and a Neon™ transfection system (parameters set to 1450V, 10ms, 3 pulses).
[0206] (2) After completing step (1), culture in complete culture medium for 12-18 hours, then replace with new complete culture medium for further culture. The total culture time after electroporation is 48 hours.
[0207] (3) After completing step (2), cells were digested and collected with trypsin, genomic DNA was extracted, and PCR amplification was performed using primers consisting of TTN-F55 and TTN-R560. Then, 1% agarose gel electrophoresis was performed.
[0208] See electrophoresis image Figure 3The 505bp band is the wild-type band (WT), and the band around 254bp (the wild-type band theoretically has a 251bp deletion) is the deletion mutation band (MT).
[0209] Gene deletion mutation efficiency = (MT gray level / MT band bp) / (WT gray level / WT band bp + MT gray level / MT band bp) × 100%. The gene deletion mutation efficiency of the first group is 19.9%, the gene deletion mutation efficiency of the second group is 39.9%, the gene deletion mutation efficiency of the third group is 79.9%, and the gene deletion mutation efficiency of the fourth group is 44.3%. No mutation occurred in the fifth group.
[0210] The results showed that the gene editing efficiency was highest when the mass ratio of the two gRNAs to the NCN protein was 1:1:4, and the actual dosage was 1 μg:1 μg:4 μg. Therefore, the optimal dosage of the two gRNAs to the NCN protein was determined to be 1 μg:1 μg:4 μg.
[0211] 3. Comparison of gene editing efficiency between NCN protein and commercial Cas9 protein
[0212] (1) Co-transfection of porcine primary fibroblasts
[0213] Cas9-A group: TTN-gRNA1, TTN-gRNA2, and commercial Cas9-A protein were co-transfected into porcine primary fibroblasts. Ratio: approximately 100,000 porcine primary fibroblasts: 1 μg TTN-gRNA1 : 1 μg TTN-gRNA2 : 4 μg Cas9-A protein.
[0214] pKG-GE4 group: TTN-gRNA1, TTN-gRNA2, and NCN protein were co-transfected into porcine primary fibroblasts. Ratio: approximately 100,000 porcine primary fibroblasts: 1 μg TTN-gRNA1 : 1 μg TTN-gRNA2 : 4 μg NCN protein.
[0215] Cas9-B group: TTN-gRNA1, TTN-gRNA2, and commercial Cas9-B protein were co-transfected into porcine primary fibroblasts. Ratio: approximately 100,000 porcine primary fibroblasts : 1 μg TTN-gRNA1 : 1 μg TTN-gRNA2 : 4 μg Cas9-B protein.
[0216] Control group: porcine primary fibroblasts were co-transfected with TTN-gRNA1 and TTN-gRNA2. Ratio: approximately 100,000 porcine primary fibroblasts: 1 μg TTN-gRNA1 : 1 μg TTN-gRNA2.
[0217] Co-transfection was performed using electroporation with a mammalian nuclear transfection kit (Neon kit, Thermofisher) and a Neon™ transfection system (parameters set to 1450V, 10ms, 3 pulses).
[0218] (2) After completing step (1), culture in complete culture medium for 12-18 hours, then replace with new complete culture medium for further culture. The total culture time after electroporation is 48 hours.
[0219] (3) After completing step (2), cells were digested and collected with trypsin, genomic DNA was extracted, and PCR amplification was performed using primers consisting of TTN-F55 and TTN-R560. Then, 1% agarose gel electrophoresis was performed.
[0220] See electrophoresis image Figure 4 The gene deletion mutation efficiency using commercial Cas9-A protein was 28.5%, that using NCN protein was 85.6%, and that using commercial Cas9-B protein was 16.6%.
[0221] The results showed that, compared with commercially available Cas9 protein, the NCN protein prepared using this invention significantly improved gene editing efficiency.
[0222] Example 2: Screening for optimal safe harbor sites in the pig genome for targeted insertion of exogenous genes
[0223] I. Constructing Donor vectors containing different safe harbor sites of the GFP gene
[0224] Plasmids PB-1G 2R 3-puro-ROSA26, PB-1G 2R 3-puro-AAVS1, PB-1G2R3-puro-H11, and PB-1G 2R 3-puro-COL1A1 were constructed. All four plasmids are circular plasmids.
[0225] The plasmid PB-1G 2R 3-puro-ROSA26 is shown in SEQ ID NO: 8. A schematic diagram of its structure can be found in [link to diagram]. Figure 5In SEQ ID NO: 8, nucleotides 9-339 form the 5' end of the ROSA26 safe harbor insertion site in the pig genome (left arm of SH1), and nucleotides 9184-10195 form the 3' end of the ROSA26 safe harbor insertion site in the pig genome (right arm of SH1). In SEQ ID NO: 8, nucleotides 346-546, 3132-3531, 6506-6706, and 8975-9175 form four different insulator regions. In SEQ ID NO: 8, nucleotides 637-1209 form the EF-1α poly(A) signal, nucleotides 1216-1935 encode the EGFP protein, nucleotides 1954-3131 form the EF-1α promoter, nucleotides 3543-4042 form the PGK promoter, nucleotides 4059-4769 encode the mCherry protein, nucleotides 4791-5015 form the bGH poly(A) signal, nucleotides 5054-6504 are the loxP-puro-loxP expression box region, nucleotides 6969-7233 form the β-globin poly(A) signal, and nucleotides 7259-8974 form the pCAG promoter.
[0226] The only difference between plasmid PB-1G 2R 3-puro-AAVS1 and plasmid PB-1G 2R 3-puro-ROSA26 is that the left arm of SH1 is replaced with the 5' end of the AAVS1 safe harbor insertion site in the pig genome (the left arm of SH2, as shown in SEQ ID NO: 9) and the right arm of SH1 is replaced with the 3' end of the AAVS1 safe harbor insertion site in the pig genome (the right arm of SH2, as shown in SEQ ID NO: 10).
[0227] The only difference between plasmid PB-1G 2R 3-puro-H11 and plasmid PB-1G 2R 3-puro-ROSA26 is that the left arm of SH1 is replaced with the 5' end of the pig genome region of the H11 safe harbor insertion site (the left arm of SH3, as shown in SEQ ID NO: 11) and the right arm of SH1 is replaced with the 3' end of the pig genome region of the H11 safe harbor insertion site (the right arm of SH3, as shown in SEQ ID NO: 12).
[0228] The only difference between plasmid PB-1G 2R 3-puro-COL1A1 and plasmid PB-1G 2R 3-puro-ROSA26 is that the left arm of SH1 is replaced with the 5' end of the pig genome region of the COL1A1 safe harbor insertion site (the left arm of SH4, as shown in SEQ ID NO: 13) and the right arm of SH1 is replaced with the 3' end of the pig genome region of the COL1A1 safe harbor insertion site (the right arm of SH4, as shown in SEQ ID NO: 14).
[0229] II. Efficient cleavage target screening for safe harbor sites in the porcine ROSA26, AAVS1, H11, and COL1A1 genomes.
[0230] Through preliminary screening, the highly efficient cleavage target of the ROSA26 safe harbor site is sgRNA. ROSA26-g3 (Cutting efficiency 38%), the highly efficient cleavage target of the AAVS1 safe harbor site is sgRNA. AAVS1-g4 (Cutting efficiency 30%), the highly efficient cleavage target of the H11 safe harbor site is sgRNA. H11-g1 (Cutting efficiency 60%), the COL1A1 safe harbor site has a highly efficient cleavage target of sgRNA. COL1A1-g3 (Cutting efficiency 56%).
[0231] The target sequence is as follows:
[0232] sgRNA ROSA26-g3 Target: 5'-GAAGGAGCAAACTGACATGG-3';
[0233] sgRNA AAVS1-g4 Target: 5'-TGCAGTGGGTCTTTGGGGAC-3';
[0234] sgRNA H11-g1 Target: 5'-TTCCAGGAACATAAGAAAGT-3';
[0235] sgRNA COL1A1-g3 Target: 5'-GCAGTCTCAGCAACCACTGA-3'.
[0236] III. Preparation of safe harbor site gRNA recombinant vector
[0237] The pKG-U6gRNA plasmid was digested with restriction endonuclease BbsI, and the vector backbone (a large linear fragment of about 3kb) was recovered.
[0238] ROSA26-g3-S and ROSA26-g3-A were synthesized separately, then mixed and annealed to obtain a double-stranded DNA molecule with sticky ends. The sticky-ended double-stranded DNA molecule was ligated to a vector backbone to obtain plasmid pKG-U6gRNA (ROSA26-g3). Plasmid pKG-U6gRNA (ROSA26-g3) expresses the sgRNA shown in SEQ ID NO: 15. ROSA26-g3 .
[0239] AAVS1-g4-S and AAVS1-g4-A were synthesized separately, then mixed and annealed to obtain a double-stranded DNA molecule with sticky ends. The sticky-ended double-stranded DNA molecule was ligated to a vector backbone to obtain plasmid pKG-U6gRNA(AAVS1-g4). Plasmid pKG-U6gRNA(AAVS1-g4) expresses the sgRNA shown in SEQ ID NO: 16. AAVS1-g4 .
[0240] H11-g1-S and H11-g1-A were synthesized separately, then mixed and annealed to obtain a double-stranded DNA molecule with sticky ends. The sticky-ended double-stranded DNA molecule was ligated to a vector backbone to obtain plasmid pKG-U6gRNA(H11-g1). Plasmid pKG-U6gRNA(H11-g1) expresses the sgRNA shown in SEQ ID NO: 17. H11-g1 .
[0241] COL1A1-g3-S and COL1A1-g3-A were synthesized separately, then mixed and annealed to obtain a double-stranded DNA molecule with sticky ends. The sticky-ended double-stranded DNA molecule was ligated to a vector backbone to obtain plasmid pKG-U6gRNA(COL1A1-g3). Plasmid pKG-U6gRNA(COL1A1-g3) expresses the sgRNA shown in SEQ ID NO: 18. COL1A1-g3 .
[0242] ROSA26-g3-S, ROSA26-g3-A, AAVS1-g4-S, AAVS1-g4-A, H11-g1-S, H11-g1-A, COL1A1-g3-S, and COL1A1-g3-A are all single-stranded DNA molecules.
[0243] ROSA26-g3-S: caccGAAGGAGCAAACTGACATGG;
[0244] ROSA26-g3-A:aaacCCATGTCAGTTTGCTCCTTC.
[0245] AAVS1-g4-S:caccgTGCAGTGGGTCTTTGGGGAC;
[0246] AAVS1-g4-A:aaacGTCCCCAAAAGACCCACTGCAc。
[0247] H11-g1-S:caccgTTCCAGGAACATAAGAAAGT;
[0248] H11-g1-A:aaacACTTTCTTATGTTCCTGGAAc。
[0249] COL1A1-g3-S:caccGCAGTCTCAGCAACCACTGA;
[0250] COL1A1-g3-A:aaacTCAGTGGTTGCTGAGACTGC。
[0251] sgRNA ROSA26-g3 (SEQ ID NO:15):
[0252] GAAGGAGCAAACUGACAUGGGguuuuaaggcuaaaaacugaaaaaaaaaaaaaaggcuaaaacuugaaaaaguggcaccgagucggugcuuuuuuuuuuuuu.
[0253] sgRNA AAVS1-g4 (SEQ ID NO:16):
[0254] UGCAGUGGGUCUUUGGGGAACguuuuaaggcuagaaaaaaaaaaaaaaaaaggcuaaaacuugaaaaaaguggcaccgagucggugcuuuuuuuuuuuuuuuuuuuuuu。
[0255] sgRNA H11-g1 (SEQ ID NO:17):
[0256] UUCCAGGAACAUAAGAAAGUGUUUUAGGCUAAUAGCAAUAGCAAUAAAAAAAAAAAAAGGUACGUACGUAAUCACUUAAAAGGCACCGAGUCGUGGUGCUUUUuuuuuuuuuuuuuuuuuuu
[0257] sgRNA COL1A1-g3 (SEQ ID NO:18):
[0258] GCAGUCUCAGCAACCACUGAguuuuagagcuagaaauagcaaguuaaaauaaggcuaguccguuaucaacuugaaaaaguggcaccgagucggugcuuuu.
[0259] IV. Electroporation of primary porcine fibroblasts with a mixture of fluorescent Donor vectors (i.e., vectors containing exogenous GFP at different safe harbor insertion sites), sgRNA vectors, and Cas9 vectors (i.e., plasmid pKG-GE3) containing homologous arms flanking different safe harbor insertion sites and detection of GFP fluorescence intensity in cells.
[0260] 1. Co-transfection
[0261] Group 1 (ROSA26 group): Plasmid PB-1G 2R 3-puro-ROSA26, plasmid pKG-U6gRNA (ROSA26-g3), and plasmid pKG-GE3 were co-transfected into porcine primary fibroblasts. The ratio was approximately 200,000 porcine primary fibroblasts: 1.26 μg plasmid PB-1G 2R 3-puro-ROSA26 : 0.82 μg plasmid pKG-U6gRNA (ROSA26-g3) : 0.92 μg plasmid pKG-GE3; that is, the molar ratio of the three plasmids was 1:3:1.
[0262] Group 2 (AAVS1 group): Plasmid PB-1G 2R 3-puro-AAVS1, plasmid pKG-U6gRNA (AAVS1-g4), and plasmid pKG-GE3 were co-transfected into porcine primary fibroblasts. The ratio was approximately 200,000 porcine primary fibroblasts: 1.26 μg plasmid PB-1G 2R 3-puro-AAVS1 : 0.82 μg plasmid pKG-U6gRNA (AAVS1-g4) : 0.92 μg plasmid pKG-GE3; that is, the molar ratio of the three plasmids was 1:3:1.
[0263] Group 3 (H11 group): Plasmids PB-1G2R3-puro-H11, pKG-U6gRNA (H11-g1), and pKG-GE3 were co-transfected into porcine primary fibroblasts. The ratio was approximately 200,000 porcine primary fibroblasts: 1.26 μg plasmid PB-1G2R3-puro-H11 : 0.82 μg plasmid pKG-U6gRNA (H11-g1) : 0.92 μg plasmid pKG-GE3; that is, the molar ratio of the three plasmids was 1:3:1.
[0264] Group 4 (COL1A1 group): Plasmid PB-1G 2R 3-puro-COL1A1, plasmid pKG-U6gRNA (COL1A1-g3), and plasmid pKG-GE3 were co-transfected into porcine primary fibroblasts. The ratio was approximately 200,000 porcine primary fibroblasts: 1.26 μg plasmid PB-1G 2R 3-puro-COL1A1 : 0.82 μg plasmid pKG-U6gRNA (COL1A1-g3) : 0.92 μg plasmid pKG-GE3; that is, the molar ratio of the three plasmids was 1:3:1.
[0265] Group 5: Primary porcine fibroblasts were electroporated using the same electroporation parameters without the addition of any plasmids.
[0266] Co-transfection was performed using electroporation. The transfection was carried out using a mammalian nuclear transfection kit (Neon kit, Thermofisher) and a Neon™ transfection system (parameters set to 1450V, 10ms, 3pulse).
[0267] 2. After completing step 1, incubate in complete culture medium for 12-24 hours, then replace with fresh complete culture medium. The total incubation time is 48 hours.
[0268] 3. After completing step 2, replace the culture medium with complete culture medium containing 1.5 μg / mL puromycin and culture for 3 weeks (replace with new complete culture medium containing 1.5 μg / mL puromycin every 2 days). Continuously observe and photograph the GFP green fluorescence. The intensity of GFP fluorescence expression is used to determine the efficiency of exogenous gene expression at the safe harbor site.
[0269] One week after puromycin screening, the fluorescence intensity of the ROSA26 and COL1A1 safe harbor sites was significantly stronger than that of the AAVS1 and H11 sites. Two weeks after puromycin screening, the fluorescence intensity from strongest to weakest was: COL1A1 > ROSA26 > H11 > AAVS1. The fluorescence intensity in the H11 group was not very uniform, while the ROSA26 group showed relatively uniform and high fluorescence intensity overall. The AAVS1 group had the weakest cell fluorescence expression, and the COL1A1 group had the most fluorescent cells and the strongest fluorescence. After three weeks of continued puromycin screening, the fluorescence intensity from strongest to weakest was: COL1A1 > ROSA26 > H11 > AAVS1. (See photos below.) Figure 6 .
[0270] V. Detection of GFP gene transcription level
[0271] To compare the differences in mRNA transcription levels after GFP gene integration into four different safe harbor sites, and to determine whether it participates in GFP expression regulation and its impact on expression levels, a pair of primers was designed at the exons of the GFP gene. Cells selected three weeks after puromycin screening in step four were used to extract total RNA, which was reverse transcribed into cDNA. The transcription levels of primary cells after GFP gene integration into the four different safe harbor sites were detected. The quantitative results obtained from the fifth group of cells (the plasmid-free control electroporation group) were used as a control. The GAPDH gene was used as an internal reference gene according to a 2... -ΔCt The method is used for calculation.
[0272] Primers used for detecting the GFP gene: F: AGATCCGCCACAACATCGAG; R: GTCCATGCCGAGAGTGATCC.
[0273] Primers used for detecting the GAPDH gene: F: GGTCGGAGTGAACGGATTTG; R: CCATTTGATGTTGGCGGGAT.
[0274] Data were analyzed using SPSS statistical software, and expressed as mean ± standard deviation. A two-way ANOVA was used for statistical analysis. 2 -ΔCt The results showed that after three weeks of puromycin screening, the GFP expression levels in the AAVS1 and H11 groups were low, while the GFP expression levels in the ROSA26 and COL1A1 groups were high. Furthermore, the GFP transcription levels in the COL1A1 and ROSA26 groups were significantly different from those in the AAVS1 and H11 groups (P<0.01). -ΔCt The values are shown in Table 1, and the results of the significance analysis are shown in Table 2. Figure 7 .
[0275] Table 1 2 -ΔCt Value information
[0276]
[0277] In summary, based on the fluorescence signal intensity after three weeks of cell culture and the results of real-time quantitative PCR of the GFP gene, the following conclusions can be drawn: among the four genomic safe harbor sites ROSA26, AAVS1, H11, and COL1A1, the COL1A1 site showed the best expression effect after the insertion of the exogenous gene.
[0278] VI. FACS detection of GFP gene protein expression level
[0279] To compare GFP protein expression after integration of the GFP gene into four different safe harbor sites, electroporated cells selected three weeks prior in step four (using puromycin) were digested with trypsin, centrifuged at 400g for 4 min, and the supernatant was discarded. Cells were resuspended in 1 mL of complete culture medium, and the cell suspension was transferred to flow cytometry tubes. GFP signal was detected in the FITC channel of a BD FACSMelody flow cytometer, and 5 × 10⁶ cells were collected. 4 Analyzing individual cells, the results are shown below. Figure 8 .
[0280] The results showed that the GFP fluorescence signal intensity was COL1A1>ROSA26>H11>AAVS1.
[0281] Therefore, based on the above results, the COL1A1 site is the most efficient porcine primary cell safe harbor site for expressing exogenous genes among the four safe harbor sites: ROSA26, AAVS1, H11, and COL1A1.
[0282] Example 3: Preparation of single-cell clones with hSRD5A2 gene expression cassettes inserted at the COL1A1 safe harbor site in the genome.
[0283] Human SRD5A2 gene (hSRD5A2 gene) information: Encodes type II 5α-reductase; Homo sapiens; Located on human chromosome 2; GeneID 6716. The amino acid sequence of human type II 5α-reductase is shown in SEQ ID NO: 19.
[0284] Through repeated experiments and research, the inventors have shown that, compared with the electroporation method using a combination of pKG-GE3 plasmid and gRNA plasmid in Example 2, the use of a combination of NCN protein and gRNA, i.e., RNP electroporation, can improve cell viability. Therefore, this example uses RNP electroporation to prepare single-cell clones with the hSRD5A2 gene expression cassette inserted at the COL1A1 safe harbor site in the genome.
[0285] I. Construction of the KAP6.1-hSRD5A2 Donor vector
[0286] The KAP6.1-hSRD5A2 Donor vector is the plasmid KAP6.1-hSRD5A2.
[0287] Plasmid KAP6.1-hSRD5A2, as shown in SEQ ID NO: 20, is a circular plasmid. A schematic diagram of its structure is shown below. Figure 9In SEQ ID NO: 20, nucleotides 9-880 represent the 5' end of the porcine genome region (left arm of SH4) of the COL1A1 safe harbor insertion site; nucleotides 887-1087 represent the insulator (named Insulator 1); nucleotides 1088-2139 represent the KAP6.1 promoter; nucleotides 2140-2904 represent the hSRD5A2 gene; nucleotides 2905-3477 represent EF1αPoly(A); nucleotides 3507-3540 represent the LoxP sequence; nucleotides 3605-4104 represent the pGK promoter; and nucleotides 4182-4985 encode the Neomycin resistance protein (abbreviated as Neo). R The protein, with nucleotides 5023-5247 being bGH Poly(A), nucleotides 5304-5337 being the LoxP sequence, nucleotides 5358-5559 being the insulator (named Insulator 5), and nucleotides 5560-6286 being the 3' end of the COL1A1 safe harbor insertion site in the porcine genome (right arm of SH4). Neomycin resistance protein is the neomycin resistance protein. Neomycin (Geneticin), also known as G418 or Geneticin.
[0288] II. Preparation of gRNA
[0289] Two highly efficient cleavage target sgRNAs from the COL1A1 safe harbor site obtained in the previous screening were selected. COL1A1-g1 (50% cleavage efficiency) and sgRNA COL1A1-g3 (Cutting efficiency 56%).
[0290] The information for the two target points is as follows:
[0291] sgRNA COL1A1-g1 Target: 5'-CTACCAAGAGAGTGACCAGC-3';
[0292] sgRNA COL1A1-g3 Target: 5'-GCAGTCTCAGCAACCACTGA-3'.
[0293] 1. Preparation of COL1A1-T7-gRNA1 and COL1A1-T7-gRNA3 transcription templates
[0294] The transcription template for COL1A1-T7-gRNA1 is a double-stranded DNA molecule, as shown in SEQ ID NO: 21.
[0295] The transcription template for COL1A1-T7-gRNA3 is a double-stranded DNA molecule, as shown in SEQ ID NO: 22.
[0296] 2. Obtain gRNA through in vitro transcription
[0297] The COL1A1-T7-gRNA1 transcription template was used for in vitro transcription using the Transcript Aid T7 High Yield Transcription Kit (Fermentas, K0441), followed by MEGA clearing. TM The COL1A1-gRNA1 was recovered and purified using a Transcription Clean-Up Kit (Thermo, AM1908). COL1A1-gRNA1 is a single-stranded RNA, as shown in SEQ ID NO: 23.
[0298] The COL1A1-T7-gRNA3 transcription template was used for in vitro transcription using the Transcript Aid T7 High Yield Transcription Kit (Fermentas, K0441), followed by MEGA clearing. TM The COL1A1-gRNA3 was recovered and purified using a Transcription Clean-Up Kit (Thermo, AM1908). COL1A1-gRNA3 is a single-stranded RNA, as shown in SEQ ID NO: 24.
[0299] COL1A1-gRNA1 (SEQ ID NO: 23):
[0300] GGCUACCAAGAGAGUGACCAGCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCCGGUGCUUUU.
[0301] COL1A1-gRNA3 (SEQ ID NO: 24):
[0302] GGGCAGUCUCAGCAACCACUGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCCGGUGCUUUU.
[0303] III. Co-transfection
[0304] Porcine primary fibroblasts were co-transfected with COL1A1-gRNA1, COL1A1-gRNA3, NCN protein, and plasmid KAP6.1-hSRD5A2. The ratio was approximately 100,000 porcine primary fibroblasts: 1 μg COL1A1-gRNA1 : 1 μg COL1A1-gRNA3 : 4 μg NCN protein : 3 μg plasmid KAP6.1-hSRD5A2. Co-transfection was performed using electroporation with a mammalian nuclear transfection kit (Neon kit, Thermofisher) and a Neon™ transfection system (parameters set to 1450V, 10ms, 3 pulses).
[0305] The functions of COL1A1-gRNA1, COL1A1-gRNA3, and NCN proteins are to create DNA double-strand breaks in porcine genomic DNA to increase the homologous recombination rate. Plasmid KAP6.1-hSRD5A2 undergoes homologous recombination with porcine genomic DNA, inserting a foreign target gene fragment (the foreign target gene fragment is the DNA molecule represented by nucleotides 881-5559 in SEQ ID NO: 20) between the left and right arms of SH4 in the porcine genomic DNA.
[0306] IV. Neomycin-based pressure screening
[0307] 1. Screening for positive cells with inserted exogenous target gene fragments.
[0308] (1) After completing step three, culture the electroporated cells in complete culture medium for 16-18 hours, and then replace with new complete culture medium for further culture. The total culture time is 48 hours.
[0309] (2) After completing step (1), replace the culture medium with a complete culture medium containing 1.5 mg / mL G418 for screening culture (replace with a new complete culture medium containing 1.5 mg / mL G418 every day) for 3 weeks.
[0310] After one week of screening and culture, a large number of cells died.
[0311] After two weeks of screening and culture, only a few cells died, while some positive clones began to divide and proliferate, and the number of cells continued to increase.
[0312] The purpose of the third week of screening culture is to ensure complete degradation of intracellular plasmids in order to eliminate false-positive cell clones.
[0313] (3) After completing step (2), collect the cells and culture them for 2 generations in complete culture medium without G418 (1 generation every 2 days) to allow the cells to recover to a good state for the next step of single-cell sorting.
[0314] 2. Single-cell sorting and scale-up culture
[0315] (1) After completing step 1, collect cells, digest them with trypsin, neutralize them with complete culture medium, centrifuge at 500g for 5min, discard the supernatant, resuspend the pellet in 1mL of complete culture medium and dilute appropriately, pick single cells with a pipette and transfer them to a 96-well plate (pre-add 100μl of complete culture medium to each well) (one group of cells per 96-well plate, one cell per well), and culture them. After 2 days of culture, replace the culture medium with complete culture medium containing 1.5mg / mL G418. Then replace the culture medium with a new one containing 1.5mg / mL G418 every 2-3 days. During this period, observe the cell growth in each well with a microscope and exclude wells without cells or with non-single-cell clones.
[0316] (2) When the cells in the wells of the 96-well plate from step (1) have filled the bottom of the wells (about 2 weeks), digest them with trypsin and collect the cells. Two-thirds of the cells are seeded into 6-well plates containing complete culture medium, and the remaining one-third of the cells are collected in 1.5 mL centrifuge tubes.
[0317] (3) When the cells in the wells of the 6-well plate in step (2) reach 50% fullness, digest them with 0.25% (Gibco) trypsin and collect the cells. Use cell cryopreservation solution (90% complete culture medium + 10% DMSO, volume ratio) to cryopreserve the cells.
[0318] V. Genome-level identification of exogenous target gene fragments inserted at the COL1A1 safe harbor site
[0319] To detect whether the exogenous target gene fragment was successfully inserted at the COL1A1 safe harbor site in the cell genome, the genomic DNA of the cells was extracted from the centrifuge tube in step 2(2) of step four. PCR amplification was performed using specific primer pairs (the specific primer pairs were: sh4-Lr-JDF1414 and sh4-Lr-JDR5965, sh4-Rr-JDF282 and sh4-Rr-JDR4723, and sh4-wt-JDF1085 and sh4-wt-JDR1560), followed by electrophoresis. Primary porcine fibroblasts were used as wild-type controls (WT).
[0320] Primer pairs consisting of sh4-Lr-JDF1414 and sh4-Lr-JDR5965 were used to identify whether the exogenous target gene fragment at the 5' end of the porcine COL1A1 safe harbor insertion site had successfully recombinated (the target sequence is 4552 bp, and an amplification product of approximately 4552 bp indicates successful recombination); primer pairs consisting of sh4-Rr-JDF282 and sh4-Rr-JDR4723 were used to identify whether the exogenous target gene fragment at the 3' end of the porcine COL1A1 safe harbor insertion site had successfully recombinated (the target sequence is 4442 bp, and an amplification product of approximately 4442 bp indicates successful recombination). (Recombination successful); the primer pair consisting of sh4-wt-JDF1085 and sh4-wt-JDR1560 was used to identify whether the exogenous target gene fragment inserted at the safe harbor site of porcine COL1A1 was homozygous or heterozygous (the genomic DNA of the wild-type control could be amplified into a 476bp fragment, but the recombinant cells could not be amplified because the inserted exogenous target gene fragment was too large; therefore, if no amplification product is shown, it means that the cell is homozygous for the inserted exogenous target gene fragment; if a 476bp amplification product is shown, it means that the cell is heterozygous or wild-type).
[0321] sh4-Lr-JDF1414: CCTGCTGTAAGTGCCGTAGT;
[0322] sh4-Lr-JDR5965: CTAGGGGCACAGCACGTC.
[0323] sh4-Rr-JDF282:AAGTTATTAGGTCTGAAGAGGAGTTT;
[0324] sh4-Rr-JDR4723: CCCATCATTCCGTCCCAGAG.
[0325] sh4-wt-JDF1085: TGCTGAGTTCTGGCTTCCTG;
[0326] sh4-wt-JDR1560:TCTACCAGAGAGTGACCAGCAG.
[0327] According to the identification results, single-cell clones numbered 1-3, 5-15, and 17-55 were all clones that successfully inserted the exogenous target gene fragment at the COL1A1 safe harbor site. Among them, single-cell clones numbered 29 and 44 were homozygous for targeted insertion, while the other single-cell clones were heterozygous for targeted insertion. See Table 2.
[0328] Table 2 Genotypes of single-cell clones
[0329]
[0330]
[0331]
[0332] The recombinant cell numbered 1 in Table 2 (heterozygous site-directed insertion type) is named recombinant cell #1. Whole-genome sequencing revealed that, compared to primary porcine fibroblasts from the same source, the only difference in the genomic DNA of recombinant cell #1 was the insertion of a foreign target gene fragment (the foreign target gene fragment being the DNA molecule represented by nucleotides 881-5559 of SEQ ID NO: 20) between the left and right arms of SH4, and it was heterozygous (i.e., in a pair of homologous chromosomes, insertion occurred on one chromosome and not on the other).
[0333] Recombinant cells numbered 29 in Table 2 (homozygous site-directed insertion type) are designated as recombinant cell #29. Recombinant cells numbered 44 in Table 2 (homozygous site-directed insertion type) are designated as recombinant cell #44. Whole-genome sequencing revealed that, compared to primary porcine fibroblasts from the same source, the only difference in the genomic DNA of recombinant cells #29 (or #44) was the insertion of a foreign target gene fragment (the foreign target gene fragment being the DNA molecule represented by nucleotides 881-5559 of SEQ ID NO: 20) between the left and right arms of SH4, and this insertion was homozygous (i.e., identical insertion occurred on both homologous chromosomes). Sequencing results of the key adaptor sequences are shown below. Figures 10 to 14 .
[0334] In summary, this invention has successfully obtained recombinant cells in which the human SRD5A2 gene is specifically integrated into the pig genome and expressed in hair follicle tissue. These recombinant cells can be used as nuclear transfer cell donors to clone and produce pigs with hair loss.
[0335] The present invention has been described in detail above. For those skilled in the art, the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments have been given, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein. Some of the essential features can be applied within the scope of the following appended claims. sequence list <110> Nanjing Qizhen Gene Engineering Co., Ltd. <120> The kit and its application in constructing recombinant porcine cells that specifically express the human type II 5α-reductase gene in hair follicle tissue. <130> GNCYX221773 <160> 25 <170> SIPOSequenceListing 1.0 <210> 1 <211> 9974 <212> DNA <213> Artificial Sequence <400> 1 tggcgaatgg gacgcgccct gtagcggcgc attaagcgcg gcgggtgtgg tggttacgcg 60 cagcgtgacc gctacacttg ccagcgccct agcgcccgct cctttcgctt tcttcccttc 120 ctttctcgcc acgttcgccg gctttccccg tcaagctcta aatcgggggc tccctttagg 180 gttccgattt agtgctttac ggcacctcga ccccaaaaaa cttgattagg gtgatggttc 240 acgtagtggg ccatcgccct gatagacggt ttttcgccct ttgacgttgg agtccacgtt 300 ctttaatagt ggactcttgt tccaaactgg aacaacactc aaccctatct cggtctattc 360 ttttgattta taagggattt tgccgatttc ggcctattgg ttaaaaaatg agctgattta 420 acaaaaattt aacgcgaatt ttaacaaaat attaacgttt acaatttcag gtggcacttt 480 tcggggaaat gtgcgcggaa cccctatttg tttatttttc taaatacatt caaatatgta 540 tccgctcatg agacaataac cctgataaat gcttcaataa tattgaaaaa ggagagtat 600 gagtattcaa cattccgtg tcgccttat tccttttt gcggcatttt gccttcctgt 660 ttttgctcac ccagaaacgc tggtgaaagt aaagatgct gaagatcagt tggtgcacg 720 agtgggttac atcgaactgg atctcacag cggtaagatc cttgagagtt ttcgccccga 780 agaacgtttt ccaatgatga gcacttta agttctgcta tgtggcgcgg tattacccg 840 tattgacgcc gggcaagagc aactcggtcg ccgcatacac tattctcaga atgacttggt 900 tgagtactca ccagtcacag aaagcatct tacggatggc atgacagtaa gagaattatg 960 cagtgctgcc ataccatga gtgatacac tgcggccac ttactctga caacgatcgg 1020 aggaccgaag gagctaccg cttttgca siacatgggg gatcatgtaa ctcgccttga 1080 tcgttgggaa ccggagctga atgaagccat accaacgac gagcgtgaca ccacgatgcc 1140 tgcagcaatg gcacaacgt tgcgcaacct atttactggc gaactactta ctctagcttc 1200 ccggcaacaa ttatagact ggatggaggc ggataagtt gcaggaccac ttctgcgctc 1260 ggcccttccg gctggctggt ttattgctga taaatctgga gccggtgagc gtgggtctcg cggtatcatt gcagcactgg ggccagatgg tagccctcc cgtatcgtag ttatctacac 1380 gacggggagt caggcaacta tggatgaacg aatagacag atcgctgaga taggtgcctc actgattaag cattggtaac tgtcagacca agtttactca fathercttt agttgattt aaaacttcat ttttaattta aaaggatcta ggtgaagatc ctttttgata atctcatgac caaatccct taacgtgagt tttcgttcca ctgagcgtca gaccccgtag aaaagatcaa aggatcttct tgagatcctt tttttctgcg cgtaatctgc tgcttgcaaa caaaaaaacc 1680 accgctacca gcggtggttt gtttgccgga tcaagagcta ccaactcttt ttccgaaggt 1740. aactggcttc agcagagcgc agataccaaa tactgtcctt ctagtgtagc cgtagttagg ccaccacttc aagaactctg tagcaccgcc tacatacctc gctctgctaa tcctgttacc agtggctgct gccagtggcg ataagtcgtg tcttaccggg ttggactcaa gacgatagtt accggataag gcgcagcggt cgggctgaac ggggggttcg tgcacacagc ccagcttgga gcgaacgacc tacaccgaac tgagatacct acagcgtgag ctatgagaaa gcgccacgct 2040 tcccgaaggg agaaaggcgg acaggtatcc ggtaagcggc agggtcggaa caggagagcg 2100 cacgagggag cttccagggg gaaacgcctg gtatctttat agtcctgtcg ggtttcgcca 2160 cctctgactt gagcgtcgat ttttgtgatg ctcgtcaggg gggcggagcc tatggaaaaa 2220 cgccagcaac gcggccttt tacggttcct ggccttttgc tggccttttg ctcacatgtt 2280 ctttcctgcg ttatcccctg attctgtgga taaccgtatt accgcctttg agtgagctga 2340 taccgctcgc cgcagccgaa cgaccgagcg cagcgagtca gtgagcgagg aagcggaaga 2400 gcgcctgatg cggtattttc tccttacgca tctgtgcggt atttcacacc catatatgg 2460 tgcactctca gtacaatctg ctctgatgcc ccatagttaa gccagtatac actccgctat 2520 cgctacgtga ctgggtcatg gctgcgcccc gacacccgcc aacacccgct gacgcgccct 2580 gacgggcttg tctgctcccg gcatccgctt acagacaagc tgtgaccgtc tccgggagct 2640 gcatgtgtca gaggttttca ccgtcatcac cgaaacgcgc gaggcagctg cggtaaagct 2700 catcagcgtg gtcgtgaagc gattcacaga tgtctgcctg ttcatccgcg tccagctcgt 2760 tgagttctc cagaagcgtt aatgtctggc ttctgataaa gcggggccatg ttaagggcgg 2820 ttttttcctg ttggtcact gatgcctccg tgtaaggggg atttctgttc atgggggtaa 2880 tgataccgat gaacgagag aggatgctca cgatacggggt tactgat gaacatgccc 2940 ggttactgga acgttgtgag ggtaacaac tggcggtag gatgcggcgg gaccagagaa 3000 aaatcactca gggtcaatgc cagcgcttcg ttatacaga tgtaggtgtt ccacagggta 3060 gccagcagca tcctgcgatg cagatccgga acataatggt gcagggcgct gacttccgcg 3120 tttccagact ttacgaaca cggaaccga agaccattca tgttgttgct caggtcgcag 3180 acgttttgca gcagcagtcg cttcacgttc gctcgcgtat cggtgattca ttctgctaac 3240 cagtaggca accccgccag cctagccggg tcctcaacga caggagcacg atcatgcgca 3300 cccgtggggc cgccatgccg gcgataatgg cctgctctc gccgaaacgt ttggtggcgg 3360 gaccagtgac gaaggcttga gcgaggggcgt gcaagattcc gataccgca agcgacaggc 3420 cgatcatcgt cgcgctccag cgaaagcggt cctcgccgaa aatgacccag agcgctgccg 3480 gcacctgtcc tacgagttgc atgataaaga agacagtcat aagtgcggcg acgatagtca 3540 tgccccgcgc ccaccggaag gagctgactg ggttgaaggc tctcaagggc atcggtcgag 3600 atcccggtgc ctaatgagtg agctaactta cattaattgc gttgcgctca ctgcccgctt 3660 tccagtcggg aaacctgtcg tgccagctgc attaatgaat cggccaacgc gcggggagag 3720 gcggtttgcg tattgggcgc cagggtggtt tttcttttca ccagtgagac gggcaacagc 3780 tgattgccct tcaccgcctg gccctgagag agttgcagca agcggtccac gctggtttgc 3840 cccagcaggc gaaaatcctg tttgatggtg gttaacggcg ggatataaca tgagctgtct 3900 tcggtatcgt cgtatcccac taccgagatg tccgcaccaa cgcgcagccc ggactcggta 3960 atggcgcgca ttgcgcccag cgccatctga tcgttggcaa ccagcatcgc agtgggaacg 4020 atgccctcat tcagcatttg catggtttgt tgaaaaccgg acatggcact ccagtcgcct 4080 tcccgttccg ctatcggctg aatttgattg cgagtgagat atttatgcca gccagccaga 4140 cgcagacgcg ccgagacaga acttaatggg cccgctaaca gcgcgatttg ctggtgaccc 4200 aatgcgacca gatgctccac gcccagtcgc gtaccgtctt catgggagaa aataatactg 4260 ttgatgggtg tctggtcaga gacatcaaga aataacgccg gaacattagt gcaggcagct 4320 tccacagcaa tggcatcctg gtcatccagc ggatagttaa tgatcagccc actgacgcgt 4380 tgcgcgagaa gattgtgcac cgccgcttta caggcttcga cgccgcttcg ttctaccatc 4440 gacaccacca cgctggcacc cagttgatcg gcgcgagatt taatcgccgc gacaatttgc 4500 gacggcgcgt gcagggccag actggaggtg gcaacgccaa tcagcaacga ctgtttgccc 4560 gccagttgtt gtgccacgcg gttgggaatg taattcagct ccgccatcgc cgcttccact 4620 ttttcccgcg ttttcgcaga aacgtggctg gcctggttca ccacgcggga aacggtctga 4680 taagagacac cggcatactc tgcgacatcg tataacgtta ctggtttcac attcaccacc 4740 ctgaattgac tctcttccgg gcgctatcat gccataccgc gaaaggtttt gcgccattcg 4800 atggtgtccg ggatctcgac gctctccctt atgcgactcc tgcattagga agcagcccag 4860 tagtagttg aggccgttga gcaccgccgc cgcaaggaat ggtgcatgca aggagatggc 4920 gcccaacagt cccccggcca cggggcctgc caccataccc acgccgaaac aagcgctcat 4980. gagcccgaag tggcgagccc gatcttcccc atcggtgatg tcggcgatat aggcgccagc 5040. aaccgcacct gtggcgccgg tgatgccggc cacgatgcgt ccggcgtaga ggatcgagat 5100 cgatctcgat cccgcgaat fathercgact cactataggg gaattgtgag cggataacaa ttcccctcta gaatattt tgtttaactt taagaaggag atatacatat gaaacaaagc actattgcac tggcactctt accgttactg tttacccctg tgacaaaagc catgagcgat aaaattattc acctgactga cgacagtttt gacacggatg tactcaaagc ggacggggcg atcctcgtcg atttctgggc agagtggtgc ggtccgtgca aaatgatcgc cccgattctg 5400. gatgaaatcg ctgacgaata tcagggcaaa ctgaccgttg caaaactgaa catcgatcaa aaccctggca ctgcgccgaa atatggcatc cgtggtatcc cgactctgct gctgttcaaa aacggtgaag tggcggcaac caaagtgggt gcactgtcta aaggtcagtt gaaagagttc ctcgacgcta acctggccgg ttctggttct ggccatatgc accatcatca tcatcatgac 5640 gatgacgata agatgcccaa aaagaaacga aaggtgggta tccacggagt cccagcagcc 5700 gacaaaaaat atagcatcgg cctggacatc ggtaccaaca gcgttggctg ggcagtgatc 5760 actgatgaat acaaagttcc atccaaaaaa tttaaagtac tgggcaacac cgaccgtcac 5820 tctatcaaaa aaaacctgat tggtgctctg ctgtttgaca gcggcgaaac tgctgaggct 5880 acccgtctga aacgtacggc tcgccgtcgc tacactcgtc gtaaaaaccg catctgttat 5940 ctgcaggaaa ttttctctaa cgaaatggca aaagttgatg atagcttctt tcatcgtctg 6000 gaagagagct tcctggtgga agaagataaa aaacacgaac gtcacccgat tttcggtaac 6060 attgtggatg aggttgccta ccacgagaaa tatccgacca tctaccatct gcgtaaaaaa 6120 ctggttgata gcactgacaa agcggatctg cgtctgatct acctggctct ggcacacatg 6180 atcaaattcc gtggtcactt cctgatcgaa ggtgatctga accctgataa ctccgacgtg 6240 gacaaactgt tcattcagct ggttcagacc tataaccagc tgttcgaaga aaacccgatc 6300 aacgcgtccg gtgtagacgc taaggcaatt ctgtctgcgc gtctgtctaa gtctcgtcgt 6360 ctggaaaacc tgattgcgca actgccaggt gaaaagaaaa acggcctgtt cggcaatctg 6420 atcgccctgt ccctgggtct gactccgaac tttaaatcca actttgacct ggcggaagat 6480 gccaagctgc agctgagcaa agatacctat gacgatgacc tggataacct gctggcacag 6540 atcggtgatc agtatgccga tctgttcctg gccgcgaaaa acctgtctga tgcgattctg 6600 ctgtctgata tcctgcgcgt taacactgaa attactaaag cgccgctgag cgcatccatg 6660 attaaacgtt acgatgaaca ccaccaggat ctgaccctgc tgaaagcgct ggtgcgtcag 6720 cagctgccgg aaaaatacaa ggagatcttc ttcgaccaga gcaaaaacgg ttacgcgggc 6780 tacattgatg gtggtgcatc tcaggaggaa ttctacaaat tcattaaacc gatcctggaa 6840 aaaatggatg gtactgaaga gctgctggtt aaactgaatc gtgaagatct gctgcgcaaa 6900 cagcgtacct tcgataacgg ttccatcccg catcagattc atctgggcga actgcacgct 6960 atcctgcgcc gtcaggaaga cttttatccg ttcctgaaag acaaccgtga gaaaattgaa 7020 aaaatcctga ccttccgtat tccgtactat gtaggtccgc tggcgcgtgg taactcccgt 7080 ttcgcttgga tgacccgcaa aagcgaagaa accatcaccc cgtggaattt cgaagaagtc 7140 gttgacaaag gcgcgtccgc gcagtctttc atcgaacgca tgacgaactt cgacaaaaac 7200 ctgccgaacg agaaagtgct gccgaaacac tctctgctgt acgagtactt cactgtgtac 7260 aacgaactga ccaaagtgaa atacgtcacc gaaggtatgc gtaaaccggc attcctgtcc 7320 ggtgagcaaa aaaaagcaat cgtggatctg ctgttcaaaa ccaaccgtaa agtaaccgtg 7380 aaacagctga aggaagacta tttcaagaaa atcgaatgtt ttgattctgt tgaaatctcc 7440 ggcgtggaag atcgcttcaa tgcgtccctg ggtacgtatc acgacctgct gaaaattatc 7500 aaagacaaag attttctgga caacgaggaa aacgaagaca tcctggagga tattgtactg 7560 accctgaccc tgttcgaaga ccgtgagatg atcgaagaac gcctgaaaac ctacgcccac 7620 ctgttcgatg acaaggtaat gaagcagctg aaacgtcgtc gttataccgg ctggggtcgt 7680 ctgtcccgta aactgatcaa tggcatccgt gataaacagt ctggcaaaac catcctggac 7740 ttcctgaaat ccgacggttt cgcgaatcgt aacttcatgc aactgattca tgacgattct 7800 ctgactttca aagaagacat ccagaaagca caggtttccg gccagggtga ctctctgcac 7860 gagcacattg ccaatctggc tggttctccg gctattaaaa agggtattct gcagactgtg 7920 aaagtagttg atgagctggt caaagtaatg ggccgtcaca agccggaaaa cattgtgatc 7980 gaaatggcac gtgaaaacca gacgacccag aaaggtcaga aaaactctcg tgaacgcatg 8040 aaacgtatcg aagaaggcat caaagaactg ggctctcaga tcctgaagga acaccctgta 8100 gaaaataccc agctgcagaa cgaaaagctg tatctgtatt acctgcagaa cggccgcgat 8160 atgtatgtgg accaggaact ggatatcaac cgcctgtccg attacgatgt agatcacatc 8220 gtgccgcaaa gcttcctgaa agacgacagc attgacaaca aagtactgac ccgttctgat 8280 aagaaccgtg gcaaatccga taacgtcccg tctgaagaag ttgttaaaaa aatgaaaaac 8340 tattggcgtc agctgctgaa cgcgaaactg atcacccagc gtaagttcga caatctgact 8400 aaagctgagc gcggtggtct gtccgaactg gataaagcgg gttttatcaa acgccagctg 8460 gttgaaaccc gtcagatcac gaagcacgtt gcgcagattc tggactctcg tatgaacacc 8520 aaatacgacg aaaacgacaa actgatccgc gaggttaagg ttatcaccct gaaaagcaaa 8580 ctggtatccg attttcgtaa agactttcag ttctacaaag tgcgcgaaat taacaactat 8640 caccacgctc acgatgcata tctgaatgca gttgttggca cggcgctgat caaaaagtat 8700 ccgaaactgg aatctgaatt cgtatacggc gattacaaag tgtatgacgt tcgtaagatg 8760 atcgcaaaat ccgagcagga aattggtaag gcgacggcga aatacttctt ttattccaat 8820 attatgaact ttttcaaaac cgaaatcacc ctggcgaatg gtgaaattcg taaacgcccg 8880 ctgatcgaaa ccaacggtga aactggtgaa atcgtttggg acaaaggccg cgacttcgcg 8940 accgtgcgta aagttctgtc tatgccgcaa gtgaacatcg tcaagaagac cgaagtacaa 9000 accggcggtt ttagcaaaga gagcattctg ccaaaacgta actccgacaa actgatcgcg 9060 cgcaagaaag actgggatcc gaaaaaatac ggtggtttcg attctccaac cgttgcttat 9120 tccgttctgg tggtagccaa agttgagaaa ggtaaaagca aaaaactgaa atccgtaaag 9180 gaactgctgg gtattactat catggagcgt agctccttcg aaaaaaaccc gatcgatttt 9240 ctggaagcga aaggctataa agaagtcaaa aaggacctga tcatcaaact gccaaaatac 9300 agcctgttcg agctggaaaa cggccgtaaa cgtatgctgg catctgcggg cgaactgcag 9360 aaaggcaacg agctggctct gccgtccaaa tacgtgaact ttctgtacct ggcctctcac 9420 tacgaaaaac tgaaaggttc cccggaagac aacgaacaga aacagctgtt cgtagagcag 9480 cacaaacact acctggacga gatcatcgaa cagatttctg aattttctaa acgtgtgatt 9540 ctggctgatg cgaatctgga taaagttctg tctgcctata acaagcatcg tgacaaaccg 9600 atccgcgaac aggctgagaa catcatccac ctgttcactc tgactaacct gggcgcgcca 9660 gcggctttca agtactttga taccaccatt gaccgcaagc gttacacctc cactaaagaa 9720 gtgctggacg cgactctgat ccaccagtcc atcaccggtc tgtacgagac ccgtatcgat 9780 ctgagccagc tgggcggtga caaaaggccg gcggccacga aaaaggccgg ccaggcaaaa 9840 aagaaaaagt gacaaagccc gaaaggaagc tgagttggct gctgccaccg ctgagcaata 9900 actagcataa ccccttgggg cctctaaacg ggtcttgagg ggttttttgc tgaaaggagg 9960 aactatatcc ggat 9974 <210> 2 <211> 1547 <212> PRT <213> Artificial Sequence <400> 2 Met Lys Gln Ser Thr Ile Ala Leu Ala Leu Leu Pro Leu Leu Phe Thr 1 5 10 15 Pro Val Thr Lys Ala Met Ser Asp Lys Ile Ile His Leu Thr Asp Asp 20 25 30 Ser Phe Asp Thr Asp Val Leu Lys Ala Asp Gly Ala Ile Leu Val Asp 35 40 45 Phe Trp Ala Glu Trp Cys Gly Pro Cys Lys Met Ile Ala Pro Ile Leu 50 55 60 Asp Glu Ile Ala Asp Glu Tyr Gln Gly Lys Leu Thr Val Ala Lys Leu 65 70 75 80 Asn Ile Asp Gln Asn Pro Gly Thr Ala Pro Lys Tyr Gly Ile Arg Gly 85 90 95 Ile Pro Thr Leu Leu Leu Phe Lys Asn Gly Glu Val Ala Ala Thr Lys 100 105 110 Val Gly Ala Leu Ser Lys Gly Gln Leu Lys Glu Phe Leu Asp Ala Asn 115 120 125 Leu Ala Gly Ser Gly Ser Gly His Met His His His His His His Asp 130 135 140 Asp Asp Asp Lys Met Pro Lys Lys Lys Arg Lys Val Gly Ile His Gly 145 150 155 160 Val Pro Ala Ala Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr 165 170 175 Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser 180 185 190 Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys 195 200 205 Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala 210 215 220 Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn 225 230 235 240 Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val 245 250 255 Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu 260 265 270 Asp Lys Lys His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu 275 280 285 Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys 290 295 300 Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala 305 310 315 320 Leu Ala His Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp 325 330 335 Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val 340 345 350 Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly 355 360 365 Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg 370 375 380 Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu 385 390 395 400 Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys 405 410 415 Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp 420 425 430 Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln 435 440 445 Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu 450 455 460 Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu 465 470 475 480 Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr 485 490 495 Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu 500 505 510 Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly 515 520 525 Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu 530 535 540 Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp 545 550 555 560 Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln 565 570 575 Ile His Leu Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe 580 585 590 Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr 595 600 605 Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg 610 615 620 Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn 625 630 635 640 Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu 645 650 655 Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro 660 665 670 Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr 675 680 685 Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser 690 695 700 Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg 705 710 715 720 Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu 725 730 735 Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala 740 745 750 Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp 755 760 765 Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu 770 775 780 Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys 785 790 795 800 Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg 805 810 815 Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly 820 825 830 Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser 835 840 845 Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser 850 855 860 Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly 865 870 875 880 Asp Ser Leu His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile 885 890 895 Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys 900 905 910 Val Met Gly Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg 915 920 925 Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met 930 935 940 Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys 945 950 955 960 Glu His Pro Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu 965 970 975 Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp 980 985 990 Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser 995 1000 1005 Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp 1010 1015 1020 Lys Asn Arg Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys 1025 1030 1035 1040 Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr 1045 1050 1055 Gln Arg Lys Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser 1060 1065 1070 Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg 1075 1080 1085 Gln Ile Thr Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr 1090 1095 1100 Lys Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr 1105 1110 1115 1120 Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr 1125 1130 1135 Lys Val Arg Glu Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu 1140 1145 1150 Asn Ala Val Val Gly Thr Ala Leu Ile Lys Tyr Pro Lys Leu Glu 1155 1160 1165 Ser Glu Phe Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met 1170 1175 1180 Ile Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe 1185 1190 1195 1200 Phe Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1205 1210 1215 Asn Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr 1220 1225 1230 Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys 1235 1240 1245 Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln 1250 1255 1260 Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp 1265 1270 1275 1280 Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly 1285 1290 1295 Phe Asp Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val 1300 1305 1310 Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu Gly 1315 1320 1325 Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile Asp Phe 1330 1335 1340 Leu Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu Ile Ile Lys 1345 1350 1355 1360 Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly Arg Lys Arg Met 1365 1370 1375 Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro 1380 1385 1390 Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu 1395 1400 1405 Lys Gly Ser Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln 1410 1415 1420 His Lys His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser 1425 1430 1435 1440 Lys Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1445 1450 1455 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile 1460 1465 1470 Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys 1475 1480 1485 Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu 1490 1495 1500 Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu 1505 1510 1515 1520 Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp Lys Arg Pro Ala Ala 1525 1530 1535 Thr Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys 1540 1545 <210> 3 <211> 1399 <212> PRT <213> Artificial Sequence <400> 3 Met Pro Lys Lys Lys Arg Lys Val Gly Ile His Gly Val Pro Ala Ala 1 5 10 15 Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr Asn Ser Val Gly 20 25 30 Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe Lys 35 40 45 Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile Gly 50 55 60 Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu Lys 65 70 75 80 Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys Tyr 85 90 95 Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser Phe 100 105 110 Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys His 115 120 125 Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr His 130 135 140 Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp Ser 145 150 155 160 Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His Met 165 170 175 Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro Asp 180 185 190 Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr Asn 195 200 205 Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala Lys 210 215 220 Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn Leu 225 230 235 240 Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn Leu 245 250 255 Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe Asp 260 265 270 Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp Asp 275 280 285 Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp Leu 290 295 300 Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp Ile 305 310 315 320 Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser Met 325 330 335 Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys Ala 340 345 350 Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe Asp 355 360 365 Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser Gln 370 375 380 Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp Gly 385 390 395 400 Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg Lys 405 410 415 Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu Gly 420 425 430 Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe Leu 435 440 445 Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile Pro 450 455 460 Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp Met 465 470 475 480 Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu Val 485 490 495 Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr Asn 500 505 510 Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser Leu 515 520 525 Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys Tyr 530 535 540 Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln Lys 545 550 555 560 Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr Val 565 570 575 Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp Ser 580 585 590 Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly Thr 595 600 605 Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp Asn 610 615 620 Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr Leu 625 630 635 640 Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala His 645 650 655 Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr Thr 660 665 670 Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp Lys 675 680 685 Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe Ala 690 695 700 Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe Lys 705 710 715 720 Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu His 725 730 735 Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly Ile 740 745 750 Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly Arg 755 760 765 His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln Thr 770 775 780 Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile Glu 785 790 795 800 Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro Val 805 810 815 Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln 820 825 830 Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg Leu 835 840 845 Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys Asp 850 855 860 Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg Gly 865 870 875 880 Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys Asn 885 890 895 Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe 900 905 910 Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp Lys 915 920 925 Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr Lys 930 935 940 His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp Glu 945 950 955 960 Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser Lys 965 970 975 Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg Glu 980 985 990 Ile Asn Asn Tyr His His Ala His Asp Ala Tyr Leu Asn Ala Val Val 995 1000 1005 Gly Thr Ala Leu Ile Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val 1010 1015 1020 Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser 1025 1030 1035 1040 Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser Asn 1045 1050 1055 Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile 1060 1065 1070 Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val 1075 1080 1085 Trp Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser Met 1090 1095 1100 Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly Phe 1105 1110 1115 1120 Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu Ile Ala 1125 1130 1135 Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe Asp Ser Pro 1140 1145 1150 Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val Glu Lys Gly Lys 1155 1160 1165 Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met 1170 1175 1180 Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys 1185 1190 1195 1200 Gly Tyr Lys Glu Val Lys Asp Leu Ile Ile Leu Pro Lys Lys 1205 1210 1215 Ser Leu Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala 1220 1225 1230 Gly Glu Leu Gln Lys Gly Asn Leu Glu Ala Leu Pro Ser Lys Tyr Val 1235 1240 1245 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser Pro 1250 1255 1260 Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys His Tyr 1265 1270 1275 1280 Leu Asp Glu Ile Ile Glu Gln Ile Served Glu Phe Served Lys Arg Val Ile 1285 1290 1295 Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys His 1300 1305 1310 Arg Asp Lys Pro With Arg Glu Gln Ala Glu Asn With His Leu Phe 1315 1320 1325 Thr Leu Thr Asn Leu Gly Ala Pro Ala Phe Lys Tyr Phe Asp Thr 1330 1335 1340 Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp Ala 1345 1350 1355 1360 Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg Ile Asp 1365 1370 1375 Leu Ser Gln Leu Gly Gly Asp Lys Arg Pro Ala Ala Thr Lys Lys Ala 1380 1385 1390 Gly Gln Ala Lys Lys Lys Lys 1395 <210> 4 <211> 225 <212> DNA <213> Artificial Sequence <400> 4 ggcttgtcgg actcttcgct attacgccag ctggcgaagg gggatgtgct gcaaggcgat 60 taagttgggt aacgccaggg ttttcccagt cacgacgtta ggaaattaat acgactcact 120 ataggagagc acagtcagcc tggcggtttt agagctagaa atagcaagtt aaaataaggc 180 tagtccgtta tcaacttgaa aaagtggcac cgagtcggtg ctttt 225 <210> 5 <211> 225 <212> DNA <213> Artificial Sequence <400> 5 ggcttgtcgg actcttcgct attacgccag ctggcgaagg gggatgtgct gcaaggcgat 60 taagttgggt aacgccaggg ttttcccagt cacgacgtta ggaaattaat acgactcact 120 ataggcttcc agaattggat ctccggtttt agagctagaa atagcaagtt aaaataaggc 180 tagtccgtta tcaacttgaa aaagtggcac cgagtcggtg ctttt 225 <210> 6 <211> 102 <212> RNA <213> Artificial Sequence <400> 6 ggagagcaca gucagccugg cgguuuuaga gcuagaaaua gcaaguuaaa auaaggcuag 60 uccguuauca acuugaaaaa guggcaccga gucggugcuu uu 102 <210> 7 <211> 102 <212> RNA <213> Artificial Sequence <400> 7 ggcuuccaga auuggaucuc cgguuuuaga gcuagaaaua gcaaguuaaa auaaggcuag 60 uccguuauca acuugaaaaa guggcaccga gucggugcuu uu 102 <210> 8 <211> 14138 <212> DNA <213> Artificial Sequence <400> 8 ggcgcgccct ctacctgctc tcggaccgt gggggtgggg ggtggaggaa ggagtgggggg 60 gtcggtcctg ctggcttgtg ggtgggaggc gcatgttctc caaaacccg cgcgagctgc 120 aatcctgagg gagctgcagt ggaggaggcg gagagaaggc cgcaccctc tccgcagggg 180 gaggggagtg ccgcaatacc tttatggg ttctctgctg cctccttttc ctaggaccg 240 ccctgggcct agaaaaatcc ctcctcccc cgcgatctcg tcatcgcctc catgtcagtt 300 tgctccttct cgattatggg cgggatctt tgccctggc gcgccccaga cccgggcctg 360 gggggcaagt cggggggcgg ggggaggtcg ggcagggtcc cctgggagga tgggacgtg 420 ctgtgccct agcggccacc agagggcacc aggacaccac tgcggtcggc tcagcggctc 480 ctgccctggt caggggcgc caggtcctgc ccctcctggg gagggcgggg ggcgagaagg 540 gcgattttaa ttaacccacg ttcacatg cacatcccag taatttggaa acatttgtt 600 660 cagttttgg cctgttttag tgacaggca tcagcacat gctgcatttc tctccagtgt 720 tgtaatcaaa gaaaccctcc catagcttta aatgatattc cttccccttc caattatgtg 780 gggggaaaac aaccctattc tccacccaga agtgttaact caagaattac attttcaaga 840 agtttccaga ttcgtaaaac cagaattaga tgtctttcac ctaaatgtct cggtgttgac 900 caaaggaaca cacaggtttc tcatttaact tttttaatgg gtctcaaaat tctgtgacaa 960 atttttggtc aagttgtttc cattaaaaag tactgatttt aaaaactaat aacttaaaac 1020 tgccacacgc aaaaaagaaa accaaagtgg tccacaaaac attctccttt ccttctgaag 1080 gttttacgat gcattgttat cattaaccag tcttttacta ctaaacttaa atggccaatt 1140 gaaacaaaca gttctgagac cgttcttcca ccactgatta agagtggggt ggcaggtatt 1200 agggataatg ctagcttact tgtacagctc gtccatgccg agagtgatcc cggcggcggt 1260 cacgaactcc agcaggacca tgtgatcgcg cttctcgttg gggtctttgc tcagggcgga 1320 ctgggtgctc aggtagtggt tgtcgggcag cagcacgggg ccgtcgccga tgggggtgtt 1380 ctgctggtag tggtcggcga gctgcacgct gccgtcctcg atgttgtggc ggatcttgaa 1440 gttcaccttg atgccgttct tctgcttgtc ggccatgata tagacgttgt ggctgttgta 1500 gttgtactcc agcttgtgcc ccaggatgtt gccgtcctcc ttgaagtcga tgcccttcag 1560 ctcgatgcgg ttcaccaggg tgtcgccctc gaacttcacc tcggcgcggg tcttgtagtt 1620 gccgtcgtcc ttgaaga tggtgcgctc ctggacgtag ccttcgggca tggcggactt 1680 gaagagtcg tgctgcttca tgtggtcggg gtagcggctg aagcactgca cgccgtaggt 1740 cagggtggtc acgagggtgg gccagggcac gggcagcttg ccggtggtgc agatgaactt 1800 cagggtcagc ttgccgtagg tggcatcgcc ctcgccctcg ccggacacgc tgaacttgtg 1860 gccgtttacg tcgccgtcca gctcgaccag gatgggcacc accccggtga acagctcctc 1920 gcccttgctc accatggtgg cgtcgaccgt acgtcacgac acctgaaatg gaaaaaa 1980 actttgaacc actgtctgag gcttgagaat gaaccaagat ccaaactcaa aaagggcaaa 2040 ttccaaggag aattacatca agtgccaagc tggcctaact tcagtctcca cccactcagt 2100 gtgggggaaac tccatcgcat aaaacccctc cccccaacct aaagacgacg tactccaaaa 2160 gctcgagaac taatcgaggt gcctggacgg cgcccggtac tccgtggagt cacatgaagc 2220 gacggctgag gacggaaagg cccttttcct ttgtgtgggt gactcacccg cccgctctcc 2280 cgagcgccgc gtcctccatt ttgagctccc tgcagcaggg ccgggaagcg gccatctttc 2340 cgctcacgca actggtgccg accgggccag ccttgccgcc cagggcgggg cgatacacgg 2400 cggcgcgagg ccaggcacca gagcaggccg gccagcttga gactaccccc gtccgattct 2460 cggtggccgc gctcgcaggc cccgcctcgc cgaacatgtg cgctgggacg cacgggcccc 2520 gtcgccgccc gcggccccaa aaaccgaaat accagtgtgc agatcttggc ccgcatttac 2580 aagactatct tgccagaaaa aaagcgtcgc agcaggtcat caaaaatttt aaatggctag 2640 agacttatcg aaagcagcga gacaggcgcg aaggtgccac cagattcgca cgcggcggcc 2700 ccagcgccca ggccaggcct caactcaagc acgaggcgaa ggggctcctt aagcgcaagg 2760 cctcgaactc tcccacccac ttccaacccg aagctcggga tcaagaatca cgtactgcag 2820 ccagtggaag taattcaagg cacgcaaggg ccataacccg taaagaggcc aggcccgcgg 2880 gaaccacaca cggcacttac ctgtgttctg gcggcaaacc cgttgcgaaa aagaacgttc 2940 acggcgacta ctgcacttat atacggttct cccccaccct cgggaaaaag gcggagccag 3000 tacacgacat cactttccca gtttaccccg cgccaccttc tctaggcacc ggttcaattg 3060 ccgacccctc cccccaactt ctcggggact gtgggcgatg tgcgctctgc ccactgacgg 3120 gcaccggagc cctagattcg attccctttg gggcaaaact caccgcctaa tcccctataa 3180 ctctaccggg gagcccggtg gagagcagac gggctgacgc tgccacctgc cggccatccc 3240 aggataggac cgccgtattc aagtcgccct caggaaggac cctcggggca ccagaggcct 3300 tcgaagcccc aatgagtgag gcaactgagg gtcgcgggtg ccattacaag gcccagccaa 3360 ggcctagagc caaggcttga accgtggggg acccccaagc cccacctgcc caggaacagc 3420 agacactggg acactttgtt tcaggtcctg cccaggcccc tcccactgtg aggctgggat 3480 ttgtcgccca gggtgcagat gagaagagtg gggaaagcag tcctgagcca ggaaattcta 3540 ccgggtaggg gaggcgcttt tcccaaggca gtctggagca tgcgctttag cagccccgct 3600 gggcacttgg cgctacacaa gtggcctctg gcctcgcaca cattccacat ccaccggtag 3660 gcgccaaccg gctccgttct ttggtggccc cttcgcgcca ccttctactc ctcccctagt 3720 caggaagttc ccccccgccc cgcagctcgc gtcgtgcagg acgtgacaaa tggaagtagc 3780 acgtctcact agtctcgtgc agatggacag caccgctgag caatggaagc gggtaggcct 3840 ttggggcagc ggccaatagc agctttgctc cttcgctttc tgggctcaga ggctgggaag 3900 gggtgggtcc gggggcgggc tcaggggcgg gctcaggggc ggggcgggcg cccgaaggtc 3960 ctccggaggc ccggcattct gcacgcttca aaagcgcacg tctgccgcgc tgttctcctc 4020 ttcctcatct ccgggccttt cgacctccta gggccaccat ggtgagcaag ggcgaggacg 4080 acaacatggc catcatcaag gagttcatgc gcttcaaggt gcacatggag ggctccgtga 4140 acggccacga gttcgagatc gagggcgagg gcgagggccg cccctacgag ggcacccaga 4200 ccgccaagct gaaggtgacc aagggcggcc ccctgccctt cgcctgggac atcctgtccc 4260 ctcagttcat gtacggctcc aaggcctacg tgaagcaccc cgccgacatc cccgactact 4320 tgaagctgtc cttccccgag ggcttcaagt gggagcgcgt gatgaacttc gaggacggcg 4380 gcgtggtgac cgtgacccag gactcctccc tgcaggacgg cgagttcatc tacaaggtga 4440 agctgcgcgg caccaacttc ccctccgacg gccccgtaat gcagaagaag accatgggct 4500 gggaggcctc ctccgagcgg atgtaccccg aggacggcgc cctgaagggc gagatcaagc 4560 agaggctgaa gctgaaggac ggcggccact acgacgccga ggtcaagacc acctacaagg 4620 ccaagaagcc cgtgcagctg cccggcgcct acaacgtcaa catcaagctg gacatcacct 4680 cccacaacga ggactacacc atcgtggaac agtacgagcg cgccgagggc cgccactcca 4740 ccggcggcat ggacgagctg tacaagtgag gatccgctga tcagcctcga ctgtgccttc 4800 tagttgccag ccatctgttg tttgcccctc ccccgtgcct tccttgaccc tggaaggtgc 4860 cactcccact gtcctttcct aataaaatga ggaaattgca tcgcattgtc tgagtaggtg 4920 tcattctatt ctggggggtg gggtggggca ggacagcaag ggggaggatt gggaagacaa 4980 tagcaggcat gctggggatg cggtgggctc tatggcttct gaggcggaaa gaacccttct 5040 gaggcggaaa gaaccagctg ccttaatata acttcgtata atgtatgcta tacgaagtta 5100 ttaggtctga agaggagttt acgtccagcc aattctgtgg aatgtgtgtc agttagggtg 5160 tggaaagtcc ccaggctccc cagcaggcag aagtatgcaa agcatgcatc tcaattagtc 5220 agcaaccagg tgtggaaagt ccccaggctc cccagcaggc agaagtatgc aaagcatgca 5280 tctcaattag tcagcaacca tagtcccgcc cctaactccg cccatcccgc ccctaactcc 5340 gcccagttcc gcccattctc cgccccatgg ctgactaatt ttttttattt atgcagaggc 5400 cgaggccgcc tctgcctctg agctattcca gaagtagtga ggaggctttt ttggaggcct 5460 aggcttttgc aaaaagctcc cgggagcttg tatatccatt ttcggcggcc gcgccaccat 5520 gaccgagtac aagcccacgg tgcgcctcgc cacccgcgac gacgtcccca gggccgtacg 5580 caccctcgcc gccgcgttcg ccgactaccc cgccacgcgc cacaccgtcg atccggaccg 5640 ccacatcgag cgggtcaccg agctgcaaga actcttcctc acgcgcgtcg ggctcgacat 5700 cggcaaggtg tgggtcgcgg acgacggcgc cgcggtggcg gtctggacca cgccggagag 5760 cgtcgaagcg ggggcggtgt tcgccgagat cggcccgcgc atggccgagt tgagcggttc 5820 ccggctggcc gcgcagcaac agatggaagg cctcctggcg ccgcaccggc ccaaggagcc 5880 cgcgtggttc ctggccaccg tcggagtctc gcccgaccac cagggcaagg gtctgggcag 5940 cgccgtcgtg ctccccggag tggaggcggc cgagcgcgcc ggggtgcccg ccttcctgga 6000 gacctccgcg ccccgcaacc tccccttcta cgagcggctc ggcttcaccg tcaccgccga 6060 cgtcgaggtg cccgaaggac cgcgcacctg gtgcatgacc cgcaagcccg gtgcctgaga 6120 attcgcggga ctctggggtt cgaaatgacc gaccaagcga cgcccaacct gccatcacga 6180 gatttcgatt ccaccgccgc cttctatgaa aggttgggct tcggaatcgt tttccgggac 6240 gccggctgga tgatcctcca gcgcggggat ctcatgctgg agttcttcgc ccaccccaac 6300 ttgtttattg cagcttataa tggttacaaa taaagcaata gcatcacaaa tttcacaaat 6360 aaagcatttt tttcactgca ttctagttgt ggtttgtcca aactcatcaa tgtatcttat 6420 catgtctgta taccgctcga ctagagcttg cggaaccctt aatataactt cgtataatgt 6480 atgctatacg aagttag gtccgctggc catctacgag ccaagactt tcaatcttt 6540 ggctgccttg gccagtagga ggcgacacga aggatttgct gctgccttgg gggatgggaa 6600 ggaacctgaa ggcattttt ccagagtggt gcagtaccac tgaggactgt tgctgtattg 6660 attaggaaaa gagacagagt aatttgcagt ttgttgatt tatactgggc tgcaggtcga 6720 gggatcttca tagagaga gggacagcta tgactgggag tagtcaggg aggaaaaa 6780 atctggctag taaaaacatgt aaggaaaatt tagggt taaagaaaaaacacaa 6840 aaaaaatat aaaaaaaatc taacctcag tcaggcttt tctatggaat aaggaatgga 6900 cagcaggggg ctgtttcata tactgatgac ctctttag ccaccttgt tcatggcagc 6960 cagcatatgg catatgttgc caactctaa accaatact cattctgatg ttttaatga 7020 tttgccctcc catatgtcct tccgagtgag agacacaaa aattccaaca cactattgca 7080 atgaaaataa atttcctttta ttagccagaa gtcagatgct caggggctt catgatgtcc 7140 ccataatttt tggcagagggg aaaagatct cagtggtatt tgtgagccag ggcattggcc 7200 acaccagcca ccaccttctg ataggcagcc tgcggtacct tacatggtgg cgaattcgtt 7260 tgccaaaatg atgagacagc acaataacca gcacgttgcc caggagctgt aggaaaaaga 7320 agaaggcatg aacatggtta gcagaggctc tagagccgcc ggtcacacgc cagaagccga 7380 accccgccct gccccgtccc ccccgaaggc agccgtcccc ctgcggcagc cccgaggctg 7440 gagatggaga aggggacggc ggcgcggcga cgcacgaagg ccctccccgc ccatttcctt 7500 cctgccggcg ccgcaccgct tcgcccgcgc ccgctagagg gggtgcggcg gcgcctccca 7560 gatttcggct ccgccagatt tgggacaaag gaagtccctg cgccctctcg cacgattacc 7620 ataaaaggca atggctgcgg ctcgccgcgc ctcgacagcc gccggcgctc cggggccgcc 7680 gcgcccctcc cccgagccct ccccggcccg aggcggcccc gccccgcccg gcacccccac 7740 ctgccgccac cccccgcccg gcacggcgag ccccgcgcca cgccccgcac ggagccccgc 7800 acccgaagcc gggccgtgct cagcaactcg gggagggggg tgcagggggg ggttacagcc 7860 cgaccgccgc gcccacaccc cctgctcacc cccccacgca cacaccccgc acgcagcctt 7920 tgttcccctc gcagcccccc cgcaccgcgg ggcaccgccc ccggccgcgc tcccctcgcg 7980 cacacgcgga gcgcacaaag ccccgcgccg cgcccgcagc gctcacagcc gccgggcagc 8040 gcgggccgca cgcggcgctc cccacgcaca cacacacgca cgcacccccc gagccgctcc 8100 cccccgcaca aagggccctc ccggagccct ttaaggcttt cacgcagcca cagaaaagaa 8160 acgagccgtc attaaaccaa gcgctaatta cagcccggag gagaagggcc gtcccgcccg 8220 ctcacctgtg ggagtaacgc ggtcagtcag agccggggcg ggcggcgcga ggcggcgcgg 8280 agcggggcac ggggcgaagg caacgcagcg actcccgccc gccgcgcgct tcgcttttta 8340 tagggccgcc gccgccgccg cctcgccata aaaggaaact ttcggagcgc gccgctctga 8400 ttggctgccg ccgcacctct ccgcctcgcc ccgccccgcc cctcgccccg ccccgccccg 8460 cctggcgcgc gccccccccc cccccgcccc catcgctgca caaaataatt aaaaaataaa 8520 taaatacaaa attgggggtg gggagggggg ggagatgggg agagtgaagc agaacgtggg 8580 gctcacctcg acccatggta atagcgatga ctaatacgta gatgtactgc caagtaggaa 8640 agtcccataa ggtcatgtac tgggcataat gccaggcggg ccatttaccg tcattgacgt 8700 caataggggg cgtacttggc atatgataca cttgatgtac tgccaagtgg gcagtttacc 8760 gtaaatagtc cacccattga cgtcaatgga aagtccctat tggcgttact atgggaacat 8820 acgtcattat tgacgtcaat gggcgggggt cgttgggcgg tcagccaggc gggccattta 8880 ccgtaagtta tgtaacgcgg aactccatat atgggctatg aactaatgac cccgtaattg 8940 attactatta ataactagtc aataatcaat gtcgtaaatg tcgtaaatgt ctcagctagt 9000 caggtagtaa aaggtgtcaa ctaggcagtg gcagagcagg attcaaattc agggctgttg 9060 tgatgcctcc gcagactctg agcgccacct ggtggtaatt tgtctgtgcc tcttctgacg 9120 tggaagaaca gcaactaaca cactaacacg gcatttacta tgggccagcc attgtacgcg 9180 ttgcttaacc tgattcttgg gcgttgtcct gcaggggatt gagcaggtgt acgaggacga 9240 gcccaatttc tctatattcc cacagtcttg agtttgtgtc acaaaataat tatagtgggg 9300 tggagatggg aaatgagtcc aggcaacacc taagcctgat tttatgcatt gagactgcgt 9360 gttattacta aagatctttg tgtcgcaatt tcctgatgaa gggagatagg ttaaaaagca 9420 cggatctact gagttttaca gtcatcccat ttgtagactt ttgctacacc accaaagtat 9480 agcatctgag attaaatatt aatctccaaa ccttaggccc cctcacttgc atccttacgg 9540 tcagataact ctcactcata ctttaagccc attttgtttg ttgtacttgc tcatccagtc 9600 ccagacatag cattggcttt ctcctcacct gttttaggta gccagcaagt catgaaatca 9660 gataagttcc accaccaatt aacactaccc atcttgagca taggcccaac agtgcattta 9720 ttcctcattt actgatgttc gtgaatattt accttgattt tcattttttt ctttttctta 9780 agctgggatt ttactcctga ccctattcac agtcagatga tcttgactac cactgcgatt 9840 ggacctgagg ttcagcaata ctccccttta tgtcttttga atacttttca ataaatctgt 9900 ttgtattttc attagttagt aactgagctc agttgccgta atgctaatag cttccaaact 9960 agtgtctctg tctccagtat ctgataaatc ttaggtgttg ctgggacagt tgtcctaaaa 10020 ttaagataaa gcatgaaaat aactgacaca actccattac tggctcctaa ctacttaaac 10080 aatgcattct atcatcacaa atgtgaaaaa ggagttccct cagtggacta accttatctt ttctcaacac ctttttcttt gcacaatttt ccacacatgc ctacaaaaag tacttatgcg gccgccataa aagttttgtt actttataga agaaatttg agttttgtt ttttttaata aataataaa cataaataaa ttgtttgttg aatttattat tagtatgtaa gtgtaaatat aataaaactt aatatctatt caataata aataacctc throwing ccgataaaac acatgcgtca attttacaca tgattatctt taacgtacgt cacaatatga ttatctttct agggttaatc tagctgcgtg ttctgcagcg tgtcgagcat cttcatctgc tccatcacgc 10500 10560. tgtaaaacac atttgcaccg cgagtctgcc cgtcctccac gggttcaaaa acgtgaatga acgaggcgcg ctcactggcc gtcgttttac aacgtcgtga ctgggaaac cctggcgtta cccaacttaa tcgccttgca gcacatcccc ctttcgccag ctggcgtaat agcgaagagg 10680. cccgcaccga tcgcccttcc caacagttgc gcagcctga tggcgaatgg gacgcgccct gtagcggcgc attaagcgcg gcgggtgtgg tggttacgcg cagcgtgacc gctacacttg 10800 ccagcgccct agcgcccgct cctttcgctt tcttcccttc ctttctcgcc acgttcgccg 10860 gctttccccg tcaagctcta aatcgggggc tccctttagg gttccgattt agtgctttac 10920 ggcacctcga ccccaaaaaa cttgattagg gtgatggttc acgtagtggg ccatcgccct 10980 gatagacggt ttttcgccct ttgacgttgg agtccacgtt ctttaatagt ggactcttgt 11040 tccaaactgg aacaacactc aaccctatct cggtctattc ttttgattta taagggattt 11100 tgccgatttc ggcctattgg ttaaaaaatg agctgattta acaaaaattt aacgcgaatt 11160 ttaacaaaat attaacgctt acaatttagg tggcactttt cggggaaatg tgcgcggaac 11220 ccctatttgt ttatttttct aaatacattc aaatatgtat ccgctcatga gacaataacc 11280 ctgataaatg cttcaataat attgaaaaag gaagagtatg agtattcaac atttccgtgt 11340 cgcccttatt cccttttttg cggcattttg ccttcctgtt tttgctcacc cagaaacgct 11400 ggtgaaagta aaagatgctg aagatcagtt gggtgcacga gtgggttaca tcgaactgga 11460 tctcaacagc ggtaagatcc ttgagagttt tcgccccgaa gaacgttttc caatgatgag 11520 cacttttaaa gttctgctat gtggcgcggt attatcccgt attgacgccg ggcaagagca 11580 actcggtcgc cgcatacact attctcagaa tgacttggtt gagtactcac cagtcacaga 11640 aaagcatctt acggatggca tgacagtaag agaattatgc agtgctgcca taaccatgag 11700 tgataacact gcggccaact tacttctgac aacgatcgga ggaccgaagg agctaaccgc 11760 ttttttgcac aacatggggg atcatgtaac tcgccttgat cgttgggaac cggagctgaa 11820 tgaagccata ccaaacgacg agcgtgacac cacgatgcct gtagcaatgg caacaacgtt 11880 gcgcaaacta ttaactggcg aactacttac tctagcttcc cggcaacaat taatagactg 11940 gatggaggcg gataaagttg caggaccact tctgcgctcg gcccttccgg ctggctggtt 12000 tattgctgat aaatctggag ccggtgagcg tggttcacgc ggtatcattg cagcactggg 12060 gccagatggt aagccctccc gtatcgtagt tatctacacg acggggagtc aggcaactat 12120 ggatgaacga aatagacaga tcgctgagat aggtgcctca ctgattaagc attggtaact 12180 gtcagaccaa gtttactcat atatacttta gattgattta aaacttcatt tttaatttaa 12240 aaggatctag gtgaagatcc ttttgataa tctcatgacc aaatccctt aacgtgagtt 12300 ttcgttccac tgagcgtcag accccgtaga aagatcaaa ggatctctt gagatccttt 12360 ttttctgcgc gtaatctgct gcttgcaac aaaaaaacca ccgctaccag cggtggtttg 12420 tttgccggat cagagctac caactttt tccgaggta actggcttca gcagagcgca 12480 gataccaaat actgtccttc tagtgtagcc gtagttaggc caccactca agaactctgt 12540 agcaccgcct acataccctcg ctctgctaat cctgttacca gtggctgctg ccagtggcga 12600 taagtcgtgt cttaccggggt tggactcag acgatagtta ccggatagg cgcagcggtc 12660 gggctgaacg gggggttcgt gcacacagcc cagcttggag cgaacgacct acaccgaact 12720 gagataccta cagcgtgagc tatgagaag cgccacgctt cccgaaggga gaaaggcgga 12780 caggtatccg gtaagcggca gggtcggaac aggagagcgc acgagggag ttccaggggg 12840 aaacgcctgg tatctttata gtcctgtcgg gtttcgccac ctctgacttg agcgtcgatt 12900 tttgtgatgc tcgtcagggg ggcggagcct atggaaaaac gccagcaacg cggccttttt 12960 acggttcctg gccttttgct ggccttttgc tcacatgttc tttcctgcgt tatcccctga 13020 ttctgtggat aaccgtatta ccgcctttga gtgagctgat accgctcgcc gcagccgaac 13080 gaccgagcgc agcgagtcag tgagcgagga agcggaagag cgcccaatac gcaaaccgcc 13140 tctccccgcg cgttggccga ttcattaatg cagctggcac gacaggtttc ccgactggaa 13200 agcgggcagt gagcgcaacg caattaatgt gagttagctc actcattagg caccccaggc 13260 tttacacttt atgcttccgg ctcgtatgtt gtgtggaatt gtgagcggat aacaatttca 13320 cacaggaaac agctatgacc atgattacgc caagcgcgcc cgccgggtaa ctcacggggt 13380 atccatgtcc atttctgcgg catccagcca ggatacccgt cctcgctgac gtaatatccc 13440 agcgccgcac cgctgtcatt aatctgcaca ccggcacggc agttccggct gtcgccggta 13500 ttgttcgggt tgctgatgcg cttcgggctg accatccgga actgtgtccg gaaaagccgc 13560 gacgaactgg tatcccaggt ggcctgaacg aacagttcac cgttaaaggc gtgcatggcc 13620 acaccttccc gaatcatcat ggtaaacgtg cgtttcgct caacgtcaat gcagcagcag 13680 tcatcctcgg caaactcttt ccatgccgct tcaacctcgc gggaaaaggc acgggcttct 13740 tcctccccga tgcccagata gcgccagctt gggcgatgac tgagccggaa aaaagacccg 13800 acgatatgat cctgatgcag ctagattaac cctagaaaga tagtctgcgt aaaattgacg 13860 catgcattct tgaaatattg ctctctctt ctaaatagcg cgaatccgtc gctgtgcatt 13920 taggacatct cagtcgccgc ttggagctcc cgtgaggcgt gcttgtcaat gcggtaagtg 13980 tcactgatt tgaactataa cgaccgcgtg agtcaaaatg acgcatgatt atcttttacg 14040 tgactttaa gatttaactc atacgataat tatattgtta tttcatgttc tacttacgtg 14100 ataacttatt atatatatat tttcttgtta tagatatc 14138 <210> 9 <211> 1069 <212> DNA <213> On sow <400> 9 gtgctgagtc cttttcccat cccacccacc tggagctccc ctcttccagt cctgagccac 60 ttgaactggc ctggtttttg ccatcctgcg ctgccctctc tccggactcg agccactgct 120 gagggcctca ggccagtcca tcctcgtctt gtctctttcg ccctgctctt tccccacctt 180 gagcgctctt aaccagcctg gcccgtgcca cctctactct gccatcgaat gctgccccac 240 tttctcgagt ccgccacttc tcccagcttc accggtaccc actgtttccc ctagtccagg 300 caggtaccac tttccctgag cgtcctcctc ctctctcctg ggcctgtgct gcttcttttc 360 ccgctctctg gcctgggccg tttcttcggc cagcccccga gccttccatg ccctttcctt 420 caggtttctg ctcttcatcc ttggtctctg ccatctgttg ccatgtaagg gtgctctttc 480 ctgagccatc gccctcaagg cgctctgctc ctcaagtgga tgcttccctc gcctggctca 540 cctcctgctc tctctcctgc ccccttcacc tgcgtgccct cctcattctc cctctgtgcc 600 acctctggcc ttgcactgta ggctctctct tggggatgtt tctccttctc cacacacttc 660 tctttcactc tgtcctcttg ctttgtgtgg gcctgcagcg ttaccctttt ttctgggcac 720 actcagagca ccctcctct tctggttctg ggccacctgt ctgtcctcgg gtcatcttgc 780 tctctctgcc tggatgccct cctgtggctt tgggcagctt ctccctcctt cagagtgcac 840 cgccagttct cctaggcccg gtcacttccc cttcccaggg gacctagac cctgctaggt 900 cctctctc cacaacctgg gcccccaaac ctttccaaaa caccttgctt tctgcctcca 960 ttggtcttgt gttccagagc cagagtcact atatgccca gaaccaggat tccctctggt 1020 tctgagggct tttatcgcat cccctgcctg gctgcagtgg gtctttggg 1069 <210> 10 <211> 260 <212> DNA <213> On sow <400> 10 gacaggccac agaagagcct ctactcctcc ctctgtcccc gaggctgtct ccctcccagt 60 cttcccagct caggccagtc cccaggcctc tcttccctgc cagagcccgt caggttcggt 120 tactttgggg cccagagagg accctgtgaa ggaagcgtgg gtaggggcac gggaatgggg 180 aggatgcctg aagaggcccc cttagccaga agaggagcag aagaggagca ggtacccaga 240 agaggagcag ttcagggaaa 260 <210> 11 <211> 540 <212> DNA <213> On sow <400> 11 aatacccac gttattggg acaaaagttg ttagggaaaa tggggcctca gagttatgat 60 tcaagtcata attctttcca tttataattt cactcgagac tctgttaact gattccttgt 120 gtgttgtatc ttactcctca gctcacaatt acttttagtt attcacctta actgtatgaa 180 taacagtgga gaaaaggatt ctaccagaat actctaatta tggttttgag tcccctttcc 240 agactgaaga ttttcagtc ttttgatct gaggtgattt ttcagtcttt tcgatctgag 300 gtgacagtct caagctcctc aattcaccca gtctcttgat acttgtccat ttagggccac 360 caaagctact ttgacttcat actagagagt caattaatga ggccattctc tgatggacag 420 gtgaagcagg caaggtgact atattttgac taaacggtag aaaacagcct gagtgttaac 480 agtgtagcct ataaaaccca gagctgccca ccctgatcta aacttccagg aacataagaa 540 <210> 12 <211> 1009 <212> DNA <213> On sow <400> 12 agtaggtcac atttcagtaa aacctggctt tgtggattga gcatgtctg tctctctcctg 60 gtacttcatt agtcccctaa gtgggatttg ctgagcaaga ctcctcatt acagaatac 120 tccagtttag aattctcgca aaggctttt gtttccacaa gtagaatcta gaagcaatc 180 tcaagtaaca acagcagaga cctgaatccc aatccatctt tcctgtgtgt cctcttttac 240 ctccttccct ttcatgttga accacagtc cttttcagt ctgaagcta gtacgaaaga 300 aatgtacaga tgtaggtacc aagcaaagcc attagccaat aactggtgag atggagctaa 360 gaggaataa aagtgttcct aagaatagca cagcagaagc tagatccaca gatcttaaa 420 cavattttggt tgagtagag taggaggaaa gaggaagct ataatgcag ttttaggag 480 ctaagagcca gataaagggt aagggcagga ggaagtgcta tctcagctaa cgagatacat 540 gaaacaacgg tggagtcca gcaggcaca gatgagttga gaagcaatca gggccagaag 600 gatgtgcaag gcctcaaat aaaaaagcac agggccacag ggaaccttat ggaattaaa 660 aggaagga tgcagtcagg agagaaaa agatgctcc ctccccatg cccaggag 720 cagctgagca gccagtactt gggaagttag tagtaataag ttggtaagag ggagttctgt 780 tcgtggctca atggttaaca aatcagacta gaaaccgtga ggttgcgggt ttgatccctg 840 gccttgctca gtgggttaag gatccggcat tgccgtgacc tgtggtgtag gtcacagacg 900 tggctcagtt cccgcattc tgtggctctg gtgtaggctg gtggctacag ctctgattag 960 acccctaggc tgggaacctc catatgccct ggaagtggcc gtagaaaag 1009 <210> 13 <211> 872 <212> DNA <213> On sow <400> 13 ggatggggac tcatgtgaat tttctaaagg tgctatttaa acggggggca cgagtgccgg 60 cttggacag ggccgctcgc tctccaccct ttctctttc ccctcggcg cctctcaccc 120 cctgaggcct ctctcccccc acgacctct ctctctcctc tgaaaccctc tcctcctcag 180 ctgcatccca ccctcgtggc ctctctctct ctctgtctgt cctgtgtcct ctctcactgg 240 gttcagagc acagatgcc aaagcacaaa agcagttttc ccctggggtg ggaggaagca 300 agagactttg tacctatttt gtatgtgtat aataatttga gatgttttta attattttga 360 ttgctggaat aaagcatgtg gaaatgacc aaaccaatct tgcactggcc tcctgatttc 420 cttccttgga gacggaggga gggggagacc tgggggaggg cgcttggggg ggggtgggct 480 ctcttctttc tgcgctcccc cccccacct ccaacacctt gacgacccct cctgcttccg 540 cttgccttc tcaggcttta acactttctc ctcgccctct cagcatgcgc atgcgcgtgc 600 ctctacctcc cccgcacatc ctggcctgcc caccctgaat ggcctggccc agcgatgcca 660 ccaactctct cgctccgtcc acggctgggg aggggggcac tctgcagggt tggggggcac 720 tgggaggctg ggttgggtga gggaggggtg cctgggcccc caccccccag caagttctct 780 ccctaggcga actggagggt cgtctggcct cttgagcctt gttgctggct ctgagctcta 840 ccaagagagt gaccagcagg accgcaccat ca 872 <210> 14 <211> 727 <212> DNA <213> On sow <400> 14 gtggttgctg agactgcgtg ggggcccaag gagacctgga gaaaggaatg cttcctgctc 60 cttcttctgg ggccccagga gagccttccc agggccttgg agaggtgctg tccagggact 120 aaccctgtgc tctaggaagg ctgcaggccc tgaccagctg ggcaggtcct gggtccctcc 180 tggccttcta agttccccaa acatgagacc tctgggtgtg gggtggcctg gggaggtcat 240 tttgcccagg ccctacctcc tgcccattcc taaccctttt taaaaatctg tgcgtcctct 300 tcttccttct tctccctccc ttcccttttc gctcaccctc tgctgctggc ctgagagccg 360 gaggccccca gggggaaggc gactggtctc ctccccagtc tcagggaagg gagacagaga 420 atccaggaag ccagaactca gcagacgaag cacccaggga cctagagatg ggttgaaaag 480 ttgacagctg tcccacctgc ctcccaaggt ctcagggcct aaacctccaa ggcaggaaag 540 gcccctgtcc ctccctgggg tccatagaaa gagggacaag tctgcacgga ccatttgctg 600 taatattaac accttggctg tcattaggta gtcttggctg ttaattatgt cctgtgataa 660 tgtattatta gcacgccgac cacatagggt agggaactgc agctagtaaa caaaagtttg 720 ttcctat 727 <210> 15 <211> 100 <212> RNA <213> Artificial Sequence <400> 15 gaaggagcaa acugacaugg guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 16 <211> 100 <212> RNA <213> Artificial Sequence <400> 16 ugcagugggu cuuuggggac guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 17 <211> 100 <212> RNA <213> Artificial Sequence <400> 17 uuccaggaac auaagaaagu guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 18 <211> 100 <212> RNA <213> Artificial Sequence <400> 18 gcagucucag caaccacuga guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcuuuu 100 <210> 19 <211> 254 <212> PRT <213> Homo sapiens <400> 19 Met Gln Val Gln Cys Gln Gln Ser Pro Val Leu Ala Gly Ser Ala Thr 1 5 10 15 Leu Val Ala Leu Gly Ala Leu Ala Leu Tyr Val Ala Lys Pro Ser Gly 20 25 30 Tyr Gly Lys His Thr Glu Ser Leu Lys Pro Ala Ala Thr Arg Leu Pro 35 40 45 Ala Arg Ala Ala Trp Phe Leu Gln Glu Leu Pro Ser Phe Ala Val Pro 50 55 60 Ala Gly Ile Leu Ala Arg Gln Pro Leu Ser Leu Phe Gly Pro Pro Gly 65 70 75 80 Thr Val Leu Leu Gly Leu Phe Cys Leu His Tyr Phe His Arg Thr Phe 85 90 95 Val Tyr Ser Leu Leu Asn Arg Gly Arg Pro Tyr Pro Ala Ile Leu Ile 100 105 110 Leu Arg Gly Thr Ala Phe Cys Thr Gly Asn Gly Val Leu Gln Gly Tyr 115 120 125 Tyr Leu Ile Tyr Cys Ala Glu Tyr Pro Asp Gly Trp Tyr Thr Asp Ile 130 135 140 Arg Phe Ser Leu Gly Val Phe Leu Phe Ile Leu Gly Met Gly Ile Asn 145 150 155 160 Ile His Ser Asp Tyr Ile Leu Arg Gln Leu Arg Lys Pro Gly Glu Ile 165 170 175 Ser Tyr Arg Ile Pro Gln Gly Gly Leu Phe Thr Tyr Val Ser Gly Ala 180 185 190 Asn Phe Leu Gly Glu Ile Ile Glu Trp Ile Gly Tyr Ala Leu Ala Thr 195 200 205 Trp Ser Leu Pro Ala Leu Ala Phe Ala Phe Phe Ser Leu Cys Phe Leu 210 215 220 Gly Leu Arg Ala Phe His His His Arg Phe Tyr Leu Lys Met Phe Glu 225 230 235 240 Asp Tyr Pro Lys Ser Arg Lys Ala Leu Ile Pro Phe Ile Phe 245 250 <210> 20 <211> 10229 <212> DNA <213> Artificial Sequence <400> 20 ggcgcgccgg atggggactc atgtgaattt tctaaaggtg ctatttaaac ggggggcacg 60 agtgccggct ttggacaggg ccgctcgctc tccacccttt cttcttcccc ctcggccgcc 120 180. tctcacccc tgaggcctct ctccccccac gaccctctct ctctcctctg aaaccctctc 240. ctcctcagct gcatcccacc ctcgtggcct ctctctctct ctgtctgtcc tgtgtcctct ctcactgggt ttcagagcac agatgcccaa agcacaaaag cagttttccc ctggggtggg 300 aggaagcaag agactttgta cctattttgt atgtgtata taatttgaga tgtttttaat tattttgatt gctggata agcatgtgga aatgaccca accaatcttg cactggcctc ctgatttcct tccttggaga cggagggagg gggagacctg ggggaggggcg cttgggggggg 480 ggtgggctct cttctttctg cgctcccccc ccccacctcc aacaccttga cgacccctcc 540 tgcttccgct tgcctttctc aggctttaac actttctcct cgccctctca gcatgcgcat 600 gcgcgtgcct ctacctcccc cgcacatcct ggcctgccca ccctgaatgg cctggcccag 660 cgatgccacc aactctctcg ctccgtccac ggctggggag gggggcactc tgcagggttg 720 gggggcactg ggaggctggg ttgggtgagg gaggggtgcc tggggcccca ccccccagca 780 agttctctcc ctaggcgaac tggagggtcg tctggcctct tgagccttgt tgctggctct 840 gagctctacc aagagagtga ccagcaggac cgcaccatca cgcgccccag acccgggcct 900 ggggggcaag tcggggggcg gggggaggtc gggcagggtc ccctgggagg atggggacgt 960 gctgtgcccc tagcggccac cagagggcac caggacacca ctgcggtcgg ctcagcggct 1020 cctgccctgg tcaggggcg ccaggtcctg cccctcctgg ggagggcggg gggcgagaag 1080 ggcgattagc ctggtaggct gcagttcatg gggtcactaa gagtcgggca tgactgagcg 1140 acttcacttt catgtatcac tttcatgcat tggagaagga aatggcaacg cactccagtg 1200 ttcttgcctg gagaatccca gggctgggg agcctggtgc actgccatct ctggggtcgc 1260 1320 1380 1440 agtggagaaa tcagatttca agaataact cctttttgca gtccttcaat agaaattgag 1500 cataaatgtg aattagtcat tggcatagac agaaaaatat aatgcatttt gctcagactt 1560 ggtttactgg aaactttaac tggttggatt atgatcaaca tcatgggaat aaaagataca 1620 ttgtagtttc aatataggaa agaaactgaa tcactgaaga agataatttg gatcaagaag 1680 ataagaatct ttgagtaaaa aggagttgtt agtcttaaga aaaaaatttt aacgtttggt 1740 gaaacaaact gaggtcaaga gcaaataaga ttaagaccaa caaatatatt tctcactata 1800 ctgaaggtgc taggtggtta aaataaaatg tgtgatctgg gacaggactg tgtaggtgtg 1860 agtctgcatc tcctctcatt caattcctta actggataag aggaatctaa actgagatgt 1920 caacacagca agcctgctga atttctctga ggtttcatct ttggttgtga acaacaagct 1980 aattagtcca gtcataaagt tagccaatgg catgaaggtg tggtgggtca cacccacact 2040 gagagcatac aaaaggccct ctgcagggag aaatgtccac actcaagtga cacttctact 2100 ctcattctct acccgagaac aacctcaaca agcaacacca tgcaggttca gtgccagcag 2160 agcccagtgc tggcaggcag cgccactttg gtcgcccttg gggcactggc cttgtacgtc 2220 gcgaagccct ccggctacgg gaagcacacg gagagcctga agccggcggc tacccgcctg 2280 ccagcccgcg ccgcctggtt cctgcaggag ctgccttcct tcgcggtgcc cgcggggatc 2340 ctcgcccggc agcccctctc cctcttcggg ccacctggga cggtacttct gggcctctc 2400 tgcctacatt acttccacag gacatttgtg tactcactgc tcaatcgagg gaggccttat 2460 ccagctatac tcattctcag aggcactgcc ttctgcactg gaaatggagt ccttcaaggc 2520 tactatctga tttactgtgc tgaataccct gatgggtggt acacagacat acggtttagc 2580 ttgggtgtct tcttatttat tttgggaatg ggaataaaca ttcatagtga ctatatattg 2640 cgccagctca ggaagcctgg agaaatcagc tacaggattc cacaaggtgg cttgtttacg 2700 tatgtttctg gagccaattt cctcggtgag atcattgaat ggatcggcta tgccctggcc 2760 acttggtccc tcccagcact tgcatttgca ttttctcac tttgtttcct tgggctgcga 2820 gctttcacc accataggtt ctacctcaag atgtttgagg actaccccaa atctcggaaa 2880 gcccttattc cattcatctt ttaaattatc cctaatacct gccaccccac tcttaatcag 2940 tggtggaaga acggtctcag aactgtttgt ttcaattggc catttaagtt tagtagtaaa 3000 agactggtta atgataacaa tgcatcgtaa aaccttcaga aggaaaggag aatgttttgt 3060 ggaccacttt ggttttcttt tttgcgtgtg gcagttttaa gttattagtt tttaaaatca 3120 gtacttttta atggaaaa cttgaccaaa aatttgtcac agaattttga gacccattaa 3180 aaaagttaaa tgagaaacct gtgtgttcct ttggtcaaca ccgagacatt taggtgaaag 3240 acatctaatt ccggttttac gaatctggaa acttcttgaa aatgtaattc ttgagttaac 3300 acttctgggt ggagaatagg gttgttttcc cccacacataa ttggaagggg aaaggaatatc 3360 atttaaagct atgggagggt ttctttgatt acaacactgg agagaaatgc agcatgttgc 3420 tgattgcctg tcactaaac aggccaaaaa ctgagtcctt gggttgcata gaaagctgtt 3480 tccgatcata ttcaataacc cttaatataa cttcgtataa tgtatgctat acgaagttat 3540 taggtctgaa gaggagttta cgtccagcca agctagcttg gctgcaggtc gtcgaaattc 3600 taccgggtag gggaggcgct tttcccaagg cagtctggag catgcgcttt agcagccccg 3660 ctgggcactt ggcgctacac aagtggcctc tggcctcgca cacattccac atccaccggt 3720 aggcgccaac cggctccgtt ctttggtggc cccttcgcgc caccttctac tcctccccta 3780. gtcaggaagt tcccccccgc cccgcagctc gcgtcgtgca ggacgtgaca aatggaagta gcacgtctca ctagtctcgt gcagatggac agcaccgctg agcaatgga gcgggtaggc ctttggggca gcggccaata gcagctttgc tccttcgctt tctgggctca gaggctggga aggggtgggt ccggggggcgg gctcaggggc gggctcaggg gcggggcggg cgcccgaagg 4020 tcctccggag gcccggcatt ctgcacgctt caaaagcgca cgtctgccgc gctgttctcc 4080 tcttcctcat ctccggggcct ttcgacctgc agcctgttga caattaatca tcggcatagt atatcggcat agtatac gacaaggtga ggactaac catgggatcg gccattgac aagatggatt gcacgcaggt tctccggccg cttgggtgga gaggctattc ggctatgact 4260. gggcacaaca gacaatcggc tgctctgatg ccgccgtgtt ccggctgtca gcgcaggggc gcccggttct ttttgtcaag accgacctgt ccggtgccct gaatgaactg caggacgagg cagcgcggct atcgtggctg gccacgacgg gcgttccttg cgcagctgtg ctcgacgttg 4440 tcactgaagc gggaagggac tggctgctat tgggcgaagt gccggggcag gatctcctgt 4500 catctcacct tgctcctgcc gagaaagtat ccatcatggc tgatgcaatg cggcggctgc 4560 atacgcttga tccggctacc tgcccattcg accaccaagc gaaacatcgc atcgagcgag 4620 cacgtactcg gatggaagcc ggtcttgtcg atcaggatga tctggacgaa gagcatcagg 4680 ggctcgcgcc agccgaactg ttcgccaggc tcaaggcgcg catgcccgac ggcgatgatc 4740 tcgtcgtgac ccatggcgat gcctgcttgc cgaatatcat ggtggaaaat ggccgctttt 4800 ctggattcat cgactgtggc cggctgggtg tggcggaccg ctatcaggac atagcgttgg 4860 ctacccgtga tattgctgaa gagcttggcg gcgaatgggc tgaccgcttc ctcgtgcttt 4920 acggtatcgc cgctcccgat tcgcagcgca tcgccttcta tcgccttctt gacgagttct 4980 tctgagggga tcaattctct agagctcgct gatcagcctc gactgtgcct tctagttgcc 5040 agccatctgt tgtttgcccc tcccccgtgc cttccttgac cctggaaggt gccactccca 5100 ctgtcctttc ctaataaaat gaggaaattg catcgcattg tctgagtagg tgtcattcta 5160 ttctgggggg tggggtgggg caggacagca aggggagaga ttgggaagac aatagcaggc 5220 atgctggggga tgcggtgggc tctatggctt ctgaggcgga aagaaccagc tggggctcga 5280 ctagagcttg cggaaccctt aatataactt cgtataatgt atgctatacg aagttattag 5340 gtccctcgag gggatccctc tctccggtct gcaggcattg gcgggtacat gcggatcata 5400 accaccagat ggcgctgttg gcctaagctc gagcacagtc cacagcctgg aggtcctggg 5460 aaagcctgac ctgaatctaa aacttctcta aactccctaa tttgattcaa aagcaaaaca 5520 gtagaaact tctgcacatc ccagcgaggt cagagttagat tggttgctga gactgcgtgg 5580 gggcccaagg agacctggag aaaggaatgc ttcctgctcc ttcttctggg gccccgaggag 5640 agccttccca gggccttgga gaggtgctgt ccagggacta accctgtgct ctaggaaggc 5700 tgcaggccct gaccagctgg gcaggtcctg ggtccctcct ggccttctaa gttcccccaaa 5760 catgagacct ctgggtgtgg ggtggcctgg ggaggtcatt ttgcccaggc cctacctcct 5820 gcccattcct aacccttttt aaaaatctgt gcgtcctctt cttccttctt ctcctccct 5880 tccctttcg ctcaccctct gctgctggcc tgagagccgg aggcccccag ggggaaggcg 5940 actggtctcc tccccagtct cagggaaggg agacagagaa tccaggaagc cagaactcag 6000 cagacgaagc acccagggac ctagagatgg gttgaaaagt tgacagctgt cccacctgcc 6060 tcccaaggtc tcagggccta aacctccaag gcaggaaagg cccctgtccc tccctggggt 6120 ccatagaaag agggacaagt ctgcacggac catttgctgt aatattaaca ccttggctgt 6180 cattaggtag tcttggctgt taattatgtc ctgtgataat gtattattag cacgccgacc 6240 acatagggta gggaactgca gctagtaaac aaaagtttgt tcctatatgc ggccgccata 6300 aaagttttgt tactttatag aagaaatttt gagtttttgt ttttttaat aaataaataa 6360 acataaataa attgtttgtt gaatttatta ttagtatgta agtgtaaata tataaaact 6420 tatatctat tcaaattaat aaataaacct cgatatacag accgataaaa cacatgcgtc 6480 aatttacac atgattatct ttaacgtacg tcacaatatg attatctttc tagggttaat 6540 ctagctgcgt gttctgcagc gtgtcgagca tcttcatctg ctccatcacg ctgtaaaaca 6600 catttgcacc gcgagtctgc ccgtcctcca cgggttcaaa aacgtgaatg aacgaggcgc 6660 gctcactggc cgtcgtttta caacgtcgtg actgggaaaa ccctggcgtt acccaactta 6720 atcgccttgc agcacatccc cctttcgcca gctggcgtaa tagcgaagag gcccgcaccg 6780 atcgcccttc ccaacagttg cgcagcctga atggcgaatg ggacgcgccc tgtagcggcg 6840 cattaagcgc ggcgggtgtg gtggttacgc gcagcgtgac cgctacactt gccagcgccc 6900 tagcgcccgc tcctttcgct ttcttccctt cctttctcgc cacgttcgcc ggctttcccc 6960 gtcaagctct aaatcggggg ctccctttag ggttccgatt tagtgcttta cggcacctcg 7020 accccaaaaa acttgattag ggtgatggtt cacgtagtgg gccatcgccc tgatagacgg 7080 tttttcgccc tttgacgttg gagtccacgt tctttaatag tggactcttg ttccaaactg 7140 gaacaacact caaccctatc tcggtctatt cttttgattt ataagggatt ttgccgattt 7200 cggcctattg gttaaaaaat gagctgattt aacaaaaatt taacgcgaat tttaacaaaa 7260 tattaacgct tacaatttag gtggcacttt tcggggaaat gtgcgcggaa cccctatttg 7320 tttatttttc taaatacatt caatatgta tccgctcatg agacaatac cctgataaat 7380 gcttcaataa tattgaaaaa ggaagagtat gagtattca cattccgtg tcgcccttat 7440 tcccttttt gcggcatttt gccttcctgt tttgctcac ccagaaacgc tggtgaaagt 7500 aaagatgct gaagatcagt tggtgcacg agtgggttac atcgactgg attchcacag 7560 cggtaagatc cttgagagtt ttcgccccga agaacgtttt ccaatgatga gcactttaa 7620 agttctgcta tgtggcgcgg tattatcccg tattgacgcc gggcaagagc aactcggtcg 7680 ccgcatacac tattctcaga atgacttggt tgagtactca ccagtcacag aaagcatct 7740 tacggatggc atgacagtaa gagaattatg cagtgctgcc ataccatga gtgataacac 7800 tgcggccaac ttactctga caacgatcgg aggaccgaag gagctaccg ctttttgca 7860 siacatgggg gatcatgtaa ctcgccttga tcgttggga ccggagctga atgaagccat 7920 accaacgac gagcgtgaca ccacgatgcc tgtagcaatg gcaaacgt tgcgcaact 7980 attaactggc gaactactta ctctagcttc ccggcaacaa ttatagact ggatggaggc 8040 ggataagtt gcaggaccac ttctgcgctc ggcccttccg gctggctggt ttattgctga 8100. taaatctgga gccggtgagc gtggttcacg cggtatcatt gcagcactgg ggccagatgg tagccctcc cgtatcgtag ttatctacac gacggggagt caggcaacta tggatgaacg aatagacag atcgctgaga taggtgcctc acttataag cattggtaac tgtcagacca agtttactca tatatacttt agttgattt aaaacttcat ttttaattta aaaggatcta ggtgaagatc ctttttgata atctcatgac caaaatccct taacgtgagt tttcgttcca ctgagcgtca gaccccgtag aaaagatca aggatcttct tgagatcctt tttttctgcg 8520. cgtaatctgc tgcttgcaaa caaaaaaacc accgctacca gcggtggttt gtttgccgga tcaagagcta ccaactcttt ttccgaaggt aactggcttc agcagagcgc agataccaaa tactgtcctt ctagtgtagc cgtagttagg ccaccacttc aagaactctg tagcaccgcc 8640 tacatacctc gctctgctaa tcctgttacc agtggctgct gccagtggcg ataagtcgtg 8700 tcttaccggg ttggactcaa gacgatagtt accggataag gcgcagcggt cgggctgaac ggggggttcg tgcacacagc ccagcttgga gcgaacgacc tacaccgaac tgagatacct 8820 acagcgtgag ctatgagaaa gcgccacgct tcccgaaggg agaaaggcgg acaggtatcc 8880 ggtaagcggc agggtcggaa caggagagcg cacgagggag cttccagggg gaaacgcctg 8940 gtatctttat agtcctgtcg ggtttcgcca cctctgactt gagcgtcgat ttttgtgatg 9000 ctcgtcaggg gggcggagcc tatggaaaaa cgccagcaac gcggccttt tacggttcct 9060 ggccttttgc tggccttttg ctcacatgtt ctttcctgcg ttatcccctg attctgtgga 9120 taaccgtatt accgcctttg agtgagctga taccgctcgc cgcagccgaa cgaccgagcg 9180 cagcgagtca gtgagcgagg aagcggaaga gcgcccaata cgcaaaccgc ctctccccgc 9240 gcgttggccg attcattaat gcagctggca cgacaggttt cccgactgga aagcgggcag 9300 tgagcgcaac gcaattaatg tgagttagct cactcattag gcaccccagg ctttacactt 9360 tatgcttccg gctcgtatgt tgtgtggaat tgtgagcgga taacaatttc acacaggaaa 9420 cagctatgac catgattacg ccaagcgcgc ccgccgggta actcacgggg tatccatgtc 9480 catttctgcg gcatccagcc aggatacccg tcctcgctga cgtaatatcc cagcgccgca 9540 ccgctgtcat taatctgcac accggcacgg cagttccggc tgtcgccggt attgttcggg9600 ttgctgatgc gcttcgggct gaccatccgg aactgtgtcc ggaaaagccg cgacgaactg gtatcccagg tggcctgaac gaacagttca ccgttaaagg cgtgcatggc cacaccttcc cgaatcatca tggtaaacgt gcgttttcgc tcaacgtcaa tgcagcagca gtcatcctcg gcaaactctt tccatgccgc ttcaacctcg cgggaaaagg cacgggcttc ttcctccccg 9840 atgcccagat agcgccagct tgggcgatga ctgagccgga aaaaagaccc gacgatatga tcctgatgca gctagattaa ccctagaaag atagtctgcg taaaattgac gcatgcattc ttgaaatatt gctctctctt tctaatagc gcgaatccgt cgctgtgcat ttaggacatc tcagtcgccg cttggagctc ccgtgaggcg tgcttgtcaa tgcggtaagt gtcactgatt 10080 ttgaactata acgaccgcgt gagtcaaaat gacgcatgat tatcttttac gtgactttta agatttaact catacgataa ttatattgtt atttcatgtt ctacttacgt gataacttat tatatatata ttttcttgtt atagatatc 10229 <210> 21 <211> 225 <212> DNA <213> Artificial Sequence <400> 21 ggcttgtcgg actcttcgct attacgccag ctggcgaagg gggatgtgct gcaaggcgat 60 taagttgggt aacgccaggg ttttcccagt cacgacgtta ggaaattaat acgactcact 120 ataggctacc aagagagtga ccagcgtttt agagctagaa atagcaagtt aaaataaggc 180 tagtccgtta tcaacttgaa aaagtggcac cgagtcggtg ctttt 225 <210> 22 <211> 225 <212> DNA <213> Artificial Sequence <400> 22 ggcttgtcgg actcttcgct attacgccag ctggcgaagg gggatgtgct gcaaggcgat 60 taagttgggt aacgccaggg ttttcccagt cacgacgtta ggaaattaat acgactcact 120 atagggcagt ctcagcaacc actgagtttt agagctagaa atagcaagtt aaaataaggc 180 tagtccgtta tcaacttgaa aaagtggcac cgagtcggtg ctttt 225 <210> 23 <211> 102 <212> RNA <213> Artificial Sequence <400> 23 ggcuaccaag agagugacca gcguuuuaga gcuagaaaua gcaaguuaaa auaaggcuag 60 uccguuauca acuugaaaaa guggcaccga gucggugcuu uu 102 <210> 24 <211> 102 <212> RNA <213> Artificial Sequence <400> 24 gggcagucuc agcaaccacu gaguuuuaga gcuagaaaua gcaaguuaaa auaaggcuag 60 uccguuauca acuugaaaaa guggcaccga gucggugcuu uu 102 <210> 25 <211> 1089 <212> DNA <213> Sus scrofa <400> 25 actttgtacc tattttgtat gtgtataata atttgagatg tttttaatta ttttgattgc 60 tggaataaag catgtggaaa tgacccaaac caatcttgca ctggcctcct gatttccttc 120 cttggagacg gagggagggg gagacctggg ggagggcgct tggggggggg tgggctctct 180 tctttctgcg ctcccccccc ccacctccaa caccttgacg acccctcctg cttccgcttg 240 cctttctcag gctttaacac tttctcctcg ccctctcagc atgcgcatgc gcgtgcctct 300 acctccccg cacatcctgg cctgcccacc ctgaatgtcc tggcccagcg atgccaccaa 360 ctctctcgct ccgtccacgg ctggggaggg gggcactctg cagggttggg gggcactggg 420 aggctgggtt gggtgaggga ggggtgcctg ggcccccacc ccccagcaag ttctctccct 480 aggcgaactg gagggtcgtc tggcctcttg agccttgttg ctggctctga gctctaccaa 540 gagagtgacc agcaggaccg caccatcagt ggttgctgag actgcgtggg ggcccaagga 600 gacctggaga aaaggaatgct tcctgctcct tcttctgggg cccagaggaga gccttcccag 660 ggccttggag agttgctgtc cagggactaa ccctgtgctc taggaaggct gcaggccctg 720 accagctggg caggtcctgg gtccctcctg gccttctaag ttccccaaac atgagacctc 780 tgggtgtggg gtggcctggg gaggtcattt tgcccaggcc ctacctcctg cccattccta 840 acccttttta aaaatctgtg cgtcctcttc ttccttcttc tccctccctt cccttttcgc 900 tcaccctctg ctgctggcct gagagccgga ggcccccagg gggaaggcga ctggtctcct 960 ccccagtctc agggaaggga gacagagaat ccaggaagcc agaactcagc agacgaagca 1020 cccagggacc tagagatgg ttgaaaagtt gacagctgtc ccacctgcct cccaaggtct 1080 caggccta 1089
Claims
1. A method for preparing recombinant porcine cells that specifically express the human type II 5α-reductase gene in hair follicle tissue, comprising the following steps: introducing a recombinant plasmid, two gRNAs, and an NCN protein into primary porcine fibroblasts, thereby integrating DNA molecule A into the COL1A1 gene of the genomic DNA of the primary porcine fibroblasts, to obtain recombinant porcine cells that specifically express the human type II 5α-reductase gene in hair follicle tissue; the base sequence of the DNA molecule A is shown in positions 881-5559 of SEQ ID NO: 20; the base sequence of the recombinant plasmid is shown in SEQ ID NO: 20; the base sequence of the first gRNA is shown in SEQ ID NO: 23, and the base sequence of the second gRNA is shown in SEQ ID NO: 24; The amino acid sequence of the NCN protein is shown in SEQ ID NO: 3; The method for preparing the NCN protein includes the following steps: (1) Plasmid pKG-GE4 was introduced into Escherichia coli BL21(DE3) to obtain recombinant bacteria; (2) The recombinant bacteria were cultured in liquid culture medium at 30°C, then IPTG was added and induced at 25°C, and then the bacterial cells were collected; (3) The collected bacterial cells were broken down to collect the crude protein solution; (4) The His6-tagged fusion protein was purified from the crude protein solution by affinity chromatography; (5) The His6-tagged fusion protein was digested with His6-tagged enterokinase, and then the His6-tagged protein was removed with Ni-NTA resin to obtain purified NCN protein. The base sequence of the plasmid pKG-GE4 is shown in SEQ ID NO:
1.
2. A kit comprising the recombinant plasmid as described in claim 1, the two gRNAs as described in claim 1, and the NCN protein as described in claim 1; the kit is used to prepare recombinant porcine cells that specifically express the human type II 5α-reductase gene in hair follicle tissue.
3. The use of the recombinant plasmid as described in claim 1, the two gRNAs as described in claim 1, and the NCN protein as described in claim 1 in the preparation kit; the kit is used to prepare recombinant pig cells that specifically express the human type II 5α-reductase gene in hair follicle tissue.
4. The use of the kit according to claim 2 in the preparation of recombinant porcine cells that specifically express the human type II 5α-reductase gene in hair follicle tissue.
Citation Information
Patent Citations
CRISPR / Cas9 system and application thereof in construction of swine-derived recombinant cells with insulin receptor substrate gene defects
CN112522255A
Kit for preparing nuclear transplantation donor cells of alopecia model pig and preparation method of kit
CN115247181A
Kit and application thereof in construction of alopecia model porcine nuclear transplantation donor cells with high expression of porcine II type 5 alpha-reductase
CN115247188A
Construction method of alopecia model porcine nuclear transplantation donor cells for expressing human II type 5 alpha-reductase
CN115247189A