Cytosine deaminase from saccharopolysporum and its base editing system targeting the chicken genome

By developing a cytosine deaminase ch-TU7-CBE derived from saccharopolysporum suitable for chicken cells, a highly efficient and stable chicken genome base editing system was constructed, solving the problem of low editing efficiency of existing systems in chicken cells and realizing efficient C-to-T editing of the chicken genome.

CN121294410BActive Publication Date: 2026-03-06YAZHOUWAN NATIONAL LABORATORY +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511843462.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-03-06
Estimated Expiration
2045-12-09

AI Technical Summary

Technical Problem

Existing base editing systems suffer from low editing efficiency, unstable expression, high toxicity, and a lack of element design optimized for the chicken cell environment in chicken cells, resulting in low efficiency in chicken genome editing and a lack of systematic development and application demonstrations.

Method used

We developed a cytosine deaminase (ch-TU7-CBE) derived from saccharpolysporum and constructed a base editing system based on ch-TU7-CBE. Combining the nucleic acid targeting domain and the cytosine deamination domain, we optimized it into a base editing tool suitable for the chicken genome.

Benefits of technology

This achievement enables efficient C-to-T editing in chicken cells, provides a stable base editing system, and offers key technical support for avian genetic improvement and functional gene research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121294410B_ABST
    Figure CN121294410B_ABST
Patent Text Reader

Abstract

This application provides a polysaccharide polyspora ( Saccharopolyspora The cytosine deaminase derived from [source name] has the amino acid sequence shown in SEQ ID NO:1. This cytosine deaminase is suitable for chicken cells and enables highly efficient C-to-T editing, providing key technical support for avian genetic improvement and functional gene research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of gene editing, specifically to saccharpolysporum (Saccharpolysporum). Saccharopolyspora Cytosine deaminase derived from ) and its base editing system targeting the chicken genome. Background Technology

[0002] Base editing technology is a genome editing tool that enables precise single-base substitution. Common cytosine base editing systems (CBEs) rely on the fusion of cytosine deaminases (such as APOBEC1) with nCas9 (D10A) to achieve C-to-T double-strand break-free mutations in the target region. These systems have been widely established in mammalian cells and model animals such as mice, demonstrating good editing efficiency and promising application prospects.

[0003] However, most of the currently available CBE systems are developed and optimized based on human or mouse cell environments, and they have obvious compatibility problems in non-mammalian cells such as chickens. Specifically, they are as follows: (1) Low editing efficiency: The C-to-T editing efficiency of commonly used cytosine deaminases (such as rAPOBEC1, PmCDA1, etc.) is significantly reduced in chicken embryos, DF-1 and other chicken-derived cells, making it difficult to reach the practical level; (2) Unstable expression or high toxicity: Some cytosine deaminases have low protein expression efficiency, insufficient localization or toxicity in chicken cells, affecting the stability of the editor; (3) Lack of element design optimized for chicken cell environment: The existing CBE systems do not consider factors such as codon optimization and transcriptional regulatory sequence adaptation, which limits their expression and function in chicken cells; (4) Few functional verification cases and lack of standardized tools: There is currently no recognized efficient CBE system suitable for chicken genome editing. Most of the existing studies are in the preliminary trial stage and lack systematic development and application demonstration.

[0004] Therefore, given the current lack of efficient and stable base editing tools in chicken-derived systems, it is imperative to conduct targeted research to develop high-quality base editing tools suitable for chicken-derived systems. Summary of the Invention

[0005] To address the problems existing in the prior art, the purpose of this application is to provide a novel saccharopolysporum (Saccharopolysporum) Saccharopolyspora A cytosine deaminase derived from ) was named ch-TU7-CBE, and a base editing system for the chicken genome based on ch-TU7-CBE was provided.

[0006] Specifically, this application relates to the following aspects:

[0007] 1. A cytosine deaminase having the amino acid sequence shown in SEQ ID NO: 1.

[0008] 2. A base-editing fusion protein comprising a nucleic acid targeting domain and a cytosine deamination domain;

[0009] The cytosine deamination domain comprises at least one cytosine deaminase as described in item 1.

[0010] 3. The base-editing fusion protein according to item 2, wherein the nucleic acid targeting domain is a TALE, ZFP, or CRISPR effector protein domain;

[0011] Preferably, the nucleic acid targeting domain is a CRISPR effector protein domain.

[0012] 4. The base-editing fusion protein according to claim 3, wherein the CRISPR effector protein is at least one of Cas9, Cpf1, Cas3, Cas8a, Cas5, Cas8b, Cas8c, Cas10d, Cse1, Cse2, Csy1, Csy2, Csy3, GSU0054, Cas10, Csm2, Cmr5, Cas10, Csx11, Csx10, Csf1, Csn2, Cas4, C2c1(Cas12b), C2c3, C2c2, Cas12c, Cas12d, Cas12e, Cas12f, Cas12g, Cas12h, Cas12i, Cas12j, Cas12l, Cas12m or other available CRISPR effector proteins;

[0013] Preferably, the CRISPR effector protein is Cas9, which is nuclease-inactivated Cas9, Cas9 nickase, or Cas9 with nuclease activity.

[0014] 5. The base-editing fusion protein according to any one of items 2-4, wherein the base-editing fusion protein further comprises at least one uracil DNA glycosylase inhibitor (UGI).

[0015] 6. The base-editing fusion protein according to any one of items 2-5, wherein the base-editing fusion protein further comprises at least one nuclear localization sequence (NLS).

[0016] 7. A base editing system for modifying the chicken genome, comprising:

[0017] The cytosine deaminase as described in item 1 or the base-editing fusion protein as described in any one of items 2-6; and / or an expression construct containing a nucleotide sequence encoding the cytosine deaminase or the base-editing fusion protein.

[0018] 8. The base editing system according to claim 7, wherein the base editing system further comprises at least one sgRNA and / or at least one expression construct containing a nucleotide sequence encoding said at least one sgRNA;

[0019] The at least one sgRNA targets at least one site in the chicken genome.

[0020] 9. A method for base editing of a chicken genome, comprising contacting the base editing system described in item 7 or 8 with a target site in the chicken genome.

[0021] 10. The method of claim 9, wherein the base editing system contacts the target site to perform deamination, resulting in the substitution of one or more nucleotides in the target site;

[0022] Preferably, the target site comprises a DNA sequence 5'-MCN-3', where M is A, T, C, or G, and N is A, T, C, or G.

[0023] 11. The method according to item 9 or 10, wherein the contact is performed in vivo or in vitro.

[0024] 12. A method for producing at least one modified chicken cell, comprising introducing the base editing system of item 7 or 8 into at least one chicken cell, thereby causing substitution of one or more nucleotides within a target site in the at least one chicken cell.

[0025] 13. A polynucleotide encoding a cytosine deaminase as described in item 1 or a base-editing fusion protein as described in any one of items 2-6.

[0026] 14. A carrier comprising the polynucleotide as described in item 13.

[0027] 15. A host cell comprising the polynucleotide as described in item 13 or the vector as described in item 14.

[0028] 16. A reagent or kit for base editing of the chicken genome, comprising the cytosine deaminase as described in item 1, the base editing fusion protein as described in any one of items 2-6, or the base editing system as described in item 7 or 8.

[0029] 17. A composition comprising the cytosine deaminase as described in item 1, the base editing fusion protein as described in any one of items 2-6, or the base editing system as described in item 7 or 8.

[0030] 18. The composition according to claim 17, wherein the composition is a pharmaceutical composition and the pharmaceutical composition further comprises a pharmaceutically acceptable carrier.

[0031] 19. The use of the cytosine deaminase according to item 1, the base-editing fusion protein according to any one of items 2-6, or the base-editing system according to item 7 or 8 in base editing of the chicken genome.

[0032] 20. The use of the cytosine deaminase according to item 1, the base-editing fusion protein according to any one of items 2-6, or the base-editing system according to item 7 or 8 in the preparation of products for base editing of the chicken genome.

[0033] 21. The use of the cytosine deaminase according to item 1, the base-editing fusion protein according to any one of items 2-6, or the base-editing system according to item 7 or 8 in the preparation of reagents for mediating base editing of the chicken genome.

[0034] 22. The use of the amino acid sequence shown in SEQ ID NO: 1 as a cytosine deaminase.

[0035] 23. The use of the amino acid sequence shown in SEQ ID NO: 1 in the preparation of cytosine deaminase.

[0036] 24. The application of the amino acid sequence shown in SEQ ID NO: 1 as a cytosine deaminase in base editing of the chicken genome.

[0037] Beneficial effects:

[0038] This application successfully screened and obtained a cytosine deaminase derived from *Saccharopolysporum* by screening and sequence functional comparison of cytosine deaminases from various sources, combined with structural prediction and functional verification. Saccharopolyspora A cytosine deaminase (ch-TU7-CBE) suitable for chicken cells and capable of efficient C-to-T editing was developed, and a CBE system with good editing activity and stability was constructed, providing key technical support for avian genetic improvement and functional gene research. Attached Figure Description

[0039] Figure 1 This is a predicted diagram of the ch-TU7-CBE structure.

[0040] Figure 2 The sequencing peak diagram shows the mutation efficiency of sgRNA1 C-to-T in DF-1 cells.

[0041] Figure 3 This is a graph quantifying the mutation efficiency of sgRNA1 C-to-T in DF-1 cells.

[0042] Figure 4The sequencing peak diagram shows the mutation efficiency of sgRNA2 C-to-T in DF-1 cells.

[0043] Figure 5 This is a graph quantifying the mutation efficiency of sgRNA2 C-to-T in DF-1 cells.

[0044] Figure 6 The results quantify the editing efficiency of sgRNA1 electrotransfection using the ch-SBG6-CBE base editor in DF-1 cells.

[0045] Figure 7 The results show the quantification of the editing efficiency of sgRNA2 electroporation using the ch-SBG6-CBE base editor in DF-1 cells.

[0046] Figure 8 This is a quantitative result of the editing efficiency of sgRNA1 electrotransfection using the ch-PTEVA-CBE base editor in DF-1 cells.

[0047] Figure 9 The results quantify the editing efficiency of sgRNA2 electroporation using the ch-PTEVA-CBE base editor in DF-1 cells.

[0048] Figure 10 The results quantify the editing efficiency of sgRNA1 electrotransfection using the BE4max base editor in DF-1 cells.

[0049] Figure 11 The results quantify the editing efficiency of sgRNA2 electrotransfection using the BE4max base editor in DF-1 cells. Detailed Implementation

[0050] The present application is further illustrated below with reference to embodiments. It should be understood that the embodiments are only used to further illustrate and explain the present application and are not intended to limit the present application.

[0051] Unless otherwise defined, technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art. While similar or identical methods and materials may be applied in experimental or practical applications, materials and methods are described herein. In case of conflict, the definitions included herein shall prevail. Furthermore, materials, methods, and examples are for illustrative purposes only and are not intended to be limiting. The present application is further described below with reference to specific embodiments, but is not intended to limit the scope of the application.

[0052] definition

[0053] As used herein, the term "base-editing fusion protein" or "fusion protein" refers to a protein that can mediate the substitution of one or more nucleotides at a target site in the genome in a sequence-specific manner. Such substitutions may be, for example, C-to-T substitutions.

[0054] As used herein, the term "target site" describes a nucleotide sequence, typically a DNA sequence, that can be edited using a base-editing fusion protein as described herein. Typically, a target site is part of the genome.

[0055] As used herein, the term "cytosine deaminase" refers to a deaminase that accepts nucleic acids, such as single-stranded DNA, as a substrate and catalyzes the deamination of cytidine or deoxycytidine to uracil or deoxyuracil, respectively.

[0056] As used herein, a “nucleic acid targeting domain” refers to a domain capable of mediating the attachment of the base-editing fusion protein to a specific target site in the genome in a sequence-specific manner (e.g., via guide RNA). In some embodiments, the nucleic acid targeting domain may include one or more zinc finger protein domains (ZFP) or transcription factor effector domains (TALE) targeting a specific target site. In some embodiments, the nucleic acid targeting domain contains at least one (e.g., one) CRISPR effector protein.

[0057] As used in this article, the term "zinc finger desmin domain (ZFP)" typically contains 3–6 individual zinc finger repeat sequences, each of which can identify a unique sequence of, for example, 3 bp. By combining different zinc finger repeat sequences, different genomic sequences can be targeted.

[0058] As used herein, the term "transcription activator-like effector domain" refers to the DNA-binding domain of a transcription activator-like effector (TALE). TALEs can be engineered to bind to virtually any desired DNA sequence.

[0059] As used herein, the term "CRISPR effector protein" generally refers to a nuclease (CRISPR nuclease) or a functional variant thereof that is present in the naturally occurring CRISPR system. The term encompasses any CRISPR-based effector protein capable of sequence-specific targeting within cells.

[0060] As used herein, a “functional variant” of a CRISPR nuclease means one that retains at least the guide RNA-mediated sequence-specific targeting ability. Preferably, the functional variant is a nuclease-inactivating variant, i.e., lacking double-stranded nucleic acid cleavage activity. However, CRISPR nucleases lacking double-stranded nucleic acid cleavage activity also encompass nickases, which form a nick in a double-stranded nucleic acid molecule but do not completely cleave the double-stranded nucleic acid. In some preferred embodiments of this application, the CRISPR effector protein described herein possesses nickase activity. In some embodiments, the functional variant recognizes a different PAM (pre-intermediate sequence adjacent motif) sequence relative to the wild-type nuclease.

[0061] "CRISPR effector proteins" can be derived from Cas9 nucleases, including Cas9 nucleases or functional variants thereof. The Cas9 nuclease can be from different species, such as those from Streptococcus pyogenes (Streptococcus pyogenes). S.pyogenes spCas9 or derived from Staphylococcus aureus ( S.aureus ) of SaCas9.

[0062] "CRISPR effector proteins" can also be derived from Cpf1 nucleases, including Cpf1 nucleases or functional variants thereof. The Cpf1 nuclease can be a Cpf1 nuclease from a different species, for example, from... Francisella novicida U112、 Acidaminococcus sp.BV3L6 and Lachnospiraceae bacterium The Cpf1 nuclease of ND2006. Available "CRISPR effector proteins" can also be derived from nucleases such as Cas3, Cas8a, Cas5, Cas8b, Cas8c, Cas10d, Cse1, Cse2, Csy1, Csy2, Csy3, GSU0054, Cas10, Csm2, Cmr5, Cas10, Csx11, Csx10, Csf1, Csn2, Cas4, C2c1 (Cas12b), C2c3, C2c2, Cas12c, Cas12d (i.e., CasY), Cas12e (i.e., CasX), Cas12f (i.e., Cas14), Cas12g, Cas12h, Cas12i, Cas12j (i.e., CasΦ), Cas12k, Cas12l, and Cas12m, including, for example, these nucleases or their functional variants.

[0063] As used herein, the terms “Cas9” or “Cas9 nuclease” or “Cas9 domain” refer to CRISPR-associated protein 9 or variants thereof, including any native Cas9 from any organism, any native Cas9 equivalent or fragment thereof, any Cas9 homolog, ortholog, or paralog from any organism, and any variant of any naturally occurring or engineered Cas9. The term Cas9 is not limited to any specific Cas9 and may be referred to as “Cas9 or variants thereof.” Exemplary Cas9 proteins are described herein.

[0064] As used herein, the term "dCas9" refers to Cas9 with nuclease inactivation or death, or a variant thereof, and includes any naturally occurring dCas9 from any organism, any naturally occurring dCas9 equivalent or functional fragment thereof, any dCas9 homolog, ortholog, or paralog from any organism, and any naturally occurring or engineered dCas9 variant. Exemplary dCas9 proteins will be described herein. Any suitable mutation that inactivates the Cas9 nuclease can be used to form dCas9, such as wild-type Streptococcus pyogenes (Streptococcus pyogenes). S.pyogenes Mutations in the D10A and H840A amino acid sequence of Cas9, or in wild-type Staphylococcus aureus ( S.aureus The D10A and N580A mutations in the Cas9 amino acid sequence.

[0065] As used herein, the terms “nCas9,” “Cas9 cleavage enzyme,” or “Cas9 cleavage enzyme” refer to Cas9 or a variant thereof that cleaves only one strand, thereby introducing a gap in the double-stranded DNA molecule rather than creating a double-strand break. This can be achieved by introducing appropriate mutations into wild-type Cas9, such as those in wild-type Streptococcus pyogenes (Streptococcus pyogenes). S.pyogenes The D10A or H840A mutation in the Cas9 amino acid sequence, or wild-type Staphylococcus aureus ( S.aureus The D10A mutation in the Cas9 amino acid sequence.

[0066] As used herein, the term "base editing system" refers to a combination of components required for base editing of nucleic acid sequences, such as genomic sequences in cells or organisms. The individual components of such a system, such as cytosine deaminases, base editing fusion proteins, and one or more sgRNAs, may exist independently or in any combination as a composition.

[0067] As used herein, the term "connector" refers to any means, entity, or portion used to join two or more entities. In some embodiments, the connector is a covalent connector. In some embodiments, the connector is a non-covalent connector. Examples of covalent connectors include connector portions covalently bonded or covalently attached to one or more proteins or domains to be joined. In some embodiments, the connector is non-covalent, such as an organometallic bond through a metal center (such as a platinum atom). Joining can be permanent or reversible. For covalent linkages, various functional groups can be used, such as amide groups, including carbonate derivatives, ethers, esters (including organic and inorganic esters), amino groups, carbamates, ureas, etc. To provide a connection, the domain can be modified by oxidation, hydroxylation, substitution, reduction, etc., to provide a coupling site. Conjugation methods are well known to those skilled in the art and are covered herein for use. Connector portions include, but are not limited to, chemical connector portions, or, for example, peptide connector portions (connector sequences). The length and type of connector can be designed as needed. In some embodiments, the connector can be selected from artificially synthesized amino acid sequences or naturally occurring polypeptide sequences. It should be understood that modifications that do not significantly reduce the function of RNA-binding domains and effector domains are preferred.

[0068] In this application, the terms "sgRNA," "guide RNA," and "CRISPR guide sequence" are used interchangeably throughout and refer to nucleic acids containing sequences that determine the specificity of CRISPR effector proteins. The sgRNA hybridizes (partially or completely complementary) to a target site in the host cell genome. The length of the sgRNA or a portion thereof hybridizing to the target site can be between 15-25 nucleotides, 18-22 nucleotides, or 19-21 nucleotides. In some embodiments, the length of the sgRNA sequence hybridizing to the target site can be 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides. In some embodiments, the length of the sgRNA sequence hybridizing to the target site is between 10-30 or 15-25 nucleotides. It should be understood that the sgRNA can be an RNA sequence or a DNA sequence corresponding to an RNA sequence.

[0069] As used in this article, the term "genome" encompasses not only chromosomal DNA, which is present in the cell nucleus, but also organelle DNA, which is present in subcellular components of the cell, such as mitochondria and plastids.

[0070] As used herein, the terms “nucleic acid,” “nucleic acid sequence,” “nucleotide sequence,” “polynucleotide,” “polynucleotide sequence,” “RNA sequence,” or “DNA sequence” refer to oligonucleotides, nucleotides, or polynucleotides, and fragments or portions thereof, and refer to DNA or RNA of genotype or synthetic origin, which may be single-stranded or double-stranded and represent sense or antisense strands. Sequences may be non-coding sequences, coding sequences, or mixtures of both. The nucleic acid sequences of this application can be prepared using standard techniques well known to those skilled in the art. Nucleotides are designated by their individual letter names as follows: “A” for adenosine or deoxyadenosine (corresponding to RNA or DNA, respectively), “C” for cytidine or deoxycytidine, “G” for guanosine or deoxyguanosine, “U” for uridine, “T” for deoxythymidine, “R” for purine (A or G), “Y” for pyrimidine (C or T), “K” for G or T, “H” for A, C, or T, “I” for inosine, and “N” for any nucleotide.

[0071] As used herein, the term “amino acid” or “amino acid sequence” refers to an oligopeptide, peptide, polypeptide, or protein sequence, or any fragment thereof, and refers to a naturally occurring or synthetic molecule. When “amino acid sequence” is described herein as referring to the amino acid sequence of a naturally occurring protein molecule, “amino acid sequence” and similar terms are not intended to limit the amino acid sequence to the complete naturally occurring amino acid sequence associated with the described protein molecule.

[0072] In this article, "amino acid" may be referred to by its name, its commonly known three-letter symbol, or a single-letter symbol recommended by the IUPAC-IUB Biochemical Nomenclature Commission.

[0073] As used herein, the percentage of “identity,” such as 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, or 99.5% identity, refers to the degree of similarity between amino acid sequences or nucleotide sequences determined by sequence alignment, which is 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, or 99.5%. For example, it is the percentage of positions with identical bases or amino acid residues determined after two sequences have as many identical residues as possible by introducing vacancies, etc. The percentage of “identity” can be determined using software programs known in the art. It is preferred to use default parameters for alignment. A preferred alignment program is BLAST. Preferred programs are BLASTN and BLASTP. Details of these programs can be found at the following internet address: blast.ncbi.nlm.nih.gov / Blast.cgi.

[0074] As used herein, the term "nuclear localization signal or sequence (NLS)" is an amino acid sequence that marks, designates, or otherwise tags a protein to facilitate its transport into the cell nucleus via nuclear transport. Typically, this signal consists of one or more short sequences of positively charged lysine or arginine residues exposed on the protein's surface. Different nuclear localization proteins may share the same NLS. NLS function as the opposite of nuclear export signals (NES), which target proteins outside the cell nucleus. Thus, a single nuclear localization signal can guide the entity it is associated with into the cell nucleus. Such sequences can be of any size and composition, for example, exceeding 25, 25, 15, 12, 10, 8, 7, 6, 5, or 4 amino acids.

[0075] As used herein, the term "uracil glycosylation inhibitor" or "UGI" refers to a protein that inhibits the base excision repair enzyme of uracil DNA glycosylation.

[0076] As used herein, the term "expression construct" refers to a vector, such as a recombinant vector, suitable for expressing a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, the expression of a nucleotide sequence can refer to the transcription of the nucleotide sequence (e.g., transcription to generate mRNA or functional RNA) and / or the translation of RNA into a precursor or mature protein. The "expression construct" of this application can be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, a translatable RNA (e.g., mRNA). The "expression construct" of this application can contain regulatory sequences and nucleotide sequences of interest from different sources, or regulatory sequences and nucleotide sequences of interest from the same source but arranged in a manner different from those typically found naturally.

[0077] As used herein, the terms “regulatory sequence” and “regulatory element” are used interchangeably and refer to a nucleotide sequence located upstream (5' non-coding sequence), midway, or downstream (3' non-coding sequence) of a coding sequence that affects the transcription, RNA processing, or stability or translation of the relevant coding sequence. Regulatory sequences may include, but are not limited to, promoters, translational leader sequences, introns, and polyadenylation recognition sequences.

[0078] As used herein, the term "operably linked" refers to the linking of a regulatory element (e.g., but not limited to, promoter sequences, transcription termination sequences, etc.) to a nucleic acid sequence (e.g., coding sequences or open reading frames) such that transcription of the nucleotide sequence is controlled and regulated by the transcriptional regulatory element. Techniques for operably linking regulatory element regions to nucleic acid molecules are known in the art.

[0079] As used herein, the term "vector" refers to a nucleic acid molecule capable of amplifying another nucleic acid linked to it. This term includes vectors as self-replicating nucleic acid structures as well as vectors integrated into the genome of a host cell that has already been introduced therein. Some vectors are capable of directing the expression of the nucleic acid to which they are operatively linked. Such vectors are referred to herein as "expression vectors."

[0080] As used herein, the term "host cell" includes, but is not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. This term includes the progeny of the original cell into which a foreign nucleic acid fragment has been introduced. An exemplary host cell includes the human embryonic kidney cell HEK293T. It should be understood that, due to natural, accidental, or intentional mutations, the progeny of a single-parent cell may not necessarily be identical to the original parent in terms of morphology or in terms of genome or total DNA complementarity.

[0081] The term "introduction" generally refers to the transfer of a foreign gene into recipient cells, such as eukaryotic or prokaryotic recipient cells. There are no particular restrictions on the method of introduction; any known transformation method that can transfer the target gene into the recipient cell is acceptable. The methods of introduction may include any of the following: (1) introducing the target gene or a recombinant vector containing the target gene into the host bacteria via chemical transformation (such as Ca ion-induced transformation, polyethylene glycol-mediated transformation, or metal cation-mediated transformation) or physical transformation (such as electroporation). (2) transducing the target gene into the host bacteria via bacteriophage transduction. (3) transferring the target gene into plant recipient cells via physical or chemical methods, such as gene gun method (also known as microparticle bombardment or biological missile method), chemical stimulation method, electroporation method, liposome-mediated method, microinjection method, laser microbeam method, pollen tube pathway method, ultrasound method, air gun method, and eddy current method. (4) Using vectors to transfer the target gene into plant recipient cells, such as Agrobacterium Ti plasmid vector (including Ti plasmid-derived vectors such as co-integration vector system and binary vector system) mediated method (Agrobacterium-mediated method), plant virus vector mediated transformation method, etc.

[0082] As used herein, the term “in vitro” refers to experiments conducted in laboratory conditions or culture media using materials, biological substances, cells and / or tissues; while the term “in vivo” refers to experiments and procedures conducted using intact multicellular organisms.

[0083] As used herein, the term "about" or "approximately," when applied to one or more values ​​of interest, refers to a value similar to a specified reference value, or a value within an acceptable margin of error for a particular value as determined by a person skilled in the art, which will depend in part on how the value is measured or determined, such as limitations of the measurement system. In some respects, the term "about" refers to a range of values ​​falling within (greater or less than) 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less of a specified reference value in any direction, unless otherwise stated or apparent from the context (unless the number exceeds 100% of the possible value). Alternatively, according to practice in the art, "about" may mean within 3 or greater than 3 standard deviations. Or, for example, with respect to biological systems or processes, the term "about" may mean within an order of magnitude of the value, preferably within 5 times, more preferably within 2 times.

[0084] Cytosine deaminase

[0085] In a first aspect, this application provides a cytosine deaminase, wherein the cytosine deaminase is capable of deaminating the cytosine base of deoxycytidine in DNA.

[0086] The cytosine deaminase of this application may be a recombinant protein, a natural protein, or a synthetic protein, preferably a recombinant protein; the protein may be a naturally purified product, a chemically synthesized product, or a product produced from a prokaryotic / eukaryotic host (e.g., bacteria, yeast, higher plants, insects, or mammalian cells) using recombinant technology.

[0087] In some embodiments, the amino acid sequence of the cytosine deaminase is shown in SEQ ID NO: 1, and the cytosine deaminase is derived from *Saccharopolysporum* (Saccharopolysporum). Saccharopolyspora ), named ch-TU7-CBE.

[0088] The amino acid sequence shown in SEQ ID NO: 1 is as follows:

[0089] MGDDVAAVAARVREAMAKLPPEAFQIAGECIDEAGAGLHPLALETNDAELAAVIGALGDAREEIDRAWQICRKVRDACADYLKIIGAAEPSAPATSAGAAGGRARHVTAKDGSQCPPEAGAVVDVLP RRVREGEAGEKTVGFVDGSVTDKFVSGRDQTWTRSILARAREVGLPPHLARFVSSHVEMKVAAMMTQTGKQHCELVINHVPCGSQPAQPPGCDQAIERFLPKGYTLTVHGTTQASRPFSKTYRGQA.

[0090] The cytosine deaminase of this application also comprises fragments, derivatives, and analogs of the ch-TU7-CBE. As used herein, the terms "fragment," "derivative," and "analyte" refer to proteins that substantially retain the same biological function or activity as the ch-TU7-CBE. The protein fragments, derivatives, or analogs of this application may be proteins in which one or more conserved or non-conserved amino acid residues (preferably conserved amino acid residues) are substituted, and such substituted amino acid residues may or may not be encoded by the genetic code, or may be proteins having substituent groups in one or more amino acid residues, or proteins formed by fusing additional amino acid sequences to this protein sequence (such as leader sequences or secretory sequences, sequences used to purify this protein, or proteoprotein sequences, or fusion proteins). As defined in this application, these fragments, derivatives, and analogs are within the scope well known to those skilled in the art.

[0091] The cytosine deaminase of this application further comprises, compared with the ch-TU7-CBE, a number of (typically 1-20, more preferably 1-10, even more preferably 1-8, 1-5, 1-3, or 1-2) amino acid deletions, insertions, and / or substitutions, and / or the addition or deletion of one or more (typically up to 20, preferably up to 10, more preferably up to 5) amino acids at the C-terminus and / or N-terminus. For example, substitution with amino acids of similar or comparable properties generally does not alter the function of the protein. Similarly, the addition of one or more amino acids at the C-terminus and / or N-terminus generally does not alter the function of the protein.

[0092] In some embodiments, the cytosine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with the amino acid sequence shown in SEQ ID NO: 1.

[0093] Base editing fusion protein

[0094] Secondly, this application provides a base editing fusion protein comprising a nucleic acid targeting domain and a cytosine deamination domain, wherein the cytosine deamination domain comprises at least one (e.g., one or two) of the cytosine deaminase described in the first aspect of this application.

[0095] In some embodiments, the nucleic acid targeting domain is a TALE, ZFP, or CRISPR effector protein domain.

[0096] In some embodiments, the nucleic acid targeting domain is a CRISPR effector protein domain. In some embodiments, the CRISPR effector protein is at least one of Cas9, Cpf1, Cas3, Cas8a, Cas5, Cas8b, Cas8c, Cas10d, Cse1, Cse2, Csy1, Csy2, Csy3, GSU0054, Cas10, Csm2, Cmr5, Cas10, Csx11, Csx10, Csf1, Csn2, Cas4, C2c1(Cas12b), C2c3, C2c2, Cas12c, Cas12d, Cas12e, Cas12f, Cas12g, Cas12h, Cas12i, Cas12j, Cas12l, Cas12m, or other available CRISPR effector proteins. In some embodiments, the CRISPR effector protein is Cas9, which is nuclease-inactivated Cas9, Cas9 nickase, or Cas9 with nuclease activity.

[0097] In some embodiments, the Cas9 is a nuclease-inactivated Cas9. The DNA cleavage domain of Cas9 is known to comprise two subdomains: the HNH nuclease subdomain and the RuvC subdomain. The HNH subdomain cleaves the strand complementary to the sgRNA, while the RuvC subdomain cleaves the non-complementary strand. Mutations in these subdomains can inactivate the nuclease activity of Cas9, thus forming a "nuclease-inactivated Cas9".

[0098] In some embodiments, the Cas9 is a Cas9 cleavage enzyme. Mutating and inactivating a subdomain of the DNA cleavage domain can impart cleavage enzyme activity to Cas9, thus obtaining a Cas9 cleavage enzyme. In some embodiments, the Cas9 cleavage enzyme is a Cas9 (D10A) cleavage enzyme, the amino acid sequence of which is shown in SEQ ID NO: 12.

[0099] The amino acid sequence shown in SEQ ID NO: 12 is as follows:

[0100]

[0101] In some embodiments, the base-editing fusion protein comprises, from the N-terminus to the C-terminus, a cytosine deamination domain and a nucleic acid targeting domain in the following order.

[0102] In some embodiments, the nucleic acid targeting domain and the cytosine deamination domain are fused via a linker.

[0103] The “linker” can be a non-functional amino acid sequence of 1 to 50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or 20-25, 25-50) or more amino acids, without secondary or higher structures.

[0104] In cells, uracil DNA glycosyltransferase catalyzes the removal of U from DNA and initiates base excision repair (BER), resulting in the repair of U:G to C:G. Therefore, without any theoretical limitations, the base editing fusion protein of this application, combined with a uracil DNA glycosyltransferase inhibitor (UGI), will be able to increase the efficiency of C-to-T base editing.

[0105] In some embodiments, the base-editing fusion protein further comprises at least one UGI (e.g., one or two). In some embodiments, the UGI is connected to other portions of the base-editing fusion protein via a connector. In some embodiments, the UGI is located at the N-terminus or C-terminus of the base-editing fusion protein, preferably the C-terminus.

[0106] In some embodiments, the base-editing fusion protein further comprises at least one nuclear localization sequence (NLS) (e.g., one or two). In some embodiments, the NLS is located at the N-terminus and / or C-terminus of the base-editing fusion protein. In some embodiments, the NLS is located between the cytosine deamination domain, the nucleic acid targeting domain, and / or the UGI. In some embodiments, the base-editing fusion protein contains about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLS at or near its C-terminus. In some embodiments, the base-editing fusion protein contains about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLS at or near its N-terminus. When more than one NLS is present, each NLS can be selected independently of the other NLS.

[0107] Furthermore, depending on the location of the DNA to be edited, the base-editing fusion protein of this application may also contain other localization sequences, such as cytoplasmic localization sequences and mitochondrial localization sequences.

[0108] In some embodiments, the amino acid sequence of the base-editing fusion protein is shown in SEQ ID NO: 3.

[0109] The sequence represented by SEQ ID NO: 3 is as follows:

[0110]

[0111] In some embodiments, the base-editing fusion protein further comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with the amino acid sequence shown in SEQ ID NO: 3.

[0112] Base editing system for modifying the chicken genome

[0113] Thirdly, this application provides a base editing system for modifying the chicken genome, comprising: the cytosine deaminase described in the first aspect of this application or the base editing fusion protein described in the second aspect of this application, and / or an expression construct containing a nucleotide sequence encoding the cytosine deaminase or the base editing fusion protein.

[0114] In some embodiments, the base editing system further comprises: at least one sgRNA and / or at least one expression construct containing a nucleotide sequence encoding said at least one sgRNA. Those skilled in the art will appreciate that if the base editing fusion protein is not based on a CRISPR effector protein, the base editing system may not require sgRNA or an expression construct encoding it.

[0115] In some embodiments, the at least one sgRNA can bind to the nucleic acid targeting domain of the base editing fusion protein, and the at least one sgRNA targets at least one target site in the chicken genome.

[0116] In some embodiments, the sgRNA is at least 5 nucleotides long. In some embodiments, the sgRNA is about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75 or more nucleotides long. In some embodiments, the sgRNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 nucleotides long.

[0117] In some embodiments, in the base editing system, the base editing fusion protein can form a complex with the sgRNA, and the complex specifically targets the target site under the mediation of the sgRNA, resulting in the substitution of one or more C-to-T at the target site.

[0118] In some embodiments, in the base editing system, the base editing fusion protein can form a complex with the sgRNA that specifically targets a target site under the mediation of the sgRNA, resulting in one or more C-to-T substitutions on a non-complementary sequence (a sequence not complementary to the sgRNA) and one or more G-to-A substitutions on a complementary sequence (a sequence complementary to the sgRNA).

[0119] Methods for base editing of the chicken genome

[0120] Fourthly, this application provides a method for base editing of the chicken genome, which includes contacting the base editing system described in the third aspect of this application with a target site in the chicken genome.

[0121] In this application, the target site to be modified can be located anywhere in the chicken genome, such as within a functional gene like a protein-coding gene, or in a gene expression regulatory region such as a promoter region or an enhancer region, thereby achieving modification of the gene function or modification of gene expression.

[0122] In some embodiments, the base editing system contacts the target site to perform deamination, resulting in the substitution of one or more nucleotides in the target site.

[0123] In some embodiments, the target site comprises the DNA sequence 5'-MCN-3', where M is A, T, C, or G, and N is A, T, C, or G; wherein the C in the middle of the 5'-MCN-3' sequence is deaminated.

[0124] In some embodiments, the deamination leads to the introduction of a mutation in the gene promoter, which results in an increase or decrease in transcription of a gene operatively linked to the gene promoter.

[0125] In some embodiments, the deamination process results in the introduction or removal of splice sites.

[0126] In some embodiments, the deamination leads to the introduction of a mutation in the gene repressor, which results in an increase or decrease in the transcription of a gene operatively linked to the gene repressor.

[0127] In some embodiments, the contact occurs within the body. In other embodiments, the contact occurs outside the body.

[0128] Methods for producing modified chicken cells

[0129] Fifthly, this application provides a method for generating at least one modified chicken cell, comprising introducing the base editing system described in the third aspect of this application into at least one chicken cell, thereby causing substitution of one or more nucleotides within a target site in the at least one chicken cell. In some embodiments, the one or more nucleotide substitutions are C-to-T substitutions.

[0130] In this application, the base editing system can be introduced into chicken cells using various methods well known to those skilled in the art. Methods for introducing the base editing system of this application into chicken cells include, but are not limited to: calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection, gene gun method, PEG-mediated protoplast transformation, and Agrobacterium-mediated transformation.

[0131] In some embodiments, the method further includes the step of screening chicken cells from the at least one chicken cell for chicken cells having one or more desired nucleotide substitutions.

[0132] In some embodiments, the method is performed in vitro. For example, the chicken cells are isolated chicken cells or chicken cells in isolated tissues or organs.

[0133] Polynucleotides, vectors and host cells

[0134] In a sixth aspect, this application provides a polynucleotide that encodes the cytosine deaminase described in the first aspect of this application or the base-editing fusion protein described in the second aspect of this application.

[0135] The polynucleotides encompass both DNA (gDNA and cDNA) and RNA molecules, as well as nucleotides, which are the basic building blocks of polynucleotides, including naturally derived nucleotides and analogs with modified sugar or base portions. The polynucleotide encoding the cytosine deaminase or base fusion protein of this application can be modified. Such modifications include the addition, deletion, or non-conserved or conserved substitution of nucleotides. The polynucleotide encoding the cytosine deaminase or base fusion protein of this application can be synthesized using conventional procedures.

[0136] In some implementations, the polynucleotide is codon-optimized, for example, for better expression in chicken cells.

[0137] In a seventh aspect, this application provides a vector comprising the nucleic acid molecule described in the sixth aspect of this application. The vector can be designed to clone and / or express the base-editing fusion protein of this application. The vector can be designed to introduce the base-editing fusion protein of this application into one or more chicken cells. The vector can be designed to express base-editing transcripts (e.g., nucleic acid transcripts, proteins, or enzymes) in prokaryotic or eukaryotic cells. For example, base-editing transcripts can be expressed in chicken cells.

[0138] In some embodiments, the vector includes cloning vectors, expression vectors, shuttle vectors, and integration vectors. In some embodiments, the vector is a viral vector (e.g., a retroviral vector, lentiviral vector, adenovirus vector, adeno-associated vector, and herpes simplex vector), and may also be a plasmid, granulosome, bacteriophage, or other type known to those skilled in the art.

[0139] Eighthly, this application provides a host cell that contains the nucleic acid molecule described in the sixth aspect of this application or the vector described in the seventh aspect of this application.

[0140] In some embodiments, the cell can be any suitable type of cell, such as a prokaryotic or eukaryotic cell, including, but not limited to, bacterial cells, fungal cells, plant cells, insect cells, or mammalian cells.

[0141] Base Editor

[0142] Ninthly, this application provides a base editor, which is constructed by integrating a multinucleotide sequence encoding the base-editing fusion protein described in the second aspect of this application into an expression vector. Other components of the base editor are known in the art.

[0143] Reagent test kit

[0144] In a tenth aspect, this application provides a reagent or kit for base editing of the chicken genome, comprising the cytosine deaminase described in the first aspect of this application, the base editing fusion protein described in the second aspect of this application, the base editing system described in the third aspect of this application, the polynucleotide described in the sixth aspect of this application, the vector described in the seventh aspect of this application, or the base editor described in the ninth aspect of this application.

[0145] Kits generally include a label indicating the intended use and / or method of use of the kit contents. The term "label" includes any written or documented material provided on or with the kit or otherwise accompanied by the kit. The kit may also contain suitable materials for constructing expression constructs in the base editing system of this application. The kit may also contain reagents suitable for introducing the base editing fusion protein of this application into cells.

[0146] Composition

[0147] In one aspect, this application provides a composition comprising the cytosine deaminase described in the first aspect of this application, the fusion protein described in the second aspect of this application, the base editing system described in the third aspect of this application, the polynucleotide described in the sixth aspect of this application, the vector described in the seventh aspect of this application, or the base editor described in the ninth aspect of this application.

[0148] In some embodiments, the composition is a pharmaceutical composition, which further comprises a pharmaceutically acceptable carrier.

[0149] Pharmaceutically acceptable carriers may include non-toxic buffers such as phosphoric acid, citric acid, and other organic acids; salts such as sodium chloride; antioxidants, including ascorbic acid and methionine; preservatives (e.g., octadecyl dimethyl benzyl ammonium chloride; hexamethonium chloride; benzalkonium chloride; benzyl chloride; phenol, butyl or benzyl alcohol; alkyl p-hydroxybenzoates). Parabens, such as methyl or propylparaben; catechol; resorcinol; cyclohexanol; 3-pentanol; and m-cresol; low molecular weight peptides (e.g., less than about 10 amino acid residues); proteins such as serum albumin, gelatin, or immunoglobulins; hydrophilic polymers such as polyvinylpyrrolidone; amino acids such as glycine, glutamine, asparagine, histidine, arginine, or lysine; carbohydrates such as monosaccharides, disaccharides, glucose, mannose, or dextrin; chelating agents such as EDTA; sugars such as sucrose, mannitol, trehalose, or sorbitol; salt-forming counterions such as sodium; metal complexes (e.g., Zn-protein complexes); and nonionic surfactants such as Tween or polyethylene glycol (PEG), etc.

[0150] application

[0151] In a twelfth aspect, this application provides the use of the cytosine deaminase described in the first aspect of this application, the base editing fusion protein described in the second aspect of this application, the base editing system described in the third aspect of this application, the polynucleotide described in the sixth aspect of this application, the vector described in the seventh aspect of this application, the base editor described in the ninth aspect of this application, the kit described in the tenth aspect of this application, or the composition described in the eleventh aspect of this application in base editing of the chicken genome.

[0152] Thirteenthly, this application provides the use of the cytosine deaminase described in the first aspect of this application, the base editing fusion protein described in the second aspect of this application, the base editing system described in the third aspect of this application, the polynucleotide described in the sixth aspect of this application, the vector described in the seventh aspect of this application, the base editor described in the ninth aspect of this application, the kit described in the tenth aspect of this application, or the composition described in the eleventh aspect of this application in the preparation of a product for base editing of the chicken genome.

[0153] In a fourteenth aspect, this application provides the use of the cytosine deaminase described in the first aspect of this application, the base editing fusion protein described in the second aspect of this application, the base editing system described in the third aspect of this application, the polynucleotide described in the sixth aspect of this application, the vector described in the seventh aspect of this application, the base editor described in the ninth aspect of this application, the kit described in the tenth aspect of this application, or the composition described in the eleventh aspect of this application in the preparation of reagents for mediating base editing of the chicken genome, including for reducing off-target effects, improving targeted editing efficiency, or improving the fidelity of targeted editing.

[0154] In a fifteenth aspect, this application provides the use of the cytosine deaminase described in the first aspect of this application, the base editing fusion protein described in the second aspect of this application, the base editing system described in the third aspect of this application, the polynucleotide described in the sixth aspect of this application, the vector described in the seventh aspect of this application, the base editor described in the ninth aspect of this application, the kit described in the tenth aspect of this application, or the composition described in the eleventh aspect of this application in the treatment of diseases in chickens.

[0155] The disease refers to a diagnosis of a point mutation-related or point mutation-caused disease, which can be corrected by the base editing fusion protein / base editing system of this application, for example, it can be used to correct any single-point T-to-C or A-to-G mutation.

[0156] In a sixteenth aspect, this application provides the use of the cytosine deaminase described in the first aspect of this application, the base editing fusion protein described in the second aspect of this application, the base editing system described in the third aspect of this application, the polynucleotide described in the sixth aspect of this application, the vector described in the seventh aspect of this application, the base editor described in the ninth aspect of this application, or the kit described in the tenth aspect of this application in the preparation of a medicament for treating chicken diseases.

[0157] The disease referred to here means a disease diagnosed as being related to or caused by point mutations.

[0158] In a seventeenth aspect, this application provides the use of the amino acid sequence shown in SEQ ID NO: 1 as a cytosine deaminase.

[0159] In an eighteenth aspect, this application provides the use of the amino acid sequence shown in SEQ ID NO: 1 in the preparation of cytosine deaminase.

[0160] Nineteenthly, this application provides the amino acid sequence shown in SEQ ID NO: 1 as an application of cytosine deaminase in base editing of the chicken genome.

[0161] Example

[0162] The following description, in conjunction with specific embodiments, illustrates the content of this application, but the scope of this application is not limited thereto. Unless otherwise specified, the reagents and instruments used in the following embodiments are all conventional reagents and instruments in the art and can be obtained commercially. The methods used are all conventional experimental methods, and those skilled in the art can undoubtedly implement the described schemes and obtain corresponding results based on the embodiments.

[0163] Example 1: Screening of candidate cytosine deaminases

[0164] In this embodiment, the inventors performed bioinformatics analysis based on existing single-base DNA deaminase (sddA) data and screened out several candidate cytosine deaminase sequences with potentially high activity.

[0165] First, the inventors extensively collected and integrated datasets from published literature, containing sequence information and related biological characteristics of multiple known cytosine deaminases. Based on these datasets, the inventors conducted systematic evolutionary analysis to reveal the similarities and differences between different cytosine deaminase sequences, and further screened cytosine deaminases with high catalytic activity as reference benchmarks. Building on this, the inventors conducted sequence mining using protein databases, focusing on novel sequences with a similarity of less than 70%. The core of this strategy is to ensure that the selected sequences have significant structural differences, avoiding overly similar sequences to known cytosine deaminases, which could potentially lead to the omission of highly active cytosine deaminases. To ensure the quality and diversity of the sequences, novelty and functional potential were particularly emphasized, with sequences exhibiting significant differences being the focus of subsequent analysis.

[0166] Subsequently, the inventors used the selected novel sequences to construct an evolutionary tree and further explored their similarities and differences in evolutionary relationships through cluster analysis. Through in-depth analysis of the clustering results, redundant sequences were removed to reduce the interference of repetitive information on the final screening results. To improve the quality of candidate cytosine deaminases, the remaining sequences were further screened and functionally predicted after redundancy removal, ultimately identifying several potentially highly active cytosine deaminase sequences as candidates for subsequent experimental verification. Among them, the inventors discovered sequences originating from *Saccharopolysporum* (…). Saccharopolyspora SddA (named ch-TU7-CBE) exhibits higher activity compared to other candidate cytosine deaminases.

[0167] Example 2: Construction of the ch-TU7-CBE base editor

[0168] The amino acid sequence of ch-TU7-CBE was confirmed by amino acid sequence alignment, and its predicted structure is as follows: Figure 1 As shown. ch-TU7-CBE was synthesized and cloned into the AncBE4max vector (Addgene #112094) with APOBEC removed by enzyme digestion, thus constructing the ch-TU7-CBE base editor.

[0169] The amino acid sequence of ch-TU7-CBE is shown in SEQ ID NO: 1:

[0170] MGDDVAAVAARVREAMAKLPPEAFQIAGECIDEAGAGLHPLALETNDAELAAVIGALGDAREEIDRAWQICRKVRDACADYLKIIGAAEPSAPATSAGAAGGRARHVTAKDGSQCPPEAGAVVDVLP RRVREGEAGEKTVGFVDGSVTDKFVSGRDQTWTRSILARAREVGLPPHLARFVSSHVEMKVAAMMTQTGKQHCELVINHVPCGSQPAQPPGCDQAIERFLPKGYTLTVHGTTQASRPFSKTYRGQA.

[0171] The nucleotide sequence of ch-TU7-CBE is shown in SEQ ID NO: 2:

[0172] ATGGGGGATGATGTTGCAGCAGTAGCCGCCAGAGTTCGTGAAGCCATGGCAAAACTGCCACCCGAGGCTTTCCAGATAGCAGGGGAGTGCATTGATGAAGCAGGCGCCGGGCTGCATCCTCTAGCTTTGGAAACGAATGATGCTGAGCTCGCAGCAGTCATAGGGGCCCTTGGTGATGCACGCGAAGAAATTGACCGGGCCTGGCAGATCTGCAGGAAAGTGCGAGACGCCTGTGCTGACTACCTGAAGATCATTGGTGCTGCAGAGCCGTCCGCTCCAGCAACATCAGCTGGTGCTGCTGGAGGCAGAGCTCGGCACGTCACTGCCAAAGACGGCAGCCAGTGCCCTCCTGAAGCTGGGGCGGTGGTGGATGTGCTTCCGAGGAGAGTGCGGGAAGGAGAAGCGGGAGAGAAGACTGTGGGCTTTGTTGATGGTTCCGTCACTGACAAATTTGTATCTGGAAGAGACCAGACCTGGACGAGGAGCATCCTGGCGAGGGCACGAGAGGTGGGATTGCCACCTCACCTGGCCCGCTTTGTTAGCAGTCATGTGGAGATGAAGGTGGCGGCTATGATGACACAGACTGGGAAGCAGCACTGCGAGCTCGTCATCAACCACGTACCCTGTGGAAGTCAGCCCGCACAGCCACCCGGCTGTGATCAAGCCATTGAGCGCTTCCTGCCAAAGGGCTATACCTTAACAGTTCATGGAACCACACAAGCATCTCGTCCTTTCTCAAAAACCTACAGAGGCCAAGCC。

[0173] Among them, the amino acid sequence of the fusion protein expressed by the ch-TU7-CBE base editor is shown in SEQ ID NO: 3:

[0174]

[0175] The nucleotide sequence of the fusion protein expressed by the ch-TU7-CBE base editor is shown in SEQ ID NO: 4:

[0176]

[0177] Example 3 Construction of sgRNA expression vector

[0178] Based on chicken ( Gallus gallus. Based on the NCBI classification number 9031) genome sequence, and according to the PAM sequence (NGG), the target region of the sgRNA for gene editing was selected, and two sgRNAs targeting chicken genes were designed and synthesized, as follows:

[0179] sgRNA1-F: 5'-ccggTAACATGCAGACCAGACAGG-3' (SEQ ID NO: 5);

[0180] sgRNA1-R: 5'-aaacCCTGTCTGGTCTGCATGTTA-3' (SEQ ID NO: 6);

[0181] sgRNA2-F: 5'-ccggCATCCGGGTGCAGCCGGTGC-3' (SEQ ID NO: 7);

[0182] sgRNA2-R: 5'-aaacGCACCGGCTGCACCCGGATG-3' (SEQ ID NO: 8).

[0183] The DNA sequences of the two pairs of single-stranded sgRNAs were annealed to form two oligonucleotide chains of sgRNAs targeting different gene loci in chickens. The two sgRNAs were then ligated into the enzyme-digested linearized 74707 vector (Addgene, #74707) to obtain two sgRNA expression vectors targeting chicken nuclear genes.

[0184] Example 4: Detection of editing efficiency in DF-1 cells

[0185] After sequencing verification of the sgRNA expression vector constructed in Example 3, the plasmid was extracted and purified by sodium acetate precipitation. The ch-TU7-CBE base editor and the two purified plasmids were then electroporated (520V (Celetrix electroporator), ch-TU7-CBE base editor to sgRNA ratio 1:1, 5 μg each) into chicken embryo fibroblasts (DF-1) (2 × 10⁻⁶ cells). 6 In DF-1 cells, the culture medium was changed 12 hours after transfection, and the genome of each cell was extracted 72 hours after transfection. PCR amplification was performed using specific primers. The amplification products were detected by agarose gel electrophoresis and then sequenced. Finally, the editing of specific gene loci in DF-1 cells was evaluated by sequencing peak diagram.

[0186] The PCR-specific primers are as follows:

[0187] sgRNA1 and sgRNA2-F: GTTCATGTTAAGTGCCCGGAACG (SEQ ID NO: 9);

[0188] sgRNA1 and sgRNA2-R: GTCATATCCTGGTTGTTAGCTGGC (SEQ ID NO: 10).

[0189] The editing results of sgRNA1 are as follows Figures 2-3 As shown, the PAM sequence recognized by sgRNA-1 is TGG, and the C4-to-T website-assessed mutation efficiency is 11.4%. 12 The mutation efficiency of the -to-T website was evaluated at 13.8% using BEAR software.

[0190] The editing results of sgRNA2 are as follows Figures 4-5 As shown, the PAM sequence recognized by sgRNA-2 is TGG, and the C5-to-T mutation efficiency was 15.1% as assessed by the website using BEAR software.

[0191] Comparison Example 1: Comparison of ch-TU7-CBE with other candidate cytosine deaminases

[0192] In addition to ch-TU7-CBE, the inventors also screened and obtained other candidate cytosine deaminases, including ch-SBG6-CBE (Uniprot No. A0A8U0SBG6 · A0A8U0SBG6_MUSPF) and ch-PTEVA-CBE (Uniprot No. A0A6P3RPW6 · A0A6P3RPW6_PTEVA). In this embodiment, the inventors compared the activities of ch-TU7-CBE, ch-SBG6-CBE, and ch-PTEVA-CBE.

[0193] Following the method in Example 2, ch-SBG6-CBE and ch-PTEVA-CBE base editors were constructed, respectively. Following the method in Example 4, the ch-SBG6-CBE and ch-PTEVA-CBE base editors, along with purified plasmids (sgRNA1 and sgRNA2), were electroporated into DF-1 cells. After transfection, PCR amplification was performed using specific primers. The amplification products were detected by agarose gel electrophoresis and then sequenced. Finally, the editing status of specific gene loci in DF-1 cells was evaluated using the sequencing peak diagram.

[0194] The editing result using the ch-SBG6-CBE base editor is as follows: Figures 6-7 As shown, its editing efficiency is significantly lower than that of the ch-TU7-CBE base editor (such as...). Figure 3 and Figure 5 As shown in the figure, this indicates that ch-SBG6-CBE has weak activity. The editing results using the ch-PTEVA-CBE base editor are as follows. Figures 8-9 As shown, its editing efficiency is higher than that of the ch-SBG6-CBE base editor, but still lower than that of the ch-TU7-CBE base editor (e.g., Figure 3 and Figure 5 (As shown).

[0195] The above results demonstrate that the ch-TU7-CBE obtained in this application has higher activity. Compared with other candidate deaminases, ch-TU7-CBE provides an optimized CBE for chicken genes, showing higher editing efficiency and applicability.

[0196] Comparison of Example 2: ch-TU7-CBE and BE4max

[0197] In this embodiment, the inventors compared the editing efficiency of the ch-TU7-CBE base editor and the BE4max base editor.

[0198] The amino acid sequence of BE4max is shown in SEQ ID NO: 11:

[0199] MKRTADGSEFESPKKKRKVSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSR YPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK.

[0200] According to the method in Example 2, a BE4max base editor was constructed, and according to the method in Example 4, the BE4max base editor and the purified plasmids (sgRNA1 and sgRNA2) were introduced into DF-1 cells by electroporation. After transfection, PCR amplification was performed using specific primers. The amplification products were detected by agarose gel electrophoresis and then sequenced. Finally, the editing status of specific gene sites in DF-1 cells was evaluated by sequencing peak diagram.

[0201] The results are as follows Figures 10-11 As shown, its editing efficiency is significantly lower than that of the ch-TU7-CBE base editor (such as...). Figure 3 and Figure 5 As shown in the figure, this further demonstrates that the ch-TU7-CBE obtained by screening in this application has higher activity.

[0202] The above description is merely a preferred embodiment of this application and is not intended to limit the application in any other way. Any person skilled in the art may make changes or modifications to the disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of this application without departing from the scope of the technical solution of this application shall still fall within the protection scope of this application.

Claims

1. Use of a base editing fusion protein in any of the following: (1) base editing of a chicken genome; (2) manufacture of a product for base editing of a chicken genome; (3) manufacture of a reagent for mediating base editing of a chicken genome; wherein the base editing fusion protein comprises a Cas9(D10A) nickase and a cytosine deaminase, the amino acid sequence of the Cas9(D10A) nickase is shown as SEQ ID NO: 12, and the amino acid sequence of the cytosine deaminase is shown as SEQ ID NO:

1.

2. The use of claim 1, wherein the base editing fusion protein further comprises at least one uracil DNA glycosylase inhibitor (UGI).

3. The use of claim 1 or 2, wherein the base editing fusion protein further comprises at least one nuclear localization sequence (NLS).

4. Use of a base editing system in any of the following: (1) base editing of a chicken genome; (2) manufacture of a product for base editing of a chicken genome; (3) manufacture of a reagent for mediating base editing of a chicken genome; wherein, the base editing system comprises a base editing fusion protein and / or an expression construct containing a nucleotide sequence encoding the base editing fusion protein; wherein the base editing fusion protein comprises a Cas9(D10A) nickase and a cytosine deaminase, the amino acid sequence of the Cas9(D10A) nickase is shown as SEQ ID NO: 12, and the amino acid sequence of the cytosine deaminase is shown as SEQ ID NO:

1.

5. The use of claim 4, wherein the base editing system further comprises at least one sgRNA and / or at least one expression construct containing a nucleotide sequence encoding the at least one sgRNA; wherein the at least one sgRNA is directed to at least one target site in the chicken genome.

6. A method of base editing a chicken genome, comprising contacting the base editing system in the use of claim 4 or 5 with a target site in a chicken genome.

7. A method of producing at least one modified chicken cell, comprising introducing the base editing system in the use of claim 4 or 5 into at least one chicken cell, thereby causing substitution of one or more nucleotides within a target site in the at least one chicken cell.

8. Use of the amino acid sequence shown as SEQ ID NO: 1 as a cytosine deaminase in base editing of a chicken genome.

Citation Information

Patent Citations

  • Cytosine deaminase, base editing system containing cytosine deaminase and application of cytosine deaminase

    CN116103271A