Novel genomic integration sites for fungal protein and fine chemical production

By introducing recombinant integrated nucleic acid molecules of specific natural genomic regions and flanking homologous arms into filamentous fungi, the problems of recombinant protein expression intensity and stability are solved, and efficient and stable recombinant protein expression is achieved, which is suitable for a variety of filamentous fungal host cells.

CN120641568APending Publication Date: 2025-09-12BASF SE
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202480009076.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-27
Filing Date
2024-01-18
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In the existing technology, the expression intensity and stability of recombinant proteins in filamentous fungi are affected by genomic position effects, and the number of beneficial genomic loci is limited, making it difficult to achieve efficient and stable recombinant expression.

Method used

A recombinant integration nucleic acid molecule comprising a specific natural genomic region and flanking homologous arms is used to integrate the recombinant nucleic acid molecule into a filamentous fungal host cell through homologous recombination, and the position effect of the natural genomic region is utilized to improve the expression intensity and stability. The recombinant integration nucleic acid molecule comprises a nucleic acid molecule flanked by different homologous arms that matches the natural genomic region.

Benefits of technology

It achieves efficient and stable recombinant protein expression in filamentous fungi, improves expression intensity and stability, and is suitable for a variety of filamentous fungal host cells, including Thermomycetes C1.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005514918550000201
    Figure BDA0005514918550000201
  • Figure BDA0005514918550000211
    Figure BDA0005514918550000211
  • Figure BDA0005514918550000212
    Figure BDA0005514918550000212
Patent Text Reader

Abstract

The present invention relates to genomic insertion sites that allow strong expression in filamentous fungi.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to genomic insertion sites that allow strong expression in filamentous fungi. Background Art

[0002] Biocatalytic and fermentative production of recombinant proteins and fine chemicals has long been practiced on an industrial scale. In recent years, demand for products derived from bioprocesses has increased. Various organisms, including bacteria and fungi, have been used in this production process. Filamentous fungi, in particular, despite being difficult to genetically transform and modify, have been increasingly used because they have proven to be highly efficient in biocatalytic and fermentative production.

[0003] Filamentous fungi have been shown to be excellent hosts for the production of a variety of proteins. Fungal strains such as Aspergillus, Trichoderma, Penicillium, and Thermothelomyces have been used in the industrial production of a variety of enzymes because they can secrete large amounts of protein into the fermentation broth. The protein secretion capacity of these fungi makes them preferred hosts for the targeted production of specific enzymes or enzyme mixtures.

[0004] Wild-type hermothelomyces thermophilus (T. thermophilus) C1 (also known as Myceliophthora thermophila, previously described as Chrysosporium lucknowense) is a thermotolerant ascomycete filamentous fungus that has been described as attractive for commercial-scale production of cellulases and other proteins. Only recently, this strain has also been shown to be efficient for the production of various fine chemicals, for example, in WO 20 / 161682, T. thermophilus C1 was disclosed to be capable of producing cannabinoids and their precursors.

[0005] WO00 / 20555 and US2012 / 0005812 disclose transformation systems for Myceliophthora thermophila C1 and describe the expression and secretion of heterologous proteins or polypeptides. A method for producing large amounts of polypeptides or proteins in an economical manner is also disclosed.

[0006] WO 15 / 004241 discloses multi-protease deficient filamentous fungal cells and methods for producing heterologous proteins.

[0007] For the expression of recombinant proteins in filamentous fungi, especially in Myceliophthora thermophila C1, some promoters that confer strong expression of recombinant proteins are known in the art (eg, Visser et al. 2011, Industrial Biotechnology 7(3), US2014 / 127788).

[0008] However, the intensity and stability of expression depend not only on the regulatory elements functionally connected to the recombinant protein to be expressed. The position of the recombinant construct in the genome also determines the intensity and stability of expression, a phenomenon known as position effect (Chen and Zhang (2016), Cell Systems 2; Bilyk et al. (2017), Microb Cell Fact 16 (5)). The number of genomic regions known in the art that are beneficial for strong and / or stable expression in Myceliophthora thermophila C1 is limited, and it is necessary to provide additional genomic loci for integrating recombinant expression constructs. One object of the present invention is to identify such beneficial genomic loci that have a positive position effect on recombinant expression. DETAILED DESCRIPTION

[0009] A first embodiment of the present invention is an expression system comprising:

[0010] a. A filamentous fungal host cell comprising a native genomic region comprising a sequence selected from the list consisting of:

[0011] i) a nucleic acid sequence comprising SEQ ID NO: 44

[0012] ii) a nucleic acid sequence that is at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 44, and

[0013] iii) a fragment of at least 500, preferably at least 750, preferably at least 1000, more preferably at least 1250, more preferably at least 1500, even more preferably at least 1750, even more preferably at least 2000, even more preferably at least 2250, even more preferably at least 2500, even more preferably at least 2750, even more preferably at least 3000, even more preferably at least 3250 consecutive nucleic acids of the nucleic acid sequence of i) or ii).

[0014] and

[0015] b. a recombinantly integrated nucleic acid molecule comprising a nucleic acid molecule flanked by different homology arms, each homology arm aligning over its entire length with at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, more preferably at least 350, even more preferably at least 400, even more preferably at least 450, even more preferably at least 500, even more preferably at least 550, even more preferably at least 600, even more preferably at least 650, even more preferably at least 700, even more preferably at least 750, even more preferably at least 800, even more preferably at least 850, even more preferably at least 900, even more preferably at least 1000, even more preferably at least 1100, even more preferably at least 1200, even more preferably at least 1300, even more preferably at least 1400, even more preferably at least 1500, even more preferably at least 1600, even more preferably at least 1700, even more preferably at least 1800, even more preferably at least 1900, even more preferably at least 2000 Preferably at least 950, even more preferably at least 1000, even more preferably at least 1050, even more preferably at least 1100, even more preferably at least 1150, even more preferably at least 1200, even more preferably at least 1250, even more preferably at least 1300, even more preferably at least 1350, even more preferably at least 1400, even more preferably at least 1450, even more preferably at least 1500 consecutive bases have at least 70% identity, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or most preferably at least 99% identity.

[0016] The nucleic acid molecule flanked by the homology arms may, for example, comprise a sequence encoding a protein or RNA molecule, a landing pad, or a recombinant expression construct. The landing pad may comprise a Cre / lox site, one or more sequences recognized by a guide RNA (which has a reduced off-target effect and guides the CRISPR / Cas molecule to this site), one or more recognition sites for a homing nuclease, etc. (Bourgeois et al. (2018) ACS Synth Biol 7(11)). The recombinant expression construct may comprise a promoter, a sequence encoding a protein or RNA molecule, and an optional terminator.

[0017] A further embodiment of the present invention is an expression system as defined above, wherein the native genomic region comprises an ORF encoding a small secreted protein, said ORF being flanked on one side by at least 250 nucleotides of non-coding DNA and on the other side by at least 250 nucleotides of non-coding DNA. Preferably, the downstream portion of the ORF is flanked by at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, preferably at least 1100, more preferably at least 1200 nucleotides of non-coding DNA. In a most preferred embodiment, the downstream portion of the ORF is flanked by at least 1296 nucleotides of non-coding DNA. Preferably, the upstream of the ORF is flanked by at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, at least 1500, at least 1600, preferably at least 1700, more preferably at least 1800 nucleotides of non-coding DNA. In the most preferred embodiment, the upstream of the ORF is flanked by at least 1847 nucleotides of non-coding DNA.

[0018] An additional embodiment of the present invention is any one of the expression systems as defined previously, wherein the small secreted protein comprises a sequence selected from the following list:

[0019] i. An amino acid sequence comprising SEQ ID NO: 46

[0020] ii. an amino acid sequence that is at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% homologous to SEQ ID NO: 46,

[0021] iii. A fragment of at least 100, preferably at least 110, more preferably at least 120, even more preferably at least 130, most preferably at least 140 consecutive amino acids of i) or ii).

[0022] Preferably, the ORF encodes protein MYCTH_2314963 (GeneBank Gene ID11512560, updated on November 28, 2019).

[0023] Another embodiment of the present invention is any one of the expression systems as defined above, wherein said native genomic region is flanked on one side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of:

[0024] a. An amino acid sequence comprising SEQ ID NO: 41

[0025] b. an amino acid sequence that is at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% homologous to SEQ ID NO: 41, and

[0026] a fragment of at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, even more preferably at least 320, most preferably at least 340 consecutive amino acids of ca) or b),

[0027] and the native genomic region is flanked on the other side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of:

[0028] d. An amino acid sequence comprising SEQ ID NO: 43

[0029] e. an amino acid sequence that is at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% homologous to SEQ ID NO: 43, and

[0030] fd) or e) of at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, even more preferably at least 325, even more preferably 350, most preferably at least 370 consecutive amino acids.

[0031] In the context of the present invention, the term "flanking" is intended to refer to the respective regions / ORFs / genes that "flank" a genomic region and define the boundaries of said genomic region. Flanking regions are not necessarily directly adjacent to the sequences defining said genomic region. For example, a genomic region may be defined by a coding sequence located within said genomic region. However, there may be intervening regions between such coding regions and the respective flanking regions defining the respective genomic region.

[0032] A further embodiment of the invention is any one of the expression systems as defined above, wherein the filamentous fungal host cell is from a species of the phylum Ascomycota, preferably Acremonium, Aspergillus, Agaricus, Aureobasidium, Cryptococcus, Corynascus, Chrysosporium, Filibasidium, Fusarium, Humicola, Magnaporthe, Monascus, Mucor, Myceliophthora, Mortierella, Nematodes, and the like. ocallimastix), Neurospora, Paecilomyces, Penicillium, Piromyces, Phanerochaete, Podospora, Pycnoporus, Rhizopus, Schizophyllum, Sordaria, Talaromyces, Rasamsonia, Thermoascus, Thermophilus, Thielavia, Tolypocladium, Trametes and Trichoderma, more preferably Aspergillus niger niger), Aspergillus oryzae, Aspergillus fumigatus, Neurospora crassa, Penicillium chrysogenum, Penicillium citrinum, Acremonium chrysogenum, Trichoderma reesei, Rasamsonia emersonii (formerly known as Talaromyces emersonii), Aspergillus sojae, Thermothelomyces heterothallica, and Thermothelomyces heterothallica (formerly known as Thermothelomyces thermophila and Chrysosporium lakenau), with Thermothelomyces thermophila being most preferred.

[0033] A further embodiment of the present invention is a method for stably integrating a recombinant polynucleotide in a filamentous fungal host cell, the method comprising the steps of:

[0034] c. Providing a filamentous fungal host cell comprising a native genomic region comprising a sequence selected from the list consisting of:

[0035] i) a nucleic acid sequence comprising SEQ ID NO: 44

[0036] ii) has at least 60%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, or more preferably at least 98% similarity to SEQ ID NO 44;

[0037] 7%, at least 98%, or at least 99% identical nucleic acid sequences, and

[0038] iii) a fragment of at least 500, preferably at least 750, preferably at least 1000, more preferably at least 1250, more preferably at least 1500, even more preferably at least 1750, even more preferably at least 2000, even more preferably at least 2250, even more preferably at least 2500, even more preferably at least 2750, even more preferably at least 3000, even more preferably at least 3250 consecutive nucleic acids of the nucleic acid sequence of i) or ii),

[0039] d. introducing a recombinant integrated nucleic acid molecule into the host cell, the recombinant integrated nucleic acid molecule comprising a nucleic acid molecule flanked by different homology arms, each homology arm aligning with at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 30 ...

[0040] 50, even more preferably at least 400, even more preferably at least 450, even more preferably at least 500, even more preferably at least 550, even more preferably at least 600, even more preferably at least 650, even more preferably at least 700, even more preferably at least 750, even more preferably at least 800, even more preferably at least 850, even more preferably at least 900, even more preferably at least 950, even more preferably at least 1000, even more preferably at least 1050, even more preferably at least 1100, even more preferably preferably at least 70% identity, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or most preferably at least 99% identity over at least 1150, even more preferably at least 1200, even more preferably at least 1250, even more preferably at least 1300, even more preferably at least 1350, even more preferably at least 1400, even more preferably at least 1450, even more preferably at least 1500 consecutive bases,

[0041] e. allowing the recombinant nucleic acid molecule to homologously recombine into the native genomic region.

[0042] Where the recombinant nucleic acid molecule comprises a coding sequence (e.g., contained in a recombinant expression construct), a preferred embodiment is that the host cell into which the recombinant nucleic acid molecule has been integrated produces the recombinant target peptide, protein and / or nucleic acid encoded by the coding region.

[0043] In a further preferred embodiment, the native genomic region prior to recombination of the recombinant nucleic acid molecule comprises a native gene encoding a small secreted protein, said native gene being flanked upstream by 250 kb of non-coding DNA and downstream by 250 kb of non-coding DNA. Preferably, the ORF of the small secreted protein is flanked downstream by at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, preferably at least 1100, more preferably at least 1200 nucleotides of non-coding DNA. In a most preferred embodiment, the ORF is flanked downstream by at least 1296 nucleotides of non-coding DNA. Preferably, the upstream of the ORF is flanked by at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, at least 1500, at least 1600, preferably at least 1700, more preferably at least 1800 nucleotides of non-coding DNA. In the most preferred embodiment, the upstream of the ORF is flanked by at least 1847 nucleotides of non-coding DNA.

[0044] In a further preferred embodiment, the landing pad or recombinant expression construct introduced into the genome of the fungal host cell partially or completely replaces the gene encoding the small secreted protein contained in the native genomic region. More preferably, it replaces all or part of the ORF encoding the small secreted protein.

[0045] A further embodiment of the present invention is the method as defined before, wherein said small secretory protein comprises a sequence selected from the list consisting of:

[0046] a. An amino acid sequence comprising SEQ ID NO: 46

[0047] b. an amino acid sequence that is at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% homologous to SEQ ID NO: 46, and

[0048] A fragment of at least 100, preferably at least 110, more preferably at least 120, even more preferably at least 130, most preferably at least 140 consecutive amino acids of ca) or b).

[0049] A further embodiment of the present invention is any of the methods as defined above, wherein said native genomic region is flanked on one side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of:

[0050] a. An amino acid sequence comprising SEQ ID NO: 41

[0051] b. an amino acid sequence that is at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% homologous to SEQ ID NO: 41, and

[0052] a fragment of at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, even more preferably at least 320, most preferably at least 340 consecutive amino acids of ca) or b),

[0053] and the native genomic region is flanked on the other side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of:

[0054] d. An amino acid sequence comprising SEQ ID NO: 43

[0055] e. an amino acid sequence that is at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% homologous to SEQ ID NO: 43, and

[0056] fd) or e) of at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, even more preferably at least 325, even more preferably 350, most preferably at least 370 consecutive amino acids.

[0057] A further embodiment of the invention is any one of the methods as defined above, wherein the filamentous fungal host cell is from a species selected from the group consisting of: Ascomycota, preferably Acremonium, Aspergillus, Agaricus, Aureobasidium, Cryptococcus, Coryneformis, Chrysosporium, Mycobasidiomycetes, Fusarium, Humicola, Magnaporthe oryzae, Monascus, Mucor, Myceliophthora, Mortierella, Neomelia, Neurospora, Paecilomyces, Penicillium, Pyrrocystis, Psoralea, Podosporium More preferably, Aspergillus niger, Aspergillus oryzae, Aspergillus fumigatus, Neurospora crassa, Penicillium chrysogenum, Penicillium citrinum, Acremonium chrysogenum, Trichoderma reesei, Rossella emersonii (formerly known as Talaromyces emersonii), Aspergillus sojae, Thermophila heterothallica and Thermophila thermophila (formerly known as Thermophila thermophila and Chrysosporium lakenau), most preferably Thermophila thermophila.

[0058] A further embodiment of the present invention is a recombinant filamentous fungal host cell comprising a recombinant nucleic acid molecule flanked by genomic DNA comprising on one side of the recombinant nucleic acid molecule a sequence selected from the list consisting of:

[0059] i) a nucleic acid sequence comprising 1296 consecutive bases from the 5′ end of SEQ ID NO: 44

[0060] ii) a nucleic acid sequence that is at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% identical to i), and

[0061] iii) a fragment of at least 100, preferably at least 200, preferably at least 300, preferably at least 400, preferably at least 500, preferably at least 600, preferably at least 700, preferably at least 800, preferably at least 900, more preferably at least 1000, more preferably at least 1100, even more preferably at least 1200, most preferably at least 1296 consecutive nucleic acids of i) or ii).

[0062] and said genomic DNA comprises on the other side of said recombinant nucleic acid molecule a sequence selected from the list consisting of:

[0063] iv) a nucleic acid sequence comprising 1847 consecutive bases from the 3′ end of SEQ ID NO: 44

[0064] v) a nucleic acid sequence that is at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% identical to iv), and

[0065] vi) a fragment of at least 100, preferably at least 200, preferably at least 300, preferably at least 400, preferably at least 500, preferably at least 600, preferably at least 700, preferably at least 800, preferably at least 900, preferably at least 1000, preferably at least 1100, more preferably at least 1200, preferably at least 1300, preferably at least 1400, preferably at least 1500, more preferably at least 1600, more preferably at least 1700, even more preferably at least 1800, most preferably at least 1847 consecutive nucleic acids of iv) or v).

[0066] A further embodiment of the present invention is a recombinant filamentous fungal host cell comprising a recombinant nucleic acid molecule flanked on one side by the XP_003662454.1 gene and on the other side by the XP_003662452.1 gene, preferably wherein the XP_003662454.1 gene is an amino acid molecule selected from the list consisting of:

[0067] a. an amino acid sequence comprising SEQ ID NO: 43,

[0068] b. an amino acid sequence that is at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% homologous to SEQ ID NO: 43, and

[0069] a fragment of at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, even more preferably at least 325, even more preferably 350, most preferably at least 370 consecutive amino acids of ca) or b),

[0070] And the XP_003662452.1 gene encodes an amino acid molecule selected from the list consisting of:

[0071] d. an amino acid sequence comprising SEQ ID NO: 41,

[0072] e. an amino acid sequence having at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% homology to SEQ ID NO: 41, and

[0073] fd) or e) of at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, even more preferably at least 320, most preferably at least 340 consecutive amino acids.

[0074] A further embodiment of the present invention is a recombinant filamentous fungal host cell as defined above, wherein the host cell is from a species selected from the group consisting of: Ascomycota, preferably Acremonium, Aspergillus, Agaricus, Aureobasidium, Cryptococcus, Coryneformis, Chrysosporium, Mycobasidiomycetes, Fusarium, Humicola, Magnaporthe oryzae, Monascus, Mucor, Myceliophthora, Mortierella, Neomelia, Neurospora, Paecilomyces, Penicillium, Pyrrocystis, Psoralea, Podosporium More preferably, Aspergillus niger, Aspergillus oryzae, Aspergillus fumigatus, Neurospora crassa, Penicillium chrysogenum, Penicillium citrinum, Acremonium chrysogenum, Trichoderma reesei, Rossella emersonii (formerly known as Talaromyces emersonii), Aspergillus sojae, Thermophilic Myceliophthora heterothallis and Thermophilic Myceliophthora (formerly known as Thermophilic Myceliophthora and Chrysosporium lakenau), most preferably Thermophilic Myceliophthora.

[0075] An additional embodiment of the present invention is a recombinantly integrated nucleic acid molecule comprising a nucleic acid molecule flanked by different homology arms, each homology arm being aligned over its entire length to at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, more preferably at least 350, even more preferably at least 400, even more preferably at least 450, even more preferably at least 500, even more preferably at least 550, even more preferably at least 600, even more preferably at least 650, even more preferably at least 700, even more preferably at least 750, even more preferably at least 800, even more preferably at least 850, even more preferably at least 900, even more preferably at least 900, even more preferably at least 900, 95%, even more preferably at least 96%, at least 97%, at least 98% or most preferably at least 99% identity over at least 50, even more preferably at least 1000, even more preferably at least 1050, even more preferably at least 1100, even more preferably at least 1150, even more preferably at least 1200, even more preferably at least 1250, even more preferably at least 1300, even more preferably at least 1350, even more preferably at least 1400, even more preferably at least 1450, even more preferably at least 1500 consecutive bases, said native genomic region comprising a sequence selected from the list consisting of:

[0076] i) a nucleic acid comprising SEQ ID NO: 44,

[0077] ii) a nucleic acid sequence that is at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 44, and

[0078] iii) a fragment of at least 500, preferably at least 750, preferably at least 1000, more preferably at least 1250, more preferably at least 1500, even more preferably at least 1750, even more preferably at least 2000, even more preferably at least 2250, even more preferably at least 2500, even more preferably at least 2750, even more preferably at least 3000, even more preferably at least 3250 consecutive nucleic acids of i) or ii).

[0079] A further embodiment is a recombinantly integrated nucleic acid molecule comprising a nucleic acid molecule flanked by different homology arms, each homology arm being aligned over its entire length to at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, more preferably at least 350, even more preferably at least 400, even more preferably at least 450, even more preferably at least 500, even more preferably at least 550, even more preferably at least 600, even more preferably at least 650, even more preferably at least 700, even more preferably at least 750, even more preferably at least 800, even more preferably at least 850, even more preferably at least 900, even more preferably at least 950, even more preferably at least 10 95%, even more preferably at least 96%, at least 97%, at least 98% or most preferably at least 99% identity over at least 100, even more preferably at least 1050, even more preferably at least 1100, even more preferably at least 1150, even more preferably at least 1200, even more preferably at least 1250, even more preferably at least 1300, even more preferably at least 1350, even more preferably at least 1400, even more preferably at least 1450, even more preferably at least 1500 consecutive bases, said native genomic region being flanked on one side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of:

[0080] 1. An amino acid sequence comprising SEQ ID NO: 41,

[0081] 2. an amino acid sequence that is at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% homologous to SEQ ID NO: 41, and

[0082] 3. a fragment of at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, even more preferably at least 320, most preferably at least 340 consecutive amino acids of 1) or 2),

[0083] and the native genomic region is flanked on the other side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of:

[0084] 4. An amino acid sequence comprising SEQ ID NO: 43,

[0085] 5. An amino acid sequence that is at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% homologous to SEQ ID NO: 43, and

[0086] 6. A fragment of at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, even more preferably at least 325, even more preferably 350, most preferably at least 370 consecutive amino acids of 4) or 5).

[0087] A further embodiment of the present invention is a vector comprising any of the recombinant nucleic acid molecules as defined above.

[0088] definition

[0089] It should be understood that the present invention is not limited to a specific method or protocol. It should also be understood that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the scope of the present invention, which is limited only by the appended claims. It must be noted that, unless the context clearly provides otherwise, as used herein and in the appended claims, the singular forms "one / kind", "and", and "said" include plural indicators. Thus, for example, reference to "vector" refers to one or more vectors and includes equivalents thereof known to those skilled in the art. The term "about" is used herein to indicate approximate, roughly, approximately, or in the region thereof. When the term "about" is used in conjunction with a numerical range, it modifies the range by extending the boundaries to above and below the stated values. Typically, the term "about" is used herein to modify a numerical value to a 20% change above and below the stated value, preferably a 10% change upward or downward (higher or lower). As used herein, the word "or" means any one member of a particular list, and also includes any combination of the members of the list. When used in this specification and the appended claims, the words "comprise", "comprising", and "include", "including", and "includes" are intended to specify the presence of one or more stated features, integers, components, or steps, but they do not exclude the presence or addition of one or more other features, integers, components, steps, or groups thereof. For the sake of clarity, certain terms used in this specification are defined and used as follows:

[0090] Coding region: As used herein, the term "coding region" or "open reading frame" or "ORF" when used in reference to a structural gene refers to the nucleotide sequence that encodes the amino acids found in the nascent polypeptide as a result of translation of the mRNA molecule. In eukaryotes, the coding region is bounded on the 5'-side by the nucleotide triplet "ATG", which encodes the initiator methionine, and on the 3'-side by one of three triplets (i.e., TAA, TAG, TGA) that specify the stop codon. In addition to containing introns, the genomic form of a gene may also include sequences located at both the 5'- and 3'-ends of the sequence present on the RNA transcript. These sequences are referred to as "flanking" sequences or regions (these flanking sequences are located 5' or 3' to the non-translated sequences present on the mRNA transcript). The 5'-flanking region may contain regulatory sequences that control or influence the transcription of the gene, such as promoters and enhancers. The 3'-flanking region may contain sequences that direct transcription termination, post-transcriptional cleavage, and polyadenylation.

[0091] Complementarity: "Complementarity" or "complementarity" refers to two nucleotide sequences comprising antiparallel nucleotide sequences that are capable of pairing with each other (by the base pairing rules) when hydrogen bonds are formed between the complementary base residues in the antiparallel nucleotide sequences. For example, the sequence 5'-AGT-3' is complementary to the sequence 5'-ACT-3'. Complementarity can be "partial" or "full." "Partial" complementarity means that one or more nucleic acid bases do not match according to the base pairing rules. "Full" or "complete" complementarity between nucleic acid molecules means that each nucleic acid base matches another base under the base pairing rules. The degree of complementarity between nucleic acid molecule chains has a significant effect on the efficiency and strength of hybridization between nucleic acid molecule chains. As used herein, the "complementary sequence" of a nucleic acid sequence refers to a nucleotide sequence whose nucleic acid molecule exhibits complete complementarity with the nucleic acid molecule of the nucleic acid sequence.

[0092] Endogenous: An "endogenous" nucleotide sequence refers to a nucleotide sequence that is present in the genome of an untransformed cell.

[0093] Enhanced expression: "Enhance" or "increase" the expression of a nucleic acid molecule in a cell are used equivalently herein and mean that after applying the method of the invention, the expression level of the nucleic acid molecule in the cell is higher than its expression in the cell before applying the method, or the expression level of the nucleic acid molecule in the cell is higher compared to a reference cell lacking the recombinant nucleic acid molecule of the invention. As used herein, the terms "enhanced" or "increased" are synonymous and mean herein higher, preferably significantly higher, expression of the nucleic acid molecule to be expressed. As used herein, "enhancement" or "increase" of the level of an agent such as a protein, mRNA or RNA means that the level is increased relative to substantially the same cells grown under substantially the same conditions lacking the recombinant nucleic acid molecule of the invention, recombinant construct of the invention or recombinant vector. As used herein, "enhancement" or "increase" of the level of an agent (e.g., preRNA, mRNA, rRNA, tRNA, snoRNA, snRNA, and / or protein products encoded thereby, expressed by a target gene) means that the level is increased by 50% or more, e.g., 100% or more, preferably 200% or more, more preferably 5-fold or more, even more preferably 10-fold or more, most preferably 20-fold or more, e.g., 50-fold, relative to a cell or organism lacking the recombinant nucleic acid molecule of the invention. The enhancement or increase can be determined by methods familiar to the skilled person. Thus, the enhancement or increase in the amount of nucleic acid or protein can be determined, for example, by immunological detection of the protein. In addition, techniques such as protein assays, fluorescence, Northern hybridization, nuclease protection assays, reverse transcription (quantitative RT-PCR), ELISA (enzyme-linked immunosorbent assay), Western blotting, radioimmunoassays (RIA) or other immunoassays, and fluorescence-activated cell analysis (FACS) can be used to measure specific proteins or RNA in cells. Depending on the type of protein product induced, its activity or effect on the cellular phenotype can also be determined. Methods for determining protein amounts are known to the skilled person. Examples which may be mentioned are: the microbiuret method (Goa J (1953) Scand J Clin Lab Invest 5: 218-222), the Folin-Ciocalteau method (Lowry OH et al. (1951) J Biol Chem 193: 265-275) or measuring the absorption of CBB G-250 (Bradford MM (1976) Analyt Biochem 72: 248-254).

[0094] Expression: "Expression" refers to the biosynthesis of a gene product, preferably the transcription and / or translation of a nucleotide sequence (e.g., an endogenous gene or a heterologous gene) in a cell. For example, in the case of a structural gene, expression involves the transcription of the structural gene into mRNA and, optionally, the subsequent translation of the mRNA into one or more polypeptides. In other cases, expression may refer solely to the transcription of a DNA-bearing RNA molecule.

[0095] Expression construct: As used herein, "expression construct" means a DNA sequence capable of directing the expression of a specific nucleotide sequence in a cell, comprising a promoter functional in the cell into which it is introduced, operably linked to a nucleotide sequence of interest, which is optionally operably linked to a termination signal. If translation is desired, it typically also contains sequences required for proper translation of the nucleotide sequence. The coding region may encode a protein of interest, but may also encode a functional RNA of interest, such as RNAa, siRNA, snoRNA, snRNA, microRNA, ta-siRNA, or any other non-coding regulatory RNA, in either the sense or antisense orientation. An expression construct comprising a nucleotide sequence of interest may be chimeric, meaning that one or more of its components is heterologous to one or more of its other components. An expression construct may also be a naturally occurring construct that has been obtained in a recombinant form for heterologous expression. However, typically, the expression construct is heterologous to the host, meaning that the specific DNA sequence of the expression construct is not naturally present in the host cell and must be introduced into the host cell or an ancestor of the host cell through a transformation event. The expression of the nucleotide sequence in the expression construct can be under the control of a constitutive promoter or an inducible promoter, which initiates transcription only when the host cell is exposed to some specific external stimulus. The promoter can also be specific for a particular developmental stage.

[0096] Exogenous: The term "exogenous" refers to any nucleic acid molecule (e.g., a gene sequence) that is introduced into the genome of a cell by experimental manipulation, and can include sequences found in that cell, so long as the introduced sequence contains some modification (e.g., point mutations, the presence of a selectable marker gene, etc.) and is therefore different from the naturally occurring sequence.

[0097] Functional connection: The term "functional connection" or "functionally connected" is to be understood as meaning, for example, the sequential arrangement of regulatory elements (e.g. promoters) and the nucleic acid sequence to be expressed and, if appropriate, further regulatory elements (e.g. terminators) such that each regulatory element can perform its intended function to allow, modify, promote or otherwise influence the expression of the nucleic acid sequence. As synonyms, the expressions "operably linked" or "operably linked" can be used. The result of the expression may depend on the arrangement of the nucleic acid sequence relative to the sense or antisense RNA. For this purpose, a direct connection in the chemical sense is not necessarily required. Genetic control sequences (e.g. enhancer sequences) can also exert their function on the target sequence from a more remote position or indeed from other DNA molecules. Preferred arrangements are those in which the nucleic acid sequence to be recombinantly expressed is located after the sequence acting as promoter, so that the two sequences are covalently linked to each other. The distance between the promoter sequence and the nucleic acid sequence to be recombinantly expressed is preferably less than 200 base pairs, particularly preferably less than 100 base pairs, very particularly preferably less than 50 base pairs. In a preferred embodiment, the nucleic acid sequence to be transcribed is located after the promoter in such a way that the transcription starting point is identical to the expected starting position of the chimeric RNA of the present invention. Functional connection and expression construct can be produced by conventional recombination and cloning techniques as described (for example, in Maniatis T, Fritsch EF and Sambrook J (1989) Molecular Cloning: A Laboratory Manual, 2nd edition, Cold Spring Harbor Laboratory, Cold Spring Harbor (NY); Silhavy et al. (1984) Experiments with Gene Fusions, Cold Spring Harbor Laboratory, Cold Spring Harbor (NY); Ausubel et al. (1987) Current Protocols in Molecular Biology, Greene Publishing Assoc. and Wiley Interscience). However, for example, a joint serving as a specific cleavage site with a restriction enzyme or a further sequence serving as a signal peptide can also be located between the two sequences. The insertion of a sequence can also result in the expression of a fusion protein. Preferably, the expression construct consisting of the connection of regulatory regions (eg promoter) and the nucleic acid sequence to be expressed can be present in vector-integrated form and inserted into the genome, for example, by transformation.

[0098] Gene: The term "gene" refers to a region operably linked to appropriate regulatory sequences that are capable of regulating the expression of a gene product (e.g., a polypeptide or functional RNA) in some manner. A gene includes untranslated regulatory regions of DNA (e.g., promoters, enhancers, repressors, etc.) preceding (upstream) and following (downstream) the coding region (open reading frame, ORF), and, where applicable, intervening sequences (i.e., introns) between the individual coding regions (i.e., exons). As used herein, the term "structural gene" is intended to mean a DNA sequence that is transcribed into mRNA, which is then translated into the characteristic amino acid sequence of a specific polypeptide.

[0099] Genome and genomic DNA: The term "genome" or "genomic DNA" refers to the heritable genetic information of a host organism. Preferably, the term genome or genomic DNA refers to the chromosomal DNA of the cell nucleus.

[0100] Heterologous: The term "heterologous" with respect to a nucleic acid molecule or DNA refers to a nucleic acid molecule that is operably linked or manipulated to become operably linked to a second nucleic acid molecule, such as a promoter, which second nucleic acid molecule, such as a promoter, is not operably linked to the nucleic acid molecule in nature (e.g., in the genome of a WT cell) or is operably linked to the nucleic acid molecule at a different position or site in nature (e.g., in the genome of a WT cell).

[0101] Preferably, the term "heterologous" with respect to a nucleic acid molecule or DNA (e.g., a promoter) refers to a nucleic acid molecule that is operably linked or manipulated to become operably linked to a second nucleic acid molecule, such as an ORF, to which the nucleic acid molecule is not operably linked in nature.

[0102] A heterologous expression construct comprising a nucleic acid molecule and one or more regulatory nucleic acid molecules (such as a promoter or transcription termination signal) linked thereto is, for example, a construct produced by experimental manipulation, wherein a) the nucleic acid molecule, b) the regulatory nucleic acid molecule, or c) both (i.e., (a) and (b)) are not located in their natural (native) genetic environment or have been modified by experimental manipulation, examples of modifications being substitutions, additions, deletions, inversions, or insertions of one or more nucleotide residues. The natural genetic environment refers to the natural chromosomal locus in the organism of origin or to its presence in a genomic library. In the case of a genomic library, the natural genetic environment of the nucleic acid molecule sequence is preferably at least partially retained. The environment flanks the nucleic acid sequence on at least one side and has a sequence length of at least 50 bp, preferably at least 500 bp, particularly preferably at least 1,000 bp, and very particularly preferably at least 5,000 bp. A naturally occurring expression construct (e.g., a naturally occurring combination of a promoter and a corresponding gene) becomes a transgenic expression construct when it is modified by non-natural, synthetic, "artificial" methods (e.g., mutagenesis). Such methods have been described (US 5,565,350; WO 00 / 15815). For example, a nucleic acid molecule encoding a protein that is operably linked to a promoter that is not the native promoter of the molecule is considered to be heterologous with respect to the promoter. Preferably, the heterologous DNA is not endogenous to or naturally associated with the cell into which it is introduced, but has been obtained from another cell or has been synthesized. Heterologous DNA also includes endogenous DNA sequences that contain certain modifications, multiple copies of non-naturally occurring endogenous DNA sequences, or DNA sequences that are not naturally associated with another DNA sequence to which they are physically linked.

[0103] Hybridization: The term "hybridization" as defined herein is a process in which substantially complementary nucleotide sequences anneal to each other. The hybridization process can occur entirely in solution, i.e., both complementary nucleic acids are in solution. The hybridization process can also occur when one of the complementary nucleic acids is fixed to a matrix (such as magnetic beads, agarose beads, or any other resin). The hybridization process can also occur when one of the complementary nucleic acids is fixed to a solid support (such as nitrocellulose or nylon membrane) or is fixed to, for example, a silica glass support (the latter being referred to as a nucleic acid array, microarray, or nucleic acid chip) by, for example, photolithography. In order for hybridization to occur, the nucleic acid molecules are typically thermally denatured or chemically denatured to melt the double strand into two single strands and / or to remove hairpins or other secondary structures from the single stranded nucleic acid.

[0104] The term "stringency" refers to the conditions under which hybridization occurs. The stringency of hybridization is affected by conditions such as temperature, salt concentration, ionic strength, and hybridization buffer composition. Typically, low stringency conditions are selected to be approximately 30°C lower than the thermal melting point (Tm) of the specific sequence under the ionic strength and pH defined. Moderate stringency conditions are temperatures 20°C lower than Tm, while high stringency conditions are temperatures 10°C lower than Tm. High stringency hybridization conditions are typically used to separate hybridization sequences with high sequence similarity to the target nucleic acid sequence. However, due to the degeneracy of the genetic code, nucleic acids can deviate from and still encode substantially the same polypeptide in sequence. Therefore, sometimes moderate stringency hybridization conditions may be needed to identify this type of nucleic acid molecules.

[0105] "Tm" is the temperature, under defined ionic strength and pH, at which 50% of the target sequence hybridizes to a perfectly matched probe. Tm depends on the solution conditions and the base composition and length of the probe. For example, longer sequences hybridize specifically at higher temperatures. The maximum hybridization rate is obtained from about 16°C below the Tm to 32°C. The presence of monovalent cations in the hybridization solution reduces the electrostatic repulsion between the two nucleic acid strands, thereby promoting hybrid formation; this effect is visible for sodium concentrations up to 0.4 M (for higher concentrations, the effect is negligible). Formamide lowers the melting temperature of DNA-DNA and DNA-RNA duplexes by 0.6 to 0.7°C per percentage formamide, and the addition of 50% formamide allows hybridization at 30 to 45°C, although the hybridization rate will be reduced. Base pair mismatches reduce the hybridization rate and the thermal stability of the duplex. On average, for large probes, the Tm is reduced by about 1°C per % base mismatch. Depending on the type of hybrid, the Tm can be calculated using the following formula:

[0106] DNA-DNA hybrid (Meinkoth and Wahl, Anal. Biochem., 138: 267-284, 1984):

[0107] Tm=81.5℃+16.6xlog[Na+]a+0.41x%[G / Cb]-500x[Lc]-1-0.61x%formamide

[0108] DNA-RNA or RNA-RNA hybrids:

[0109] Tm=79.8+18.5(log10[Na+]a)+0.58(%G / Cb)+11.8(%G / Cb)2-820 / L oligo-DNA or oligo-RNA hybrid:

[0110] For <20 nucleotides: Tm = 2 (1n)

[0111] For 20-35 nucleotides: Tm = 22 + 1.46 (1n)

[0112] a or for other monovalent cations, but is accurate only in the range 0.01–0.4 M.

[0113] bAccurate only for %GC in the range of 30% to 75%.

[0114] c L = length of the duplex in base pairs.

[0115] d oligo, oligonucleotide; ln, effective primer length = 2×(number of G / C) + (number of A / T).

[0116] Non-specific binding can be controlled using any of a number of known techniques, for example, by blocking the membrane with a solution containing protein, adding heterologous RNA, DNA, and SDS to the hybridization buffer, and treating with RNase. For non-related probes, a series of hybridizations can be performed by changing one of the following: (i) gradually lowering the annealing temperature (e.g., from 68° C. to 42° C.) or (ii) gradually lowering the formamide concentration (e.g., from 50% to 0%). The skilled person is aware of the various parameters that can be changed during hybridization, and these parameters will maintain or change stringency conditions.

[0117] In addition to hybridization conditions, the specificity of hybridization is usually also dependent on the function of post-hybridization washing. In order to remove the background produced by non-specific hybridization, the sample is washed with a dilute salt solution. The key factors of this washing include the ionic strength and temperature of the final wash solution: the lower the salt concentration and the higher the wash temperature, the higher the stringency of the washing. Washing conditions are usually carried out at or below hybridization stringency. The signal produced by positive hybridization is at least twice the background signal. Usually, suitable stringent conditions for nucleic acid hybridization assays or gene amplification detection procedures are as described above. More stringent or less stringent conditions can also be selected. The technician knows the various parameters that can be changed during washing, and these parameters will maintain or change stringency conditions.

[0118] For example, typical high stringency hybridization conditions for DNA hybrids longer than 50 nucleotides include hybridization at 65°C in 1x SSC or at 42°C in 1x SSC and 50% formamide, followed by a wash in 0.3x SSC at 65°C. Examples of moderate stringency hybridization conditions for DNA hybrids longer than 50 nucleotides include hybridization at 50°C in 4x SSC or at 40°C in 6x SSC and 50% formamide, followed by a wash in 50°C in 2x SSC. The length of the hybrid is the expected length of the hybridizing nucleic acids. When nucleic acids of known sequence are hybridized, the hybrid length can be determined by aligning the sequences and identifying conserved regions as described herein. 1x SSC is 0.15M NaCl and 15mM sodium citrate; hybridization and wash solutions may additionally contain 5x Denhardt's reagent, 0.5-1.0% SDS, 100 μg / ml denatured fragmented salmon sperm DNA, and 0.5% sodium pyrophosphate. Another example of high stringency conditions is hybridization in 0.1x SSC (containing 0.1 SDS and optionally 5x Denhardt's reagent, 100 μg / ml denatured fragmented salmon sperm DNA, 0.5% sodium pyrophosphate) at 65°C, followed by a wash in 0.3x SSC at 65°C.

[0119] For the purpose of defining stringency levels, reference may be made to Sambrook et al. (2001) Molecular Cloning: a laboratory manual, 3rd ed., Cold Spring Harbor Laboratory Press, CSH, New York or to Current Protocols in Molecular Biology, John Wiley & Sons, NY (1989 and annual updates).

[0120] "Identity": When used in relation to comparing two or more nucleic acid or amino acid molecules, "identity" means that the sequences of the molecules share a degree of sequence similarity such that the sequences are partially identical.

[0121] When compared with the parent enzyme or nucleic acid molecule, enzyme variants can be defined by their sequence identity.Sequence identity is usually provided with " sequence identity % " or " identity % ".In order to determine the percent identity between two amino acid sequences in the first step, a paired sequence alignment is generated between the two sequences, wherein the two sequences are compared (that is, paired global alignment) over their full length. This alignment is generated by a program implementing Needleman and Wunsch algorithm (J.Mol.Biol. (1979) 48, 443-453 pages), preferably by using the program " NEEDLE " (European Molecular Biology Open Software Suite (EMBOSS)) with program default parameters (gap opening=10.0, gap extension=0.5 and matrix=EBLOSUM62). For purposes of the present invention, a preferred alignment is an alignment from which the highest sequence identity can be determined.

[0122] The following example is intended to illustrate two nucleotide sequences, but the same calculations apply to protein sequences:

[0123] Seq A: AAGATACTG Length: 9 bases

[0124] Seq B: GATCTGA Length: 7 bases

[0125] Therefore, the shorter sequence is sequence B.

[0126] Generating a pairwise global alignment showing the two sequences over their full length results in

[0127]

[0128] The "I" symbol in the alignment represents the same residue (which means a base for DNA or an amino acid for protein). The number of the same residues is 6.

[0129] The "-" sign in the alignment indicates a gap. The number of gaps introduced by alignment within Seq B is 1. The number of gaps introduced by alignment at the Seq B boundary is 2, and the number at the Seq A boundary is 1.

[0130] The aligned sequences are shown with an alignment length of 10 over their full length.

[0131] Generating a pairwise alignment according to the invention which displays shorter sequences over their entire length thus results in:

[0132]

[0133] Generating a pairwise alignment according to the invention showing sequence A over its entire length thus results in:

[0134]

[0135] Generating a pairwise alignment according to the invention showing sequence B over its entire length thus results in:

[0136]

[0137] The alignment length of the shorter sequence is shown to be 8 over its entire length (there is one gap which is taken into account in the alignment length of the shorter sequence).

[0138] Therefore, the alignment length showing Seq A over its full length would be 9 (meaning Seq A is a sequence of the invention).

[0139] Therefore, the alignment length showing Seq B over its full length would be 8 (meaning Seq B is a sequence of the invention).

[0140] After aligning the two sequences, in a second step, the identity value is determined based on the resulting alignment. For the purposes of this specification, the percent identity is calculated as follows: % identity = (identical residues / length of the alignment region showing the corresponding sequences of the invention over its entire length) * 100. Thus, the sequence identity relevant to the comparison of two amino acid sequences according to this embodiment is calculated by dividing the number of identical residues by the length of the alignment region showing the corresponding sequences of the invention over its entire length. This value is multiplied by 100 to obtain the "% identity". According to the examples provided above, the % identity is: for Seq A, which is a sequence of the invention, (6 / 9) * 100 = 66.7%; for Seq B, which is a sequence of the invention, (6 / 8) * 100 = 75%.

[0141] The terms "introducing" or "introduction" and the like, with reference to the introduction of a donor DNA molecule into a target site of a target DNA, refer to any introduction of the sequence of the donor DNA molecule into the target region, such as by physical integration of the donor DNA molecule or a portion thereof into the target region, or introduction of the sequence of the donor DNA molecule or a portion thereof into the target region, wherein the donor DNA serves as a template for a polymerase.

[0142] Isogenic: Genetically identical organisms (such as fungi) except that they may differ by the presence or absence of heterologous DNA sequences.

[0143] Isolated: As used herein, the term "isolated" means that a substance has been artificially removed and exists separately from its original natural environment and is therefore not a natural product. An isolated substance or molecule (such as a DNA molecule or enzyme) can exist in a purified form or can exist in a non-natural environment, such as in a transgenic host cell. For example, a naturally occurring polynucleotide or polypeptide present in a living cell is not isolated, but the same polynucleotide or polypeptide separated from some or all coexisting substances in the natural system is isolated. Such a polynucleotide can be part of a vector and / or such a polynucleotide or polypeptide can be part of a composition and will be isolated because such a vector or composition is not part of its original environment. Preferably, when the term "isolated" is used for a nucleic acid molecule, as in "isolated nucleic acid sequence", it refers to a nucleic acid sequence that has been identified and separated from at least one contaminating nucleic acid molecule that is normally associated with it in its natural source. An isolated nucleic acid molecule is a nucleic acid molecule that exists in a form or environment that is different from the form or environment in which it exists in nature. In contrast, a non-isolated nucleic acid molecule is a nucleic acid molecule that exists in its naturally occurring state, such as DNA and RNA. For example, a given DNA sequence (e.g., a gene) is present on a host cell chromosome adjacent to neighboring genes; an RNA sequence, such as a specific mRNA sequence encoding a specific protein, is present in the cell as a mixture with many other mRNAs encoding multiple proteins. However, an isolated nucleic acid sequence comprising, for example, SEQ ID NO: 1 includes, for example, such nucleic acid sequences in cells that normally contain SEQ ID NO: 1, wherein the nucleic acid sequence is located in a chromosomal or extrachromosomal position different from that of the natural cell, or is otherwise flanked by nucleic acid sequences that are different from those found in nature. An isolated nucleic acid sequence may exist in single-stranded or double-stranded form. When an isolated nucleic acid sequence is used to express a protein, the nucleic acid sequence will contain at least a portion of the sense strand or coding strand (i.e., the nucleic acid sequence may be single-stranded). Alternatively, it may contain both the sense strand and the antisense strand (i.e., the nucleic acid sequence may be double-stranded).

[0144] Minimal promoter: A promoter element, particularly a TATA element, that is inactive or has greatly reduced promoter activity in the absence of upstream activation. In the presence of appropriate transcription factors, the minimal promoter functions to allow transcription.

[0145] Non-coding: The term "non-coding" refers to nucleic acid sequences that do not encode part or all of an expressed protein. Non-coding sequences include, but are not limited to, functional RNA, introns, enhancers, promoter regions, 3' untranslated regions, and 5' untranslated regions.

[0146] Nucleic acid and nucleotide: The terms "nucleic acid" and "nucleotide" refer to naturally occurring or synthetic or artificial nucleic acids or nucleotides. The terms "nucleic acid" and "nucleotide" include deoxyribonucleotides or ribonucleotides or any nucleotide analogs and polymers or hybrids thereof in single-stranded or double-stranded, sense or antisense form. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences, as well as the explicitly indicated sequence. The term "nucleic acid" is used interchangeably herein with "gene," "cDNA," "mRNA," "oligonucleotide," and "polynucleotide." Nucleotide analogs include nucleotides with modifications in the chemical structure of the base, sugar and / or phosphate, including but not limited to 5-position pyrimidine modification, 8-position purine modification, modification at the cytosine exocyclic amine, substitution of 5-bromo-uracil, etc.; and 2'-position sugar modification, including but not limited to sugar-modified ribonucleotides, wherein the 2'-OH is replaced by a group selected from the following items: H, OR, R, halogen, SH, SR, NH2, NHR, NR2 or CN. Short hairpin RNA (shRNA) can also contain non-natural elements such as non-natural bases (e.g., inosine and xanthine), non-natural sugars (e.g., 2'-methoxyribose) or non-natural phosphodiester bonds (e.g., methylphosphonate, phosphorothioate and peptide bonds).

[0147] Nucleic acid sequence: The phrase "nucleic acid sequence" refers to a single-stranded or double-stranded polymer of deoxyribonucleotide or ribonucleotide bases read from the 5' end to the 3' end. It includes chromosomal DNA, self-replicating plasmids, infectious polymers of DNA or RNA, and DNA or RNA that plays a primary structural role. "Nucleic acid sequence" also refers to a consecutive list of abbreviations, letters, characters or words representing nucleotides. In one embodiment, a nucleic acid can be a "probe", which is a relatively short nucleic acid, generally less than 100 nucleotides in length. Typically, nucleic acid probes are about 50 nucleotides to about 10 nucleotides in length. The "target region" of a nucleic acid is a portion of the nucleic acid that is identified as being of interest. The "coding region" of a nucleic acid is a portion of a nucleic acid that is transcribed and translated in a sequence-specific manner to produce a specific polypeptide or protein when placed under the control of appropriate regulatory sequences. The coding region is considered to encode such a polypeptide or protein.

[0148] Oligonucleotide: The term "oligonucleotide" refers to an oligomer or polymer of ribonucleic acid (RNA) or deoxyribonucleic acid (DNA), or mimetics thereof, as well as oligonucleotides with non-naturally occurring portions that function similarly. Such modified or substituted oligonucleotides are often preferred over the native forms due to desirable properties, such as enhanced cellular uptake, enhanced affinity for nucleic acid targets, and increased stability in the presence of nucleases. An oligonucleotide preferably comprises two or more nucleotide monomers covalently coupled to one another by a bond (e.g., a phosphodiester bond) or a substituted bond.

[0149] Overhang: An "overhang" is a relatively short single-stranded nucleotide sequence on the 5'- or 3'-hydroxyl terminus of a double-stranded oligonucleotide molecule (also called an "extension," "protruding end," or "sticky end").

[0150] Polypeptide: The terms "polypeptide," "peptide," "oligopeptide," "polypeptide," "gene product," "expression product," and "protein" are used interchangeably herein to refer to a polymer or oligomer of consecutive amino acid residues.

[0151] "Precisely" with respect to the introduction of a donor DNA molecule in a target region means that the sequence of the donor DNA molecule is introduced into the target region without any InDels, duplications or other mutations compared to the unaltered DNA sequence of the target region not contained in the sequence of the donor DNA molecule.

[0152] Promoter: The terms "promoter" or "promoter sequence" are equivalent and, as used herein, refer to a DNA sequence that is capable of controlling the transcription of a target nucleotide sequence into RNA when operably linked to a target nucleotide sequence. The promoter is located 5' (i.e., upstream) of the transcription start site of the target nucleotide sequence, controls the transcription of the target nucleotide sequence into RNA, and provides a site for RNA polymerase and other transcription factors to specifically bind to initiate transcription. The promoter comprises, for example, at least 10 kb, such as 5 kb or 2 kb, near the transcription start site. It may also comprise at least 1500 bp, preferably at least 1000 bp, near the transcription start site. In a further preferred embodiment, the promoter comprises at least 50 bp, such as at least 25 bp, near the transcription start site. The promoter does not comprise exon and / or intron regions or a 5' untranslated region. The promoter may, for example, be heterologous or homologous to the corresponding cell. A polynucleotide sequence is "heterologous" to an organism or a second polynucleotide sequence if the polynucleotide sequence is derived from a foreign bacterial species or, if from the same bacterial species, is modified from its original form. For example, a promoter operably linked to a heterologous coding sequence refers to a coding sequence from a different bacterial species than the one from which the promoter is derived, or, if from the same bacterial species, to a coding sequence not naturally associated with the promoter (e.g., a genetically engineered coding sequence or an allele from a different ecotype or species). Suitable promoters can be derived from genes of the host cell in which expression should occur or from a pathogen derived from such a host cell. The activity of a promoter can be assessed by, for example, operably linking a reporter gene to the promoter sequence to create a reporter construct, introducing the reporter construct into the genome of the cell, and detecting expression of the reporter gene (e.g., detecting the activity of mRNA, protein, or protein encoded by the reporter gene).

[0153] Purified: As used herein, the term "purified" refers to a molecule (nucleic acid or amino acid sequence) that is removed, isolated, or separated from its natural environment. A "substantially purified" molecule is at least 60% free, preferably at least 75% free, and more preferably at least 90% free from other components with which it is naturally associated. A purified nucleic acid sequence can be an isolated nucleic acid sequence.

[0154] Recombinant: The term "recombinant" with respect to nucleic acid molecules refers to nucleic acid molecules produced by recombinant DNA techniques. Recombinant nucleic acid molecules may also include molecules that do not themselves occur in nature but have been artificially modified, altered, mutated or otherwise manipulated. Preferably, a "recombinant nucleic acid molecule" is a non-naturally occurring nucleic acid molecule that differs in sequence from a naturally occurring nucleic acid molecule by at least one nucleic acid. A "recombinant nucleic acid molecule" may also include a "recombinant construct" comprising, preferably operably linked, the sequence of a nucleic acid molecule that does not occur naturally in that order. Preferred methods for producing such recombinant nucleic acid molecules may include cloning techniques, directed or non-directed mutagenesis, synthesis or recombination techniques.

[0155] Sense: The term "sense" is understood to mean a nucleic acid molecule having a sequence that is complementary or identical to a target sequence, such as a sequence that binds to a protein transcription factor and is involved in the expression of a given gene. According to a preferred embodiment, the nucleic acid molecule comprises a target gene and an element that allows the expression of said target gene.

[0156] Significant increase or decrease: for example, an increase or decrease in enzyme activity or gene expression that is greater than the error range inherent in the measurement technology, preferably an increase or decrease of about 2-fold or more, more preferably about 5-fold or more, most preferably about 10-fold or more, of the control enzyme activity or expression in control cells.

[0157] Small nucleic acid molecules: "Small nucleic acid molecules" are understood to be molecules consisting of nucleic acids or derivatives thereof, such as RNA or DNA. They can be double-stranded or single-stranded and are between about 15 bp and about 30 bp, for example, between 15 bp and 30 bp, more preferably between about 19 bp and about 26 bp, for example, between 19 bp and 26 bp, even more preferably between about 20 bp and about 25 bp, for example, between 20 bp and 25 bp. In a particularly preferred embodiment, the oligonucleotide is between about 21 bp and about 24 bp, for example, between 21 bp and 24 bp. In a most preferred embodiment, the small nucleic acid molecules are between about 21 bp and about 24 bp, for example, 21 bp and 24 bp.

[0158] Substantially complementary: In its broadest sense, the term "substantially complementary", when used herein with respect to nucleotide sequences related to a reference or target nucleotide sequence, means nucleotide sequences having a percentage of identity of at least 60%, more desirably at least 70%, more desirably at least 80% or 85%, preferably at least 90%, more preferably at least 93%, still more preferably at least 95% or 96%, still more preferably at least 97% or 98%, still more preferably at least 99% or most preferably 100% between the substantially complementary nucleotide sequence and the exact complement of the reference or target nucleotide sequence (the latter being equivalent to the term "identical" in this context). Preferably, identity to the reference sequence is assessed over a length of at least 19 nucleotides, preferably at least 50 nucleotides, more preferably over the entire length of the nucleic acid sequence (if not otherwise specified below). Sequence comparisons are performed based on the algorithm of Needleman and Wunsch (Needleman and Wunsch (1970) J Mol. Biol. 48: 443-453; as defined above) using the default GAP analysis in the University of Wisconsin GCG, SEQWEB GAP application. A nucleotide sequence that is "substantially complementary" to a reference nucleotide sequence hybridizes to the reference nucleotide sequence under low stringency conditions, preferably medium stringency conditions, and most preferably high stringency conditions (as defined above).

[0159] As used herein, "target site" means a location in a genome where a double-strand break or one or a pair of single-strand breaks (nicks) are induced using recombinant technology such as Zn fingers, TALENs, restriction endonucleases, homing endonucleases, RNA-guided nucleases, RNA-guided nickases (such as CRISPR / Cas nucleases or nickases), etc.

[0160] Transgene: As used herein, the term "transgene" refers to any nucleic acid sequence that is introduced into the genome of a cell through experimental manipulation. A transgene can be an "endogenous DNA sequence" or a "heterologous DNA sequence" (i.e., "foreign DNA"). The term "endogenous DNA sequence" refers to a nucleotide sequence that naturally occurs in the cell into which it is introduced, provided that it does not contain some modification relative to the naturally occurring sequence (e.g., point mutations, the presence of a selectable marker gene, etc.).

[0161] Transgenic: The term transgenic when referring to an organism means transformed, preferably stably transformed, with a recombinant DNA molecule, preferably comprising a suitable promoter operably linked to a DNA sequence of interest.

[0162] Vector: As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid molecule to which it is linked. One type of vector is a genomic integrating vector or "integrating vector," which can be integrated into the chromosomal DNA of a host cell. Another type of vector is an episomal vector, a nucleic acid molecule capable of extrachromosomal replication. Vectors that are capable of directing the expression of genes to which they are operably linked are referred to herein as "expression vectors." In this specification, "plasmid" and "vector" are used interchangeably unless the context clearly indicates otherwise. Expression vectors designed for producing RNA as described herein in vitro or in vivo can contain sequences recognized by any RNA polymerase, including mitochondrial RNA polymerase, RNA pol I, RNA pol II, and RNA pol III. These vectors can be used to transcribe the desired RNA molecules in cells according to the present invention.

[0163] Wild-type: With respect to an organism, polypeptide or nucleic acid sequence, the term "wild-type," "native," or "naturally derived" means that the organism is naturally occurring or obtainable in at least one naturally occurring organism that has not been altered, mutated, or otherwise manipulated by man. BRIEF DESCRIPTION OF THE DRAWINGS

[0164] Figure 1 Shown are the total specific phytase activities in U / g protein of the different phytase producing strains.

[0165] A: UV18#100f Δpyr5 Δalp1 Δku70 cbh1::Pchi-phytase-Tcbh1 (overexpression of the phytase gene driven by the chi1 promoter; expression cassette integrated at the cbh1 locus)

[0166] B: UV18#100f Δpyr5 Δalp1 Δku70 cbh1::PTEF-phytase-Tcbh1 (overexpression of the phytase gene driven by the TEF promoter; expression cassette integrated at the cbh1 locus)

[0167] C: UV18#100f Δpyr5 Δalp1 Δku70 MYCTH:XP_003662453.1::Pchi-phytase-Tcbh1 (overexpression of the phytase gene driven by the chi1 promoter; expression cassette integrated at the MYCTH:XP_003662453.1 locus)

[0168] D: UV18#100f Δpyr5 Δalp1 Δku70 MYCTH:XP_003662453.1::PTEF-phytase-Tcbh1 (overexpression of the phytase gene driven by the TEF promoter; expression cassette integrated at the MYCTH:XP_003662453.1 locus).

[0169] Figure 2 Shown are the total specific glucanase activities in U / g protein of the different glucanase producing strains.

[0170] A: UV18#100f Δpyr5 Δalp1 Δku70 cbh1::Pchi-glucanase-Tcbh1 (overexpression of the glucanase gene driven by the chi1 promoter; expression cassette integrated at the cbh1 locus)

[0171] B: UV18#100f Δpyr5 Δalp1 Δku70 cbh1::PTEF-glucanase-Tcbh1 (overexpression of the glucanase gene driven by the TEF promoter; expression cassette integrated at the cbh1 locus)

[0172] C: UV18#100fΔpyr5Δalp1Δku70 MYCTH:XP_003662453.1::Pchi-glucanase-Tcbh1 (overexpression of the glucanase gene driven by the chi1 promoter; expression cassette integrated at the MYCTH:XP_003662453.1 locus)

[0173] D: UV18#100f Δpyr5 Δalp1 Δku70 MYCTH:XP_003662453.1::PTEF-glucanase-Tcbh1 (overexpression of the glucanase gene driven by the TEF promoter; expression cassette integrated at the MYCTH:XP_003662453.1 locus).

[0174] Figure 3 Shown are the total specific phytase activities in U / g protein of the different phytase producing strains.

[0175] A: UV18#100f Δpyr5 Δalp1 Δku70 cbh1::Pchi-phytase-Tcbh1 (overexpression of the phytase gene driven by the chi1 promoter; expression cassette integrated at the cbh1 locus)

[0176] B: UV18#100f Δpyr5 Δalp1 Δku70 cbh1::Pchi-phytase-Tcbh1 MYCTH:XP_003662453.1::nat1 (overexpression of the phytase gene driven by the chi1 promoter; expression cassette integrated at the cbh1 locus; MYCTH:XP_003662453.1 locus replaced by the nat1 selection marker cassette)

[0177] C: UV18#100fΔpyr5Δalp1Δku70 MYCTH:XP_003662453.1::Pchi-phytase-Tcbh1 (overexpression of the phytase gene driven by the chi1 promoter; expression cassette integrated at the MYCTH:XP_003662453.1 locus)

[0178] D: UV18#100fΔpyr5Δalp1Δku70 MYCTH:XP_003662453.1::Pchi-phytase-Tcbh1 cbh1::nat1 (overexpression of the phytase gene driven by the chi1 promoter; expression cassette integrated at the MYCTH:XP_003662453.1 locus; the cbh1 locus replaced by the nat1 selection marker cassette).

[0179] Figure 4 Shown are the total specific phytase activities in U / g protein of the different phytase producing strains.

[0180] A: UV18#100f Δpyr5 Δalp1 Δku70 cbh1::Pchi-phytase-Tcbh1 (overexpression of the phytase gene driven by the chi1 promoter; expression cassette integrated at the cbh1 locus)

[0181] B: UV18#100f Δpyr5 Δalp1 Δku70 MYCTH:XP_003662453.1::Pchi-phytase-Tcbh1 (overexpression of the phytase gene driven by the chi1 promoter; expression cassette integrated at the MYCTH:XP_003662453.1 locus)

[0182] C: UV 18#100f Δpyr5 Δalp1 Δku70 Δpep4 Δprt1 Δprt2 Δprt3 Δprt4 Δlam2 Δtre1 ΔmycA Δgla1 Δprt5 Δprt6 cbh1::Pchi-phytase-Tcbh1 (overexpression of the phytase gene driven by the chi1 promoter; expression cassette integrated at the cbh1 locus)

[0183] C: UV18#100f Δpyr5 Δalp1 Δku70 Δpep4 Δprt1 Δprt2 Δprt3 Δprt4 Δlam2 Δtre1 ΔmycA Δgla1 Δprt5 Δprt6 MYCTH:XP_003662453.1::Pchi-phytase-Tcbh1 (overexpression of the phytase gene driven by the chi1 promoter; expression cassette integrated at the MYCTH:XP_003662453.1 locus).

[0184] Figure 5 The genomic organization of the genomic integration site of the present invention in Myceliophthora thermophila C1 is shown.

[0185] Example

[0186] Chemicals and common methods

[0187] Unless otherwise stated, the cloning procedure carried out for the purposes of the present invention is as described, including restriction enzyme digestion, agarose gel electrophoresis, nucleic acid purification, nucleic acid connection, transformation, selection and cultivation of bacterial cells (Sambrook et al., 1989). Using Sanger technology (Sanger et al., 1977), the sequence analysis of recombinant DNA is carried out with a laser fluorescence DNA sequencer (Applied Biosystems, Foster City, CA, USA). Unless otherwise stated, chemicals and reagents are available from Sigma Aldrich (Sigma Aldrich, St.Louis, USA), Promega (Madison, WI, USA), Duchefa (Haarlem, The Netherlands) or Invitrogen (Carlsbad, CA, USA). Restriction endonucleases are from New England Biolabs (Ipswich, MA, USA). Oligonucleotides are synthesized by Integrated DNA Technologies (Coralville, IA, USA).

[0188] Example 1

[0189] Transformation of Thermomycetes

[0190] Several methods for transforming Myceliophthora thermophila are described in the literature (WO 00 / 20555, US 2012 / 0005812, Verdoes et al. (2007) Industrial Biotechnology 3(1): 48-57).

[0191] The cells were grown in a 100 ml shake flask at a concentration of 0.7-1 × 10 5 Protoplasts of the thermophilic Myceliophthora strain were prepared by inoculating 25 ml of a pre-culture of standard fungal growth medium with spores / ml and incubating at 37°C and 250 rpm for 24 hours. A main culture was prepared by inoculating 100 ml of standard fungal growth medium with 20 ml of the pre-culture in a 500 ml shake flask and incubating at 37°C and 250 rpm for 24 hours. Mycelia were harvested by filtration through a sterile cell strainer (VWR) and washed with 100 ml of 2000 mosmol / L NaCl / CaCl2 (0.6 M NaCl and 0.27 M CaCl2 * HO). 1 g of washed mycelia was transferred to a 100 ml flask. With mycelium and 150mg VinoTaste Pro solution (at 2000mosmol / L NaCl / CaCl2, be 2.5mg / ml) and 10mg Yatalase solution (at 2000mosmol / L NaCl / CaCl2, be 0.625mg / ml) and 10ml 2000mosmol / LNaCl / CaCl2 mixing.Mycelium suspension is hatched 50min-70min at 30 ℃ and 70rpm place, until visible protoplastis under the microscope.Filter into the 50ml sterile tube and gather in the crops protoplastis by sterile cell strainer.After 25ml ice-cold STC solution (1.2M sorbitol, 50mM CaCl2, 35mM NaCl, 10mM Tris / HCl pH7.5) is added in the filtrate, by centrifugal (1200 × g, 10min, 4 ℃) gather in the crops protoplastis. The protoplasts were washed again in 50 ml STC and resuspended in 0.5 ml-1.2 ml STC.

[0192] For transformation, 5 μg-10 μg linearized DNA, 1 tbsp 0.5M aurintricarboxylic acid (ATA) and 100 μl protoplast suspension were mixed and incubated on ice for 30 min-40 min. Then 1.7 mL PEG solution (60% PEG4000 [polyethylene glycol], 50 mM CaCl 2 , 35 mM NaCl, 10 mM Tris / HCl pH 7.5) was added and mixed gently. After incubation at 4° C. for 30 min, the tube was filled with 11 ml STC solution, centrifuged (900 × g, 10 min, 4° C.), and the supernatant was discarded. The precipitate was resuspended in the remaining STC and plated on a selective culture medium plate known in the art. After the plate was incubated at 37° C. for 3-6 days, the transformant was picked and re-streaked on the selective culture medium.

[0193] Selective medium plates

[0194] If the amdS gene is used as a selectable marker, select positive transformants using an enriched minimal medium containing uridine and uracil as well as acetamide (sucrose is added only in the case of protoplast plating):

[0195]

[0196] If selecting for clones that have lost acetamidase function, use an enriched minimal medium containing uridine and uracil as well as fluoroacetamide:

[0197]

[0198]

[0199] If the pyr5 gene is used as a selection marker, select positive transformants using enriched minimal medium without uridine and uracil (sucrose is added only in case of protoplast plating):

[0200]

[0201] If the nat1 gene is used as an additional selection marker, select positive transformants using enriched minimal medium without uridine and uracil and containing nourseothricin (clonNAT) (sucrose is added only in case of protoplast plating):

[0202]

[0203] Example 2

[0204] Generation of expression plasmids for integration at the cbh1 locus

[0205] Expression construct with phytase and chi1 promoter

[0206] A synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) (SEQ ID NO: 1) encoding a synthetic phytase from a bacterial source (disclosed in WO 2012 / 143862; as phytase PhV-99; SEQ ID NO. 2) was used to construct a phytase expression plasmid. To secrete the phytase, a signal sequence encoding a signal peptide derived from Myceliophthora thermophila was added to the mature sequence of the phytase. A promoter sequence amplified from the upstream region of the Chi 1 gene encoding Myceliophthora thermophila (MYCTH: XP_003663544.1) and a terminator sequence amplified from the downstream region of the Cbh1 gene encoding Myceliophthora thermophila (MYCTH: XP_003660789.1) were used as regulatory elements to drive phytase expression. Using standard cloning techniques, the expression plasmid MT2701 (SEQ ID NO: 7) was constructed based on the standard E. coli cloning vectors MT940 (SEQ ID NO: 4) and MT944 (SEQ ID NO: 5). Plasmid MT940 consists of the pMB1 replication origin, kan resistance, the pyr5 gene, upstream and downstream regions of the Cbh1 gene encoding Myceliophthora thermophila for homologous recombination, and lacZ for blue / white spot selection. Plasmid MT944 consists of the pUC replication origin, kan resistance, and lacZ for blue / white spot selection. The expression plasmid contains the upstream region of cbh1 (bases 1913-3345), the chi1 promoter sequence (bases 3350-5162), the phytase containing signal sequence (bases 5165-6478), the cbh1 terminator sequence (bases 6483-6689), and the downstream region of cbh1 (bases 7981-9610).

[0207] The plasmid was digested with NotI to remove the vector backbone and the fragment containing the phytase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was subsequently used for transformation.

[0208] Expression construct with phytase and TEF1 promoters

[0209] A synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) (SEQ ID NO: 1) encoding a synthetic phytase from a bacterial source (disclosed in WO 2012 / 143862; as phytase PhV-99; SEQ ID NO. 2) was used to construct a phytase expression plasmid. To secrete the phytase, a signal sequence encoding a signal peptide derived from Myceliophthora thermophila was added to the mature sequence of the phytase. A promoter sequence amplified from the upstream region of the TEF1 encoding gene (MYCTH: XP_003660173.1) of Myceliophthora thermophila and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH: XP_003660789.1) of Myceliophthora thermophila were used as regulatory elements to drive phytase expression. Using standard cloning techniques, expression plasmid MT2703 (SEQ ID NO: 8) was constructed based on the standard E. coli cloning vectors MT940 (SEQ ID NO: 4) and MT944 (SEQ ID NO: 5). Plasmid MT940 consists of the pMB1 replication origin, kan resistance, the pyr5 gene, upstream and downstream regions of the Cbh1 gene encoding Myceliophthora thermophila for homologous recombination, and lacZ for blue / white spot selection. Plasmid MT944 consists of the pUC replication origin, kan resistance, and lacZ for blue / white spot selection. The expression plasmid contains the upstream region of cbh1 (bases 1913-3345), the TEF1 promoter sequence (bases 3350-5847), the phytase containing signal sequence (bases 5850-7163), the cbh1 terminator sequence (bases 7168-7374), and the downstream region of cbh1 (bases 8666-10295).

[0210] The plasmid was digested with NotI to remove the vector backbone and the fragment containing the phytase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was subsequently used for transformation.

[0211] Expression construct with glucanase and chi1 promoter

[0212] The synthetic gene encoding the glucanase from Rosalia emersonii (GeneArt, ThermoFisherScientific Inc., USA) (SEQ ID NO.3) was used to construct the glucanase expression plasmid. In order to secrete the glucanase, a signal sequence encoding the signal peptide derived from Myceliophthora thermophila was added to the mature sequence of the glucanase. A promoter sequence amplified from the upstream region of the Chi 1 coding gene (MYCTH: XP_003663544.1) of Myceliophthora thermophila and a terminator sequence amplified from the downstream region of the Cbh1 coding gene (MYCTH: XP_003660789.1) of Myceliophthora thermophila were used as regulatory elements to drive glucanase expression. Using standard cloning techniques, the expression plasmid MT2705 (SEQ ID NO: 9) was constructed based on the E. coli standard cloning vectors MT940 (SEQ ID NO: 4) and MT944 (SEQ ID NO: 5). Plasmid MT940 consists of the pMB1 origin of replication, kan resistance, the pyr5 gene, upstream and downstream regions of the Cbh1 gene encoding Myceliophthora thermophila for homologous recombination, and lacZ for blue / white spot selection. Plasmid MT944 consists of the pUC origin of replication, kan resistance, and lacZ for blue / white spot selection. The expression plasmid contains the upstream region of cbh1 (bases 1913-3345), the chi1 promoter sequence (bases 3350-5162), the glucanase containing signal sequence (bases 5165-6151), the cbh1 terminator sequence (bases 6156-6362), and the downstream region of cbh1 (bases 7654-9283).

[0213] The plasmid was digested with NotI to remove the vector backbone, and the fragment containing the glucanase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was subsequently used for transformation.

[0214] Expression construct with glucanase and TEF1 promoters

[0215] A synthetic gene encoding a glucanase from Rosalia emersonii (GeneArt, ThermoFisher Scientific Inc., USA) (SEQ ID NO. 3) was used to construct a glucanase expression plasmid. To secrete the glucanase, a signal sequence encoding a signal peptide derived from Myceliophthora thermophila was added to the mature sequence of the glucanase. A promoter sequence amplified from the upstream region of the TEF1 coding gene (MYCTH: XP_003663544.1) from Myceliophthora thermophila and a terminator sequence amplified from the downstream region of the Cbh1 coding gene (MYCTH: XP_003660789.1) from Myceliophthora thermophila were used as regulatory elements to drive glucanase expression. Using standard cloning techniques, expression plasmid MT2707 (SEQ ID NO: 10) was constructed based on the standard E. coli cloning vectors MT940 (SEQ ID NO: 4) and MT944 (SEQ ID NO: 5). Plasmid MT940 consists of the pMB1 origin of replication, kan resistance, the pyr5 gene, upstream and downstream regions of the Cbh1 gene from Myceliophthora thermophila for homologous recombination, and lacZ for blue / white spot selection. Plasmid MT944 consists of the pUC origin of replication, kan resistance, and lacZ for blue / white spot selection. The expression plasmid contains the upstream region of cbh1 (bases 1913-3345), the TEF1 promoter sequence (bases 3350-5847), the glucanase containing signal sequence (bases 5850-6836), the cbh1 terminator sequence (bases 6841-7047), and the downstream region of cbh1 (bases 8339-9968).

[0216] The plasmid was digested with NotI to remove the vector backbone, and the fragment containing the glucanase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was subsequently used for transformation.

[0217] The plasmid was digested with NotI to remove the vector backbone, and the fragment containing the glucanase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was subsequently used for transformation.

[0218] Example 3

[0219] Generation of expression plasmids for integration at the MYCTH:XP 003662453.1 locus

[0220] Expression construct with phytase and chi1 promoter

[0221] A synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) (SEQ ID NO: 1) encoding a synthetic phytase from a bacterial source (disclosed in WO 2012 / 143862; as phytase PhV-99; SEQ ID NO. 2) was used to construct a phytase expression plasmid. To secrete the phytase, a signal sequence encoding a signal peptide derived from Myceliophthora thermophila was added to the mature sequence of the phytase. A promoter sequence amplified from the upstream region of the Chil-encoding gene (MYCTH: XP_003663544.1) from Myceliophthora thermophila and a terminator sequence amplified from the downstream region of the Cbh1-encoding gene (MYCTH: XP_003660789.1) from Myceliophthora thermophila were used as regulatory elements to drive phytase expression. Using standard cloning techniques, expression plasmid MT2709 (SEQ ID NO: 11) was constructed based on the E. coli standard cloning vectors MT944 (SEQ ID NO: 5) and MT1266 (SEQ ID NO: 6). Both plasmids consist of a pUC origin of replication, kan resistance, and lacZ for blue / white spot screening. The expression plasmid contains the upstream region of MYCTH: XP_003662453.1 (bases 1913-3412), the chi1 promoter sequence (bases 3417-5229), the phytase containing signal sequence (bases 5232-6545), the cbh1 terminator sequence (bases 6550-6756), and the downstream region of MYCTH: XP_003662453.1 (bases 8048-9518).

[0222] The plasmid was digested with NotI to remove the vector backbone and the fragment containing the phytase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was subsequently used for transformation.

[0223] Expression construct with phytase and TEF1 promoters

[0224] A synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) (SEQ ID NO: 1) encoding a synthetic phytase from a bacterial source (disclosed in WO 2012 / 143862; as phytase PhV-99; SEQ ID NO. 2) was used to construct a phytase expression plasmid. To secrete the phytase, a signal sequence encoding a signal peptide derived from Myceliophthora thermophila was added to the mature sequence of the phytase. A promoter sequence amplified from the upstream region of the TEF1 encoding gene (MYCTH: XP_003663544.1) from Myceliophthora thermophila and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH: XP_003660789.1) from Myceliophthora thermophila were used as regulatory elements to drive phytase expression. Using standard cloning techniques, the expression plasmid MT2711 (SEQ ID NO: 12) was constructed based on the standard E. coli cloning vectors MT944 (SEQ ID NO: 5) and MT1266 (SEQ ID NO: 6). Both plasmids consist of a pUC origin of replication, kan resistance, and lacZ for blue / white spot selection. The expression plasmid contains the upstream region of MYCTH: XP_003662453.1 (bases 1913-3412), the TEF1 promoter sequence (bases 3417-5914), the phytase containing signal sequence (bases 5917-7230), the cbh1 terminator sequence (bases 7235-7441), and the downstream region of MYCTH: XP_003662453.1 (bases 8733-10203).

[0225] The plasmid was digested with NotI to remove the vector backbone and the fragment containing the phytase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was subsequently used for transformation.

[0226] Expression construct with glucanase and chi1 promoter

[0227] The synthetic gene encoding the glucanase from Rosalia emersonii (GeneArt, ThermoFisherScientific Inc., USA) (SEQ ID NO.3) was used to construct the glucanase expression plasmid. In order to secrete the glucanase, a signal sequence encoding the signal peptide derived from Myceliophthora thermophila was added to the mature sequence of the glucanase. A promoter sequence amplified from the upstream region of the Chi 1 coding gene (MYCTH: XP_003663544.1) of Myceliophthora thermophila and a terminator sequence amplified from the downstream region of the Cbh1 coding gene (MYCTH: XP_003660789.1) of Myceliophthora thermophila were used as regulatory elements to drive glucanase expression. Using standard cloning techniques, expression plasmid MT2713 (SEQ ID NO: 13) was constructed based on the E. coli standard cloning vectors MT944 (SEQ ID NO: 5) and MT1266 (SEQ ID NO: 6). Both plasmids consist of a pUC origin of replication, kan resistance, and lacZ for blue / white selection. The expression plasmid contains the upstream region of MYCTH:XP_003662453.1 (bases 1913-3412), the chi1 promoter sequence (bases 3417-5229), the glucanase containing signal sequence (bases 5232-6218), the cbh1 terminator sequence (bases 6223-6429), and the downstream region of MYCTH:XP_003662453.1 (bases 7721-9191).

[0228] The plasmid was digested with NotI to remove the vector backbone, and the fragment containing the glucanase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was subsequently used for transformation.

[0229] Expression construct with glucanase and TEF1 promoters

[0230] A synthetic gene encoding a glucanase from Rosalia emersonii (GeneArt, ThermoFisherScientific Inc., USA) (SEQ ID NO. 3) was used to construct a glucanase expression plasmid. To secrete the glucanase, a signal sequence encoding a signal peptide derived from Myceliophthora thermophila was added to the mature sequence of the glucanase. A promoter sequence amplified from the upstream region of the TEF1 coding gene (MYCTH: XP_003663544.1) from Myceliophthora thermophila and a terminator sequence amplified from the downstream region of the Cbh1 coding gene (MYCTH: XP_003660789.1) from Myceliophthora thermophila were used as regulatory elements to drive glucanase expression. Using standard cloning techniques, expression plasmid MT2715 (SEQ ID NO: 14) was constructed based on the E. coli standard cloning vectors MT944 (SEQ ID NO: 5) and MT1266 (SEQ ID NO: 6). Both plasmids consist of a pUC origin of replication, kan resistance, and lacZ for blue / white selection. The expression plasmid contains the upstream region of MYCTH:XP_003662453.1 (bases 1913-3412), the TEF1 promoter sequence (bases 3417-5914), the glucanase containing signal sequence (bases 5917-6903), the cbh1 terminator sequence (bases 6908-7114), and the downstream region of MYCTH:XP_003662453.1 (bases 8406-9876).

[0231] The plasmid was digested with NotI to remove the vector backbone, and the fragment containing the glucanase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was subsequently used for transformation.

[0232] Example 4

[0233] Generation of plasmids for deletion of the cbh1 locus and the MYCTH: XP_003662453.1 locus Targeting MYCTH: Deletion construct for the XP 003662453.1 locus

[0234] The synthetic nourseothricin acetyltransferase gene (nat1) selection marker cassette (SEQ ID NO: 15) was fused to the upstream and downstream regions of the MYCTH: XP_003662453.1 locus from Myceliophthora thermophila to carry out homologous recombination. Standard cloning techniques were used to construct the KO plasmid MT2599 (SEQ ID NO: 16) based on the E. coli standard cloning vector MT944 (SEQ ID NO: 5). It consists of a pUC origin of replication, kan resistance, and lacZ for blue / white spot screening.

[0235] The plasmid was digested with NotI to remove the vector backbone, and the fragment containing the nat1 selectable marker cassette was isolated from an agarose gel. Only the isolated DNA fragment was subsequently used for transformation.

[0236] Deletion constructs targeting the cbh1 locus

[0237] A synthetic nourseothricin acetyltransferase gene (nat1) selection marker cassette (SEQ ID NO: 15) was fused to the upstream and downstream regions of the Cbh1 coding gene from Myceliophthora thermophila for homologous recombination. KO plasmid MT2699 (SEQ ID NO: 17) was constructed based on the E. coli standard cloning vector MT944 (SEQ ID NO: 5) using standard cloning techniques. It consists of a pUC origin of replication, kan resistance, and lacZ for blue / white spot selection.

[0238] The plasmid was digested with NotI to remove the vector backbone, and the fragment containing the nat1 selectable marker cassette was isolated from an agarose gel. Only the isolated DNA fragment was subsequently used for transformation.

[0239] Example 5

[0240] Mycelium thermophila strain UV18#100f Δpyr5 Δalp1 Δku70 Δpep4 Δprt1 Δprt2 Δ prt3 Generation of Δprt4 Δlam2 Δtre1 ΔmycA Δgla1 Δprt5 Δprt6

[0241] As described in Example 1, the Myceliophthora thermophila host strain UV1 8#100f Δpyr5 Δalp 1 Δku70 was co-transformed with two deletion fragments isolated from plasmids MT185 (SEQ ID No. 18) and MT186 (SEQ ID No. 19) to delete the gene pep4. Enriched minimal medium for amdS selection was used for incubation. After restreaking on enriched minimal medium with fluoroacetamide for counterselection, transformants were analyzed by PCR for correct integration of the deletion cassette into the target locus, loss of the intact gene, and removal of the amdS marker. The resulting strain UV18#100f Δpyr5 Δalp1 Δku70 Δpep4 was modified in successive transformation rounds and the following genes were deleted in the same manner; prt1 gene: using two deletion fragments isolated from plasmids MT111 (SEQ ID No. 20) and MT112 (SEQ ID NO. 21), the amdS marker was removed to obtain the strain UV18#100f Δpyr5 Δalp1 Δku70 Δpep4 Δprt1; prt2 gene: using two deletion fragments isolated from plasmids MT368 (SEQ ID No. 22) and MT453 (SEQ ID NO. 23), the amdS marker was removed to obtain the strain UV18#100f Δpyr5 Δalp1 Δku70 Δpep4 Δprt1 Δprt2; prt3 gene: using two deletion fragments isolated from plasmids MT392 (SEQ ID No. 24) and MT393 (SEQ ID The strain UV18#100f Δpyr5 Δalp1 Δku70 Δpep4 Δprt1 Δprt2 Δprt3 was obtained after removing the amdS marker; the prt4 gene: two deletion fragments isolated from plasmids MT386 (SEQ ID No. 26) and MT649 (SEQ ID NO. 27) were used, and the amdS marker was removed to obtain the strain UV18#100f Δpyr5 Δalp1 Δku70 Δpep4 Δprt1 Δprt2 Δprt3 Δprt4; the lam2 gene: two deletion fragments isolated from plasmids MT706 (SEQ ID No. 28) and MT707 (SEQ ID NO. 29) were used, and the amdS marker was removed to obtain the strain UV18#100f Δpyr5 Δalp1 Δku70 Δpep4 Δprt1 Δprt2 Δprt3 Δprt4 Δlam2; tre1 gene: The Δprt4 Δlam2; tre1 gene was obtained from plasmids MT710 (SEQ ID No. 30) and MT711 (SEQ ID NO.31) were used to remove the amdS marker and obtain the strain UV18#100f Δpyr5 Δalp1 Δku70 Δpep4 Δprt1 Δprt2 Δprt3 Δprt4 Δlam2 Δtre1; the mycA gene: two deletion fragments isolated from plasmids MT738 (SEQ ID No. 32) and MT739 (SEQ ID NO. 33) were used to remove the amdS marker and obtain the strain UV18#100f Δpyr5 Δalp1 Δku70 Δpep4 Δprt1 Δprt2 Δprt3 Δprt4 Δlam2 Δtre1 ΔmycA; the gla1 gene: two deletion fragments isolated from plasmids MT605 (SEQ ID No. 34) and MT606 (SEQ ID NO. 35) were used to remove the amdS marker and obtain the strain UV18#100f Δpyr5 Δalp1 Δku70 Δpep4 Δprt1 Δprt2 Δprt3 Δprt4 Δlam2 Δtre1 ΔmycA Δgla1; prt5 gene: two deletion fragments isolated from plasmids MT396 (SEQ ID No.36) and MT397 (SEQ ID NO.37) were used, and the amdS marker was removed to obtain strain UV18#100f Δpyr5Δalp1 Δku70 Δpep4 Δprt1 Δprt2 Δprt3 Δprt4 Δlam2 Δtre1 ΔmycA Δgla1Δprt5; prt6 gene: two deletion fragments isolated from plasmids MT388 (SEQ ID No.38) and MT389 (SEQ ID NO.39) were used. Finally, the positively tested clone was designated as UV18#100f Δpyr5 Δalp1 Δku70 Δpep4 Δprt1 Δprt2 Δprt3 Δprt4 Δlam2 Δtre1 ΔmycA Δgla1 Δprt5 Δprt6.

[0242] Example 6

[0243] A strain of Myceliophthora thermophila having a phytase or glucanase expression construct integrated at the cbh1 locus produce

[0244] To express phytase, the Myceliophthora thermophila host strain UV18#100f Δpyr5 Δalp1 Δku70 from the C1 lineage (construct described in detail in WO 2017 / 093450), which has uracil auxotrophy, reduced protease activity and an impaired non-homologous end joining (NHEJ) repair system, was transformed with NotI-digested and isolated phytase expression constructs from plasmids MT2701 (SEQ ID NO: 7) and MT2703 (SEQ ID NO: 8) as described in Example 1 (see Example 2). As described in Example 1, the Myceliophthora thermophila host strain UV18#100fΔpyr5Δalp1Δku70Δpep4Δprt1Δprt2Δprt3Δprt4Δlam2Δtre1ΔmycAΔgla1Δprt5Δprt6 from the C1 lineage, which has a uracil auxotrophy, reduced protease activity and an impaired non-homologous end joining (NHEJ) repair system, was transformed with a NotI-digested and isolated phytase expression construct from plasmid MT2701 (SEQ ID NO: 7) (see Example 2). To express the glucanase, the thermophilic Myceliophthora host strain UV18#100f Δpyr5 Δalp1 Δku70 (a construct described in detail in WO2017 / 093450) was transformed with NotI digested and isolated glucanase expression constructs (see Example 2) from plasmids MT2705 (SEQ ID NO: 9) and MT2707 (SEQ ID NO: 10) as described in Example 1. Transformants were incubated at 37°C on enriched minimal medium for pyr5 selection for 3-6 days to select for restored uracil prototrophy by complementing the pyr5 deletion with a pyr5 marker as known in the art. Colonies were restreaked and checked for integration of the phytase or glucanase expression cassette using PCR using primer pairs specific for the phytase or glucanase expression cassette and the cbh1 locus as known in the art. Transformants that tested positive for the expression of either the phytase or glucanase construct at the cbh1 locus were selected for further characterization. Phytase-producing strain UV18#100f Δpyr5 Δalp1 Δku70 cbh1::Pchi-phytase-Tcbh1 were transformed with NotI digestion and the isolated nat1 selection marker cassette (see Example 4) from plasmid MT2599 (SEQ ID NO: 16) as described in Example 1. Transformants were incubated at 37°C on enriched minimal medium for nat1 selection for 3-6 days to select for resistance to nourseothricin (clonNAT) by integrating the nat1 marker known in the art.Colonies were restreaked and checked for integration of the nat1 selectable marker and deletion of the MYCTH:XP_003662453.1 locus using PCR using primer pairs known in the art that are specific for the nat1 and MYCTH:XP_003662453.1 loci.

[0245] Example 7

[0246] with a phytase or glucanase expression construct integrated at the MYCTH: XP_003662453.1 locus Production of Thermomycetic Pseudomonas thermophila strains

[0247] To express phytase, the Myceliophthora thermophila host strain UV18#100f Δpyr5 Δalp1 Δku70 from the C1 lineage (construct described in detail in WO 2017 / 093450), which has uracil auxotrophy, reduced protease activity and an impaired non-homologous end joining (NHEJ) repair system, was transformed with NotI-digested and isolated phytase expression constructs from plasmids MT2709 (SEQ ID NO: 11) and MT2711 (SEQ ID NO: 12) as described in Example 1 (see Example 3). As described in Example 1, the Myceliophthora thermophila host strain UV18#100f Δpyr5 Δalp1 Δku70 Δpep4 Δprt1 Δprt2 Δprt3 Δprt4 Δlam2 Δtre1 ΔmycA Δgla1 Δprt5 Δprt6 from the C1 lineage, which has a uracil auxotrophy, reduced protease activity and an impaired non-homologous end joining (NHEJ) repair system, was transformed with a NotI-digested and isolated phytase expression construct from plasmid MT2709 (SEQ ID NO: 11) (see Example 3). To express the glucanase, the thermophilic Myceliophthora host strain UV18#100f Δpyr5 Δalp1 Δku70 (a construct described in detail in WO2017 / 093450) was transformed with NotI digested and isolated glucanase expression constructs (see Example 3) from plasmids MT2713 (SEQ ID NO: 13) and MT2715 (SEQ ID NO: 14) as described in Example 1. Transformants were incubated at 37°C on enriched minimal medium for pyr5 selection for 3-6 days to select for restored uracil prototrophy by complementing the pyr5 deletion with a pyr5 marker as known in the art. Colonies were restreaked and checked for integration of the phytase or glucanase expression cassette using PCR using primer pairs specific for the phytase or glucanase expression cassette and the MYCTH: XP_003662453.1 locus as known in the art. Transformants that tested positive for the phytase or glucanase expression construct at the MYCTH: XP_003662453.1 locus were selected for further characterization. Phytase producing strain UV18#100f Δpyr5 Δalp1 Δku70 MYCTH: XP_003662453.1::Pchi-phytase-Tcbh1 was transformed with NotI digested and isolated nat1 selection marker cassette (see Example 4) from plasmid MT2699 (SEQ ID NO: 17) as described in Example 1.Transformants were incubated at 37° C. for 3-6 days on enriched minimal medium for nat1 selection to select for resistance to nourseothricin (clonNAT) through integration of the nat1 marker known in the art. Colonies were restreaked and checked for integration of the nat1 selectable marker and deletion of the cbh1 locus using PCR using primer pairs known in the art specific for the nat1 and cbh1 loci.

[0248] Example 8

[0249] Phytase activity assay

[0250] Phytase activity is measured in microtiter plates. The supernatant containing phytase is diluted in reaction buffer (250mM sodium acetate, 1mM CaCl , 0.01% Tween 20, pH 5.5). 10 μ l enzyme solution and 140 μ l substrate solution (6mM sodium phytate (Sigma P3168) in reaction buffer) are hatched at 37 ℃ for 1 h. By adding 150 μ l trichloroacetic acid solution (15% w / w) quenching reaction. In order to detect the released phosphate, the reaction solution of 20 μ l quenching is used 280 μ l freshly prepared color reagent (60mM L-ascorbic acid (Sigma A7506), 2.2mM ammonium molybdate tetrahydrate, 325mM H SO ) process, and hatch 20min at 37 ℃, measure the absorption at 820nm place subsequently. For the blank, the substrate buffer was incubated at 37°C by itself, and only after quenching with trichloroacetic acid was 10 μl of enzyme sample added. The color reaction was performed similarly to the rest of the measurement. The amount of released phosphate was determined using a calibration curve using the color reaction with phosphate solutions of known concentrations.

[0251] Example 9

[0252] Glucanase activity assay

[0253] The activity of glucanase was determined using the DNS method in a microtiter plate. The supernatant containing glucanase was diluted in a reaction buffer (100 mM sodium acetate, 0.005% Tween 20, pH 4.5). Low-viscosity carboxymethyl cellulose (CMC) was used as a substrate and dissolved in a reaction buffer at a concentration of 4% under stirring while heating. PCR cycler was used to mix 35 μl of the diluted enzyme solution with 35 μl of substrate solution at 40°C for 30 minutes. The reaction was quenched by adding 105 μl of DNS solution (16 g / L NaOH, 10 g / L dinitrosalicylic acid, 300 g / L potassium sodium tartrate tetrahydrate). After heating at 80°C for 20 min, the sample was cooled to 4°C and 150 μl of the solution was transferred to a flat-bottom microtiter plate. The resulting dark orange color was detected at 540 nm. Enzyme activity was determined by subtracting an appropriate blank value and using a series of glucose solutions of known concentration and reported as U / mL enzyme solution, where U is defined as μmol reducing ends released per minute of reaction time.

[0254] Example 10

[0255] Analysis of phytase or glucanase production in small-scale cultures

[0256] The resulting mutant strains were fermented in small-scale cultures, and the supernatants were analyzed. Agar-grown Myceliophthora thermophila strain was inoculated into 1 ml of culture medium in a 96-deep-well microtiter plate, as shown in Table 1. The strains were fermented at 37°C on a microtiter plate shaker at 900 rpm and 80% humidity for 72 hours. 300 μl of the 72-hour preculture was transferred to 700 μl of culture medium in a 96-deep-well microtiter plate, as shown in Table 2. The strains were fermented at 37°C on a microtiter plate shaker at 900 rpm and 80% humidity for 96 hours.

[0257] At the end of the culture, the cell-free supernatant was harvested and assayed for phytase activity (see Example 8) or glucanase activity (see Example 9). Figure 1 As can be seen in Figure 3, phytase production was higher when the expression cassette was integrated at the MYCTH:XP_003662453.1 locus compared to when the expression cassette was integrated at the cbh1 locus, regardless of whether expression was driven by the chi1 promoter or the TEF1 promoter. Figure 2 As can be seen in Figure 3, glucanase production was higher when the expression cassette was integrated at the MYCTH:XP_003662453.1 locus compared to when the expression cassette was integrated at the cbh1 locus, regardless of whether expression was driven by the chi1 promoter or the TEF1 promoter. Figure 3As can be seen in Figure 2, after the integration of the expression cassette at the cbh1 locus, the phytase production was not related to the presence of the MYCTH:XP_003662453.1 locus or the replacement of the MYCTH:XP_003662453.1 locus with the nat1 selection marker cassette; similarly, after the integration of the expression cassette at the MYCTH:XP_003662453.1 locus, the phytase production was not related to the presence of the cbh1 locus or the replacement of the cbh1 locus with the nat1 selection marker cassette. Figure 4 As can be seen, compared with the use of the thermophila host strain UV18#100fΔpyr5Δalp1Δku70, the phytase production was higher when the thermophila host strain UV18#100fΔpyr5Δalp1Δku70Δpep4Δprt1Δprt2Δprt3Δprt4Δ1am2Δtre1ΔmycAΔg1a1Δprt5Δprt6 was used; similarly, compared with the use of the thermophila host strain UV18#100fΔpyr5Δalp1Δku70, the phytase production was higher when the thermophila host strain UV18#100fΔpyr5Δalp1Δku70Δpep4Δprt1Δprt2Δprt3Δprt4Δlam2Δtre When the expression cassette was integrated at the MYCTH:XP_003662453.1 locus, the phytase production was higher when the expression cassette was integrated at the MYCTH:XP_003662453.1 locus.

[0258] Table 1: Culture medium

[0259]

[0260] Table 2: Culture medium

[0261]

Claims

1. An expression system, comprising: a. A filamentous fungal host cell comprising a native genomic region comprising a sequence selected from the list consisting of: i) a nucleic acid sequence comprising SEQ ID NO: 44 ii) a nucleic acid sequence that is at least 60% identical to SEQ ID NO: 44, and iii) a fragment of at least 500 consecutive nucleic acids of the nucleic acid sequence of i) or ii), and b. A recombinantly integrated nucleic acid molecule comprising a nucleic acid molecule flanked by different homology arms, each homology arm having at least 70% identity to at least 100 consecutive bases of the native genomic region.

2. The expression system of claim 1 , wherein the native genomic region comprises a native gene encoding a small secreted protein, the native gene being flanked on one side by at least 250 nucleotides of non-coding DNA and on the other side by at least 250 nucleotides of non-coding DNA.

3. The expression system according to claim 1 or 2, wherein the small secretory protein comprises a sequence selected from the following list: i. An amino acid sequence comprising SEQ ID NO: 46 ii. an amino acid sequence having at least 50% homology to SEQ ID NO: 46, iii. A fragment of at least 100 consecutive amino acids of i) or ii).

4. The expression system of claims 1 to 3, wherein the native genomic region is flanked by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of: a. An amino acid sequence comprising SEQ ID NO: 41 b. an amino acid sequence having at least 50% homology to SEQ ID NO: 41, and a fragment of at least 100 consecutive amino acids of ca) or b), And the other side of the native genomic region is a sequence encoding the following amino acid molecule, said amino acid molecule comprising a sequence selected from the list consisting of: d. An amino acid sequence comprising SEQ ID NO: 43 e. an amino acid sequence having at least 50% homology to SEQ ID NO: 43, and fd) or e) a fragment of at least 100 consecutive amino acids.

5. The expression system according to claim 1 , wherein the filamentous fungal host cell is from a species of the phylum Ascomycota, preferably Acremonium, Aspergillus, Agaricus, Acropora, Cryptococcus, Coryneformis, Chrysosporium, Mycobasidiomycetes, Fusarium, Humicola, Magnaporthe oryzae, Monascus, Mucor, Myceliophthora, Mortierella, Neomelia, Neurospora, Paecilomyces, Penicillium, Pyrocystis, Psoralea, Podocarpus, Podocarpus, Pycnoporus More preferably, Aspergillus niger, Aspergillus oryzae, Aspergillus fumigatus, Neurospora crassa, Penicillium chrysogenum, Penicillium citrinum, Acremonium chrysogenum, Trichoderma reesei, Rossella emersonii (formerly known as Talaromyces emersonii), Aspergillus sojae, Thermomycetes heterothallicus and Thermomycetes (formerly known as Thermomycetes and Chrysosporium lakenau), most preferably Thermomycetes.

6. A method for stably integrating a recombinant polynucleotide into a filamentous fungal host cell, the method comprising the steps of: a. Providing a filamentous fungal host cell comprising a native genomic region comprising a sequence selected from the list consisting of: i) a nucleic acid sequence comprising SEQ ID NO: 44 ii) a nucleic acid sequence having at least 60% identity to SEQ ID NO 44, and iii) a fragment of at least 500 consecutive nucleic acids of the nucleic acid sequence of i) or ii), b. introducing into said host cell a recombinantly integrated nucleic acid molecule comprising a nucleic acid molecule flanked by different homology arms, each homology arm having at least 70% identity to at least 100 consecutive bases of said native genomic region, c. allowing the recombinant nucleic acid molecule to homologously recombine into the native genomic region.

7. The method of claim 6, wherein the small secretory protein comprises a sequence selected from the list consisting of: a. An amino acid sequence comprising SEQ ID NO: 46 b. an amino acid sequence having at least 50% homology to SEQ ID NO: 46, and A fragment of at least 100 consecutive amino acids of ca) or b).

8. The method of claim 6 or 7, wherein the native genomic region is flanked on one side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of: a. An amino acid sequence comprising SEQ ID NO: 41 b. an amino acid sequence having at least 50% homology to SEQ ID NO: 41, and a fragment of at least 100 consecutive amino acids of ca) or b), and the native genomic region is flanked on the other side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of: d. An amino acid sequence comprising SEQ ID NO: 43 e. an amino acid sequence having at least 50% homology to SEQ ID NO: 43, and fd) or e) a fragment of at least 100 consecutive amino acids.

9. The method of claim 6 to 8, wherein the filamentous fungal host cell is from a species selected from the group consisting of: Ascomycota, preferably Acremonium, Aspergillus, Agaricus, Aureobasidium, Cryptococcus, Coryneformis, Chrysosporium, Mycobasidiomycetes, Fusarium, Humicola, Magnaporthe oryzae, Monascus, Mucor, Myceliophthora, Mortierella, Neomelia, Neurospora, Paecilomyces, Penicillium, Pyrocystis, Psoralea, Podocarpus, Pycnoporus, Rhizopus, Schizophyllum, Sordella, Talaromyces, Larssenella, Thermoascus, Thermobacterium, Thielavia, Tolylcopersicon, Trametes and Trichoderma, more preferably Aspergillus niger, Aspergillus oryzae, Aspergillus fumigatus, Neurospora crassa, Penicillium chrysogenum, Penicillium citrinum, Acremonium chrysogenum, Trichoderma reesei, Rossella emersonii (formerly known as Talaromyces emersonii), Aspergillus sojae, Thermophila heterothallica and Thermophila thermophila (formerly known as Thermophila thermophila and Chrysosporium lakenau), most preferably Thermophila thermophila.

10. A recombinant filamentous fungal host cell comprising a recombinant nucleic acid molecule flanked by genomic DNA, said genomic DNA comprising on one side of said recombinant nucleic acid molecule a sequence selected from the list consisting of: i) a nucleic acid sequence comprising 1296 consecutive bases from the 5′ end of SEQ ID NO: 44 ii) a nucleic acid sequence that is at least 60% identical to i), and iii) a fragment of at least 100 consecutive nucleic acids of a) or b), and said genomic DNA comprises on the other side of said recombinant nucleic acid molecule a sequence selected from the list consisting of: iv) a nucleic acid sequence comprising 1847 consecutive bases from the 3′ end of SEQ ID NO: 44 v) a nucleic acid sequence that is at least 60% identical to iv), and vi) a fragment of at least 100 consecutive nucleic acids of a) or b).

11. A recombinant filamentous fungal host cell comprising a recombinant nucleic acid molecule flanked on one side by the XP_003662454.1 gene and on the other side by the XP_003662452.1 gene.

12. The recombinant filamentous fungal host cell of claim 11, wherein the XP_003662454.1 gene encodes an amino acid molecule selected from the list consisting of: a. an amino acid sequence comprising SEQ ID NO: 43, b. an amino acid sequence having at least 50% homology to SEQ ID NO: 43, and a fragment of at least 100 consecutive amino acids of ca) or b), And the XP_003662452.1 gene encodes an amino acid molecule selected from the list consisting of: d. an amino acid sequence comprising SEQ ID NO: 41, e. an amino acid sequence having at least 50% homology to SEQ ID NO: 41, and fd) or e) a fragment of at least 100 consecutive amino acids.

13. The recombinant filamentous fungal host cell according to claims 10 to 12, wherein the host cell is from a species selected from the group consisting of: Ascomycota, preferably Acremonium, Aspergillus, Agaricus, Acropora, Cryptococcus, Coryneformis, Chrysosporium, Mycobasidiomycetes, Fusarium, Humicola, Magnaporthe oryzae, Monascus, Mucor, Myceliophthora, Mortierella, Neomelia, Neurospora, Paecilomyces, Penicillium, Pyrocystis, Psoralea, Podosporium More preferably, Aspergillus niger, Aspergillus oryzae, Aspergillus fumigatus, Neurospora crassa, Penicillium chrysogenum, Penicillium citrinum, Acremonium chrysogenum, Trichoderma reesei, Rossella emersonii (formerly known as Talaromyces emersonii), Aspergillus sojae, Thermophilic Myceliophthora heterothallis and Thermophilic Myceliophthora (formerly known as Thermophilic Myceliophthora and Chrysosporium lakenau), most preferably Thermophilic Myceliophthora.

14. A recombinant nucleic acid molecule flanked by different homology arms, each homology arm having at least 70% identity to at least 100 consecutive bases of a native genomic region of a filamentous fungus, the native genomic region comprising a sequence selected from the list consisting of: i) a nucleic acid comprising SEQ ID NO: 44, ii) a nucleic acid sequence that is at least 60% identical to SEQ ID NO: 44, and iii) a fragment of at least 500 consecutive nucleic acids of i) or ii).

15. A recombinant nucleic acid molecule flanked by different homology arms, each homology arm having at least 70% identity to at least 100 consecutive bases of a native genomic region of a filamentous fungus, wherein one side of the native genomic region is flanked by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of: a. an amino acid sequence comprising SEQ ID NO: 41, b. an amino acid sequence having at least 50% homology to SEQ ID NO: 41, and a fragment of at least 100 consecutive amino acids of ca) or b), and the native genomic region is flanked on the other side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of: d. an amino acid sequence comprising SEQ ID NO: 43, e. an amino acid sequence having at least 50% homology to SEQ ID NO: 43, and fd) or e) a fragment of at least 100 consecutive amino acids.

16. A vector comprising the recombinant nucleic acid molecule according to claim 14 or 15.

Citation Information

Patent Citations

  • Insect protective garment

    US20120005812A1

  • Transformation system in the field of filamentous fungal hosts

    US20140127788A1

  • Compounds and methods for site directed mutations in eukaryotic cells

    US5565350A

  • RAC-like genes from maize and methods of use

    WO2000015815A1

  • Transformation system in the field of filamentous fungal hosts

    WO2000020555A2