Engineered purine nucleoside phosphorylase variant enzymes

By providing engineered purine nucleoside phosphorylase (PNP) enzymes and related polynucleotides and polypeptides, the problem of insufficient inhibition efficiency of HIV reverse transcriptase in the prior art is solved, potential improvements in AIDS treatment are achieved, and a new method for producing drug compounds is provided.

CN112601543BActive Publication Date: 2025-06-13CODEXIS INC
View PDF 96 Cites 0 Cited by

Patent Information

Application Number
CN201980055238.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-22
Filing Date
2019-07-02
Publication Date
2025-06-13
Estimated Expiration
2039-07-02

AI Technical Summary

Technical Problem

The prior art has problems with insufficient efficiency in inhibiting HIV reverse transcriptase, and it is difficult to effectively improve the therapeutic effect on acquired immunodeficiency syndrome (AIDS).

Method used

Engineered purine nucleoside phosphorylase (PNP) enzymes, polypeptides with PNP activity, and polynucleotides encoding these enzymes, as well as vectors and host cells containing these polynucleotides and polypeptides, are provided for the production of PNP enzymes and to explore their applications in the production of pharmaceutical compounds.

Benefits of technology

By engineering PNP enzymes, it may be possible to inhibit HIV reverse transcriptase more effectively, thereby improving the therapeutic effect on AIDS and providing a new method for drug compound production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The present invention provides engineered purine nucleoside phosphorylase (PNP) enzymes, polypeptides having PNP activity, polynucleotides encoding these enzymes, as well as vectors and host cells comprising these polynucleotides and polypeptides. Methods for producing PNP enzymes are also provided. The present invention also provides compositions comprising PNP enzymes, and methods of using the engineered PNP enzymes. The present invention is particularly useful for the production of pharmaceutical compounds.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Provisional Patent Application Serial No. 62 / 695,507, filed on Jul. 9, 2018, and U.S. Provisional Patent Application Serial No. 62 / 822,263, filed on Mar. 22, 2019, the entire contents of both U.S. Provisional Patent Application Serial Nos. are incorporated herein by reference for all purposes. Field of the Invention

[0002] The present invention provides engineered purine nucleoside phosphorylase (PNP) enzymes, polypeptides having PNP activity, polynucleotides encoding these enzymes, as well as vectors and host cells containing these polynucleotides and polypeptides. Methods for producing PNP enzymes are also provided. The present invention also provides compositions containing PNP enzymes, as well as methods of using the engineered PNP enzymes. The present invention is particularly useful for the production of pharmaceutical compounds.

[0003] Reference to a Sequence Listing, Table, or Computer Program

[0004] A formal copy of the sequence listing is submitted concurrently with the specification as an ASCII formatted text file via EFS-Web, named "CX2-172WO1_ST25.txt", created on Jun. 25, 2019, and having a size of 1,547 kilobytes. The sequence listing submitted via EFS-Web is part of the specification and is incorporated herein by reference in its entirety. Background of the Invention

[0006] Retroviruses, such as the human immunodeficiency virus (HIV), are the causative agents of acquired immunodeficiency syndrome (AIDS), a complex disease involving the progressive destruction of the immune system of affected individuals and the degeneration of the central and peripheral nervous systems. A common feature of retroviral replication is the reverse transcription of the viral RNA genome by a virus-encoded reverse transcriptase to produce a DNA copy of the HIV sequences required for viral replication. Some compounds, such as MK-8591, are known reverse transcriptase inhibitors and can be used to treat AIDS and similar diseases. Although there are some compounds known to inhibit HIV reverse transcriptase, there is still a need in the art for additional compounds that can more effectively inhibit this enzyme and thus improve the treatment of AIDS.

[0007] Due to their similarity to natural nucleosides used in DNA synthesis, nucleoside analogues such as MK-8591 (Merck) are effective inhibitors of the reverse transcriptase of HIV. The binding of the reverse transcriptase to these analogues stalls DNA synthesis by inhibiting the progressive nature of the reverse transcriptase. The stalling of the enzyme leads to premature termination of the DNA molecule, rendering it ineffective. However, the production of nucleoside analogues by standard chemical synthesis techniques can be challenging due to their chemical complexity. SUMMARY OF THE INVENTION

[0009] The present invention provides engineered purine nucleoside phosphorylase (PNP) enzymes, polypeptides having PNP activity, and polynucleotides encoding these enzymes, as well as vectors and host cells comprising these polynucleotides and polypeptides. Methods for producing PNP enzymes are also provided. The present invention also provides compositions comprising PNP enzymes, and methods of using the engineered PNP enzymes. The present invention is particularly useful for the production of pharmaceutical compounds.

[0010] The present invention provides an engineered purine nucleoside phosphorylase, the engineered purine nucleoside phosphorylase comprising a polypeptide sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO:2 or a functional fragment thereof, wherein the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions, and wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:2. In some embodiments, the engineered purine nucleoside phosphorylase comprises a polypeptide sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO:2. In some embodiments, the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions at one or more positions selected from the following in the polypeptide sequence: 2, 12, 12 / 42, 12 / 74 / 80 / 101 / 162 / 214 / 215, 12 / 74 / 101, 12 / 80 / 162 / 210 / 215, 12 / 80 / 210, 12 / 162, 12 / 165 / 173 / 210 / 214, 12 / 187, 12 / 210, 20 / 90, 21, 25, 42, 45, 65, 69, 72, 91, 95 / 178 / 199, 105, 111, 115, 155, 162, 164, 171 / 176 / 199, 176 / 178 / 199, 176 / 199, 177, 178, 178 / 199, 179, 181, 184, 199, 202, 204, 207 and 212, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:2.In some additional embodiments, the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions selected from: 2N, 2P, 2S, 2T, 12A, 12A / 42D, 12A / 74A / 80E / 101V / 162T / 214S / 215S, 12A / 74A / 101V, 12A / 80E / 162T / 210G / 215S, 12A / 80E / 210G, 12A / 162T, 12A / 165P / 173N / 210G / 214S, 12A / 210G, 12F, 12K / 187H, 12L, 12S, 20S / 90L, 21R, 25L, 42D, 45T, 45W, 65A, 65P, 65S, 65T, 69K, 72A, 72L, 72M, 72T, 72V, 91T, 95I / 178A / 199A, 105L, 111M, 115G, 115R, 115V, 155H, 162A, 162R, 164E, 164V, 171L / 176V / 199A, 176V / 178A / 199A, 176V / 199A, 177V, 178A, 178A / 199A, 179T, 181L, 184S, 199A, 202V, 204A, 207C and 212A, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:2.In still some other embodiments, the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions selected from: A2N, A2P, A2S, A2T, D12A, D12A / N42D, D12A / T74A / D80E / L101V / S162T / T214S / A215S, D12A / T74A / L101V, D12A / D80E / S162T / H210G / A215S, D12A / D80E / H210G, D12A / S162T, D12A / G165P / K173N / H210G / T214S, D12A / H210G, D12F, D12K / Y187H, D12L, D12S, P20S / G90L, G21R, R25L, N42D, G45T, G45W, M65A, M65P, M65S, M65T, S69K, I72A, I72L, I72M, I72T, I72V, S91T, V95I / G178A / T199A, V105L, C111M, K115G, K115R, K115V, F155H, S162A, S162R, D164E, D164V, M171L / I176V / T199A, I176V / G178A / T199A, I176V / T199A, L177V, G178A, G178A / T199A, V179T, M181L, A184S, T199A, T202V, S204A, I207C, and Q212A, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:2.

[0011] The present invention provides an engineered purine nucleoside phosphorylase, the engineered purine nucleoside phosphorylase comprising a polypeptide sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID NO:6 or a functional fragment thereof, wherein the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions, and wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:6. In some embodiments, the engineered purine nucleoside phosphorylase comprises a polypeptide sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID NO:6. In some embodiments, the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions at one or more positions selected from the following in the polypeptide sequence: 2, 2 / 65, 12 / 42 / 65, 12 / 65, 12 / 65 / 74 / 80 / 101 / 162 / 214 / 215, 12 / 65 / 74 / 101, 12 / 65 / 80 / 210, 12 / 65 / 162, 12 / 65 / 165 / 173 / 210 / 214, 21 / 65, 38, 42, 42 / 155, 42 / 177, 45 / 65, 54, 65 / 72, 65 / 91, 65 / 105, 65 / 115, 65 / 202, 65 / 212, 80, 80 / 95, 80 / 155, 80 / 175, 84, 91, 91 / 115, 95, 95 / 101, 95 / 155, 95 / 155 / 215, 95 / 212, 95 / 212 / 215, 101, 101 / 105, 101 / 187, 101 / 212, 105 / 155 / 212 / 215, 108, 155, 155 / 177 / 204, 155 / 184 / 212 / 215, 155 / 212, 162 / 199, 175, 177, 184 / 212 / 215, 199, 212, 212 / 215 and 215, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:6.In some additional embodiments, the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions selected from the following: 2P, 2P / 65M, 2S, 2S / 65M, 2T, 12A / 42D / 65M, 12A / 65M, 12A / 65M / 74A / 80E / 101V / 162T / 214S / 215S, 12A / 65M / 74A / 101V, 12A / 65M / 80E / 210G, 12A / 65M / 162T, 12A / 65M / 165P / 173N / 210G / 214S, 12F / 65M, 12L / 65M, 21R / 65M, 38E, 42D, 42D / 155H, 42D / 177V, 45W / 65M, 54D, 65M / 72M, 65M / 91T, 65M / 105L, 65M / 115R, 65M / 115V, 65M / 202V, 65M / 212A, 80E, 80E / 95I, 80E / 155H, 80E / 175V, 84E, 91T, 91T / 115R, 95I, 95I / 101V, 95I / 155H, 95I / 155H / 215S, 95I / 212A, 95I / 212A / 215S, 101V, 101V / 105L, 101V / 187H, 101V / 212A, 105L / 155H / 212A / 215S, 108I, 155H, 155H / 177V / 204A, 155H / 184S / 212A / 215S, 155H / 212A, 162T / 199A, 175V, 177V, 184S / 212A / 215S, 199A, 212A, 212A / 215S, and 215S, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:6.In still some other embodiments, the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions selected from the following: A2P, A2P / A65M, A2S, A2S / A65M, A2T, D12A / N42D / A65M, D12A / A65M, D12A / A65M / T74A / D80E / L101V / S162T / T214S / A215S, D12A / A65M / T74A / L101V, D12A / A65M / D80E / H210G, D12A / A65M / S162T, D12A / A65M / G165P / K173N / H210G / T214S, D12F / A65M, D12L / A65M, G21R / A65M, R38E, N42D, N42D / F155H, N42D / L177V, G45W / A65M, K54D, A65M / I72M, A65M / S91T, A65M / V105L, A65M / K115R, A65M / K115V, A65M / T202V, A65M / Q212A, D80E, D80E / V95I, D80E / F155H, D80E / G175V, K84E, S91T, S91T / K115R, V95I, V95I / L101V, V95I / F155H, V95I / F155H / A215S, V95I / Q212A, V95I / Q212A / A215S, L101V, L101V / V105L, L101V / Y187H, L101V / Q212A, V105L / F155H / Q212A / A215S, M108I, F155H, F155H / L177V / S204A, F155H / A184S / Q212A / A215S, F155H / Q212A, S162T / T199A, G175V, L177V, A184S / Q212A / A215S, T199A, Q212A, Q212A / A215S, and A215S, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:6.

[0012] The present invention provides an engineered purine nucleoside phosphorylase, the engineered purine nucleoside phosphorylase comprising a polypeptide sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID NO: 126 or a functional fragment thereof, wherein the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions, and wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 126. In some embodiments, the engineered purine nucleoside phosphorylase comprises a polypeptide sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID NO: 126. In some embodiments, the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions at one or more positions selected from the following in the polypeptide sequence: 2, 2 / 80 / 95, 2 / 80 / 95 / 155, 2 / 80 / 95 / 155 / 199 / 212, 2 / 80 / 95 / 199, 2 / 80 / 95 / 199 / 212, 2 / 80 / 101 / 155 / 212, 2 / 80 / 101 / 212, 2 / 80 / 155 / 177 / 215, 2 / 80 / 155 / 215, 2 / 80 / 175 / 199, 2 / 80 / 175 / 199 / 212 / 215, 2 / 80 / 177, 2 / 80 / 199 / 212 / 215, 2 / 80 / 215, 2 / 95, 2 / 95 / 155, 2 / 95 / 155 / 199, 2 / 95 / 155 / 199 / 212 / 215 / 223, 2 / 95 / 155 / 199 / 215, 2 / 95 / 155 / 215, 2 / 95 / 175, 2 / 95 / 199, 2 / 95 / 199 / 212 / 215, 2 / 95 / 212 / 215, 2 / 95 / 215, 2 / 101 / 155, 2 / 155, 2 / 155 / 177 / 212, 2 / 155 / 199, 2 / 199, 2 / 199 / 212 / 215, 2 / 212 / 215, 2 / 215, 28, 39, 42, 45, 53, 54 / 173, 57 / 175, 75, 80, 80 / 95 / 101 / 155 / 199, 80 / 95 / 101 / 155 / 199 / 215, 80 / 95 / 101 / 199, 80 / 95 / 155, 80 / 95 / 155 / 175 / 199 / 212 / 215, 80 / 95 / 155 / 199, 80 / 95 / 155 / 199 / 212, 80 / 95 / 155 / 199 / 215, 80 / 95 / 155 / 212 / 215, 80 / 95 / 177,80 / 95 / 177 / 199, 80 / 95 / 177 / 212 / 215, 80 / 95 / 215, 80 / 101, 80 / 155, 80 / 155 / 177 / 199, 80 / 155 / 177 / 212 / 215, 80 / 155 / 199, 80 / 199, 80 / 212, 84, 85, 95, 95 / 155, 95 / 155 / 177, 95 / 155 / 177 / 199, 95 / 155 / 177 / 199 / 212, 95 / 155 / 199, 95 / 155 / 212 / 215, 95 / 177 / 199, 95 / 199, 95 / 212, 95 / 212 / 215, 95 / 215, 98, 101, 101 / 155 / 199, 101 / 155 / 199 / 212, 101 / 177 / 212, 101 / 215, 119, 123, 124, 129, 130, 131, 132, 136, 137, 139, 140, 143, 144, 148, 148 / 175, 150, 151, 152, 153, 155, 155 / 199, 155 / 212 / 215, 160, 170 / 203, 173, 175 / 212, 177, 177 / 199, 191, 191 / 237, 196, 199, 199 / 212, 210, 212, 212 / 215, 216 and 220, wherein the amino acid positions of the polypeptide sequences are numbered with reference to SEQ ID NO:126. In some additional embodiments, the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions selected from: 2S, 2S / 80E / 95I, 2S / 80E / 95I / 155H / 199A / 212A, 2S / 80E / 95I / 199A, 2S / 80E / 95I / 199A / 212A, 2S / 80E / 101V / 155H / 212A, 2S / 80E / 101V / 212A, 2S / 80E / 155H / 177V / 215S, 2S / 80E / 175V / 199A / 212A / 215S, 2S / 80E / 215S, 2S / 95I, 2S / 95I / 155H / 199A, 2S / 95I / 155H / 199A / 212A / 215S / 223H, 2S / 95I / 155H / 199A / 215S, 2S / 95I / 175V, 2S / 95I / 212A / 215S, 2S / 95I / 215S, 2S / 101V / 155H, 2S / 155H / 199A, 2S / 199A / 212A / 215S, 2S / 212A / 215S, 2S / 215S, 2T, 2T / 80E / 95I / 155H, 2T / 80E / 155H / 215S, 2T / 80E / 175V / 199A, 2T / 80E / 177V,2T / 80E / 199A / 212A / 215S, 2T / 95I, 2T / 95I / 155H, 2T / 95I / 155H / 199A, 2T / 95I / 155H / 215S, 2T / 95I / 199A, 2T / 95I / 199A / 212A / 215S, 2T / 95I / 212A / 215S, 2T / 95I / 215S, 2T / 155H, 2T / 155H / 177V / 212A, 2T / 155H / 199A, 2T / 199A, 2T / 215S, 28H, 39L, 42H, 42S, 45A, 45C, 53L, 54N / 173G, 57A / 175D, 75S, 80E, 80E / 95I / 101V / 155H / 199A, 80E / 95I / 101V / 155H / 199A / 215S, 80E / 95I / 101V / 199A, 80E / 95I / 155H, 80E / 95I / 155H / 175D / 199A / 212A / 215S, 80E / 95I / 155H / 199A, 80E / 95I / 155H / 199A / 212A, 80E / 95I / 155H / 199A / 215S, 80E / 95I / 155H / 212A / 215S, 80E / 95I / 177V, 80E / 95I / 177V / 199A, 80E / 95I / 177V / 212A / 215S, 80E / 95I / 215S, 80E / 101V, 80E / 155H, 80E / 155H / 177V / 199A, 80E / 155H / 177V / 212A / 215S, 80E / 155H / 199A, 80E / 199A, 80E / 212A, 84E, 85A, 85T, 95I, 95I / 155H, 95I / 155H / 177V, 95I / 155H / 177V / 199A, 95I / 155H / 177V / 199A / 212A, 95I / 155H / 199A, 95I / 155H / 212A / 215S, 95I / 177V / 199A, 95I / 199A, 95I / 212A, 95I / 212A / 215S, 95I / 215S, 98D, 98Y, 101I, 101V, 101V / 155H / 199A, 101V / 155H / 199A / 212A, 101V / 177V / 212A, 101V / 215S, 119M, 119T, 119V, 123G, 123M, 123S, 123T, 124R, 129G, 129M, 129S, 130P, 130T, 131M, 131R, 132A, 132C, 132L, 132N, 132S, 136A, 136C, 136D, 136E, 136G, 136I, 136L, 136S136V, 137E, 137Q, 137W, 139A, 139D, 139G, 139S, 139T, 140G, 143C, 143E, 143G, 143R, 143V, 143Y, 144F, 144H, 144L, 144R, 144T, 144Y, 148F, 148G, 148I / 175D, 148M, 148R / 175D, 150E, 150M, 150Y, 151F, 151H, 151L, 151N, 151Q, 152A, 152N, 152S, 153C, 153G, 153I, 153L, 153P, 153R, 153S, 153T, 153Y, 155H, 155H / 199A, 155H / 212A / 215S, 160L, 170R / 203I, 173C, 173E, 173F, 173G, 173H, 173M, 173Q, 173S, 173V, 173W, 175V / 212A, 177A, 177G, 177T, 177V, 177V / 199A, 191F, 191G, 191P, 191T / 237N, 191V, 191W, 191Y, 196V, 199A, 199A / 212A, 210G, 210R, 212A, 212A / 215S, 212W, 216C, 216F, 216L, 216R, 216W, 216Y and 220F, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:126. In yet some additional embodiments, the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions selected from: A2S, A2S / D80E / V95I, A2S / D80E / V95I / F155H / T199A / Q212A, A2S / D80E / V95I / T199A, A2S / D80E / V95I / T199A / Q212A, A2S / D80E / L101V / F155H / Q212A, A2S / D80E / L101V / Q212A, A2S / D80E / F155H / L177V / A215S, A2S / D80E / G175V / T199A / Q212A / A215S, A2S / D80E / A215S, A2S / V95I, A2S / V95I / F155H / T199A, A2S / V95I / F155H / T199A / Q212A / A215S / N223H, A2S / V95I / F155H / T199A / A215S, A2S / V95I / G175V, A2S / V95I / Q212A / A215S, A2S / V95I / A215S, A2S / L101V / F155H, A2S / F155H / T199A,A2S / T199A / Q212A / A215S, A2S / Q212A / A215S, A2S / A215S, A2T, A2T / D80E / V95I / F155H, A2T / D80E / F155H / A215S, A2T / D80E / G175V / T199A, A2T / D80E / L177V, A2T / D80E / T199A / Q212A / A215S, A2T / V95I, A2T / V95I / F155H, A2T / V95I / F155H / T199A, A2T / V95I / F155H / A215S, A2T / V95I / T199A, A2T / V95I / T199A / Q212A / A215S, A2T / V95I / Q212A / A215S, A2T / V95I / A215S, A2T / F155H, A2T / F155H / L177V / Q212A, A2T / F155H / T199A, A2T / T199A, A2T / A215S, Y28H, E39L, N42H, N42S, G45A, G45C, Y53L, K54N / K173G, K57A / G175D, K75S, D80E, D80E / V95I / L101V / F155H / T199A, D80E / V95I / L101V / F155H / T199A / A215S, D80E / V95I / L101V / T199A, D80E / V95I / F155H, D80E / V95I / F155H / G175D / T199A / Q212A / A215S, D80E / V95I / F155H / T199A, D80E / V95I / F155H / T199A / Q212A, D80E / V95I / F155H / T199A / A215S, D80E / V95I / F155H / Q212A / A215S, D80E / V95I / L177V, D80E / V95I / L177V / T199A, D80E / V95I / L177V / Q212A / A215S, D80E / V95I / A215S, D80E / L101V, D80E / F155H, D80E / F155H / L177V / T199A, D80E / F155H / L177V / Q212A / A215S, D80E / F155H / T199A, D80E / T199A, D80E / Q212A, K84E, K85A, K85T, V95I, V95I / F155H, V95I / F155H / L177V, V95I / F155H / L177V / T199A, V95I / F155H / L177V / T199A / Q212A, V95I / F155H / T199A,V95I / F155H / Q212A / A215S, V95I / L177V / T199A, V95I / T199A, V95I / Q212A, V95I / Q212A / A215S, V95I / A215S, H98D, H98Y, L101I, L101V, L101V / F155H / T199A, L101V / F155H / T199A / Q212A, L101V / L177V / Q212A, L101V / A215S, I119M, I119T, I119V, D123G, D123M, D123S, D123T, H124R, I129G, I129M, I129S, A130P, A130T, D131M, D131R, F132A, F132C, F132L, F132N, F132S, R136A, R136C, R136D, R136E, R136G, R136I, R136L, R136S, R136V, N137E, N137Q, N137W, V139A, V139D, V139G, V139S, V139T, D140G, K143C, K143E, K143G, K143R, K143V, K143Y, A144F, A144H, A144L, A144R, A144T, A144Y, D148F, D148G, D148I / G175D, D148M, D148R / G175D, R150E, R150M, R150Y, V151F, V151H, V151L, V151N, V151Q, G152A, G152N, G152S, N153C, N153G, N153I, N153L, N153P, N153R, N153S, N153T, N153Y, F155H, F155H / T199A, F155H / Q212A / A215S, F160L, V170R / V203I, K173C, K173E, K173F, K173G, K173H, K173M, K173Q, K173S, K173V, K173W, G175V / Q212A, L177A, L177G, L177T, L177V, L177V / T199A, A191F, A191G, A191P, A191T / D237N, A191V, A191W, A191Y, K196V, T199A, T199A / Q212A, H210G, H210R, Q212A, Q212A / A215S, Q212W, A216C, A216F, A216L, A216R, A216W, A216Y, and T220F, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 126.,

[0013] The present invention provides an engineered purine nucleoside phosphorylase, wherein the engineered purine nucleoside phosphorylase comprises a polypeptide sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID NO: 242 or a functional fragment thereof, wherein the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions, and wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 242. In some embodiments, the engineered purine nucleoside phosphorylase comprises a polypeptide sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID NO: 242. In some embodiments, the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions at one or more positions selected from the following in the polypeptide sequence: 7, 12, 27, 28, 31, 32, 35, 37, 38, 52, 57, 85, 94, 97, 98, 100, 102, 133, 136, 143, 144, 145, 148, 149, 150, 162, 169, 170, 172, 173, 195 / 199, 196, 199, 207, 208, 208 / 238, 209, 210, 211, 212, 213, 214, 215, 217, 219, 220, 221, 223, 224, 227, 235 and 238, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 242.In some additional embodiments, the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions selected from: 7Y, 12A, 27F, 27S, 28A, 28G, 28L, 28T, 31L, 31T, 31W, 32G, 32Q, 32V, 35R, 37L, 38L, 38Y, 52V, 52Y, 57L, 57Y, 85A, 94T, 97A, 97S, 98A, 98C, 98D, 98N, 100A, 100Y, 102L, 133A, 133W, 133Y, 136K, 143R, 144H, 144R, 144T, 145A, 145F, 145H, 145Q, 145S, 148I, 148S, 149P, 150H, 150V, 162M, 169H, 169R, 169S, 170C, 172R, 173R, 195S / 199A, 196A, 199A, 207L, 208F, 208H, 208H / 238P, 208K, 208S, 208T, 209F, 209G, 209H, 209L, 209S, 209W, 210F, 211N, 211Q, 211S, 211T, 211Y, 212G, 212R, 212V, 213S, 214A, 214H, 214V, 215G, 215H, 215P, 215S, 217A, 217D, 217M, 217Q, 219A, 220A, 220G, 220R, 220S, 221S, 223A, 223L, 223V, 224A, 224G, 224K, 224N, 227T, 235M and 238P, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:242.In still some other embodiments, the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions selected from: N7Y, D12A, K27F, K27S, Y28A, Y28G, Y28L, Y28T, E31L, E31T, E31W, T32G, T32Q, T32V, E35R, A37L, R38L, R38Y, T52V, T52Y, K57L, K57Y, K85A, A94T, P97A, P97S, H98A, H98C, H98D, H98N, K100A, K100Y, R102L, D133A, D133W, D133Y, R136K, K143R, A144H, A144R, A144T, L145A, L145F, L145H, L145Q, L145S, D148I, D148S, A149P, R150H, R150V, S162M, D169H, D169R, D169S, V170C, E172R, K173R, A195S / T199A, K196A, T199A, I207L, R208F, R208H, R208H / K238P, R208K, R208S, R208T, T209F, T209G, T209H, T209L, T209S, T209W, H210F, E211N, E211Q, E211S, E211T, E211Y, A212G, A212R, A212V, T213S, T214A, T214H, T214V, A215G, A215H, A215P, A215S, E217A, E217D, E217M, E217Q, Q219A, T220A, T220G, T220R, T220S, T221S, N223A, N223L, N223V, D224A, D224G, D224K, D224N, K227T, L235M and K238P, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:242.

[0014] The present invention provides an engineered purine nucleoside phosphorylase, wherein the engineered purine nucleoside phosphorylase comprises a polypeptide sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID NO: 684 or a functional fragment thereof, wherein the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions, and wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 684. In some embodiments, the engineered purine nucleoside phosphorylase comprises a polypeptide sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID NO: 684. In some embodiments, the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions at one or more positions selected from the following in the polypeptide sequence: 10, 10 / 133 / 227, 10 / 145, 10 / 145 / 227, 18, 20, 26, 29, 39 / 98, 39 / 133 / 227, 60, 62, 63, 74, 89, 94, 97, 97 / 98, 108, 126, 133, 135, 145, 161, 162, 165, 167, 168, 177, 199, 208, 210, 212, 224 and 227, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 684. In some additional embodiments, the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions selected from the following: 10P, 10P / 133G / 227E, 10P / 145K, 10P / 145K / 227E, 18M, 20A, 20G, 26S, 29L, 29T, 39Q / 98D, 39Q / 133G / 227E, 60T, 62A, 63S, 74S, 74V, 89I, 89T, 94S, 94T, 97D, 97D / 98D, 108A, 108Q, 108S, 108V, 126L, 133G, 135L, 145Q, 161H, 162F, 165A, 165C, 165H, 165L, 165P, 165R, 165S, 165T, 167L, 168L, 168T, 177M, 199S, 208F, 208K, 208L, 208V, 210M, 212M, 212S, 212V, 224G and 227E, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 684.In still some other embodiments, the polypeptide sequence of the engineered purine nucleoside phosphorylase comprises at least one substitution or set of substitutions selected from: M10P, M10P / D133G / K227E, M10P / L145K, M10P / L145K / K227E, L18M, P20A, P20G, A26S, I29L, I29T, E39Q / H98D, E39Q / D133G / K227E, V60T, G62A, H63S, T74S, T74V, V89I, V89T, A94S, A94T, P97D, P97D / H98D, M108A, M108Q, M108S, M108V, F126L, D133G, V135L, L145Q, Y161H, S162F, G165A, G165C, G165H, G165L, G165P, G165R, G165S, G165T, M167L, F168L, F168T, V177M, T199S, R208F, R208K, R208L, R208V, H210M, A212M, A212S, A212V, D224G, and K227E, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 684.

[0015] In some additional embodiments, the engineered purine nucleoside phosphorylase comprises a polypeptide sequence that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to the sequence of at least one engineered purine nucleoside phosphorylase variant listed in Tables 4.1, 5.1, 6.1, 7.1 and / or 8.1. In some additional embodiments, the engineered purine nucleoside phosphorylase comprises a polypeptide sequence that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to SEQ ID NO:2. In yet some additional embodiments, the engineered purine nucleoside phosphorylase comprises a polypeptide sequence that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to the sequence of at least one engineered purine nucleoside phosphorylase variant listed in the even-numbered sequences of SEQ ID NOs:6-1002. In some additional embodiments, the engineered purine nucleoside phosphorylase comprises the polypeptide sequence listed in at least one of the even-numbered sequences of SEQ ID NOs:6-1002. In some additional embodiments, the engineered purine nucleoside phosphorylase comprises at least one improved property as compared to wild-type Escherichia coli purine nucleoside phosphorylase. In some additional embodiments, the improved property comprises improved activity towards a substrate. In some embodiments, the substrate comprises Compound 4. In some embodiments, the substrate comprises Compound 3. In some embodiments, the substrate comprises Compound 8. In some additional embodiments, the improved property comprises improved production of Compound 1. In some additional embodiments, the improved property comprises improved production of Compound 10. In yet some additional embodiments, the engineered purine nucleoside phosphorylase is purified. The present invention also provides a composition comprising at least one engineered purine nucleoside phosphorylase provided herein. In some embodiments, the composition comprises one engineered purine nucleoside phosphorylase provided herein.

[0016] The present invention also provides polynucleotide sequences that encode at least one engineered purine nucleoside phosphorylase provided herein. In some embodiments, the polynucleotide sequences have at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, 5, 125, 241, and / or 683. In some additional embodiments, the polynucleotide sequences of the engineered purine nucleoside phosphorylase contain at least one substitution at one or more positions. In some additional embodiments, the polynucleotide sequences encoding at least one engineered purine nucleoside phosphorylase or a functional fragment thereof have at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, 5, 125, 241, and / or 683. In still some additional embodiments, the polynucleotide sequences are operably linked to a control sequence. In still some additional embodiments, the polynucleotide sequences are codon-optimized. In some additional embodiments, the polynucleotide sequences include the polynucleotide sequences listed in the odd-numbered sequences of SEQ ID NO: 5 - 1001.

[0017] The present invention also provides expression vectors that contain at least one polynucleotide sequence provided herein. The present invention also provides host cells that contain at least one expression vector provided herein. In some embodiments, the host cells contain at least one polynucleotide sequence provided herein. The present invention also provides a method for producing an engineered purine nucleoside phosphorylase in a host cell, the method comprising culturing the host cell under suitable conditions to produce at least one engineered purine nucleoside phosphorylase provided herein. In some embodiments, the method further comprises recovering at least one engineered purine nucleoside phosphorylase from the culture and / or the host cell. In some additional embodiments, the method further comprises a step of purifying the at least one engineered purine nucleoside phosphorylase.

[0018] Description of the Invention

[0019] The present invention provides engineered purine nucleoside phosphorylase (PNP) enzymes, polypeptides having PNP activity, polynucleotides encoding these enzymes, as well as vectors and host cells containing these polynucleotides and polypeptides. Methods for producing PNP enzymes are also provided. The present invention also provides compositions containing PNP enzymes, as well as methods of using the engineered PNP enzymes. The present invention is particularly useful for the production of pharmaceutical compounds.

[0020] Unless otherwise defined, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Generally, the nomenclature used herein and the experimental procedures in cell culture, molecular genetics, microbiology, organic chemistry, analytical chemistry, and nucleic acid chemistry described below are those well known and commonly employed in the art. Such techniques are well known and are described in many textbooks and reference works well known to those skilled in the art. Standard techniques or their modified forms are used for chemical synthesis and chemical analysis. All patents, patent applications, articles, and publications mentioned herein (both above and below) are hereby expressly incorporated herein by reference.

[0021] Although any suitable methods and materials similar or equivalent to those described herein can be used in the practice of the present invention, some methods and materials are described herein. It should be understood that the present invention is not limited to the specific methods, protocols, and reagents described, as these can vary depending on how those skilled in the art use them. Accordingly, the terms defined below are more fully described by reference to the present invention as a whole.

[0022] It should be understood that the foregoing general description and the following detailed description are merely exemplary and illustrative, and not restrictive of the present invention. The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described. Numerical ranges include the numbers defining the range. Thus, each numerical range disclosed herein is intended to include every narrower numerical range that falls within such broader numerical range, as if such narrower numerical ranges were all clearly written out herein. It is also intended that each maximum (or minimum) numerical limitation disclosed herein include every lower (or higher) numerical limitation, as if such lower (or higher) numerical limitations were clearly written out herein.

[0023] Abbreviations and Definitions

[0024] The abbreviations for the amino acids used in genetic coding are conventional and are as follows: alanine (Ala or A), arginine (Arg or R), asparagine (Asn or N), aspartic acid (Asp or D), cysteine (Cys or C), glutamic acid (Glu or E), glutamine (Gln or Q), histidine (His or H), isoleucine (Ile or I), leucine (Leu or L), lysine (Lys or K), methionine (Met or M), phenylalanine (Phe or F), proline (Pro or P), serine (Ser or S), threonine (Thr or T), tryptophan (Trp or W), tyrosine (Tyr or Y), and valine (Val or V).

[0025] When using three-letter abbreviations, an amino acid can be of the L-configuration or D-configuration with respect to the α-carbon (Cα) unless specifically preceded by "L" or "D", or it is clear from the context in which the abbreviation is used. For example, "Ala" represents alanine without specifying the configuration with respect to the α-carbon, while "D-Ala" and "L-Ala" represent D-alanine and L-alanine, respectively. When using single-letter abbreviations, capital letters represent amino acids of the L-configuration with respect to the α-carbon, and lowercase letters represent amino acids of the D-configuration with respect to the α-carbon. For example, "A" represents L-alanine and "a" represents D-alanine. When a polypeptide sequence is presented as a string of single-letter or three-letter abbreviations (or a mixture thereof), the sequence is presented in the amino (N)-to-carboxyl (C) direction according to conventional convention.

[0026] The abbreviations used for genetically encoded nucleosides are conventional and are as follows: adenosine (A); guanosine (G); cytidine (C); thymidine (T); and uridine (U). Unless specifically described, an abbreviated nucleoside can be a ribonucleoside or a 2'-deoxyribonucleoside. Nucleosides can be specified as ribonucleosides or 2'-deoxyribonucleosides either individually or collectively. When a nucleic acid sequence is presented as a string of single-letter abbreviations, the sequence is presented in the 5'-to-3' direction according to conventional convention, and the phosphates are not shown.

[0027] Referring to the present invention, the technical and scientific terms used in the description herein will have the meanings commonly understood by those of ordinary skill in the art, unless specifically defined otherwise. Accordingly, the following terms are intended to have the following meanings.

[0028] Unless the context clearly indicates otherwise, as used herein, the singular forms "a", "an", and "the" include plural referents. Thus, for example, a reference to "a polypeptide" includes more than one polypeptide.

[0029] Similarly, "comprise", "comprises", "comprising", "include", "includes", and "including" are interchangeable and are not intended to be limiting. Accordingly, as used herein, the term "comprising" and its cognates are used in their inclusive sense (i.e., equivalent to the term "including" and its corresponding cognates).

[0030] It should also be understood that in instances where the description of various embodiments uses the term "comprising", those of skill in the art will understand that in some specific instances, the language "consisting essentially of" or "consisting of" can optionally be used to describe the embodiments.

[0031] As used herein, the term "about" means an acceptable error for a particular value. In some instances, "about" means within 0.05%, 0.5%, 1.0%, or 2.0% of a given value range. In some instances, "about" means within 1, 2, 3, or 4 standard deviations of a given value.

[0032] As used herein, the "EC" number refers to the enzyme nomenclature of the Nomenclature Committee of the International Union of Biochemistry and Molecular Biology (NC-IUBMB). This IUBMB biochemical classification is an enzyme numerical classification system based on the chemical reactions catalyzed by enzymes.

[0033] As used herein, "ATCC" refers to the American Type Culture Collection, whose biological deposit collection includes genes and strains.

[0034] As used herein, "NCBI" refers to the National Center for Biological Information in the United States and the sequence databases provided therein.

[0035] As used herein, the enzyme "phosphopentomutase" ("PPM") is an enzyme that catalyzes the reversible isomerization of ribose 1-phosphate to ribose 5-phosphate and the reversible isomerization of related compounds such as deoxyribose phosphate and analogs of ribose phosphate and deoxyribose phosphate.

[0036] As used herein, the enzyme "purine nucleoside phosphorylase" ("PNP") is an enzyme that catalyzes the reversible phosphorylation of purine ribonucleosides and related compounds (such as deoxyribonucleosides and analogs of ribonucleosides and deoxyribonucleosides) to free purine bases and ribose-1-phosphate (and its analogs). "Protein", "polypeptide", and "peptide" are used interchangeably herein to denote a polymer of at least two amino acids covalently linked by amide bonds, regardless of length or post-translational modification (such as glycosylation or phosphorylation). This definition includes D-amino acids and L-amino acids, as well as mixtures of D-amino acids and L-amino acids, and polymers containing D-amino acids and L-amino acids and mixtures of D-amino acids and L-amino acids.

[0037] "Amino acid" is referred to herein by its commonly known three-letter symbol or by the single-letter symbol recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Similarly, nucleotides can be referred to by their commonly accepted single-letter codes.

[0038] As used herein, "hydrophilic amino acid or residue" refers to an amino acid or residue having a side chain that exhibits a hydrophobicity less than zero according to the normalized consensus hydrophobicity scale of Eisenberg et al. (Eisenberg et al., J. Mol. Biol., 179:125 - 142

[1984] ). Genetically encoded hydrophilic amino acids include L-Thr (T), L-Ser (S), L-His (H), L-Glu (E), L-Asn (N), L-Gln (Q), L-Asp (D), L-Lys (K), and L-Arg (R).

[0039] As used herein, "acidic amino acid or residue" refers to a hydrophilic amino acid or residue having a side chain that exhibits a pKa value less than about 6 when the amino acid is incorporated in a peptide or polypeptide. Due to the loss of a hydrogen ion, acidic amino acids typically have a negatively charged side chain at physiological pH. Genetically encoded acidic amino acids include L-Glu (E) and L-Asp (D).

[0040] As used herein, "basic amino acid or residue" refers to a hydrophilic amino acid or residue having a side chain that exhibits a pKa value greater than about 6 when the amino acid is incorporated in a peptide or polypeptide. Due to the association with a hydronium ion, basic amino acids typically have a positively charged side chain at physiological pH. Genetically encoded basic amino acids include L-Arg (R) and L-Lys (K).

[0041] As used herein, "polar amino acid or residue" refers to a hydrophilic amino acid or residue having a side chain that is uncharged at physiological pH but has at least one bond in which an electron pair shared by two atoms is held more closely by one of the atoms. Genetically encoded polar amino acids include L-Asn (N), L-Gln (Q), L-Ser (S), and L-Thr (T).

[0042] As used herein, "hydrophobic amino acid or residue" refers to an amino acid or residue having a side chain that exhibits a hydrophobicity greater than zero according to the normalized consensus hydrophobicity scale of Eisenberg et al. (Eisenberg et al., J. Mol. Biol., 179:125 - 142

[1984] ). Genetically encoded hydrophobic amino acids include L-Pro (P), L-Ile (I), L-Phe (F), L-Val (V), L-Leu (L), L-Trp (W), L-Met (M), L-Ala (A), and L-Tyr (Y).

[0043] As used herein, "aromatic amino acid or residue" refers to a hydrophilic or hydrophobic amino acid or residue having a side chain that includes at least one aromatic or heteroaromatic ring. Genetically encoded aromatic amino acids include L-Phe (F), L-Tyr (Y), and L-Trp (W). Although L-His (H) is sometimes classified as a basic residue due to the pKa of its heteroaromatic nitrogen atom, or as an aromatic residue because its side chain includes a heteroaromatic ring, in this document, histidine is classified as a hydrophilic residue or as a "constrained residue" (see below).

[0044] As used herein, "constrained amino acid or residue" refers to an amino acid or residue having a constrained geometry. In this document, constrained residues include L-Pro (P) and L-His (H). Histidine has a constrained geometry because it has a relatively small imidazole ring. Proline has a constrained geometry because it also has a five-membered ring.

[0045] As used herein, "nonpolar amino acid or residue" refers to a hydrophobic amino acid or residue having a side chain that is uncharged at physiological pH and has bonds in which an electron pair is commonly shared by two atoms and is usually held equally by each of the two atoms (i.e., the side chain is not polar). Genetically encoded nonpolar amino acids include L-Gly (G), L-Leu (L), L-Val (V), L-Ile (I), L-Met (M), and L-Ala (A).

[0046] As used herein, "aliphatic amino acid or residue" refers to a hydrophobic amino acid or residue having an aliphatic hydrocarbon side chain. Genetically encoded aliphatic amino acids include L-Ala (A), L-Val (V), L-Leu (L), and L-Ile (I). It is notable that cysteine (or "L-Cys" or "[C]") is uncommon because it can form disulfide bridges with other L-Cys (C) amino acids or other amino acids containing a sulfonyl or mercapto group. "Cysteine-like residues" include cysteine and other amino acids containing a mercapto moiety that can be used to form disulfide bridges. The ability of L-Cys (C) (and other amino acids having a -SH side chain) to exist in a peptide in a reduced free -SH or oxidized disulfide-bridged form affects whether L-Cys (C) contributes a net hydrophobic or hydrophilic character to the peptide. Although L-Cys (C) exhibits a hydrophobicity of 0.29 according to Eisenberg's normalized consensus scale (Eisenberg et al., 1984, supra), it should be understood that for the purposes of this disclosure, L-Cys (C) is classified into its own unique group.

[0047] As used herein, "small amino acid or residue" refers to an amino acid or residue having a side chain that includes a total of three or fewer carbons and / or heteroatoms (excluding the α-carbon and hydrogen). According to the above definition, small amino acids or residues can be further classified as aliphatic, nonpolar, polar, or acidic small amino acids or residues. Genetically encoded small amino acids include L-Ala (A), L-Val (V), L-Cys (C), L-Asn (N), L-Ser (S), L-Thr (T), and L-Asp (D).

[0048] As used herein, "hydroxyl-containing amino acid or residue" refers to an amino acid that contains a hydroxyl (-OH) moiety. Genetically encoded hydroxyl-containing amino acids include L-Ser (S), L-Thr (T), and L-Tyr (Y).

[0049] As used herein, "polynucleotide" and "nucleic acid" refer to two or more nucleotides covalently linked together. A polynucleotide can consist entirely of ribonucleotides (i.e., RNA), consist entirely of 2'-deoxyribonucleotides (i.e., DNA), or contain a mixture of ribonucleotides and 2'-deoxyribonucleotides. Although nucleotides are typically linked together via standard phosphodiester linkages, a polynucleotide can include one or more non-standard linkages. A polynucleotide can be single-stranded or double-stranded, or can include both single-stranded and double-stranded regions. In addition, although polynucleotides typically contain naturally occurring coding nucleobases (i.e., adenine, guanine, uracil, thymine, and cytosine), it can contain one or more modified and / or synthetic nucleobases, such as, for example, inosine, xanthine, hypoxanthine, etc. In some embodiments, such modified or synthetic nucleobases are nucleobases that encode an amino acid sequence.

[0050] As used herein, "nucleoside" refers to a glycosylamine that contains a nucleobase (i.e., nitrogenous base) and a 5-carbon sugar (e.g., ribose or deoxyribose). Non-limiting examples of nucleosides include cytidine, uridine, adenosine, guanosine, thymidine, and inosine. In contrast, the term "nucleotide" refers to a glycosylamine that contains a nucleobase, a 5-carbon sugar, and one or more phosphate groups. In some embodiments, a nucleoside can be phosphorylated by a kinase to produce a nucleotide.

[0051] As used herein, "nucleoside diphosphate" refers to a glycosylamine that contains a nucleobase (i.e., nitrogenous base), a 5-carbon sugar (e.g., ribose or deoxyribose), and a diphosphate (i.e., pyrophosphate) moiety. In some embodiments herein, "nucleoside diphosphate" is abbreviated as "NDP". Non-limiting examples of nucleoside diphosphates include cytidine diphosphate (CDP), uridine diphosphate (UDP), adenosine diphosphate (ADP), guanosine diphosphate (GDP), thymidine diphosphate (TDP), and inosine diphosphate (IDP). In some instances, the terms "nucleoside" and "nucleotide" may be used interchangeably.

[0052] As used herein, "coding sequence" refers to the portion of a nucleic acid (e.g., a gene) that encodes the amino acid sequence of a protein.

[0053] As used herein, the terms "biocatalysis", "biocatalytic", "biotransformation", and "biosynthesis" refer to the use of enzymes to effect chemical reactions on organic compounds.

[0054] As used herein, "wild-type" and "naturally occurring" refer to the form found in nature. For example, a wild-type polypeptide or polynucleotide sequence is the sequence present in an organism, which can be isolated from a natural source and has not been intentionally modified by human manipulation.

[0055] As used herein, when used in reference to a cell, nucleic acid, or polypeptide, "recombinant", "engineered", "variant", and "non-naturally occurring" refer to a material that has been modified in a manner that does not occur in nature or to a material corresponding to the natural or native form of such a material. In some embodiments, a cell, nucleic acid, or polypeptide is identical to a naturally occurring cell, nucleic acid, or polypeptide but is produced or derived from synthetic materials and / or by using recombinant techniques. Non-limiting examples include, among others, recombinant cells that express genes not found in cells expressing the natural (non-recombinant) form or express natural genes that are otherwise expressed at different levels.

[0056] The term "percent sequence identity (%)" is used herein to refer to the comparison between polynucleotides or polypeptides and is determined by comparing the two best aligned sequences in a comparison window, where the portion of the polynucleotide or polypeptide sequence in the comparison window may include additions or deletions (i.e., gaps) as compared to the reference sequence for the optimal alignment of the two sequences. The percentage can be calculated as follows: determine the number of positions in the two sequences where the same nucleic acid base or amino acid residue occurs to yield the number of matching positions, divide the number of matching positions by the total number of positions in the comparison window, and multiply the result by 100 to obtain the percent sequence identity. Optionally, the percentage can be calculated as follows: determine the number of positions in the two sequences where the same nucleic acid base or amino acid residue occurs or where a nucleic acid base or amino acid residue is aligned with a gap to yield the number of matching positions, divide the number of matching positions by the total number of positions in the comparison window, and multiply the result by 100 to obtain the percent sequence identity. Those skilled in the art understand that there are many established algorithms available for aligning two sequences. The optimal alignment of the sequences to be compared can be performed by any suitable method, including but not limited to the local homology algorithm of Smith and Waterman (Smith and Waterman, Adv. Appl. Math., 2:482

[1981] ), the homology alignment algorithm of Needleman and Wunsch (Needleman and Wunsch, J. Mol. Biol., 48:443

[1970] ), the similarity search method of Pearson and Lipman (Pearson and Lipman, Proc. Natl. Acad. Sci. USA 85:2444

[1988] ), the computerized implementations of these algorithms (e.g., GAP, BESTFIT, FASTA, and TFASTA in the GCG Wisconsin software package), or by visual inspection, as known in the art. Examples of algorithms suitable for determining percent sequence identity and sequence similarity include but are not limited to the BLAST and BLAST 2.0 algorithms, described by Altschul et al. (see Altschul et al., J. Mol. Biol., 215:403-410

[1990] ; and Altschul et al., Nucl. Acids Res., 3389-3402

[1977] ). Software for performing BLAST analysis is available to the public through the National Center for Biotechnology Information website. The algorithm involves first identifying high-scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that match or satisfy a positive-valued threshold score T when aligned with words of the same length in the database sequences. T is referred to as the neighborhood word-scoring threshold (see, Altschul et al., supra).These initial word hits serve as seeds to initiate the search for longer HSPs that contain them. The word hits are then extended in both directions along each sequence until no further increase in the cumulative alignment score can be achieved. For nucleotide sequences, the cumulative score is calculated using parameters M (reward score for matching residue pairs; always >0) and N (penalty score for mismatched residues; always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction is stopped when: the cumulative alignment score has decreased by amount X from its maximum achieved value; the cumulative score reaches 0 or less due to the accumulation of one or more negatively scored residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses the following as defaults: word size (W) of 11, expectation (E) of 10, M = 5, N = -4, and comparison of both strands. For amino acid sequences, the BLASTP program uses the following as defaults: word size (W) of 3, expectation (E) of 10, and the BLOSUM62 scoring matrix (see, Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89:10915

[1989] ). Exemplary determination of sequence alignment and % sequence identity can be made using the BESTFIT or GAP programs in the GCG Wisconsin software package (Accelrys, Madison WI) with the default parameters provided.

[0057] As used herein, a "reference sequence" is a defined sequence used as a basis for sequence and / or activity comparison. A reference sequence can be a subset of a larger sequence, e.g., a segment of a full-length gene or polypeptide sequence. Generally, a reference sequence is at least 20 nucleotides or amino acid residues in length, at least 25 residues in length, at least 50 residues in length, at least 100 residues in length, or the full length of a nucleic acid or polypeptide. Because two polynucleotides or polypeptides can each (1) include sequences that are similar between the two sequences (i.e., a portion of the complete sequence), and (2) can also include sequences that are divergent between the two sequences, sequence comparison between two (or more) polynucleotides or polypeptides is generally performed by comparing the sequences of the two polynucleotides or polypeptides in a "comparison window" to identify and compare local regions of sequence similarity. In some embodiments, a "reference sequence" can be based on a primary amino acid sequence, where the reference sequence is a sequence that can have one or more variations in the primary sequence.

[0058] As used herein, "comparison window" refers to a conceptual segment of at least about 20 contiguous nucleotide positions or amino acid residues, wherein a sequence can be compared to a reference sequence of at least 20 contiguous nucleotides or amino acids, and wherein the portion of the sequence in the comparison window can include 20% or fewer additions or deletions (i.e., gaps) as compared to the reference sequence (which does not contain additions or deletions) for optimal alignment of the two sequences. The comparison window can be longer than 20 contiguous residues and optionally includes windows of 30, 40, 50, 100 or longer.

[0059] As used herein, when used in the context of numbering a given amino acid or polynucleotide sequence, "corresponds to", "refers to" or "relative to" means numbering the residues of a designated reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. In other words, the residue numbering or residue position of a given polymer is designated with respect to the reference sequence, rather than by the actual numerical position of the residues within the given amino acid or polynucleotide sequence. For example, a given amino acid sequence, such as the amino acid sequence of engineered purine nucleoside phosphorylase, can be optimized for residue matching between the two sequences by introducing gaps to align with the reference sequence. In such cases, despite the gaps, the residues in the given amino acid or polynucleotide sequence are numbered with respect to the reference sequence to which it is aligned.

[0060] As used herein, "substantial identity" means a polynucleotide or polypeptide sequence that has at least 80% sequence identity, at least 85% identity, at least 89% to 95% sequence identity, or more typically at least 99% sequence identity as compared to a reference sequence in a comparison window of at least 20 residue positions, typically in a window of at least 30 - 50 residues, wherein the percentage of sequence identity is calculated by comparing the reference sequence and the sequence containing a total of 20% or fewer deletions or additions to the reference sequence in the comparison window. In some specific embodiments applied to polypeptides, the term "substantial identity" means that when optimally aligned using default gap weights by a program such as GAP or BESTFIT, two polypeptide sequences share at least 80% sequence identity, preferably at least 89% sequence identity, at least 95% sequence identity or more (e.g., 99% sequence identity). In some embodiments, the residue positions that are not identical in the sequences being compared differ by conservative amino acid substitutions.

[0061] As used herein, "amino acid difference" and "residue difference" refer to a difference in an amino acid residue at a position in a polypeptide sequence relative to the amino acid residue at the corresponding position in a reference sequence. In some cases, the reference sequence has a histidine tag, but the numbering remains the same relative to an equivalent reference sequence without the histidine tag. The position of an amino acid difference herein is typically referred to as "Xn", where n refers to the corresponding position in the reference sequence on which the residue difference is based. For example, "a residue difference at position X93 compared to SEQ ID NO:4" refers to a difference in the amino acid residue at the polypeptide position corresponding to position 93 of SEQ ID NO:4. Thus, if the reference polypeptide of SEQ ID NO:4 has serine at position 93, "a residue difference at position X93 compared to SEQ ID NO:4" refers to an amino acid substitution of any residue other than serine at the polypeptide position corresponding to position 93 of SEQ ID NO:4. In most examples herein, a specific amino acid residue difference at a position is indicated as "XnY", where "Xn" designates the corresponding position as described above, and "Y" is the single-letter identifier of the amino acid found in the engineered polypeptide (i.e., the residue different from that in the reference polypeptide). In some examples (e.g., in the tables presented in the Examples), the present invention also provides specific amino acid differences represented by the conventional symbol "AnB", where A is the single-letter identifier of the residue in the reference sequence, "n" is the number of the residue position in the reference sequence, and B is the single-letter identifier of the residue substitution in the sequence of the engineered polypeptide. In some examples, the polypeptides of the present invention may contain one or more amino acid residue differences relative to the reference sequence, which are indicated by a column of designated positions where residue differences exist relative to the reference sequence. In some embodiments, when more than one amino acid can be used at a specific residue position of a polypeptide, the various amino acid residues that can be used are separated by " / " (e.g., X307H / X307P or X307H / P). The slash can also be used to indicate more than one substitution within a given variant (i.e., when more than one substitution is present in a given sequence such as in a combinatorial variant). In some embodiments, the present invention includes engineered polypeptide sequences containing one or more amino acid differences, which include conservative amino acid substitutions or non-conservative amino acid substitutions. In some additional embodiments, the present invention provides engineered polypeptide sequences containing both conservative amino acid substitutions and non-conservative amino acid substitutions.

[0062] As used herein, "conservative amino acid substitution" refers to the replacement of a residue with a different residue having a similar side chain, and thus generally includes the replacement of an amino acid in a polypeptide with an amino acid in the same or a similar amino acid-defined category. By way of example and not limitation, in some embodiments, an amino acid having an aliphatic side chain is replaced with another aliphatic amino acid (e.g., alanine, valine, leucine, and isoleucine); an amino acid having a hydroxyl side chain is replaced with another amino acid having a hydroxyl side chain (e.g., serine and threonine); an amino acid having an aromatic side chain is replaced with another amino acid having an aromatic side chain (e.g., phenylalanine, tyrosine, tryptophan, and histidine); an amino acid having a basic side chain is replaced with another amino acid having a basic side chain (e.g., lysine and arginine); an amino acid having an acidic side chain is replaced with another amino acid having an acidic side chain (e.g., aspartic acid or glutamic acid); and / or a hydrophobic amino acid or a hydrophilic amino acid is replaced with another hydrophobic amino acid or hydrophilic amino acid, respectively.

[0063] As used herein, "non-conservative substitution" refers to the replacement of an amino acid in a polypeptide with an amino acid having significantly different side chain properties. Non-conservative substitutions can be made with amino acids from defined groups rather than within a group, and affect: (a) the structure of the peptide backbone in the region of substitution (e.g., proline replacing glycine), (b) charge or hydrophobicity, or (c) side chain volume. By way of example and not limitation, exemplary non-conservative substitutions can be the replacement of an acidic amino acid with a basic or aliphatic amino acid; the replacement of an aromatic amino acid with a small amino acid; and the replacement of a hydrophilic amino acid with a hydrophobic amino acid.

[0064] As used herein, "deletion" refers to the modification of a polypeptide by removing one or more amino acids from a reference polypeptide. Deletions can include the removal of 1 or more amino acids, 2 or more amino acids, 5 or more amino acids, 10 or more amino acids, 15 or more amino acids, or 20 or more amino acids, up to 10% of the total number of amino acids making up the reference enzyme or up to 20% of the total number of amino acids, while retaining enzyme activity and / or retaining the improved properties of the engineered purine nucleoside phosphorylase. Deletions can involve internal portions and / or terminal portions of the polypeptide. In various embodiments, deletions can include contiguous segments or can be non-contiguous. Deletions in an amino acid sequence are typically denoted by "-".

[0065] As used herein, "insertion" refers to the modification of a polypeptide by adding one or more amino acids to a reference polypeptide. Insertions can be made into an internal portion of the polypeptide or to the carboxyl or amino terminus. As used herein, insertions include fusion proteins as are known in the art. Insertions can be contiguous segments of amino acids or can be separated by one or more amino acids in a naturally occurring polypeptide.

[0066] The term "amino acid substitution set" or "substitution set" refers to a set of amino acid substitutions in a polypeptide sequence as compared to a reference sequence. A substitution set can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more amino acid substitutions. In some embodiments, the substitution set refers to the set of amino acid substitutions present in any of the variant purine nucleoside phosphorylases listed in the tables provided in the examples.

[0067] "Functional fragment" and "bioactive fragment" are used interchangeably herein and refer to a polypeptide that has an amino-terminal deletion and / or a carboxyl-terminal deletion and / or an internal deletion, but wherein the remaining amino acid sequence is the same as the corresponding positions in the sequence to which it is compared (e.g., the full-length engineered purine nucleoside phosphorylase of the present invention), and retains substantially all of the activity of the full-length polypeptide.

[0068] As used herein, an "isolated polypeptide" refers to a polypeptide that is substantially separated from other contaminants with which it is naturally associated (e.g., proteins, lipids, and polynucleotides). The term includes polypeptides that have been removed or purified from their natural environment or expression system (e.g., within a host cell or via in vitro synthesis). Recombinant purine nucleoside phosphorylase polypeptides can be present intracellularly, in cell culture medium, or prepared in various forms (such as lysates or isolated preparations). Thus, in some embodiments, a recombinant purine nucleoside phosphorylase polypeptide can be an isolated polypeptide.

[0069] As used herein, a "substantially pure polypeptide" or "purified protein" refers to a composition in which the polypeptide species is the major species present (i.e., it is more abundant than any other individual macromolecular species in the composition on a molar or weight basis), and a composition is generally substantially purified when the target species constitutes at least about 50% on a molar or % weight basis of the macromolecular species present. However, in some embodiments, a composition comprising purine nucleoside phosphorylase comprises less than 50% pure (e.g., about 10%, about 20%, about 30%, about 40% or about 50%) purine nucleoside phosphorylase. Generally, a substantially pure purine nucleoside phosphorylase composition constitutes about 60% or more, about 70% or more, about 80% or more, about 90% or more, about 95% or more, and about 98% or more on a molar or % weight basis of all the macromolecular species present in the composition. In some embodiments, the target species is purified to substantial homogeneity (i.e., contaminant species cannot be detected in the composition by conventional detection methods), where the composition consists essentially of a single macromolecular species. Solvent species, small molecules (<500 daltons), and elemental ion species are not considered macromolecular species. In some embodiments, an isolated recombinant purine nucleoside phosphorylase polypeptide is a substantially pure polypeptide composition.

[0070] As used herein, "improved enzyme property" refers to at least one improved property of an enzyme. In some embodiments, the present invention provides an engineered purine nucleoside phosphorylase polypeptide that exhibits an improvement in any enzyme property as compared to a reference purine nucleoside phosphorylase polypeptide and / or a wild-type purine nucleoside phosphorylase polypeptide and / or another engineered purine nucleoside phosphorylase polypeptide. Thus, the level of "improvement" can be determined and compared among various purine nucleoside phosphorylase polypeptides, including wild-type and engineered purine nucleoside phosphorylases. Improved properties include, but are not limited to, properties such as increased protein expression, increased thermoactivity, increased thermal stability, increased pH activity, increased stability, increased enzyme activity, increased substrate specificity or affinity, increased specific activity, increased resistance to substrate or end-product inhibition, increased chemical stability, improved chemoselectivity, improved solvent stability, increased tolerance to acidic pH, increased tolerance to proteolytic activity (i.e., decreased sensitivity to proteolysis), decreased aggregation, increased solubility, and altered temperature profile. In additional embodiments, the term is used to refer to at least one improved property of a purine nucleoside phosphorylase. In some embodiments, the present invention provides an engineered purine nucleoside phosphorylase polypeptide that exhibits an improvement in any enzyme property as compared to a reference purine nucleoside phosphorylase polypeptide and / or a wild-type purine nucleoside phosphorylase polypeptide and / or another engineered purine nucleoside phosphorylase polypeptide. Thus, the level of "improvement" can be determined and compared among various purine nucleoside phosphorylase polypeptides, including wild-type and engineered purine nucleoside phosphorylases.

[0071] As used herein, "increased enzymatic activity" and "enhanced catalytic activity" refer to improved properties of an engineered polypeptide that can be expressed as an increase in specific activity (e.g., product produced / time / weight of protein) or an increase in the percentage of conversion of a substrate to a product (e.g., the percentage of conversion of an initial amount of substrate to product using a specified amount of enzyme over a specified period of time) compared to a reference enzyme. In some embodiments, the term refers to improved properties of the engineered purine nucleoside phosphorylase polypeptides provided herein, which can be expressed as an increase in specific activity (e.g., product produced / time / weight of protein) or an increase in the percentage of conversion of a substrate to a product (e.g., the percentage of conversion of an initial amount of substrate to product using a specified amount of purine nucleoside phosphorylase over a specified period of time) compared to a reference purine nucleoside phosphorylase. In some embodiments, these terms are used to refer to the improved purine nucleoside phosphorylases provided herein. Exemplary methods for determining the enzymatic activity of the engineered purine nucleoside phosphorylases of the invention are provided in the Examples. Any property related to enzymatic activity can be affected, including the typical enzymatic properties Km, Vmax, or kcat, and alterations in them can result in increased enzymatic activity. For example, the improvement in enzymatic activity can be from about 1.1-fold the enzymatic activity of the corresponding wild-type enzyme to up to 2-fold, 5-fold, 10-fold, 20-fold, 25-fold, 50-fold, 75-fold, 100-fold, 150-fold, 200-fold, or greater enzymatic activity compared to a naturally occurring purine nucleoside phosphorylase or another engineered purine nucleoside phosphorylase from which the purine nucleoside phosphorylase polypeptide is derived.

[0072] As used herein, "conversion" refers to the enzymatic conversion (or bioconversion) of one or more substrates into one or more corresponding products. "Percentage of conversion" refers to the percentage of substrate that is converted to product under specified conditions over a given period of time. Thus, the "enzymatic activity" or "activity" of a purine nucleoside phosphorylase polypeptide can be expressed as the "percentage of conversion" of substrate to product over a specific period of time.

[0073] An enzyme with "generalist properties" (or "generalist enzymes") refers to an enzyme that exhibits improved activity towards a broad range of substrates compared to the parental sequence. A generalist enzyme does not have to exhibit improved activity towards every possible substrate. In some embodiments, the invention provides purine nucleoside phosphorylase variants with generalist properties, as they exhibit similar or improved activity towards a broad range of spatially and electronically diverse substrates compared to the parental gene. In addition, the generalist enzymes provided herein are engineered to improve metabolite / product production across a broad range of diverse molecules.

[0074] The term "stringent hybridization conditions" is used herein to refer to conditions under which a nucleic acid hybrid is stable. As is known to those of skill in the art, the stability of a hybrid is reflected in the melting temperature (Tm) of the hybrid. Generally, the stability of a hybrid is a function of ionic strength, temperature, G / C content, and the presence of chaotropic agents. The Tm value of a polynucleotide can be calculated using known methods for predicting melting temperature (see, e.g., Baldino et al., Meth. Enzymol., 168:761-777

[1989] ; Bolton et al., Proc. Natl. Acad. Sci. USA 48:1390

[1962] ; Bresslauer et al., Proc. Natl. Acad. Sci. USA 83:8893-8897

[1986] ; Freier et al., Proc. Natl. Acad. Sci. USA 83:9373-9377

[1986] ; Kierzek et al., Biochem., 25:7840-7846

[1986] ; Rychlik et al., Nucl. Acids Res., 18:6409-6412

[1990] (erratum, Nucl. Acids Res., 19:698

[1991] ); Sambrook et al., supra); Suggs et al., 1981, in Developmental Biology Using Purified Genes Brown et al. [eds.], pp. 683-693, Academic Press, Cambridge, MA

[1981] ; and Wetmur, Crit. Rev. Biochem. Mol. Biol. 26:227-259

[1991] ). In some embodiments, the polynucleotide encodes a polypeptide disclosed herein and hybridizes, under defined conditions, such as moderately stringent or highly stringent conditions, to a complementary sequence of a sequence encoding an engineered purine nucleoside phosphorylase of the invention.

[0075] As used herein, "hybridization stringency" refers to the hybridization conditions in nucleic acid hybridization, such as washing conditions. Generally, the hybridization reaction is carried out under conditions of lower stringency, followed by washing at different but higher stringencies. The term "moderate stringency hybridization" refers to conditions that permit binding of a target DNA to a complementary nucleic acid having about 60% identity, preferably about 75% identity, about 85% identity, and greater than about 90% identity to the target polynucleotide. Exemplary moderate stringency conditions are equivalent to hybridization in 50% formamide, 5×Denhart's solution, 5×SSPE, 0.2% SDS at 42°C, followed by washing in 0.2×SSPE, 0.2% SDS at 42°C. "High stringency hybridization" generally refers to conditions that are about 10°C or less different from the melting temperature Tm as determined for a defined polynucleotide sequence under solution conditions. In some embodiments, high stringency conditions refer to those conditions that permit hybridization of only those nucleic acid sequences that form stable hybrids at 65°C in 0.018 M NaCl (i.e., if the hybrid is unstable at 65°C in 0.018 M NaCl, it is unstable under the high stringency conditions contemplated herein). High stringency conditions can be provided, for example, by hybridization in conditions equivalent to 50% formamide, 5×Denhart's solution, 5×SSPE, 0.2% SDS at 42°C, followed by washing in 0.1×SSPE and 0.1% SDS at 65°C. Another high stringency condition is hybridization in conditions equivalent to hybridization in 5X SSC containing 0.1% (w / v) SDS at 65°C and washing in 0.1x SSC containing 0.1% SDS at 65°C. Other high stringency hybridization conditions as well as moderate stringency conditions are described in the references cited above.

[0076] As used herein, "codon-optimized" refers to the alteration of the codons of a polynucleotide encoding a protein to those codons that are preferentially used in a particular organism such that the encoded protein is efficiently expressed in the organism of interest. Although the genetic code is degenerate, i.e., most amino acids are represented by several codons referred to as "synonyms" or "synonymous" codons, it is well known that the codon usage of a particular organism is non-random and biased for particular codon triplets. Such codon usage bias may be higher with respect to a given gene, genes having a common function or ancestral origin, highly expressed proteins versus low copy number proteins, and the aggregated protein-coding regions of the genome of an organism. In some embodiments, a polynucleotide encoding purine nucleoside phosphorylase can be codon-optimized for optimized production in a selected host organism for expression.

[0077] As used herein, "preferred", "optimal", and "high codon usage bias" codons, when used singly or in combination, can interchangeably refer to codons that are used at a higher frequency than other codons encoding the same amino acid in a protein-coding region. Preferred codons can be determined based on a single gene, a group of genes of common function or origin, codon usage in highly expressed genes, codon frequencies in the aggregated protein-coding regions of an entire organism, codon frequencies in the aggregated protein-coding regions of related organisms, or combinations thereof. Codons whose frequency increases with the level of gene expression are generally the optimal codons for expression. Various methods for determining codon frequencies (e.g., codon usage, relative synonymous codon usage) and codon preferences and the effective number of codons used in a gene in a particular organism are known, including multivariate analyses such as using cluster analysis or correlation analysis (see, e.g., GCG CodonPreference, Genetics Computer Group Wisconsin Package; CodonW, Peden, University of Nottingham; McInerney, Bioinform., 14:372-73

[1998] ; Stenico et al., Nucl. Acids Res., 22:2437-46

[1994] ; and Wright, Gene 87:23-29

[1990] ). Codon usage tables for many different organisms are available (see, e.g., Wada et al., Nucl. Acids Res., 20:2111-2118

[1992] ; Nakamura et al., Nucl. Acids Res., 28:292

[2000] ; Duret et al., supra; Henaut and Danchin, in Escherichia coli and Salmonella Neidhardt et al. (eds.), ASM Press, Washington D.C., p. 2047-2066

[1996] ). Data sources for obtaining codon usage can depend on any available nucleotide sequence capable of encoding a protein. These data sets include nucleic acid sequences actually known to encode expressed proteins (e.g., complete protein-coding sequences - CDSs), expressed sequence tags (ESTs), or predicted coding regions of genomic sequences (see, e.g., Mount, Bioinformatics:Sequence and Genome Analysis, Chapter 8, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y.

[2001] ; Uberbacher, Meth. Enzymol., 266:259-281

[1996] ; and Tiwari et al., Comput. Appl. Biosci., 13:263-270

[1997] ).

[0078] As used herein, "control sequence" includes all components that are necessary or advantageous for the expression of a polynucleotide and / or polypeptide of the invention. Each control sequence may be native or foreign to the nucleic acid sequence encoding the polypeptide. Such control sequences include, but are not limited to, leader sequences, polyadenylation sequences, propeptide sequences, promoter sequences, signal peptide sequences, initiation sequences, and transcription terminator sequences. At a minimum, the control sequences include a promoter and transcription and translation termination signals. For the purpose of introducing specific restriction sites, the control sequences may be provided with linkers that facilitate the ligation of the control sequences to the coding region of the nucleic acid sequence encoding the polypeptide.

[0079] "Operably linked" is defined herein as a configuration in which a control sequence is placed at an appropriate position (i.e., in a functional relationship) relative to a polynucleotide of interest such that the control sequence directs or regulates the expression of the polynucleotide of interest and / or the polypeptide.

[0080] "Promoter sequence" refers to a nucleic acid sequence recognized by a host cell for the expression of a polynucleotide of interest, such as a coding sequence. The promoter sequence contains transcriptional control sequences that mediate the expression of the polynucleotide of interest. The promoter may be any nucleic acid sequence that shows transcriptional activity in the selected host cell, including mutant, truncated, and hybrid promoters, and may be obtained from genes encoding extracellular or intracellular polypeptides that are homologous or heterologous to the host cell.

[0081] The phrase "suitable reaction conditions" refers to those conditions (e.g., ranges of enzyme loading, substrate loading, temperature, pH, buffer, cosolvent, etc.) in an enzymatic conversion reaction solution under which the purine nucleoside phosphorylase polypeptide of the invention is capable of converting a substrate into a desired product compound. Some exemplary "suitable reaction conditions" are provided herein.

[0082] As used herein, "loading", such as in "compound loading" or "enzyme loading", refers to the concentration or amount of a component in a reaction mixture at the start of the reaction.

[0083] As used herein, in the context of an enzymatic conversion reaction process, a "substrate" refers to a compound or molecule acted upon by an engineered enzyme provided herein (e.g., an engineered purine nucleoside phosphorylase polypeptide).

[0084] As used herein, an "increased" yield of a product (e.g., a deoxyribose phosphate analog) resulting from a reaction occurs when a particular component (e.g., purine nucleoside phosphorylase) present during the reaction causes more of the product to be produced as compared to a reaction carried out under the same conditions with the same substrate and other substituents but in the absence of the component of interest.

[0085] If the amount of a particular enzyme is less than about 2%, about 1%, or about 0.1% (wt / wt) compared to other enzymes involved in a catalytic reaction, the reaction is said to be "substantially free" of that enzyme.

[0086] As used herein, "fractionating" a liquid (e.g., a culture broth) refers to applying a separation process (e.g., salt precipitation, column chromatography, size exclusion, and filtration) or a combination of such processes to provide a solution in which the percentage of the desired protein in the solution is greater than the percentage in the initial liquid product.

[0087] As used herein, a "starting composition" refers to any composition that contains at least one substrate. In some embodiments, the starting composition contains any suitable substrate.

[0088] As used herein, in the context of an enzymatic conversion process, a "product" refers to a compound or molecule that results from the action of an enzyme polypeptide on a substrate.

[0089] As used herein, "equilibration" as used herein refers to the process of determining the steady-state concentration of chemical species in a chemical or enzymatic reaction (e.g., the interconversion of two species A and B), including the interconversion of stereoisomers, as determined by the forward and reverse rate constants of the chemical or enzymatic reaction.

[0090] As used herein, "alkyl" refers to a saturated hydrocarbon group having from 1 to 18 carbon atoms (including the termini), straight-chain or branched-chain, more preferably from 1 to 8 carbon atoms (including the termini), and most preferably from 1 to 6 carbon atoms (including the termini). An alkyl having a specified number of carbon atoms is indicated in parentheses (e.g., (C 1 -C 4 ) alkyl refers to an alkyl having from 1 to 4 carbon atoms).

[0091] As used herein, "alkenyl" refers to a group having from 2 to 12 carbon atoms (including the termini), straight-chain or branched-chain, containing at least one double bond but optionally containing more than one double bond.

[0092] As used herein, "alkynyl" refers to a straight-chain or branched-chain group having from 2 to 12 carbon atoms (including the termini), containing at least one triple bond but optionally containing more than one triple bond, and additionally optionally containing one or more double-bonded moieties.

[0093] As used herein, "heteroalkyl", "heteroalkenyl", and "heteroalkynyl" refer to alkyl, alkenyl, and alkynyl as defined herein, wherein one or more carbon atoms are each independently replaced by the same or different heteroatoms or heteroatom groups. Heteroatoms and / or heteroatom groups that can replace carbon atoms include, but are not limited to, -O-, -S-, -S-O-, -NR α -, -PH-, -S(O)-, -S(O)2-, -S(O)NR α -, -S(O) 2 NR α -, etc., including combinations thereof, wherein each R α is independently selected from hydrogen, alkyl, heteroalkyl, cycloalkyl, heterocycloalkyl, aryl, and heteroaryl.

[0094] As used herein, "alkoxy" refers to the group -OR β , wherein R β is an alkyl group as defined above, including optionally substituted alkyl groups as also defined herein.

[0095] As used herein, "aryl" refers to an unsaturated aromatic carbocyclic group having from 6 to 12 carbon atoms (including the termini) with a single ring (e.g., phenyl) or more than one fused ring (e.g., naphthyl or anthracenyl). Exemplary aryl groups include phenyl, pyridyl, naphthyl, etc.

[0096] As used herein, "amino" refers to the group -NH 2 . Substituted amino refers to the groups -NHR δ , NR δ R δ and NR δ R δ R δ , wherein each R δ is independently selected from substituted or unsubstituted alkyl, cycloalkyl, cycloheteroalkyl, alkoxy, aryl, heteroaryl, heteroarylalkyl, acyl, alkoxycarbonyl, sulfanyl, sulfinyl, sulfonyl, etc. Representative amino groups include, but are not limited to, dimethylamino, diethylamino, trimethylammonium, triethylammonium, methylsulfonylamino, furanyl-oxy-sulfamino, etc.

[0097] As used herein, "oxo" refers to =O.

[0098] As used herein, "oxy" refers to the divalent group -O-, which may have various substituents to form different oxy groups, including ethers and esters.

[0099] As used herein, "carboxy" refers to -COOH.

[0100] As used herein, "carbonyl" refers to -C(O)-, which may have various substituents to form different carbonyl groups, including acids, acyl halides, aldehydes, amides, esters, and ketones.

[0101] As used herein, "alkoxycarbonyl" refers to -C(O)OR ε , where R ε is an alkyl group as defined herein, which may be optionally substituted.

[0102] As used herein, "aminocarbonyl" refers to -C(O)NH 2 . Substituted aminocarbonyl refers to -C(O)NR δ R δ , where the amino group NR δ R δ is as defined herein.

[0103] As used herein, "halogen" and "halo" refer to fluorine, chlorine, bromine, and iodine.

[0104] As used herein, "hydroxy" refers to -OH.

[0105] As used herein, "cyano" refers to -CN.

[0106] As used herein, "heteroaryl" refers to an aromatic heterocyclic group having from 1 to 10 carbon atoms (including the termini) and from 1 to 4 heteroatoms (including the termini) selected from oxygen, nitrogen, and sulfur within the ring. Such heteroaryl groups may have a monocyclic ring (e.g., pyridyl or furyl) or more than one fused ring (e.g., indolizinyl or benzothienyl).

[0107] As used herein, "heteroarylalkyl" refers to an alkyl group substituted by a heteroaryl (i.e., a heteroaryl-alkyl-group), preferably having from 1 to 6 carbon atoms (including the termini) in the alkyl portion and from 5 to 12 ring atoms (including the termini) in the heteroaryl portion. Such heteroarylalkyl groups are exemplified by pyridylmethyl and the like.

[0108] As used herein, "heteroarylalkenyl" refers to an alkenyl group substituted by a heteroaryl (i.e., a heteroaryl-alkenyl-group), preferably having from 2 to 6 carbon atoms (including the termini) in the alkenyl portion and from 5 to 12 ring atoms (including the termini) in the heteroaryl portion.

[0109] As used herein, "heteroarylalkynyl" refers to an alkynyl group substituted with a heteroaryl (i.e., a heteroaryl-alkynyl-group), preferably having from 2 to 6 carbon atoms (including the termini) in the alkynyl moiety and from 5 to 12 ring atoms (including the termini) in the heteroaryl moiety.

[0110] As used herein, "heterocycle", "heterocyclic" and interchangeably "heterocyclic hydrocarbyl" refer to a saturated or unsaturated group having a single ring or more than one fused ring, having from 2 to 10 carbon ring atoms (including the termini) and from 1 to 4 heteroatoms (including the termini) selected from nitrogen, sulfur or oxygen within the ring. Such heterocyclic groups can have a single ring (e.g., piperidinyl or tetrahydrofuranyl) or more than one fused ring (e.g., dihydroindolyl, dihydrobenzofuran or quinuclidinyl). Examples of heterocycles include, but are not limited to, furan, thiophene, thiazole, oxazole, pyrrole, imidazole, pyrazole, pyridine, pyrazine, pyrimidine, pyridazine, indolizine, isoindole, indole, indazole, purine, quinolizine, isoquinoline, quinoline, phthalazine, naphthylpyridine, quinoxaline, quinazoline, cinnoline, pteridine, carbazole, carboline, phenanthridine, acridine, phenanthroline, isothiazole, phenazine, isoxazole, phenoxazine, phenothiazine, imidazolidine, imidazoline, piperidine, piperazine, pyrrolidine, indoline, etc.

[0111] As used herein, "membered ring" means including any cyclic structure. The number preceding the term "membered" represents the number of skeletal atoms making up the ring. Thus, for example, cyclohexyl, pyridine, pyran and thiopyran are 6-membered rings, and cyclopentyl, pyrrole, furan and thiophene are 5-membered rings.

[0112] Unless otherwise indicated, the positions occupied by hydrogen in the foregoing groups may be further substituted by substituents such as, but not limited to, the following: hydroxy, oxo, nitro, methoxy, ethoxy, alkoxy, substituted alkoxy, trifluoromethoxy, haloalkoxy, fluorine, chlorine, bromine, iodine, halogen, methyl, ethyl, propyl, butyl, alkyl, alkenyl, alkynyl, substituted alkyl, trifluoromethyl, haloalkyl, hydroxyalkyl, alkoxyalkyl, thio, alkylthio, acyl, carboxy, alkoxycarbonyl, formamido, substituted formamido, alkylsulfonyl, alkylsulfinyl, alkylsulfonylamino, sulfonamido, substituted sulfonamido, cyano, amino, substituted amino, alkylamino, dialkylamino, aminoalkyl, acylamino, amidino, amidoximo, hydroxamoyl, phenyl, aryl, substituted aryl, aryloxy, arylalkyl, arylalkenyl, arylalkynyl, pyridyl, imidazolyl, heteroaryl, substituted heteroaryl, heteroaryloxy, heteroarylalkyl, heteroarylalkenyl, heteroarylalkynyl, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloalkyl, cycloalkenyl, cycloalkylalkyl, substituted cycloalkyl, cycloalkyloxy, pyrrolidinyl, piperidinyl, morpholino, heterocycle, (heterocycle)oxy and (heterocycle)alkyl; and the preferred heteroatoms are oxygen, nitrogen and sulfur. It is understood that where there are open valences on these substituents, they may be further substituted by alkyl, cycloalkyl, aryl, heteroaryl and / or heterocyclic groups, and where there are open valences on carbon, they may be further substituted by halogen and oxygen-, nitrogen- or sulfur-bonded substituents, and where there are more than one such open valences, these groups may be joined to form a ring either by direct bond formation or by formation of a bond with a new heteroatom (preferably oxygen, nitrogen or sulfur). It is also understood that the above substitutions may be made provided that replacement of hydrogen by a substituent does not confer unacceptable instability on the molecules of the present invention and is otherwise chemically reasonable.

[0113] As used herein, the term "culturing" refers to the growth of a microbial cell population under any suitable conditions (e.g., using a liquid, gel or solid medium).

[0114] Recombinant polypeptides can be produced using any suitable method known in the art. Genes encoding wild-type polypeptides of interest can be cloned into vectors such as plasmids and expressed in a desired host such as Escherichia coli and the like. Variants of recombinant polypeptides can be produced by various methods known in the art. In fact, there are a variety of different mutagenesis techniques well known to those skilled in the art. In addition, mutagenesis kits are also available from many commercial molecular biology suppliers. The methods can be used to make specific substitutions at defined amino acids (site-directed), specific (region-specific) or random mutations in local regions of the gene, or random mutagenesis within the entire gene (e.g., saturation mutagenesis). Many suitable methods for generating enzyme variants are known to those skilled in the art, including but not limited to, site-directed mutagenesis of single-stranded DNA or double-stranded DNA using PCR, cassette mutagenesis, gene synthesis, error-prone PCR, shuffling, and chemical saturation mutagenesis, or any other suitable method known in the art. Mutagenesis and directed evolution methods can be readily applied to polynucleotides encoding enzymes to generate variant libraries that can be expressed, screened, and assayed. Any suitable mutagenesis and directed evolution method can be used in the present invention and is well known in the art (see, for example, U.S. Patent Nos. 5,605,793, 5,811,238, 5,830,721, 5,834,252, 5,837,458, 5,928,905, 6,096,548, 6,117,679, 6,132,970, 6,165,793, 6,180,406, 6,251,674, 6,265,201, 6,277,638, 6,287,861, 6,287,862, 6,291,242, 6,297,053, 6,303,344, 6,309,883, 6,319,713, 6,319,714, 6,323,030, 6,326,204, 6,335,160, 6,335,198, 6,344,356, 6,352,859, 6,355,484, 6,358,740, 6,358,742, 6,365,377, 6,365,408, 6,368,861, 6,372,497, 6,337,186, 6,376,246, 6,379,964, 6,387,702, 6,391,552, 6,391,640, 6,395,547, 6,406,855, 6,406,910, 6,413,745, 6,413,774, 6,420,175, 6,423,542, 6,426,224, 6,436,675, 6,444,468, 6,455,253, 6,479,652, 6,482,647, 6,483,011, 6,484,105, 6,489,146, 6,500,617, 6,500,639, 6,506,U.S. Patent Nos. 6,026,506, 6,506,603, 6,518,065, 6,519,065, 6,521,453, 6,528,311, 6,537,746, 6,573,098, 6,576,467, 6,579,678, 6,586,182, 6,602,986, 6,605,430, 6,613,514, 6,653,072, 6,686,515, 6,703,240, 6,716,631, 6,825,001, 6,902,922, 6,917,882, 6,946,296, 6,961,664, 6,995,017, 7,024,312, 7,058,515, 7,105,297, 7,148,054, 7,220,566, 7,288,375, 7,384,387, 7,421,347, 7,430,477, 7,462,469, 7,534,564, 7,620,500, 7,620,502, 7,629,170, 7,702,464, 7,747,391, 7,747,393, 7,751,986, 7,776,598, 7,783,428, 7,795,030, 7,853,410, 7,868,138, 7,783,428, 7,873,477, 7,873,499, 7,904,249, 7,957,912, 7,981,614, 8,014,961, 8,029,988, 8,048,674, 8,058,001, 8,076,138, 8,108,150, 8,170,806, 8,224,580, 8,377,681, 8,383,346, 8,457,903, 8,504,498, 8,589,085, 8,762,066, 8,768,871, 9,593,326, and all related U.S. and PCT and non-U.S. corresponding applications; Ling et al., Anal. Biochem., 254(2):157-78

[1997] ; Dale et al., Meth. Mol. Biol., 57:369-74

[1996] ; Smith, Ann. Rev. Genet., 19:423-462

[1985] ; Botstein et al., Science, 229:1193-1201

[1985] ; Carter, Biochem. J., 237:1-7

[1986] ; Kramer et al., Cell, 38:879-887

[1984] ; Wells et al., Gene, 34:315-323

[1985] ; Minshull et al., Curr. Op. Chem. Biol.,3:284 - 290

[1999] ; Christians et al., Nat. Biotechnol., 17:259 - 264

[1999] ; Crameri et al., Nature, 391:288 - 291

[1998] ; Crameri, et al., Nat. Biotechnol., 15:436 - 438

[1997] ; Zhang et al., Proc. Nat. Acad. Sci. U.S.A., 94:4504 - 4509

[1997] ; Crameri et al., Nat. Biotechnol., 14:315 - 319

[1996] ; Stemmer, Nature, 370:389 - 391

[1994] ; Stemmer, Proc. Nat. Acad. Sci. USA, 91:10747 - 10751

[1994] ; WO 95 / 22625; WO 97 / 0078; WO 97 / 35966; WO 98 / 27230; WO 00 / 42651; WO 01 / 75767; and WO 2009 / 152336, which are hereby incorporated by reference in their entirety).

[0115] In some embodiments, enzyme clones obtained after mutagenesis treatment are screened by subjecting the enzyme preparation to a defined temperature (or other assay conditions) and measuring the amount of enzyme activity remaining after heat treatment or other suitable assay conditions. Clones containing polynucleotides encoding the polypeptide are then isolated from the gene, sequenced to identify nucleotide sequence alterations (if any), and used to express the enzyme in a host cell. Measuring enzyme activity from the expression library can be performed using any suitable method known in the art (e.g., standard biochemical techniques such as HPLC analysis).

[0116] After variant generation, any desired properties can be screened for (e.g., high or increased activity, or low or decreased activity, increased thermal activity, increased thermal stability, and / or acidic pH stability, etc.). In some embodiments, a “recombinant purine nucleoside phosphorylase polypeptide” (also referred to herein as an “engineered purine nucleoside phosphorylase polypeptide”, “variant purine nucleoside phosphorylase”, “purine nucleoside phosphorylase variant”, and “purine nucleoside phosphorylase combinatorial variant”) can be used. In some embodiments, a “recombinant purine nucleoside phosphorylase polypeptide” (also known as an “engineered purine nucleoside phosphorylase polypeptide”, “variant purine nucleoside phosphorylase”, “purine nucleoside phosphorylase variant”, and “purine nucleoside phosphorylase combinatorial variant”) can be used.

[0117] As used herein, "vector" is a DNA construct used to introduce a DNA sequence into a cell. In some embodiments, the vector is an expression vector operably linked to appropriate control sequences capable of effecting the expression of a polypeptide encoded in the DNA sequence in a suitable host. In some embodiments, an "expression vector" has a promoter sequence operably linked to a DNA sequence (e.g., a transgene) to drive expression in a host cell, and in some embodiments, also contains a transcription terminator sequence.

[0118] As used herein, the term "expression" includes any step involved in polypeptide production, including but not limited to, transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses the secretion of the polypeptide from the cell.

[0119] As used herein, the term "production" refers to the production of a protein and / or other compound from a cell. It is intended that the term encompasses any step involved in polypeptide production, including but not limited to, transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses the secretion of the polypeptide from the cell.

[0120] As used herein, if an amino acid or nucleotide sequence (e.g., a promoter sequence, signal peptide, terminator sequence, etc.) is not associated in nature with another sequence to which it is operably linked, the two sequences are "heterologous". For example, a "heterologous polynucleotide" is any polynucleotide introduced into a host cell by laboratory techniques, and includes a polynucleotide removed from a host cell, subjected to laboratory manipulation, and then reintroduced into the host cell.

[0121] As used herein, the terms "host cell" and "host strain" refer to a suitable host for an expression vector containing the DNA provided herein (e.g., a polynucleotide encoding a purine nucleoside phosphorylase variant). In some embodiments, the host cell is a prokaryotic or eukaryotic cell that has been transformed or transfected with a vector constructed using recombinant DNA techniques known in the art.

[0122] The term "analogue" means a polypeptide having more than 70% sequence identity, but less than 100% sequence identity, to a reference polypeptide (e.g., more than 75%, 78%, 80%, 83%, 85%, 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity). In some embodiments, an analogue means a polypeptide that contains one or more non-naturally occurring amino acid residues (including but not limited to homoarginine, ornithine, and norvaline) as well as naturally occurring amino acids. In some embodiments, an analogue also includes one or more D-amino acid residues and non-peptide linkages between two or more amino acid residues.

[0123] The term "effective amount" means an amount sufficient to produce the desired result. One of ordinary skill in the art can determine what an effective amount is by using routine experimentation.

[0124] The terms "isolated" and "purified" are used to refer to a molecule (e.g., an isolated nucleic acid, polypeptide, etc.) or other component that has been removed from at least one other component with which it is naturally associated. The term "purified" does not require absolute purity, but is intended as a relative definition.

[0125] As used herein, "stereoselectivity" refers to the preferential formation of one stereoisomer over another in a chemical or enzymatic reaction. Stereoselectivity can be partial, in which case the formation of one stereoisomer is favored over the other, or stereoselectivity can be complete, in which case only one stereoisomer is formed. When the stereoisomers are enantiomers, the stereoselectivity is referred to as enantioselectivity, i.e., the fraction of one enantiomer in the sum of both (usually reported as a percentage). Alternatively, it is commonly reported in the art as the enantiomeric excess ("e.e.") (usually as a percentage) calculated according to the formula: [major enantiomer - minor enantiomer] / [major enantiomer + minor enantiomer]. In the case where the stereoisomers are diastereomers, the stereoselectivity is referred to as diastereoselectivity, i.e., the fraction of one diastereomer in a mixture of two diastereomers (usually reported as a percentage), and is usually alternatively reported as the diastereomeric excess ("d.e."). Enantiomeric excess and diastereomeric excess are types of stereoisomeric excess.

[0126] As used herein, the terms "regioselective" and "regioselective reaction" refer to a reaction in which one direction of bond formation or cleavage occurs preferentially over all other possible directions. If the discrimination is complete, the reaction can be completely (100%) regioselective, if the reaction product at one site is favored over the reaction products at other sites, the reaction can be substantially regioselective (at least 75%), or partially regioselective (x%, where the percentage depends on the reaction setup of interest).

[0127] As used herein, "chemoselectivity" refers to the preferential formation of one product over another in a chemical or enzymatic reaction.

[0128] As used herein, "pH-stable" refers to a purine nucleoside phosphorylase polypeptide that maintains similar activity (e.g., more than 60% to 80%) after exposure to high or low pH (e.g., 4.5 - 6 or 8 to 12) for a period of time (e.g., 0.5 - 24 hours) compared to the untreated enzyme.

[0129] As used herein, "thermostable" refers to a purine nucleoside phosphorylase polypeptide that maintains a similar activity (e.g., more than 60% to 80%) after exposure to the same elevated temperature (e.g., 40°C to 80°C) for a period of time (e.g., 0.5 h - 24 h) compared to the wild-type enzyme exposed to the same elevated temperature.

[0130] As used herein, "solvent-stable" refers to a purine nucleoside phosphorylase polypeptide that maintains a similar activity (more than e.g., 60% to 80%) after exposure to the same concentration of the same solvent (ethanol, isopropanol, dimethyl sulfoxide [DMSO], tetrahydrofuran, 2-methyltetrahydrofuran, acetone, toluene, butyl acetate, methyl tert-butyl ether, etc.) for a period of time (e.g., 0.5 h - 24 h) compared to the wild-type enzyme exposed to different concentrations (e.g., 5% - 99%) of the solvent.

[0131] As used herein, "thermostable and solvent-stable" refers to a purine nucleoside phosphorylase polypeptide that is both thermostable and solvent-stable.

[0132] As used herein, "optional" and "optionally" mean that the subsequent described event or circumstance may or may not occur, and mean that the description includes instances when the event or circumstance occurs and instances when the event or circumstance does not occur. One of ordinary skill in the art will understand that for any molecule described as including one or more optional substituents, only those compounds that are spatially achievable and / or synthetically feasible are intended to be included.

[0133] As used herein, "optionally substituted" refers to all subsequent modifiers in one or a series of chemical groups. For example, in the term "optionally substituted arylalkyl", the "alkyl" part and the "aryl" part of the molecule may or may not be substituted, and for a series of "optionally substituted alkyl, cycloalkyl, aryl, and heteroaryl", the alkyl group, cycloalkyl group, aryl group, and heteroaryl group may each independently be substituted or may not be substituted. DETAILED DESCRIPTION OF THE INVENTION

[0135] The present invention provides engineered purine nucleoside phosphorylase (PNP) enzymes, polypeptides having PNP activity, polynucleotides encoding these enzymes, as well as vectors and host cells comprising these polynucleotides and polypeptides. Methods for producing PNP enzymes are also provided. The present invention also provides compositions comprising PNP enzymes, and methods of using the engineered PNP enzymes. The present invention is particularly useful for the production of pharmaceutical compounds.

[0136] In some embodiments, the present invention provides enzymes suitable for generating nucleoside analogs such as MK-8591 (Merck). The present invention was developed to address the potential use of enzymes for generating these nucleoside analogs. However, it has been determined that one challenge with this approach is that wild-type enzymes are unlikely to be the best choice for the required substrate analogs to generate all of the necessary intermediates. Additionally, each enzyme in the synthetic pathway requires some engineering to make it compatible with alternative substrates and the processes used to synthesize the desired nucleoside analogs.

[0137] In some embodiments, the present invention provides enzymes that can be used to generate compounds, ultimately resulting in an in vitro enzymatic synthesis method for the unnatural nucleoside analog shown in Compound (1).

[0138]

[0139] Unnatural nucleosides are important building blocks for many important classes of drugs, including those used to treat cancer and viral infections. There are at least a dozen nucleoside analog drugs on the market or in clinical trials (Jordheim et al., Nat. Rev. Drug Discovery 12:447 - 464

[2013] ). One method for preparing Compound (1) is through the coupling of ethynyl ribose-1-phosphate (Compound (3)) and fluoroadenine (Compound (2)) catalyzed by purine nucleoside phosphorylase (PNP), as shown in Scheme I.

[0140]

[0141] Scheme I. Reaction Catalyzed by Purine Nucleoside Phosphorylase (PNP)

[0142] Deoxyribose-1-phosphate compounds, such as Compound (3), can be difficult to prepare. However, the corresponding deoxyribose-5-phosphate compounds can be prepared by the coupling of acetaldehyde and D-glyceraldehyde-3-phosphate (or analogs thereof) catalyzed by the enzyme 2-deoxyribose-5-phosphate aldolase (DERA) (Barbas et al., J. Am. Chem. Soc. 112:2013 - 2014

[1990] ). After the formation of the deoxyribose-1-phosphate analog, it can be converted or isomerized to the corresponding deoxyribose-5-phosphate analog required for Scheme I through the action of the enzyme phosphopentose mutase (PPM), as shown in Scheme II.

[0143]

[0144] Scheme II

[0145] The equilibrium positions of the PNP and PPM reactions shown in Scheme I generally favor the reactants (compounds 2 and 4) over the products (compound 1 and inorganic phosphate). One way to drive the reaction to higher conversion is to remove the inorganic phosphate formed in the coupling step. This can be achieved by reacting the inorganic phosphate with a disaccharide such as sucrose in the presence of the enzyme sucrose phosphorylase (SP) (see, for example, U.S. Patent No. 7,229,797). This reaction produces glucose-1-phosphate and fructose, which is highly advantageous and can drive the overall reaction shown in Scheme III.

[0146]

[0147] Scheme III. Overall reaction scheme for the production of compound (1)

[0148] Purine nucleoside phosphorylase (PNP) enzymes have been isolated and / or recombinantly expressed from many sources, including Escherichia coli (Xie, Xixian et al., Biotechnol Lett 33:1107-1112

[2011] , Lee et al., Protein Expr. Purif. 22:180-188

[2001] ), Bacillus subtilis 168, species of Pseudoalteromonas sp. XM2107 (Xie, Xixian et al., Biotechnol Lett 33:1107-1112

[2011] ), Bacillus halodurans Alk36 (Visser et al., Extremophiles 14:185-192

[2010] ), Plasmodium falciparum (Schnick et al., Acta Cryst. D 61:1245-1254

[2005] ), and humans (Silva et al., Protein Expr. Purif. 27:158-164

[2003] ), among others. Crystal structures of several PNPs are also available, such as those from Escherichia coli (Mao et al., Structure 5:1373-1383

[1997] ) and humans (Canduri et al., Biochem. Biophys. Res. Commun. 26:335-338

[2005] ). These enzymes catalyze the reversible phosphorolysis of (2-deoxy) purine nucleosides to free base and (2'-deoxy) ribose-1-phosphate. A trimeric form specific for 6-oxopurine nucleosides is present in higher organisms and prokaryotes, while a hexameric form active towards both 6-oxopurine nucleosides and 6-aminopurine nucleosides is only present in lower organisms (Bennett et al., J Biol Chem 278:47110-47118

[2003] ). The activity of wild-type PNP enzymes towards non-natural nucleosides has been demonstrated (Schnick et al., Acta Cryst. D 61:1245-1254

[2005] , Visser et al., Extremophiles 14:185-192

[2010] , Birmingham et al., Nat. Chem. Biol. 10:392-399

[2014] ), but this activity is generally not high enough for the production of many non-natural nucleosides, such as compound (1), in commercial quantities.

[0149] Due to the poor activity of PNP towards unnatural substrates for the preparation of unnatural and therapeutically useful nucleosides, there is a need for engineered PNP with improved activity and capable of operating under typical industrial conditions. The present invention addresses this need and provides engineered PNP suitable for use in these reactions under industrial conditions.

[0150] Engineered PNP polypeptide

[0151] The present invention provides engineered PNP polypeptides, polynucleotides encoding the polypeptides, methods for preparing the polypeptides, and methods for using the polypeptides. When describing the polypeptide, it should be understood that it also describes the polynucleotide encoding the polypeptide. In some embodiments, the present invention provides engineered, non-naturally occurring PNP enzymes having improved properties compared to wild-type PNP enzymes. Any suitable reaction conditions can be used in the present invention. In some embodiments, methods are used to analyze the improved properties of the engineered polypeptide for the isomerization reaction. In some embodiments, as further described below and in the examples, the reaction conditions are varied according to the concentration or amount of the engineered PNP, one or more substrates, one or more buffers, one or more solvents, pH, conditions including temperature and reaction time, and / or the conditions under which the engineered PNP polypeptide is immobilized on a solid support.

[0152] In some embodiments, the reaction conditions are supplemented with additional reaction components or additional techniques. In some embodiments, these include taking measures to stabilize the enzyme or prevent enzyme inactivation, reduce product inhibition, and shift the reaction equilibrium towards the desired product formation.

[0153] In some additional embodiments, any of the methods described above for converting a substrate compound to a product compound may further include one or more steps selected from: extraction, separation, purification, crystallization, filtration, and / or lyophilization of one or more product compounds. Methods, techniques, and protocols for extracting, separating, purifying, and / or crystallizing one or more products from a biocatalytic reaction mixture produced by the methods provided herein are known to those of ordinary skill in the art and / or obtained by routine experimentation. In addition, illustrative methods are provided in the examples below.

[0154] Engineered PNP polynucleotides, expression vectors, and host cells encoding the engineered polypeptides

[0155] The present invention provides polynucleotides encoding the engineered enzyme polypeptides described herein. In some embodiments, the polynucleotide is operably linked to one or more heterologous regulatory sequences that control gene expression to produce a recombinant polynucleotide capable of expressing the polypeptide. In some embodiments, an expression construct comprising at least one heterologous polynucleotide encoding one or more engineered enzyme polypeptides is introduced into a suitable host cell to express one or more corresponding enzyme polypeptides.

[0156] As will be apparent to the skilled person, the availability of protein sequences and the knowledge of the codons corresponding to the various amino acids provide a description of all polynucleotides capable of encoding the subject polypeptides. The degeneracy of the genetic code, where the same amino acid is encoded by alternative or synonymous codons, allows the preparation of a vast number of nucleic acids, all of which encode an engineered enzyme (e.g., PNP) polypeptide. Accordingly, the present invention provides methods and compositions for selecting combinations based on possible codon options for generating each and every possible variant form of a preparable enzyme polynucleotide encoding an enzyme polypeptide described herein, and all such variant forms are considered to be specifically disclosed for any polypeptide described herein, including the amino acid sequences presented in the examples (e.g., in the respective tables).

[0157] In some embodiments, the codons are preferably optimized for utilization by the selected host cell for protein production. For example, the preferred codons used in bacteria are typically used for expression in bacteria. Thus, a codon-optimized polynucleotide encoding an engineered enzyme polypeptide contains preferred codons at about 40%, 50%, 60%, 70%, 80%, or greater than 90% of the codon positions in the full-length coding region.

[0158] In some embodiments, the enzyme polynucleotide encodes an engineered polypeptide having enzymatic activity and the properties disclosed herein, wherein the polypeptide comprises an amino acid sequence having at least 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to a reference sequence selected from the SEQ ID NOs provided herein, or an amino acid sequence of any variant (e.g., those provided in the examples), and one or more residue differences (e.g., at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acid residue positions) compared to the amino acid sequence of one or more reference polynucleotides or any variant disclosed in the examples. In some embodiments, the reference polypeptide sequences are selected from SEQ ID NO:2, 6, 126, 242, and / or 684.

[0159] In some embodiments, the polynucleotide is capable of hybridizing under highly stringent conditions to a reference polynucleotide sequence selected from any of the polynucleotide sequences provided herein or its complementary sequence, or to a polynucleotide sequence encoding any variant enzyme polypeptide provided herein. In some embodiments, the polynucleotide capable of hybridizing under highly stringent conditions encodes an enzyme polypeptide comprising an amino acid sequence having one or more residue differences compared to the reference sequence.

[0160] In some embodiments, the isolated polynucleotides encoding any of the engineered enzyme polypeptides herein are manipulated in various ways to facilitate the expression of the enzyme polypeptides. In some embodiments, the polynucleotides encoding the enzyme polypeptides constitute an expression vector in which there is one or more control sequences to regulate the expression of the enzyme polynucleotide and / or polypeptide. Depending on the expression vector used, manipulation of the isolated polynucleotide prior to insertion of the polynucleotide into the vector may be desirable or necessary. Techniques for modifying polynucleotides and nucleic acid sequences using recombinant DNA methods are well known in the art. In some embodiments, the control sequences include, among others, a promoter, a leader sequence, a polyadenylation sequence, a propeptide sequence, a signal peptide sequence, and a transcription terminator. In some embodiments, a suitable promoter is selected based on the choice of host cell. For bacterial host cells, suitable promoters for directing transcription of the nucleic acid constructs of the present disclosure include, but are not limited to, promoters obtained from: the Escherichia coli lac operon, the Streptomyces coelicolor agarase gene (dagA), the Bacillus subtilis levansucrase gene (sacB), the Bacillus licheniformis α-amylase gene (amyL), the Bacillus stearothermophilus maltogenic amylase gene (amyM), the Bacillus amyloliquefaciens α-amylase gene (amyQ), the Bacillus licheniformis penicillinase gene (penP), the Bacillus subtilis xylA and xylB genes, and the prokaryotic β-lactamase gene (see, e.g., Villa-Kamaroff et al., Proc. Natl. Acad. Sci. USA 75:3727-3731

[1978] ), as well as the tac promoter (see, e.g., DeBoer et al., Proc. Natl. Acad. Sci. USA 80:21-25

[1983] ).Exemplary promoters for filamentous fungal host cells include, but are not limited to, promoters obtained from the following genes: Aspergillus oryzae TAKA amylase, Rhizomucor miehei aspartic proteinase, Aspergillus niger neutral α-amylase, Aspergillus niger acid stable α-amylase, Aspergillus niger or Aspergillus awamori glucoamylase (glaA), Rhizomucor miehei lipase, Aspergillus oryzae alkaline protease, Aspergillus oryzae triose phosphate isomerase, Aspergillus nidulans acetamidase, and Fusarium oxysporum trypsin-like protease (see, e.g., WO 96 / 00787), as well as the NA2-tpi promoter (a hybrid of the promoters from the Aspergillus niger neutral α-amylase gene and the Aspergillus oryzae triose phosphate isomerase gene), and mutants, truncated, and hybrid promoters thereof. Exemplary yeast cell promoters may be from the following genes: Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae galactokinase (GAL1), Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP), and Saccharomyces cerevisiae 3-phosphoglycerate kinase. Other useful promoters for yeast host cells are known in the art (see, e.g., Romanos et al., Yeast 8:423-488

[1992] ).

[0161] In some embodiments, the control sequence is also a suitable transcription terminator sequence (i.e., a sequence recognized by the host cell to terminate transcription). In some embodiments, the terminator sequence is operably linked to the 3' end of the nucleic acid sequence encoding the enzyme polypeptide. Any suitable terminator that is functional in the selected host cell can be used in the present invention. Exemplary transcription terminators for filamentous fungal host cells can be obtained from the following genes: Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Aspergillus niger α-glucosidase, and Fusarium oxysporum trypsin-like protease. Exemplary terminators for yeast host cells can be obtained from the following genes: Saccharomyces cerevisiae enolase, Saccharomyces cerevisiae cytochrome C (CYC1), and Saccharomyces cerevisiae glyceraldehyde-3-phosphate dehydrogenase. Other useful terminators for yeast host cells are known in the art (see, e.g., Romanos et al., supra).

[0162] In some embodiments, the control sequence is also a suitable leader sequence (i.e., the untranslated region of the mRNA that is important for translation by the host cell). In some embodiments, the leader sequence is operably linked to the 5' end of the nucleic acid sequence encoding the enzyme polypeptide. Any suitable leader sequence that is functional in the selected host cell can be used in the present invention. Exemplary leader sequences for filamentous fungal host cells are obtained from the genes: Aspergillus oryzae TAKA amylase and Aspergillus nidulans triose phosphate isomerase. Suitable leader sequences for yeast host cells are obtained from the genes: Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae alpha-factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP).

[0163] In some embodiments, the control sequence is also a polyadenylation sequence (i.e., a sequence that is operably linked to the 3' end of the nucleic acid sequence and that, when transcribed, is recognized by the host cell as a signal to add polyadenosine residues to the transcribed mRNA). Any suitable polyadenylation sequence that is functional in the selected host cell can be used in the present invention. Exemplary polyadenylation sequences for filamentous fungal host cells include, but are not limited to, the genes: Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Fusarium oxysporum trypsin-like protease, and Aspergillus niger alpha-glucosidase. Useful polyadenylation sequences for yeast host cells are known (see, e.g., Guo and Sherman, Mol. Cell. Bio., 15:5983-5990

[1995] ).

[0164] In some embodiments, the control sequence is also a signal peptide (i.e., a coding region encoding an amino acid sequence that is linked to the amino terminus of a polypeptide and directs the encoded polypeptide into the secretory pathway of a cell). In some embodiments, the 5' end of the coding sequence of the nucleic acid sequence inherently comprises a signal peptide coding region that is naturally linked in translation reading frame to a segment of the coding region encoding the secreted polypeptide. Alternatively, in some embodiments, the 5' end of the coding sequence comprises a signal peptide coding region that is foreign to the coding sequence. Any suitable signal peptide coding region that directs the expressed polypeptide into the secretory pathway of a selected host cell can be used for the expression of one or more engineered polypeptides. Effective signal peptide coding regions for bacterial host cells include, but are not limited to, those obtained from the genes of Bacillus NCIB 11837 maltogenic amylase, Bacillus stearothermophilus α-amylase, Bacillus licheniformis subtilisin, Bacillus licheniformis β-lactamase, Bacillus stearothermophilus neutral protease (nprT, nprS, nprM), and Bacillus subtilis prsA. Additional signal peptides are known in the art (see, e.g., Simonen and Palva, Microbiol. Rev., 57:109-137

[1993] ). In some embodiments, effective signal peptide coding regions for filamentous fungal host cells include, but are not limited to, those obtained from the genes of Aspergillus oryzae TAKA amylase, Aspergillus niger neutral amylase, Aspergillus niger glucoamylase, Rhizomucor miehei aspartic proteinase, Humicola insolens cellulase, and Humicola lanuginosa lipase. Useful signal peptides for yeast host cells include, but are not limited to, those from the genes of Saccharomyces cerevisiae α-factor and Saccharomyces cerevisiae invertase.

[0165] In some embodiments, the control sequence is also a propeptide-encoding region encoding an amino acid sequence located at the amino-terminus of the polypeptide. The resulting polypeptide is referred to as a "proenzyme", "propolypeptide" or "zymogen". The propolypeptide can be converted to the mature active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the propolypeptide. The propeptide-encoding region can be obtained from any suitable source of genes including but not limited to: Bacillus subtilis alkaline protease (aprE), Bacillus subtilis neutral protease (nprT), Saccharomyces cerevisiae α-factor, Rhizomucor miehei aspartic protease, and Myceliophthora thermophila lactase (see, e.g., WO 95 / 33836). When both a signal peptide and a propeptide region are present at the amino-terminus of the polypeptide, the propeptide region is located immediately adjacent to the amino-terminus of the polypeptide and the signal peptide region is located immediately adjacent to the amino-terminus of the propeptide region.

[0166] In some embodiments, regulatory sequences are also utilized. These sequences facilitate regulation of polypeptide expression relative to host cell growth. Examples of regulatory systems are those that cause the expression of a gene to be turned on or off in response to a chemical or physical stimulus, including the presence of a regulatory compound. In prokaryotic host cells, suitable regulatory sequences include but are not limited to the lac, tac, and trp operon systems. In yeast host cells, suitable regulatory systems include but are not limited to the ADH2 system or the GAL1 system. In filamentous fungi, suitable regulatory sequences include but are not limited to the TAKA α-amylase promoter, Aspergillus niger glucoamylase promoter, and Aspergillus oryzae glucoamylase promoter.

[0167] In another aspect, the invention relates to a recombinant expression vector comprising a polynucleotide encoding an engineered enzyme polypeptide and, depending on the type of host to be introduced therewith, one or more expression control regions such as promoters and terminators, origins of replication, etc. In some embodiments, the various nucleic acids and control sequences described herein are ligated together to produce a recombinant expression vector that includes one or more convenient restriction sites to allow insertion or substitution of the nucleic acid sequence encoding the enzyme polypeptide at such sites. Alternatively, in some embodiments, the nucleic acid sequences of the invention are expressed by inserting the nucleic acid sequence or a nucleic acid construct comprising the sequence into a suitable vector for expression. In some embodiments involving the production of an expression vector, the coding sequence is positioned in the vector such that the coding sequence is operably linked to the appropriate control sequences for expression.

[0168] The recombinant expression vector can be any suitable vector (e.g., plasmid or virus) that can be conveniently subjected to recombinant DNA procedures and cause the expression of the enzyme polynucleotide sequence. The choice of vector usually depends on the compatibility of the vector with the host cell into which the vector is to be introduced. The vector can be a linear plasmid or a closed circular plasmid.

[0169] In some embodiments, the expression vector is an autonomously replicating vector (i.e., a vector that exists as an extrachromosomal entity, whose replication is independent of chromosomal replication, such as a plasmid, an extrachromosomal element, a minichromosome, or an artificial chromosome). The vector can contain any means for ensuring self-replication. In some alternative embodiments, the vector is one that, when introduced into a host cell, is integrated into the genome and replicated along with one or more chromosomes into which it is integrated. Additionally, in some embodiments, a single vector or plasmid is utilized, or two or more vectors or plasmids and / or transposons that together contain the total DNA to be introduced into the genome of the host cell are used.

[0170] In some embodiments, the expression vector contains one or more selectable markers that allow for easy selection of transformed cells. A "selectable marker" is a gene whose product provides antimicrobial or virus resistance, resistance to heavy metals, prototrophy to auxotrophs, etc. Examples of selectable markers for bacteria include, but are not limited to, the dal gene from Bacillus subtilis or Bacillus licheniformis, or markers that confer antibiotic resistance such as ampicillin, kanamycin, chloramphenicol, or tetracycline resistance. Suitable markers for yeast host cells include, but are not limited to, ADE2, HIS3, LEU2, LYS2, MET3, TRP1, and URA3. Selectable markers for use in filamentous fungal host cells include, but are not limited to, amdS (acetamidase; e.g., from Aspergillus nidulans or Aspergillus oryzae), argB (ornithine carbamoyltransferase), bar (phosphinothricin acetyltransferase; e.g., from Streptomyces hygroscopicus), hph (hygromycin phosphotransferase), niaD (nitrate reductase), pyrG (orotidine-5'-phosphate decarboxylase; e.g., from Aspergillus nidulans or Aspergillus oryzae), sC (sulfate adenylyltransferase), and trpC (anthranilate synthase), and their equivalents.

[0171] In another aspect, the present invention provides a host cell comprising at least one polynucleotide encoding at least one engineered enzyme polypeptide of the present invention, the polynucleotide being operably linked to one or more control sequences for expressing one or more engineered enzymes in the host cell. Host cells suitable for use in expressing polypeptides encoded by the expression vectors of the present invention are well known in the art and include, but are not limited to, bacterial cells such as Escherichia coli, Vibrio fluvialis, Streptomyces, and Salmonella typhimurium cells; fungal cells such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC accession number 201178)); insect cells such as Drosophila S2 and Spodoptera Sf9 cells; animal cells such as CHO, COS, BHK, 293, and Bowes melanoma cells; and plant cells. Exemplary host cells also include various strains of Escherichia coli (e.g., W3110 (ΔfhuA) and BL21). Examples of selectable markers for bacteria include, but are not limited to, the dal gene from Bacillus subtilis or Bacillus licheniformis, or markers conferring antibiotic resistance such as ampicillin, kanamycin, chloramphenicol, and / or tetracycline resistance.

[0172] In some embodiments, the expression vectors of the present invention contain elements that permit the vector to integrate into the genome of the host cell or to replicate autonomously in the cell independent of the genome. In some embodiments involving integration into the genome of the host cell, the vector relies on the nucleic acid sequence encoding the polypeptide or any other element of the vector for integrating the vector into the genome by homologous or non-homologous recombination.

[0173] In some alternative embodiments, the expression vector contains additional nucleic acid sequences for directing integration into the genome of the host cell by homologous recombination. The additional nucleic acid sequences enable the vector to integrate into the genome of the host cell at one or more precise locations in one or more of the chromosomes. To increase the likelihood of integration at a precise location, the integration element preferably contains a sufficient number of nucleotides, such as 100 to 10,000 base pairs, preferably 400 to 10,000 base pairs, and most preferably 800 to 10,000 base pairs, which are highly homologous to the corresponding target sequence to enhance the likelihood of homologous recombination. The integration element can be any sequence homologous to the target sequence in the genome of the host cell. In addition, the integration element can be a non-coding or coding nucleic acid sequence. In another aspect, the vector can integrate into the genome of the host cell by non-homologous recombination.

[0174] For autonomous replication, the vector can also contain an origin of replication, such that the vector is capable of autonomous replication in the host cell under discussion. Examples of bacterial origins of replication are the P15A ori, which permits replication in Escherichia coli, or the origins of replication of plasmids pBR322, pUC19, pACYCl77 (which plasmid has the P15A ori), or pACYC184, and the origins of replication of pUB110, pE194, or pTA1060, which permit replication in Bacillus. Examples of origins of replication for use in yeast host cells are the 2μm origin of replication, ARS1, ARS4, the combination of ARS1 and CEN3, and the combination of ARS4 and CEN6. The origin of replication can be an origin of replication having a mutation that renders it temperature-sensitive for function in the host cell (see, for example, Ehrlich, Proc. Natl. Acad. Sci. USA 75:1433

[1978] ).

[0175] In some embodiments, more than one copy of the nucleic acid sequence of the invention is inserted into the host cell to increase production of the gene product. The increase in the copy number of the nucleic acid sequence can be obtained by integrating at least one additional copy of the sequence into the host cell genome, or by including an amplifiable selectable marker gene in the nucleic acid sequence, where cells containing amplified copies of the selectable marker gene and thus additional copies of the nucleic acid sequence can be selected by culturing the cells in the presence of a suitable selection agent.

[0176] Many expression vectors for use in the present invention are commercially available. Suitable commercial expression vectors include, but are not limited to, the p3xFLAGTM TM expression vector (Sigma-Aldrich Chemicals), which includes a CMV promoter and an hGH polyadenylation site for expression in mammalian host cells, as well as a pBR322 origin of replication and an ampicillin resistance marker for amplification in Escherichia coli. Other suitable expression vectors include, but are not limited to, pBluescriptII SK(-) and pBK-CMV (Stratagene), and plasmids derived from pBR322 (Gibco BRL), pUC (Gibco BRL), pREP4, pCEP4 (Invitrogen), or pPoly (see, for example, Lathe et al., Gene 57:193-201

[1987] ).

[0177] Thus, in some embodiments, a vector comprising a sequence encoding at least one variant purine nucleoside phosphorylase is transformed into a host cell to permit propagation of the vector and expression of one or more variant purine nucleoside phosphorylases. In some embodiments, the variant purine nucleoside phosphorylase is post-translationally modified to remove the signal peptide and, in some cases, may be cleaved after secretion. In some embodiments, the transformed host cells described above are cultured in a suitable nutrient medium under conditions that permit expression of one or more variant purine nucleoside phosphorylases. Any suitable medium for culturing host cells can be used in the present invention, including but not limited to minimal media or complex media containing suitable supplements. In some embodiments, the host cells are grown in HTP medium. Suitable media are available from a number of commercial suppliers or may be prepared according to published formulations (e.g., in the catalog of the American Type Culture Collection).

[0178] In another aspect, the invention provides a host cell comprising a polynucleotide encoding an improved purine nucleoside phosphorylase polypeptide provided herein, the polynucleotide operably linked to one or more control sequences for expressing purine nucleoside phosphorylase in the host cell. Host cells for expressing the purine nucleoside phosphorylase polypeptide encoded by the expression vectors of the invention are well known in the art and include, but are not limited to, bacterial cells such as Escherichia coli, Bacillus megaterium, Lactobacillus kefir, Streptomyces spp., and Salmonella typhimurium cells; fungal cells such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC accession number 201178)); insect cells such as Drosophila S2 and Spodoptera Sf9 cells; animal cells such as CHO, COS, BHK, 293, and Bowes melanoma cells; and plant cells. Suitable media and growth conditions for the host cells described above are well known in the art.

[0179] The polynucleotide for expressing purine nucleoside phosphorylase can be introduced into cells by a variety of methods known in the art. Techniques include, among others, electroporation, biolistic particle bombardment, liposome-mediated transfection, calcium chloride transfection, and protoplast fusion. The various methods for introducing polynucleotides into cells are known to those of skill in the art.

[0180] In some embodiments, the host cell is a eukaryotic cell. Suitable eukaryotic host cells include, but are not limited to, fungal cells, algal cells, insect cells, and plant cells. Suitable fungal host cells include, but are not limited to, Ascomycota, Basidiomycota, Deuteromycota, Zygomycota, Fungi imperfecti. In some embodiments, the fungal host cell is a yeast cell and a filamentous fungal cell. The filamentous fungal host cells of the present invention include all filamentous forms of Eumycotina and Oomycota. Filamentous fungi are characterized by a vegetative mycelium, in which the cell wall is composed of chitin, cellulose, and other complex polysaccharides. The filamentous fungal host cells of the present invention are morphologically different from yeasts.

[0181] In some embodiments of the present invention, the filamentous fungal host cell is any suitable genus and species, including but not limited to: Achlya, Acremonium, Aspergillus, Aureobasidium, Bjerkandera, Ceriporiopsis, Cephalosporium, Chrysosporium, Cochliobolus, Corynascus, Cryphonectria, Cryptococcus, Coprinus, Coriolus, Diplodia, Endothis, Fusarium, Gibberella, Gliocladium, Humicola, Hypocrea, Myceliophthora, Mucor, Neurospora, Penicillium, Podospora, Phlebia, Piromyces, Pyricularia, Rhizomucor, Rhizopus, Schizophyllum, Scytalidium, Sporotrichum, Talaromyces, Thermoascus, Thielavia, Trametes, Tolypocladium, Trichoderma, Verticillium, and / or Volvariella, and / or sexual or asexual forms, as well as synonyms, basionyms or taxonomic equivalents thereof.

[0182] In some embodiments of the present invention, the host cell is a yeast cell, including but not limited to cells of species of the genus Candida, Hansenula, Saccharomyces, Schizosaccharomyces, Pichia, Kluyveromyces or Yarrowia. In some embodiments of the present invention, the yeast cell is Hansenula polymorpha, Saccharomyces cerevisiae, Saccharomyces carlsbergensis, Saccharomyces diastaticus, Saccharomyces norbensis, Saccharomyces kluyveri, Schizosaccharomyces pombe, Pichia pastoris, Pichia finlandica, Pichia trehalophila, Pichia kodamae, Pichia membranaefaciens, Pichia opuntiae, Pichia thermotolerans, Pichia salictaria, Pichia quercuum, Pichia pijperi, Pichia stipitis, Pichia methanolica, Pichia angusta, Kluyveromyces lactis, Candida albicans or Yarrowia lipolytica.

[0183] In some embodiments of the present invention, the host cell is an algal cell, such as Chlamydomonas (e.g., Chlamydomonas reinhardtii) and Phormidium (Phormidium species ATCC 29409).

[0184] In some other embodiments, the host cell is a prokaryotic cell. Suitable prokaryotic cells include, but are not limited to, Gram-positive, Gram-negative, and Gram-variable bacterial cells. Any suitable bacterial organism can be used in the present invention, including, but not limited to, Agrobacterium, Alicyclobacillus, Anabaena, Anacystis, Acinetobacter, Acidothermus, Arthrobacter, Azobacter, Bacillus, Bifidobacterium, Brevibacterium, Butyrivibrio, Buchnera, Campestris, Campylobacter, Clostridium, Corynebacterium, Chromatium, Coprococcus, Escherichia, Enterococcus, Enterobacter, Erwinia, Fusobacterium, Faecalibacterium, Francisella, Flavobacterium, Geobacillus, Haemophilus, Helicobacter, Klebsiella, Lactobacillus, Lactococcus, Ilyobacter, Micrococcus, Microbacterium, Mesorhizobium, Methylobacterium, Methylobacterium, Mycobacterium, Neisseria, Pantoea, Pseudomonas, Prochlorococcus, Rhodobacter, Rhodopseudomonas, Rhodopseudomonas, Roseburia,Rhodospirillum, Rhodococcus, Scenedesmus, Streptomyces, Streptococcus, Synecoccus, Saccharomonospora, Staphylococcus, Serratia, Salmonella, Shigella, Thermoanaerobacterium, Tropheryma, Tularensis, Temecula, Thermosynechococcus, Thermococcus, Ureaplasma, Xanthomonas, Xylella, Yersinia, and Zymomonas. In some embodiments, the host cell is a species of: Agrobacterium, Acinetobacter, Azotobacter, Bacillus, Bifidobacterium, Buchnera, Geobacillus, Campylobacter, Clostridium, Corynebacterium, Escherichia, Enterococcus, Erwinia, Flavobacterium, Lactobacillus, Lactococcus, Pantoea, Pseudomonas, Staphylococcus, Salmonella, Streptococcus, Streptomyces, or Zymomonas. In some embodiments, the bacterial host strain is non-pathogenic to humans. In some embodiments, the bacterial host strain is an industrial strain. Industrial strains of many bacteria are known and applicable to the present invention. In some embodiments of the present invention, the bacterial host cell is a species of Agrobacterium (e.g., A. radiobacter, A. rhizogenes, and A. rubi). In some embodiments of the present invention, the bacterial host cell is a species of Arthrobacter (e.g., A. aurescens, A. citreus, A. globiformis, A. hydrocarboglutamicus, A. mysorens, A. nicotianae, A. paraffineus, A. protophonniae, A. roseoparqffinus, A. sulfureus, and A. ureafaciens). In some embodiments of the present invention, the bacterial host cell is a species of Bacillus (e.g., B. thuringensis, B. anthracis,Bacillus megaterium, Bacillus subtilis, Bacillus lentus, Bacillus circulans, Bacillus pumilus, Bacillus lautus, Bacillus coagulans, Bacillus brevis, Bacillus firmus, B. alkaophius, Bacillus licheniformis, Bacillus clausii, Bacillus stearothermophilus, Bacillus halodurans, and Bacillus amyloliquefaciens). In some embodiments, the host cell is an industrial Bacillus strain, including but not limited to Bacillus subtilis, Bacillus pumilus, Bacillus licheniformis, Bacillus megaterium, Bacillus clausii, Bacillus stearothermophilus, or Bacillus amyloliquefaciens. In some embodiments, the Bacillus host cell is Bacillus subtilis, Bacillus licheniformis, Bacillus megaterium, Bacillus stearothermophilus, and / or Bacillus amyloliquefaciens. In some embodiments, the bacterial host cell is a Clostridium species (e.g., Clostridium acetobutylicum, Clostridium tetani E88, Clostridium lituseburense, C. saccharobutylicum, Clostridium perfringens, and Clostridium beijerinckii). In some embodiments, the bacterial host cell is a Corynebacterium species (e.g., Corynebacterium glutamicum and Corynebacterium acetoacidophilum). In some embodiments, the bacterial host cell is an Escherichia species (e.g., Escherichia coli). In some embodiments, the host cell is Escherichia coli W3110. In some embodiments, the bacterial host cell is an Erwinia species (e.g., Erwinia uredovora, Erwinia carotovora, Erwinia ananas, Erwinia herbicola, E. punctata, and E. terreus). In some embodiments, the bacterial host cell is a Pantoea species (e.g., Pantoea citrea and Pantoea agglomerans). In some embodiments, the bacterial host cell is a Pseudomonas species (e.g., Pseudomonas putida, Pseudomonas aeruginosa,P. mevalonii and Pseudomonas species D-0l 10 (P. sp. D-0l 10). In some embodiments, the bacterial host cell is a Streptococcus species (e.g., S. equisimiles, S. pyogenes, and S. uberis). In some embodiments, the bacterial host cell is a Streptomyces species (e.g., S. ambofaciens, S. achromogenes, S. avermitilis, S. coelicolor, S. aureofaciens, S. aureus, S. fungicidicus, S. griseus, and S. lividans). In some embodiments, the bacterial host cell is a Zymomonas species (e.g., Z. mobilis and Z. lipolytica).

[0185] Many prokaryotic and eukaryotic strains useful in the present invention are readily available to the public from a number of culture collection centers, such as the American Type Culture Collection (ATCC), the Deutsche Sammlung von Mikroorganismen und Zellkulturen GmbH (DSM), the Centraalbureau Voor Schimmelcultures (CBS), and the Agricultural Research Service Patent Culture Collection, Northern Regional Research Center (NRRL).

[0186] In some embodiments, the host cell is genetically modified to have characteristics of improved protein secretion, protein stability, and / or other properties desirable for protein expression and / or secretion. The genetic modification can be achieved by genetic engineering techniques and / or classical microbial techniques (e.g., chemical or UV mutagenesis and subsequent selection). In fact, in some embodiments, a combination of recombinant modification and classical selection techniques is used to generate the host cell. Using recombinant techniques, nucleic acid molecules can be introduced, deleted, inhibited, or modified in a manner that results in increased production of one or more purine nucleoside phosphorylase variants in the host cell and / or in the culture medium. For example, knocking out Alp1 function produces protease-deficient cells, and knocking out pyr5 function produces cells with a pyrimidine-deficient phenotype. In one genetic engineering method, homologous recombination is used to induce targeted gene modification by specifically targeting a gene in vivo to inhibit the expression of the encoded protein. In an alternative method, siRNA, antisense, and / or ribozyme techniques can be used to inhibit gene expression. A variety of methods for reducing protein expression in cells are known in the art, including but not limited to deletion of all or part of the gene encoding the protein, and site-specific mutagenesis to disrupt the expression or activity of the gene product. (See, for example, Chaveroche et al., Nucl. Acids Res., 28:22e97

[2000] ; Cho et al., Molec. Plant Microbe Interact., 19:7-15

[2006] ; Maruyama and Kitamoto, Biotechnol Lett., 30:1811-1817

[2008] ; Takahashi et al., Mol. Gen. Genom., 272:344-352

[2004] ; and You et al., Arch. Microbiol., 191:615-622

[2009] , each of which is incorporated herein by reference). Random mutagenesis can also be used, followed by screening for the desired mutations (see, for example, Combier et al., FEMS Microbiol. Lett., 220:141-8

[2003] ; and Firon et al., Eukary. Cell 2:247-55

[2003] , both of which are incorporated by reference).

[0187] Introduction of the vector or DNA construct into the host cell can be accomplished using any suitable method known in the art, including but not limited to calcium phosphate transfection, DEAE-dextran-mediated transfection, PEG-mediated transformation, electroporation, or other commonly used techniques known in the art. In some embodiments, the Escherichia coli expression vector pCK100900i can be used (see, U.S. Patent No. 9,714,437, which is hereby incorporated by reference).

[0188] In some embodiments, the engineered host cells of the invention (i.e., "recombinant host cells") are cultured in a conventional nutrient medium that has been appropriately modified to activate a promoter, select for transformants, or amplify the purine nucleoside phosphorylase polynucleotide. The culture conditions, such as temperature, pH, etc., are those previously used with the host cells selected for expression and are well known to those skilled in the art. As described, many standard references and textbooks are available for the culturing and production of many cells, including those of bacterial, plant, animal (especially mammalian), and archebacterial origin.

[0189] In some embodiments, cells expressing the variant purine nucleoside phosphorylase polypeptide of the invention are grown under batch or continuous fermentation conditions. A typical "batch fermentation" is a closed system in which the composition of the medium is set at the start of the fermentation and is not subject to artificial change during the fermentation. A variation of the batch system is "fed-batch fermentation", which can also be used in the present invention. In this variation, substrates are added incrementally as the fermentation proceeds. Fed-batch systems are useful when catabolite repression may inhibit the metabolism of the cells and when a limited amount of substrate in the medium is desired. Batch fermentation and fed-batch fermentation are common and well known in the art. "Continuous fermentation" is an open system in which a defined fermentation medium is continuously added to a bioreactor and an equal amount of conditioned medium is removed simultaneously for processing. Continuous fermentation generally maintains the culture at a constant high density, where the cells are mainly in the logarithmic phase of growth. Continuous fermentation systems strive to maintain steady-state growth conditions. Methods for regulating nutrients and growth factors for continuous fermentation processes and techniques for maximizing the rate of product formation are well known in the field of industrial microbiology.

[0190] In some embodiments of the invention, cell-free transcription / translation systems can be used to produce one or more variant purine nucleoside phosphorylases. Several systems are commercially available and the methods are well known to those skilled in the art.

[0191] The present invention provides methods for preparing variant purine nucleoside phosphorylase polypeptides or bioactive fragments thereof. In some embodiments, the method comprises: providing a host cell transformed with a polynucleotide encoding an amino acid sequence having at least about 70% (or at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99%) sequence identity to SEQ ID NO: 2, 6, 126, 242, 684 and comprising at least one mutation provided herein; culturing the transformed host cell in a medium under conditions for expression of the encoded variant purine nucleoside phosphorylase polypeptide; and optionally recovering or isolating the expressed variant purine nucleoside phosphorylase polypeptide, and / or recovering or isolating the medium containing the expressed variant purine nucleoside phosphorylase polypeptide. In some embodiments, the method further provides for optionally lysing the transformed host cell after expression of the encoded purine nucleoside phosphorylase polypeptide, and optionally recovering and / or isolating the expressed variant purine nucleoside phosphorylase polypeptide from the cell lysate. The present invention also provides a method for preparing a variant purine nucleoside phosphorylase polypeptide, the method comprising culturing a host cell transformed with a variant purine nucleoside phosphorylase polynucleotide under conditions suitable for producing the variant purine nucleoside phosphorylase polypeptide, and recovering the variant purine nucleoside phosphorylase polypeptide. Generally, purine nucleoside phosphorylase polypeptides are recovered or isolated from the host cell culture medium, the host cell, or both, using protein recovery techniques well known in the art, including those described herein. In some embodiments, host cells are collected by centrifugation, disrupted by physical or chemical means, and the resulting crude extract is retained for further purification. Microbial cells used for protein expression can be disrupted by any convenient method, including but not limited to freeze-thaw cycles, sonication, mechanical disruption, and / or use of cell lysing agents, as well as many other suitable methods known to those skilled in the art.

[0192] Engineered purine nucleoside phosphorylase expressed in host cells can be recovered from the cells and / or the medium using any one or more of the techniques known in the art for protein purification, including, among others, lysozyme treatment, sonication, filtration, salting out, ultracentrifugation, and chromatography. A suitable solution for lysing and efficiently extracting proteins from bacteria such as E. coli is available under the trade name CelLytic B TM(Sigma-Aldrich) and are commercially available. Thus, in some embodiments, the resulting polypeptide is recovered / separated and optionally purified by any of a variety of methods known in the art. For example, in some embodiments, the polypeptide is separated from the nutrient medium by conventional procedures, including but not limited to centrifugation, filtration, extraction, spray drying, evaporation, chromatography (e.g., ion exchange, affinity, hydrophobic interaction, chromatofocusing, and size exclusion), or precipitation. In some embodiments, a protein refolding step is used, as needed, to complete the construction of the mature protein. Additionally, in some embodiments, high performance liquid chromatography (HPLC) is employed in the final purification step. For example, in some embodiments, methods known in the art can be used in the present invention (see, e.g., Parry et al., Biochem. J., 353:117

[2001] ; and Hong et al., Appl. Microbiol. Biotechnol., 73:1331

[2007] , both incorporated herein by reference). In fact, any suitable purification method known in the art can be used in the present invention.

[0193] Chromatographic techniques for separating purine nucleoside phosphorylase polypeptides include but are not limited to, reverse phase chromatography, high performance liquid chromatography, ion exchange chromatography, gel electrophoresis, and affinity chromatography. Conditions for purifying a particular enzyme depend in part on factors such as net charge, hydrophobicity, hydrophilicity, molecular weight, molecular shape, etc., and are known to those skilled in the art.

[0194] In some embodiments, affinity techniques can be used to isolate the improved purine nucleoside phosphorylase. For affinity chromatography purification, any antibody that specifically binds to the purine nucleoside phosphorylase polypeptide can be used. To generate antibodies, various host animals can be immunized by injection with purine nucleoside phosphorylase, including but not limited to rabbits, mice, rats, etc. The purine nucleoside phosphorylase polypeptide can be attached to a suitable carrier such as BSA by means of a side chain functional group or a linker attached to the side chain functional group. Depending on the host species, various adjuvants can be used to enhance the immune response, including but not limited to Freund's (complete and incomplete), mineral gels such as aluminum hydroxide, surface active substances such as lysolecithin, pluronic polyols, polyanions, peptides, oil emulsions, keyhole limpet hemocyanin, dinitrophenol, and potentially useful human adjuvants such as BCG (Bacillus Calmette-Guerin) and Corynebacterium parvum.

[0195] In some embodiments, a purine nucleoside phosphorylase variant is prepared and used in the form of cells expressing the enzyme, as a crude extract, or as an isolated or purified product. In some embodiments, the purine nucleoside phosphorylase variant is prepared as a lyophilized product, in powder form (e.g., acetone powder), or as an enzyme solution. In some embodiments, the purine nucleoside phosphorylase variant is in a substantially pure product form.

[0196] In some embodiments, the purine nucleoside phosphorylase polypeptide is attached to any suitable solid substrate. Solid substrates include, but are not limited to, a solid phase, a surface, and / or a membrane. Solid supports include, but are not limited to, organic polymers such as polystyrene, polyethylene, polypropylene, polyvinyl fluoride, polyethyleneoxy, and polyacrylamide, as well as copolymers and grafts thereof. The solid support can also be inorganic, such as glass, silica, controlled pore glass (CPG), reversed-phase silica, or a metal such as gold or platinum. The configuration of the substrate can be in the form of beads, spheres, particles, granules, gels, membranes, or surfaces. The surface can be flat, substantially flat, or non-flat. The solid support can be porous or non-porous and can have swelling or non-swelling characteristics. The solid support can be configured in the form of pores, depressions, or other containers, vessels, features, or locations. More than one support can be configured in an array at multiple locations that can be addressed by automated delivery of reagents or by detection methods and / or instruments.

[0197] In some embodiments, immunological methods are used to purify the purine nucleoside phosphorylase variant. In one method, antibodies against the variant purine nucleoside phosphorylase polypeptide (e.g., against a polypeptide comprising any of SEQ ID NO: 2, 6, 126, 242, and / or 684, and / or an immunogenic fragment thereof) generated using conventional methods are immobilized on beads, mixed with cell culture medium under conditions where the variant purine nucleoside phosphorylase binds, and precipitated. In a related method, immunochromatography can be used.

[0198] In some embodiments, the variant purine nucleoside phosphorylase is expressed as a fusion protein that includes a non-enzymatic portion. In some embodiments, the variant purine nucleoside phosphorylase sequence is fused to a purification facilitating domain. As used herein, the term "purification facilitating domain" refers to a domain that mediates the purification of a polypeptide to which it is fused. Suitable purification domains include, but are not limited to, metal chelating peptides, histidine-tryptophan modules that permit purification on immobilized metal, glutathione-binding sequences (e.g., GST), hemagglutinin (HA) tags (corresponding to epitopes derived from influenza hemagglutinin protein; see, e.g., Wilson et al., Cell 37:767

[1984] ), maltose binding protein sequences, FLAG epitopes used in the FLAGS extension / affinity purification system (e.g., a system available from Immunex Corp.), and the like. One expression vector contemplated for use in the compositions and methods described herein provides for the expression of a fusion protein that includes a polypeptide of the invention fused to a polyhistidine region separated by an enterokinase cleavage site. The histidine residues facilitate purification on IMIAC (immobilized metal ion affinity chromatography; see, e.g., Porath et al., Prot. Exp. Purif., 3:263-281

[1992] ), while the enterokinase cleavage site provides a means for separating the variant purine nucleoside phosphorylase polypeptide from the fusion protein. The pGEX vector (Promega) can also be used to express an exogenous polypeptide as a fusion protein with glutathione S-transferase (GST). Generally, such fusion proteins are soluble and can be readily purified from lysed cells by adsorption to ligand-agarose beads (e.g., glutathione-agarose in the case of a GST-fusion protein) followed by elution in the presence of free ligand.

[0199] Accordingly, in another aspect, the invention provides a method of generating an engineered enzyme polypeptide, wherein the method includes culturing a host cell capable of expressing a polynucleotide encoding the engineered enzyme polypeptide under conditions suitable for expression of the polypeptide. In some embodiments, the method further includes the step of isolating and / or purifying the enzyme polypeptide as described herein.

[0200] Suitable media and growth conditions for host cells are well known in the art. Any suitable method for introducing a polynucleotide for expressing an enzyme polypeptide into a cell is contemplated for use in the invention. Suitable techniques include, but are not limited to, electroporation, biolistic particle bombardment, liposome-mediated transfection, calcium chloride transfection, and protoplast fusion.

[0201] Various features and embodiments of the invention are illustrated in the following representative examples, which are intended to be illustrative and not limiting.

[0202] Experiments

[0203] The following examples, including experiments and the results obtained, are provided for illustrative purposes only and should not be construed as limiting the invention. In fact, many of the reagents and equipment described below have various suitable sources. It is not intended to limit the invention to any particular source of any reagent and equipment item. In some embodiments, the polypeptide sequence includes a histidine tag (i.e., 6 histidine residues) at the C-terminus. The sequence listing submitted herewith contains the polypeptide sequence without this histidine tag.

[0204] In the experimental disclosures below, the following abbreviations are used: M (moles per liter); mM (millimoles per liter), uM and μM (micromoles per liter); nM (nanomoles per liter); mol (mole); gm and g (gram); mg (milligram); ug and μg (microgram); L and l (liter); ml and mL (milliliter); cm (centimeter); mm (millimeter); um and μm (micrometer); sec. (second); min(s) (minute); h(s) and hr(s) (hour); U (unit); MW (molecular weight); rpm (revolutions per minute); psi and PSI (pounds per square inch); °C (degrees Celsius); RT and rt (room temperature); CV (coefficient of variation); CAM and cam (chloramphenicol); PMBS (polymyxin B sulfate); IPTG (isopropyl β-D-l-thiogalactopyranoside); LB (lysogeny broth); TB (terrific broth); SFP (shake flask powder); CDS (coding sequence); DNA (deoxyribonucleic acid); RNA (ribonucleic acid); nt (nucleotide; polynucleotide); aa (amino acid; polypeptide); Escherichia coli W3110 (a commonly used laboratory strain of E. coli available from the Coli Genetic Stock Center [CGSC], New Haven, CT); HTP (high throughput); HPLC (high performance liquid chromatography); HPLC-UV (HPLC-ultraviolet visible detector); 1H NMR (proton nuclear magnetic resonance spectroscopy); FIOPC (fold improvement over positive control); Sigma and Sigma-Aldrich (Sigma-Aldrich, St. Louis, MO); Difco (Difco Laboratories, BD Diagnostic Systems, Detroit, MI); Microfluidics (Microfluidics, Westwood, MA); Life Technologies (Life Technologies, part of Fisher Scientific, Waltham, MA); Amresco (Amresco, LLC, Solon, OH); Carbosynth (Carbosynth, Ltd., Berkshire, UK); Varian (Varian Medical Systems, Palo Alto, CA); Agilent (Agilent Technologies, Inc., Santa Clara, CA); Infors (Infors USA Inc.,Annapolis Junction, MD); and Thermotron (Thermotron, Inc., Holland, MI).

[0205] Example 1

[0206] Preparation of wet cell pellet containing HTP PNP

[0207] The parental gene of the PNP (SEQ ID NO:2) enzyme used to generate the variants of the present invention was obtained from the Escherichia coli genome and cloned into the pCK110900 vector. Escherichia coli W3110 cells were transformed with the corresponding plasmid containing the PNP-encoding gene and plated on LB agar plates containing 1% glucose and 30 μg / ml chloramphenicol (CAM), and grown overnight at 37 °C. Single colonies were picked and inoculated into 180 μl of LB containing 1% glucose and 30 μg / mL chloramphenicol and placed in the wells of a 96-well shallow-well microtiter plate. The plate was sealed with an 2 O-permeable seal and the cultures were grown overnight at 30 °C, 200 rpm and 85% humidity. Then, 10 μl of each cell culture was transferred to the wells of a 96-well deep-well plate containing 390 μl of TB and 30 μg / ml CAM. The deep-well plate was sealed with an 2 O-permeable seal and incubated at 30 °C, 250 rpm and 85% humidity until an OD 600 of 0.6 - 0.8 was reached. Then the cell cultures were induced by adding isopropylthiogalactoside (IPTG) to a final concentration of 1 mM and incubated with shaking at 30 °C overnight at 250 rpm. Then the cells were pelleted by centrifugation at 4000 rpm for 10 min. The supernatant was discarded and the pellet was frozen at -80 °C before lysis.

[0208] Example 2

[0209] Preparation of cell lysate containing HTP PNP

[0210] The frozen pellet prepared as described in Example 1 was lysed with 200 μl of lysis buffer containing 100 mM triethanolamine buffer pH 7.5, 1 mg / mL lysozyme, 0.5 mg / mL PMBS and 5 mM MnCl 2 . The lysis mixture was shaken at room temperature for 2 hours. Then the plate was centrifuged at 4000 rpm and 4 °C for 15 min. Then the supernatant was used as the clarified lysate for biocatalytic reactions to determine the activity level.

[0211] Example 3

[0212] Preparation of freeze-dried lysate from shake flask (SF) cultures

[0213] A single colony containing the desired gene, picked from an LB agar plate with 1% glucose and 30 μg / ml CAM and incubated overnight at 37°C, was transferred to 6 ml of LB with 1% glucose and 30 μg / ml CAM. The culture was grown at 30°C, 250 rpm for 18 h and subcultured at approximately 1:50 into 250 ml of TB containing 30 μg / ml CAM to reach a final OD of approximately 0.05. 600 The subculture was grown at 30°C, 250 rpm for approximately 195 minutes to reach an OD between 0.6 - 0.8. 600 and induced with 1 mM IPTG. The subculture was then grown at 30°C, 250 rpm for 20 h. The subculture was centrifuged at 4000 rpm for 20 min. The supernatant was discarded and the pellet was resuspended in 35 ml of 25 mM triethanolamine buffer pH 7.5 with 5 mM MnCl 2 . Cells were lysed at 18,000 psi using a Processor System (Microfluidics). The lysate was centrifuged (10,000 rpm x 60 min), and then the supernatant was frozen and lyophilized to yield shake flask (SF) enzyme powder.

[0214] Example 4

[0215] Improved purine nucleoside phosphorylase variants for the production of Compound 1

[0216] For these experiments, SEQ ID NO:2 was selected as the parental enzyme. A library of engineered genes was generated using well-established techniques such as saturation mutagenesis, and recombination of previously identified beneficial mutations. Polypeptides encoded by each gene were produced in HTP as described in Example 1, and clarified lysates were produced as described in Example 2.

[0217] For each enzyme, the clarified cell lysate was diluted 8-fold in 50 mM TEoA, 5 mM MnCl 2 , pH 7.5. Each 100 μL reaction was carried out in a 96-well shallow-well microtiter plate with 50% (v / v) diluted lysate, 30 mM Compound 4, 36 mM Compound 2, 5 g / L PPM (SEQ ID NO:1004), 100 mM TEoA buffer and 5.0 mM MnCl 2 at pH 7.5. The plate was heat-sealed and incubated at 45°C and on an Infors Stir overnight at 500 RPM in the oscillator. Remove the plate and quench by adding 1 volume of 1:1 DMSO:1M KOH, mix well until the compound dissolves, then dilute 10-fold into 25:75 v:v acetonitrile:0.1M TEoA, pH 7, and then analyze.

[0218] The activity relative to SEQ ID NO:2 was calculated as the percentage of conversion of the product formed by the variant enzyme compared to the percentage of conversion produced by SEQ ID NO:2. The percentage of conversion was quantified by dividing the area of the product peak by the sum of the areas of the substrate peak and the product peak determined by HPLC analysis.

[0219]

[0220]

[0221]

[0222] Example 5

[0223] Improved purine nucleoside phosphorylase variants for the production of compound 10

[0224]

[0225] An alternative base-exchange high-throughput screening method was developed to improve the robustness of the assay. By incorporating 500 mM potassium bromide into the screening conditions, improvements in the potassium bromide tolerance of PNP variants were also screened. For these experiments, SEQ ID NO:6 was selected as the parental enzyme. A library of engineered genes was generated using well-established techniques such as saturation mutagenesis and recombination of previously identified beneficial mutations. Each gene-encoded polypeptide was produced in HTP as described in Example 1, and a clarified lysate was produced using 400 μL instead of 200 μL of lysis buffer as described in Example 2.

[0226] Each 100 μL reaction was carried out in a 96-well deep-well microtiter plate (2 mL volume) with 5 μL of lysate, 13 mM of compound (8), 15 mM of compound (9), 500 mM potassium bromide, and 5 mM ammonium phosphate in 100 mM TEoA buffer at pH 7.5. The plate was heat-sealed and incubated at 40 °C and in Infors Stir overnight at 600 RPM in an oscillator. Remove the plate and quench by adding 1 volume of acetonitrile, mix well until the compound dissolves, then dilute 5-fold by adding 800 μL of 100 mM TEoA buffer at pH 7.5. Transfer 40 μL of the diluted quenched sample to a 96-well Millipore filter plate (0.45 μm pore size) pre-filled with 160 μL of 100 mM TEoA buffer at pH 7.5, mix and centrifuge at 4000 rpm for 5 min at 4 °C, then analyze the eluate by HPLC.

[0227] The activity relative to SEQ ID NO:6 was calculated as the percentage of conversion of compound 10 formed by the variant enzyme compared to the percentage of conversion produced by SEQ ID NO:6. The percentage of conversion was quantified by dividing the area of the product peak by the sum of the areas of the substrate peak and the product peak determined by HPLC analysis as described in Table 9.3.

[0228]

[0229]

[0230] Example 6

[0231] Improved purine nucleoside phosphorylase variants for the production of compound 10

[0232]

[0233] An alternative base-exchange high-throughput screening method was developed to improve the robustness of the assay. By incorporating 500 mM potassium bromide in the screening conditions, the improvement of potassium bromide tolerance of PNP variants was also screened. For these experiments, SEQ ID NO:126 was selected as the parental enzyme. A library of engineered genes was generated using well-established techniques such as saturation mutagenesis and recombination of previously identified beneficial mutations. Each gene-encoded polypeptide was produced in HTP as described in Example 1, and a clarified lysate was produced using 400 μL instead of 200 μL of lysis buffer as described in Example 2.

[0234] Each 100 μL reaction was carried out in a 96-well deep-well microtiter plate (2 mL volume) with 5 μL of lysate, 13 mM compound (8), 15 mM compound (9), 500 mM potassium bromide and 5 mM ammonium phosphate in 100 mM TEoA buffer at pH 7.5. The plate was heat-sealed and incubated at 40 °C and in Infors Stir overnight at 600 RPM in an oscillator. Remove the plate and quench by adding 1 volume of acetonitrile, mix well until the compound dissolves, then dilute 5-fold by adding 800 μL of 100 mM TEoA buffer at pH 7.5. Transfer 40 μL of the diluted quenched sample to a 96-well Millipore filter plate (0.45 μm pore size) pre-filled with 160 μL of 100 mM TEoA buffer at pH 7.5, mix and centrifuge at 4000 rpm for 5 min at 4 °C, then analyze the eluate by HPLC.

[0235] The activity relative to SEQ ID NO:126 was calculated as the percentage of conversion of compound 10 formed by the variant enzyme compared to the percentage of conversion produced by SEQ ID NO:126. The percentage of conversion was quantified by dividing the area of the product peak by the sum of the areas of the substrate peak and the product peak determined by HPLC analysis as described in Table 9.3.

[0236]

[0237]

[0238]

[0239]

[0240]

[0241]

[0242] Example 7

[0243] Improved purine nucleoside phosphorylase variants for the production of compound 1

[0244] For these experiments, SEQ ID NO:242 was chosen as the parental enzyme. A library of engineered genes was generated using well-established techniques such as saturation mutagenesis and recombination of previously identified beneficial mutations. Each gene-encoded polypeptide was produced in HTP as described in Example 1 and a clarified lysate was produced as described in Example 2 using 400 μL instead of 200 μL of lysis buffer.

[0245] For each enzyme, the clarified cell lysate was in 100 mM TEoA, 5 mM MnCl 2, Dilute 128-fold in pH 7.5. Each 100 μL reaction was carried out in a 96-well shallow-well microtiter plate with 20 μL of diluted lysate, 98 mM compound 4, 196 mM compound 2, 10 g / L PPM fermentation powder (SEQ ID NO: 1006), 0.25 g / L sucrose phosphorylase SUP001 (EC 2.4.1.7, Alloscardovia omnicolens SP154, GenBank accession number WP_021617468.1), 196 mM sucrose, 100 mM potassium sulfate, 100 mM TEoA buffer, and 5.0 mM MnCl 2 Performed at pH 7.5. The plate was heat-sealed and incubated at 40 °C and stirred overnight at 800 RPM in an Infors oscillator. The plate was removed and quenched by adding 200 μL of 1:1 DMSO:1 M KOH, and mixed well until the compound was dissolved. Transfer 10 μL of the diluted quenched sample to a 96-well Millipore filter plate (0.45 μm pore size) pre-filled with 190 μL of a 75:25 mixture of 100 mM TEoA buffer:acetonitrile at pH 7.5, mix and centrifuge at 4000 rpm for 5 min at 4 °C, and then analyze the eluate by HPLC.

[0246] The activity relative to SEQ ID NO: 242 was calculated as the percentage of conversion of compound 1 formed by the variant enzyme compared to the percentage of conversion produced by SEQ ID NO: 242. The percentage of conversion was quantified by dividing the area of the product peak by the sum of the areas of the substrate peak and the product peak determined by HPLC analysis as described in Table 9.2.

[0247]

[0248]

[0249]

[0250] Example 8

[0251] Improved purine nucleoside phosphorylase variants for the production of compound 1

[0252] For these experiments, SEQ ID NO: 684 was selected as the parental enzyme. A library of engineered genes was generated using well-established techniques such as saturation mutagenesis and recombination of previously identified beneficial mutations. Each gene-encoded polypeptide was produced in HTP as described in Example 1, and a clarified lysate was produced with 400 μL instead of 200 μL of lysis buffer as described in Example 2.

[0253] For each enzyme, the clarified cell lysate was diluted 32-fold in 100 mM TEoA, 5 mM MnCl 2 , pH 7.5. Each 100 μL reaction was in a 96-well shallow-well microtiter plate with 20 μL of the diluted lysate, 98 mM compound 4, 196 mM compound 2, 10 g / L PPM46 fermentation powder (SEQ ID NO: 514), 0.25 g / L sucrose phosphorylase SUP001, 196 mM sucrose, 100 mM potassium sulfate, 100 mM TEoA buffer, and 5.0 mM MnCl 2 at pH 7.5. The plate was heat-sealed and incubated at 40 °C and stirred overnight at 800 RPM in an Infors shaker. The plate was removed and quenched by adding 300 μL of 1:1 DMSO:1 M KOH and mixed well until the compounds were dissolved. 10 μL of the diluted quenched sample was transferred to a 96-well Millipore filter plate (0.45 μm pore size) pre-filled with 190 μL of a 75:25 mixture of 100 mM TEoA buffer:acetonitrile at pH 7.5, mixed, and centrifuged at 4000 rpm for 5 min at 4 °C, and then the eluate was analyzed by HPLC.

[0254] The activity relative to SEQ ID NO: 684 was calculated as the percentage of compound 1 conversion formed by the variant enzyme compared to the percentage of conversion produced by SEQ ID NO: 684. The percentage of conversion was quantified by dividing the area of the product peak by the sum of the areas of the substrate peak and the product peak determined by HPLC analysis as described in Table 9.2.

[0255]

[0256]

[0257] Example 9

[0258] Analytical Method

[0259] This example provides a method for collecting the data provided in the above examples. The data obtained as described in Example 4 was collected using the analytical method in Table 9.1. The data obtained as described in Example 6 was collected using the analytical method in Table 9.2. The data obtained as described in Example 5 was collected using the analytical method in Table 9.3. The method provided in this example can be used to analyze the variants produced using the present invention. However, this is not intended to limit the present invention to the methods described herein, as other suitable methods are known to those skilled in the art.

[0260]

[0261]

[0262]

[0263]

[0264] For all purposes, all publications, patents, patent applications, and other documents cited in this application are hereby incorporated by reference in their entirety to the same extent as if each individual publication, patent, patent application, or other document were specifically and individually indicated to be incorporated by reference for all purposes.

[0265] Although various specific embodiments have been shown and described, it will be understood that various changes may be made without departing from the spirit and scope of the invention.

Claims

1. An engineered purine nucleoside phosphorylase, wherein the engineered purine nucleoside phosphorylase is composed of a polypeptide sequence having 1 to 4 substitutions selected from the following, which is different from SEQ ID NO: 6 consisting of: A2S / T / P, R38E, N42D, K54D, D80E, K84E, S91T, V95I, L101V, V105L, M108I, K115R, F155H, S162T, G175V, L177V, A184S, Y187H, T199A, S204A, Q212A, and A215S, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

6.

2. An engineered purine nucleoside phosphorylase, wherein the engineered purine nucleoside phosphorylase is composed of the polypeptide sequence of SEQ ID NO: 126, or is composed of a polypeptide sequence having 1 to 7 substitutions selected from the following, which is different from SEQ ID NO: 126 consisting of: A2S / T, Y28H, E39L, N42H / S, G45A / C, Y53L, K54N, K57A, K75S, D80E, K84E, K85A / T, V95I, H98Y / D, L101I / V, I119M / V / T, D123T / G / M / S, H124R, I129M / S / G, A130P / T, D131R / M, F132L / A / C / S / N, R136I / L / V / C / E / D / S / A / G, N137W / E / Q, V139D / T / S / A / G, D140G, K143E / C / V / R / G / Y, A144Y / T / L / F / R / H, D148R / M / I / G / F, R150E / Y / M, V151F / H / L / N / Q, G152A / N / S, N153I / Y / T / P / R / C / S / L / G, F155H, F160L, V170R, K173E / W / H / S / F / G / Q / V / C / M, G175V / D, L177V / A / T / G, A191G / Y / W / F / P / T / V, K196V, T199A, V203I, H210R / G, Q212A / W, A215S, A216W / L / Y / F / C / R, T220F, N223H, and D237N, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

126.

3. An engineered purine nucleoside phosphorylase, wherein the engineered purine nucleoside phosphorylase is composed of the polypeptide sequence of SEQ ID NO: 242, or is composed of a polypeptide sequence having 1 to 2 substitutions selected from the following, which is different from SEQ ID NO: 242 consisting of: N7Y, D12A, K27F / S, Y28A / G / L / T, E31L / T / W, T32G / Q / V, E35R, A37L, R38L / Y, T52V / Y, K57L / Y, K85A, A94T, P97A / S, H98A / C / D / N, K100A / Y, R102L, D133A / W / Y, R136K, K143R, A144H / R / T, L145A / F / H / Q / S, D148I / S, A149P, R150H / V, S162M, D169H / R / S, V170C, E172R, K173R, A195S, K196A, T199A, I207L, R208F / H / K / S / T, T209F / G / H / L / S / W, H210F, E211N / Q / S / T / Y, A212G / R / V, T213S, T214A / H / V, A215G / H / P / S, E217A / D / M / Q, Q219A, T220A / G / R / S, T221S, N223A / L / V, D224A / G / K / N, K227T, L235M and K238P, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

242.

4. An engineered purine nucleoside phosphorylase, which is composed of the polypeptide sequence of SEQ ID NO: 684, or a polypeptide sequence that differs from SEQ ID NO: 684 by 1 - 3 substitutions selected from the following consisting of: M10P, L18M, P20A / G, A26S, I29L / T, E39Q, V60T, G62A, H63S, T74S / V, V89I / T, A94S / T, P97D, H98D, M108A / Q / S / V, F126L, D133G, V135L, L145K / Q, Y161H, S162F, G165A / C / H / L / P / R / S / T, M167L, F168L / T, V177M, T199S, R208F / K / L / V, H210M, A212M / S / V, D224G and K227E, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

684.

5. The engineered purine nucleoside phosphorylase according to any one of claims 1 - 4, wherein the engineered purine nucleoside phosphorylase is composed of the polypeptide sequence listed in any one of the even - numbered sequences of SEQ ID NOs: 124 - 200 and 242 - 1002.

6. The engineered purine nucleoside phosphorylase according to any one of claims 1-4, wherein compared with the wild-type Escherichia coli ( E.coli ) purine nucleoside phosphorylase, the engineered purine nucleoside phosphorylase comprises at least one improved property, wherein the at least one improved property comprises improved activity towards a substrate, improved production of compound 1 and / or improved production of compound 10 .

7. The engineered purine nucleoside phosphorylase according to claim 5, wherein compared with the wild-type Escherichia coli ( E.coli ), the engineered purine nucleoside phosphorylase comprises at least one improved property, and the at least one improved property comprises improved activity towards substrates, improved production of compound 1 and / or improved production of compound 10 .

8. The engineered purine nucleoside phosphorylase according to claim 6, wherein the substrate comprises compound 3 .

9. The engineered purine nucleoside phosphorylase according to claim 7, wherein the substrate comprises compound 3 .

10. The engineered purine nucleoside phosphorylase according to claim 6, wherein the substrate comprises compound 4 .

11. The engineered purine nucleoside phosphorylase according to claim 7, wherein the substrate comprises compound 4 .

12. The engineered purine nucleoside phosphorylase according to claim 6, wherein the substrate comprises compound 8 .

13. The engineered purine nucleoside phosphorylase according to claim 7, wherein the substrate comprises compound 8 .

14. The engineered purine nucleoside phosphorylase according to any one of claims 1 - 4 and 7 - 13, wherein the engineered purine nucleoside phosphorylase is purified.

15. The engineered purine nucleoside phosphorylase according to claim 5, wherein the engineered purine nucleoside phosphorylase is purified.

16. The engineered purine nucleoside phosphorylase according to claim 6, wherein the engineered purine nucleoside phosphorylase is purified.

17. A composition comprising the engineered purine nucleoside phosphorylase according to any one of claims 1-16.

18. A polynucleotide encoding at least one engineered purine nucleoside phosphorylase according to any one of claims 1-16.

19. The polynucleotide according to claim 18, wherein the polynucleotide is operably linked to a control sequence.

20. The polynucleotide according to any one of claims 18-19, wherein the polynucleotide is codon-optimized.

21. The polynucleotide according to any one of claims 18-19, wherein the polynucleotide comprises a polynucleotide sequence listed in any one of the odd-numbered sequences of SEQ ID NO: 123-199 and 241-1001.

22. The polynucleotide according to claim 20, wherein the polynucleotide comprises a polynucleotide sequence listed in any one of the odd-numbered sequences of SEQ ID NO: 123-199 and 241-1001.

23. An expression vector comprising at least one polynucleotide according to any one of claims 18-22.

24. A host cell comprising at least one expression vector according to claim 23, wherein the cell is not a plant cell.

25. A host cell comprising at least one polynucleotide according to any one of claims 18-22, wherein the cell is not a plant cell.

26. A method for producing an engineered purine nucleoside phosphorylase in a host cell, the method comprising culturing the host cell according to claim 24 and / or 25 under suitable conditions to produce at least one engineered purine nucleoside phosphorylase.

27. The method according to claim 26, further comprising recovering at least one engineered purine nucleoside phosphorylase from the culture and / or the host cell.

28. The method according to claim 26 or 27, further comprising the step of purifying the at least one engineered purine nucleoside phosphorylase.

Citation Information

Patent Citations

  • Methods for in vitro recombination

    US5605793A

  • Methods for generating polynucleotides having desired characteristics by iterative selection and recombination

    US5811238A

  • DNA mutagenesis by random fragmentation and reassembly

    US5830721A

  • End-complementary polymerase reaction

    US5834252A

  • Methods and compositions for cellular and metabolic engineering

    US5837458A