A new class of miniature viral polymerases for untemplated synthesis of natural or modified nucleic acids
PolARVs offer a new mechanism for untemplated nucleic acid synthesis, addressing the replication challenge of linear dsDNA genomes in archaeal viruses, with applications in nucleic acid synthesis and potential for engineered improvements.
Patent Information
- Application Number
- PCT/EP2025/070409
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-16
- Filing Date
- 2025-07-16
- Publication Date
- 2026-01-22
AI Technical Summary
The replication of linear double-stranded DNA genomes in archaeal viruses is not well understood, as the existing archaeal replication machinery cannot faithfully replicate these genomes without encountering the end-replication problem, and there is a lack of characterized DNA polymerases in these viruses.
Identification and characterization of a new family of archaeal virus polymerases (PolARV) with conserved motifs and a unique structure, including a C-terminal loop, which exhibit de novo nucleic acid synthesis activity, enabling untemplated nucleic acid synthesis.
PolARVs provide a novel mechanism for nucleic acid synthesis, particularly suitable for DNA or RNA printing applications, with potential for improved properties through targeted engineering of the C-terminal loop to enhance performance.
Smart Images

Figure IMGF000017_0001 
Figure IMGF000081_0001 
Figure IMGF000083_0001
Abstract
Description
[0001]A NEW CLASS OF MINIATURE VIRAL POLYMERASES FOR UNTEMPLATED SYNTHESIS OF NATURAL OR MODIFIED NUCLEIC ACIDS FIELD OF THE INVENTION 5 The invention relates to a new family of polymerases from archaeal viruses (PolARV), which constitute the shortest catalytic core described so far, and their uses for nucleic acid synthesis, in particular untemplated synthesis of natural or modified nucleic acids. BACKGROUND OF THE INVENTION The mechanisms of genome replication are far more diverse in viruses compared to cellular 10 organisms and it has been suggested that viruses represent evolution’s workshop for exploring, refining, and selecting genome replication and expression circuits. Furthermore, viral enzymes often display unique specificities or superior performance compared to the cellular counterparts and hence hold a great potential for development of new tools in biomedicine and biotechnology. Viruses infecting Archaea represent one of the least understood parts of 15 the virosphere. Most of the archaeal viruses characterized thus far have been isolated from extreme environments, including those with temperatures of 80-100ºC (hyperthermophiles), low pH (acidophiles) or nearly saturating salt concentrations (hyperhalophiles). These viruses display a range of unique virion morphologies, resembling lemons, bottles or droplets, and carry DNA genomes that encode proteins with very few (<25%) identifiable homologs in 20 sequence databases. Thus, many aspects of the life cycles of archaeal viruses are poorly understood, with even such key stages as viral genome replication remaining enigmatic. In particular, many unrelated groups of archaeal viruses infecting hyperthermophilic, acidophilic and hyperhalophilic hosts, carry linear double-stranded (ds)DNA genomes with inverted terminal repeats, but unlike bacterial or eukaryotic viruses with similar genome organizations 25 (e.g., phage phi29 or adenovirus), do not encode recognizable DNA polymerases. Given that the archaeal replication machinery cannot faithfully replicate linear dsDNA genomes without succumbing to the so-called end-replication problem, replication of these linear virus genomes presents an unresolved mystery in virology and, more generally, the field of molecular biology. Understanding of this unusual replication mechanism might uncover new enzymes 30 with industrial potential. SUMMARY OF THE INVENTION The inventors have identified a new protein family (PolARV) conserved across multiple groups of archaeal viruses with linear dsDNA genomes and characterized a set of enzymatic activities displayed by representative members of this new protein family (SBV1 polymerase 5 [SBV1pol or polSBV] and SOV polymerase [polSoV]). The PolARV family has a great industrial potential since it includes small polymerases having de novo (i.e., untemplated) nucleic acid synthesis activity from natural or modified nucleotides suitable for DNA or RNA printing applications. In addition, these small polymerases comprise a small C-terminal loop predicted to control the access of incoming nucleotides to the active site that can be easily 10 targeted to engineer new polymerases with improved properties. The invention relates to an isolated polymerase of archaeal viruses (PolARV) or a variant thereof with nucleotidyl transferase activity. One aspect of the invention relates to an isolated polymerase of archaeal viruses (PolARV) comprising a catalytic core domain having a structure comprising three alpha-helices (alpha1 15 to alpha3) packed against a beta-sheet consisting of three beta-strands (beta1 to beta3), said catalytic core domain having an amino acid sequence comprising the following conserved motifs: (a) DX1D / N in positions 9 to 11, wherein X1 in position 10 is L, F, M, I, Y, V, K, W, H, C, or R; 20 (b) H in position 48; (c) D, S, E, S, C or N in position 71; (d) D, C, N or H in position 72; (e) R, H, K or L in position 75; (f) R, K, A, L, G, Q, S, Y, C, H or T in position 82; and 25 (g) F, Y, W, L, H, M, or N in position 96; the indicated positions being determined by alignment with Sulfolobales Beppu Virus 1 DNA polymerase (SBV1pol) of SEQ ID NO: 1. In some embodiments: - the motif in (a) is DX1D; - the motif in (d) is D; - the motif in (e) is R, H or K; preferably R; and - the motif in (g) is F, W, Y or L; optionally wherein X1 is L, F, I, Y, W or H; the motif in (c) is D or E; and / or the motif in (f) is 5 R. In some embodiments, the motifs in (a), (b), (d), (e) and, optionally the motif in (f), are in the polymerase catalytic site. In some embodiments, the motif in (g) is in the C-terminal region. In some embodiments, the motif in (a) is in beta1, the motif in (b) is in beta3, the motif in (d) 10 is in alpha 2 and the motif in (e) is in alpha3. In some embodiments, the core structure of the polymerase catalytic core domain consists of a sequence of 90 to 100 amino acids. In some embodiments, the catalytic core domain further comprises a C-terminal loop; preferably forming the C-terminus of the polymerase. In some particular embodiments, the 15 C-terminal loop comprises the motif in (g). In some particular embodiments, the C-terminal loop is a flexible loop; preferably comprising an alpha-helix and optionally further comprising a beta-strand; more preferably wherein the C-terminal loop encloses the polymerase catalytic site. In some particular embodiments, the C-terminal loop consists of a sequence of up to 50 amino acids; preferably 20 to 30 amino acids; more preferably around 25 amino acids. 20 In some particular embodiments, the polymerase consists of a catalytic core domain; preferably a catalytic core domain comprising a C-terminal flexible loop. In some embodiments, the polymerase consists of a sequence from about 100 to about 300 amino acids; preferably from about 100 to about 150 amino acids. In some embodiments, the catalytic core domain is preceded by an additional domain which 25 may have a structural or functional role, as in the case of polymerase of SIFV (Sulfolobus_islandicus_filamentous_virus__SIFV0030; SEQ ID NO: 8). In some embodiments, the polymerase is from a group of archaeal viruses chosen from: Acidianus filamentous virus 1 (AFV1), Acidianus filamentous virus 2 (AFV2), Sulfolobus islandicus filamentous virus (SIFV), Hyperthermophilic Archaeal Virus 1 (HAV1), Sulfolobales Beppu virus 1 (SBV1) and Haloarcula hispanica virus SH1 (SH1) and / or which is from a group of PolARV chosen from: AFV1-like, AFV2-like, SIFV-like, HAV1-like, SBV1-like, SH1-like and MetaB. In some embodiments, the polymerase has an amino acid 5 sequence which clusters with the sequences of one group of PolARVs chosen from AFV1- like, AFV2-like, SIFV-like, HAV1-like, SBV1-like, SH1-like and MetaB, based on pairwise sequence similarity. The invention also relates to a variant of the polymerase according to the present disclosure, which has nucleotidyl transferase activity, and said variant comprising at least one mutation, 10 in particular at least one substitution; preferably in the C-terminal loop and / or preferably comprising one or more substitutions at positions selected from the group consisting of S42, S43, N44, H46, I95, R94, F96, T97, K99, M100, N101, K102, E103, V105, F107, and optionally E13, D71 and / or R82; the indicated positions being determined by alignment with SEQ ID NO: 1. In some embodiments, the PolARV variant comprises at least one of the 15 following substitutions: E13Q; H46N; E13Q and H46N; F96Y; M100A; M100K; N101A and V105A. In some embodiments, the polymerase or derived variant according to the present disclosure comprises an amino acid sequence having at least 85 % identity with any one of SEQ ID NO: 1 to 339. In some preferred embodiments, the polymerase or derived variant according to the 20 present disclosure comprises an amino acid sequence having at least 85 % identity with any one of SEQ ID NO: 1 to 272, 274 to 283, 285 to 288, 290 to 297, 299 to 313 and 316 to 339; more preferably SEQ ID NO: 1 or 3; still more preferably SEQ ID NO: 3. In some preferred embodiments, the polymerase or derived variant according to the present disclosure comprises an amino acid sequence having at least 85% identity with the sequence 25 from positions 7 to 83 of SEQ ID NO: 1. In some embodiments, the polymerase or derived variant is a recombinant polymerase. In some embodiments, the polymerase or derived variant has terminal transferase activity on natural or modified nucleotides. Another aspect of the invention relates to an expression vector for the recombinant production of the polymerase, or derived variant according to the present disclosure in a host cell, comprising a nucleic acid encoding said polymerase or derived variant. A method for nucleic acid synthesis comprising incubating the polymerase, or derived variant 5 according to the present disclosure with at least one oligonucleotide primer and at least one nucleotide under conditions that allow incorporation of the nucleotide(s) on the primer. In some preferred embodiment, the method is for untemplated nucleic acid synthesis, wherein the method is performed in the absence of nucleic acid template. A kit for nucleic acid synthesis comprising the polymerase, or derived variant according to 10 the present disclosure; and optionally, at least one nucleotide, reaction buffer, and / or oligonucleotide primer; preferably further comprising at least one modified nucleotide, in particular at least one 3’-OH blocked modified nucleotide. In some preferred embodiment, the kit is for untemplated nucleic acid synthesis, wherein the kit does not comprise a nucleic acid template. 15 DETAILED DESCRIPTION OF THE INVENTION The present invention provides a new family of polymerases from archaeal viruses (PolARV) which are particularly useful for nucleic acid synthesis. The invention encompasses, isolated polymerases of archaeal viruses (PolARV); variants of PolARV with improved properties (i.e., engineered PolARV); and derived nucleic acids and vectors. The invention provides methods 20 of making and using said PolARVs and variants thereof. The invention encompasses the use of PolARVs and variants thereof for nucleic acid synthesis in various methods and applications. PolARV polymerases One aspect of the invention relates to an isolated polymerase of archaeal viruses, named 25 PolARV or PolARV polymerase. PolARV are nucleotidyl transferases. PolARV family is characterized by conserved motifs. Sulfolobales Beppu Virus 1 (SBV1) polymerase of SEQ ID NO: 1 and Sulfurisphaera ohwakuensis virus (SOV) polymerase of SEQ ID NO: 3 are representative members of the PolARV family. Representative members of PolARVs such as SBV1 polymerase and SOV polymerase are characterized by unique structural features associated with relevant enzymatic activities. PolARV sequences are also clustered in different groups based on pairwise sequence similarity (Figure 2 and Table 1). A polymerase refers to a protein capable of primer extension, which means able to incorporate nucleotides to the 3’ end (3’ hydroxyl end) of a nucleotide primer in vitro or in vivo, using a 5 DNA or RNA template as a guide to direct each incorporation event. A nucleotidyl transferase or terminal transferase is a template-independent polymerase that catalyzes the addition of nucleotides to the 3’ end of a nucleotide primer in vitro or in vivo, in the absence of nucleic acid template. The nucleotidyl transferase activity of a PolARV according to the invention may be tested by standard primer extension assays as disclosed in 10 the Examples. As used herein, a “motif” refers to one amino acid residue or several successive amino acid residues (amino acid sequence) that is conserved across polymerases of archaeal viruses with dsDNA genomes. The conserved motifs include residues that are important or essential for the structure or function (i.e. polymerase enzymatic activity) of PolARV. In particular, the 15 inventors have identified seven motifs ((a) to (g)) that are conserved across at least 339 PolARV amino acid sequences (Figure 1). In the following description, the residues are designated by the standard one letter amino acid code and the indicated positions are determined by alignment with SEQ ID NO: 1. One skilled in the art can easily determine the positions in another PolARV, by alignment with the 20 reference sequence using appropriate software available in the art such as BLAST, CLUSTALW, or MAFFT (with G-INS-1 option) and others. The terms “catalytic core domain”, “catalytic domain” or “catalytic core” are used interchangeably to designate the domain containing the catalytic site of PolARV. All the residues of the catalytic site of PolARVs are present in a single domain (catalytic domain). 25 The catalytic domain of PolARVs has a structure comprising three alpha-helices (alpha1 to alpha3) packed against a beta-sheet consisting of three beta-strands (beta1 to beta3) with a β1α1β2β3α2α3 topology where the beta-strands 1 and 2 are anti-parallel to the beta-strand 3. The three alpha-helices (alpha1 to alpha3) and the three beta-strands (beta1 to beta3) form the core structure of PolARV catalytic domain. The catalytic core advantageously further 30 comprises a C-terminal loop as disclosed herein. “a”, “an”, and “the” include plural referents, unless the context clearly indicates otherwise. As such, the term “a” (or “an”), “one or more” or “at least one” can be used interchangeably herein; unless specified otherwise, “or” means “and / or”. One aspect of the invention relates to an isolated PolARV polymerase comprising a catalytic 5 core domain, wherein the catalytic core domain has: (i) a structure comprising three alpha- helices (alpha1 to alpha3) packed against a beta-sheet consisting of three beta-strands (beta1 to beta3) and (ii) an amino acid sequence comprising the following conserved motifs: (a) DX1D / N in positions 9 to 11, wherein X1 in position 10 is L, F, M, I, Y, V, K, W, H, C, or R; 10 (b) H in position 48; (c) D, S, E, S, C or N in position 71; (d) D, C, N or H in position 72; (e) R, H, K or L in position 75; (f) R, K, A, L, G, Q, S, Y, C, H or T in position 82; and 15 (g) F, Y, W, L, H, M, or N in position 96; the indicated positions being determined by alignment with Sulfolobales Beppu Virus 1 (SBV1) polymerase of SEQ ID NO: 1. All the listed residues are listed from the most to the less frequent residues at each indicated position. Based on the above definition, PolARV may be characterized by the following sequence 20 pattern (I): <x(3,230)-D-[LFMIYVKWHCR]-[DN]-x(22,65)-H-x(17,41)-[DSECN]-[DCNH]-x(2)- [RHKL]-x(6)-[RKALGQSYCHT]-x(7,22)-[FYWLHMN]-x(3,185)> where x is any amino acid and numbers in parentheses indicate the range; variation within conserved amino acid positions is provided within square brackets, with the most conserved 25 amino acid residues listed first and the indicated residues are in positions 9, 10, 11 (motif (a)); 48 (motif (b)); 71 (motif (c)); 72 (motif (d)); 75 (motif (e)); 82 (motif (f)) and 96 (motif (g)), after alignment with SBV1pol amino acid sequence of SEQ ID NO: 1 used as reference. Therefore, in some embodiments, the motif in (a) is separated from the N-terminus by 3 to 230 amino acids; the motifs in (a) and (b) are separated by 22 to 65 amino acids; the motifs in 30 (b) and (c) are separated by 17 to 41 amino acids; the motifs in (d) and (e) are separated by 2 amino acids; the motifs in (e) and (f) are separated by 6 amino acids; the motifs in (f) and (g) are separated by 7 to 22 amino acids; and the motif in (g) is separated from the C-terminus by 3 to 185 amino acids. In some embodiments: 5 - the motif in (a) is DX1D; - the motif in (d) is D; - the motif in (e) is R, H or K; preferably R; and / or - the motif in (g) is F, W, Y or L. The PolARV may comprise 1, 2, 3 or all of said particular motifs in (a), (d), (e) and (g) above. 10 In some particular embodiments, the PolARV comprises at least said particular motifs in (a), (d), (e) above. In some more particular embodiments, the PolARV comprises at least the particular motifs in (a), (d), (e) and (g) above. Optionally, X1 is L, F, I, Y, W or H; the motif in (c) is D or E; and / or the motif in (f) is R. In some preferred embodiments, the PolARV comprises said particular motifs in (a), (d), (e) and (f) above; preferably in (a), (d), (e), (f) and 15 (g) above; more preferably the PolARV comprises said particular motifs in (a), (c), (d), (e) and (f) above, still more preferably said particular motifs in (a), (c) (d), (e), (f) and (g) above. In some embodiments, the PolARV has the following preferred sequence pattern (II): <x(3,230)-D-[LFMIYVKWHCR]-[D]-x(22,65)-H-x(17,41)-[D / E]-x(2,3)-[R]-x(6)- [RKALGQSYCHT]-x(7,22)-[FYWL]-x(3,185)> 20 in which the indicated residues are in positions 9, 10, 11 (motif (a)); 48 (motif (b)); 72 (motif (d)); 75 (motif (e)); 82 (motif (f)) and 96 (motif (g)) after alignment with SBV1pol amino acid sequence of SEQ ID NO: 1 used as reference. In some embodiments, when said particular motif in (f) is different from R or K, then the PolARV further comprises R in positions +1 to +5 or -1 to -5 relative to said motif in (f); preferably in positions +1, +2, -1 to -2 relative to 25 said motif in (f). In some preferred embodiments, said particular motif in (f) is R or K, preferably R. In some particular embodiments, the PolARV comprises particular motifs in (c) and (d) chosen from DD or ED. Some particular examples of PolARV comprise the following motifs: (a) DX1D in positions 9 to 11, wherein X1 in position 10 is L, F, I, Y, W or H; 30 (b) H in position 48; (c) D or E in position 71; (d) D in position 72; (e) R in position 75; (f) R position 82; and 5 (g) F, Y, or W in position 96. A preferred example of PolARV is composed of a catalytic domain comprising a N-terminal catalytic core having a β1α1β2β3α2α3 topology where the beta-strands 1 and 2 are anti- parallel to the beta-strand 3 and a flexible C-terminal loop; and said catalytic domain comprises the following motifs: 10 (a) DX1D in positions 9 to 11, wherein X1in position 10 is L, F, I, Y, W or H; (b) H in position 48; (c) D or E in position 71; (d) D in position 72; (e) R in position 75; 15 (f) R position 82; and (g) F, Y, or W in position 96; wherein the motifs in (a), (b), (d), (e), and optionally the motif in (f), are in the catalytic site of the catalytic core domain and the motif in (g) is in the C-terminal region, preferably wherein the catalytic core has a sequence of about 90 to 100 amino acids; and / or the PolARV consists 20 of a sequence from about 100 to about 150 amino acids; in particular the PolARV sequence has at least 85 % identity with SEQ ID NO: 1 or SEQ ID NO: 3. In some particular embodiments, the motifs in (a), (b), (d), (e), and optionally the motif in (f), are in the catalytic site of the catalytic core domain; preferably wherein the motif in (a) is DX1D; the motif in (d) is D; and the motif in (e) is R, H or K; preferably R. 25 In some embodiments, the motif in (g) is in the C-terminal region. PolARV catalytic core comprises three alpha-helices (alpha1 to alpha3) packed against a beta- sheet consisting of three beta-strands (beta1 to beta3), as illustrated in Figure 4-A to D Figure 5-A and 5-B; Figure 9-A to C. More particularly, the motif in (a) is in beta1, the motif in (b) is in beta3, the motif in (d) is in alpha 2 and the motif in (e) is in alpha3; preferably 30 wherein the motif in (a) is DX1D; the motif in (d) is D; and / or the motif in (e) is R, H or K; preferably R. This structure of a representative PolARV such as SBV1 polymerase is significantly different from that found in the replicative primases-polymerases (prim-pols) of the archaeo-eukaryotic primase (AEP) family. In all archaea-eukaryotic primase-polymerases, whose structures have been reported, the active site shows a bi-partite active site composed of 5 two domains illustrated in Figure 5-C: the N-terminal alpha-beta domain contains critical residues, such as steric gate residues that control the selectivity of the enzyme, and a C-terminal domain that hosts the catalytic-metals chelating residues. In PolARV, the canonical N-terminal domain of AEP is systematically absent. The 3D structure of the SBV1 and SOV polymerases determined experimentally by using X-ray crystallography shows that 10 the motifs (a) to (f) are clustered within a compact catalytic site. In the examples of SBV1 and SOV, the structure shows that the active site is completed by a C-terminal loop that hosts motif (g) and is strategically positioned against the active site. The C-terminal residues 95– 111 stack against the catalytic site in SBV1 polymerase, while this region in SOV polymerase is more disordered, leaving the site more accessible. AlphaFold further predicts a third 15 alternative conformation for this region in SOV polymerase, with the C-terminal residues forming a two-stranded β-hairpin (Figure 9-C). The four conserved catalytic motifs that are characteristic of the AEP superfamily are also present in representative PolARVs such as SBV1 and SOV polymerases. The motif I (DXD) and motif III (D) are on beta1 and alpha2, respectively. Motifs II and IV contain positively charged residues and are located on beta3 20 and alpha3, respectively (Figure 5-A-B and Figure 9-B). The three alpha-helices (alpha1 to alpha3) and the beta-sheet consisting of three beta-strands (beta1 to beta3) which form the core structure of PolARV catalytic domain have a sequence of about 90 to 100 amino acids. The catalytic domain of some PolARVs may have a larger amino acid sequence due to sequence insertion(s) between the different structural elements of 25 the catalytic domain. In particular, the insertions can occur between the motifs in (a) and (b); between the motifs in (b) and (c); and between the motifs in (f) and (g). In some embodiments, the catalytic core further comprises a C-terminal loop; preferably forming the C-terminus of the polymerase. In some particular embodiments, the C-terminal loop comprises the motif in (g); preferably, the motif in (g) is F, Y, W or L. In some particular 30 embodiments, the C-terminal loop is a flexible loop; particularly wherein the flexible loop is a deviated loop. In some more particular embodiments, the flexible loop comprises an alpha- helix and optionally further comprising a beta-strand; preferably alpha4 or alpha4 and beta4. In some more particular embodiments, the flexible loop comprises a two-stranded β-hairpin composed ofbeta4 and beta5. In some particular embodiments, the C-terminal loop consists 5 of a sequence of up to 50 amino acids; preferably 20 to 30 amino acids; more preferably around 25 amino acids. In some particular embodiments, the C-terminal loop corresponds to positions 92 to 116 or 95 to 111; said positions being determined by alignment with SEQ ID NO: 1. In some particular embodiments, the N-terminus of the loop is connected to the C- terminus of the polymerase catalytic core domain as disclosed herein. In some more particular 10 embodiments, the C-terminal loop encloses the polymerase catalytic site; particularly the loop is packed against the catalytic site and contains an aromatic (F, Y or W) or bulky hydrophobic (L) residue at position 96 (corresponding to motif (g)), blocking directly access to the catalytic active site residues). This C-terminal loop that encloses the catalytic site is thus likely to play a role in catalysis or regulation and to be important for the incorporation of the nucleotide or 15 DNA substrate (Figure 4C-4D). In some embodiments, the PolARV has an N-terminal, and / or C-terminal extension, respectively at the N-terminus and C-terminus of the catalytic core domain. In some particular embodiments, the catalytic core domain is preceded by an additional domain which may have a structural or functional role, as in the case of polymerase of SIFV (Sulfolobus_islandicus_filamentous_virus_SIFV0030; SEQ ID NO: 8). 20 In some other embodiments, the polymerase consists of a catalytic core domain composed of a catalytic core formed of three alpha-helices (alpha1 to alpha3) packed against a beta-sheet consisting of three beta-strands (beta1 to beta3) with a β1α1β2β3α2α3 topology where the beta-strands 1 and 2 are anti-parallel to the beta-strand 3. In some particular embodiments, the PolARV catalytic core domain further comprises a C-terminal flexible loop. preferably The 25 C-terminal flexible loop may comprise an alpha-helix (alpha4) and a beta-strand (beta4); this structure is found in representative members of PolARVs such as SBV1 polymerase. Alternatively, the C-terminal flexible loop may comprise a two-stranded β-hairpin composed of beta4 and beta5; this structure is found in representative members of PolARVs such as SOV polymerase. 30 In some embodiments, the polymerase consists of a sequence from about 100 to about 300 amino acids; preferably from about 100 to about 150 amino acids. In some embodiments, the polymerase forms a dimer. In some particular embodiments, residues 57-87 are involved in dimerization, the indicated positions being determined by alignment with SBV1 polymerase of SEQ ID NO: 1. The PolARVs according to the invention may be isolated from any archaeal virus with linear 5 dsDNA genome. In some embodiments, the PolARV polymerase is from a group of archaeal viruses chosen from: Acidianus filamentous virus 1 (AFV1), Acidianus filamentous virus 2 (AFV2), Sulfolobus islandicus filamentous virus (SIFV), Hyperthermophilic Archaeal Virus 1 (HAV1), Sulfolobales Beppu virus 1 (SBV1) and Haloarcula hispanica virus SH1 (SH1). In some embodiments, the PolARV polymerase is from a group of PolARV chosen from: AFV1- 10 like, AFV2-like, SIFV-like, HAV1-like, SBV1-like, SH1-like and MetaB groups. Representative SBV1-like PolARVs include SEQ ID NO: 1, 2, 3, 14, 15, 16 and 17; a representative AFV1-like PolARV is SEQ ID NO: 4; Representative AFV2-like PolARVs include SEQ ID NO: 5 and 11; Representative SH1-like PolARVs include SEQ ID NO: 6, 7, 12 and 13; Representative SIFV-like PolARVs include SEQ ID NO: 8, 9, and 18 to 23; a 15 representative HAV1-like PolARV is SEQ ID NO: 10. Other PolARVs of the AFV1-like, AFV2-like, SIFV-like, HAV1-like, SBV1-like, SH1-like and MetaB groups are listed in Table 1. In some particular embodiments, said PolARV is the polymerase of an archaeal virus chosen from : Sulfolobales_Beppu_virus_1__SBV1_gp38 (SBV1-like; SEQ ID NO: 1); 20 Metallosphaera_turreted_icosahedral_virus__c109 (SBV1-like; SEQ ID NO: 2); Sohwa_SBV1-like_ITR_37 (SBV1-like (SEQ ID NO: 3); Captovirus_AFV1__AFV1_ORF144 (AFV1-like; SEQ ID NO: 4); Acidianus_filamentous_virus_2__gp50 (AFV2-like; SEQ ID NO: 5); Haloarcula_hispanica_virus_SH1__ORF41 (SH1-like; SEQ ID NO: 6); 25 Haloarcula_hispanica_icosahedral_virus_2__gp30 (SH1-like; SEQ ID NO: 7); Sulfolobus_islandicus_filamentous_virus__SIFV0030 (SIFV-like; SEQ ID NO: 8; Acidianus_filamentous_virus_9__AFV9_gp28 (SIFV-like; SEQ ID NO: 9); Hyperthermophilic_Archaeal_Virus_1__HAV1_gp16 (HAV1-like; SEQ ID NO: 10; Sulfolobales_Beppu_filamentous_virus_3__HOU83_gp44 (AFV2-like; SEQ ID NO: 11); 30 Haloarcula_californiae_icosahedral_virus_1__BGV91_gp36 (SH1-like; SEQ ID NO: 12); Haloarcula_hispanica_virus_PH1__gp35 (SH1-like; SEQ ID NO: 13); Sulfolobus_polyhedral_virus_3__SPV3_ORF40 (SBV1-like; SEQ ID NO: 14; Metallosphaera_turreted_icosahedral_virus__c109 (SBV1-like; SEQ ID NO: 15; Sulfolobus_polyhedral_virus_3__SPV3_ORF13 (SBV1-like; SEQ ID NO: 16); Metallosphaera_turreted_icosahedral_virus_3__MTIV3_ORF14 (SBV1-like; SEQ ID NO: 5 17); Sulfolobus_islandicus_filamentous_virus_2 (SIFV-like; SEQ ID NO: 18); Acidianus_filamentous_virus_3__AFV3_gp30(SIFV-like; SEQ ID NO: 19); Acidianus_filamentous_virus_6__AFV6_gp31 (SIFV-like; SEQ ID NO: 20); Acidianus_filamentous_virus_8__AFV8_gp26 (SIFV-like; SEQ ID NO: 21); 10 Acidianus_filamentous_virus_7__AFV7_gp24 (SIFV-like; SEQ ID NO: 22) and Saccharolobus_shibatae_filamentous_virus_3 (SIFV-like; SEQ ID NO: 23); preferably the PolARV is chosen from SEQ ID NO: 1 to 10 or SEQ ID NO: 1, 2 and 4 to 11; in particular SEQ ID NO: 1 or 3. In some embodiments, the PolARV has a sequence which clusters with the sequences of one 15 group of PolARVs chosen from AFV1-like, AFV2-like, SIFV-like, HAV1-like, SBV1-like, SH1-like and MetaB, based on pairwise sequence similarity. The clustering of PolARV sequences based on pairwise sequence similarity is illustrated in Figure 2. The clustering of PolARV sequences based on pairwise sequence similarity may be obtained using appropriate algorithms and / or softwares for clustering amino acid sequences that are well-known in the 20 art. A particular example is CLANS (https: / / github.com / inbalpaz / CLANS or http: / / ftp.tuebingen.mpg.de / pub / protevo / CLANS / ) as disclosed in the Examples. In some embodiments, CLANS is used with default parameters: BLOSUM62 matrix, E-value cutoff of 1e-4. In some particular embodiments, the PolARV has a sequence which clusters with the 25 sequences of AFV1-like PolARVs, in particular SEQ ID NO: 123, 125, 126, 131, 129, 130, 175, 4, 124, 128, 127, 273, 277 and 298, as shown in Table 1. In some particular embodiments, the PolARV has a sequence which clusters with the sequences of AFV2-like PolARVs, in particular SEQ ID NO: 37, 5, 11, 38, 143, 142, 39, 58, 40, 164, 54, 41, 48, 161, 61, 62, 173, 52, 141, 152, 167, 172, 56, 162, 60, 163, 57, 165, 176, 30 145, 59, 149, 170, 44, 53, 43, 166, 146, 148, 49, 140, 174, 150, 42, 151, 46, 201, 275, 77, 285, 301, 45, 289, 66, 136, 68, 135, 267, 72, 293, 296, 65, 319, 265, 316, 282, 268, 257, 280, 287, 312, 187, 230, 243, 191, 252, 254, 261, 263, 309, 313, 310, 314, 315, 79, 278 and 318, as shown in Table 1. In some particular embodiments, the PolARV has a sequence which clusters with the sequences of SIFV-like PolARVs, in particular SEQ ID NO: 47, 115, 112, 250, 307, 106, 5 240, 107, 110, 117, 235, 109, 198, 111, 113, 116, 118, 94, 18, 8, 95, 102, 100, 98, 99, 96, 101, 232, 9, 19, 20, 21, 22, 103, 23, 104, 322, 321, 105, 320, 108, 236, 114, 189, 251, 192, 229, 234, 237, 242, 239, 303, 195, 249, 197, 202, 295, 194, 196, 247, 200 and 241, as shown in Table 1. In some particular embodiments, the PolARV has a sequence which clusters with the 10 sequences of HAV1-like PolARVs, in particular SEQ ID NO: 10, 119, 120, 121, 122 and 304 as shown in Table 1. In some particular embodiments, the PolARV has a sequence which clusters with the sequences of SBV1-like PolARVs, in particular SEQ ID NO: 80, 334, 1, 3, 324, 81, 207, 208, 238, 244, 188, 210, 217, 33, 218, 327, 90, 220, 223, 221, 226, 224, 209, 215, 219, 336, 222, 15 205, 211, 82, 84, 330, 85, 204, 89, 93, 325, 328, 337, 83, 91, 14, 88, 206, 338, 2, 216, 15, 331, 339, 335, 16, 17, 323, 86, 87, 212, 214, 213, 255, 326, 332, 92 and 329, as shown in Table 1. In some particular embodiments, the PolARV has a sequence which clusters with the sequences of SH1-like PolARVs, in particular SEQ ID NO: 24, 12, 26, 28, 177, 31, 178, 32, 35, 183, 180, 185, 181, 169, 171, 179, 33, 157, 139, 6, 13, 7, 133, 155, 34, 132, 25, 51, 50, 20 147, 160, 27, 30, 168, 184, 154, 36, 153, 159, 156, 144, 158, 182, 29, 138, 55, 134, 137, 233, 256 and 276, as shown in Table 1. In some particular embodiments, the PolARV has a sequence which clusters with the sequences of MetaB PolARVs, in particular SEQ ID NO: 64, 266, 274, 291, 69, 70, 288, 281, 294, 292, 305, 253, 264, 299, 67, 190, 203, 269, 270 and 71, as shown in Table 1. 25 In some embodiments, the PolARV polymerase comprises an amino acid sequence having at least 85 % identity with any one of SEQ ID NO: 1 to 339. The PolARV may comprise a sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100 % identity with any one of SEQ ID NO: 1 to 339. In some preferred embodiments, the PolARV polymerase comprises an amino acid sequence having at least 85 % identity with any one of SEQ ID NO: 1 to 272, 274 to 283, 285 to 288, 290 to 297, 299 to 313 and 316 to 339. All these sequences share the preferred sequence pattern (II) as disclosed herein. The PolARV may comprise a sequence having at least 85%, 5 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100 % identity with any one of SEQ ID NO: 1 to 272, 274 to 283, 285 to 288, 290 to 297, 299 to 313 and 316 to 339. In some more preferred embodiments, the PolARV polymerase comprises an amino acid sequence having at least 85 % identity with any one of SEQ ID NO: 1 to 65, 67 to 135, 137 10 to 272, 274, 275 and 277 to 283, 285 to 288, 290 to 297, 299 to 313 and 316 to 339. The PolARV may comprise a sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100 % identity with any one of SEQ ID NO: 1 to 65, 67 to 135, 137 to 272, 274, 275 and 277 to 283, 285 to 288, 290 to 297, 299 to 313 and 316 to 339. In some even more preferred embodiments, the PolARV polymerase 15 comprises an amino acid sequence having at least 85 % identity with SEQ ID NO: 1 or SEQ ID NO: 3; still more preferably SEQ ID NO: 3. In some preferred embodiments, the PolARV polymerase comprises an amino acid sequence having at least 85% identity with the N-terminal region of the catalytic core domain (positions 7 to 83 of SEQ ID NO: 1) corresponding to the three alpha helices packed against the beta- 20 sheet comprising the conserved motifs (a) to (f), which means the catalytic core domain without the C-terminal flexible loop. The percent amino acid sequence or nucleotide sequence identity is defined as the percent of amino acid residues or nucleotides in a Compared Sequence that are identical to the Reference Sequence after aligning the sequences and introducing gaps if necessary, to achieve the 25 maximum sequence identity and not considering any conservative substitutions for amino acid sequences as part of the sequence identity. Alignment for purposes of determining percent amino acid sequence identity can be achieved in various ways known to a person of skill in the art, for instance using publicly available computer software such as the GCG (Genetics Computer Group, Program Manual for the GCG Package, Version 7, Madison, Wisconsin) 30 pileup program, or any of sequence comparison algorithms such as BLAST (Altschul et al., J. Mol. Biol., 1990, 215, 403-), FASTA or CLUSTALW. When using such software, the default parameters, are preferably used. The BLASTP program uses as default a word length (W) of 3 and an expectation (E) of 10. One aspect of the invention relates to an isolated PolARV polymerase having an amino acid sequence which clusters with the sequences of one group of PolARVs chosen from AFV1- 5 like, AFV2-like, SIFV-like, HAV1-like, SBV1-like, SH1-like and MetaB, based on pairwise sequence similarity. The clustering of PolARV sequences based on pairwise sequence similarity is illustrated in Figure 2. The clustering of PolARV sequences based on pairwise sequence similarity may be obtained using appropriate algorithms and / or softwares for clustering amino acid sequences that are well-known in the art. A particular example is 10 CLANS or as disclosed in the Examples. In some embodiments, CLANS is used with default parameters: BLOSUM62 matrix, E-value cutoff of 1e-4. In some embodiments, the PolARV has a sequence which clusters with the sequences of 15 AFV1-like PolARVs, in particular SEQ ID NO: 123, 125, 126, 131, 129, 130, 175, 4, 124, 128, 127, 273, 277 and 298, as shown in Table 1. In some embodiments, the PolARV has a sequence which clusters with the sequences of AFV2-like PolARVs, in particular SEQ ID NO: 37, 5, 11, 38, 143, 142, 39, 58, 40, 164, 54, 41, 48, 161, 61, 62, 173, 52, 141, 152, 167, 172, 56, 162, 60, 163, 57, 165, 176, 145, 59, 149, 20 170, 44, 53, 43, 166, 146, 148, 49, 140, 174, 150, 42, 151, 46, 201, 275, 77, 285, 301, 45, 289, 66, 136, 68, 135, 267, 72, 293, 296, 65, 319, 265, 316, 282, 268, 257, 280, 287, 312, 187, 230, 243, 191, 252, 254, 261, 263, 309, 313, 310, 314, 315, 79, 278 and 318, as shown in Table 1. In some embodiments, the PolARV has a sequence which clusters with the sequences of SIFV- like PolARVs, in particular SEQ ID NO: 47, 115, 112, 250, 307, 106, 240, 107, 110, 117, 235, 25 109, 198, 111, 113, 116, 118, 94, 18, 8, 95, 102, 100, 98, 99, 96, 101, 232, 9, 19, 20, 21, 22, 103, 23, 104, 322, 321, 105, 320, 108, 236, 114, 189, 251, 192, 229, 234, 237, 242, 239, 303, 195, 249, 197, 202, 295, 194, 196, 247, 200 and 241, as shown in Table 1. In some embodiments, the PolARV has a sequence which clusters with the sequences of HAV1-like PolARVs, in particular SEQ ID NO: 10, 119, 120, 121, 122 and 304 as shown in 30 Table 1. In some embodiments, the PolARV has a sequence which clusters with the sequences of SBV1-like PolARVs, in particular SEQ ID NO: 80, 334, 1, 3, 324, 81, 207, 208, 238, 244, 188, 210, 217, 33, 218, 327, 90, 220, 223, 221, 226, 224, 209, 215, 219, 336, 222, 205, 211, 82, 84, 330, 85, 204, 89, 93, 325, 328, 337, 83, 91, 14, 88, 206, 338, 2, 216, 15, 331, 339, 335, 5 16, 17, 323, 86, 87, 212, 214, 213, 255, 326, 332, 92 and 329, as shown in Table 1. In some embodiments, the PolARV has a sequence which clusters with the sequences of SH1- like PolARVs, in particular SEQ ID NO: 24, 12, 26, 28, 177, 31, 178, 32, 35, 183, 180, 185, 181, 169, 171, 179, 33, 157, 139, 6, 13, 7, 133, 155, 34, 132, 25, 51, 50, 147, 160, 27, 30, 168, 184, 154, 36, 153, 159, 156, 144, 158, 182, 29, 138, 55, 134, 137, 233, 256 and 276, as shown 10 in Table 1. In some embodiments, the PolARV has a sequence which clusters with the sequences of MetaB PolARVs, in particular SEQ ID NO: 64, 266, 274, 291, 69, 70, 288, 281, 294, 292, 305, 253, 264, 299, 67, 190, 203, 269, 270 and 71, as shown in Table 1. In some particular embodiments, said PolARV polymerase having an amino acid sequence 15 which clusters with the sequences of one group of PolARVs as disclosed herein further comprises a catalytic core domain, wherein the catalytic core domain has: (i) a structure comprising three alpha-helices (alpha1 to alpha3) packed against a beta-sheet consisting of three beta-strands (beta1 to beta3) and (ii) an amino acid sequence comprising the conserved motifs in (a) to (g) as disclosed herein; preferably, the structure further comprises a C-terminal 20 loop as disclosed herein. The PolARV polymerase according to the invention is capable of incorporating various nucleotides, including deoxyribonucleotides, ribonucleotides and modified nucleotides. Modified nucleotides may carry modifications, either on the 2’ hydroxyl group, such as 2’-O- methyl-dGTP, 2′-O-methyl-CTP, 2′-fluoro-CTP, 2′-fluoro-GTP and others; or on the 3’ 25 hydroxyl group, such as 3’-O-amino-dNTPs, 3’-O-azido-dNTPs and others. Modified nucleotides include in particular 3’OH blocked modified nucleotides such as without limitation 3’-O-amino-dNTPs, 3’-O-azido-dNTPs and other 3’OH blocked modified nucleotides. 3’OH blocked modified nucleotides such as reversible terminators are well- known in the art (Review for example in Fei Chen et al., Genomics Proteomics 30 Bioinformatics, 2013, 11, 34-40; see also EP2607369B1). Modified ribonucleotides (rNTPs) that can be incorporated by PolARVs according to the invention include in particular pseudouridine-UTP (^^-UTP) and 5-methyl-CTP which are known to enhance mRNA stability and are used particularly in RNA printing applications. Modified nucleotides include also locked nucleic acid (LNA) such as modified LNA as disclosed herein. In addition, the PolARV polymerase according to the invention has terminal nucleotidyl transferase activity (or 5 nucleotidyl transferase activity), which means that it is able to synthesize a nontemplated sequence of nucleic acids (de novo nucleic acid synthesis). In some embodiments, the PolARV is a recombinant PolARV. Another aspect of the invention relates to a PolARV variant, preferably a PolARV variant having improved properties (i.e., an engineered PolARV). In particular, the PolARV variant 10 may have at least one mutation which improves recombinant PolARV production, which improves recombinant PolARV solubility, and / or which improves recombinant PolARV enzymatic activity (which means that the PolARV variant has an improved production yield, solubility, and / or enzymatic activity). In particular embodiments, the engineered PolARV has improved ability to incorporate modified nucleotides such as modified dNTPs and / or modified 15 rNTPs. As used herein, the term “PolARV variant” refers to a polypeptide comprising an amino acid sequence having at least 70% sequence identity with the native sequence of a PolARV from archaeal viruses according to the present disclosure; preferably having at least 75%, 80%, 85%, 90%, 95%, 98% or 99% sequence identity with the native sequence. The term variant 20 includes fragments. As used herein, the term “fragment” with respect to a PolARV protein, refers to a polypeptide having a sequence of at least 20 consecutive amino acids from a PolARV protein sequence. A PolARV fragment is also named truncated form of PolARV. The term “variant” refers to a functional variant having the activity of the native sequence (i.e. 25 polymerase activity, in particular nucleotidyl transferase activity). Functional fragments of the native sequence or variant thereof are also encompassed by the present disclosure. The polymerase activity of a variant or fragment may be assessed using methods well-known by the skilled person such as those disclosed herein. In some embodiments, the term "variant" refers to a polypeptide having an amino acid 30 sequence that differs from a native sequence by the substitution, insertion and / or deletion of less than 100, 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10 or 5 amino acids. In a preferred embodiment, the variant differs from the native sequence by one or more conservative substitutions, preferably by less than 50, 40, 30, 25, 20, 15, 10 or 5 conservative substitutions. Conservative substitutions are substitutions of one amino acid with another having similar 5 chemical or physical properties (size, charge or polarity), which substitution generally does not adversely affect the biochemical, biophysical and / or biological properties of the protein. Examples of conservative substitutions may be within the following groups : Group 1-small aliphatic, non-polar or slightly polar residues (A, S, T, P, G); Group 2-polar, negatively charged residues and their amides (D, N, E, Q); Group 3-polar, positively charged residues 10 (H, R, K); Group 4-large aliphatic, nonpolar residues (M, L, I, V, C); and Group 5-large, aromatic residues (F, Y, W); or within the groups of basic amino acids (R, K, H), acidic amino acids (D, E), polar amino acids (Q, N), hydrophobic amino acids (M, L, I, V), aromatic amino acids (F, W, Y), and small amino acids (G, A, S, T). In some embodiments, the PolARV variant, comprises at least one mutation, preferably at 15 least one substitution, in the C-terminal loop. In some embodiments, the PolARV variant, comprises one or more mutations, preferably substitutions, at positions selected from the group consisting of S42, S43, N44, H46, I95, F96, T97, M100, N101, K102, E103, V105, F107, and optionally D71 and R82; or S42, S43, N44, H46, I95, R94, F96, T97, K99, M100, N101, K102, E103, V105, F107; and optionally E13, 20 D71, R82; in particular E13 in combination with H46; preferably E13; H46; E13 and H46; F96; M100; N101; and V105, the indicated positions being determined by alignment with SEQ ID NO: 1. The PolARV variant preferably comprises an alanine substitution or another amino acid substitution. In some embodiments, the PolARV variant comprises at least one of the following substitutions: E13Q; H46N; E13Q and H46N; F96Y; M100A; M100K; N101A 25 and V105A In particular embodiments, the PolARV variant comprises at least one of the following substitutions: E13Q and H46N; F96Y; M100A and M100K. In some embodiments, the PolARV variant comprises one or more mutations, preferably substitutions, at other positions. In some embodiments, the polymerase variant comprises an amino acid sequence having at 30 least 85 % identity with any one of SEQ ID NO: 1 to 339. The PolARV variant may comprise a sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100 % identity with any one of SEQ ID NO: 1 to 339. In some preferred embodiments, the polymerase variant comprises an amino acid sequence having at least 85 % identity with any one of SEQ ID NO: 1 to 272, 274 to 283, 285 to 288, 290 to 297, 299 to 313 and 316 to 339. The PolARV variant may comprise a sequence having at least 5 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity with any one of SEQ ID NO: 1 to 272, 274 to 283, 285 to 288, 290 to 297, 299 to 313 and 316 to 339. In some more preferred embodiments, the PolARV variant comprises an amino acid sequence having at least 85 % identity with any one of SEQ ID NO: 1 to 65, 67 to 135, 137 to 272, 274, 275 and 277 to 283, 285 to 288, 290 to 297, 299 to 313 and 316 to 10 339. The PolARV variant may comprise a sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100 % identity with any one of SEQ ID NO: 1 to 65, 67 to 135, 137 to 272, 274, 275 and 277 to 283, 285 to 288, 290 to 297, 299 to 313 and 316 to 339. In some even more preferred embodiments, the PolARV variant comprises an amino acid sequence having at least 85 % identity with SEQ ID NO: 1 15 or SEQ ID NO: 3; still more preferably SEQ ID NO: 3. In some embodiments, the PolARV polymerase or variant thereof forms a dimer. The PolARV polymerase or variant thereof as disclosed herein may further comprise an additional heterologous sequence. The heterologous sequence is a sequence different from the sequence naturally present in the native PolARV sequence. The heterologous sequence may 20 be for example another protein moiety to form a fusion protein or a tag. The heterologous sequence may be added at the N-terminus or C-terminus of the PolARV sequence. In some embodiments, the heterologous sequence, such as a tag or protein moiety, is added at the N- terminus of the PolARV sequence. A tag is usually of up to 50 amino acids. A tag may be added for the purification, the detection (antibody epitope or label) or the coupling to a 25 molecule or agent of interest. In some embodiments, the tag is a purification tag suitable for affinity purification such as polyhistidine tag or streptavidine tag including double streptavidine tag. Polyhistidine tag usually comprises at least 5 histidines which bind to metal matrices comprising nickel or cobalt. Streptavidin tag binds specifically to ligands such as streptactin. The tag may be removable by chemical agents or by enzymatic means such as 30 proteases (TEV protease, Thrombin, Factor Xa or Enteropeptidase). In some embodiments, the protein moiety is Maltose Binding Protein (MBP); MBP is used to increase the solubility of recombinant proteins expressed in Escherichia coli. In some embodiments, the protein moiety is another enzymatic domain, endowed with the RNA or DNA synthetic, modifying or degrading activities. In some embodiments, the protein moiety is a nucleic acid binding domain, which can recruit the fusion enzymes to RNA or DNA in a sequence specific or non- 5 specific manner. The PolARV polymerase and variant thereof according to the invention can be made by routine techniques in the art, in particular by expression of a recombinant DNA in a suitable cell system (eukaryotic or prokaryotic) and screened for activity (i.e. polymerase activity) using the assays described herein or other similar assays. 10 Nucleic acid, vector, cell The invention relates also to an isolated nucleic acid comprising a nucleotide sequence encoding the PolARV polymerase (native sequence or variant thereof) in expressible form. The nucleic acid encoding the PolARV polymerase in expressible form refers to a nucleic acid molecule which, upon expression in a cell or a cell-free system, results in a functional protein. 15 The nucleic acid may be recombinant, synthetic or semi-synthetic nucleic acid which is expressible in the recombinant cell. The nucleic acid may be DNA, RNA, or mixed molecule, either single- and / or double-stranded which may further be modified and / or included in any suitable expression vector. The nucleic acid may comprise a coding sequence which is optimized for the host in which the PolARV construct is expressed. 20 In some embodiments said nucleic acid comprises the sequenceSEQ ID NO: 340. The coding sequence is operably linked to appropriate regulatory sequence(s) for its expression in the host cell (recombinant cell). Such sequences which are well-known in the art include in particular a promoter, and further regulatory sequences capable of further controlling the expression of a transgene, such as without limitation, enhancer or activator, 25 terminator, kozak sequence and intron (in eukaryote), ribosome-binding site (RBS) (in prokaryote). In some particular embodiments, the coding sequence is operably linked to a promoter. The promoter may be a ubiquitous, constitutive or inducible promoter that is functional in the recombinant cell. As used herein, the terms "vector" and "expression vector" mean the vehicle by which a DNA or RNA sequence (e.g. a foreign gene) can be introduced and maintained into a host cell, so as to transform the host and promote expression (e.g. transcription and translation) of the introduced sequence. The recombinant vector can be a vector for eukaryotic or prokaryotic 5 expression, such as a plasmid, a phage for bacterium introduction, a YAC able to transform yeast, a transposon, a mini-circle, a viral vector, or any other expression vector. The vector may be a replicating vector such as a replicating plasmid. The replicating vector such as replicating plasmid may be a low-copy or high-copy number vector or plasmid. Another aspect of the invention relates to an expression vector for the recombinant production 10 of a PolARV according to the present disclosure in a host cell, comprising a nucleic acid encoding said PolARV according to the present disclosure. The PolARV may be a native PolARV from archaeal virus or a variant thereof such as engineered PolARV according to the present disclosure. The nucleic acid according to the invention is prepared by the conventional methods known15 in the art. For example, it is produced by amplification of a nucleic sequence by PCR or RT- PCR, by screening genomic DNA libraries by hybridization with a homologous probe, or else by total or partial chemical synthesis. The recombinant vectors are constructed and introduced into host cells by the conventional recombinant DNA techniques, which are known in the art. A further aspect of the invention provides a host cell comprising the nucleic acid or 20 recombinant vector. Prokaryote cell is in particular bacteria or archaea. In some embodiments, the prokaryotic cell is a bacterial cell, in particular an E. coli cell. In other embodiments, the prokaryotic cell is an archaeal cell, in particular Saccharolobus islandicus, Saccharolobus solfataricus, Sulfolobus acidocaldarius, Haloferax volcanii or Haloarcula hispanica. Another aspect of the invention relates to a method of production of the PolARV according to 25 the present disclosure. In some embodiments, the method comprises: (i) culturing the host cell of the present disclosure for expression of said PolARV by the host cell; (ii) recovering the PolARV from the culture medium or host cells; and (iii) purifying said PolARV. In some embodiments, the method of production of the PolARV is a cell-free method. Cell-free methods for protein production are disclosed for example in Levine et al., ACS Synth Biol. 30 2020 Oct 16;9(10):2765-2774. doi: 10.1021 / acssynbio.0c00283. The PolARV may be a native PolARV from archaeal virus or a variant thereof such as engineered PolARV according to the present disclosure. Uses of PolARv polymerases The invention also encompasses the use of the PolARv polymerase and variant thereof 5 according to the present disclosure for nucleic acid synthesis, as well as methods of using the same and kits thereof. The PolARV polymerase according to the invention has terminal nucleotidyl transferase activity, which means that it is able to synthesize a nontemplated sequence of nucleic acids; it has de novo nucleic acid synthesis activity. In one embodiment, the invention relates to a 10 method for nucleic acid synthesis comprising incubating the PolARV polymerase or variant thereof according to the present disclosure with at least one oligonucleotide primer and at least one nucleotide under conditions that allow incorporation of the nucleotide(s) on the primer. Such conditions are well-known in the art and disclosed in the Examples of the present application. 15 The oligonucleotide primer and at least one nucleotide may comprise natural deoxyribonucleotides (dATP, dGTP, dCTP, dTTP, dUTP) or ribonucleotides (ATP, GTP, CTP, UTP), modified deoxyribonucleotides or ribonucleotides or combination thereof. In some embodiments, the nucleotide is a 3’OH modified nucleotide comprising a 3’OH blocking moiety such as without limitation 3’-O-amino-dNTPs, 3’-O-azido-dNTPs and other 20 3’OH blocked modified nucleotides. Nucleotides modified with a 3’OH blocking moiety that block 3’extension such as reversible terminators are well-known in the art (Review for example in Fei Chen et al., Genomics Proteomics Bioinformatics, 2013, 11, 34-40; see also EP2607369B1). In some embodiments, the nucleotide is a 2’ OH modifed nucleotide such as without limitation 25 2’-O-methyl-dGTP, 2′-O-methyl-CTP, 2′-fluoro-CTP, 2′-fluoro-GTP and others. In some embodiments, the nucleotide is a modified ribonucleotide (rNTP) such as without limitation pseudouridine-UTP (^^-UTP) and 5-methyl-CTP which are known to enhance mRNA stability. In some embodiments, the nucleotide is a locked nucleic acid (LNA) nucleotide, such as an unblocked LNA-containing nucleotide, a modified LNA nucleotide including without limitation 3’-O-benzyl-LNA-TTP (LNA-Benz), 3’-O-pivaloyl-LNA-TTP (LNA-Piv), 3’-O- azidomethyl-LNA-TTP (LNA-azidomethyl), and a-phosphorothioate-LNA-TTP (LNA- 5 Thiophos) . In some preferred embodiment, the method is for untemplated nucleic acid synthesis, wherein the method is performed in the absence of nucleic acid template. For nucleic acid amplification, the incubation mixture further comprises at least one nucleic acid template. The nucleic acid template is any target nucleic acid of interest. The nucleic acid 10 template may be DNA, RNA or mixed nucleic acid. Due to their de novo nucleic acid synthesis activity and their ability to incorporate various nucleotides including modified nucleotides, the PolARV and variant thereof according to the invention may be used in a wide variety of applications and techniques. Non-limiting examples include enzymatic and chemo-enzymatic synthesis of nucleic acids, in particular 15 enzymatic synthesis of nucleic acids such as DNA printing (as initially disclosed in Palluk et al., Nat. Biotechnol., 2018, 36, 645-650) and RNA printing; RNA synthesis, functionalized nucleic acid synthesis; synthetic biology (artificial genetic polymers (XNAs) such as aptamers); unnatural base pairs. The present invention also encompasses a kit for nucleic acid synthesis, comprising a PolARV 20 polymerase or variant thereof according to the present disclosure, and optionally, at least one nucleotide(s), reaction buffer, and / or oligonucleotide primer; preferably further comprising at least one modified nucleotide, in particular at least one 3’-OH blocked modified nucleotide. In some embodiments, the reaction buffer comprises manganese and / or cobalt ions that are able to stimulate the PolARV activity. In particular, the reaction buffer comprises 1mM or 25 2mM manganese and / or cobalt ions. In some preferred embodiment, the kit is for untemplated nucleic acid synthesis, wherein the kit does not comprise a nucleic acid template. The practice of the present invention will employ, unless otherwise indicated, conventional techniques which are within the skill of the art. Such techniques are explained fully in the literature. The invention will now be exemplified with the following examples, which are not limitative, 5 with reference to the attached drawings in which: FIGURE LEGENDS Figure 1: Sequence alignment of ten representative members of the PolARV family encoded by archaeal viruses Amino acid residues absolutely conserved in this set of viruses are highlighted on the black 10 background, whereas other residues are highlighted based on their conservation (Blosum62 matrix). Histogram of per residue conservation is shown under the alignment. AZI75987: Sulfolobales_Beppu_virus_1 (SEQ ID NO: 1); YP_003737: Captovirus_AFV1 (SEQ ID NO: 4); YP_009817885: Sulfolobales_Beppu_filamentous_virus_3 (SEQ ID NO: 11); 15 YP_009408151: Metallosphaera_turreted_icosahedral_virus (SEQ ID NO: 2); YP_271898: Haloarcula_hispanica_virus_SH1 (SEQ ID NO: 6); YP_005352816: Haloarcula_hispanica_icosahedral_virus_2 (SEQ ID NO: 7); YP_003773414: Hyperthermophilic_Archaeal_Virus_1__HAV1 (SEQ ID NO: 10); YP_001496975: Acidianus_filamentous_virus_2 (SEQ ID NO: 5); 20 NP_445695: Sulfolobus_islandicus_filamentous_virus (SEQ ID NO: 8); and YP_001798546: Acidianus_filamentous_virus_9__AFV9 (SEQ ID NO: 9). Figure 2: Clustering of PolARV sequences based on pairwise sequence similarity Each circle (node) represents a sequence and is colored based on its source database, with verified archaeal viruses shown in black. Edges (lines) in this network connect sequences 25 which share significant sequence similarity. Major clusters including archaeal virus representatives are circled and labelled. Figure 3A-C: Production and purification of recombinant SBV1pol. A. Schematic view of SBV1pol construct used in this study. B. SBV1pol is purified to homogeneity after expression in E. coli. C. After 14-His cleavage, SBV1pol has a molecular 30 weight of less than 15 kDa. Figure 4A-D: Structural characterization of SBV1pol (A) AlphaFold model prediction of SBV1pol. (B) Experimental structure of SBV1pol determined by x-ray crystallography at 1.85 Å resolution. (C) Overlay of the AlphaFold model and the experimental structure reveals an alpha-helix 4 facing the polymerase active site, and 5 whose structure has not been predicted using AlphaFold. (D) Modeling of an incoming nucleotide inside SBV1 catalytic core. This model helps mapping possible residues to be mutated (S42, S43, N44, H46, I95, T97, M100, N101, E103, V105, F107), in order to broaden the substrate selectivity of the enzyme to modified nucleotides. Figure 5A-C: (A) Front (left) and side (right) views of SBV1pol showing the positions of the 10 catalytic residues indicated in black sticks. These catalytic residues constitute the four conserved catalytic AEPs motifs. (B) Structural diagram of SBV1pol highlighting the position of the four conserved motifs. (C) Comparison of SBV1pol to four representative Primase- Polymerases (Prim-Pols) from different species. SBV1pol differs from all other Prim-Pols by 2 main features: First, all Prim-Pols possess an N-terminal selectivity domain (gray square) 15 completely absent in SBV1pol. This selectivity domain contains a conserved residue (called steric gate) responsible for the sugar selectivity. Second, motif IV is different in SBV1 and other polymerases. Figure 6A-D: Nucleotide incorporation assays (A) In vitro nucleotide (dNTPs) incorporation assay indicating terminal transferase activity 20 for SBV1pol (polSBV). (B) In vitro nucleotide incorporation (dNTPs) assay using three different ions: manganese, magnesium and cobalt. Manganese and cobalt stimulate SBV1pol activity whereas magnesium has no obvious effect. (C) Mutations of the catalytic residues severely reduce SBV1pol activity. The triple mutant (DDD-AAA) corresponds to mutations D9A-D11A-D72A. (D) Activity profiles of PolSBV wild-type and variant enzymes with 25 mutation(s) in the catalytic core domain, either in the N-terminal region or the C-terminal flexible loop. Figure 7A-D: Modified nucleotide assays In vitro nucleotide incorporation assays showing that SBV1pol is capable of incorporating ribonucleotides (A &D), 2’-O-blocked nucleotides (B&D) as well as 3'-O-blocked nucleotides 30 (C&D). Figure 8A-D. Biochemical and functional characterization of polSoV (A). Size-exclusion chromatography profile of purified polSoV using a Superdex S7510 / 300 G column (left), and corresponding SDS-PAGE analysis showing a single band at ~18 kDa (right). (B) Ion screen assay of polSoV incubated with a 19-mer FAM-labeled ssDNA in the 5 presence of dTTP or UTP at 37^°C for 3 hours. (C) Temperature-dependent activity assay of polSoV performed from 25^°C to 80^°C in the presence of dNTPs and a 20-mer FAM-labeled ssDNA. (D) Optimization of the DNA-to-nucleotide ratio using a 19-mer ssDNA and two concentrations of dTTP or UTP, incubated with polSoV at 37^°C for 3 hours. Figure 9A-D:.X-ray structure and structural features of apo-polSoV 10 (A) Overall structure of apo-polSoV showing chain A and chain B forming a dimer. (B) Structure of the polSoV monomer with a close-up view of the active site highlighting the conserved catalytic motifs: motif I (D9–x–D11), motif II (H48), motif III (D72), and motif IV (R75). (C) AlphaFold prediction of the polSoV monomer revealing a β-like structure spanning residue 95–111. (D) Comparative activity assay of SoV and SBV using a 20-mer FAM-labeled 15 ssDNA in the presence of the four native dNTPs. Figure 10A-B: Characterization of polSoV with dTTP and dGTP (A) Activity assay of polSoV with a 20-mer FAM-labeled ssDNA in the presence of the four native dNTPs, showing nucleotide preference in the following order: dCTP > dATP > dTTP > dGTP. (B) Activity assay of polSoV with a 20-mer FAM-labeled ssDNA in the presence of 20 the four native rNTPs. Figure 11A-C: 2′-selectivity of polSoV (A) Activity assay of polSoV in the presence of a 19-mer FAM-labeled ssDNA and either the four canonical deoxyribonucleotides or ribonucleotides. (B) Time-course incorporation (1, 2, and 3 hours) of 2′-fluoro-dCTP or dATP using an 18-mer FAM-labeled ssRNA template. 25 (C) Incorporation of 3′-O-NH₂-modified nucleotides in the presence of a 20-mer FAM-labeled ssDNA. Figure 12A-C: Biotechnological potential of polSoV (A) Incorporation of a broad range of modified nucleotides by polSoV in the presence of a 19- mer FAM-labeled ssDNA (gel image on the left), with the corresponding list of nucleotide 30 analogs shown in the table on the right. (B) Efficient incorporation of base-modified nucleotides by polSoV. (C) Comparison of ion screen assays using a 19-mer FAM-labeled ssDNA and an 8-mer FAM-labeled ssRNA substrate, demonstrating polSoV activity on an RNA template. Figure 13: SDS-PAGE analysis of purified AFV1 showing a single band at approximately 5 20^kDa. EXAMPLES Example 1: Identification of the PolARV family Comparative genomic analysis of evolutionarily unrelated archaeal viruses with linear dsDNA genomes revealed that many of them share a small protein (100-150 amino acids) of unknown 10 function, suggesting that these proteins might be involved in a key step of linear genome replication. To collect a representative set of related proteins iterative psi-blast searches were performed queried with the protein of Sulfolobales Beppu Virus 1 (SBV1; SEQ ID NO: 1), an unclassified hyperthermophilic archaeal virus previously discovered through metagenomics, against the non-redundant protein database (GenBank, NCBI). After 13 iterations with the 15 inclusion E-value threshold of 1e-3, no new sequences were identified and the search converged on 155 protein sequences. These included protein homologs from multiple archaeal virus families, unclassified archaeal viruses as well as diverse bacterial and archaeal sequences discovered through metagenomics. To further enrich the set of homologs, additional homology searches were performed against the metagenomics databases, namely, MGnify 20 (https: / / www.ebi.ac.uk / metagenomics / sequence-search / search / phmmer) and IMG / VR (https: / / img.jgi.doe.gov / vr / ) queried with the protein sequences of homologs encoded by the following archaeal viruses: SBV1 (AZI75987), Acidianus filamentous virus 2 (AFV2, YP_001496975; SEQ ID NO: 5), Hyperthermophilic Archaeal Virus 1 (HAV1, YP_003773414; SEQ ID NO: 10), Acidianus filamentous virus 1 (AFV1, YP_003737; SEQ 25 ID NO:4) and Haloarcula hispanica virus SH1 (YP_271898; SEQ ID NO: 6). Following the clustering of the protein dataset to eliminate redundancy and removal of false positives, the final dataset contained 339 protein sequences (SEQ ID NO: 1 to 339). The collected protein sequences were aligned using MAFFT (with G-INS-1 option) and analysis of the multiple sequence alignment revealed a set of conserved sequence motifs, 30 bearing similarity to motifs characteristic of enzymes in the Archaeo-Eukaryotic Primase (AEP) superfamily (Figure 1), a vast assemblage of enzymes capable of nucleotide polymerization activities and distributed across all three domains of life and their viruses. Importantly, even after 13 iterations of psi-blast, no matches to previously characterized enzymes were obtained, suggesting that the assembled dataset represents a new family of viral 5 proteins, which is referred to as PolARV. Clustering based on pairwise sequence similarity using CLANS with default parameters (BLOSUM62 E-value cutoff of 1e-4), confirmed the homology between the proteins in our dataset and further revealed several clusters, most of which included representative archaeal viruses 10 (Figure 2). The PolARV family includes highly divergent representatives. Although sequence similarity within clusters is relatively high, between the clusters the pairwise identities can be rather low, ~20-30%. Nevertheless, all clusters are reliably interconnected through, as shown in Figure 2 and further confirmed through analysis of the conserved motifs identified in PolARV members. Analysis of the sequence alignment of all 339 family members showed 15 that they can be characterized by the following shared sequence pattern: <x(3,230)-D-[LFMIYVKWHCR]-[DN]-x(22,65)-H-x(17,41)-[DSECN]-[DCNH]-x(2)- [RHKL]-x(6)-[RKALGQSYCHT]-x(7,22)-[FYWLHMN]-x(3,185)> where x is any amino acid and numbers in parentheses indicate the range; variation within conserved amino acid positions is provided within square brackets, with the most conserved 20 amino acid residues listed first, and the indicated residues are in positions 9, 10, 11, 48, 71,72, 75, 82 and 96 after alignment with SBV1pol amino acid sequence of SEQ ID NO: 1 used as reference. The majority of PolARVs (333 out of 339) can be characterized by the following preferred shared sequence pattern: 25 <x(3,230)-D-[LFMIYVKWHCR]-[D]-x(22,65)-H-x(17,41)-[D / E]-x(2,3)-[R]-x(6)- [RKALGQSYCHT]-x(7,22)-[FYWL]-x(3,185)> in which the indicated residues are in positions 9, 10, 11, 48, 72, 75, 82 and 96 after alignment with SBV1pol amino acid sequence of SEQ ID NO: 1 used as reference. This sequence pattern includes invariable residues (D9, D11, H48, R75) in the catalytic motifs I, II and IV and a conservative D to E substitution in motif III (position 72) together with a bulky hydrophobic residue (F, Y, W, L) in the loop region (position 96). In addition, position 82 is dominated by R (309 / 333) and K (7 / 333). In other proteins, although the exact residue 5 at that position is different, there is usually an R or K close by. For instance, in 4 proteins which have A, G or T in that position, R is present 1 residue down, which will likely play the role of R82 of SBV1pol (and probably occupies the same position in the structure). The 333 sequences which include the preferred shared sequence pattern are SEQ ID NO: 1 to 272, 274 to 283, 285 to 288, 290 to 297, 299 to 313 and 316 to 339. 10 Thus, the PolARV family can be defined using a combination of the conserved pattern provided above and amino acid similarity-based clustering with other members of the family. Example 2: Protein production and purification Using the innovative bioinformatic approaches presented above, 10 distinct representatives of the PolARV family (SEQ ID NO: 1 to 10) were selected and cloned. 15 The genes of the PolARV homologs were all cloned as follow: the open reading frame (ORF) of the PolARV homolog encoded by Sulfolobales Beppu virus 1 (SBV1) (a virus which grows optimally at 80ºC and pH 2) was optimized for expression in E. coli and synthesized by Genart (ThermoFischer) and inserted into a pRSFDuet plasmid (Novagen) with a TEV-cleavable N- terminal 14^×^His tag (Figure 3-A). So far, focus was mainly on SBV1pol, the PolARV family 20 member encoded by SBV1. SBV1pol is expressed in E. coli strain BL21-DE3* at 37^°C in LB medium supplemented with 100^μg^mL−1of kanamycin. Protein expression was induced by adding 1^mM IPTG. Following expression, cell pellets were resuspended in lysis buffer [50 mM Hepes (pH 7.0), 300 mM KCl, 10 mM imidazole, and 1 protein inhibitor cocktail tablet (Thermo scientific)], and cells were sonicated. The resulting lysate was then mixed with 2 μl 25 of benzonase (25 kU of stock, activity per microliter, Sigma-Aldrich) and MgCl2 to a final concentration of 10 mM and left on ice for 20 min. The lysate was then centrifuged (30,000g for 20 min). The supernatant was purified using a HisTrap FF crude Ni– nitrilotriacetic acid affinity column previously equilibrated with lysis buffer and eluted using the lysis buffer containing 1 M imidazole. Eluted SBV1 was bound to a Heparin HP cation 30 exchange column in buffer A [50 mM Hepes (pH 7.0), 50 mM KCl] and eluted using a linear gradient of buffer A with 2 M KCl (Figure 3-B). The 14-Histidine tag was cleaved over-night by addition of 3% (w / w) TEV protease with 0.5 mM final concentration of EDTA (Figure 3- C). Finally, the protein is extensively dialyzed into a final buffer of 20 mM Hepes (pH 7.0), 150 mM KCl before being stored at −80°C for further use. Overall, SBV1pol could be purified 5 to homogeneity with a final yield of 1 mg per liter of culture. Example 3: SBV1pol: a small polymerase with a minimal conserved AEP core 1. Material and methods X-ray crystallography for structure determination of SBV1pol For crystallization experiments, SBV1pol was concentrated using Millipore concentrators 10 with 10 kDa cut-off membranes through centrifugation at 5,000 revolutions per minute (RPM). To remove any potential aggregates, the sample was further centrifuged at 13,500 RPM for 10 minutes. The protein was concentrated until a final concentration of 2 mg / mL was reached. Crystallization screens were performed at the HTS Crystallization Platform (PFX, Pasteur Institute) following the sitting drop vapor diffusion method in 400 nL drops (1:1 15 protein to reservoir ratio) and with visualization at 18°C. Single crystals grew in 3 to 7 days in ^0.2 M Lithium sulfate, 0.1 M Tris Hydrochloride (pH 8.5), 30% PEG-400 v / v^ and were frozen in 33% glycerol. Diffraction data were collected at the Proxima1 beamline at synchrotron SOLEIL (St. Aubin). The dataset was indexed and integrated using the XDSME package (XDS Made Easier, https: / / github.com / legrandp / xdsme) and the CCP4 suite. The 20 dataset is at a 1.85-Å resolution. The structure was resolved by molecular replacement, with the Phaser Phenix software (https: / / phenix-online.org), using an AlphaFold prediction model of SBV1pol (Jumper, J., Evans, R., Pritzel, A. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021)). The (^4)-loop-(^4) of SBV1pol were built manually using the Coot program (CCP4 package). Four SBV1pol monomers were 25 placed in the asymmetric unit. Refinement was done using Phenix or Buster [global Phasing, Bricogne G., Blanc E., Brandl M., Flensburg C., Keller P., Paciorek W., Roversi P, Sharff A., Smart O.S., Vonrhein C., Womack T.O. (2017). BUSTER version 2.10.4. Cambridge, United Kingdom: Global Phasing Ltd.]. Enzymatic activity assay Nucleotide incorporation assays were carried out in the presence of a 20-mer FAM-labeled DNA substrate. The sequence of the DNA substrate used in the activity assays shown in figure 6C is: 5'GTGCATGCAGAGCGGGGCTC -3' (SEQ ID NO: 341). In all other activity assays 5 the sequence of the DNA substrate is: 5'-TTTTTTTTTTTTTTTTTTCC-3' (SEQ ID NO: 342). Reactions were performed in 10 µL of commercial buffer (Thermopol 10X, NEB), 2^mM MnCl2, and 250^µM of each dNTPs in the presence of 4 µM SBV1pol or the different mutants and 1 µM of the DNA substrate. Reactions were conducted at 60^°C for one hour^and quenched by addition of 10^µl of formamide and 20^mM EDTA, before heating at 95^°C for 10^min. 10 Products were resolved on 25% polyacrylamide gels. Gels were run for 3 hours at 3000 V then scanned with a Mode Imager Typhoon 9500 (GE Healthcare). 2. Results To define the structural properties of SBV1pol, x-ray crystallography and computational modeling with AlphaFold (AF) were pursued. AF was used to predict the structure of SBV1 15 and the predicted model was used as a reference during molecular replacement in the experimental approach. The program succeeded in predicting with high confidence the initial model for SBV1pol, albeit with some variations compared to the experimental structure (Figure 4-A). In parallel, crystals for SBV1pol were obtained. The crystal structure was determined by molecular replacement to 1.85-Å resolution (Figure 4-B), and the x-ray map 20 showed the presence of four SBV1pol molecules in the asymmetric unit. An overlay of the predicted model and x-ray structure illustrates the difference between the two approaches. In particular, the β-hairpin predicted by AF was resolved as a helix with an extended random coil region in the x-ray structure (Figure 4-C). It was possible to confidently build residues 94 to 111 of the SBV1pol within the x-ray map, thus improving the structural characterization of 25 SBV1pol. To gain a more in-depth understanding of the structure of the SBV1pol, these experimental data were compared to previously resolved structures of AEP polymerases from different species (Figure 5-C). Structural analysis of SBV1pol shows that it shares a structural core with the archaeo-eukaryotic replicative primases-polymerases (prim-pols). This shared core 30 includes three strands (^1-^3) and two helices (^1-^2). The prim-pols and SBV1pol also share a third helix (noted ^3 in SBV1pol) that is often preceded by a large domain insertion (^100 residues) in other AEP prim-pols. In all cases, these helices are packed against the ^- strands to form the catalytic core of the enzyme. Indeed, the four conserved catalytic motifs that are characteristic of the AEP superfamily are also present in SBV1pol. The motif I (D-x- 5 D) and motif III (D) are on ^1 and ^2, respectively. Motifs II and IV contain positively charged residues and are located on ^3 and ^3, respectively (see Figure 5-A and 5-B). Third, all AEPs have a large C-terminal domain of ^60 residues that is replaced by a deviated loop of ^20 residues in SBV1pol. This loop in SBV1pol contains one short helix (^4) and one short ^- strand (^4). This loop is packed against the catalytic site and contains a bulky aromatic residue 10 (F96) blocking directly the access to the catalytic active site residues. Notably, the vast majority of the PolARV members contain an aromatic (F, Y or W) or bulky hydrophobic (L) residue at the equivalent position, suggesting an important role in catalysis or regulation. Moreover, this loop was not visible in one of the four SBV1pol molecules present in an asymmetric unit due to different crystal contacts. This observation strongly suggests the 15 flexibility of the loop. It was hypothesized that this region of the SBV1pol encloses the catalytic site and is thus likely to be important for the incorporation of the nucleotide or DNA substrate. To assess the molecular basis of SBV1pol (polSBV) interaction with a DNA substrate, mutations in the core catalytic domain, either within the N-terminal region comprising the 20 three alpha helices packed against the beta-sheet or within the C-terminal flexible loop, were designed. For the catalytic core, a double mutation for motif I (D9A+D11A) and three single codon mutations (H48A; D72A; R75A) for motifs II, III and IV, respectively were introduced. For the flexible loop, mutations R94A, F96A, F96D, F96S, F96Y, K99A and K102A were also introduced. Two additional mutants D71A and R82A, corresponding to the motif in (c) 25 and (f) which are also clustered in the compact catalytic site were also expressed to evaluate their influence on SBV1pol activity. In addition, the modelling of a nucleotide in the crystallographic structure allows to speculate a first list of possible mutations to be tested at positions E13, S42, S43, N44, H46, I95, T97, M100, N101, E103, V105, F107, in order to broaden the substrate selectivity of the enzyme to modified nucleotides (Figure 4-D). E13 is 30 located on beta1 near the catalytic site; S42, S43, N44 and H46 are located on the loop between beta2 and beta3 also close to the catalytic site; and the other residues are in the C-terminal loop. Interestingly, the modelling shows that E13 / H46 are coordinating a manganese ion. Additional mutants at these positions were thus produced and tested (E13Q / H46N; N44K, I95G, M100A, N101A, V105A, F107A). All mutants were cloned, expressed and purified following the experimental procedure described above for the wild type protein. The impact 5 of these mutations was evaluated using in vitro nucleotide incorporation assays (Figure 6-C and 6-D; Table 2). Table 2: Comparative enzymatic activity of polSBV wild-type and variants polSBV_Variants Enzymatic activity polSBV-WT ++ polSBV-E13Q / H46N ++ polSBV-N44K - polSBV-R82A - polSBV-I95G - polSBV-F96A + polSBV-F96S + polSBV-F96D - polSBV-F96Y ++ polSBV-M100A ++ polSBV-M100K +++ polSBV-N101A + polSBV-V105A + polSBV-F107A - In this experimental setup, the mutations in the four conserved motifs (I-IV) exhibit a severely reduced catalytic activity in the presence of a 20-mer DNA substrate (Figure 6-C). These data 10 establish the importance of the corresponding residues for the enzymatic activity of SBV1pol as predicted from the sequence and structure analysis. The loss of catalytic activity observed with the variant R82A confirms the importance of the arginine residue at position 82 (corresponding to motif in (f)) for the enzymatic activity of SBV1pol (Table 2). More generally, this position is dominated by R (309 / 333) and K (7 / 333). 15 In other PolARV proteins, although the exact residue at that position is different, there is usually an R or K close by. For instance, in 4 proteins which have A, G or T in that position, R is present 1 residue down, which will likely play the role of R82 of SBV1pol (and probably occupies the same position in the structure). The decreased activity observed with the variants F96A, F96S and F96D as opposed to a similar activity to wt observed with the variant F96Y confirms the importance of an aromatic 5 (F, Y or W) or bulky hydrophobic (L) residue at position 96 (corresponding to motif (g)) to control the access of incoming nucleotides to the active site (Figure 6-D and Table 2). The activity observed with the variants at the other selected positions predicted to broaden the substrate selectivity of the enzyme to modified nucleotides such as E13, Q46, F96, M100, N101 and V105 (E13Q / H46N, F96Y, M100A, M100K, N101A, V105A; Figure 6-D and 10 Table 2); in particular E13Q / H46N, F96Y, M100A and M100K indicates that they play a significant role in nucleotide incorporation overall and thus may be able to incorporate non- canonical (i.e. modified) nucleotides. Collectively, these data are in agreement with the sequence and structural analyses and highlight the importance of the C-terminal loop along with other key residues of the catalytic 15 core domain for the incorporation of the nucleotide. These preliminary results suggest that the C-terminal loop (92-116) and other key residues of the catalytic core domain or in proximity of the catalytic site present a promising target for engineering to improve the efficiency of SBV1pol for modified deoxyribonucleotides and ribonucleotides. The structure also showed important other structural differences between SBV1pol and AEP 20 polymerases. First, all AEP polymerases contain an N-terminal domain of ^100 residues that closes the polymerase active site (Figure 5-C). More specifically, this domain grasps the ribose moiety of the incoming nucleotide thereby restricting access to the active site of deoxyribonucleotides or ribonucleotides depending on the selectivity of the enzyme. The N- terminal domain in previously known AEP members always contains a conserved residue that 25 is known as the steric gate. This conserved position (can be a histidine, tyrosine, glutamic acid, aspartic acid) controls the sugar moiety of the incoming nucleotide by selectively inserting a deoxyribonucleotide or a ribonucleotide. The corresponding N-terminal domain is completely absent in SBV1pol and most of its homologs, resulting in a widely open active site compared to other AEP polymerases. Instead, SBV1pol apparently evolved a short loop of 30 ^15 residues that can control the access of incoming nucleotides to the active site. This short loop is more easily amenable to engineering than the large compact N-terminal domain found in all other AEP polymerases. This structure therefore establishes that SBV1 shows distinctive features compared to all other AEP polymerases. Example 4: SBV1pol is a terminal transferase 5 1. Material and methods Nucleotide incorporation assays were carried out in the presence of a 20-mer FAM-labeled DNA substrate. Reactions were performed in 10 µL of commercial buffer (Thermopol 10X, NEB), 2^mM MnCl2, and 250^µM of each dNTPs in the presence of 4 µM SBV1pol and 1 µM of the DNA substrate. Reactions were conducted at 60^°C for one hour^and quenched by 10 addition of 10^µl of formamide and 20^mM EDTA, before heating at 95^°C for 10^min. Products were resolved on 25% polyacrylamide gels. Gels were run for 3 hours at 3000 V then scanned with a Mode Imager Typhoon 9500 (GE Healthcare). 2. Results To assess the role of SBV1pol, nucleotide incorporation assays were carried out in the 15 presence of a 20-mer FAM-labeled DNA substrate. SBV1pol can efficiently incorporate deoxyribonucleotides (Figure 6-A) and ribonucleotides (Figure 7-A) with a strong preference to deoxyribonucleotides. SBV1pol shows a strong terminal nucleotidyltransferase activity, being able to repeatedly incorporate dozens of deoxyribonucleotides (dNTPs) from a pool containing all four canonical dNTPs: dATP, dGTP, dCTP and dTTP. However, it is also able 20 to incorporate each of the four dNTPs individually with a marked preference for dCTP (Figure 6-A). The activity of SBV1pol was evaluated in presence of three different ions: manganese, magnesium and cobalt (Figure 6-B). Manganese and cobalt stimulate SBV1pol activity whereas magnesium has no obvious effect. In addition, SBV1pol is capable of incorporating multiple rounds of ribonucleotides (rNTPs), either from a mixture or from individual rNTP 25 pools (Figure 7-A). Importantly, SBV1pol can incorporate modified nucleotides carrying modifications, either on the 2’ hydroxyl group, such as 2’-O-methyl-dGTP, 2′-O-methyl-CTP, 2′-fluoro-CTP, and 2′-fluoro-GTP (Figure 7-B and 7-D) or on the 3’ hydroxyl group including 3’-O-amino-dNTPs and 3’-O-azido-dNTPs (Figure 7-C). These 3’-O-blocked nucleotides act as chain terminators and are key components in next-generation DNA synthesis technologies, 30 often referred to as DNA printing. Their incorporation enables controlled, stepwise DNA synthesis through temporary termination and subsequent removal of the blocking group. The ability of SBV1pol to incorporate such modified nucleotides highlights its potential for further engineering towards biotechnological DNA printing applications. SBV1pol also incorporates modified rNTPs such as pseudouridine-UTP (^^-UTP) and 5-methyl-CTP (Figure 7D), which 5 are known to enhance mRNA stability. The incorporation of rNTPs by SBV1pol in general shows that it also holds potential for RNA printing applications. Example 5: polSoV and polSBV: two close homologs encoded by viruses infecting thermoacidophilic archaea The functional and structural characterization of another PolARV was performed. PolSoV is 10 encoded by SoV, a virus closely related to SBV1, infecting a thermoacidophilic archaeon Sulfurisphaera ohwakuensis (Sulfolobales order). Notably, polSoV shares 84% sequence identity with polSBV and 85.7% identity with the N-terminal region of the catalytic core domain (positions 7 to 83 of SEQ ID NO: 1) corresponding to the three alpha helices packed against the beta-sheet comprising the conserved motifs (a) to (f), which means the catalytic 15 core domain without the C-terminal flexible loop. 1. Material and methods 1.1 Production and purification polSoV The gene encoding polSoV, was codon-optimized for expression in Escherichia coli and synthesized by GenArt (Thermo Fisher Scientific). The synthetic gene was cloned into the 20 pRSFDuet plasmid (Novagen), incorporating an N-terminal 14×His tag followed by a TEV protease cleavage site. The recombinant plasmid was transformed into E. coli strain BL21 (DE3)*, and cells were grown at 37^°C in LB medium supplemented with 100^µg^mL⁻¹ kanamycin. Protein expression was induced by adding 1^mM IPTG. After induction, cells were harvested and resuspended in lysis buffer [50^mM Tris-HCl (pH^8.0), 500^mM NaCl, 25 20^mM imidazole, 10% glycerol, and one protease inhibitor cocktail tablet (Thermo Scientific)]. Cells were lysed by sonication, and the lysate was treated with benzonase (2^μL of 25^kU stock, Sigma-Aldrich) and MgCl₂ to a final concentration of 10^mM, followed by incubation on ice for 20 minutes. The lysate was clarified by centrifugation at 30,000×g for 20 minutes. The supernatant was loaded onto a HisTrap FF crude Ni–NTA affinity column 30 (Cytiva) pre-equilibrated with lysis buffer. The bound protein was eluted with lysis buffer containing 500^mM imidazole. The eluate was further purified using a HiTrap Heparin HP cation exchange column (Cytiva). The protein was bound in buffer A [50^mM Tris-HCl (pH^8.0), 50^mM NaCl, 10% glycerol] and eluted with a linear gradient of NaCl up to 2^M. The 14×His tag was removed by overnight digestion with TEV protease (3% w / w). The 5 cleaved protein was dialyzed extensively into storage buffer [20^mM Tris-HCl (pH^8.0), 250^mM NaCl, 5% glycerol], then stored at −80^°C for future use. 1.2 X-ray crystallography for polSoV structure determination For crystallization experiments, pol SoV was concentrated using Millipore concentrators with 10 kDa cut-off membranes through centrifugation at 5,000 revolutions per minute (RPM). To 10 remove any potential aggregates, the sample was further centrifuged at 13,500 RPM for 10 minutes. The protein was concentrated until a final concentration of 2 mg / mL was reached. Crystallization screens were performed at the HTS Crystallization Platform (PFX, Pasteur Institute) following the sitting drop vapor diffusion method in 400 nL drops (1:1 protein to reservoir ratio) and with visualization at 18°C. Single crystals grew in 7 days in ^0.2 M 15 Ammoniuem Sulfate, 26% PEG-4000 v / v^ and were frozen in 33% ethylene glycol. Diffraction data were collected at the Massif beamline at ESRF (Grenoble). The dataset was indexed and integrated using the XDSME package (XDS Made Easier, https: / / github.com / legrandp / xdsme) and the CCP4 suite. The dataset is at a 1-Å resolution. The structure was resolved by molecular replacement, with the Phaser Phenix 20 software (https: / / phenix-online.org), using an AlphaFold prediction model of polSoV (Jumper, J., Evans, R., Pritzel, A. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021)). The C-terminal loop of polSoV was built manually using the Coot program (CCP4 package). Four polSoV monomers were placed in the asymmetric unit. Refinement was done using Phenix or Buster [global Phasing, Bricogne G., 25 Blanc E., Brandl M., Flensburg C., Keller P., Paciorek W., Roversi P, Sharff A., Smart O.S., Vonrhein C., Womack T.O. (2017). BUSTER version 2.10.4. Cambridge, United Kingdom: Global Phasing Ltd.]. 30 2. Results 2.1 Production and purification polSoV PolSoV was purified to apparent homogeneity with a final yield of ~1^mg per liter of culture. Size exclusion chromatography showed that polSoV eluted as a single species, as confirmed 5 by SDS-PAGE (Figure^8-A). 2.2 Optimization of assay conditions for polSoV activity As polSBV's enzymatic properties have been previously extensively characterized, this study focused on establishing optimized in vitro activity assays specifically for polSoV. The study began by systematically examining the conditions that best support polSoV's activity, 10 including metal ion preference, temperature range, and the DNA / nucleotide ratio. Among the tested ions (Figure 8-B), manganese proved to be the most effective cofactor, significantly enhancing terminal transferase activity. Temperature variations showed no discernible impact, confirming polSoV’s robust activity across a wide range of conditions (Figure 8-C). Additionally, it was determined that a higher nucleotide-to-DNA ratio substantially enhances 15 its activity, consistent with the behavior of terminal transferase enzymes (Figure 8-D). Overall, these findings support the strict terminal transferase nature of polSoV. 2.3 Structure of polSoV apoenzyme PolSoV was crystallized and its X-ray structures determined at 1.1 Å resolution. In the asymmetric unit, four monomers of polSoV are observed. Their symmetric arrangement leads 20 to the formation of a dimer (Figure 9-A), a feature that was also previously observed in polSBV. polSoV display a conserved minimal core characteristic of archaeo-eukaryotic primases (AEPs), featuring a derived RRM-like fold (Lipps et al., Nat. Struct. Mol. Biol., 2004, 11, 157-162; doi: 10.1038), with a β1α1β2β3α2 topology. This core harbors the four conserved motifs typical of AEPs: motif I (D9-x-D11), motif II (H48), motif III (D72), and 25 motif IV (R75) (Figure 9-B). All previously characterized enzymes of the AEP superfamily feature a bipartite active site distributed across the N-terminal α / β domain and the core RRM-like fold. However, the obtained structure of polSoV corresponds exclusively to the core domain, entirely lacking the N-terminal domain, highlighting a key divergence from the typical AEP architecture. However, instead of the N-terminal domain, PolSoV features a C-terminal loop that is strategically positioned against the active site. Moreover, this loop (residues 95–111) adopts distinct conformations in polSoV and polSBV. The C-terminal residues 95–111 stack against the catalytic site in polSBV, while this region in polSoV is more disordered, leaving the site 5 more accessible. AlphaFold further predicts a third alternative conformation for this region, with the C-terminal residues forming a two-stranded β-hairpin (beta4-beta5) (Figure 9-C). It was then investigated whether the observed structural variations between polSBV and polSoV impact their respective enzymatic activities. It was performed in vitro nucleotide incorporation assays using recombinant polSBV-His (His-tagged), polSBV (cleaved His-tag), 10 polSoV-His, and polSoV with a single-stranded DNA substrate. The results revealed two key findings. First, polSoV displayed significantly higher activity than polSBV, as indicated by the longer product lengths, suggesting greater efficiency in nucleotide incorporation (Figure 9-D). Second, the tag-cleaved versions of both enzymes exhibited higher activity than their His-tagged counterparts, likely due to reduced metal ion quenching by the His-tag 15 present in proximity of motif I. Overall, these combined structural and functional insights provide a comprehensive understanding of how polSoV specifically binds nucleotides, highlighting conserved mechanisms essential for its catalytic activity and selectivity. 2.4 Molecular Insights into Nucleotide Binding Specificity PolSoV shows a strong terminal nucleotidyltransferase activity, being able to repeatedly 20 incorporate dozens of deoxyribonucleotides (dNTPs) from a pool containing all four canonical dNTPs: dATP, dGTP, dCTP and dTTP. However, it is also able to incorporate each of the four dNTPs individually a clear preference for dCTP, followed by dATP, dTTP, and dGTP in descending order of favorability (Figure 10-A). In addition, polSoV can incorporate multiple rounds of ribonucleotides (rNTPs), either from a mixture or from individual rNTP pools 25 (Figure 10-B). Collectively, these findings provide critical insights into the nucleotide binding preferences of polSoV and establish a foundation for further exploration of nucleotide interactions within the active site. 2.5 Functional Basis of 2' Moiety Selectivity in polSoV To explore whether polSoV exhibits 2’-sugar selectivity, a series of nucleotide incorporation 30 assays was performed using nucleotides with different 2' sugar modifications. At this stage, the activity assay conditions were further optimized to achieve complete conversion of the initial DNA substrate into reaction products, thereby enabling a more accurate assessment of polSoV’s efficiency in incorporating modified nucleotides.The incorporation assays revealed a clear preference for deoxynucleotides (dNTPs) over ribonucleotides (rNTPs), confirming 5 that polSoV favors substrates with a 2'-deoxyribose sugar (Figure 11-A). Interestingly, when examining the rNTP assays more closely, it was observed a slight preference for rATP over rCTP— a pattern that differs from the selectivity observed with deoxynucleotides, where dCTP is favored. Another surprising observation was that polSoV incorporated 2'-fluoro- dCTP more efficiently than rCTP (Figure 11-B). This finding seems contradictory, as both10 nucleotides possess a modification at the 2' position, yet the enzyme clearly favors the 2'- fluoro derivative over the ribose form. As for 3'-O-NH2 nucleotides (Figure 11-C), the results showed that polSoV could only incorporate one nucleotide, highlighting a clear selectivity against bulky modifications at the 2' position. 2.6 Broad Substrate Tolerance of polSoV: Nucleotide Engineering and RNA 15 Synthesis In the preceding sections, the unique functional characteristics of polSoV were described. Notably, its ability to efficiently incorporate a modified nucleotide: 2'-fluoro-dCTP we was surprising. This unexpected observation prompted the inventors to explore whether polSoV possesses a broader capacity to incorporate a wider variety of modified nucleotides. 20 The results show that polSoV can also incorporate dTTPαS, a nucleotide analog in which one of the non-bridging oxygen atoms in the triphosphate chain is replaced by sulfur (Figure 12- A). This analog is commonly used in various molecular biology techniques, particularly in DNA sequencing and nucleic acid labeling, facilitating detection and visualization in a range of assays. 25 Then, modifications on the sugar moiety were investigated, including the incorporation of locked nucleic acid (LNA) nucleotides. LNAs contain one or more nucleotide building blocks in which an extra methylene bridge locks the ribose in a rigid conformation. Due to their exceptional binding affinity, high mismatch discrimination, low toxicity, and increased metabolic stability, LNAs are widely utilized in antisense technologies and are particularly 30 valuable for in vivo applications that are not accessible via RNA interference. Remarkably, polSoV was able to incorporate at least one unblocked LNA-containing nucleotide, but also modified LNA nucleotides including 3’-O-benzyl-LNA-TTP (LNA-Benz), 3’-O-pivaloyl- LNA-TTP (LNA-Piv), 3’-O-azidomethyl-LNA-TTP (LNA-azidomethyl), and a- phosphorothioate-LNA-TTP (LNA-Thiophos). PolSoV can also incorporate 3’-Phosphate- 5 dTTP. Such 3’-O-blocked nucleotides are known to be important terminators that can be used in next-generation DNA synthesis approaches, known as DNA printing, as their 3’-O protecting group will terminate DNA synthesis, allowing controlled addition of a next desired nucleotide following the removal of the blockage (Figure 12-A). Potentially, this could be extended to XNAs (xenonucleic acids) such as LNA, a feat which is not accessible to most 10 existing template-independent polymerases. Furthermore, polSoV could incorporate at least one 2'-O-methyl-TTP, a modification often used to stabilize short RNA molecules. We also tested modifications on the nucleobase, where polSoV successfully incorporated up to four 2’-ara-CTP molecules as well as several N6- modified ATP analogs (Figure 12-B). 15 Together, these results motivate further exploration of polSoV’s substrate tolerance. This includes testing additional modified nucleotides and engineering specific mutants to enhance its activity. It was next investigated whether polSoV possesses the ability to elongate single-stranded RNA (Figure 12-C). To explore this possibility, we performed an ion screen assay analogous to the 20 one previously conducted with DNA (top gel image). Remarkably, polSoV was able to efficiently elongate an 18-mer FAM-labeled ssRNA substrate in the presence of dTTP and UTP. Among the metal ions tested, manganese and cobalt most strongly enhanced this RNA extension activity. This result reveals yet another intriguing feature of polSoV—its capacity to elongate RNA templates—thereby opening promising avenues for its use in RNA printing 25 and other RNA-based biotechnological applications. Example 6: Production of AFV1pol and SIFVpol The Acidianus filamentous virus 1 gene encoding AFV1pol (amino acid sequence SEQ ID NO: 4) and the Sulfolobus_islandicus_filamentous_virus gene encoding SIFVpol (amino acid sequence SEQ ID NO: 8) were codon-optimized for expression in Escherichia coli and 30 synthesized by GenArt (Thermo Fisher Scientific). For each polARV, the synthetic gene was cloned into the pRSFDuet plasmid (Novagen), incorporating an N-terminal 14×His tag followed by a TEV protease cleavage site. The recombinant plasmid was transformed into E. coli strain BL21 (DE3)*, and cells were grown at 37^°C in LB medium supplemented with 100^µg^mL⁻¹ kanamycin. Protein expression was induced by adding 1^mM IPTG. After induction, cells were harvested and resuspended in lysis buffer [50^mM Tris-HCl (pH^8.0), 500^mM NaCl, 20^mM imidazole, 20% glycerol, and one protease inhibitor cocktail tablet (Thermo Scientific)]. Cells were lysed by sonication, and the lysate was treated with benzonase (2^μL of 25^kU stock, Sigma-Aldrich) and MgCl₂ to a final concentration of 10^mM, followed by incubation on ice for 20 minutes. The lysate was clarified by centrifugation at 20,000×g for 20 minutes. The supernatant was loaded onto a HisTrap FF crude Ni–NTA affinity column (Cytiva) pre-equilibrated with lysis buffer. The bound protein was eluted with lysis buffer containing 500^mM imidazole. The eluate was further purified using a HiTrap Heparin HP cation exchange column (Cytiva). The protein was bound in buffer A [50^mM Tris-HCl (pH^8.0), 50^mM NaCl, 20% glycerol] and eluted with a linear gradient of NaCl up to 1^M. Size exclusion chromatography showed that AFV1pol eluted as a single species, as confirmed by SDS-PAGE (Figure^13). Overall, AFV1pol was purified to apparent homogeneity with a final yield of ~3^mg per liter of culture. With the purification protocol for AFV1pol now established, the next step is to develop and optimize activity assays to enable its functional and structural characterization. Sequences disclosed in the present application SEQ ID NO: 1 (SBV1-like)116aa AZI75987__Sulfolobales_Beppu_virus_1__SBV1_gp38 MVMNIIYIDWDYEITEEEAKERVNKYVKSLKIKPLYIYYRKSSNGHAHLKLVFINDIDVFEHFMIRAIFYDDPK RIRADLMRYYLFRDETKVNRIFTAKMNKEGVKFVGEWKLIWS SEQ ID NO: 2 (SBV1-like)109aa YP_009408151__Metallosphaera_turreted_icosahedral_virus__c109 MITLDHDVPYDGYRHSAKGKLVVEFCKSFELRCWYRRSSTGNVHVAIDVDVDFLRELEIRAVLLDDPMRILNDL RRHAFHMPTGRLWDVKVDREGKRQAGQWEVLFSPE SEQ ID NO: 3 (SBV1-like)116a Sohwa_SBV1-like_ITR_37 (SBV-like) MTMNIIYIDWDYEMTEEEIKEKVKKCVNSLRIKPLYIYYRKSSNNHVHLKLVFINDIDIFEHFMIRAMFYDDPK RIRADLMRYYLTRNEAKINRIFTVKWNKEGVKFAGEWKLIWS SEQ ID NO:4 (AFV1-like)144aa YP_003737__Captovirus_AFV1 (Acidianus_filamentous_virus 1)AFV1_ORF144 MTIVNILRVDIDQPFDYLDKQFYGNLTLRKLLVWRIFYVSKVFSQHVESLEFRKSLSGNIHVLVTINPGIQRTL VPLAQFLMGDDIMRTFMNLKRKGKGNYLFSYGNTDADRFLAERQIARKQKALNRKKSKTKNGEKNGEGKS SEQ ID NO: 5 (AFV2-like)116aa YP_001496975__Acidianus_filamentous_virus_2__gp50 MFQLDFDFKDLSRWDNFYRVYKAIHELSMICKPVVKETHKGYHIYCDIELSPEKIMNLRYYFGDDIWRIMYDEQ RMIFAPHLFDVLYQEKEVFTITPHGIHVDENYHEYDVTDKVL SEQ ID NO: 6 (SH1-like)163aa YP_271898__Haloarcula_hispanica_virus_SH1__ORF41 MNPEHVERPDERFASRITVDLDGDKYTHTQFDLEAVKVWHNLQNSADYVDVHVSSSGLGLHFVAWFEQPLRFHE EVAIRRSSGDDPRRVWMDCQRWLNGLYTDVLFEEKDTREFDKERDFATVYDALAFVAEHRTDDADRMKRLANEG HRAAPDLARKARREA SEQ ID NO: 7 (SH1-like)157aa YP_005352816__Haloarcula_hispanica_icosahedral_virus_2 ((HHIV-2)_gp30 MSDRPDRDYESRITVDLDGDKYTHGQFDLKATQIWHNLQDSADHVDVHVSSSGLGMHFVAWFEQDLQFAEEVAL RRASGDDPRRIWMDCQRWLNGLYTDVLFESKDTRRFDKERGFATVYEALAFINEYRSDDHERVKSVAQDGHRGD PELARRADL SEQ ID NO: 8 (SIFV-like)265aa NP_445695__Sulfolobus_islandicus_filamentous_virus (SIFV)__SIFV0030 MLIAQVIKKGWYIDFKVSSLVDIPTKYLLYTKTIGEEKYGTDAMCHIVFPLRRMEPSFIRGGYNLPTIDKRKRP NQGDGGGIRGHDIGSVWYSSIIFIRGRGGMVWGCSGTTDYMVPPVSGRRGRVYMVDAKQILKLDIDVDVNFYDP NWLLQKKLDMLHALGYQEEEAWWEYSPSGKHIHVIIVLKDPISTKELFDLQFLLGDDHKRVYFNYLRYSVMKED AVHFNVLYTYKKSLTFSDKLKAIFRHWFKSKQYSKNLRLGQTT SEQ ID NO: 9 (SIFV-like)252aa YP_001798546__Acidianus_filamentous_virus_9__AFV9_gp28 MLIAQVVKKGWYIDFKVDKLDDIPARYIPYIRSIGEEKYGTNSVCYIIIRLRRVEPPAFRGSNWLLTAGREEGE RVDKGNRGGVGRPASGFIRHSGFGIADGRGDMVWWCSGVQNNVVSPIPGRRGRIYMVDVRQILKIDIDVSVNYY DPNWLLEKKSQQLRALGYDYDDAWWEYSPSGKHIHILIILKDPVTVKELFDLQFLLGDDPKRVEFNYLRYSVMG EDAIHFNVLYTYKKKLSLKDKIRTIIRHYF SEQ ID NO: 10 (HAV1-like)141aa YP_003773414__Hyperthermophilic_Archaeal_Virus_1__HAV1_gp16 MKIIRLDYDGNISEEQIREWIKYASEIMGIEVKRITIFPSRNGGRHVYVFTDDVTDIQAFLFRYYANEDRKRIN ADIRRWKYNYPIEYSYLFYAYIKKPKAFNTLPLDFELWDYLLANALRYYPLYNENPFPASRWNPFPP SEQ ID NO: 11 (AFV2-like)132aa YP_009817885__Sulfolobales_Beppu_filamentous_virus_3__HOU83_gp44 MLKLDFDFSTLSDWDNFYRFTKAVLELSSYAGVKVTVKETAKGYHIYADLDIDPKKAIPLRYYYFDDDYRIAYD EERVDQVPYLTDVLYQEKLVVTLSITAKRNSIRRRKSIKEHYIEREVNPLSVIEANSI SEQ ID NO: 12 (SH1-like)186 aa YP_009272856__Haloarcula_californiae_icosahedral_virus_1__BGV91_gp36 MSAPASGPVPPGLSVEGSADELPTSPGVKRPIRKHYADRFTVDLDDHGDLHSRAVRAWYYLQRVADAVDVHVSS SGGGLHLVAYLETPMPFHKKVEHRRTAGDDPRRVDMEIQRWHAGLQVDVIFQQKDGGPGEGAKVKDRCFADVWD ALDAIEAHNADDYDRVKTLANEGHKGAPDLARRARWDA SEQ ID NO: 13 (SH1-like) 163aa YP_007761624__Haloarcula_hispanica_virus_PH1__gp35 MTPEHVERPDQEFESRITVDLDGDKYTHTQFDLEAVKVWHNLQNTADYVDVHVSSSGLGLHFVAWFEQPLRFHE EVAIRRSSGDDPRRVWMDCQRWLNGLYTDVLFEQKDTRRFDKERDFATVYDALAFIAANRTDDADRMKRLANEG HRGAPDLARKARRTA SEQ ID NO: 14 (SBV1-like) 162aa WHA35270__Sulfolobus_polyhedral_virus_3__SPV3_ORF40 MLGILFTPVLFVWKMILAYVLNVGHFLVWVICVVEINLIDEASRFTDVVTIDLDGNECEKFIQSKNYETIIKFA TKNEIVKEVYYRCSANNHVHVKVVLKEKVDFLSQLIIRSLLNDDPYRIRADLKRYSVGGEVNRLWDFKIEDGKV KKAGQWIKIYPTQS SEQ ID NO: 15 (SBV1-like) 109aa ASO67407__Metallosphaera_turreted_icosahedral_virus__c109 MITLDHDIPFDQYRQSAKGQLVYEFCRMFRVRCWYRRSASGNVHVAMDMNVDFLRELEIRAILLDDPMRILNDL RRHAFHVPTGRLWDVKVDREGKRQAGQWEVLFSPE SEQ ID NO: 16 (SBV1-like) 127aa WHA35244__Sulfolobus_polyhedral_virus_3__SPV3_ORF13 MIVTLDYDVKFEDFPKTFLFALVSRYCEGKECYVRRSANGRTHVKIMEDVEDFYKRITIRSLLFDDPARISNDI RRHAMGKPTDRLFDIKGYDGVVKHAGEWIPLSFLASPPTVNICQSKKTSRKRR SEQ ID NO: 17 (SBV1-like) 134aa WHA35156__Metallosphaera_turreted_icosahedral_virus_3__MTIV3_ORF14 MNTLRRWGRTSMSSWRVPIDPVHPYVTLDWDVRYDTRHSTLKWRSMVAFLTGAGIHEVMTRHSAGGNTHVIIKW TYDLTFWERLVIRSILQDDPYRISADIRRYYRGLPVERLFGIKATPGVLRKAGPWTRETF SEQ ID NO: 18 (SIFV-like)265aa AOS58382__Sulfolobus_islandicus_filamentous_virus_2 MLIVQVVKKGWYIDYKVSSLGDIPAKFIPYTKTIGEEKYGTDAMCYIVFPLRRMEPSFIRGGYNLRTFGERERA NQGDGGGIRGYDPGLVWYSSIIFARGRGGMVWGCSGTTDSLVPQVSGRRGGIYMVDAKQILKLDIDVDVNFYDP NWLLEKKLQQLHALGYQEEDAWWEFSPSGKHIHIIIVLKDPVSTKELLDLQFLLGDDHKRVYFNYLRYSVMGED AVHFNVLYTYKKSLTFSDKLKAILRHWFRSKQYSKNRNDGQTT SEQ ID NO: 19 (SIFV-like)267aa >YP_001604372__Acidianus_filamentous_virus_3__AFV3 (Betalipothrixvirus acidiani)_gp30 MLVVQVVKPGFYIDYKVRSLEDLPWHFIPFVKTIGEEKYGTDDMCYVILRLRRVESSDFRGGDYLRTITGGERE GAGEGNGRRAGGPDPGFVWGANSPSPHRRGDLVWWCSREKDYLVSSLSFGREGVYMVDARQVVKLDVDVDVTWY DPNWLLEKKLSLLHALGYKEEKAWWELSPSGRHVHLIVVLQSPLSTKELFDLQFLLGDDPKRVEFNYLRYNAIG EDAVHFNVLYAYKKSLTRVDKLRAIFRHWSRSLQYSKKRRDGQMT SEQ ID NO: 20 (SIFV-like)267aa YP_001604189__Acidianus_filamentous_virus_6__AFV6 (Betalipothrixvirus pozzuoliense)_gp31 MLIAQVVKPGFYIDYKVRSLEDIPWYFVHFVKTIGEEKYGTNNMCYIILRLRRVEPPNLRGGDYLRTITGGERE GAGEGNGRRAGGSDPGFVWSANSSLVNRRGDLVWWCSREKDYMVPSLSFGREGVYVVDARQVVKLDVDVDVAWY DPNWLLEKKLSLLHALGYREEKAWWELSPSGRHVHLIIVLPSPLSTKELFDLQFLLGDDPKRVEFNYLRYSAIG EDAVHFNVLYAYKKSLTRVDKLRAIFRHWARSLQYSKKRNDGQMT SEQ ID NO: 21 (SIFV-like)267aa YP_001604307__Acidianus_filamentous_virus_8__AFV8 (Betalipothrixvirus puteoliense)__gp26 MLIAQVVKPGYYIDYKVRSLDDIPWHFIPFVRTVGKEKYGTDSVCYVIIRLRRMESSDLCGGDYLRAIAGGKRE GADKGDGGGVGGPDPGFVRGASSSSSHGRGGLVWWCSRKEDYLVPSLSFGREGVHVVDARQVVKLDVDVDVTWY DPNWLLEKKLALLHALGYREEKAWWELSPSGKHIHLIIVLPSPLSTKELFDLQFLLGDDPKRVEFNYLRYSAIG EDAIHFNVLYTYKKSLTRVDKLRAIFRHWARSLQYSKKRNDGQIT SEQ ID NO: 22 (SIFV-like)267aa YP_001604248__Acidianus_filamentous_virus_7__AFV7 (Betalipothrixvirus pezzuloense)__gp24 MLIAQVVKPGFYIDYKVRSLEDIPWHFIPFVKTIGEEKYGTNDMCYVIIRLRRMESSDIRGGDYLRATAGGKGE GAGEGNDGGIGGSNSGLVRSASSPSSHRRGDLVWWCSRKEDYLVPSLSFGREGVHVVDARQVVKLDVDVDVTWY DPNWLLEKKLSLLHALGYREEKAWWELSPSGRHVHLIIVLPSPLSTKELFDLQFLLGDDPKRVEFNYLRYNAIG EDAIHFNVLYAYKKSLTRVDKLRAIFRHWARSLQYSKKRNEGQMT SEQ ID NO: 23 (SIFV-like)194aa WDS52708__Saccharolobus_shibatae_filamentous_virus_3 MRKVVKTIDKKRGETSDILPEGRTCGQDYGNVWNRCSFDFYRSGGILRRGRRDKDHVESRILIRDEIVGGYVDA RRVLKVDVDIPTSLYSYREMFEKRTSMLRKLGYQIEDARLEFSSSGEHVHYVIVLKEPLANAKEVYDLQVLLGD DQKRATFNYVRYKAIGDNSLHFNVLFKRKMKITRWQKIKLYLSKIV SEQ ID NO: 24 IMG__Ga0102674_10004311 VSAPASGPVPPGLSVEGDGEELPTSPGVKRPIRQHYADRFTVDLDDHGDLHSRAFRAWYYLQRVADAVDVHVSS SGGGLHFVAYLEATMPFHKKVEHRRTAGDDPRRVDMEIQRWHAGLQVDVIFQQKDGQPGEGVKVKDRRFRDVWD ALDAIDAHNANDYDRMKRLANEGHRAAPDLARRADL SEQ ID NO: 25 IMG__Ga0102674_1000553 MSADPRESYRGRVTVDIDGYVGNFELVAVRAYYNLARLADEVEVHISSSGEGIHLIGWFSEECDFPTRLKLRRT LGDDQNRLQFDLERFMAGVYTGVLWSQKDRRDSPENAPENRPGKDRDFADIHDALRHINMSTTTDHDGLKRLAD RGHKGAPELAHHAGRWER SEQ ID NO: 26 IMG__Ga0102676_1001754 MTAPSDPERPIVENYADRFTVDLDAGSDLERRAVRSWHYLGDVADAVDVHVSSSGTGLHFIAYLKESMPFHEKV GHRRTAGDDARRIDMEIQRWHAGLRVDVVFSQKQGPEGEGTQIKDRRFRDVWDALDHIKAQRDDTERMNRLANQ GHKGDPELARRAWRVEA SEQ ID NO: 27 IMG__Ga0102672_10003114 MHHTPYHLRYTVLYHYPAKAMILDRITLDIDYYERNFRLKALKAYHGLEALDFVDEVIVHISTSGRGLHIQGHL SEVLDDNERFALRRSLNDDDKRADLDEQRGAVGHATDIYWDEKEGNEMEREEFEDICTALDRLEKTRAPDVSRV KALAQHGCKGVHDQHGLNRASKGEPTEERAHR SEQ ID NO: 28 IMG__Ga0136447_100068140 MTENSYRGVKRPVREHYTDRFTVDLDAKGGGLEDRAVRSWYYLRDVADAVDVHVSSSGKGLHLVAYLTEPMAFY RKIEHRRAAGDDARRVDMEIQRWQAGLQVDVIFQDKDGGSGEDVNVKERRFRDVWDALDHIRAQRDDVARLDRL AEYGHKGDARLARHARRRDA SEQ ID NO: 29 IMG__Ga0136589_100160111 MTAAASRITVDIDSHVPNPRLRLLRAKQGLERAGAFEIEYATSSSGTGFHIVGYFDEPLARETQFQIRENLNDD PNRLQMDRQRAERGLPINTMWDHKDGNEGSRQVFETVDDVIRFAETNTQSTKGRLNAIQNRGHKAIIDEEIPKM STITAK SEQ ID NO: 30 IMG__Ga0224631_10016602 MILDRITIDIDHYVPSHRLKALETYHALKRMDIVDDVVVHISTGGSGLHIEGHLNEVLSDDERLALRRSLHDDD QRTHLDEERGSVGHATDIYWSQKDGNDGERERMPGIWAALDRLEATRATVHARVKALAQHGRKGCWDTHGLNRP SLVEEGH SEQ ID NO: 31 IMG__Ga0224631_100397216 MSDDVLAYDAWRDPDEFDDVDDDPLVHRYEQRVTVDFDGKHHATAGAFRFEATRTWHNLRDVADRVDVHVSTGQ HGLHFVAWFEADLDFEGQMSVRRSNGDDPRRVDMDRQRFEQLGGRFTDVLFNQKGGRSATKERRFRDVYDALDY IEAQRDDYDRMRRLANDGHKGAPDLARRADL SEQ ID NO: 32 IMG__Ga0224631_10241806 MTSTTSDDSAALVDEFQQRVTIDLDGKDYPSFEAFRYAAVETWHNLHHSSDRVDVAVSSGQHGLHLVAWFHDDV AFYKQVKLRRAHGDDPRRIDMDCQRWLELGGRFSDVLFASKDDRDETKTRYRDVYDALDYIAARRDDLDRLKRL AVEGHRAAPELARRADR SEQ ID NO: 33 IMG__Ga0326458_10183535 MSAADAPLHEHDAGAWYCPDVGAPFGPDEKDDDADACPWCGGPIETADDEGDEFGPDRPDRQFASRITIDLDAD KYDRDGFDLRATRTWHNLQNDADYVDVHVSSSGLGLHFVAWYEQPLRFHEEVAARRTHGDDPRRTWMDCQRWLN GLYTDVLFEEKDSRPMEKERRFSTVYDALDYIREHRRNDTKRMKALAEHGHRAAPELARRARRDA SEQ ID NO: 34 IMG__Ga0399984_004358_4291_4782 MSAVDSGSPTGKADQYRGRITVDLDAKEYSHNNFRTRAVRTWHNLQHDADDVEVHVSSSGRGLHFVAWFEQAIP FHVEIGLRRMHGDDQRRVGMDCERWYNGLYTGVLFDEKSSRPMTKERGFVDVYAALDFIAAQRDDHDRMNRLAN RGHKGAPSLATRSDL SEQ ID NO: 35 IMG__Ga0496842_01694_3578_4054 MTPALVDEYQQRVTVDLDGKDYPSFEAFRYAAVETWHNLQHASDRVDVGVSSGQHGLHFVAWFRDDVAFYKQVK LRRAHGDDARRDDMDCQRWLELGGRFSDVLFTSKGDRDETKTRFPDVYDALDYIAAQRDDLDRLKRLAVAGHRA APELARRADR SEQ ID NO: 36 IMG__Ga0102942_100299915 VSESRLTVDIDDDRAGWEMDAVRAYYGLRRVATSVEVRISSSGSGIHLVGWFDEPLDADEATRLRHTLGDDGKR IELDGMRDRVGHATNVLWTRKSASDGVDTDFTDIHDALAHIGMTA SEQ ID NO: 37 AFV2___MGYP000084726888 MIQLDFDFKDLSKWDNFYKFYKAMYELKLIGTVTIKQTHKGYHIYCSADIPPEKALSLRYYFNDDEWRIRYDEQ RIQQSSYLYDTLYEEKEVFVLTRHGKIVLEHYKERDYADN SEQ ID NO: 38 AFV2__MGYP000028702348 MKLDFDFKDLDKWENFYKLYKAIHELKLVGKVEIKETYKGYHVYASYVTDPDKELNLRYYFGDDPIRIAYDEER LRQAPHLFDVLYRHKITGKVVATRKGLLILTDEKYDEVQVNE SEQ ID NO: 39 AFV2__MGYP000108320997 MKLDFDFKDLDKWENFYRLYKAIYELKLVGEVEIKETYKGYHVYSSYETDPDKELNLRYYFGDDPIRIAYDEER LRQAPHLFDVLYKHKTTGKIVVWSRRILIITGEKYDELQVNG SEQ ID NO: 40 AFV2__MGYP000580124568 MRIDFDFRDLSKWENFYRFFKAISELKLVGDIRVTKTFKGYHIYSSYSTDPEKELTLREYFGDDTIRIMYDEER AKKIPHLFDVLYNYKVAGTIRRIGNKTVLVKVEEYEEEEINV SEQ ID NO: 41 AFV2__MGYP000627308355 MRIDFDFRDLEKWNNFYRFFKAVYELRLVGAVRITKTFKGYHIYSSFSSDPEKELTLREYFGDDTIRIMYDEER MRKVPHLFDVLYNYKVAGSLKRIGRKTVLVKVEEYEEEEVYV SEQ ID NO: 42 AFV2__MGYP000291149273 MPIPPLIGSLMLGIDIDHEPTFSECIHLHNLRRLAKSVQVKKTRKGYHFYFDIDVTPEVGLLLRYIVDDPLRVS FDEARLKAYPELVDVLFEAKLTFNLVVRPDGSTFMKDVTFYAEQDVNFSDLCF SEQ ID NO: 43 AFV2__MGYP000462165471 IAEVIFMYIGVDVDNRDWNAIYTICKSKYLAKWVGGTYRITRTRKGFHVRLYLPFEITPESAIVIRQYLGDDPW RIDYDIARLKVARDLLDVLYEEKATGTYYESSGKAIIGTYYKEVDVDIC SEQ ID NO: 44 AFV2__MGYP000476917579 MLGIDLDDVNWYILCKAKYELVYLFGKDNVKIYKTYKGYHIRINVDIPFDKALSLRIYFYDDDKRLLYDEERQR IGSVSLVDVLYVEKIAGRAYVDKYGNVTKKYSEKYKEEDVTESVCT SEQ ID NO: 45 AFV2__MGYP000728171762 MVSRSLSERMAWKCVTVDIDIEEGSRWETEYRLVKALYMYSHIPRRIVTTRHGIHVYFEGIELSDLRELLKVRA LFGDDERRIEFDERLRERCLHLNVLFQTSEIGEVAVSYTHLTLPTN SEQ ID NO: 46 AFV2__MGYP000459219318 MNLDIDIKPYGLKYITQSELKLPMKWEDFSSICIPLHMIKIISNNNYTVKETFKGFHIYDDMHNADYNRDMPLR LQYDDQWRVIYDEDRYRHNNDDIEVLYKQKVTLNVVYEYTGDGNVVMHYRYDEEYKEVDSYIQDWCF SEQ ID NO: 47 AFV2__ IMG__Ga0187308_1489313 MTATRRLELKHSVLKLDLDFKPPEEWFREWLRTRLIILRAEGLKATKTCIHPTERGYHIWIHLNKPIPYEEQIK LQFLLGDDHQRAYFNQQRMGFTRFRSLFNILFEHKRQLTEEERRIETPQIHPQRNQR SEQ ID NO: 48 IMG__JGI20129J14369_10008762 MRIDFDFLDLSKWENFYRFFKAVFELKMVGTVRITKTFKGYHIHSSYDTDPYKELTLREYFGDDSIRIMYDEER AKKIPHLFDTLYNYKVAGTFKVIKGRTVLIKTEEYEEEEVYV SEQ ID NO: 49 IMG__Ga0081529_1386669 MYIGVDVDNRDWDAIYTICKSKYLAKWVGGNYRITRTKKGFHVRLYLPFEITPESAIAIRQYLGDDPWRIDYDI ARLKVARDLLDVLYEEKVTGTYYEATGKAIIDTYYREVDVDIC SEQ ID NO: 50 IMG__Ga0102676_1002116 MSADPRESYRGRVTVDIDGYVGNFELVAVRAYYNLARLADEVEVHISSSGEGIHLIGWFSEECDFPTRLKLRRT LGDDQNRLQFDLERFMAGVYTGVLWSQKDRRDSPENAPENRPGKDRDFADIHDALRHINMSTTTDHDRLKRLAD RGHKGAPELAHYAGRWER SEQ ID NO: 51 IMG__Ga0102677_1001224 MSADPRESYRGRVTVDIDGYVGNFELVAVRAYYNLARLADEVEVHISSSGEGIHLIGWFSEECDFPTRLKLRRT LGDDQNRLQFDLERFMAGVYTGVLWSQKDRRDSPENAPENRPGKDRDFADIHDALRHINMSTTTDHDRLKRLAD RGHKGAPELAHHAGRWER SEQ ID NO: 52 IMG__Ga0105118_10000023 MRIDFDFQDLSRWDNFYRFFKAVFELKMIGDVRVTKTFKGFHIYSSYDTEPEKELTLREYFGDDSIRIMYDEER AKKIPHLFDVLYNYKVAGTFKVIKGRTVLIKTEEYEEEEVYV SEQ ID NO: 53 IMG__Ga0167616_10013348 MLGIDLDGVDWYILCKAKHELIYLFGKDNVKVYKTYKGYHIRVNVYIPFDKALLLRIYFYDDDKRLLYDEERQR IGSVSLVDVLYTEKIAGRVRVDKYGNVMKKYSEKYKEEDVTESVCT SEQ ID NO: 54 IMG__Ga0187310_149505 MRIDFDFKDLSKWENFYRFFKAINELRLVGAVRVTKTFKGYHIYSTYITDPEKELTLREYFGDDSIRIMYDEER AKKVPHLFDVLYNYKVAGSIRRIGRHTVFVKTEEYEEEEVYV SEQ ID NO: 55 IMG__SL_4KL_010_1000010851 MTAPTTRITIDIDSHVQNPRLRLLRATNGLQRAGAYDVEVATSTSGTGFHVVGFFDEHVPLETQFQVRRNLNDD PNRIQMDRQRALRGLPINTMWTTKDGNEGDRQVFDTVDDAIEYAETNCRTDKQRLNSIQNHGHKAIRDGQIPRM SAVNAL SEQ ID NO: 56 IMG__JGI20127J14776_10019163 MRIDFDFRDLSKWENFYRFFKAISELKLVGDVRVTKTFKGYHIYSSYSTDPEKELTLREYFGDDSIRIMYDEER AKKIPHLFDVLYNYKVAGSVKVIKGRTVLVKVEEYEEEEVYV SEQ ID NO: 57 IMG__Ga0080005_12223821 MRIDFDFRDLSKWENFYRFFKAVYELRLVGAVRVAKTFKGYHIYSSFSSDPEKELTLREYFGDDSIRIMYDEER AKKIPHLFDVLYNYKVAGTFKVVKGRTALVKVEEYEEEEVYV SEQ ID NO: 58 IMG__Ga0080005_1555837 MKLDFDFKDLDKWENFYKLYKAIYELKLVGEVEIKETYKGYHVYSSYETDPDKELNLRYYFGDDPIRIAYDEER LRQAPHLFDVLYKHKTTGKIVVWSRRILIITGEKYDELQVNG SEQ ID NO: 59 IMG__Ga0080006_114576042 VNLDFDFLDLEKWENFYKFMKAIHELRYLGFSVRVYRTFKGYHVRATHPLETGRELPLRIYLGDDPFRIDYDVF RERNNAPVLLDTLYIEKISVVHNRRTEYYQEEEIYAYD SEQ ID NO: 60 IMG__Ga0081473_1309783 MRIDFDFQDLSRWENFYRFFKAVYELRLVGAVRVTKTFKGYHIYSSYNTDPEKELTLREYFGDDSIRIMYDEER AKKIPHLFDVLYNYKVAGTVKVIKGRTVLVKTEEYEEEEVYV SEQ ID NO: 61 IMG__Ga0081473_1342719 MRIDFDFLDLSRWDNFYRFFKAVYELRLVGDVRVTKTFKGYHIYSSYNTDPEKELTLREYFGDDQIRVMYDEER AKKIPHLFDVLYNYKVAGTFKVVKGKTVLVKTEEYEEEEVYV SEQ ID NO: 62 IMG__Ga0079048_10004117 MRIDFDFLDLSRWDNFYRFFKAVYELRLVSTVRVTKTFKGYHIHSSYDTDPYKELTLREYFGDDQIRVMYDEER AKKIPHLFDTLYNYKVAGTFKVVKGRTVLIKTEEYEEEEVYV SEQ ID NO: 63 MCW7078991__Ca_Methanoxibalbensis_ujae MSKRINEYDIKVDLDISKCEGDVVKTAEEICRRMYFIEIYTPVRFGDIKVYETKRGFHLYIDLKEPAYLKKNKA FIVILQLLLMSDWKREVFNLSRVMSMFFLNVEYENWNILFYCKRDTKGNYSIERRTYLSIMFEKILQNYEKIGE TIFDQNEKEVIIDK SEQ ID NO: 64 MAR65923__Crocinitomicaceae_bacterium MTQGLSFDWDGVLRGEPLVEDALNSLAEDFGWTRVFYRTSSSGTGLHILIAELSLDMNLEQSLHPISLSQETIM DYRKRFAEPPWNLECRGRFISDSARSQAGMRTSRVFTVKNDDLSMPWKNIGPRRS SEQ ID NO: 65 HIP90237__Ca_Nanopusillus_sp. MERKTRVVKVDYDSKKLSNYFYINFTKLFANLWEKKGVAIKKIEVLESRKGFHIRLTLNKDITVEESLLFALML YSDPIREVYNYCRYLWDKDITFNFFAKRKMVFFGDGRVVLSRERVTKRSRKLKRKLIRIVMEIEQGKCSKYLGK YF SEQ ID NO: 66 RLE60575__Thermoprotei_archaeon MNYKNLRDEKLIRRITLDIDSKNPLYRLWAMLILRLFNFRDFEIRRSPSGKGYHLVAWHPIGFRLERLLKIRRL AGDDEMRIKLDSMAERQIQVLFDTKVVEEVKLIWSE SEQ ID NO: 67 MAR17766__Paracoccaceae_bacterium MDATLTFDIDDVEDMQVFRQFAEKVATTYDNVWIRKSSSGDGFHLKITGQTDYDEKTGRMIVADKLFNAEDVIS LRDNEEEECRGRLTGDRGRLKVGLQVGRLFGVKSGKSAGEWLPIEAFFKDESILNH SEQ ID NO: 68 MDD3492290__Ca_Thermoplasmatota_archaeon MRIPRRIGRVTCDIDSLDEGALAHAIAVVRQHGFIRDIVFNISSSGNGFHLVAWHRGRGVFKRRLLKIREEAGD DPVRIRLDGLGHRQINVLFTRKVKK SEQ ID NO: 69 MDA8718416__bacterium MWLTFDWDDVSITDPIVETAINELSVRYPLRVWYRISSSGTGLHLVIAELKWDASTGTMQVTPIDFTESETFRI RTEFSAPPWGLECRGRLISDSVRSKNGYRTGRLFSAKNQNLAGGWILYER SEQ ID NO: 70 MCP4124081__Bacteroidota_bacterium MWLTFDWDDVSITDPIVETAINELAARYPLRVWYRISSSGTGLHLVIAELKWDASTGTMQVTPVDFTESETFRI RTEFSAPPWGLECRGRLISDSVRSKNGYRTGRLFSAKNQNLAGGWILYER SEQ ID NO: 71 RLG16669__Ca_Pacearchaeota_archaeon MILFKNKEYFEWVENWKKGLSPLNFDCDSFKKAFEVGLKLFLFFDEVEWRFSSSKRGVHFRVLKNGKQLYIETE KSIYLRMVFGDDPERIAWDLFKYYNNDGPIARLFDTKNGKKASEWKKLDFKSLLLFLKLSRNEKDEKV SEQ ID NO: 72 RLI19273__Ca_Bathyarchaeota_archaeon MRVNVDIDSKNQILLLKVYFWMAQFCKYLEAFETTNGYHVIGYGFPDMTKEEFYEFRRFLGDDEIRVWLDETLH GKPEQVLFSRRVHMEDRKPRIRHRLWSMLWKPFWAKLPARKPYIR SEQ ID NO: 73 MCB7128533__Ca_Bathyanammoxibius_sp. MSFDMARLGKLNLDRWPPKYRVATDYELQLDIDSEQELGFFKMMYQRFAKDLAAQGLASFKPYIVKESRKPGHY HVTVPLEHPLTAFDRGFLQACFGSDRRRELRTWGRVKLGLPDPIVFTIHDLEKPK SEQ ID NO: 74 MDA8161414__Desulfobacteraceae_bacterium MKQASVSVNRNGKHVVRVRDWGRGSGKAGRFLKWGLVEEFHEPGKIMVDIDDKKNMLNLRKSDHIDSRRRKLIG ISSVRRLWIRGISIQSLAHTIGLSVQWIRWDRTRHGWHIVIKVRQKLTLAETIAAQAILGSDPARERLNLARCI SLRKRPSKYWEARANILYSRKAEHG SEQ ID NO: 75 MCL5439375__Ca_Thermoplasmatota_archaeon MRITIDIDKEYNDKNIALGIKAIERLGISLKRVSARRSAHDRIHFDIEYIDELGYWLEEKIDYSDIADMYLILV RLAAGDDIKRVRSDALKWIRGGEIKHWIPAKEDYEKMKKEVRNEV SEQ ID NO: 76 MCL4336290__Ca_Thermoplasmatota_archaeon MKVTVDIDRECNDKNIALAMEAIERVGIDLKRVSVRRSAHDRIHIDIEYPDDLGYWVEDRIDYSDIADMYLILI RLAAGDDIRRVRGDAIRMLKGEEISHWIPAKEDYEKMREKGK SEQ ID NO: 77 MCD6421754__bacterium MITIDFDKPGKTDLDMIFFVLRLKNICDKKRCFIGESHKGYHVYIMGINSDKDGGLRALLMDDPARLWFDDYRI YKNIGYWRNTLFSVKNNGDYLSKEKSLSLKDFMHNILYLDDSFVKIIHRNKTRKMMGGDN SEQ ID NO: 78 MCK5386793__Gammaproteobacteria_bacterium MTEGLVTEFLINNVSRTGIDLDDASEFEFLKAYYNARHMFPDNEVRAYKSSGGDGYHIEIVGVKSRLSIRRTLG DCADRIKYAILRSSTVYGYDELSTGDPIVDDVLFSCKSRVFCGTKYTRRHTQQRVRLDHKSVIAQGFWI SEQ ID NO: 79 RLI98721__Ca_Aenigmarchaeota_archaeon MRVTVDLDENSKRLARWVFFNFIGIFSLPNFSYLKNFVPKIWKTRRGWHFSLNHLRISFEEACMYRLLLNDDRK RVRFDFESVHKPKQILFSKKDGYKKKEVSPEELI SEQ ID NO: 80 SBV1_MGYP000276409847 MKEIYLDYDLRYDSHLDSIIAARFAARISKIISKYGSVSIFKRKSSSGNTHYKLVFDNDISVFDHFIIRAALCD DRDRVYIDLKRYFLNGEQEINRLFDGKVILSLTQVSASDSGNWVPVNLDDMIKAIKFYIKKHIDNGEEKYVQLL RDIP SEQ ID NO: 81 SBV1_MGYP000087678355 MNNEEKIKEIYVDIDLKRELFRASKTWELIRKNLIEIEKRYGKFVLYIRTSSSGNVHLKIEFENAVTVLDMFMI RSLMHDDIYRIGIDLRRLFLQGSTEINRIFDMKARGQQILQASDWELLGWR SEQ ID NO: 82 SBV1_MGYP000606667282 MIKGTFKVITLDWDWNCTDFDPTTNKLIKSISEHPNVKEIWYRCSANEHIHIKVVLRQSIDFWESIYLRAKWDD DANRIRMDVIRYWEHGEVMRLWDAKLVCEEKEGCVLRKAGPWVKINV SEQ ID NO: 83 SBV1_MGYP000488706528 MRDVTLDWDWPWEDFRPEENRLIKAIAGDPRVLWIKARRSANGKIHIWVKLDRDVDFWESLYLRSIWDDDANRI RMDIIRYWEGGETGRLWDVKVDCEGERCEVKKAGRWIEIYRREKR SEQ ID NO: 84 SBV1_MGYP000149601838 MMPRTYFTVVDIFKAITLDWDWPCSDFNPATNNLISSIKRHSNVKEIWYRCSANGHIHVRVFLHEPVDFWESVY LRSIWDDDANRIRMDVIRYWESGEVMRLWDAKLICDEKRCVLRKAGPWRKL SEQ ID NO: 85 SBV1_MGYP000583074339 MREITLDWDWPCKKFVPEESRIIRQIRMHEKVKEIWWRCSANKHLHIKVVLTEDVDFWESIYLRAKWDDDANRI RMDIIRYWENMEYMRLWDEKLVCGKRCVLRKAGPWRKL SEQ ID NO: 86 SBV1_MGYP000536778559 MNGNDYIVGLDYDIKFEDFFAKIDLFKQNFEILKKFIKGKYEIKEIKWRPSSRGHTHTALIFDKPVNKFDSLII RAIMYDDAYRIRADLKRMAEGGKFDILFDRKIKIYKGKREEYVAGEWQSFPIEVK SEQ ID NO: 87 SBV1_MGYP000016906733 MNGNDYIVGLDFDVKFQDFFAKADLFMQNLEVLKKFIKNKYELKEIKWRPSSRGHTHVSLIFNKPVNKIDSLII RAIMYDDAYRIRADLKRMAENGKFDVLFDKKIRIYKGKKEEYVAGEWQSFPIKVN SEQ ID NO: 88 SBV1_MGYP000568329260 MITLDHDVPYAEYRQSAKGQLVLEFCRTFKVRCWYRMSASGNVHVSIDVKVDFLRELEIRAVLLDDPMRILNDL RRHAFHMCTGRLWDAKIDREGKKTAGPWEVLYDPQS SEQ ID NO: 89 SBV1_MGYP000049344190 MVSGTFRYITLDWDWSCKDFEPSNSRLFQQIISDERVKRVWFRCSAHGKIHLRVELSEPVDFWESLYLRSIWDD DANRIRMDIIRYWNDGEVMRLWDSKADCIVREDGKIDCIFLKAGEWKLIYQR SEQ ID NO: 90 SBV1_MGYP000000301035 MNLFDNRSVGSQVTIDLDVFFINYIRSADQVPPDERIGINLFTTRDRIKVLEKRVGMRIEAKYRRSSSGNVHVR LFFPCEVSVLDAFMIRAWMMDDQTRLALDMARYLKSNDLNEMNCCFDEKGDI SEQ ID NO: 91 SBV1_MGYP000645889754 METMRDITLDWDWYFDKYDFEKSYLIKAIKSDPRVVWIKVRKSAHGKIHVWIRLRDEIDFWESLYLRSIWDDDA NRIRMDIVRWWEGGEVGRLWDVKVDCKGDRCEMKKAGKWITIYQREASSSSAV SEQ ID NO: 92 SBV1_MGYP000108321946 MYSQLKFRLNPIYIDIDYMFDVYSFPLPFNAYQVSSKGIHLKLDICEPITLYEHFYIRLLYLDDFNRIKNDLRR LSKGLPEFNRIFKAKLKGGRWIKYGEWMYVRKRCIKKEEILNDSYEFSAEVKYHLLPLVLEILGQLNSLELSD SEQ ID NO: 93 SBV1_MGYP000403190332 MIKGTFDYITLDWDWDCRDFEPTTNKLFQQIISDERVKRVWFRCSAHGKIHLYIELSESVDFWESLYLRAKWND DANRIRMDIIRFWQDGEVMRLWDSKADCVVDNDGKINCVFLKAGEWKLVFQR SEQ ID NO: 94 AFV1_MGYP000243964497 MLIAQVVKPNWYIDFKVNSLEDIPIRYIPYIRSIGEEKYGTDIMCYIVIPLRRMEPPAFRGGDYLSTFEEGEGE TTGKRNGGGIGGHDPRPLWYSGSSFARGRGGLVWGCSGKESDMVSQLSGRWGRIYMVNQLQLLKLDIDVDTIFY DPNWLLEKKLEMLRSLGYQEEDAWWEYSPSGKHIHIFIVLEDPIPVKELFDLQFLLGDDQKRVYFNYLRYSVMK EDAIHFNVLYTYKKSLTLSDKVKAVLRHWFKSRQYSKKRIEGQTT SEQ ID NO: 95 AFV1_MGYP000515246319 MLVVQIVKKGWYIDFKVNSIADIPTKFIPCIREIGEEKYGTDNMCYIILPLRRMEPSIIRGRDYLPAVTEGEGK GTSEGNGGGVGGHDSRPLRDSSTTIARGRGDMVWWCSGEKSNMVPQISWRRGVIRMVSVNQIVKLDIDVDVDFY DPNWLLEKKLQMLHSLGYQEDDAWWEYSPSGKHIHILIVLKDPISTKELFDLQFLLGDDHKRVYFNYLRYSVMG EDAIHFNVLYAYKKSLTFSDKLRAILRHWFKSKQYSKK SEQ ID NO: 96 AFV1_MGYP000706928381 MLIAQVVKKGWYIDYKVNDLDKVPIKFIPYVRAFGKEKYGTNSMCYVIVRLRRMGSSNFCGSTRLRTLRPEEGE GAGQGINGGIGRPNFRPFWHSSPSIINRRGDILWWCSRETSDLVSPLSGRWGRVYMVDQRQLIKLDIDVDVNFY DPNLILEKKLEMLRSLGYQEEDAWWEYSPSKKHIHVLIVLTDPIPVKELFDLQFLLGDDPKRAEFNYLRYSVMG EDAVHFNVLYTYKKSLTFSDKVKAILRHWSKSLQYLKK SEQ ID NO: 97 AFV1_MGYP000417934545 MLIAQVVKKGWYIDFKVNSLDDVPAKFIPYIREVGEEKYGTNDMCYIILPIRRMVTSIVYRGYDLLTFTKGEGE GASEGDGGGVRRYDSRPLWDSSSSFAGGRGDMVWWCSREKGDMVSPIPGRWGIIRMVNVNQIVKLDIDVDVDFY DPNWLLEKKLEMLHSLGYQEDDAWWEPSPSGKHIHVLIVLKNPITTKELFDLQFLLGDDHKRVYFNYLRYSVMG EDAIHFNVLYAYKKSLTFSDKLKAILRHWFKSKQYSKKCKEGQTT SEQ ID NO: 98 AFV1_MGYP000320634216 MLVIQVLRPNWYIDFKVRSIEQVPVKFIPYIKTIGEEKYGTDDMCYVIVPLRRMEPPTFYRRDYLPAEGTQEGE TANQGDGGGAGGRDFGLVWNPSFAVAGRRGGVVWWCCGEKGNVVSPLPWRRGGVRMVNVNQIVKLDIDVDVDFY DPQWLLEKKLEMLHSLGYQEEDAWWEYSPSGKHVHIIIVLKDPVPTKELFDLQFLLGDDHKRVYFNYLRYSVMG ENAIHFNVLYAYKKSLTLSDKVKAILRHWFKSKQYSKKCRDGQTT SEQ ID NO: 99 AFV1_MGYP000161397819 MLIAQAVKKGWYIDFKVNSLDCIPTRFIPYIRSIGEEKYGTYNMCYIIVPLRRLKSSSFRRGYDLSTEETNGGE EANKANGGGARRDDFGSLRDPSPIVTGRRGDMVWWCDRETSSVVSSLPGRWRGLHMVNVNQLLKLDIDVDVVFY DPQWLLEKKLQQLHSLGYQEEDAWWEYSPSGKHVHVLIVLKDPVSTKELFDLQFLLGDDQKRVYFNYLRYSVMK EDAVHFNVLYTYKKSLTFSDKVLAILRHWFRSRQYSKKRIEGHTT SEQ ID NO: 100 AFV1_MGYP000031653108 MLVVQIVKPNWYIDFKVNSLDDVPVKFIPFIKTIGEEKYGTDNMCYIIIPLRRVVASVIRRGYDLPTFTEGERE GTGKGNGGGTRRYDIGLVRDSSVTIARRRGDMVWGCYGEKDSLVSPIPGGRRGIYMVDQRQIVKLDIDVDTIFY DPQWLLEKKLQQLHALGYQEDDAWWEPSPSGKHIHILIVLKDPVSTKELFDLQFLLGDDHKRAYFNYLRYSVMG EDAIHFNVLYAYKKSLTFSDKVKAILRHWFKSKQYSKKCKEGQTM SEQ ID NO: 101 AFV1_MGYP000220373578 MLIAQVIKKGWYIDYKVRSIGDIPTSYIPFIEKIGEEKYGTDNMCYIILRLRRMGTSDFRRGTGLRTVESEEQE ATDSRINGGIGGSDFGPLRGSGASLTDGRGGFLWWCGGKEDSMVSPLSTRWGRVYVVDQRQLIKLDIDVDINFY DPNLILEKKLEMLRSLGYKEEDAWWEHSPSKKHIHILIALTDPIPVKELFDLQFLLGDDHKRAEFNYLRYSVMG ENAIHFNVLYKYKRSLTLSDKLKAILRHWFRSKQYSKKRREGQIT SEQ ID NO: 102 AFV1_MGYP000603722359 MLVVQVVKPNWYIDYKVNSLADIPVKYLPYVREIGEEKYGTNSVCYIELPIRRMEPSIIRRGYNLPTFTEGEGK GASEGNGRRTRRYDSGLVRDSSASFTGGRRSVVWWCSREKDYMVPSIPGRWGTIRMVNVNQIVKLDIDVDTIFY DPNWLLEKKLEMLHSLGYKEDDAWWEYSPSGKHIHILIILKDPVSTKELFDLQFLLGDDQKRVYFNYLRYSVMK EDAIHFNVLYAYKKSLTFSDKLRAILRHWFKSKQYSKKCREGQTT SEQ ID NO: 103 AFV1_MGYP000636154455 KRGWYIDYKVKDIDDIPVRHILYIKTIGYEKHGTNCVSYIELRLPGVESSHLCGSGRVLKFRKREGKRPNRGDR GGTHGSAIGSLRSPSPWVVKRRGGMVRRGTGVRYLVVPQGSETEKGVSVVNQLRLIKLDIDVDIQYYKPEWLLT QKVKLVESLGYEVERAWWEISPSGRHVHILIMLKRPLSVKELFDLQFVLGDDPKRSEYNYLRYSVMGEDAIHFN VLYAYKKPFTLSDKLRAILRHWFKSLQYSKKRKGEQIT SEQ ID NO: 104 AFV1_MGYP000683332791 MLYDDKMYVCDVRKKYWYIDIIIKHIEELSFSDIINLKEAGVTSEPLPRIELRLRRRESPFECRNTGVLSSDNV GEEIEGNAGGLSYGENGGTVRVPSDPIDNGLGNLVREWRNKGGTGGMGTRERSNKGRNFYDRVLKIDVDVPVEY YDPDKLLQDKLKMLRSLGYTEANAYWKPSSSGHHIHIVIELTNDVELSTLFYLQFMLGDDHKRSTLNFFRLKYF PDKAKYFNVLFTEKVKITRRRKLIFLIRHIIRKVF SEQ ID NO: 105 AFV1_MGYP000535888559 MTNIIKLDYDIKHEYFQKYINFDEFIKTRIAILETLGYKVKKWEFSETRRGYHLIIEIDKDLPLQRIFELQFLL GDDHNRVNYNFFRLENWGENYAKYFNLLFTKKFKRK SEQ ID NO: 106 AFV1_MGYP000633814126 MERSNVLKIDLDVKVSQKLLEKWLETRKLILEHLGYTITKIRYVETEKGYHFWIHLKENLEPKEVAELQFLLGD DHNRARYNFLRLKFRTFHEFNVLFNRKKRIERPQY SEQ ID NO: 107 AFV1_MGYP000742917628 MSKSERFNVLKVDLDLKVPQKLLQKWINTRKWLLKRLGYTVTKVNYCETEKGYHFWFHVKEDLTPQEIAELQFL LGDDHNRARYNFWRLQINTFDKFNLLFNYKKKRNRESR SEQ ID NO: 108 AFV1_MGYP000618462941 MTKIVKVDIDLSHKNFDFLTSENEFINTRIAIIEKLGYKVTNYKILKSKHGYHFYFYINKDITMKEMIKMQFLL GDDVHRVAYNIMRLQNFGEKYALYFNILFSKKFKRKKK SEQ ID NO: 109 AFV1_MGYP000521756141 MQVRQNTPQELEVLKIDIDIPWDIFKRLKYTWVATRRAILSTLGYELADWIIHRSQNGKVHAWITVEVRKPLSV WEKAELQFLLGDDHNRSKLNFLRAEKTPEKFDSFNILFSKKAVSYTHLRAHETSLHLVCRLLL SEQ ID NO: 110 AFV1_MGYP000508602690 MSMSTNTLTNILKIDIDVKVSRTELELWLCTRKLLLKHLGYTVKKVKVTNSKKGYHIWIHLNESCTRYEIALLQ FLLGDDHRRCYYNFLRCDLRHANTFNILFNFKITPNE SEQ ID NO: 111 AFV1_MGYP000241629041 MRQNTPQELEVLKIDIDIPIDIFKEVEAEWIITRRAILWVLCYKVRDIIIHKSQNGKVHAWITIETSKPLEPKE KAMLQFLLGDDHNRAKLNFLRATKTPEKFDSFNILFSKKVVKNGGDNK SEQ ID NO: 112 AFV1_MGYP000341894015 MSRVLKIDVDLPYELFIMLYKEWLETREAILSRYGLRVTEYHITRSPSGKTHVYLYVDRPLNARETAVLQFLCG DDHRRALFNFARLAYGEKIFHAFNILFSEKKHVEAGKDPIVELALRGLKHD SEQ ID NO: 113 AFV1_MGYP000088290850 MVTIAKIDVDVKMDNVLLVQFLGTRLAIIEHLGYHLKDVRWRETTHGYHFWFEIEENLTDNELANLQFLLGDDQ IRSKYNFMRVKGKAFRDFNALFSKKVKKRRTSLVRKLIIALWVVWSWLKERLS SEQ ID NO: 114 AFV1_MGYP000462780078 MNTLLIDVDWRPPFKEWIDKWVELRKEMLRRLGVEFEDIVVSYTDRGFHVWVWLKKDLPPEKVLELQFLLGDDP VRCKINYFRLKRGVFERYNLLFSKVIYRRPPDEKCMQCKLWRIYVEEVSGVKPPSYAVEFDIKPEEEEKVVKAC IDISARDPTFEYRYQRGKLIIYSKDRDQAYKRGIWFKANVLGDMKRYFRVREIGSNYEV SEQ ID NO: 115 AFV1_MGYP000496809245 MKVQVLKIDLDIKPPEDWLSEWQETRQLILRSRGLEVTKMFRHETEKGWHIWIHLSQPISFDEAMKLQFLLGDD HQRVYFTLQRRGFTKFKHLFNILFSKKLKQDGES SEQ ID NO: 116 AFV1_MGYP000368427839 MGDRSLVILKIDIDTKVPRKLLNEWIETRRVILQYLGYNVRNIIITETTHGYHAWIYIEDHVDSKRKAELQFLL DDDHRRAQLNLMRAELNAFDEFNALFSKKIKKKVSLLKVLKIWIKVKVKAFSNLVTSKLRK SEQ ID NO: 117 AFV1_MGYP000170857367 MTKTTYFKVDVDVKMSDKLLDRWITTRIKMLEWLGFQLVDFNFRETEKGYHFWFGVQGEYSPKTIAKTQFLLGD DQTRCKFNFLRLEADCFHEFNCLFNKKLKLKNKKRKGG SEQ ID NO: 118 AFV1_MGYP000262268821 MGNRDMVILKIDVDVKMSKQLLREFMETRRVILSYLGYNVRSIVAQETEHGYHFWIDIEDHEELYVDQKRKAEL QFLLGDDQRRAKYNFFRAELGAFDTFNALFSEKIKEKVSLWKVIKIWLKVKLKLLLNSVIGKLRKW SEQ ID NO: 119 HAV1_MGYP000438579838 MKIIRMDYDGNVSEDQIREWIRYASEIMGVEIKRIAIFPSRNGGKHVYVFTNNISDIQSFLFRYYANEDRKRIN ADIRRWKYNYPIEYSYLFYAYIKKPKVFNTLPLDFELWDYLLAITLRYYPLYDAP SEQ ID NO: 120 HAV1_MGYP000751157063 MKIIRLDYDGNISEAQIKELIKYASEITGIKVNKITIFSSRNGGKHVYVFTDNVFDIQAFIFRYYANEDRKRIN ADTRRWKYGYPIEYSYLFYAYIKKPRIFSTLPLDFEFWDYLLAVTVKYYPLYSENTF SEQ ID NO: 121 HAV1_MGYP000016906087 MKLIKLDYDGMDRKEIAEWIDFASEVTGIKVRRITMFPSRNGGTHVYVYVPNDVTDIQIFMFKYYANEDRRRVN ADVRRWKYGYPKDYSYLFYAYINKPKASVSIPNEFEFFINVLAIANEYYPLY SEQ ID NO: 122 HAV1_MGYP000181291329 MQICTNTIHAVHVMIVRLDYDDKNEDEVIEWINYASEVMRIKVKDVILRKSRNSGIHVYVYVSGSLTPIQQFLF RYYAGEDRKRINADIKRYNVGYPLGIDFLFYQYVSKPCSKLPHASEFEYFIHVLAVAAELSQNAKHG SEQ ID NO: 123 IMG_Ga0073352_14605 MTIVNVLRIDVDHPFDYLEKTFYGNTTLRNLLVWRIFYVCKVFSQRVESLEFRRSLSGNIHIYVTISPGIQQTL APLIQFMMGDDIMRTMMNLKRKGKGNYLFSYGNSDADKLLAERQMARKQKALNRKKNKSEIGEKNGERENGKEG GK SEQ ID NO: 124 IMG_Ga0080003_10016094 MSTVNILRVDIDKPFDYLDKLFYRNMTLRNLLVWRIHFTCKVFSQTVHTIEFRRSLSGNIHVYITIQPGIQRTL VPLIQFMMGDDIMRTFMNLKRKGKGNYLFSYGNTDVDRFLAERQTARKQKRKQKALNRKESRRNLGEKNG SEQ ID NO: 125 IMG_Ga0081529_1168807 MTIVNVLRVDIDKPFDYLDKVFYGNLTLRNLLVWRIFYVSKVFSQKVESLEFRRSLSGNIHIYVTISPGIQRTL VPLVQFIMGDDIMRAMMNLKRKGKGNYLFSYGNTDVDRFLAERQVKRKQKALNRKENRGKKGEENG SEQ ID NO: 126 IMG_Ga0081474_13087620 MTIVNVLRVDIDKPFDYLDKVFYGNLTLRNLLVWRIFYVSKVFSQKVESLEFRRSLSGNIHILVVISPGIEQTL SPLVQFLMGDDIMRVAMNLKRKGKGNYLFSYGDSSADRFLAERQVKRKQKALNRKENRGKTGEENG SEQ ID NO: 127 IMG_Ga0187308_109244 MTTVSTLRIDIDQPFDYLDKVFYRNMTLRNLLVWRIFYTSKVFSQHIDAIEFRRSLSENIHIYVIIQPGIQQTL VPLVQFMMGDDIMRTFMNLKRKGKGNYLFSYGDSDADKVLKERQTARKQKRKQKALNRKESKTENGEKNG SEQ ID NO: 128 IMG_Ga0209120_100168315 MSTVNILRVDIDKPFDYLDKLFYRNMTLRNLLVWRIHFTCKVFSQTVHTIEFRRSLSGNIHVYITIQPGIQQTL LPLVQFLMGDDIMRVMMNLKRKGKGNYLFSYGNTDVDRFLAERQTARKQKRKQKALNRKESRRNLGEKNG SEQ ID NO: 129 IMG_Ga0208662_10058918 MTTANILRIDIDKPFDYLDKVFYGNLTLRNLLVWRIFYVSKVFSQKVESLEFRRSLSGNIHILVVISPGIQQTL LPLVQFMMGDDATRTMMNLKRNGRGNYLFSYGDSSADRFLAERQVKRKQKALNRKENRGKTGEENG SEQ ID NO: 130 IMG_Ga0208312_1004087 MTIVNVLRVDIDKPFDYLDKVFYVNLTLRMTLRNLLVWRIFYVSKVFSQKVESLEFRRSLSGNIHILVVISPGI QQTLLPLVQFMMGDDATRTMMNLKRKGKGNYLFSYGNNNADKLLAERQMVRKQKALNRKENRGKTGEENG SEQ ID NO: 131 IMG_Ga0208429_1005046 MTIVNVLRVDIDKPFDYLDKVFYVNLTLRMTLRNLLVWRIFYVSKVFSQKVESLEFRRSLSGNIHILVVISPGI EQTLSPLVQFLMGDDIMRVAMNLKRKGKGNYLFSYGDSSADRFLAERQVKRKQKALNRKENRGKTGEENG SEQ ID NO: 132 SH1_MGYP000659110386 AVSDSGEFATGPANPSDYGARATVDLDWKEYPASQFRMRAIRTWHNLRDAADLVDVHVSSGGMGLHFVAWFEDD VPFHEQIAIRRATGDDPRRVEMDCERWLNGLYTDVLFETKEERPQTKERRFADVYDALDYIYAQRDDHDRMRRL ANEGRRGAPDLVHRQEVDA SEQ ID NO: 133 SH1_MGYP000305394855 MTDHADHRGRITVDLDAKDYSRNQFRLRATRAYSQLADDADEVEVHVSSSGLGLHLVAWFEDPLEFHEEIALRR QHGDDPRRVDMDIKRWFQGLYTGVLFEEKGDKEKERRFSGIFDALDWIDAQRDDANRVKRLANDGHKGAPDLVA RANL SEQ ID NO: 134 SH1_MGYP000332456958 MTATPSPPTIEYRTRVTIDIDHYVSTPRLRLLRAYHNLLHAGATEVKVAVSSSRRGYHVEGLFAEELSDDEQVS IRRNLADDSNRIYMDVERSAIGHTSNVMWFTKSTNDHVRQQFDTPEDAINYVSETRKGDPQRVRAIANHGHKAM NDSSMPRPSNVTNDRK SEQ ID NO: 135 SH1_MGYP000023411091 MKKDLCGEDMYRRITVDIDILDEYEVANALKVLKENGFDKKITLRLSPSGKGYHIIAWSDKGVPLRELLKIREQ AGDDPARIWFDSMTHREINVLFDNKRYRFITMEQLLDDNGSQSVVRYNTFYETYDDMERVLKEKGTWLSDEDFR FLKSLYLYKVIEYNRYSEMPRLIQKKKHGEILQHIFMEFLKYFNIKKFVSHNKIQIEQKTFKKKNLTMLFHIMP LTRGEYLVAIRKLVQSGKIKRKRLNNHGRGRYYYEISRN SEQ ID NO: 136 SH1_MGYP000445089181 MNYKNLRDERLIRRITLDIDSKNPLYRLWAMLILRLFNFRDFEIKRSPSGKGYHLVAWHPIGFRLERLLKIRRL AGDDEMRIKLDSMAERQIQVLFDTKVVEEVKLIWSD SEQ ID NO: 137 SH1_MGYP000597192124 MSHAPRDASFTRPSWTGERRVTIDLDDADPCDLYRAVYGLRQAGAHEVEARVSASGDGAHVRAWFDDDAVTADG VETLRLAHGDHPRRTWMDRDHDAKPQQVMFTSKPGGRAGPWRTDPHVVVDELRRRAEHLERNDKATPGARNWKP SEQ ID NO: 138 IMG__ADL20m3uS_00618990 MTAAASRITVDIDSHVPNPRLRLLRAKQGLERAGAFEIEYATSSSGTGFHIVGYFDEPLARETQFQIRENLNDD PNRLRMDRQRAERGLPINTMWDHKDGNEGSRQVFETVDDVIRFAETNTQSTKGRLNAIQNRGHKAIIDEEIPKM STITAK SEQ ID NO: 139 IMG__Ga0102672_10015612 MNPEHVERPDQRFESRITVDLDGDKYTHTQFDLEAVKVWHNLQNAADYVDVHVSSSGLGLHFVAWYEQPLQFHE EVAARRSHGDDPRRTWMDCQRWLNGLYTDVLFEEKDTRPMAKERRFSTVYDALDFIRERRSDDYDRMKGLANDG HRAAPDLARRADL SEQ ID NO: 140 IMG__Ga0099838_1161291 MYIGVDVDNRDWDAIYTICKSKYLAKWVGGTYRITRTRKGFHVRLYLPFEITPESAIAIRQYLGDDPWRIDYDV ARLKVARDLLDVLYEEKATGTYYEATGKAII SEQ ID NO: 141 IMG__Ga0167616_100072111 MRIDFDFQDLSRWDNFYRFFKAVFELKMIGDVRVTKTFKGFHIYSSYDTEPEKDLTLREYFGDDSIRIMYDEER AKKIPHLFDVLYNYKVAGTFKVIKGRTVLIKTEEYEEEEVYV SEQ ID NO: 142 IMG__Ga0167615_10012449 MKLDFDFKDLHKWENFYKLYKAIYELKLVGDVEVKETYKGYHVYASYATDPDKELNLRYYFGDDPIR IMYDEERLRQSPHLFDVLYKHKITGKVIATRTGLLILTDEKYDEVQVNE SEQ ID NO: 143 IMG__Ga0187308_1489313 MKLDFDFKDLDKWENFYKLYKAIYELKLVGKVEVKETYKGYHIYASYVTDPDKELNLRYYFGDDPIRIAYDEER LRQAPHLFDVLYKHKITGKVVIGSTGLLIMTDEKYDEVQVNE SEQ ID NO: 144 IMG__Ga0187336_10001376 MAHGSRVTLDIDAYIEDFELTATRAAYQLEHLADKVEVRVSSSGEGVHIIAWFEEQLGKDAKRRLRRTLGDDAK RLELDKRRWRFRQTDNVLWTRKENGEADQDFDDIDDALDYIRDARDPRERLKYAVQKGLVV SEQ ID NO: 145 IMG__Ga0187310_1274816 MRIDFDFQDLNKWDNFYRFFKAVFELKMIGKVRVTRTFKGYHIYSSYSTDPEKELTLREYFGDDSIRIMYDEER AKKIPHLFDVLYEYKVAGSIKLMKGKTVLIKNEEYEEEEVYV SEQ ID NO: 146 IMG__Ga0187310_130514 MYIGVDVDNRDWDAIYTICKSKYLAKWVGGTYRITRTRKGFHVRLYLPFEITPESAIAIRQYLGDDPWRIDYDI ARLKVARDLLDVLYEEKVTGTYYESSGIAIIDTYYKEVDVDIC SEQ ID NO: 147 IMG__Ga0256680_100253924 MTDRDECRGRVTIDIDGYVDNFELTALRAYHNLRRAAEHVDVHVSSGEGGLHIIGYFESPKSFPERIHLRRMLG DDQKRINIDIQRFKNGVYTGVLWDQKSRNGSGTKHRDFEDIHDALAFVNRDKNLSPERIKRFAEYGHKRAPELT RHAEGL SEQ ID NO: 148 IMG__Ga0209012_10033876 MYIGVDVDNRDWDAIYTVCKSKYLAKWVGGTYRITRTRKGFHVRLYLPFEVTPESAIAIRQYLGDDPWRIDYDI ARLKVARDLLDVLYEEKVTGTYYESSGKAIIDTYYREVDVDIC SEQ ID NO: 149 IMG__Ga0208314_1017149 VNLDFDFLDLGKWENFYKFMKAIHELKYLGFSVRVYRTFKGYHVRATHPLETGRELPLRIYLGDDPFRIDYDIF RERNNAPVLLDTLYVEKLSVVHNRRTEYYQEEEIYAYD SEQ ID NO: 150 IMG__Ga0208313_1011137 MYIGVDVDNRDWDAIYTICKSKYLAKWVGGTYRITRTRKGFHVRLYLPFEITPESAIAIRQYLGDDPWRIDYDI ARLKVARDLLDVLYEEKATGIYYEATGKTIIDTYYREVDVDIC SEQ ID NO: 151 IMG__Ga0208429_1000144 MLGIDIDHEPTFNECIHLHNLRRLAKSVQVKKTRKGYHFYFDIDVPPEVGLLLRYIVDDPLRVSFDEARLKAYP ELVDVLFEAKLTFNLVVRPDGSTFMKDVTFYAEQDVNFSDLCF SEQ ID NO: 152 IMG__Ga0392390_01197_9457_9807 MRIDFDFQDLRRWDNFYRFFKAIFELKMVGDVRVTKTFKGYHIYSSYMTDPEKELTLREYFGDDSIRIMYDEER AKKIPHLFDVLYNYKVAGTFKVVNGRTVLVKTEEYEEEEIYV SEQ ID NO: 153 IMG__Ga0401378_005373_968_1327 MSESRVTVDIDDDATDWEMDAVRAYYGLRRAASRVEVRISSSGSGIHLVGWFDERLDETQTDRLRRTLGDDGKR IELDGMRTRVGHATNVLWTRKSASAGVDTDFDNIHDALAHIGMTA SEQ ID NO: 154 IMG__Ga0401358_0001079_16217_16594 MASKPRADRLTLDYDAYVPSMALRAIRGYHGLLRDADAVDVYVSTSHRGIHLVGWFDELLDDETRERIRRNLGD DTRRVNLDTDRGRVGHTTDVTWTRKAKNGSQERHQFTDVYDALTFIEQRSL SEQ ID NO: 155 IMG__Ga0496801_00511_5410_5898 MNDACSSSRDELEDREGRITIDLDAKDYSASQFRLKATQIWHNLKDDASDVEVHVSSSGLGLHFVAWFDESLQF YEEVVLRRQHCDDPRRIDMDVQRWLQGLYTGVLFEEKSNRAHEKERRFRDIYDALDWIDNQRDDASRIQRLAQD GHKGAPDLAPRADL SEQ ID NO: 156 IMG__Ga0496842_00441_8484_8903 MSYQSRVTVDIDSHHLDWRLRAVRAYYELARTADEVDIRISSSGEGLHLIGWFSSRLAAEGKTRLRRHLADDSN RVRLDELRGGVGHTTNVCWTSKYVDSDPQHADGDFADVWDALDHIEVEDTPERRLKRALEAGLVV SEQ ID NO: 157 IMG__Ga0496842_04513_4513_5154 MSAADAPLHEYDAGAWYCPDVGAPFGPDEKDDDADACPWCGGPIETAEEQGDEFGPDRPDRQFDSRITIDLDAD KYDRDGFDLRATRTWHNLQNDADYVDVHVSSSGLGLHFVAWYEQPLRFHEEVAARRTHGDDPRRTWMDCQRWLN GLYTDVLFEEKDSRPMEKERRFSTVYDALDYIREHRRNDTKRMEALAEHGHRAAPELARRARRDA SEQ ID NO: 158 IMG__Ga0496812_03007_1093_1512 VTRVDRVTVDVDGTFPDYEQRALTAYYNLRASGAEAVEVRVSSSHEGIHLIAWYDERLGREARLRIRESVADDY NRVRMDHERGRHGHTTGVCWREKSSRDGEAFRVAGIDEALSYTLSRGLALSRWDRPAVAGREHAR SEQ ID NO: 159 IMG__SL_4KL_010_BRINEDRAFT_1000424716 MSRQSRVTLDIDVDTYSWVSKALRAYYGLSRAADEVEVRISSSGSGLHMVGWFDDALDDDQKTNLRRTLGDDAK RLELDTIRSHIGHTTNALWTAKGEGGGQIDASFDHIEDALSHISMTNVSMTDEAKARVTGNPRHAMLPK SEQ ID NO: 160 IMG__ProkHueca_100083828 MGDTPTAGTCRGRVTIDLDGYADRFRFKAIRAYHALQRRAERVEVSISSSGEGLHLVAWFDRSLSFAEKLNIRR ELGDDGKRIQIDRERARHGVYTGVLWSRKSSNPGGRKDRDFADVYAALDHMDATNRSDAERMNRLANGGHKAEP SMARAAHGEVNSR SEQ ID NO: 161 IMGVR__JGI20129J51889_10012011 MRIDFDFLDLSKWENFYRFFKAVFELKMVGTVRITKTFKGYHIYSSYDTDPYKELTLREYFGDDQIRVMYDEER AKKIPHLFDVLYNYKVAGTFKVIKGRTVLIKTEEYEEEEVYV SEQ ID NO: 162 IMG__Ga0080005_1344263 MRIDFDFRDLSKWENFYRFFKAISELKLVGDVRVTKTFKGYHIYSSYSTDPEKELTLREYFGDDQIRIMYDEER AKKVPHLFDVLYNYKVAGSVKVIKGRTVLVKVEEYEEEEVYV SEQ ID NO: 163 IMG__Ga0080005_1348823 VRIDFDFQDLERWENFYRFFKAVNELRLVGAVRVTKTFKGYHIYSSYSTDPKKELTLREYFGDDTIRIMYDEER AKKIPHLFDVLYNYKVAGSVKVIKGRTVLVKAEEYEEEEVYV SEQ ID NO: 164 IMG__Ga0080005_1369049 MRIDFDFRDLSKWDNFYRFFKAISELKLVGDIRVTKTFKGYHIYSSYSTDPEKELTLREYFGDDTIRIMYDEER AKKIPHLFDVLYNYKVAGTIRRIGNKTVLVKVEEYEEEEINV SEQ ID NO: 165 IMG__Ga0080005_1380401 MRIDFDFRDLSKWENFYRFFKAVNELRLVGAVRVTKTFKGYHIYSSFVTDPKKELTLREYFGDDTIRIMYDEER MRKVPHLFDVLYNYKVAGTVKVIKGRTILVKVEEYEEEEVYV SEQ ID NO: 166 IMG__Ga0080004_113468213 MYIGIDVDNRDWNAIYTICKSKYLAKWVGGTYRITRTRKGFHVRLYLPFEITPESAIAIRQYLGDDPWRIDYDI ARLKVARDLLDVLYEEKATGTYYESSGKAIIDTYYKEVDVDVC SEQ ID NO: 167 IMG__Ga0081534_1024119 MRIDFDFKDLSKWDNFYRFFKAVYELKMVGDVRVTKTFKGFHIYSSYNTDPDKELTLREYFGDDSIRIMYDEER AKKVPHLFDVLYNYKVAGTFKVIKGRTVLVKTEEYEEEEVYV SEQ ID NO: 168 IMG__Ga0102672_1002282 MLMYVGMGKGMIPDRITIDLDAYAHNFTLKALRAYYQCSRLDVTEEVIIHISTGGKGLHIEAHFNEILSDEERY RIRRTLADDNKRTDLDEQRGAEGHATDIFWSEKAGNDHERERMDGIWAALDRLESTRASDPSKVKALALHGRKA AWDTHGSFDRPSRAEAE SEQ ID NO: 169 IMG__Ga0102672_1005301 MTEYESRITVDLDGDKYSRTGFESRAVQTWHNLKASADRVDVAVSSGGFGLHFVAWFRDSLTFDAQLQIRRAAG DDPRRVDMDKQRYLNGLFTDVLFREKDGRENTRTEYRDVWDALDAIDGRRDDHDRMKQVANAGHKADPSIARRV SEQ ID NO: 170 IMG__Ga0105118_100000516 VNLDFDFLDLGKWENFYKFMKAIHELKYLGFSVRVYRTFKGYHVRATHPLETGRELPLRIYLGDDPFRIDYDVF RERNNAPVLLDTLYVEKLSVINNRRTEYYREEEIYAYD SEQ ID NO: 171 IMG__Ga0256680_100317414 MTMTRESHTERREEFESRVTIDFDGDKARNLRFEAAKAYHNLREVADEVDVHVSSGQGGLHFVAWFRESLPFHR KIAIRRGHGDDPRRTDMDIQRWQNGIFTGVLFQEKGHAEDARKERRFEDVHDALDYIENRRDDAERMNRLANRG HKGEPELIRHATEVNT SEQ ID NO: 172 IMG__Ga0208447_1001963 MRIDFDFLDLSRWDNFYRFFKAVYELKMVGDVRVTKTFKGYHIYSPYDTDPKKELTLREYFGDDSIRIMYDEER MRKVPHLFDVLYNYKVAGTFKVIKGRTVLVKTEEYEEEEVYV SEQ ID NO: 173 IMG__Ga0208447_1005623 MRIDFDFKDLSRWDNFYRFFKAVYELKMVGAVRVTKTFKGYHIYSSYVTDPYKELTLREYFGDDQIRVMYDEER AKKIPHLFDVLYNYKVAGTFKVIKGRTVLIKTEEYEEEEVYV SEQ ID NO: 174 IMG__Ga0208313_1005772 MYIGVDVDNRDWDAIYTICKSKYLAKWVGGTYRITRTRKGFHVRLYLPFEITPESAIAIRQYLGDDQWRIDYDI ARLKVARDLLDVLYEEKATGTYYEATGKAIIDTYYKEVDVDIC SEQ ID NO: 175 IMG__Ga0208683_1020107 MTTANILRIDIDKPFDYLDKVFYGNLTLRNLLVWRIFYVSKVFSQKVESLEFRRSLSGNIHILVVISPGIQQTL LPLVQFMMGDDATRTMMNLKRKGKGNYLFSYGNNNADKLLAERQMVRKQKALNRKENRGKTGEENG SEQ ID NO: 176 IMG__Ga0326767_000714_3833_4183 MRIDFDFQDLDKWDNFYRFFKAVYELRLVGAVRVTKTFKGYHVYSSFVTDPEKELTLREYFGDDQIRVMYDEER ARKIPRLFDVLYNYKVAGTFKVVNGRTVLVKTEEYEEEEIYV SEQ ID NO: 177 IMG__Ga0496842_00083_20984_21505 MSESWRSPQSPHVERPVVDNYADRWTVDLDGKDGNLEQRAIRAWHYCRRVADGVEIHVSSSGEGLHLLAYHRQP VPFHKKIEHRRAAGDDDRRIDMEIQRWHAGLEVDVVFQQKDSPEGIETVTKERRYADVYDALDAVRANRSDPAE RMRRLANDGHKGAPDLARRARRMEP SEQ ID NO: 178 IMG__Ga0496842_00289_5763_6239 MSRHERYQSRITVDLDGDHSEDFRLRAIRTWHFLDGVADDVKVHVSTGQQGLHFVAWFREALEFHEQISIRRQA ADDARRIDMDIQRWLQVGPEFTDVLFNQKGDRENVKERRFSDVYDALDYVDAYGSDDADRVRRLANDGHKGAPD LARKGPKGWA SEQ ID NO: 179 IMG__Ga0496802_01417_11191_11682 MTMTRDSTTERRDEFESRVTIDFDGDKARNFRFEAVKIYHNLREVADDVEVHVSSGQGGLHFVAWFRDELPFHE KIAIRRAHGDDPRRTDMDVQRWQNGIFTGVLFQEKGHAEDARKERRFADVHDALDYIYSRRDDAERMNRLANRG HKGEPELIRHATEVK SEQ ID NO: 180 IMG__Ga0496803_05564_3452_3949 MTTTTSDDSAALVDEYKQRVTIDLDGKDYPSFEAFRLAAVETWHNLQHASDRVDVAVSSGQHGLHLVAWFHDDV AFYKQVKLRRAHGDDARRDDMDCQRWLELGGRFSDVLFMSKGDRDETKTRFRDVYDALDYIASQRSDLDRLKRL AVEGHGAAPELARRADR SEQ ID NO: 181 IMG__Ga0496806_05944_3404_3946 MSAEPDAGADGDVDASSEPLAEQYQQRVTIDMDGDHYRSFRAFRLAATRIWHNLEAVADRVDVDVSTGQEGLHF VAWFRDDLEMFEQVAIRRAHQDDPRRIDMDVQRWRQLGGRYSDVLFEVKGGRDTRKERRFRDVYDALDYIAGRR SDHARMKRLAIDGHQGDPRLARRARRESEGKR SEQ ID NO: 182 IMG__Ga0496807_00343_7880_8299 VTRVDRVTVDVDGKFPDYEQRALTAYYNLRASGAEAVEVRVSSSQEGVHLIAWYDERLGQDARMRIRESVADDW NRVRMDHERGRHGHTTGVCWREKSSRDGEAFRVAGIDEALSYTLTRGLGLSRWERPAVAGREHAR SEQ ID NO: 183 IMG__Ga0496809_06171_4229_4705 MTPALVDEYQQRVTIDLDGKDYPSFEAFRYAAVETWHNLQHASDRVDVAVSSGQHGLHLVAWFHDDVAFYKQVK LRRAHGDDARRDDMDCQRWLELGGRFSDVLFTSKGDRDETKTRFPDVYDALDYIAAQRDDLDRLKRLAVEGHRA APELARRADQ SEQ ID NO: 184 IMG__Ga0496810_00105_3713_4180 MIADRLTLDIDSYIPNFRLVAIRAWYALQRHDEVTEVTTHVSTSGEGIHLTAHLTTRLDQDTRMQLRRTLGDDQ KRVDLDIERGRVGHATDICWTQKAGNDDERREMRDIWAALDHIEHQRASAHSRVKALSQHGHKAVWDTHGINRA SLAEGIK SEQ ID NO: 185 IMG__Ga0496810_03341_6336_6818 MTPPAALVDEYEQRVTIDLDGKDYPSFEAFRYAAVETWHNLQHASDRVDVAVSSGQHGLHLVAWFEADVPFHEQ IEIRREYGDDPRRIDMDCQRWLELGGRFSDVLFCSKGGRDETKTRYRDVYDALDYIASQRSDLDRLKRLAVEGH HAAPELARRADR SEQ ID NO: 186 MCP4902934__bacterium MSERVLKLDLDVPTLDFEELSERVDWTLSLIRRRATTIGLSESPSRHGWHVRIELNRGVSAMRAVALQAILGSD PKREAFNLARVSDWPQLSPLARLRWNVLFSRKVKL SEQ ID NO: 187 RLJ03953__Ca_Aenigmarchaeota_archaeon MNRVGIDLDYYNLPSVIELKRRILREQRPRGLTQVLVFQTKHGYHLELIYDRDISAEENFQIREQYGDCKKRME YSKKRYDLIGDGYDILFQMKEGVWRRRVWV SEQ ID NO: 188 MCY0881404__Bacillota_bacterium MERKEVIKGGSGKMINKPISSHALYTTPWYTFVYVDIDKPYKTIKYDNEWKGIYRGIQQLTEFNIDIRPSSHGN THMRVYRKDGNLINFVDSMKIRALLHDDPYRFEYDWIRNYYEGIGETNRIFSMKITKDNICKAGAWIPLKQYLE SD SEQ ID NO: 189 MCD6138641__Deltaproteobacteria_bacterium MGNILKIDVDLRIDRWGWIDKYKKFVAAGLKSLGYGVEKIVVRESDSKKGLHIWIYLDREVDDRTKNMLQFLCC DDRTRVRINYYRIEAGVKNWNKLFSKVLYRRPLEPPCSECRLIKYLKEVEERVDNEVCGGKGKGG SEQ ID NO: 190 MBT4661558__Ca_Marinimicrobia_bacterium MMRLTFDWDSVGLEDIMVQKGLAKLSQDFPNNTIYYRISASGTGIHAIISPKNSTPAPIEIEDEDALEYRREMV AFGLEDKYRLAHDEARVGTGLPTAQLWEWKKGKKAGEWIKYVE SEQ ID NO: 191 RLF38074__Thermoplasmata_archaeon MNRVGLDLDYYDLPSVIELKRRILKEEEQNGLTQVLVFKTKHGYHLELIYDRDIPPEENFLIREKYGDCERRLE YSQRRYMLLGDCYDILFHEKKGFLRRRVWI SEQ ID NO: 192 MCW1312749__Ca_Parvarchaeum_tengchongense MKQVFKVDVDLVVPESWLKLWIETRKAILKELGIEPVEINIHKTERGFHAWVHGKVKKKLTPTECNFVQWVLGD DTGRVRINQRRIACGMSWRTFNKLFSEVIWRKKHKCNCDIHKKILKRMEEGRVALLKIMEKEKIKGISLPNIEK NRRLLCR SEQ ID NO: 193 MCC6021120__Thermoproteaceae_archaeon MNAGVLLVDVDEVTLPLCREEYELAARRVLEAFGYSAAGFVYRLSSSRGVHVAVWVEPEPELQYVPLLQYLLGS DARRECINYERIRRGIDINVLFSGRRRVSGGGGLKCRPELMCPGLLRMIELAGKIGVERVELRGGCARG SEQ ID NO: 194 MDO8135783__Ca_Njordarchaeum_guaymaensis MQRGVGDVDVDKVGVLKIDKDCFVDGDWLEDYIRLMVHACRFHGVTVLSVKVCRSRVKGFHLYIEISPAIDSEL ANRLQWLLGDDSARVDFNRARIRSGLNEWNKLFEEPRKRLRTIYKHTDNSFRGRYA SEQ ID NO: 195 RLI87637__Archaeoglobales_archaeon MCSSWKDVIYYEPKKKATQTFKIDIDCHLNEEEIDLFVKTRKAILSSLGFTLIDYRYFQTERGMHFWFIAHGEL LDDKTKNLYQFLLGDDHHRFDINRRRIERGIPWQKANILFSRVIDRKRNYDATFELQTYKKYLDKIEGDKNDFI NYAIKVMLND SEQ ID NO: 196 MBN1357569__Ca_Bathyarchaeota_archaeon MLRRNGNSRESEEVSELKIDKDCFPDEDWIPFYVNVITKTCSDFSVTVNSISVCESKKKGLHFYIRIKPAINSE LANMLQWLLGDDSRRVDFNRARVKSGLKEWNKLFEVRENVLRTIYEKGEMLQ SEQ ID NO: 197 OYT25530__Thermofilum_sp._ex4484_82 MKKIIFKTDLDITLNEEWLKEWVRTRKLILKNLGFKVEDVVVKPSSKRGHHFWWHCMSEKELSDMEIVKVQFLL GDCIGRTLVNIKRVKRGYPMSRGNKLFSLVLWRKDPNEMKENFNKLLDELKEGKKLTKKERRFIVRYASNLQKI VEKYSELMKEGQEILKGG SEQ ID NO: 198 RLG58786__Ca_Geothermarchaeota_archaeon MKIDIDIPWDIFKRLKYTWVATRRAILSTLGYELADWIIHRSQNGKVHAWITVEVRKPLSVWEKAELQFLLGDD HNRSKLNFLRAEKTPEKFDSFNILFSKKVVKSERRKGGDN SEQ ID NO: 199 RLG75071__Thermoprotei_archaeon MSLKEYELKLDFDFKIYNLSEFAKQFVDKLRFLERFYGIYFKSIEVYETNKGYHFYLKVTTERKLNNKDIVVLQ LALGSDYKRELFNWLRVRRGQKFKHWNVLFKCKYENGLKVSEERKTESAIRLEKLVSKLYFNGYKKA SEQ ID NO: 200 MDO8135069__Ca_Njordarchaeum_guaymaensis MDSEGVLKIDRDIHVDADWIKDFICLAIAVCERYGITVVSIKMCQSQHKWLHFYIHITPQIEPEFANRLQFLLG DDARRVDFNRARIRSHLTGWNKLFEQPSSRLTTLYQLPGRRSRSVRVTPW SEQ ID NO: 201 MDA8040848__Pirellulales_bacterium MKITKIGIDIDNYASDPINFERYVDRCKSRARMFGKVYRSRKTKNGYHLFVKLDKSVNFWRSIELRYYCGDDPR RCFYDIMRYRTGGRFIDTLFDRKKRFKRSKDQGSGHGE SEQ ID NO: 202 RLG09439__Ca_Pacearchaeota_archaeon MKTIFKTDLDVKLNKKEMELWIETRKLILKYLGFKVEKVIVRNSSRGHHFWWHVKGKKLKAQEINFIQLLLGDD MGRYYINKRRIERGVKNWNKLFSEVLWKRKILPLVFVCGDGSILNFKKLKAYVRL SEQ ID NO: 203 MCP4121364__Bacteroidota_bacterium MMRLTFDWDSVGLEDIMVQKGLAKLSQDFPNNTIYYRLSASGTGIHAIISPKNSTPAPVEIEDEDALEYRREMV AFGLEDKYRLAHDEARVGTRLPTAQLWEWKNGKKAGEWIKYVE SEQ ID NO: 204 Ga0208683_1033688 MKEITLDWDWPCKKFIPDENRIIRQIRMHEKVKEIWWRCSANKHLHIKVVLTEDVDFWESIYLRSKWDDDANRI RMDIIRYWENMEYMRLWDEKLICGKRCVLRKAGPWRKI SEQ ID NO: 205 Ga0209450_1000786710 MSTTKKVVSCKKVIDRTSKLTRFTVDIDENFDEWRSSNSFLTIRKRITDEIADVWKAQIRRSSGGHVHIRIYLN HSISLFQSLCLRAFLDDDPHRLACDLDRFYRTGFIEDTGRCFDEKYTKGKLRLAGKWERFI SEQ ID NO: 206 Ga0208429_1003805 MITLDHDVKYEEYRNSAKGKLVLEFCRAFELKCWYRRSSSGNVHVAIDVDVDFLRELEIRAILLDDPMRILNDL RRHAFHVPTGRLWDAKVSRDGKKTAGEWEVLYVPE SEQ ID NO: 207 Ga0188176_1007518 MSFELYVDIDQDFDKFLESRTWEYIKENWKYVMHLYQYFSVMLYYRRSSSGHVHLMLKSVHQPSVLEQFQIRAL LHDDPWRIGIDLRRLTVHGKNEINRIFSMKVKNGVVYRVGPWIDITHMIEGLK SEQ ID NO: 208 Ga0188171_1015374 MSIELYVDIDQDFDKFLKSRTFGYIKENWKYVTDIIGKYIWSEHFDVFLYYRRSSSGNVHLKLEIPDNVEILDQ FQIRALLHDDPWRIGIDLRRLAIQGQSEINRIFEMKVKDGKVFRVGEWRDISSEVVK SEQ ID NO: 209 Ga0315296_100011799 MYAMVYRNSKGADIDMSLKPLFPNSSVGNTVTVDLDVWMEYFLRDQNEIPPDEKSGITIGYIQEQIVKLECKLE MPIIAHYRRSSKGHIHLRLLFTDEITVFDAFLLRSCLFDDLTRHSLDERRYALWGSLHEMNKCFDHKGEPGGKV YSSGPWISLNIGRDNLTGEALIDFQNYMKWRRDHPLKNQKEEGQTALGLHGLGS SEQ ID NO: 210 Ga0315295_100440255 MNCRHEIKEIYVDVDVKYKNYMSDGLPSAIIRGLLQWLPMRFKLYIRKSPNKRVHLKLVPEYSIPLFTTFQIRA LLHDDPFRIRQDLARFHLTGDVTKTGRIFDEKFCDGVIKKAGKWIELF SEQ ID NO: 211 Ga0334893_10071887 MEKNEQKTAVRSITIDLDYPYTEYIESNHFKMIQKTTSEWQDVRLREIRTSPSGRVHLRIWFDKPITLLASFCR RAILGDDPFRLACDLARLELYGDEKIGRTFDVKYVDGQQKRAGNWRIF SEQ ID NO: 212 Ga0334901_10052473 MRKKTRVEIDHDIPYDDYDTGYILRWIKKLPYDVLELYIRRSCGGNTHLALIVVGDLSPIDQMLIRAVMHDDKR RLRGDMERYLLDSPLFGLLFDAKYSSQTGLITRAGEWMEVTL SEQ ID NO: 213 Ga0334899_100372210 MRKKTRVEIDHDIPYEEYDTGYILRWIQKLPYDVLELYIRRSCGGNTHVALIVVGDLSPIDQMLIRAVMHDDKR RLRGDMERYLLDSPLFGLLFDVKYSSQTGMITRAGEWMEVTL SEQ ID NO: 214 Ga0334885_101513210 MRKKTRVEIDHDIPYDDYDTGYILRWIKKLPYDVLELYIRRSCGGNTHLALIVVGDLSPIDQMLIRAVMHDDKR RLRGDMERYLLDSPLFGLLFDVKYSSQTGLITRAGEWMEVTL SEQ ID NO: 215 Ga0326728_100744798 MTGLFQNKIIGRFCTVDVDTWFEYFLRDQKDTPKDEKPGITLGLIREQIVKLEAKLECPVNVVYRRSSKGHVHL RVTFPHDITILDAFLLRSCLFDDLVRQNLDERRYALWGSLDEMNKCFDSKADTEGIHKSGPWIPINKGRDDLAG DAKADWLQYWDSIGREYHKPENHINLVKSHWKELSAPQKNELLQELVRGIVESDLQQELATE SEQ ID NO: 216 Ga0326767_001510_2616_2945 MITLDHDIKYEEYRQSAKGQLVYEFCKTFHVRCWYRRSASGNVHVAIDMNVDFLRELEIRAILLDDPGRILNDL RRHAFHVPTGRLWDVKVDRQGKRQAGEWEILYSPE SEQ ID NO: 217 Ga0373184_0007613_3188_3532 MSNEVTLDWDIQHEDVRDSVSWRVLMKYMKDNEDGAIGADMRRSSSGNVHVRIRYSKELDAIKVYQIRALLRDD LYRLRLDLIRDYNKGETNRLWDFKIKNGEIHEAGSWEVII SEQ ID NO: 218 Ga0373184_0008556_2796_3155 MEPTNIITLDHDIEFDTYKESVTFDILKKWYLEWKQHIVGIWYRKSSNGNVHIKIEWVQPLDPLRWFELRAWLR DDVYRLRIDMCRDYQGDQVNRLWDEKYNSATQELKKAEDWQVLEF SEQ ID NO: 219 Ga0374055_0065819_5559_6191 MTGLFQNKIIGRFCTVDVDTWFEYFLRDQKDTPKDEKPGITLGLIREQIVKLEAKLECPVNVVYRRSSKGHVHL RVTFPHDITILDAFLLRSCLFDDLVRQNLDERRYALWGSLDEMNKCFDSKADTEGIHKSGPWIPINKGRGDLTG EAKADWLQYWDSIGGEYHKPENHINLVKSHWKELSAPQKNELLQELVRGIVESDLQQELATE SEQ ID NO: 220 Ga0373627_0013648_3727_4185 MNLFDNRSVGSQVTIDLDVFFTNYIRSPDQVPPDERIGINLFTTRDRIKILEKRVGMRIEAKYRRSSSGNVHVR LFFPCEVTVLDAFMIRAWMFDDQTRLALDMARYLKSNDLNEMNRCFDEKGDINGSKKAGPWIPIDEIPKVQEDP QRKL SEQ ID NO: 221 Ga0378408_0071754_317_856 MIPLFDNRSVGREIKVDLDVYYEAFLLTDKDLQEHFGREAINPDSPVYVGLNRQEMIYWRDRLVKERKVTPIKT EYRRSSNGHIHLKLVMSQEVSVLDGFMLRAWFLDDKTRLELDLKRYLLTNDLNEMNRCFDEKANADGMKKCGPW IPLDCEPKGTPEKSELISRLMQRINPQRTIT SEQ ID NO: 222 Ga0480576_007648_3290_3688 MDALADTKSDVIYVDIDQVWEAWSPRLWSLFLARCTQAQERAGFSIRKIQARKSSSGKIHAKIIFEDPVSIDVT FLVRAFLGDDPFRLAADMVRWVKYGKGAEVNRIFDTKFIDNEVRRAGPWIDLGLLARP SEQ ID NO: 223 Ga0484980_0000265_10787_11335 MIFNNQSIGKEVTIDLDTFFFYFVRPADEVPDGEKQGMTLQTVRDRLKRLCSIIGHDPIVAQYRRSSNDHAHVR LQFVPDLSVLDGFMIRAFMLDDQTRLELDLARYLLTGSLHEMNRCFDEKATVDGTKHSGPWISLQVDRDQYARD ALKDWSLYLPRWEIYKKNHPELFGKTATGQKELI SEQ ID NO: 224 MDD3967103__Ca_Marinimicrobia_bacterium MIPLFENRSVGWEIKVDLDVYYEAFLLTDKDLQEHFGREAINPDSPVYVGLNRQEMIYWRDRLVKQQGITPLKT EYRKSSNGHIHLKLVMSQEMSVLDGFMLRAWFLDDKTRLELDLKRYLMTNDLNEMNRCFDEKANADGLKKCGPW IPLDCEPKGIPEKSELISRLMQRINPQRTL SEQ ID NO: 225 MCK4329446__candidate_division_WOR-3_bacterium MMTDADILLDIDQKDVDAKEIAERLRFITMFTSFEFVRLVVYETRKGFHLYFWTSKPKPSAKDKVMIQLALGSD YRREIFNYLRICMTPLPKKWNVLFQSKYDKEGNKVSEEVHSGRTLKLEEEIMALYETKEVVV SEQ ID NO: 226 MDD3095778__Ca_Marinimicrobia_bacterium MIPLFDNRSVGREIKVDLDVYYEAFLLTDKDLQEHFGREAINPDSPVYVGLNRQEMIYWRDRLVKERGITPIKT EYRRSSNGHIHLKLVMSQEMSVLDGFMLRAWFLDDKTRLELDLKRYLLTNDLNEMNRCFDEKANADGMKKCGPW IPLDCEPKGIPEKSELISRLMQRINPQRTIT SEQ ID NO: 227 MCK4328908__candidate_division_WOR-3_bacterium MMTDADILLDIDKKDVDAKELAERLRFITMFTSFEFVRLVVYETRKGFHLYFWTDKPKPSAKDKVMIQLALGSD YRREIFNYLRICMTPLPKKWNVLFQSKYDKEGNKVSEEVHSGKTVKLEEEIMALYETKEVVV SEQ ID NO: 228 MCD6148170__bacterium MNKYTFLIDIDYKPTNPNKFAEEFTNKIKFVEDILKIFVEEVEVFETKKGIHIYVYASSERKISDEEIVVIQLA LGSDYKREIFNWSRVISNPKPKHWNVLFKSKERITKLSKMLTILINGKLDGLGKDL SEQ ID NO: 229 RLF07372__Thermoprotei_archaeon MRHVFKIDCDLKVPESWLKLWIKTRKAILRSLGLKPLEINVRNSERGFHVWIHTESEHRLTPTECNFVQWLLGD DQGRVIINQRRIARGMSWKEFNKLFSEVIWRTKHKCNCKIHQKILRQRKEGVVEFYKLIEKEEGGGNSLPHLYR AKRGSRKSKGKA SEQ ID NO: 230 HDL60173__candidate_division_WOR-3_bacterium MNRIGIDLDYANPPSVIELKRKILKEQKQNGLTQVLVFQTRHGYHLELIYDRDISAEENFQIREQYGDCKKRME YSKKRYELIGGGYDILFQMKEGVWRRRVWDEIESQNQK SEQ ID NO: 231 MBD3191301__Ca_Heimdallarchaeota_archaeon MTNLVKNTIIKLDLDGEESFELFLERGWICKYLGLRITKAEVFTTKNGYHVYLHTKNTLPYERILLIESLLGDD YRRVLYNLLRVVMGCTDFDVLFQEKWQLTRLGKVKVSSREEYHPALSEGLEMVLDQELTEKVNVLAVGGEARSV SEQ ID NO: 232 PSO06369__Ca_Marsarchaeota_G2_archaeon_BE_D MLIAQVIKRGWYIDYKVRFIGDVPTRYIPFIEKIGEEKYGTDNMCYIILRLRRMGTSDFRGGTGLRTVGSEEEE ATDSRINGGIGGSDFGSLRGSGAPLTDGRGGFLWWCVGKEDSVVSPLSTRWGRVYVVDQRQLIKLDIDVDVHFY DPNLILEKKLEMLRSLGYKEEDAWWEHSPSKKHIHILIVLTDPIPVKELFDLQFLLGDDPKRAEFNYLRYSVMG ENAIHFNVLYKYKRSLTFSDKFKAILRHWFRSKQYSKKRKEGQIT SEQ ID NO: 233 MCZ7405142__Ca_Methanoperedens_sp. MRVTLDIDNQKIERVRRKAILLKRAFPYAEVRYRISSSGKGGHVEMYERYGLVDLEMSYNIRRLLGDHAARIRI DLMRGKNDNPMLPLQVLFDYKVIDGVKKEAGKWCYVYE SEQ ID NO: 234 RLE65063__Thermoprotei_archaeon MVVLKIDIDIKVPEEWIDDWIGSRIAILNWYFDSDEVVQDIVVKPSSKRGYHAWIHLDAGDMPPGEANKLQWLC GDDETRVAINQRRIERGVPWKEANVLFSRVLKRKEYKNEQCENCGLRRAIESLFGECLR SEQ ID NO: 235 MBS7270390__Ca_Jordarchaeia_archaeon MESTLKIDIDIPELKKDKTLLQTWIQTRKAILEKLQIPIQKIRYTETQKGYHFWITINCQPTDKGICDLQFLLG DDQTRCRYNYLRLEAGCFKQFNVLFNKKLKNHNKKQTPLLLRLLKPLRNLPHPF SEQ ID NO: 236 MCL7344718__Ca_Aramenus_sulfurataquae MQKTKVIKLDLDLHHEFEERYHLLADFIKTREILLKALGYEVEEIVTETSQHGYHIIIILNNDITIEEAFRLQF LLGDDIHRVNFNFTRLALFGDETANYLNLLFTRKYKLRHRKA SEQ ID NO: 237 MCD6147782__bacterium MSLLKIDIDHKEPFEKWKDEWIETRKKLLEFLGFKVEYMEIYKSGSGRGYHVYIKIDKDIPDEEINKLQFLLGD DQTRVVIAMRRIERGVKYWNVLFSKVLWKRSDKEDLERALRLIEESNLNEFDKEWLKDYVEMIYNSLKKFTEAL KDVG SEQ ID NO: 238 MBX8640802__Ca_Sysuiplasma_jiujiangense MQYKELYIDIDQNFRQFKKSRTFEYIKENWKYVIDILLKYQGKEYLNILLPLWYRKSSSGNVHLRLIVPSEMSV LDQFKVRALLHDDPWRMGIDLRRLAIQGEEEINRIFEMKVKNGVVYKVGPWIDISEMMIQ SEQ ID NO: 239 MCM8802828__Ca_Omnitrophota_bacterium MKETVLKIDIDYKPNQYWLNKWIKTRKFILERMKCKVLRTNIFETKRGFHAYFVIDKKLNDNQLNMLQFLLGDD ATRVKINEWRIKRGIRRWNKLFSKILWRKKAKAVRCWYCGNVIPLRD SEQ ID NO: 240 RLF99547__Nitrososphaerota_archaeon MKLERSNVLKIDLDVKVSQKLLEKWLETRKLILEHLGYTITKIRYVETEKGYHFWIHLKENLEPKEVAELQFLL GDDHNRARYNFLRLKFRTFHEFNVLFNRKKRIERPQY SEQ ID NO: 241 TSA43629__archaeon MRRRHEKRILKTNVLKIDKDCHVSPELNLEYVRTILQTCRKYEPKVLWVKSSRSRHGMHFYIKIKPALEPHVAN NLQYLLGDDAKRVAFNRARIESGSDEWNKLFERANARLTTVYRNYYIRSAGSDIEIQRYRAPTE SEQ ID NO: 242 OYT41350__Ca_Aenigmarchaeota_archaeon_ex4484_224 MSLLKIDIDHKEPFDKWKKEWIETRKKILEFFGFKVEDIVIYESGSKRGYHIYIKIDKEIPDEEINKLQFLLGD DLTRVVINMRRIERGVGYWNVLFSKILRKRSDKEDLKKAINLIEKSNLNEYEKEWLKDYVEMLYRSIKKFTEVL K SEQ ID NO: 243 MCD6477621__Ca_Aenigmarchaeota_archaeon MNRIGIDLDYANLPSVIELKRKILKEQKQNGLTQVLVFQTRHGYHLELIYDRTISAEENFQIREQYGDCEKRME YSKKRYNLIGDSYDILFQMKEGVWRRRVWV SEQ ID NO: 244 MDG7037036__Nitrososphaerota_archaeon MRNNPLFNPIFVDIDVKYNAFMKSRLKFLIDANLKKIRGIDEVWIRSSSHGNVHLKLVMVKPVDFCMMMQIRAL CHDDHYRLGLDLRRYYLQGANEVNRIFDVKNEGVAGEWRRYEY SEQ ID NO: 245 RLI81048__Archaeoglobales_archaeon MITDADILVDIDKKDVSAEEIAERMFFVSVFTPFEFVRVRVKETAKGFHIYLWCADVKPSPTDKVVIQLILGSD YRRELFNYLRVCGRERAEKWNVLFATKYDGDGNRISRERTTAKSIQLEEEIFALYRTMSESESESESESEGA SEQ ID NO: 246 RLG54494__Ca_Korarchaeota_archaeon MKTLKLDLDGKNGLDVFLERAWIMKYMGLKVVAVRCSHTTNGYHLELDLDNEIDDIKAVFMQLALGSDYRREVC NLLRIERGCKDWNILFKRKFKINKLGQRVKVSEEKYDPELSQKILDILQLGE SEQ ID NO: 247 MCJ7425184__Ca_Bathyarchaeota_archaeon MRRKDVSPSVEKVSTLKIDKDCFVDRDWIENYIHVVLIPVCRSFEVTVNSIRMCNSRRRGLHFYIEISPPIDDD SANRLQWLLGDDSRRVDFNRARVESGLDEWNKLFEVPGRRLRAVYRDAKYEKRPM SEQ ID NO: 248 RLE63676__Thermoprotei_archaeon MNKYTFLIDIDYKPTNPNKFAEEFTNKIKFVEDILKVFVEEVEVFETRKGIHIYVYASSERKISDEEIVVIQLA LGSDYKREIFNWSRVISNPKPKHWNVLFKSKEKITKLSRMLTILINNKLDGLGKDL SEQ ID NO: 249 RLI99014__Ca_Aenigmarchaeota_archaeon MCSSWKDIIYYKPKKKATQTFKIDIDCHLNEEEIDLFVKTRKAILSSLGFTLIDYRYFQTERGMHFWFIAHGEL LDDKTKNLYQFLLGDDHHRFDINRRRIERGIPWQKANILFSRIIDRKRNYDCEFKLRTYKKYLEKIEGDKNDFI NYAIKVMLYD SEQ ID NO: 250 RLJ02943__Ca_Aenigmarchaeota_archaeon MGYRYVLKIDIDVPFLDKTALSTWEETRRVILKHLGVRPIGFKYARTKHGWHVWVDIDSDYPLNDYYLAFLQFL LGDDHRRATFNLARAEAGSFKVFNVLFSKKLRQKWPMERLIPYVLKLITAWSLFEVVKELEEDVEL SEQ ID NO: 251 RLJ07022__Ca_Aenigmarchaeota_archaeon MSENVLKIDVDLRIDKWGWIDSYKKFIIAGLRSLGYGVKKIIVKESDSKKGIHIWVHLDKKVDDRTKNMLQFLC CDDKTRVRINYYRIEAGIKNWNKLFSKVLYRKPLEPPCSECKLIKYLKEVEVDVDNEIRGEGKG SEQ ID NO: 252 RLF44718__Thermoplasmata_archaeon MRVGLDLDYTPLKDTLELKERILKEQKANGLTQILIFQTKHGYHLELIYNRPVAVEENFKLREKYGDCKKRREY SERRYNLIGNDYDILFQVKEGFWRKRVW SEQ ID NO: 253 MBP98872__Ca_Poribacteria_bacterium MGVLTFDWDDVVIDNDIVQQALSQLADSFGPERVWYRISSSGQGLHVLVGELDDSYHLRPIAVDSDDSFAWRSL FHDPPFELECGGRLRADNERQAHGFPVGRLFSHKDGLVAGEWQLYEVIP SEQ ID NO: 254 RKY01503__Spirochaetota_bacterium MIQRIGIDLDYTPLKDTLELKDRILKEQKANGLTQILVFQTKHGYHLELIYNRPVTVEENFRLREKYRDCKKRM EFSKKRYEIIKNNYDILFQIKEGFWRKRIWV SEQ ID NO: 255 MDD3091792__Methanoregulaceae_archaeon MRKKTRVEIDHDIPYEEYDTGYILRWIKKLPYDVLELYIRRSCGGNTHVALIVVGDLSPIDQMLIRAVMHDDKR RLRGDMERYLLDSPLFGLLFDAKYSSKTGIITRAGEWMEVTL SEQ ID NO: 256 NLM31067__Methanomicrobiales_archaeon MRVTIDIDSNRAFDVLRSWFTLRAITGKDPIGRVSSSGTGAHILVSGVPISQQTAIVIRRVCGDDYARIAFDEE SRGKPLQILFDEKHGRRAGIWHSDVENLIHELCGGRYLEGN SEQ ID NO: 257 MCG2864465__Vulcanisaeta_sp. MFWENNIYEVKIDIDLNKQDPCWVGVRLQKIYTFLSGICWYVGIYRTANGFHIRGYLREPISKNTDLVLRAVLN DDRERWGLDLWRATSGKRQFDILFTYRVKRYDKEHRVGGEVPITLEEALKIAEELCTKKTGLNLTS SEQ ID NO: 258 MCL7390232__Ca_Geocrenenecus_arthurdayi MNSKVLVDIDYKYEDRERMVKMVSDFAFKMYFVTDVLGVRFTGIKVYETSKGIHIYLDAESERPLTPLEIIIIQ LALGSDYKRELYNLRRARAWLDGEELENNWNTLFKYKYKDGKLVSKETVTEFSKFFEEYVMRKYNSLLRGESE SEQ ID NO: 259 MBP1449906__Thermoproteus_sp. MIYHVTADIDGITLPFPRERYEAWARERIRRLGNDITHFGYKLSSGSGIHVTAWLKNPISMDLVPIYQLAIGSD VKRECMNYFRFKYGVDINVLFQYKRANRIDWSVKVCGSERRFMGRICPGLLQLLELYKKELGLDLLQRLF SEQ ID NO: 260 MCD6436466__Clostridiales_bacterium MLKGLFKDKLKIDIDSKKFPIKEFFDKYNSLKELVDMEILEIKRSKNNGFHVKIVLNDNIDDVEKILLQMFLCS DKNREFMNYLRYKAGAKMSQWSFLFDEKYI SEQ ID NO: 261 NIA03798__Nitrospiraceae_bacterium MNRIGIDLDNVSLERAIKLKYIILNEEKKNGLTQILVFQTKHGYHLELIYNRDITVRENFMVRKEYGDCEYRQM FSKERYELLAGGYDILFYVKDNHWRRRVWA SEQ ID NO: 262 RLF40497__Thermoplasmata_archaeon MELRLDVDGKQKGSKRIKKVIELLGLKVRYWEIYKTNNGWHHYIGVDNKLTDLEVVLVQALMGSDFKRECFNYL RVKSGKFSYDDWNVLFKRKYEVDLVNGDVKLVSREIKVGVKL SEQ ID NO: 263 MCX6822129__Ca_Aenigmarchaeota_archaeon MQKNCVGIDKDSISDEEAKALEKKIIKEQMCQVIVYKTKRGFHFILIFDKDISKKENFEIRNKYGDCLERVRRS LLRSEMPGVPYDTLFSIKDGHWRKRVE SEQ ID NO: 264 MAY90185__Rickettsiales_bacterium MGVLTFDWDDVAIDNDIVQQALSQLAESFGPERVWYRVSSSGQGLHVLVGELDDSYHLRPIAVDSVDSFAWRSR FHDPPFELECGGRLRADNERQAHGFPVGRLFSHKDGLASGEWQLYEVIE SEQ ID NO: 265 HDD68945__Ca_Korarchaeota_archaeon MFDRVKRKVSYIGIDIDSPKVFDLLLIYMRARKLFPHSKIEVFISPSGRGFHLKIWKRCTILENIFYRAAVGDD PERLRLSIAKLFLNPDEQFFDILFDVKFKKKSRKIDLEKLLNGVNLSEKSIEEIREMAEKLEGKIKMKECWVTC IAFSGEEFRERLKVICEDIAIRDPSFKYKIYQSYFPKHDYVLVIFSDNKNQAYQRGEWFRRVIKKELGKDINYW VKIKKTK SEQ ID NO: 266 HCX21312__Cytophagales_bacterium MSGLTFDWDDVNFDNPKVQEALKHLCKIFDNKVWYRISSSGSGLHVIIAELSYDSLFGMILNPVVMPTTEQFEY RKQFAEPPWNLECPGRFNSDQVRSSEGFRTSRVFTSKNGNTAGGWMNVSMEIAEANEDE SEQ ID NO: 267 MCD6148449__bacterium MRVTIDIDGKGVLAKWKVALDFLVLKVLSRGNAYVRRSANNKFHLKAHGLPISFRTSLLIRALLGDDKMRIKFD LERKKKPKQILWSWKDGKYAGRWSKCLRSVLFVNTQ SEQ ID NO: 268 RLI98591__Ca_Aenigmarchaeota_archaeon MVFCSKEIKKLLDKKVDYIGFDIDGGNVSKLLRLYFALKKAFPRSKIDVYVSSSRRGFHVIVRKKVSVLENLYW RALLGDDNIRISLNLRKMFSNPNESFNDVLFDIKKGKHRVKINLEKILAKHSGLVKKYLEHKRWEDLIALSDLV RMELPVIKKWIVCMPFSEEKFFEIEEICESCGFDYSIFQSYYPDSDHLLVVFSKARDDAVRIGNFFKKELGLSF WVKEIY SEQ ID NO: 269 MBP98594__Ca_Poribacteria_bacterium MRLTFDWDYVGLEDKLVQDGLAKLSQDFPKYDVYYRISANGNGIHAIISPKDSTPTPIEMEDEDALDYRRKMVD FGLEDNWRLITDELRVDKGMPTSQLWEWKDGKQAGEWVKYVE SEQ ID NO: 270 MCK5223505__Ca_Calescamantes_bacterium MRLTFDWDSVGLEDIIVQDGLAKLSNDFPLYDVYYRISASGTGIHAIISPKHSEIPIPIEMENKEALDYRRQMV IFGLEDEWRLKGDIVRVERGLPTSQLWEWKDGKQAGEWTKYVE SEQ ID NO: 271 RKZ03316__Ca_Fermentibacteria_bacterium MKTDEPFSETLKLDIDIRDFRLVKKIFTQRCSFVLNVLKIWPMGLRVYSTKKGYHIYFDIKGVYTSFDICFLQL ALGSDYKREVFNFKRFSEELGKEWNVLFKEKYDAKGRLLSRECAEPSLSGELFEAVRDVVNYRHIQGGD SEQ ID NO: 272 PMP88411__Caldisphaera_sp. MKTEILIDIDKKSLKEFYERYIFVQKYLKFKLLGYEIAETKKGYHVRLIVDLPYEYSDKDIVLLQLLLGDDWKR ATINYFRVIHNLDDWNVLFRKKYRIFKAGNLFKLASKEKCIGCLHGDVS SEQ ID NO: 273 WP_287912851__Thermofilum_sp. MSNIVRVDIDIPFKVFMRLFKRPYENAVRLIANELGLKIESIRYYESEHGNTHIYITLDKELNNWDYIVLQFLF YNDAKKTFHNLRRLIGLGDPVDLMFHFIR SEQ ID NO: 274 MBP04108__Euryarchaeota_archaeon MSSGLTFDWDDTDISNPKVMKALSVLCTEFGPRVWVRTSSSGTGLHVLIGELTYDTLFGIVMCPVPMDLATQMM WRKVFSEKPWNLECVGRLFSDQVRSAEGFRTSRVFKSKNGMTSDKWISAETLDLPIVGELHEEE SEQ ID NO: 275 MCC5990101__Thermosphaera_sp. MAWHIMVWFTSSTSKSYYDYPYAVHFFTNRKLSCETALRRSIEIRDADIENLTSRFVGYEELPVNVPRDLYVAV EFKPTDFGSYNLSVIANTIPSDVAEEIAKYRMSYLTVDIDEYKPRVASSTLEELATAIDVAESMAEEVRVFRTR GGYHIRAKLKAPLGFEKLMELRQKAWDDPERVRIDYLYHEHGLTFLTNLLFNEKCTFENGKPTCVEERELPLNN ITVAREAFNFGFDPTYGQTLKLSVGQVEAVIGLCRVVLYGRPNVVTKELASRVKKEIERMYDNEETVATLSKAY GDAKTFTLRSVRTIDLGYIVYVITPKELMGSLIGRGGSKVKSAERTLGKRIVVVDRESRDAIEAYADLVRIAVR RALST SEQ ID NO: 276 MCC7570564__Ca_Micrarchaeota_archaeon MSEHRQAYRASQITLDIDNPTGVRILSVWYNMQIIGSVEGRISASGNGVHIRCTTKTQLTYAEMLLIRQQHADD ALRIWYDETAPESKPVMILFDSKHGQSADDWLSDPIQLIDRYGAML SEQ ID NO: 277 PMQ00718__Dictyoglomus_sp._NZ13-RE01 MKAQKVKSMRIVWNVLKVDIDKDFETMIKNKPHYEEYVNWILEKLGYKVEQILYRKSTNGKTHVYIFLNKAPES WKEYVFLLLACGDDIGRFNVNMMRLKMFGDPLIKFFAVKLRSKN SEQ ID NO: 278 HDJ96623__Ca_Aenigmarchaeota_archaeon MLTNRITIDIDGNFFVFVHSFIRCLILFSKVPRYRKTSKGYHLWIYLDRLITEKELYKYRLLLFDDRKRVKLDM RCRLKPKQVLFDEKKITEFDGKNVRVIKHQVFPCCVTTVKRTGE SEQ ID NO: 279 OGM08252__Ca_Woesebacteria_bacterium_RBG_13_36_22 MYDYKKETKKSLKEKANKNKNVTRFANARILLIDIDSEEDFRRWKMEIEQFEPILNFPKYKVEVSKGGLPHRHI TVYLKTPLDIWKRIALQFCLGSDLKRETMNCYRQLVGRAANIVFFEKKDE SEQ ID NO: 280 MDT7969454__Vulcanisaeta_sp. MAMFWENDIYEVKIDIDLNKQDPCWVGVRLQKIYTLLSGICWYVGIYRTANGFHIRGYLREPINRDTDLVLRAV LNDDRERWGLDLWRITSGQRQFDILFTYRVKRYDREHKVGSEVPITLEDALKVVEELCMSKRKD SEQ ID NO: 281 MDA9917593__Gammaproteobacteria_bacterium MDSGLTFDWDKVNIEDGLVNLELTELFEYFGEGKVWYRISSSGNGLHVIIADMYYDPPTNKILLRPIPMDSKIQ MDFRKKSLLECKGRFISDIHRIRNGLRTSRVFIVKNGSKSSDWKTYYPKNHKH SEQ ID NO: 282 RLI59238__Ca_Thorarchaeota_archaeon MEKINYIGIDIDSSDVFELLKRYILARKLFPDSEIITRISPSGHGYHMIIRLNKRITPFENFLYRALLDDDPYR LLLAMKKFAIDGERWYDLIFTHKYGNKSQELDLDRLLKNYDIEEIIKNWGEFGTMDKIEKISKGIKKKIGIKEV WMTCFAFKTNALREKLKGICEDISLKDESFKFKIYQNFQPEWDFILVIFSDSKDKAFQRGAWFLKNCLDEEDLG NVKKFGKDKYYWVKKRVEK SEQ ID NO: 283 MDE2104359__Patescibacteria_group_bacterium MSTGKIQPGRREQTLLAVRKGQVTVGDPAARRGNRSNKRARHCGSRAGFEQRQCVLRRWGRVEELANPRHLLLD YDRAKTPSLRDVYRVARIAGFGVEWVRHDRTRRGWHVVVRTDRKLLPAEQVALQAILGSDPRRELLNIMRVIAI RRRDPGVEWRCRWNMLFSGKLKEQGGNR SEQ ID NO: 284 MCP5004615__Planctomycetota_bacterium MTCIDNETTATNSERHPRRPQPPRGIRGTVAYPPLYIPAMADLETLTQTDVAEVTEYEDLDIEDVEKRAKDFNC SVYSAKEDELLLDIDNENDYRRLGDRITSLKDRDIIVHIRDQWRSKSNKWHVVMHLKDHILRSLPEKLLLQVCL GSDPKHVSVAYQRWLMGQEVYNLLFKPIRTTEGGNLRTAEQPVPF SEQ ID NO: 285 WP_174591886__Methanocella_conradii MSFIDLGKRGDELVDEIGVDIDTKNPLINLMVQVNASNLGVVHSYETRKGWHYRVKLKREITLKDAFKLRQYLG DDMYRMLIDAMRIRMGLTIYDVLFTAKKQLK SEQ ID NO: 286 NPA98176__Thermoproteota_archaeon MKLKLDIDFIDKDNLGVQYLKVVGWEIGRRLALIDAFLCISPEAIKVRETKHGFHIYIDTLTNFDVTPELIVAL QLLLKSDWKREMFNLRRILHDKKHNGKVSPDWNVLFMEKKALANDSKEERTELSDTIAEAIINGYEEQMELLID D SEQ ID NO: 287 MCG2892315__Vulcanisaeta_sp. MFWENDIHEVKIDIDLNKQDPCWVGVRLQKIYTFLSGICWYVGIYRTANGFHIRGYLREPINRDTDLVLRAVLN DDRERWGLDLWRMTSGQRQFDILFTYRIKRYDREHKVGSEVPITLEDALKVVEELCIEKKR SEQ ID NO: 288 MCP4122157__Bacteroidota_bacterium MYLTFDWDDVEQDEPLVQTALQELQHRFGVGCVWYRISSSGEGIHVIIANLSWNSNLGAMEITPKNFEDDFTLT VRKEFSEPPWGLECKGRLISDSVRTKNGFRTGRIFAAKNDKSAGEWLPYV SEQ ID NO: 289 PMP89296__Caldisphaera_sp. MSYEMYYVKRYCEKTTLRVYHRLPTLRLYIPFFISTKLHNALGETTAILLNGHDITNYVYYVYDKRRKQGRYVL NSYITAVLELTTAEEYTFCVDKIIPAKPQYRYEYYFSYEKENEESANDRKFDISFYSTSKLSFNETEERARQSV EDALGSDYWDVVIGINQGVAVPKWLGFETVPNVQWDLKEDTEVDFPYNILDLQGHYSPKAWCNLTNNMIKELLE YKLSMITVDYDNKNLTELETIIGNFLTKHPDLKWREIRIYETNKGYHVYIYLETIIKVRNLMSYRHELGDDLNR ILIDFYKIFQQKFPTFEEAMAFLTKNPYLNNPCAIFTINTLFKSKFSIKHEKGKIEYIKVSEESLVKLIKKH SEQ ID NO: 290 MCD6432010__Ca_Bathyarchaeota_archaeon MQKKGTATTEKIGPKSTLLIKLDYDYKDQRKIIEEMDLIRKMLGRMFNLKSFEVKETRRGYHVKLKIEPLINLD NKDIIIIQLLLGSDKYREFFNWIRVKGNQKHWNVLFDQKVEVKDIKTQKT SEQ ID NO: 291 MDA8718627__bacterium MITFKSVIQVMELDTMSSGLTFDWDDTDLSNPKVLKALSVLCTEFGPRVWVRTSSSGTGLHVLIGELTYDTLFG IVMSPVPMKLSEQMMWRKVFSEKPWNLECVGRLFSDQVRSAEGFRTSRVFKSKNGMTSANWISAETLDLPTGSE AYEEE SEQ ID NO: 292 HAI42127__Maribacter_sp. MQHTLLNRYFKEGDDMAEFSGLTFDWDEVSIDDVKVQKELQDLCNEFGEEYVWFRESSSKTGLHVMIAEIQLDP KTMDFIIVPLPMSTEEQMMYREKTDIECRGRFFSDLFRKKMGLRTSRVFSTKNGKQVGKWRRFK SEQ ID NO: 293 MCD6240054__Ca_Bathyarchaeota_archaeon MRVTVDKDYPSDVELLKTYYNMKYFLGKEPEIRVSKSGKGYHFIVRRLRISFETSLHLRRFFGDDETRVFLDEM GYGKPRQVLFLPLLRLTHPNFLLDPQDQPLKIPIHTVKKLLYSRKRHKYIGDGNGHTSTRTSEKS SEQ ID NO: 294 MBT4662373__Ca_Marinimicrobia_bacterium MDSGLTFDWDGVSIEEDLVNIELTELFDHFGEGKVWYRVSSSGKGLHIMIADMYYDPPTNKILLRPIAIDSKTQ MDFRKKSLLECKGRFISDIHRMKNGLRTSRVFVVKNGSRSSDWMGYYPNKI SEQ ID NO: 295 MDI6905679__Ca_Bathyarchaeia_archaeon MAPLTYLKIDYDLPFTLIPNWVKQAKVCTAEHILSLYNIKMETYSFTQSKNGNTHLTIWIKDEIPDETKAWLQW IIGDDCGRTYYNLKRIKKGIKNWNMLNRVHR SEQ ID NO: 296 OYD16924__candidate_division_WOR-3_bacterium_JGI_Cruoil_03_51_56 MNKDWRRRRIGVDLDRCTRLELLRVYFNACYYFPGADIEVKETSKGYHIRIHKQHRLEENLDVRRALLDDPVRI KFDEMRMVKPELHSWLDTLFRFKVKNGKIISREEPCNVLAEAFYTKLPCRKPSFLG SEQ ID NO: 297 MCA9560694__Myxococcales_bacterium MANGNRMKYGSLLDDNPFEIASRLGLDVVIPTEHELFVDIDDESDLLVLDSQIETLNRNLAGQWVDPGAKPAAI RTRTTRSRGGNCHAVVLMPFAVTPLERVLLQAALGSDRRRELLSYLRIKFATDRPPTLFFEVPSERRDQP SEQ ID NO: 298 PNV81826__Fervidicoccus_sp. MITPRDLLLIDIDFSYQNFMEYIYDFYMKHVSLILEKFNLKAESVVVKKSTNGNTHIMIKLDKEIDYETMLHVL LALGCDLGLVSVSLMRLKNFGDPLVKQFSRKLRVRDK SEQ ID NO: 299 MBO66741__Acidiferrobacteraceae_bacterium MLTFDWDDVGISHPRVQKALNRLIEEFGKWYVYVRMSSSTNGLHVVIAEKTYDEALGKTILTAIPLEPEQSQQW RTKFAEEPWLLECKGRLESDRPRAQVGLAVGRLFGQKNGDSCGPWVTAARALQEESVIQELQDEIL SEQ ID NO: 300 WP_304329813__Ca_Culexarchaeum_yellowstonense MSEDLVFTDEILLDFDNSHPGKALDVERIGYIARKLRLEFKEIKAYRTAKGFHVYMKLKEKVHPITAVLIQALM GSDYAREAYNAIRVYNLLAFPEKYPPYMWQAWNVLYKEKYVDGKLVSKEEYDEELTKKLWEVVKRHDEE SEQ ID NO: 301 MBT5170908__Ca_Nitrosopelagicus_sp. MIDVDRKVESIGIDIDGFDKPFLNKVQHNAAEYGNIQTTRSKNGHHIKIDLFVPTTLKNSLWIRFYLNDDPLRI LFDAMRIISDNDNIDVLWDEKQKKSINGIITQS SEQ ID NO: 302 RLC42946__Ca_Coatesbacteria_bacterium MRFCEDKYEVKVDLDIKKDESEVVKTAEEICRRMYFIEIYTPIRFGNVEVYETRRGFHLYIEVKEPAYLKKNKA FIVALQLLLMSDWKREVFNLSRVMSMFFLNVDYENWNILFYCKRNADGKYSTERRTYLSIMLEQILRSYETVGE TIFDNEVSNE SEQ ID NO: 303 MCK4329772__candidate_division_WOR-3_bacterium MNKQEIIEKARTAILKGEKEVDGYYIYTSKVDLDMKLDKNEIPLWEQTRQAILEFMYFVVIDINTSETERGLHT VINWLSEEGETDHSLNFIQMLLGDDVVRVKINERRIEKGISFERANILFSKILKRYPHEQDSRINIALERELGI IK SEQ ID NO: 304 WP_287913325__Thermofilum_sp. MKLIKLDYDGMSRKEIAEWIDFASEVVHIKIKRITMFPSRNGGTHVYVYTTNDVTDIQVFLFKYYANEDRRRVN ADIRRWKYRYPKDYSYLFYAYINKPKQSITLPNEFEFFIHVLAIANEHYPLG SEQ ID NO: 305 RLI48674__Ca_Bathyarchaeota_archaeon MSSPNGLTFDWDDVGLEDKTVQEALSWLNFNFGQGNVWYRLSSSGDGLHIIIGRMVIDPKTLHRYIEPIPMAAE DQISYRKKMAKDPWNLECRGRFISDTARKLGGRNTSRIFIVKNENISGDWNCWITETMV SEQ ID NO: 306 MBS7248391__Ca_Jordarchaeia_archaeon MRKRKFTLTKQELEELYWNQRLSTLEIAKKIGVHHSTVRYWMIKYGVERRHQKASMSRKEYVKKYYQEHKNELR KKNKEYKMLHKARLRKISVEKYKKEREKWLGEVPKITKEISRKAEVLAKSILEHEGFKEVTSFTRQSVFDFFAK KDGKLCGIDVTTAMRKPIRFQQIQMANYFDLKFYILFIKPDFTKYRIVEVPKSYLKKKIEYKSLNVHSLKDFKD IQTEIIRLDIDSSLLPLWEDDWRKAQTAIIESFGYLLKKTIIRQSGHPLPWKEGEKPAKGKGYHVWHHILVPHQ LDDLEKLKLQFLLGSDPGRCWINYLRIKRGINNWDKIFGYVLWRRPLPEQCQKCSLRVNLEKIAEEEEVVPDEG QP SEQ ID NO: 307 RLI87194__Archaeoglobales_archaeon MTKQDFWARGWPYEYTLKIDLDVPFLTEGDLYLWVETRIAILNRLNLLLDGWNYARTKHGWHFWFKIRAQRSLT DRELALLQLLLGDDHRRATFNLARAEAGSFKVFNVLFSKKLRKKWPMEKLILHVLRLIIAWSLFETVRELHEEV EL SEQ ID NO: 308 RLG16657__Ca_Pacearchaeota_archaeon MVRNGKKKILKKGSEFEIKIDYDKKIEGLDLIEKMVGRIFEIKGFEVSRSTNKGIHLRIRAKSLISIDEKDIII IQLLLGSDRYREFFNWLRVKGGLKRWNVLFKKKLG SEQ ID NO: 309 RLF34561__Thermoplasmata_archaeon MIVGLDWDGISLEDAKRRQEKIEQAENVQTILFETKHGFHLELIYPTPVSVEEGFKIREKYDDCKTRMVIGRKR FNVTGTGHDILFTMKNGFIRKRVW SEQ ID NO: 310 MDH5806834__Ca_Verstraetearchaeota_archaeon MLNEFKTYDLHHYLAIFRKEKLYSHDYECIHIFSPNKIPFDEVIEIGEREFFGEWRVVGRETFQISKPIERRYI ILLNLFYENFGRRLEVKGDKLPDEILKELLYIKKSRLTLDIDLLGIYDADFSLPELLVKIEEFLKKFCFQWKFY KSTRGFHLRASLIQPMEIFEIFKIRKMLYDNPERLWFDEFKIREGLFFLSDILFNEKLFKEDGRIGHYLEEEVD IYNEKFPYRTFITFGKNLINEIVFDTIKEKIKYELHNKLEKRDIRLIDMYENEFIIEVKLKHVKKIEKIIAFEL ADILHEEMKKGVEVNPYRILFYHPDTQIRVNSNLFKTTIIEKNGTTMILIKVPTALMERFCGKDDKYIISLSSI FNVPIIPIPY SEQ ID NO: 311 MBW8002588__Planctomycetota_bacterium MKKLKKLLRQTKTGLHEYIVRGDELVKDNPDNYIVPDANQILIDIDGEGQYTLFNERLEILEEFYEFEYSVKPS SSGVPHRHVSVIFRCEFTVPEKLFLQSFLASDHMRDIMSFVQFQAGDKIPILLRKVTDG SEQ ID NO: 312 RLG89972__Ca_Hecatellales_archaeon MRITLDIDGPAWKAWAAFYTHVSLSNKVEIYKTRTGFHVIGYGAPVETPEQVIRVRRWLGDDPVRIDLDEALVK AGKPFQILWTKKNDFQVKLLEVVENRNLD SEQ ID NO: 313 MBS3748192__Ca_Thermoplasmatota_archaeon MVCLDWDNISLDEAFKRINIIENKEHIQTLLFETKHGFHVYLFFKREISPFENFRIREKYWDCPLRLQYSKARF ETTGKDYDILFTMKSDFFEKLIRS SEQ ID NO: 314 MCX8169763__Ca_Methanomethyliaceae_archaeon MLSETYDVHHYVAIFRKEKLYSYDYDCIHIFSPDKMSLEEVISIGEREVFGEWRVVGRETLQLSKPLKKRYVIL LSFLYENFGRRMEVRGNTLPDDIIKELIHLKKSRLTLDINLLGVYDTEFSLPELLMKIEEFLKRFCFQWKFYKS TRGFHLRALLIQPMEILEILKIRKMFYDNPERVWFDEFKVKEGVAFLSDVLFNEKMFKDDGRVCYYAENEADIY NEKFPYRTFITLGKSLINDIVFDSIKEKIKHELREKLEKRDIRLIDMYENEFIIEVKLKDVKKIEKIIAYELAD IMHEEMKKGVEVNPSRILFYHPDNNIRFNSNLFKVKVLERDGTTLAIIKVPNSLMEKFVGKDGKYITSLSSLLN VKIIPEDLFKD SEQ ID NO: 315 MCS7098305__Ca_Methanomethyliaceae_archaeon MLSETYDVHHYVAIFRKEKLYSYDYDCIHIFSPDKMSLEEVISIGEREVFGEWRVVGRETLQLSKPLKKRYVIL LSFLYENFGRRMEVRGNTLPDDIIKELIHLKKSRLTLDINLLGVYDTEFSLPELLMKIEEFLKRFCFQWKFYKS TRGFHLRALLIQPMEILEILKIRKMFYDNPERVWFDEFKVKEGVAFLSDVLFNEKMFKNNGRICYYEEKEADIY NEKFPYRTFITLGRSLINDIVFDSIKEKIKHELRERLEKRDMRLIDMYENEFIVEVKLKDVRKIENIIAYELAD IIHEEMKKGVEVNPSRILFYHPDNNIRFNSNLFKVKVLERDGTTLAIVKVPNNLMEKFVGKDGKYISSLFSLLN VKIIPEAEEKNS SEQ ID NO: 316 MBO3801777__Ca_Brockarchaeota_archaeon MISIDIDEFNVFTLYKTIWKGLKLGLKFERAEISPSGNGFHVIFSDEVDDLENIIYRAILDDDPYRLRYSIKRY SMGGQVDICFSVKNKKKSKKIEEKLLNLEELKKAERIEEVIKLAENSEVNKLKPEFYITVIPFDGEELREKVAK IIEDIEAKDESFKASIYRSYLKDYYYVIVIKSPDKNQAYQRGEWFLKILKKEFQYDTFYWVKHS SEQ ID NO: 317 PWT76403__Bacteroidota_bacterium MKGLCYETIEISKIQDGVMTDYEVRNSPNSQRALIEAEQEGLLVVFPEPNQLQLDFDTEHQYNVYRELYPIIDK YYSIMEEQVTPSRSGLPHRHVTVTLGITLNNYQRIAIQACLGSDRVRELLSVIQEDNKDPHPTLFLEKKPAQLT EGEVPSELLLGE SEQ ID NO: 318 MCW4047126__Ca_Bathyarchaeota_archaeon MWHVMVYSQGLESPYSRPYAVHLFTREKPAPGEAFRLADEILTNAFPKWRREEDRLVYVGYENIPSVDADNFYY VATKLDVDNFEFKVEGIASTVPQQLASAIEEYRRKHVTLDYDDDEYTRARVIDALEYLTELGCRTRVFRTQRGY HVRAELPSPLSLEKILETREKLEEDYARISVDKAYLQKQLGFLTNLLFNSKCWVTANNTVECYEEKEVDPLSIT TIRVETISIPLPQMRIELPKGVVEIDERRIKFIGRFTEKDVKAITTSIEDNLWEYAYVRMREDSDIKSKLVKAY GKISSSLARIAEKCDVRVENGTVVIHVPEHLSEYVGRLIGKQGQNIRAVEGELGMRIRIERTSPPPEDVTIRKR LQELLKNLAEG SEQ ID NO: 319 HDD44451__Ca_Desulfofervidus_auxilii MQPEIDYITVDIDNLDFYRLFYAYLNAKEILKDNIKEIKTFWSPSGRGFHLKIFLKKPIPLIDNIIFRAVLLDD PHRIRYSLKKYFMGVYSGVDICFDYKNGKYEKPFEMPENLTKDNLEELAKEYEEKYKDRETKWEILEFNDDKLI DFFHEIGQIVKKYGGTYKIIRNPYRNISKWLFILYYD SEQ ID NO: 320 SCRFV1_ORF36 VESERAMQTKNKGSRIIKLDIDLLHIYEEKYHFLEDFIMTRKVILQSLGYKVKHIEIKKSNKGYHIWIEIDKFI PMIEILKLQFLLGDDINRVNYNMMRYNEWGEEISNLFNILFTKKYKNKKKNK SEQ ID NO: 321 SCRFV2_ORF35 MHVIDVRVPRSYVDVLAERPEDLPWWAVIFAREAGEVAMRDDIDRLPECEPSYICRSDSVLPINNERKRGKTTI AGCVGRPGCRIIWTTSNPYLARFWILVRGRNKKVCDNNMASREGRLTDRDVRDNTLKIDIDVDLEYYDPEKVLN DKVTMLKSLGYDVKDAYFLYSSSHHHIHIIITLTEYYPISTIFYLQFILGDDPDRANLNFYRLKYFPSVAKYFN VLFVKKEKITWRWKFHVLAKRWKKWLHN SEQ ID NO: 322 S7_428_Lipothrix_ITR_41 MESRDGVSEGRNVRDRVLKLDVDVPTEYYDPVRVLFDKLTMLKALGYEYSDATWKLSPSGHHIHIIITLNRDIP LDELFYLQFVLGDDPKRATFNFLRLTHFPGNAKYFNVLFEKKVKVSWWLKLKIMLKKVF SEQ ID NO: 323 S7_60_MTIV-like__ITR_19 MNELVGGNRKVEHKFRDPRESPNDVITLDLDYPITKFESSMYKALLEAFAESYPGVEAIYIRKSSSGNIHVSIH LKHPVSFWDRVVIRAWLHDDPSRIVNDIKRFTKGEPIERLFSTKGTDGVFRHAGEWKKIFPK SEQ ID NO: 324 Sohwa_mini-SBV MSDHKIYNVYVDWDEEFNDEQIKQRIKERLELLNQKPDRIYYRRSSNNHVHLNLVFDNGVDILEHFMIRAIFLD DPKRIRADLMRYYYTRDIKDINRIFDMKIHIEHDNNKILSIKFVSEWVQWTDER SEQ ID NO: 325 JGI20127J14776_100174810 MIKGTFDYITLDWDWDCKQFEPTTNKLFQQIISDPRVKRVWFRCSAHSKIHLRVELSEPVDFWESLYLRSIWDD DANRIRMDIIRYWNDGEVMRLWDAKIDCVENNEGKINCMILKAGEWRLIYQR SEQ ID NO: 326 Draft_1000754117 MVCGWCVEMNNTRVEIDHDIAYDEYDIQYVSRWIQKLPYKVLEMHIRRSCGGNTHVALVIDGDISQIDQMLIRA VMHDDTRRLRGDMERYLLDSPVFGLLFDIKYNSRSGTISEAGQWIRLI SEQ ID NO: 327 MLSBCLC_1001962422 MRLTDEIFLDIDVDYEEFIKNHFDELQRQIDVVEKQYNITTKAIRKSSSGRVHIALVIPGRKDLDIVNQFDLMS IRAFLGDDEARILADLRRYFISRDYGHVNRIFDHKFKKGKAWTAGPWVMF SEQ ID NO: 328 Ga0080005_12541116 MIKGTFDYITLDWDWDCKQFEPTTNKLFQQIISDPRVKRIWFRCSAHGKIHLRVELSEPVDFWESLYLRSIWND DANRIRMDIVRFWQDGEVMRLWDSKADCVVKEDGKIDCIFLKAGEWRLIYQR SEQ ID NO: 329 Ga0080003_100118612 MYSQLKFRLNPIYIDIDYMFDVYSFPLPFNAYRVSSKGIHLKLDICEPINLYEHFMIRLLYSDDFNRIKNDLRR LSKGLPEFNRIFKAKLKGGRWVKYGEWMNVRKSCIKKEEILNDSYEFSAEVKYHLLPLVLEILGQLNSLELFD SEQ ID NO: 330 Ga0081534_1005148 MFRVITLDWDWPCVEFNPATNKLIKSISEHPNVKEIWYRCSANGHIHVKVILKEPVDFWESIYLRSIWDDDANR IRMDIVRYWENGEVMRLWDAKLICGEKGCVLREAGPWVKINA SEQ ID NO: 331 Ga0081474_1307306 VITLDHDIPYDQYRQSAKGQLVLEFCRTFRVRCWYRRSASGNVHVSIDMSVDFLRELEIRAILLDDPGRILNDL RRHAFHVPTGRLWDVKVDRTGKRQAGEWEVLYSPPE SEQ ID NO: 332 Ga0116237_100511718 MTTVIRIDHDISFSRYDTAYVKDWIKKLPYTILSTEMRESTGGNTHIKVIIKENISPIDTLMLRAVLHDDSRRI RGDLERYLFESPLFDILFDAKRENGITRYAGEWIAI SEQ ID NO: 333 Ga0255987_1000663510 MRDSIDLDWDIMYETVRSSVSWRVLNQYLEKTIDKVVFCYMRRSSSGNTHIRLVFENELTEMFKYQIRALLRDD VYRMRLDLIRSYTSKETNRLWDWKIKEGEWKGASEWEKIWERKAIG SEQ ID NO: 334 Ga0209012_100464615 MKEIYLDYDLRYDSHLDSIIAARFAARISKIISKYGSVSIFKRKSSSGNTHYKLVFENDISVFDHFIIRAALCD DRDRVYIDLKRYFLNGEQEINRLFDGKVILSLTQVSASDTGNWIPVNLDDMIKAIKFYIKKHIDNGEEKYVQLL RDIP SEQ ID NO: 335 Ga0209012_10100439 VITLDHDIPYDQYLASAKGQLVVEFCRMFQVRCWYRKSASGNTHVVINLDVDFLRELEIRAILLDDAMRILNDL RRHAYHLPTGRLWDVKVSREGKKEAGPWEILYSPPE SEQ ID NO: 336 Ga0209605_100626214 MITLFDNQSIGKCICVDWDLYYSDFLKEGIPAAISLVRYLKDVILSGEEVRVYHQQSPGGNTHLAIKFDRPLSV LDGFMIRAWMGDDSQRLRLDMARYAKTGSLYEMNRCFQNKIKVTGGQGTLFNAGEWVGLDIPVADPRDIPTTSQ AHTMIIDLVRKREHENRGGS SEQ ID NO: 337 Ga0208448_10013312 MVKGTFNYITLDWDWDCRDFEPSSNRLFQQIISDSRVKRVWFRCSAHGKIHLLIQLSEPVDFWESLYLRAKWND DANRIRMDIIRYWEDGEIMRLWDAKIECVTKENGEINCMILKAGDWRLIYER SEQ ID NO: 338 Ga0208313_1013614 MITLDHDIKYEEYRQSAKGKLVLEFCRTFELRCWYRRSASGNTHVTIDLDVDFLRELEIRAVLLDDPMRILNDL RRHAFHMPTGRLWDVKVTREGTHKAGEWEILYSPPE SEQ ID NO: 339 Ga0208683_1010195 MITLDHDIKYEEYLSTAKGQLVVEFCKMFRIRCWYRRSASGNVHIAIDLDVDFLRELEIRAILLDDPGRIMNDL RRHAFHVPTGRLWDVKVDRAGKREAGQWEILYDPQS SEQ ID NO: 340 SBV1 polymerase coding sequence ATGGTGATGAACATTATCTACATTGATTGGGACTATGAGATTACGGAGGAAGAAGCGAAAGAACGGGTTAACAA GTATGTGAAAAGCCTGAAAATCAAACCGTTGTACATCTACTATCGCAAATCCAGTAATGGTCATGCACATCTGA AACTGGTGTTCATCAATGACATTGATGTCTTTGAGCACTTCATGATTCGTGCGATCTTTTACGATGATCCAAAA CGCATTCGTGCTGACTTAATGCGCTATTATCTGTTTCGTGATGAAACCAAAGTCAATCGCATTTTCACTGCCAA GATGAACAAAGAAGGCGTAAAGTTTGTTGGGGAATGGAAACTCATTTGGTCGTAA Table 1 : Characteristics of the PolARV sequences SEQ Group Group Protein-accession number Length, aa ID # NO: 75 1 MCL5439375__Ca_Thermoplasmatota_archaeon 119 76 1 MCL4336290__Ca_Thermoplasmatota_archaeon 116 78 1 MCK5386793__Gammaproteobacteria_bacterium 143 37 2 AFV2-like AFV2___MGYP000084726888 114 5 2 AFV2-like YP_001496975__Acidianus_filamentous_ 116 virus_2__gp50 11 2 AFV2-like YP_009817885__Sulfolobales_Beppu_filamentous_ 132 virus_3__HOU83_gp44 38 2 AFV2-like AFV2__MGYP000028702348 116 143 2 AFV2-like IMG__Ga0187308_1489313 116 142 2 AFV2-like IMG__Ga0167615_10012449 116 39 2 AFV2-like AFV2__MGYP000108320997 116 58 2 AFV2-like IMG__Ga0080005_1555837 116 40 2 AFV2-like AFV2__MGYP000580124568 116 164 2 AFV2-like IMG__Ga0080005_1369049 116 54 2 AFV2-like IMG__Ga0187310_149505 116 41 2 AFV2-like AFV2__MGYP000627308355 116 48 2 AFV2-like IMG__JGI20129J14369_10008762 116 161 2 AFV2-like IMGVR__JGI20129J51889_10012011 116 61 2 AFV2-like IMG__Ga0081473_1342719 116 62 2 AFV2-like IMG__Ga0079048_10004117 116 173 2 AFV2-like IMG__Ga0208447_1005623 116 52 2 AFV2-like IMG__Ga0105118_10000023 116 141 2 AFV2-like IMG__Ga0167616_100072111 116 152 2 AFV2-like IMG__Ga0392390_01197_9457_9807 116 167 2 AFV2-like IMG__Ga0081534_1024119 116 172 2 AFV2-like IMG__Ga0208447_1001963 116 56 2 AFV2-like IMG__JGI20127J14776_10019163 116 162 2 AFV2-like IMG__Ga0080005_1344263 116 60 2 AFV2-like IMG__Ga0081473_1309783 116 163 2 AFV2-like IMG__Ga0080005_1348823 116 57 2 AFV2-like IMG__Ga0080005_12223821 116 165 2 AFV2-like IMG__Ga0080005_1380401 116 176 2 AFV2-like IMG__Ga0326767_000714_3833_4183 116 145 2 AFV2-like IMG__Ga0187310_1274816 116 2 AFV2-like RLF38074__Thermoplasmata_archaeon 104 2 AFV2-like RLF44718__Thermoplasmata_archaeon 102 2 AFV2-like RKY01503__Spirochaetota_bacterium 105 2 AFV2-like NIA03798__Nitrospiraceae_bacterium 104 2 AFV2-like MCX6822129__Ca_Aenigmarchaeota_archaeon 101 2 AFV2-like RLF34561__Thermoplasmata_archaeon 98 2 AFV2-like MBS3748192__Ca_Thermoplasmatota_archaeon 98 2 AFV2-like MDH5806834__Ca_Verstraetearchaeota_archaeon 380 2 AFV2-like MCX8169763__Ca_Methanomethyliaceae_archaeo 381 n 2 AFV2-like MCS7098305__Ca_Methanomethyliaceae_archaeon 382 2 AFV2-like RLI98721__Ca_Aenigmarchaeota_archaeon 108 2 AFV2-like HDJ96623__Ca_Aenigmarchaeota_archaeon 118 2 AFV2-like MCW4047126__Ca_Bathyarchaeota_archaeon 381 3 SH1-like IMG__Ga0102674_10004311 184 3 SH1-like YP_009272856__Haloarcula_californiae_icosahedral 186 _virus_1__BGV91_gp36 3 SH1-like IMG__Ga0102676_1001754 165 3 SH1-like IMG__Ga0136447_100068140 168 3 SH1-like IMG__Ga0496842_00083_20984_21505 173 3 SH1-like IMG__Ga0224631_100397216 179 3 SH1-like IMG__Ga0496842_00289_5763_6239 158 3 SH1-like IMG__Ga0224631_10241806 165 3 SH1-like IMG__Ga0496842_01694_3578_4054 158 3 SH1-like IMG__Ga0496809_06171_4229_4705 158 3 SH1-like IMG__Ga0496803_05564_3452_3949 165 3 SH1-like IMG__Ga0496810_03341_6336_6818 160 3 SH1-like IMG__Ga0496806_05944_3404_3946 180 3 SH1-like IMG__Ga0102672_1005301 148 3 SH1-like IMG__Ga0256680_100317414 164 3 SH1-like IMG__Ga0496802_01417_11191_11682 163 3 SH1-like IMG__Ga0326458_10183535 213 3 SH1-like IMG__Ga0496842_04513_4513_5154 213 3 SH1-like IMG__Ga0102672_10015612 161 3 SH1-like YP_271898__Haloarcula_hispanica_virus_SH1__ORF 163 41 3 SH1-like YP_007761624__Haloarcula_hispanica_virus_PH1__ 163 gp35 3 SH1-like YP_005352816__Haloarcula_hispanica_icosahedral_ 157 virus_2__gp30 3 SH1-like SH1_MGYP000305394855 152 3 SH1-like IMG__Ga0496801_00511_5410_5898 162 3 SH1-like IMG__Ga0399984_004358_4291_4782 163 3 SH1-like SH1_MGYP000659110386 167 3 SH1-like IMG__Ga0102674_1000553 166 3 SH1-like IMG__Ga0102677_1001224 166 3 SH1-like IMG__Ga0102676_1002116 166 3 SH1-like IMG__Ga0256680_100253924 154 3 SH1-like IMG__ProkHueca_100083828 161 3 SH1-like IMG__Ga0102672_10003114 180 3 SH1-like IMG__Ga0224631_10016602 155 3 SH1-like IMG__Ga0102672_1002282 165 3 SH1-like IMG__Ga0496810_00105_3713_4180 155 3 SH1-like IMG__Ga0401358_0001079_16217_16594 125 3 SH1-like IMG__Ga0102942_100299915 119 3 SH1-like IMG__Ga0401378_005373_968_1327 119 3 SH1-like IMG__SL_4KL_010_BRINEDRAFT_1000424716 143 3 SH1-like IMG__Ga0496842_00441_8484_8903 139 3 SH1-like IMG__Ga0187336_10001376 135 3 SH1-like IMG__Ga0496812_03007_1093_1512 139 3 SH1-like IMG__Ga0496807_00343_7880_8299 139 3 SH1-like IMG__Ga0136589_100160111 154 3 SH1-like IMG__ADL20m3uS_00618990 154 3 SH1-like IMG__SL_4KL_010_BRINEDRAFT_1000010851 154 3 SH1-like SH1_MGYP000332456958 164 3 SH1-like SH1_MGYP000597192124 148 3 SH1-like MCZ7405142__Ca_Methanoperedens_sp. 112 3 SH1-like NLM31067__Methanomicrobiales_archaeon 115 3 SH1-like MCC7570564__Ca_Micrarchaeota_archaeon 120 4 MetaB MAR65923__Crocinitomicaceae_bacterium 129 4 MetaB HCX21312__Cytophagales_bacterium 133 4 MetaB MBP04108__Euryarchaeota_archaeon 138 4 MetaB MDA8718627__bacterium 153 4 MetaB MDA8718416__bacterium 124 4 MetaB MCP4124081__Bacteroidota_bacterium 124 4 MetaB MCP4122157__Bacteroidota_bacterium 124 4 MetaB MDA9917593__Gammaproteobacteria_bacterium 127 4 MetaB MBT4662373__Ca_Marinimicrobia_bacterium 125 4 MetaB HAI42127__Maribacter_sp. 138 4 MetaB RLI48674__Ca_Bathyarchaeota_archaeon 133 4 MetaB MBP98872__Ca_Poribacteria_bacterium 123 4 MetaB MAY90185__Rickettsiales_bacterium 123 4 MetaB MBO66741__Acidiferrobacteraceae_bacterium 140 4 MetaB MAR17766__Paracoccaceae_bacterium 130 4 MetaB MBT4661558__Ca_Marinimicrobia_bacterium 117 4 MetaB MCP4121364__Bacteroidota_bacterium 117 4 MetaB MBP98594__Ca_Poribacteria_bacterium 116 4 MetaB MCK5223505__Ca_Calescamantes_bacterium 117 4 MetaB RLG16669__Ca_Pacearchaeota_archaeon 142 5 SBV1-like SBV1_MGYP000276409847 152 5 SBV1-like YP_009408151__Metallosphaera_turreted_icosahe 109 dral_virus__c109 5 SBV1-like Ga0326767_001510_2616_2945 109 5 SBV1-like ASO67407__Metallosphaera_turreted_icosahedral_ 109 virus__c109 5 SBV1-like Ga0081474_1307306 110 5 SBV1-like Ga0208683_1010195 110 5 SBV1-like Ga0209012_10100439 110 5 SBV1-like WHA35244__Sulfolobus_polyhedral_virus_3__SPV3 127 _ORF13 5 SBV1-like WHA35156__Metallosphaera_turreted_icosahedral 134 _virus_3__MTIV3_ORF14 5 SBV1-like S7_60_MTIVlike__ITR_19 136 5 SBV1-like SBV1_MGYP000536778559 129 5 SBV1-like SBV1_MGYP000016906733 129 5 SBV1-like Ga0334901_10052473 116 5 SBV1-like Ga0334885_101513210 116 5 SBV1-like Ga0334899_100372210 116 5 SBV1-like MDD3091792__Methanoregulaceae_archaeon 116 5 SBV1-like Draft_1000754117 122 5 SBV1-like Ga0116237_100511718 110 5 SBV1-like SBV1_MGYP000108321946 147 5 SBV1-like Ga0080003_100118612 147 6 SIFV-like AFV2__MGYP000520397373 131 6 SIFV-like MGYP000496809245 108 6 SIFV-like MGYP000341894015 125 6 SIFV-like RLJ02943__Ca_Aenigmarchaeota_archaeon 140 6 SIFV-like RLI87194__Archaeoglobales_archaeon 150 6 SIFV-like MGYP000633814126 109 6 SIFV-like RLF99547__Nitrososphaerota_archaeon 111 6 SIFV-like MGYP000742917628 112 6 SIFV-like MGYP000508602690 111 6 SIFV-like MGYP000170857367 112 6 SIFV-like MBS7270390__Ca_Jordarchaeia_archaeon 128 6 SIFV-like MGYP000521756141 137 6 SIFV-like RLG58786__Ca_Geothermarchaeota_archaeon 114 6 SIFV-like MGYP000241629041 122 6 SIFV-like MGYP000088290850 127 6 SIFV-like MGYP000368427839 135 6 SIFV-like MGYP000262268821 140 6 SIFV-like MGYP000243964497 267 6 SIFV-like AOS58382__Sulfolobus_islandicus_filamentous_viru 265 s_2 6 SIFV-like NP_445695__Sulfolobus_islandicus_filamentous_vir 265 us__SIFV0030 6 SIFV-like MGYP000515246319 260 6 SIFV-like MGYP000417934545 267 6 SIFV-like MBN1357569__Ca_Bathyarchaeota_archaeon 126 6 SIFV-like MCJ7425184__Ca_Bathyarchaeota_archaeon 129 6 SIFV-like MDO8135069__Ca_Njordarchaeum_guaymaensis 124 6 SIFV-like TSA43629__archaeon 138 7 AFV1-like IMG_Ga0073352_14605 150 7 AFV1-like IMG_Ga0081529_1168807 140 7 AFV1-like IMG_Ga0081474_13087620 140 7 AFV1-like IMG_Ga0208429_1005046 144 7 AFV1-like IMG_Ga0208662_10058918 140 7 AFV1-like IMG_Ga0208312_1004087 144 7 AFV1-like IMG__Ga0208683_1020107 140 7 AFV1-like YP_003737__Captovirus_AFV1__AFV1_ORF144 144 7 AFV1-like IMG_Ga0080003_10016094 144 7 AFV1-like IMG_Ga0209120_100168315 144 7 AFV1-like IMG_Ga0187308_109244 144 7 AFV1-like WP_287912851__Thermofilum_sp. 103 7 AFV1-like PMQ00718__Dictyoglomus_sp._NZ13RE01 118 7 AFV1-like PNV81826__Fervidicoccus_sp. 111 8 MetaA MCW7078991__Ca_Methanoxibalbensis_ujae 162 8 MetaA RLC42946__Ca_Coatesbacteria_bacterium 158 8 MetaA MCK4329446__candidate_division_WOR3_bacteriu 136 m 8 MetaA MCK4328908__candidate_division_WOR3_bacteriu 136 m 8 MetaA RLI81048__Archaeoglobales_archaeon 146 8 MetaA RLG75071__Thermoprotei_archaeon 141 8 MetaA MCD6148170__bacterium 130 8 MetaA RLE63676__Thermoprotei_archaeon 130 8 MetaA MCD6432010__Ca_Bathyarchaeota_archaeon 124 8 MetaA RLG16657__Ca_Pacearchaeota_archaeon 109 8 MetaA MCL7390232__Ca_Geocrenenecus_arthurdayi 147 8 MetaA RKZ03316__Ca_Fermentibacteria_bacterium 143 8 MetaA PMP88411__Caldisphaera_sp. 123 8 MetaA MBD3191301__Ca_Heimdallarchaeota_archaeon 148 8 MetaA RLG54494__Ca_Korarchaeota_archaeon 126 8 MetaA MDA8161414__Desulfobacteraceae_bacterium 173 8 MetaA MDE2104359__Patescibacteria_group_bacterium 176 8 MetaA RLF40497__Thermoplasmata_archaeon 116 8 MetaA WP_304329813__Ca_Culexarchaeum_yellowstonen 143 se 8 MetaA MCA9560694__Myxococcales_bacterium 144 8 MetaA MCP4902934__bacterium 109 8 MetaA NPA98176__Thermoproteota_archaeon 149 8 MetaA MCC6021120__Thermoproteaceae_archaeon 143 8 MetaA MBP1449906__Thermoproteus_sp. 144 8 MetaA MCB7128533__Ca_Bathyanammoxibius_sp. 129 8 MetaA OGM08252__Ca_Woesebacteria_bacterium_RBG_1 124 3_36_22 8 MetaA MBW8002588__Planctomycetota_bacterium 133 8 MetaA PWT76403__Bacteroidota_bacterium 160 8 MetaA MCD6436466__Clostridiales_bacterium 104 8 MetaA MBS7248391__Ca_Jordarchaeia_archaeon 372 8 MetaA MCP5004615__Planctomycetota_bacterium 193 9 HAV1-like HAV1_MGYP000438579838 129 9 HAV1-like YP_003773414__Hyperthermophilic_Archaeal_Virus 141 _1__HAV1_gp16 9 HAV1-like HAV1_MGYP000751157063 131 9 HAV1-like HAV1_MGYP000016906087 126 9 HAV1-like WP_287913325__Thermofilum_sp. 126 9 HAV1-like HAV1_MGYP000181291329 141
Claims
CLAIMS 1. An isolated polymerase of archaeal viruses (PolARV) comprising a catalytic core domain having a structure comprising three alpha-helices (alpha1 to alpha3) packed against a beta-sheet consisting of three beta-strands (beta1 to beta3) and further having an amino acid sequence comprising the following conserved motifs: (a) DX1D / N in positions 9 to 11, wherein X1in position 10 is L, F, M, I, Y, V, K, W, H, C, or R; (b) H in position 48; (c) D, S, E, S, C or N in position 71; (d) D, C, N or H in position 72; (e) R, H, K or L in position 75; (f) R, K, A, L, G, Q, S, Y, C, H or T in position 82; and (g) F, Y, W, L, H, M, or N in position 96; the indicated positions being determined by alignment with Sulfolobales Beppu Virus 1 (SBV1) polymerase of SEQ ID NO: 1, and wherein the PolARV has nucleotidyl transferase activity.
2. The polymerase of claim 1, wherein: - the motif in (a) is DX1D; - the motif in (d) is D; - the motif in (e) is R, H or K; preferably R; and - the motif in (g) is F, W, Y or L; optionally, wherein X1 is L, F, I, Y, W or H; the motif in (c) is D or E; and / or the motif in (f) is R.
3. The polymerase of claim 1 or claim 2, wherein, the motifs in (a), (b), (d) (e), and optionally the motif in (f), are in the catalytic site of the polymerase catalytic core domain; preferably wherein the motif in (a) is in beta1, the motif in (b) is in beta3, the motif in (d) is in alpha 2 and the motif in (e) is in alpha3.
4. The polymerase of any one of claims 1 to 3, wherein the core structure of the polymerase catalytic core domain consists of a sequence of 90 to 100 amino acids.
5. The polymerase of any one of claim 1 to 4, wherein the catalytic core domain further comprises a C-terminal loop; wherein the motif in (g) is in the C-terminal region; and / or wherein the C-terminal region is a flexible loop that encloses the polymerase catalytic site.
6. The polymerase of any one of claims 1 to 5, which consists of the catalytic core domain; preferably a catalytic core domain comprising a C-terminal flexible loop.
7. The polymerase of any one of claims 1 to 6, which consists of a sequence from about 100 to about 300 amino acids; preferably from about 100 to about 150 amino acids.
8. The polymerase of any one of claims 1 to 7, which is from a group of archaeal viruses chosen from: Acidianus filamentous virus 1 (AFV1), Acidianus filamentous virus 2 (AFV2), Sulfolobus islandicus filamentous virus (SIFV), Hyperthermophilic Archaeal Virus 1 (HAV1), Sulfolobales Beppu virus 1 (SBV1) and Haloarcula hispanica virus SH1 (SH1) and / or which is from a group of PolARV chosen from: AFV1-like, AFV2- like, SIFV-like, HAV1-like, SBV1-like, SH1-like and MetaB groups.
9. The polymerase of any one of claims 1 to 8, which has an amino acid sequence which clusters with the sequences of one group of PolARVs chosen from AFV1-like, AFV2- like, SIFV-like, HAV1-like, SBV1-like, SH1-like and MetaB, based on pairwise sequence similarity.
10. A variant of the polymerase according to any one of claims 1 to 9, which has nucleotidyl transferase activity and which comprises at least one amino acid mutation; preferably one or more amino acid substitutions, at positions selected from the group consisting of S42, S43, N44, H46, I95, R94, F96, T97, K99, M100, N101, K102, E103, V105, F107, and optionally E13, D71 and R82, the indicated positions being determined by alignment with SEQ ID NO:
1.
11. The variant according to claim 10, which comprises at least one of the following substitutions: E13Q; H46N; E13Q and H46N; F96Y; M100A; M100K; N101A and V105A.
12. The polymerase or derived variant of any one of claims 1 to 11, which comprises an amino acid sequence having at least 85% identity with any one of SEQ ID NO: 1 to 339.
13. The polymerase or derived variant of claim 12, which comprises an amino acid sequence having at least 85% identity with any one of SEQ ID NO: 1 to 272, 274 to 283, 285 to 288, 290 to 297, 299 to 313 and 316 to 339; preferably SEQ ID NO: 1 or 3.
14. The polymerase or derived variant of any one of claims 1 to 13, which comprises an amino acid sequence having at least 85% identity with the sequence from positions 7 to 83 of SEQ ID NO:
1.
15. The polymerase or derived variant of any one of claims 1 to 14, which is a recombinant polymerase.
16. The polymerase or derived variant of any one of claims 1 to 15, which has terminal transferase activity on natural or modified nucleotides.
17. An expression vector for the recombinant production of the polymerase or derived variant of any one of claims 1 to 16 in a host cell, comprising a nucleic acid encoding said polymerase or derived variant.
18. A method for nucleic acid synthesis comprising incubating the polymerase, or derived variant of any one of claims 1 to 16 with at least one oligonucleotide primer and at least one nucleotide under conditions that allow incorporation of the nucleotide(s) on the primer; preferably wherein the nucleotide(s) comprise at least one modified nucleotide, in particular at least one 3’-OH blocked modified nucleotide and / or preferably wherein the method is for untemplated nucleic acid synthesis.
19. A kit for nucleic acid synthesis, comprising a polymerase or derived variant of any one of claims 1 to 16; and optionally, at least one nucleotide, reaction buffer, and / or oligonucleotide primer; preferably further comprising at least one modified nucleotide, in particular at least one 3’-OH blocked modified nucleotide and / or preferably wherein the kit is for untemplated nucleic acid synthesis.
Citation Information
Patent Citations
Modified nucleotides for polynucleotide sequencing
EP2607369B1
Thermostable DNA polymerase of the archaeal ampullavirus ABV and its applications
WO2007132358A2