Nucleic acid analysis method and device

The method of site-specific mutagenesis libraries with parallel sequencing and error detection addresses inefficiencies in nucleic acid variant analysis, ensuring accurate sequencing and analysis of properties like protein stability and function.

JP7721531B2Active Publication Date: 2025-08-12BASF SE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022537357
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-18
Filing Date
2020-12-17
Publication Date
2025-08-12
Estimated Expiration
2040-12-17

AI Technical Summary

Technical Problem

Existing methods for analyzing nucleic acid variants are time-consuming, laborious, and inefficient, particularly when using massively parallel sequencing, as they often lose the correlation between sequence and mutant clones, and cannot guarantee that sequences originate from specific clones.

Method used

A method involving site-specific mutagenesis libraries with unique mutation sites, allowing for the creation of a mixture of probe nucleic acids that are sequenced in parallel, using a sequencing device with error detection and a deconvolution program to identify specific library members based on defined mutation sites.

Benefits of technology

Enables fast and accurate sequencing of nucleic acid variants while maintaining a one-to-one correspondence between sequences and library members, allowing for efficient analysis of properties like protein stability and function without the need for individual sequencing reactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007721531000001
    Figure 0007721531000001
  • Figure 0007721531000002
    Figure 0007721531000002
  • Figure 0007721531000003
    Figure 0007721531000003
Patent Text Reader

Abstract

The present invention relates to materials and methods for nucleic acid and / or protein analysis, including materials and methods for generating and / or analyzing mutant libraries. In particular, the present invention relates to materials and methods for correlating the properties of mutants of a target nucleic acid with the respective sequences of the mutants. The present invention is particularly useful for analyzing saturation mutagenesis libraries and for analyzing sensitivity to chemical or physical conditions, such as temperature and solvent stability, and for analyzing production characteristics, such as yield or cellular localization.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to materials and methods for nucleic acid and / or protein analysis, including materials and methods for generating and / or analyzing mutant libraries. In particular, the present invention relates to materials and methods for correlating the properties of mutants of a target nucleic acid with the respective sequences of the mutants. The present invention is particularly useful for analyzing saturation mutagenesis libraries and for analyzing sensitivity to chemical or physical conditions, such as temperature and solvent stability, and for analyzing production characteristics, such as yield or cellular localization. [Background technology]

[0002] The properties of nucleic acids and proteins encoded by nucleic acids mainly depend on the sequence of the nucleic acid.To fully understand the properties that may be achieved by a selected target nucleic acid, it is necessary to create mutants of the target nucleic acid and analyze the mutants.Therefore, the creation of mutants has been practiced in the art for a long time.For the rapid creation of a large number of mutants, random mutagenesis and site-directed mutagenesis have been developed.

[0003] Basically, there are two methods for analyzing variants of target nucleic acid. In the first method, for all variants in a library or a subset of generated libraries that are determined to be interesting by the characteristics identified through characterization, the sequence of each variant is determined individually. This method is very time-consuming, laborious, and expensive. In particular, this method cannot take advantage of massively parallelized sequencing, which can reduce the workload and time required to sequence a large number of variants. The advantage of this method is that the very variants that are sequenced can still be used for further analysis, for example, further characterization analysis.

[0004] Another approach is described by Deng et al. (J Mol. Biol. 2012, 150-167; doi:10.1016 / j.jmb.2012.09.014). In this method, a population of mutants is first generated. The population is then subjected to selective pressure to increase the occupancy of mutants favorable to each selective pressure. Finally, nucleic acid segments containing the target nucleic acid are isolated from the entire population obtained after selection and sequenced in parallel. The sequences thus obtained are counted, and the frequency of the mutated amino acids is then correlated with the applied selective pressure. If each amino acid change is advantageous for overcoming those selective pressures, the frequency of a particular mutation is expected to be higher than average.

[0005] An advantage of such methods is that massively parallel sequencing techniques can be used to facilitate the analysis of mutation frequency patterns. However, this method necessarily implies that single mutant clones are not analyzed. Instead, only the mutation frequency pattern for the entire population is obtained. If the utilization of any specific mutant is desired, such mutant must be constructed de novo. By mixing unpredictably mutated nucleic acids (e.g., generated by using degenerate primers) in parallel sequencing, e.g., Illumina-type sequencing or 454-type sequencing, the correlation between the sequence and the mutant clone is lost. The sequencing device (sequencer) only returns sequences and cannot separate the nucleic acids of individual mutants. Such separation requires individual sequencing reactions, thereby eliminating the advantage of parallel sequencing. Furthermore, this method is limited to the analysis of properties that can be translated into selective pressure. For many properties of interest, such as pH stability or allergenicity of proteins, simple methods for applying selective pressure are not available. Also, because it cannot usually be guaranteed that all surviving members of a population contain a mutation in the target nucleic acid, many sequences obtained by parallel sequencing will be wild-type sequences and therefore will be of little or no informational value.

[0006] Vandewalle et al., Characterization of genome-wide ordered sequence-tagged Mycobacterium mutant libraries by Cartesian Pooling-Coordinate Sequencing, Nature Communications 2015, doi:10.1038 / ncomms8106, Baym et al., Rapid construction of a whole-genome transposon insertion collection for Shewanella oneidensis by Knockout Sudoku, Nature Communications 2016, doi:10.1038 / ncomms13270, or Dale et al., Comprehensive Functional Analysis of the Enterococcus faecalis Core Genome Using an Ordered, Sequence-Defined Collection of Insertional Mutations in Strain OG1RF, MSystems Methods for randomly mutagenizing the whole genome, for example by transposon integration to create a non-specific mutant library, are also known, as described by [End Page 2018] (doi:10.1128 / mSystems.00062-18). In such techniques, transposon integration is performed on the genome, and the resulting clones are pooled according to a pooling scheme to facilitate parallel sequencing. However, these techniques cannot guarantee that sequence reads originate from a specific clone, as several clones may contain the same mutation. Therefore, such methods are inefficient. Summary of the Invention [Problem to be solved by the invention]

[0007] The object of the present invention is to reduce or overcome the disadvantages of the prior art described above. In particular, it is an object of the present invention to provide materials and methods for analyzing variants of a target nucleic acid without sacrificing the variants in order to obtain the respective mutated nucleic acid sequences of the target nucleic acid. Further objects, particularly advantages, of the present invention are described below. [Means for solving the problem]

[0008] Accordingly, the present invention provides a method for analyzing a target nucleic acid, comprising the steps of: i) providing isolated members of a set of two or more site-specific mutagenesis libraries, each site-specific mutagenesis library member comprising a target nucleic acid mutated at one or more library-specific mutation sites, the mutation sites of the two or more site-specific mutagenesis libraries being different from one another; ii) selecting one member of each site-specific mutagenesis library; iii) obtaining, for each member, a probe nucleic acid of each mutated target nucleic acid that contains at least the nucleotide at each mutation site and adjacent nucleotides to identify each member; iv) mixing the probe nucleic acids into a mixture; and v) sequencing the probe nucleic acids of the mixture obtained in step iv) in parallel. The present invention provides a method comprising:

[0009] The present invention also provides i) generating a library of variants for each mutation site of a target nucleic acid, wherein the variants contain mutations of the target nucleic acid only at the mutation site(s) specific to each library; and ii) analyzing the library by the method according to the invention The present invention provides a method for site-saturation mutagenesis, comprising:

[0010] Furthermore, the present invention provides a) a constraint database containing definitions of permissible mutation sites in variants of the target nucleic acid; and b) a sequencing device for sequencing nucleic acids in parallel; and c) a monitoring program that indicates or suppresses output of erroneous sequences that do not comply with the constraint definitions in the constraint database when the sequencing device is used to sequence nucleic acids that comply with the constraint definitions; A sequence analysis device comprising:

[0011] The present invention also provides a) a definition database of library members, where the library members are variants of a target nucleic acid and the positions of the variants of the target nucleic acid are specific to each library; b) a sequencing device for sequencing nucleic acids in parallel; and c) a deconvolution program that identifies library members according to the library member definition based on the sequences obtained by the sequencing instrument. and optionally a sequence analysis device as described above.

[0012] The present invention also provides a monitoring program for the sequence analysis device described herein, and / or - a deconvolution program for the sequence analysis device described herein, and / or - a correlation program for the sequence analysis device described herein The present invention provides a sequence analysis apparatus computer program, The present invention also relates to the following: [Item 1] 1. A method for analyzing a target nucleic acid, comprising: i) providing isolated members of a set of two or more site-specific mutagenesis libraries, each site-specific mutagenesis library member comprising a target nucleic acid mutated at one or more library-specific mutation sites, the mutation sites of the two or more site-specific mutagenesis libraries being different from one another; ii) selecting one member of each site-specific mutagenesis library; iii) obtaining, for each member, a probe nucleic acid of each mutated target nucleic acid that contains at least the nucleotide at each mutation site and adjacent nucleotides to identify each member; iv) mixing the probe nucleic acids into a mixture; and v) sequencing the probe nucleic acids of the mixture obtained in step iv) in parallel. A method comprising: [Item 2] - repeating steps ii) and iii) to select at least one additional member of the one or more site-specific mutagenesis libraries; - in each round of repetition of steps ii) and iii), before step iv), the probe nucleic acid is labeled with a round-specific nucleic acid tag, The method according to item 1. [Item 3] 3. The method according to item 1 or 2, wherein for at least one member selected in step ii), a property other than its sequence that depends on the target nucleic acid is determined, preferably the same property is determined for each selected member. [Item 4] the property is a property of the target nucleic acid or a protein encoded by the target nucleic acid; - Expression, - secondary, tertiary or quaternary structure, folding efficiency, aggregation, multimerization, - temperature stability, pH stability, solvent stability, detergent stability, protease stability, binding stability, solubility, protease activity, storage stability, storage stability in detergents, residual activity after storage, stability against proteolysis, - kinetic parameters, changes in enzyme activity at low temperatures, preferably at temperatures of 30°C or less, in particular at temperatures below 20°C; - Substrate, substrate specificity, enantioselectivity, cofactor dependency, inhibitor specificity, changes in inhibition kinetics, - immunogenicity, toxicity, allergenicity, cleaning performance, - Location in cytoplasm, spore, cell membrane, organelle lumen or membrane, cell wall, excreta Item 4. The method according to item 3, wherein the method is selected from one or more of the following: [Item 5] (a) Each member: - only one mutation site, or - Two or more non-adjacent mutation sites wherein one or more mutation sites consist of one or more adjacent nucleotides preferably 1 to 9 nucleotides in length, even more preferably 2 to 9 nucleotides in length, even more preferably 2 to 6 nucleotides in length, even more preferably 3 to 9 nucleotides in length, even more preferably 3 to 6 nucleotides in length, even more preferably 3 to 4 nucleotides in length, even more preferably 3 nucleotides in length; and (b) the mutation sites in two or more libraries are - Partially overlapping, or - The method according to any one of items 1 to 4, wherein the mutation sites of all libraries are preferably non-overlapping, more preferably non-overlapping. [Item 6] The nucleotide at the mutation site of one or more libraries is (a) differs from the target nucleic acid sequence at one or more arbitrary nucleotides; and / or (b) differs from the target nucleic acid sequence by one or more degenerate nucleotides independently selected from [AT], [CG], [AC], [GT], [AG], [CT], [CGT], [AGT], [ACT], or [ACG]; or (c) selected from a predetermined set of nucleotides or nucleotide sequences, preferably comprising or consisting of a codon where the mutation site has been replaced with another codon selected from the predetermined set of codons; 6. The method according to any one of items 1 to 5. [Item 7] i) generating a library of mutants for each mutation site of a target nucleic acid, the mutants comprising mutations of the target nucleic acid only at each library-specific mutation site; and ii) analyzing the library by the method according to any one of items 1 to 6 site-saturation mutagenesis, including [Item 8] a) a constraint database containing definitions of permissible mutation sites in variants of the target nucleic acid; and b) a sequencer for sequencing nucleic acids in parallel; and c) a monitoring program that indicates or suppresses output of erroneous sequences that do not comply with the constraint definitions in the constraint database when the sequencer is used to determine the sequence of a nucleic acid that complies with the constraint definitions; A sequence analysis device comprising: [Item 9] a) a definition database of library members, where the library members are variants of a target nucleic acid and the positions of the variants of the target nucleic acid are specific to each library; b) a sequencer for sequencing nucleic acids in parallel; and c) a deconvolution program that identifies library members according to the library member definition based on the sequences obtained by the sequencer; and optionally the sequence analysis device according to item 8. [Item 10] - a property database containing, for at least one library member, a property characteristic other than its sequence that depends on the target nucleic acid; and - A correlation program that generates paired assignments between library member sequences and characteristic properties of each library member. Item 10. The sequence analysis device according to Item 9, further comprising: [Item 11] - a monitoring program for a sequence analyzer according to item 7, and / or - a deconvolution program for a sequence analysis apparatus according to item 8, and / or - a correlation program for a sequence analysis apparatus according to item 9. A sequence analysis device computer program. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 is a schematic diagram of a method for analyzing a target nucleic acid using the mutant library of Example 1 described herein. [Figure 2A] FIG. 1 is a schematic diagram of a method for analyzing a target nucleic acid using the mutant library of Example 2 described herein. [Figure 2B] Continuation of Figure 2A. DETAILED DESCRIPTION OF THE INVENTION

[0014] The present invention provides a method for analyzing target nucleic acids using a mutant library. A library according to the present invention is a mixture of microorganisms or other vessels for nucleic acids, containing one or more mutation sites of a target nucleic acid. Each mutated target nucleic acid vessel, particularly its microbial clone, is referred to as a "member" of the library. It is preferred, but not essential, that the library consist of or consist essentially of mutants of the target nucleic acid. Preferably, the proportion of unmutated target nucleic acid members in the library is less than 20%, and even more preferably less than 10%, of the total library. The library can be implemented as a mixture of clones in a microbial host strain, which is not limited to a particular type of microorganism and can be selected according to the needs of one skilled in the art. Preferably, the microbial host strain is selected from any of the microorganisms of the genera Escherichia, Bacillus, Corynebacterium, Lactobacillus, Lactococcus, Saccharomyces, Pichia, or Yarrowia.

[0015] The target nucleic acid can be any nucleic acid, as long as the host strain can survive after mutation of the target nucleic acid, for example, by supplementing the growth medium with media components required for possible mutation of the target nucleic acid. For example, if the target nucleic acid encodes an enzyme essential for the metabolism of the microbial host, the host strain should be maintained in a medium supplemented with the respective substance that can no longer be synthesized by the metabolism of the microbial host on its own.

[0016] In particular, the target nucleic acid may be a transcription effector nucleic acid, such as a promoter, a terminator, an interfering nucleic acid (e.g., RNAi), a guide nucleic acid (e.g., a single guide RNA or crRNA and tracer RNA for CRISPR applications), or a transcription factor binding site. The target nucleic acid may also be a tRNA or rRNA. It is particularly preferred that the target nucleic acid encodes a protein. The protein may be of any type or length, for example, an enzyme.

[0017] Each mutation site in the library can be one or more nucleotides, preferably one or more contiguous nucleotides. Particularly preferred mutation site definitions are described below.

[0018] A method for analyzing a target nucleic acid according to the present invention includes providing isolated members of a set of two or more site-specific mutagenesis libraries, each of which contains a target nucleic acid mutated at one or more mutation sites, the mutation sites being different from one another. Thus, the target nucleic acid is mutated according to a predetermined pattern of mutation sites, such that the mutation sites are specific to each library of target nucleic acid variants, i.e., each library contains mutations at known positions. For example, if the target nucleic acid encodes a protein, one library may contain target nucleic acid variants mutated only at position 1 of the peptide sequence, while another library contains variants mutated only at position 2 of the amino acid sequence of the protein encoded by the target nucleic acid. As discussed below, it is preferred, but not essential, that the mutation sites in each library not overlap, so long as the mutation sites are specific to each library. Providing isolated members of a set of two or more site-specific mutagenesis libraries as defined herein allows for accurate assignment of every mutant sequence of the target nucleic acid to a single member of a library. In this way, the present invention can simultaneously utilize the advantages of individual sequencing, i.e., the ability to maintain sequenced clones in an isolated form, and the advantages of massively parallel sequencing, particularly with regard to the speed with which many sequences can be obtained without the need for the time-consuming and error-prone preparation of many individual sequencing reactions.

[0019] The method for analyzing target nucleic acids using mutant libraries further includes the step of selecting one member from each site-specific mutagenesis library. A typical method for selecting library members in the art involves inoculating microbial hosts that constitute the libraries and selecting one of the resulting colonies. By selecting only one member per library, the present invention avoids multiple reads of the same sequence. Because each library is characterized by a limited number of nucleotides at which mutations can occur, many clones with identical mutations may be detected if the target nucleic acid regions of all library members are sequenced in parallel. Furthermore, as described above with respect to population analysis techniques, it may be impossible to maintain a correspondence between a specific member of the library and the specific sequence of the mutated target nucleic acid. The method of the present invention ensures that the selected colonies contain the same genetic information. Such selected colonies from one generated library can be mixed with one or more colonies from other generated libraries with non-overlapping mutations, still allowing for identification of the recovered mutations.

[0020] The method for analyzing a target nucleic acid using a mutant library according to the present invention further includes obtaining, for each member, a probe nucleic acid for each mutated target nucleic acid, containing at least the nucleotides at each mutation site and adjacent nucleotides for identifying each member. In this way, sequencing of the nucleic acids of selected members can be limited to a desired segment with high information content. In particular, when length is a critical factor in parallel sequencing, the size of the probe nucleic acid can be adjusted accordingly. Thus, the probe nucleic acid can comprise or consist of the entire length of the target nucleic acid. In this case, the probe nucleic acid necessarily contains nucleotides at one or more library-specific mutation sites and all adjacent nucleotides. Therefore, since only one of the selected members may differ from the target nucleic acid sequence at the mutation site that characterizes the library of selected members, the selected members can be identified by the sequence of the probe nucleic acid. Since only one member per library is selected, identifying the correct library is essentially the same as identifying the selected members of the library. The probe nucleic acid may also be shorter than the length of the target nucleic acid, as long as the library-specific mutation site and adjacent nucleotides are present in the probe nucleic acid, allowing for identification of each member. For example, if a protein contains short amino acid sequence repeats and the mutation site(s) are adjacent to or contained within one of the repeats, the probe nucleic acid must contain additional flanking nucleotides that define each repeat of the amino acid sequence motive. Preferably, each probe nucleic acid has a length of at least 15 nucleotides, even more preferably at least 30 nucleotides, even more preferably at least 40 nucleotides, even more preferably at least 50 nucleotides, even more preferably at least 60 nucleotides, even more preferably at least 75 nucleotides, even more preferably at least 90 nucleotides, even more preferably at least 120 nucleotides, and / or comprises the full length of the target sequence.

[0021] The method of analyzing target nucleic acids using a mutant library according to the present invention further comprises the step of mixing probe nucleic acids into a single mixture. In this way, the method according to the present invention can avoid the need to individually sequence the probe nucleic acids and instead use the advantages offered by parallel sequencing methods.

[0022] In a further step, the method for analyzing target nucleic acids using the mutant library of the present invention includes sequencing all probe nucleic acids in the mixture obtained in step iv) in parallel. The method of the present invention is not limited to a specific type of parallel sequencing. Preferably, the parallel sequencing is next-generation sequencing, such as Illumina-type sequencing-by-synthesis, 454-type pyrosequencing, nanopore sequencing, or SMRT sequencing. A particular advantage of the present invention is that the methods described herein allow for the use of parallel sequencing methods that are generally independent of the library members used to construct the mixture to be sequenced, while maintaining a one-to-one correspondence between every sequence obtained by parallel sequencing and each member of all libraries selected in step ii). Thus, the methods of the present invention allow for the use of particularly fast and reliable sequencing methods that typically have a low cost per sequenced nucleotide. Furthermore, as described above, the methods of the present invention leave the production of sequences, as well as the library members themselves, to the discretion of the skilled artisan. In the example described above, the skilled artisan can again select selected colonies of the site-mutagenesis library, grow the microbial hosts thus obtained, and subject them to further analysis.

[0023] A further advantage of the methods of the present invention is that essentially any method for generating variants is applicable, as long as the induced mutations are restricted to the mutation sites of the respective libraries. Examples of such site-directed mutagenesis techniques are described in WO201926211 and WO2009152336.

[0024] Furthermore, the method of the present invention advantageously allows for sequencing error detection during or after sequencing of the probe nucleic acid. Due to the definition of the library in which mutations are limited to library-specific sites, the sequence obtained in step v) can only deviate from the sequence of the target nucleic acid in a very restricted manner. For example, if a sequence containing mutations at two sites, at least one of which is not provided according to the mutation site definition of a single library, is detected, this sequence is considered to be a sequence containing an error. Therefore, if the correct sequence at a particular position is in doubt, it is also possible to correct the sequence. Because only certain nucleotides are allowed for each probe nucleic acid at each position, if the sequencing reaction shows two or more alternative readings at a particular position, it is possible to unambiguously select the correct option.

[0025] In the method of analyzing a target nucleic acid using a mutant library according to the present invention, steps ii) and iii) are preferably repeated to select at least one additional member of one or more site-specific mutagenesis libraries, and in each round of the repetition of steps ii) and iii), the probe nucleic acid is labeled with a round-specific nucleic acid tag before step iv). A unique advantage of the method of the present invention is that it does not rely on tagging each member of each library to be sequenced. Instead, tags must be added to the probe nucleic acids according to the total number of selected members per library only if two or more members per library are analyzed. Preferably, a skilled artisan first selects one member of each site-specific mutagenesis library, obtains its corresponding probe nucleic acid, and then combines these probe nucleic acids to form a first mixture. The same nucleic acid tag is then attached to all probe nucleic acids in this first mixture. In the next round, a skilled artisan again selects one member per library, obtains each probe nucleic acid, and combines them to form a second mixture. A tag nucleic acid different from that attached in the first round is then attached to all probe nucleic acids in the second mixture. The repetition of member selection, obtaining probe nucleic acids, and binding of round-specific tags can then be repeated as necessary. The tagged mixtures from each round of repetition are then combined into one mixture as described in step iv) of the method according to the present invention. Optionally, tags are not bound in the first round, i.e., nucleic acid tags are bound only to the probe nucleic acids of the second and further members selected from the library.

[0026] Thus, the present invention allows for an increase in the number of analyzed members of each library with each additional reaction per iteration. It is therefore a particular advantage that the methods of the present invention facilitate the simultaneous sequencing of multiple target nucleic acid variants without compromising the one-to-one relationship between selected members and sequences.

[0027] The concentration of probe nucleic acids in the mixture(s) for sequencing is preferably selected to be high enough to achieve oversampling, i.e., two or more copies of each probe nucleic acid (optionally comprising a tag as described above) are read during sequencing step v). Preferably, each probe nucleic acid is sequenced at least three times (3-fold oversampling), even more preferably at least five times (5-fold oversampling), and even more preferably at least 10 times (10-fold oversampling) in step v).

[0028] Preferably, a total of up to 100,000 probe nucleic acids, optionally including the tags described above, more preferably at least 5 different nucleic acids, even more preferably 5 to 70,000 nucleic acids, even more preferably at least 25 different nucleic acids, even more preferably 25 to 65,000 nucleic acids, even more preferably at least 50 different nucleic acids, and even more preferably 50 to 10,000 nucleic acids, are sequenced in parallel in step v). The number of nucleic acids to be sequenced in parallel is selected depending on the performance of the sequencing method used in step v), the required liquid volume of each sample, and the volume of the screening plate.

[0029] In the method of analyzing a target nucleic acid using a mutant library of the present invention, a target nucleic acid-dependent characteristic is preferably determined for at least one member selected in step ii), and preferably the same characteristic is determined for each selected member. Because it is pointless to resequence a nucleic acid whose sequence is already known, the characteristic can be any characteristic other than the mutated sequence of the target nucleic acid. A particular advantage of the method of the present invention is that any measurable characteristic can be analyzed for library members, as long as the characteristic does not interfere with the creation of the library. For example, if the host strain is grown in a medium that compensates for any loss of function caused by mutation of an essential target nucleic acid, mutants of nucleic acids important for the survival of the host strain can be generated and analyzed. In particular, the method of the present invention is not limited to analyzing the effect of mutations on the survival of the host strain under selection conditions. Therefore, preferably, the characteristic analyzed is not antibiotic resistance.

[0030] Preferably, the property is a property of the target nucleic acid itself, and even more preferably, a property of the protein encoded by the target nucleic acid. The property of the target nucleic acid itself typically requires that the target nucleic acid be operably linked to a reporter gene. For example, if the target nucleic acid is a promoter, the target nucleic acid may be fused to a reporter gene that enhances the expression strength of the promoter under selected culture conditions. Furthermore, two or more properties may be analyzed in parallel in the same or different analytical reactions.

[0031] Preferably, the property is - Expression, - secondary, tertiary or quaternary structure, folding efficiency, aggregation, multimerization, - temperature stability, pH stability, solvent stability, detergent stability, protease stability, binding stability, solubility, protease activity, storage stability, storage stability in detergents, residual activity after storage, stability against proteolysis, - changes in kinetic parameters, changes in enzyme activity at low temperatures, preferably at temperatures of 30°C or less, in particular at temperatures below 20°C; - Substrate, substrate specificity, enantioselectivity, cofactor dependency, inhibitor specificity, changes in inhibition kinetics, - immunogenicity, toxicity, allergenicity, cleaning performance, - Location in the cytoplasm, spore, cell membrane, organelle lumen or membrane, cell wall, excretion is selected from one or more of:

[0032] The list of preferred properties described above demonstrates the versatility of the method for analyzing target nucleic acids using the mutant libraries of the present invention, and thus the method is essentially applicable to any mutational analysis of the function of any protein, for example, in a host cell.

[0033] Particularly preferred properties are described below by providing appropriate definitions.

[0034] "Enzyme properties" include, but are not limited to, catalytic activity per se, substrate / cofactor specificity, product specificity, increased stability over time, thermal stability, pH stability, chemical stability, and improved stability under storage conditions.

[0035] The term "substrate specificity" reflects the range of substrates that can be catalytically converted by an enzyme.

[0036] "Enzyme activity" refers to at least one catalytic effect exerted by an enzyme. In one embodiment, enzyme activity is expressed as units per milligram of enzyme (specific activity) or substrate molecules transformed per minute per enzyme molecule (molecular activity). Enzyme activity can be defined by the actual function of the enzyme, such as the proteolytic activity of proteases by catalyzing the hydrolytic cleavage of peptide bonds, or the lipolytic activity of lipases by hydrolytic cleavage of ester bonds.

[0037] Enzyme activity may change during storage of the enzyme or during use in a process. The term "enzyme stability" relates to the retention of enzyme activity as a function of time during storage or operation.

[0038] To determine and quantify the change over time in the catalytic activity of an enzyme stored or used under certain conditions, the "initial enzyme activity" is measured under defined conditions at time 0 (100%) and at a specific time point thereafter (x%). By comparing the measured values, the extent of potential loss of enzyme activity can be determined. The extent of enzyme activity loss determines enzyme stability or instability.

[0039] Parameters that influence the enzymatic activity and / or storage and / or operational stability of the enzyme are, for example, pH, temperature and the presence of oxidizing substances.

[0040] "pH stability" refers to the ability of a protein to function at a particular pH. Generally, most enzymes function under conditions of slightly higher or lower pH. A substantial change in pH stability is evidenced by a change (lengthening or shortening) in the half-life of the enzyme activity of at least about 5% or more compared to the enzyme activity at the enzyme's optimum pH.

[0041] The terms "thermostable" and "thermostable" refer to the ability of a protein to function at a particular temperature. Generally, most enzymes have a limited temperature range in which they function. In addition to enzymes that work at mid-range temperatures (e.g., room temperature), there are also enzymes that can work at very high or very low temperatures. Thermostability can be characterized by what is known as the T50 value (also called half-life, see above). T50 indicates the temperature at which, after a certain period of heat inactivation, there is still 50% residual activity compared to a reference sample that has not been subjected to heat treatment. A substantial change in thermostability is evidenced by a change (lengthening or shortening) of the half-life of the enzyme activity by at least about 5% or more when exposed to a given temperature.

[0042] The terms "thermostable" and "thermostable" refer to the ability of a protein to function after exposure to a particular temperature, e.g., very high or very low. A thermotolerant protein may not function at the temperature of exposure, but will function upon return to a favorable temperature.

[0043] "Oxidative stability" refers to the ability of a protein to function under oxidizing conditions, particularly in the presence of various concentrations of HO, peracids, and other oxidizing agents. A substantial change in oxidative stability is evidenced by a change (lengthening or shortening) in the half-life of the enzyme activity by at least about 5% or more compared to the enzyme activity present in the absence of the oxidizing compound.

[0044] "Proteolytic stability" refers to the ability of a protein to resist proteolysis. Proteolysis is the breakdown of proteins (e.g., enzymes) into peptides or amino acids. Enzymatic proteolysis is catalyzed by proteases, enzymes with proteolytic activity. Non-enzymatically induced proteolysis can be caused by extreme pH and / or high temperature. Proteolytic stability in this specification is distinguished from protease stabilization.

[0045] Enzyme storage stability is usually compromised over time in aqueous solutions. This can be avoided by storing the enzyme under non-aqueous conditions. When non-aqueous conditions are not applicable, for example in compositions that naturally contain water, different or additional strategies must be applied. Stabilization of proteolytic enzymes (proteases) by inhibition is a common technique to prevent proteolysis (proteolysis) of proteins (e.g., enzymes) into peptides or amino acids (which could, for example, inactivate the enzyme's functionality). Generally, protease stabilization relies on reversible inhibition of the enzyme.

[0046] "Half-life of enzyme activity" is a measure of the time required for enzyme activity to decay to one-half (50%) of its initial value.

[0047] "Enzyme inhibitors" slow down enzyme activity by several mechanisms, outlined below. Inhibitor binding can be either reversible or irreversible. Irreversible inhibitors typically bind covalently to the enzyme by modifying key amino acids required for enzyme activity. Reversible inhibitors typically bind non-covalently (hydrogen bonding, hydrophobic interactions, ionic bonding). Four general types of reversible inhibitors are known: (1) The substrate and the inhibitor compete for the use of the enzyme active site (competitive inhibition). (2) The inhibitor binds to the substrate-enzyme complex (uncompetitive inhibition). (3) Inhibitor binding reduces enzyme activity but does not affect substrate binding (noncompetitive inhibition). (4) The inhibitor can bind to the enzyme simultaneously with the substrate (mixed inhibition). It is assumed that the use of enzyme inhibitors, especially reversible inhibitors, stabilizes enzymes. Stabilization of enzymes can result from the temporary inhibition of the catalytic activity of the enzyme compared to the catalytic activity of the same enzyme when not inhibited. Preferably, proteases are inhibited in their proteolytic activity. Due to the inhibition of the proteolytic activity of at least one protease, other enzymes and the protease itself can be stabilized due to the prevention of their proteolysis, resulting in the maintenance of the catalytic activity of the other enzymes.

[0048] Many methods for immobilizing enzymes or fragments thereof or nucleic acids on solid supports are known to those skilled in the art. Some examples of such methods include, for example, electrostatic droplet formation, electrochemical means, adsorption, covalent bonding, cross-linking, chemical reactions or processes, encapsulation, entrapment, calcium alginate, or poly(2-hydroxyethyl methacrylate). Similar methods are described in "Methods in Enzymology, Immobilized Enzymes and Cells, Part C, 1987, Academic Press, edited by SP Colowick and NO Kaplan, Vol. 136" and "Immobilization of Enzymes and Cells, 1997, Humana Press, edited by GF Bickerstaff, Series: Methods in Biotechnology, edited by JM Walker."

[0049] The nucleic acids, enzymes, or collections or cocktails of nucleic acids or enzymes can be immobilized by attachment to a solid support. The solid support can be placed in or removed from a process tank or other vessel in which the nucleic acids or enzymes are repeatedly used. In one embodiment, the solid support is selected from the group of gels, resins, polymers, ceramics, glass, microelectrodes, and / or any combination thereof.

[0050] In the methods of the present invention, each library member preferably contains only one mutation site, and more preferably contains two or more non-adjacent mutation sites. In the present invention, a mutation site consists of one or more adjacent nucleotides. The definition of a mutation site may vary from library to library, such that some library members contain only one mutation site and other library members contain two or more non-adjacent mutation sites. When the number of mutation sites is limited to one per library, this facilitates systematic analysis of single point mutations in a target nucleic acid. This is particularly useful for generating a complete transcription factor binding matrix of transcription factor binding sites in a target nucleic acid or analyzing the effect of point mutations on protein function. Alternatively, one skilled in the art may determine that each library contains two or more non-adjacent mutation sites. This is useful, for example, for analyzing the effect of multiple simultaneous amino acid substitutions on a protein encoded by a target nucleic acid, for example, when modifying the active center of an enzyme. For example, if the mutation sites are one codon long and separated by one codon, the mutated amino acids may typically rely on the same beta-sheet structure. Also, if the mutation site is one codon long and separated by two or three codons, the amino acids affected by the mutation will usually be found facing the same direction in an alpha helix structure, so those skilled in the art can modify the definition of the mutation site as needed.

[0051] Preferably, each mutation site is 1 to 9 nucleotides in length, more preferably 2 to 9 nucleotides in length, even more preferably 2 to 6 nucleotides in length, even more preferably 3 to 9 nucleotides in length, even more preferably 3 to 6 nucleotides in length, even more preferably 3 to 4 nucleotides in length, and even more preferably 3 nucleotides in length.

[0052] More preferably, the mutation sites of two or more libraries at most partially overlap, even more preferably do not overlap, and most preferably the mutation sites of all libraries do not overlap. The method of the present invention requires that the mutation sites, optionally with the aid of nucleic acid tags as described above, allow identification of the libraries based on the sequences obtained in step v). If the mutation sites of all libraries are identical, then in effect only one library has been provided in step i).

[0053] It is particularly preferred that the mutation sites of all libraries do not overlap.In this way, any given nucleotide position in the sequence of target nucleic acid can only be mutated in one library.Therefore, the detection of any mutation in the sequence of probe nucleic acid will immediately identify the corresponding library, so that the detection of sequencing error becomes particularly easy, and if any additional mutation outside the range of one or more mutation sites is detected during sequencing, this sequence will be the sequence containing error.In addition, as described above, when it is uncertain which of two nucleotide readings is correct reading, it can select the correct reading by considering all other mutations and allowed mutations according to the definition of each library.

[0054] In the methods of the present invention, the nucleotides at the mutation sites of one or more libraries preferably (a) differ from the target nucleic acid sequence by one or more arbitrary nucleotides, and / or (b) differ from the target nucleic acid sequence by one or more degenerate nucleotides independently selected from [AT], [CG], [AC], [GT], [AG], [CT], [CGT], [AGT], [ACT], or [ACG], or (c) are selected from a predetermined set of nucleotides or nucleotide sequences, and preferably, the mutation site comprises or consists of a codon replaced with another codon selected from a predetermined set of codons. In the context of the present invention, capital letters in square brackets indicate that the respective nucleotides are interchangeable with each other. Thus, for example, "[AT]" means "A or T."

[0055] As described above, the introduction of random mutations at defined positions in a target nucleic acid sequence is well described in the art and can be performed without any undue burden to the skilled artisan. Thus, the method according to the invention can advantageously be applied using conventional techniques and materials.

[0056] Preferably, one, two, or more, preferably all, nucleotides at the mutation sites of the library differ from the target nucleic acid sequence at at least one position by one or more degenerate nucleotides as described above. The use of degenerate nucleotides is well established in the prior art; see, for example, the International Patent Application Publications cited above. By using such partially degenerate nucleotides, the method according to the present invention allows for the detection of one or more sequencing errors in the synthesis of the mutant library. Each partially degenerate nucleotide excludes one or more nucleotides at each position. For example, if the mutation site is defined by the sequence NN[AT], the occurrence of either A or T at the last position of the mutation site is considered to be the result of a sequencing or synthesis error. Furthermore, if the mutation site is defined by NN[CG], any undesired random introduction of a stop codon can be limited to an amber stop codon.

[0057] Even more preferably, the nucleotide at the mutation site is selected from a predetermined set of nucleotides or nucleotide sequences, as described, for example, in U.S. Patent Application Publication No. 20110165627. The methods of the present invention preferably allow for the definition of a mutation site such that the mutation site comprises or consists of a codon replaced with another codon selected from the predetermined set of codons. Such methods do not rely on the use of degenerate primers and are free from the risk of randomly introducing unwanted stop codons. Alternatively, each amino acid at a given position in the sequence of a protein encoded by a target nucleic acid can be replaced with any other amino acid (and stop codon, if desired) by simple synthesis of the target nucleic acid, without the need for reliance on degenerate nucleotides. Furthermore, because, for all libraries, each amino acid at any mutation site is coded for by only one codon, it is possible to detect sequencing or synthesis errors whenever a codon not included in the predetermined set of codons is detected at the mutation site. The codons belonging to the predetermined set of codons may be defined in any way. Preferably, the set of codons comprises or consists of the most preferred codons for every amino acid according to the codon preferences of the microbial host. Codons may also be selected to optimize Levenshtein distance.

[0058] The methods of the present invention can be implemented in a sequence analysis instrument or sequence analysis device computer program. In the present invention, the computer program can be a single monolithic computer program or can include modules that can be executed simultaneously or sequentially.

[0059] Therefore, the present invention provides a sequence analysis device equipped with a constraint database containing definitions of permissible mutation sites in variants of target nucleic acids. As described above, each mutation site consists of one or more adjacent nucleotides, and the definition of the mutation site allows for the identification of different libraries. Therefore, the constraint database of the sequence analysis device of the present invention reflects the definition of the library of the method of the present invention.

[0060] The sequence analysis apparatus of the present invention also includes a sequencing device for sequencing nucleic acids in parallel. Preferred sequencing methods are described above in connection with the method for analyzing target nucleic acids using a mutant library according to the present invention, and one skilled in the art can select a sequencing device according to any parallel sequencing method.

[0061] When a sequence analyzer according to the present invention is equipped with the above-described constraint database, it also includes a monitoring program that, when used to sequence a nucleic acid that conforms to the constraint definitions in the constraint database, indicates sequences containing errors that do not conform to the constraint definitions in the constraint database or suppresses the output of such erroneous sequences. In this way, the sequence analyzer can advantageously detect erroneous sequences and indicate to the operator that a particular member of a particular library requires further sequencing or refuse to present the sequence of this member. In particular, the monitoring program can detect unexpected mutations outside the range of each mutation site in the library by counting the frequency of every amino acid at every position in the sequence. As described above, the probe nucleic acid can contain mutations only at the mutation site, and therefore, at any given position in the probe nucleic acid, the mutated nucleotide can be the most frequent nucleotide. If the nucleotide at a given position is not the most frequent nucleotide at this position, the nucleotide is definitely considered to be part of the mutation site or indicates a sequencing error.

[0062] The present invention also provides a sequence analysis device comprising a database of library member definitions, where the library members are variants of a target nucleic acid, and the positions of the variants in the target nucleic acid are specific to each library. In this case, the sequence analysis device also comprises a deconvolution program that identifies library members according to the library member definitions based on sequences obtained by the sequencing device. As described above, the method of the present invention preserves a one-to-one correlation between library members and sequences obtained during the parallel sequencing process. Thus, the sequence analysis device of the present invention allows for the implementation and utilization of some or all of the advantages provided by the method of analyzing a target nucleic acid using a variant library of the present invention.

[0063] The sequence analysis device of the present invention preferably further comprises a database containing target nucleic acid-dependent characteristic properties for at least one library member, and a correlation program that generates pairwise assignments between library member sequences and characteristic properties for each library member. Thus, the sequence analysis device makes it possible to associate, for example, in a record, the measured strength of a characteristic, or any other property of a characteristic, for example if the characteristic is measured according to a primary or classification definition, with each library member sequence and thus with each library member.

[0064] The present invention also provides a sequence analyzer computer program which is a monitoring program for the sequence analyzer described above, and / or a deconvolution program for the sequence analyzer described above, and / or a correlation program for the sequence analyzer described above.

[0065] The present invention is further described below in the accompanying examples and figures, none of which are intended to limit the scope of the claims.

[0066] [Example] [Example 1] Single-selection library Figure 1 is an exemplary schematic diagram, characterizing several features individually and in combination with other features that are preferred according to the present invention. In this example, a set of mutant libraries (11, 12, 13) is constructed by first separately mutating a target nucleic acid at specific mutation sites (1, 2, 3). In the figure, the mutation sites are one nucleotide long, and each mutation is made by replacing the respective wild-type nucleotide with a degenerate [ACGT] nucleotide. As noted above, the present invention is not limited to such single-base substitutions; mutation sites may consist of, for example, three consecutive nucleotide substitutions at library-specific mutation sites, resulting in a codon substitution for each library and, therefore, an amino acid substitution in the protein encoded by the target nucleic acid.

[0067] Mutant libraries (11, 12, 13) consisting of microbial host cells each containing a mutated target nucleic acid are plated onto a suitable medium plate, and one library member colony (21, 22, 23) is selected from each plated library. The selected colonies are analyzed for a target nucleic acid-dependent property, e.g., enzyme activity if the target nucleic acid encodes an enzyme (not shown). If the nature of the property meets the needs of those skilled in the art, e.g., because the enzyme activity is significantly higher or lower than the wild-type enzyme activity, probe nucleic acids (31, 32, 33) are generated from each selected member. The probe nucleic acid can be generated, for example, by PCR amplification of the mutated target nucleic acid. As described above, the property does not need to be determined before selecting each member for probe nucleic acid generation; the property can be determined later.

[0068] The amounts of probe nucleic acids (31, 32, 33) are adjusted to ensure that multiple probe nucleic acids (31, 32, 33) are sequenced during sequencing. In the case shown in Figure 1, a 5-fold oversampling is desired.

[0069] The prepared probe nucleic acids (31, 32, 33) are then combined into a sequencing mixture (40) and sequenced in a parallel sequencing instrument (50), which generates a sequence listing (60).

[0070] For the purposes of this example, the wild-type target nucleic acid consists of seven consecutive T nucleotides. Thus, the location of the mutated nucleotide in the sequence (60) obtained from the sequencing device (50) is readily apparent. The sequences can be deconvoluted according to their respective mutation sites. For example, the first sequence contains a nucleotide other than T at position 2. The first library 11 consists of variants at this position. Therefore, the first sequence is considered to belong to the mutated target nucleic acid of member 21. The second sequence contains a nucleotide other than T at position 6. The third library 13 consists of variants at this position. Therefore, the second sequence belongs to the mutated target nucleic acid of member 23 of the third library. The third nucleic acid is identical to the first nucleic acid, a result of the probe nucleic acid concentration adjustment (31, 32, 33) that results in oversampling. The next two sequences are also identical due to oversampling and contain a nucleotide other than T at position 4. The second library 12 consists of variants at this position. Therefore, the last two sequences shown in FIG. 1 belong to the mutated target nucleic acid of member 22 of the second library.

[0071] [Example 2] Combinatorial Libraries and Multiple Rounds of Selection and Probe Nucleic Acid Generation In Figure 2A, target nucleic acids are mutated at multiple library-specific single nucleotide positions as in Example 1, and only two of such mutated target nucleic acids are shown (1, 2). Mutated target nucleic acids are included in each mutant library (11, 12, 13, 14, 15, 16).

[0072] Two of the mutant libraries are then mixed to create a combinatorial library (19, 19', 19''). A combinatorial library is composed of mutant libraries in which the mutation sites in the parent libraries are different from each other and preferably do not overlap.

[0073] As in Example 1, the combinatorial mutant libraries (19, 19', 19'') are plated onto suitable media plates. Plating is shown in FIG. 2A for only one combinatorial library (19). Individual clones (21, 21', 21'') are selected from the plated combinatorial library (19) before, after, or without determining a selected property that depends on the target nucleic acid (see Example 1, mutatis mutandis). As shown in FIG. 2A, probe nucleic acids (31, 31', 31'') are generated separately from each of the selected combinatorial library members (21, 21', 21''). Because the combinatorial library (19) is composed of members of two parent libraries (11, 12), the resulting probe nucleic acids (31, 31', 31'') are mutated (31, 31') at the respective mutation site(s) of one parent library (11) or mutated (31'') at the respective mutation site(s) of the other parent library (12). If the combinatorial library (19) is composed of three or more parent libraries, then probe nucleic acids can be obtained that are mutated at the mutation site of each of the additional parent library(ies).

[0074] Again, the amount of probe nucleic acid (31, 31', 31'') is adjusted as in Example 1 to achieve 5-fold oversampling in the subsequent sequencing step.

[0075] FIG. 2B further expands upon FIG. 2A, showing that the corresponding probe nucleic acids (32, 32', 32''', 33, 33', 33'') are obtained from additional combinatorial libraries (created as libraries 19' and 19'' in FIG. 2A). The probe nucleic acid adjustment step is shown only for combinatorial library 19 in FIG. 2A and is not shown in FIG. 2B for simplicity, although adjustment is performed for all combinatorial libraries to achieve 5-fold oversampling.

[0076] After adjusting the amount of probe nucleic acid (not shown in FIG. 2B), one probe nucleic acid (31) from each combinatorial library (19) is combined with one probe nucleic acid (32, 33) from each of the other combinatorial libraries (19', 19"), respectively, to obtain a mixture of probe nucleic acids (39). This combination operation is then repeated to combine one probe nucleic acid (31', 31") with another probe nucleic acid (32', 33', 32", 33") from an additional combinatorial library (19', 19") to produce an additional mixture of probe nucleic acids (39', 39'). Thus, FIG. 2B shows a total of three repetitions of steps ii) and iii) according to the above description of the present invention.

[0077] An individual nucleic acid tag (not shown) is then linked to each mixed probe nucleic acid (39, 39', 39''), for example by ligation or chemical coupling. Thus, all probe nucleic acids in the first mixture 39 have a first nucleic acid tag attached, probe nucleic acids in the second mixture 39' have a second nucleic acid tag attached, and so on.

[0078] The specifically tagged mixed probe nucleic acids (39, 39', 39") are then mixed to obtain a single mixture of tagged probe nucleic acids (40). The tagged mixed probe nucleic acids (40) are then sequenced (not shown) as described in Example 1 to obtain a complete sequence listing. The sequences are deconvoluted in a two-step process. In the first step, the sequences are sorted according to the individual nucleic acid tags attached to the probe nucleic acid sequences. Since the individual nucleic acid sequence tags are specific for each round of repetition of steps ii) and iii) as described above, it is possible to identify each probe nucleic acid mixture (39, 39', 39"). Furthermore, from the location of the mutation site(s) in each probe nucleic acid sequence, it is possible to determine, in a second deconvolution step, the parent library (11, 12) and the corresponding member to which each probe nucleic acid (31, 31', 31", 32, 32', 32", 33, 33', 33") originates.

Claims

1. 1. A method for analyzing a target nucleic acid, comprising: i) providing isolated members of a set of two or more site-specific mutagenesis libraries, each site-specific mutagenesis library member comprising a target nucleic acid mutated at one or more library-specific mutation sites, the mutation sites of the two or more site-specific mutagenesis libraries being different from one another; ii) selecting one member of each site-specific mutagenesis library; iii) obtaining, for each member, a probe nucleic acid of each mutated target nucleic acid that contains at least the nucleotide at each mutation site and adjacent nucleotides to identify each member; iv) mixing the probe nucleic acids into a mixture; and v) sequencing the probe nucleic acids of the mixture obtained in step iv) in parallel. A method comprising:

2. - repeating steps ii) and iii) to select at least one additional member of the one or more site-specific mutagenesis libraries; - in each round of repetition of steps ii) and iii), before step iv), the probe nucleic acid is labeled with a round-specific nucleic acid tag, The method of claim 1.

3. 3. The method of claim 1 or 2, wherein for at least one member selected in step ii), a characteristic other than its sequence that depends on the target nucleic acid is determined.

4. A method according to claim 1 or 2, wherein for each member selected in step ii), a characteristic other than its sequence that depends on the target nucleic acid is determined, and the same characteristic is determined for each selected member.

5. the property is a property of the target nucleic acid or a protein encoded by the target nucleic acid; - Expression, - secondary, tertiary or quaternary structure, folding efficiency, aggregation, multimerization, - temperature stability, pH stability, solvent stability, detergent stability, protease stability, binding stability, solubility, protease activity, storage stability, storage stability in detergents, residual activity after storage, stability against proteolysis, - kinetic parameters, changes in enzyme activity at low temperatures, at or below 30°C; - Substrate, substrate specificity, enantioselectivity, cofactor dependency, inhibitor specificity, changes in inhibition kinetics, - immunogenicity, toxicity, allergenicity, cleaning performance, - Location in cytoplasm, spore, cell membrane, organelle lumen or membrane, cell wall, excreta 4. The method of claim 3, wherein the method is selected from one or more of the following:

6. (a) Each member: - only one mutation site, or - Two or more non-adjacent mutation sites wherein (a) one or more mutation sites are 2 to 9 nucleotides in length; and (b) the mutation sites of the two or more libraries do not overlap.

7. The nucleotide at the mutation site of one or more libraries is (a) differs from the target nucleic acid sequence at one or more arbitrary nucleotides; and / or (b) differs from the target nucleic acid sequence by one or more degenerate nucleotides independently selected from [AT], [CG], [AC], [GT], [AG], [CT], [CGT], [AGT], [ACT], or [ACG]; or 7. The method of any one of claims 1 to 6, wherein (c) the nucleotide or nucleotide sequence is selected from a predetermined set of nucleotides or nucleotide sequences.

8. i) generating a library of mutants for each mutation site of a target nucleic acid, the mutants comprising mutations of the target nucleic acid only at each library-specific mutation site; and ii) analyzing the library by the method of any one of claims 1 to 7 site-saturation mutagenesis, including

Citation Information

Patent Citations

  • Eukaryotic expression libraries based on double lox recombination and methods of use

    JP2004514444A