Means and methods for providing a CRISPR pad and uses thereof

CN122847533APending Publication Date: 2026-09-29MAX PLANCK GESELLSCHAFT ZUR FOERDERUNG DER WISSENSCHAFTEN EV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480083735.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-06
Filing Date
2024-11-06
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

考虑到核酸成功整合的这些前提条件(例如,为了在需要进行此整合的每个物种中实现异源生产),这给研究群体带来了沉重的负担,导致在每个物种中的大量工作

Benefits of technology

CRISPRpad可插入宿主细胞或宿主生物体的基因组中,从而生成一个编辑盒,该编辑盒可被方便地靶向,用于插入目的基因(GOI)和/或调控GOI的表达。例如,CRISPRpad可通过同源重组插入宿主细胞或生物体中,例如使用最5'和3'的原间隔子序列作为同源臂。本发明的CRISPRpad具有特别的优势,每个原间隔子均可被RNA引导的核酸酶靶向,从而允许将一个或多个GOI靶向插入编辑盒中,和/或(例如使用CRISPRa或CRISPRi)动态地空间调控其表达。这可以通过例如以下方式实现:使用最5'或3'的原间隔子作为靶位点将GOI引入编辑盒,并单独地靶向其余原间隔子以调控GOI的表达。由此,可以通过CRISPRpad的全长来空间调控GOI的表达。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

This invention relates to nucleic acids that provide non-natural target sites (docking elements) for site-directed nucleases, and methods for generating and using them. Furthermore, the invention provides elements for generating said nucleic acids, namely, sequences / sequence fragments and protospacers that are low-representational in one or more reference genomes, and methods for generating said elements. The invention also provides genomes and cells containing said nucleic acids (docking elements).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to nucleic acids that provide new-to-nature target sites (landing pads) for site-directed nucleases, and methods for producing and using them. Furthermore, the invention provides elements for producing said nucleic acids, namely, sequences / sequence fragments and protospacers that are low-representational in one or more reference genomes, and methods for generating said elements. The invention also provides genomes and cells containing said nucleic acids (landing pads). Background Technology

[0002] For heterologous production, DNA can be integrated into naturally occurring gene loci within the cellular genome. When selecting an integration site, its impact on the host, or the influence of the surrounding DNA environment on heterologous construct expression, must be considered. Furthermore, transferring gene constructs to new / different host platforms requires repeating the processes of designing expression constructs, selecting integration loci, and using host-specific tools for integration. Given these prerequisites for successful nucleic acid integration (e.g., achieving heterologous production in every species where this integration is required), this places a heavy burden on the research community, resulting in a significant amount of work across each species.

[0003] Therefore, it is necessary to overcome the cumbersome steps of designing nucleic acid integration separately for each species. The technical problem of this invention is to provide a universal site-directed endonuclease docking element, that is, a docking element that can be applied without considering the species itself, for example, for integrating nucleic acids into a specific species. Summary of the Invention

[0004] This technical problem is solved by providing the embodiments characterized in the claims and provided below. Specifically, this technical problem and the above difficulties can be overcome by providing a CRISPRpad containing underrepresented sequences in one or more reference genomes.

[0005] The CRISPRpads of this invention, along with low-representation nucleic acid sequences and protospacers, address the aforementioned challenges by providing synthetic orthogonal genome docking elements usable across species. CRISPRpads can be constructed from low-representation sequence fragments assembled into synthetic protospacers, which are specifically selected based on their absence in the target organism's genome and their optimized binding to guide RNA (gRNA) and / or single guide RNA (sgRNA). Therefore, these docking elements (CRISPRpads) can be integrated into the genomes of many different organisms, providing a unified genome working platform across different species and avoiding the burden of designing separate insertion strategies for each species.

[0006] This invention addresses the problem of finding unique target sites (e.g., protospacers) when excising, for example, transgenes from organisms in the laboratory. This leads to off-target effects when attempting to excise transgenes. Currently, site-directed nucleases such as the CRISPR / Cas system are primarily used to target existing sequences, such as endogenous sequences in the genome of an organism. The present invention overcomes these problems, particularly the problem of inducing single- or double-stranded DNA breaks in any organism without introducing off-target effects, through the means and methods provided herein. Specifically, inducing single- or double-stranded DNA breaks without introducing off-target effects can be achieved, in particular, by using docking elements that contain sequences that are low-representational in the reference genome (one or more). When targeting these docking elements, for example using a CRISPR / Cas system, for the introduction or excision of target nucleic acids, no off-target effects occur because the docking elements are selected to contain sequences that are as different as possible from the reference genome, i.e., nucleic acid sequences that are low-representational in the genome of the cell to be edited. The methods and techniques provided in this paper address the aforementioned problems, for example, by allowing the introduction and excision of transgenes into docking elements without concern for off-target effects or the need for separately designed target sites. Therefore, CRISPRpads are particularly suitable for targeted editing, such as inserting target nucleic acids and / or transcribedly modifying any inserted target nucleic acid. Once inserted into the genome of the target cell, it can serve as a docking element, for example, in a CRISPR / Cas system, thereby avoiding off-target effects in the host cell. Furthermore, the low-representation sequences and protospacers of this invention facilitate the generation of the aforementioned docking elements. By generating docking elements from the low-representation sequences and protospacers of this invention, off-target effects can be minimized because the resulting docking element sequence will not appear in the host genome of the target cell. In addition, this invention also provides means and methods for generating protospacers from low-representation sequences in a reference genome, and for selecting protospacers that are as different as possible from any sequence in the reference genome. By using a combination of multiple reference genomes, docking elements, protospacers, or low-representation sequences can be specifically selected to make them universal across multiple species, for example, when the docking element contains the least representative sequence among multiple reference genomes (such as multiple reference genomes for mammals). Such docking elements can be advantageously integrated into the host genome of mammalian species, thus avoiding off-target effects when editing targeting the docking element (e.g., using CRISPR / Cas), because the docking element sequence and its contained protospacers are low-representation in mammalian species.As an exemplary embodiment of the present invention, CRISPRpads, consisting of 28 or 30 rationally designed protospacers containing sequences with low representativeness in the respective host cell genomes, have been integrated into the genomes of four industrially relevant microbial hosts: *Saccharomyces cerevisiae* and *Yarrowia lipolytica*, and the bacteria *Escherichia coli* and *Bacillus subtilis*. As a proof of concept, a GFP reporter cassette has been integrated into the CRISPRpads, and dynamic expression experiments have been performed in all four hosts using dCas9. As shown in Example 2, CRISPRpads have been successfully integrated into host cells of *Saccharomyces cerevisiae* and *Yarrowia lipolytica*, as well as the bacteria *Escherichia coli* and *Bacillus subtilis*. Furthermore, it has been demonstrated that dCas9 (CRISPRa and / or CRISPRi) can successfully target the CRISPRpads. Specifically, the expression of the GFP reporter gene has been regulated by targeting different spacers in the CRISPRpads inserted into the host cells using CRISPRa or CRISPRi.

[0007] Furthermore, an exemplary CRISPRpad of the present invention has been inserted into an industrially relevant mammalian cell line—Chinese hamster ovary K1 (CHO K1) cells. As shown in Example 2, as a proof of concept, a corresponding cell line containing the CRISPRpad of the present invention and GFP was generated, demonstrating that any target cell line (e.g., any mammalian cell line) can be modified by introducing the docking element of the present invention, thereby allowing targeted editing within the docking element, avoiding off-target effects, and achieving specific insertion and expression of target nucleic acids (e.g., transgenes), or specific insertion and expression regulation of target nucleic acids inserted into the CRISPRpad. Therefore, the present invention also provides genomes and cell lines containing the low-representation sequences, protospacers, and / or docking elements of the present invention, which are particularly advantageous for the specific targeted production of target proteins (e.g., therapeutic proteins and antibodies).

[0008] The above content is illustrated in the appended embodiments.

[0009] Therefore, the present invention has the following advantages: CRISPRpads can be inserted into the genome of a host cell or organism to generate an editing cassette that can be conveniently targeted for inserting a target gene (GOI) and / or regulating GOI expression. For example, CRISPRpads can be inserted into host cells or organisms via homologous recombination, using, for example, the 5' and 3' protospacer sequences as homologous arms. The CRISPRpads of this invention have a particular advantage: each protospacer can be targeted by an RNA-guided nuclease, allowing for the targeted insertion of one or more GOIs into the editing cassette and / or (e.g., using CRISPRa or CRISPRi) dynamic spatial regulation of their expression. This can be achieved, for example, by introducing the GOI into the editing cassette using the 5' or 3' protospacer as a target site and individually targeting the remaining protospacers to regulate GOI expression. Thus, GOI expression can be spatially regulated using the full length of the CRISPRpad.

[0010] Another advantage of CRISPRpad is that the protospacers are "new-to-nature," meaning they are not present in the host cell or host organism. Specifically, the protospacers are assembled from selected sequences that are as different as possible from any sequences present in one or more reference genomes. Therefore, there is no risk of off-target effects when targeting any protospacer in the CRISPRpad within the host cell or host organism.

[0011] Another advantage of CRISPRpads is that each protospacer is unique and selected to be as different as possible from other protospacers within the CRISPRpad. Furthermore, the CRISPRpad itself is selected to not contain sequences that could lead to recombination within the pad. Therefore, any protospacer / target site within the CRISPRpad can be targeted individually without the risk of off-target effects or unwanted recombination within the CRISPRpad.

[0012] As described in this article, these advantages can be achieved by assembling protospacers from sequences or sequence fragments that are low-representational in one or more genomes. The orthogonality of CRISPRpad sequences to natural reference genome sequences allows CRISPRpads to be applied to different species without the need to design separate strategies for inserting GOIs or regulating GOI expression.

[0013] Therefore, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes (e.g., wherein the nucleic acid is a docking element, preferably a CRISPRpad). In other words, in a preferred embodiment, the present invention relates to a docking element, preferably a CRISPRpad. The docking element / CRISPRpad described herein comprises one or more protospacers, particularly wherein the one or more protospacers comprise or consist of one or more sequences that are low-representation sequences in one or more reference genomes. The docking element / CRISPRpad preferably comprises a protospacer sequence neighbor motif (PAM) located at the 3' end of the one or more protospacers. The nucleic acid described herein (specifically, the protospacers and / or docking elements / CRISPRpads) is preferably non-naturally occurring or engineered. The present invention also provides a non-naturally occurring or engineered nucleic acid comprising one or more protospacers (preferably two or more protospacers) and a protospacer sequence neighbor motif (PAM) located at the 3' end of the one or more protospacers (preferably a PAM located at the 3' end of each of the two or more protospacers). The present invention also provides a nucleic acid comprising two or more protospacers and a protospacer sequence neighbor motif (PAM) located at the 3' end of each protospacer, wherein each protospacer is unique within the nucleic acid.Furthermore, none of these protospatial septa are endogenously present in one or more reference genomes from the following organisms: *Saccharomyces cerevisiae*, *Yarrowia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces pombe*, *Pichiapastoris*, *Komagataella phaffii*, *Kluveromyces lactis*, *Candida antarctica*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, and *Bacillus pumilus*. *Pumilus*, *Paenibacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechocystis* sp. (PCC 6803), *Synechococcus* sp. (PCC 7002), *Synechococcus elongatus* (PCC 7942), *Anabaena* sp. (PCC 7120), *Synechococcus elongatus* UTEX 2973, *Streptomyces coelicolor*, *Streptomyces lividans*, *Clostridium acetobutylicum*, *Oryza sativa* (rice), *Zea mays* (corn), *Nicotiana* (tobacco). The genomes of *Taccamus spp.*, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells. This invention also provides a nucleic acid.It comprises alternating sequences of protospacers and a protospacer sequence neighbor motif (PAM) located at the 3' end of each protospacer. The present invention also provides a nucleic acid comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) located at the 3' end of the one or more protospacers, wherein the one or more protospacers consist of 3-9 consecutive fragments of 4 or 5 nucleotides selected from any one of Tables 4-142.

[0014] The present invention will now be described in more detail.

[0015] CRISPR / Cas

[0016] The CRISPR / CRISPR-associated protein (Cas) system provides a novel source of nucleases and endonucleases, including CRISPR / Cas9. The CRISPR system is a prokaryotic system that provides prokaryotes with resistance against foreign genetic elements such as bacteriophages. As demonstrated in *Streptococcus pyogenes*, Cas9, guided by a double strand formed between a mature activating tracrRNA and a target crRNA, introduces site-specific double-stranded DNA (dsDNA) breaks (DSBs) into invading cognate DNA. Cas9 is a multi-domain enzyme that uses its HNH nuclease domain to cleave the target strand (defined as the strand complementary to the spacer sequence of the crRNA) and its RuvC-like domain to cleave the non-target strand. Upon introduction of DSBs, the cell's natural DNA repair mechanisms come into play. DNA repair primarily involves two methods: Non-homologous end joining (NHEJ): This method often results in small insertions or deletions (indels) at the cleavage site. These indels can disrupt gene function or lead to loss of gene expression.

[0017] Homologous Directed Repair (HDR): This method can be used for more precise gene editing. For example, a donor polynucleotide (donor template) can be provided, which the cell can use to repair the DSB, causing a specific modification / insertion.

[0018] DNA cleavage specificity is determined by two parameters: the variable spacer-derived sequence of the crRNA targeting the protospacer sequence, and the short sequence immediately adjacent to the 3' end (downstream) of the protospacer on the non-target DNA strand, i.e., the protospacer sequence neighbor motif (PAM). The term "protospacer" as used herein corresponds to a nucleic acid site that can be targeted by an RNA-guided nuclease. The protospacer may contain a nucleic acid sequence complementary to the gRNA. Therefore, the protospacer of this invention can be targeted by RNA-guided nucleases. The RNA-guided nuclease in the context of this invention is preferably a CRISPR-associated protein (Cas). Exemplary Cas proteins that may be used in the context of this invention include Cas9, SpCas9, SaCas9, NmeCas9, CjCas9, StCas9, Cpf1 (also known as Cas12), such as LbCpf1 or AsCpf1, AacCas12b, BhCas12b v4, C2c2 (also known as Cas13), Cas12k, Cas12e, Cas14, Cas3, or any synthetic Cas variant. The specific Cas protein used in the context of this invention may depend on the specific requirements of the application, such as the desired PAM sequence. However, in the context of this invention, the RNA-guided nuclease is preferably Cas9, more preferably Streptococcus pyogenes Cas9 (SpCas9).

[0019] Therefore, in one aspect, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein adjacent sequences form a protospacer, and wherein the protospacer contains a target site for an RNA-guided endonuclease (preferably a CRISPR protein).

[0020] On the other hand, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein adjacent sequences form a protospacer, and wherein the protospacer contains a target site for an RNA-guided endonuclease, wherein the RNA-guided endonuclease comprises a Cas protein, preferably Cas9, most preferably Streptococcus pyogenes Cas9 (SpCas9).

[0021] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence adjacent motif (PAM) located at the 3' end of each protospacer, wherein the protospacer contains a target site for an RNA-guided endonuclease (preferably a CRISPR protein).

[0022] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence adjacent motif (PAM) located at the 3' end of each protospacer, wherein the protospacer contains a target site for an RNA-guided endonuclease, wherein the RNA-guided endonuclease contains a Cas protein, preferably Cas9, most preferably Streptococcus pyogenes Cas9 (SpCas9).

[0023] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, The RNA-guided endonuclease is a CRISPR protein.

[0024] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, The RNA-guided endonuclease contains Cas protein, preferably Cas9, and most preferably Streptococcus pyogenes Cas9 (SpCas9).

[0025] Protospacer neighbor motif (PAM)

[0026] Each Cas protein has a specific PAM sequence that recognizes that sequence in the target DNA. Typically, the PAM sequence at the 3' end of each protospacer depends on either an RNA-guided nuclease used to insert the CRISPRpad into the host cell or host organism, or an RNA-guided nuclease used to target a CRISPRpad already integrated into the host cell or host organism's genome. For example, the commonly used *Streptococcus pyogenes* Cas9 (SpCas9) requires a PAM sequence corresponding to 5'-NRG-3', where R includes A or G, and N can be any nucleotide selected from C, G, A, or T, with N immediately adjacent to the 3' end of the target nucleic acid sequence targeted by the spacer sequence. The preferred PAM sequence for SpCas9 is 5'-NGG-3'. Other synthetic or naturally occurring Cas nucleases may require different PAMs. An overview is shown in Table 1 below.

[0027] Table 1. Overview of PAM for Exemplary Cas Proteins

[0028] Wherein N can be any nucleotide selected from C, G, A or T, where R can include A or G, and where N, where Y can include T or C, where W can include A or T, and where V can include G, C or A.

[0029] Typically, the PAM can be included in the nucleic acid of the docking element of the present invention in different forms, or can be introduced into the method of the present invention in different ways. For example, the PAM can be included in the nucleic acid or docking element as a nucleic acid sequence containing only the PAM and no other sequences, located at the 3' end of each protospacer. Then, this PAM can be followed by another protospacer with a 3' PAM, and so on. The PAM can also be included in the nucleic acid or docking element as a follow-up linker, located at the 3' end of each protospacer. The linker is preferably a nucleic acid linker. For example, such a linker can contain 1 to 20 nucleotides located at the 3' end of the PAM, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides located at the 3' end of the PAM sequence. The linker preferably contains 7 nucleotides adjacent to the 3' end of the PAM sequence. Preferably, the linker can contain 13 nucleotides adjacent to the 3' end of the PAM sequence.

[0030] PAM may also be included in a linker, preferably a nucleic acid linker. Such linkers may be located at the 3' end of each protospacer and contain 3 to 20 nucleotides, for example, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides at the 3' end of the protospacer. Such linkers containing PAM preferably contain 3 to 10 nucleotides.

[0031] PAM can also be contained in a low-representation sequence fragment located at the 3' end of each protospacer. For example, PAM can be contained in a low-representation sequence of 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, or 10mer, preferably a 4mer or 5mer sequence. Such sequence fragments can be specifically selected to contain PAM. For example, if SpCas9 PAM is required, a 4mer sequence starting with 5'-NRG-3' can be selected. This 4mer can then be contained in the docking element or nucleic acid of the present invention, located at the 3' end of the protospacer. This also applies to any 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, or 10mer sequence mentioned herein.

[0032] The PAM sequence can also be included in the protospacer of the present invention. For example, a protospacer can be selected to include a 5'-NRG-3' sequence at its 5' end. This protospacer can then be included in the nucleic acid or docking element of the present invention at the 3' end of another protospacer, thereby forming a PAM at the 3' end of said other protospacer. For this purpose, the partial protospacer of the present invention can also be used. For example, the partial protospacer can contain 3 to 20 nucleotides, such as 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in the protospacer of the present invention. Such a partial protospacer can be selected to include 5' NRG-3' at any position in its sequence, and then it can be included in the nucleic acid or docking element of the present invention at the 3' end of the protospacer. The remaining or overlapping sequences of the partial protospacer can be discarded. In the nucleic acid or docking element of the present invention, it is advantageous to use the protospacer of the present invention containing PAM at the 3' end of the protospacer because the protospacer of the present invention is specifically selected to be as different as possible from the reference genome. Therefore, when using the protospacer of the present invention containing PAM, there is no risk of off-target effects in the corresponding genome of the cell to which the nucleic acid or docking element of the present invention is to be introduced.

[0033] With adaptive modifications, the above content also applies to the methods of the present invention for generating docking elements and for generating non-natural protospacers. Therefore, in the methods described above, the nucleic acid sequence containing PAM can be used to introduce PAM into the 3' end of the protospacer.

[0034] Therefore, one aspect of the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein adjacent sequences form a protospacer, wherein the nucleic acid further comprises: (a) The protospacer sequence neighbor motif (PAM) located at the 3' end of each protospacer, or (b) A protospacer adjacent motif (PAM) located at the 3' end of each protospacer, and a linker of 1 to 10 nucleotides adjacent to the 3' end of the PAM, preferably a linker of 7 nucleotides adjacent to the 3' end of the PAM, or (c) A linker of 3 to 10 nucleotides, preferably a linker of 10 nucleotides, located at the 3' end of each protospacer and containing a protospacer adjacent motif (PAM). (d) A 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, or 10mer sequence, preferably a 4mer or 5mer sequence, located at the 3' end of each protospacer and containing a protospacer adjacent motif (PAM). (e) Another protospacer or another incomplete protospacer located at the 3' end of each protospacer and containing a protospacer sequence adjacent motif (PAM), wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the protospacer, preferably 10 nucleotides of the protospacer.

[0035] On the other hand, the present invention provides a nucleic acid comprising one or more protospacers of the present invention, preferably at least two protospacers of the present invention, more preferably all of them being protospacers of the present invention, and a protospacer sequence adjacent motif (PAM) at the 3' end of each protospacer, and optionally a linker of 1 to 10 nucleotides adjacent to the 3' end of the motif, preferably a linker of 7 nucleotides adjacent to the 3' end of the motif, wherein optionally, the PAM is included in the following: (a) A linker of 3 to 10 nucleotides (preferably 10 nucleotides) located at the 3' end of each protospacer, or (b) A 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, or 10mer sequence located at the 3' end of each protospacer, preferably a 4mer or 5mer sequence, or (c) Another protospacer or another incomplete protospacer located at the 3' end of each protospacer, wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the protospacer, preferably 10 nucleotides of the protospacer.

[0036] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence adjacent motif (PAM) located at the 3' end of each protospacer, and a linker of 1 to 10 nucleotides optionally adjacent to the 3' end of the PAM, preferably a linker of 7 nucleotides adjacent to the 3' end of the PAM, wherein the PAM is included in the following: (a) A linker of 3 to 10 nucleotides (preferably 10 nucleotides) located at the 3' end of each protospacer, or (b) A 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, or 10mer sequence located at the 3' end of each protospacer, preferably a 4mer or 5mer sequence, or (c) Another protospacer or another incomplete protospacer located at the 3' end of each protospacer, wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the protospacer, preferably 10 nucleotides of the protospacer.

[0037] On the other hand, the present invention provides a method for generating docking elements, the method comprising: (i) Identify sequences that are underrepresented in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Introduce a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, or A protospacer adjacent motif (PAM) and a linker of 1 to 10 nucleotides adjacent to the 3' end of the PAM are introduced at the 3' end of each protospacer, preferably a linker of 7 nucleotides adjacent to the 3' end of the PAM, or A linker of 3 to 10 nucleotides (preferably 10 nucleotides) containing a protospacer adjacent motif (PAM) is introduced at the 3' end of each protospacer, or A 3-mer, 4-mer, 5-mer, 6-mer, 7-mer, 8-mer, 9-mer, or 10-mer sequence containing the neighboring motif (PAM) of the protospacer sequence is introduced at the 3' end of each protospacer, preferably a 4-mer or 5-mer sequence, or Another protospacer (ii) or another incomplete protospacer (ii) containing a protospacer sequence adjacent motif (PAM) is introduced at the 3' end of each protospacer, wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the protospacer (ii), preferably 10 nucleotides of the protospacer (ii). (iv) Assembling two or more protospacers into a nucleic acid encoding it, wherein each protospacer is unique within that nucleic acid. (v) Synthesize nucleic acids encoding the two or more protospacers to generate docking elements.

[0038] On the other hand, the present invention provides a method for generating non-natural atomic spacers, the method comprising: (i) Extract sequences from one or more reference genomes. (ii) Identify sequences that are underrepresented in one or more reference genomes. (iii) Assemble low-representation sequences into one or more original spacers. (iv) Introduce a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, or A protospacer adjacent motif (PAM) and a linker of 1 to 10 nucleotides adjacent to the 3' end of the PAM are introduced at the 3' end of each protospacer, preferably a linker of 7 nucleotides adjacent to the 3' end of the PAM, or A linker of 3 to 10 nucleotides (preferably 10 nucleotides) containing a protospacer adjacent motif (PAM) is introduced at the 3' end of each protospacer, or A 3-mer, 4-mer, 5-mer, 6-mer, 7-mer, 8-mer, 9-mer, or 10-mer sequence containing the neighboring motif (PAM) of the protospacer sequence is introduced at the 3' end of each protospacer, preferably a 4-mer or 5-mer sequence, or Another protospacer (ii) or another incomplete protospacer (ii) containing a protospacer sequence adjacent motif (PAM) is introduced at the 3' end of each protospacer, wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the protospacer (ii), preferably 10 nucleotides of the protospacer (ii). (v) Optionally, one or more protospacers containing PAM are synthesized into nucleic acids.

[0039] Although the nucleic acids of the docking elements of the present invention preferably contain PAM at the 3' end of each protospacer, PAM may not be necessary when using catalytically inactivated Cas (e.g., dCas9 in the CRISPRa or CRISPRi systems) to target the protospacers in the nucleic acids or docking elements of the present invention. This is because PAM allows for the realization of SSB or DSB, which are not induced by these systems. Therefore, any nucleic acid, docking element, or protospacer of the present invention, or the method of generating them, may not contain PAM, or may not include a step of introducing PAM.

[0040] Catalytically inactivated Cas proteins

[0041] As described above, Cas9 can complex with gRNA and induce DSB at sites complementary to the spacer sequence of the gRNA. However, in this invention, a catalytically inactivated form of Cas9, such as catalytically inactivated Cas9, can also be used. This approach is advantageous when expression regulation is required (e.g., in the docking elements of this invention), for example, when using a CRISPR system to inactivate or disable the expression of a target gene. Several catalytically inactivated forms of Cas9 are available for use in this invention, including a CRISPR interference form (CRISPRi) and a CRISPR activation form (CRISPRa). Catalytically inactivated forms of Cas9 lack the nuclease activity of wild-type Cas9; that is, they can still hybridize with the original spacer but do not induce DSB or SSB. Catalytically inactivated forms of Cas9 can be generated through specific mutations. For example, dead Cas9 (dead Cas9, dCas9) contains D10A and H840A mutations. The D10A mutation affects the RuvC domain of Cas9, while the H840A mutation affects the HNH domain. The RuvC and HNH domains are responsible for the endonuclease activity of Cas9. These mutations prevent Cas9 from inducing DSB. Exemplary catalytically inactivated forms of Cas9 are dCas9, nickase Cas9 (nCas9), Cas9 nickase (Cas9n), FokI-dCas9, or xCas9. In this invention, the preferred catalytically inactivated form of Cas9 is dead Cas9 (dead Cas9, dCas9). To regulate expression, such as the expression of a target gene, the catalytically inactivated Cas protein can be fused with a transcription activator to induce the expression of the target gene. This system is called the CRISPR activation (CRISPRa) system. In short, the catalytically inactivated Cas protein can be fused with one or more transcription activators, which, upon binding to a target gene or its promoter region, can activate the transcription of that gene, thereby causing an increase in the production of the gene mRNA and ultimately the corresponding protein (without introducing SSB or DSB). In this invention, exemplary transcriptional activators include Vp16 transcriptional activator, Vp64 transcriptional activator, p65 transcriptional activator, Rta transcriptional activator, co-activator mediator (SAM), SunTag activation system, Moontag, HSF1, HSFA6b, AvrXa10, DOF1, DREB1, DREB2, dCas9-TV, p300, CBP, SoxS, SoxS variant (R93A), or combinations thereof. Exemplary combinations of transcriptional activators are VPR comprising a combination of VP64, p65, and Rta, or the SuntagTgVPH system comprising a combination of the SunTag activator system and VP64, p65, and HSF1.

[0042] Similarly, to regulate expression, such as the expression of a target gene, a catalytically inactivated Cas protein can be fused with a transcriptional repressor to reduce or eliminate the expression of the target gene. This system is known as the CRISPR interference (CRISPRi) system. In short, a catalytically inactivated Cas protein can be fused with one or more transcriptional repressors, which, upon binding to a target gene or its promoter region, can inhibit the transcription of that gene, thereby reducing or eliminating the production of the gene's mRNA and ultimately the corresponding protein (without introducing SSB or DSB). Exemplary transcriptional repressors in this invention are: Kruppel-associated box (KRAB), Mxi1, MeCP2 (methyl-CpG-binding protein 2), SIN3A (switch-independent 3A), HDAC (histone deacetylase), Groucho / TLE (transduction protein-like enhancer), NCOR (nuclear receptor co-repressor), NCoR2 (nuclear receptor co-repressor 2), lysine-specific demethylase 1 (LSD1), retinoic acid and thyroid hormone receptor silencing mediator (SMRT), or combinations thereof. An exemplary combination of transcriptional repressors is: the NCOR / SMRT system combining NCOR and SMRT, or the Sin3-HDAC complex.

[0043] Therefore, in one aspect, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein adjacent sequences form a protospacer, and wherein the protospacer contains a target site of a Cas protein, preferably Cas9, most preferably Streptococcus pyogenes Cas9 (SpCas9), wherein the Cas protein comprises a catalytically inactivated Cas protein, such as dead Cas9 (dCas9) containing D10A and H840A mutations.

[0044] On the other hand, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein adjacent sequences form a protospacer, and wherein the protospacer contains a target site of a catalytically inactivated Cas protein, wherein the catalytically inactivated Cas protein comprises a transcription activator, such as Vp16 transcription activator, Vp64 transcription activator, p65 transcription activator, Rta transcription activator, co-activator mediator (SAM), SunTag activator system, Moontag, HSF1, HSFA6b, AvrXa10, DOF1, DREB1, DREB2, dCas9-TV, p300, CBP, or combinations thereof.

[0045] On the other hand, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein adjacent sequences form a protospacer, and wherein the protospacer contains a target site of a catalytically inactivated Cas protein, wherein the catalytically inactivated Cas protein comprises a transcriptional repressor such as Kruppel-associated box (KRAB), Mxi1, MeCP2 (methyl-CpG-binding protein 2), SIN3A (switch-independent 3A), HDACs (histone deacetylases), Groucho / TLE (transduction protein-like enhancer), NCOR (nuclear receptor co-repressor), NCoR2 (nuclear receptor co-repressor 2), lysine-specific demethylase 1 (LSD1), retinoic acid and thyroid hormone receptor silencing mediator (SMRT), or combinations thereof.

[0046] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence adjacent motif (PAM) located at the 3' end of each protospacer, wherein the protospacer contains a target site for an RNA-guided endonuclease, wherein the RNA-guided endonuclease contains a Cas protein, preferably Cas9, most preferably Streptococcus pyogenes Cas9 (SpCas9), and wherein the Cas protein contains a catalytically inactivated Cas protein, such as dead Cas9 (dCas9) containing D10A and H840A mutations.

[0047] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) located at the 3' end of each protospacer, wherein the protospacer contains a target site for a catalytically inactivated Cas protein, wherein the catalytically inactivated Cas protein contains a transcription activator, such as Vp16 transcription activator, Vp64 transcription activator, p65 transcription activator, Rta transcription activator, co-activator mediator (SAM), SunTag activator system, Moontag, HSF1, HSFA6b, AvrXa10, DOF1, DREB1, DREB2, dCas9-TV, p300, CBP, or combinations thereof.

[0048] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) located at the 3' end of each protospacer, wherein the protospacer contains a target site for a catalytically inactivated Cas protein, wherein the catalytically inactivated Cas protein comprises a transcriptional repressor such as Kruppel-associated box (KRAB), Mxi1, MeCP2 (methyl-CpG-binding protein 2), SIN3A (switch-independent 3A), HDACs (histone deacetylases), Groucho / TLE (transduction protein-like enhancer), NCOR (nuclear receptor co-repressor), NCoR2 (nuclear receptor co-repressor 2), lysine-specific demethylase 1 (LSD1), retinoic acid and thyroid hormone receptor silencing mediator (SMRT), or combinations thereof.

[0049] Guide RNA (gRNA)

[0050] Guide RNA (gRNA) is part of the CRISPR / Cas system that guides a nuclease to a target defined by complementarity with a spacer sequence within the gRNA. The guide RNA may contain at least one spacer sequence that hybridizes to the target nucleic acid sequence and a CRISPR repeat sequence. In type II systems, the gRNA also contains a second RNA, called a tracrRNA sequence. In type II guide RNA (gRNA), the CRISPR repeat sequence and the tracrRNA sequence hybridize to form a double strand. In type V guide RNA (gRNA), the crRNA forms a double strand. In both systems, the double strand can bind to a site-directed nuclease (e.g., Cas9), thereby forming a complex between the guide RNA and the site-directed nuclease. Through the binding of the gRNA to the site-directed polypeptide, the gRNA provides targeting specificity to the complex. Thus, the gRNA nucleic acid can guide the activity of the site-directed nuclease. The guide RNA (gRNA) can be a bimolecular guide RNA or a single-molecule guide RNA (sgRNA). For the purposes of this invention, the gRNA is preferably sgRNA. A bimolecular guide RNA may contain two RNA strands. The first strand contains an optional spacer extension sequence, a spacer sequence, and a minimal CRISPR repeat sequence in the 5' to 3' direction. The second strand may contain a minimal tracrRNA sequence (complementary to the minimal CRISPR repeat sequence), a 3' tracrRNA sequence, and an optional tracrRNA extension sequence. In the type II system, the single-molecule guide RNA (sgRNA) may contain an optional spacer extension sequence, a spacer sequence, a minimal CRISPR repeat sequence, a single-molecule guide linker, a minimal tracrRNA sequence, a 3' tracrRNA sequence, and an optional tracrRNA extension sequence in the 5' to 3' direction. The optional tracrRNA extension may contain elements that contribute to additional functions of the guide RNA, such as stability. The single-molecule guide linker can link the minimal CRISPR repeat sequence and the minimal tracrRNA sequence to form a hairpin structure. The optional tracrRNA extension may contain one or more hairpin structures. The sgRNA may contain a 20-nucleotide spacer sequence at the 5' end of the sgRNA sequence. The sgRNA may contain a spacer sequence of less than 20 nucleotides at the 5' end of the sgRNA sequence. Preferably, the sgRNA comprises a spacer sequence of at least 16 nucleotides at the 5' end of the sgRNA sequence. The sgRNA may comprise a spacer sequence of more than 20 nucleotides at the 5' end of the sgRNA sequence. The sgRNA may comprise a variable-length spacer sequence of 16-35 nucleotides at the 5' end of the sgRNA sequence.

[0051] The spacer sequence hybridizes with a sequence in the target nucleic acid. In the context of this invention, the spacer sequence hybridizes with the protospacer of this invention or a protospacer contained in the docking element or nucleic acid of this invention. Therefore, the spacer of the gRNA or sgRNA of this invention can target the protospacer of this invention already integrated into the genome of a cell or organism, or a protospacer contained in the docking element or nucleic acid of this invention, and interact with the protospacer in a sequence-specific manner through hybridization (i.e., base pairing). The nucleotide sequence of the spacer can vary depending on the sequence of the targeted protospacer.

[0052] However, the present invention also provides gRNA or sgRNA containing a spacer sequence that targets a genomic sequence in the genome of a cell or organism, such as an endogenous genomic sequence. This is for the purpose of introducing the original spacer or nucleic acid or docking element of the present invention into the genome of a cell / organism, for example, through homologous recombination.

[0053] In the CRISPR / Cas system disclosed herein, the spacer sequence can be designed to hybridize with the target nucleic acid located at the 5' end of the PAM of the Cas9 enzyme used in the system. The spacer sequence can be perfectly matched with the target sequence, or mismatches may exist. When targeting the protospacer of the present invention or the protospacer contained in the docking element or nucleic acid of the present invention, a perfect match is preferred, i.e., the gRNA or sgRNA contains a spacer sequence complementary to the protospacer sequence. This avoids any potential off-target effects in the docking element or nucleic acid of the present invention, or in the genome of the cell / organism. Specifically, the gRNA or sgRNA may contain only spacer sequences complementary to the unique protospacer in the docking element or nucleic acid of the present invention. This avoids any competition or off-target effects in the docking element or nucleic acid of the present invention. However, the spacer sequence contained in the gRNA or sgRNA of the present invention can also tolerate some mismatches with its target sequence and / or the protospacer of the present invention.

[0054] As used herein, "substantial complementarity" means that the gRNA or sgRNA contains a spacer sequence capable of hybridizing with its target sequence, wherein the hybridization is not complete (e.g., complementarity less than 100%) and includes mismatches. In some examples, the percentage of complementarity between the spacer sequence and the target nucleic acid and / or the original spacer of the present invention is at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, or 100%. In some examples, the complementarity percentage between the spacer sequence and the target nucleic acid and / or the original spacer of the present invention is at most about 30%, at most about 40%, at most about 50%, at most about 60%, at most about 65%, at most about 70%, at most about 75%, at most about 80%, at most about 85%, at most about 90%, at most about 95%, at most about 97%, at most about 98%, at most about 99%, or 100%. Preferably, the complementarity percentage between the spacer sequence and the target nucleic acid and / or the original spacer of the present invention is 100% at the 5' end of the target sequence of the complementary strand of the target nucleic acid. The complementarity percentage between the spacer sequence and the target nucleic acid at about 20 consecutive nucleotides can be at least 60%. The length of the spacer sequence and the target nucleic acid can differ by 1 to 6 nucleotides, which can be considered as one or more bulges. Preferably, the complementarity percentage between the spacer sequence and the target nucleic acid and / or the original spacer of the present invention is 100%, that is, the gRNA or sgRNA of the present invention contains a spacer sequence that is complementary to the target nucleic acid and / or the original spacer of the present invention or the original spacer contained in the docking element or nucleic acid of the present invention.

[0055] The length of the spacer sequence that hybridizes with the target nucleic acid can be at least about 6 nucleotides (nt). The spacer subsequence can be at least about 6 nt, at least about 10 nt, at least about 15 nt, at least about 18 nt, at least about 19 nt, at least about 20 nt, at least about 25 nt, at least about 30 nt, at least about 35 nt, or at least about 40 nt, about 6 nt to about 80 nt, about 6 nt to about 50 nt, about 6 nt to about 45 nt, about 6 nt to about 40 nt, about 6 nt to about 35 nt, about 6 nt to about 30 nt, about 6 nt to about 25 nt, about 6 nt to about 20 nt, about 6 nt to about 19 nt, about 10 nt to about 50 nt, about 10 nt to about 45 nt, about 10 nt to about 40 nt, about 10 nt to about 35 nt, about 10 nt to about 30 nt, about 10 nt to about 25 nt, about 10 nt to about 20 nt, about 10 nt to about 19 nt. The spacer sequence may contain 20 nucleotides. In some embodiments, the spacer may contain 19 nucleotides to about 25 nucleotides, about 19 nucleotides to about 30 nucleotides, about 19 nucleotides to about 35 nucleotides, about 19 nucleotides to about 40 nucleotides, about 19 nucleotides to about 45 nucleotides, about 19 nucleotides to about 50 nucleotides, or about 20 nucleotides to about 60 nucleotides. Preferably, the spacer sequence may contain 20 nucleotides. In some embodiments, the spacer may contain 19 nucleotides. In some embodiments, the spacer may contain 18 nucleotides. In some embodiments, the spacer may contain 21 or 22 nucleotides. Preferably, the complementarity percentage between the 20-nucleotide spacer sequence and the target nucleic acid and / or the protospacer of the present invention is 100%, that is, the gRNA or sgRNA of the present invention contains a spacer sequence that is complementary to the target nucleic acid and / or the protospacer of the present invention or the protospacer contained in the docking element or nucleic acid of the present invention over the entire length of the 20 nucleotides.

[0056] Spacer sequences can be designed or selected using computer programs. These programs can use variables such as predicted melting temperature, secondary structure formation, predicted annealing temperature, sequence identity, genomic background, chromatin accessibility, GC percentage, genomic frequency (e.g., the frequency of identical or similar sequences that differ at one or more sites due to mismatches, insertions, or deletions), methylation status, and the presence of SNPs.

[0057] Therefore, in one aspect, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein adjacent sequences form a protospacer, and wherein the protospacer contains a target site for an RNA-guided endonuclease, and wherein the RNA-guided endonuclease is complexed with a guide RNA (gRNA), preferably a single-molecule guide RNA (sgRNA).

[0058] On the other hand, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein adjacent sequences form a protospacer, and wherein the protospacer contains a target site for an RNA-guided nuclease, and wherein the RNA-guided nuclease is complexed with a guide RNA (gRNA), preferably a single-molecule guide RNA (sgRNA), wherein the gRNA or sgRNA contains a spacer sequence substantially complementary to the protospacer sequence of the nucleic acid, preferably wherein the gRNA or sgRNA contains a spacer sequence complementary to the protospacer sequence of the nucleic acid.

[0059] On the other hand, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein adjacent sequences form a protospacer, and wherein the protospacer contains a target site for an RNA-guided nuclease, and wherein the RNA-guided nuclease is complexed with a guide RNA (gRNA), preferably a single-molecule guide RNA (sgRNA), wherein the gRNA or sgRNA contains a spacer sequence substantially complementary to the protospacer sequence of the nucleic acid, preferably wherein the gRNA or sgRNA contains a spacer sequence substantially complementary to the unique protospacer sequence of the nucleic acid, preferably wherein the gRNA or sgRNA contains a spacer sequence complementary to the unique protospacer sequence of the nucleic acid.

[0060] On the other hand, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein adjacent sequences form a protospacer, and wherein the protospacer contains a target site for an RNA-guided endonuclease, and wherein the RNA-guided endonuclease is complexed with a guide RNA (gRNA) or a single-molecule guide RNA (sgRNA) comprising a spacer selected from SEQ ID NO: 180-237.

[0061] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence adjacent motif (PAM) located at the 3' end of each protospacer, wherein the protospacer contains a target site for an RNA-guided endonuclease, and wherein the RNA-guided endonuclease is complexed with a guide RNA (gRNA), preferably a single-molecule guide RNA (sgRNA).

[0062] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence adjacent motif (PAM) located at the 3' end of each protospacer, wherein the protospacer contains a target site for an RNA-guided nuclease, and wherein the RNA-guided nuclease is complexed with a guide RNA (gRNA), preferably a single-molecule guide RNA (sgRNA), wherein the gRNA or sgRNA contains a spacer sequence substantially complementary to the protospacer sequence of the docking element, preferably wherein the gRNA or sgRNA contains a spacer sequence complementary to the protospacer sequence of the docking element.

[0063] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence adjacent motif (PAM) located at the 3' end of each protospacer, wherein the protospacer contains a target site for an RNA-guided nuclease, and wherein the RNA-guided nuclease is complexed with a guide RNA (gRNA), preferably a single-molecule guide RNA (sgRNA), wherein the gRNA or sgRNA contains a spacer sequence substantially complementary to the unique protospacer sequence of the docking element, preferably wherein the gRNA or sgRNA contains a spacer sequence complementary to the unique protospacer sequence of the docking element.

[0064] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence adjacent motif (PAM) located at the 3' end of each protospacer, wherein the protospacer contains a target site for an RNA-guided endonuclease, and wherein the RNA-guided endonuclease is complexed with a guide RNA (gRNA) or a single-molecule guide RNA (sgRNA) comprising a spacer selected from SEQ ID NO: 180-237.

[0065] RNA-guided nucleases and gRNA or sgRNA can be applied separately to cells or organisms. Alternatively, RNA-guided nucleases can be pre-complexed with one or more guide RNAs, or with one or more crRNAs in combination with tracrRNA. This pre-complexed substance can then be applied to cells or organisms. This pre-complexed substance is called a ribonucleoprotein particle (RNP). The RNA-guided nuclease in the RNP can be, for example, the Cas9 endonuclease or Cpf1 endonuclease described herein, or a catalytically inactivated Cas protein. The RNA-guided nuclease may have one or more nuclear localization signals (NLS) attached to its N-terminus, C-terminus, or both simultaneously. For example, the Cas9 endonuclease or catalytically inactivated Cas9 may have two NLSs attached, one at the N-terminus and the other at the C-terminus. This NLS can be any NLS known in the art, such as the SV40 NLS. The weight ratio of gRNA or sgRNA to the RNA-guided nuclease in the RNP can be 1:1. For example, the weight ratio of sgRNA to Cas9 endonuclease or catalytically inactivated Cas9 in RNP can be 1:1.

[0066] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, Wherein, the one or more gRNAs are one or more sgRNAs.

[0067] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, The RNA-guided endonuclease is pre-complexed with one or more gRNAs or one or more sgRNAs to form a ribonucleoprotein (RNP) complex.

[0068] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, The RNA-guided endonuclease or the nucleic acid encoding the RNA-guided endonuclease is formulated in liposomes or lipid nanoparticles, and the liposomes or lipid nanoparticles further contain one or more gRNAs or nucleic acids encoding the one or more gRNAs.

[0069] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, Each of (i), (ii) and (iii) is formulated in a liposome or a lipid nanoparticle, or (i), (ii) and (iii) are formulated together in a single liposome or lipid nanoparticle.

[0070] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or sgRNAs or nucleic acids encoding them, The one or more gRNAs or sgRNAs contain spacer sequences substantially complementary to loci in the host cell genome, preferably wherein the one or more gRNAs or sgRNAs contain spacer sequences complementary to loci in the host cell genome.

[0071] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or sgRNAs or nucleic acids encoding them, Wherein, the one or more gRNAs or sgRNAs contain a spacer sequence substantially complementary to the original spacer sequence of the nucleic acid (i), preferably wherein the one or more gRNAs or sgRNAs contain a spacer sequence complementary to the original spacer sequence of the nucleic acid (i).

[0072] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or sgRNAs or nucleic acids encoding them, The gRNA or sgRNA contains a spacer sequence substantially complementary to the unique protospacer of nucleic acid (i), preferably wherein the gRNA or sgRNA contains a spacer sequence complementary to the unique protospacer of nucleic acid (i).

[0073] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) At least two gRNAs or sgRNAs or nucleic acids encoding them.

[0074] On the other hand, the present invention provides a system comprising: (i) The nucleic acid or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) At least two gRNAs or sgRNAs or nucleic acids encoding them, Wherein, the at least two gRNAs or sgRNAs contain spacer sequences substantially complementary (preferably complementary) to loci in the host cell genome, and wherein one of the at least two gRNAs or sgRNAs contains a spacer sequence substantially complementary (preferably complementary) to a unique protospacer at the 5' end of the nucleic acid or docking element of (i), and wherein the other of the at least two gRNAs or sgRNAs contains a spacer sequence substantially complementary (preferably complementary) to a unique protospacer at the 3' end of the nucleic acid of (i).

[0075] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or sgRNAs containing a spacer selected from SEQ ID No: 180-237.

[0076] docking components

[0077] As described above, the nucleic acid of the present invention has a particular advantage because it contains a sequence different from the reference genome sequence. This avoids off-target effects when targeting the nucleic acid of the present invention in vivo (e.g., using a site-directed nuclease) and allows the use of the nucleic acid in multiple species without the risk of introducing sequences that may correspond to endogenous sequences (e.g., sequences in the species genome). This is achieved by the nucleic acid of the present invention containing sequences that are low-representative in one or more reference genomes. Low-representative sequences in one or more reference genomes can be selectively and specifically assembled to obtain a nucleic acid that, when targeted (e.g., using a site-directed nuclease), does not cause any off-target effects in vivo.

[0078] Therefore, one aspect of the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes.

[0079] The nucleic acids of this invention may contain CRISPRpads / docking elements. In this document, the terms "CRISPRpad" and "docking element" are used interchangeably. A docking element is a nucleic acid containing a target site for a nuclease, such as a target site for an RNA-guided nuclease, like a protospacer. Generally, a docking element is an artificial / synthetic nucleic acid containing one or more protospacers. The docking element preferably contains a PAM sequence located at the 3' end of each protospacer. The docking element can be integrated into an organism, for example, into the genome, and can then be used as a target site, such as a target site for a nuclease, to insert and / or excise a target gene / transfer, or to regulate the expression of a target gene / transfer, similar to a safe harbor locus. An advantage of the docking element of this invention is that it contains a low-representation sequence in one or more reference genomes, thus preventing off-target effects when targeting the docking element.

[0080] Therefore, one aspect of the present invention provides a docking element comprising one or more primary spacers and a primary spacer sequence neighbor motif (PAM) located at the 3' end of each primary spacer. The one or more primary spacers may be any primary spacers provided herein.

[0081] The smallest unit of the docking element of this invention can be a primary spacer followed by a PAM located at the 3' end of the primary spacer. The primary spacer can contain at least 15 nucleotides, such as 16 or 20 nucleotides, or be composed of them. Therefore, the smallest unit can contain a primary spacer of 16 or 20 nucleotides and a PAM of 3 nucleotides, such as SpCas9 PAM. This results in a smallest unit of 19 or 23 nucleotides.

[0082] Therefore, in one aspect, the present invention provides a docking element comprising a primary spacer and a PAM at the 3' end of the primary spacer.

[0083] In a preferred aspect, the present invention provides a docking element comprising at least two protospacers and a PAM located at the 3' end of each of the two protospacers. Thus, the docking element can be designed to include alternating sequences, i.e., a protospacer followed by a 3' PAM, then another protospacer, then another 3' PAM, and so on. The docking element of the present invention does not impose a particular limitation on the number of protospacers, which can be adjusted according to the intended use and the organism. For example, when regulatory expression is required rather than simply inserting and / or excising a target gene, the docking element of the present invention can use more protospacers. The number of protospacers used can also depend on the organism in which the docking element of the present invention will be used. For example, some organisms may require shorter or longer sequences to allow insertion via homologous recombination. Therefore, the number of protospacers can be adjusted based on the desired application.

[0084] Therefore, in one aspect, the present invention provides a docking element comprising about 2 to 60 primary spacers, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 primary spacers. In a preferred aspect, the present invention provides a docking element comprising about 30 primary spacers. In another preferred aspect, the present invention provides a docking element comprising 13 primary spacers.

[0085] The size of the docking element generally depends on the number of protospacers it contains. For example, as described herein, the smallest unit of the docking element of the present invention can be a protospacer followed by a PAM at the 3' end of the protospacer. However, preferably, the docking element contains at least two protospacers. This is advantageous because at least two protospacers allow for different target sites and provide sequences that can be used for homologous recombination, for example, when inserted into a target gene encoded in a donor polynucleotide. There is no particular limitation on the size of the docking element of the present invention. The size of the docking element can be adjusted according to the application. For example, more protospacers may be advantageous for spatially regulating the expression of a target gene. Conversely, for simple insertion / removal, one or two protospacers may be sufficient.

[0086] Therefore, in one aspect, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, wherein the docking element comprises about 19-2500 nucleotides, preferably about 600 nucleotides, and most preferably about 840 or 900 nucleotides. In some embodiments, the docking element comprises about 23-2500 nucleotides.

[0087] Generally, each protospacer contained in the docking element of the present invention is unique within the docking element. This is particularly advantageous in avoiding any competing or off-target effects within the docking element, especially when targeting a single protospacer within the docking element. Furthermore, this also helps to avoid the formation of secondary structures when processing the docking element (e.g., transforming / inserting it into a organism).

[0088] Therefore, in one aspect, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, wherein the sequence of each protospacer is unique within the docking element.

[0089] The protospacers included in the docking elements of this invention are preferably as different as possible from any other sequence within the docking element. Therefore, the protospacers included in the docking elements of this invention preferably do not have substantial sequence identity with any other sequence within the docking element. As used herein, the term "X% sequence identity" or "X% of sequence identity" is intended to describe the degree of sequence similarity / homology between two nucleotide sequences, expressed as: the percentage of nucleotides in the first sequence that are identical to corresponding nucleotides in the second sequence when the two sequences are aligned to achieve maximum correspondence. This value is determined by performing an optimal comparison of the two sequences, which may involve introducing gaps in either sequence to achieve optimal alignment. Alignment can be performed using various sequence comparison algorithms or programs known in the art, such as BLASTn, CLUSTALW, or Smith-Waterman. The "percentage of sequence identity" value is calculated by dividing the number of identical nucleotide positions by the total number of nucleotides in the shorter sequence (or the defined comparison segment), and then multiplying by 100. It is used to quantitatively represent the similarity / homology between two nucleotide sequences. Therefore, the sequence identity between the primary spacers included in the docking element and any other sequence in the docking element is less than 33.3% (i.e., a repeating decimal of 3), for example, less than 34%. Preferably, the sequence identity between the primary spacers included in the docking element and any other sequence in the docking element is as low as possible. For example, the sequence identity between the primary spacers included in the docking element and any other sequence in the docking element may be less than 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, or less than 5%. In other words, the sequence identity of the primary spacers included in the docking element of the present invention with any other sequence in the docking element may be no greater than 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6% or greater than 5%.

[0090] Therefore, in one aspect, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, wherein any protospacer does not have more than 34% sequence identity with any other sequence in the docking element.

[0091] The protospacer of the present invention is generated and selected to be as different as possible from any sequence in one or more reference genomes. Therefore, in a preferred aspect, the protospacer included in the docking element of the present invention is not present in one or more reference genomes, i.e., the one or more reference genomes do not contain any sequence identical to the protospacer included in the docking element. Preferably, the protospacer included in the docking element of the present invention does not have substantial sequence identity with any sequence in one or more reference genomes. For example, the sequence of the protospacer included in the docking element may have less than 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 56%, 55%, 54%, 53%, 52%, 51%, 50% of any other sequence in one or more reference genomes. %, 49%, 48%, 47%, 46%, 45%, 44%, 43%, 42%, 41%, 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6% or less of sequence identity. Preferably, the sequence of the protospacer contained in the docking element may have sequence identity with any other sequence in one or more reference genomes of less than 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, or less than 5%. Most preferably, the sequence of the protospacer contained in the docking element may have sequence identity with any other sequence in one or more reference genomes of less than 40%. Preferably, the sequence identity of the protospacer contained in the docking element is as low as possible.In other words, the original spacers included in the docking elements of this invention, and sequences in one or more reference genomes, do not have a concentration greater than 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, or 41%. Sequence identity of 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84% or greater than 85%. Preferably, the protospacer included in the docking element of the present invention does not have a sequence identity greater than 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or greater than 40% with sequences in one or more reference genomes. Most preferably, the protospacer included in the docking element of the present invention does not have a sequence identity greater than 40% with sequences in one or more reference genomes.

[0092] Therefore, in one aspect, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, wherein none of the protospacers are present in one or more reference genomes.

[0093] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, wherein the protospacers do not have more than 40% sequence identity with sequences in one or more reference genomes.

[0094] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, wherein none of the protospacers are present in any of one or more prokaryotic and / or eukaryotic reference genomes.

[0095] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, wherein none of the protospacers are present in one or more reference genomes from the following organisms: *Saccharomyces cerevisiae*, *Yarrowia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces pombe*, *Pichiapastoris*, *Komagataella phaffii*, *Kluveromyces lactis*, *Candida antarctica*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, and *Bacillus thuringiensis*. *Bacillus thuringiensis*, *Bacillus pumilus*, *Paenibacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechocystis* sp. (PCC 6803), *Synechococcus* sp. (PCC 7002), *Synechococcus elongatus* (PCC 7942), and *Anabaena* sp.The genomes of PCC 7120, *Synechococcus slenderus* UTEX 2973, *Streptomyces coelicolor*, *Streptomyces lividans*, *Clostridium acetobutylicum*, rice (*Oryza sativa*), maize (*Zea mays*), tobacco (*Nicotiana tabacum*), *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells were collected.

[0096] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, wherein none of the protospacers are present in one or more reference genomes from: Saccharomyces cerevisiae, Yarrowialipolytica, Bacillus subtilis, and / or Escherichia coli.

[0097] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, wherein none of the protospacers are present in a combination of reference genomes from: Saccharomyces cerevisiae, Yarrowialipolytica, Bacillus subtilis, and Escherichia coli.

[0098] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, wherein the protospacer does not have more than 40% sequence identity with any sequence of any one of one or more prokaryotic and / or eukaryotic reference genomes.

[0099] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, wherein the protospacer does not have more than 40% sequence identity with any sequence from one or more reference genomes of the following: *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, and *Synechococcus elongata* UTEX 2973. The genomes of Streptomyces azure, Streptomyces cerevisiae, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells.

[0100] The docking element of the present invention may also include specific protospacers that have been identified as particularly advantageous in the present invention, such as protospacers that are low in representativeness in one or more reference genomes and unique within the docking element. Exemplary such protospacers are shown in SEQ ID NO: 122-179 and 307-349038.

[0101] Therefore, in one aspect, the present invention provides a docking element comprising one or more primary spacers and a primary spacer sequence neighbor motif (PAM) at the 3' end of each primary spacer, wherein the one or more primary spacers are selected from SEQ ID NO: 122-179 and 307-349038.

[0102] On the other hand, the present invention provides a docking element comprising a protospacer listed in any one of SEQ ID NO: 122-179 and 307-349038, a protospacer sequence adjacent motif (PAM) at the 3' end of each protospacer, and a spacer of 1 to 10 nucleotides (preferably a spacer of 7 nucleotides adjacent to the 3' end of the PAM) optionally adjacent to the 3' end of the PAM.

[0103] The present invention also provides exemplary docking elements / CRISPRpads that have been generated according to the present invention. These exemplary docking elements are specifically designed based on different reference genomes, such as *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK cells. Genomes of 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells. Exemplary docking elements are provided in SEQ ID NO: 1 and 2. For example, the docking elements defined in SEQ ID NO: 1 and 2 are generated based on a combination of reference genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis, and Escherichia coli.

[0104] Therefore, one aspect of the present invention provides a docking element comprising or composed of any one of the nucleic acids listed in SEQ ID NO: 1 and 2.

[0105] On the other hand, the present invention provides a docking element comprising or composed of any one of the nucleic acids listed in SEQ ID NO: 349039 – 4572163 and / or 4572167 – 4687816.

[0106] Generally, the docking elements of the present invention are encoded in nucleic acids, such as DNA or RNA. The docking elements of the present invention are preferably DNA. This is advantageous because DNA can be readily inserted into host cells. The nucleic acid encoding or containing the docking element of the present invention may also contain homologous arms that may be attached to the 5' and 3' ends of the nucleic acid encoding or containing the docking element. The homologous arm is a nucleic acid sequence, such as DNA, that is complementary to the nucleic acid at the target site to facilitate insertion via homologous recombination. Thus, the homologous arm may contain a nucleic acid sequence homologous to a locus in the host cell genome, wherein the locus is the location where the docking element or the nucleic acid encoding the docking element is to be inserted. The present invention does not particularly limit the length of the homologous arm, which may depend on the specific application / use and the locus to be inserted. The length of the homologous arm generally depends on the length required for homologous recombination in the cell where homologous recombination is performed. For example, in *Saccharomyces cerevisiae* and *Escherichia coli* cells, a homologous arm of 35 nucleotides is usually sufficient for targeted insertion via homologous recombination, while in *Bacillus subtilis* it is typically 500 bp, and in *Yarrowia lipolytica* it is typically 1000-2000 bp. According to the present invention, the desired locus is not particularly limited, and any locus requiring the insertion of a docking element or a nucleic acid encoding a docking element can be used. The specific locus may also depend on the application. For example, if it is necessary to express a target gene without affecting surrounding endogenous genes, the desired locus may be a safe harbor locus. However, if it is necessary to knock out an endogenous gene and simultaneously introduce a docking element or a nucleic acid encoding that docking element, the desired locus may also be that endogenous gene. Therefore, the selection of the desired locus may depend on the specific application. According to the present invention, an exemplary desired locus is: Ribosomal RNA (rRNA) loci: Genes encoding ribosomal RNA (rRNA), including 18S, 5.8S, and 25S rRNA, are typically among the most highly expressed genes in cells. These genes are crucial for ribosome assembly and protein synthesis.

[0107] Translation machinery: Genes encoding ribosomal proteins (RPs) and translation factors are highly expressed to support the protein synthesis machinery. Examples include genes encoding ribosomal proteins such as the RPL and RPS genes.

[0108] Glycolysis and fermentation: In glucose-rich environments, genes involved in glycolysis and fermentation pathways tend to be highly expressed. These include genes such as ENO1, ENO2, TPI1, PFK1, and PFK2.

[0109] Heat shock genes: When cells are exposed to heat stress, heat shock genes such as HSP70 and HSP90 are highly expressed to protect cells from protein denaturation.

[0110] Cell cycle regulators: Genes involved in regulating the cell cycle, such as those encoding cyclins and cyclin-dependent kinases (CDKs), are also highly expressed at specific stages of the cell cycle.

[0111] Nutrient transporters: Genes that encode nutrient transporters, such as glucose (e.g., the HXT gene) and amino acid transporters, are highly expressed when these nutrients are scarce, enabling cells to access available resources.

[0112] Stress response genes: Cells can upregulate stress response genes under various stress conditions, including oxidative stress (e.g., SOD1, SOD2), DNA damage (e.g., RAD51), and osmotic stress (e.g., HOG pathway genes).

[0113] Mitochondrial genes: genes that encode proteins involved in mitochondrial respiration and energy production, and are highly expressed when cells need to increase energy output.

[0114] Secretion pathway genes: Genes involved in protein secretion, including genes encoding endoplasmic reticulum (ER) and Golgi apparatus components, are highly expressed when cells actively secrete proteins.

[0115] Generally, desired loci can include highly expressed loci, loci with high genome accessibility, loci for spatially specific expression, loci encoding selection markers, or safe harbor loci. Exemplary desired loci that can be used in this invention include the dppF- locus in *Escherichia coli*, the pksX- locus in *Bacillus subtilis*, the PDC6- locus in *Saccharomyces cerevisiae*, and the URA3- locus in *Yarrowia lipolytica*. Loci with high genome accessibility can be loci with an open chromatin conformation.

[0116] Therefore, in one aspect, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, wherein the docking element is encoded in a nucleic acid, preferably deoxyribonucleic acid (DNA).

[0117] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, wherein the docking element further comprises 5' and 3' homologous arms for insertion into a host cell (e.g., into the host cell genome), the homologous arms comprising sequences homologous to a desired locus in the host cell genome.

[0118] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, wherein the docking element further comprises 5' and 3' homologous arms for insertion into a host cell (e.g., into the host cell genome), the homologous arms comprising sequences homologous to desired loci in the host cell genome, and wherein the length of the one or more protospacers is related to the length of the homologous region required for homologous recombination in the host cell.

[0119] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, wherein the docking element further comprises 5' and 3' homologous arms for insertion into a host cell (e.g., into the host cell's genome), the homologous arms comprising sequences homologous to a desired locus in the host cell's genome, and wherein the desired locus comprises a highly expressed locus, or a locus with high genome accessibility, or a locus for specific spatial expression in the host cell, or a locus encoding a selection marker, or a safe harbor locus.

[0120] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, wherein the docking element further comprises 5' and 3' homologous arms for insertion into a host cell (e.g., into the host cell genome), the homologous arms comprising sequences homologous to a desired locus in the host cell genome, and wherein the desired locus comprises the dppF-locus in Escherichia coli host cells, the pksX-locus in Bacillus subtilis host cells, the PDC6-locus in Saccharomyces cerevisiae host cells, and / or the URA3-locus in Yersinia lipolytica host cells.

[0121] The docking element of the present invention may also include, for example, other elements besides one or more protospacers and PAMs and / or homologous arms described above. This is advantageous, for example, for expressing a target gene (however, the target gene may also be inserted into the docking element after its integration into the cell or cellular genome), or for including a selection marker to verify successful docking element integration. Examples of other elements that may be included in the docking element include promoters, target genes, or selection markers. Generally, other elements can be introduced at any location on the docking element. However, it is advantageous to introduce them at the 5' or 3' end of the docking element. This ensures, for example, in the case of a target gene, that the entire length of the docking element is available for spatial regulation of GOI expression. The same applies to selection markers, which are advantageously placed at the 5' or 3' end of the docking element so that the entire length of the docking element, i.e., all available protospacers, can be utilized as target sites for site-directed nucleases. The target gene may also be combined with a promoter, such as its endogenous promoter (i.e., the endogenous promoter of the target gene) or a synthetic promoter. The present invention is not limited to any particular promoter; the promoter used depends on the desired application. For example, when a large amount of gene product needs to be produced, a high-expression promoter can be used. Exemplary promoters that can be used according to the invention are the TEF1 minimal promoter, LEU, TRE, the tomato SlDFR gene promoter (pSlDFR), CMV, TPGI, TRE3G, MYOD1, and / or HBG1 promoters. Non-limiting examples of suitable eukaryotic promoters (promoters that are functional in eukaryotic cells) include the cytomegalovirus (CMV) immediate early promoter, the herpes simplex virus (HSV) thymidine kinase promoter, the SV40 early and late promoters, long terminal repeat (LTR) sequences of retroviruses, and the mouse metallothionein-1 promoter. The selection of a suitable promoter is within the skill of those skilled in the art. In general, any target gene can be used. This also depends on the desired application, such as protein production or identification of specific gene functions. The target gene can be a fluorophore, a selection marker, a monoclonal antibody, an interferon, and / or a growth hormone. According to the invention, exemplary fluorophores are GFP, RFP, YFP, BFP, or derivatives thereof. According to the present invention, exemplary growth factors are EGF, bFGF, FGF2, HGF, TGF, and PDGF. The target gene can be a gene encoding an enzyme cascade reaction for the production of secondary metabolites, such as a modular nonribosomal peptide synthase or a polyketide synthase. Such target genes can be integrated into the nucleic acid or docking element (CRISPRpad) of the present invention and can be controlled in a variable manner via CRISPRa and / or CRISPRi. This is advantageous in overcoming the problem that such enzyme cascade reaction clusters are often silenced (i.e., not expressed) in their original host.

[0122] Therefore, in one aspect, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence adjacent motif (PAM) at the 3' end of each protospacer, wherein the nucleic acid further comprises a promoter and a target gene, optionally located at the 5' or 3' end.

[0123] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence adjacent motif (PAM) at the 3' end of each protospacer, wherein the nucleic acid further comprises a promoter and a target gene, optionally located at the 5' end or the 3' end, and wherein the promoter is selected from the TEF1 minimal promoter, LEU, TRE, tomato SlDFR gene promoter (pSlDFR), CMV, TPGI, TRE3G, MYOD1 and / or HBG1 promoter.

[0124] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence adjacent motif (PAM) at the 3' end of each protospacer, wherein the nucleic acid further comprises a promoter and a target gene, optionally located at the 5' or 3' end, and wherein the target gene encodes a fluorophore, a selection marker, a monoclonal antibody, an interferon, and / or a growth hormone.

[0125] On the other hand, the present invention provides a docking element comprising one or more protospacers and a protospacer sequence adjacent motif (PAM) at the 3' end of each protospacer, wherein the nucleic acid further comprises a promoter and a target gene, optionally located at the 5' or 3' end, and wherein the target gene encodes GFP, RFP, YFP, BFP or derivatives thereof, and / or growth factors EGF, bFGF, FGF2, HGF, TGF, PDGF.

[0126] The docking elements of this invention may also comprise or consist of the specific sequences provided herein. For example, the inventors have generated specific sequence fragments, protospacers, and docking elements contained in one or more reference genomes, such as low-representation sequences in the reference genomes of the following: *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC7942, *Anabaena* species PCC 7120, and *Synechococcus elongata* UTEX 2973. The genomes of Streptomyces azure, Streptomyces cerevisiae, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells.

[0127] Therefore, in one aspect, the present invention provides a docking element comprising one or more primary spacers and a primary spacer sequence neighbor motif (PAM) at the 3' end of each primary spacer, wherein the one or more primary spacers comprise a 4mer selected from any of Tables 4-47 or Tables 92-116.

[0128] On the other hand, the present invention provides a docking element comprising one or more primary spacers and a primary spacer sequence neighbor motif (PAM) at the 3' end of each primary spacer, wherein the one or more primary spacers comprise a 5mer selected from any of Tables 48-91 or 117-142.

[0129] In another aspect, the present invention provides a docking element comprising one or more primary spacers and a primary spacer sequence neighbor motif (PAM) at the 3' end of each primary spacer, wherein the one or more primary spacers are selected from SEQ ID NO: 122-179 and 307-349038.

[0130] Reference genome

[0131] This invention utilizes one or more reference genomes to identify low-representation sequences within said one or more reference genomes. This ensures that the protospacers and docking elements of this invention do not contain any sequences that may be identical to or substantially sequence-identical to genomic sequences in an organism / cell. This results in protospacers and CRISPRpads that do not cause off-target effects and can be used independently of the species. Typically, according to this invention, a single reference genome can be used to identify low-representation sequences. However, multiple reference genomes can also be combined. This approach can be advantageous when a universal protospacer or docking element is required. For example, reference genomes of species intended to be used as host cells can be combined to ensure that the generated protospacer or docking element is not present in the combination of these genomes and does not have substantial sequence identity with the combination of these genomes. Such protospacers or docking elements can be safely used in all species whose genomes have been combined. For example, the inventors have combined the genomes of *Saccharomyces cerevisiae*, *Yersinia lipophila*, *Bacillus subtilis*, and *Escherichia coli* to generate protospacers of SEQ ID NO: 122-179 and docking elements of SEQ ID NO: 1 and 2. These sequences can be safely used as docking elements in any of *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, and *Escherichia coli* because the protosepta and docking elements are as different as possible from any genomic sequence in the combination. Therefore, no off-target effects occur in these species. Generally, there is no limit to the number of reference genomes used; any number of reference genomes can be used in this invention. The number of reference genomes used can also depend on the desired application. For example, if protosepta or docking elements are needed for a single species, a single reference genome is sufficient. If protosepta or docking elements are needed for multiple species, or if universal protosepta or docking elements are required, multiple reference genomes need to be combined. The inventors have found that low-representation sequences are often associated with phylogenetic relationships. Therefore, for example, if protosepta or docking elements are needed for mammals, multiple mammalian reference genomes can be combined to identify low-representation sequences.

[0132] According to the present invention, the reference genome can be a prokaryotic reference genome. According to the present invention, the reference genome can be a eukaryotic reference genome. An exemplary reference genome according to the present invention may be one or more of the following: *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK cells. Genomes of 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells. The inventors have used, for example, *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, and / or *Escherichia coli* as reference genomes to identify low-representation sequences. Therefore, the reference genome according to the present invention is preferably one or more of the genomes of *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, and / or *Escherichia coli*. Specifically, the present invention has used combinations of reference genomes. Therefore, the reference genome according to the present invention is most preferably a combination of the genomes of *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, and *Escherichia coli*.

[0133] Therefore, one aspect of the present invention provides a nucleic acid comprising a low-representation sequence in one or more prokaryotic and / or eukaryotic reference genomes.

[0134] On the other hand, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein the one or more reference genomes comprise one or more of the following: *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongatus* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongatus* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO cells). Genomes of K1 cells, HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells.

[0135] In a preferred aspect, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein the one or more reference genomes comprise one or more of the genomes of Saccharomyces cerevisiae, Yersinia lipophila, Bacillus subtilis, and / or Escherichia coli.

[0136] On the other hand, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein the one or more reference genomes comprise a combination of two or more prokaryotic and / or eukaryotic reference genomes.

[0137] On the other hand, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein the one or more reference genomes comprise a combination of two or more of the following: *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO cells). Genomes of K1 cells, HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells.

[0138] On the other hand, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein the one or more reference genomes comprise a combination of the genomes of Saccharomyces cerevisiae, Yersinia lipophila, Bacillus subtilis, and Escherichia coli.

[0139] With adaptive modifications, the above content also applies to one or more reference genomes involved in the methods according to the present invention (including methods for generating docking elements, methods for generating non-natural protospacers, and computer-based methods for generating non-natural protospacers). Therefore, in the method, one or more reference genomes can be one or more prokaryotic and / or eukaryotic reference genomes, preferably *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli* genomes, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO cells). Genomes of K1 cells, HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells, preferably one or more of Saccharomyces cerevisiae, Yersinia lipolyticis, Bacillus subtilis, and / or Escherichia coli. A combination of reference genomes may also be used in this method. Therefore, in the method, one or more reference genomes can be a combination of two or more prokaryotic and / or eukaryotic reference genomes, preferably a combination of the following: *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO cells). The genomes of K1 cells, HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells are preferred, and one or more of the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and Escherichia coli are preferred, with the most preferred combination of the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and Escherichia coli.

[0140] Underrepresented sequences or sequence fragments

[0141] Generally, the sequences contained in the nucleic acids, protospacers, or docking elements of this invention can correspond to sequences found in one or more reference genomes. Therefore, sequences that are low-representative in one or more reference genomes are typically nucleic acid sequences present in said reference genomes. As used herein, the term "sequence" or "sequence fragment" refers to a nucleic acid sequence present in one or more reference genomes. In the context of this invention, the terms "low-representative sequence," "lacking-representative sequence," "low-representative sequence fragment," or "lacking-representative sequence fragment" refer to nucleic acid sequences or fragments that are absent or have a low count in one or more reference genomes. As used herein, “low count” means that, compared with all other sequences or sequence fragments of the same length (e.g., 4mer) in one or more reference genomes, the count of the sequence or sequence fragment (i.e., the number of times the sequence or sequence fragment appears in the one or more reference genomes) is less than 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or less than 1%, preferably less than 40%. These sequences can be fragments of one or more reference genomes, such as consecutive fragments of 3 to 20 nucleotides, for example, sequences of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides from one or more reference genomes. These sequences may also be referred to herein as 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, 10mer, 11mer, 12mer, 13mer, 14mer, 15mer, 16mer, 17mer, 18mer, 19mer, or 20mer, respectively. For the purposes of this invention, the sequence or sequence fragment is preferably 4mer or 5mer. These sequences may also be referred to herein as “sequence fragments” of one or more reference genomes. For example, the reference genome can be fragmented into single 4mers to generate sequence fragments. Thus, the first four nucleotides of the genome form the first 4mer, and the second to fifth nucleotides form the second 4mer. This process is repeated throughout the genome or combinations of genomes used. For example, a reference genome could begin with the sequence ATTTGTAAAGGCACACAAAATCCAAA. This sequence contains the following 4mers: Table 2: Examples of 4mer sequence fragments

[0142] All 4mer sequences generated from a reference genome or a combination of multiple genomes can then be sorted according to their counts in the reference genome. For example, this can be achieved by counting the number of times each 4mer appears in the stated genome. In the example in Table 2 above, the 14th and 23rd 4mer sequences appear twice, while the rest appear only once. When this principle is applied to the entire genome or a combination of several genomes, the counts will be higher, and the sorting results may resemble those shown in Table 3.

[0143] Table 3: Examples of 4mer sequence fragment frequencies

[0144] In the example in Table 3, the first through fifth 4mers have the lowest counts in the one or more reference genomes. These are particularly significant in the context of this invention. Conversely, the sixth through tenth 4mers have the highest counts in the one or more reference genomes. These 4mers will be less significant because using them may generate sequences substantially complementary to sequences in the one or more reference genomes. Applying the above to the reference genome will generate a sorted list of 256 4mers, which can be categorized by count. The sequences or sequence fragments with the lowest counts in this list are particularly advantageous in this invention.

[0145] Generally, the present invention preferably includes sequence fragments with low counts. The foregoing applies to all sequences / sequence fragments available according to the present invention, regardless of their length. In particular, the foregoing also applies to sequences or sequence fragments of 3 to 20 nucleotides (e.g., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides) described herein. Therefore, although nucleic acid sequence fragments are present in one or more reference genomes, the protospacers, nucleic acids, and docking elements provided and used in the present invention are the least representative sequence fragments. This avoids the generation of sequences in nucleic acids, protospacers, or docking elements that are complementary to or highly similar to the one or more reference genomes. This means that the sequences / sequence fragments included in the nucleic acids, protospacers, or docking elements according to the present invention are sequences with low counts in the one or more reference genomes compared to all other sequences / sequence fragments of the same length in the one or more reference genomes. For example, if the sequence is a 4mer (as shown in the examples in Tables 2 and 3 above), the 4mer with the lowest count in one or more reference genomes is preferably included in the nucleic acid, protospacer, or docking element of the present invention. Those skilled in the art will understand that "count" can also be expressed as a relative frequency with a comparison value or reference value. Therefore, a low-representative sequence or sequence fragment can be a sequence or sequence fragment with a low count in one or more reference genomes compared to all other sequences or sequence fragments in those genomes. Preferably, a low-representative sequence or sequence fragment is a sequence or sequence fragment with the lowest count in one or more reference genomes compared to all other sequences or sequence fragments in those genomes. For example, a low-representative sequence or sequence fragment can be a sequence or sequence fragment that ranks in the lowest 40% of the k-mer list after sorting by their respective counts. For example, using 4mers, there are 256 different sequences; after sorting them according to their counts, the lowest 40% are selected, thus approximately 104 4mers are selected as low-representative, and the rest are discarded. Similarly, for 5mer, with a threshold of 40%, 409 out of 1024 5mers were selected as low representativeness, and the rest were discarded.Therefore, a low-representation sequence or sequence fragment can be a sequence or sequence fragment that, in terms of the number of occurrences, falls within the range of 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% in the one or more reference genomes compared to all other sequences or sequence fragments of the same length (e.g., 4mer). This means, for example, that if, in one or more reference genomes, a 4mer falls within the range of below 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% in terms of count, then that 4mer is considered low-representative. A low-representation sequence or sequence fragment can be one that, by count, falls within the range of 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or 0.1% compared to all other sequences or sequence fragments of the same length (e.g., 4mer) in one or more reference genomes. Preferably, a low-representation sequence or sequence fragment can be one that, by count, falls within the range of 40% compared to all other sequences or sequence fragments of the same length (e.g., 4mer) in one or more reference genomes and / or combinations thereof. As used herein, “count” refers to the number of times a particular sequence or sequence fragment (e.g., the 4mer example in Tables 2 and 3 above) appears in one or more reference genomes and / or combinations of reference genomes.

[0146] In addition to counting, low-representation sequences can also be defined alternatively or additionally based on their total frequency in one or more reference genomes. The "total frequency" in one or more reference genomes refers to the number of times a sequence or sequence fragment appears in the entirety of the one or more reference genomes, compared to the number of times the most frequently occurring sequence or sequence fragment appears in the one or more reference genomes. For example, the number of times 4mer GGCC appears in the *E. coli* genome can be compared to the number of times the most frequently occurring 4mer appears in that reference genome. For example, if the maximum frequency of 4mer in the reference genome is 0.00577, then the highest frequency of 4mer that can be selected as low-representation could be 0.00403, which is 70% of that maximum frequency in the genome (0.00403 × 100 / 0.00577). For example, this averages to 70% (for 4mer) and 30% (for 5mer) for the 4mers and 5mers provided in Table 4-142, respectively. Therefore, a low-representation sequence or sequence fragment can be a sequence or sequence fragment whose total frequency, relative to the highest observed frequency of sequences or sequence fragments of the same length, is less than 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, or 10% in one or more reference genomes. Preferably, a low-representation sequence or sequence fragment can be a sequence or sequence fragment whose total frequency, relative to the highest observed frequency of sequences or sequence fragments of the same length, is less than 75% or less than 30%.

[0147] The foregoing, with appropriate modifications, applies to sequences or sequence fragments of any length. In particular, the foregoing applies to sequences or sequence fragments of 3 to 20 nucleotides, such as 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides as described herein.

[0148] The nucleic acids, protospacers, or docking elements of the present invention contain low-representation sequences / sequence fragments to ensure that, when inserted into a host organism, the nucleic acids, protospacers, or docking elements do not contain sequences substantially complementary to any sequence in the host organism's genome.

[0149] Table 4-142 lists exemplary sequences or sequence fragments that may be included in the nucleic acids, protospacers, or docking elements of the present invention.

[0150] Therefore, in one aspect, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein a sequence is low-representation in the one or more reference genomes if, in terms of frequency of occurrence, it falls below 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%, preferably below 40%, compared to all other sequences or sequence fragments of the same length.

[0151] On the other hand, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein the sequence is low-representation in the one or more reference genomes if the total frequency of a sequence in the one or more reference genomes is less than 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, or 10% relative to the highest observed frequency of sequences or sequence fragments of the same length.

[0152] In a preferred aspect, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein a sequence is low-representation in the one or more reference genomes if the total frequency of a sequence in the one or more reference genomes is less than 70% or 30% of the highest observed frequency of sequences or sequence fragments of the same length.

[0153] On the other hand, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein the (low-representation) sequence is a fragment of the sequence of the one or more reference genomes, preferably wherein the fragment is a continuous fragment of 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides of the sequence of the one or more reference genomes, and preferably wherein the fragment is a continuous fragment of 4 or 5 nucleotides of the sequence of the one or more reference genomes.

[0154] On the other hand, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein the (low-representation) sequence is 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer or 10mer, preferably wherein the sequence is 4mer or 5mer.

[0155] On the other hand, the present invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes, wherein the sequence is a 4mer and selected from any of Tables 4-47 or 92-116, or wherein the sequence is a 5mer and selected from any of Tables 48-91 or 117-142.

[0156] The present invention also provides low-representation sequences or low-representation sequence fragments that are particularly advantageous for generating the nucleic acids, protospacers and docking elements of the present invention.

[0157] Therefore, one aspect of the present invention provides a nucleic acid comprising or consisting of one or more sequences that are low-representational in one or more reference genomes.

[0158] On the other hand, the present invention provides a nucleic acid comprising or composed of one or more sequences that are low-representative in one or more reference genomes, wherein a sequence is low-representative in the one or more reference genomes if, compared with all other sequences or sequence fragments of the same length, its frequency of occurrence is less than 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2% or less, preferably less than 40%, in percentile terms.

[0159] On the other hand, the present invention provides a nucleic acid comprising or composed of one or more sequences that are low-representative in one or more reference genomes, wherein in the one or more reference genomes, a sequence is low-representative if the total frequency of a sequence is less than 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, or 10%, preferably less than 70% or 30%, relative to the highest observed frequency of sequences or sequence fragments of the same length.

[0160] On the other hand, the present invention provides a nucleic acid comprising or composed of one or more low-representation sequences in one or more reference genomes, wherein the (low-representation) sequence is a fragment of the sequence of the one or more reference genomes, preferably a continuous fragment of 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides of the sequence of the one or more reference genomes, and preferably a continuous fragment of 4 or 5 nucleotides of the sequence of the one or more reference genomes.

[0161] On the other hand, the present invention provides a nucleic acid comprising or composed of one or more low-representation sequences in one or more reference genomes, wherein the (low-representation) sequences are 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer or 10mer, preferably wherein the sequences are 4mer or 5mer.

[0162] On the other hand, the present invention provides a nucleic acid comprising or composed of one or more sequences that are low-representational in one or more reference genomes, wherein the sequences are 4mer and selected from any of Tables 4-47 or Tables 92-116, or wherein the sequences are 5mer and selected from any of Tables 48-91 or Tables 117-142.

[0163] On the other hand, the present invention provides a nucleic acid comprising or composed of one or more sequences of low representativeness in one or more reference genomes, wherein the one or more reference genomes comprise one or more prokaryotic and / or eukaryotic reference genomes, preferably *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC7942, *Anabaena* species PCC 7120, and *Synechococcus elongata* UTEX 2973. The genomes of Streptomyces azureus, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, with the most preferred being one or more of the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli.

[0164] On the other hand, the present invention provides a nucleic acid comprising or composed of one or more sequences of low representativeness in one or more reference genomes, wherein the one or more reference genomes comprise a combination of two or more prokaryotic and / or eukaryotic reference genomes, preferably a combination of two or more of the following: *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, and *Synechococcus elongata* UTEX2973. The genomes of Streptomyces azureense, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells are preferred, and one or more of the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and Escherichia coli are preferred, with the most preferred combination of the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and Escherichia coli.

[0165] Original spacer

[0166] Generally, the nucleic acid or docking element of the present invention comprises one or more protospacers. A protospacer is a polynucleotide sequence, i.e., a nucleic acid, that provides a recognition site / target site for RNA-guided nucleases (e.g., endonucleases) by complementing a spacer in gRNA. The protospacer sequence can bind to / hybridize with a spacer sequence of gRNA containing a complementary sequence. Those skilled in the art will understand that when DNA and RNA hybridize, T hybridizes with U, not A. Therefore, artificial protospacer sequences can be used to generate target sites for RNA-guided nucleases. The protospacers of the present invention may comprise or consist of one or more low-representation sequences or sequence fragments or nucleic acids encoding them as described above. Therefore, one aspect of the present invention provides a protospacer comprising or consisting of one or more nucleic acids according to the present invention, wherein said nucleic acid comprises or consists of one or more low-representation sequences in one or more reference genomes.

[0167] The protospacer may contain two or more low-representation nucleic acids according to the invention. Typically, the protospacer contains two or more low-representation nucleic acids of the invention, wherein said low-representation nucleic acids (e.g., sequence fragments) are adjacent to each other. For example, the protospacer may contain 3-7 (e.g., 3, 4, 5, 6, or 7) adjacent low-representation nucleic acids of the invention. The number of adjacent low-representation nucleic acids contained in the protospacer may depend on (i) the RNA-guided nuclease used to target the protospacer, and / or (ii) the length of nucleic acid required for homologous recombination in the host cell. For example, the protospacer of the invention may contain 3-7 (e.g., 3, 4, 5, 6, or 7) adjacent 4mer or 5mer low-representation nucleic acids of the invention, or composed thereof. Thus, the protospacer may contain 15-35 nucleotides, such as 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides, or composed thereof. Preferably, the protospacer contains at least 16 nucleotides. Most preferably, the protospacer contains 20 nucleotides or is composed of them. Preferably, SpCas9 is used to target the protospacer of the present invention. Thus, the protospacer may contain or be composed of four adjacent 5-mer low-representation nucleic acids of the present invention. Such a protospacer contains or is composed of 20 nucleotides. Generally, the protospacer of the present invention is not present in one or more reference genomes or the genome of a host cell or host organism (unless the cell or organism already contains the artificial protospacer and / or the nucleic acid or docking element containing the protospacer). This ensures that when the protospacer of the present invention is targeted in a host cell or host organism containing the protospacer of the present invention and / or a CRISPR pad, no off-target effects occur in the host cell or host organism. To further avoid any potential off-target effects, the protospacer of the present invention is selected to be as different as possible from one or more reference genomes. Therefore, sequences similar to the protospacer should also not be present in one or more reference genomes and / or host cells or host organisms.Therefore, the protospacer of the present invention, with respect to any sequence in one or more reference genomes and / or any sequence in a host cell or host organism (excluding cells or organisms that already contain the artificial protospacer and / or contain the nucleic acid or docking element of the protospacer), has a value less than 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 5 6%, 55%, 54%, 53%, 52%, 51%, 50%, 49%, 48%, 47%, 46%, 45%, 44%, 43%, 42%, 41%, 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6% or less of sequence identity. Preferably, the protospacer of the present invention has a sequence identity of less than 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, or less than 5% with any sequence in one or more reference genomes and / or with any sequence in a host cell or host organism (excluding cells or organisms that have already contained an artificial protospacer and / or with nucleic acids or docking elements containing the protospacer). Most preferably, the protospacer of the present invention has a sequence identity of less than 40% with any sequence in one or more reference genomes and / or with any sequence in a host cell or host organism (excluding cells or organisms that have already contained an artificial protospacer and / or with nucleic acids or docking elements containing the protospacer). Although the present invention is not limited to any primary spacer, exemplary primary spacers of the present invention are shown in SEQ ID NO: 122-179 and 307-349038.

[0168] A primary spacer may contain 20 nucleotides. A primary spacer may contain fewer than 20 nucleotides. A primary spacer may contain more than 20 nucleotides. A primary spacer may contain at least 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 30 nucleotides or more. A primary spacer may contain at most 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 30 nucleotides or more.

[0169] Therefore, in one aspect, the present invention provides a protospacer comprising or composed of one or more nucleic acids of the present invention, wherein the one or more nucleic acids comprise or composed of one or more sequences that are low-representative in one or more reference genomes, preferably the protospacer comprises or composed of two or more nucleic acids of the present invention, preferably wherein the two or more nucleic acids are adjacent, more preferably, two or more adjacent 4mer or 5mer nucleic acids.

[0170] On the other hand, the present invention provides a protospacer containing or composed of one or more nucleic acids of the present invention, wherein the protospacer contains or composed of 3-7 adjacent 4mer or 5mer nucleic acids, preferably wherein the protospacer contains or composed of 5 adjacent 4mers, or preferably wherein the protospacer contains or composed of 4 adjacent 5mers.

[0171] On the other hand, the present invention provides a protospacer that contains or is composed of one or more nucleic acids of the present invention, wherein the protospacer contains or is composed of about 15-35 nucleotides, preferably about 20 nucleotides.

[0172] On the other hand, the present invention provides a protospacer comprising or composed of one or more nucleic acids of the present invention, wherein the protospacer is not present in one or more reference genomes of the following: *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO cells). The genomes of K1 cells, HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, preferably wherein the protospacer is not present in any of the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and Escherichia coli.

[0173] On the other hand, the present invention provides a protospacer comprising or composed of one or more nucleic acids of the present invention, wherein the protospacer does not have a sequence higher than 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 56%, 55%, or 54% with any sequence in any of one or more prokaryotic and / or eukaryotic reference genomes. 53%, 52%, 51%, 50%, 49%, 48%, 47%, 46%, 45%, 44%, 43%, 42%, 41%, 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, or 5%, preferably not exceeding 40% sequence identity. The reference genome is preferably one or more of the following: *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK cells. Genomes of 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells.In the most preferred embodiment, the protospacer does not have a sequence higher than 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 56%, 55%, 54%, 53%, 52%, 51%, 50%. %, 49%, 48%, 47%, 46%, 45%, 44%, 43%, 42%, 41%, 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6% or 5%, preferably not exceeding 40% sequence identity.

[0174] On the other hand, the present invention provides a protospacer comprising or composed of one or more nucleic acids of the present invention, wherein the protospacer is selected from SEQ ID NO: 122-179 and 307-349038.

[0175] On the other hand, the present invention provides a primary spacer comprising 4mer selected from any one of Tables 4-47 or Tables 92-116.

[0176] On the other hand, the present invention provides a primary spacer comprising 5mer selected from any of Tables 48-91 or Tables 117-142.

[0177] On the other hand, the present invention provides a primary spacer composed of five 4mers selected from any one of Tables 4-47 or Tables 92-114.

[0178] On the other hand, the present invention provides an atomic spacer composed of four 5mers selected from any one of Tables 48-91 or Tables 117-142.

[0179] Nucleic acid

[0180] This invention provides a nucleic acid comprising a low-representation sequence in one or more reference genomes. In this document, such nucleic acid may be referred to as a “dock element” and / or a “CRISPRpad”.

[0181] According to the present invention, "nucleic acid" is a polymer of nucleotides, such as polynucleotides. Nucleic acid can be deoxyribonucleic acid (DNA), the sequence of which is defined by four nucleobases: cytosine (C), guanine (G), adenine (A), or thymine (T). Therefore, one aspect of the present invention provides a DNA containing a low-representation sequence in one or more reference genomes.

[0182] The nucleic acid of the present invention can also be contained in a vector. Therefore, one aspect of the present invention provides a vector comprising the nucleic acid of the present invention, a protospacer, and / or a docking element. This vector can be an expression vector, such as a recombinant expression vector. The recombinant expression vector can be a viral construct, such as a recombinant adeno-associated virus construct, a recombinant adenovirus construct, a recombinant lentivirus construct, a recombinant retrovirus construct, etc. Suitable expression vectors include, but are not limited to, viral vectors (e.g., viral vectors based on vaccinia virus, poliovirus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, human immunodeficiency virus, retrovirus, etc.). The vector can also be a vector optimized for introducing nucleic acid into the cellular genome.

[0183] Those skilled in the art are familiar with a variety of suitable expression vectors, many of which are commercially available. The following vectors are provided as examples suitable for eukaryotic host cells: pXT1, pSGS (Stratagene), pSVK3, pBPV, pMSG, and pSVLSV40 (Phanmacia). However, any other vector may be used as long as it is compatible with the host cell. Depending on the host / vector system used, a variety of suitable transcriptional and translational regulatory elements may be used in the expression vector, including constitutive and inducible promoters, transcriptional enhancer elements, transcription terminators, etc. The vector may also be a plasmid, such as a DNA plasmid.

[0184] Cells and Genomes

[0185] The present invention also provides an organismal genome comprising the nucleic acids, protospacers, or docking elements of the present invention. For example, the nucleic acids, protospacers, or docking elements have been introduced into the genome. Typically, the genome can be a genome as described in the "Reference Genome" section above. These genomes can be corresponding because: when the genome is used as a reference genome as described herein, the nucleic acids, protospacers, or docking elements of the present invention can be specifically designed for use in the genome.

[0186] Therefore, one aspect of the present invention provides a genome comprising the nucleic acids, protospacers, or docking elements of the present invention.

[0187] On the other hand, the present invention provides a genome comprising the nucleic acids, protospacers, or docking elements of the present invention, wherein the genome is a prokaryotic or eukaryotic genome.

[0188] On the other hand, the present invention provides a genome comprising the nucleic acid, protospacer, or docking element of the present invention, wherein the genome is a genome of Bacillus, Escherichia, Saccharomyces, Yarrowia, Clostridium, or Corynebacterium.

[0189] On the other hand, the present invention provides a genome containing the nucleic acids, protospacers, or docking elements of the present invention, wherein the genome is *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, and CHO cells (e.g., CHO cells). Genomes of K1 cells, HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, preferably wherein the genomes are Saccharomyces cerevisiae, Yersinia lipolyticis, Bacillus subtilis and / or Escherichia coli genomes.

[0190] On the other hand, the present invention provides a genome comprising one or more protospacers selected from SEQ ID NO: 122-179 and 307-349038.

[0191] On the other hand, the present invention provides a genome that includes the docking element shown in SEQ ID NO: 1 or 2.

[0192] On the other hand, the present invention provides a genome comprising one or more docking elements selected from any one of SEQ ID NO: 349039 – 4572163 and / or 4572167 – 4687816.

[0193] On the other hand, the present invention provides an Escherichia coli genome containing the nucleic acid shown in SEQ ID NO: 1, 2 or 3.

[0194] On the other hand, the present invention provides a Saccharomyces cerevisiae genome containing the nucleic acid shown in SEQ ID NO: 1, 2 or 4.

[0195] On the other hand, the present invention provides a Yersinia lipophila genome containing the nucleic acid shown in SEQ ID NO: 1, 2, 5 or 6.

[0196] On the other hand, the present invention provides a Bacillus subtilis genome containing the nucleic acid shown in SEQ ID NO: 1, 2, 106 or 107.

[0197] On the other hand, the present invention provides a CHO K1 cell genome containing nucleic acids shown in any one of SSEQ ID NO: 2140375-2213124, 4572162, 4572163 and / or 4572167-4687816.

[0198] On the other hand, the present invention provides a CHO K1 cell line CHO-BXB-SV40 cell genome, which contains nucleic acids listed in any one of SEQ ID NO: 2140375 – 2213124, 4572162, 4572163 and / or 4572167 – 4687816.

[0199] On the other hand, the present invention provides a CHO K1 cell line CHO-BXB-SV40 cell genome containing the nucleic acids shown in SEQ ID NO: 4572162 or 4572163. The CHO K1 cell line CHO-BXB-SV40 preferably corresponds to catalog number INS-SF1029 provided by InSCREENeX GmbH of Braunschweig, Germany.

[0200] Example 2 illustrates an exemplary method for generating cells containing the docking element of the present invention. The docking element containing GFP has been integrated into the genome of a CHO K1 cell. GFP is introduced as a proof of concept to validate... Figure 15The integration is illustrated. Those skilled in the art will recognize that a corresponding cell line without GFP can be constructed, for example, a cell line containing the CRISPRpad of SEQ ID NO: 4572162, the construction of which can be referred to the exemplary description in Example 2 regarding the integration of the CRISPRpad of SEQ ID NO: 4572163. Those skilled in the art can also construct a CHO K1 cell line containing any docking element of SEQ ID NO: 349039 – 4572163 and / or 4572167 – 4687816, preferably any docking element of SEQ ID NO: 2140375 – 2213124, 4572162 and / or 4572167 – 4687816, according to the method described in Example 2. Furthermore, those skilled in the art can generate any target cell, cell line, or organism containing any docking element of the present invention by integrating the docking element of the present invention into the genome, or any other self-replicating genetic element present in the cell, cell line, or organism. Example 2 provides several exemplary methods for inserting docking elements into different target organisms. Mammalian cell lines, such as the CHO K1 cell line, which incorporate the docking element of the present invention, are preferred here. These cell lines can be used to produce therapeutic proteins or antibodies. Specifically, when the docking element of the present invention is inserted into such cell lines, it enables the targeted and dynamic expression of target nucleic acids (e.g., nucleic acids encoding therapeutic proteins or antibodies).

[0201] This invention also provides cells comprising the nucleic acids, protospacers, or docking elements of the invention. For example, the nucleic acids, protospacers, or docking elements may be or have been introduced into the cells. In this document, such cells may be referred to as “cells” or “host cells.” Such cells are particularly advantageous because the nucleic acids, protospacers, or docking elements of the invention can be targeted, for example, to insert into a target gene and / or regulate the expression of a target gene without causing off-target risks in the cell genome. Generally, cells or host cells in the sense of this invention can be cells, host cells, or strains used in fermentation and / or production technologies. For example, cells or host cells in the sense of this invention can be host cells / strains used for amino acid production / amino acid fermentation, including but not limited to *Escherichia coli* or *Corynebacterium*, such as *Corynebacterium glutamicum*. These host cells / strains may also carry gene modifications to achieve enhanced and / or advantageous metabolic fermentation / amino acid production. Such cells / strains with (identified) improvements are known in the art, typically involving the introduction of feedback resistance genes, additional amplification of feedback resistance biosynthetic genes, and / or enhancement of promoters of genes positively regulating amino acid production. However, those skilled in the art will also recognize that cells / strains containing silent genes (e.g., genes encoding negative feedback regulators) can be used for corresponding formation and / or production technologies. According to the invention, such production / fermentation host cells / strains can also (additionally) benefit from the introduction of the nucleic acids, protospacers, and / or docking elements described herein. These nucleic acids, protospacers, or docking elements can be stably integrated into the host cell's genome. Since the nucleic acids, protospacers, or docking elements of the invention are preferably introduced into cells already considered in the selection of low-representative sequences or sequence fragments, the host cell can correspond to one or more reference genomes. Therefore, the host cell can correspond to the organism mentioned in the "Reference Genome" section above. The cells or host cells of the invention can also be edited cells, i.e., non-wild-type cells. Such cells can be cells containing one or more mutations and / or one or more insertions / deletions. Specifically, the cells or host cells of the invention can be edited cells, such as commercially available cells for the production of peptides, monoclonal antibodies, interferons, and / or growth hormones. This is advantageous because the nucleic acids, protospacers, or docking elements of the present invention can be introduced into such cells, thereby providing a template for genome editing without the risk of any off-target effects in the cellular genome. For example, cells or host cells in the sense of the present invention are *Escherichia coli* BL21(DE3), *Bacillus subtilis* 168, *Bacillus subtilis* WB800, *Saccharomyces cerevisiae* S288C, CHO-S, CHO-DG44, and CHO-K1.

[0202] The nucleic acids, protospacers, or docking elements of the present invention may be included in desired loci. Desired loci in the sense of the present invention have been described in detail in the "Docking Elements" section above. The features described in that section, particularly those relating to desired loci, are applicable, with necessary modifications, to host cells containing the nucleic acids, protospacers, or docking elements of the present invention.

[0203] Therefore, one aspect of the present invention provides a host cell that contains or is transformed with the nucleic acid, protospacer, or docking element of the present invention.

[0204] On the other hand, the present invention provides a host cell containing or transformed with the nucleic acid, protospacer, or docking element of the present invention, wherein the nucleic acid, protospacer, or docking element is integrated into the genome of the host cell.

[0205] On the other hand, the present invention provides a host cell that contains or is transformed with the nucleic acid, protospacer, or docking element of the present invention, wherein the host cell is a prokaryotic cell or a eukaryotic cell.

[0206] On the other hand, the present invention provides a host cell containing or transformed with the nucleic acid, protospacer, or docking element of the present invention, wherein the host cell is a cell of the genera Bacillus, Escherichia, Saccharomyces, Yarrowia, Clostridium, or Corynebacterium.

[0207] On the other hand, the present invention provides a host cell containing or transformed with the nucleic acids, protospacers, or docking elements of the present invention, wherein the host cell is *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, or CHO cells (e.g., CHO cells). K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, preferably wherein the host cells are Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli cells.

[0208] On the other hand, the present invention provides a host cell containing or transformed with the nucleic acid, protospacer, or docking element of the present invention, wherein the nucleic acid is contained in a high-expression locus, or a locus with high genome accessibility, or a locus for specific spatial expression in the host cell, or a locus encoding a selection marker, or a safe harbor locus.

[0209] On the other hand, the present invention provides a host cell containing or transformed with the nucleic acid, protospacer, or docking element of the present invention, wherein the nucleic acid, protospacer, or docking element is contained in the dppF-locus of Escherichia coli host cell, the pksX-locus of Bacillus subtilis host cell, the PDC6-locus of Saccharomyces cerevisiae host cell, and / or the URA3-locus of Yeast lipolytica host cell.

[0210] On the other hand, the present invention provides a host cell comprising one or more primary spacers selected from SEQ ID NO: 122-179 and 307-349038.

[0211] On the other hand, the present invention provides a host cell comprising the docking element shown in SEQ ID NO: 1 or 2.

[0212] On the other hand, the present invention provides a host cell comprising one or more docking elements selected from any one of SEQ ID NO: 349039-4572163 and / or 4572167-4687816.

[0213] On the other hand, the present invention provides an *E. coli* host cell containing the nucleic acid shown in SEQ ID NO: 1, 2, or 3. *E. coli* strain BW25113, containing the docking element as defined in SEQ ID NO: 1, was deposited internationally at the German Center for Microbiology and Cell Depository (DSMZ) in Braunschweig, Germany, on October 30, 2023, under the name *Escherichia coli* S26901.3, with accession number DSM34807, in accordance with the Budapest Treaty.

[0214] On the other hand, the present invention provides a brewer's yeast host cell containing the nucleic acid shown in SEQ ID NO: 1, 2 or 4.

[0215] On the other hand, the present invention provides a Yersinia lipolytica host cell containing the nucleic acid shown in SEQ ID NO: 1, 2, 5 or 6.

[0216] On the other hand, the present invention provides a Bacillus subtilis host cell containing the nucleic acid shown in SEQ ID NO: 1, 2, 106 or 107.

[0217] On the other hand, the present invention provides a CHO K1 host cell comprising the nucleic acid shown in any one of SSEQ ID NO: 2140375-2213124, 4572162, 4572163 and / or 4572167-4687816.

[0218] On the other hand, the present invention provides a CHO K1 cell line CHO-BXB-SV40 host cell containing nucleic acids listed in any one of SEQ ID NO: 2140375 – 2213124, 4572162 and / or 4572163.

[0219] On the other hand, the present invention provides a CHO K1 cell line CHO-BXB-SV40 host cell containing the nucleic acid shown in SEQ ID NO: 4572162 or 4572163. In a preferred aspect, the present invention provides a CHO K1 cell line CHO-BXB-SV40 host cell containing the nucleic acid shown in SEQ ID NO: 4572163. The CHO K1 cell line CHO-BXB-SV40 preferably corresponds to catalog number INS-SF1029 provided by InSCREENeX GmbH of Braunschweig, Germany.

[0220] The CHO K1 cell line CHO-BXB-SV40, containing the docking element defined in SEQ ID NO: 4572163, was deposited internationally on October 21, 2024, under the name T949 clone 6, in accordance with the Budapest Treaty at the German Center for Microbiology and Cell Depository (DSMZ) in Braunschweig, Germany, with accession number DSM ACC3380.

[0221] On the other hand, the present invention provides a host cell, which is deposited in DSMZ with accession number DSM ACC3380.

[0222] method

[0223] The present invention also provides methods for generating the docking elements / nucleic acids, protospacers, and cells containing the present invention. Furthermore, the present invention provides methods for regulating the expression of target genes in cells containing the nucleic acids, protospacers, docking elements, vectors, or systems of the present invention.

[0224] Generally, the method for generating the docking element of the present invention includes: (i) Identify sequences that are underrepresented in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Introduce a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer. (iv) Assembling two or more protospacers into a nucleic acid encoding them, wherein the sequence of each protospacer is unique within the nucleic acid. (v) Synthesize nucleic acids encoding the two or more protospacers to generate docking elements.

[0225] Step (i): Identify sequences that are underrepresented in one or more reference genomes.

[0226] The low-representative sequences in (i) can be identified as described in the "Low-representative sequences or sequence fragments" section above. That is, by counting their number or total frequency in one or more reference genomes. Therefore, the features described in the "Low-representative sequences or sequence fragments" section above are adapted to the method for generating docking elements / nucleic acids or protospacers of the present invention. Furthermore, the features of the docking elements / nucleic acids or protospacers of the present invention described in the "Docking Elements," "Protospacers," and "Nucleic Acids" sections are also adapted to the method for generating docking elements / nucleic acids or protospacers. Example 1 below describes in detail how to identify low-representative sequences in one or more reference genomes. Step (i) of the method for generating docking elements / nucleic acids or step (ii) of the method for generating protospacers of the present invention may further include one or more of the following steps or a combination or all of them.

[0227] Create a lookup dictionary for forward and reverse translations between short sequences of length 5 or 3 and integer values. For example: AAAAA = 0; AAAAC = 1, AAAAG = 2 [...].

[0228] Read one or more reference genomes from files such as GenBank or FASTA. These files can be compressed using gzip. When using multiple reference genomes, they can be concatenated before being used as input references.

[0229] If one or more reference genomes are circular genomes, the start region is copied and added to the end of the reference sequence, for example, to linearize it.

[0230] Create a mask for all non-A, C, G, or T positions in the reference sequence, for example, to prevent them from being used as low-representation sequences or sequence fragments.

[0231] Iterate through all positions in the reference genome, extracting all possible 5mers and 3mers starting from each position in the reference sequence, excluding 5mers and 3mers that overlap with previously masked positions, and transform them according to the lookup table / dictionary mentioned above to save memory. Preserve the specific order of the kmers according to the original spacer length, and use the starting position as a position indicator. For example, the sequence GAAAGAGCTTCGGAGACATATTA will be transformed into a list of numbers as follows: [514, 159, 418, 76, 1084]; where GAAAG = 514, AGCTT = 159, and so on.

[0232] The converted reference genome can be stored in a NumPy array for quick access.

[0233] Generate the inverse complementary sequence of the reference sequence, and repeat the masking and transformation process.

[0234] Prepare a mismatch matrix, which consists of all possible pairings of all possible 5mers. The score is the product of the activity scores for each single base pairing, which are derived from Doench JG et al., Optimized sgRNA design to maximize activity and minimize off-target effects of CRISPR-Cas9. Nature biotechnology. 2016 Feb 1;34(2):184-91., doi:10.1038 / nbt.3437. For example, o Sequence 1 = CGAGC o Sequence 2 = CTAAG Starting score = 1.0 Position 1: C – C -> Score = 1.0 1.0 (This calculation can be omitted to save time) Position 2: G – T -> Score = 1.0 0.571428571 = 0.571428571 Position 3: A – A -> Score = 0.571428571 1.0 = 0.571428571 Position 4: G – A -> Score = 0.571428571 0.642857143 = 0.36734693858163264 Position 5: C – G -> Score = 0.36734693858163264 0.388888889 = 0.14285714282256237 The mismatch score for these two 5mers is 0.14285714282256237. The absolute position of each base pair in the 20 nucleotides is crucial. Therefore, the 5mer mismatch matrix needs to be calculated four times: the first 5 nucleotide (1 - 5) = the first 5mer, the second 5 nucleotide (6 – 10) = the second 5mer, and so on. To avoid redundant calculations, the matrix is ​​stored as a compressed table and reloaded.

[0235] Create all possible 4mers (sequences of length 4 bp) and count their occurrences in one or more reference genomes, as described in Tables 2 and 3 in the examples.

[0236] Step (ii): Assemble the low-representative sequences into one or more primary spacers.

[0237] Step (ii) of the method for generating docking elements / nucleic acids or step (iii) of the method for generating protospacers may further include one or more of the following steps or a combination or all of them.

[0238] Based on the lowest assembly count of 4mer (from the previous step), generate 20mer (original spacers) until the total number of designs exceeds a predefined threshold, such as 50,000. This step can be performed as follows (one iteration per CPU, running in parallel): o Use the 4mer counting dictionary generated above If a 4mer exists with a count of 0, then randomly select one; otherwise, sort the 4mers by their counts and randomly select one from the 10 rarest 4mers. o Update the 4mer count dictionary for that single CPU accordingly, adding a new observation, i.e., incrementing the count of the selected 4mer by 1. After this step, the length of the original spacer design is 4 bp. o Repeat this process, selecting four more 4-mers, until the optimal length (e.g., 20 bp) is reached. o Repeat the entire process until 100 unique primary spacer designs are created for this complete single-CPU iteration. Note: When running in parallel on multiple CPUs, all CPUs use the same starting 4mer counting dictionary as the starting point in this iteration.

[0239] After a complete iteration, all designs can be aggregated and duplicates removed.

[0240] The 4-mer counting dictionary, which has not yet been adjusted outside of each independent iteration, can be updated with new count values ​​that include all new designs. These new counts will include redundant designs to avoid generating them repeatedly. For example, if two CPUs run in parallel and both generate designs such as ACGT CGGC TACA TCGG AGGC, then the count in the counting dictionary for each of these k-mers will be increased by 2 instead of 1.

[0241] The above iterations can be repeated on one or more CPUs.

[0242] New designs can be generated incrementally until any of the following conditions are met: o No new and unique primary spacer designs have been produced in the last 10 iterations, or o The last 7 iterations have not produced any unique new designs, and more than 30,000 unique designs have been generated, with a GC content of more than 40% in the reference sequence, or More than 50,000 unique designs have been generated.

[0243] A threshold can be set, which is half the total number of unique designs or 100 times the number of required designs, for example: 50,000 / 2 = 25,000.

[0244] Each 20bp primary spacer design can be divided into four 5mers, and the highest observation frequency (mismatch matrix) of these four 5mers can be determined based on one or more reference sequences, as mentioned above.

[0245] All designs can be grouped according to their maximum 5mer reference frequency.

[0246] These groups are designed to be sorted according to their maximum 5mer reference frequency.

[0247] This generates a filtered list of designs, starting with the group of designs with the lowest maximum frequency and continuing up to the threshold defined above. All designs in that group that exceed the threshold will be included; for example, there could be 25,435 designs instead of 25,000.

[0248] A new threshold could be adopted, which corresponds to half the current number of unique designs or 70 times the number of required designs, for example, 25,000 / 2 = 12,500.

[0249] This process can be repeated using a newly defined threshold for the 5mer frequency of the inverse complementary sequence of the reference sequence.

[0250] Exclude primary spacer designs that exceed a predetermined GC threshold, for example, exclude primary spacers with GC content exceeding 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, and 15%, preferably exclude primary spacers with GC content exceeding 65%.

[0251] Exclude spacer designs that exceed a predetermined homopolymer length threshold, for example, exclude spacers with homopolymer lengths exceeding 2, 3, 4, 5, 6, 7, 8, 9, or 10, preferably exclude spacers with homopolymer lengths of 4 (i.e., 4 bp) or more.

[0252] The off-target activity of all remaining protospacer designs was analyzed by multiplying each 5mer in each protospacer design with its corresponding reference counterpart at each position on one or more reference genomes, based on the mismatch matrix described above. This step is computationally intensive because, for example, with 10,000 designs, there are 10,000 off-target activities. Four 5mers, targeting all locations across a 2 Mb genome, will yield 10,000 4 2,000,000 multiplications = 80,000,000,000 multiplications. Only the maximum value of these products is recorded. This result can be used to evaluate the off-target activity of the protospacer.

[0253] By using pairwise Hamming distance filtering to maximize the diversity of the 20mer design pool: All designs are compared with all other designs; that is, for 10,000 designs, 10,000 comparisons will be performed.2 One value; Hamming distance: The difference between two strings of equal length, calculated by counting the number of positions where corresponding symbols differ; o Next, Ward link calculation (Ward link: a hierarchical clustering method that minimizes the total variance within a cluster, aiming to find the two clusters that minimize the increase in total variance within the cluster after merging), for example, implemented using the Python library SciPy; Next, a dendrogram is calculated based on the link results, for example, also using SciPy; Then, the dendrogram output can be used to extract clusters of similar original spacer designs and return the designs grouped accordingly. Each diverse design becomes an independent cluster with only one member (i.e., the diverse design itself). For each cluster, the primary spacer design with the lowest off-target activity can be selected, and other designs within the cluster can be discarded.

[0254] The output is a sorted list of original interval sublists.

[0255] Although the following steps for evaluating the secondary structure of the protospacer are not essential in the method for generating docking elements / nucleic acids of the present invention, these steps may optionally be included. Step (ii) of the method for generating docking elements / nucleic acids of the present invention or step (ii) of the method for generating protospacers may therefore further include one or more of the following steps, or a combination or all of them.

[0256] We used the Vienna RNAfold software package to predict folds and analyze fold structures to minimize secondary structures that interfere with the activity of protospacers (Lorenz et al., ViennaRNA Package 2.0 Algorithms for Molecular Biology, 6:126, 2011, doi:10.1186 / 1748-7188-6-26).

[0257] The remaining 20mer candidate designs were evaluated and ranked based on the following criteria: Firstly, the design avoids folding: If the last 5 nucleotides contain at least 3 G or C, the first sorting score is: On-site activity score (1 in the new design) - Off-site activity. If there are fewer than 3 Gs or Cs, the first ranking score is: in situ activity score (1 in the new design). 0.8 - Dislocation activity o Next is the design that folds successfully (therefore ranking lower): For these designs, the in-situ-dislocation activity score was calculated, and an additional 0.8 factor was included if there were fewer than 3 Gs or Cs in the last 5 nucleotides.

[0258] This method allows for the screening out of designs with high exotropy activity while maximizing in-situ activity. Furthermore, high GC content at the 3' end of each design is also given priority.

[0259] Similarly, the output here is a sorted list of original interval sublists.

[0260] Step (iii): Introduce a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer.

[0261] As described in the section "Protospacer Sequence Neighbor Motif (PAM)" above, a PAM can be introduced. The features described in that section, with adaptive modifications, are applicable to the methods of the present invention for generating docking elements / nucleic acids and for generating protospacers. In some embodiments, step (iii) includes: (a) Introduce a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, or (b) Introduce a protospacer adjacent motif (PAM) and a linker of 1 to 10 nucleotides adjacent to the 3' end of the PAM at the 3' end of each protospacer, preferably a linker of 7 nucleotides adjacent to the 3' end of the PAM, or (c) Introduce a linker of 3 to 10 nucleotides (preferably 10 nucleotides) containing a protospacer adjacent motif (PAM) at the 3' end of each protospacer, or (d) Introduce a 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, or 10mer sequence containing the neighboring motif (PAM) of the protospacer sequence at the 3' end of each protospacer, preferably a 4mer or 5mer sequence, or (e) Introduce at the 3' end of each protospacer another (ii) protospacer or another (ii) incomplete protospacer containing a protospacer sequence adjacent motif (PAM), wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the protospacer of (ii), preferably 10 nucleotides of the protospacer of (ii).

[0262] In some embodiments, a PAM may also be introduced in step (iv) (see below) during the assembly of two or more protospacers into a nucleic acid. Table 1 shows exemplary PAMs that may be introduced in the methods of the present invention for generating docking elements / nucleic acids and for generating protospacers.

[0263] Step (iv): Assemble two or more protospacers into the nucleic acid encoding them.

[0264] Step (iv) of the method for generating docking elements / nucleic acids of the present invention may further include one or more of the following steps, or a combination or all of them.

[0265] If step (ii) produces 2000 or more primitive spacers, the lowest-ranked half of all designs from step (ii) can be discarded to reduce the number of primitive spacers and discard those with poor rankings. If there are fewer than 2000 primitive spacers, no primitive spacers are discarded.

[0266] The remaining primary spacers can be divided into two groups: matrix & connector Connectors can be further divided into positive connector and Reverse Linkers. Typically, protospacers in the matrix assembly can serve as target sites for RNA-guided nucleases, while protospacers in the linker assembly, or a portion thereof, can be used in… matrix PAM is introduced at the 3' end of the original spacer. o matrix Including the original spacer design, at positions 2 and 3 of the second 5mer segment. No GG At positions 3 and 4 of the third 5mer segment No CC ; o Forward connection subgroup Including the original spacer design, which has positions 2 and 3 in the second 5mer segment. GG (When classifying, it takes precedence over reverse connectors; that is, if both conditions are met, the design becomes a forward connector.) o Reverse connection subgroup Including the original spacer design, which has positions 3 and 4 in the third 5mer segment. CC (When classifying, forward connectors take precedence; that is, if both conditions are met, the design becomes a reverse connector.) In this way, the design can be divided into protospacers containing a suitable PAM sequence for CRISPR / Cas to function after the complete protospacer. In other words, the protospacer design can be a linker element containing PAM to connect the individual protospacers together, while providing PAM at the 3' end of the protospacer. Alternatively, PAM can also be introduced as detailed in the "Protospacer Sequence Proximity Motif (PAM)" section above. The features described in that section, with adaptive modifications, are applicable to the method of the present invention for generating docking elements / nucleic acids. See also the description of step (iii) above in this regard.

[0267] For each matrix design and each linker design (forward and reverse), matrix + linker pairings are recorded, requiring the fourth 5mer of the matrix to perfectly match the first 5mer of the linker. Linker designs in the reverse linker set are converted to their reverse complementary sequences before the matching step. This overlap avoids the accidental introduction of extra sequences that could form unwanted protospacers.

[0268] If a matrix design does not ultimately match any linker, the matrix design can be recorded as a right-side dead end (this applies to the last element of a CRISPRpad because: adjacent to this last matrix design, any nucleotide sequence can be designed that requires PAM for the dead end design).

[0269] For the left-side connectors, the same process can be repeated, which means comparing all current pairings consisting of matrix + right-side connectors with all connector designs to generate the following structure: Linker + Matrix + Linker This structure can be called Full pairing .

[0270] However, to check for proper overlap, the first 5mer segment of the matrix needs to perfectly match the fourth 5mer segment of the connector. Furthermore, before checking for matching 5mer segments, reverse complementary sequences can be generated only for the forward connectors.

[0271] Similarly, for matrix + connector combinations where a suitable left connector cannot be identified, it can be considered a left dead end.

[0272] Since only successful "matrix-connector" combinations have been checked so far, the next step is to iterate through all right-hand "dead ends" (matrix designs only) and check for matches when any connector is used as a left-hand connector, attempting to add connectors on the left to construct a "connector + matrix" structure. For a match to occur, the fourth 5mer segment of the connector needs to perfectly match the first 5mer segment of the (dead end) matrix. As before, before matching, the reverse complementary sequence of the forward connectors is generated.

[0273] Next, if the connectors connected to the left and right dead ends do not match any complete pairing on their respective sides, they are removed, for example: matrix + connector <mismatch> connector + matrix + connector, etc.

[0274] Now, starting with all left-hand dead-end matrix + connector pairs, iterate through all complete pairs to generate all possible combinations, and repeat this process until there are no unused complete pairs and no more unique combinations can be derived. Here, each complete pair can only be used once per CRISPRpad. Adjacent connectors need to perfectly match their two adjacent 5mer segments. For example, if connector 1 is 1a-1b-1c-1d and connector 2 is 2a-2b-2c-2d, then 1c = 2a and 1d = 2b. The reason behind this is that 1a (a 5mer segment) is already identical to the last 5mer segment of the previous protospacer. Similarly, 2d is identical to the first 5mer segment of the next protospacer.

[0275] For example, the following structure can be created: matrix – incomplete connector – incomplete connector – matrix – incomplete connector – incomplete connector – matrix…

[0276] This assembly process can run in parallel, collecting / assembling up to 10 CRISPRpads from each process into nucleic acids. Each left dead end initiates a process. Therefore, for example, starting with 20 left dead ends, a design of 200 CRISPRpads can be generated.

[0277] CRISPRpads shorter than 90% of the longest CRISPRpad design length can be filtered out.

[0278] Entropy Filter: Based on normalized Shannon entropy, using 5mer of the original spacer design contained in each specific CRISPRpad, a value between 0 (low entropy or value) and 1 (high entropy or value) is generated. All designs are then sorted, allowing further processing of the 100 CRISPRpad designs with the highest entropy.

[0279] Next, each design can be scored based on the highest off-target activity of any protospacer contained within that design, considering the entire sequence of the CRISPRpad itself. The highest off-target activity of all protospacers in the CRISPRpad across one or more reference genomes can be calculated. This value can be used as a threshold. The off-target activity of all protospacers relative to that CRISPRpad can then be compared to this threshold, and the number of protospacer designs that do not exceed this threshold can be counted. This avoids deterioration of off-target activity due to the use of the CRISPRpad design in question. Next, the ratio of the number of remaining protospacer designs to the total number of protospacer designs in the corresponding CRISPRpad can be calculated. The final score is based on the following formula: o 1 / F S / N o F = the proportion of the original spacer design remaining after threshold filtering o S = the sum of the off-target activities of all remaining protospacers designed in this CRISPRpad after threshold filtering. o N = Number of remaining primary spacers The lower the score, the better the CRISPRpad design.

[0280] Based on its 5mer, the distance matrix between all CRISPRpad designs can be calculated to improve / maximize its diversity.

[0281] The program parameters allow for the selection of CRISPRpad designs based on the number of design requests. First, the CRISPRpad with the lowest score is selected. Then, selections are made for each subsequent design (until the required number of CRISPRpad designs is reached). The distances and average distances of all available designs relative to the selected designs are calculated, and the design with the largest distance from the selected CRISPRpad design is selected.

[0282] Output a sorted list of CRISPRpad sequences.

[0283] Step (v): Synthesize nucleic acids encoding two or more protospacers.

[0284] Typically, nucleic acids encoding two or more protospacers (i.e., docking elements) can be synthesized by any method known in the art. Those skilled in the art are familiar with various techniques for synthesizing synthetic / artificial nucleic acids. See, for example, Hoose, Alex, et al., "DNA synthesis technologies to close the gene writing gap." Nature Reviews Chemistry 7.3 (2023): 144-161, the full text of which is incorporated herein by reference. The above applies to any method provided herein that includes, or optionally includes, a nucleic acid synthesis step.

[0285] Therefore, in one aspect, the present invention provides a method for generating a docking element, the method comprising: (i) Identify sequences that are underrepresented in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Introduce a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, or A protospacer adjacent motif (PAM) and a linker of 1 to 10 nucleotides adjacent to the 3' end of the PAM are introduced at the 3' end of each protospacer. A linker of 7 nucleotides adjacent to the 3' end of the PAM is preferred. A linker of 3 to 10 nucleotides, preferably 10 nucleotides, is introduced at the 3' end of each protospacer. A 3-mer, 4-mer, 5-mer, 6-mer, 7-mer, 8-mer, 9-mer, or 10-mer sequence containing the neighboring motif (PAM) of the protospacer sequence is introduced at the 3' end of each protospacer, preferably a 4-mer or 5-mer sequence, or At the 3' end of each protospacer, another (ii) protospacer or another (ii) incomplete protospacer containing a protospacer sequence adjacent motif (PAM) is introduced, wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the (ii) protospacer, preferably 10 nucleotides of the (ii) protospacer. (iv) Assembling two or more protospacers into a nucleic acid encoding them, preferably wherein the sequence of each protospacer is unique in the nucleic acid. (v) Synthesize nucleic acids encoding the two or more protospacers to generate docking elements.

[0286] In one embodiment, a sequence is considered low-representative in one or more reference genomes if, in terms of frequency of occurrence, it falls below 40% compared to other sequences or sequence fragments of the same length; and / or if, in one or more reference genomes, the total frequency of a sequence is below 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, or 10%, preferably below 70% or 30%, compared to the highest observed frequency of sequences or sequence fragments of the same length. In one embodiment, the (low-representative) sequence is a fragment of the sequence in one or more reference genomes, preferably a continuous fragment of 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides of the sequence in one or more reference genomes, and preferably a continuous fragment of 4 or 5 nucleotides of the sequence in one or more reference genomes. In some embodiments, the (low-representative) sequence is 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, or 10mer, preferably a 4mer or 5mer sequence fragment. In some embodiments, the sequence is 4mer and selected from any of Tables 4-47 or 92-116. In some embodiments, the sequence is 5mer and selected from any of Tables 48-91 or 117-142. In some embodiments, the protospacer contains about 15-35 nucleotides, for example 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides. In a preferred embodiment, the protospacer contains about 20 nucleotides. In some embodiments, the length of one or more protospacers is related to the length of the homologous region required for homologous recombination in the host cell. In some embodiments, step (ii) includes assembling 3-7 (e.g., 3, 4, 5, 6, or 7) 4mers or 5mers into a protospacer. In a preferred embodiment, step (ii) includes assembling 5 4mers or 4 5mers into a 20-nucleotide protospacer. In some embodiments, step (ii) includes assembling 5 4mers selected from any of Tables 4-47 or Tables 92-116 into one or more protospacers, wherein each protospacer consists of 5 4mers. In some embodiments, step (ii) includes assembling 4 5mers selected from any of Tables 48-91 or Tables 117-142 into one or more protospacers, wherein each protospacer consists of 4 5mers.In some embodiments, step (iv) includes assembling approximately 2 to 60 protospacers, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 protospacers, into a nucleic acid encoding the protospacer. In a preferred embodiment, step (iv) includes assembling approximately 30 protospacers into a nucleic acid encoding the protospacer. In the most preferred embodiment, step (iv) includes assembling approximately 13 protospacers into a nucleic acid encoding the protospacer. In some embodiments, none of the protospacers has a sequence identity greater than 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, or greater than 5%, preferably greater than 34%, with another sequence within the docking element. In some embodiments, none of the protospacers is present in one or more reference genomes. In some implementations, no protospacer has a sequence greater than 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 56%, 55%, 54%, 53%, 52%, 51%, 50%, 4 9%, 48%, 47%, 46%, 45%, 44%, 43%, 42%, 41%, 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, or greater than 5% sequence identity. In a preferred embodiment, no protospacer has more than 40% sequence identity with sequences in one or more reference genomes. In some embodiments, two or more protospacers are assembled into a nucleic acid comprising about 38-2500 nucleotides (preferably about 600 nucleotides, most preferably about 840 or 900 nucleotides).In some embodiments, two or more protospacers are assembled into a nucleic acid comprising about 46-2500 nucleotides. In some embodiments, one or more assembled protospacers are selected from SEQ ID NO: 122-179 and 307-349038. In some embodiments, two or more protospacers are assembled into a nucleic acid shown in any one of SEQ ID NO: 1-2 or 349039-4572163 and / or 4572167-4687816. In some embodiments, two or more protospacers shown in SEQ ID NO: 122-149 and 150-179 are respectively assembled into nucleic acid sequences shown in SEQ ID NO: 1 and 2. In some embodiments, no protospacer is present in one or more prokaryotic and / or eukaryotic reference genomes, preferably one or more of the following reference genomes: *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, and *Synechococcus elongata* UTEX 2973. The genomes of Streptomyces azureense, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells, preferably, none of the protosepta are present in any of the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis, and Escherichia coli.In some embodiments, any protospacer does not share more than 10% sequence identity with any sequence in one or more prokaryotic and / or eukaryotic reference genomes, preferably one or more of the following: *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, and *Synechococcus elongata* UTEX 2973. The genomes of *Streptomyces azureense*, *Streptomyces limoninus*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells are preferred. Ideally, any protospacer from any of these cells should not have a sequence higher than 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 56%, 5 5%, 54%, 53%, 52%, 51%, 50%, 49%, 48%, 47%, 46%, 45%, 44%, 43%, 42%, 41%, 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, or higher than 5%, preferably higher than 40% sequence identity. In some embodiments, the nucleic acid is deoxyribonucleic acid (DNA). In some embodiments, the method further includes attaching 5' and 3' homologous arms to the nucleic acid for insertion into a host cell, for example, into the genome of a host cell, wherein the homologous arms contain sequences homologous to desired sites in the host cell genome. Homologous arms are described in detail in the "Dock Elements" section above. The aforementioned features, after adaptive modifications, are applicable to the method of generating docking elements / nucleic acids of the present invention.In some embodiments, the desired locus includes a highly expressed locus, or a locus with high genome accessibility, or a locus for spatial expression in a specific host cell, or a locus encoding a selection marker or a safe harbor locus. Which loci can be used to integrate docking elements / nucleic acids is described in detail in the "Docking Elements" section above. The features described therein, particularly those relating to loci (e.g., desired loci), are adaptively modified to suit the docking element / nucleic acid generation method of the present invention. In some implementations, the host cells are *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC7120, *Synechococcus elongata* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK cells. 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells, preferably wherein the host cells comprise *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, and / or *Escherichia coli* cells. Host cells suitable for generating the docking elements / nucleic acids of the present invention have been described in detail in the "Cells and Genomes" section above. The features described herein are adaptively modified for use in the docking element / nucleic acid generation method of the present invention. In some embodiments, the desired loci are the dppF locus in *E. coli* host cells, the pksX locus in *Bacillus subtilis* host cells, the PDC6 locus in *Saccharomyces cerevisiae* host cells, and / or the URA3 locus in *Yersinia lipolytica*. In some embodiments, the method further includes attaching a nucleic acid sequence encoding a promoter and a target gene to the 5' or 3' end of the nucleic acid of the present invention. Other elements that may be included in the docking elements / nucleic acids of the present invention are described in detail in the "Docking Elements" section above. These other elements may be introduced / attached in the method for generating the docking elements / nucleic acids. Therefore, the features described herein are adaptively modified to suit the docking element / nucleic acid generation method of the present invention. In some embodiments, the promoter is selected from the TEF1 minimal promoter, LEU, TRE, tomato SlDFR gene promoter (pSlDFR), CMV, TPGI, TRE3G, MYOD1, and / or HBG1 promoter. In some embodiments, the target gene encodes a fluorophore, a selection marker, a monoclonal antibody, interferon, and / or growth hormone.In some embodiments, the target gene encodes GFP and / or growth factors EGF, bFGF, FGF2, HGF, TGF, and PDGF. In some embodiments, the nucleic acid comprises the nucleic acid of the present invention or the docking element of the present invention. In some embodiments, the protospacer comprises the protospacer of the present invention. In some embodiments, the low-representation sequence comprises the nucleic acid of the present invention, such as the nucleic acid sequences / fragments shown in Table 4-142.

[0287] On the other hand, the present invention also provides a nucleic acid generated by a method for generating docking elements.

[0288] The present invention also provides a method for generating protospacers. This method is particularly advantageous because it can generate non-natural (i.e., protospacers that are as different as possible from one or more reference genomes) protospacers. These protospacers can be used for various purposes, such as generating unique target sites or for docking elements, such as the nucleic acids or docking elements of the present invention. Generally, the method for generating the protospacers of the present invention corresponds to the docking element generation method described in detail above, except that the protospacers are not assembled into nucleic acids containing two or more protospacers (step (iv) above), and the introduction of PAMs is optional. The introduction of PAMs can be optional because the protospacers can be used in the context of pre-existing PAMs, such as in specific genomic loci or synthetic DNA, or they can subsequently be included in nucleic acids with pre-existing PAMs. Therefore, the features described in steps (i), (ii), and optional steps (iii) and / or (v) of the method for generating docking elements of the present invention are equally applicable to the method for generating the protospacers of the present invention with adaptive modifications.

[0289] Therefore, in one aspect, the present invention provides a method for generating primary spacers, the method comprising: (i) Identify sequences that are underrepresented in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Introduce a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, or A protospacer adjacent motif (PAM) and a linker of 1 to 10 nucleotides adjacent to the 3' end of the PAM are introduced at the 3' end of each protospacer. A linker of 7 nucleotides adjacent to the 3' end of the PAM is preferred. A linker of 3 to 10 nucleotides, preferably 10 nucleotides, is introduced at the 3' end of each protospacer. A 3-mer, 4-mer, 5-mer, 6-mer, 7-mer, 8-mer, 9-mer, or 10-mer sequence containing the neighboring motif (PAM) of the protospacer sequence is introduced at the 3' end of each protospacer, preferably a 4-mer or 5-mer sequence, or At the 3' end of each protospacer, another (ii) protospacer or another (ii) incomplete protospacer containing a protospacer sequence adjacent motif (PAM) is introduced, wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the (ii) protospacer, preferably 10 nucleotides of the (ii) protospacer. (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids.

[0290] On the other hand, the present invention provides a method for generating primary spacers, the method comprising: (i) Identify sequences that are underrepresented in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Introduce a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids.

[0291] In some implementations, step (iii) includes: (a) Introduce a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, or (b) Introduce a protospacer adjacent motif (PAM) and a linker of 1 to 10 nucleotides adjacent to the 3' end of the PAM at the 3' end of each protospacer, preferably a linker of 7 nucleotides adjacent to the 3' end of the PAM, or (c) Introduce a linker of 3 to 10 nucleotides, preferably a linker of 10 nucleotides, at the 3' end of each protospacer containing a protospacer adjacent motif (PAM). (d) Introduce a 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, or 10mer sequence containing the neighboring motif (PAM) of the protospacer sequence at the 3' end of each protospacer, preferably a 4mer or 5mer sequence, or (e) Introduce at the 3' end of each protospacer another (ii) protospacer or another (ii) incomplete protospacer containing a protospacer sequence adjacent motif (PAM), wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the protospacer of (ii), preferably 10 nucleotides of the protospacer of (ii).

[0292] Table 1 shows exemplary PAMs that can be incorporated into the methods for generating docking elements / nucleic acids and the methods for generating protospacers in this invention.

[0293] On the other hand, the present invention provides a method for generating primary spacers, the method comprising: (i) Identify sequences that are underrepresented in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids. If, in one or more reference genomes, a sequence, in terms of the frequency of its occurrence, falls within the range of 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, or 10% or less compared to other sequences or sequence fragments of the same length, then preferably 40%. If the total frequency of a sequence is below 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, or 10%, preferably 70% or 30%, compared to the highest observed frequency of sequences or sequence fragments of the same length in one or more reference genomes, then the sequence is low representative in one or more reference genomes.

[0294] On the other hand, the present invention provides a method for generating primary spacers, the method comprising: (i) Identify sequences that are underrepresented in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids. The (low-representative) sequence is a fragment of the sequence of the one or more reference genomes, preferably a continuous fragment of 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides of the sequence of the one or more reference genomes, and more preferably a continuous fragment of 4 or 5 nucleotides of the sequence of the one or more reference genomes.

[0295] On the other hand, the present invention provides a method for generating primary spacers, the method comprising: (i) Identify sequences that are underrepresented in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids. The (low-representative) sequence is 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, or 10mer, preferably 4mer or 5mer.

[0296] On the other hand, the present invention provides a method for generating primary spacers, the method comprising: (i) Identify sequences that are underrepresented in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, one or more protospacers containing PAM are synthesized into nucleic acids. The sequence is 4mer and is selected from either Table 4-47 or Table 92-116.

[0297] On the other hand, the present invention provides a method for generating primary spacers, the method comprising: (i) Identify sequences that are underrepresented in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids. The sequence is 5mer and is selected from either Table 48-91 or 117-142.

[0298] On the other hand, the present invention provides a method for generating primary spacers, the method comprising: (i) Identify 4mers or 5mers that are underrepresented in one or more reference genomes. (ii) Assemble low-representation 4mer or 5mer into one or more primary spacers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids.

[0299] In a preferred aspect, the present invention provides a method for generating primary spacers, the method comprising: (i) Identify 4mers that are underrepresented in one or more reference genomes. (ii) Assemble five low-representation 4mers selected from either Table 4-47 or Table 92-116 into one or more primary spacers, wherein each primary spacer consists of five 4mers.

[0300] (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and

[0301] (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids.

[0302] In another preferred aspect, the present invention provides a method for generating primary spacers, the method comprising: (i) Identify 4mers that are underrepresented in one or more reference genomes. (ii) Assemble four low-representation 5mers, each selected from either Table 48-91 or Table 117-142, into one or more primary spacers, wherein each primary spacer consists of four 5mers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids.

[0303] With regard to the method for generating the protospacers of this invention, one or more reference genomes may be any reference genome described in the "Reference Genomes" section herein. The features described herein are adaptively modified to suit the method for generating the protospacers of this invention. Therefore, in some implementations, one or more reference genomes in the method for generating protospacers can be *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC7120, *Synechococcus elongata* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK cells. The genomes of 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, or combinations thereof, preferably wherein the genomes are Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli genomes, more preferably wherein the genomes are a combination of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and Escherichia coli genomes.

[0304] This invention also provides methods for generating cells comprising the nucleic acids, docking elements, protospacers, vectors, or systems of this invention. Introducing the nucleic acids, docking elements, protospacers, vectors, or systems of this invention into cells, for example, by introducing docking elements into the cell genome, is particularly advantageous. The docking elements within the cell can then be targeted by RNA-guided endonucleases, for example, to introduce a target gene and / or regulate the expression of the target gene (e.g., using CRISPRa or CRISPRi), without the risk of off-target effects. Furthermore, the nucleic acids, docking elements, protospacers, vectors, or systems of this invention can be introduced into cells of different species because they are specifically selected to contain protospacers that are as different as possible from the endogenous sequences in the genome. Typically, cells comprising the nucleic acids, docking elements, protospacers, vectors, or systems of this invention can be generated by administering the nucleic acids, docking elements, protospacers, vectors, or systems of this invention to cells. The nucleic acids, docking elements, protospacers, or vectors of this invention can then be integrated into the cell genome, for example, through homologous recombination. Example 2 illustrates an exemplary method for generating cells containing the docking elements of the present invention, wherein several types of cells containing the docking elements of SEQ ID NO: 1 and 2 have been generated, for example, E. coli cells containing the docking element of SEQ ID NO: 1 have been generated. Methods for introducing nucleic acids into host cells are known in the art, and these methods can be used to introduce nucleic acids / docking elements into cells. Suitable methods include, for example, viral or bacteriophage infection, transfection, conjugation, protoplast fusion, liposome transfection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery, etc. Methods for generating cells containing the nucleic acids, docking elements, protospacers, vectors, or systems of the present invention can be in vitro or ex vivo methods.Generally, any cell can be modified to include the docking element of the present invention; however, the cells are preferably Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis, Escherichia coli, Schizosoma sacchari, Pichia pastoris, Saccharomyces phagnum, Kluyveromyces lactis, Candida antarctica, Pseudomonas putida, Corynebacterium glutamicum, Bacillus licheniformis, Bacillus thuringiensis, Bacillus pumilus, Bacillus polymyxa, Lactococcus lactis, Agrobacterium tumefaciens, Synechococcus species PCC 6803, Synechococcus species PCC 7002, Synechococcus elongata PCC 7942, Anabaena species PCC 7120, Synechococcus elongata UTEX 2973, Streptomyces azure, Streptomyces cerevisiae, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK cells. 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, with Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli cells being the most preferred. The present invention also provides cells that have been generated, obtained, or are available through methods for producing cells comprising the nucleic acids, docking elements, protospacers, carriers, or systems of the present invention.

[0305] Therefore, in one aspect, the present invention provides a method for generating cells comprising the nucleic acids, docking elements, protospacers, vectors, or systems of the present invention, wherein the method comprises applying the nucleic acids, protospacers, docking elements, vectors, or systems of any of the present invention to the cells.

[0306] On the other hand, the present invention provides a method for generating cells containing docking elements, wherein the method includes applying the nucleic acid, docking element, vector or system of the present invention to the cells.

[0307] On the other hand, the present invention provides a method for generating cells comprising the nucleic acids, docking elements, protospacers, vectors, or systems of the present invention, wherein the method comprises applying the nucleic acids, protospacers, docking elements, vectors, or systems of the present invention to cells, wherein the nucleic acids, protospacers, docking elements, or vectors are integrated into the genome of the host cell.

[0308] On the other hand, the present invention provides a method for generating cells comprising the nucleic acids, docking elements, protospacers, carriers, or systems of the present invention, wherein the method comprises applying the nucleic acids, protospacers, docking elements, carriers, or systems of any one of the present invention to cells, wherein the cells are *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC7942, *Anabaena* species PCC 7120, and *Synechococcus elongata* UTEX 2973. Streptomyces azureense, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, preferably wherein the host cells are Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli cells.

[0309] On the other hand, the present invention provides an in vitro or ex vivo method for generating cells comprising the nucleic acid, docking element, protospacer, carrier or system of the present invention, wherein the method comprises applying the nucleic acid, protospacer, docking element, carrier or system of any of the present invention to the cells.

[0310] On the other hand, the present invention provides an in vitro or ex vivo method for generating cells comprising the nucleic acid, docking element, protospacer, vector, or system of the present invention, wherein the method comprises applying the nucleic acid, protospacer, docking element, vector, or system of the present invention to the cells, wherein the nucleic acid, protospacer, docking element, or vector is integrated into the genome of the host cell.

[0311] On the other hand, the present invention provides an in vitro or ex vivo method for generating cells comprising the nucleic acids, docking elements, protospacers, carriers, or systems of the present invention, wherein the method comprises applying the nucleic acids, protospacers, docking elements, carriers, or systems of any one of the present invention to cells, wherein the cells are *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, and *Synechococcus elongata* UTEX 2973. Streptomyces azureense, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, preferably wherein the host cells are Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli cells.

[0312] On the other hand, the present invention provides a cell that can be obtained, acquired, or generated by a method for generating a cell comprising the nucleic acid, docking element, protospacer, carrier, or system of the present invention, or by an in vitro or ex vivo method for generating a cell comprising the nucleic acid, docking element, protospacer, carrier, or system of the present invention.

[0313] As described above, the present invention can be used to regulate the expression of a target gene in cells. Therefore, the present invention also provides a method for regulating the expression of a target gene in cells containing the nucleic acid, protospacer, docking element, vector, or system of the present invention. Regulated expression refers to increasing or decreasing the expression of a target gene compared to, for example, endogenous expression. In the context of transgenes, regulated expression refers to increasing or decreasing expression compared to expression without any regulation (e.g., expression induced by a promoter for expressing the transgene). Typically, the expression of an endogenous target gene can be regulated by introducing the nucleic acid or docking element of the present invention into the vicinity of the target gene in the genome (e.g., at its transcription start site or upstream of the promoter). The expression of any target gene (including transgenes) can be achieved by the following steps: first, introducing the nucleic acid or docking element of the present invention into the desired locus, and then introducing the target gene into the nucleic acid or docking element of the present invention through homologous recombination. The target gene is preferably inserted into the 5' or 3' end of the nucleic acid or docking element of the present invention, for example, into the position of the protospacer closest to the 5' or 3' end within the docking element, or into the position of the two outermost protospacers at the 5' or 3' end within the docking element. This is advantageous because, after insertion, the remaining protospacer in the nucleic acid or docking element can be targeted to regulate the expression of the integrated target gene, for example, via CRISPRa or CRISPRi. Therefore, a method for regulating the expression of a target gene in a cell may include the following steps: First, introducing the target gene (e.g., encoded by a donor polynucleotide) into the nucleic acid or docking element of the present invention contained in the cell. The “donor polynucleotide” or “donor template” is a nucleic acid that can be inserted into or will be inserted into, for example, the cellular genome, or into the nucleic acid / docking element of the present invention contained in said genome. Site-directed polypeptides, such as DNA endonucleases, can introduce double-strand breaks or single-strand breaks into the nucleic acid (e.g., genomic DNA). Double-strand breaks can stimulate endogenous DNA repair pathways in the cell (e.g., homology-dependent repair (HDR), non-homologous end joining (NHEJ), alternative non-homologous end joining (A-NHEJ), or microhomology-mediated end joining (MMEJ)). NHEJ repairs broken target nucleic acids without a homologous template. This can sometimes lead to small deletions or insertions (indels) at the break site of the target nucleic acid, potentially disrupting or altering gene expression. HDR, also known as homologous recombination (HR), can occur when a homologous repair template or donor is available. The homologous donor template has a sequence homologous to the sequence flanking the break site of the target nucleic acid (=homologous arm). Cells typically use sister chromatids as repair templates. However, for genome editing purposes, repair templates are often provided in the form of exogenous nucleic acids, such as plasmids, duplex oligonucleotides, single-stranded oligonucleotides, double-stranded oligonucleotides, or viral nucleic acids (=donor polynucleotides).For exogenous donor templates, a common practice is to introduce additional nucleic acid sequences (e.g., transgenes or target endogenous genes) or modifications (e.g., single- or multi-base changes or deletions) between homologous flanking regions to integrate the additional or altered nucleic acid sequence into a target locus, such as a cellular genome or a nucleic acid / docking element of the present invention contained within said genome. Therefore, in some cases, homologous recombination is used to insert a donor polynucleotide into a target nucleic acid break site in a cellular genome or a nucleic acid / docking element of the present invention contained within said genome. Hereinafter, the exogenous polynucleotide sequence is referred to as a donor polynucleotide or donor or donor sequence or polynucleotide donor template or donor template. In some embodiments, a donor polynucleotide, a portion of a donor polynucleotide, a copy of a donor polynucleotide, or a portion of a copy of a donor polynucleotide is inserted into a target nucleic acid break site, such as a cellular genome or a nucleic acid / docking element of the present invention contained within said genome. In some embodiments, the donor polynucleotide is an exogenous polynucleotide sequence, i.e., a sequence not naturally present at the target nucleic acid break site, such as an endogenous target gene or any transgene.

[0314] The method for regulating the expression of a target gene in cells containing the nucleic acid, protospacer, docking element, vector, or system of the present invention can be performed in vitro or ex vivo.

[0315] Therefore, one aspect of the present invention provides a method for regulating the expression of a target gene in a cell, said cell comprising the nucleic acid, protospacer, docking element, vector, or system of the present invention, the method comprising introducing into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, wherein the gRNA is substantially complementary, preferably complementary, to one or more protospacers of the nucleic acids, protospacers, docking elements, vectors, or systems of the present invention.

[0316] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, wherein the gRNAs are substantially complementary, preferably complementary, to one or more spacers in the nucleic acid or docking element of the present invention.

[0317] On the other hand, the present invention provides an in vitro or ex vivo method for regulating the expression of a target gene in cells, wherein the cells contain the nucleic acid, protospacer, docking element, vector, or system of the present invention, and the method includes introducing the following into the cells: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, wherein the gRNA is substantially complementary, preferably complementary, to one or more protospacers of the nucleic acids, protospacers, docking elements, vectors or systems of the present invention.

[0318] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, wherein the gRNA is substantially complementary, preferably complementary, to one or more spacers in the nucleic acid or docking element of the present invention. The donor polynucleotide contains homologous arms at its 5' and 3' ends for insertion into a nucleic acid or docking element, the homologous arms containing sequences homologous to the nucleic acid or docking element.

[0319] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, wherein the gRNA is substantially complementary, preferably complementary, to one or more spacers in the nucleic acid or docking element of the present invention. The donor polynucleotide includes homologous arms at its 5' and 3' ends for insertion into a nucleic acid or docking element. The homologous arms contain sequences homologous to the nucleic acid or docking element, preferably sequences homologous to at least two protospacers contained in the nucleic acid or docking element.

[0320] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, wherein the gRNA is substantially complementary, preferably complementary, to one or more spacers in the nucleic acid or docking element of the present invention. The donor polynucleotide includes homologous arms at its 5' and 3' ends for insertion into a nucleic acid or docking element. The homologous arms contain sequences homologous to one or more protospacers in the nucleic acid or docking element, preferably sequences homologous to at least two protospacers (preferably two protospacers) contained in the nucleic acid or docking element.

[0321] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, wherein the gRNA is substantially complementary, preferably complementary, to one or more spacers in the nucleic acid or docking element of the present invention. The length of the homologous arm is related to the length of the homologous region required for homologous recombination in the cell.

[0322] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, wherein the gRNA is substantially complementary, preferably complementary, to one or more spacers in the nucleic acid or docking element of the present invention. The donor polynucleotide also includes promoters, such as the TEF1 minimal promoter, LEU, TRE, tomato SlDFR gene promoter (pSlDFR), CMV, TPGI, TRE3G, MYOD1 and / or HBG1 promoter.

[0323] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, wherein the gRNA is substantially complementary, preferably complementary, to one or more spacers in the nucleic acid or docking element of the present invention. The cells mentioned are *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, and CHO cells (e.g., CHO). The host cells are K1 cells, HeLa cells, HEK293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, preferably Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli cells.

[0324] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding the gRNA, wherein the gRNA is substantially complementary, preferably complementary, to one or more protospacers in the nucleic acid or docking element of the present invention. Wherein, the one or more gRNAs are substantially complementary to the unique protospacer of the nucleic acid or docking element, preferably complementary; preferably, the one or more gRNAs are substantially complementary to the unique protospacer at the 5' or 3' end of the nucleic acid or docking element, preferably complementary.

[0325] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding the gRNA, wherein the gRNA is substantially complementary, preferably complementary, to one or more protospacers in the nucleic acid or docking element of the present invention. (iii) comprises two gRNAs or sgRNAs or nucleic acids encoding them, wherein the first of the two gRNAs or sgRNAs comprises a spacer sequence substantially complementary, preferably complementary, to the unique protospacer at the 5' end of the nucleic acid or docking element, and wherein the second of the two gRNAs or sgRNAs comprises a spacer sequence substantially complementary, preferably complementary, to the unique protospacer at the 3' end of the nucleic acid or docking element.

[0326] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding the gRNA, wherein the gRNA is substantially complementary, preferably complementary, to one or more protospacers in the nucleic acid or docking element of the present invention. This method also includes introducing the following into the cells: (a) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (b) One or more gRNAs or nucleic acids encoding the gRNA, wherein the gRNA is substantially complementary, preferably complementary, to one or more protospacers in the nucleic acid or docking element, wherein the one or more protospacers are different from the protospacers in (iii). (a) and (b) are introduced into the cell after (i) is integrated into the nucleic acid or docking element at the location targeted by (iii).

[0327] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding the gRNA, wherein the gRNA is substantially complementary, preferably complementary, to one or more protospacers in the nucleic acid or docking element of the present invention. This method also includes introducing the following into the cells: (a) Catalytically inactivated Cas or nucleic acid encoding the Cas, and (b) One or more gRNAs or nucleic acids encoding the gRNA, wherein the gRNA is substantially complementary, preferably complementary, to a unique protospacer in the nucleic acid or docking element, wherein the unique protospacer is different from the protospacer in (iii). (a) and (b) are introduced into the cell after (i) is integrated into the nucleic acid or docking element at the location targeted by (iii).

[0328] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding the gRNA, wherein the gRNA is substantially complementary, preferably complementary, to one or more protospacers in the nucleic acid or docking element of the present invention. The gRNA is not substantially complementary to the endogenous sequence in the cell genome, preferably not complementary.

[0329] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding the gRNA, wherein the gRNA is substantially complementary, preferably complementary, to one or more protospacers in the nucleic acid or docking element of the present invention. Wherein, the gRNA does not have a common affinity for endogenous sequences within the cell genome greater than 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 56%, 55%, 54%, 53%, 52%, 51%, 50%, 49%, 48%, 47%. %, 46%, 45%, 44%, 43%, 42%, 41%, 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6% or greater than 5%, preferably greater than 40% sequence identity.

[0330] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding the gRNA, wherein the gRNA is substantially complementary, preferably complementary, to one or more protospacers in the nucleic acid or docking element of the present invention. The RNA-guided endonuclease is a Cas protein, preferably Cas9, and most preferably Streptococcus pyogenes Cas9 (SpCas9).

[0331] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) The Cas protein or the nucleic acid encoding the Cas protein, and (iii) One or more gRNAs or nucleic acids encoding the gRNA, wherein the gRNA is substantially complementary, preferably complementary, to one or more protospacers in the nucleic acid or docking element of the present invention. The Cas protein mentioned includes catalytically inactivated Cas proteins, such as deadCas9 (dCas9) containing D10A and H840A mutations.

[0332] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) Catalytically inactivated Cas protein or nucleic acid encoding the Cas protein, and (iii) One or more gRNAs or nucleic acids encoding the gRNA, wherein the gRNA is substantially complementary, preferably complementary, to one or more protospacers in the nucleic acid or docking element of the present invention. The catalytically inactivated Cas protein includes transcription activators such as Vp16 transcription activator, Vp64 transcription activator, p65 transcription activator, Rta transcription activator, co-activator mediator (SAM), SunTag activation system, Moontag, HSF1, HSFA6b, AvrXa10, DOF1, DREB1, DREB2, dCas9-TV, p300, CBP, or combinations thereof.

[0333] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) Catalytically inactivated Cas protein or nucleic acid encoding the Cas protein, and (iii) One or more gRNAs or nucleic acids encoding the gRNA, wherein the gRNA is substantially complementary, preferably complementary, to one or more protospacers in the nucleic acid or docking element of the present invention. The catalytically inactivated Cas proteins mentioned therein include transcriptional repressors such as Krüppel-associated box (KRAB), Mxi1, MeCP2 (methyl-CpG-binding protein 2), SIN3A (switch-independent 3A), HDACs (histone deacetylases), Groucho / TLE (transduction protein-like enhancers), NCOR (nuclear receptor co-repressor), NCoR2 (nuclear receptor co-repressor 2), lysine-specific demethylase 1 (LSD1), retinoic acid and thyroid hormone receptor silencing mediator (SMRT), or combinations thereof.

[0334] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more sgRNAs or nucleic acids encoding the sgRNA, wherein the sgRNA is substantially complementary to, preferably complementary to, one or more protospacers in the nucleic acid or docking element of the present invention.

[0335] On the other hand, the present invention provides a method for regulating the expression of a target gene in a cell, the cell containing the nucleic acid or docking element of the present invention, the method comprising introducing the following into the cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or sgRNAs or nucleic acids encoding them, wherein the gRNA is substantially complementary, preferably complementary, to one or more spacers in the nucleic acid or docking element of the present invention. The gRNA or sgRNA contained a spacer selected from SEQ ID No: 180-237.

[0336] system

[0337] The present invention also provides a system for introducing, for example, the nucleic acid, protospacer, or docking element of the present invention into cells. This system generally comprises the nucleic acid, protospacer, or docking element of the present invention (e.g., in the form of a donor polynucleotide), an RNA-guided endonuclease or nucleic acid encoding the enzyme, and one or more gRNAs or nucleic acids encoding the gRNAs. The system can be used to introduce, for example, the nucleic acid, protospacer, or docking element of the present invention into cells, for example, in the form of a donor polynucleotide, wherein the RNA-guided endonuclease forms a complex with the one or more gRNAs and induces one or more DSBs, thereby allowing homologous recombination at target sites in the cellular genome with the nucleic acid, protospacer, or docking element of the present invention in the form of a donor polynucleotide. The RNA-guided endonuclease is preferably a CRISPR protein, such as the Cas protein described in the "CRISPR / Cas" section above. Preferably, the RNA-guided endonuclease comprises a Cas protein, more preferably Cas9, and most preferably SpCas9. The one or more gRNAs can be any gRNA defined in the "Guide RNA (gRNA)" section above. The gRNA is preferably sgRNA. The one or more gRNAs or sgRNAs can form complexes with RNA-guided endonucleases to form the ribonucleoprotein (RNP) complex described in the "Guide RNA (gRNA)" section above. The RNA-guided endonuclease and / or the one or more gRNAs can also be included as nucleic acids in the system of the present invention, i.e., encoded within a nucleic acid. The RNA-guided endonuclease and one or more gRNAs can each be encoded in different nucleic acids, or they can all be encoded in one, i.e., a single nucleic acid. The nucleic acid encoding the RNA-guided endonuclease can be DNA or RNA. For example, the RNA-guided endonuclease can be encoded in mRNA. The nucleic acid encoding one or more gRNAs can be DNA or RNA. For example, one or more gRNAs can exist in the system of the present invention as RNA pre-complexed with the RNA-guided endonuclease. The nucleic acids, protospacers, or docking elements of the present invention can exist in the system in the form of nucleic acids (e.g., DNA forming donor polynucleotides). The nucleic acids, protospacers, or docking elements of the present invention can also be encoded in an adeno-associated virus (AAV) vector. RNA-guided endonucleases or nucleic acids encoding such endonucleases can be formulated alone or together with the gRNA and / or nucleic acids, protospacers, or docking elements (e.g., in the form of donor polynucleotides) of the present invention in liposomes or lipid nanoparticles. This is described in detail in the "Guide RNA (gRNA)" section above. The features described therein are adapted for use in the systems of the present invention. Homologous arms may also be included when the nucleic acids, protospacers, or docking elements of the present invention are included in the system. Homologous arms are described in detail in the "Docking Elements" section above.The features described herein, adapted for use in the system of the present invention, particularly the features described relating to homologous arms, are adapted for use in nucleic acids, protospacers, or docking elements included in the system of the present invention. As previously stated, the system of the present invention can be used to introduce nucleic acids, protospacers, or docking elements into a cell genome. Nucleic acids, protospacers, or docking elements can generally be introduced into any desired locus. For example, nucleic acids, protospacers, or docking elements may contain homologous arms containing sequences homologous to the desired locus. The “Docking Elements” section above describes in detail which loci can be used to integrate nucleic acids, protospacers, or docking elements. The features described in this section, particularly those relating to loci such as desired loci, are adapted for use in the system of the present invention, particularly in nucleic acids, protospacers, or docking elements included in the system. The system of the present invention can be applied to cells / host cells. The “Cells and Genomes” section above describes in detail host cells that can be used in this regard. The features described herein are adapted for use in the system of the present invention.

[0338] Therefore, one aspect of the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them.

[0339] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, The RNA-guided endonuclease is a CRISPR protein.

[0340] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, The RNA-guided endonuclease contains Cas protein, preferably Cas9, and most preferably Streptococcus pyogenes Cas9 (SpCas9).

[0341] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, Wherein, the one or more gRNAs are one or more sgRNAs.

[0342] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or sgRNAs or nucleic acids encoding them, In this process, an RNA-guided endonuclease is pre-complexed with one or more gRNAs or one or more sgRNAs to form a ribonucleoprotein (RNP) complex.

[0343] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, The nucleic acid encoding RNA-guided endonucleases is deoxyribonucleic acid (DNA).

[0344] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, The nucleic acid encoding RNA-guided endonucleases is ribonucleic acid (RNA).

[0345] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, The RNA encoding RNA-guided endonucleases is mRNA.

[0346] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, In this case, the nucleic acid, protospacer, or docking element of (i) is encoded in an adeno-associated virus (AAV) vector.

[0347] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, In this process, RNA-guided endonucleases or the nucleic acids encoding them are formulated in liposomes or lipid nanoparticles.

[0348] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, In this embodiment, an RNA-guided endonuclease or the nucleic acid encoding it is formulated in a liposome or lipid nanoparticle, wherein the liposome or lipid nanoparticle further contains one or more gRNAs or nucleic acids encoding the one or more gRNAs.

[0349] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, Each of (i), (ii) and (iii) is formulated separately in liposomes or lipid nanoparticles, or (i), (ii) and (iii) are formulated together in a single liposome or lipid nanoparticle.

[0350] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or sgRNAs or nucleic acids encoding them, The one or more gRNAs or sgRNAs contain spacer sequences that are substantially complementary (preferably complementary) to loci in the host cell genome.

[0351] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or sgRNAs or nucleic acids encoding them, The one or more gRNAs or sgRNAs contain spacer sequences that are substantially complementary (preferably complementary) to loci in the host cell genome, and the one or more gRNAs or sgRNAs contain spacer sequences that are substantially complementary (preferably complementary) to the original spacer sequence of the nucleic acid (i).

[0352] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, The gRNA or sgRNA contains a spacer sequence that is substantially complementary (preferably complementary) to the unique protospacer of the nucleic acid, protospacer, or docking element of (i).

[0353] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or sgRNAs or nucleic acids encoding them, Wherein, the one or more gRNAs or sgRNAs contain spacer sequences substantially complementary (preferably complementary) to loci in the host cell genome, and wherein the gRNA or sgRNA contains spacer sequences substantially complementary (complementary) to unique protospacers of the nucleic acid, protospacer, or docking element of (i).

[0354] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or sgRNAs or nucleic acids encoding them, (iii) contains at least two gRNAs or sgRNAs or nucleic acids encoding them.

[0355] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or sgRNAs or nucleic acids encoding them, (iii) contains two gRNAs or sgRNAs or the nucleic acids encoding them.

[0356] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or sgRNAs or nucleic acids encoding them, Wherein, (iii) comprises at least two gRNAs or sgRNAs or nucleic acids encoding the gRNAs or sgRNAs, wherein the at least two gRNAs or sgRNAs comprise spacer sequences substantially complementary (preferably complementary) to loci in the host cell genome, wherein one of the at least two gRNAs or sgRNAs comprises a spacer sequence substantially complementary (preferably complementary) to a unique protospacer at the 5' end of the nucleic acid, protospacer, or docking element of (i), and wherein the other of the at least two gRNAs or sgRNAs comprises a spacer sequence substantially complementary (preferably complementary) to a unique protospacer at the 3' end of the nucleic acid, protospacer, or docking element of (i).

[0357] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or sgRNAs or nucleic acids encoding them, Wherein, (iii) comprises two gRNAs or sgRNAs or nucleic acids encoding the gRNAs or sgRNAs, wherein both gRNAs or sgRNAs contain spacer sequences substantially complementary (preferably complementary) to loci in the host cell genome, wherein the first of the two gRNAs or sgRNAs contains a spacer sequence substantially complementary (preferably complementary) to a unique protospacer at the 5' end of the nucleic acid, protospacer, or docking element of (i), and wherein the second of the two gRNAs or sgRNAs contains a spacer sequence substantially complementary (preferably complementary) to a unique protospacer at the 3' end of the nucleic acid, protospacer, or docking element of (i).

[0358] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or sgRNAs or nucleic acids encoding them, The one or more gRNAs or sgRNAs contain spacer sequences that are substantially complementary (preferably complementary) to loci in the host cell genome, and the loci are high-expression loci, or loci with high genome accessibility, or loci for specific spatial expression in the host cell, or loci encoding selection markers or safe harbor loci.

[0359] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or sgRNAs or nucleic acids encoding them, The one or more gRNAs or sgRNAs contain spacer sequences substantially complementary (preferably complementary) to loci in the host cell genome, wherein the host cell is *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, and *Synechococcus elongata* UTEX 2973. Streptomyces azureense, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, preferably wherein the host cells are Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli cells.

[0360] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or sgRNAs or nucleic acids encoding them, The one or more gRNAs or sgRNAs contain spacer sequences that are substantially complementary (preferably complementary) to loci in the host cell genome, wherein the loci are the dppF locus in Escherichia coli host cells, the pksX locus in Bacillus subtilis host cells, the PDC6 locus in Saccharomyces cerevisiae host cells, and / or the URA3 locus in Yersinia lipolytica host cells.

[0361] On the other hand, the present invention provides a system comprising: (i) The nucleic acid, protospacer, or docking element of the present invention, (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them, The gRNA or sgRNA contained a spacer selected from SEQ ID No: 180-237.

[0362] Computer implementation method

[0363] The method provided herein can also be implemented by a computer. Therefore, this invention provides a computer-implemented method for generating docking elements. Typically, the computer-implemented method of this invention for generating docking elements may include steps (i), (ii), (iii), and (iv) of the docking element generation method described in detail in the "Method" section above. The features described therein, particularly those relating to steps (i), (ii), (iii), and (iv) of the docking element generation method, are adapted to the computer-implemented docking element generation method.

[0364] Therefore, one aspect of the present invention provides a computer-implemented method for generating docking elements, the method comprising: (i) Identify sequences that are underrepresented in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Introduce a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, or A protospacer adjacent motif (PAM) and a linker of 1 to 10 nucleotides adjacent to the 3' end of the PAM are introduced at the 3' end of each protospacer. A linker of 7 nucleotides adjacent to the 3' end of the PAM is preferred. A linker of 3 to 10 nucleotides, preferably 10 nucleotides, is introduced at the 3' end of each protospacer. A 3-mer, 4-mer, 5-mer, 6-mer, 7-mer, 8-mer, 9-mer, or 10-mer sequence containing the neighboring motif (PAM) of the protospacer sequence is introduced at the 3' end of each protospacer, preferably a 4-mer or 5-mer sequence, or At the 3' end of each protospacer, another (ii) protospacer or another (ii) incomplete protospacer containing a protospacer sequence adjacent motif (PAM) is introduced, wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the (ii) protospacer, preferably 10 nucleotides of the (ii) protospacer. (iv) Assemble two or more protospacers into a nucleic acid that encodes them, wherein the sequence of each protospacer is unique in the nucleic acid.

[0365] On the other hand, the present invention provides a data processing apparatus comprising means for performing a method for generating docking elements.

[0366] On the other hand, the present invention provides a computer program containing instructions that, when executed by a computer, cause the computer to perform a method for generating docking elements.

[0367] Another aspect of the present invention provides a computer-readable storage medium containing instructions that, when executed by a computer, cause the computer to perform a method for generating docking elements.

[0368] The output of the method is a nucleic acid encoding a docking element. This nucleic acid can then be further used, for example, through synthesis, to obtain a physical product of the docking element.

[0369] Furthermore, the primary spacer generation method of the present invention can be implemented using a computer. Generally, the computer-implemented primary spacer generation method of the present invention may include steps (i), (ii), and optional (iii) and / or (iv) of the primary spacer generation method described in detail in the "Method" section above. The features described therein, particularly those relating to steps (i), (ii), (iii), and (iv) of the primary spacer generation method, are adapted to the computer-implemented primary spacer generation method.

[0370] Therefore, the present invention provides a computer-implemented method for generating primary spacers, comprising: (i) Identify low-representation sequences in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, or A protospacer adjacent motif (PAM) and a linker of 1 to 10 nucleotides adjacent to the 3' end of the PAM are introduced at the 3' end of each protospacer. A linker of 7 nucleotides adjacent to the 3' end of the PAM is preferred. A linker of 3 to 10 nucleotides, preferably 10 nucleotides, is introduced at the 3' end of each protospacer. A 3-mer, 4-mer, 5-mer, 6-mer, 7-mer, 8-mer, 9-mer, or 10-mer sequence containing the neighboring motif (PAM) of the protospacer sequence is introduced at the 3' end of each protospacer, preferably a 4-mer or 5-mer sequence, or Introduce at the 3' end of each protospacer another (ii) protospacer containing a protospacer sequence adjacent motif (PAM) or another (ii) incomplete protospacer, wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the protospacer of (ii), preferably 10 nucleotides of the protospacer of (ii).

[0371] On the other hand, the present invention provides a computer-implemented method for generating primitive spacers, the method comprising: (i) Identify low-representation sequences in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids.

[0372] In some implementations, step (iii) includes: (a) Introduce a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, or (b) Introduce a protospacer adjacent motif (PAM) and a linker of 1 to 10 nucleotides adjacent to the 3' end of the PAM at the 3' end of each protospacer, preferably a linker of 7 nucleotides adjacent to the 3' end of the PAM, or (c) Introduce a linker of 3 to 10 nucleotides, preferably a linker of 10 nucleotides, at the 3' end of each protospacer containing a protospacer adjacent motif (PAM). (d) Introduce a 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, or 10mer sequence containing the neighboring motif (PAM) of the protospacer sequence at the 3' end of each protospacer, preferably a 4mer or 5mer sequence, or (e) Introduce at the 3' end of each protospacer another (ii) protospacer or another (ii) incomplete protospacer containing a protospacer sequence adjacent motif (PAM), wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the (ii) protospacer, preferably 10 nucleotides of the (ii) protospacer.

[0373] Table 1 shows exemplary PAMs that can be incorporated into the computer-implemented docking element generation method and the original spacer generation method of the present invention.

[0374] On the other hand, the present invention provides a computer-implemented method for generating primitive spacers, the method comprising: (i) Identify low-representation sequences in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids. If, in one or more reference genomes, a sequence, in terms of the frequency of its occurrence, falls within the range of 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, or 10% or less compared to other sequences or sequence fragments of the same length, then preferably 40%. If the total frequency of a sequence is below 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, or 10%, preferably 70% or 30%, compared to the highest observed frequency of sequences or sequence fragments of the same length in one or more reference genomes, then the sequence is low representative in one or more reference genomes.

[0375] On the other hand, the present invention provides a computer-implemented method for generating primitive spacers, the method comprising: (i) Identify low-representation sequences in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids. The (low-representative) sequence is a fragment of the sequence of the one or more reference genomes, preferably a continuous fragment of 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides of the sequence of the one or more reference genomes, and more preferably a continuous fragment of 4 or 5 nucleotides of the sequence of the one or more reference genomes.

[0376] On the other hand, the present invention provides a computer-implemented method for generating primitive spacers, the method comprising: (i) Identify low-representation sequences in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids. The (low-representative) sequence is 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, or 10mer, preferably 4mer or 5mer.

[0377] On the other hand, the present invention provides a computer-implemented method for generating primitive spacers, the method comprising: (i) Identify low-representation sequences in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids. The sequence is 4mer and is selected from either Table 4-47 or Table 92-116.

[0378] On the other hand, the present invention provides a computer-implemented method for generating primitive spacers, the method comprising: (i) Identify low-representation sequences in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids. The sequence is 5mer and is selected from either Table 48-91 or Table 117-142.

[0379] On the other hand, the present invention provides a computer-implemented method for generating primitive spacers, the method comprising: (i) Identify low-representation 4mers or 5mers in one or more reference genomes. (ii) Assemble low-representation 4mer or 5mer into one or more primary spacers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids.

[0380] In a preferred aspect, the present invention provides a computer-implemented method for generating primitive spacers, the method comprising: (i) Identify low-representation 4mers in one or more reference genomes. (ii) Assemble five low-representation 4mers, each selected from either Table 4-47 or Table 92-116, into one or more primary spacers, wherein each primary spacer consists of five 4mers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids.

[0381] In another preferred aspect, the present invention provides a computer-implemented method for generating primary spacers, the method comprising: (i) Identify low-representation 4mers in one or more reference genomes. (ii) Assemble four low-representation 5mers, each selected from either Table 48-91 or Table 117-142, into one or more primary spacers, wherein each primary spacer consists of four 5mers. (iii) Optionally, a protospacer sequence neighbor motif (PAM) is introduced at the 3' end of each protospacer, and (iv) Optionally, the one or more protospacers containing PAM are synthesized into nucleic acids.

[0382] With regard to the computer implementation method of the present invention for generating protospacers and the computer implementation method of the present invention for generating docking elements, one or more reference genomes may be any reference genome described in the "Reference Genome" section herein. The features described herein are adaptively modified for use in the protospacer generation method of the present invention. Therefore, in some implementations, one or more reference genomes in the method for generating protospacers can be *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, maize, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK cells. The genomes of 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, or combinations thereof, preferably wherein the genomes are Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli genomes, more preferably wherein the genomes are a combination of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and Escherichia coli genomes.

[0383] On the other hand, the present invention provides a data processing apparatus comprising means for performing a method for generating primitive spacers.

[0384] On the other hand, the present invention provides a computer program containing instructions that, when executed by a computer, cause the computer to perform a method for generating primitive spacers.

[0385] Another aspect of the present invention provides a computer-readable storage medium containing instructions that, when executed by a computer, cause the computer to perform a method for generating primitive spacers.

[0386] The output of the method is a protospacer, i.e., a nucleic acid. The nucleic acid / protospacer can be further used, for example, assembled into docking elements as described above, or synthesized to obtain a physical product of the protospacer, which can be used to generate target sites in a cellular genome or any synthetic / artificial nucleic acid construct.

[0387] Unless otherwise defined, all technical terms, symbols, and other scientific terms used herein are intended to have the meaning commonly understood by one of ordinary skill in the art to which this invention pertains. In some instances, for clarity and / or ease of reference, terms with their commonly understood meanings are defined herein, and these definitions should not necessarily be construed as indicating a difference from the meaning commonly understood in the art. The techniques and procedures described or referenced herein are generally well known to those skilled in the art and are typically performed using conventional methods. Where applicable, procedures involving the use of commercially available kits and reagents are generally performed according to the manufacturer's prescribed protocols and conditions, unless otherwise stated.

[0388] As used herein, unless the context otherwise indicates, the singular forms “a” and “the” include plural indicators. Unless otherwise specified, the terms “including,” “for example,” etc., are intended to indicate a non-restrictive inclusion relationship.

[0389] As used herein, unless the context clearly indicates otherwise, the term "or" is generally used in its common meaning, including "and / or". The term "and / or" refers to one or all of the listed elements, or any combination of two or more of the listed elements.

[0390] As used herein, the term “comprising” also specifically includes implementations that are “composed of” and “mainly composed of”, unless otherwise expressly stated.

[0391] As used herein, the term "about" indicates and covers the value and the range above and below it. In some embodiments, the term "about" indicates a specified value ±10%, ±5%, or ±1%. In some embodiments, where applicable, the term "about" indicates a specified value ± one standard deviation of that value.

[0392] In some embodiments, any of the methods or uses described herein may be performed in vitro and / or in vitro or in vivo. In some embodiments, any method steps described herein may be performed in vitro and / or in vitro or in vivo. In a preferred embodiment, the methods or uses described herein are performed in vitro and / or in vitro.

[0393] The content disclosed within the context of the methods described herein, after necessary modifications, shall be deemed a disclosure for the corresponding use. The content disclosed within the context of the uses described herein, after necessary modifications, shall be deemed a disclosure of the corresponding methods. The content disclosed within the context of the methods of the invention described herein is applicable to the corresponding uses, and vice versa.

[0394] On the one hand, the method of the present invention is not a method for treating human or animal bodies through therapy. On the other hand, the method of the present invention is not a method for altering the genetic characteristics of the human reproductive system. On the one hand, the method of the present invention is an in vitro or ex vivo method. On the other hand, the method of the present invention is non-therapeutic or non-medical. On the other hand, the method of the present invention is not a method for altering the genetic characteristics of animals and may cause suffering to animals without providing any substantial medical benefit to humans or animals.

[0395] Those skilled in the art will understand that, unless otherwise stated, the polynucleotide sequences listed in this application will use "T" in representative DNA sequences; however, when the sequence represents RNA (e.g., mRNA), the "T" should be replaced with "U". Therefore, any RNA polynucleotide encoded by DNA identified by a specific sequence identifier may also contain a corresponding RNA (e.g., mRNA) sequence encoded by that DNA, wherein each "T" in the DNA sequence is replaced with "U".

[0396] The present invention also provides specific low-representation sequences or sequence fragments, primary spacers, and docking elements as described below.

[0397] Low-representative sequences or sequence fragments identified according to the present invention are listed in Tables 4-142. Tables 4-47 contain low-representative 4mers for each species. Tables 48-91 contain low-representative 5mers for each species. Tables 92-116 contain low-representative 4mers for species grouped according to the phylogenetic relationships between low-representative 4mers. Tables 117-142 contain low-representative 5mers for species grouped according to the phylogenetic relationships between low-representative 4mers.

[0398] The protospacers generated according to the present invention are listed in SEQ ID NO: 307-349038. The number of protospacers generated has reached a preset threshold of 100,000. Only the 10,000 protospacers with the lowest scores are provided. The lower the score, the better the performance of the protospacer in the sense of the present invention. An overview of the score distribution is as follows. Figure 14 As shown.

[0399] SEQ ID NO: 307-10306 is a protospacer generated using a low-representation sequence of the Agrobacterium tumefaciens genome according to the present invention.

[0400] SEQ ID NO: 10307 - 20306 are protospacers generated using a low-representation sequence of the Bacillus licheniformis genome according to the present invention.

[0401] SEQ ID NO: 20307 – 30306 are protospacers generated using a low-representation sequence of the Bacillus pumilus genome according to the present invention.

[0402] SEQ ID NO: 30307 – 40306 are protospacers generated using low-representation sequences of the Bacillus subtilis genome according to the present invention.

[0403] SEQ ID NO: 40307 – 50306 are protospacers generated using a low-representation sequence of the Bacillus thuringiensis genome according to the present invention.

[0404] SEQ ID NO: 50307 – 60306 are protospacers generated using low-representation sequences of the CHO cell line genome according to the present invention.

[0405] SEQ ID NO: 60307 – 70306 are protospacers generated using a low-representation sequence of the Clostridium acetobutylicum genome according to the present invention.

[0406] SEQ ID NO: 70307 – 80306 are protospacers generated using a low-representation sequence of the Corynebacterium glutamicum genome according to the present invention.

[0407] SEQ ID NO: 80307 – 90306 are protospacers generated using low-representation sequences of the Escherichia coli genome according to the present invention.

[0408] SEQ ID NO: 90307 - 91781 is a protospacer generated using a low-representation sequence of the Kluyveromyces lactis genome according to the present invention.

[0409] SEQ ID NO: 91782 - 97217 is a protospacer generated using a low-representation sequence of the Komagataella pastoris genome according to the present invention.

[0410] SEQ ID NO: 97218 - 101735 is a protospacer generated using a low-representation sequence of the Komagataellaphaffii genome according to the present invention.

[0411] SEQ ID NO: 101736 - 111735 is a protospacer generated using a low-representation sequence of the Lactococcus lactis genome according to the present invention.

[0412] SEQ ID NO: 111736 - 121735 is a protospacer generated using a low-representation sequence of the genome of *Moesziomyces antarcticus* according to the present invention.

[0413] SEQ ID NO: 121736 - 131735 are protospacers generated using low-representation sequences of the mouse (Mus musculus) genome according to the present invention.

[0414] SEQ ID NO: 131736 - 132457 is a protospacer generated using a low-representation sequence of the Nicotiana abentonhamiana genome according to the present invention.

[0415] SEQ ID NO: 132458 - 142457 is a protospacer generated from a low-representation sequence of the genome of the Nostoc sp. PCC 7120 species according to the present invention.

[0416] SEQ ID NO: 142458 - 152457 are protospacers generated using low-representation sequences of the rice (Oryza sativa) genome according to the present invention.

[0417] SEQ ID NO: 152458 - 162457 are protospacers generated using a low-representation sequence of the genome of Paenibacillus polymyxa according to the present invention.

[0418] SEQ ID NO: 162458 - 172457 is a protospacer generated from a low-representation sequence of the genome of the dermatologic alginate species PCC 7002, according to the present invention.

[0419] SEQ ID NO: 172458 - 182457 is a protospacer generated using a low-representation sequence of the genome of *Pseudomonas putida* according to the present invention.

[0420] SEQ ID NO: 182458 - 184035 is a protospacer generated using a low-representation sequence of the Saccharomyces cerevisiae genome according to the present invention.

[0421] SEQ ID NO: 184036 - 184792 is a protospacer generated using a low-representation sequence of the genome of *Schizosaccharomyces pombe* according to the present invention.

[0422] SEQ ID NO: 184793 - 194792 is a protospacer generated using a low-representation sequence of the SF9 cell line genome according to the present invention.

[0423] SEQ ID NO: 194793 - 204792 is a protospacer generated using a low-representation sequence in the genome of Streptococcus pyogenes according to the present invention.

[0424] SEQ ID NO: 204793 - 214792 is a protospacer generated using a low-representation sequence from the genome of Streptomyces coelicolor, according to the present invention.

[0425] SEQ ID NO: 214793 - 224792 is a protospacer generated using a low-representation sequence in the genome of Streptomyces lividans according to the present invention.

[0426] SEQ ID NO: 224793 - 234792 is a protospacer generated from a low-representation sequence of the genome of Synechococcus elongatus PCC 7942 according to the present invention.

[0427] SEQ ID NO: 234793 - 244792 is a protospacer generated from a low-representation sequence in the genome of Synechococcus elongatus UTEX 2973 according to the present invention.

[0428] SEQ ID NO: 244793 - 254792 is a protospacer generated from a low-representation sequence of the genome of a Synechocystis sp. PCC 6803 species according to the present invention.

[0429] SEQ ID NO: 254793 - 264792 is a protospacer generated using a low-representation sequence of the Vero cell line genome according to the present invention.

[0430] SEQ ID NO: 264793 - 274792 is a protospacer generated using a low-representation sequence from the genome of Yarrowia lipolytica according to the present invention.

[0431] SEQ ID NO: 274793 - 284792 is a protospacer generated using a low-representation sequence from the Zea mays genome according to the present invention.

[0432] SEQ ID NO: 284793 - 294792 are protospacers generated using low-representation sequences from the Arabidopsis thaliana genome according to the present invention.

[0433] SEQ ID NO: 294793 - 304792 is a protospacer generated using a low-representation sequence of the Hi5 cell line genome according to the present invention.

[0434] SEQ ID NO: 304793 - 305278 is a protospacer generated using a low-representation sequence of the tobacco (Nicotiana tabacum) genome according to the present invention.

[0435] SEQ ID NO: 305279 - 315278 is a protospacer generated using low-representation sequences from the genomes of the Streptomyces genus group: Streptomyces azure and Streptomyces leucosus, according to the present invention.

[0436] SEQ ID NO: 315279 - 325228 is a protospacer generated according to the present invention using low-representation sequences of the genomes of the genus *Komagata*: *Komagata Pasteurella* and *Komagata phaf*.

[0437] SEQ ID NO: 325229 - 335228 is a protospacer generated according to the present invention using low-representation sequences of the genomes of mammalian groups: cell lines A549, CHO, HEK293T, HeLa, K562, Vero, Caki-2, Homo sapiens, and house mouse.

[0438] SEQ ID NO: 335229 - 345228 are protospacers generated according to the present invention using low-representation sequences from the genomes of the genus Synechococcus: Synechococcus PCC7942 and Synechococcus UTEX 2973.

[0439] SEQ ID NO: 345229 – 349038 is a protospacer generated using low-representation sequences of the genomes of yeast groups: Kluyveromyces lactis, Saccharomyces cerevisiae, and Schizosaccharomyces cerevisiae, according to the present invention.

[0440] The docking elements generated according to the present invention are listed in SEQ ID NO: 349039-4572163 and 4572167-4687816. These docking elements are generated based on the original spacers of SEQ ID NO: 307-349038. These docking elements have not been scored or optimized for their internal diversity. The number of docking elements generated has reached a preset threshold of 150,000. Currently, only 50% of the generated docking elements are provided.

[0441] SEQ ID NO: 349039 - 513588 is a docking element generated using the above-described Agrobacterium tumefaciens protospacer according to the present invention.

[0442] SEQ ID NO: 513589 - 701269 is a docking element generated using the above-described Corynebacterium glutamicum protospacer according to the present invention.

[0443] SEQ ID NO: 701270 - 756269 is a docking element generated using the above-described house mouse primitive spacer according to the present invention.

[0444] SEQ ID NO: 756270 - 762125 is a docking element generated using the above-described Saccharomyces cerevisiae protospacer according to the present invention.

[0445] SEQ ID NO: 762126 - 906175 is a docking element generated using the above-mentioned slender Synechococcus UTEX 2973 protospacer according to the present invention.

[0446] SEQ ID NO: 906176 - 1044228 is a docking element generated using the above-mentioned Bacillus licheniformis protospacer according to the present invention.

[0447] SEQ ID NO: 1044229 - 1176578 is a docking element generated using the above-mentioned Escherichia coli protospacer according to the present invention.

[0448] SEQ ID NO: 1176579 - 1176948 is a docking element generated using the above-described Benjamin tobacco prime spacer according to the present invention.

[0449] SEQ ID NO: 1176949 - 1178535 is a docking element generated using the above-mentioned Schizosaccharomyces cerevisiae protoseptum according to the present invention.

[0450] SEQ ID NO: 1178536 - 1263774 is a docking element generated using the protospacer of the above-mentioned Synechocystis species PCC 6803 according to the present invention.

[0451] SEQ ID NO: 1263775 - 1456237 is a docking element generated using the above-described Bacillus pumilus protospacer according to the present invention.

[0452] SEQ ID NO: 1456238 - 1481804 is a docking element generated using the above-described Kluyveromyces lactis protospacer according to the present invention.

[0453] SEQ ID NO: 1481805 - 1635504 is a docking element generated using the above-mentioned Nostoc species PCC 7120 primary spacer according to the present invention.

[0454] SEQ ID NO: 1635505 - 1744724 is a docking element generated using the above-described SF9 cell line protospacer according to the present invention.

[0455] SEQ ID NO: 1744725 - 1793524 is a docking element generated using the above-described Vero cell line protospacer according to the present invention.

[0456] SEQ ID NO: 1793525 - 1891424 is a docking element generated using the above-described Bacillus subtilis protospacer according to the present invention.

[0457] SEQ ID NO: 1891425 - 1972624 is a docking element generated using the above-described Pasteurella mongolica protospacer according to the present invention.

[0458] SEQ ID NO: 1972625 - 2140374 is a docking element generated using the above-described protospacer of the *Komagata* genus according to the present invention.

[0459] SEQ ID NO: 2140375 - 2213124 is a docking element generated using the above-described mammalian group protospacer according to the present invention.

[0460] SEQ ID NO: 2213125 - 2348774 is a docking element generated using the above-mentioned Streptomyces genus protospacer according to the present invention.

[0461] SEQ ID NO: 2348775 - 2498706 is a docking element generated using the above-mentioned Synechococcus group protospacer according to the present invention.

[0462] SEQ ID NO: 2498707 - 2624030 is a docking element generated using the above-described Streptococcus pyogenes protoseptum according to the present invention.

[0463] SEQ ID NO: 2624031 - 2718930 is a docking element generated using the above-described Yersinia lipolytica protospacer according to the present invention.

[0464] SEQ ID NO: 2718931 - 2898458 is a docking element generated using the above-mentioned Bacillus thuringiensis protospacer according to the present invention.

[0465] SEQ ID NO: 2898459 - 3024721 is a docking element generated using the above-described yeast group protospacer according to the present invention.

[0466] SEQ ID NO: 3024722 - 3097571 is a docking element generated using the above-described Phaeocotyledon protospacer according to the present invention.

[0467] SEQ ID NO: 3097572 - 3281728 is a docking element generated using the above-mentioned polymyxa protospacer according to the present invention.

[0468] SEQ ID NO: 3281729 - 3402478 is a docking element generated using the above-mentioned Streptomyces cerevisiae protospacer according to the present invention.

[0469] SEQ ID NO: 3402479 - 3535870 is a docking element generated using the above-mentioned lactococcus protospacer according to the present invention.

[0470] SEQ ID NO: 3535871 - 3717942 is a docking element generated using the protospacer of the above-mentioned Synechococcus species PCC7002 according to the present invention.

[0471] SEQ ID NO: 3717943 - 3874892 is a docking element generated using the above-mentioned Streptomyces cerevisiae protospacer according to the present invention.

[0472] SEQ ID NO: 3874893 - 4006652 is a docking element generated using the above-mentioned Clostridium acetone-butanol protospatial spacer according to the present invention.

[0473] SEQ ID NO: 4006653 - 4274652 is a docking element generated using the above-mentioned spacer of *Ustilago maydis* according to the present invention.

[0474] SEQ ID NO: 4274653 - 4426752 is a docking element generated using the above-mentioned *Pseudomonas putida* protospatial spacer according to the present invention.

[0475] SEQ ID NO: 4426753 - 4572161 is a docking element generated using the above-mentioned slender Synechococcus PCC 7942 protospacer according to the present invention.

[0476] SEQ ID NO: 4572162 and 4572163 are docking elements generated according to the present invention for integration into the CHOK1 cell line. SEQ ID NO: 4572162 and 4572163 contain the same docking element, the only difference being that SEQ ID NO: 4572163 also contains a sequence encoding GFP. This GFP sequence is included as a proof of concept to verify that the "docking element" has been successfully integrated into the cell (e.g., detected by flow cytometry).

[0477] SEQ ID NO: 4572167 - 4687816 is a docking element generated using the above-described CHO cell line genomic prostomes according to the present invention.

[0478] The following reference genomes have been used to identify low-representation sequences or sequence fragments.

[0479] agrobacterium_tumefaciens.gb.gz

[0480] Locus: CP033031 3027766 bp DNA Circular BCT 21-OCT-2018

[0481] Definition: Agrobacterium tumefaciens strain 12D1, chromosomal circular, complete sequence.

[0482] Login ID: CP033031

[0483] Version: CP033031.1

[0484] DBLINK BioProject: PRJNA494485

[0485] BioSample: SAMN10169604.

[0486] bacillus_licheniformis.gb.gz

[0487] Locus: CP014842 4136986 bp DNA Circular BCT 29-MAR-2017

[0488] Definition: Chromosome 14 of Bacillus licheniformis strain SCDB, complete sequence

[0489] Login ID: CP014842

[0490] Version: CP014842.1

[0491] DBLINK BioProject: PRJNA293170

[0492] BioSample: SAMN03998272.

[0493] bacillus_pumilus.gb.gz

[0494] Locus: NZ_PTXV01000001 1799065 bp DNA Linear CON 08-DEC-2022

[0495] Definition: Whole genome shotgun sequence of Bacillus pumilus strain Ha06YP001 Contig_01.

[0496] Login ID: NZ_PTXV01000001 NZ_PTXV01000000

[0497] Version: NZ_PTXV01000001.1

[0498] DBLINK BioProject: PRJNA224116

[0499] BioSample: SAMN08568348.

[0500] bacillus_thuringiensis.gb.gz

[0501] Locus: CM000753 6260142 bp DNA circular CON 12-MAR-2015

[0502] Definition: Shotgun sequence of the entire genome of Bacillus thuringiensis ATCC 10792 (Berlin serotype).

[0503] Login ID: CM000753 ACNF01000000

[0504] Version: CM000753.1

[0505] DBLINK BioProject: PRJNA29723

[0506] BioSample: SAMN00738287.

[0507] CHO_cell_line.gb.gz

[0508] Locus: RAZU02000001 275698159 bp DNA Linear ROD 06-JUN-2020

[0509] Definition: Shotgun sequence of the entire genome of the Chinese hamster (Cricetulus griseus) strain 17A / GY, chromosome 1 chr1_0.

[0510] Login ID: RAZU02000001 RAZU02000000

[0511] Version: RAZU02000001.1

[0512] DBLINK BioProject: PRJNA389969

[0513] BioSample: SAMN07140313.

[0514] clostridium_acetobutylicum.gb.gz

[0515] Locus: NZ_CP030018 3957627 bp DNA circular CON 13-APR-2023

[0516] Definition: Chromosome of Clostridium acetone-butanol strain LJ4

[0517] Login ID: NZ_CP030018

[0518] Version: NZ_CP030018.1

[0519] DBLINK BioProject: PRJNA224116

[0520] BioSample: SAMN09240565

[0521] Assembly: GCF_003254825.1.

[0522] corynebacterium_glutamicum.gb.gz

[0523] Locus: CP004048 3350619 bp DNA Circular BCT 31-JAN-2014

[0524] Definition: Corynebacterium glutamicum SCgG2, complete genome

[0525] Login ID: CP004048

[0526] Version: CP004048.1

[0527] DBLINK BioProject: PRJNA175241

[0528] BioSample: SAMN02603964.

[0529] Hi5 cell line

[0530] Locus: CM010751 23865964 bp DNA Linear CON 24-SEP-2018

[0531] Definition: Shotgun sequence of the entire genome of the isolated ovarian cell line Hi5 from the white-spotted armyworm.

[0532] Login ID: CM010751 NKQN01000000

[0533] Version: CM010751.1

[0534] DBLINK BioProject: PRJNA336361

[0535] BioSample: SAMN07304761.

[0536] kluyveromyces_lactis.gb.gz

[0537] Locus: CR382121 1062590 bp DNA Linear PLN 27-FEB-2015

[0538] Definition: Complete sequence of chromosome A of Kluyveromyces lactis strain NRRL Y-1140

[0539] Login ID: CR382121

[0540] Version: CR382121.1

[0541] DBLINK BioProject: PRJNA13835

[0542] BioSample: SAMEA3138170.

[0543] Pasteurella mongolica

[0544] Locus: CP121226 2798491 bp DNA Linear PLN 10-APR-2023

[0545] Definition: Chromosome 1 of Pasteurella morganii isolate EAMORG09

[0546] Login ID: CP121226

[0547] Version: CP121226.1

[0548] DBLINK BioProject: PRJNA942376

[0549] BioSample: SAMN33688472.

[0550] komagataella_phaffii.gb.gz

[0551] Locus: FN392319 2798491 bp DNA Linear PLN 27-FEB-2015

[0552] Definition: Pichia pastoris GS115, chromosome 1, complete sequence

[0553] Login ID: FN392319

[0554] Version: FN392319.1

[0555] DBLINK BioProject: PRJEA37871

[0556] BioSample: SAMEA2272385.

[0557] lactococcus_lactis.gb.gz

[0558] Locus: CP059048 2426597 bp DNA Circular BCT 13-OCT-2021

[0559] Definition: Chromosome and complete genome of Lactococcus lactis strain LAC460

[0560] Login ID: CP059048

[0561] Version: CP059048.1

[0562] DBLINK BioProject: PRJNA645372

[0563] BioSample: SAMN15502559.

[0564] moesziomyces_antarcticus.gb.gz

[0565] Locus: DF830264 1009 bp DNA Linear CON 29-OCT-2014

[0566] Definition: Pseudozyma antarctica DNA, scaffold: scaffold_v10197, strain JCM10317, whole genome shotgun sequence.

[0567] Login ID: DF830264 BBIZ01000000

[0568] Version: DF830264.1

[0569] DBLINK BioProject: PRJDB2910

[0570] BioSample: SAMD00018595.

[0571] Little House Mouse (downloaded on 12.8.2020)

[0572] Description: Genome Reference Consortium Mouse Construct 39

[0573] Name of the creature: House mouse (or house mouse)

[0574] Seed Name: Strain: C57BL / 6J

[0575] BioProject: PRJNA20689

[0576] Submitted by: Genome Reference Consortium

[0577] Date: 2020 / 06 / 24

[0578] Alias: mm39

[0579] Assembly level: chromosomes

[0580] Genome representativeness: complete

[0581] RefSeq Category: Reference Genome

[0582] GenBank Assembly Registry Number: GCA_000001635.9 (Latest)

[0583] RefSeq assembly login number: GCF_000001635.27 (latest).

[0584] Ben's Tobacco

[0585] Locus: CBMM010000001 1211 bp DNA Linear PLN 27-MAY-2014

[0586] Definition: Data from the Tobacco Benedict WGS project CBMM00000000, contig Ni_ben_LEAF_contig_1, whole genome shotgun sequence.

[0587] Login ID: CBMM010000001 CBMM010000000

[0588] Version: CBMM010000001.1.

[0589] nicotiana_tabacum.gb.gz

[0590] Locus: NW_015787227 26924 bp DNA Linear CON 03-MAY-2016

[0591] Definition: Unlocalized genome scaffold of tobacco cultivar TN90, Ntab-TN90 Ntab-TN90_scaffold10, whole genome shotgun sequence.

[0592] Login ID: NW_015787227

[0593] Version: NW_015787227.1

[0594] DBLINK BioProject: PRJNA319578

[0595] BioSample: SAMN02316627.

[0596] nostoc_sp_pcc_7120.gb.gz

[0597] Locus: NZ_RSCN01000001 428901 bp DNA Linear CON 10-FEB-2023

[0598] Definition: Nostoc species PCC 7120 = FACHB-418 strain PCC 7120 sequence 001, whole genome shotgun sequence.

[0599] Login ID: NZ_RSCN01000001 NZ_RSCN01000000

[0600] Version: NZ_RSCN01000001.1

[0601] DBLINK BioProject: PRJNA224116

[0602] BioSample: SAMN10102199.

[0603] oryza_sativa.gb.gz

[0604] Locus: NC_029256 43270923 bp DNA Linear CON 07-AUG-2018

[0605] Definition: Nipponbare, a japonica rice variety, chromosome 1, IRGSP-1.0

[0606] Login ID: NC_029256

[0607] Version: NC_029256.1

[0608] DBLINK BioProject: PRJNA122

[0609] BioSample: SAMD00000397.

[0610] paenibacillus_polymyxa.gb.gz

[0611] Locus: CP040829 5703931 bp DNA Circular BCT 12-JUN-2019

[0612] Definition: Chromosome and complete genome of Bacillus polymyxa strain ZF129

[0613] Login ID: CP040829

[0614] Version: CP040829.1

[0615] DBLINK BioProject: PRJNA545384

[0616] BioSample: SAMN11890480.

[0617] picosynechococcus_sp_pcc_7002.gb.gz

[0618] Locus: NC_010475 3008047 bp DNA circular CON 08-OCT-2023

[0619] Definition: Species PCC 7002 of Synechococcus dermataceae, complete genome.

[0620] Login ID: NC_010475

[0621] Version: NC_010475.1

[0622] DBLINK BioProject: PRJNA224116

[0623] BioSample: SAMN01081740.

[0624] pseudomonas_putida.gb.gz

[0625] Locus: AP013070 6156701 bp DNA Circular BCT 07-OCT-2016

[0626] Definition: DNA of *Pseudomonas putida* NBRC 14164, complete genome.

[0627] Login ID: AP013070

[0628] Version: AP013070.1

[0629] DBLINK BioProject: PRJDB191

[0630] BioSample: SAMD00061028.

[0631] schizosaccharomyces_pombe.gb.gz

[0632] Locus: CU329670 5579133 bp DNA Linear PLN 27-FEB-2015

[0633] Definition: Chromosome I of *Schizosaccharomyces cerevisiae*, complete sequence

[0634] Login ID: CU329670 AL009197 AL009227 AL021046 AL021809 AL021813 AL021817

[0635] AL031180 AL034486 AL034565 AL034583 AL035064 AL035248 AL035254

[0636] AL035439 AL096845 AL109734 AL109770 AL109820 AL109951 AL109988

[0637] AL110469 AL110509 AL117210 AL117390 AL121732 AL121741 AL121745

[0638] AL121770 AL122032 AL132667 AL132675 AL132714 AL132769 AL132779

[0639] AL132798 AL132828 AL132839 AL133154 AL133225 AL133302 AL133357

[0640] AL133442 AL133498 AL135751 AL136078 AL136235 AL136499 AL136521

[0641] AL136538 AL137130 AL138666 AL138854 AL139315 AL157734 AL157811

[0642] AL157872 AL157917 AL158056 AL159180 AL159951 AL162531 AL162631

[0643] AL163031 AL163071 AL163191 AL163481 AL163529 AL353014 AL353860

[0644] AL355252 AL355452 AL355632 AL356333 AL356335 AL357232 AL358272

[0645] AL360054 AL360094 AL390095 AL390274 AL390814 AL391713 AL391744

[0646] AL391746 AL391783 AL441621 AL441624 AL512491 AL512493 AL512496

[0647] AL512549 AL512562 AL583902 AL590562 AL590582 AL590602 AL590605

[0648] AL672256 AL691405 Z49811 Z50142 Z50728 Z54096 Z54142 Z54285 Z54308

[0649] Z54328 Z54354 Z54366 Z56276 Z64354 Z66568 Z67757 Z67961 Z68136

[0650] Z68144 Z68166 Z68887 Z69086 Z69380 Z69944 Z70043 Z70721 Z81312

[0651] Z81317 Z94864 Z95334 Z97185 Z98056 Z98849 Z98944 Z99091 Z99126

[0652] Z99292 Z99568 Z99753

[0653] Version: CU329670.1

[0654] DBLINK BioProject: PRJNA13836

[0655] BioSample: SAMEA3138176.

[0656] SF9.gb.gz

[0657] Locus: NJHR02000001 30299 bp DNA Linear INV 31-AUG-2021

[0658] Definition: Shotgun sequence of the whole genome of Spodoptera frugiperda isolate Sf91.

[0659] Login ID: NJHR02000001 NJHR02000000

[0660] Version: NJHR02000001.1

[0661] DBLINK BioProject: PRJNA380964

[0662] BioSample: SAMN06658935.

[0663] streptomyces_coelicolor.gb.gz

[0664] Locus: NZ_CP050522 8585093 bp DNA Linear CON 24-JUL-2023

[0665] Definition: Chromosome and complete genome of Streptomyces cerevisiae strain M1154 / pAMX4 / pGP1416.

[0666] Login ID: NZ_CP050522

[0667] Version: NZ_CP050522.1

[0668] DBLINK BioProject: PRJNA224116

[0669] BioSample: SAMN14414788.

[0670] streptomyces_lividans.gb.gz

[0671] Locus: NZ_CP009124 8345283 bp DNA Linear CON 12-MAR-2023

[0672] Definition: Chromosome TK24 of *Streptomyces pulveratum*, complete genome.

[0673] Login ID: NZ_CP009124

[0674] Version NZ_CP009124.1

[0675] DBLINK BioProject: PRJNA224116

[0676] BioSample: SAMN02947304.

[0677] synechococcus_elongatus_pcc_7942.gb.gz

[0678] Locus: CP000100 2695903 bp DNA Circular BCT 27-APR-2022

[0679] Definition: Synechococcus slenderus PCC 7942 = FACHB-805 strain PCC 7942 chromosome, complete genome

[0680] Login ID: CP000100 AADZ01000000 AADZ01000001 AADZ01000002 AADZ01000003AADZ01000004

[0681] Version: CP000100.1

[0682] DBLINK BioProject: PRJNA10645

[0683] BioSample: SAMN02598254.

[0684] synechococcus_elongatus_utex_2973.gb.gz

[0685] Locus: NZ_CP006471 2690418 bp DNA circular CON 14-MAY-2023

[0686] Definition: Chromosome 2973 of Synechococcus slenderus, complete genome

[0687] Login ID: NZ_CP006471

[0688] Version: NZ_CP006471.1

[0689] DBLINK BioProject: PRJNA224116

[0690] BioSample: SAMN03278348.

[0691] synechocystis_sp_pcc_6803.gb.gz

[0692] Locus: NZ_CP073017 3571181 bp DNA circular CON 14-AUG-2023

[0693] Definition: Chromosome PCC 6803, complete genome of a species in the genus *Synthia*.

[0694] Login ID: NZ_CP073017

[0695] Version: NZ_CP073017.1

[0696] DBLINK BioProject: PRJNA224116

[0697] BioSample: SAMN18375591.

[0698] Vero.gb.gz

[0699] Locus: JACDXN010000001 81790585 bp DNA Linear PRI 06-NOV-2020

[0700] Definition: Whole genome shotgun sequence of the green monkey (Chlorocebus sabaeus) strain WHO RCB 10-87 scaffold-1.

[0701] Login ID: JACDXN010000001 JACDXN010000000

[0702] Version: JACDXN010000001.1

[0703] DBLINK BioProject: PRJNA644395

[0704] BioSample: SAMN15458746.

[0705] zea_mays.gb.gz

[0706] Locus: NC_050096 308452471 bp DNA Linear CON 01-SEP-2020

[0707] Definition: Whole genome shotgun sequence of corn cultivar B73 chromosome 1, Zm-B73-REFERENCE-NAM-5.0.

[0708] Login ID: NC_050096

[0709] Version: NC_050096.1

[0710] DBLINK BioProject: PRJNA655717

[0711] BioSample: SAMEA5569141.

[0712] streptococcus_pyogenes.gb.gz

[0713] Locus: NC_002737 1852433 bp DNA circular CON 23-FEB-2023

[0714] Definition: Streptococcus pyogenes M1 GAS, complete sequence

[0715] Login ID: NC_002737 NZ_AE006472-NZ_AE006638 ...

Claims

1. A nucleic acid contained in a low-representation sequence in one or more reference genomes.

2. The nucleic acid as described in claim 1, wherein, If a sequence falls below 40% in one or more reference genomes in terms of frequency of occurrence compared to other sequences or sequence fragments of the same length, then the sequence is considered low-representative in the one or more reference genomes; and / or if, in one or more reference genomes, the total frequency of a sequence is below 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, or 10%, preferably 70% or 30%, compared to the highest observed frequency of sequences or sequence fragments of the same length, then the sequence is considered low-representative in one or more reference genomes.

3. The nucleic acid of claim 1 or 2, wherein the (low-representative) sequence is a fragment of the sequence of the one or more reference genomes, preferably wherein the fragment is a continuous fragment of 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides of the sequence of the one or more reference genomes, and preferably wherein the fragment is a continuous fragment of 4 or 5 nucleotides of the sequence of the one or more reference genomes.

4. The nucleic acid according to any one of claims 1 to 3, wherein the (low-representative) sequence is 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer or 10mer, preferably wherein the sequence is 4mer or 5mer.

5. The nucleic acid of claim 4, wherein the sequence is 4mer and selected from any of Tables 4-47 or 92-116, or wherein the sequence is 5mer and selected from any of Tables 48-91 or 117-142.

6. The nucleic acid according to any one of claims 1-5, wherein adjacent sequences form a protospacer, preferably wherein adjacent 4mer or 5mers form a protospacer, optionally wherein the GC content of the protospacer is less than 65%, and further optionally wherein the protospacer does not contain homopolymers longer than 4 bp.

7. The nucleic acid according to any one of claims 4-6, wherein 3-7 adjacent 4mers or 5mers form a protospacer, preferably 5 adjacent 4mers form a protospacer, or preferably 4 adjacent 5mers form a protospacer.

8. The nucleic acid of claim 6 or 7, wherein the nucleic acid further comprises: (a) The protospacer sequence neighbor motif (PAM) located at the 3' end of each protospacer, or (b) A protospacer adjacent motif (PAM) located at the 3' end of each protospacer and a linker of 1 to 10 nucleotides adjacent to the 3' end of the PAM, preferably a linker of 7 nucleotides adjacent to the 3' end of the PAM, or (c) A linker of 3 to 10 nucleotides, preferably a linker of 10 nucleotides, located at the 3' end of each protospacer and containing a protospacer adjacent motif (PAM). (d) A 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, or 10mer sequence, preferably a 4mer or 5mer sequence, located at the 3' end of each protospacer and containing a protospacer adjacent motif (PAM). (e) Another protospacer or another incomplete protospacer located at the 3' end of each protospacer containing a protospacer sequence adjacent motif (PAM), wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the protospacer, preferably 10 nucleotides of the protospacer.

9. The nucleic acid according to any one of claims 6-8, wherein the nucleic acid comprises one or more, preferably at least two protospacers.

10. The nucleic acid according to any one of claims 6-9, wherein the nucleic acid comprises about 2-60 spacers, preferably about 30, and most preferably about 13.

11. The nucleic acid according to any one of claims 6-10, wherein the protospacer comprises about 15-35 nucleotides, preferably about 20 nucleotides.

12. The nucleic acid according to any one of claims 4-11, wherein five adjacent 4mers or four adjacent 5mers form a 20-nucleotide spacer.

13. The nucleic acid according to any one of claims 6-12, wherein the sequence of each protospacer is unique within the nucleic acid.

14. The nucleic acid according to any one of claims 6-13, wherein, No protospacer has more than 34% sequence identity with another sequence within the nucleic acid.

15. The nucleic acid according to any one of claims 6-14, wherein, No prospacer is present in the one or more reference genomes.

16. The nucleic acid according to any one of claims 6-15, wherein, No protospacer has more than 40% sequence identity with any of the sequences in the one or more reference genomes.

17. The nucleic acid according to any one of claims 1-16, wherein the nucleic acid comprises about 19-2500 nucleotides, preferably about 600 nucleotides.

18. The nucleic acid according to any one of claims 6-17, wherein the one or more protospacers are selected from SEQ ID NO: 122-179 and 307-349038.

19. The nucleic acid of any one of claims 6-18, wherein the nucleic acid comprises one or more of the original spacers shown in SEQ ID NO: 122-179 and 307-349038.

20. The nucleic acid according to any one of claims 6-19, wherein the nucleic acid comprises a protospacer as shown in any one of SEQ ID NO: 122-179 and 307-349038, a protospacer sequence adjacent motif (PAM) at the 3' end of each protospacer, and a linker of 1 to 10 nucleotides optionally adjacent to the 3' end of the PAM, preferably a linker of 7 nucleotides adjacent to the 3' end of the PAM.

21. The nucleic acid according to any one of claims 1-20, wherein the nucleic acid comprises or is composed of the nucleic acid shown in or composed of any one of SEQ ID NO: 1-2 or 349039 – 4572163 and / or 4572167 – 4687816.

22. The nucleic acid according to any one of claims 1-21, wherein the one or more reference genomes comprise one or more prokaryotic and / or eukaryotic reference genomes, preferably one or more of the following: *Saccharomyces cerevisiae*, *Yarrowia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces pombe*, *Pichia pastoris*, *Komagataella phaffii*, *Kluveromyces lactis*, *Candida antarctica*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, and *Bacillus thuringiensis*. *Bacillus thuringiensis*, *Bacillus pumilus*, *Paenibacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechocystis* sp. PCC 6803, *Synechococcus* sp. PCC 7002, *Synechococcus elongatus* PCC 7942, and *Anabaena* sp. PCC 7120.The genomes of *PCC7120*, *Synechococcus elongatus* UTEX 2973, *Streptomyces coelicolor*, *Streptomyces lividans*, *Clostridium acetobutylicum*, rice (*Oryza sativa*), maize (*Zea mays*), tobacco (*Nicotiana tabacum*), *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells, with a preference for one or more genomes from *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, and / or *Escherichia coli*.

23. The nucleic acid according to any one of claims 1-22, wherein the one or more reference genomes comprise a combination of two or more prokaryotic and / or eukaryotic reference genomes, preferably a combination of two or more of the following: *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973. The genomes of Streptomyces azureense, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells are preferred, and one or more of the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and Escherichia coli are preferred, with the most preferred combination of the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and Escherichia coli.

24. The nucleic acid according to any one of claims 6-23, wherein no protospacer is present in any one or more of the prokaryotic and / or eukaryotic reference genomes (preferably one or more of the following reference genomes): *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973. The genomes of Streptomyces azureense, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells, preferably none of which are present in the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis, and Escherichia coli.

25. The nucleic acid according to any one of claims 6-24, wherein any protospacer does not have more than 40% sequence identity with any sequence of any one of one or more prokaryotic and / or eukaryotic reference genomes (preferably one or more of the following reference genomes): *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973 The genomes of Streptomyces azureense, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells are preferred, with any protospacer among them having no more than 40% sequence identity with any sequence of any of the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis, and Escherichia coli.

26. The nucleic acid according to any one of claims 1-25, wherein the nucleic acid is deoxyribonucleic acid (DNA).

27. The nucleic acid of any one of claims 1-26, wherein the nucleic acid further comprises 5' and 3' homologous arms for insertion into a host cell (e.g., into the genome of the host cell), the homologous arms comprising sequences homologous to a desired locus in the host cell genome.

28. The nucleic acid of any one of claims 27, wherein the desired locus comprises a highly expressed locus, or a locus with high genome accessibility, or a locus for spatial expression in a specific host cell, or a locus encoding a selection marker or a safe harbor locus.

29. The nucleic acid of any one of claims 27 or 28, wherein the host cell comprises *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK cells. 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, preferably wherein the host cells comprise Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli cells.

30. The nucleic acid of any one of claims 27-29, wherein the desired locus comprises the dppF locus in Escherichia coli host cells, the pksX locus in Bacillus subtilis host cells, the PDC6 locus in Saccharomyces cerevisiae host cells, and / or the URA3 locus in Yersinia lipolytica.

31. The nucleic acid of any one of claims 6-30, wherein the length of the one or more protospacers is related to the length of the homologous region required for homologous recombination in the host cell.

32. The nucleic acid according to any one of claims 1-31, wherein the nucleic acid further comprises a promoter and a target gene located at the 5' or 3' end.

33. The nucleic acid of claim 32, wherein the promoter is selected from the TEF1 minimal promoter, LEU, TRE, tomato SlDFR gene promoter (pSlDFR), CMV, TPGI, TRE3G, MYOD1 and / or HBG1 promoter.

34. The nucleic acid of claim 32 or 33, wherein the target gene encodes a fluorophore, a selection marker, a monoclonal antibody, an interferon, and / or a growth hormone.

35. The nucleic acid according to any one of claims 32-34, wherein the target gene encodes GFP and / or growth factors EGF, bFGF, FGF2, HGF, TGF, PDGF.

36. The nucleic acid according to any one of claims 6-35, wherein the protospacer contains a target site of an RNA-guided endonuclease (preferably a CRISPR protein).

37. The nucleic acid of claim 36, wherein the RNA-guided endonuclease comprises a Cas protein, preferably Cas9, most preferably Streptococcus pyogenes Cas9 (SpCas9).

38. The nucleic acid of claim 37, wherein the Cas protein comprises a catalytically inactivated Cas protein, such as dead Cas9 (dCas9) containing D10A and H840A mutations.

39. The nucleic acid of claim 38, wherein the catalytically inactivated Cas protein comprises a transcription activator, such as Vp16 transcription activator, Vp64 transcription activator, p65 transcription activator, Rta transcription activator, co-activator mediator (SAM), SunTag activation system, Moontag, HSF1, HSFA6b, AvrXa10, DOF1, DREB1, DREB2, dCas9-TV, p300, CBP, or a combination thereof.

40. The nucleic acid of claim 39, wherein the catalytically inactivated Cas protein comprises a transcriptional repressor, such as Krüppel-associated box (KRAB), Mxi1, MeCP2 (methyl-CpG-binding protein 2), SIN3A (switch-independent 3A), HDACs (histone deacetylases), Groucho / TLE (transduction protein-like enhancers), NCOR (nuclear receptor co-repressor), NCoR2 (nuclear receptor co-repressor 2), lysine-specific demethylase 1 (LSD1), retinoic acid and thyroid hormone receptor silencing mediator (SMRT), or a combination thereof.

41. The nucleic acid according to any one of claims 36-40, wherein the RNA-guided endonuclease is complexed with guide RNA (gRNA), preferably with single-molecule guide RNA (sgRNA).

42. The nucleic acid of claim 41, wherein the gRNA or sgRNA comprises a spacer sequence substantially complementary to the original spacer sequence of the nucleic acid.

43. The nucleic acid of claim 41 or 42, wherein the gRNA or sgRNA comprises a spacer sequence substantially complementary to the unique protospacer sequence of the nucleic acid.

44. The nucleic acid according to any one of claims 41-43, wherein the gRNA or sgRNA comprises a spacer selected from SEQ ID NO: 180-237.

45. The nucleic acid of any one of claims 1-44, wherein the nucleic acid is a docking element (CRISPRpad).

46. ​​A nucleic acid comprising or consisting of one or more sequences (low-representation sequences / components) that are low-representational in one or more reference genomes.

47. The nucleic acid of claim 46, wherein, If a sequence falls below 40% in one or more reference genomes in terms of frequency of occurrence compared to other sequences or sequence fragments of the same length, then the sequence is considered low-representative in the one or more reference genomes; and / or if, in one or more reference genomes, the total frequency of a sequence is below 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, or 10%, preferably 70% or 30%, compared to the highest observed frequency of sequences or sequence fragments of the same length, then the sequence is considered low-representative in one or more reference genomes.

48. The nucleic acid of claim 46 or 47, wherein the (low-representative) sequence is a fragment of the sequence of the one or more reference genomes, preferably a continuous fragment of 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides of the sequence of the one or more reference genomes, preferably a continuous fragment of 4 or 5 nucleotides of the sequence of the one or more reference genomes.

49. The nucleic acid according to any one of claims 46-48, wherein the (low-representative) sequence is 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer or 10mer, preferably wherein the sequence is 4mer or 5mer.

50. The nucleic acid of claim 49, wherein the sequence is 4mer and selected from any of Tables 4-47 or 92-116, or wherein the sequence is 5mer and selected from any of Tables 48-91 or 117-142.

51. A primary spacer comprising or composed of one or more nucleic acids of any one of claims 46-50, preferably two or more nucleic acids of any one of claims 46-50, preferably wherein the two or more nucleic acids are adjacent, more preferably wherein two or more 4-mer or 5-mer nucleic acids are adjacent.

52. The original spacer of claim 51, wherein the original spacer comprises or consists of 3-7 adjacent 4mer or 5mer nucleic acids, preferably wherein the original spacer comprises or consists of 5 adjacent 4mers, or preferably wherein the original spacer comprises or consists of 4 adjacent 5mers.

53. The primary spacer as claimed in claim 51 or 52, wherein the primary spacer comprises or consists of about 15-35 nucleotides, preferably about 20 nucleotides.

54. The protospacer of any one of claims 51-53, wherein the protospacer is not present in the one or more reference genomes.

55. The protospacer of any one of claims 51-54, wherein the protospacer has less than 10% sequence identity with a sequence in one or more reference genomes.

56. The primary spacer as claimed in any one of claims 51-55, wherein the primary spacer is selected from SEQ ID NO: 122-179 and 307-349038.

57. A nucleic acid comprising one or more protospacers according to any one of claims 51-56, preferably at least two protospacers according to any one of claims 51-56, more preferably all protospacers according to any one of claims 51-56, and a protospacer sequence adjacent motif (PAM) at the 3' end of each protospacer, and a linker of 1 to 10 nucleotides adjacent to the 3' end of said PAM, preferably a linker of 7 nucleotides adjacent to the 3' end of said PAM, optionally wherein said PAM is included in the following: (a) A linker consisting of 3 to 10 nucleotides, preferably 10 nucleotides, at the 3' end of each primary spacer, or (b) A 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, or 10mer sequence at the 3' end of each protospacer, preferably a 4mer or 5mer sequence, or (c) Another protospacer or another incomplete protospacer at the 3' end of each protospacer, wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the protospacer, preferably 10 nucleotides of the protospacer.

58. A docking element (CRISPRpad) comprising one or more protospacers and a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, preferably comprising one or more protospacers of any one of claims 51-56 or the nucleic acid of claim 57, preferably wherein the docking element comprises at least two protospacers of any one of claims 51-56.

59. The docking element (CRISPRpad) as claimed in claim 58, comprising about 2 to 60 primary spacers, preferably about 30, and most preferably about 13.

60. The docking element (CRISPRpad) as described in claim 58 or 59, wherein the sequence of each primary spacer is unique within the docking element (CRISPRpad).

61. The docking element (CRISPRpad) as described in any one of claims 58-60, wherein any primary spacer does not have more than 34% sequence identity with another sequence within the docking element (CRISPRpad).

62. The docking element (CRISPRpad) according to any one of claims 58-61, wherein the docking element (CRISPRpad) comprises about 19-2500 nucleotides, preferably about 600 nucleotides.

63. The docking element (CRISPRpad) according to any one of claims 58-62, comprising, or consisting of, the protospacers shown in SEQ ID NO: 122-179 and 307-349038, a protospacer sequence adjacent motif (PAM) at the 3' end of each protospacer, and a linker of 1 to 10 nucleotides adjacent to the 3' end of the PAM, preferably a linker of 7 nucleotides adjacent to the 3' end of the PAM, or thereof, optionally wherein the PAM is included in the following: (a) A linker consisting of 3 to 10 nucleotides, preferably 10 nucleotides, at the 3' end of each primary spacer, or (b) A 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, or 10mer sequence at the 3' end of each protospacer, preferably a 4mer or 5mer sequence, or (c) Another protospacer or another incomplete protospacer at the 3' end of each protospacer, wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the protospacer, preferably 10 nucleotides of the protospacer.

64. The docking element (CRISPRpad) according to any one of claims 58-63, comprising or composed of the nucleic acid shown in any one of SEQ ID NO: 1-2 or 349039 – 4572163 and / or 4572167 – 4687816.

65. The docking element (CRISPRpad) according to any one of claims 58-64, wherein the one or more reference genomes comprise one or more prokaryotic and / or eukaryotic reference genomes, preferably including: *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongatus* PCC 7942, and *Anabaena* species PCC 7120, *Synechococcus slenderus* UTEX2973, *Streptomyces azure*, *Streptomyces limonensis*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cell genomes, with the most preferred being one or more of the genomes of *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis* and / or *Escherichia coli*.

66. The docking element (CRISPRpad) as described in any one of claims 58-65, wherein the one or more reference genomes comprise a combination of two or more prokaryotic and / or eukaryotic reference genomes, preferably a combination of two or more of the following: *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, and *Synechococcus elongata* UTEX 2973. The genomes of Streptomyces azureense, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells are preferred, and one or more of the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and Escherichia coli are preferred, with the most preferred combination of the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and Escherichia coli.

67. The docking element (CRISPRpad) as described in any one of claims 58-66, wherein none of the protospacers are present in any one of the one or more prokaryotic and / or eukaryotic reference genomes (preferably one or more of the following reference genomes): *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973. The genomes of Streptomyces azureense, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells, preferably none of which are present in the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis, and Escherichia coli.

68. The docking element (CRISPRpad) as described in any one of claims 58-67, wherein any protospacer does not have more than 40% sequence identity with any sequence in any one or more of the following prokaryotic and / or eukaryotic reference genomes (preferably one or more of the following reference genomes): *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973. The genomes of Streptomyces azureense, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells are preferred, with any protospacer among them having no more than 40% sequence identity with any sequence of any of the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis, and Escherichia coli.

69. The docking element (CRISPRpad) according to any one of claims 58-68, wherein the docking element is encoded in a nucleic acid, preferably deoxyribonucleic acid (DNA).

70. The docking element (CRISPRpad) of any one of claims 58-69, wherein the docking element further comprises 5' and 3' homologous arms for insertion into a host cell (e.g., insertion into the host cell's genome), the homologous arms comprising sequences homologous to a desired locus in the host cell's genome.

71. The docking element (CRISPRpad) of claim 70, wherein the desired locus includes a highly expressed locus, or a locus with high genome accessibility, or a locus for spatial expression in a specific host cell, or a locus encoding a selection marker or a safe harbor locus.

72. The docking element (CRISPRpad) as described in claim 70 or 71, wherein the host cell comprises *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC6803, *Synechococcus* species PCC7002, *Synechococcus elongata* PCC7942, *Anabaena* species PCC7120, *Synechococcus elongata* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO cells). K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, preferably wherein the host cells comprise Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli cells.

73. The docking element (CRISPRpad) according to any one of claims 70-72, wherein the desired locus comprises the dppF locus in Escherichia coli host cells, the pksX locus in Bacillus subtilis host cells, the PDC6 locus in Saccharomyces cerevisiae host cells, and / or the URA3 locus in Yersinia lipolytica.

74. The docking element (CRISPRpad) of any one of claims 58-73, wherein the length of the one or more protospacers is related to the length of the homologous region required for homologous recombination in the host cell.

75. The docking element (CRISPRpad) of any one of claims 58-74, wherein the nucleic acid further comprises a promoter and a target gene, optionally located at the 5' or 3' end.

76. The docking element (CRISPRpad) of claim 75, wherein the promoter is selected from the TEF1 minimal promoter, LEU, TRE, tomato SlDFR gene promoter (pSlDFR), CMV, TPGI, TRE3G, MYOD1 and / or HBG1 promoter.

77. The docking element (CRISPRpad) as described in claim 75 or 76, wherein the target gene encodes a fluorophore, a selection marker, a monoclonal antibody, interferon, and / or growth hormone.

78. The docking element (CRISPRpad) according to any one of claims 75-77, wherein the target gene encodes GFP and / or growth factors EGF, bFGF, FGF2, HGF, TGF, PDGF.

79. The docking element (CRISPRpad) of any one of claims 58-78, wherein the protospacer contains a target site of an RNA-guided endonuclease (preferably a CRISPR protein).

80. The docking element (CRISPRpad) of claim 79, wherein the RNA-guided endonuclease comprises a Cas protein, preferably Cas9, most preferably Streptococcus pyogenes Cas9 (SpCas9).

81. The docking element (CRISPRpad) of claim 80, wherein the Cas protein comprises a catalytically inactivated Cas protein, such as dead Cas9 (dCas9) containing D10A and H840A mutations.

82. The docking element (CRISPRpad) of claim 81, wherein the catalytically inactivated Cas protein comprises a transcription activator, such as Vp16 transcription activator, Vp64 transcription activator, p65 transcription activator, Rta transcription activator, co-activator mediator (SAM), SunTag activation system, Moontag, HSF1, HSFA6b, AvrXa10, DOF1, DREB1, DREB2, dCas9-TV, p300, CBP, or a combination thereof.

83. The docking element (CRISPRpad) of claim 81, wherein the catalytically inactivated Cas protein comprises a transcriptional repressor, such as Krüppel-associated box (KRAB), Mxi1, MeCP2 (methyl-CpG-binding protein 2), SIN3A (switch-independent 3A), HDACs (histone deacetylases), Groucho / TLE (transduction protein-like enhancers), NCOR (nuclear receptor co-repressor), NCoR2 (nuclear receptor co-repressor 2), lysine-specific demethylase 1 (LSD1), retinoic acid and thyroid hormone receptor silencing mediator (SMRT), or a combination thereof.

84. The docking element (CRISPRpad) according to any one of claims 79-83, wherein the RNA-guided endonuclease is complexed with guide RNA (gRNA), preferably with single-molecule guide RNA (sgRNA).

85. The docking element (CRISPRpad) of claim 84, wherein the gRNA or sgRNA comprises a spacer sequence substantially complementary to the original spacer sequence of the docking element.

86. The docking element (CRISPRpad) of claim 84 or 85, wherein the gRNA or sgRNA comprises a spacer sequence substantially complementary to the unique protospacer sequence of the docking element.

87. The docking element (CRISPRpad) of any one of claims 84-86, wherein the gRNA or sgRNA comprises a spacer selected from SEQ ID NO: 180-237.

88. A vector comprising the nucleic acid of any one of claims 1-50 and 57, the protospacer of any one of claims 51-56, or the docking element of any one of claims 58-87.

89. The use of the nucleic acid as described in any one of claims 1-50, 57 or the docking element as described in any one of claims 58-87 as a docking element / CRISPRpad in one or more prokaryotic and eukaryotic cells / organisms, preferably wherein said one or more cells / organisms are *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC7942, *Anabaena* species PCC 7120, and *Synechococcus elongata* UTEX 2973. Streptomyces azureense, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, preferably wherein one or more of the cells / organisms are Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli cells / organisms.

90. A method for generating docking elements, the method comprising: (i) Identify sequences that are underrepresented in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Introduce a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, or A protospacer adjacent motif (PAM) and a linker of 1 to 10 nucleotides adjacent to the 3' end of the PAM are introduced at the 3' end of each protospacer. A linker of 7 nucleotides adjacent to the 3' end of the PAM is preferred. A linker of 3 to 10 nucleotides, preferably 10 nucleotides, is introduced at the 3' end of each protospacer. A 3-mer, 4-mer, 5-mer, 6-mer, 7-mer, 8-mer, 9-mer, or 10-mer sequence containing the neighboring motif (PAM) of the protospacer sequence is introduced at the 3' end of each protospacer, preferably a 4-mer or 5-mer sequence, or At the 3' end of each protospacer, another (ii) protospacer or another (ii) incomplete protospacer containing a protospacer sequence adjacent motif (PAM) is introduced, wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the (ii) protospacer, preferably 10 nucleotides of the (ii) protospacer. (iv) Assembling two or more protospacers into a nucleic acid encoding them, wherein the sequence of each protospacer is unique within the nucleic acid. (v) Synthesize nucleic acids encoding the two or more protospacers to generate docking elements.

91. The method of claim 44, wherein, If a sequence falls below 40% in one or more reference genomes in terms of frequency of occurrence compared to other sequences or sequence fragments of the same length, then the sequence is considered low-representative in the one or more reference genomes; and / or if, in one or more reference genomes, the total frequency of a sequence is below 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, or 10%, preferably 70% or 30%, compared to the highest observed frequency of sequences or sequence fragments of the same length, then the sequence is considered low-representative in one or more reference genomes.

92. The method of claim 90 or 91, wherein the (low-representative) sequence is a fragment of the sequence of the one or more reference genomes, preferably wherein the fragment is a continuous fragment of 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides of the sequence of the one or more reference genomes, and preferably wherein the fragment is a continuous fragment of 4 or 5 nucleotides of the sequence of the one or more reference genomes.

93. The method of any one of claims 90-92, wherein the (low-representative) sequence is 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, or 10mer, preferably wherein the sequence fragment is 4mer or 5mer.

94. The method of any one of claims 90-93, wherein the sequence is 4mer and selected from any one of Tables 4-47 or Tables 92-116, or wherein the sequence is 5mer and selected from any one of Tables 48-91 or Tables 117-142.

95. The method of any one of claims 90-94, wherein the protospacer comprises about 15-35 nucleotides, preferably about 20 nucleotides.

96. The method of any one of claims 90-95, wherein (ii) comprises assembling 3-7 4mers or 5mers into a protospacer, preferably, wherein (ii) comprises assembling 5 4mers or 4 5mers into a 20-nucleotide protospacer.

97. The method of any one of claims 90-96, wherein (ii) comprises assembling five 4mers selected from any one of Tables 4-47 or Tables 92-116 into one or more primary spacers, wherein each primary spacer consists of five 4mers, and / or wherein (ii) comprises assembling four 5mers selected from any one of Tables 48-91 or Tables 117-142 into one or more primary spacers, wherein each primary spacer consists of four 5mers.

98. The method of any one of claims 90-97, wherein the method further comprises introducing a linker of 1 to 10 nucleotides adjacent to the 3' end of the PAM, preferably introducing a linker of 7 nucleotides adjacent to the 3' end of the PAM.

99. The method of any one of claims 90-98, wherein (iv) comprises assembling about 2-60 protospacers into a nucleic acid encoding thereof, preferably, wherein (iv) comprises assembling about 30 protospacers into a nucleic acid encoding thereof, and most preferably, wherein (iv) comprises assembling about 13 protospacers into a nucleic acid encoding thereof.

100. The method of any one of claims 90-99, wherein any primary spacer does not have more than 34% sequence identity with another sequence within the docking element.

101. The method of any one of claims 90-100, wherein none of the original spacers are present in the one or more reference genomes.

102. The method of any one of claims 90-101, wherein any protospacer does not have more than 10% sequence identity with any sequence in the one or more reference genomes.

103. The method of any one of claims 90-102, wherein the two or more primary spacers are assembled into a nucleic acid comprising about 19-2500 nucleotides, preferably about 600 nucleotides.

104. The method of any one of claims 90-103, wherein the one or more assembled protospacers are protospacers selected from SEQ ID NO: 122-179 and 307-349038.

105. The method of any one of claims 90-104, wherein the two or more protospacers are assembled into nucleic acids represented by any one of SEQ ID NO: 1-2 or 349039-4572163 and / or 4572167-4687816.

106. The method according to any one of claims 90-105, wherein two or more protospacers shown in SEQ ID NO: 122-149 or 150-179 are respectively assembled into nucleic acid sequences shown in SEQ ID NO: 1 or 2.

107. The method of any one of claims 90-106, wherein the one or more reference genomes comprise one or more prokaryotic and / or eukaryotic reference genomes, preferably including: *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, and *Synechococcus elongata* UTEX. 2973, Streptomyces azureus, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHOK1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cell genomes, with the most preferred being one or more of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli.

108. The method of any one of claims 90-107, wherein the one or more reference genomes comprise a combination of two or more prokaryotic and / or eukaryotic reference genomes, preferably a combination of: *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, and *Synechococcus elongata* UTEX2973. The genomes of Streptomyces azureense, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells are preferred, and one or more of the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and Escherichia coli are preferred, with the most preferred combination of the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and Escherichia coli.

109. The method of any one of claims 90-108, wherein no protospacer is present in any one of one or more prokaryotic and / or eukaryotic reference genomes (preferably the one or more reference genomes of the following): *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973. The genomes of Streptomyces azureense, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells, preferably none of which are present in the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis, and Escherichia coli.

110. The method according to any one of claims 90-109, wherein any protospacer does not have more than 10% sequence identity with any sequence of any of one or more prokaryotic and / or eukaryotic reference genomes (preferably one or more of the following reference genomes): *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973. The genomes of Streptomyces azureense, Streptomyces limonensis, Clostridium acetobutyricum, rice, corn, tobacco, Nicotiana benthamiana, Arabidopsis thaliana, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells, and / or Hi5 cells are preferred, with any protospacer among them having no more than 40% sequence identity with any sequence of any of the genomes of Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis, and Escherichia coli.

111. The method according to any one of claims 90-110, wherein the nucleic acid is deoxyribonucleic acid (DNA).

112. The method of any one of claims 90-111, wherein the method further comprises attaching 5' and 3' homologous arms to the nucleic acid for insertion into a host cell, for example, into the genome of a host cell, said homologous arms comprising sequences homologous to a desired locus in the host cell genome.

113. The method of claim 112, wherein the desired locus includes a highly expressed locus, or a locus with high genome accessibility, or a locus for spatial expression in a specific host cell, or a locus encoding a selection marker or a safe harbor locus.

114. The method of claim 112 or 113, wherein the host cell is *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK cells. 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, preferably wherein the host cells comprise Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli cells.

115. The method of any one of claims 112-114, wherein the desired locus is the dppF locus in Escherichia coli host cells, the pksX locus in Bacillus subtilis host cells, the PDC6 locus in Saccharomyces cerevisiae host cells, and / or the URA3 locus in Yersinia lipolytica.

116. The method of any one of claims 90-115, wherein the length of the one or more protospacers is related to the length of the homologous region required for homologous recombination in the host cell.

117. The method of any one of claims 90-116, wherein the method further comprises attaching a nucleic acid sequence encoding a promoter and a target gene to the 5' or 3' end of the nucleic acid.

118. The method of claim 117, wherein the promoter is selected from the TEF1 minimal promoter, LEU, TRE, tomato SlDFR gene promoter (pSlDFR), CMV, TPGI, TRE3G, MYOD1 and / or HBG1 promoter.

119. The method of claim 117 or 118, wherein the target gene encodes a fluorophore, a selection marker, a monoclonal antibody, interferon, and / or growth hormone.

120. The method of any one of claims 117-119, wherein the target gene encodes GFP and / or growth factors EGF, bFGF, FGF2, HGF, TGF, PDGF.

121. The method of any one of claims 90-120, wherein the nucleic acid comprises the nucleic acid of any one of claims 1-45 or 57, or the docking element of any one of claims 58-87.

122. The method of any one of claims 90-120, wherein the primary spacer comprises the primary spacer of any one of claims 51-56.

123. The method of any one of claims 90-120, wherein the low-representation sequence comprises the nucleic acid of any one of claims 46-50.

124. A nucleic acid produced by the method of any one of claims 90-123.

125. A genome comprising the nucleic acid of any one of claims 1-50 and 57, the protospacer of any one of claims 51-56, or the docking element of any one of claims 58-87.

126. The genome of claim 125, wherein the genome is a prokaryotic or eukaryotic genome.

127. The genome of claim 125 or 126, wherein the genome is a genome of Bacillus, Escherichia, Yeast, Yersinia, Clostridium, or Corynebacterium.

128. The genome according to any one of claims 125-127, wherein the genome is *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK cells. The genomes of 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, preferably wherein the genomes are Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli genomes.

129. The genome of claim 127 or 128, wherein the genome is a Bacillus subtilis genome containing the nucleic acid shown in SEQ ID NO: 1, 2, 106 or 107.

130. The genome of claim 127 or 128, wherein the genome is an Escherichia coli genome containing the nucleic acid shown in SEQ ID NO: 1, 2 or 3.

131. The genome of claim 127 or 128, wherein the genome is a Saccharomyces cerevisiae genome containing the nucleic acid shown in SEQ ID NO: 1, 2 or 4.

132. The genome of claim 127 or 128, wherein the genome is a Yersinia lipophila genome containing the nucleic acid shown in SEQ ID NO: 1, 2, 5 or 6.

133. The genome according to any one of claims 125-128, wherein the genome is a CHO K1 cell genome, preferably a CHO-BXB-SV40 cell genome, comprising nucleic acids represented by SEQ ID NO: 2140375 – 2213124 and / or 4572167 – 4687816, preferably any one of SEQ ID NO: 4572162 or 4572163.

134. A host cell comprising or transformed with any one of claims 1-50 and 57, any one of claims 51 to 56, or any one of claims 58 to 87, a docking element.

135. The host cell of claim 134, wherein, The nucleic acid as described in any one of claims 1-50 and 57, the protospacer as described in any one of claims 51 to 56, or the docking element as described in any one of claims 58 to 87 is integrated into the genome of the host cell.

136. The host cell as claimed in claim 134 or 135, wherein the host cell is a prokaryotic cell or a eukaryotic cell.

137. The host cell according to any one of claims 134-136, wherein the host cell is a cell of Bacillus, Escherichia, Yeast, Yersinia, Clostridium, or Corynebacterium.

138. The host cell according to any one of claims 134-137, wherein the host cell is *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK cells. 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, preferably wherein the host cells are Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli cells.

139. The host cell of any one of claims 134-138, wherein the nucleic acid, protospacer, or docking element is contained in a high-expression locus, or a locus with high genome accessibility, or a locus for specific spatial expression in the host cell, or a locus encoding a selection marker or a safe harbor locus.

140. The host cell according to any one of claims 134-138, wherein the host cell is a CHO K1 cell, preferably a CHO-BXB-SV40 cell, comprising the nucleic acid shown in any one of SEQ ID NO: 2140375 – 2213124 and / or 4572167 – 4687816, preferably SEQ ID NO: 4572162 or 4572163.

141. The host cell according to any one of claims 134-140, wherein the host cell is a cell deposited in DSMZ with accession number DSM ACC3380.

142. The host cell of any one of claims 134-139, wherein the nucleic acid, protospacer, or docking element is contained in the dppF locus of the *Escherichia coli* host cell, the pksX locus of the *Bacillus subtilis* host cell, the PDC6 locus of the *Saccharomyces cerevisiae* host cell, and / or the URA3 locus of the *Yarrowia lipolytica* host cell.

143. The host cell according to any one of claims 134-142, wherein the host cell is a cell deposited in DSMZ with accession number DSM 34807.

144. A system comprising: (i) the nucleic acid as described in any one of claims 1-50 and 57, the protospacer as described in any one of claims 51 to 56, or the docking element as described in any one of claims 58 to 87. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding them.

145. The system of claim 144, wherein the RNA-guided endonuclease is a CRISPR protein.

146. The system of claim 144 or 145, wherein the RNA-guided endonuclease comprises a Cas protein, preferably Cas9, most preferably Streptococcus pyogenes Cas9 (SpCas9).

147. The system of any one of claims 144-146, wherein the one or more gRNAs are one or more sgRNAs.

148. The system according to any one of claims 144-147, wherein, RNA-guided endonucleases are pre-complexed with one or more gRNAs or one or more sgRNAs to form ribonucleoprotein (RNP) complexes.

149. The system of any one of claims 144-148, wherein the nucleic acid encoding the RNA-guided endonuclease is deoxyribonucleic acid (DNA).

150. The system of any one of claims 144-149, wherein the nucleic acid encoding the RNA-guided endonuclease is ribonucleic acid (RNA).

151. The system of claim 150, wherein the RNA encoding the RNA-guided endonuclease is mRNA.

152. The system as claimed in any one of claims 144-151, wherein, (i) The nucleic acid, protospacer, or docking element is encoded in an adeno-associated virus (AAV) vector.

153. The system as described in any one of claims 144-151, wherein, The RNA-guided endonuclease or the nucleic acid encoding the RNA-guided endonuclease is formulated in liposomes or lipid nanoparticles.

154. The system of claim 153, wherein, The liposomes or lipid nanoparticles also contain one or more gRNAs or nucleic acids encoding one or more gRNAs.

155. The system of claim 153, wherein, Each of (i), (ii) and (iii) is formulated separately in liposomes or lipid nanoparticles, or (i), (ii) and (iii) are formulated together in a single liposome or lipid nanoparticle.

156. The system according to any one of claims 144-155, wherein, The one or more gRNAs or sgRNAs contain spacer sequences substantially complementary to loci in the host cell genome.

157. The system according to any one of claims 144-156, wherein, The one or more gRNAs or sgRNAs contain a spacer sequence that is substantially complementary to the protospacer of the nucleic acid, protospacer, or docking element of (i).

158. The system of any one of claims 144-157, wherein the gRNA or sgRNA comprises a spacer sequence that is substantially complementary to the unique spacer of the nucleic acid, protospacer, or docking element of (i).

159. The system according to any one of claims 144-158, wherein (iii) comprises at least two gRNAs or sgRNAs or nucleic acids encoding them.

160. The system of claim 159, wherein the at least two gRNAs or sgRNAs comprise spacer sequences substantially complementary to loci in the host cell genome, and wherein one of the at least two gRNAs or sgRNAs comprises a spacer sequence substantially complementary to a unique protospacer at the 5' end of the nucleic acid, protospacer, or docking element of (i), and wherein the other of the at least two gRNAs or sgRNAs comprises a spacer sequence substantially complementary to a unique protospacer at the 3' end of the nucleic acid, protospacer, or docking element of (i).

161. The system of any one of claims 156-160, wherein the locus is a high-expression locus, or a locus with high genome accessibility, or a locus for spatial expression in a specific host cell, or a locus encoding a selection marker or a safe harbor locus.

162. The system according to any one of claims 156-161, wherein the host cell is *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK cells. 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, preferably wherein the host cells are Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli cells.

163. The system according to any one of claims 156-162, wherein the locus is the dppF locus in Escherichia coli host cells, the pksX locus in Bacillus subtilis host cells, the PDC6 locus in Saccharomyces cerevisiae host cells, and / or the URA3 locus in Yersinia lipolytica.

164. The system according to any one of claims 144-163, wherein the gRNA or sgRNA comprises a spacer selected from SEQ ID NO: 180-237.

165. A method of generating a cell comprising a docking element, wherein the method comprises applying a nucleic acid of any one of claims 1-50 and 57, a protospacer of any one of claims 51-56, a docking element of any one of claims 58-87, a vector of claim 88, or a system of any one of claims 144-164 to the cell, optionally wherein the nucleic acid, protospacer, docking element, or vector is integrated into the genome of the cell.

166. An in vitro or ex vivo method for generating cells comprising docking elements, wherein the method comprises applying a nucleic acid of any one of claims 1-50 and 57, a protospacer of any one of claims 51-56, a docking element of any one of claims 58-87, a vector of claim 88, or a system of any one of claims 144-164 to the cells, optionally wherein the nucleic acid, protospacer, docking element, or vector is integrated into the genome of the cells.

167. The method of claim 165 or 166, wherein the cells are *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK cells. 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, preferably wherein the cells are Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli cells.

168. A cell that can be obtained, acquired, or generated by any one of claims 165-167.

169. A method for regulating the expression of a target gene in a cell, said cell comprising a nucleic acid of any one of claims 1-50 and 57, a protospacer of any one of claims 51 to 56, a docking element of any one of claims 58 to 87, a vector of claim 88, or a system of any one of claims 144-164, said method comprising introducing into said cell: (i) Donor polynucleotides (donor templates) encoding the target gene. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding the gRNA, wherein the gRNA is substantially complementary to one or more protospacers of the nucleic acid of any one of claims 1-44, or to one or more protospacers of the nucleic acid of any one of claims 51-56, the docking element of any one of claims 58-87, the vector of claim 88, or the system of any one of claims 144-164.

170. An in vitro or ex vivo method for regulating the expression of a target gene in cells, said cells comprising a nucleic acid of any one of claims 1-50 and 57, a protospacer of any one of claims 51-56, a docking element of any one of claims 58-87, a vector of claim 88, or a system of any one of claims 144-164, said method comprising introducing into said cells: (i) A donor polynucleotide (donor template) that encodes the target gene and contains the following. (ii) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (iii) One or more gRNAs or nucleic acids encoding the gRNA, wherein the gRNA is substantially complementary to one or more protospacers of the nucleic acid of any one of claims 1-44, a protospacer of any one of claims 51-56, a docking element of any one of claims 58-87, a vector of claim 88, or a system of any one of claims 144-164.

171. The method of claim 169 or 170, wherein the donor polynucleotide comprises homologous arms at its 5' and 3' ends for insertion into the nucleic acid of any one of claims 1-43, the protospacer of any one of claims 51-56, the docking element of any one of claims 58-87, the vector of claim 88, or the nucleic acid of any one of claims 144-164, wherein the homologous arms comprise sequences homologous to the nucleic acid of any one of claims 1-44, the protospacer of any one of claims 51-56, the docking element of any one of claims 58-87, the vector of claim 88, or the nucleic acid of any one of claims 144-164, preferably comprising sequences homologous to at least two protospacers contained in the nucleic acid of any one of claims 1-44, the docking element of any one of claims 58-87, the vector of claim 88, or the nucleic acid of any one of claims 144-164.

172. The method of claim 171, wherein the length of the homologous arm is related to the length of the homologous region required for homologous recombination in the cell.

173. The method according to any one of claims 169-172, wherein, Donor polynucleotides also contain promoters, such as the TEF1 minimal promoter, LEU, TRE, tomato SlDFR gene promoter (pSlDFR), CMV, TPGI, TRE3G, MYOD1 and / or HBG1 promoter.

174. The method according to any one of claims 169-173, wherein the cells are *Saccharomyces cerevisiae*, *Yersinia lipolytica*, *Bacillus subtilis*, *Escherichia coli*, *Schizosaccharomyces cerevisiae*, *Pichia pastoris*, *Saccharomyces phagnum*, *Kluyveromyces lactis*, *Candida antarcticus*, *Pseudomonas putida*, *Corynebacterium glutamicum*, *Bacillus licheniformis*, *Bacillus thuringiensis*, *Bacillus pumilus*, *Bacillus polymyxa*, *Lactococcus lactis*, *Agrobacterium tumefaciens*, *Synechococcus* species PCC 6803, *Synechococcus* species PCC 7002, *Synechococcus elongata* PCC 7942, *Anabaena* species PCC 7120, *Synechococcus elongata* UTEX 2973, *Streptomyces azure*, *Streptomyces cerevisiae*, *Clostridium acetobutyricum*, rice, corn, tobacco, *Nicotiana benthamiana*, *Arabidopsis thaliana*, CHO cells (e.g., CHO K1 cells), HeLa cells, HEK cells. 293 cells, Vero cells, 3T3 cells, Sf9 cells, Sf21 cells and / or Hi5 cells, preferably wherein the host cells are Saccharomyces cerevisiae, Yersinia lipolytica, Bacillus subtilis and / or Escherichia coli cells.

175. The method according to any one of claims 169-174, wherein, The one or more gRNAs are substantially complementary to the unique 5' or 3' end spacer of the nucleic acid of any one of claims 1-45, the docking element of any one of claims 58-87, the vector of claim 88, or the system of any one of claims 144-164.

176. The method of any one of claims 169-175, wherein the method further comprises introducing into the cell: (a) RNA-guided endonucleases or nucleic acids encoding such endonucleases, and (b) One or more gRNAs or nucleic acids encoding the gRNA, wherein, The gRNA is substantially complementary to one or more protospacers in the nucleic acid of any one of claims 1-44, the docking element of any one of claims 58-87, the vector of claim 88, or the nucleic acid of any one of claims 144-164, wherein the one or more protospacers are different from the protospacers described in (iii). (a) and (b) are introduced into the cell after (i) is integrated into the nucleic acid of any one of claims 1-44, the docking element of any one of claims 58-87, the vector of claim 88, or the nucleic acid of any one of claims 144-164 at the site targeted by (iii).

177. The method of any one of claims 169-176, wherein the gRNA is not substantially complementary to an endogenous sequence within the cell genome.

178. The method of any one of claims 169-177, wherein the gRNA does not have more than 40% sequence identity with an endogenous sequence within the cell genome.

179. The method according to any one of claims 169-178, wherein the RNA-guided endonuclease is a Cas protein, preferably Cas9, and most preferably Streptococcus pyogenes Cas9 (SpCas9).

180. The method of claim 179, wherein the Cas protein comprises a catalytically inactivated Cas protein, such as dead Cas9 (dCas9) containing D10A and H840A mutations.

181. The method of claim 180, wherein the catalytically inactivated Cas protein comprises a transcriptional activator, such as Vp16 transcriptional activator, Vp64 transcriptional activator, p65 transcriptional activator, Rta transcriptional activator, co-activator mediator (SAM), SunTag activation system, Moontag, HSF1, HSFA6b, AvrXa10, DOF1, DREB1, DREB2, dCas9-TV, p300, CBP, or a combination thereof.

182. The method of claim 180, wherein the catalytically inactivated Cas protein comprises a transcriptional repressor, such as Krüppel-associated box (KRAB), Mxi1, MeCP2 (methyl-CpG-binding protein 2), SIN3A (switch-independent 3A), HDACs (histone deacetylases), Groucho / TLE (transduction protein-like enhancers), NCOR (nuclear receptor co-repressor), NCoR2 (nuclear receptor co-repressor 2), lysine-specific demethylase 1 (LSD1), retinoic acid and thyroid hormone receptor silencing mediator (SMRT), or a combination thereof.

183. The method of any one of claims 169-182, wherein the gRNA is sgRNA.

184. The method of any one of claims 169-183, wherein the gRNA or sgRNA comprises a spacer selected from SEQ ID NO: 180-237.

185. A method for generating primary spacers, the method comprising: (i) Identify sequences that are underrepresented in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Introduce a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, or A protospacer adjacent motif (PAM) and a linker of 1 to 10 nucleotides adjacent to the 3' end of the PAM are introduced at the 3' end of each protospacer. A linker of 7 nucleotides adjacent to the 3' end of the PAM is preferred. A linker of 3 to 10 nucleotides, preferably 10 nucleotides, is introduced at the 3' end of each protospacer. A 3-mer, 4-mer, 5-mer, 6-mer, 7-mer, 8-mer, 9-mer, or 10-mer sequence containing the neighboring motif (PAM) of the protospacer sequence is introduced at the 3' end of each protospacer, preferably a 4-mer or 5-mer sequence, or At the 3' end of each protospacer, another (ii) protospacer or another (ii) incomplete protospacer containing a protospacer sequence adjacent motif (PAM) is introduced, wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the (ii) protospacer, preferably 10 nucleotides of the (ii) protospacer. (iv) Optionally, the one or more protospacers containing the PAM are synthesized into nucleic acids.

186. A computer-implemented method for generating primary spacers, comprising: (i) Identify sequences that are underrepresented in one or more reference genomes. (ii) Assemble low-representation sequences into one or more original spacers. (iii) Introduce a protospacer sequence neighbor motif (PAM) at the 3' end of each protospacer, or A protospacer adjacent motif (PAM) and a linker of 1 to 10 nucleotides adjacent to the 3' end of the PAM are introduced at the 3' end of each protospacer. A linker of 7 nucleotides adjacent to the 3' end of the PAM is preferred. A linker of 3 to 10 nucleotides, preferably 10 nucleotides, is introduced at the 3' end of each protospacer. A 3-mer, 4-mer, 5-mer, 6-mer, 7-mer, 8-mer, 9-mer, or 10-mer sequence containing the neighboring motif (PAM) of the protospacer sequence is introduced at the 3' end of each protospacer, preferably a 4-mer or 5-mer sequence, or Introduce at the 3' end of each protospacer another (ii) protospacer containing a protospacer sequence adjacent motif (PAM) or another (ii) incomplete protospacer, wherein the incomplete protospacer preferably contains 3 to 20 nucleotides of the protospacer of (ii), preferably 10 nucleotides of the protospacer of (ii).

187. The method of claim 185 or 186, wherein, If a sequence falls below 40% in one or more reference genomes in terms of frequency of occurrence compared to other sequences or sequence fragments of the same length, then the sequence is considered low-representative in the one or more reference genomes; and / or if, in one or more reference genomes, the total frequency of a sequence is below 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, or 10%, preferably 70% or 30%, compared to the highest observed frequency of sequences or sequence fragments of the same length, then the sequence is considered low-representative in one or more reference genomes.

188. The method of any one of claims 185-187, wherein the (low-representative) sequence is a fragment of the sequence of the one or more reference genomes, preferably wherein the fragment is a continuous fragment of 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides of the sequence of the one or more reference genomes, and preferably wherein the fragment is a continuous fragment of 4 or 5 nucleotides of the sequence of the one or more reference genomes.

189. The method of any one of claims 185-188, wherein the (low-representative) sequence is 3mer, 4mer, 5mer, 6mer, 7mer, 8mer, 9mer, or 10mer, preferably wherein the sequence is 4mer or 5mer.

190. The method of claim 189, wherein the sequence is 4mer and selected from any of Tables 4-47 or Tables 92-116, or wherein the sequence is 5mer and selected from any of Tables 48-91 or Tables 117-142.

191. The method of any one of claims 185-190, wherein (ii) comprises assembling low-representation 4mer or 5mer sequences into one or more primary spacers, preferably wherein (ii) comprises assembling five 4mers selected from any one of Tables 4-47 or Tables 92-116 into one or more primary spacers, wherein each primary spacer consists of five 4mers, and / or wherein (ii) comprises assembling four 5mers selected from any one of Tables 48-91 or Tables 117-142 into one or more primary spacers, wherein each primary spacer consists of four 5mers.

192. A data processing apparatus comprising means for performing the method of any one of claims 186-191.

193. A computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 186-191.

194. A computer-readable storage medium containing instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 186-191.

Citation Information

Patent Citations

  • Method for producing a filter element provided with a sealing part

    EP3730201A1