Stem-loop adapter, polynucleotide library for grna production, and methods for using same

The stem-loop adapter and PCR-based method efficiently produce gRNA libraries from small amounts of genome material, addressing the inefficiencies of current methods by enabling rapid and cost-effective genome and epigenome editing, including non-coding region analysis.

WO2026071097A1PCT designated stage Publication Date: 2026-04-02YOSHIDA KEISUKE +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Current methods for producing gRNA libraries are costly, time-consuming, and inefficient, particularly when dealing with small amounts of genome material, and do not effectively address the need for comprehensive genome editing and phenotypic analysis.

Method used

A stem-loop type adapter is combined with PCR technology to produce a polynucleotide library for gRNA production, utilizing adapters A and B to facilitate rapid, simple, and inexpensive generation of gRNA libraries from small amounts of polynucleotide fragments.

Benefits of technology

Enables the easy, rapid, and cost-effective production of diverse gRNA libraries spanning the entire target genome, allowing for comprehensive genome and epigenome editing, including analysis of non-coding regions and identification of gene mutations associated with phenotypic changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000109_0000
    Figure 00000109_0000
  • Figure 00000110_0000
    Figure 00000110_0000
  • Figure 00000111_0000
    Figure 00000111_0000
Patent Text Reader

Abstract

[Problem] To provide: an adapter for binding to polynucleotide fragments that is used for preparing a genome-wide gRNA library for genome editing systems; a gRNA library that is obtained by using the adapter; and uses of the gRNA library. [Solution] An adapter having a structure described in A). A) A structure having i) to iii). i) A polynucleotide chain having sequences i)-1 to i)-3. i)-1: Sequence corresponding to the forward primer. i)-2: Sequence of the recognition site of a restriction enzyme capable of excising a gRNA spacer sequence from the 5' end of a polynucleotide fragment bound to the adapter. i)-3: Sequence of the recognition site of a restriction enzyme capable of cleaving an adapter having the structure of A) from the polynucleotide fragment. ii) A polynucleotide chain having the sequence complementary to i). iii) A single-stranded polynucleotide chain for formation of a loop connecting the 5' end of i) and the 3' end of ii) of the double strand resulting from the binding of i) and ii).
Need to check novelty before this filing date? Find Prior Art

Description

Stem-loop adapter, polynucleotide library for gRNA production, and method of using the same.

[0001] This invention relates to a polynucleotide fragment adapter for efficiently producing gRNA libraries for use in genome editing technology, a polynucleotide library for gRNA production produced using the same, and methods for utilizing the same.

[0002] Traditionally, identifying the causative genes for cancer and other diseases has been widely done by examining the genes of actual patients. However, due to limitations in the amount of patient information available, accurately determining the precise location of gene mutations is not always easy.

[0003] Furthermore, even if the causative gene can be estimated, it is currently difficult to clarify the specific details of which mutations in that gene may be the cause. For example, in the case of cancer, as mentioned above, it is difficult to clarify which gene mutations may be the cause, so although many gene mutations have been reported in cancer patients, only a small number have been identified as causative factors (driver mutations) of cancer. Gene mutations that may be causative factors of cancer can be used as biomarkers and are useful for individual diagnosis and treatment selection, so there is a high clinical need for them. However, currently, it is difficult to say that there are enough reliable biomarkers that meet clinical needs.

[0004] In recent years, the development of the CRISPR-Cas9 system has opened the way to investigating the phenotypes of genes with artificially induced single nucleotide mutations and examining the correlation between disease and gene mutations. However, to increase the variety of single nucleotide mutations, various gRNAs are needed to evenly guide Cas9, the deaminases used in combination with it, and other "editing enzymes" throughout the entire target genome.

[0005] However, designing and creating various gRNAs across the entire target genome based on genome sequence information and other data would require considerable expense and time, making it impractical.

[0006] On the other hand, there have been reports of methods for creating gRNAs based on randomly fragmented genome sequences and using them for genome editing with the CRISPR-Cas9 system (Non-Patent Literature 1). However, this method is intended only for experimental and research-level gRNA production and does not take into account the practical issues of cost and effort that arise when comprehensively editing the genomes of various target genes.

[0007] In other words, the technology described in Non-Patent Document 1 does not take into account any efforts to rapidly and easily produce a sufficient amount of gRNA from a very small amount of genome for use in confirming phenotypic changes in cells or model animals. Furthermore, the technology described in Non-Patent Document 1 only confirms that gRNA targeting a specific genomic region has been produced, and does not take into account the execution of comprehensive genome editing, such as single nucleotide mutation introduction, or the use of the analysis results thereof.

[0008] Japanese Patent Application No. 2019-510221

[0009] A Molecular Chipper technology for CRISPR sgRNA library generation and functional mapping of noncoding regions (NATURE COMMUNICATIONS | 7:11178 | DOI: 10.1038 / ncomms11178 |) Preclinical assessment of combination therapy of EGFR tyrosine kinase inhibitors in a highly heterogeneous tumor model. Oncogene. 2022 Apr;41(17):2470-2479. doi: 10.1038 / s41388-022-02263-4. Epub 2022 Mar 19. PMID: 35304574; PMCID: PMC9033582.Pan-KRAS inhibitor disables oncogenic signaling and tumor growth. (Nature | Vol 619 | 6 July 2023)

[0010] The present inventors have developed a gRNA production technique using a novel stem-loop type adapter, and have further discovered that by appropriately combining this with PCR technology and the like, it is possible to produce a large amount of genome editing gRNA from a much smaller amount of polynucleotide fragment material than before, leading to the present invention. The objective of this invention is to establish a polynucleotide fragment adapter and a method for using it that enables the easy, simple, rapid, and inexpensive production of gRNA libraries.

[0011] The above-mentioned objective is achieved by the following invention.

[0012] (First Invention) An adapter for preparing a gRNA library from a group of double-stranded polynucleotide fragments, which is a stem-loop type adapter having the structure of A) below.

[0013] A) Structures having the polynucleotide chains described in i) to iii) below:

[0014] i) A polynucleotide chain having the sequences i)-1 to i)-3 below i)-1: Forward primer corresponding sequence i)-2: Recognition site sequence of a restriction enzyme capable of cleaving a gRNA spacer sequence from the 5' end of the polynucleotide fragment bound to the adapter i)-3: Recognition site sequence of a restriction enzyme capable of cleaving an adapter having the structure of A) above from the polynucleotide fragment

[0015] ii) A polynucleotide chain having a sequence complementary to the polynucleotide chain in i) above.

[0016] iii) A polynucleotide chain having a single-stranded polynucleotide sequence for loop formation, which connects the 5' end of the polynucleotide chain of i) and the 3' end of the polynucleotide chain of ii) in a double-stranded structure formed by the complementary linkage of the polynucleotide chains of i) and ii).

[0017] (Second invention)

[0018] A polynucleotide library for gRNA production comprising a group of double-stranded polynucleotide fragments, wherein the polynucleotide fragments comprise at least one of the polynucleotide fragments described in 1) to 7) below.

[0019] 1) A polynucleotide fragment to which the adapter described in claim 1 is attached to both ends,

[0020] 2) In the polynucleotide fragment of 1) above, the loop formed by the single-stranded polynucleotide chain of the adapter iii) above is cleaved, and the polynucleotide fragment is amplified by polymerase chain reaction (PCR) using a label-bound forward primer.

[0021] 3) The polynucleotide fragments obtained by further cleaving the polynucleotide fragments of 2) with a restriction enzyme that recognizes the sequence of i)-2,

[0022] 4) A polynucleotide fragment in which an adapter having the structure of B) below is further attached to the restriction enzyme cleavage site of the polynucleotide fragment of 3),

[0023] B) Structures having the polynucleotide chains of iv) and v) below: iv) A polynucleotide chain having the sequences of iv)-1 to iv)-3 below: iv)-1: A scaffold sequence for the Cas protein or its mutant protein iv)-2: A recognition site sequence for a restriction enzyme that can cleave a portion of the adapter having the structure of B) from a polynucleotide fragment iv)-3: A sequence corresponding to a reverse primer v) A polynucleotide chain having a sequence complementary to the polynucleotide chain of iv) above

[0024] 5) The polynucleotide fragments from 4) above, purified using a label attached to a forward primer,

[0025] 6) A polynucleotide fragment obtained by PCR using the following two primers, vi)-1 and vi)-2, from the polynucleotide fragment of 4) or 5): vi)-1: Forward primer corresponding to the sequence of i)-1 vi)-2: Reverse primer corresponding to the sequence of iv)-3

[0026] 7) A polynucleotide fragment obtained by cleaving the polynucleotide fragment of 6) at the restriction enzyme recognition sites of the sequences i)-3 and iv)-2.

[0027] The second invention, in another aspect, encompasses the following invention: a polynucleotide library for gRNA production comprising a group of double-stranded polynucleotide fragments, wherein the polynucleotide fragments comprise at least one of the polynucleotide fragments described in 8) to 11) below.

[0028] 8) A polynucleotide fragment in which an adapter A' having the structure A' below is bound to one end of a gRNA spacer sequence, and an adapter B having the structure B) below is bound to the other end.

[0029] A') Structure having the polynucleotide chains of i) and ii) below: i) A polynucleotide chain having the sequences of i)-1 to i)-3 below i)-1: Forward primer corresponding sequence i)-2: Recognition site sequence of a restriction enzyme capable of cleaving the gRNA spacer sequence from the 5' end of the polynucleotide fragment bound to the adapter i)-3: Recognition site sequence of a restriction enzyme capable of cleaving the adapter having the structure of A') from the polynucleotide fragment ii) A polynucleotide chain having a sequence complementary to the polynucleotide chain of i) above

[0030] B) Structures having the polynucleotide chains of iv) and v) below: iv) A polynucleotide chain having the sequences of iv)-1 to iv)-3 below: iv)-1: A scaffold sequence for the Cas protein or its mutant protein iv)-2: A recognition site sequence for a restriction enzyme that can cleave a portion of the adapter having the structure of B) from a polynucleotide fragment iv)-3: A sequence corresponding to a reverse primer v) A polynucleotide chain having a sequence complementary to the polynucleotide chain of iv) above

[0031] 9) The polynucleotide fragments from 8) above, purified using labels attached to forward primers,

[0032] 10) A polynucleotide fragment obtained by amplifying the polynucleotide fragment of 8) or 9) by PCR using two types of primers of the following vi)-1 and vi)-2: vi)-1: A forward primer corresponding to the sequence of i)-1 vi)-2: A reverse primer corresponding to the sequence of iv)-3

[0033] 11) A polynucleotide fragment obtained by cleaving the polynucleotide fragment of 10) at the restriction enzyme recognition sites of the sequences of i)-3 and iv)-2

[0034] (Third Invention) In the library described in the second invention, in the aspect where the polynucleotide fragment contains the polynucleotide fragment of 1), each of the polynucleotide fragments is cleaved at the single-stranded polynucleotide sequence portion for forming an adapter inner loop at both ends, and amplified by polymerase chain reaction (PCR) using a labeled binding forward primer. A polynucleotide library for preparing gRNA, characterized in that the single-stranded polynucleotide sequence for forming a loop is cleaved as described above, whereby the loop is cleaved or opened.

[0035] (Fourth Invention) A polynucleotide library for preparing gRNA, characterized in that each polynucleotide fragment of the library described in the third invention is cleaved by a restriction enzyme for i)-2

[0036] (Fifth Invention) A polynucleotide library for preparing gRNA, characterized in that an adapter having the following structure B is bound to the restriction enzyme cleavage surface of each polynucleotide fragment in the library described in the fourth invention

[0037] B) An adapter (Adapter B) having the following sequences of iv) and v)

[0038] iv) A polynucleotide chain having the following sequences of iv)-1 to iv)-3: iv)-1: A scaffold sequence for a Cas protein or its mutant protein iv)-2: A recognition site sequence of a restriction enzyme capable of cleaving a part of the adapter of B) from the polynucleotide fragment iv)-3: A reverse primer corresponding sequence

[0039] v) Polynucleotide chain having a sequence complementary to iv)

[0040] (Sixth Invention) A polynucleotide library for gRNA production, characterized in that the library described in the fifth invention is purified using a label bound to a forward primer as an indicator.

[0041] (Seventh Invention) This is a polynucleotide library for gRNA production, characterized in that the library described in the sixth invention is amplified by PCR using two types of primers, vi)-1 and vi)-2, as described below.

[0042] vi)-1: Forward primer corresponding to i)-1 vi)-2: Reverse primer corresponding to iv)-3

[0043] (Eighth Invention) A polynucleotide library for gRNA production, characterized in that each polynucleotide fragment in the library described in the seventh invention is cleaved at the restriction enzyme recognition sites described in i)-3 and iv)-2 below.

[0044] i)-3: Recognition site sequence of restriction enzyme capable of cleaving adapter A) from a polynucleotide fragment. iv)-2: Recognition site sequence of restriction enzyme capable of cleaving a portion of adapter B) from a polynucleotide fragment.

[0045] (Ninth Invention) A gRNA expression vector library characterized in that each polynucleotide for gRNA production in the library described in the second invention is introduced into a vector capable of expressing an exogenous gene.

[0046] (The Tenth Invention) See below and <ii>These are genome-edited or epigenome-edited cultured cells characterized by the introduction of [a specific technology / method].

[0047] Vectors included in the gRNA expression vector library described in the ninth invention <ii>A vector expressing the Cas protein or its mutant protein.

[0048] (Invention 11) The cultured cells described in Invention 10 are further modified as follows: <iii>This is a genome-edited or epigenome-edited cultured cell characterized by the introduction of a vector.

[0049] <iii>A vector expressing an enzyme having genome editing activity or epigenome editing activity.

[0050] (The twelfth invention) An adapter set for preparing a gRNA library from a group of double-stranded polynucleotide fragments, characterized by comprising adapter A having the structure of A) below and adapter B having the structure of B) below.

[0051] A) Structures having the polynucleotide chains described in i) to iii) below:

[0052] i) A polynucleotide chain having the sequences i)-1 to i)-3 below i)-1: Forward primer corresponding sequence i)-2: Recognition site sequence of a restriction enzyme capable of cleaving a gRNA spacer sequence from the 5' end of the polynucleotide fragment bound to the adapter i)-3: Recognition site sequence of a restriction enzyme capable of cleaving an adapter having the structure of A) above from the polynucleotide fragment

[0053] ii) A polynucleotide chain having a sequence complementary to the polynucleotide chain in i) above.

[0054] iii) A polynucleotide chain having a single-stranded polynucleotide sequence for loop formation, which connects the 5' end of the polynucleotide chain of i) and the 3' end of the polynucleotide chain of ii) in a double-stranded structure formed by the complementary linkage of the polynucleotide chains of i) and ii).

[0055] B) Structures having the polynucleotide chains of iv) and v) below:

[0056] iv) A polynucleotide chain having the sequences iv)-1 to iv)-3 below: iv)-1: A scaffold sequence for the Cas protein or its mutant protein iv)-2: A recognition site sequence of a restriction enzyme capable of cleaving a portion of the adapter having the structure of B) from the polynucleotide fragment iv)-3: A sequence corresponding to a reverse primer

[0057] v) A polynucleotide chain having a sequence complementary to the polynucleotide chain of iv) above.

[0058] (The thirteenth invention) A kit for preparing a gRNA library, characterized by comprising the following <X> and <Y>.

[0059] <X> Adapter according to the first invention or adapter set according to the twelfth invention

[0060] <Y> At least one of the enzymes listed below <Y>-1 to <Y>-10

[0061] <Y>-1: Enzyme for converting the target gene into a polynucleotide fragment <Y>-2: Enzyme for binding adapter A of the first invention or adapter A of the adapter set described in the twelfth invention to the polynucleotide fragment for gRNA production <Y>-3: Enzyme for binding adapter B) described in the twelfth invention to the polynucleotide fragment for gRNA production <Y>-4: Enzyme for cleaving the single-stranded polynucleotide sequence for loop formation in adapter A <Y>-5: Restriction enzyme capable of cleaving the gRNA spacer sequence from the polynucleotide fragment for gRNA production bound to the sequence derived from adapter A <Y>-6: Restriction enzyme for cleaving and removing the sequence derived from adapter A after loop removal <Y>-7: Restriction enzyme for cleaving and removing the reverse primer corresponding sequence in adapter B <Y>-8: Heat-stable polymerase <Y>-9: Cas protein or its mutant protein <Y>-10: Enzyme having genome editing activity or epigenome editing activity

[0062] (Invention 14) A method for preparing a polynucleotide library for gRNA production from a group of polynucleotide fragments, characterized by comprising the following steps (I) to (X).

[0063] (I) A step of smoothing both ends of the polynucleotide fragment as necessary. (II) A step of attaching adapter A to both ends of the polynucleotide fragment. (III) A step of cleaving at least one single-stranded polynucleotide chain of adapter A attached to the polynucleotide fragment in step (II) below. (IV) A step of amplifying the polynucleotide fragment obtained in step (III) by performing PCR using a forward primer to which a label has been attached. (V) A step of treating the polynucleotide fragment amplified in step (IV) with a restriction enzyme that recognizes the sequence i)-2 below. (VI) A step of attaching adapter B below to the end of the polynucleotide fragment treated in step (V) that was cleaved by the restriction enzyme (usually the 3' end of the polynucleotide chain below i). (VII) A step of purifying and / or concentrating the polynucleotide fragment obtained in step (VI) using the label attached to the forward primer as an indicator. (VIII) A step of amplifying the polynucleotide fragment obtained in step (VII) by performing PCR using the forward primer and reverse primer. (IX) A step of treating the polynucleotide fragment obtained in step (VIII) with restriction enzymes that recognize the sequences i)-3 and iv)-2 below. (X) A step of inserting the polynucleotide fragment after treatment with restriction enzymes in step (VIII) into a vector capable of expressing an exogenous gene to produce a gRNA expression vector.

[0064] A) A stem-loop adapter A for gRNA library preparation having a structure with polynucleotide chains as described in i) to iii) below:

[0065] i) A polynucleotide chain having the sequences i)-1 to i)-3 below i)-1: Forward primer corresponding sequence i)-2: Recognition site sequence of a restriction enzyme capable of cleaving a gRNA spacer sequence from the 5' end of a polynucleotide fragment bound to the adapter i)-3: Recognition site sequence of a restriction enzyme capable of cleaving an adapter having the structure of A) from the polynucleotide fragment

[0066] ii) A polynucleotide chain having a sequence complementary to the polynucleotide chain in i) above.

[0067] iii) A polynucleotide chain having a single-stranded polynucleotide sequence for loop formation, which connects the 5' end of the polynucleotide chain of i) and the 3' end of the polynucleotide chain of ii) in a double-stranded structure formed by the complementary linkage of the polynucleotide chains of i) and ii).

[0068] B) Adapter B having a structure having the polynucleotide chains described in iv) and v) below:

[0069] iv) A polynucleotide chain having the sequences iv)-1 to iv)-3 below: iv)-1: A scaffold sequence for the Cas protein or its mutant protein iv)-2: A recognition site sequence for a restriction enzyme capable of cleaving a portion of the adapter having the structure of B) from a polynucleotide fragment iv)-3: A sequence corresponding to a reverse primer v) A polynucleotide chain having a sequence complementary to the polynucleotide chain of iv) above

[0070] (The Fifteenth Invention) A method for identifying gene mutations that are candidates for causing changes in the phenotype to be analyzed, characterized by comprising the following steps (XIV) to (XVI).

[0071] (XIV) A step of editing the genome of genes in cultured cells by culturing the cultured cells described in the tenth invention. (XV) A step of selecting cultured cells or non-human animal models with altered phenotypes to be analyzed from the cultured cells obtained in step (XIV) or from non-human animal models transplanted with them. (XVI) A step of comparing the genes extracted from the cultured cells or non-human animal models selected in step (XV) with the genes before genome editing.

[0072] (The sixteenth invention) A method for determining the potential influence of a genome on the phenotype, characterized by comprising the steps S1) to S4) below.

[0073] S1) A step of selecting gene mutations that are candidates for causing phenotypic changes to be analyzed using the method described in the fifteenth invention. S2) A step of creating a list of candidate gene mutations enumerated and / or ranked based on the results of step S1. S3) A step of comparing the information obtained from the genome with the list from step S2. S4) A step of determining the likelihood of the phenotypic change to be analyzed.

[0074] (Invention No. 17) The method according to the invention No. 16, further characterized by including step S5 below.

[0075] S5) A process to make a decision by combining at least one of the following pieces of information with the result of process S4: S5)-1: Metabolome analysis information data S5)-2: Proteome analysis information data

[0076] (Invention No. 18) A determination system comprising: an information acquisition device that stores information on gene mutations obtained by the method for identifying gene mutations of Invention No. 15, and acquires gene information of the genome contained in a sample; and a determination device that compares the gene information acquired by the information acquisition device with the gene mutation information stored in the memory device, and determines the possibility of a phenotypic change to be analyzed based on the comparison result.

[0077] (Invention No. 19) A biomarker for determining EGFR gene hyperfunction, comprising an EGFR protein having any of the following amino acid mutations, or a polynucleotide having an EGFR gene mutation capable of causing any of the following amino acid mutations.

[0078] p.Q32R, p.G109E, p.S116F, p.P195S, p.A237V, p.S306=, p.R324H, p.G331=, p.F376L, p.H433Q, p.K454E, p.W 477*, p.C555R, p.G614S, p.P631S, p.A647T, p.L655=, p.R675W, p.I744V, p.T751I, p.P794L, p.V802A, p.H80 5R, p.C818R, p.P848=, p.K852R, p.K875R, p.A882V, p.Q894=, p.T940=, p.F968=, p.M987R, p.P992S, p.F997V , p.E1005=, p.D1014V, p.L1034I, p.C1058Y, p.D1084N, p.S1096=, p.R1100S, p.P1108=, p.A1181T, p.V1200=

[0079] (Twentieth Invention) A biomarker for determining hyperfunction of the KRAS gene, comprising a KRAS protein having any of the following amino acid mutations, or a polynucleotide having a KRAS gene mutation capable of causing any of the following amino acid mutations.

[0080] p.T2I, p.V7A, p.V7G, p.V8E, p.V9=, p.V9D, p.G10=, p.G10V, p.A11P, p.A11V , p.G12=, p.V14A, p.G15=, p.G15C, p.S17T, p.A18D, p.A18V, p.T20S, p.I21K, p.Q22L, p.Q22H, p.I24N, p.Q25*, p.H27Y, p.D30E, p.P34Q, p.P34S, p.P34T, p.T35=, p.I36L, p.S39=, p.S39Y, p.K42=, p.V44E, p.E49K, p.C51*, p.C51=, p .D57Y, p.G60=, p.E62D, p.Y64N, p.A66S, p.R68W, p.L79I, p.C80S, p.K88*, p .S89*, p.D92Y, p.I93F, p.R97I, p.E98*, p.E98=, p.K104N, p.P110S, p.V112I , p.V114I, p.G115R, p.D119N, p.P121H, p.S122F, p.R123*, p.R123I, p.D126V , p.Q129*, p.A134T, p.A134V, p.Y137=, p.G138R, p.P140S, p.R149K, p.R151T

[0081] (Invention 21) A biomarker for determining BRAF gene enhancement, comprising a BRAF protein having any of the following amino acid mutations, or a polynucleotide having a BRAF gene mutation capable of causing any of the following amino acid mutations.

[0082] p.W48*, p.L64=, p.S122F, p.K206E, p.L312P, p.T401=, p.Q494*, p.R509Q, p.R509L, p.L514P, p.L514I, p.Y519=, p.T521K, p.A526V, p.W531C, p.H540Q, p.L553R, p.A561S, p.L567=, p.A569S, p.I572F, p.H574N , p.L577I, p.N581I, p.N581S, p.L584I, p.L588R, p.L588F, p.F595I, p.F595L, p.G596C, p.G596V, p.G59 6S, p.A598=, p.V600L, p.V600E, p.S602Y, p.R603=, p.W604C, p.W604*, p.S605R, p.G606=, p.G606V, p.Q6 09=, p.E611K, p.Q612E, p.S614Y, p.M620I, p.V624=, p.R626I, p.D629Y, p.D629G, p.P632L, p.F635I, p. Q636H, p.Y647*, p.E648Q, p.M650V, p.Q653K, p.Q653*, p.S657*, p.I659=, p.R662S, p.D663N, p.D663A, p .F667V, p.R671L, p.R671*, p.L674=, p.R682W, p.P705Q, p.E715*, p.R719P, p.H725Q, p.S727I, p.S732=, p.N734S, p.R735Q, p.T740K, p.E741=, p.C748S, p.C748Y, p.I755=, p.G758=, p.G758E, p.G759E, p.G759R

[0083] (Invention 22) A method for testing for EGFR gene hyperfunction, comprising comparing information on the EGFR protein or its gene based on a biological sample from a subject with information consisting of a dataset containing the biomarker of Invention 19 to detect whether there are matching mutations.

[0084] (Invention 23) A method for testing for hyperfunction of the KRAS gene, comprising comparing information on the KRAS protein or its gene based on a biological sample from a subject with information consisting of a dataset containing the biomarker of Invention 20, to detect whether there are matching mutations.

[0085] (Invention 24) A method for testing BRAF gene hyperfunction, comprising comparing information on the BRAF protein or its gene based on a biological sample from a subject with information consisting of a dataset containing the biomarker of Invention 21, to detect whether there are matching mutations.

[0086] By using the "adapter," "polynucleotide library for gRNA production," "vector library for gRNA expression," or "method for producing a polynucleotide library for gRNA expression" of the present invention, a diverse range of gRNA libraries, with target sites spanning the entire target genome, can be produced simply and quickly, even from very small amounts of raw materials. The libraries obtained in this manner can be used for so-called genome editing (e.g., editing of base sequences such as cleavage, mutation introduction, and prime editing), as well as epigenome editing (control of genome function (activation, suppression), etc.). In particular, by performing genome editing or epigenome editing in the "genome-edited or epigenome-edited cultured cells" of the present invention into which the vector contained in the "vector library" of the present invention has been introduced, it is possible to identify gene mutations that are candidate factors for changes (worsening or improvement) in various phenotypes, such as cancer and other diseases with a large number of patients, as well as rare diseases, changes in physical condition and / or constitution, or drug resistance, more rapidly than in conventional methods. Furthermore, since genome editing sites are not limited to exon regions, it becomes possible to analyze various changes in gene function caused by mutations in non-coding regions (of proteins). In addition, by performing prime editing using the gRNA expression vector library of the present invention, it becomes possible to rapidly and easily impart new functions to genes.

[0087] Figure 1 shows an example of the sequence of "Adapter A" used in the present invention (Adap_A). Sequence ID 1 shows the sense strand of the double-stranded portion of "Adapter A," and Sequence ID 2 shows the antisense strand of the double-stranded portion. Sequence ID 9 shows the single-stranded nucleotide sequence for loop formation. Figure 2 shows an example of the sequence of "Adapter B" used in the present invention (Adap_B_v1). Sequence ID 3 shows the sense strand of "Adapter B," and Sequence ID 4 shows the antisense strand. In the figure, "N" represents one of "A," "T," "G," or "C," and "Adap_B_v1" has "NN," so Figure 2 means that there are 16 different Adapter B combinations that cover all of these base combinations. Figure 3 shows an example of the sequence of "Adapter B" used in the present invention (Adap_B_v2). Sequence ID 5 shows the sense strand of "Adapter B," and Sequence ID 6 shows the antisense strand. In the figure, "N" represents one of "A", "T", "G", or "C", and "Adap_B_v2" has "NN", meaning that Figure 3 represents 16 different adapter B sequences that cover all of these base combinations. Figure 4 is a diagram showing an example of the sequence of "adapter B" used in the present invention (Adap_B_v3). Sequence ID 7 is the sense strand of "adapter B", and sequence ID 8 is the antisense strand. In the figure, "N" represents one of "A", "T", "G", or "C", and "Adap_B_v3" has "NN", meaning that Figure 4 represents 16 different adapter B sequences that cover all of these base combinations. Figure 5 is a schematic diagram showing the procedure for the "method for producing a polynucleotide fragment library for gRNA production" of the present invention. Figure 6 is a plasmid map (schematic diagram) showing an example of a vector into which the "polynucleotide for gRNA production" of the present invention has been introduced. Figure 7 shows the length of the "spacer sequence" in a gRNA expression vector prepared using the adapter A and B set of the present invention, as confirmed by a next-generation sequencer. Figure 8 shows the results of genome editing of the KRAS gene using the gRNA expression vector of the present invention, and the single nucleotide substitution sites and mutation rates investigated by a next-generation sequencer. Figure 9 shows the results of confirming gene mutations in cancerous cells for KRAS and EGFR.Y > X: Single nucleotide mutations located to the upper left of the line X = Y in the figure are considered to be highly likely to be deeply related to carcinogenesis. Figure 10 shows the single nucleotide mutations and amino acid changes that are considered to be representative causes of carcinogenesis, as confirmed in Figure 9, and the Y / X value (for each mutation, the percentage change in mutation rate after culture compared to before culture in a low-adhesion dish: Enrichment score). Since known single nucleotide mutations involved in carcinogenesis have also been detected, it is shown that the method using the gRNA library of the present invention can be used for the detection of cancer-causing genes. Figure 11 is a diagram showing an example of the "determination system" of the present invention, which has a program that executes each step of the "determination method" of the present invention. Figure 12 shows the results of using the "polynucleotide library for gRNA production" of the present invention for "control of target gene function (activation)" by epigenome editing. Figure 13 shows the results of using the "polynucleotide library for gRNA production" of the present invention for "deletion of gene function due to deletion of DNA sequence" by genome editing. Figure 14 shows the results of investigating the effect of differences in the structure of "adapter A" on the gRNA library production efficiency. Figure 15 is a diagram of the system of the present invention. Figure 16 is a list of EGFR protein mutations identified by the present invention. Figure 17 is a list of KRAS protein mutations identified by the present invention. Figure 18 is a list of KRAS protein mutations identified by the present invention. Figure 19 is a list of KRAS protein mutations identified by the present invention. Figure 20 is a list of BRAF protein mutations identified by the present invention. Figure 21 is a list of BRAF protein mutations identified by the present invention. Figure 22 is a graph showing that single nucleotide mutations identified by the present invention result in enhanced cancer cell-specific proliferative capacity.

[0088] The present invention will be described in detail below, but before that, the definitions of each term and abbreviation, the sequence indicated by the sequence number, etc. will be explained.

[0089] Definitions of Terms

[0090] (Definition of Adapter) In the initial stage of creating a polynucleotide for gRNA production, as described in the first invention of this invention, the adapter A) used to attach to both ends of a polynucleotide fragment is called "Adapter A." In the process of creating a "polynucleotide for gRNA production," as described in the fifth invention, the adapter B) used to attach to the 3' end after the "spacer sequence" has been excised is called "Adapter B." (These may be written as "Adap_A" and "Adap_B," respectively.)

[0091] Definition of gRNA) In this invention, unless otherwise specified, "gRNA" means "gRNA in the broad sense" as described below, but in some cases it may also mean "gRNA in the narrow sense" as described below.

[0092] "gRNA in the broad sense": A sequence that includes the "spacer sequence" described below, and the "scaffold sequence" (tracrRNA) for Cas protein binding, etc.

[0093] "gRNA in the narrow sense": A "spacer sequence" that binds complementarily to the editing target site on the target genome in the CRISPR-Cas system.

[0094] Furthermore, "gRNA in a broad sense" also includes "prime editing gRNA (pegRNA)" used to introduce new sequences into the genome rather than replacing existing genome sequences. In addition to the above, prime editing gRNA includes the following sequences.

[0095] Editing (insertion) sequence (Editing Template) Primer binding sequence (PBS) that leads to the Editing Template

[0096] Definition of Cas Protein) In this invention, "Cas protein (CRISPR-associated protein) or its variant protein" is a general term for enzymes used in the CRISPR-Cas system, and generally refers to "Cas9." However, in this invention, it includes all of the following Cas9, Cas12, Cas13, or their variants. In the description of the invention below, "Cas9" may be used as a representative example of "Cas protein," but this is not intended to limit the explanation to "Cas9." Basically, it applies to all enzymes belonging to "Cas protein or its variant protein."

[0097] (Definition of genome editing activity) In this invention, "genome editing activity" means the activity of changing the base sequence of DNA itself.

[0098] (Definition of Epigenome Editing Activity) In the present invention, "epigenome editing activity" refers to an activity that performs some kind of editing on a region containing targeted DNA without changing the DNA base sequence itself. Specifically, it means an activity that can control (activate or suppress, etc.) genome function by, for example, adding or removing methyl groups or acetyl groups to target sites of DNA or histones.

[0099] Note that "genome editing" and "epigenome editing" are sometimes collectively referred to as "(epi)genome editing."

[0100] (Definition of Genome) The term "genome" generally means "the smallest unit of genetic information that defines a particular species," but in this invention, it may mean "the polynucleotide itself (the totality of genes or individual genes)" that forms the basis of said "genetic information."

[0101] 《Definition of Abbreviations》 The following abbreviations may be used for each term used in this invention. Forward primer: Fwd primer (or sometimes simply written as Fwd) Reverse primer: Rev primer (or sometimes simply written as Rev) Sequence notation: "A" represents adenine, "T" represents thymine, "G" represents guanine, "C" represents cytosine, and "U" represents uracil.

[0102] 《Sequence Number》 Note that the underlined portion in the sequence below indicates the "MS2 Stemloop sequence". The "MS2 Stemloop sequence" is a sequence necessary to guide functional proteins other than Cas proteins to the genome-binding region of the gRNA-Cas protein. Note that all sequence numbers are represented from the 5' end to the 3' end.

[0103] Sequence ID 1: GTAAAACGACGGCCAGTCAGCAGGGATCCG

[0104] Sequence ID 2: CGGATCCCTGCTGACTGGCCGTCGTTTTAC

[0105] Sequence ID 3: NN GTTTTAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTTCTCGAGGTCATAGCTGTTTCCTG CCCC

[0106] Sequence ID 4: CAGGAAACAGCTATGAC CTCGAG AAAAAAGCACCGACTCGGTGCCACTTTTTCAAGTTGATAACGGACTAGCCTTATTTTAACTTGCTATTTCTAGCTCTAAAAC

[0107] SEQ ID NO: 5: NN GTTTTAGAGCTAGGCCAACATGAGGATCACCCATGTCTGCAGGGCCTAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGGCCAACATGAGGATCACCCATGTCTGCAGGGCCAAGTGGCACCGAGTCGGTGCTTTTTTTCTCGAGGTCATAGCTGTTTCCTGCCCC

[0108] SEQ ID NO: 6: GGGGCAGGAAACAGCTATGACCTCGAGAAAAAAAGCACCGACTCGGTGCCACTTGGCCCTGCAGACATGGGTGATCCTCATGTTGGCCAAGTTGATAACGGACTAGCCTTATTTTAACTTGCTAGGCCCTGCAGACATGGGTGATCCTCATGTTGGCCTAGCTCTAAAAC

[0109] SEQ ID NO: 7: NN GTTTAAGAGCTAaGCCAACATGAGGATCACCCATGTCTGCAGGGCaTAGCAAGTTTAAATAAGGCTAGTCCGTTATCAACTTGGCCAACATGAGGATCACCCATGTCTGCAGGGCCAAGTGGCACCGAGTCGGTGCTTTTTTT CTCGAG GTCATAGCTGTTTCCTG CCCC

[0110] SEQ ID NO: 8: CAGGAAACAGCTATGACCTCGAGAAAAAAAGCACCGACTCGGTGCCACTTGGCCCTGCAG ACATGGGTGATCCTCATGT TGGCCAAGTTGATAACGGACTAGCCTTATTTAAACTTGCTAtGCCCTGCAGACATGGGTGATCCTCATGTTGGCtTAGCTCTTAAAC

[0111] Sequence ID 9: GTCGTATCCAGTGC AGGGUCCGAGGTATTC GCACTGGATACGAC

[0112] Furthermore, in sequence numbers 3, 5, and 7 above, "N" represents one of "A", "T", "G", or "C", and each sequence containing "NN" means that there are 16 possible combinations of these sequences.

[0113] [[The "Adapter (Adapter A)" of the present invention]] The adapter of the present invention will be described in detail below.

[0114] [Aspect 1 of "Adapter A" of the present invention] The stem-loop type adapter for gRNA library preparation of the present invention is an adapter for preparing a gRNA library from a double-stranded polynucleotide fragment group, and is characterized by having the structure A) described below.

[0115] A) an adapter having the sequence A) i) to iii) (Adapter A)

[0116] i) Polynucleotide chains having the sequences i)-1 to i)-3 below

[0117] i)-1: Forward primer corresponding sequence i)-2: Recognition site sequence of restriction enzyme capable of cleaving the gRNA spacer sequence from the 5' end of the polynucleotide fragment bound to adapter A i)-3: Recognition site sequence of restriction enzyme capable of cleaving adapter A from the polynucleotide fragment

[0118] ii) A polynucleotide chain having a complementary sequence to i)

[0119] iii) A single-stranded polynucleotide sequence for loop formation, connecting the ends opposite to the side where the polynucleotide fragments of the double-stranded polynucleotide sequence consisting of complementaryly linked i) and ii) are bound.

[0120] <Role of Adapter A> In this invention, the "stem-loop type adapter for gRNA library production" has the role of introducing restriction enzymes that allow for the simple and reliable excision of polynucleotides of a suitable length for gRNA from genome-derived "polynucleotide fragments" that serve as raw materials for gRNA. It also acts as a protecting group that protects each "polynucleotide fragment" until the excised "polynucleotides for gRNA production" are amplified to a substantially usable concentration (number), and is ultimately detached from the "polynucleotides for gRNA production".

[0121] Furthermore, a "stem-loop adapter" refers to a double helix formed by base pairing between two regions within the same chain, usually regions with complementary nucleotide sequences, and where a loop region that does not form a complementary pair exists at one end of the double helix.

[0122] <Target of Adapter A> As described above, "Adapter A" binds to both ends of each fragment in the "double-stranded polynucleotide fragment group". Examples of this "double-stranded polynucleotide fragment group" include a) and b) below.

[0123] a) Various target genes (disease genes such as cancer genes, disease suppressor genes, genomic regions that control the activation of disease genes, etc.) extracted from living organisms and targeted for genome editing and epigenome editing, their cDNA, genes amplified using genetic engineering techniques, and genes artificially synthesized from the sequence information of these target genes, are randomly cut by enzymes, etc., resulting in "assemblies of different polynucleotide fragments."

[0124] (i) A collection of different partial sequences artificially created from the outset as "fragments" based on known genome sequence information, etc.

[0125] Furthermore, considering the possibility of single nucleotide polymorphisms, etc., it may be acceptable to use genes obtained from living organisms as raw materials, particularly genes from multiple individuals.

[0126] In phrases like "different polynucleotide fragments" or "different subsequences," "different" means that they differ in "length" and / or "base sequence," etc.

[0127] Furthermore, each fragment within the above group may be simply referred to as "(raw material) polynucleotide fragment" below.

[0128] While there are no particular limitations on the method of random cleavage, the use of fragmentase is preferred. Fragmentase fragments the entire target genome into "polynucleotide fragments" consisting of 50 to 1000 base pairs evenly by fragmenting double-stranded DNA in a time-dependent manner. This reduces the risk of fragmentation being biased towards specific regions, as is the case with restriction enzymes that require recognition of specific regions, and also eliminates the time required for specific region recognition.

[0129] Furthermore, since the ends of the "polynucleotide fragments" prepared with Fragmentase usually contain a mixture of protruding and blunt ends, it is preferable to perform a blunting (fill-in, or blunting) treatment on both ends of the "polynucleotide fragment" using a known method before binding to "Adapter A" of the present invention, which has a smooth 3' binding surface.

[0130] <i) Explanation>

[0131] This i) is an array represented, for example, by sequence number 1 in Figure 1, and each of the constituent sequences will be explained below.

[0132] i)-1: Regarding forward primer corresponding sequences

[0133] i)-1's "forward primer compatible sequence" is the sequence that forms the basis for designing the "forward primer" used to amplify the polynucleotide fragment to which this stem-loop type adapter is attached at both ends by PCR or other methods, and means the "same sequence" as the "forward primer". During PCR, the "forward primer" binds to the complementary strand side of ii) ("adapter A").

[0134] Specifically, examples include, but are not limited to, M13 forward primers (M13_Fwd primer) and other universal primer sequences.

[0135] The specific sequence of the M13_Fwd primer is as described in Figure 1 (Sequence No. 1) as the constituent sequence of "Adapter A".

[0136] By pre-integrating such PCR primer-compatible sequences into the adapter, it becomes possible to create gRNA libraries from extremely small amounts of DNA fragments. Furthermore, by amplifying only the DNA fragments containing restriction enzyme recognition sequences, it contributes to improved production efficiency.

[0137] Furthermore, in Non-Patent Document 1, it is sufficient to produce gRNAs that cover the entire genome, and no techniques for producing gRNAs from extremely small amounts of DNA fragments are used.

[0138] i)-2: Recognition site sequence of restriction enzyme capable of cleaving the gRNA spacer sequence from the 5' end of a polynucleotide fragment bound to adapter A.

[0139] A "spacer" is a sequence consisting of approximately 17 to 23 base pairs that designates the target site of the target gene (genome) to be edited within the gRNA. By binding the gRNA to the corresponding genomic location via this spacer, epigenome editing at the desired location becomes possible.

[0140] In other words, the "restriction enzyme capable of cleaving the gRNA spacer sequence" described in i)-2 (hereinafter sometimes referred to as "restriction enzyme for i)-2") refers to a restriction enzyme that can cleave a portion of the gRNA slightly away from the recognition site, specifically a portion downstream of the recognition site by the length of the gRNA spacer (approximately 17 to 23 base pairs from the 5' end of the "raw material polynucleotide fragment"), in addition to the lengths described in i)-3 and i)-4 below. A type III restriction enzyme is a preferred example.

[0141] However, if the "i)-4: Sequence for the transcription start site of the gRNA transcription promoter" described later is further included near the 3' end of "Adapter A", the number of base pairs corresponding to that start site may be excluded from the above "17 to 23 base pairs".

[0142] Since type III restriction enzymes have a cleavage site located 10 to 20 base pairs away from the recognition site of the restriction enzyme, by incorporating such a restriction enzyme recognition site into an appropriate position (near the 3' end) within "Adapter A," it becomes possible to easily design "Adapter A" that is suitable for excising a "gRNA spacer sequence" of the above length from a "polynucleotide fragment."

[0143] Examples of type III restriction enzymes include EcoP1I, EcoP15I, HinfIII, and PsrII. However, EcoP15I is preferred because its cleavage site is approximately 25 to 27 bases ahead of the 17 to 23 bases required for the "spacer sequence," and it is commercially available and easily obtainable.

[0144] Furthermore, the length of the gRNA spacer is usually preferably 17 to 23 base pairs, and more preferably 18 to 20, but it is not necessarily limited to these values.

[0145] i)-3: Recognition site sequences of restriction enzymes capable of cleaving adapter A from polynucleotide fragments.

[0146] i) The restriction enzyme that recognizes -3 (hereinafter sometimes referred to as "restriction enzyme for i)-3") is not particularly limited, but in order to avoid making the design of "Adapter A" too redundant and complex, it is preferable to use a type II restriction enzyme in which the recognition site and the cleavage site are not too far apart.

[0147] Examples of type II restriction enzymes include BamHI, PstI, HindIII, EcoRI, XhoI, SalI, and SmaI, which have protruding ends at the cleavage site, and HaeIII, which have blunt ends at the cleavage site. However, if it is anticipated that polynucleotide fragments for gRNA production will be introduced into a vector after cleaving "Adapter A", then enzymes with protruding ends at the cleavage site are preferable so that they can bind specifically and strongly through phosphodiester bonds between the polynucleotide fragments that adhere (including the meaning of binding or linking) and hydrogen bonds between the complementary chains. In particular, for the efficient construction of the "gRNA expression vector library" of the present invention described later, enzymes with a recognition site of 6 bases or more and high sequence specificity are preferred, and BamHI is a good example of such an enzyme.

[0148] Furthermore, in i), i)-1, i)-2, and i)-3 are basically arranged in this order from the 5' side.

[0149] In particular, the restriction enzyme recognition site i)-3 must be downstream of the restriction enzyme recognition site i)-2. This is because if the restriction enzyme recognition site i)-2 remains after the cleavage of "adapter A," it may impair the function of the gRNA.

[0150] <Explanation of ii)>

[0151] This ii) is a polynucleotide chain having a sequence complementary to i), as represented, for example, by Sequence ID No. 2 in Figure 1.

[0152] <Explanation of iii>

[0153] The "single-stranded polynucleotide sequence for loop formation (hereinafter sometimes referred to as the "loop sequence")" in iii) preferably incorporates a mechanism that allows for the cleavage or removal of this loop sequence before amplification by PCR or the like. For example, at least a portion of the loop sequence may be a "sequence that can be subjected to enzymatic degradation."

[0154] Examples of sequences that can be enzymatically degraded include sequences such as iii)-1 and iii)-2 below.

[0155] iii)-1: Uridine

[0156] If uridine is present in the DNA sequence, it is possible to cleave single-stranded DNA by excising the uridine portion, for example, using "USER(registered trademark) II Enzyme" as described below.

[0157] USER (Registered Trademark) II Enzyme (New England Biolabs, Inc.): A mixture of UDG and endonuclease VIII.

[0158] UDG (Uracil DNA Glycosylase): An enzyme that hydrolyzes the N-glycosidic bonds of uracil-containing DNA, releasing uracil and forming a depyrimidine site. Endonuclease VIII: An enzyme that can cleave the phosphodiester bonds at both ends of the depyrimidine site, thereby cleaving the DNA strand.

[0159] iii)-2: Sequence corresponding to the recognition region of restriction enzymes that act on single-stranded DNA

[0160] Examples of restriction enzymes that act on single-stranded DNA include, but are not limited to, the following.

[0161] SI nuclease (manufactured by Takara Bio Inc.): Recognizes and excavates single-stranded segments. RsaI restriction enzyme (manufactured by Nippon Gene Inc.): Recognition sequence 5'...GT▼AC...3' (▼ indicates the cleavage site)

[0162] The length of the "loop sequence" in iii) is not particularly limited, but it must be long enough to properly form a loop (usually about 4 bases or more) and not impair the original function of the adapter. Specifically, examples include the loop-forming sequences listed below, which are commonly used when forming stem-loop structures. Since the length of these sequences is about 19 to 50 bases, it is considered preferable that the length be about 10 to 60 bases.

[0163] https: / / www.ncbi.nlm.nih.gov / pmc / articles / PMC5997974 / table / t0005 / https: / / respiratory-research.biomedcentral.com / articles / 10.1186 / s12931-017-0528-7 / tables / 2

[0164] In the present invention, by pre-integrating a loop array like iii) into "Adapter A", the following effects can be achieved.

[0165] - It prevents the adapter from binding to the polynucleotide fragment in the opposite direction to the one it was originally designed for. - It prevents or suppresses self-ligation between adapters, dramatically improving amplification efficiency in subsequent PCR and other processes.

[0166] [Other embodiments of the "Adapter A" of the present invention] Other embodiments of the "Adapter A" of the present invention include those in which i) further includes the following i)-4.

[0167] i)-4: Sequence for the transcription start site of the gRNA transcription promoter

[0168] i)-4: Regarding the transcription start site sequence of the gRNA transcription promoter, a promoter is necessary to express gRNA from the "gRNA production polynucleotide library" of the present invention prepared using the "adapter A" of the present invention, etc., but some promoters require a specific nucleotide as the transcription start site. On the other hand, the 5' end of each polynucleotide fragment in the "gRNA production polynucleotide library" after the "adapter A" has been cleaved does not necessarily contain a specific nucleotide that serves as the transcription start site.

[0169] Therefore, in order to ensure that each polynucleotide fragment with various 5' ends in the prepared "polynucleotide library for gRNA production" is reliably transcribed and gRNA is expressed, it is preferable to pre-insert a nucleotide sequence corresponding to a transcription start site near the 3' end of "adapter A" so that a transcription start site sequence is present at the 5' end of all "polynucleotides for gRNA production" even after cleaving "adapter A" with restriction enzymes.

[0170] Furthermore, it is necessary to insert the nucleotide sequence corresponding to this transcription start site downstream of the restriction enzyme recognition site of "i)-3" of "Adapter A" so that it remains at the 5' end of the polynucleotide for gRNA production even after "Adapter A" is removed.

[0171] Furthermore, "near the 3' end" does not necessarily mean strictly the 3' end; it means that there may be several sequences downstream of it. This is because gRNA expressed through transcription can bind sufficiently complementary to the target site in the target genome through an intrinsic spacer sequence of about 20 base pairs derived from the genome. Therefore, even if there are some sequences downstream of the transcription start site that do not consider complementarity with the target site, it is thought that the gRNA can still function.

[0172] However, in order to maintain the gRNA function while keeping the overall length of the "spacer sequence" within an appropriate range, it is preferable that the insertion site of the transcription start site not be too far from the 3' end (upstream). Furthermore, in order to maintain the specificity of the gRNA and eliminate off-target effects as much as possible, it is preferable that the transcription start site be located at the 3' end of "adapter A".

[0173] Specifically, for example, if the promoter used for gRNA proliferation is the U6 promoter, it is preferable to have a purine base as the transcription start site, so it is preferable to attach the "G" or "A" sequence downstream of the restriction enzyme recognition site in i)-3 and near the 3' end of "adapter A".

[0174] Furthermore, the "Adapter A" of the present invention described above exhibits the following excellent effects when used by attaching it to both ends of a polynucleotide fragment derived from a target genome.

[0175] 1) During PCR, the forward primer can also function as a reverse primer. (Only one type of primer needs to be prepared.) 2) "Spacer sequences for gRNA" can be excised from both ends of the polynucleotide fragment (doubling the variety of gRNAs). 3) Increases the number of gRNAs in the gRNA library. 4) Suitable for use with restriction enzymes such as EcoP15I, which have improved cleavage efficiency when two opposing recognition sites are present.

[0176] [Effects of the "Adapter A" of the present invention] Furthermore, because the "Adapter A" of the present invention has one end blocked by a stem-loop, self-ligation between "Adapter A" units is prevented or suppressed when preparing the "polynucleotide library for gRNA production" of the present invention, as described later, making it possible to bind "Adapter A" to both ends of a polynucleotide fragment more easily and reliably. The "Adapter A" of the present invention is useful for preparing the "polynucleotide library for gRNA production" of the present invention, as described later, but this library can also be used for genome editing such as single-nucleotide editing and prime editing, as well as epigenome editing.

[0177] [[The "Polynucleotide Library for gRNA Production" of the Present Invention]] [Aspect 1 of the Polynucleotide Library for gRNA Production of the Present Invention] The "Polynucleotide Library for gRNA Production" of Aspect 1 of the Present Invention is characterized in that the above-mentioned adapter (adapter A) of the Present Invention is bound to both ends of each "polynucleotide fragment" in a double-stranded polynucleotide fragment group.

[0178] The binding of the "polynucleotide fragment" and "adapter A" can be carried out by known methods, such as using a DNA ligase like NEB Quick ligase.

[0179] Furthermore, if both ends of each "polynucleotide fragment" are not blunt ends, it is preferable to blunt them beforehand using restriction enzymes such as SmaI or HaeIII to match the ends of "adapter A".

[0180] The terminal structures and sequences of "polynucleotide fragments" fragmented by fragmentation enzymes, etc., are inconsistent and unclear from fragment to fragment, making it impractical to process the ends of "adapter A" individually to match each fragment. Furthermore, since "adapter A" of the present invention is meticulously designed, including restriction enzyme recognition sites and promoter transcription start sites, processing "adapter A" for binding to polynucleotide fragments is undesirable in the first place.

[0181] Furthermore, the "adapter A" of the present invention, as described above, has one end blocked by a stem loop, which prevents self-ligation between "adapter A" units and allows for simpler and more reliable binding of "adapter A" to both ends of a "polynucleotide fragment." Therefore, the "polynucleotide library for gRNA production (Aspect 1)" of the present invention, which uses this as a material, has a significantly higher concentration of "polynucleotide fragments" with "adapter A" bound to both ends compared to when an adapter without a loop is used (see Figure 14).

[0182] Furthermore, since each "polynucleotide fragment" has the same "adapter A" attached to both ends, PCR has the advantage that the forward primer can also function as a reverse primer.

[0183] [Aspect 2 of the Polynucleotide Library for gRNA Production of the Present Invention] Another aspect of the "Polynucleotide Library for gRNA Production of the Present Invention" is obtained by cleaving each polynucleotide fragment obtained above, to which "Adapter A" is bound at both ends, at the loop sequence portions at both ends, and amplifying it by polymerase chain reaction (PCR) using a label-bound forward primer.

[0184] The method of "cutting the loop sequence" is not particularly limited, but for example, if a "sequence that can be enzymatically degraded" is incorporated within the "loop sequence" as described above, it can be cut with the enzymes described in the explanations of iii)-1 and iii)-2 above.

[0185] A "label" is some kind of marker, which serves as an indicator for sorting, concentrating, etc., a labeled substance.

[0186] Specifically, examples include the following, which are commonly used as gene markers.

[0187] Biotin

[0188] Fluorescent dyes (GFP (Green Fluorescent Protein), RFP (Red Fluorescent Protein), YFP (Yellow Fluorescent Protein), CFP (Cyan fluorescent protein), EGFP (Enhanced Green Fluorescent Protein), Alexa Fluor dye (Alexa Fluor is a brand name, manufactured by Thermo Fisher Scientific), Qdot label (Qdot is a brand name, manufactured by Thermo Fisher Scientific), etc.)

[0189] Enzymes (HRP: Horseradish peroxidase, AP: Alkaline phosphatase)

[0190] When concentrating and purifying DNA amplified by PCR, a substance with "specific and strong binding affinity" to the labeling and detection substance is preferred over "visibility." Therefore, among the above, biotin is a preferred choice.

[0191] Avidin and streptavidin, used in biotin purification, have a very strong affinity for biotin, and because each monomer forming the tetramer binds to one biotin molecule, the recovery efficiency of biotin-labeled DNA is very high.

[0192] The label may be attached to the 5' end, 3' end, or non-terminal portion of the primer, but attachment to the 5' end is preferred because it carries less risk of interfering with binding to the template.

[0193] Since each "polynucleotide fragment" in the "polynucleotide library for gRNA production (Aspect 1)" of the present invention has the same "adapter A" bound to both ends, the above-mentioned label-binding forward primer can also be used as a reverse primer in PCR.

[0194] Furthermore, since the loop sequence connects the double-stranded polynucleotide sequence of adapter A, it is bound to the end of that double-stranded sequence. Therefore, the forward primer corresponding sequence is positioned downstream of the loop sequence.

[0195] Since the above primers recognize the forward primer-corresponding sequence "i)-1 portion" downstream of the loop sequence portion, the "loop sequence" of iii) is completely removed from the amplicon (amplified polynucleotide) from the second cycle onward in PCR, etc., and each "polynucleotide fragment" in the "polynucleotide library for gRNA production" of this embodiment retains the "adapter A-derived sequence" including i) and ii), with the iii) at both ends removed.

[0196] Furthermore, it is preferable to improve the purity of the polynucleotide library of "Aspect 2 of the polynucleotide library for gRNA production" by appropriately incorporating steps such as "electrophoresis" and "filtration," which will be explained later in "Method for producing a polynucleotide library for gRNA production." These steps will remove polynucleotides that are extremely long or short for gRNA production.

[0197] [Aspect 3 of the Polynucleotide Library for gRNA Production of the Present Invention] Another aspect of the "Polynucleotide Library for gRNA Production" of the Present Invention is obtained in which each "Polynucleotide Fragment" in the "Polynucleotide Library for gRNA Production (Aspect 2)" obtained above is cleaved by the restriction enzyme i)-2.

[0198] Each of the above "polynucleotide fragments" is cleaved by the "i)-2 restriction enzyme" to a length that allows for the expression of a polynucleotide of appropriate length as a "spacer sequence" in gRNA after transcription.

[0199] Furthermore, as shown in "Aspect 1 of the polynucleotide library for gRNA production" above, this library uses "polynucleotide fragments" to which the "stem-loop type adapter (adapter A)" of the present invention is attached to both ends as starting material, resulting in a gRNA library with a wider variety of options than conventional libraries. Using this library, it may be possible to perform (epi)genome editing at more random locations.

[0200] This is because, by using "adapter A," which has a loop sequence, self-lagging and reverse binding of "adapter A" are prevented, and the loss of the number and variation of the "polynucleotide fragments with adapters attached to both ends" is already suppressed at the "library configuration 1." Furthermore, in "library configuration 3," the "spacer sequences" of the gRNA can be excised from "both ends" of the "polynucleotide fragment," and theoretically, twice the number and variety of gRNAs can be produced compared to when "adapter A" is attached to only one end. This is a significant advantage from the perspective of producing gRNAs for comprehensive (epi)genome editing of the target gene.

[0201] Furthermore, in the "Adapter A" portion, the loop portion is removed at the stage of "Library Embodiment 2" described above, resulting in the "Adapter A-derived sequence." However, even at this "Library Embodiment 3" stage, each fragment cut from both sides of the "raw material polynucleotide fragment" still retains the "Adapter A-derived sequence" at its respective 5' end.

[0202] In other words, from a single "polynucleotide fragment," two types of polynucleotides for gRNA production can be obtained by cleaving with restriction enzyme i)-2, with the "adapter A-derived sequence" remaining at the 5' end.

[0203] Furthermore, these two types, through transcription, express "spacer sequences" that bind to the "sense strand" and "antisense strand" of the genome, which are the raw materials for the "polynucleotide fragments," respectively. This is because the "adapter A-derived sequence" bound to the 3' side (of the sense strand) of the "(double-stranded) polynucleotide fragment" functions to excavate the "spacer sequence" from the 5' side of the antisense strand by the restriction enzyme i)-2.

[0204] When genome editing mutates a specific base in one of the polynucleotide strands of a double-stranded DNA, the corresponding polynucleotide strand is automatically converted to a complementary base to the mutated base with a certain probability by the cell's mismatch repair mechanism. Therefore, genome editing can target either strand of the double-stranded DNA.

[0205] At this stage, the "adapter A-derived sequence" at the 5' end is retained because, even after the loop is removed, the "adapter A-derived sequence" still contains a sequence corresponding to the forward primer for further PCR.

[0206] Furthermore, the fact that the "adapter A-derived sequence" remains on one side offers the advantage that, during the ligation of "adapter B" (described later), "adapter B" can be bound only to the 3' end of the "polynucleotide fragment" as designed.

[0207] Next, in another aspect of the present invention, a "polynucleotide library for gRNA production" of the present invention will be described, in which an "adapter B" is further attached to the 3' side of the "spacer sequence".

[0208] [Aspect 4 of the Polynucleotide Library for gRNA Production of the Present Invention] Another aspect of the "Polynucleotide Library for gRNA Production of the Present Invention" is that an adapter having the structure B) below is further attached to the restriction enzyme cleavage surface of each polynucleotide fragment in the "Polynucleotide Library for gRNA Production" obtained above.

[0209] B) An adapter (adapter B) having the arrangement of iv) and v) below.

[0210] iv) Polynucleotide chains having the sequences iv)-1 to iv)-3 below

[0211] iv)-1: Scaffold sequence for Cas protein or its variant protein iv)-2: Recognition site sequence of restriction enzyme capable of cleaving a portion of adapter B from a polynucleotide fragment iv)-3: Reverse primer corresponding sequence

[0212] v) Polynucleotide chain having a sequence complementary to iv)

[0213] <Explanation of iv)> This iv) is a sequence represented, for example, sequence number 3 in Figure 2, sequence number 5 in Figure 3, sequence number 7 in Figure 4, etc., and each of the sequences that make up the sequence will be explained below.

[0214] iv)-1: Regarding scaffold sequences for Cas proteins or their variant proteins

[0215] iv)-1-a: Definition of Cas Proteins "Cas proteins" are a general term for enzymes that have DNA (or RNA) cleavage activity used when editing the genome in the CRISPR-Cas9 system. When used together with the "gRNA expression vector library" of the present invention, they can randomly cleave genomic DNA (or RNA). There are several types depending on the number of strands of double-stranded DNA (or RNA) that can be cleaved and differences in PAM sequences (described later), etc. Specifically, "Cas9", "Cas12", "Cas13", etc.

[0216] Examples of "Cas12" include "Cas12a (Cpf1)".

[0217] Examples of "Cas13" include "Cas13a (C2c2)".

[0218] iv)-1-b: Definition of mutant proteins For each of the above "Cas proteins", there are various "mutant proteins" that differ in the following respects, for example.

[0219] iv)-1-b1: "DNA cleavage activity" iv)-1-b2: The necessity and sequence of "PAM" as a cleavage marker, i.e., "PAM dependency" iv)-1-b3: "Presence or absence of an additional domain that binds to the Cas protein" iv)-1-b4: Multiple combinations of the above differences

[0220] Examples of "additional domains" in b3 include "domains with epigenome editing activity" (such as transcriptional activation domains).

[0221] Furthermore, the above-mentioned mutant proteins include those referred to as "improved Cas9" proteins, such as "Cas9-VPR," which will be discussed later.

[0222] iv)-1-b1 example: For example, for the wild-type "Cas9" which can cleave double strands of DNA, examples of "mutant proteins" with different DNA cleavage activity include "nCas9" which cleaves only one of the two strands of DNA, and "dCas9" which has lost all DNA cleavage ability (nuclease activity).

[0223] In the present invention, when using a separate "deaminase" or other enzyme for genome editing or epigenome editing, dCas9 may be used to prevent unintended mutations due to extraneous cleavage. However, it is preferable to use nCas9 because even a single cleavage improves the enzyme's access to the target site.

[0224] Furthermore, prime editing requires the use of nickases such as nCas9 or Cas12 (which cut single strands of DNA) to cut the genome and insert a new sequence.

[0225] iv)-1-b2 example: Among the "Cas proteins," for example, "Cas9" generally requires a marker-like sequence called a "PAM" sequence (5'-NGG-3') near the site to be edited in the target gene. However, in order to make the most of the "polynucleotide belonging to the polynucleotide group for gRNA production that can bind evenly throughout the entire length of the target gene" produced in the present invention and to enable more comprehensive (epi)genome editing, it is preferable to appropriately select or combine, as needed, the use of Cas9 mutants that depend on a different PAM sequence than Cas9, or Cas9 mutants that have low dependence on PAM.

[0226] Furthermore, the most commonly used "Cas9" in genetic engineering is SpCas9 (Cas9 derived from Streptococcus pyogenes), and the "PAM" of this enzyme is 5'-NGG-3'.

[0227] Examples of Cas9 variants that depend on a different PAM sequence than "Cas9" include the following:

[0228] <SpRY (SpCas9 mutant)> PAM sequence: 5'-NNN-3' (effectively non-PAM required)

[0229] <SpG (SpCas9 mutant)> PAM sequence: 5'-NGN-3'

[0230] <eNme2Cas9> PAM sequence: Almost not required ((near)PAM-less)

[0231] <SpCas9-NG> PAM sequence: 5'-NG-3'

[0232] <FnCas12a (Cpf1): Derived from Francisella novicida> PAM sequence: 5'-TTTV-3'

[0233] <SaCas9 (Cas9 derived from Staphylococcus aureus)> PAM sequence: 5'-NNGRRT-3'

[0234] In the above, "N" means that it can be any of "A", "T", "G", or "C". "R" means that it can be either "A" or "G". "V" means that it can be any of "A", "G", or "C" other than "T".

[0235] The Cas9 mutant used in this invention is preferably one that is not dependent on PAM or is not PAM-requiring. For example, among the above, "SpRY" is preferred because it is almost completely PAM-requiring.

[0236] Examples of iv)-1-b3 include those in which a transcriptional activation domain is bound to a Cas protein, as illustrated in example iv)-1-b4 below.

[0237] iv)-1-b4 (combination of b1 to b3): (b1 + b2) Examples of mutants in which both "DNA cleavage activity" and "PAM dependence" differ from the wild-type Cas protein include, for example, the dead-type non-PAM-requiring Cas9 (dSpRY).

[0238] (b1+b3) Examples of mutants that differ from the wild-type Cas protein in both "DNA cleavage activity" and "presence or absence of an additional domain that binds to the Cas protein" include transcription-activating enzymes such as "CRISPR-ON (CRISPRa, CRISPR activation) enzyme" which has a "transcription activation domain" attached to dCas9, and specifically, the following are examples.

[0239] "dCas9-VPR": A transcriptional activator enzyme in which the highly potent transcriptional activating domain VPR (a fusion of VP64, p65, and Rta) is attached to dCas9.

[0240] This "dCas9-VPR," when used in conjunction with the "gRNA expression vector library" of the present invention, can activate specific genes (or groups of genes).

[0241] (b1 + b2 + b3) Examples of mutants that differ from the wild-type Cas protein in terms of "DNA cleavage activity," "PAM dependence," and "presence or absence of an additional domain that binds to the Cas protein" include dSpRY-VPR.

[0242] Incidentally, Cas proteins and their variants have been discovered or developed one after another in recent years, and the Cas proteins and their variant proteins used in this invention are not limited to those exemplified above.

[0243] In the following, "Cas or its variant proteins" may be referred to as "Cas proteins, etc." or simply "Cas, etc."

[0244] iv)-1-c: Definition of scaffolding sequence In iv)-1, "(Scaffolding sequence for Cas protein or its variant protein)" means DNA that expresses "Scaffolding sequence (RNA) such as Cas" through transcription.

[0245] "Scaffold sequences (RNA) for Cass, etc." are also called "scaffold sequences" or "tracrRNA (trans-activating crRNA)," and are approximately 80-base RNA sequences within (broadly defined) gRNA that function as scaffold sequences for binding to Cass, etc. The specific sequence is determined by the type of Cas, etc., and many such sequences have been reported, but it is not necessarily limited to these publicly known sequences.

[0246] Various specific "scaffolding arrangements (such as Cas)" have been developed and publicly known ones can be used, but examples include the following.

[0247] In Figure 2 (Sequence IDs 3, 4), the partial sequence shown as tracrRNA is "Adapter B (Adap_B_v1)".

[0248] In Figure 3 (SEQ ID NOs. 5, 6), the partial sequence shown as tracrRNA is "Adapter B (Adap_B_v2)".

[0249] In Figure 4 (Sequence IDs 7, 8), the partial sequence shown as tracrRNA is "Adapter B (Adap_B_v3)".

[0250] When Cas is SpCas9 or its variant (SpRY, etc.), a tracrRNA like the one shown in Figure 2 (SEQ ID NOs: 3, 4) is preferred as the scaffold sequence.

[0251] Furthermore, if the other "Cas protein or its variant" is an epigenetic modifying enzyme that controls gene transcription activity, such as a DNA methyltransferase or DNA demethylase, it is preferable to appropriately select from tracrRNAs such as those shown in Figure 3 (SEQ ID NOs. 5, 6) or Figure 4 (SEQ ID NOs. 7, 8), which contain the following MS2 sequence within the tracrRNA sequence.

[0252] An MS2 sequence is a sequence that can recruit MBP (MS2 Stem Loop Binding Protein) tagged proteins to a target region of the target genome.

[0253] In other words, by using epigenetic modifying enzymes involved in regulating MBP-tagged gene activity in combination with gRNAs containing the MS2 sequence within the tracrRNA sequence, it becomes possible to more reliably induce "Cas proteins, etc." at target sites and perform genome editing and epigenetic editing.

[0254] iv)-2: Recognition site sequences of restriction enzymes capable of cleaving a portion of adapter B from polynucleotide fragments.

[0255] "iv)-2: Recognition site sequence of a restriction enzyme capable of cleaving a portion of adapter B from a polynucleotide fragment" refers to the restriction enzyme recognition site sequence of "adapter B" described below, which is downstream of the "scaffold sequence for the Cas protein or its mutant protein" in iv)-1 and is for removing the "reverse primer compatible sequence" in iv)-3.

[0256] While there are no particular limitations on the restriction enzyme used for iv)-2, it is preferable to use a type II restriction enzyme whose recognition site and cleavage site are not too far apart, in order to avoid making the design of "Adapter B" too redundant and complex.

[0257] Type II restriction enzymes include those described in the section "i)-3 Restriction Enzymes" above. However, if the enzyme is to be introduced into a vector after cleavage, it is preferable to use one with a protruding end at the cleavage site so that it can bind specifically and strongly, as in the case of i)-3. In particular, to efficiently construct the "gRNA expression vector library" of the present invention described later, it is even more preferable to use one with a recognition site of 6 bases or more and high sequence specificity.

[0258] Furthermore, in order to reliably introduce the "gRNA production polynucleotide" into the vector in the intended direction (5'→3' direction), it is preferable to select a restriction enzyme for "iv)-2" that has a different recognition site sequence from the restriction enzyme for "i)-3".

[0259] For example, if BamHI, as shown in Figure 1, is selected as the restriction enzyme for "i)-3", it is preferable to select another type II restriction enzyme (with protruding ends), such as XhoI, as shown in Figures 2, 3, and 4, as the restriction enzyme for "iv)-2".

[0260] While it is possible to select the reversed enzymes for each of these, it is preferable to select them considering the order of restriction enzyme recognition sites in the multi-cloning site of the existing vector and the direction of introduction of the foreign gene (polynucleotide for gRNA production) to be introduced.

[0261] Furthermore, it is preferable that the recognition site of iv)-2 is located upstream of the reverse primer-corresponding sequence of iv)-3 so that the entire nucleotide sequence of the gRNA does not become unnecessarily long.

[0262] iv)-3: Regarding the sequences corresponding to reverse primers

[0263] "iv)-3: Reverse primer compatible sequence" refers to a sequence to which a "reverse primer," used to amplify polynucleotides for gRNA production by PCR, can bind. This sequence has sequences complementary to the "reverse primer" in order from the 3' end, and the "reverse primer" binds to the sense strand of adapter B during PCR.

[0264] Specifically, examples include, but are not limited to, the M13 reverse primer (M13_Rev primer) and other universal primer sequences.

[0265] The specific sequences corresponding to the M13_Rev primer are shown as "M13_Rev" in sequence number 3 in Figure 2, sequence number 5 in Figure 3, sequence number 7 in Figure 4, etc.

[0266] By pre-integrating such a reverse primer-compatible sequence into "Adapter B" used in the present invention, the amount of gRNA can be further increased in the present invention.

[0267] Furthermore, the above iv)-1, iv)-2, and iv)-3 are basically placed within iv) in this order.

[0268] Furthermore, in order to bind "Adapter B" to the downstream side (3' end of the "spacer sequence") of each "polynucleotide fragment" in the "gRNA production polynucleotide library" obtained in the above embodiment 3, it is preferable to position the 5' end of "Adapter B" to correspond to the cut surface (3' end) when the "spacer sequence" is excised.

[0269] In other words, for example, if the 3' end of the "spacer array" is a smooth end, it is preferable that the 5' end of "adapter B" also be a smooth end, or if it is a protruding end, it is preferable that it be a protruding end that can form a complementary pair with it.

[0270] Furthermore, if the enzyme used to excavate the spacer sequence is a type III restriction enzyme that forms overhanging ends, the sequences of the 3' overhanging ends of the "spacer sequence" vary depending on the "raw material polynucleotide fragment." Therefore, it is preferable to prepare multiple types of "adapter B" each with an overhanging cross-section corresponding to all of these random ends. In other words, if the 3' end of the "spacer sequence" is, for example, a two-base overhanging end, it is preferable to use "adapter B" containing equimolar proportions of each adapter, with each adapter having 16 different (= 4 × 4) binding sequences attached to its 5' end.

[0271] Furthermore, it is preferable to add an array to the 3' end of "Adapter B" that prevents it from being connected to the 3' end of the "Spacer Array," so that "Adapter B" can be connected to the "Spacer Array" in the correct direction.

[0272] For example, if the 3' end of the antisense strand of the "spacer array" is a protruding end, by making the 3' end of the sense strand of "adapter B" a protruding end in addition to the 5' end of the sense strand of "adapter B", it is possible to prevent the 3' end of "adapter B" from accidentally binding to the 3' end of the "spacer array".

[0273] There are no particular restrictions on the specific sequence to be added to the 3' end of the sense strand of "Adapter B," but from the viewpoint of structural stability, CCCC and the like are preferred. There are no particular restrictions on the length either, but about 4 bases is preferred as a necessary and sufficient length.

[0274] <Explanation of v)> This v) is a polynucleotide chain having a sequence complementary to iv), represented, for example, by sequence number 4 in Figure 2, sequence number 6 in Figure 3, sequence number 8 in Figure 4, etc.

[0275] However, when introducing polynucleotides for gRNA production into the vector in a later step, the restriction enzyme cuts used do not require complementary strands for the protruding ends, so the lengths of iv) and v) do not need to be exactly the same.

[0276] [Aspect 5 of the Polynucleotide Library for gRNA Production of the Present Invention] Another aspect of the "Polynucleotide Library for gRNA Production" of the Present Invention is obtained by enriching the "Polynucleotide Library for gRNA Production" obtained above using a label bound to a forward primer as an indicator.

[0277] Methods for purification and concentration using labels as indicators can be conventional methods in genetic engineering, and are not particularly limited as long as they use labels for purification or concentration. Examples include affinity column chromatography using a substance that interacts with the label as a ligand, a method of mixing and stirring with ligand-bound beads and recovering with a magnet, or a method of extracting bands that interact with the ligand from multiple bands separated by gel electrophoresis, etc.

[0278] Examples of ligands include avidin and streptavidin when biotin is the label, but streptavidin is preferred due to its nonspecific binding and higher affinity for biotin.

[0279] Furthermore, if the label is a fluorescent dye or enzyme, antibodies against them can be used as ligands.

[0280] Furthermore, using beads is preferable to affinity column chromatography. This is because the bead method is simpler to implement, results in less sample DNA loss, and affinity column chromatography, which utilizes biotin-avidin binding, requires more careful elution conditions than usual.

[0281] Beads include metal beads and magnetic beads. Metal beads can be recovered using filters, while magnetic beads can be recovered using magnets.

[0282] Examples of magnetic beads include Thermo's dyna beads.

[0283] [Aspect 6 of the Polynucleotide Library for gRNA Production of the Present Invention] Another aspect of the "Polynucleotide Library for gRNA Production of the Present Invention" is obtained by amplifying the "Polynucleotide Library for gRNA Production" obtained above by a second PCR using two types of primers, vi)-1 and vi)-2, described below.

[0284] vi)-1: Forward primer corresponding to i)-1 vi)-2: Reverse primer corresponding to iv)-3

[0285] In this second PCR, it is not necessary to attach a label such as biotin to the primer vi)-1 described above, so the amplified polynucleotides no longer have any label attached.

[0286] This second PCR further amplifies and enriches the "gRNA production polynucleotide library," easily increasing the yield of the "gRNA expression vector library" to the amount required for E. coli vector production.

[0287] The reason for performing PCR in two separate steps will be explained later in the "Method for Preparing a Gene Library for gRNA Production" of this invention.

[0288] [7th embodiment of the polynucleotide library for gRNA production of the present invention] Another embodiment of the "polynucleotide library for gRNA production" of the present invention is obtained in which each "polynucleotide fragment" in the "polynucleotide library for gRNA production" obtained above is cleaved at the restriction enzyme recognition sites i)-3 and iv)-2 below.

[0289] i)-3: Recognition site sequence of restriction enzyme capable of cleaving "adapter A" from "polynucleotide fragment" iv)-2: Recognition site sequence of restriction enzyme capable of cleaving a portion of "adapter B" from "polynucleotide fragment"

[0290] i) Restriction enzymes for recognition of -3 are restriction enzymes that can cleave and remove the aforementioned "adapter A-derived sequence" remaining at the 5' end of the polynucleotide, and specifically include the type II restriction enzymes described below.

[0291] Examples of type II restriction enzymes include BamHI, PstI, HindIII, EcoRI, XhoI, SalI, SmaI, and HaeIII, as mentioned above, with BamHI being preferred.

[0292] The restriction enzyme for iv)-2 recognition refers to a restriction enzyme that can cleave and remove a portion of the aforementioned "adapter B" (iv)-3, etc., that remains at the 3' end of the polynucleotide. Specifically, as with the restriction enzyme for i)-3 recognition, a type II restriction enzyme is preferred.

[0293] However, it is preferable to select a restriction enzyme for iv)-2 recognition that has a different recognition site sequence from the restriction enzyme for i)-3 recognition. For example, if BamHI is selected as the restriction enzyme for i)-3 recognition, it is preferable to select XhoI or the like as another type II restriction enzyme (with protruding ends) for iv)-2 recognition.

[0294] [Effects of the "gRNA Production Polynucleotide Library" of the Present Invention] gRNAs produced using the library of the present invention can be used for genome editing such as single-nucleotide substitutions evenly throughout the entire genome, and can also be used for epigenome editing. Furthermore, the gRNA polynucleotide (DNA) of the library of the present invention can also be further modified to include an Editing Template DNA sequence to be inserted into the genome, which can be attached to the 3' end to create "pegRNA DNA" for use in prime editing. Furthermore, the gRNA production polynucleotide library of the present invention may include a group of polynucleotide fragments to which "adapter B" is bound to a spacer sequence. In addition, these polynucleotide fragments may further include an i)-3 recognition restriction enzyme and an i)-1 forward primer corresponding sequence on the 5' side of the spacer sequence.

[0295] [[The gRNA Expression Vector Library of the Present Invention]] The "gRNA Expression Vector Library" of the present invention is characterized in that each polynucleotide for gRNA production in the library of the present invention described above is introduced into a vector capable of expressing an exogenous gene.

[0296] Examples of vectors include viral vectors and plasmid vectors, but plasmid vectors are preferred because they are easy to handle, easy to introduce into the host, and relatively inexpensive.

[0297] The introduction site into the vector can be any site commonly used in genetic engineering. If the vector is a plasmid vector, examples include multi-cloning sites (MCS) and polylinkers. However, multi-cloning sites are preferred because they offer a wider range of restriction enzyme options and facilitate the introduction of the polynucleotides for gRNA production according to the present invention in the correct orientation.

[0298] If the vector is a plasmid, it naturally contains other sequences necessary for it to function as a plasmid vector, such as the ori sequence and promoter sequence, in addition to the MCS sequence. Furthermore, in some cases, it may also contain drug resistance gene sequences or other transcription factor sequences.

[0299] In particular, it is preferable that the promoter includes the following promoters, which have high transcriptional activity for gRNA genes.

[0300] vii) Promoter for gRNA expression As the promoter for vii), a general promoter can be used, but a preferred one is, for example, "a promoter that has stable high activation ability regardless of cell type".

[0301] Specific examples of vii) are listed below, but are not limited to these: vii)-1: U6 promoter vii)-2: H1 promoter vii)-3: CMV promoter vii)-4: EF1α promoter vii)-5: T7 promoter vii)-6: CaMV35S promoter vii)-7: PC1SV promoter vii)-8: CAG promoter

[0302] Among the above, vii)-1, vii)-3, vii)-8, etc. are preferred, but vii)-1 is particularly preferred because, in addition to its high short-chain RNA expression efficiency, its promoter function is activated only when gRNA and Cas protein are co-expressed.

[0303] Furthermore, as described above in the explanation of "Primer A" of the present invention, when using a promoter that requires a transcription start site, it is preferable to pre-insert a nucleotide sequence corresponding to the transcription start site near the 3' end of "Adapter A" so that all of the polynucleotides for gRNA production belonging to the library that are incorporated into the plasmid can be reliably expressed.

[0304] Therefore, when using the U6 promoter of vi)-1, as described above, it is preferable to attach the "G" or "A" sequence downstream of the restriction enzyme recognition site of i)-3, near the 3' end of "adapter A".

[0305] [Effects of the "gRNA Expression Vector Library" of the present invention] The "gRNA Expression Vector Library" of the present invention makes it possible to easily express gRNAs corresponding to the entire genome.

[0306] [The "Genome Editing or Epigenome Editing Cultured Cells" of the present invention]

[0307] In the present invention, "genome-edited or epigenome-edited cultured cells" refers to cultured cells that have been subjected to the following: , <ii> , <iii>Cells into which such substances have been introduced, in most cases in which (epi)genome editing has already begun, and in some cases has been completed, but in a time-specific and / or site-specific manner. <ii>It may also be cells that have not yet begun editing, and are intended to express the protein introduced into vector <III>.

[0308] [Aspect 1 of the genome-edited or epigenome-edited cultured cells of the present invention] The genome-edited or epigenome-edited cultured cells of the present invention are as follows: and <ii>It is characterized by the introduction of [a specific feature / feature].

[0309] Vectors included in the gRNA expression vector library of the present invention described above <ii>A vector expressing the Cas protein or its mutant protein.

[0310] In this specification, cultured cells refer to cells that have been separated from an organism and undergo division and proliferation in a culture medium. Cultured cells can be those used for the proliferation of foreign cells in genetic engineering, and are not particularly limited, but examples include primary cells established from plants and animals, and tumor-derived cultured cells. Among these, cultured cells derived from humans, mice, rats, and flies are preferred because their genome sequences have been deciphered and the functions of each gene have been largely elucidated.

[0311] The Cas protein or its variant proteins are as described above.

[0312] [Aspect 2 of the genome-edited or epigenome-edited cultured cells of the present invention] Another aspect of the genome-edited or epigenome-edited cultured cells of the present invention is that, during the culture described above, the following further steps are taken: <iii>Examples include those that incorporate vectors.

[0313] <iii>A vector expressing an enzyme having genome editing activity or epigenome editing activity.

[0314] <iii>Examples of "enzymes with genome editing activity" in this context include the following: <iii>-1 and even <iii>-3, and also, as for "enzymes with epigenome editing activity", the following: <iii>Examples include -4, but this list is not limited to these.

[0315] <iii>-1: Base Editors (Single-nucleotide Mutation Inducing Enzymes) ABE (Adenine Base Editor / Adenine Deaminase): An enzyme that converts "A" to "G" CBE (Cytidine Base Editor / Cytidine Deaminase): An enzyme that converts "C" to "T" CGBE (C-to-G Base Editor): An enzyme that converts C to G ACBE (A-to-C Base Editor): An enzyme that converts A to C

[0316] Furthermore, single-nucleotide mutagenesis enzymes have been developed on a daily basis in recent years, and currently, it is possible to convert any combination of bases, and any of these can be used together with the polynucleotide library for gRNA production of the present invention.

[0317] <iii>-2: Reverse transcriptase M-MLV RT (Moloney Murine Leukemia Virus Reverse Transcriptase) "PrimeScript RTase" (manufactured by Takara Bio, a new RT developed based on M-MLV RT)

[0318] <iii>Reverse transcriptases like the one described in -2 are enzymes used to extend a new sequence using Editing Template DNA as a template, and are used when performing "prime editing" or the like with the gRNA vector library prepared in this invention.

[0319] <iii>-3: Restriction enzymes Acc65I: (Recognition site and cleavage site: G / GTACC) BsiWI: (Recognition site and cleavage site: C / GTACG) BsrGI: (Recognition site and cleavage site: G / GTACC) RsaI: (Recognition site and cleavage site: GT / CGAC) AciI: (Recognition site and cleavage site: C / CGC) HpaII: (Recognition site and cleavage site: C / CGG) NarI: (Recognition site and cleavage site: GG / CGCC)

[0320] <iii>Restriction enzymes like those listed in -3 can be used in the following cases because they can cut at two locations on the genome, in addition to the site where Cas9 normally cuts: • To make gene knockout more reliable • To insert an "external polynucleotide fragment" whose two ends each have sequences corresponding to different cleavage sites on the genome.

[0321] <iii>-4: Enzymes with epigenome editing activity. "Enzymes with epigenome editing activity" include enzymes that enhance or suppress the functions of the genome without altering the genome's base sequence, such as enzymes that regulate the transcriptional activity of the genome.

[0322] Specifically, examples include enzymes in which a "transcriptional activation domain" is attached to a Cas protein or its variant, such as dCas9-VPR and dSpRY-VPR, which are types of Cas9 variants mentioned above.

[0323] Furthermore, the above , <ii> , <iii>These may be all separate vectors, vectors introduced in any combination, or vectors introduced into the same vector.

[0324] , <ii> , <iii>Since each of these vectors has variations, using separate vectors increases the flexibility of their combinations, while in some cases it is preferable to use fewer types and numbers of vectors introduced into the cells.

[0325] Reasons why it is preferable to introduce fewer types of vectors include, for example, the following:

[0326] 1) Reduced preparation work and costs for vectors (they can also be purchased from the market) 2) Simplified vector delivery process 3) Reduced stress on the cells being delivered (prevention or suppression of excessive vector delivery) 4) Improved cell viability and delivery efficiency

[0327] Therefore, when preparing the cultured cells of the present invention, , <ii> , <iii>The combination of elements to be inserted into the same vector depends on the type of target genome, the function of the genome to be investigated, and the intended (epi)genome editing content. , <ii> , <iii>It is preferable to make an appropriate selection based on an overall assessment of the combination of enzymes contained within.

[0328] Examples of cases where it is preferable to use the same vector include, for example, <ii>and <iii>Examples include combinations of the above.

[0329] Because, This requires the creation of a different "gRNA library expression vector" for each genome used as raw material, <ii>or <iii>This is because, when the purpose is the same, such as genome editing or epigenome editing, it can be used universally, making it convenient to create a vector that includes both.

[0330] Furthermore, it is preferable to select and use cultured cells of the present invention based on markers such as drug resistance genes introduced into the vector, ensuring that each vector has been introduced into the cells.

[0331] Drug resistance markers such as puromycin (puro), plastocidine (bsd), hygromycin (hyg), and zeosin (sheble) can be used.

[0332] [Effects of the "Cultured Cells" of the Present Invention] By culturing the cultured cells of the present invention described above, "gRNA libraries," "Cas proteins, etc.," "enzymes with genome editing activity," "enzymes with epigenome editing activity," etc. are expressed within the cultured cells, and gene editing and gene regulation of target sites in the target genome are initiated. Therefore, a collection of various gRNAs ("gRNA library") spanning the entire length of the target genome is expressed within these cultured cells, and the ability to easily and reliably produce such a "gRNA library" is a major advantage of the "Cultured Cells" of the present invention.

[0333] [[Adapter set for gRNA library preparation according to the present invention]] The adapter set for gRNA library preparation according to the present invention is for preparing a gRNA library from a double-stranded polynucleotide fragment group and is characterized by including the above-described "adapter A" and "adapter B".

[0334] Furthermore, as described above, "Adapter A" and "Adapter B" included in the adapter set of the present invention are bound to the "polynucleotide fragment" at different times during the gRNA production process.

[0335] Specifically, "Adapter A" is attached to both ends of the "polynucleotide fragment" in the initial stage of creating gRNA from the "polynucleotide fragment." However, when the "gRNA spacer sequence" is excised from the "polynucleotide fragment," only one end (the 3' end) is removed, and the other end (the 5' end) is removed after serving as a forward primer in the second PCR just before being introduced into the "gRNA expression vector."

[0336] Adapter B excises the gRNA spacer sequence from the polynucleotide fragment and attaches it to the 3' end. The reverse primer-compatible sequence in Adapter B is removed by restriction enzyme treatment after the second PCR, but the scaffold sequence of Adapter B remains bound to the spacer sequence in the vector expressing the gRNA.

[0337] [Effects of the "Adapter Set" of the present invention] By using the "Adapter Set" of the present invention, the "polynucleotide fragment library for gRNA production" of the present invention, described later, can be produced simply and efficiently from regions throughout the entire genome.

[0338] [[The gRNA library preparation kit of the present invention]] The gRNA library preparation kit of the present invention is characterized by comprising the following <X> and <Y>.

[0339] <X> The "Adapter A" or "Adapter Set" of the present invention described above <Y> At least one of the enzymes listed below <Y>-1 to <Y>-10

[0340] <Y>-1: Enzyme for converting the target gene into a polynucleotide fragment <Y>-2: Enzyme for binding "Adapter A" to the polynucleotide fragment for gRNA production <Y>-3: Enzyme for binding "Adapter B" to the polynucleotide fragment for gRNA production <Y>-4: Enzyme for cleaving the single-stranded polynucleotide sequence for loop formation within "Adapter A" <Y>-5: Restriction enzyme capable of excising the "spacer sequence" of gRNA from the polynucleotide fragment for gRNA production (the 5' end) bound to the "sequence derived from Adapter A" <Y>-6: Restriction enzyme for cleaving and removing the "sequence derived from Adapter A" after loop removal <Y>-7: Restriction enzyme for cleaving and removing the reverse primer-compatible sequence in "Adapter B" <Y>-8: Heat-stable polymerase <Y>-9: Cas protein or its mutant protein <Y>-10: Enzyme having genome editing activity or epigenome editing activity

[0341] Examples of <Y>-1 include the aforementioned Fragmentase, and specifically, for example, "NEB fragmentase" manufactured by NEB Corporation can be used.

[0342] Examples of <Y>-2 include the aforementioned DNA ligase, and specifically, for example, "NEB Quick ligase" manufactured by NEB Corporation can be used.

[0343] Examples of <Y>-3 include DNA ligases, and among them, enzymes that can efficiently bind "adapter B" to the 3' end of "polynucleotide fragments" having various sequences are preferred, such as "Ligation Mix" manufactured by Takara Bio Inc.

[0344] Examples of <Y>-4 include enzymes that excavate uridine moieties, such as "USER(registered trademark) II Enzyme" (New England Biolabs, Inc.), as described in the section above under "Sequences that can be enzymatically degraded in loop sequences: iii)-1,-2", and restriction enzymes for single-stranded DNA, such as SI nuclease (manufactured by Takara Bio Inc.) and RsaI.

[0345] Examples of <Y>-5 include the aforementioned "i) restriction enzyme for recognition of -2 (e.g., type III restriction enzyme such as EcoP15I)." Furthermore, "adapter A-derived sequence" refers to the "partial sequence of adapter A" from which the loop sequence has been removed during the PCR process, as described above.

[0346] Examples of <Y>-6 include the above-mentioned "i) restriction enzymes for recognizing -3 (e.g., type II restriction enzymes such as BamHI)."

[0347] Examples of <Y>-7 include the aforementioned "restriction enzymes for iv)-2 recognition (e.g., type II restriction enzymes such as XhoI)."

[0348] <Y>-8 can be any heat-resistant polymerase that can be used in PCR. Examples include polymerases belonging to Family A (polymerase I (Pol I type)), such as Taq polymerase, and polymerases belonging to Family B (α type), such as KOD polymerase. If amplification speed is a priority, a Family A polymerase without 3' end proof-reading function can be used. However, in this invention, it is preferable to use a Family B polymerase that has proof-reading activity, as it allows for the amplification of the "spacer sequence" with a more accurate sequence.

[0349] <Y>-9 includes Cas9 and other Cas proteins that have genome-cutting function in (epi)genome editing using the CRISPR-Cas system, as well as the various Cas proteins mentioned above or their variants, and also includes scaffold proteins for other enzymes with (epi)genome editing activity, such as <Y>-10, specifically those mentioned above.

[0350] <Y>-10 is as described above. <iii>-1 and even <iii>Examples include various enzymes with single-nucleotide substitution function, reverse transcriptases, restriction enzymes, and enzymes with epigenome editing activity, as explained in section 4.

[0351] Furthermore, <Y>-9 and <Y>-10 may be forms that are already included in the vector.

[0352] [Effects of the "gRNA Library Preparation Kit" of the present invention] By using the "gRNA Library Preparation Kit" of the present invention, a gRNA library can be easily prepared, and cultured cells can be efficiently genome edited using it.

[0353] [[Method for preparing a polynucleotide library for gRNA production according to the present invention]] The "method for preparing a polynucleotide library for gRNA production" of the present invention is a method for preparing a polynucleotide library for gRNA production from a group of polynucleotide fragments, characterized by including the following steps (I) to (X).

[0354] Incidentally, the polynucleotide fragment group used as raw material can be either one that has been treated with Fragmentase or other methods, or one that has been artificially synthesized from scratch. However, if it is treated with Fragmentase, the length of polynucleotide fragments suitable for gRNA production can be produced by controlling the treatment time.

[0355] Preferred polynucleotide fragment lengths are 50 to 500 base pairs, more preferably 50 to 300 base pairs, and even more preferably 50 to 200 base pairs.

[0356] Furthermore, if the DNA fragment has been treated with Fragmentase, it is preferable that the DNA before treatment with Fragmentase is not contaminated with fragments longer than the original DNA size.

[0357] (I) A step of smoothing both ends of each polynucleotide fragment as necessary (II) A step of attaching the adapter A) below to both ends of the polynucleotide fragment (III) A step of cleaving at least one single-stranded polynucleotide sequence portion for loop formation in the adapter of (II) (IV) PCR using a forward primer to which the label has been attached (1 st (V) A step of amplifying polynucleotide fragments by performing PCR (2) (VI) A step of treating each amplified polynucleotide fragment with the restriction enzyme described in i)-2 below (VI) A step of attaching the adapter described in B) below to the end of the polynucleotide fragment treated in step (V) that has been cleaved by the restriction enzyme (usually the 3' end of the polynucleotide chain described in i) below (VII) A step of purifying and / or concentrating the polynucleotides using the label attached to the forward primer as an indicator (VIII) PCR (2) using the forward primer and reverse primer nd (IX) A step of amplifying polynucleotide fragments by PCR (IX) A step of restriction enzyme treatment at positions i)-3 and iv)-2 below (X) A step of inserting the polynucleotides after restriction enzyme treatment into a polynucleotide expression vector to prepare a gRNA expression vector

[0358] A) A stem-loop adapter (Adapter A) for gRNA library preparation having the sequences i) to iii) below.

[0359] i) Polynucleotide chains having the sequences i)-1 to i)-3 below

[0360] i)-1: Forward primer corresponding sequence i)-2: Recognition site sequence of restriction enzyme capable of cleaving the gRNA spacer sequence from the 5' end of the polynucleotide fragment bound to adapter A i)-3: Recognition site sequence of restriction enzyme capable of cleaving adapter A from the polynucleotide fragment

[0361] ii) A polynucleotide chain having a complementary sequence to i)

[0362] iii) A single-stranded polynucleotide sequence for loop formation, connecting the 5' end of i) and the 3' end of ii) in a double-stranded structure consisting of complementaryly linked i) and ii).

[0363] B) An adapter (Adapter B) for coupling downstream of a spacer array, having the arrangements described in iv) and v) below.

[0364] iv) A polynucleotide chain having the sequence iv)-1 to iv)-3 below.

[0365] iv)-1: Scaffold sequence for the Cas protein or its variant protein iv)-2: Recognition site sequence of a restriction enzyme capable of cleaving a portion of the adapter of B) from a polynucleotide fragment iv)-3: Reverse primer corresponding sequence

[0366] v) Polynucleotide chain having a sequence complementary to iv)

[0367] Furthermore, while each of the above steps is basically as described in the respective embodiments of the "polynucleotide library for gRNA production" of the present invention, the following supplementary information is provided.

[0368] Supplement to step (I): The phrase "as needed" in step (I) means that this step is necessary if "non-blunt end cuts are produced" during the cleavage of polynucleotides.

[0369] This step is unnecessary when fragmenting polynucleotides using only enzymes that cleave them smoothly, or when artificially creating polynucleotide fragments with blunt ends from the start. However, when using fragmentases that can efficiently fragment the genome, the cleavage surface is not always a "smooth surface," so it is preferable to include such a blunting step.

[0370] Supplementary information regarding step (VIII): The forward primer used in the second PCR step of step (VIII) can be the same sequence as the forward primer used in the first step, but it is not necessary to add any labels such as biotin.

[0371] Supplementary information on steps (IV) and (VIII): The reason for performing PCR in two separate steps, as in steps (IV) and (VIII), is as follows.

[0372] First reason: By using each primer, it is possible to reliably enrich DNA fragments that have the desired structure as described below.

[0373] Step (IV) First PCR: Polynucleotide fragments with sequences corresponding to Fwd (forward) primers at both ends. Step (VIII) Second PCR: Polynucleotide fragments with sequences corresponding to either Fwd primer or Rev (reverse) primer at each end.

[0374] The second reason is that the more polynucleotide fragments with Fwd primer-compatible sequences bound to both ends (i.e., EcoP15I recognition sequences also bound to both ends) are concentrated by the first PCR, the more accurately the specific cleavage effect of EcoP15I can be exerted in the subsequent enzymatic processing step (V). This increases the efficiency of producing polynucleotide fragments containing sequences of approximately 20 bp, which are suitable as spacer sequences for gRNA, and as a result, the concentration of polynucleotide fragments of the desired length in the library finally obtained by the "method for producing a polynucleotide fragment library for gRNA production" of the present invention increases.

[0375] The EcoP15I, currently the most widely used commercially, is chosen because, as mentioned above, its cutting efficiency increases when there are two opposing recognition areas.

[0376] The third reason is that by using a different Rev primer during the second PCR compared to the first, the 5' and 3' ends can be treated with different types of restriction enzymes (BamH1 and HindIII sites), improving the efficiency of insertion into the vector and also increasing the effectiveness of preventing the incorporation of unexpected sequences.

[0377] The fourth reason is that "Adapter B" used in step (VI) contains expensive synthetic oligonucleotides for scaffolding, so in order to reduce the amount used, it is preferable to concentrate the material obtained in the previous steps into a group that is likely to produce the gRNA of the desired base length before performing the second PCR.

[0378] Supplementary information regarding step (X): After performing step (X), it is preferable to further include steps such as introducing the vector into competent cells such as E. coli to grow the vector, or using drug resistance markers introduced into the vector, such as plasmids, to reliably select a vector in which the target gene has been reliably introduced. Introduction into competent cells and selection using drug resistance markers can be carried out by conventional methods.

[0379] Supplementary information on other additional steps: In the "method for preparing a polynucleotide library for gRNA production" of the present invention, for example, by providing the following steps, either individually or in combination, "(XI) electrophoresis step" or "(XII) filtration step," extremely long or short polynucleotides for gRNA production can be removed.

[0380] Step (XI) Electrophoresis Step Electrophoresis steps include gel electrophoresis. This method uses the network in the support gel to separate polynucleotide fragments by size and measure the molecular weight of each. By cutting out the band with the desired molecular weight, polynucleotide fragments of a suitable length for gRNA production can be obtained. Gels used as supports include agarose gel electrophoresis (AGE) and polyacrylamide gel electrophoresis (PAGE).

[0381] Process (XII) Filtration Process Filtration fixation can be performed using filtration processes such as gel filtration. Gel filtration chromatography, a type of column chromatography, is a method of separating samples based on differences in molecular weight by utilizing the pores in porous gel beads packed into a column. Polynucleotides with smaller molecular weights penetrate the pores and elute more slowly, while polynucleotides with larger molecular weights elute more quickly. The greater the difference in size, the clearer the separation can be.

[0382] Furthermore, there are no particular restrictions on the timing, order, number of times, or type of gel used for "Step (XI) Electrophoresis" and "Step (XII) Filtration," and these can be selected as appropriate. However, it is efficient to perform these steps at appropriate times, such as between Step (IV) and Step (V), between Step (VIII) and Step (IX), or between Step (IX) and Step (X).

[0383] Regarding the type of gel used in electrophoresis in step (XI), for electrophoresis performed between steps (IV) and (V), an inexpensive AGE gel that allows for rough separation is acceptable. However, for later stages, such as electrophoresis performed between steps (IX) and (X), a higher degree of purification (the proportion of polynucleotide fragments for gRNA production with spacer sequences of approximately 20 bp) is required, so PAGE is preferable.

[0384] Experiments by the inventors have confirmed that PAGE significantly improves the degree of purification compared to AGE (data is not provided).

[0385] Furthermore, the molecular weight of the band extracted after electrical transfer by PAGE naturally differs depending on the length of the adapter B (Adap_B) to which it is bonded. For example, if Adap_B is v1 as shown in Figure 2, the base length of the extracted band is 109 bp. If Adap_B is v2 in Figure 3 or v3 in Figure 4, the base length of the extracted band is 176 bp.

[0386] Thus, in the "method for preparing a polynucleotide library for gRNA production" of the present invention, the purity of the gRNA library having the desired length can be further increased by performing the above-mentioned "step (XI) electrophoresis" and "step (XII) filtration" multiple times at appropriate stages.

[0387] Furthermore, by following the "method for preparing a polynucleotide library for gRNA production" of the present invention, the following step (XIII) can be performed to prepare a "gRNA library".

[0388] (XIII) The step of expressing the gRNA expression vector obtained in (X) in suitable cultured cells or a non-human animal model.

[0389] Furthermore, this process (XIII) is as follows: <ii>or <iii>The process of introducing and expressing the vector may be carried out simultaneously.

[0390] <ii>Cas protein or its mutant protein expression vector <iii>A vector expressing an enzyme having genome editing activity or epigenome editing activity.

[0391] [Effects of the "Method for Preparing a Polynucleotide Library for gRNA Production" of the present invention] The "Method for Preparing a Polynucleotide Library for gRNA Production" of the present invention makes it possible to prepare a gRNA library from a very small amount of DNA fragment by using the "Adapter A" of the present invention and performing multiple PCR amplifications.

[0392] [[The "Method for Identifying Gene Mutations" of the Present Invention]] The "Method for Identifying Gene Mutations that Candidate Factors Causing Changes in the Phenotype to be Analyzed" of the Present Invention is characterized by including the following steps (XIV) to (XVI).

[0393] (XIV) A step of editing the genome of genes in cultured cells by culturing vector-introduced cultured cells of the present invention. (XV) A step of selecting cultured cells or non-human animal models with altered phenotypes to be analyzed from the cultured cells of (XIV) or non-human animal models transplanted with them. (XVI) A step of comparing the genes extracted from the cultured cells or non-human animal models selected in (XV) with the genes before genome editing.

[0394] Furthermore, the "selection" in (XV) can be carried out, for example, as follows:

[0395] (XV)-1: The cell population into which random mutations were introduced in (XIV), or the non-human animal model into which these mutations were transplanted, is cultured under conditions in which only cells with a specific cellular phenotype can proliferate or survive. (XV)-2: The mutation rate of each mutation site in the cell population before culture (control group if drug-treated) and the cell population after culture (treatment group if drug-treated) is analyzed using NGS (next-generation sequencing). (XV)-3: The mutation rate of each mutation site is calculated as the ratio of the count (number of times the target single nucleotide substitution at that mutation site was detected in NGS analysis) to the number of reads (number of times that mutation site itself was read in NGS analysis) at that site. (XV)-4: The change in mutation rate between the two groups is calculated as fold change (relative change in mutation rate between the two groups, FC value), and the count / read count for each mutation site in both groups is calculated as the P value and FDR (False Discovery Rate) using Fisher's exact test. (XV)-5: For fold change and FDR, mutation sites that meet a specific threshold are selected as candidates for functional single nucleotide variants.

[0396] Desired thresholds include, for example, fold change > 1, 5, and FDR < 0.05.

[0397] Furthermore, it is preferable to adjust the optimal threshold value as needed depending on the experimental conditions.

[0398] Examples of "phenotypes to be analyzed" include, but are not limited to, the following.

[0399] 1) Disease 2) Constitution and / or physical condition 3) Drug resistance

[0400] The "changes" in this phenotype include not only "negative changes" such as the onset of disease, worsening of symptoms, and other such changes, but also "positive changes" such as reduced susceptibility to disease, improvement in physical condition and constitution, and other such changes.

[0401] Furthermore, gene mutations that contribute to "resistance to disease" can be identified by adding, for example, the following step between steps (XIV) and (XV).

[0402] The process of placing cultured cells obtained in (XIV-XV)(XIV), or a non-human animal model transplanted with them, under conditions that induce disease symptoms.

[0403] The conditions that trigger disease symptoms can be set by following known methods such as "methods for altering the phenotype of cultured cells" or "methods for creating disease model animals."

[0404] Specifically, "diseases" include all diseases in which DNA mutations are involved in the disease phenotypic system, such as cancer, congenital metabolic disorders, and phenotypes showing resistance to therapeutic drugs due to variants. Cancer and metabolic disorders, in particular, are preferred because they have a large number of patients and therefore high demand. Conversely, rare diseases are also preferred because it is difficult to obtain genetic information from patients, making them valuable as an alternative means of research.

[0405] [Effects of the "Method for Identifying Gene Mutations" of the Present Invention] The "Method for Identifying Gene Mutations" of the Present Invention makes it possible to identify gene mutations that cause various diseases without the need for large quantities of biological samples from many patients.

[0406] [[The present invention's "method for determining the potential influence of a genome on phenotypes"]] [Aspect 1 of the determination method of the present invention] The present invention's "method for determining the potential influence of a genome on phenotypes" is characterized by including the following steps S1) to S4).

[0407] S1) A step of selecting gene mutations that are candidates for causing phenotypic changes to be analyzed using the method of the present invention described above. S2) A step of creating a list of candidate gene mutations enumerated and / or ranked based on the results of S1. S3) A step of comparing the information obtained from the genome with the list in S2. S4) A step of determining the likelihood of the phenotypic change to be analyzed.

[0408] Examples of "phenotypes" include those described in the section on "method for identifying gene mutations" of the present invention.

[0409] The term "genome" includes not only those extracted from living organisms, but also those artificially synthesized based on genomic information.

[0410] [Aspect 2 of the determination method of the present invention] Another aspect of the "method for determining the effect of a genome on the phenotype" of the present invention further includes the step S5 described below.

[0411] S5) A process of making an overall judgment by combining at least one of the following pieces of information with the results of S4: S5)-1: Metabolome analysis information S5)-2: Proteome analysis information

[0412] For metabolome and proteome analysis information, publicly available sources can be used, such as information obtained from KomicMarket (Kazusa Omics Data Market), MassBase, KNApSAcK, and Metabolonet.

[0413] [Effects of the "Method for Determining the Potential Influence of a Genome on the Phenotype" of the Present Invention] By using the "determination method" of the Present Invention, it is possible to create tools such as "lists" and "systems" for determining the potential influence of a genome on the phenotype by comprehensively judging genome-specific results (S4) or those results together with information obtained from various analytical data (S5).

[0414] In other words, by statistically analyzing the relationship between metabolome analysis information and proteome analysis information and its effects on various phenotypes such as constitution, physical condition, or health status, and by analyzing the relationship between these analysis results and genomic information, it is thought that a tool can be created to more accurately determine the "possibility of the genome's influence on phenotypes" from the judgment results of the above-described embodiment 1.

[0415] [[The "System for Determining the Potential Influence of a Genome on Phenotypes" of the Present Invention]] The "System for Determining the Potential Influence of a Genome on Phenotypes" of the Present Invention is characterized by having a program that executes each step of the "Determination Method" of the Present Invention described above.

[0416] Examples of system embodiments include a determination system comprising: a storage means for storing information on gene mutations that are candidates for causing phenotypic changes to be analyzed, obtained by the "method for identifying gene mutations" of the present invention described above; an information acquisition means for acquiring gene information of the genome contained in a sample; and a determination means for comparing the gene information acquired in the information acquisition step with the gene mutation information stored in the storage means, and determining the possibility of phenotypic changes to be analyzed based on the comparison result.

[0417] The information stored in the memory means is information on gene mutations corresponding to changes in the phenotype to be analyzed, and includes gene mutation information obtained by the gene mutation identification method of the present invention.

[0418] Here, there are no particular restrictions on the sample as long as it is a biological sample from which genetic information can be obtained, and examples include blood, bodily fluids, and various somatic cells, but these are not limited to these.

[0419] The above-described determination system can be configured to include, for example, an input means, an output means, and a calculation means in addition to a memory means, and to implement a program that sequentially executes the following: writing the genetic information of the sample input via the input means to the internal memory (information acquisition), reading the mutation information stored in the memory means, comparing the acquired genetic information of the sample with the mutation information, determining whether there are any matches in the comparison results, and outputting the determination result. This program can then be executed by the calculation device, thereby enabling the construction of a system that automatically determines the possibility of phenotypic change.

[0420] Figure 15 is a diagram showing the configuration of the system of the present invention.

[0421] System 1 of the present invention comprises an array data storage unit 3, an analysis data storage unit 5, a data input unit 7, a data matching unit 10, a data determination unit 11, and a determination result display unit 13.

[0422] The sequence data storage unit 3 can store, for example, a list (dataset) of mutation information, which is a list of gene mutation sequences created by the "method for determining the potential influence of a genome on the phenotype" of the present invention and / or a list of mutations ranked according to their likelihood of affecting the phenotype. This list can be arbitrarily constructed, and known sequence information can be added.

[0423] The analysis data storage unit 5 can store, for example, metabolome analysis information data and proteome analysis information data. For the analysis information data, publicly known data can be used, and information obtained from the database described above may be stored.

[0424] The data input unit 7 is an input means for inputting data into the system. Part or all of the sequence data (genetic information) of a subject analyzed using a next-generation sequencer or the like is input into the system via the data input unit 7. The input in the data input unit 7 may be configured to receive data via a network, and receiving means may be provided for this purpose.

[0425] The data matching unit 10 compares the list stored in the sequence data storage unit 3 with the sequence of the subject entered in the data input unit 7, and performs calculations to extract gene mutations that are candidates for factors causing changes in the phenotype to be analyzed.

[0426] The data determination unit 11 performs calculations to determine the likelihood of the subject's genome influencing the phenotypic changes being analyzed, based on a selection list that enumerates gene mutations in the subject's sequence data and information stored in the analysis data storage unit 5.

[0427] The display unit 13 displays the above determination result.

[0428] Based on the above configuration, this system performs calculations in the following order.

[0429] When some or all of the sequence data information of the subject is input from the data input unit 7,

[0430] The data matching unit 10 compares the subject sequence data with a list stored in the sequence data storage unit 3. Prior to this comparison, the data matching unit 10 compares the subject sequence data with a reference gene sequence and assigns base numbers to the subject sequence data. The comparison is performed in correspondence with the base numbers. Through the comparison of the two, gene mutations in the subject sequence data are extracted, a gene mutation list is created, and transmitted to the data determination unit 11.

[0431] The data determination unit 11, upon receiving the gene mutation list, performs analysis based on the data stored in the data storage unit 5, taking into account the results described above. At this time, the data to be analyzed can be appropriately selected. For example, metabolome analysis information can be used to determine gene mutations that affect metabolites in the body, while proteome analysis can be used to determine gene mutations that affect proteins in the body. Of course, it is also possible to combine these two methods for determination. The system determines the potential influence of the subject's genome on the phenotypic changes being analyzed.

[0432] The judgment result is displayed on the judgment result display unit 13.

[0433] [Effects of the "System" of the Present Invention] The "System" of the Present Invention makes it possible to efficiently, simply, and reliably determine the potential influence of the genome on the phenotype, even in situations where patient information is limited.

[0434] (Description of biomarkers) According to the present invention, the following protein mutations and their corresponding gene mutations have been identified as biomarkers useful for specific applications.

[0435] Biomarkers are biological indicators used for diagnosing diseases, determining treatment efficacy, predicting prognosis, estimating future onset risks, and predicting drug safety and efficacy for individuals. Specifically, they include test items, substances in vivo, gene mutations, protein expression levels, etc. For example, some genes can trigger the cancerization process. When these genes undergo mutations (DNA sequence changes or epigenetic modifications), they are abnormally activated or inactivated, causing cells to lose control and proliferate without limit. Such mutations that promote the tumor-forming ability of cancer are called driver mutations and can serve as indicators of cancerization.

[0436] The biomarker of the present invention is identified in the EGFR (epidermal growth factor receptor protein) gene, BRAF (B-raf proto-oncogene, serine / threonine kinase) gene, and KRAS (Kirsten rat sarcoma virus oncogene homologue) gene in which driver mutations are detected in lung cancer, and it is a biomarker for determining the hyperfunction of these genes.

[0437] The wild-type EGFR gene sequence and its corresponding amino acid sequence are described in SEQ ID NOs: 10 and 11 in the sequence listing, respectively. The wild-type KRAS gene sequence and its corresponding amino acid sequence are described in SEQ ID NOs: 12 and 13 in the sequence listing, respectively. The wild-type BRAF gene sequence and its corresponding amino acid sequence are described in SEQ ID NOs: 14 and 15 in the sequence listing, respectively.

[0438] In this specification, amino acid mutations are represented by the notation "p.X1 n X2". The "p" is a symbol indicating that this notation refers to a protein. When the start methionine is set at position 1 (the first position, n = 1) in the amino acid sequence of the protein, it means that the amino acid X1 (wild type) at the nth position has changed to amino acid X2. For example, p.I744V indicates that isoleucine (I) at amino acid position 744 has changed to valine (V).

[0439] Also, when an amino acid mutation is represented as "p.X1 n =", it means that although the amino acid X1 (wild type) itself at the n-th position has not changed, a so-called "silent mutation" has occurred, where the corresponding codon has changed. Furthermore, when an amino acid mutation is represented as "p.X1 n *", it indicates that the codon encoding the amino acid X1 (wild type) at the n-th position has been replaced with a stop codon (nonsense mutation).

[0440] In addition, among the following biomarkers, the protein includes not only the mature protein but also a precursor protein having a signal sequence or a domain having a specific function or a three-dimensional structural unit.

[0441] As a biomarker for determining the enhanced function of the EGFR gene, those containing an EGFR protein having any of the following amino acid mutations or a polynucleotide having an EGFR gene mutation capable of causing any of the following amino acid mutations were identified.

[0442] p.Q32R, p.G109E, p.S116F, p.P195S, p.A237V, p.S306=p., R324H, p.G331=, p.F376L, p.H433Q, p.K454E, p.W 477*, p.C555R, p.G614S, p.P631S, p.A647T, p.L655=, p.R675W, p.I744V, p.T751I, p.P794L, p.V802A, p.H80 Details of these markers, including 5R, p.C818R, p.P848=, p.K852R, p.K875R, p.A882V, p.Q894=, p.T940=, p.F968=, p.M987R, p.P992S, p.F997V, p.E1005=, p.D1014V, p.L1034I, p.C1058Y, p.D1084N, p.S1096=, p.R1100S, p.P1108=, p.A1181T, and p.V1200=, are shown in Figures 9, 10, and 16. Of these, p.Q32R, p.G109E, p.R324H, p.F376L, p.K454E, Ep.L655=, p.R675W, p.T751I, p.K852R, p.K875R, p.A882V, p.T940=, p.F968=, p.M987R, p.P992S, p.F997V, p.E1005=, p.D1014V, p.L1034I, p.D1084N, p.S1096=, p.R1100S, p.P1108=, and p.V1200= are preferred markers as they have an FC value of 4.0 or higher as shown in Figure 16.

[0443] To determine the hyperfunction of the KRAS gene, we identified KRAS proteins containing any of the following amino acid mutations, or polynucleotides containing KRAS gene mutations that can produce any of the following amino acid mutations, as biomarkers.

[0444] p.T2I, p.V7A, p.V7G, p.V8E, p.V9=, p.V9D, p.G10=, p.G10V, p.A11P, p.A11V, p.G12=, p.V14A, p.G15=, p.G15C, p.S17T, p.A18D, p.A18V, p.T20S, p.I21K, p.Q22L, p.Q22H, p.I24N, p.Q25*, p.H27Y, p.D30E, p.P34Q, p.P34S, p.P34T, p.T35=, p.I36L, p.S39=, p.S39Y, p.K42=, p.V44E, p.E49K, p.C51*, p.C51=, p.D57Y, p.G60=, p.E62D, p.Y64N, p.A66S, p.R68W, p.L79I, p.C80S, p.K88*, p.S89*, p.D92Y, p.I93F, p.R97I, p.E98*, p.E98=, p.K104N, p.P110S, p.V112I, p.V114I, p.G115R, p.D119N, p.P121H, p.S122F, p.R123*, p.R123I, p.D126V, p.Q129*, p.A134T, p.A134V, p.Y137=, p.G138R, p.P140S, p.R149K, p.R151T

[0445] Details of these markers are shown in FIGS. 9, 10, 17 to 19. Among these, p.T2I, p.V7G, p.V9=, p.V9D, p.G10V, p.A11P, p.A11V, p.G12=, p.V14A, p.G15C, p.A18D, p.Q22L, p.Q25*, p.T35=, p.I36L, p.G60=, p.R68W, p.L79I, p.C80S, p.K88*, p.S89*, p.D92Y, p.I93F, p.R97I, p.E98*, p.E98=, p.K104N, p.P110S, p.V112I, p.D119N, p.P121H, p.R123*, p.R123I, p.D126V, p.Q129*, p.A134T, p.A134V, p.Yi37=, p.G138R, p.P140S, p.R149K, p.R151T are those with an FC value of 4.0 or more shown in FIGS. 17 to 19 and are preferred markers.

[0446] It should be noted that there seems to be a typo in "p.Yi37=" in the translation of line , which should probably be "p.Y137=".Furthermore, p.V7A, p.V8E, p.V9D, p.A11P, p.A11V, p.G12=, p.V14A, p.G15C, p.A18V, p.Q22L, p.Q25*, p.G60=, p.C80S, and p.R149K have been detected in both HEK293 cells and KMST6 cells, and are therefore preferred markers from a different perspective than those mentioned above. Moreover, among these, p.A11P, p.Q22L, p.T35=, p.C80S, p.Y137=, and p.R151T showed an FC value of infinite (Inf), making them even more preferred markers.

[0447] To determine BRAF gene hyperfunction, we identified BRAF proteins containing any of the following amino acid mutations, or polynucleotides containing BRAF gene mutations that can produce any of the following amino acid mutations, as biomarkers.

[0448] p.W48*、p.L64=、p.S122F、p.K206E、p.L312P、p.T401=、p.Q494*、p.R509Q、p.R509L、p.L514P、p.L514I、p.Y519=、p.T521K、p.A526V、p.W531C、p.H540Q、p.L553R、p.A561S、p.L567=、p.A569S、p.I572F、p.H574N、p.L577I、p.N581I、p.N581S、p.L584I、p.L588R、p.L588F、p.F595I、p.F595L、p.G596C、p.G596V、p.G596S、p.A598=、p.V600L、p.V600E、p.S602Y、p.R603=、p.W604C、p.W604*、p.S605R、p.G606=、p.G606V、p.Q609=、p.E611K、p.Q612E、p.S614Y、p.M620I、p.V624=、p.R626I、p.D629Y、p.D629G、p.P632L、p.F635I、p.Q636H、p.Y647*、p.E648Q、p.M650V、p.Q653K、p.Q653*、p.S657*、p.I659=、p.R662S、p.D663N、p.D663A、p.F667V、p.R671L、p.R671*、p.L674=、p.R682W、p.P705Q、p.E715*、p.R719P、p.H725Q、p.S727I、p.S732=、p.N734S、p.R735Q、p.T740K、p.E741=、p.C748S、p.C748Y、p.I755=、p.G758=、p.G758E、p.G759E、p.G759R

[0449] Details of these markers are shown in Figures 20 and 21. Among them are p.W48*, p.L64=, p.S122F, p.K206E, p.L312P, p.T401=, p.R509L, p.L514I, p.T521K, p.A526V, p.W531C, p.H540Q, p.L553R, p.A561S, p.L567=, p.A569S, p.I572F, p.H574N, and p.L57 7I, p.N581I, p.L584I, p.L588R, p.L588F, p.F595I, p.F595L, p.G596C, p.G596V, p.A598=, p .V600L, p.V600E, p.S602Y, p.R603=, p.W604C, p.W604*, p.G606=, p.Q609=, p.E611K, p.Q612 E, p.S614Y, p.M620I, p.V624=, p.R626I, p.D629Y, p.D629G, p.P632L, p.F635I, p.Q636H, p. Y647*, p.E648Q, p.M650V, p.Q653K, p.S657*, p.I659=, p.R662S, p.D663N, p.D663A, p.F667V p.R671L, p.P705Q, p.E715*, p.R719P, p.H725Q, p.S727I, p.S732=, p.T740K, p.E741=, p.C748S, p.C748Y, p.I755=, p.G758=, and p.G758E are preferred markers as they have an FC value of 4.0 or higher, as shown in Figures 20 to 21.

[0450] Furthermore, p.T521K, p.H540Q, p.A561S, p.I572F, p.L588R, p.Y647*, p.E648Q, p.E715*, and p.H725Q exhibited an FC value of infinite (Inf), making them more preferable markers.

[0451] By detecting EGFR, KRAS, and BRAF proteins or genes with these mutations in a subject, it is possible to determine the hyperfunction of the corresponding gene. These biomarkers can be found in the subject's sample in the form of proteins or their motifs (polypeptides), DNA, RNA, or even polynucleotide fragments thereof that have the aforementioned mutations. Therefore, by using these polypeptides as antigens, antibodies that target and capture the corresponding biomarkers can be produced. It is also possible to form these polynucleotides into DNA probes and use them.

[0452] Furthermore, using one or more pieces of information (mutation information) from the biomarkers mentioned above, a dataset (list) can be constructed for use in determining the enhancement of the above-mentioned gene function. Specifically, such a list may consist of all the mutation information mentioned above, or it may consist of one or more arbitrary combinations. An arbitrary combination may be a set of the preferred biomarkers mentioned above.

[0453] Furthermore, such lists may be constructed by adding information on known biomarkers. Examples of known biomarkers include, but are not limited to, amino acid mutations in the EGFR protein such as p.G63R, p.V148M, p.T790M, and L858R, in the KRAS protein such as G12V, G13C, Q61H, L19F, and Q22K, and in the BRAF protein such as p.D594G and p.G596R. By using lists constructed by adding known biomarkers, it becomes possible to comprehensively capture the diverse factors contributing to the hyperfunction of genes that differ from one subject to another. Figures 16 to 21 show examples of such lists.

[0454] (Cancer screening method) This method involves comparing information on EGFR protein, KRAS protein, BLAF protein, or their genes based on a biological sample from a subject with information on mutations in EGFR protein, KRAS protein, BLAF protein and / or their genes to detect whether there are matching mutations. The mutation information includes those listed in Figures 16 to 21, and is compiled into a list and stored, for example, in a computer.

[0455] This testing method can also be performed using the system shown in Figure 15 above. In this case, the information in the list shown in Figures 16 to 21 is stored in the sequence data storage unit 3. Then, information on EGFR protein, KRAS protein, BLAF protein, or their genes based on a biological sample from the subject is input to the data input unit 7. This information is then compared with the information in the list shown in Figures 16 to 21 by the data matching unit 10, and the data determination unit 11 extracts information such as the risk of developing cancer corresponding to the matching mutation. The results are then displayed in the determination result display unit 13. Newly discovered information similar to that in Figures 16 to 21, as well as additional information, can be input and stored in the sequence data storage unit 3.

[0456] The present invention will be described in more detail below based on examples. Abbreviations used in the examples will be explained below.

[0457] Buffer solution: Buf. Enzyme: Enz. Triton X-100: TX-100

[0458] [Example 1: Stem-Loop Adapter of the Present Invention (Adap_A)] Using a standard method in genetic engineering, a stem-loop adapter of the present invention (Adapter A) was fabricated, which contains the complementary pairs of Sequence IDs 1 and 2, as described in Figure 1 and the above-mentioned explanation of Sequence IDs, and a single-stranded polynucleotide sequence for loop formation (Sequence ID 9) connecting them.

[0459] Furthermore, the arrangement of the loop portion in the "stem-loop type adapter" of this embodiment 1 is as follows.

[0460] 5'→3': GTCGTATCCAGTGC AGGGUCCGAGGTATTC GCACTGGATACGAC (Sequence No. 9)

[0461] Furthermore, in this loop-forming single-stranded polynucleotide sequence, the sequences from both ends up to the 14th position are such that they can form complementary pairs.

[0462] [Example 2: The "Adapter Set" of the Present Invention (Adap_A & Adap_B)] The "Adapter Set" of the Present Invention was created by combining "Adapter A" from Example 1 with "Adapter B" described below.

[0463] (Adapter B) Using a standard method in genetic engineering, adapter B was created, which has the complementary strand sequences of sequence numbers 3 and 4 as described in Figure 2 and the above-mentioned explanation of the sequence numbers.

[0464] [Example 3: "gRNA Library Preparation Kit" of the present invention (Adap_A & Adap_B & various enzymes)] The "gRNA Library Preparation Kit" of the present invention was prepared by combining the "adapter set" of Example 2 with the enzymes <Y>-1 to <Y>-8 below.

[0465] <Y>-1: NEB fragmentase (DNA fragmentase) This enzyme, manufactured by NEB, is used to break down target genes into polynucleotide fragments.

[0466] <Y>-2: NEB Quick ligase This enzyme is used to bind "Adapter A" in the "Adapter Set" to the polynucleotide fragment for gRNA production.

[0467] <Y>-3: Ligation Mix This enzyme is used to bind "Adapter B" in the "Adapter Set" to the polynucleotide fragment for gRNA production.

[0468] <Y>-4:USERII Enzyme This enzyme is used to cleave the single-stranded polynucleotide sequence for loop formation within "Adapter A".

[0469] <Y>-5: EcoP15I This enzyme is a restriction enzyme that can excise the spacer sequence of gRNA from a polynucleotide fragment that binds to "Adapter A", that is, it is a restriction enzyme used to cleave the 3'-side of any base from the 17th to the 23rd from the 5'-side, for example, of the polynucleotide fragment.

[0470] <Y>-6: BamHI This enzyme is a restriction enzyme for cleaving and removing "Adapter A".

[0471] <Y>-7: XhoI This enzyme is a restriction enzyme for cleaving and removing a part of "Adapter B".

[0472] <Y>-8: DNA Polymerase For example, it is a reagent kit for PCR in which a group of dNTPs, metal ions (MgCl2), buffer solution, etc., which are raw materials for DNA amplification, are combined with a heat-resistant polymerase.

[0473] In addition to the above, <Y>-9: Cas protein or its mutant protein <Y>-10: An enzyme having genome editing activity or epigenome editing activity may be combined.

[0474] [Example 4: "Polynucleotide library for gRNA preparation" of the present invention (polynucleotide fragment with Adap_A added to both ends)] Regarding the genes of the following six proteins, which are known to have gene mutations in lung cancer patients, the "polynucleotide library for gRNA preparation" of the present invention was prepared by the following procedures of Step 1-1 to 1-3 (see Figure 5) respectively.

[0475] EGFR (Epidermal Growth Factor Receptor Protein) gene, BRAF (B-raf proto-oncogene, serine / threonine kinase) gene, KRAS (Kirsten rat sarcoma virus oncogene homologue) gene, MDM2 (Mouse double minute protein 2) gene, MET (Mesenchymal epithelial transition protein) gene, CDK4 (Cyclin-dependent kinase-4) gene

[0476] Step 1-1: Fragmentation of DNA samples The prepared cDNAs of EGFR, BRAF, KRAS, MDM2, MET, and CDK4 were randomly cut to 100-200 bp using NEB fragmentase to produce polynucleotide fragments.

[0477] The target DNA used and the mixing ratio of each reagent are as shown in List 1 below.

[0478] [List 1] DNA sample (60 ng) 2.0 μl, 10× fragmentase Buf. 2.0 μl, NEB fragmentase 2.0 μl, ddH2O 14.0 μl, Total 20.0 μl

[0479] In the above, "10×" means that the reagent is intended to be diluted 10 times with another reagent before use.

[0480] Step 1-2: Blurring of DNA ends The protruding ends of each polynucleotide fragment produced in the previous step were blunted using a blunting enzyme.

[0481] The mixing ratios of the reagents used are as shown in List 2 below.

[0482] [List 2] Fragmented DNA 11.4 μl 10×Blunting Buf. 1.5 μl dNTP 1.5 μl Blunting Enz. 0.6 μl Total 15.0 μl

[0483] In the above, "10×" means that the reagent is intended to be diluted 10 times with another reagent before use.

[0484] Furthermore, blunting enzymes are enzymes that smooth the overhang portion of a single-stranded enzyme that remains after cleavage by restriction enzymes.

[0485] Step 1-3: Ligation of Adap_A Adapter A (a stem-loop DNA containing EcoP15I, BamHI, and M13_Fwd primer sequences) prepared in Example 1 was attached to both ends of each polynucleotide fragment that had been blunt-ended in Step 1-2.

[0486] The mixing ratios of the reagents used are as shown in List 3 below.

[0487] [List 3] Blunted DNA 15.0 μl Adap_A (0.6μg / μl) 1.0 μl 2×reaction Buf. 17.5 μl NEB Quick ligase 1.5 μl Total 35.0 μl

[0488] In the above, "2×" means that the reagent is intended to be diluted twice with another reagent before use. NEB Quick ligase is a DNA fragment ligation kit manufactured by NEB.

[0489] (Results) Because one end of "Adapter A" of the present invention is blocked by a stem loop, self-ligation between "Adapter A" units is prevented, allowing "Adapter A" to be attached to both ends of the polynucleotide fragment of each gene more reliably and efficiently.

[0490] [Example 5: "Polynucleotide Library for gRNA Production" of the present invention (Biotin-labeled Adap_A + polynucleotide fragment + Biotin-labeled Adap_A)] For each of the six genes mentioned above, a "polynucleotide library for gRNA production" of the present invention (Example 5) was prepared according to the procedure in Steps 1-4 below (see Figure 5).

[0491] Step 1-4: Decomposition of uridine in the hinge part and PCR Using USERII Enzyme, the uridine contained in the loop part of "Adapter A" in each polynucleotide fragment with "Adapter A" bound to both ends prepared in Example 4 of PCR was decomposed, and a DNA fragment having "Adapter A" at both ends was concentrated (amplified) by PCR using the biotinylated M13_Fwd primer.

[0492] The compounding ratios of the reagents used etc. are as shown in List 4 below.

[0493] [List 4] Adap_A-ligated DNA obtained in Example 4 2.0 μl KAPA HiFi 25.0 μl 100 μM M13-biotin primer 1.0 μl USERII Enzyme 1.0 μl ddH2O 21.0 μl Total 50.0 μl

[0494] Incidentally, the above-mentioned KAPA HiFi is a DNA polymerase-containing kit for PCR. Also, the above-mentioned USERII Enzyme is an enzyme that induces cleavage of the DNA loop by decomposing uridine in the stem loop.

[0495] Also, PCR was carried out according to the following procedure.

[0496] (1 st PCR procedure) PCR: 1 st -1) 30 min at 37°C PCR: 1 st -2) 3 min at 95°C PCR: 1 st -3) 20 sec at 98°C PCR: 1 st -4) 15 sec at 6°C PCR: 1 st -5) 15 sec at 72°C PCR: 1 st -6) 1 min at 72°C PCR: 1 st -7) ∞ at 4°C

[0497] Incidentally, PCR: 1 st -3) to PCR: 1 st -5) were repeated 25 - 30 cycles. Also, "∞" means maintaining the same temperature until the PCR apparatus is stopped

[0498] After PCR, the resulting polynucleotide library for gRNA preparation was purified using the following procedure.

[0499] (Electrophoresis Procedure) The following conditions were used for agarose gel electrophoresis, and bands of approximately several hundred bp were extracted: 1.5% AGE, 24 min, 100V

[0500] (Gel filtration procedure) The extracted 200-500 bp bands were purified using a gel extraction kit and eluted with 31.0 μl of ddH2O.

[0501] For the gel extraction kit, we used the "QIAquick Gel Extraction Kit" manufactured by QIAGEN.

[0502] Since each polynucleotide in the library prepared in Example 4 above has "adapter A" bound to both ends, the forward primer also acted as a reverse primer in the PCR in Example 5.

[0503] Furthermore, in the PCR performed after loop cleavage, a forward primer corresponding to the sequence downstream of the loop was used; therefore, in the gRNA polynucleotide library obtained in this example, the loop sequence was completely eliminated.

[0504] (Results) Example 5 yielded a polynucleotide fragment library for gRNA production with an even higher concentration than the library in Example 4.

[0505] [Example 6: "Polynucleotide Library for gRNA Production" of the Present Invention (Biotin-labeled Adap_A + Spacer Sequence (approx. 20 bp))] The libraries obtained in Example 5 for the six genes described above were processed according to the procedure in Step 2-1 below (see Figure 5) to produce the "Polynucleotide Library for gRNA Production" of the Present Invention (Example 6).

[0506] Step 2-1: EcoP15I treatment The DNA of adapter sequence A was treated with the restriction enzyme EcoP15I so that a DNA molecule of approximately 20 base pairs was attached to it.

[0507] The mixing ratios of the reagents used are as shown in List 5 below.

[0508] [List 5] Adap_A- DNA (100ng) obtained in Example 5: 11.0 μl, 10×NEB2.1: 1.5 μl, 10mM ATP: 1.5 μl, EcoP15I: 1.0 μl, Total: 15.0 μl

[0509] The above composition ratio was incubated (at 37°C for 40 minutes). Afterward, the EcoP15I reaction was stopped using the following: 0.6 μl of 0.5 M EDTA.

[0510] NEB2.1 is a restriction enzyme buffer solution manufactured by NEB.

[0511] Furthermore, in the above, "10×" means that the reagent is intended to be diluted 10 times with another reagent before use.

[0512] (Results) In the library of Example 5 described above, "adapter A" is bound to both ends of each polynucleotide fragment. Therefore, by the restriction enzyme treatment described above in Example 6, it is theoretically possible to produce two gRNA polynucleotides from a single polynucleotide fragment, further improving the variety of the gRNA library.

[0513] [Example 7: "Polynucleotide Library for gRNA Production" of the Present Invention (Biotin-labeled Adap_A + Spacer Sequence + Adap_B)] From the libraries of Example 6 relating to the six genes described above, the "Polynucleotide Library for gRNA Production" of the Present Invention (Example 7) was prepared by following the procedure in Steps 2-2 to 2-3 below (see Figure 5).

[0514] Step 2-2: Using an EcoP15I-cut DNA purification column, DNA of the expected fragment length was purified from the library of Example 6.

[0515] The mixing ratios of the reagents used are as shown in List 6 below.

[0516] [List 6] *Select-a-Size DNA Clean & Concentrator Kits were used. Select-a-Size DNA Binding Buffer 500 μl 95% EtOH 200 μl + EcoP15I-cut DNA obtained in Example 6 (before purification) 100 μl (scale up the 15 μl obtained in Step 2-1 with Elution Buff. 85 μl) Total 800 μl

[0517] Furthermore, "Select-a-Size DNA Clean & Concentrator Kits" are purification kits manufactured by ZYMO Research that allow for the rapid and easy purification of DNA fragments within a specific chain length range from restriction enzyme-treated samples.

[0518] Step 2-3: Adap_B Ligation The "Adapter B" shown in Figure 2 (a DNA fragment containing the gRNA scaffold sequence and the M13_Rev primer sequence) was ligated to the above DNA.

[0519] The mixing ratios of the reagents used are as shown in List 7 below.

[0520] [List 7] EcoP15I cut-DNA (after purification) 11.0 μl Adap_B (0.6 μg) 1.0 μl Ligation Mix 12.0 μl Total 24.0 μl

[0521] The above composition was incubated (at 16°C for 1 hour (or at 16°C Overnight)).

[0522] The "or" in the above statement indicates that the processing time differed for each of the six genes, but it is believed that this difference in processing time had little effect on the results.

[0523] [Example 8: "Polynucleotide Library for gRNA Production" of the Present Invention (Purified Library from Example 7)] From the libraries obtained in Example 7 for the six genes described above, the "Polynucleotide Library for gRNA Production" of the Present Invention (Example 8) was prepared by following the procedure in Step 2-4 below (see Figure 5).

[0524] Step 2-4: Purification of Adap_B ligated DNA The DNA fragment from Example 7, to which Biotin-labeled "Adapter A" was attached, was purified using streptavidin beads.

[0525] The mixing ratios of the reagents used are as shown in List 8 below.

[0526] [List 8] Adap_B ligated DNA obtained in Example 7: 24.0 μl, S.Avidin Mg Beads: 10.0 μl, TNET Buf.: 366.0 μl, Total: 400.0 μl

[0527] The solution of the above composition was nutated by inversion back and forth for 30 minutes at room temperature, then washed twice with 0.8 ml of TNET Buff to recover the purified biotin-bound polynucleotide (those that did not bind to the avidin beads were removed).

[0528] Furthermore, the TNET Buf. mentioned above refers to a buffer solution having the following composition.

[0529] 10mM Tris-HCl (pH 8.0) 1mM EDTA 0.5M NaCl 0.1% TX-100

[0530] [Example 9: "Polynucleotide Library for gRNA Production" of the present invention (PCR amplification of Adap_A + spacer sequence + Adap_B)] From the libraries of Example 8 relating to the six genes described above, the "Polynucleotide Library for gRNA Production" of the present invention (Example 9) was prepared according to the procedure in Step 3-1 below (see Figure 5).

[0531] Step 3-1: PCR amplification of the insert. The "polynucleotide library for gRNA production" obtained in Example 8 was further amplified by a second PCR using M13_Fwd and M13_Rev primers.

[0532] The mixing ratios of the reagents used are as shown in List 9 below.

[0533] [List 9] Adap_B-ligated DNA obtained in Example 8: 10.0 μl, KAPA HiFi: 25.0 μl, 100 μM InsAmp_F: 0.1 μl, 100 μM InsAmp_R: 0.1 μl, ddH2O: 14.8 μl. Total: 50.0 μl × 2

[0534] Furthermore, KAPA HiFi, as mentioned above, is a PCR-compatible DNA polymerase kit. The "×2" in "Total" indicates that two sets of the sample and reagents with the above composition ratio were used for PCR.

[0535] (2 nd PCR procedure) PCR: 2 nd -1) 3 min at 95℃ PCR: 2 nd -2) 20 sec at 98℃ PCR: 2 nd -3) 15 sec at 62℃ PCR: 2 nd -4) 15 sec at 72℃ PCR: 2 nd -5) 1 min at 72℃ PCR: 2 nd -6)∞ at 4℃

[0536] Furthermore, PCR: 2 nd -2) ~ PCR: 2 nd Step (-4) was repeated 14 times. "∞" means that the same temperature is maintained until the PCR machine is stopped, as mentioned above.

[0537] After PCR, the resulting "polynucleotide library for gRNA production" was purified using the following procedure.

[0538] (Gel filtration procedure) A total of 100 μl of reaction was purified using the QIAGEN "QIAquick PCR Purification Kit" and eluted with 32.0 μl of ddH2O.

[0539] This "gel filtration" (purification) allowed us to reduce the volume of reaction solution during the restriction enzyme treatment in the next step.

[0540] [Example 10: "Polynucleotide Library for gRNA Production" of the present invention (Spacer sequence + scaffold sequence (excision of Rev primer portion of Adap_A and Adap_B)] From the library of Example 9 relating to the six genes described above, the "Polynucleotide Library for gRNA Production" of the present invention (Example 10) was prepared according to the procedure in Step 3-2 below (see Figure 5).

[0541] Step 3-2: Restriction enzyme treatment of the inserts: The M13_Rev sequence region in the "Adapter A" sequence and the "Adapter B" sequence was separated from the gRNA production sequences (spacer sequence + scaffold sequence) in the library of Example 9 by BamHI (for Adap_A) and XhoI (for Adap_B).

[0542] The mixing ratios of the reagents used are as shown in List 10 below.

[0543] [List 10] gRNA Amplicon DNA obtained in Example 9: 30.0 μl, 10× CutSmart: 3.6 μl, BamHI: 1.2 μl, XhoI: 1.2 μl, Total: 36.0 μl

[0544] The above composition ratio was used for incubation (at 37°C for 90 minutes to Overnight).

[0545] In the above, "CutSmart" refers to a restriction enzyme buffer solution manufactured by NEB that can be used with any type of restriction enzyme.

[0546] (Electrophoresis procedure) Place 36 μl of enzyme-treated solution into a Thermo Novex electrophoresis machine. TM Electrophoresis (PAGE) was performed for 40 minutes using "TBE Gels, 10%", and a 109 bp band was extracted.

[0547] (Gel filtration procedure) Purification was performed using the "QIAquick PCR Purification Kit" mentioned above.

[0548] (Results) Through the above purification process, an average of 300 ng, and at least approximately 30 ng, of gRNA amplicon DNA was obtained from the six target genes.

[0549] (Discussion) This yield is sufficient for the (epi)genome editing operations described later, and considering that the 60 ng of DNA sample used in Example 4 was processed through numerous steps, it represents an astonishing amplification rate.

[0550] [Example 11: The "gRNA expression vector library" of the present invention (polynucleotide + vector from Example 10)] Using the library obtained in Example 10 for the six target genes described above, the "gRNA expression vector library" of the present invention (Example 11) was prepared by following the procedure in Steps 3-3 to 3-4 below (see Figure 5).

[0551] Step 3-3: Insertion into hU6 vector The sequences for each gRNA in the polynucleotide library for gRNA production prepared in Example 10 were inserted into the vector described later.

[0552] The mixing ratios of the reagents used are as shown in List 11 below.

[0553] [List 11] DNA Amplicon (30 ng) obtained in Example 10: 4.0 μl, Cut U6 Vector (100 ng): 2.0 μl, Ligation Mix: 6.0 μl, Total: 12.0 μl

[0554] Each sample was mixed according to the above-mentioned composition ratio, incubated (at 16°C for 1 hour, or at 16°C overnight), and the gRNA sequence was inserted into the pEXA2J2-pU6MCS vector.

[0555] As mentioned above, "or" means that the processing time differed depending on the target gene.

[0556] Furthermore, the "Cut U6 vector" in the above composition is an expression vector containing a U6 promoter, obtained by treating the multi-cloning site of the pEXA2J2-U6MCS vector with restriction enzymes XhoI and BamHI.

[0557] Step 3-4: E. coli Transformation The gRNA expression plasmid vector obtained in Step 3-3 above was transformed into competent cells (E. coli) according to a standard method to prepare a plasmid library.

[0558] Figure 6 shows a schematic plasmid map of the prepared vector.

[0559] The sequences included in the "guide RNA expression vector library" were comprehensively analyzed using a next-generation sequencer. The base lengths of the relevant parts of the spacer region were calculated, and the results of plotting their proportions within the whole are shown in Figure 7.

[0560] (Results) As can be seen from Figure 7, the length of the "spacer sequences" in the prepared "gRNA expression vector library" ranged from 17 to 22 base pairs. Theoretically, when using the adapter of Example 1, all "spacer sequences" should be exactly 20 base pairs long. This range may be due to errors in the reaction conditions, such as variations in the Stra activity (cleavage precision) of the restriction enzyme, or the fact that the fragmented raw material polynucleotide fragments also contained the same sequences as the recognition sequences of i)-2 and i)-3.

[0561] (Discussion) The above results demonstrate that when using the "adapter set" of the present invention, most of the "spacer sequences" of the produced gRNA can be produced to the appropriate length.

[0562] [Example 12: Genome-edited cultured cells of the present invention (gRNA vector + nCas9 vector + genome editing vector)] Using the vector library obtained in Example 11 for the six target genes described above, "genome-edited cultured cells" of the present invention (Example 12) were prepared according to the following procedure.

[0563] Plasmid libraries consisting of the plasmid vectors described in P) and Q) below were prepared, and these were transfected into HEK293 cells (a cell line derived from human embryonic kidney cells) to prepare "genome-edited cultured cells."

[0564] P) "(I) Plasmid vector for gRNA expression": The "gRNA expression vector library" prepared in Example 11.

[0565] Q) Two types of vectors that serve both as "(II) Cas9 or its mutant protein expression vectors" and "(III) vectors expressing enzymes with genome editing activity":

[0566] Q-1: We purchased the following plasmid from Addgene, which is a type of Cas9 gene expression plasmid (SpRY) that does not require a PAM sequence, and then introduced a drug resistance gene (puromycin resistance gene) into it for use.

[0567] "pCMV-T7-ABEmax(7.10)-SpRY-P2A-EGFP (RTW5025)" (Plasmid #140003)

[0568] This plasmid vector contains <ii>nCas9 <iii>As such, it contains all the enzymes of the Adenine Base Editor (ABE) for the conversion of A (adenine) to G (guanine).

[0569] Q-2: SpRY_CBE_Puro Similar to Q-1, the following plasmid, a type of Cas9 gene expression plasmid (SpRY) that does not require a PAM sequence, was purchased from Addgene and used after introducing a drug resistance gene (puromycin resistance gene).

[0570] "pCAG-CBE4max-SpRY-P2A-EGFP (RTW5133)" (Plasmid #139999)

[0571] This plasmid vector contains, <ii>nCas9 <iii>These enzymes include single-nucleotide convertases (CBEs) for the conversion of C (cytosine) to U (uracil) {or T (thymine)}.

[0572] In this Example 12, the gRNA library of the present invention was expressed in the cultured cells into which various vectors had been introduced, and genome editing involving various single-nucleotide substitutions occurred simultaneously by the expressed nCas9 and two types of deaminase enzymes.

[0573] For example, Figure 8 shows the results of investigating the single nucleotide substitution sites and mutation rate of the KRAS gene using a next-generation sequencer after genome editing, as a representative example.

[0574] The survey methodology is as follows:

[0575] The DNA sequence of the protein-coding exon region of KRAS was cloned by PCR and converted into a gRNA library expression vector using the technology of the present invention. This library vector, along with SpRY_CBE and SpRY_ABE expression vectors, was introduced into normal cultured cells. Subsequently, mRNA was recovered from the cells, and the cloned KRAS gene sequence was analyzed using a next-generation sequencer to calculate the introduced mutation sites and introduction rates.

[0576] (Results) As can be seen from Figure 8, genome editing using the gRNA of the present invention successfully introduced mutations evenly throughout the entire exon region of the KRAS gene, which consists of a total of approximately 564 base pairs.

[0577] (Discussion) The above results demonstrate that comprehensive genome editing can be easily performed using the gRNA expression vector of the present invention.

[0578] [Example 13: Method for identifying gene mutations that cause changes in the phenotype to be analyzed according to the present invention] Using cells whose genomes were edited in the same manner as in Example 12 for the six target genes mentioned above, the method for identifying gene mutations that cause changes in the phenotype to be analyzed according to the present invention (Example 13) was carried out according to the following Steps 4-1 to 4-2.

[0579] Step 4-1: Extraction of Phenotypic Changed Cells For each of the six target genes (EGFR, BRAF, KRAS, MDM2, MET, CDK4) that underwent genome editing in the above example, cells showing phenotypic changes (cancerous changes) were extracted from cultured cells in which each gene had been genome edited.

[0580] For extraction, a method was used to select cells that had proliferated even when cultured in a low-binding dish.

[0581] Step 4-2: Confirmation of gene (and amino acid) sequences in extracted cells. Mutations in each target gene were analyzed in the cancerous cell population obtained in 4-1. While DNA mutations that do not change amino acids (synonymous mutations) may be associated with cancer, in most cases amino acid mutations are present. Therefore, here we mainly focused on single nucleotide mutations corresponding to "amino acid mutations" that have a track record of being investigated for their association with cancer in human samples, from a comprehensive single nucleotide mutation range across the entire genome. We also confirmed whether other single nucleotide mutations that may be associated with cancer could be similarly derived from the genome editing results according to the present invention.

[0582] Specifically, the target of the review was human cancer samples that have been detected and are listed in databases such as the COSMIC database.

[0583] Figures 9 and 10 show the results of investigating single nucleotide substitution sites using next-generation sequencing for representative examples (KRAS and EGFR) of the six target genes mentioned above, along with the corresponding amino acids.

[0584] Furthermore, as shown in Figure 10, single nucleotide variant sites detected in cancer tissue and thought to be associated with carcinogenesis may not have corresponding amino acid variants (e.g., "G331=").

[0585] The procedure for comparing cultures before and after using a "low-adhesion dish," and the reason why this comparison can confirm the presence of cancer-related gene mutations, are described in detail below.

[0586] The dots in Figure 9 indicate the locations of mutations in the corresponding amino acids of the target genes (KRAS and EGFR) detected in the patient's cancer tissue.

[0587] Similar to the experiment in Example 12 (Figure 8) described above, random mutations were introduced into the KRAS and EGFR gene regions of the oncogenes in HEK293, which are normal human cells, using the method of the present invention.

[0588] Normally, cells cannot proliferate unless they adhere to a scaffold structure, whereas cancer cells possess "scaffold-independent cell proliferation activity," which allows them to proliferate without relying on a scaffold. Therefore, while normal HEK293 cells cannot proliferate without a scaffold structure, when HEK293 cells with random mutations introduced into the KRAS and EGFR gene regions are cultured under scaffold-independent conditions, proliferating cells appear. This is thought to be because the introduction of random mutations leads to gene mutations that contribute to cancer development, causing the cells to become cancerous.

[0589] In other words, Figures 9 and 10 show the results of analyzing the amino acids and gene mutation sites contained in these "cells that are thought to have proliferated due to cancer."

[0590] In other words, each dot, X-axis, and Y-axis in Figure 9 represents the following:

[0591] Each dot: DNA mutation at a specific site on the protein where the target gene is expressed. The amino acid changes induced by each mutation are also shown. X-axis: Amino acid mutation rate (%) for each specific site before low-adhesion dish culture. Y-axis: Amino acid mutation rate (%) for each specific site after low-adhesion dish culture.

[0592] Furthermore, "amino acid mutation rate" refers to the proportion of cells in the entire cell population that have undergone random mutations that possess that particular mutation site.

[0593] If the amino acid mutation rate at a specific site remains unchanged before and after low-adhesion dish culture (X=Y), or if the mutation rate after culture is lower (Y<X), then these amino acid mutations are unlikely to be associated with cancer. On the other hand, if the mutation rate after culture is higher (Y>X: to the upper left of the X=Y line in the figure), then these amino acid mutations are likely to be deeply associated with cancer.

[0594] Furthermore, the amino acid mutations, their corresponding single nucleotide mutations, and the Y / X value (for each mutation, the percentage change in mutation rate after culture compared to before culture in a low-adhesion dish: Enrichment score) are shown in accordance with Figure 10.

[0595] (Results) As shown in Figures 9 and 10 as representative examples, amino acid (and base) mutations that were significantly increased in the KRAS gene and EGFR gene in cancerous cells were identified.

[0596] Furthermore, these results showed a statistically significant correlation when compared with the results of gene mutation testing using genes derived from actual cancer patients (GILA (Growth in Low Attachment), from Jesse Boehm, Broad Institute, Boehum group) (not shown).

[0597] Although not shown in the diagram, similar results were obtained for the other four target genes (BRAF, MDM2, MET, CDK4).

[0598] (Discussion) The above results demonstrate that the analysis of artificial genome editing results using the "polynucleotide library for gRNA production" of the present invention is a practical method for rapidly and easily identifying gene mutations that are candidate factors in the alteration of various "phenotypes," including diseases such as cancer. Furthermore, the present invention demonstrates that the amount of "polynucleotide library for gRNA production" required for such identification experiments can be produced from a very small amount of DNA.

[0599] [Example 14: Method for determining the potential influence of a genome on phenotype according to the present invention] This can be determined by the determination system (Figure 11) described in Example 15, which will be described later.

[0600] [Example 15: System for Determining the Potential Influence of a Genome on Phenotypes according to the Present Invention] The "determination system" of the present invention, which has a program for executing the "determination method" of the present invention, is constructed, for example, as shown in Figure 11.

[0601] [Example 16: Method of using the "gRNA expression vector library" of the present invention (activation of target gene)] From genomic DNA (10 ng) of mouse cells, the sequence of the promoter region of the "Pparg (Peroxisome proliferator-activated receptor γ) gene," which is involved in liposuction, was cloned by PCR as the target gene. Using the technology of the present invention, a "polynucleotide library for gRNA production" (300 ng) and a "gRNA expression vector library" (100 μg) were prepared for this promoter region.

[0602] This "gRNA expression vector library" and the "dSpRY-VPR (a Cas9 enzyme with deficient DNA cleavage activity and linked to a transcription-activating enzyme) expression vector" were introduced into mouse fibroblast cells (3T3), and the subsequent activation of the Pparg gene was quantified by RT-qPCR.

[0603] Quantitative results were evaluated by the relative expression level of the Pparg gene, with the housekeeping gene "Hprt (Hypoxanthine-guanine phosphoribosyltransferase)" used as an internal control. The results are shown in Figure 12.

[0604] In Figure 12, "gRL" refers to a "gRNA expression vector library."

[0605] (Results) As can be seen from Figure 12, it was observed that the expression level of Pparg increased by approximately 250 times when both the "gRNA expression vector library" and the "dSpRY-VPR expression vector" were introduced.

[0606] (Discussion) It has been proven that the "gRNA expression vector library" of the present invention, which was prepared using the "adapter A" of the present invention from a very small amount of DNA, has the "quality" and "quantity" to be used not only for "genome editing" but also for "epigenome editing (genome function control (transcriptional activation))."

[0607] [Example 17: Method of using the "gRNA expression vector library" of the present invention (comprehensive disruption of target genes)] Similar to Example 16, the sequence of the "promoter region of the Pparg gene" involved in liposuction was cloned by PCR from genomic DNA (10 ng) of mouse cells as the target gene. Using the technology of the present invention, a "polynucleotide library for gRNA production" (300 ng) and a "gRNA expression vector library" (100 μg) were prepared for this promoter region.

[0608] This "gRNA expression vector library" and the "SpRY (SpCas9 mutant that does not require a PAM sequence with DNA cleavage activity) expression vector" were introduced into mouse fibroblast cells (3T3), and subsequent DNA sequence insertions and deletions in the promoter region of the Pparg gene were evaluated by PCR. The results are shown in Figure 13.

[0609] In Figure 13, "gRL" refers to a "gRNA expression vector library."

[0610] (Results) As can be seen from Figure 13, the introduction of both the "gRNA expression vector library" and the "SpRY expression vector" resulted in some of the sequences amplified by PCR being dispersed into sequences longer or shorter than the length designed by the primers (approximately 800 bp). In other words, it was observed that the number of target genes with their original function decreased as a result of insertions and deletions of target gene sequences.

[0611] (Discussion) It has been proven that the "gRNA expression vector library" of the present invention, which was prepared using "Adapter A" of the present invention from a very small amount of DNA, has sufficient "quality" and "quantity" to be used for comprehensive disruption of target genes, i.e., "genome editing". Furthermore, this comprehensive disruption is thought to be the result of "errors in the break repair function" caused by repeated DNA breaks, which led to multiple insertions of several bases into the target region or deletions of the break region (Indel: Insertion and Deletion).

[0612] [Test Example 1: Comparison of Polynucleotide Library Production Efficiency for gRNA Production Based on Differences in Adapter Structure] OligoDNA was prepared to produce two types of "Adapter A'" ("Adap_A_Δloop" and "Adap_A Δprimer"), which have different structures from the "Adapter A (Adap_A)" of the present invention, and the difference in "production efficiency" of the "polynucleotide library for gRNA production" was investigated.

[0613] (Comparative Example 1) "Adap_A Δloop" refers to "Adapter A'" of the present invention, which lacks "iii: single-stranded polynucleotide sequence for loop formation" (it is not a stem-loop type).

[0614] (Comparative Example 2) "Adap_A Δprimer" means "Adapter A'" of the "Adapter A" of the present invention, which lacks "i)-1: Forward primer corresponding sequence".

[0615] Furthermore, "Adapter A" of the present invention, which was manufactured in the same manner as in Example 1, was used as a comparative example.

[0616] As described below, gRNA libraries were prepared in accordance with Examples 4 to 9.

[0617] Each sequence of "Adapter A" or "Adapter A'" was ligated into a sample containing the same DNA fragment, and the experiment was carried out similarly. After PCR following ligation of "Adapter B," which contains the tracrRNA sequence, electrophoresis was performed, and the efficiency of library preparation was evaluated from the generation rate of the target PCR band.

[0618] The evaluation was performed before the "adapter A-derived sequence (24 bp)," the "reverse primer-compatible sequence (17 bp)" of "adapter B," and the restriction enzyme-removed portion (5 bp) were removed by restriction enzymes (between Step 3-1 and Step 3-2 of the above procedure). Therefore, in the case of Comparative Example 1, these sequences were evaluated using a band of "155 bp (Comparative Example 1)," which is longer than the "109 bp" band of Example 10 after removal (109 + 24 + 17 + 5 = 155). In the case of Comparative Example 2, the band should be shorter than that of Comparative Example 1 because there is no "forward primer-compatible sequence."

[0619] The evaluation results are shown in Figure 14.

[0620] (Results) As can be seen from Figure 14, when using "Adapter A" of the present invention, which has both a "stem-loop formation sequence" and a "primer-compatible sequence," the target band is clear. In contrast, when using "Adapter A'" of Comparative Example 1, which lacks a stem-loop structure, the signal of the target band is not clear at all, similar to the background (bands around 140-120 bp), and it was found that no yield suitable for use in specific (epi)genome editing etc. could be obtained.

[0621] Furthermore, when using "Adapter A'" from Comparative Example 2, which lacked a primer binding sequence, the desired band could not be obtained at all.

[0622] (Discussion) From the results of the above test examples, it was found that "iii): single-stranded polynucleotide sequence for loop formation" and "i)-1: forward primer corresponding sequence" are essential for "Adapter A" of the present invention.

[0623] [Example 18: Identification of Biomarkers] Cells were prepared with genome editing in the same manner as in Example 12, using EGFR, BRAF, and KRAS as target genes. Specifically, a polynucleotide library for gRNA expression was prepared for each of the oncogenes KRAS, EGFR, and BRAF, expressing gRNAs that comprehensively cover the exon region. Then, a group of gRNA expression plasmid vectors contained in each polynucleotide library, along with three types of vectors, CBE-SpRY-PURO and ABE-SpRY-PURO, were introduced into the cells. In this example, in addition to HEK293, a normal human cell, human cell KMST-6 was also used. As a result, genome-edited cultured cells with random mutations introduced into the KRAS gene, genome-edited cultured cells with random mutations introduced into the EGFR gene, and genome-edited cultured cells with random mutations introduced into the BRAF gene were obtained.

[0624] Subsequently, using the same experimental method as in Example 13, we performed a "method for identifying gene mutations that cause changes in the phenotype to be analyzed." More specifically, we used 1.0 × 10⁶ HEK293 or KMST-6 cells each containing a random single nucleotide mutation in the KRAS, EGFR, and BRAF genes. 6 The cells were cultured for 5 days in a 100 mm low-adhesion dish (#4615; Corning). The culture conditions were established according to the previously reported GILA assay 51 (PMID: 27723082). All relatively large spheroid clumps were collected from the resulting spheroids. mRNA was then recovered from the spheroid clumps (cancer cells), and the KRAS, EGFR, and BRAF gene regions were amplified by RT-PCR. Sequence analysis was then performed using a next-generation sequencer to verify single nucleotide variant sites. FC and FDR values ​​were determined for HEK293 and KMST-6 comparison samples that had not undergone genome editing. Mutations with an FC value > 1.5 and FDR < 0.01 were used as biomarkers for driver mutations (see Figures 16-21). Furthermore, for biomarkers of driver mutations that are more likely to cause hyperfunction, FC values ​​of 2 or higher may be selected, and even FC values ​​of 4.0 or higher may be selected.

[0625] The FC (Fold Change) value is a well-known index that shows how much the mutation rate changes between different conditions. Here, the mutation rate refers to the fraction of each mutation pattern included in the entire cell population. The FDR value is an index used to adjust for the rate of false positives that occur when performing multiple comparisons, and a lower FDR value indicates higher statistical reliability of the FC value.

[0626] Figures 16 to 21 show a list of driver mutations identified as biomarkers in this embodiment. Mutations registered as carcinogenic in the US OncoKB database, which is used for cancer diagnosis, are indicated as "Oncogenic" or "Likely Oncogenic" in the annotation column.

[0627] Here, we investigated whether the newly identified driver mutations actually exhibited driver activity. Specifically, after individually introducing specific single-nucleotide mutations into cultured HEK293 cells using genome editing, we examined the adhesion-independent proliferation ability, a characteristic of cancer cells, using low-adhesion dish culture.

[0628] As a result, it was statistically confirmed that the single nucleotide mutations identified in this study led to an enhancement of cancer cell-specific proliferative capacity (see Figure 22). In Figure 22, to the right of the mutation name shown in the legend, "*" indicates a p-value < 0.05 and "***" indicates a p-value < 0.001, indicating a statistically significant enhancement of proliferative capacity. From the above, it was confirmed that these mutations can serve as biomarkers for driver mutations.

[0629] [Summary of Examples and Test Examples] The above examples revealed that the production of gRNA (from the preparation of DNA fragments of the target region to the creation of a gRNA expression vector library), which conventionally took about one month, can be achieved in an astonishingly short period of about three days. Furthermore, it was found that the gRNA of the present invention can be subjected to rapid and reliable comprehensive genome editing (single nucleotide substitution) by using the adapter of the present invention (see Figure 8, etc.).

[0630] Moreover, from an extremely small amount of raw material (input) of about 60 ng, an astonishing amount of "polynucleotide library for gRNA production" (output) of "300 ng, or at least 30 ng" as shown in Example 10 can be obtained. Considering that this amount can be further increased by plasmid amplification after vector insertion, it is a sufficient amount to perform various evaluation tests on multiple target genes.

[0631] Furthermore, it has been found that, based on the analysis results of the effects of such genome editing, it is possible to more accurately predict or determine the potential impact of the genome on disease, physical condition, health status, and other phenotypic changes. In addition, by combining this with other information (metabolome analysis information, proteome analysis information, etc.), even more detailed determinations become possible.

[0632] Furthermore, since gRNA can be used not only for single-nucleotide substitutions by base editors such as deaminase, but also for editing that involves base deletion and insertion, it has been sufficiently demonstrated that gRNA libraries prepared using the adapter of the present invention can be used not only for single-nucleotide substitutions, but also for "genome editing" such as the introduction of base sequences by so-called prime editing using reverse transcriptase (RT), and "epigenome editing," which includes genome function control using enzymes that have genome function control activity without altering the base sequence.

[0633] By using the "adapter," "polynucleotide library for gRNA production," "vector library for gRNA expression," or "method for producing a polynucleotide library for gRNA expression" of the present invention, a diverse range of gRNA libraries, with target sites spanning the entire target genome, can be easily and quickly produced even from very small amounts of raw materials. The "gRNA library" obtained in this manner can be used in any system that applies the CRISPR-Cas9 system, including so-called genome editing (including cleavage, mutagenesis, prime editing, and other base sequence editing) as well as epigenome editing.

[0634] In particular, by performing genome editing or epigenome editing on the "genome-edited or epigenome-edited cultured cells" of the present invention that incorporate the "vector library" of the present invention, it is possible to rapidly identify gene mutations that cause changes (worsening or improvement) in various phenotypes, including not only diseases with a large number of patients, such as cancer, but also rare diseases, changes in physical condition and / or constitution, or drug resistance.

[0635] Furthermore, since genome editing sites are not limited to exon regions, it becomes possible to analyze various changes in gene function caused by mutations in non-coding regions (of proteins). In addition, by performing prime editing using the gRNA expression vector library of the present invention, it becomes possible to rapidly and easily impart new functions to genes.< / iii> < / ii> < / iii> < / ii> < / iii> < / ii> < / iii> < / ii> < / iii> < / iii> < / iii> < / ii> < / iii> < / ii> < / iii> < / ii> < / iii> < / ii> < / iii> < / ii> < / iii> < / ii> < / iii> < / iii> < / iii> < / iii> < / iii> < / iii> < / iii> < / iii> < / iii> < / iii> < / iii> < / iii> < / ii> < / ii> < / ii> < / iii> < / ii> < / iii> < / iii> < / ii> < / ii>

Claims

An adapter for constructing a gRNA library from a group of double-stranded polynucleotide fragments, having the structure described in A) below, a stem-loop type adapter. A) Structures having the polynucleotide chains described in i) to iii) below: i) Polynucleotide chains having the sequences i)-1 to i)-3 below i)-1: Forward primer corresponding sequence i)-2: Recognition site sequence of a restriction enzyme capable of cleaving the gRNA spacer sequence from the 5' end of a polynucleotide fragment bound to an adapter. i)-3: Recognition site sequence of a restriction enzyme capable of cleaving an adapter having the structure described in A) from a polynucleotide fragment. ii) A polynucleotide chain having a sequence complementary to the polynucleotide chain in i) above. iii) A polynucleotide chain having a single-stranded polynucleotide sequence for loop formation, which connects the 5' end of the polynucleotide chain of i) and the 3' end of the polynucleotide chain of ii) in a double-stranded structure formed by the complementary linkage of the polynucleotide chains of i) and ii). A polynucleotide library for gRNA production comprising a group of double-stranded polynucleotide fragments, wherein the polynucleotide fragments comprise at least one of the polynucleotide fragments described in 1) to 7) below. 1) A polynucleotide fragment to which the adapter described in claim 1 is attached to both ends, 2) In the polynucleotide fragment of 1) above, the loop formed by the single-stranded polynucleotide chain of the adapter iii) above is cleaved, and the polynucleotide fragment is amplified by polymerase chain reaction (PCR) using a label-bound forward primer. 3) The polynucleotide fragments obtained by further cleaving the polynucleotide fragments of 2) with a restriction enzyme that recognizes the sequence of i)-2, 4) A polynucleotide fragment in which an adapter having the structure of B) below is further attached to the restriction enzyme cleavage site of the polynucleotide fragment of 3), B) Structures having the polynucleotide chains of iv) and v) below: iv) Polynucleotide chains having the sequences iv)-1 to iv)-3 below iv)-1: Scaffold sequence for Cas protein or its variant protein iv)-2: Recognition site sequence of a restriction enzyme capable of cleaving a portion of an adapter having the structure described in B) from a polynucleotide fragment. iv)-3: Reverse primer corresponding sequence v) A polynucleotide chain having a sequence complementary to the polynucleotide chain of iv) above. 5) The polynucleotide fragments from 4) above, purified using a label attached to a forward primer, 6) A polynucleotide fragment obtained by PCR using the two primers vi)-1 and vi)-2 described below, from the polynucleotide fragment of 4) or 5), vi)-1: Forward primer corresponding to the sequence of i)-1 above vi)-2: Reverse primer corresponding to the sequence of iv)-3 above. 7) A polynucleotide fragment obtained by cleaving the polynucleotide fragment of 6) at the restriction enzyme recognition sites of the sequences i)-3 and iv)-2. A polynucleotide library for gRNA production comprising a group of double-stranded polynucleotide fragments, wherein the polynucleotide fragments comprise at least one of the polynucleotide fragments described in 8) to 11) below. 8) A polynucleotide fragment in which an adapter A' having the structure A' below is bound to one end of a gRNA spacer sequence, and an adapter B having the structure B) below is bound to the other end. A') Structure having the polynucleotide chains described in i) and ii) below: i) Polynucleotide chains having the sequences i)-1 to i)-3 below i)-1: Forward primer corresponding sequence i)-2: Recognition site sequence of a restriction enzyme capable of cleaving the gRNA spacer sequence from the 5' end of a polynucleotide fragment bound to an adapter. i)-3: Recognition site sequence of a restriction enzyme capable of cleaving an adapter having the structure described in A') from a polynucleotide fragment. ii) A polynucleotide chain having a sequence complementary to the polynucleotide chain in i) above. B) Structures having the polynucleotide chains of iv) and v) below: iv) Polynucleotide chains having the sequences iv)-1 to iv)-3 below iv)-1: Scaffold sequence for Cas protein or its variant protein iv)-2: Recognition site sequence of a restriction enzyme capable of cleaving a portion of an adapter having the structure described in B) from a polynucleotide fragment. iv)-3: Reverse primer corresponding sequence v) A polynucleotide chain having a sequence complementary to the polynucleotide chain of iv) above. 9) The polynucleotide fragments from 8) above, purified using labels attached to forward primers, 10) A polynucleotide fragment obtained by PCR using the two primers vi)-1 and vi)-2 described below, from the polynucleotide fragment of 8) or 9), vi)-1: Forward primer corresponding to the sequence of i)-1 above vi)-2: Reverse primer corresponding to the sequence of iv)-3 above. 11) A polynucleotide fragment obtained by cleaving the polynucleotide fragment of 10) at the restriction enzyme recognition sites of the sequences i)-3 and iv)-2. A gRNA expression vector library in which polynucleotide fragments in the library according to claim 2 or 3 are introduced into a vector capable of expressing an exogenous gene. the below described and <ii> Genome-edited or epigenome-edited cultured cells into which the following has been introduced.< / ii> Vectors included in the gRNA expression vector library according to claim 4 <ii> A vector expressing the Cas protein or its mutant protein.< / ii> Furthermore, see below <iii> A genome-edited or epigenome-edited cultured cell according to claim 5, into which the vector has been introduced.< / iii> <iii> A vector expressing an enzyme having genome editing activity or epigenome editing activity.< / iii> An adapter set for constructing a gRNA library from a group of double-stranded polynucleotide fragments, comprising adapter A having the structure described in A) below and adapter B having the structure described in B) below. A) Structures having the polynucleotide chains described in i) to iii) below: i) Polynucleotide chains having the sequences i)-1 to i)-3 below i)-1: Forward primer corresponding sequence i)-2: Recognition site sequence of a restriction enzyme capable of cleaving the gRNA spacer sequence from the 5' end of a polynucleotide fragment bound to an adapter. i)-3: Recognition site sequence of a restriction enzyme capable of cleaving an adapter having the structure described in A) from a polynucleotide fragment. ii) A polynucleotide chain having a sequence complementary to the polynucleotide chain in i) above. iii) A polynucleotide chain having a single-stranded polynucleotide sequence for loop formation, which connects the 5' end of the polynucleotide chain of i) and the 3' end of the polynucleotide chain of ii) in a double-stranded structure formed by the complementary linkage of the polynucleotide chains of i) and ii). B) Structures having the polynucleotide chains of iv) and v) below: iv) Polynucleotide chains having the sequences iv)-1 to iv)-3 below iv)-1: Scaffold sequence for Cas protein or its variant protein iv)-2: Recognition site sequence of a restriction enzyme capable of cleaving a portion of an adapter having the structure described in B) from a polynucleotide fragment. iv)-3: Reverse primer corresponding sequence v) A polynucleotide chain having a sequence complementary to the polynucleotide chain of iv) above. A kit for preparing a gRNA library, including the following <X> and <Y>. <X> Adapter according to claim 1 or adapter set according to claim 7 <Y> At least one of the enzymes listed below <Y>-1 to <Y>-10 <Y>-1: Enzyme for converting target genes into polynucleotide fragments <Y>-2: An enzyme for binding adapter A of the adapter A described in claim 1 or the adapter set described in claim 7 to a polynucleotide fragment for gRNA production. <Y>-3: An enzyme for binding adapter B of the adapter set according to claim 7 to a polynucleotide fragment for gRNA production. <Y>-4: Enzyme for cleaving the single-stranded polynucleotide sequence for loop formation within adapter A. <Y>-5: Restriction enzyme capable of cleaving a gRNA spacer sequence from a polynucleotide fragment for gRNA production bound to a sequence derived from adapter A. <Y>-6: Restriction enzyme for cleaving and removing the sequence derived from adapter A after loop removal. <Y>-7: Restriction enzyme for cleaving and removing the reverse primer-corresponding sequence in adapter B. <Y>-8: Heat-resistant polymerase <Y>-9: Cas protein or its variant protein <Y>-10: Enzymes having genome editing activity or epigenome editing activity A method for preparing a polynucleotide library for gRNA production from a group of polynucleotide fragments, comprising the following steps (I) to (X). (I) A step of smoothing both ends of the polynucleotide fragment as needed. (II) The step of attaching the adapter A described below to both ends of the polynucleotide fragment. (III) A step of cleaving at least one of the single-stranded polynucleotide chains of adapter A, which was attached to the polynucleotide fragment in step (II), as described in iii) below. (IV) A step to amplify the polynucleotide fragment obtained in step (III) by performing PCR using a forward primer to which a label has been attached. (V) A step in which the polynucleotide fragment amplified in step (IV) is treated with a restriction enzyme that recognizes the sequence i)-2 below. (VI) A step of adding the adapter described below to the ends of the polynucleotide fragment processed in step (V) that have been cleaved by the restriction enzyme. (VII) A step of purifying and / or concentrating the polynucleotide fragment obtained in step (VI) using the label bound to the forward primer as an indicator. (VIII) A step to amplify the polynucleotide fragment obtained in step (VII) by performing PCR using forward primers and reverse primers. (IX) A step in which the polynucleotide fragment obtained in step (VIII) is treated with restriction enzymes that recognize the sequences i)-3 and iv)-2 below. (X) A step in which the polynucleotide fragments treated with restriction enzymes in step (VIII) are inserted into a vector capable of expressing foreign genes to produce a gRNA expression vector. A) A stem-loop adapter A for gRNA library preparation having a structure with polynucleotide chains as described in i) to iii) below: i) Polynucleotide chains having the sequences i)-1 to i)-3 below i)-1: Forward primer corresponding sequence i)-2: Recognition site sequence of a restriction enzyme capable of cleaving the gRNA spacer sequence from the 5' end of a polynucleotide fragment bound to an adapter. i)-3: Recognition site sequence of a restriction enzyme capable of cleaving an adapter having the structure described in A) from a polynucleotide fragment. ii) A polynucleotide chain having a sequence complementary to the polynucleotide chain in i) above. iii) A polynucleotide chain having a single-stranded polynucleotide sequence for loop formation, which connects the 5' end of the polynucleotide chain of i) and the 3' end of the polynucleotide chain of ii) in a double-stranded structure formed by the complementary linkage of the polynucleotide chains of i) and ii). B) Adapter B having a structure having the polynucleotide chains described in iv) and v) below: iv) A polynucleotide chain having the sequence iv)-1 to iv)-3 below. iv)-1: Scaffold sequence for Cas protein or its variant protein iv)-2: Recognition site sequence of a restriction enzyme capable of cleaving a portion of an adapter having the structure described in B) from a polynucleotide fragment. iv)-3: Reverse primer compatible sequence v) A polynucleotide chain having a sequence complementary to the polynucleotide chain of iv) above. A method for identifying candidate gene mutations that could be factors in the phenotypic changes to be analyzed, comprising the steps (XIV) to (XVI) below. (XIV) A step of editing the genome of genes in cultured cells by culturing the cultured cells described in claim 5. (XV) A step to select cultured cells or non-human animal models with altered phenotypes from the cultured cells obtained in step (XIV) or from non-human animal models transplanted with those cells. (XVI) A step to compare the genes extracted from cultured cells or non-human animal models selected in step (XV) with the genes before genome editing. A method for determining the potential impact of a genome on its phenotype, comprising the steps S1) to S4) described below. S1) A step of selecting a gene mutation that is a candidate factor for phenotypic change to be analyzed by the method described in claim 10. S2) A step of creating a list of candidate gene mutations enumerated and / or ranked based on the results of step S1. S3) A step to match the information obtained from the genome with the list in step S2. S4) Steps to determine the possibility of phenotypic changes to be analyzed. Furthermore, the method according to claim 11, further comprising the step S5 described below. S5) A process to make a decision by combining at least one of the following pieces of information with the result of process S4. S5)-1: Metabolome analysis information data S5)-2: Proteome Analysis Information Data The device comprises a storage device for storing information on gene mutations obtained by the method for identifying gene mutations according to claim 10, An information acquisition device that obtains genetic information of the genome contained in a sample, A determination system comprising a determination device that compares genetic information obtained by the information acquisition device with genetic mutation information stored in the memory device, and determines the possibility of a phenotypic change to be analyzed based on the comparison result.   A biomarker for determining EGFR gene hyperfunction, comprising an EGFR protein having any of the following amino acid mutations, or a polynucleotide having an EGFR gene mutation capable of causing any of the following amino acid mutations. p.Q32R, p.G109E, p.S116F, p.P195S, p.A237V, p.S306=, p.R324H, p.G331=, p.F376L, p.H433Q, p.K454E, p.W 477*, p.C555R, p.G614S, p.P631S, p.A647T, p.L655=, p.R675W, p.I744V, p.T751I, p.P794L, p.V802A, p.H80 5R, p.C818R, p.P848=, p.K852R, p.K875R, p.A882V, p.Q894=, p.T940=, p.F968=, p.M987R, p.P992S, p.F997V , p.E1005=, p.D1014V, p.L1034I, p.C1058Y, p.D1084N, p.S1096=, p.R1100S, p.P1108=, p.A1181T, p.V1200= A biomarker for determining KRAS gene hyperfunction, comprising a KRAS protein having any of the following amino acid mutations, or a polynucleotide having a KRAS gene mutation capable of causing any of the following amino acid mutations. p.T2I, p.V7A, p.V7G, p.V8E, p.V9=, p.V9D, p.G10=, p.G10V, p.A11P, p.A11V , p.G12=, p.V14A, p.G15=, p.G15C, p.S17T, p.A18D, p.A18V, p.T20S, p.I21K, p.Q22L, p.Q22H, p.I24N, p.Q25*, p.H27Y, p.D30E, p.P34Q, p.P34S, p.P34T, p.T35=, p.I36L, p.S39=, p.S39Y, p.K42=, p.V44E, p.E49K, p.C51*, p.C51=, p .D57Y, p.G60=, p.E62D, p.Y64N, p.A66S, p.R68W, p.L79I, p.C80S, p.K88*, p .S89*, p.D92Y, p.I93F, p.R97I, p.E98*, p.E98=, p.K104N, p.P110S, p.V112I , p.V114I, p.G115R, p.D119N, p.P121H, p.S122F, p.R123*, p.R123I, p.D126V , p.Q129*, p.A134T, p.A134V, p.Y137=, p.G138R, p.P140S, p.R149K, p.R151T A biomarker for determining BRAF gene hyperfunction, comprising a BRAF protein having any of the following amino acid mutations, or a polynucleotide having a BRAF gene mutation capable of causing any of the following amino acid mutations. p.W48*, p.L64=, p.S122F, p.K206E, p.L312P, p.T401=, p.Q494*, p.R509Q, p.R509L, p.L514P, p.L514I, p.Y519=, p.T521K, p.A526V, p.W531C, p.H540Q, p.L553R, p.A561S, p.L567=, p.A569S, p.I572F, p.H574N , p.L577I, p.N581I, p.N581S, p.L584I, p.L588R, p.L588F, p.F595I, p.F595L, p.G596C, p.G596V, p.G59 6S, p.A598=, p.V600L, p.V600E, p.S602Y, p.R603=, p.W604C, p.W604*, p.S605R, p.G606=, p.G606V, p.Q6 09=, p.E611K, p.Q612E, p.S614Y, p.M620I, p.V624=, p.R626I, p.D629Y, p.D629G, p.P632L, p.F635I, p. Q636H, p.Y647*, p.E648Q, p.M650V, p.Q653K, p.Q653*, p.S657*, p.I659=, p.R662S, p.D663N, p.D663A, p .F667V, p.R671L, p.R671*, p.L674=, p.R682W, p.P705Q, p.E715*, p.R719P, p.H725Q, p.S727I, p.S732=, p.N734S, p.R735Q, p.T740K, p.E741=, p.C748S, p.C748Y, p.I755=, p.G758=, p.G758E, p.G759E, p.G759R A method for testing for EGFR gene hyperfunction, comprising comparing information on the EGFR protein or its gene based on a biological sample from a subject with information consisting of a dataset containing the biomarker of claim 14 to detect whether there are matching mutations. A method for testing for hyperfunction of the KRAS gene, comprising comparing information on the KRAS protein or its gene based on a biological sample from a subject with information consisting of a dataset containing the biomarker of claim 15 to detect whether there are matching mutations. A method for testing for BRAF gene hyperfunction, comprising comparing information on the BRAF protein or its gene based on a biological sample from a subject with information consisting of a dataset containing the biomarker of claim 16 to detect whether there are matching mutations.