Nucleic acid construct capable of selectively detecting or killing mismatch repair-deficient cell, and use thereof

JPWO2024106448A5Pending Publication Date: 2025-09-01
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024558906
Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2024-05-23
Publication Date
2025-09-01

AI Technical Summary

Technical Problem

Current methods for diagnosing and treating mismatch repair-deficient cancers, such as Lynch syndrome, are inefficient due to low coincidence rates between microsatellite instability and immunostaining results, and existing treatments like immune checkpoint inhibitors come with severe side effects, necessitating a more precise and safer diagnostic and therapeutic approach.

Method used

A nucleic acid construct that enables protein expression through two or more single-strand annealing (SSA) recombination reactions, allowing for rapid detection of mismatch repair activity and targeted therapy by expressing proteins that can reduce cell survival or serve as diagnostic markers.

Benefits of technology

Enables efficient detection of mismatch repair defects in cancer cells and potentially lethal side effects from immune checkpoint inhibitors are avoided, providing a more specific and safer therapeutic option for mismatch repair-deficient cancers.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The nucleic acid construct according to the present invention has such a structure that a protein can be expressed by the occurrence of at least two times of a recombinant reaction by single-strand annealing (SSA), is useful for the diagnosis and treatment of mismatch repair (MMR) deficient cancer, is applicable to the searching of a factor capable of regulating a reaction of MMR or SSA, an inhibitor or an activator for the reaction of MMR or SSA, and is also applicable to the prediction of whether of not a gene mutation identified in a cancer patient is a mutation that deteriorates an MMR activity. A conventional SSA construct that has been developed in the past by the present inventors has such a structure that a protein is expressed by the occurrence of one time of a SSA reaction. According to the construct of the present invention, the difference in the protein expression amount (activity) between the case where an MMR deficient is present and the case where an MMR deficient is absent occurs more greatly compared with that in the conventional SSA construct. Therefore, the construct of the present invention is expected to act on an MMR deficient cell more specifically and provide, for example, a treatment means having fewer adverse side effects.
Need to check novelty before this filing date? Find Prior Art

Description

Nucleic acid construct capable of selectively detecting or killing mismatch repair-deficient cells and use thereof

[0001] The present invention relates to a nucleic acid construct and a therapeutic or diagnostic agent for mismatch repair-deficient cancer, which comprises the nucleic acid construct.

[0002] Mismatch repair is one of the mechanisms by which cells repair DNA damage and plays an important role in maintaining genome stability (Non-Patent Documents 1 and 2). In human cells, proteins encoded by genes such as MSH2, MSH6, MLH1, and PMS2 are involved in mismatch repair, and cells lacking these factors lose mismatch repair activity. Abnormalities in factors that affect the gene expression of these mismatch repair proteins can also cause mismatch repair deficiency.

[0003] Mismatch repair deficiency causes various cancers (when base mismatches in DNA remain unrepaired, genomic mutations accumulate, which causes high carcinogenicity). Lynch syndrome (also known as hereditary nonpolyposis colorectal cancer (HNPCC)) is well-known as a familial (hereditary) tumor, but deficiencies in mismatch repair factors have also been reported in many sporadic (non-hereditary) cancers (especially colorectal cancer and endometrial cancer) (Non-Patent Document 3).

[0004] Mismatch repair-deficient cancers are diagnosed by microsatellite instability (MSI) or immunohistochemistry for the four factors. However, the results of both methods are not 100% consistent (Non-Patent Document 4). Furthermore, the latter method, in particular, may miss mismatch repair deficiencies caused by abnormalities other than the four factors (the same applies to genetic testing for the four factors). Both methods require several days to several weeks for diagnosis.

[0005] With regard to treatment, immune checkpoint inhibitors (such as anti-PD-1 antibodies) have recently been found to have a high response rate for mismatch repair-deficient cancers, attracting considerable attention (Non-Patent Documents 3 and 5). However, immune checkpoint inhibitors have been reported to cause autoimmune diseases due to immunosuppression, some of which can be fatal, and therefore the development of safer and less expensive treatments is also important.

[0006] WO 2021 / 162121 A1WO 2021 / 162120 A1

[0007] Jiricny J (2006) The multifaceted mismatch-repair system. Nat Rev Mol Cell Biol7(5):335-346.Harfe BD and Jinks-Robertson S (2000) DNA mismatch repair and genetic instability. Annu Rev Genet 34:359-399.Le DT et al. Science. 2017 Jul 28;357(6349):409-413Hampel H et al. N Engl J Med 352:1851-1860, 2015Le DT et al. N Engl J Med 2015; 372: 2509-2520.Stark JM et al. Mol Cell Biol. 2004 Nov;24(21):9305-9316.Mendez-Dorantes C et al. Genes Dev. 2018 Apr 1;32(7-8):524-536.Ochiai H et al. Genes Cells. 2010 Aug;15(8):875-885.

[0008] "If we could easily and quickly detect the presence or absence of mismatch repair activity in cells, tissues, and individuals, it would be an extremely useful method for diagnosing and treating mismatch repair-deficient cancers, including Lynch syndrome, regardless of the type of cancer. For example, in cancer treatment with immune checkpoint inhibitors, if we could diagnose in advance whether mismatch repair is normal or not, we could avoid the risk of administering the drug to patients with a low response rate and causing serious side effects."

[0009] The present inventors previously developed a nucleic acid construct as a system for detecting the presence or absence of mismatch repair activity. The construct is constructed by dividing a gene sequence encoding a protein into two fragments, and then incorporating the two fragments into an expression vector with homologous regions (Patent Document 1). This construct allows the gene sequence to be reproduced by a single-strand annealing (SSA) recombination reaction between a pair of homologous regions, enabling protein expression. When the construct is introduced into cells with impaired mismatch repair activity, the SSA mechanism is activated, reproducing the gene sequence and expressing the protein.

[0010] The present inventors have also developed a construct with a structure that allows the gene sequence encoding a protein to be reproduced by homologous recombination (HR), thereby enabling protein expression (Patent Document 2).

[0011] The object of the present invention is to provide a new means that enables simple and rapid detection of the presence or absence of mismatch repair activity, and that has superior characteristics to previously developed constructs as a useful means for diagnosing and treating mismatch repair-deficient cancers.

[0012] As a result of further intensive research, the inventors of the present application succeeded in developing a nucleic acid construct with a structure that enables protein expression when SSA-mediated recombination reactions occur two or more times, and discovered that this nucleic acid construct is useful for the diagnosis and treatment of mismatch repair-deficient cancers, thereby completing the present invention.

[0013] That is, the present invention relates to a nucleic acid construct that enables protein expression by two or more SSA reactions and uses thereof, and includes the following aspects.

[0014] (Nucleic acid construct of first aspect) [1] A nucleic acid construct comprising: a promoter region; a 5' region of a gene sequence encoding protein A; a 3' region of a gene sequence encoding protein A; and at least one complementary region comprising: homologous region α, which uses a region of the 5' region comprising at least the 3'-end portion as a substrate for a recombination reaction by single-strand annealing, and homologous region β, which uses a region of the 3' region comprising at least the 5'-end portion as a substrate for the recombination reaction, on one nucleic acid molecule or on two or more different nucleic acid molecules; wherein homologous region α and homologous region β are contained in the same or different complementary regions, and in the latter case, at least one complementary region comprises one or more sets of additional homologous regions and substrates that cause the recombination reaction; the homology between each homologous region and each region serving as its substrate is 40% or more and less than 100%; and wherein the recombination reaction forms a nucleic acid having a base sequence that encodes protein A or protein B having the same activity as protein A, and protein A or protein B is expressed. [2] The nucleic acid construct according to [1], wherein the amino acid sequence of protein B has 80% or more but less than 100% sequence identity with the amino acid sequence of protein A. [3] The nucleic acid construct according to [1] or [2], wherein the promoter region, the 5' region, and the 3' region are located on the same nucleic acid molecule. [4] The nucleic acid construct according to claim [1] or [2], wherein the promoter region and the 5' region are located on one nucleic acid molecule, and the 3' region is located on another nucleic acid molecule. [5] The nucleic acid construct according to any one of [1] to [4], wherein a poly(A) addition signal is functionally linked downstream of the 3' region.[6] A nucleic acid construct according to [3], which is composed of two nucleic acid molecules and includes one complementary region, wherein nucleic acid molecule 1 includes a promoter region, the 5'-side region, and the 3'-side region; and nucleic acid molecule 2 includes a complementary region including homologous region α and homologous region β; the homology between homologous region α and a region of the 5'-side region that is a substrate for the recombination reaction and that includes at least the 3'-end portion is 40% or more but less than 100%, and the homology between homologous region β and a region of the 3'-side region that is a substrate for the recombination reaction and that includes at least the 5'-end portion is 40% or more but less than 100%; when the 5'-side region, complementary region, and 3'-side region are arranged in this order, each homologous region and its substrate overlap with each other to encode the amino acid sequence of protein A or protein B; and protein A or protein B is expressed by the recombination reaction that occurs between each homologous region and its substrate.[7] The nucleic acid construct according to [3], which is composed of n nucleic acid molecules and contains (n-1) complementary regions, wherein nucleic acid molecule 1 contains a promoter region, the 5'-side region, and the 3'-side region; nucleic acid molecule 2 contains a first complementary region containing a first homologous region and a second homologous region; mth nucleic acid molecule m contains a (m-1)th substrate region and a (m-1)th complementary region containing the mth homologous region; the first homologous region is homologous region α, and the homology between the homologous region and a region containing at least the 3'-end portion of the 5'-side region, which is its substrate in the recombination reaction, is 40% or more but less than 100%; and the (m-1)th homologous region uses the (m-1)th substrate region as a substrate for the recombination reaction, and the homology between the two is 40% or more but less than 100%; a nucleic acid construct in which the nth homologous region in nucleic acid molecule n is homologous region β, the homology between the homologous region and a region of the 3'-side region that is its substrate in the recombination reaction and that includes at least the 5'-end portion is 40% or more but less than 100%, the 5'-side region, the first complementary region, the (m-1)th complementary region, and the 3'-side region, when arranged in this order, encode the amino acid sequence of protein A or protein B with mutual overlap between each homologous region and its substrate, and protein A or protein B is expressed by the recombination reaction that occurs between each homologous region and its substrate (wherein integer n is a constant satisfying 3≦n≦10, and integer m is a variable satisfying 3≦m≦n). [8] The nucleic acid construct according to [6] or [7], wherein nucleic acid molecule 1 is a linear nucleic acid molecule comprising a promoter region and the 5' region downstream of the 3' region in this order, or a circular nucleic acid molecule comprising the 5' region and the 3' region downstream of a promoter region in this order. [9] The nucleic acid construct according to [8], wherein the circular nucleic acid molecule has a cleavage site between the 5' region and the 3' region.

[10] The nucleic acid construct according to [9], wherein the cleavage site is a restriction enzyme recognition site.

[11] The nucleic acid construct according to [4], which is composed of three nucleic acid molecules and includes one complementary region, wherein nucleic acid molecule 1-1 includes a promoter region and the 5'-side region, nucleic acid molecule 1-2 includes the 3'-side region, and nucleic acid molecule 2 includes a complementary region including homologous region α and homologous region β, the homology between homologous region α and a region of the 5'-side region that is a substrate for the recombination reaction and that includes at least the 3'-end portion is 40% or more and less than 100%, and the homology between homologous region β and a region of the 3'-side region that is a substrate for the recombination reaction and that includes at least the 5'-end portion is 40% or more and less than 100%, and when the 5'-side region, complementary region, and 3'-side region are arranged in this order, each homologous region and its substrate overlap with each other to encode the amino acid sequence of protein A or protein B, and protein A or protein B is expressed by the recombination reaction that occurs between each homologous region and its substrate.

[12] The nucleic acid construct according to [4], which is composed of (n+1) nucleic acid molecules and contains (n-1) complementary regions, wherein nucleic acid molecule 1-1 contains a promoter region and the 5'-side region, nucleic acid molecule 1-2 contains the 3'-side region, nucleic acid molecule 2 contains a first complementary region containing a first homologous region and a second homologous region, the mth nucleic acid molecule m contains the (m-1)th substrate region and the (m-1)th complementary region containing the mth homologous region, the first homologous region is homologous region α, and the homology between the homologous region and a region containing at least the 3'-end portion of the 5'-side region, which is its substrate in the recombination reaction, is 40% or more and less than 100%, and the (m-1)th homologous region uses the (m-1)th substrate region as a substrate for the recombination reaction, and the homology between them is 40% or more and less than 100%, a nucleic acid construct in which the nth homologous region in nucleic acid molecule n is homologous region β, the homology between the homologous region and a region of the 3'-side region that is its substrate in the recombination reaction and that includes at least the 5'-end portion is 40% or more but less than 100%, the 5'-side region, the first complementary region, the (m-1)th complementary region, and the 3'-side region, when arranged in this order, encode the amino acid sequence of protein A or protein B with mutual overlap between each homologous region and its substrate, and protein A or protein B is expressed by the recombination reaction that occurs between each homologous region and its substrate (wherein integer n is a constant satisfying 3≦n≦10, and integer m is a variable satisfying 3≦m≦n).

[13] The nucleic acid construct according to [3], which is composed of one circular nucleic acid molecule and contains one complementary region, wherein the circular nucleic acid molecule contains the 5'-side region, the cleavage site, and the 3'-side region in this order downstream of a promoter region, and a complementary region containing homologous region α and homologous region β upstream of the promoter region and downstream of the 3'-side region, wherein the homology between a region of the 5'-side region comprising at least the 3'-end portion and homologous region α which uses the region as a substrate for the recombination reaction is 40% or more but less than 100%, and the homology between a region of the 3'-side region comprising at least the 5'-end portion and homologous region β which uses the region as a substrate for the recombination reaction is 40% or more but less than 100%, and when arranged in this order, the 5'-side region, complementary region, and 3'-side region encode the amino acid sequence of protein A or protein B with mutual overlap between each homologous region and its substrate, and protein A or protein B is expressed by the recombination reaction which occurs between each homologous region and its substrate.

[14] The nucleic acid construct according to [3], which is composed of one circular nucleic acid molecule and contains (n-1) complementary regions, wherein the circular nucleic acid molecule contains the 5'-side region, the cleavage site, and the 3'-side region downstream of the promoter region, in this order, and contains (n-1) complementary regions upstream of the promoter region and downstream of the 3'-side region, the first complementary region contains a first homologous region and a second homologous region, the (m-1)th complementary region contains the (m-1)th substrate region and the mth homologous region, the first homologous region is homologous region α which uses a region of the 5'-side region comprising at least the 3'-end portion as a substrate for the recombination reaction, and the homology between the substrate and the first homologous region is 40% or more but less than 100%, and the (m-1)th homologous region uses the (m-1)th substrate region as a substrate for the recombination reaction, and the homology between them is 40% or more but less than 100%, a nucleic acid construct according to

[13] or

[14] , wherein the nth homologous region of the (n-1)th complementary region is a homologous region β that uses a region of the 3'-side region comprising at least the 5'-end portion as a substrate for the recombination reaction, the homology between the substrate and the nth homologous region being 40% or more but less than 100%, and the 5'-side region, the first complementary region, the (m-1)th complementary region, and the 3'-side region, when arranged in this order, encode the amino acid sequence of protein A or protein B with mutual overlap between each homologous region and its substrate, and the recombination reaction occurring between each homologous region and its substrate results in expression of protein A or protein B (wherein integer n is a constant satisfying 3≦n≦10, and integer m is a variable satisfying 3≦m≦n).

[15] The nucleic acid construct according to

[13] or

[14] , wherein the cleavage site is a restriction enzyme recognition site.

[16] A nucleic acid construct according to [3], which is composed of one linear nucleic acid molecule and contains one complementary region, wherein the linear nucleic acid molecule contains, from upstream to downstream, the 3' region, the complementary region, the promoter region, and the 5' region, in this order; the complementary region includes homologous region α and homologous region β; the homology between a region of the 5' region comprising at least the 3'-end portion and homologous region α, which uses the region as a substrate for the recombination reaction, is 40% or more and less than 100%; the homology between a region of the 3' region comprising at least the 5'-end portion and homologous region β, which uses the region as a substrate for the recombination reaction, is 40% or more and less than 100%; and when the 5' region, complementary region, and 3' region are arranged in this order, each homologous region and its substrate overlap with each other to encode the amino acid sequence of protein A or protein B, and protein A or protein B is expressed by the recombination reaction occurring between each homologous region and its substrate.

[17] The nucleic acid construct according to [3], which is composed of one linear nucleic acid molecule and contains (n-1) complementary regions, wherein the linear nucleic acid molecule contains, from upstream to downstream, the 3' region, the promoter region, and the 5' region in this order, and contains (n-1) complementary regions between the 3' region and the promoter region, the first complementary region contains a first homologous region and a second homologous region, the (m-1)th complementary region contains the (m-1)th substrate region and the mth homologous region, the first homologous region is a homologous region α that uses a region of the 5' region containing at least the 3' end portion as a substrate for the recombination reaction, and the homology between the substrate and the first homologous region is 40% or more but less than 100%, and the (m-1)th homologous region uses the (m-1)th substrate region as a substrate for the recombination reaction, and the homology between them is 40% or more but less than 100%, a nucleic acid construct in which the nth homologous region of the (n-1)th complementary region is a homologous region β that uses a region of the 3'-side region that includes at least the 5'-end portion as a substrate for the recombination reaction, the homology between the substrate and the nth homologous region being 40% or more but less than 100%, and the 5'-side region, the first complementary region, the (m-1)th complementary region, and the 3'-side region, when arranged in this order, encode the amino acid sequence of protein A or protein B with overlap between each homologous region and its substrate, and protein A or protein B is expressed by the recombination reaction occurring between each homologous region and its substrate (wherein integer n is a constant in the range of 3≦n≦10, and integer m is a variable in the range of 3≦m≦n).

[18] The nucleic acid construct according to any one of [1] to

[17] , wherein the chain length of all of the homologous regions is at least 20 bases.

[19] The nucleic acid construct according to any one of [1] to

[18] , wherein the gene sequence encoding protein A is a gene sequence encoding a protein that has the effect of reducing cell viability or a protein whose expression in cells can be detected.

[20] The nucleic acid construct according to

[19] , wherein the gene sequence encoding protein A is a sequence of a suicide gene, a DNA damage-inducing gene, a DNA repair inhibitor gene, a luciferase gene, a fluorescent protein gene, a cell surface antigen gene, a secreted protein gene, or a membrane protein gene.

[21] A therapeutic agent for mismatch repair-deficient cancer, comprising the nucleic acid construct of any one of [1] to

[20] , wherein the gene sequence encoding protein A is a gene sequence encoding a protein having the effect of reducing cell viability.

[22] The therapeutic agent of

[21] , wherein the gene sequence encoding protein A is a sequence of a suicide gene, a DNA damage-inducing gene, or a DNA repair-inhibiting gene.

[23] A diagnostic agent for mismatch repair-deficient cancer, comprising the nucleic acid construct of any one of [1] to

[20] .

[24] The diagnostic agent of

[23] , wherein the gene sequence encoding protein A is a gene sequence encoding a protein whose expression in cells can be detected.

[25] A companion diagnostic agent for predicting the effect of an anticancer drug on mismatch repair-deficient cancer, comprising the nucleic acid construct of any one of [1] to

[20] .

[26] The companion diagnostic agent of

[25] , wherein the anticancer drug is an immune checkpoint inhibitor.

[27] The companion diagnostic agent according to

[25] or

[26] , wherein the gene sequence encoding protein A is a gene sequence encoding a protein whose expression in cells can be detected.

[28] A method for treating mismatch repair-deficient cancer, comprising administering to a patient having mismatch repair-deficient cancer the nucleic acid construct according to any one of [1] to

[20] , wherein the gene sequence encoding protein A is a gene sequence encoding a protein whose activity reduces cell viability.

[29] A method for diagnosing mismatch repair-deficient cancer, comprising: introducing the nucleic acid construct according to any one of [1] to

[20] into cancer cells of a cancer patient; and measuring the expression of protein A or protein B.

[30] The method according to

[29] , wherein the expression of protein A or protein B is measured by directly or indirectly measuring the activity of protein A or protein B.

[31] The method according to

[29] or

[30] , wherein the cancer cells are cells isolated from the cancer patient, and the introduction of the nucleic acid construct into the cancer cells is performed ex vivo.

[32] The method of

[29] or

[30] , wherein protein A or protein B is a protein whose intracellular expression can be detected as a signal, the introduction of a nucleic acid construct into the cancer cells is carried out by administering the nucleic acid construct to the cancer patient, and whether a signal from the protein is detected from the cancer lesion is examined.

[33] The method of

[29] or

[30] , wherein protein A or protein B is a secretory protein whose intracellular expression can be detected, the introduction of a nucleic acid construct into the cancer cells is carried out by administering the nucleic acid construct to the cancer patient, and the activity of the protein in blood isolated from the patient after administration of the nucleic acid construct is measured.

[34] A method for predicting the efficacy of an anticancer drug for mismatch repair-deficient cancer, comprising: introducing the nucleic acid construct of any one of [1] to

[20] into cancer cells of a cancer patient; and measuring the expression of protein A or protein B.

[35] The method of

[34] , wherein the anticancer drug is an immune checkpoint inhibitor.

[36] The method of

[34] or

[35] , wherein the cancer cells are cells isolated from the patient, and the introduction of the nucleic acid construct into the cancer cells is carried out ex vivo.

[37] The method of

[34] or

[35] , wherein protein A or protein B is a protein whose expression in cells can be detected as a signal, the introduction of the nucleic acid construct into the cancer cells is carried out by administering the nucleic acid construct to the patient, and whether a signal from the protein is detected from the cancer lesion is examined.

[38] The method of

[34] or

[35] , wherein protein A or protein B is a secretory protein whose expression in cells can be detected, the introduction of the nucleic acid construct into the cancer cells is carried out by administering the nucleic acid construct to the cancer patient, and the activity of the protein in blood isolated from the patient after administration of the nucleic acid construct is measured.

[0015] (Nucleic acid construct of second aspect)

[39] A nucleic acid construct comprising: a promoter region, two substrate regions α and β, a gene sequence encoding protein X, and at least one complementary region comprising two homologous regions that use each of the two substrate regions as substrates for a recombination reaction by single-stranded annealing, arranged on one nucleic acid molecule or on two or more different nucleic acid molecules, wherein the two homologous regions are contained in the same or different complementary regions, and in the latter case, the at least one complementary region comprises one or more sets of additional homologous regions and substrates that cause the recombination reaction, and the homology between corresponding substrate regions and homologous regions is 40% or more and less than 100%, and protein X is expressed by the recombination reaction.

[40] The nucleic acid construct according to

[39] , wherein the promoter region, substrate region α, substrate region β, and the gene sequence are arranged on the same nucleic acid molecule.

[41] The nucleic acid construct according to

[39] , wherein the promoter region and substrate region α are located on one nucleic acid molecule, and the substrate region β and the gene sequence are located on another nucleic acid molecule.

[42] The nucleic acid construct according to any one of

[39] to

[41] , wherein a poly A addition signal is operably linked downstream of the gene sequence encoding protein X.

[43] The nucleic acid construct according to

[40] , which is composed of two nucleic acid molecules and includes one complementary region, wherein nucleic acid molecule 1 includes a promoter region, substrate region α, substrate region β, and a gene sequence encoding protein X; nucleic acid molecule 2 includes one complementary region including a first homologous region and a second homologous region; the first homologous region uses substrate region α as a substrate for the recombination reaction, and the homology between the two is 40% or more and less than 100%; and the second homologous region uses substrate region β as a substrate for the recombination reaction, and the homology between the two is 40% or more and less than 100%; and protein X is expressed by the recombination reaction occurring between each homologous region and its substrate.

[44] The nucleic acid construct according to

[40] , which is composed of n nucleic acid molecules and contains (n-1) complementary regions, wherein nucleic acid molecule 1 contains a promoter region, substrate region α, substrate region β, and the gene sequence; nucleic acid molecule 2 contains a first complementary region containing a first homologous region and a second homologous region; the mth nucleic acid molecule m contains the (m-1)th substrate region and the (m-1)th complementary region containing the mth homologous region; the first homologous region uses substrate region α as a substrate for the recombination reaction, and the homology therebetween is 40% or more but less than 100%; the (m-1)th homologous region uses the (m-1)th substrate region as a substrate for the recombination reaction, and the homology therebetween is 40% or more but less than 100%; and the nth homologous region in nucleic acid molecule n uses substrate region β as a substrate for the recombination reaction, and the homology therebetween is 40% or more but less than 100%; A nucleic acid construct in which protein X is expressed by the recombination reaction occurring between each homologous region and its substrate (wherein integer n is a constant satisfying 3≦n≦10, and integer m is a variable satisfying 3≦m≦n).

[45] The nucleic acid construct according to

[43] or

[44] , wherein nucleic acid molecule 1 is a circular nucleic acid molecule comprising, downstream of a promoter region, substrate region α, substrate region β, and a gene sequence encoding protein X, in that order, or a linear nucleic acid molecule comprising, from upstream to downstream, substrate region β, a gene sequence encoding protein X, a promoter region, and substrate region α, in that order.

[46] The nucleic acid construct according to

[45] , wherein the circular nucleic acid molecule comprises a cleavage site between substrate region α and substrate region β.

[47] The nucleic acid construct according to

[46] , wherein the cleavage site is a restriction enzyme recognition site.

[48] ​​A nucleic acid construct according to

[41] , which is composed of three nucleic acid molecules and includes one complementary region, wherein nucleic acid molecule 1-1 includes a promoter region and substrate region α, nucleic acid molecule 1-2 includes substrate region β and the gene sequence, the first homologous region uses substrate region α as a substrate for the recombination reaction, and the homology between the two is 40% or more and less than 100%, and the second homologous region uses substrate region β as a substrate for the recombination reaction, and the homology between the two is 40% or more and less than 100%, and protein X is expressed by the recombination reaction occurring between each homologous region and its substrate.

[49] The nucleic acid construct according to

[41] , which is composed of (n+1) nucleic acid molecules and contains (n-1) complementary regions, wherein nucleic acid molecule 1-1 contains a promoter region and a substrate region α, nucleic acid molecule 1-2 contains a substrate region β and the gene sequence, nucleic acid molecule 2 contains a first complementary region containing a first homologous region and a second homologous region, the mth nucleic acid molecule m contains the (m-1)th substrate region and the (m-1)th complementary region containing the mth homologous region, the first homologous region uses substrate region α as a substrate for the recombination reaction, and the homology between them is 40% or more and less than 100%, and the (m-1)th homologous region uses the (m-1)th substrate region as a substrate for the recombination reaction, and the homology between them is 40% or more and less than 100%, A nucleic acid construct in which the nth homologous region of nucleic acid molecule n uses substrate region β as a substrate for the recombination reaction, the homology between the two being 40% or more but less than 100%, and protein X is expressed by the recombination reaction occurring between each homologous region and its substrate (wherein integer n is a constant in the range of 3≦n≦10, and integer m is a variable in the range of 3≦m≦n).

[50] The nucleic acid construct according to

[40] , which is composed of one circular nucleic acid molecule and includes one complementary region, wherein the circular nucleic acid molecule includes, downstream of a promoter region, a substrate region α, a cleavage site, a substrate region β, and a gene sequence encoding protein X, in this order, and includes a complementary region including a first homologous region and a second homologous region, located upstream of the promoter region and downstream of the gene sequence, wherein the first homologous region uses substrate region α as a substrate for the recombination reaction, and the homology between the first homologous region and the second homologous region is 40% or more but less than 100%, and the second homologous region uses substrate region β as a substrate for the recombination reaction, and the homology between the first homologous region and the second homologous region is 40% or more but less than 100%, and protein X is expressed by the recombination reaction occurring between each homologous region and its substrate.

[51] The nucleic acid construct according to

[40] , which is composed of one circular nucleic acid molecule and comprises (n-1) complementary regions, wherein the circular nucleic acid molecule comprises, downstream of a promoter region, a substrate region α, a cleavage site, a substrate region β, and a gene sequence encoding protein X, in this order, and comprises (n-1) complementary regions upstream of the promoter region and downstream of the gene sequence, wherein the first complementary region comprises a first homologous region and a second homologous region, wherein the (m-1)th complementary region comprises the (m-1)th substrate region and the mth homologous region, wherein the first homologous region uses substrate region α as a substrate for the recombination reaction, and the homology between the first homologous region and the second homologous region is 40% or more but less than 100%, and wherein the (m-1)th homologous region uses the (m-1)th substrate region as a substrate for the recombination reaction, and the homology between the first homologous region and the second homologous region is 40% or more but less than 100%, A nucleic acid construct, wherein the nth homologous region of the (n-1)th complementary region uses substrate region β as a substrate for the recombination reaction, and the homology between the two is 40% or more and less than 100%, and protein X is expressed by the recombination reaction occurring between each homologous region and its substrate (wherein integer n is a constant in the range of 3≦n≦10, and integer m is a variable in the range of 3≦m≦n).

[52] The nucleic acid construct according to either one of

[50] or

[51] , wherein the cleavage site is a restriction enzyme recognition site.

[53] The nucleic acid construct according to

[40] , which is composed of one linear nucleic acid molecule and includes one complementary region, wherein the linear nucleic acid molecule includes, from upstream to downstream, a substrate region β, a gene sequence encoding protein X, a complementary region, a promoter region, and a substrate region α, in this order; the complementary region includes a first homologous region and a second homologous region; the first homologous region uses substrate region α as a substrate for the recombination reaction, and the homology between the first homologous region and the second homologous region is 40% or more and less than 100%; and the second homologous region uses substrate region β as a substrate for the recombination reaction, and the homology between the second homologous region and the substrate is 40% or more and less than 100%; and protein X is expressed by the recombination reaction occurring between each homologous region and its substrate.

[54] The nucleic acid construct according to

[40] , which is composed of one linear nucleic acid molecule and comprises (n-1) complementary regions, wherein the linear nucleic acid molecule comprises, from upstream to downstream, a substrate region β, a gene sequence encoding protein X, a promoter region, and a substrate region α, in this order, and comprises (n-1) complementary regions between the gene sequence and the promoter region, a first complementary region comprises a first homologous region and a second homologous region, the (m-1)th complementary region comprises the (m-1)th substrate region and the mth homologous region, the first homologous region uses substrate region α as a substrate for the recombination reaction, and the homology between the first homologous region and the second homologous region is 40% or more but less than 100%, and the (m-1)th homologous region uses the (m-1)th substrate region as a substrate for the recombination reaction, and the homology between the first homologous region and the second homologous region is 40% or more but less than 100%, a nucleic acid construct in which the nth homologous region of the (n-1)th complementary region uses substrate region β as a substrate for the recombination reaction, the homology between the two being 40% or more and less than 100%, and protein X is expressed by the recombination reaction occurring between each homologous region and its substrate (wherein integer n is a constant in the range of 3≦n≦10, and integer m is a variable in the range of 3≦m≦n).

[55] The nucleic acid construct of any one of

[39] to

[54] , in which protein X is a protein that has the effect of reducing cell viability or a protein whose expression in cells can be detected.

[56] The nucleic acid construct of

[55] , in which the gene sequence encoding protein X is the sequence of a suicide gene, a DNA damage-inducing gene, a DNA repair inhibitor gene, a luciferase gene, a fluorescent protein gene, a cell surface antigen gene, a secreted protein gene, or a membrane protein gene.

[57] The nucleic acid construct according to any one of

[39] to

[56] , wherein the base sequence generated by the recombination reaction between each substrate region and each homologous region is a gene sequence encoding protein Y, and the recombination reaction results in expression of protein X and protein Y.

[58] The nucleic acid construct according to

[57] , wherein a sequence enabling polycistronic expression is disposed between substrate region β and the gene sequence encoding protein X.

[59] The nucleic acid construct according to

[58] , wherein the sequence enabling polycistronic expression is an IRES sequence or a 2A peptide coding sequence.

[60] The nucleic acid construct according to any one of

[57] to

[59] , wherein one of protein X and protein Y is a protein that reduces cell viability, and the other is a protein whose expression in cells can be detected.

[61] The nucleic acid construct according to

[60] , wherein one of the gene sequence encoding protein X and the gene sequence encoding protein Y is the sequence of a suicide gene, a DNA damage-inducing gene, or a DNA repair-inhibiting gene, and the other is the sequence of a luciferase gene, a fluorescent protein gene, a cell surface antigen gene, a secreted protein gene, or a membrane protein gene.

[62] A therapeutic agent for mismatch repair-deficient cancer, comprising the nucleic acid construct according to any one of

[39] to

[61] , wherein the gene sequence encoding protein X is the gene sequence encoding a protein that reduces cell viability.

[63] The therapeutic agent according to

[62] , wherein the gene sequence encoding protein X is the sequence of a suicide gene, a DNA damage-inducing gene, or a DNA repair-inhibiting gene.

[64] An agent for detecting and treating mismatch repair-deficient cancer, comprising the nucleic acid construct according to

[60] or

[61] .

[65] A diagnostic agent for mismatch repair-deficient cancer, comprising the nucleic acid construct of any one of

[39] to

[61] .

[66] The diagnostic agent of

[65] , wherein the gene sequence encoding protein X is a gene sequence encoding a protein whose intracellular expression can be detected.

[67] A companion diagnostic agent for predicting the effect of an anticancer drug on mismatch repair-deficient cancer, comprising the nucleic acid construct of any one of

[39] to

[61] .

[68] The companion diagnostic agent of

[67] , wherein the anticancer drug is an immune checkpoint inhibitor.

[69] The companion diagnostic agent of

[67] or

[68] , wherein the gene sequence encoding protein X is a gene sequence encoding a protein whose intracellular expression can be detected.

[70] A method for treating mismatch repair-deficient cancer, comprising administering to a patient with mismatch repair-deficient cancer the nucleic acid construct of any one of

[39] to

[61] , wherein the gene sequence encoding protein X is a gene sequence encoding a protein whose intracellular expression can be detected.

[71] A method for detecting and treating mismatch repair-deficient cancer, comprising administering to a cancer patient the nucleic acid construct according to

[60] or

[61] , wherein one of protein X and protein Y is a protein whose intracellular expression is detectable as a signal, and examining whether a signal of the protein is detected from a cancer lesion.

[72] A method for detecting and treating mismatch repair-deficient cancer, comprising administering to a cancer patient the nucleic acid construct according to

[60] or

[61] , wherein one of protein X and protein Y is a secreted protein whose intracellular expression is detectable, and measuring the activity of the protein in blood isolated from the patient after administration.

[73] A method for diagnosing mismatch repair-deficient cancer, comprising introducing the nucleic acid construct according to any one of

[39] to

[61] into cancer cells of a cancer patient; and measuring the expression of protein X.

[74] The method according to

[73] , wherein the expression of protein X is measured by directly or indirectly measuring the activity of protein X.

[75] The method of

[73] or

[74] , wherein the cancer cells are cells isolated from the cancer patient, and the introduction of a nucleic acid construct into the cancer cells is carried out ex vivo.

[76] The method of

[73] or

[74] , wherein protein X is a protein whose expression in cells can be detected as a signal, the introduction of the nucleic acid construct into the cancer cells is carried out by administering the nucleic acid construct to the cancer patient, and whether a signal from the protein is detected from a cancer lesion is examined.

[77] The method of

[73] or

[74] , wherein protein X is a secreted protein whose expression in cells can be detected, the introduction of the nucleic acid construct into the cancer cells is carried out by administering the nucleic acid construct to the cancer patient, and the activity of the protein in blood isolated from the patient after administration of the nucleic acid construct is measured.

[78] A method for predicting the efficacy of an anticancer drug against mismatch repair-deficient cancer, comprising: introducing the nucleic acid construct of any one of

[39] to

[61] into cancer cells of a cancer patient; and measuring the expression of protein X.

[79] The method described in

[78] , wherein the anticancer drug is an immune checkpoint inhibitor.

[80] The method of

[78] or

[79] , wherein the cancer cells are cells isolated from the patient, and the introduction of a nucleic acid construct into the cancer cells is carried out ex vivo.

[81] The method of

[78] or

[79] , wherein protein X is a protein whose expression in cells can be detected as a signal, and the introduction of the nucleic acid construct into the cancer cells is carried out by administering the nucleic acid construct to the patient, and whether a signal from the protein is detected from a cancer lesion is examined.

[82] The method of

[78] or

[79] , wherein protein X is a secretory protein whose expression in cells can be detected, and the introduction of the nucleic acid construct into the cancer cells is carried out by administering the nucleic acid construct to the cancer patient, and the activity of the protein in blood isolated from the patient after administration of the nucleic acid construct is measured.

[0016] The present invention provides a novel nucleic acid construct that can detect mismatch repair deficiency using two or more SSA reactions. The nucleic acid construct of the present invention allows the presence or absence of mismatch repair activity to be detected in a very short time using a transient expression system. Mismatch repair-deficient cancer cells can be efficiently killed by using a gene sequence encoding a protein that reduces cell viability, such as a suicide gene. Furthermore, mismatch repair-deficient cancers can be detected by using a gene sequence encoding a protein whose expression in cells can be detected, such as a luciferase gene or a fluorescent protein gene. The present invention significantly contributes to the diagnosis and treatment of various mismatch repair-deficient cancers, including Lynch syndrome, regardless of the type of cancer. Furthermore, the nucleic acid construct of the present invention can be used to discover factors, inhibitors, and activators that positively or negatively regulate mismatch repair and single-strand annealing reactions, and to predict whether genetic mutations identified in cancer patients impair mismatch repair activity.

[0017] The SSA construct previously developed by the present inventors (Patent Document 1) was a structure in which a normal gene sequence was formed and protein was expressed by a single SSA reaction. In contrast, the construct of the present invention is a structure that requires two or more SSA reactions to enable protein expression. The construct of the present invention is more likely to produce a larger difference in protein expression level (activity) depending on the presence or absence of mismatch repair deficiency than conventional SSA constructs (see Examples below). Therefore, it is expected that the construct can act more specifically on mismatch repair-deficient cells, providing a therapeutic method with, for example, fewer side effects.

[0018] This figure illustrates the structure of the two-part, two-time SSA Nluc construct (construct of the first embodiment) prepared in Example AI. The two nucleic acid molecules constituting the construct (A), the SceNluc fragment containing the promoter region (B), and the iNluc fragment serving as the complementary region, with (C) and without (D) additional sequences at both ends. This figure illustrates the structure of the pIRES plasmid vector used in the examples (partially modified from the structure diagram of Clontech's pIRES Vector Information). This figure illustrates the structure of the pUC19 plasmid vector used in the examples (source: Takara Bio's pUC19 DNA data sheet). This figure illustrates the sequence of the iNluc fragment prepared in Example AI, with the homology between the homologous region and the substrate reduced to 95%, 90%, 80%, and 70%. As shown in the figure, silent mutations were introduced scattered throughout the homologous sequence. (Continued from Figure 4-1) This figure illustrates the sequence of the iNluc fragment prepared in Example AI, with the homology between the homologous region and the substrate reduced to 95%, 90%, 80%, and 70%. As shown in the figure, silent mutations were introduced scattered throughout the homologous sequence. This is the sequence of the iNluc fragment prepared in Example AI, with homology between the homologous region and the substrate reduced to 99-96% and 94%. As shown in the figure, silent mutations were introduced scattered throughout the homologous sequence. (Continued from Figure 4-3) This is the sequence of the iNluc fragment prepared in Example AI, with homology between the homologous region and the substrate reduced to 99-96% and 94%. As shown in the figure, silent mutations were introduced scattered throughout the homologous sequence. Example of luciferase assay results in Nalm-6 cells transfected with a bipartite, twice-SSA Nluc construct and its MSH2-reverted cells. An iNluc fragment without additional sequence (homology 100%, 95%, 90%, 80%, 70%) was cotransfected with pCMV-SceNluc, and luciferase activity was measured over time. Example of luciferase assay results for Nalm-6 cells transfected with two-part, two-times SSA-type Nluc constructs and their MSH2-reverted cells. Luciferase activity was measured 4 hours after transfection with the constructs (homology 100-94%, 90%, 80%, 70%).Example of luciferase assay results for Nalm-6 cells transfected with a two-part, two-times SSA-type Nluc construct and their MSH2-reverted cells. Luciferase activity 24 hours after transfection with the construct (homology 100-94%, 90%, 80%, 70%). Example of luciferase assay results for Nalm-6 cells transfected with a two-part, two-times SSA-type Nluc construct and their MSH2-reverted cells. Luciferase activity was compared 24 hours after transfection between an iNluc fragment with an additional sequence (pUC-iNluc / AhdI, left side of graph) and an iNluc fragment without the additional sequence (pUC-iNluc / BamHI + EcoRI, right side of graph). Example of luciferase assay results for HCT116 cells transfected with a two-part, two-times SSA-type Nluc construct and their MLH1-reverted cells. Luciferase activity 4 hours after transfection with constructs (homology 100-94%, 90%, 80%, 70%). Example results of luciferase assay in HCT116 cells transfected with a two-part, two-times SSA-type Nluc construct and its MLH1-reverted cells. Luciferase activity 24 hours after transfection with a two-part, two-times SSA-type Nluc construct and its MLH1-reverted cells. Luciferase activity was compared 24 hours after transfection between iNluc fragments with additional sequences (pUC-iNluc / AhdI, left side of graph) and iNluc fragments without additional sequences (pUC-iNluc / BamHI + EcoRI, right side of graph). 1 shows the structure of the one-molecule, double-SSA integrated Nluc construct (construct of the first embodiment) prepared in Example A-II. This figure shows the luciferase activity 24 hours after the integrated Nluc construct was introduced into Nalm-6 cells and their MSH2-reverted cells. This figure shows the structure of the two-part, two-SSA NlucP construct (construct of the first embodiment) prepared in Example A-III. This figure shows the luciferase activity 24 hours after the two-part, two-SSA NlucP construct (homology 100%, 98%, 96%, 94%, 90%) was introduced into Nalm-6 cells and their MSH2-reverted cells.Figure 13 illustrates the structure and mechanism of the three-part, three-times SSA-type Nluc construct (construct of the first embodiment) prepared in Example A-IV. Figure 14 illustrates the structure and mechanism of the four-part, three-times SSA-type Nluc construct (construct of the first embodiment) prepared in Example A-IV. Figure 15 illustrates the sequence of the iNluc3-1 fragment prepared in Example A-IV, in which the homology between the homologous region and the substrate was reduced to 90% and 80% (sequence of part A in Figure 13). As shown in the figure, silent mutations were introduced scattered throughout the homologous sequence. Figure 13 illustrates the sequences of the iNluc3-1 fragment and iNluc3-2 fragment prepared in Example A-IV, in which the homology between the homologous region and the substrate was reduced to 90% and 80% (sequence of part B in Figure 13). As shown in the figure, silent mutations were introduced scattered throughout the homologous sequence. This figure shows the sequence of the iNluc3-2 fragment prepared in Example A-IV, in which the homology between the homologous region and the substrate was reduced to 90% and 80% (the sequence of part C in Figure 13). As shown in the figure, silent mutations were introduced scattered throughout the homologous sequence. This figure shows an example of the results of a luciferase assay in Nalm-6 cells transfected with a three-part, three-times SSA-type Nluc construct and their MSH2-reverted cells (results 4 hours after gene transfection). This figure shows an example of the results of a luciferase assay in Nalm-6 cells transfected with a three-part, three-times SSA-type Nluc construct and their MSH2-reverted cells (results 24 hours after gene transfection). This figure shows the results of luciferase activity measurement 4 and 24 hours after transfection of HCT116 cells and their MLH1-reverted cells with a three-part, three-times SSA-type Nluc construct. This figure shows the results of luciferase activity measurement 24 hours after transfection of Nalm-6 cells and their MSH2-reverted cells with a four-part, three-times SSA-type Nluc construct. 1 shows the results of measuring luciferase activity 24 hours after transfection of a 4-division, 3-SSA Nluc construct into HCT116 cells and their MLH1-reverted cells. 2 shows the results of measuring luciferase activity 4 hours after transfection of a 2-division, 2-SSA GeNL construct into Nalm-6 cells and their MSH2-reverted cells.

[0033] Figure 1 shows the results of measuring luciferase activity 24 hours after transfection of a two-part, two-times SSA-type GeNL construct into Nalm-6 cells and their MSH2-reverted cells. Figure 2 shows the structure of a two-part, two-times SSA-type Nluc-homeo DTA-linked construct (construct of the second embodiment) prepared in Example B, which expresses two types of proteins in a polycistronic manner by SSA. Figure 3 shows the results of comparing the viability between Nalm-6 cells and their MSH2-reverted cells 96 hours after transfection of a two-part, two-times SSA-type Nluc-homeo DTA-linked construct. Figure 4 shows the results of comparing the viability between HCT116 cells and their MLH1-reverted cells 72 hours after transfection of a two-part, two-times SSA-type Nluc-homeo DTA-linked construct. Figure 5 shows the structure of a two-part, two-times SSA-type Nluc-homeo TK-linked construct (construct of the second embodiment) prepared in Example C, which expresses two types of proteins in a polycistronic manner by SSA. Figure 27-1 shows the results of comparing the survival rates of Nalm-6 cells and their MSH2-reverted cells 96 hours after transfection with a two-part, two-times SSA Nluc-homeoTK ligated construct. Figure 27-1 shows the structure of a two-part, two-times SSA DTA construct (construct of the first embodiment) prepared in Example DI. Figure 27-1 shows the sequence of an iDTA2 fragment prepared in Example DI, in which the homology between the homologous region and the substrate was reduced to 99-94%, 90%, 80%, and 70%. As shown in the figure, silent mutations were introduced scattered throughout the homologous sequence. (Continued from Figure 27-1) Figure 27-1 shows the sequence of an iDTA2 fragment prepared in Example DI, in which the homology between the homologous region and the substrate was reduced to 99-94%, 90%, 80%, and 70%. As shown in the figure, silent mutations were introduced scattered throughout the homologous sequence. 1 shows the results of comparing the survival rates between Nalm-6 cells and their MSH2-reverted cells 96 hours after introduction of a 2-split, 2-SSA type DTA construct. 2 shows the results of comparing the survival rates between HCT116 cells and their MLH1-reverted cells 72 hours after introduction of a 2-split, 2-SSA type DTA construct. 3 shows a diagram illustrating the structure and mechanism of a 3-split, 3-SSA type DTA construct (construct of the first embodiment) prepared in Example D-II. 4 shows a diagram illustrating the structure and mechanism of a 4-split, 3-SSA type DTA construct (construct of the first embodiment) prepared in Example D-II.

[0049] Figures 30 and 31 show the sequences of iDTA3-1 fragments prepared in Example D-II, in which the homology between the homologous region and the substrate was reduced to 98%, 94%, and 90% (the sequences in part A in Figures 30 and 31). As shown in the figures, silent mutations were introduced so as to be scattered throughout the homologous sequence. Figures 30 and 31 show the sequences of iDTA3-1 fragments and iDTA3-2 fragments prepared in Example D-II, in which the homology between the homologous region and the substrate was reduced to 98%, 94%, and 90% (the sequences in part B in Figures 30 and 31). As shown in the figures, silent mutations were introduced so as to be scattered throughout the homologous sequence. Figures 30 and 31 show the sequences of iDTA3-2 fragments prepared in Example D-II, in which the homology between the homologous region and the substrate was reduced to 98%, 94%, and 90% (the sequences in part C in Figures 30 and 31). As shown in the figures, silent mutations were introduced so as to be scattered throughout the homologous sequence. The left panel shows the results of comparing the viability of Nalm-6 cells and their MSH2-reverted cells 96 hours after transfection with a 3-split, 3-times SSA-type DTA construct. The right panel shows the results of comparing the viability of HCT116 cells and their MLH1-reverted cells 72 hours after transfection with the construct. The 4-split, 3-times SSA-type DTA construct was transfected into Nalm-6 cells and their MSH2-reverted cells, and the results are compared for viability (96 hours after transfection). This figure explains the structure and mechanism of the 4-split, 4-times SSA-type TK construct (construct of the first embodiment) prepared in Example EI. These are the sequences of the iTK3-1, iTK3-2, and iTK3-3 fragments prepared in Example EI, with homology between the homologous region and the substrate reduced to 98%, 94%, and 90% (sequences of parts A and D in Figure 35 ). As shown, silent mutations were introduced scattered throughout the homologous sequence. These are the sequences of the iTK3-1 fragment and the iTK3-2 fragment prepared in Example EI, in which the homology between the homologous region and the substrate was reduced to 98%, 94%, and 90% (the sequences shown in part B in Figure 35). As shown in the figure, silent mutations were introduced so as to be scattered throughout the homologous sequences. These are the sequences of the iTK3-2 fragment and the iTK3-3 fragment prepared in Example EI, in which the homology between the homologous region and the substrate was reduced to 98%, 94%, and 90% (the sequences shown in part C in Figure 35). As shown in the figure, silent mutations were introduced so as to be scattered throughout the homologous sequences.40 shows the results of comparing the survival rates of Nalm-6 cells and their MSH2-reverted cells 96 hours after transfection with a 4-split, 4-times SSA-type TK construct.

[0049]

[0100]

[0111]

[0112]

[0113]

[0114]

[0115]

[0116]

[0116]

[0117]

[0118]

[0119]

[0120]

[0121]

[0122]

[0123]

[0124]

[0125]

[0126]

[0127]

[0128]

[0129]

[0130]

[0131]

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143]

[0144]

[0145]

[0146]

[0147]

[0148]

[0149]

[0150]

[0151]

[0152]

[0153]

[0154]

[0155]

[0156]

[0157]

[0158]

[0159]

[0160]

[0161]

[0162]

[0163]

[0164]

[0165]

[0166]

[0167]

[0168]

[0169]

[0170]

[0171]

[0172] These are the sequences of the iTK3-1 and iTK303-2 fragments prepared in Example E-II, in which the homology between the homologous region and the substrate was reduced to 98%, 94%, and 90% (the sequences in part B in Figure 40). As shown in the figure, silent mutations were introduced scattered throughout the homologous sequences. These are the sequences of the iTK303-2 and iTK3-3 fragments prepared in Example E-II, in which the homology between the homologous region and the substrate was reduced to 98%, 94%, and 90% (the sequences in part C in Figure 40). As shown in the figure, silent mutations were introduced scattered throughout the homologous sequences. These are the results of comparing the survival rates of Nalm-6 cells and their MSH2-reverted cells 96 hours after transfection with the 4-fourth SSA-type TK30 construct. These are the results of comparing the survival rates of HCT116 cells and their MLH1-reverted cells 72 hours after transfection with the 4-fourth SSA-type TK30 construct.

[0049] Figure 4 shows the structure and mechanism of the 4-part, 4-SSA CDUPRT construct (construct of the first embodiment) prepared in Example E-III. Figure 4 shows the sequences of the iCDUPRTA3-1 fragment, iCDUPRTA3-2 fragment, and iCDUPRTA3-3 fragment prepared in Example E-III, in which the homology between the homologous region and the substrate was reduced to 98%, 94%, and 90%, respectively (sequences of portions A and D in Figure 44). As shown in the figure, silent mutations were introduced so as to be scattered throughout the homologous sequence.These are the sequences of the iCDUPRTA3-1 and iCDUPRTA3-2 fragments prepared in Example E-III, in which the homology between the homologous region and the substrate was reduced to 98%, 94%, and 90% (the sequences shown in part B in Figure 44). As shown in the figure, silent mutations were introduced scattered throughout the homologous sequences. These are the sequences of the iCDUPRTA3-2 and iCDUPRTA3-3 fragments prepared in Example E-III, in which the homology between the homologous region and the substrate was reduced to 98%, 94%, and 90% (the sequences shown in part C in Figure 44). As shown in the figure, silent mutations were introduced scattered throughout the homologous sequences. These are the results of comparing the viability of Nalm-6 cells and their MSH2-reverted cells 96 hours after transfection with a 4-fold SSA-type CDUPRT construct. These are the results of comparing the viability of HCT116 cells and their MLH1-reverted cells 72 hours after transfection with a 4-fold SSA-type CDUPRT construct.

[0019] In the present invention, the term "single-strand annealing recombination reaction" refers to a recombination reaction that occurs between homologous nucleotide sequences via a Rad51-independent pathway, which is different from Rad51-dependent standard homologous recombination. This recombination reaction does not require 100% homology between the homologous nucleotide sequences, as long as there is a certain level of homology. In particular, in the nucleic acid construct of the present invention, the homology between the homologous nucleotide sequences is 40% or more but less than 100%.

[0020] In the present invention, the unit of nucleic acid chain length is expressed as a "base", but when the nucleic acid is a double-stranded nucleic acid, the term "base" means a "base pair".

[0021] In the present invention, "homology" of a base sequence has the same meaning as "sequence identity" of a base sequence. The two base sequences to be compared are aligned so that as many bases as possible match, and the number of matched bases is divided by the total number of bases, expressed as a percentage. During the alignment, gaps may be inserted as needed into one or both of the two sequences to be compared. Such sequence alignment can be performed using well-known programs such as BLAST, FASTA, and CLUSTAL W. When gaps are inserted, the total number of bases is calculated by counting each gap as one base. If the total number of bases counted in this way differs between the two sequences to be compared, the homology (%) is calculated by dividing the number of matched bases by the total number of bases in the longer sequence. The "sequence identity" of amino acid sequences is calculated in a similar manner.

[0022] The nucleic acid construct of the present invention may be either DNA or RNA, but is preferably prepared using DNA from the viewpoint of the stability of the nucleic acid molecule. Furthermore, the nucleic acid construct of the present invention may partially contain a nucleotide analog. Examples of nucleotide analogs include, but are not limited to, cross-linked nucleic acids such as LNA, ENA, and PNA, and modified bases such as Super T (registered trademark) and Super G (registered trademark).

[0023] The nucleic acid construct of the present invention can be broadly divided into two embodiments. The nucleic acid construct of the first embodiment is a construct in which a gene sequence encoding a protein is formed and the protein can be expressed by a recombination reaction by single-strand annealing (SSA) (hereinafter sometimes abbreviated as "SSA reaction"). The nucleic acid construct of the second embodiment is a construct that originally contains a gene sequence encoding a protein, but which becomes operably linked to a promoter by an SSA reaction, allowing the gene sequence to express a protein. The first embodiment will be described below.

[0024] [Structure of the Nucleic Acid Construct of the First Aspect] The nucleic acid construct of the first aspect comprises a promoter region, a 5' region of a gene sequence encoding protein A (hereinafter simply referred to as the "5' region"), a 3' region of a gene sequence encoding protein A (hereinafter simply referred to as the "3' region"), and at least one complementary region, all located on a single nucleic acid molecule or on two or more different nucleic acid molecules. The promoter region and the 5' region are operably linked and contained on the same nucleic acid molecule, while the other elements may be located on the same nucleic acid molecule or on different nucleic acid molecules. Focusing on the location of the three elements other than the complementary region (promoter region, 5' region, and 3' region), the nucleic acid construct of the first aspect can be classified into two types: a construct in which these three elements are located on the same nucleic acid molecule, and a construct in which the promoter region and 5' region are located on one nucleic acid molecule and the 3' region is located on another nucleic acid molecule. The complementary region may be located on a nucleic acid molecule separate from the other elements, or on the same nucleic acid molecule as at least one of the other elements.

[0025] The 5' and 3' regions may be composed of two regions obtained by dividing the entire gene sequence encoding protein A, or may be composed of partial regions at both ends of the gene sequence lacking the intermediate region. In the latter case, one or more complementary regions contain a coding sequence that complements the intermediate region, and a gene sequence with the complemented intermediate region is generated by the SSA reaction.

[0026] The size of the 5' region must be set so that a protein fragment capable of exerting the activity of protein A is not expressed from the 5' region in cells lacking SSA activity. The size of the 5' region may be set appropriately depending on the type of protein A and is not particularly limited. Generally, the size of the 5' region can be set to about 60 / 100 or less of the full-length protein A, for example, 50 / 100 or less, 40 / 100 or less, 30 / 100 or less, or 25 / 100 or less. In a case where protein A is a fusion protein of protein X, which has the effect of reducing cell viability or emitting a signal, and protein Y, which has the function of enhancing the effect of protein X, and protein Y is fused to the C-terminus of protein X, the size of the 5' region can be set to about 60 / 100 or less of the protein X portion, for example, 50 / 100 or less, 40 / 100 or less, 30 / 100 or less, or 25 / 100 or less. Alternatively, the cleavage site can be located upstream of the nucleotide sequence encoding the amino acids that form the active center so that the 5' region is the region N-terminal of the active center of protein A (when protein A is a fusion protein in which protein Y is fused to the C-terminus of protein X, the active center of protein X) and is the 5' region, thereby preventing the protein A fragment consisting of the 5' region from exhibiting activity. There is no particular lower limit to the size of the 5' region, but it is usually set to a size of 50 bases or more.

[0027] A poly A addition signal may be operably linked downstream of the 3' region. "Operatively linked" here means that a poly A addition signal is located downstream of the 3' region so that an mRNA molecule having a poly A tail is expressed from the nucleotide sequence formed after the recombination reaction by single-stranded annealing, and no other elements, such as a complementary region, are located between the 3' region and the poly A addition signal. However, mRNAs that function properly without poly A, such as histone mRNAs, are known, and it is possible to omit the poly A addition signal from the nucleic acid construct of the present invention.

[0028] At least one complementary region includes one homologous region (homologous region α) that uses a region of the 5'-side region including at least the 3'-end portion as a substrate for the SSA reaction, and one homologous region (homologous region β) that uses a region of the 3'-side region including at least the 5'-end portion as a substrate for the SSA reaction.

[0029] Homologous region α and homologous region β are contained in the same or different complementary regions. In the case of a double SSA type construct in which a protein is expressed by two SSA reactions, there is one complementary region, and this one complementary region contains homologous region α and homologous region β. In the case of a triple or more SSA type construct in which a protein is expressed by three or more SSA reactions, there are two or more complementary regions, and homologous region α and homologous region β are contained in different complementary regions. The two or more complementary regions in a triple or more SSA type construct will contain, in addition to homologous region α and homologous region β, one or more pairs of further homologous regions and their substrates for causing an SSA reaction between the complementary regions.

[0030] The homology between each homologous region and its substrate is not particularly limited, as long as it is sufficient for single-stranded annealing to occur between the two, and is in the range of 40% to 100%. If the homology between the homologous region and the substrate is 100%, recombination by homologous recombination (HR) reaction will be more prevalent than SSA reaction, making it difficult to observe differences in protein expression due to the presence or absence of mismatch repair factors such as MSH2 (see, for example, A-1-3 in the Examples below). Therefore, in the constructs of the present invention, the homology between each homologous region and its substrate is less than 100%. The upper limit of homology may be, for example, 99% or less, 98% or less, 97% or less, 96% or less, or 95% or less, and the lower limit may be, for example, 45% or more, 50% or more, 55% or more, 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, or 90% or more.

[0031] In order to facilitate the SSA reaction, it is preferable that the mismatched bases between the homologous region and its substrate are scattered throughout the homologous region or its substrate. Generally, it is sufficient that there are no more than 10 consecutive mismatched bases. The number of consecutive mismatched bases in the homologous region or its substrate may be, for example, 9 or less, 8 or less, 7 or less, 6 or less, 5 or less, 4 or less, 3 or less, or 2 or less, and single mismatched bases may be scattered throughout.

[0032] A homologous region or substrate having mismatched bases is conveniently and preferably designed by introducing silent mutations based on the gene sequence encoding protein A, but it is also possible to introduce mutations that change the encoded amino acid sequence. In the latter case, the gene sequence generated by the SSA reaction may encode protein B having an amino acid sequence different from that of protein A, but this is acceptable as long as protein B has the same activity as protein A. Such protein B has a sequence identity with protein A of 80% or more, for example, 85% or more, 90% or more, 95% or more, or 98% or more.

[0033] Conservative substitutions, i.e., substitutions with amino acids with similar chemical properties, are likely to not impair the properties and activity of proteins. Amino acids with similar side chains have similar chemical properties. Amino acids can be grouped based on side chain similarity, for example, into groups with aliphatic side chains (glycine, alanine, valine, leucine, isoleucine), aliphatic hydroxyl side chains (serine, threonine), amide-containing side chains (asparagine, glutamine), aromatic side chains (phenylalanine, tyrosine, tryptophan), basic side chains (arginine, lysine, histidine), acidic side chains (aspartic acid, glutamic acid), and sulfur-containing side chains (cysteine, methionine). Substitutions with other amino acids from the same group are considered conservative substitutions. If the mismatched base is not a silent mutation, a mutation resulting in a conservative amino acid substitution is preferably selected.

[0034] The length of each homologous region and its substrate may be sufficient to allow SSA-mediated recombination in cells, for example, 4 or more bases, 10 or more bases, or 15 or more bases, or 20 or more bases, 30 or more bases, 40 or more bases, or 50 or more bases. Constructs that generate SSA with a 20-base homologous region are known (WO 2021 / 162121 A1). The upper limit is not particularly limited, but is usually 10,000 bases or less, and can be, for example, 5,000 bases or less, 1,000 bases or less, 700 bases or less, 500 bases or less, or 200 bases or less.

[0035] The promoter is not particularly limited, and may be any promoter that can exert promoter activity in cells (typically mammalian cells such as human cells). Generally, promoters that constitutively exert strong promoter activity in cells into which the nucleic acid construct of the present invention is introduced are preferably used, but inducible promoters that exert promoter activity under certain conditions may also be used. In the examples below, the human cytomegalovirus promoter, which constitutively exerts strong promoter activity in cells, is used, but this is not limiting. Promoters are not limited to virus-derived promoters, and any sequence that exhibits promoter activity in cells, preferably human cells, may be used, regardless of origin.

[0036] The gene sequence encoding protein A or protein B (hereinafter sometimes abbreviated as protein A / B) produced by the SSA reaction may be a sequence containing the full-length coding region of a naturally occurring gene, or may be a sequence consisting of a portion of the coding region or a modified sequence. For example, a naturally occurring gene sequence may be used in the present invention that encodes only the region encoding the domain necessary for protein activity (e.g., a sequence encoding a protein fragment from which domains not necessary for the protein's activity, such as a domain required for membrane localization, have been removed). Furthermore, variously modified genes for fluorescent proteins and luminescent proteins are widely used commercially, and gene sequences encoding such non-natural proteins may also be used. Furthermore, a gene sequence encoding a fusion protein in which two or more proteins are fused may also be used as the gene sequence encoding protein A / B. In other words, protein A / B may be a fusion protein. A typical example of a gene sequence encoding a fusion protein in the present invention is a gene sequence encoding a fusion protein in which a protein whose activity can be measured directly or indirectly using, for example, a phenotypic change caused in cells by the activity of the protein as an indicator is fused with a protein that has the effect of enhancing the activity of the protein.

[0037] Protein A / B is preferably a protein whose expression level can be measured quickly and easily. For example, rather than measuring the amount of protein produced or accumulated, it is preferable to use as protein A / B a protein whose activity can be measured directly or indirectly using as an indicator, for example, a phenotypic change caused in cells by the activity of the protein. Examples of such proteins include proteins whose expression in cells can be detected by signals such as luminescence from a reaction with a substrate or fluorescence from the protein itself. Another example is a protein that has the effect of reducing cell viability, in which case protein activity can be measured indirectly as a reduction in cell viability.

[0038] The term "measuring the expression (expression level) of protein A / B" includes measuring the production or accumulation levels of protein A / B and measuring the activity of protein A / B. As described above, in the present invention, it is preferable to employ as protein A / B a protein whose expression can be measured by directly or indirectly measuring the activity of protein A / B, i.e., a protein whose activity can be measured directly or indirectly.

[0039] Specific examples of genes encoding proteins that have the effect of reducing cell viability include, but are not limited to, suicide genes, DNA damage-inducing genes, and DNA repair inhibitor genes.

[0040] Suicide genes include genes that encode proteins that are themselves toxic or damaging to cells, genes that encode proteins that act on other compounds to produce toxic substances, and genes that encode natural or non-natural proteins that exert cytotoxicity by inhibiting endogenous proteins through a dominant-negative effect. Specific examples of suicide genes include the diphtheria toxin A fragment gene (DT-A) (the DT-A protein itself is cytotoxic and kills cells), the herpes virus-derived thymidine kinase gene (HSV-tk) (the HSV-tk protein acts on ganciclovir or 5-iodo-2'-fluoro-2'deoxy-1-beta-D-arabino-furanosyl-uracil (FIAU) to convert it into a toxic substance with DNA synthesis inhibitory activity; when this gene is used as a suicide gene, cells must be treated with ganciclovir or FIAU), the p53 gene (overexpression induces cell growth cycle arrest or apoptosis, resulting in cell death), and the Fcy1 gene from budding yeast (encoding cytosine deaminase (CD). CD deaminates 5-fluorocytosine (5-FC) to convert it into 5-fluorouracil. 5-Fluorouracil is metabolized within the cell and converted into a substance that inhibits DNA synthesis and RNA synthesis, thereby inducing cell death (Austin EA and Huber BE. Mol Pharmacol. 1993 Mar; 43(3): 380-387.), and variants of these genes with enhanced activity. Specific examples of variants include the TK30 gene, which is a mutant of the HSV-TK gene (see Example E-II below), and the CD::UPRT gene, which is a fusion gene of the Fcy1 gene and the Fur1 gene derived from budding yeast (see Example E-III below).The TK30 gene is a mutant of the HSV-TK gene that has been modified by introducing mutations into amino acids near the active site of the HSV-TK gene (specifically, substitutions of alanine 152 with valine, leucine 159 with isoleucine, isoleucine 160 with leucine, phenylalanine 161 with alanine, alanine 168 with tyrosine, and leucine 169 with phenylalanine) to enhance its activity against ganciclovir. This gene is known to exhibit stronger cytotoxicity than wild-type HSV-TK (Kokoris MS et al. Gene Ther. 1999, Aug;6(8):1415-1426.). The Fur1 gene used for the CD::UPRT gene encodes uracil phosphoribosyltransferase (UPRT). UPRT is known to catalyze the metabolic pathway by which 5-fluorouracil is converted into a toxic substance within cells, and its use in combination with CD enhances the cytotoxicity of 5-fluorocytosine and provides a bystander effect (Tiraby M et al. FEMS Microbiol Lett. 1998 Oct 1;167(1): 41-49., Bourbeau D et al. J Gene Med. 2004 Dec;6(12):1320-1332.).In Example E-III below, the CD::UPRT gene is a fusion gene formed by linking the full-length codon-optimized nucleotide sequence (SEQ ID NO: 229) encoding CD (NP_015387, SEQ ID NO: 230) with the nearly full-length codon-optimized nucleotide sequence (SEQ ID NO: 231) encoding UPRT (NP_011996, SEQ ID NO: 232) via an alanine codon, thereby encoding a fusion protein (SEQ ID NO: 186) in which the full-length CD and the sequence from residue 3 to the C-terminal residue of UPRT are linked via a single alanine residue. However, the CD::UPRT gene is not limited to this. A known Fcy1 gene sequence (e.g., NM_001184159 (SEQ ID NO: 233)) or Fur1 gene sequence (e.g., NM_001179258 (SEQ ID NO: 234)) can be codon-optimized as appropriate depending on the animal species of the cells into which the construct is to be introduced, and then linked via a linker residue, if desired, to freely design a CD::UPRT gene. Further specific examples of suicide genes include natural nucleases such as restriction enzymes and meganucleases, and artificial nucleases such as ZFN, TALEN, and CRISPR systems (these are also specific examples of DNA damage-inducible genes). However, suicide genes are not limited to these specific examples.

[0041] Genes encoding proteins whose intracellular expression can be detected include, but are not limited to, luciferase genes (including genes encoding secreted luciferases), fluorescent protein genes (genes encoding natural or artificial fluorescent proteins, including GFP, CFP, OFP, RFP, YFP, and modified versions thereof), cell surface antigen genes, secreted protein genes, and membrane protein genes. Luciferase genes and fluorescent protein genes encode proteins whose intracellular expression can be detected as signals such as luminescence or fluorescence, and are preferred examples of genes that can be used in the present invention. Similar to suicide genes, modified versions of the above-mentioned genes with enhanced activity can also be used. A specific example is the gene encoding nanolantern (NL), which combines the chemiluminescent protein Nluc with a fluorescent protein. It is known that NL enhances luminescence intensity through Forester resonance energy transfer (FRET) between Nluc and the fluorescent protein (Suzuki K et al. Nat Commun. 2016, Dec 14;7:13718). To date, NLs using fluorescent proteins such as cyan, green, yellow-green, orange, and red have been reported, but NLs can also be designed using other fluorescent proteins.

[0042] Typical examples of the nucleic acid construct of the first aspect include the following: (1) A construct in which the promoter region, 5' region, and 3' region are located on the same nucleic acid molecule, and at least one complementary region is entirely contained in a separate nucleic acid molecule; (2) A construct in which the promoter region and the 5' region are located on one nucleic acid molecule, and the 3' region is located on another nucleic acid molecule, and at least one complementary region is entirely contained in a separate nucleic acid molecule; and (3) A nucleic acid construct in which the promoter region, 5' region, 3' region, and all complementary regions are located on the same nucleic acid molecule.

[0043] The nucleic acid construct of the first aspect also includes the following constructs, which are variants of (1) to (3): (4) (Variant of (1)(3)) A portion of at least one complementary region is located on the same nucleic acid molecule as the promoter region, etc., and the remaining complementary regions are located on one or more nucleic acid molecules other than the promoter region. (5) (Variant of (2)) At least one complementary region is located on either or both of a nucleic acid molecule including a 5' region and a nucleic acid molecule including a 3' region. (6) (Variant of (2)) A portion of at least one complementary region is located on either or both of a nucleic acid molecule including a 5' region and a nucleic acid molecule including a 3' region, and the remaining complementary regions are located on one or more nucleic acid molecules other than the promoter region.

[0044] Among the nucleic acid constructs of (1), a construct consisting of two nucleic acid molecules and including one complementary region has the following configuration: nucleic acid molecule 1 includes a promoter region, a 5' region, and a 3' region; nucleic acid molecule 2 includes a complementary region including homologous region α and homologous region β; the homology between homologous region α and a region of the 5' region that is its substrate in the SSA reaction and that includes at least the 3'-end portion is 40% or more but less than 100%; the homology between homologous region β and a region of the 3' region that is its substrate in the SSA reaction and that includes at least the 5'-end portion is 40% or more but less than 100%; when arranged in this order, the 5' region, complementary region, and 3' region encode the amino acid sequences of proteins A / B with overlap between each homologous region and its substrate, and proteins A / B are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0045] Among the nucleic acid constructs of (1), a construct consisting of three or more nucleic acid molecules and containing two or more complementary regions can be expressed as follows, using an integer n that is a constant in the range of 3≦n≦10 and an integer m that is a variable in the range of 3≦m≦n: n is preferably 9 or less, 8 or less, 7 or less, 6 or less, 5 or less, or 4 or less.

[0046] A nucleic acid construct is composed of n nucleic acid molecules and includes (n-1) complementary regions, wherein nucleic acid molecule 1 includes a promoter region, a 5' region, and a 3' region; nucleic acid molecule 2 includes a first complementary region including a first homologous region and a second homologous region; and mth nucleic acid molecule m includes an (m-1)th substrate region and an (m-1)th complementary region including the mth homologous region. The first homologous region is the homologous region α described above, and the homology between this homologous region and a region including at least the 3'-end portion of the 5' region, which serves as a substrate for the SSA reaction, is 40% or more but less than 100%. The (m-1)th homologous region uses the (m-1)th substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more but less than 100%. The nth homologous region in nucleic acid molecule n is the aforementioned homologous region β, and the homology between this homologous region and a region of the 3'-side region that is its substrate in an SSA reaction and includes at least the 5'-end portion is 40% or more but less than 100%. When the 5'-side region, the first complementary region, the (m-1)th complementary region, and the 3'-side region are arranged in this order (the complementary regions are arranged in ascending order of their numbers), each homologous region and its substrate overlap with each other to encode the amino acid sequence of protein A / B, and protein A / B is expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0047] For example, when n=3, a nucleic acid construct (1) consisting of three nucleic acid molecules (nucleic acid molecules 1 to 3) and containing two complementary regions has the above-mentioned configuration, in which: nucleic acid molecule 3 contains a second complementary region containing a second substrate region and a third homologous region; the second homologous region uses the second substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%; the third homologous region is the above-mentioned homologous region β, and the homology between this homologous region and a region including at least the 5'-end portion of the 3'-side region, which is its substrate in the SSA reaction, is 40% or more and less than 100%; when arranged in this order, the 5'-side region, first complementary region, second complementary region, and 3'-side region encode the amino acid sequences of proteins A / B, with each homologous region and its substrate overlapping each other; and proteins A / B are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0048] When n=4, the nucleic acid construct (1) is composed of four nucleic acid molecules (nucleic acid molecules 1 to 4) and includes three complementary regions, and in the above-mentioned configuration, nucleic acid molecule 3 includes a second complementary region including a second substrate region and a third homologous region, nucleic acid molecule 4 includes a third complementary region including a third substrate region and a fourth homologous region, the second homologous region uses the second substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%, the third homologous region uses the third substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%, the fourth homologous region is the above-mentioned homologous region β, and the homology between the homologous region and a region including at least the 5'-end portion of the 3'-side region, which is the substrate for the SSA reaction, is 40% or more and less than 100%, When the 5' region, first complementary region, second complementary region, third complementary region, and 3' region are arranged in this order, they encode the amino acid sequences of proteins A and B with overlapping between each homologous region and its substrate, and proteins A and B are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0049] In the nucleic acid construct (1), nucleic acid molecule 1 may be a linear nucleic acid molecule comprising a promoter region and a 5' region downstream of the 3' region in this order, or a circular nucleic acid molecule comprising a 5' region and a 3' region downstream of the promoter region in this order. When a poly(A) addition signal is included, nucleic acid molecule 1 may be a linear nucleic acid molecule comprising a poly(A) addition signal, a promoter region, and a 5' region downstream of the 3' region in this order, or a circular nucleic acid molecule comprising a 5' region, a 3' region, and a poly(A) addition signal downstream of the promoter region in this order.

[0050] It is preferable that all nucleic acid molecules containing a complementary region other than nucleic acid molecule 1 are linear molecules for ease of production and use.

[0051] When nucleic acid molecule 1 is a circular molecule, it can be introduced into cells as a circular molecule and used. However, it is usually cleaved between the 5' and 3' regions and linearized before being introduced into cells. Therefore, it is preferable to have a cleavage site between the 5' and 3' regions. Typical examples of cleavage sites include, but are not limited to, restriction enzyme recognition sites. In addition to cleavage by restriction enzymes, cleavage can also be achieved using genome editing techniques that cause DNA strand cleavage. A specific example is cleavage by the CRISPR / Cas system, i.e., cleavage by a complex of guide RNA and a Cas protein such as Cas9. When using this complex, a PAM sequence is introduced at an appropriate position between the 5' and 3' regions. In the case of Cas9 derived from Streptococcus pyogenes type II, the most commonly used known CRISPR / Cas system, the PAM sequence is 5'-NGG (N is A, T, G, or C). The complex cleaves the nucleic acid construct several bases upstream of the PAM sequence. In the case of cleavage by the CRISPR / Cas system, the PAM sequence and the actual cleavage site located several bases upstream thereof constitute the cleavage site in the present invention. A stop codon may be added to the 3' end of the 5' region, which reliably prevents expression of a protein having protein A activity from the [5' region] + [cleavage site (optional)] + [3' region] portion of the circular molecule in cells lacking SSA activity. However, since a double-stranded break is required for the SSA reaction, and cleavage typically occurs between the 5' and 3' regions when the nucleic acid construct of the present invention containing the circular molecule is introduced into cells for purposes such as evaluating SSA activity, adding a stop codon to the 5' region is not essential.

[0052] Among the nucleic acid constructs of (2), a construct consisting of three nucleic acid molecules and including one complementary region has the following configuration: Nucleic acid molecule 1-1 includes a promoter region and a 5' region. Nucleic acid molecule 1-2 includes a 3' region. Nucleic acid molecule 2 includes a complementary region including homologous region α and homologous region β. The homology between homologous region α and a region of the 5' region that is its substrate in the SSA reaction and that includes at least the 3'-end portion is 40% or more and less than 100%. The homology between homologous region β and a region of the 3' region that is its substrate in the SSA reaction and that includes at least the 5'-end portion is 40% or more and less than 100%. When arranged in this order, the 5' region, complementary region, and 3' region encode the amino acid sequences of proteins A / B with overlap between each homologous region and its substrate, and proteins A / B are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0053] Among the nucleic acid constructs of (2), constructs consisting of four or more nucleic acid molecules and containing two or more complementary regions can be expressed as follows, using an integer n that is a constant in the range of 3≦n≦10 and an integer m that is a variable in the range of 3≦m≦n: n is preferably 9 or less, 8 or less, 7 or less, 6 or less, 5 or less, or 4 or less.

[0054] A nucleic acid construct is composed of (n+1) nucleic acid molecules and includes (n-1) complementary regions, wherein nucleic acid molecule 1-1 includes a promoter region and a 5'-side region, nucleic acid molecule 1-2 includes a 3'-side region, nucleic acid molecule 2 includes a first complementary region including a first homologous region and a second homologous region, and the mth nucleic acid molecule m includes an (m-1)th substrate region and an (m-1)th complementary region including the mth homologous region. The first homologous region is the homologous region α described above, and the homology between this homologous region and a region including at least the 3'-end portion of the 5'-side region, which serves as a substrate for the SSA reaction, is 40% or more but less than 100%. The (m-1)th homologous region uses the (m-1)th substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more but less than 100%. The nth homologous region in nucleic acid molecule n is the aforementioned homologous region β, and the homology between this homologous region and a region of the 3'-side region that is its substrate in an SSA reaction and includes at least the 5'-end portion is 40% or more but less than 100%. When the 5'-side region, the first complementary region, the (m-1)th complementary region, and the 3'-side region are arranged in this order (the complementary regions are arranged in ascending order of their numbers), each homologous region and its substrate overlap with each other to encode the amino acid sequence of protein A / B, and protein A / B is expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0055] For example, when n=3, a nucleic acid construct (2) consisting of four nucleic acid molecules (nucleic acid molecules 1-1, 1-2, 2, and 3) and containing two complementary regions has the above-mentioned configuration, in which: nucleic acid molecule 3 contains a second complementary region containing a second substrate region and a third homologous region; the second homologous region uses the second substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%; the third homologous region is the above-mentioned homologous region β, and the homology between this homologous region and a region including at least the 5'-end portion of the 3'-side region, which is its substrate in the SSA reaction, is 40% or more and less than 100%; when arranged in this order, the 5'-side region, first complementary region, second complementary region, and 3'-side region encode the amino acid sequences of proteins A / B, with each homologous region and its substrate overlapping each other; and proteins A / B are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0056] When n=4, the nucleic acid construct (2) is composed of five nucleic acid molecules (nucleic acid molecules 1-1, 1-2, 2, 3, and 4) and includes three complementary regions, and in the above-mentioned configuration, nucleic acid molecule 3 includes a second complementary region including a second substrate region and a third homologous region, nucleic acid molecule 4 includes a third complementary region including a third substrate region and a fourth homologous region, the second homologous region uses the second substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%, the third homologous region uses the third substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%, the fourth homologous region is the above-mentioned homologous region β, and the homology between the homologous region and a region including at least the 5'-end portion of the 3'-side region, which is the substrate for the SSA reaction, is 40% or more and less than 100%, When the 5' region, first complementary region, second complementary region, third complementary region, and 3' region are arranged in this order, they encode the amino acid sequences of proteins A and B with overlapping between each homologous region and its substrate, and proteins A and B are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0057] It is preferable that all of the nucleic acid molecules constituting the nucleic acid construct (2) are prepared as linear nucleic acid molecules for ease of production and use, but some or all of them may be prepared as circular molecules. When the nucleic acid construct (2) contains a poly(A) addition signal, nucleic acid molecule 1-2 containing the 3'-region contains a poly(A) addition signal downstream of the 3'-region.

[0058] The nucleic acid construct (3) is composed of a single circular or linear nucleic acid molecule. A construct (3) composed of a single circular nucleic acid molecule and containing a single complementary region has the following configuration: The circular nucleic acid molecule contains a 5' region, a cleavage site, and a 3' region, in this order, downstream of the promoter region. The complementary region, including homologous region α and homologous region β, is contained upstream of the promoter region and downstream of the 3' region. In other words, on a single circular nucleic acid molecule, each element is arranged from upstream to downstream in the following order: [promoter]-[5' region]-[cleavage site]-[3' region]-[complementary region]. The complementary region may be arranged with homologous region α facing upstream or homologous region β facing upstream. When a poly(A) addition signal is included, the poly(A) addition signal is arranged operably linked to the 3' region, and the complementary region is therefore arranged downstream of the poly(A) addition signal. The homology between a region of the 5'-side region comprising at least the 3'-end portion and homologous region α, which uses this region as a substrate for the SSA reaction, is 40% or more but less than 100%. The homology between a region of the 3'-side region comprising at least the 5'-end portion and homologous region β, which uses this region as a substrate for the SSA reaction, is 40% or more but less than 100%. When the 5'-side region, complementary region, and 3'-side region are arranged in this order with the complementary region oriented from complementary region α to complementary region β, they encode the amino acid sequences of proteins A / B with overlap between each homologous region and its substrate, and proteins A / B are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0059] The cleavage site is as explained in (1) nucleic acid construct, and includes a restriction enzyme recognition site and the like.

[0060] In constructs in which the complementary region is contained on the same nucleic acid molecule as other elements, it is preferable to excise the complementary region from the nucleic acid molecule and use it as a separate nucleic acid molecule when introducing it into a cell. Therefore, in the nucleic acid construct of (3), it is preferable to locate cleavage sites on both sides of the complementary region on the nucleic acid molecule. If the cleavage sites between the 5' and 3' regions and these cleavage sites are all the same type, all cleavage sites can be easily cleaved in a single cleavage treatment (e.g., by treatment with a single type of restriction enzyme).

[0061] Among the nucleic acid constructs (3), a construct consisting of one circular nucleic acid molecule and containing two or more complementary regions can be expressed as follows, using an integer n that is a constant in the range of 3≦n≦10 and an integer m that is a variable in the range of 3≦m≦n: n is preferably 9 or less, 8 or less, 7 or less, 6 or less, 5 or less, or 4 or less.

[0062] A nucleic acid construct is composed of one circular nucleic acid molecule and contains (n-1) complementary regions, and has the following configuration: the circular nucleic acid molecule contains a 5' region, a cleavage site, and a 3' region downstream of the promoter region, in that order, and (n-1) complementary regions upstream of the promoter region and downstream of the 3' region. When there are two or more complementary regions, the order of their arrangement is not particularly limited, and they may be in numerical order or in any order. The orientation of the complementary regions is also not particularly limited. The first complementary region contains a first homologous region and a second homologous region, and the (m-1)th complementary region contains the (m-1)th substrate region and the mth homologous region. The first homologous region is homologous region α, which uses a region of the 5' region including at least the 3'-end portion as a substrate for the SSA reaction, and the homology between the substrate and the first homologous region is 40% or more but less than 100%. The (m-1)th homologous region uses the (m-1)th substrate region as a substrate for the SSA reaction, and the homology between them is 40% or more but less than 100%. The (n-1)th homologous region in the (n-1)th complementary region is a homologous region β that uses a region of the 3'-side region including at least the 5'-end portion as a substrate for the SSA reaction, and the homology between the substrate and the nth homologous region is 40% or more but less than 100%. When the 5'-side region, the first complementary region, the (m-1)th complementary region, and the 3'-side region are arranged in this order (the complementary regions are arranged in ascending order of number), the homologous regions and their substrates overlap with each other to encode the amino acid sequences of proteins A / B, and proteins A / B are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0063] For example, when n=3, a nucleic acid construct (3) consisting of one circular nucleic acid molecule and containing two complementary regions has the above-mentioned configuration, in which: the circular nucleic acid molecule contains two complementary regions, i.e., a first and a second complementary region, located upstream of the promoter region and downstream of the 3'-side region; the second complementary region contains a second substrate region and a third homologous region; the second homologous region uses the second substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%; the third homologous region is the above-mentioned homologous region β, and the homology between the homologous region and a region of the 3'-side region that is its substrate for the SSA reaction and includes at least the 5'-end portion is 40% or more and less than 100%; When the 5' region, first complementary region, second complementary region, and 3' region are arranged in this order, they encode the amino acid sequences of proteins A and B, with overlapping between each homologous region and its substrate, and proteins A and B are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0064] When n=4, the nucleic acid construct (3) is composed of one circular nucleic acid molecule and contains three complementary regions, and in the above-mentioned configuration, the circular nucleic acid molecule contains three complementary regions, i.e., the first, second, and third complementary regions, located upstream of the promoter region and downstream of the 3'-side region, the second complementary region contains a second substrate region and a third homologous region, the third complementary region contains a third substrate region and a fourth homologous region, the second homologous region uses the second substrate region as a substrate for the SSA reaction, and the homology between them is 40% or more and less than 100%, the third homologous region uses the third substrate region as a substrate for the SSA reaction, and the homology between them is 40% or more and less than 100%, the fourth homologous region is the above-mentioned homologous region β, and the homology between said homologous region and a region of the 3'-side region that is its substrate in the SSA reaction and includes at least the 5'-end portion, is 40% or more and less than 100%, When the 5' region, first complementary region, second complementary region, third complementary region, and 3' region are arranged in this order, they encode the amino acid sequences of proteins A and B with overlapping between each homologous region and its substrate, and proteins A and B are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0065] A construct (3) composed of one linear nucleic acid molecule and containing one complementary region has the following configuration: the linear nucleic acid molecule contains, from upstream to downstream, the 3' region, complementary region, promoter region, and 5' region, in this order. The complementary region includes homologous region α and homologous region β. The orientation of the complementary region is not limited, as with the construct (3) composed of a circular nucleic acid molecule. When a poly(A) addition signal is contained, the complementary region is located downstream of the 3' region + poly(A) addition signal. The homology between a region of the 5' region comprising at least the 3'-end portion and homologous region α, which uses this as a substrate for the SSA reaction, is 40% or more but less than 100%. The homology between a region of the 3' region comprising at least the 5'-end portion and homologous region β, which uses this as a substrate for the SSA reaction, is 40% or more but less than 100%. When the 5' region, complementary region, and 3' region are arranged in this order with the complementary region oriented from complementary region α to complementary region β, they encode the amino acid sequences of proteins A and B with overlapping regions between each homologous region and its substrate. Proteins A and B are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0066] As with the construct (3) consisting of a circular molecule, it is preferable to place cleavage sites on both sides of the complementary region.

[0067] Among the nucleic acid constructs of (3), a construct consisting of one linear nucleic acid molecule and containing two or more complementary regions can be expressed as follows, using an integer n that is a constant in the range of 3≦n≦10 and an integer m that is a variable in the range of 3≦m≦n: n is preferably 9 or less, 8 or less, 7 or less, 6 or less, 5 or less, or 4 or less.

[0068] The nucleic acid construct is composed of a single linear molecule and contains (n-1) complementary regions, and has the following configuration: the linear nucleic acid molecule contains, from upstream to downstream, a 3' region, a promoter region, and a 5' region, in this order, and contains (n-1) complementary regions between the 3' region and the promoter region. As with the circular molecule, the order and orientation of the complementary regions are not limited. The first complementary region contains a first homologous region and a second homologous region, and the (m-1)th complementary region contains an (m-1)th substrate region and an mth homologous region. The first homologous region is homologous region α, which uses a region of the 5' region including at least the 3'-end portion as a substrate for the SSA reaction, and the homology between the substrate and the first homologous region is 40% or more but less than 100%. The (m-1)th homologous region uses the (m-1)th substrate region as a substrate for the SSA reaction, and the homology between them is 40% or more but less than 100%. The (n-1)th homologous region in the (n-1)th complementary region is a homologous region β that uses a region of the 3'-side region including at least the 5'-end portion as a substrate for the SSA reaction, and the homology between the substrate and the nth homologous region is 40% or more but less than 100%. When the 5'-side region, the first complementary region, the (m-1)th complementary region, and the 3'-side region are arranged in this order (the complementary regions are arranged in ascending order of number), the homologous regions and their substrates overlap with each other to encode the amino acid sequences of proteins A / B, and proteins A / B are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0069] For example, when n=3, a nucleic acid construct (3) consisting of one linear nucleic acid molecule and containing two complementary regions has the above-mentioned configuration, in which: the linear nucleic acid molecule contains two complementary regions, i.e., a first and a second complementary region, between the 3'-side region and the promoter region; the second complementary region contains a second substrate region and a third homologous region; the second homologous region uses the second substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%; the third homologous region is the above-mentioned homologous region β, and the homology between the homologous region and a region of the 3'-side region that is its substrate in the SSA reaction and includes at least the 5'-end portion is 40% or more and less than 100%; When the 5' region, first complementary region, second complementary region, and 3' region are arranged in this order, they encode the amino acid sequences of proteins A and B, with overlapping between each homologous region and its substrate, and proteins A and B are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0070] When n=4, the nucleic acid construct (3) which is composed of one linear nucleic acid molecule and contains three complementary regions has the above-mentioned configuration, wherein the linear nucleic acid molecule contains three complementary regions, namely, the first, second and third complementary regions, between the 3'-side region and the promoter region, the second complementary region contains a second substrate region and a third homologous region, the third complementary region contains a third substrate region and a fourth homologous region, the second homologous region uses the second substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%, the third homologous region uses the third substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%, the fourth homologous region is the above-mentioned homologous region β, and the homology between the homologous region and a region of the 3'-side region which is the substrate for the SSA reaction and includes at least the 5'-end portion is 40% or more and less than 100%, When the 5' region, first complementary region, second complementary region, third complementary region, and 3' region are arranged in this order, they encode the amino acid sequences of proteins A and B with overlapping between each homologous region and its substrate, and proteins A and B are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0071] Nucleic acid constructs (4) to (6), which are variants of (1) to (3), are possible configurations for constructs having two or more complementary regions. Specific examples of (4) to (6) for constructs having up to four complementary regions are given below. All of these can be designed in accordance with the nucleic acid constructs (1) to (3).

[0072] Examples of the configuration of the nucleic acid construct (4), which is a variant of (1) and (3): There are two complementary regions, and one complementary region is located on the same nucleic acid molecule (nucleic acid molecule 1) as the promoter region, etc., and the remaining complementary region is located on another nucleic acid molecule. There are three complementary regions, and two complementary regions are located on nucleic acid molecule 1 and the remaining complementary region is located on another nucleic acid molecule. There are three complementary regions, and one complementary region is located on nucleic acid molecule 1 and the remaining two complementary regions are (i) located on one nucleic acid molecule, or (ii) located separately on two nucleic acid molecules. There are four complementary regions, and one complementary region is located on nucleic acid molecule 1 and the remaining three complementary regions are (i) located on one nucleic acid molecule, (ii) located separately on two nucleic acid molecules, or (iii) located one each on three nucleic acid molecules. There are four complementary regions, two of which are located on nucleic acid molecule 1 and the remaining two complementary regions are either (i) located on one nucleic acid molecule, or (ii) separated onto two nucleic acid molecules. There are four complementary regions, three of which are located on nucleic acid molecule 1 and the remaining one complementary region is located on another nucleic acid molecule.

[0073] An example of the configuration of the nucleic acid construct (5), which is a variant of (2): There are two complementary regions, (i) one complementary region is located in either the nucleic acid molecule containing the 5' region (nucleic acid molecule 1-1) or the nucleic acid molecule containing the 3' region (nucleic acid molecule 1-2), and one complementary region is located in the other, or (ii) all of the two complementary regions are located in either nucleic acid molecule 1-1 or nucleic acid molecule 1-2. There are three complementary regions, (i) one complementary region is located in either nucleic acid molecule 1-1 or nucleic acid molecule 1-2, and two complementary regions are located in the other, or (ii) all of the three complementary regions are located in either nucleic acid molecule 1-1 or nucleic acid molecule 1-2. There are four complementary regions, and (i) one complementary region is located in either nucleic acid molecule 1-1 or nucleic acid molecule 1-2, and three complementary regions are located in the other; (ii) two complementary regions are located in each of nucleic acid molecule 1-1 and nucleic acid molecule 1-2; or (iii) all four complementary regions are located in either nucleic acid molecule 1-1 or nucleic acid molecule 1-2.

[0074] An example of the configuration of the nucleic acid construct (6), which is a variant of (2): There are two complementary regions, one of which is located on either nucleic acid molecule 1-1 or nucleic acid molecule 1-2, and one of which is located on a nucleic acid molecule other than nucleic acid molecule 1-1 and nucleic acid molecule 1-2. There are three complementary regions, one of which is located on either nucleic acid molecule 1-1 or nucleic acid molecule 1-2, and the remaining two complementary regions are located on a nucleic acid molecule other than nucleic acid molecule 1-1 and nucleic acid molecule 1-2, or are located on separate nucleic acid molecules. There are three complementary regions, two of which are located (i) on nucleic acid molecule 1-1, (ii) on nucleic acid molecule 1-2, or (iii) one each on nucleic acid molecules 1-1 and 1-2, and the remaining one complementary region is located on a nucleic acid molecule other than nucleic acid molecules 1-1 and 1-2. There are four complementary regions, one of which is located on either nucleic acid molecule 1-1 or nucleic acid molecule 1-2, and the remaining three complementary regions are (i) located on a single nucleic acid molecule other than nucleic acid molecules 1-1 and 1-2, (ii) located separately on two nucleic acid molecules other than nucleic acid molecules 1-1 and 1-2, or (iii) located one each on three nucleic acid molecules other than nucleic acid molecules 1-1 and 1-2. There are four complementary regions, two of which are located (i) on nucleic acid molecule 1-1, (ii) on nucleic acid molecule 1-2, or (iii) on nucleic acid molecules 1-1 and 1-2, and the remaining two complementary regions are located on a single nucleic acid molecule other than nucleic acid molecule 1-1 and nucleic acid molecule 1-2, or located on separate nucleic acid molecules. There are four complementary regions, three of which are located (i) on nucleic acid molecule 1-1, (ii) on nucleic acid molecule 1-2, or (iii) separately on nucleic acid molecules 1-1 and 1-2, and the remaining complementary region is located on a nucleic acid molecule other than nucleic acid molecules 1-1 and 1-2.

[0075] The 5' region, 3' region, and complementary region can be prepared by PCR amplification using appropriate primers from a known plasmid incorporating a gene encoding protein A, or from cultured cells or a cDNA library of an organism that expresses protein A. Nucleic acid molecules with reduced homology between the homologous region and the substrate can be prepared by PCR using primers containing mutations. Alternatively, nucleic acid molecules with specified sequences can be prepared by well-known chemical synthesis methods. Other essential and optional elements, such as promoter regions, can be prepared from known plasmids.

[0076] Circular nucleic acid molecules can be prepared by incorporating the essential elements of the present invention and, optionally, optional elements into a plasmid vector. Linear nucleic acid molecules can be prepared entirely by chemical synthesis or artificial gene synthesis, or by linking the essential elements of the present invention and, optionally, optional elements using fusion PCR or other methods. Alternatively, they can be prepared as circular nucleic acid molecules and then cleaved at an appropriate site (e.g., a cleavage site located in the circular molecule). Furthermore, the nucleic acid construct of the present invention may be prepared by incorporating the essential elements and, optionally, optional elements into a viral vector. The nucleic acid construct of the present invention incorporated into a vector may contain, in addition to the elements of the present invention described above, common elements depending on the type of vector, such as a replication origin, a translation origin, and a selectable marker gene for drug resistance or the like. The same applies to the circular nucleic acid molecule and linear nucleic acid molecule in the nucleic acid construct of the second aspect.

[0077] [Various uses of the nucleic acid construct of the first embodiment] The nucleic acid construct of the first embodiment can be preferably used as a therapeutic or diagnostic agent for mismatch repair-deficient cancer, or as a companion diagnostic agent for predicting the effect of an anticancer drug on mismatch repair-deficient cancer. In the following description of various uses of the construct of the first embodiment, the protein encoded by the gene sequence formed by the SSA reaction is referred to as "protein A," but when the gene sequence for protein B is formed, "protein A" should be read as "protein B."

[0078] Mismatch repair deficient cancers have been reported in various cancers. Specific examples include colorectal cancer, endometrial cancer, gastric cancer, esophageal cancer, ovarian cancer, breast cancer, pancreatic cancer, hepatobiliary cancer, urinary tract cancer, small intestine cancer, prostate cancer, liver cancer, brain tumor, small intestine cancer, cervical cancer, neuroendocrine tumor, bile duct cancer, uterine sarcoma, thyroid cancer, retroperitoneal or peritoneal sarcoma, uveal melanoma, lung cancer, other malignant tumors of the female reproductive organs, cancer of unknown primary, and malignant soft tissue tumor (Non-Patent Document 3, etc.), and it is expected that mismatch repair deficient cancers will be discovered in various other cancers (including cancers of various animal species, including non-human animals) in the future. It has also been reported that molecularly targeted cancer therapies such as EGFR inhibitors and BRAF inhibitors reduce mismatch repair activity (Science 20 Dec 2019: Vol. 366, Issue 6472, pp. 1473-1480, DOI: 10.1126 / science.aav4474). Furthermore, a well-known example of mismatch repair-deficient cancer is familial (hereditary) Lynch syndrome (also known as hereditary nonpolyposis colorectal cancer (HNPCC)). Mismatch repair-deficient cancers that are targeted when the nucleic acid constructs of the present invention are used as therapeutic or diagnostic agents include various cancers, including those specifically mentioned above. In the present invention, the term "colorectal cancer" encompasses not only non-hereditary colorectal cancer but also Lynch syndrome (hereditary nonpolyposis colorectal cancer). Furthermore, in the present invention, the term "cancer" is synonymous with malignant tumor and includes sarcomas as well as malignant tumors (cancer in the narrow sense) arising from epithelial cells. The term "mismatch repair deficient" means that mismatch repair activity is reduced compared to cells with normal mismatch repair activity (e.g., non-MMR deficient control cells described below), and is not limited to complete deficiency of mismatch repair.

[0079] When the nucleic acid construct of the present invention is used as a therapeutic agent for mismatch repair-deficient cancer, the gene sequence encoding protein A may be a gene sequence encoding a protein that reduces cell viability. In cancers in which mismatch repair is not deficient, single-stranded annealing of the nucleic acid construct is blocked by the mismatch repair mechanism, preventing the gene sequence encoding protein A from being replicated within the cell and protein A from being expressed. On the other hand, in cancers in which mismatch repair is deficient, single-stranded annealing of the nucleic acid construct occurs within the cell, allowing the gene sequence encoding protein A to be replicated within the cell, resulting in the expression of protein A and efficient cancer cell killing. In cases where the cancer is not mismatch repair-deficient, the effect of killing cancer cells is not achieved. Therefore, prior to administering the therapeutic agent of the present invention, it may be possible to determine whether the cancer is mismatch repair-deficient, for example, using the diagnostic agent of the present invention, and then administer the therapeutic agent of the present invention to patients diagnosed with mismatch repair-deficient cancer. In cases where protein A is a protein that acts on other compounds to generate toxic substances, the target compound is also administered to the patient. For example, when protein A is HSV-tk, ganciclovir, FIAU, or the like may be administered in combination with the nucleic acid construct of the present invention.

[0080] When the nucleic acid construct of the present invention is used as a diagnostic agent for mismatch repair-deficient cancer, the nucleic acid construct of the present invention is introduced into cancer cells of a cancer patient, and the expression of protein A is measured. As described above, the measurement of protein A expression is preferably performed by directly or indirectly measuring the activity of protein A. As the gene sequence encoding protein A employed in the 5' and 3' regions, a gene sequence encoding a protein whose expression in cells can be detected is preferably used, but a gene sequence encoding a protein that has the effect of reducing cell viability can also be used. When mismatch repair is not deficient in a cell, single-stranded annealing of the nucleic acid construct is blocked by the mismatch repair mechanism, so the gene sequence encoding protein A is not reproduced in the cell, and protein A activity is not detected. On the other hand, when mismatch repair is deficient in a cell, single-stranded annealing of the nucleic acid construct occurs, and the gene sequence encoding protein A is reproduced in the cell, allowing protein A activity to be detected.

[0081] In one embodiment of the diagnostic agent for mismatch repair-deficient cancer, the cancer cells are cells isolated from a cancer patient, and the nucleic acid construct is introduced into the cancer cells ex vivo. Among the usage modes described below, (a) and (c) are specific examples of this mode. In this mode, either a protein whose expression in cells can be detected or a protein that reduces cell viability can be used as protein A, with the former being more preferred. In this mode, the cancer patient includes patients with solid cancers and patients with blood cancers. Cancer tissue cells isolated from cancer patients by biopsy or surgery, or, in the case of leukemia, cancer cells in the blood collected from cancer patients, can be used. If desired, the cancer cells may be concentrated before the nucleic acid construct is introduced. The introduction of the nucleic acid construct may be transient.

[0082] After introducing the nucleic acid construct, the expression of protein A in the cancer cells, preferably the activity of protein A, is measured. When protein A is a protein whose expression in cells can be detected, the expression (activity) of protein A can be measured by measuring the signal of protein A. When protein A is a protein that reduces cell viability, the expression (activity) of protein A can be measured by measuring the viability of the cancer cells after introducing the nucleic acid construct. When measuring the expression of protein A in cancer cells, it is preferable to use cells that have normal mismatch repair activity or are not deficient in mismatch repair activity as an MMR-non-deficient control, introduce the nucleic acid construct into the cancer cells and the control cells, and then compare the expression of protein A between the cancer cells and the control cells.

[0083] Preferred examples of control cells include non-cancer cells (cells that are not deficient in mismatch repair activity) collected from the same cancer patient. Also, known cell lines that are not deficient in mismatch repair activity can be used as control cells. Many known wild-type mammalian cell lines are known that do not have mutations in MMR-related genes (e.g., MSH2, MSH6, MLH1, PMS2, EPCAM, MSH3, PMS1, MLH3, ARID1A, SETD2, EXO1, RPA1, RPA2, RPA3, POLD, LIG1, LIG3, etc.), and any of these known lines can be used as MMR-non-deficient control cells. Specific examples of known mammalian cell lines having normal mismatch repair activity include, but are not limited to, human cell lines such as HT1080 (human fibrosarcoma cell line), U2OS (human osteosarcoma-derived cell line), HeLa (human cervical carcinoma-derived cell line), MCF-7 (human breast adenocarcinoma-derived cell line), HAP1 (human chronic myeloid leukemia-derived cell line), HEK293 (human fetal kidney-derived cell line), TIG-7, TIG-3 (human lung-derived cell lines), iPS cells (human induced pluripotent stem cells; established from normal human cells), and ES cells (human embryonic stem cells).

[0084] If desired, cells deficient in mismatch repair activity may be used as an additional control. Examples of such MMR-deficient control cells include cell lines derived from mismatch repair-deficient cancers. Various cell lines derived from mismatch repair-deficient cancers with mutations in the MMR-related genes described above are known, and any of these cells can be used as MMR-deficient control cells. Alternatively, cell lines created by introducing mutations into one or more MMR-related genes can also be used. Since gene knockout technology is well established, MMR-deficient cells from various animal species can be prepared by knocking out one or more MMR-related genes in cultured mammalian cells with normal mismatch repair activity. Specific examples of known MMR-deficient cell lines include human cell lines such as Nalm-6 (a human B-cell leukemia-derived cell line), LoVo (a human colon adenocarcinoma-derived cell line), Jurkat (a human T-cell leukemia-derived cell line), and C-33 A (a human cervical carcinoma-derived cell line), which are MSH2-deficient; HCT116 (a human colon adenocarcinoma-derived cell line), SK-OV-3 (a human ovarian carcinoma-derived cell line), and DU-145 (a human prostate carcinoma-derived cell line), which are MLH1-deficient; and DLD-1 (a human colon adenocarcinoma-derived cell line), which is MSH6-deficient. These cell lines are also available from Horizon Discovery (the following gene-deficient cells are available as the genetically modified HAP1 cell line (a human chronic myeloid leukemia-derived cell line): MSH2-deficient cell line, MLH1-deficient cell line, MSH6-deficient cell line, and PMS2-deficient cell line).Specific examples of non-human animal cell lines include MMR-deficient mice (frozen embryos) available from the NCI Mouse Repository (Msh2-deficient mice, Mlh1-deficient mice, and Msh6-deficient mice); the canine cancer cell line TYLER2 (Dodd G. Sledge et al., Proceedings of the 103rd Annual Meeting of the American Association for Cancer Research; 2012 Mar 31-Apr 4; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2012;72(8 Suppl):Abstract nr 5258. doi:1538-7445.AM2012-5258); and the MMR-deficient canine cell line Angus (MSH6-deficient) (Sunetra Das et al., Mol Cancer Ther. Author manuscript; available in PMC 2020 Feb 1. Published in final edited form as: Mol Cancer Ther. 2019 Aug;18(8): 1460-1471. Published online 2019 Jun 7. doi: 10.1158 / 1535-7163.MCT-18-1346), etc., but are not limited to these.

[0085] If the expression level / activity level of protein A in cancer cells is higher than the expression level / activity level of protein A in MMR-non-deficient control cells, it indicates that the cancer of the cancer patient is mismatch repair-deficient cancer.If protein A is a protein whose expression in cells can be detected as a signal, if a higher protein A signal is detected in cancer cells than in MMR-non-deficient control cells, it indicates that the cancer of the cancer patient is mismatch repair-deficient cancer.If protein A is a protein that has the effect of reducing cell viability, if the viability of cancer cells is lower than the viability of MMR-non-deficient control cells, it indicates that the cancer of the cancer patient is mismatch repair-deficient cancer.

[0086] If the expression level / activity level of protein A in cancer cells is not higher (same as or lower) than the expression level / activity level of protein A in MMR-non-deficient control cells, it is indicated that the cancer of the cancer patient is not mismatch repair-deficient. If protein A is a protein whose expression in cells can be detected as a signal, if a signal is detected in cancer cells at the same level or lower than that in MMR-non-deficient control cells, it is indicated that the cancer of the cancer patient is not mismatch repair-deficient. If protein A is a protein that has the effect of reducing cell viability, if the viability of cancer cells is at least the same as that of MMR-non-deficient control cells, it is indicated that the cancer of the cancer patient is not mismatch repair-deficient.

[0087] In another embodiment of the diagnostic agent for mismatch repair-deficient cancer, protein A is a protein whose intracellular expression can be detected as a signal, and the nucleic acid construct is introduced into cancer cells by administering the nucleic acid construct to a cancer patient, and then detecting a protein A signal from the cancer lesion is examined. Among the usage modes described below, (b) is a specific example of this mode. The nucleic acid construct may be administered locally into or near the tumor of the cancer patient, or systemically, with local administration being preferred. In this mode, the cancer patient typically has a solid tumor. To detect a protein A signal from the cancer lesion in the cancer patient's body, it is preferable to use a protein that generates a signal without reacting with a substrate, such as a fluorescent protein. If a protein A signal is detected from the cancer lesion (e.g., if the signal from the cancer lesion is clearly higher than the signal from the non-cancerous site surrounding the tumor), this indicates that the cancer in the cancer patient is mismatch repair-deficient. If a protein A signal is not detected from the cancer lesion, or if a low signal similar to that from the non-cancerous site surrounding the tumor is detected, this indicates that the cancer in the cancer patient is likely not mismatch repair-deficient. If desired, cancer cells may be collected from the patient and the nucleic acid construct may be introduced in vitro to further confirm whether the cancer is mismatch repair deficient.

[0088] In yet another embodiment, protein A is a secreted protein whose expression in cells can be detected, the nucleic acid construct is introduced into cancer cells by administering the nucleic acid construct to a cancer patient, and the activity of protein A in blood isolated from the patient after administration of the nucleic acid construct is measured. Among the use embodiments described below, (d) is a specific example of this embodiment. The nucleic acid construct may be administered locally into or near the tumor of the cancer patient, or systemically, with local administration being more preferred. In this embodiment, the cancer patient typically has a solid tumor. If the patient's cancer is mismatch repair deficient, secreted protein A is expressed from the protein A gene reproduced in the cancer cells and secreted outside the cancer cells, allowing the activity of protein A to be detected using a blood sample from the cancer patient. Since the activity of protein A secreted into the blood is detected in vitro, proteins that generate a signal upon reaction with a substrate can also be preferably used, and a typical example is a secreted luminescent enzyme, particularly secreted luciferase. Gene sequences encoding secreted proteins such as secreted luciferase may be prepared by directly using known secreted protein coding sequences, or by adding a sequence encoding a secretory signal to the protein coding sequence. Examples of negative control samples include blood samples collected from patients before administration of the nucleic acid construct, and culture supernatants obtained by introducing the nucleic acid construct into animal cell lines that are not deficient in mismatch repair. Positive controls include culture supernatants obtained by introducing the nucleic acid construct into cell lines deficient in mismatch repair activity (e.g., cell lines derived from mismatch repair-deficient cancers). Detection of a protein A signal in a blood sample (e.g., a higher protein A signal than in the negative control sample) indicates that the patient's cancer is mismatch repair-deficient. Detection of a protein A signal in a blood sample, particularly a signal not higher than the negative control (i.e., a signal at or below the level of the negative control sample), indicates that the patient's cancer is not likely to be mismatch repair-deficient.If desired, cancer cells may be collected from the patient and the nucleic acid construct may be introduced in vitro to further confirm whether the cancer is mismatch repair deficient.

[0089] By using the diagnostic technique for mismatch repair-deficient cancer according to the present invention, patients diagnosed with mismatch repair-deficient cancer are likely to be effective against immune checkpoint inhibitors, and therefore can be preferably administered with immune checkpoint inhibitors. Patients diagnosed with cancer other than mismatch repair-deficient cancer are likely to experience severe side effects despite insufficient cancer treatment with immune checkpoint inhibitors, and therefore can be preferably administered with anticancer drugs other than immune checkpoint inhibitors.

[0090] The diagnostic agent of the present invention may be embodied in the following manner, but is not limited thereto.

[0091] (a) Cancer cells collected from a cancer patient are cultured and transfected with the nucleic acid construct of the present invention to determine whether a protein A signal can be detected (or whether a cytocidal effect of protein A can be observed). If desired, non-cancerous cells (e.g., peripheral blood cells) collected from the same cancer patient can be used as a control. If protein A is luciferase, a luciferase substrate is added and the detection reaction is performed. If protein A is a fluorescent protein, the fluorescent signal can be detected using a luminometer or the like. If a protein A signal is detected (especially if the protein A signal is higher than that of control non-cancerous cells), the patient's cancer can be determined to be mismatch repair deficient. If a protein A signal is not detected or is detected at a level similar to that of control non-cancerous cells, the patient's cancer can be determined to be non-mismatch repair deficient. Methods for measuring cell viability are well known in the art and can be easily measured using commercially available kits, etc. If the survival rate of cancer cells into which the nucleic acid construct has been introduced is reduced (especially if the survival rate is lower than that of control non-cancer cells), the patient's cancer can be determined to be mismatch repair deficient. If no reduction in the survival rate of cancer cells is observed or the survival rate is similar to that of control non-cancer cells, the patient's cancer can be determined to not be mismatch repair deficient.

[0092] (b) The nucleic acid construct of the present invention is administered to a patient, and whether a protein A signal is detected in the cancer lesion is examined. If a signal is detected, the patient's cancer is determined to be mismatch repair deficient; if a signal is not detected, the patient's cancer may not be mismatch repair deficient. In this embodiment, a protein whose expression can be detected in cells is used as protein A, but it is necessary to use a protein with sufficiently low biotoxicity.

[0093] (c) Cancer cells present in the blood are isolated from the patient, concentrated as necessary, and transfected with the nucleic acid construct of the present invention to examine whether a signal from Protein A is detected (or whether a cell-killing effect (reduced cell viability) due to Protein A is observed) (in the case of leukemia).

[0094] (d) A nucleic acid construct employing a gene sequence encoding secreted luciferase as the gene sequence encoding protein A is administered to a patient, and then a blood sample is collected and examined for the presence or absence of a luciferase reaction in the blood sample. If the patient's cancer is mismatch repair deficient, secreted luciferase is produced from the luciferase gene replicated within the cancer cells and secreted outside the cells, allowing the luciferase reaction to be detected using the blood sample. If the patient's cancer is not mismatch repair deficient, secreted luciferase is not produced, and therefore no luciferase reaction is detected in the blood sample.

[0095] The diagnostic agent of the present invention is also useful as a companion diagnostic agent for predicting the efficacy of anticancer drugs against mismatch repair-deficient cancers. Examples of anticancer drugs include immune checkpoint inhibitors. Specific examples of immune checkpoint inhibitors include, but are not limited to, anti-PD-1 antibodies such as nivolumab, pembrolizumab, spartalizumab, and cemiplimab; anti-PD-L1 antibodies such as avelumab, atezolizumab, and durvalumab; and anti-CTLA-4 antibodies such as ipilimumab and tremelimumab. Specific examples of the use of the companion diagnostic agent are the same as those of the diagnostic agent described above. That is, when using the nucleic acid construct of the present invention as a companion diagnostic agent to predict the efficacy of an anticancer drug against mismatch repair-deficient cancers, similar to the diagnostic technique for mismatch repair-deficient cancers described above, the nucleic acid construct of the present invention can be introduced into cancer cells of a cancer patient, and the expression of protein A (preferably protein A activity) can be measured to determine whether the patient has mismatch repair-deficient cancer. When a patient is diagnosed with mismatch repair-deficient cancer, it can be determined that the patient is likely to achieve the desired anti-cancer effect from anti-cancer drugs for mismatch repair-deficient cancer, such as immune checkpoint inhibitors (i.e., it can be predicted that anti-cancer drugs for mismatch repair-deficient cancer, such as immune checkpoint inhibitors, will be effective in this patient). Therefore, immune checkpoint inhibitors can be preferably administered to such cancer patients for cancer treatment. However, there have been reports of severe side effects of immune checkpoint inhibitors. By determining whether a patient has mismatch repair-deficient cancer in advance, it is possible to avoid administering immune checkpoint inhibitors to patients who are likely to experience poor therapeutic effects or side effects.

[0096] Furthermore, the nucleic acid construct of the present invention can be used to search for factors, inhibitors, or activators that positively or negatively regulate mismatch repair or single-strand annealing reactions, and to predict whether a gene mutation identified in a cancer patient is a pathogenic mutation that impairs mismatch repair activity. By using a nucleic acid construct with 100% homology between homologous regions as a control in combination with a nucleic acid construct with low homology, drugs or mutations that affect mismatch repair activity can be identified. For example, in predicting a pathogenic mutation, if a mutation in a certain gene Q (preferably an MMR-associated gene) is identified in the cancer cells of a cancer patient, an expression vector expressing a mutant form of gene Q having the same mutation and an expression vector expressing a wild-type form of gene Q without the mutation can be constructed, and these expression vectors can be introduced into gene Q-deficient cells (which may be cells generated by knocking out gene Q, or, if a known cell line known to completely lack the function of gene Q exists, such a known cell line can be used), and the nucleic acid construct of the present invention (a 100% homologous construct and a low-homology construct) can be introduced, and the expression level of protein A in the cells can be measured. By comparing the expression of protein A between cells into which a 100% homologous nucleic acid construct has been introduced and cells into which a low-homology nucleic acid construct has been introduced, and between cells into which mutant gene Q has been introduced and cells into which wild-type gene Q has been introduced, it is possible to predict whether the mutation is a pathogenic mutation that impairs mismatch repair activity.

[0097] When the agent of the present invention is administered to a patient, the route of administration may be oral or parenteral, although parenteral administration such as intravenous, intraarterial, subcutaneous, or intramuscular administration is generally preferred. Systemic administration may be used, or administration may be into or near a tumor, or to a lymph node draining the tumor. The dosage is not particularly limited, but may be about 1 pg to 10 g, for example, about 0.01 mg to 100 mg, of nucleic acid construct per day per patient. Daily administration may be once or in divided doses. Administration may be daily, or every few days or weeks.

[0098] The dosage form of the agent of the present invention to be administered to a patient is not particularly limited, and can be formulated by appropriately mixing additives such as pharmaceutically acceptable carriers, diluents, and excipients with the nucleic acid construct of the present invention depending on the administration route. Examples of dosage forms include parenteral preparations such as drip infusions, injections, suppositories, and inhalants, and oral preparations such as tablets, capsules, granules, powders, and syrups. Formulation methods and usable additives are well known in the field of pharmaceutical formulations, and any method and additive can be used.

[0099] In the case of a diagnostic agent used in vitro, the amount used when treating cells may be, for example, about 10 pg to 1 mg, e.g., about 0.1 ng to 100 μg, of nucleic acid construct per 1 million cells. The dosage form may be a liquid agent to which additives useful for improving the stability of the nucleic acid construct, etc., have been added as desired, or the nucleic acid construct may be in powder form, for example, lyophilized.

[0100] When the nucleic acid construct of the present invention is used as a therapeutic agent, diagnostic agent, or the like, the target cancer patients include cancer patients of various mammals such as humans, dogs, cats, ferrets, hamsters, mice, rats, horses, pigs, etc. The nucleic acid construct of the present invention can be applied to various animal species without changing its configuration, except for changing the origin of the promoter as necessary.

[0101] [Structure of the Nucleic Acid Construct of the Second Aspect] The nucleic acid construct of the second aspect is a construct that originally contains a gene sequence encoding a protein, but becomes operably linked to a promoter upon the occurrence of an SSA reaction, enabling the gene sequence to express the protein. This construct comprises a promoter region, two substrate regions α and β, a gene sequence encoding protein X (hereinafter sometimes referred to as gene sequence X), and at least one complementary region containing two homologous regions that use each of the substrate regions α and β as substrates for the SSA reaction, located on a single nucleic acid molecule or on two or more different nucleic acid molecules.

[0102] The substrate region α is located downstream of the promoter region, and these are contained on the same nucleic acid molecule, but the other elements may be on the same nucleic acid molecule as these or on different nucleic acid molecules. Focusing on the location of the four elements other than the complementary region (promoter region, substrate region α, substrate region β, and gene sequence X), the nucleic acid construct of the second embodiment includes two types: a construct in which these four elements are located on the same nucleic acid molecule, and a construct in which the promoter region and substrate region α are located on one nucleic acid molecule and the substrate region β and gene sequence X are located on another nucleic acid molecule. The complementary region may be located on a nucleic acid molecule separate from the other elements, or it may be located on the same nucleic acid molecule as at least one of the other elements.

[0103] The definitions, conditions, and preferred examples of protein X and the gene sequence X encoding it are the same as those of protein A and the gene sequence encoding it in the nucleic acid construct of aspect 1. Briefly again, preferred examples of protein X include proteins that have the effect of reducing cell viability, or proteins whose expression in cells can be detected, and specific preferred examples of gene sequence X encoding such proteins include sequences of suicide genes, DNA damage-inducing genes, DNA repair inhibitor genes, luciferase genes, fluorescent protein genes, cell surface antigen genes, secreted protein genes, or membrane protein genes.

[0104] A poly A addition signal may be functionally linked downstream of gene sequence X. When linked, no other elements such as a complementary region are located between gene sequence X and the poly A addition signal. As with the nucleic acid construct of the first aspect, the poly A addition signal is not an essential element and can be omitted.

[0105] The conditions and preferred examples of the promoter region are also the same as those for the nucleic acid construct of the first aspect.

[0106] At least one complementary region includes one homologous region (hereinafter referred to as homologous region α) that uses substrate region α as a substrate for the SSA reaction, and one homologous region (hereinafter referred to as homologous region β) that uses substrate region β as a substrate for the SSA reaction. Similar to the complementary regions in the first nucleic acid construct, homologous region α and homologous region β are contained in the same or different complementary regions. In the case of a double SSA construct in which a protein is expressed by two SSA reactions, there is one complementary region, and this single complementary region includes homologous region α and homologous region β. In the case of a triple or more SSA construct in which a protein is expressed by three or more SSA reactions, two or more complementary regions are included, and homologous region α and homologous region β are contained in different complementary regions. The two or more complementary regions in a triple or more SSA construct include, in addition to homologous region α and homologous region β, one or more pairs of additional homologous regions and their substrates for causing an SSA reaction between the complementary regions.

[0107] The homology between each homologous region and each substrate region serving as its substrate is 40% or more but less than 100%, and the upper and lower limits can be selected from the values ​​described above. As described in the first embodiment, it is preferable that mismatched bases are scattered throughout the homologous region or its substrate.

[0108] The chain length of each homologous region and substrate region may basically be 4 to 10,000 bases, similar to the first embodiment, and the upper and lower limits thereof can be selected from the values ​​described in the explanation of the first embodiment.

[0109] The nucleotide sequence generated by the SSA reaction between each homologous region and its substrate may be a nucleotide sequence that does not encode a protein (Type I of the second embodiment). In this case, the substrate regions α and β and at least one complementary region may be designed based on any nucleotide sequence that does not encode a protein. Naturally occurring non-coding sequences, such as repetitive sequences such as SINEs found in the genome of living organisms, may be used, or artificial nucleotide sequences not present in genomes may be used. In Type I constructs, the non-coding sequence generated as a result of the SSA reaction is used as part of the 5' untranslated region between the translation initiation point downstream of the promoter and the start codon of gene sequence X. In the nucleic acid construct of the second embodiment constructed using a non-coding sequence, the SSA reaction causes gene sequence X to become operably linked to the promoter region, and protein X is expressed from gene sequence X.

[0110] Furthermore, the nucleotide sequence generated by the SSA reaction between each homologous region and its substrate may be a nucleotide sequence encoding a protein (Type II of the second embodiment). Hereinafter, the gene sequence generated by the SSA reaction in a Type II construct and the protein encoded thereby will be referred to as gene sequence Y and protein Y, respectively. This type of construct can also be said to be a construct having a structure in which gene sequence X is placed downstream of the 3' region of the construct of the first embodiment, and the substrate regions α and β and at least one complementary region can be designed in the same manner as the 5' region, 3' region, and at least one complementary region of the nucleic acid construct of the first embodiment. More precisely, substrate region α in Type II of the second embodiment corresponds to a region including at least the 3'-end portion of the 5'-side region that serves as a substrate for the SSA reaction by homologous region α in the construct of the first embodiment (i.e., the entire 5'-side region or a partial region including the 3'-end), and substrate region β corresponds to a region including at least the 5'-end portion of the 3'-side region that serves as a substrate for the SSA reaction by homologous region β in the construct of the first embodiment (i.e., the entire 3'-side region or a partial region including the 5'-end). When arranged in this order, substrate region α, at least one complementary region, and substrate region β encode the amino acid sequence of protein Y with mutual overlap between each homologous region and its substrate, and gene sequence Y is produced as a result of the SSA reaction.

[0111] In a type II construct, gene sequence Y is generated upstream of gene sequence X by the SSA reaction, and protein Y and protein X are expressed by operably linking the promoter region with gene sequence Y and the downstream gene sequence X. Protein Y and protein X may be expressed in the form of a fusion protein by providing a nucleotide sequence encoding an appropriate polypeptide such as a linker peptide between substrate region β and gene sequence X, or protein Y and protein X may be expressed polycistronically by utilizing a sequence that enables polycistronic expression. In the former case, a protease cleavage sequence that is cleaved by a protease (such as furin) present in cells may be inserted between protein Y and protein X, and the protein may be designed so that after expression in the form of a fusion protein, individual proteins are cleaved by the protease to produce the individual proteins.

[0112] Various sequences that enable polycistronic expression are known, including IRES sequences and 2A peptide coding sequences, and any of these may be used. Specific examples of IRES sequences include the IRES sequence shown in SEQ ID NO: 170 and the IRES2 sequence shown in SEQ ID NO: 171 (both of which are derived from the internal ribosome entry site of encephalomyocarditis virus (ECMV)). 2A peptide coding sequences include P2A coding sequences (SEQ ID NOs: 164 and 165), E2A coding sequences (SEQ ID NOs: 166 and 167), and F2A coding sequences (SEQ ID NOs: 168 and 169). When an IRES sequence is used, a stop codon must be added to gene sequence Y; on the other hand, when a 2A peptide coding sequence is used, a stop codon must not be added to gene sequence Y.

[0113] The conditions and preferred examples of gene sequence Y and the protein Y encoded thereby are also the same as those for protein A and the gene sequence encoding it in the nucleic acid construct of aspect 1. In a Type II construct, it is preferred that one of protein X and protein Y is a protein that has the effect of reducing cell viability, and the other is a protein whose expression in cells can be detected, and it is preferred that one of gene sequence X and gene sequence Y is the sequence of a suicide gene, a DNA damage-inducing gene, or a DNA repair-inhibiting gene, and the other is the sequence of a luciferase gene, a fluorescent protein gene, a cell surface antigen gene, a secreted protein gene, or a membrane protein gene.

[0114] Typical examples of the nucleic acid construct of the second aspect include the following: (1) A construct in which the promoter region, substrate region α, substrate region β, and gene sequence X are arranged on the same nucleic acid molecule, and at least one complementary region is entirely contained in a separate nucleic acid molecule; (2) A construct in which the promoter region and substrate region α are arranged on one nucleic acid molecule, and substrate region β and gene sequence X are arranged on another nucleic acid molecule, and at least one complementary region is entirely contained in a separate nucleic acid molecule; and (3) A nucleic acid construct in which the promoter region, substrate region α, substrate region β, gene sequence X, and all complementary regions are arranged on the same nucleic acid molecule.

[0115] The nucleic acid construct of the second aspect also includes the following constructs, which are variants of (1) to (3): (4) (Variant of (1)(3)) A portion of at least one complementary region is located on the same nucleic acid molecule as the promoter region, etc., and the remaining complementary regions are located on one or more separate nucleic acid molecules. (5) (Variant of (2)) At least one complementary region is located on either or both of a nucleic acid molecule comprising substrate region α and a nucleic acid molecule comprising substrate region β. (6) (Variant of (2)) A portion of at least one complementary region is located on either or both of a nucleic acid molecule comprising substrate region α and a nucleic acid molecule comprising substrate region β, and the remaining complementary regions are located on one or more separate nucleic acid molecules.

[0116] Among the constructs of the second aspect (1), a construct consisting of two nucleic acid molecules and including one complementary region has the following configuration: nucleic acid molecule 1 includes a promoter region, substrate region α, substrate region β, and gene sequence X; nucleic acid molecule 2 includes one complementary region including a first homologous region and a second homologous region; the first homologous region uses substrate region α as a substrate for the SSA reaction, and the homology between the two is 40% or more but less than 100%; the second homologous region uses substrate region β as a substrate for the SSA reaction, and the homology between the two is 40% or more but less than 100%; in the case of a Type I construct, protein X is expressed by an SSA reaction occurring between each homologous region and its substrate; in the case of a Type II construct, substrate region α, complementary region, and substrate region β, when arranged in this order, encode the amino acid sequence of protein Y with overlap between each homologous region and its substrate, and protein X and protein Y are expressed by an SSA reaction occurring between each homologous region and its substrate.

[0117] Among the constructs of the second aspect (1), a construct consisting of three or more nucleic acid molecules and containing two or more complementary regions can be expressed as follows, using an integer n that is a constant in the range of 3≦n≦10 and an integer m that is a variable in the range of 3≦m≦n: n is preferably 9 or less, 8 or less, 7 or less, 6 or less, 5 or less, or 4 or less.

[0118] A nucleic acid construct is composed of n nucleic acid molecules and includes (n-1) complementary regions, wherein nucleic acid molecule 1 includes a promoter region, substrate region α, substrate region β, and gene sequence X, nucleic acid molecule 2 includes a first complementary region including a first homologous region and a second homologous region, and the mth nucleic acid molecule m includes the (m-1)th substrate region and the (m-1)th complementary region including the mth homologous region. The first homologous region uses substrate region α as a substrate for an SSA reaction, and the homology between the two is 40% or more and less than 100%, the (m-1)th homologous region uses the (m-1)th substrate region as a substrate for an SSA reaction, and the homology between the two is 40% or more and less than 100%, and the nth homologous region in nucleic acid molecule n uses substrate region β as a substrate for an SSA reaction, and the homology between the two is 40% or more and less than 100%. In the case of a type I construct, protein X is expressed by an SSA reaction occurring between each homologous region and its substrate. In the case of a type II construct, when substrate region α, the first complementary region, the (m-1)th complementary region, and substrate region β are arranged in this order (complementary regions are arranged in ascending order of their numbers), they encode the amino acid sequence of protein Y with overlapping between each homologous region and its substrate, and protein X and protein Y are expressed by an SSA reaction occurring between each homologous region and its substrate.

[0119] For example, when n=3, a nucleic acid construct (1) consisting of three nucleic acid molecules (nucleic acid molecules 1 to 3) and containing two complementary regions has the above-described configuration: nucleic acid molecule 3 contains a second complementary region including a second substrate region and a third homologous region; the second homologous region uses the second substrate region as a substrate for the SSA reaction, with a homology between the two being 40% or more and less than 100%; and the third homologous region uses substrate region β as a substrate for the SSA reaction, with a homology between the two being 40% or more and less than 100%. In the case of a Type I construct, protein X is expressed by an SSA reaction occurring between each homologous region and its substrate. In the case of a Type II construct, when substrate region α, first complementary region, second complementary region, and substrate region β are arranged in this order, they encode the amino acid sequence of protein Y with overlapping between each homologous region and its substrate, and protein X and protein Y are expressed by an SSA reaction occurring between each homologous region and its substrate.

[0120] When n=4, the nucleic acid construct (1) is composed of four nucleic acid molecules (nucleic acid molecules 1 to 4) and includes three complementary regions. In the above-mentioned configuration, nucleic acid molecule 3 includes a second complementary region including a second substrate region and a third homologous region, nucleic acid molecule 4 includes a third complementary region including a third substrate region and a fourth homologous region, the second homologous region uses the second substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%, the third homologous region uses the third substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%, and the fourth homologous region uses substrate region β as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%. In the case of a Type I construct, protein X is expressed by the SSA reaction occurring between each homologous region and its substrate. In the case of a type II construct, when the substrate region α, the first complementary region, the second complementary region, the third complementary region, and the substrate region β are arranged in this order, they encode the amino acid sequence of protein Y with overlapping between each homologous region and its substrate, and protein X and protein Y are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0121] In the construct of (1) of the second aspect, nucleic acid molecule 1 may be a circular nucleic acid molecule comprising, in this order, substrate region α, substrate region β, and gene sequence X downstream of the promoter region, or a linear nucleic acid molecule comprising, in this order from upstream to downstream, substrate region β, gene sequence X, the promoter region, and substrate region α. ​​When a poly(A) addition signal is included, nucleic acid molecule 1 may be a circular nucleic acid molecule comprising, in this order from upstream to downstream, substrate region α, substrate region β, gene sequence X, and poly(A) addition signal downstream of the promoter region, or a linear nucleic acid molecule comprising, in this order from upstream to downstream, substrate region β, gene sequence X, the poly(A) addition signal, the promoter region, and substrate region α.

[0122] When nucleic acid molecule 1 is a circular molecule, it is usually cleaved between substrate regions α and β to linearize it and then introduced into cells for use, so it preferably has a cleavage site between substrate regions α and β. Examples of the cleavage site are the same as those in the construct of the first aspect.

[0123] It is preferable that all nucleic acid molecules containing a complementary region other than nucleic acid molecule 1 are linear molecules for ease of production and use.

[0124] Among the constructs of the second aspect (2), a construct consisting of three nucleic acid molecules and including one complementary region has the following configuration: Nucleic acid molecule 1-1 includes a promoter region and substrate region α. ​​Nucleic acid molecule 1-2 includes substrate region β and gene sequence X. The first homologous region uses substrate region α as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%. The second homologous region uses substrate region β as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%. In the case of a Type I construct, protein X is expressed by the SSA reaction occurring between each homologous region and its substrate. In the case of a Type II construct, when the substrate region α, complementary region, and substrate region β are arranged in this order, each homologous region and its substrate overlap with each other to encode the amino acid sequence of protein Y, and protein X and protein Y are expressed by the SSA reaction occurring between each homologous region and its substrate.

[0125] Among the constructs of the second aspect (2), a construct composed of four or more nucleic acid molecules and containing two or more complementary regions can be expressed as follows, using integer n, which is a constant such that 3≦n≦10, and integer m, which is a variable such that 3≦m≦n. n is preferably 9 or less, 8 or less, 7 or less, 6 or less, 5 or less, or 4 or less. A nucleic acid construct composed of (n+1) nucleic acid molecules and containing (n−1) complementary regions, wherein nucleic acid molecule 1-1 contains a promoter region and a substrate region α, nucleic acid molecule 1-2 contains a substrate region β and a gene sequence X, nucleic acid molecule 2 contains a first complementary region containing a first homologous region and a second homologous region, and the mth nucleic acid molecule m contains an (m−1)th complementary region containing an (m−1)th substrate region and the mth homologous region. The first homologous region uses substrate region α as a substrate for the SSA reaction, and the homology between them is 40% or more but less than 100%; the (m-1)th homologous region uses substrate region (m-1) as a substrate for the SSA reaction, and the homology between them is 40% or more but less than 100%; and the nth homologous region in nucleic acid molecule n uses substrate region β as a substrate for the SSA reaction, and the homology between them is 40% or more but less than 100%. In the case of a Type I construct, protein X is expressed by an SSA reaction occurring between each homologous region and its substrate. In the case of a Type II construct, when substrate region α, the first complementary region, the (m-1)th complementary region, and substrate region β are arranged in this order (the complementary regions are arranged in ascending order of their numbers), they encode the amino acid sequence of protein Y with overlap between each homologous region and its substrate, and protein X and protein Y are expressed by an SSA reaction occurring between each homologous region and its substrate.

[0126] For example, when n=3, a nucleic acid construct (2) consisting of four nucleic acid molecules (nucleic acid molecules 1-1, 1-2, 2, and 3) and containing two complementary regions has the above-mentioned configuration, where nucleic acid molecule 3 contains a second complementary region containing a second substrate region and a third homologous region, the second homologous region uses the second substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%, and the third homologous region uses substrate region β as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%. In the case of a Type I construct, protein X is expressed by the SSA reaction that occurs between each homologous region and its substrate. In the case of a type II construct, when the substrate region α, the first complementary region, the second complementary region, and the substrate region β are arranged in this order, they encode the amino acid sequence of protein Y with overlapping between each homologous region and its substrate, and protein X and protein Y are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0127] When n=4, the nucleic acid construct (2) is composed of five nucleic acid molecules (nucleic acid molecules 1-1, 1-2, 2, 3, and 4) and includes three complementary regions, and in the above-mentioned configuration, nucleic acid molecule 3 includes a second complementary region including a second substrate region and a third homologous region, nucleic acid molecule 4 includes a third complementary region including a third substrate region and a fourth homologous region, the second homologous region uses the second substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%, the third homologous region uses the third substrate region as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%, and the fourth homologous region uses substrate region β as a substrate for the SSA reaction, and the homology between the two is 40% or more and less than 100%. In the case of a Type I construct, protein X is expressed by the SSA reaction occurring between each homologous region and its substrate. In the case of a type II construct, when the substrate region α, the first complementary region, the second complementary region, the third complementary region, and the substrate region β are arranged in this order, they encode the amino acid sequence of protein Y with overlapping between each homologous region and its substrate, and protein X and protein Y are expressed by the SSA reaction that occurs between each homologous region and its substrate.

[0128] It is preferable that all of the nucleic acid molecules constituting the nucleic acid construct of the second aspect (2) are prepared as linear nucleic acid molecules for ease of production and use, but some or all of them may be prepared as circular molecules. When the nucleic acid construct of (2) contains a poly(A) addition signal, the nucleic acid molecule 1-2 containing the gene sequence X contains the poly(A) addition signal downstream of the gene sequence X.

[0129] Nucleic acid constructs (4) to (6), which are variants of constructs (1) to (3), can have configurations that include two or more complementary regions. Constructs (4) to (6) of the first embodiment are exemplified as specific configurations of constructs with up to four complementary regions, but in the explanation of these specific examples, the "5'-side region" is replaced with "substrate region α" and the "3'-side region" is replaced with "substrate region β," respectively, to provide specific examples of constructs with up to four complementary regions of constructs (4) to (6) of the second embodiment.

[0130] [Various uses of the nucleic acid construct of the second embodiment] A type I construct in which the sequence generated by the SSA reaction is a non-coding sequence can be used in the same manner as the construct of the first embodiment. In the description of [Various uses of the nucleic acid construct of the first embodiment], "protein A" can be read as "protein X."

[0131] A type II construct in which the sequence produced by the SSA reaction is a coding sequence and two types of proteins X and Y are expressed as a result of the SSA reaction can be used in the same manner as the construct of the first embodiment.

[0132] Furthermore, in a type II construct, if one of proteins X and Y is a protein that reduces cell viability and the other is a protein whose expression in cells can be detected, detection and killing of mismatch repair-deficient cells can be achieved with a single construct. Such dual-type constructs can be used as agents for detecting and treating mismatch repair-deficient cancers. In Examples B and C below, the gene sequence Y generated by the SSA reaction is a gene sequence encoding a protein whose expression in cells can be detected, and the gene sequence X originally incorporated into the nucleic acid construct is a gene sequence encoding a protein that reduces cell viability. However, the reverse configuration is also possible. The step of detecting mismatch repair-deficient cancers when using a dual-type construct can be carried out in the same manner as the detection of mismatch repair-deficient cancers using the nucleic acid construct of the first embodiment.

[0133] The present invention will be described in more detail below with reference to examples. However, the present invention is not limited to the following examples. In the following examples, the size of each DNA fragment may be expressed excluding the stop codon added to the 3' end.

[0134] Example AI: Split DR-Nluc-homeo (SceNluc + iNluc) (two-split, two-times SSA type) (Figure 1) A construct was constructed in which the normal Nluc gene was formed by two-point (two-times) single-strand annealing (SSA), and luciferase activity was evaluated.

[0135] AI-1. Vector Construction. The 513-bp SceNluc fragment used in pCMV-SceNluc (SEQ ID NO: 9; DNA in which the region from positions 223 to 243 in the Nluc gene (SEQ ID NO: 7) was replaced with a stop codon (TGA) plus an I-SceI recognition sequence, with the stop codon TAA added to the 3' end) and the 459-bp iNluc fragment used in pUC-iNluc (SEQ ID NO: 12; the region from positions 16 to 474 in the Nluc gene (SEQ ID NO: 7) with the stop codon TAA added to the 3' end) were amplified by PCR using the pNL1.1[Nluc] Vector (Promega, Madison, WI, USA) as a template. PrimeSTAR HS DNA Polymerase (Takara Bio) or Tks Gflex DNA Polymerase (Takara Bio) were used as the polymerase for PCR. The primers used are listed in Table 1 below.

[0136] The SceNluc fragment was amplified separately into two fragments: the 5'SceNluc fragment and the 3'SceNluc fragment. The 5'SceNluc fragment was amplified using Nluc-Fw (SEQ ID NO: 1) and Sce-5'Nluc-Rv (SEQ ID NO: 2). A DNA fragment with a structure in which a sequence for the In-Fusion reaction was added to the 5' end of the 5'SceNluc fragment and a stop codon (TGA) + I-SceI recognition sequence was added to the 3' end was amplified. The 3'SceNluc fragment was amplified using Sce-3'Nluc-Fw (SEQ ID NO: 3) and Nluc-Rv (SEQ ID NO: 4). A DNA fragment with a stop codon (TGA) + I-SceI recognition sequence was added to the 5' end of the 3'SceNluc fragment and a stop codon (TAA) + I-SceI recognition sequence was added to the 3' end was amplified.

[0137] The iNluc fragment was amplified using iNluc-Fw (SEQ ID NO: 5) and iNluc-Rv (SEQ ID NO: 6), resulting in a DNA fragment with a stop codon (TGA) linked to the 3' end of the iNluc fragment and sequences for the In-Fusion reaction added to both ends.

[0138] Next, pIRES (Takara Bio, Figure 2) was digested with NheI and XbaI to remove the IRES region and recover a fragment containing the CMV promoter and poly(A) addition signal. Using the In-Fusion HD Cloning Kit (Takara Bio), the 5'SceNluc fragment and the 3'SceNluc fragment were inserted downstream of the CMV promoter of the recovered fragment, yielding pCMV-SceNluc, which is a CMV promoter-SceNluc fragment-poly(A) addition signal ligation. The SceNluc fragment (SEQ ID NO: 9) contains a stop codon TGA introduced upstream of the I-SceI recognition sequence and two stop codons TAG and TAA within the I-SceI recognition sequence. The amino acid sequences encoded by positions 1 to 225 (the region up to the first stop codon) and positions 235 to 516 (the region downstream of the third stop codon) of the SceNluc fragment are shown in SEQ ID NOs: 10 and 11, respectively (however, a polypeptide consisting of the latter sequence is not expressed).

[0139] pUC-iNluc was obtained by digesting pUC19 (Takara Bio, Figure 3) with SmaI in the multicloning site and inserting the iNluc fragment using the DNA Ligation Kit <Mighty Mix> (Takara Bio).

[0140]

[0141] To create a vector with reduced homology between homologous sequences (between the homologous region and the substrate), DNA fragments were prepared by artificial gene synthesis, in which silent mutations were introduced into the regions of positions 1 to 207 and 229 to 459 of the iNluc sequence (459 bp, SEQ ID NO: 12). The silent mutations had 99% (4 sites; 2 / 207 at the 5' end and 2 / 231 at the 3' end), 98% (8 sites; 4 / 207 at the 5' end and 4 / 231 at the 3' end), 97% (13 sites; 6 / 207 at the 5' end and 7 / 231 at the 3' end), 96% (17 sites; 8 / 207 at the 5' end and 9 / 231 at the 3' end), 95% (23 sites; 11 / 207 at the 5' end and 12 / 231 at the 3' end), 94% (26 sites; 12 / 207 at the 5' end and 14 / 231 at the 3' end), 90% (44 sites; 21 / 207 at the 5' end and 23 / 231 at the 3' end), and 80% (94 sites; Silent mutations were introduced so that the 5' end (44 / 207 sites), 3' end (50 / 231 sites), and 70% (140 sites; 5' end (66 / 207 sites), 3' end (74 / 231 sites) were distributed as evenly as possible throughout the sequence (Figures 4-1 to 4-4). The sequences of iNluc fragments with 95%, 90%, 80%, and 70% homology (Figures 4-1 and 4-2) are shown in SEQ ID NOs: 14–17, and the sequences of iNluc fragments with 99%, 98%, 97%, 96%, and 94% homology (Figures 4-3 and 4-4) are shown in SEQ ID NOs: 18–22, respectively. The vectors into which nine artificially synthesized iNluc fragments with different mutations were cloned were digested with BamHI (Takara Bio) and EcoRI (Takara Bio), and the fragments containing the iNluc with each mutation were excised. Next, pUC19 (Takara Bio) was digested with BamHI and EcoRI, and each iNluc fragment carrying a mutation was inserted using the DNA Ligation Kit <Mighty Mix> (Takara Bio). This yielded pUC-iNluc (99% homology), pUC-iNluc (98% homology), pUC-iNluc (97% homology), pUC-iNluc (96% homology), pUC-iNluc (95% homology), pUC-iNluc (94% homology), pUC-iNluc (90% homology), pUC-iNluc (80% homology), and pUC-iNluc (70% homology), respectively.

[0142] pCMV-SceNluc was purified using the Qiagen Plasmid Plus Midi Kit (Qiagen KK) and linearized by digestion with I-SceI (New England Biolabs, Ipswich, MA, USA) before transfection. pUC19 and pUC-iNluc were purified using the Qiagen Plasmid Plus Midi Kit (Qiagen KK) and linearized by digestion with AhdI or BamHI and EcoRI. The iNluc-containing fragment was purified using the Wizard SV Gel and PCR Clean-Up System (Promega) before use. Digestion of pUC-iNluc with AhdI yielded DNA with an additional sequence of 1 kb or more ligated to both ends of the iNluc fragment (Figure 1C), whereas digestion with BamHI and EcoRI yielded DNA without the additional sequence (Figure 1D).

[0143] AI-2. Cells, Gene Transfection, and Luciferase Assay. The human pre-B cell line Nalm-6 (a cancer cell line deficient in mismatch repair) and its derivatives were cultured at 37°C in a 5% CO2 incubator (So et al. (2004) Genetic interactions between BLM and DNA ligase IV in human cells. J. Biol. Chem. 279: 55433-55442). Msh2 revertant cells (Msh2+ cells) were obtained by knocking in MSH2 cDNA into the human pre-B cell line Nalm-6 (a cancer cell line deficient in mismatch repair (MMR); Msh2-). The specific method was as previously reported (Restoration of mismatch repair functions in human cell line Nalm-6, which has high efficiency for gene targeting. Suzuki T, Ukai A, Honma M, Adachi N, Nohmi T. PLoS One. 2013 Apr 15;8(4):e61189. doi: 10.1371 / journal.pone.0061189.). In Msh2+ cells, the mismatch repair mechanism was restored by knocking in MSH2. Constructs that show large differences in luciferase activity and cell killing effect (i.e., recombination efficiency) between Msh2- and Msh2+ cells can be evaluated as constructs that are sensitive to the presence or absence of mismatch repair. Nalm-6 cells (wild-type and its Msh2-revertant cells) were transfected by electroporation as previously reported (Saito et al. (2017) Dual loss of human POLQ and LIG4 abolishes random integration. Nat. Commun. 8: 16112).

[0144] The human colon cancer cell line HCT116 (Horizon Discovery, Cambridge, UK) and its MLH1 revertant cells (MLH1+ HCT116 cells) (Horizon Discovery) were cultured at 37°C in a 5% CO2 incubator. The culture medium used was Dulbecco's modified Eagle's medium (Nissui Pharmaceutical, Tokyo, Japan) supplemented with calf serum (Global Life Science Technologies Japan, Tokyo, Japan; 50 ml added per 500 ml of Eagle's MEM) and L(+)-glutamine (Fujifilm Wako Pure Chemical Industries, Osaka, Japan; added to a final concentration of 2 mM). Transfection was performed using jetPEI (Polyplus-transfection, Illkirch, France) according to the manufacturer's protocol.

[0145] Luciferase assays in Nalm-6 cells (wild-type and its Msh2-revertant cells) were performed as follows. 6 Each construct (1 μg) was transfected into 5 x 10 Nalm-6 cells. 5 The cells were seeded into a 12-well dish at 1000 cells / ml, and luciferase activity (RLU value) was measured immediately after gene transfection, 4 hours later, and 24 hours later using the Nano-Glo Luciferase Assay System (Promega).

[0146] Luciferase assays in HCT116 cells (wild-type and its MLH1-reverted cells) were performed as follows. 4 Cells were seeded into 24-well dishes, cultured overnight, and then transfected with 1 μg of each construct. Luciferase activity (RLU) was measured immediately, 4 hours, and 24 hours after transfection using the Nano-Glo Luciferase Assay System (Promega).

[0147] AI-3. Results The construct constructed in this example was designed to generate a normal Nluc gene when homologous recombination (HR) occurs between the SceNluc and iNluc fragments or when single-strand annealing (SSA) occurs at two sites on the 5' and 3' ends of each fragment. HR occurs when the homology in the homologous regions is sufficiently high. However, reducing the homology in the homologous regions prevents HR, and normal Nluc gene is generated only when SSA occurs. Because the SSA reaction is inhibited by MMR factors, the presence or absence of MMR factors is thought to result in differences in Nluc gene expression levels. The SSA construct previously developed by the inventors of the present application (WO 2021 / 162121 A1; hereinafter sometimes referred to as the conventional construct) was designed to form a normal Nluc gene through a single SSA reaction, but the construct of this example requires SSA to occur at two locations (two times) to form a normal Nluc gene, and therefore it is expected that the difference in Nluc gene expression levels due to the presence or absence of MMR factors will be greater than with the conventional construct.

[0148] Figures 5, 6-1, and 6-2 show examples of luciferase assay results in Nalm-6 cells transfected with a two-part, two-times SSA construct and their MSH2-reverted counterparts. Figure 5 shows the results of measuring luciferase activity over time after cotransfection of an iNluc fragment lacking additional sequences with pCMV-SceNluc. Constructs with reduced homology in the homologous region (95% to 70%) exhibited high Luc activity in MSH2-deficient (MMR-deficient WT Nalm-6 cells). Figure 6-1 shows luciferase activity 4 hours after transfection (A) and 24 hours after transfection (B). Constructs with reduced homology in the homologous region (98% or less) clearly demonstrated differences in luciferase activity between MSH2-deficient (MMR-deficient) and non-deficient cells. The difference in luciferase activity was more pronounced 24 hours after transfection than 4 hours after transfection. In particular, the construct with 90% homology in the homologous region showed a 170-fold difference in luciferase activity, while the construct with 80% homology showed a 62-fold difference in luciferase activity. Figure 6-2 shows a graph comparing the luciferase activity 24 hours after gene transfection between the iNluc fragment with an additional sequence (pUC-iNluc / AhdI, left side of the graph) and the iNluc fragment without an additional sequence (pUC-iNluc / BamHI + EcoRI, right side of the graph). The iNluc fragment without an additional sequence showed a larger difference in luciferase activity depending on whether or not the MSH2 deficiency (MMR deficiency) was present. In particular, the constructs with 80% and 90% homology in the homologous region showed a significant difference of more than 200-fold in luciferase activity.

[0149] In substrates with 100% homology in the homologous region, recombination by the HR reaction is dominant over recombination by the SSA reaction, making it difficult to observe differences in luciferase activity due to the presence or absence of MSH2 deficiency (MMR deficiency).On the other hand, in substrates with reduced homology, Nluc gene formation by the SSA reaction proceeds under MSH2 deficiency (MMR deficiency), which is thought to result in differences in luciferase activity due to the presence or absence of MSH2 deficiency (MMR deficiency).

[0150] Figure 7-1 shows the results of transfection of two SSA constructs into HCT116, a cancer cell line with MMR deficiency due to MLH1 deficiency, and its MLH1 revertant cells. Luciferase activity was measured 4 hours (A) or 24 hours (B). Figure 7-2 compares luciferase activity 24 hours after transfection between the iNluc fragment with an additional sequence (pUC-iNluc / AhdI, left side of the graph) and the iNluc fragment without the additional sequence (pUC-iNluc / BamHI + EcoRI, right side of the graph). Similar to the presence or absence of mismatch repair deficiency due to MSH2, the presence or absence of mismatch repair deficiency due to MLH1 also confirmed that the difference in gene expression (luciferase activity) was enhanced by reducing the homology of the homologous region. Furthermore, even in the MLH1-deficient system, the difference in luciferase activity between the iNluc fragment without the long additional sequence and the absence or presence of the deletion was greater, although not as significant as in the MSH2-deficient system. This result suggests that the constructs of the present invention are not affected by the cause of the mismatch repair deficiency.

[0151] Example A-II: Integrated DR-Nluc-homeo (double SSA in one molecule) (Figure 8) A double SSA construct, pCMV-SceNluc-iNluc, in which the SceNluc gene and the iNluc gene are carried in a single vector, was constructed and its luciferase activity was evaluated.

[0152] A-II-1. Vector Preparation pCMV-SceNluc prepared in AI-1 was digested with BamHI, and the iNluc fragment was inserted downstream of the poly(A) using the In-Fusion HD Cloning Kit (Takara Bio). This gave pCMV-SceNluc-iNluc, which is a ligation of [CMV promoter]-[SceNluc fragment]-[poly(A) addition signal]-[iNluc fragment].

[0153] Next, to construct pCMV-SceNluc-iNluc, in which the SceNluc gene and the iNluc gene with each mutation were carried in a single vector, the vector into which the iNluc gene with each mutation had been cloned was digested with BamHI and AhdI (New England Biolabs), and fragments containing iNluc with each mutation were excised. Next, pCMV-SceNluc was digested with BamHI and AhdI, and fragments containing the SceNluc gene were excised. Finally, using the DNA Ligation Kit <Mighty Mix>, the iNluc fragments with each mutation were inserted downstream of the poly(A) addition signal to obtain pCMV-SceNluc-iNluc (95% homology), pCMV-SceNluc-iNluc (90% homology), pCMV-SceNluc-iNluc (80% homology), and pCMV-SceNluc-iNluc (70% homology), each of which ligated the CMV promoter, SceNluc fragment, poly(A) addition signal, and iNluc fragment with each mutation.

[0154] A-II-2. Cells and gene transfection, luciferase assay. As described in AI-2.

[0155] A-II-3. Results Figure 9 shows the luciferase activity 24 hours after transfection of the integrated construct into Nalm-6 cells and their MSH2-reverted counterparts. Although the difference in luciferase activity between the presence and absence of MSH2 deficiency (MMR deficiency) was somewhat less apparent than with the two-part construct, the integrated construct was still able to detect the presence or absence of MMR deficiency.

[0156] Example A-III: Split NlucP-homeotransfectant (SceNlucP + iNluc) (Bipartite, Double SSA Type) (Figure 10). To create a construct sensitive to changes in the expression level of the normal Nluc gene, we constructed a bipartite expression vector for the NlucP gene, in which the stability of the Nluc protein (i.e., increased susceptibility to degradation) was reduced by fusing the Nluc gene with a PEST sequence (proteolysis signal). The PEST sequence (hPEST, SEQ ID NO: 27) derived from mouse ornithine decarboxylase, codon-optimized for use in human cells, was used as the proteolysis signal. The hPEST sequence was generated by artificial gene synthesis (phPEST). For iNluc, we used pUC-iNluc constructs with 100%, 98%, 96%, 94%, and 90% homology, as prepared in Example AI, which were linearized by digestion with BamHI and EcoRI.

[0157] A-III-1. Vector Construction. The 513-bp Nluc fragment used in pCMV-NlucP was obtained by PCR amplification using pNL1.1[Nluc] Vector (Promega, Madison, WI, USA), and the 120-bp hPEST fragment (SEQ ID NO: 26, with a stop codon added to the 3' end) was obtained by PCR amplification using hPEST as a template. PrimeSTAR HS (trade name) DNA Polymerase (Takara Bio) was used as the PCR polymerase. The primers used are listed in Table 2 below. Nluc-Fw (SEQ ID NO: 1) and NlucP-Rv (SEQ ID NO: 23) were used to amplify the Nluc fragment, resulting in a DNA fragment with sequences for the In-Fusion reaction added to both ends of the Nluc fragment. The hPEST fragment was amplified using hPEST-Fw (SEQ ID NO: 24) and hPEST-Rv (SEQ ID NO: 25). A DNA fragment containing a stop codon (TAA) at the 3' end of the hPEST fragment and sequences for the In-Fuson reaction at both ends was amplified. Next, pIRES (Takara Bio, Figure 2) was digested with NheI and XbaI to remove the IRES region and recover a fragment containing the CMV promoter and poly(A) addition signal. Using the In-Fusion HD Cloning Kit (Takara Bio), the Nluc fragment and hPEST fragment were inserted downstream of the CMV promoter of the recovered fragment to obtain pCMV-NlucP, which is a CMV promoter-NlucP fragment-hPEST fragment-poly(A) addition signal ligation. The nucleotide sequence of the NlucP fragment (NlucP gene) and the amino acid sequence of the NlucP protein it encodes are shown in SEQ ID NOs: 28 and 29.

[0158] Next, to construct a vector (pCMV-SceNlucP) expressing the SceNlucP gene (SEQ ID NO: 30), which is a fusion of the SceNluc gene (SEQ ID NO: 9, see AI-1) and the hPEST sequence, pCMV-SceNluc prepared in AI-1 was digested with PpuMI and XbaI to remove the 3' end of SceNluc, including the termination codon, and recover a fragment containing the CMV promoter, the 5' end of SceNluc, and a poly(A) addition signal. Next, pCMV-NlucP was digested with PpuMI and XbaI to recover a fragment containing the 3' end of Nluc and the hPEST sequence. The two recovered fragments were ligated using the DNA Ligation Kit (Mighty Mix) (Takara Bio) to obtain pCMV-SceNlucP, which is composed of the CMV promoter, SceNlucP fragment, hPEST fragment, and poly(A) addition signal. The reading frame of the SceNlucP gene (SEQ ID NO: 30) contains a stop codon TGA introduced upstream of the I-SceI recognition sequence and two stop codons TAG and TAA within the I-SceI recognition sequence. The amino acid sequences encoded by positions 1 to 225 of the SceNlucP gene (the region up to the first stop codon) and positions 235 to 639 (the region downstream of the third stop codon) are shown in SEQ ID NOs: 31 and 32, respectively (however, a polypeptide consisting of the latter sequence is never expressed).

[0159] pCMV-NlucP and pCMV-SceNlucP were purified using a Qiagen Plasmid Plus Midi Kit (Qiagen KK). pCMV-NlucP was used in its circular form, while pCMV-SceNlucP was linearized by digestion with I-SceI (New England Biolabs, Ipswich, MA, USA) before transfection.

[0160]

[0161] A-III-2. Cells and gene transfection, luciferase assay. As described in AI-2.

[0162] A-III-3. Results Figure 11 shows the results of luciferase assays (24 hours after gene transfection) in Nalm-6 cells transfected with the constructed constructs and their MSH2-reverted cells. Constructs with homology in the homologous region of 98% or less exhibited high luciferase activity in MSH2-deficient (MMR-deficient WT Nalm-6 cells). In particular, constructs with homology in the homologous region of 94% exhibited a 154-fold difference in luciferase activity depending on whether or not MSH2 was present.

[0163] Example A-IV: Split SA-Nluc-v3 (3-split & 4-split, 3-times SSA type) (Fig. 12, Fig. 13) A split construct in which a normal Nluc gene is formed by SSA at three sites (3 times) was constructed, and luciferase activity was evaluated.

[0164] A-IV-1-1. Construction of Vectors (3-Part, 3-Time SSA Type) (Figure 12) The 280-bp SceNluc3 fragment (SEQ ID NO: 47, DNA in which the region from positions 151 to 416 in the Nluc gene sequence was replaced with a stop codon (TGA) plus the I-SceI recognition sequence, with a stop codon added to the 3' end) used in pCMV-SceNluc3, the 220-bp iNluc3-1-100 fragment (the region from positions 101 to 320 in the Nluc gene sequence; SEQ ID NO: 48), and the 196-bp iNluc3-2-100 fragment (the region from positions 271 to 466 in the Nluc gene sequence; SEQ ID NO: 49) were amplified by PCR using the pNL1.1[Nluc] Vector (Promega, Madison, WI, USA) as a template. PrimeSTAR HS (trade name) DNA Polymerase (Takara Bio) was used as the polymerase for PCR. The primers used are shown in Table 3 below. The SceNluc3 fragment was amplified by dividing it into two fragments: the 5'SceNluc3 fragment and the 3'SceNluc3 fragment. Nluc-Fw (SEQ ID NO: 1) and 5'SceNluc3-Rv (SEQ ID NO: 33) were used to amplify the 5'SceNluc3 fragment. A sequence for the In-Fusion reaction was added to the 5' end of the 5'SceNluc3 fragment, and a DNA fragment with a structure in which a "stop codon (TGA) + I-SceI recognition sequence" was added to the 3' end was amplified. The 3'SceNluc3 fragment was amplified using 3'SceNluc3-Fw (SEQ ID NO: 34) and Nluc-Rv (SEQ ID NO: 4), resulting in a DNA fragment with an I-SceI recognition sequence added to the 5' end of the 3'SceNluc3 fragment and a stop codon (TAA) and a sequence for the In-Fusion reaction added to the 3' end. The iNluc3-1-100 fragment (SEQ ID NO: 48) was amplified using iNluc3-1-100-Fw (SEQ ID NO: 35) and iNluc3-1-100-Rv (SEQ ID NO: 36). The 50-bp nucleotide sequence at the 5' end of the iNluc3-1-100 fragment amplified by this PCR is 100% homologous to the 50-bp nucleotide sequence at the 3' end of pCMV-SceNluc3 cleaved with I-SceI, and the 50-bp nucleotide sequence at the 3' end is 100% homologous to the 5' end of the iNluc3-2-100 fragment.For the iNluc3-2-100 fragment (SEQ ID NO: 49), iNluc3-2-100-Fw (SEQ ID NO: 37) and iNluc3-2-100-Rv (SEQ ID NO: 38) were used. The 50-bp nucleotide sequence from the 5' end of the iNluc3-2-100 fragment amplified by this PCR is 100% homologous to the 3' end of the iNluc3-1-100 fragment, and the 50-bp nucleotide sequence from the 3' end is 100% homologous to the 50-bp nucleotide sequence from the 5' end of pCMV-SceNluc3 cleaved with I-SceI.

[0165] Next, pIRES (Takara Bio, Figure 2) was digested with NheI and XbaI to remove the IRES region and recover a fragment containing the CMV promoter and poly(A) addition signal. Using the In-Fusion HD Cloning Kit (Takara Bio), the 5'SceNluc3 fragment and the 3'SceNluc3 fragment were inserted downstream of the CMV promoter in the recovered fragment, yielding pCMV-SceNluc3, which is a CMV promoter-SceNluc3 fragment-poly(A) addition signal ligation construct.

[0166] The iNluc3-1 and iNluc3-2 fragments, which have reduced homology between homologous sequences, were amplified by PCR using the pNL1.1[Nluc] Vector (Promega, Madison, WI, USA) as a template. The PCR polymerase used was Tks Gflex (trade name) DNA Polymerase (Takara Bio). Silent mutations were introduced to achieve 90% (5 sites) and 80% (10 sites) homology, respectively (Figures 14-1 to 14-3). The primers used are listed in Table 3 below. For amplification of the iNluc3-1-90 fragment (90% homology, SEQ ID NO: 50), iNluc3-1-90-Fw (SEQ ID NO: 39) and iNluc3-1-90-Rv (SEQ ID NO: 40) were used. For amplification of the iNluc3-1-80 fragment (80% homology, SEQ ID NO: 51), iNluc3-1-80-Fw (SEQ ID NO: 41) and iNluc3-1-80-Rv (SEQ ID NO: 42) were used. iNluc3-2-90-Fw (SEQ ID NO: 43) and iNluc3-2-90-Rv (SEQ ID NO: 44) were used to amplify the uc3-2-90 fragment (90% homology, SEQ ID NO: 52), and iNluc3-2-80-Fw (SEQ ID NO: 45) and iNluc3-2-80-Rv (SEQ ID NO: 46) were used to amplify the iNluc3-2-80 fragment (80% homology, SEQ ID NO: 53).

[0167] pCMV-SceNluc3 was purified using the Qiagen Plasmid Plus Midi Kit (Qiagen KK) and linearized by digestion with I-SceI (New England Biolabs, Ipswich, MA, USA) or I-SceI and AhdI (New England Biolabs) before transfection. The iNluc3-1-100, iNluc3-2-100, iNluc3-1-90, iNluc3-1-80, iNluc3-2-90, and iNluc3-2-80 fragments were purified using the Wizard SV Gel and PCR Clean-Up System (Promega) before transfection.

[0168]

[0169] A-IV-1-2. Construction of vector (four-part, three-times SSA type) (Figure 13) pCMV-SceNluc3 was cleaved at two sites with I-SceI and AhdI (the AhdI recognition sequence is present in the ampicillin resistance gene in the backbone (vector sequence)) to split it into two molecules, which were then combined with various iNluc3-1 and iNluc3-2 fragments to form four-part constructs.

[0170] A-IV-2. Cells and gene transfection, luciferase assay. As described in AI-2.

[0171] A-IV-3. Results The three-part and four-part constructs constructed in this example were constructed so that a normal Nluc gene would be generated when single-strand annealing (SSA) occurred at three sites: between the 3' side of the SceNluc3 fragment and the 5' side of the iNluc3-1 fragment, between the 3' side of the iNluc3-1 fragment and the 5' side of the iNluc3-2 fragment, and between the 3' side of the iNluc3-2 fragment and the 5' side of the SceNluc3 fragment (Figures 12 and 13).

[0172] Since the SSA reaction is inhibited by MMR factors when the homology of the homologous region is reduced, it is thought that the presence or absence of MMR factors will result in differences. The SSA construct constructed in Example AI is designed to form a normal Nluc gene through two SSA reactions, and the difference in Nluc gene expression levels due to the presence or absence of MMR factors was greater than in the conventional SSA construct (WO 2021 / 162121 A1). However, the construct of this example requires SSA to occur at three locations (three times) to form a normal Nluc gene, so it is expected that the difference in Nluc gene expression levels due to the presence or absence of MMR factors will be even greater than the construct in Example AI.

[0173] Figure 15 shows an example of the results of luciferase assays in Nalm-6 cells transfected with the three-part split-three-times SSA construct constructed in this example (Figure 12) and their MSH2-reverted cells (results obtained 4 hours after gene transfection). High Luc activity was observed in MSH2-deficient (MMR-deficient wild-type Nalm-6 cells) constructs with reduced homology in the homologous region (90% and 80%).

[0174] Figure 16 shows the luciferase activity 24 hours after transfection of the 3-split-3 SSA construct. The difference in luciferase activity between constructs with reduced homology in the homologous region and those with or without MSH2 deficiency (MMR deficiency) was more pronounced after 4 hours. In particular, a 102-fold difference in luciferase activity occurred in the construct with 80% homology in the homologous region.

[0175] Figure 17 shows the results of transfecting the 3-split-3 SSA construct into HCT116, a cancer cell line with MMR deficiency due to MLH1 deficiency, and its MLH1 revertant cells, and measuring luciferase activity 4 and 24 hours later. Similar to the presence or absence of mismatch repair deficiency due to MSH2, the presence or absence of mismatch repair deficiency due to MLH1 also confirmed that the difference in gene expression (luciferase activity) widens as a result of reducing the homology of the homologous region.

[0176] The four-part, three-times SSA construct was transfected into Nalm-6 cells and their MSH2-reverted counterparts, and luciferase activity was measured 24 hours later. The results are shown in Figure 18. Although the absolute luciferase activity was reduced to less than one-tenth of that of the three-part, three-times SSA construct, in which pCMV-SceNluc3 was cleaved with I-SceI alone, the difference in luciferase activity between constructs with 90% homology in the homologous region was 53-fold greater depending on whether the MSH2 deficiency (MMR deficiency) was present or not. When the four-part, three-times SSA construct was transfected into HCT116 cells and their MLH1-reverted counterparts, the absolute luciferase activity was also reduced to less than one-tenth of that of the three-part, three-times SSA construct. However, the difference in luciferase activity was greater when the homology in the homologous region was reduced (Figure 19).

[0177] Example AV: Split SA-GeNL-v2 (SceGeNL + iNluc) (Two-Part, Two-Sequence SSA) (Figure 20). The GeNL (Green enhanced nano-lantern) gene encodes a fusion protein of the chemiluminescent protein Nluc and the green fluorescent protein mNeonGreen (Shaner NC et al. Nat Methods. 2013, May 10(5), 407-409). The luminescence intensity of the GeNL protein has been shown to be approximately 1.8-fold higher than that of Nluc due to Forester resonance energy transfer (FRET) between Nluc and mNeonGreen (Suzuki K et al. Nat Commun. 2016, Dec 14;7:13718). Therefore, by using the GeNL gene sequence as the gene sequence formed by SSA in the nucleic acid construct of the present invention, the presence or absence of MMR deficiency can be detected with higher sensitivity than using a chemiluminescent gene alone. Specific examples are shown below.

[0178] AV-1. Vector Construction AV-1-1. GeNL Expression Vector (Figure 20, top panel) To create a construct expressing the GeNL gene, a partial GeNL gene fragment was constructed by artificial gene synthesis (pGeNL) by linking the XhoI recognition sequence to the 5' upstream of the mNeonGreen gene (the region from positions 1 to 678 of the mNeonGreen gene sequence shown in SEQ ID NO: 172) and the region from positions 16 to 79 of the Nluc gene (SEQ ID NO: 7) (including the EcoNI recognition sequence) to the 3' downstream. The constructed pGeNL was digested with XhoI and EcoNI, and the fragment containing the partial GeNL gene fragment was recovered. Next, pCMV-Nluc (Saito S, et al. Nature Commun. 8:16112, 2017) was digested with XhoI and EcoNI to remove the 5' untranslated region of the Nluc gene and positions 1 to 73 of the Nluc gene, resulting in a fragment containing the CMV promoter, poly(A) addition signal, and a portion of the Nluc gene. Using the Mighty Mix DNA Ligation Kit (Takara Bio), a partial GeNL gene fragment was inserted into the fragment, resulting in a GeNL expression vector (pCMV-GeNL) (Figure 20, top panel) consisting of the CMV promoter, GeNL gene fragment, and poly(A) addition signal. The nucleotide sequence of the GeNL gene and the amino acid sequence of the encoded GeNL protein are shown in SEQ ID NOs: 174 and 175, respectively. The GeNL protein prepared here has a structure in which residues 1 to 226 of the mNeonGreen protein (SEQ ID NO: 173) and residues 6 to 171 of the Nluc protein (SEQ ID NO: 8) are linked by two amino acid residues (GF).

[0179] AV-1-2. Split SA-GeNL-v2 (Figure 20, bottom panel) To prepare pCMV-SceGeNL, pCMV-SceNluc prepared in Example AI was digested with EcoNI and XbaI to recover a fragment containing the SceNluc gene fragment. Next, pCMV-GeNL was digested with EcoNI and XbaI to remove positions 743 to 1182 and the termination codon in the GeNL gene, recovering a fragment containing the CMV promoter, poly(A) addition signal, and a portion of the GeNL gene. Using the DNA Ligation Kit "Mighty Mix" (Takara Bio), the SceNluc gene fragment was inserted into the recovered fragment, yielding a vector (pCMV-SceGeNL) expressing SceGeNL, which is linked to the CMV promoter, SceGeNL gene fragment, and poly(A) addition signal. The SceGeNL gene is a DNA structure in which the region from positions 892 to 912 in the GeNL gene has been replaced with a "stop codon (TGA) + I-SceI recognition sequence." The nucleotide sequence of the SceGeNL gene is shown in SEQ ID NO: 176, the amino acid sequence encoded by positions 1 to 894 (the region up to the first stop codon) is shown in SEQ ID NO: 177, and the amino acid sequence encoded by positions 904 to 1185 (the region downstream of the third stop codon) is shown in SEQ ID NO: 178 (however, a polypeptide consisting of this sequence is not expressed).

[0180] pUC-iNluc, which has sequences homologous to positions 685 to 891 and 913 to 1143 in the SceGeNL gene, and pUC-iNluc (90% homology) and pUC-iNluc (80% homology), which have reduced homology between the homologous sequences, were constructed as described in AI-1.

[0181] pCMV-SceGeNL was purified using the Qiagen Plasmid Plus Midi Kit (Qiagen KK) and linearized by digestion with I-SceI (New England Biolabs, Ipswich, MA, USA) before transfection. pUC19, pUC-iNluc, pUC-iNluc (90% homology), and pUC-iNluc (80% homology) were purified using the Qiagen Plasmid Plus Midi Kit (Qiagen KK) and linearized by digestion with BamHI and EcoRI. The iNluc-containing fragment was purified using the Wizard SV Gel and PCR Clean-Up System (Promega) before use. Digestion of pUC-iNluc with BamHI and EcoRI yielded DNA without additional sequences (Figure 20, bottom panel).

[0182] AV-2. Cells and gene transfection, luciferase assay. As described in AI-2.

[0183] AV-3. Results. The construct constructed in this example was designed to generate a normal GeNL gene upon homologous recombination (HR) between the SceGeNL and iNluc fragments or single-strand annealing (SSA) at two sites on the 5' and 3' ends of each fragment. HR occurs when the homology between the homologous regions is sufficiently high, but reducing the homology between the homologous regions prevents HR, resulting in normal GeNL gene generation only upon SSA. Because the SSA reaction is inhibited by MMR factors, the presence or absence of MMR factors is thought to result in differences in GeNL gene expression levels. Examples of luciferase assay results for Nalm-6 cells transfected with the constructed construct and their MSH2-reverted cells are shown in Figures 21-1 and 21-2. Figure 21-1 shows luciferase activity 4 hours after gene transfection, and Figure 21-2 shows luciferase activity 24 hours after gene transfection. Constructs with reduced homology in the 90% and 80% homologous regions demonstrated high Luc activity in MSH2-deficient (MMR-deficient wild-type Nalm-6 cells). Because the luminescence intensity of GeNL protein is approximately 1.8-fold higher than that of Nluc protein, constructs using GeNL are expected to detect luciferase activity with higher sensitivity than Nluc. Indeed, as shown in Figure 21-1, with the GeNL construct, there was a clear difference in luciferase activity between wild-type and MSH2-reverted cells at 4 hours after transfection, with a difference of approximately 5-fold. In contrast, with the Nluc construct, the difference in luciferase activity between MSH2-deficient and non-deficient cells was approximately 2-fold at 2 hours after transfection. Twenty-four hours after transfection of the constructs with reduced homology in the homologous region (90% and 80%), a clear difference in luciferase activity was observed between the wild-type strain and the MSH2-revertant cells, regardless of whether the GeNL or Nluc construct was used (Fig. 21-2). In particular, the GeNL construct with 80% homology in the homologous region showed a 130-fold difference in luciferase activity between the presence and absence of MSH2 deficiency.This confirmed that, like the construct using Nluc, the construct using GeNL can also be used to detect the presence or absence of MMR deficiency in cells.

[0184] Example B: Split DR-Nluc-homeo DTA-linked (SceNluc + iNluc) (two splits, two SSAs) (Figure 22) A two-split construct was constructed in which a normal Nluc gene was formed by two SSAs (two SSAs), and Nluc and a DTA linked downstream of it were expressed in a polycistronic manner, and the cytocidal effect was evaluated.

[0185] To prepare pCMV-Nluc-2A-DTA, a DNA fragment linking the Nluc gene (SEQ ID NO: 7) and a sequence (SEQ ID NO: 54) encoding the 2A peptide (SEQ ID NO: 55) derived from Thosea asigna virus was synthesized by artificial gene synthesis (pNluc-2A, a plasmid constructed by inserting the artificially synthesized Nluc-2A fragment into the EcoRI-HindIII site of pUC57 (GenScript Japan, Tokyo, Japan)). The artificially synthesized pNluc-2A was digested with BamHI to linearize it. pCMV-DTA (a construct expressing the normal DT-A gene (SEQ ID NO: 56) under the control of the CMV promoter; see Example DI for the preparation method) was digested with BamHI to obtain a DTA gene fragment. This fragment was then ligated to the linearized fragment using the Mighty Mix DNA Ligation Kit to obtain pNluc-2A-DTA. The constructed pNluc-2A-DTA was digested with NheI and MfeI to excise a DNA fragment in which the Nluc and DTA genes were linked via a 2A peptide sequence. Next, pIRES (Takara Bio, Figure 2) was digested with NheI and MfeI to excise a fragment containing the CMV promoter and poly(A) addition signal. Using the Mighty Mix DNA Ligation Kit, the Nluc-2A-DTA fragment was inserted downstream of the CMV promoter to obtain pCMV-Nluc-2A-DTA, which is composed of the CMV promoter, Nluc, 2A peptide sequence, DTA, and poly(A) addition signal. The nucleotide sequence of the Nluc-2A peptide sequence-DTA fragment is shown in SEQ ID NO: 58, and the amino acid sequence encoded by this is shown in SEQ ID NO: 59.

[0186] Furthermore, to prepare a vector in which the SceNluc gene and DTA gene were linked via a 2A peptide sequence, pCMV-Nluc-2A-DTA was digested with EcoNI and PpuMI to cleave the Nluc gene at two sites, removing the central portion of the gene, and recovering a fragment containing the CMV promoter, the Nluc gene with the central portion deleted (5' and 3' ends of the Nluc gene), and 2A-DTA. Next, pCMV-SceNluc prepared in AI-1 was digested with EcoNI and PpuMI to cleave the SceNluc gene at two sites (upstream and downstream of the I-SceI recognition sequence), recovering a partial fragment of the SceNluc gene containing the I-SceI recognition sequence. Finally, using the DNA Ligation Kit (Mighty Mix), a partial fragment of the SceNluc gene was inserted into the central deletion site of the Nluc gene from the recovered fragment, yielding pCMV-SceNluc-2A-DTA, which ligated [CMV promoter]-[SceNluc]-[2A peptide sequence]-[DTA]-[poly(A) addition signal]. The nucleotide sequence of the [SceNluc]-[2A peptide sequence]-[DTA] portion is shown in SEQ ID NO: 60, the amino acid sequence encoded by positions 1 to 225 (up to the first stop codon) is shown in SEQ ID NO: 61, and the amino acid sequence encoded by positions 235 to 1209 (downstream of the third stop codon) is shown in SEQ ID NO: 62 (although a polypeptide consisting of this sequence is not expressed). Positions 1 to 513 of SEQ ID NO: 60 represent the SceNluc gene sequence, positions 550 to 603 represent the 2A peptide coding sequence, and positions 619 to 1209 represent the DTA gene sequence. Furthermore, residues 106 to 123 in SEQ ID NO: 62 are the 2A peptide sequence. During intracellular protein translation, the 2A peptide skips a peptide bond between the C-terminal glycine and proline (between Gly at position 122 and Pro at position 123 in SEQ ID NO: 62). Therefore, by including a 2A peptide sequence between two protein sequences, two proteins can be expressed polycistronically from one ORF.

[0187] The prepared plasmids were purified using a Qiagen Plasmid Plus Midi Kit (Qiagen KK), and then pCMV-Nluc-2A-DTA was used as a circular plasmid, while pCMV-SceNluc-2A-DTA was used for transfection after linearization by digestion with I-SceI (New England Biolabs, Ipswich, MA, USA). For iNluc, pUC-iNluc with 100% homology and reduced homology prepared in Example AI were linearized by digestion with BamHI and EcoRI.

[0188] B-2. Cells, gene transfer, and evaluation of cytotoxicity. For details on cells and gene transfer, see AI-2. above. The cytotoxicity was evaluated as follows: 2 x 10 6 Each construct (1 μg) was transfected into 1 x 10 Nalm-6 cells (wild-type and its Msh2-revertant cells). 5 The cells were seeded into a 24-well dish at 1000 cells / ml and cultured for 96 hours, after which the cell viability was measured using CellTiter-Glo (CellTiter-Glo® Luminescent Cell Viability Assay, Promega).

[0189] B-3. ​​Results: This construct results in the formation of a normal Nluc gene through recombination, leading to transcription of the DTA gene, and the expression of the Nluc and DTA proteins via the ribosomal skip mechanism of the 2A peptide. By examining Nluc activity using a luciferase assay, cells that would eventually die due to the action of the DTA protein can be detected.

[0190] Figure 23-1 shows the comparison of viability between Nalm-6 cells and their MSH2-reverted counterparts 96 hours after construct transfection, and Figure 23-2 shows the comparison of viability between HCT116 cells and their MLH1-reverted counterparts 72 hours after construct transfection. Constructs with reduced homology in the homologous region were confirmed to selectively kill MMR-deficient cells (WT Nalm-6 cells and WT HCT116 cells). When a construct with 90% homology was transfected, approximately 20% of MMR-deficient cells (WT Nalm-6 cells and WT HCT116 cells) survived, but there was little toxicity to MMR-normal cells (MSH2-reverted cells (Nalm-6 cells) and MLH1-reverted cells (HCT116 cells)), suggesting their potential as anticancer therapeutics with minimal side effects.

[0191] Example C: Split DR-Nluc-homeoTK-linked (SceNluc + iNluc) (two splits, two SSAs) (Figure 24) A two-split construct was constructed in which a normal Nluc gene was formed by two SSAs (two SSAs), and Nluc and HSV-TK linked downstream of it were expressed in a polycistronic manner, and the cytocidal effect was evaluated.

[0192] C-1. Vector Construction To construct a vector (pCMV-TK) for expressing the HSV-TK gene, the ORF of the HSV-TK gene was amplified by PCR using an HSV-TK gene fragment (GenBank: V00470.1, Kobayashi et al. (2001) Decreased topoisomerase IIalpha expression confers increased resistance to ICRF-193 as well as VP-16 in mouse embryonic stem cells. Cancer Lett. 166(1): 71-77; the nucleotide sequence of the coding region of the HSV-TK gene and the amino acid sequence of the encoded TK protein are shown in SEQ ID NOs: 63 and 64) as a template. The PCR polymerase used was PrimeSTAR HS™ DNA Polymerase (Takara Bio), and the primers used were TK-Fw (TAATACGACTCACTATAGGCTAGCCTCGAGATCCACCGGTCATGGCTTCGTACCCCGGCCATC, SEQ ID NO: 65) and TK-Rv (CCGCCCCGACTCTAGAGTCGCGGTCAGTTAGCCTCCCCCATCTCC, SEQ ID NO: 66). Next, pIRES (Takara Bio, Figure 2) was digested with NheI and XbaI to recover a fragment containing the CMV promoter and polyadenylation signal. Using the In-Fusion HD Cloning Kit (Takara Bio), the TK gene fragment was inserted downstream of the CMV promoter to obtain pCMV-TK, which is a CMV promoter-HSV-TK gene fragment-polyadenylation signal ligation. The resulting plasmid was purified using the Qiagen Plasmid Plus Midi Kit (Qiagen KK) and used for transfection.

[0193] To construct pCMV-Nluc-2A-TK, pCMV-Nluc-2A-DTA (prepared in Section B-1) was digested with AgeI and MfeI to remove the DT-A gene and poly(A) addition signal, yielding a fragment containing the CMV promoter, Nluc gene, and 2A peptide sequence. Next, pCMV-TK was digested with AgeI and MfeI to yield a fragment containing the HSV-TK gene and poly(A) addition signal. The HSV-TK gene and poly(A) addition signal were inserted downstream of the 2A peptide sequence using the Mighty Mix DNA Ligation Kit, yielding pCMV-Nluc-2A-TK, which is composed of the CMV promoter, Nluc, 2A peptide sequence, HSV-TK gene, and poly(A) addition signal.

[0194] To prepare a vector in which the SceNluc gene and HSV-TK gene were linked via a 2A peptide sequence, pCMV-SceNluc-2A-DTA prepared in B-1 was digested with AgeI and MfeI to remove the DT-A gene and poly(A) addition signal, and a fragment containing the CMV promoter, SceNluc gene, and 2A peptide sequence was recovered. Next, pCMV-TK was digested with AgeI and MfeI to recover a fragment containing the HSV-TK gene and poly(A) addition signal. The HSV-TK gene and poly(A) addition signal were then inserted downstream of the 2A peptide sequence using the Mighty Mix DNA Ligation Kit to obtain pCMV-SceNluc-2A-TK, which is composed of the CMV promoter, SceNluc, 2A peptide sequence, HSV-TK gene, and poly(A) addition signal. The nucleotide sequence of the DNA fragment (SceNluc-2A-TK portion) in which the SceNluc gene, 2A peptide sequence, and TK gene are linked is shown in SEQ ID NO: 67, and the amino acid sequences encoded by positions 1 to 225 (the region up to the first stop codon) and positions 235 to 1,749 (the region downstream of the third stop codon) are shown in SEQ ID NOs: 68 and 69, respectively (however, a polypeptide consisting of the amino acid sequence of SEQ ID NO: 69 is not expressed). In SEQ ID NO: 67, positions 1 to 513 represent the SceNluc gene sequence, positions 550 to 603 represent the 2A peptide-coding sequence, and positions 619 to 1,749 represent the TK gene sequence. Furthermore, residues 106 to 123 of SEQ ID NO: 69 represent the 2A peptide sequence. During intracellular protein translation, the 2A peptide causes a peptide bond skip between its C-terminal glycine and proline (between Gly at position 122 and Pro at position 123 in SEQ ID NO: 69). Therefore, by including a 2A peptide sequence between two protein sequences, two proteins can be expressed polycistronically from a single ORF.

[0195] The prepared plasmids were purified using a Qiagen Plasmid Plus Midi Kit (Qiagen KK), and then pCMV-Nluc-2A-TK was used in its circular form, while pCMV-SceNluc-2A-TK was used for transfection after linearization by digestion with I-SceI (New England Biolabs, Ipswich, MA, USA). For iNluc, pUC-iNluc with 100% homology and reduced homology prepared in Example AI were linearized by digestion with BamHI and EcoRI.

[0196] C-2. Cells, gene transfection, and evaluation of cell killing effect For cells and gene transfection, see AI-2 above. Cell killing effect was evaluated as follows: 2 x 10 6 Each construct (1 μg) was transfected into 1 x 10 Nalm-6 cells (wild-type and its Msh2-reverted cells). After culturing the transfected cells for 48 hours, each construct (1 μg) was transfected again and 1 x 10 cells were cultured. 5 Cells were seeded into 24-well dishes at 1000 cells / ml, and ganciclovir (Fujifilm Wako Pure Chemical Industries, Ltd.) was added to the medium at a final concentration of 500 nM. After 96 hours of culture, cell viability was measured using CellTiter-Glo (CellTiter-Glo® Luminescent Cell Viability Assay, Promega).

[0197] C-3. Results Figure 25 shows the results of comparing the survival rates of Nalm-6 cells and their MSH2-reverted counterparts 96 hours after construct transfection. The survival rates are shown relative to the survival rate of cells not transfected with DNA (No DNA) cultured in the presence of ganciclovir (GANC), which was set at 100%. Constructs with 98% to 94% homology in the homologous region were confirmed to be able to selectively kill MMR-deficient cells (WT Nalm-6 cells and WT HCT116 cells).

[0198] Example DI: Split DR-DTA-homeo (SceDTA + iDTA) (two splits, two SSAs) (Figure 26) A construct was constructed in which the normal DT-A gene was formed by two SSAs (two times), and the cytocidal effect was evaluated.

[0199] DI-1. Vector Construction The 588-bp SceDTA fragment used in pCMV-Sce-DTA (DNA in which the region from positions 133 to 153 in the DT-A gene was replaced with a stop codon (TGA) plus an I-SceI recognition sequence, with a stop codon added to the 3' end; SEQ ID NO: 76) was obtained by PCR amplification using pMC1DT-ApA (KURABO, Osaka, Japan) as a template. PrimeSTAR HS (trade name) DNA Polymerase (Takara Bio) was used as the polymerase for PCR. The primers used are listed in Table 4 below. The SceDTA fragment was amplified separately into two fragments: the 5'SceDTA fragment and the 3'SceDTA fragment. The 5'SceDTA fragment was amplified using DTA-Fw (SEQ ID NO: 70) and Sce-5'DTA-Rv (SEQ ID NO: 71). This resulted in a DNA fragment with a sequence for the In-Fusion reaction added to the 5' end of the 5'SceDTA fragment and a stop codon (TGA) plus an I-SceI recognition sequence added to the 3' end. The 3'SceDTA fragment was amplified using Sce-3'DTA-Fw (SEQ ID NO: 72) and DTA-Rv (SEQ ID NO: 73). This resulted in a DNA fragment with a sequence for the In-Fusion reaction added to the 5' end of the 3'SceDTA fragment and a stop codon (TGA) and an I-SceI recognition sequence added to the 3' end. Next, pIRES (Takara Bio, Figure 2) was digested with NheI and XbaI to remove the IRES region, and a fragment containing the CMV promoter and poly(A) addition signal was recovered. Using the In-Fusion HD Cloning Kit (Takara Bio), the 5'SceDTA fragment and the 3'SceDTA fragment were inserted downstream of the recovered CMV promoter to obtain pCMV-Sce-DTA, which is a CMV promoter-SceDTA fragment-poly(A) addition signal ligation. The amino acid sequences encoded by positions 1 to 135 (the region up to the first stop codon) and positions 145 to 591 (the region downstream of the third stop codon) of the SceDTA fragment (SEQ ID NO: 76) are shown in SEQ ID NOs: 77 and 78, respectively (although a polypeptide consisting of the latter sequence will not be expressed).

[0200] The 588-bp DT-A gene fragment used in pCMV-DTA was obtained by PCR amplification using pMC1DT-ApA (KURABO, Osaka, Japan) as a template. PrimeSTAR HS™ DNA Polymerase (Takara Bio) was used as the PCR polymerase. Using DTA-Fw (SEQ ID NO: 70) and DTA-Rv (SEQ ID NO: 73) as primers, a DNA fragment containing sequences for the In-Fusion reaction at the 5' and 3' ends was amplified. Next, pIRES (Takara Bio, Figure 2) was digested with NheI and XbaI to remove the IRES region and recover a fragment containing the CMV promoter and poly(A) addition signal. Using the In-Fusion HD Cloning Kit (Takara Bio), the DTA gene fragment was inserted downstream of the CMV promoter in the recovered fragment, yielding pCMV-DTA, which contains the CMV promoter, DTA gene, and poly(A) addition signal.

[0201] The 221-bp iDTA2 fragment (positions 33 to 253 in the DT-A gene, SEQ ID NO: 79), which has a homology to pCMV-Sce-DTA, was amplified by PCR using pCMV-DTA as a template. PrimeSTAR HS (trade name) DNA Polymerase (Takara Bio) was used as the polymerase, and iDTA2-Fw (SEQ ID NO: 74) and iDTA2-Rv (SEQ ID NO: 75) were used as primers. Next, pUC19 (Takara Bio) was digested with SmaI in the multicloning site, and the iDTA2 fragment was inserted using the DNA Ligation Kit (Mighty Mix) (Takara Bio) to obtain pUC-iDTA2.

[0202]

[0203] To create a vector with reduced homology between homologous sequences, DNA fragments were created by artificial gene synthesis, with silent mutations introduced into the regions 1-100 and 122-221 of the iDTA2 sequence (221 bp). The silent mutations were 99% (2 sites; 1 / 100 on the 5' side and 1 / 100 on the 3' side), 98% (4 sites; 2 / 100 on the 5' side and 2 / 100 on the 3' side), 97% (6 sites; 3 / 100 on the 5' side and 3 / 100 on the 3' side), 96% (8 sites; 4 / 100 on the 5' side and 4 / 100 on the 3' side), 95% (10 sites; 5 / 100 on the 5' side and 5 / 100 on the 3' side), 94% (12 sites; 6 / 100 on the 5' side and 6 / 100 on the 3' side), 90% (20 sites; 10 / 100 on the 5' side and 10 / 100 on the 3' side), and 80% (20 sites; 5 / 100 on the 5' side). Silent mutations were introduced so that the nucleotide sequence was distributed as evenly as possible throughout the entire sequence, with 99%, 98%, 97%, 96%, 95%, 94%, 90%, 80%, and 70% homology (62 sites; 5' 20 / 100, 3' 20 / 100). The sequences of iDTA2 fragments with 99%, 98%, 97%, 96%, 95%, 94%, 90%, 80%, and 70% homology are shown in SEQ ID NOs: 81–89, respectively. The vector containing the nine artificially synthesized iDTA2 fragments with different mutations was digested with BamHI (Takara Bio) and EcoRI (Takara Bio), and the fragments containing iDTA2 with each mutation were excised. Next, pUC19 (Takara Bio) was digested with BamHI and EcoRI, and each iDTA2 fragment carrying a mutation was inserted using the DNA Ligation Kit <Mighty Mix> (Takara Bio) to obtain pUC-iDTA2 (99% homology), pUC-iDTA2 (98% homology), pUC-iDTA2 (97% homology), pUC-iDTA2 (96% homology), pUC-iDTA2 (94% homology), pUC-iDTA2 (90% homology), pUC-iDTA2 (80% homology), and pUC-iDTA2 (70% homology), respectively.

[0204] pCMV-Sce-DTA was purified using the Qiagen Plasmid Plus Midi Kit (Qiagen KK) and linearized by digestion with I-SceI (New England Biolabs) before transfection. pCMV-DTA was purified using the Qiagen Plasmid Plus Midi Kit (Qiagen KK) and used in its circular form for transfection. pUC-iDTA2 was purified using the Qiagen Plasmid Plus Midi Kit (Qiagen KK) and linearized by digestion with BamHI and EcoRI. The iDTA2-containing fragment was purified using the Wizard SV Gel and PCR Clean-Up System (Promega) before use.

[0205] DI-2. Cells, gene transfer, and evaluation of cytotoxicity. For details on cells and gene transfer, see AI-2. Evaluation of cytotoxicity using Nalm-6 cells (wild-type and its Msh2-reverted cells) was performed as follows. 6 Each construct (1 μg) was transfected into 1 x 10 cells. After culturing the transfected cells for 48 hours, each construct (1 μg) was transfected again and 1 x 10 cells were cultured. 5 The cells were seeded into a 24-well dish at 5 x 10 cells / ml and cultured for 96 hours. Afterwards, cell viability was measured using CellTiter-Glo (CellTiter-Glo® Luminescent Cell Viability Assay, Promega). The cytotoxic effect on HCT116 cells (wild-type and its Mlh1 revertant cells) was evaluated as follows. 4Cells were seeded into 24-well dishes, cultured overnight, and then transfected with 1 μg of each construct. After 24 hours of culture, the transfected cells were transfected again with 1 μg of each construct. After 72 hours of culture, cell viability was measured using CellTiter-Glo (CellTiter-Glo® Luminescent Cell Viability Assay, Promega).

[0206] DI-3. Results Figure 28 shows the results of a comparison of the survival rates between Nalm-6 cells and their MSH2-reverted cells 96 hours after construct transfection. Figure 29 shows the results of a comparison of the survival rates between HCT116 cells and their MLH1-reverted cells 72 hours after construct transfection. It was confirmed that even when the DT-A gene was used as a recombination substrate, MMR-deficient cells (WT Nalm-6 cells and WT HCT116 cells) could be selectively killed by reducing the homology of the homologous region.

[0207] Example D-II: Split SA-DTA-v3 (3-split & 4-split, 3-time SSA type) (Figures 30, 31) A split construct was constructed in which a normal DT-A gene was formed by SSA at three locations (3 times), and the cytocidal effect was evaluated.

[0208] D-II-1-1. Construction of vectors (three-part, three-times SSA type) (Figure 30) The 288-bp SceDTA3 fragment (DNA in which the region from positions 131 to 464 in the DT-A gene sequence was replaced with a stop codon (TGA) plus the I-SceI recognition sequence, with a stop codon added to the 3' end; SEQ ID NO: 108), the 251-bp iDTA3-1-100 fragment (the region from positions 81 to 331 in the DT-A gene sequence; SEQ ID NO: 112), and the 283-bp iDTA3-2-100 fragment (the region from positions 232 to 514 in the DT-A gene sequence; SEQ ID NO: 113) used in pCMV-SceDTA3 were amplified by PCR using pMC1DT-ApA (KURABO, Osaka, Japan) as a template. PrimeSTAR HS (trade name) DNA Polymerase (Takara Bio) was used as the polymerase for PCR. The primers used are shown in Table 5 below. The SceDTA3 fragment was amplified separately into two fragments: the 5'SceDTA3 fragment and the 3'SceDTA3 fragment. The 5'SceDTA3 fragment was amplified using DTA-Fw (SEQ ID NO: 70) and 5'SceDTA3-Rv (SEQ ID NO: 90). A sequence for the In-Fusion reaction was added to the 5' end of the 5'SceDTA3 fragment, and a DNA fragment with a structure in which a stop codon (TGA) + I-SceI recognition sequence was added to the 3' end was amplified. The 3'SceDTA3 fragment was amplified using 3'SceDTA3-Fw (SEQ ID NO: 91) and DTA-Rv (SEQ ID NO: 73), resulting in a DNA fragment with a termination codon (TGA) + I-SceI recognition sequence added to the 5' end of the 3'SceDTA3 fragment and a sequence for In-Fusion reaction added to the 3' end. The iDTA3-1-100 fragment was amplified using iDTA3-1-100-Fw (SEQ ID NO: 92) and iDTA3-1-100-Rv (SEQ ID NO: 93). The 50-bp nucleotide sequence at the 5' end of the iDTA3-1-100 fragment amplified by this PCR is 100% homologous to the 50-bp nucleotide sequence at the 3' end of pCMV-SceDTA3 digested with I-SceI, and the 100-bp nucleotide sequence at the 3' end is 100% homologous to the 5' end of the iDTA3-2-100 fragment.The iDTA3-2-100 fragments used were iDTA3-2-100-Fw (SEQ ID NO: 94) and iDTA3-2-100-Rv (SEQ ID NO: 95). The 5'-terminal 100-bp nucleotide sequence of the iDTA3-2-100 fragment amplified by this PCR is 100% identical to the 3'-terminal nucleotide sequence of the iDTA3-1-100 fragment, and the 3'-terminal 50-bp nucleotide sequence is 100% identical to the 5'-terminal 50-bp of pCMV-SceDTA3 digested with I-SceI.

[0209] Next, pIRES (Takara Bio, Figure 2) was digested with NheI and XbaI to remove the IRES region, and a fragment containing the CMV promoter and poly(A) was recovered. Using the In-Fusion HD Cloning Kit (Takara Bio), the 5'SceDTA3 fragment and the 3'SceDTA3 fragment were inserted downstream of the CMV promoter of the recovered fragment to obtain pCMV-SceDTA3, which contains the CMV promoter, SceDTA3 fragment, and poly(A) addition signal. The amino acid sequences encoded by positions 1 to 135 (the region up to the first stop codon), 136 to 156 (the region up to the second stop codon), and 157 to 288 (the region downstream of the third stop codon) of the SceDTA3 fragment (SEQ ID NO: 108) are shown in SEQ ID NOs: 109 to 111, respectively (note that polypeptides consisting of the sequences of SEQ ID NOs: 110 to 111 will not be expressed).

[0210] The iDTA3-1 and iDTA3-2 fragments, which had reduced homology between homologous sequences, were obtained by PCR amplification using pCMV-DTA prepared in Example DI as a template. The PCR polymerase used was Tks Gflex (trade name) DNA Polymerase (Takara Bio). Silent mutations were introduced to achieve homology of 98% (one site for a 50-bp homologous region and two sites for a 100-bp homologous region), 94% (three sites for a 50-bp homologous region and six sites for a 100-bp homologous region), and 90% (five sites for a 50-bp homologous region and ten sites for a 100-bp homologous region) (Figures 32-1 to 32-3). The primers used are listed in Table 5 below. iDTA3-1-98-Fw (SEQ ID NO: 96) and iDTA3-1-98-Rv (SEQ ID NO: 97) were used to amplify the iDTA3-1-98 fragment (98% homology, SEQ ID NO: 114), iDTA3-1-94-Fw (SEQ ID NO: 98) and iDTA3-1-94-Rv (SEQ ID NO: 99) were used to amplify the iDTA3-1-94 fragment (94% homology, SEQ ID NO: 115), and iDTA3-1-90-Fw (SEQ ID NO: 100) and iDTA3-1-90-Rv (SEQ ID NO: 101) were used to amplify the iDTA3-1-90 fragment (90% homology, SEQ ID NO: 116). iDTA3-2-98-Fw (SEQ ID NO: 102) and iDTA3-2-98-Rv (SEQ ID NO: 103) were used to amplify the iDTA3-2-98 fragment (98% homology, SEQ ID NO: 117). iDTA3-2-94-Fw (SEQ ID NO: 104) and iDTA3-2-94-Rv (SEQ ID NO: 105) were used to amplify the iDTA3-2-94 fragment (94% homology, SEQ ID NO: 118). iDTA3-2-90-Fw (SEQ ID NO: 106) and iDTA3-2-90-Rv (SEQ ID NO: 107) were used to amplify the iDTA3-2-90 fragment (90% homology, SEQ ID NO: 119).

[0211] pCMV-SceDTA3 was purified using the Qiagen Plasmid Plus Midi Kit (Qiagen KK) and linearized by digestion with I-SceI (New England Biolabs, Ipswich, MA, USA) before transfection. The iDTA3-1-100, iDTA3-2-100, iDTA3-1-98, iDTA3-1-94, iDTA3-1-90, iDTA3-2-98, iDTA3-2-94, and iDTA3-2-90 fragments were purified using the Wizard SV Gel and PCR Clean-Up System (Promega) before transfection.

[0212]

[0213] D-II-1-2. Construction of vector (four-part, three-times SSA type) (Figure 31) pCMV-SceDTA3 was cleaved at two sites with I-SceI and AhdI (the AhdI recognition sequence is present in the ampicillin resistance gene in the backbone (vector sequence)) to split it into two molecules, which were then combined with various iDTA3-1 and iDTA3-2 fragments to form four-part constructs.

[0214] D-II-2. Cells and gene transfer, evaluation of cell killing effect See DI-2.

[0215] D-II-3. Results The three-part and four-part constructs constructed in this example were constructed so that a normal Nluc gene would be generated when single-strand annealing (SSA) occurred at three sites: between the 3' side of the SceDTA3 fragment and the 5' side of the iDTA3-1 fragment, between the 3' side of the iDTA3-1 fragment and the 5' side of the iDTA3-2 fragment, and between the 3' side of the iDTA3-2 fragment and the 5' side of the SceDTA3 fragment (Figures 30 and 31).

[0216] The left side of Figure 33 shows the results of comparing the survival rates of Nalm-6 cells and their MSH2-reverted cells 96 hours after transfection with a three-fold SSA construct. The right side of Figure 33 shows the results of comparing the survival rates of HCT116 cells and their MLH1-reverted cells 72 hours after transfection with the construct. Constructs with reduced homology in the homologous region showed differences in survival rates depending on the presence or absence of MMR factors. In particular, transfection with a construct with 94% homology showed almost no toxicity to MMR-normal cells (MSH2-reverted cells (Nalm-6 cells) and MLH1-reverted cells (HCT116 cells)). This indicates that the constructs of the present invention can selectively kill MMR-deficient cells (WT Nalm-6 cells and WT HCT116 cells).

[0217] The four-part, three-times SSA construct was introduced into Nalm-6 cells and their MSH2-reverted counterparts, and the survival rates were compared 96 hours later. The results are shown in Figure 34. Although the killing effect was lower than that of the three-part construct, the 98% homology construct resulted in a 3.3-fold difference in survival rate depending on whether the MSH2 deficiency (MMR deficiency) was present. These results demonstrate that MMR-deficient cells can be selectively killed even when the number of gene fragments required for normal DTA gene formation is increased.

[0218] Example EI: Split SA-TK-v3 (four-split, four-times SSA type) (Figure 35) A split construct was constructed in which the normal HSV-TK gene was formed by four SSA sites (four times), and its cytocidal effect was evaluated.

[0219] EI-1. Construction of vector (Fig. 35) The 330 bp SceTK3 fragment used in pCMV-SceTK3 (DNA structure in which the region from positions 137 to 968 in the HSV-TK gene sequence was replaced with a "stop codon (TGA) + I-SceI recognition sequence," with a stop codon added to the 3' end, SEQ ID NO: 148), the 362 bp iTK3-1-100 fragment (the region from positions 87 to 448 in the HSV-TK gene sequence, SEQ ID NO: 152), the 412 bp iTK3-1-100 fragment (the region from positions 87 to 448 in the HSV-TK gene sequence, SEQ ID NO: 153), the 412 bp iTK3-1-100 fragment (the region from positions 87 to 448 in the HSV-TK gene sequence, SEQ ID NO: 154), the 412 bp iTK3-1-100 fragment (the region from positions 87 to 448 in the HSV-TK gene sequence, SEQ ID NO: 155), the 412 bp iTK3-1-100 fragment (the region from positions 87 to 448 in the HSV-TK gene sequence, SEQ ID NO: 156), the 412 bp iTK3-1-100 fragment (the region from positions 87 to 448 in the HSV-TK gene sequence, SEQ ID NO: 157), the 412 bp iTK3-1-100 fragment (the region from positions 87 to 448 in the HSV-TK gene sequence, SEQ ID NO: 158), the 412 bp iTK3-1-100 fragment (the region from positions 87 to 448 in the HSV-TK gene sequence, SEQ ID The 358-bp iTK3-2-100 fragment (the region from positions 349 to 760 in the HSV-TK gene sequence; SEQ ID NO: 153) and the 358-bp iTK3-3-100 fragment (the region from positions 661 to 1018 in the HSV-TK gene sequence; SEQ ID NO: 154) were amplified by PCR using the HSV-TK gene fragment (GenBank: V00470.1, Kobayashi et al. (2001) Decreased topoisomerase IIalpha expression confers increased resistance to ICRF-193 as well as VP-16 in mouse embryonic stem cells. Cancer Lett. 166(1): 71-77; the nucleotide sequence of the coding region of the HSV-TK gene and the amino acid sequence of the encoded TK protein are shown in SEQ ID NOs: 63 and 64, respectively) as a template. PrimeSTAR HS (trade name) DNA Polymerase (Takara Bio) was used as the polymerase for PCR. The primers used are shown in Table 6 below. The SceTK3 fragment was amplified separately into two fragments: the 5'SceTK3 fragment and the 3'SceTK3 fragment. The 5'SceTK3 fragment was amplified using Sce-5'TK-Fw (SEQ ID NO: 120) and 5'SceTK3-Rv (SEQ ID NO: 121), resulting in the amplification of a DNA fragment with a sequence for the In-Fusion reaction added to the 5' end of the 5'SceTK3 fragment and a "stop codon (TGA) + I-SceI recognition sequence" added to the 3' end.The 3'SceTK3 fragment was amplified using 3'SceTK3-Fw (SEQ ID NO: 122) and Sce-3'TK-Rv (SEQ ID NO: 123), resulting in a DNA fragment with a termination codon (TGA) + I-SceI recognition sequence added to the 5' end of the 3'SceTK3 fragment and a sequence for In-Fusion reaction added to the 3' end. The iTK3-1-100 fragment was amplified using iTK3-1-100-Fw (SEQ ID NO: 124) and iTK3-1-100-Rv (SEQ ID NO: 125). The 50-bp nucleotide sequence from the 5' end of the iTK3-1-100 fragment amplified by this PCR is 100% identical to the 50-bp nucleotide sequence from the 3' end of pCMV-SceTK3 digested with I-SceI, and the 100-bp nucleotide sequence from the 3' end is 100% identical to the 5' end of the iTK3-2-100 fragment (iTK3-2-100-Fw (SEQ ID NO: 126) and iTK3-2-100-Rv (SEQ ID NO: 127) were used as the iTK3-2-100 fragments. The 5'-terminal 100-bp nucleotide sequence of the iTK3-2-100 fragment amplified by this PCR is 100% identical to the 3'-terminal nucleotide sequence of the iTK3-1-100 fragment, and the 3'-terminal 100-bp nucleotide sequence is 100% identical to the 5'-terminal nucleotide sequence of the iTK3-3-100 fragment. The iTK3-3-100 fragments used were iTK3-3-100-Fw (SEQ ID NO: 128) and iTK3-3-100-Rv (SEQ ID NO: 129). The 100-bp nucleotide sequence at the 5' end of the iTK3-3-100 fragment amplified by this PCR is 100% homologous to the 3'-end nucleotide sequence of the iTK3-3-100 fragment, and the 50-bp nucleotide sequence at the 3' end is 100% homologous to the 50-bp nucleotide sequence at the 5' end of pCMV-SceTK3 digested with I-SceI.

[0220] Next, pIRES (Takara Bio, Figure 2) was digested with NheI and BamHI to recover a fragment containing the CMV promoter. Using the In-Fusion HD Cloning Kit (Takara Bio), the 5'SceTK3 fragment and the 3'SceTK3 fragment were inserted downstream of the CMV promoter of the recovered fragment to obtain pCMV-SceTK3, which contains the CMV promoter, SceTK3 fragment, and poly(A) addition signal. The amino acid sequences encoded by positions 1 to 141 (the region up to the first stop codon), 142 to 162 (the region up to the second stop codon), and 163 to 330 (the region downstream of the third stop codon) of the SceTK3 fragment (SEQ ID NO: 148) are shown in SEQ ID NOs: 149 to 151, respectively (note that polypeptides consisting of the sequences of SEQ ID NOs: 150 to 151 will not be expressed).

[0221] The iTK3-1, iTK3-2, and iTK3-3 fragments, which have reduced homology between homologous sequences, were amplified by PCR using the HSV-TK gene fragment as a template. The PCR polymerase used was Tks Gflex (trade name) DNA Polymerase (Takara Bio). Silent mutations were introduced to achieve homology of 98% (one site for a 50-bp homologous region and two sites for a 100-bp homologous region), 94% (three sites for a 50-bp homologous region and six sites for a 100-bp homologous region), and 90% (five sites for a 50-bp homologous region and ten sites for a 100-bp homologous region) (Figures 36-1 to 36-3). The primers used are listed in Table 6. iTK3-1-98-Fw (SEQ ID NO: 130) and iTK3-1-98-Rv (SEQ ID NO: 131) were used to amplify the iTK3-1-98 fragment (98% homology, SEQ ID NO: 155). iTK3-1-94-Fw (SEQ ID NO: 132) and iTK3-1-94-Rv (SEQ ID NO: 133) were used to amplify the iTK3-1-94 fragment (94% homology, SEQ ID NO: 156). iTK3-1-90-Fw (SEQ ID NO: 134) and iTK3-1-90-Rv (SEQ ID NO: 135) were used to amplify the iTK3-1-90 fragment (90% homology, SEQ ID NO: 157). iTK3-2-98-Fw (SEQ ID NO: 136) and iTK3-2-98-Rv (SEQ ID NO: 137) were used to amplify the iTK3-2-98 fragment (98% homology, SEQ ID NO: 158). iTK3-2-94-Fw (SEQ ID NO: 138) and iTK3-2-94-Rv (SEQ ID NO: 139) were used to amplify the iTK3-2-94 fragment (94% homology, SEQ ID NO: 159). iTK3-2-90-Fw (SEQ ID NO: 140) and iTK3-2-90-Rv (SEQ ID NO: 141) were used to amplify the iTK3-2-90 fragment (90% homology, SEQ ID NO: 160).iTK3-3-98-Fw (SEQ ID NO: 142) and iTK3-3-98-Rv (SEQ ID NO: 143) were used to amplify the iTK3-3-98 fragment (98% homology, SEQ ID NO: 161). iTK3-3-94-Fw (SEQ ID NO: 144) and iTK3-3-94-Rv (SEQ ID NO: 145) were used to amplify the iTK3-3-94 fragment (94% homology, SEQ ID NO: 162). iTK3-3-90-Fw (SEQ ID NO: 146) and iTK3-3-90-Rv (SEQ ID NO: 147) were used to amplify the iTK3-3-90 fragment (90% homology, SEQ ID NO: 163).

[0222] pCMV-SceTK3 was purified using the Qiagen Plasmid Plus Midi Kit (Qiagen KK) and linearized by digestion with I-SceI (New England Biolabs, Ipswich, MA, USA) before transfection. The iTK3-1-100, iTK3-2-100, iTK3-3-100, iTK3-1-98, iTK3-1-94, iTK3-1-90, iTK3-2-98, iTK3-2-94, iTK3-2-90, iTK3-3-98, iTK3-3-94, and iTK3-3-90 fragments were purified using the Wizard SV Gel and PCR Clean-Up System (Promega) before transfection.

[0223]

[0224] EI-2. Cells, gene transfer, and evaluation of cytotoxicity. For details on cells and gene transfer, see AI-2. Evaluation of cytotoxicity in Nalm-6 cells (wild-type and its Msh2-reverted cells) was performed as follows. 6 Each construct (1 μg) was transfected into 1 x 10 cells. After culturing the transfected cells for 48 hours, each construct (1 μg) was transfected again and 1 x 10 cells were cultured. 5Cells were seeded into 24-well dishes at 5 x 10 cells / ml, and ganciclovir (Fujifilm Wako Pure Chemical Industries, Ltd.) was added to the medium at a final concentration of 500 nM. After 96 hours of culture, cell viability was measured using CellTiter-Glo (CellTiter-Glo® Luminescent Cell Viability Assay, Promega). The cytotoxic effect on HCT116 cells (wild-type and its Mlh1 revertant cells) was evaluated as follows. 5 x 10 4 Cells were seeded into 24-well dishes, cultured overnight, and then transfected with 1 μg of each construct. After 24 hours of culture, the transfected cells were retransfected with 1 μg of each construct. Ganciclovir (Fujifilm Wako Pure Chemical Industries, Ltd.) was added to the medium at a final concentration of 1 μM. After 72 hours of culture, cell viability was measured using CellTiter-Glo (CellTiter-Glo® Luminescent Cell Viability Assay, Promega).

[0225] EI-3. Results. Figure 37 shows the results of comparing the viability of Nalm-6 cells and their MSH2-reverted counterparts 96 hours after construct transfection. Figure 38 shows the results of comparing the viability of HCT116 cells and their MLH1-reverted counterparts 72 hours after construct transfection (relative values ​​are shown, with the viability of untransfected cells (No DNA) cultured in the presence of ganciclovir (GANC) set at 100%). In both cases, the 98% and 94% homology constructs, which reduced the homology of the homologous region, showed differences in viability depending on the presence or absence of MMR factors. In particular, the 94% homology construct showed almost no toxicity to MMR-normal cells (MSH2-reverted Nalm-6 cells and MLH1-reverted HCT116 cells), demonstrating that the constructs of the present invention can selectively kill MMR-deficient cells (WT Nalm-6 cells and WT HCT116 cells).

[0226] Example E-II: Split SA-TK30 v3 (four-split, four-times SSA type) (Figure 40) The TK30 gene is a mutant of the HSV-TK gene in which the effect of ganciclovir is enhanced by introducing a mutation into an amino acid near the active center of the HSV-TK gene, and it exhibits stronger cytotoxicity than wild-type HSV-TK (Kokoris MS et al. Gene Ther. 1999, Aug;6(8):1415-1426.). In this example, the construct of the present invention was prepared using the TK30 gene as a suicide gene, and its cytocidal effect was evaluated.

[0227] To create the construct pCMV-TK30 expressing the TK30 gene, mutations were introduced into each codon of the HSV-TK gene, substituting alanine at position 152 with valine, leucine at position 159 with isoleucine, isoleucine at position 160 with leucine, phenylalanine at position 161 with alanine, alanine at position 168 with tyrosine, and leucine at position 169 with phenylalanine (Figure 39). The HSV-TK gene fragment containing these mutations (TK30 gene, SEQ ID NO: 179) was constructed by artificial gene synthesis from the SphI recognition sequence site 5' upstream of the mutation site (position 386 in the HSV-TK gene) to the BspEI recognition sequence site 3' downstream (position 629 in the HSV-TK gene). The constructed pTK30 was digested with SphI and BspEI to recover a fragment containing the TK30 gene. Next, pCMV-TK was digested with SphI and BspEI to remove positions 391 to 624 in the HSV-TK gene, resulting in recovery of a fragment containing the CMV promoter, poly(A) addition signal, and a portion of the HSV-TK gene. The TK30 fragment was then inserted using the Mighty Mix DNA Ligation Kit (Takara Bio) to obtain a TK30 expression vector (pCMV-TK30) consisting of the CMV promoter, TK30 fragment, and poly(A) addition signal. The amino acid sequence of the TK30 protein encoded by the TK30 fragment (SEQ ID NO: 179) is shown in SEQ ID NO: 180.

[0228] pCMV-SceTK3, in which the CMV promoter, SceTK3 fragment, and poly(A) addition signal were ligated, was prepared as described in EI-1.

[0229] Next, the 362-bp iTK3-1-100 fragment (region from positions 87 to 448 in the TK30 gene sequence; SEQ ID NO: 152), the 412-bp iTK303-2-100 fragment (region from positions 349 to 760 in the TK30 gene sequence; SEQ ID NO: 181), and the 358-bp iTK3-3-100 fragment (region from positions 661 to 1018 in the TK30 gene sequence; SEQ ID NO: 154) were amplified by PCR using pCMV-TK30 as a template. The polymerase used for PCR was Tks Gflex (trade name) DNA Polymerase (Takara Bio). The iTK3-1-100 fragment was amplified using iTK3-1-100-Fw (SEQ ID NO: 124) and iTK3-1-100-Rv (SEQ ID NO: 125). The 50-bp nucleotide sequence from the 5' end of the iTK3-1-100 fragment amplified by this PCR is 100% identical to the 50-bp nucleotide sequence from the 3' end of pCMV-SceTK3 digested with I-SceI, and the 100-bp nucleotide sequence from the 3' end is 100% identical to the 5' end of the iTK303-2-100 fragment. iTK3-2-100-Fw (SEQ ID NO: 126) and iTK3-2-100-Rv (SEQ ID NO: 127) were used to amplify the iTK303-2-100 fragment. The 5'-terminal 100-bp nucleotide sequence of the iTK303-2-100 fragment amplified by this PCR is 100% identical to the 3'-terminal nucleotide sequence of the iTK3-1-100 fragment, and the 3'-terminal 100-bp nucleotide sequence is 100% identical to the 5'-terminal nucleotide sequence of the iTK3-3-100 fragment. The iTK3-3-100 fragment was amplified using iTK3-3-100-Fw (SEQ ID NO: 128) and iTK3-3-100-Rv (SEQ ID NO: 129). The 5'-terminal 100-bp nucleotide sequence of the iTK3-3-100 fragment amplified by this PCR is 100% identical to the 3'-terminal nucleotide sequence of the iTK303-2-100 fragment, and the 3'-terminal 50-bp nucleotide sequence is 100% identical to the 5'-terminal 50-bp of pCMV-SceTK3 digested with I-SceI (Figure 40).

[0230] The iTK3-1, iTK303-2, and iTK3-3 fragments, which have reduced homology between homologous sequences, were amplified by PCR using the HSV-TK gene fragment as a template. The PCR polymerase used was Tks Gflex (trade name) DNA Polymerase (Takara Bio). Silent mutations were introduced to achieve homology of 98% (one site for a 50-bp homologous region and two sites for a 100-bp homologous region), 94% (three sites for a 50-bp homologous region and six sites for a 100-bp homologous region), and 90% (five sites for a 50-bp homologous region and ten sites for a 100-bp homologous region) (Figures 41-1 to 41-3). The primers used were those listed in Table 6. iTK3-1-98-Fw (SEQ ID NO: 130) and iTK3-1-98-Rv (SEQ ID NO: 131) were used to amplify the iTK3-1-98 fragment (98% homology, SEQ ID NO: 155). iTK3-1-94-Fw (SEQ ID NO: 132) and iTK3-1-94-Rv (SEQ ID NO: 133) were used to amplify the iTK3-1-94 fragment (94% homology, SEQ ID NO: 156). iTK3-1-90-Fw (SEQ ID NO: 134) and iTK3-1-90-Rv (SEQ ID NO: 135) were used to amplify the iTK3-1-90 fragment (90% homology, SEQ ID NO: 157). iTK3-2-98-Fw (SEQ ID NO: 136) and iTK3-2-98-Rv (SEQ ID NO: 137) were used to amplify the iTK303-2-98 fragment (98% homology, SEQ ID NO: 182). iTK3-2-94-Fw (SEQ ID NO: 138) and iTK3-2-94-Rv (SEQ ID NO: 139) were used to amplify the iTK303-2-94 fragment (94% homology, SEQ ID NO: 183). iTK3-2-90-Fw (SEQ ID NO: 140) and iTK3-2-90-Rv (SEQ ID NO: 141) were used to amplify the iTK303-2-90 fragment (90% homology, SEQ ID NO: 184).iTK3-3-98-Fw (SEQ ID NO: 142) and iTK3-3-98-Rv (SEQ ID NO: 143) were used to amplify the iTK3-3-98 fragment (98% homology, SEQ ID NO: 161). iTK3-3-94-Fw (SEQ ID NO: 144) and iTK3-3-94-Rv (SEQ ID NO: 145) were used to amplify the iTK3-3-94 fragment (94% homology, SEQ ID NO: 162). iTK3-3-90-Fw (SEQ ID NO: 146) and iTK3-3-90-Rv (SEQ ID NO: 147) were used to amplify the iTK3-3-90 fragment (90% homology, SEQ ID NO: 163).

[0231] pCMV-SceTK3 was purified using the Qiagen Plasmid Plus Midi Kit (Qiagen KK) and linearized by digestion with I-SceI (New England Biolabs, Ipswich, MA, USA) before transfection. The iTK3-1-100, iTK303-2-100, iTK3-3-100, iTK3-1-98, iTK3-1-94, iTK3-1-90, iTK303-2-98, iTK303-2-94, iTK303-2-90, iTK3-3-98, iTK3-3-94, and iTK3-3-90 fragments were purified using the Wizard SV Gel and PCR Clean-Up System (Promega) before transfection.

[0232] E-II-2. Cells and gene transfer, evaluation of cytotoxicity For cells and gene transfer, see AI-2 above. For evaluation of cytotoxicity, see EI-2 above.

[0233] E-II-3. Results The four-part, four-times SSA construct constructed in this example was constructed so that single-strand annealing (SSA) would produce a normal TK30 gene (modified HSV-TK gene) at four sites: the 3' side of the SceTK3 fragment and the 5' side of the iTK3-1 fragment, the 3' side of the iTK3-1 fragment and the 5' side of the iTK303-2 fragment, the 3' side of the iTK303-2 fragment and the 5' side of the iTK3-3 fragment, and the 3' side of the iTK3-3 fragment and the 5' side of the SceTK3 fragment (Figure 40).

[0234] The construct constructed in this example (Figure 40) was transfected into Nalm-6 cells and their MSH2-reverted counterparts, and the results of comparing their survival rates after 96 hours are shown in Figure 42 (relative values ​​are shown, with the survival rate of untransfected cells (no DNA) cultured in the presence of ganciclovir (GANC) set at 100%). Constructs with reduced homology in the homologous region (98%, 94%, and 90%) showed differences in survival rate depending on the presence or absence of MSH2. In particular, transfection with the 94% homology construct showed almost no toxicity to MSH2-reverted cells, resulting in an approximately 13-fold difference in survival rate.

[0235] Figure 43 shows the results of comparing the survival rates after 72 hours of transfection of HCT116 cells and their MLH1-reverted counterparts with the construct constructed in this example (Figure 40). As with the presence or absence of MSH2, the presence or absence of MLH1 also confirmed that the difference in survival rates widened as the homology in the homologous region decreased.

[0236] These results confirmed that even when TK30 was used as the suicide gene, the 4-split 4-SSA construct was able to selectively kill MMR-deficient cells (WT Nalm-6 cells and WT HCT116 cells).

[0237] Example E-III: Split SA-CDUPRT v3 (Four-Split, Four-Sequence SSA Form) (Figure 44) In this example, the CD::UPRT gene was used as a gene encoding a protein that reduces cell viability. The CD::UPRT gene is a fusion gene of the Fcy1 gene (encoding CD (cytosine deaminase)) derived from budding yeast and the Fur1 gene (encoding UPRT (uracil phosphoribosyltransferase)) derived from budding yeast. CD deaminates 5-fluorocytosine (5-FC) to convert it to 5-fluorouracil. 5-Fluorouracil is metabolized intracellularly and converted into a substance that inhibits DNA and RNA synthesis, thereby inducing cell death (Austin EA and Huber BE. Mol Pharmacol. 1993 Mar; 43(3): 380-387.). UPRT is known to catalyze the metabolic pathway by which 5-fluorouracil is converted into a toxic substance within cells, and its use in combination with CD enhances the cytotoxicity of 5-fluorocytosine and provides a bystander effect (Tiraby M et al. FEMS Microbiol Lett. 1998 Oct 1;167(1): 41-49., Bourbeau D et al. J Gene Med. 2004 Dec;6(12):1320-1332.).

[0238] E-III-1. Vector Construction. To construct pCMV-CDUPRT, which expresses a fusion gene (CD::UPRT gene) of the budding yeast Fcy1 gene (CD) and the budding yeast Fur1 gene (UPRT), the codons of the CD::UPRT gene were optimized for use in human cells as previously reported (Ho YK et al. Sci Rep. 2020 Aug 31;10(1):14257). Furthermore, XhoI and XbaI recognition sequences were added 5' upstream and 3' downstream, respectively, to facilitate restriction enzyme cloning. The CD::UPRT gene sequence (SEQ ID NO: 185; encoding a fusion protein (SEQ ID NO: 186) consisting of the full-length CD and residues 3-216 of UPRT linked by a single alanine residue) was generated by artificial gene synthesis (pCDUPRT). The constructed pCDUPRT was digested with XhoI and XbaI to recover a fragment containing the CD::UPRT gene. Next, pIRES (Takara Bio) was digested with XhoI and XbaI to remove the IRES region and recover a fragment containing the CMV promoter and poly(A) addition signal. The CD::UPRT fragment was then inserted using the DNA Ligation Kit (Mighty Mix) (Takara Bio) to obtain a CD::UPRT expression vector (pCMV-CDUPRT) in which the CMV promoter, CD::UPRT fragment, and poly(A) addition signal were ligated.

[0239] The 400-bp SceCDUPRT3 fragment used in pCMV-SceCDUPRT3 (DNA in which the region from positions 169 to 917 in the CD::UPRT gene sequence was replaced with a stop codon (TGA) + cgcg + I-SceI recognition sequence, with a stop codon added to the 3' end; SEQ ID NO: 187) was obtained by PCR amplification using pCMV-CDUPRT as a template. PrimeSTAR HS (trade name) DNA Polymerase (Takara Bio) was used as the polymerase for PCR. The primers used are listed in Table 7 below. The SceCDUPRT3 fragment was amplified separately into two fragments: the 5'SceCDUPRT3 fragment and the 3'SceCDUPRT3 fragment. The 5'SceCDUPRT fragment was amplified using Sce-5'CDUPRT-Fw (SEQ ID NO: 189) and Sce-5'CDUPRT3-Rv (SEQ ID NO: 190) to amplify a DNA fragment with a structure in which a sequence for In-Fusion reaction was added to the 5' end of the 5'SceCDUPRT3 fragment and a "stop codon (TGA) + cgcg + I-SceI recognition sequence" was added to the 3' end. The 3'SceCDUPRT3 fragment was amplified using Sce-3'CDUPRT3-Fw (SEQ ID NO: 191) and Sce-3'CDUPRT-Rv (SEQ ID NO: 192) to amplify a DNA fragment with a structure in which an "I-SceI recognition sequence" was added to the 5' end of the 3'SceCDUPRT3 fragment and a sequence for In-Fusion reaction was added to the 3' end. Next, pIRES (Takara Bio) was digested with NheI and BamHI to remove the IRES region and poly(A) addition signal, and a fragment containing the CMV promoter was recovered. Using In-Fusion Snap Assembly Master Mix (Takara Bio, formerly known as the In-Fusion HD Cloning Kit), the 5'SceCDUPRT3 fragment and the 3'SceCDUPRT3 fragment were inserted downstream of the CMV promoter in the recovered fragment, yielding pCMV-SceCDUPRT3, which is a CMV promoter-SceCDUPRT3 fragment-poly(A) addition signal ligation. The SceCDUPRT3 fragment contains a stop codon TGA introduced upstream of the I-SceI recognition sequence, and multiple stop codons downstream due to frameshifts.The amino acid sequence encoded by positions 1 to 168 (the region up to the first stop codon) of the SceCDUPRT3 fragment (SEQ ID NO: 187) is shown in SEQ ID NO: 188.

[0240] Next, the 256-bp iCDUPRT3-1-100 fragment (region from positions 119 to 374 in the CD::UPRT gene sequence; SEQ ID NO: 217), the 399-bp iCDUPRT3-2-100 fragment (region from positions 275 to 673 in the CD::UPRT gene sequence; SEQ ID NO: 218), and the 394-bp iCDUPRT3-3-100 fragment (region from positions 574 to 967 in the CD::UPRT gene sequence; SEQ ID NO: 219) were amplified by PCR using pCMV-CDUPRT as a template. The polymerase used for PCR was Tks Gflex (trade name) DNA Polymerase (Takara Bio). The iCDUPRT3-1-100 fragment (SEQ ID NO: 217) was amplified using iCDUPRT3-1-100-Fw (SEQ ID NO: 193) and iCDUPRT3-1-100-Rv (SEQ ID NO: 194). The 50-bp nucleotide sequence from the 5' end of the iCDUPRT3-1-100 fragment amplified by this PCR is 100% identical to the 100-bp nucleotide sequence from the 3' end of pCMV-SceCDUPRT3 cleaved with I-SceI, and the 100-bp nucleotide sequence from the 3' end is 100% identical to the 5' end of the iCDUPRT3-2-100 fragment. The iCDUPRT3-2-100 fragment (SEQ ID NO: 218) was amplified using iCDUPRT3-2-100-Fw (SEQ ID NO: 195) and iCDUPRT3-2-100-Rv (SEQ ID NO: 196). The 5'-terminal 100-bp nucleotide sequence of the iCDUPRT3-2-100 fragment amplified by this PCR is 100% identical to the 3'-terminal nucleotide sequence of the iCDUPRT3-1-100 fragment, and the 3'-terminal 100-bp nucleotide sequence is 100% identical to the 5'-terminal nucleotide sequence of the iCDUPRT3-3-100 fragment. iCDUPRT3-3-100-Fw (SEQ ID NO: 197) and iCDUPRT3-3-100-Rv (SEQ ID NO: 198) were used to amplify the iCDUPRT3-3-100 fragment (SEQ ID NO: 219).The 100-bp nucleotide sequence at the 5' end of the iCDUPRT3-3-100 fragment amplified by this PCR is 100% identical to the 3'-end nucleotide sequence of the iCDUPRT3-2-100 fragment, and the 50-bp nucleotide sequence at the 3' end is 100% identical to the 100-bp nucleotide sequence at the 5' end of pCMV-SceCDUPRT3 cleaved with I-SceI (Figure 44).

[0241] The iCDUPRT3-1, iCDUPRT3-2, and iCDUPRT3-3 fragments, which have reduced homology between homologous sequences, were amplified by PCR using pCMV-CDUPRT as a template. Tks Gflex (trade name) DNA Polymerase (Takara Bio) was used as the polymerase for PCR. Silent mutations were introduced to achieve homology of 98% (one site for a 50-bp homologous region and two sites for a 100-bp homologous region), 94% (three sites for a 50-bp homologous region and six sites for a 100-bp homologous region), and 90% (five sites for a 50-bp homologous region and ten sites for a 100-bp homologous region) (Figures 45-1 to 45-3). The primers used are listed in Table 7 below. iCDUPRT3-1-98-Fw (SEQ ID NO: 199) and iCDUPRT3-1-98-Rv (SEQ ID NO: 200) were used to amplify the iCDUPRT3-1-98 fragment (98% homology, SEQ ID NO: 220). iCDUPRT3-1-94-Fw (SEQ ID NO: 201) and iCDUPRT3-1-94-Rv (SEQ ID NO: 202) were used to amplify the iCDUPRT3-1-94 fragment (94% homology, SEQ ID NO: 221). iCDUPRT3-1-90-Fw (SEQ ID NO: 203) and iCDUPRT3-1-90-Rv (SEQ ID NO: 204) were used to amplify the iCDUPRT3-1-90 fragment (90% homology, SEQ ID NO: 222). iCDUPRT3-2-98-Fw (SEQ ID NO: 205) and iCDUPRT3-2-98-Rv (SEQ ID NO: 206) were used to amplify the iCDUPRT3-2-98 fragment (98% homology, SEQ ID NO: 223), iCDUPRT3-2-94-Fw (SEQ ID NO: 207) and iCDUPRT3-2-94-Rv (SEQ ID NO: 208) were used to amplify the iCDUPRT3-2-94 fragment (94% homology, SEQ ID NO: 224), and iCDUPRT3-2-90-Fw (SEQ ID NO: 209) and iCDUPRT3-2-90-Rv (SEQ ID NO: 210) were used to amplify the iCDUPRT3-2-90 fragment (90% homology, SEQ ID NO: 225).iCDUPRT3-3-98-Fw (SEQ ID NO: 211) and iCDUPRT3-3-98-Rv (SEQ ID NO: 212) were used to amplify the iCDUPRT3-3-98 fragment (98% homology, SEQ ID NO: 226), iCDUPRT3-3-94-Fw (SEQ ID NO: 213) and iCDUPRT3-3-94-Rv (SEQ ID NO: 214) were used to amplify the iCDUPRT3-3-94 fragment (94% homology, SEQ ID NO: 227), and iCDUPRT3-3-90-Fw (SEQ ID NO: 215) and iCDUPRT3-3-90-Rv (SEQ ID NO: 216) were used to amplify the iCDUPRT3-3-90 fragment (90% homology, SEQ ID NO: 228).

[0242] pCMV-SceCDUPRT3 was purified using the Qiagen Plasmid Plus Midi Kit (Qiagen KK) and linearized by digestion with I-SceI (New England Biolabs, Ipswich, MA, USA) before transfection. The iCDUPRT3-1-100, iCDUPRT3-2-100, iCDUPRT3-3-100, iCDUPRT3-1-98, iCDUPRT3-1-94, iCDUPRT3-1-90, iCDUPRT3-2-98, iCDUPRT3-2-94, iCDUPRT3-2-90, iCDUPRT3-3-98, iCDUPRT3-3-94, and iCDUPRT3-3-90 fragments were purified using the Wizard SV Gel and PCR Clean-Up System (Promega) before transfection.

[0243]

[0244] E-III-2. Cells, gene transfection, and evaluation of cytotoxicity. For details on cells and gene transfection, see AI-2. Evaluation of cytotoxicity in Nalm-6 cells and their Msh2-reverted cells was performed as follows. 2 x 10 6Each construct (1 μg) was transfected into 1 x 10 cells. After transfection, the cells were cultured for 48 hours in a medium containing 5-fluorocytosine (Fujifilm Wako Pure Chemical Industries, Ltd.) at a final concentration of 100 μM, and then transfected again with each construct (1 μg). The transfected cells were cultured at 1 x 10 cells. 5 The cells were seeded into a 24-well dish at 5 x 10 cells / ml, and 5-fluorocytosine was added to the medium at a final concentration of 100 μM. After 96 hours of culture, cell viability was measured using CellTiter-Glo (CellTiter-Glo® Luminescent Cell Viability Assay, Promega). The cytocidal effect on HCT116 cells and their MLH1-reverted cells (MLH1+ HCT116 cells) was evaluated as follows. 4 Cells were seeded into 24-well dishes, cultured overnight, and then transfected with 1 μg of each construct. After 24 hours of culture, the transfected cells were retransfected with 1 μg of each construct, and 5-fluorocytosine was added to the medium at a final concentration of 100 μM. After 72 hours of culture, cell viability was measured using CellTiter-Glo (CellTiter-Glo® Luminescent Cell Viability Assay, Promega).

[0245] E-III-3. Results The four-part, four-times SSA construct constructed in this example was constructed so that when single-strand annealing (SSA) occurs at four sites, namely, the 3' side of the SceCDUPRT3 fragment and the 5' side of the iCDUPRT3-1 fragment, the 3' side of the iCDUPRT3-1 fragment and the 5' side of the iCDUPRT3-2 fragment, the 3' side of the iCDUPRT3-2 fragment and the 5' side of the iCDUPRT3-3 fragment, and the 3' side of the iCDUPRT3-3 fragment and the 5' side of the SceCDUPRT3 fragment, a normal CD::UPRT gene (a fusion gene of the Fcy1 gene (CD) derived from budding yeast and the Fur1 gene (UPRT) derived from budding yeast) is generated (Figure 44).

[0246] Figure 46 shows the results of a comparison of the survival rates between Nalm-6 cells and their MSH2-reverted counterparts 96 hours after transfection with the construct constructed in this example. Figure 47 shows the results of a comparison of the survival rates between HCT116 cells and their MLH1-reverted counterparts 72 hours after transfection with the construct. (All figures are relative values, with the survival rate of untransfected cells (no DNA) cultured in the presence of 5-FC (100 μM) set at 100%). It was confirmed that the presence or absence of MMR factors resulted in differences in survival rates for constructs with reduced homology in the homologous region. In particular, when a construct with 94% homology was introduced, almost no toxicity was observed against MMR-normal cells (MSH2-reverted cells (Nalm-6 cells) and MLH1-reverted cells (HCT116 cells)). This indicates that the 4-split 4-SSA construct using CD::UPRT as the suicide gene can selectively kill MMR-deficient cells (WT Nalm-6 cells and WT HCT116 cells).

Claims

1. A diagnostic agent for mismatch repair deficient cancer or a companion diagnostic agent for predicting the effect of an anticancer agent on mismatch repair deficient cancer, comprising a nucleic acid construct, wherein the nucleic acid construct is A promoter region; a 5' region of a gene sequence encoding protein A; a 3' region of a gene sequence encoding protein A; one complementary region including a homologous region α, which uses a region including at least the 3'-end portion of the 5'-side region as a substrate for a recombination reaction by single-strand annealing, and a homologous region β, which uses a region including at least the 5'-end portion of the 3'-side region as a substrate for the recombination reaction; are located on one nucleic acid molecule or on two or more different nucleic acid molecules, the homology between each homologous region and each region that serves as its substrate is 40% or more but less than 100%; The agent is a nucleic acid construct, which is formed by the recombination reaction with a nucleic acid having a base sequence encoding protein A or protein B having the same activity as protein A, and which expresses protein A or protein B.

2. 84. The agent according to claim 1 or 83, wherein the amino acid sequence of protein B has a sequence identity of 80% or more but less than 100% with the amino acid sequence of protein A.

3. The agent according to claim 1 or 83, wherein the promoter region, the 5' region and the 3' region are located on the same nucleic acid molecule.

4. 84. The agent of claim 1 or 83, wherein the promoter region and the 5' region are located on one nucleic acid molecule, and the 3' region is located on another nucleic acid molecule.

5. The agent according to claim 1 or 83, wherein a poly A addition signal is operably linked downstream of the 3' region.

6. The agent according to claim 3, wherein the nucleic acid construct is composed of two nucleic acid molecules and includes one complementary region, the nucleic acid molecule 1 comprises a promoter region, the 5' side region, and the 3' side region; Nucleic acid molecule 2 comprises a complementary region comprising homologous region α and homologous region β, the homology between homologous region α and a region of the 5'-side region that is a substrate in the recombination reaction and that includes at least the 3'-end portion is 40% or more but less than 100%, and the homology between homologous region β and a region of the 3'-side region that is a substrate in the recombination reaction and that includes at least the 5'-end portion is 40% or more but less than 100%, The agent, wherein when the 5' region, complementary region, and 3' region are arranged in this order, they encode the amino acid sequence of protein A or protein B with overlapping between each homologous region and its substrate, and protein A or protein B is expressed by the recombination reaction that occurs between each homologous region and its substrate.

7. A promoter region; a 5' region of a gene sequence encoding protein A; a 3' region of a gene sequence encoding protein A; (n-1) complementary regions, including a homologous region α having a region including at least the 3'-end portion of the 5'-side region as a substrate for a recombination reaction by single-strand annealing, and a homologous region β having a region including at least the 5'-end portion of the 3'-side region as a substrate for the recombination reaction; are arranged on n nucleic acid molecules, the nucleic acid molecule 1 comprises a promoter region, the 5' side region, and the 3' side region; nucleic acid molecule 2 comprises a first complementary region comprising a first homologous region and a second homologous region; the mth nucleic acid molecule m comprises an (m-1)th substrate region and an (m-1)th complementary region comprising an mth homologous region; the first homologous region is homologous region α, and the homology between the homologous region and a region of the 5′-side region that is a substrate for the homologous region in the recombination reaction and includes at least the 3′-end portion is 40% or more but less than 100%; the (m-1) homologous region uses the (m-1) substrate region as a substrate for the recombination reaction, and the homology therebetween is 40% or more but less than 100%; the nth homologous region in the nucleic acid molecule n is a homologous region β, and the homology between the homologous region and a region comprising at least the 5'-end portion of the 3'-side region, which is a substrate for the homologous region in the recombination reaction, is 40% or more but less than 100%; A nucleic acid construct in which, when arranged in this order, the 5' region, the first complementary region, the (m-1)th complementary region, and the 3' region encode the amino acid sequence of protein A or protein B having the same activity as protein A, with overlapping between each homologous region and its substrate, and the recombination reaction occurring between each homologous region and its substrate forms a nucleic acid having a base sequence encoding protein A or protein B, resulting in expression of protein A or protein B (wherein integer n is a constant satisfying 3≦n≦10, and integer m is a variable satisfying 3≦m≦n).

8. The agent according to claim 6, wherein nucleic acid molecule 1 is a linear nucleic acid molecule comprising a promoter region and the 5' region in this order downstream of the 3' region, or a circular nucleic acid molecule comprising the 5' region and the 3' region in this order downstream of a promoter region.

9. The agent according to claim 8 , wherein the circular nucleic acid molecule has a cleavage site between the 5′ region and the 3′ region.

10. The agent according to claim 9, wherein the cleavage site is a restriction enzyme recognition site.

11. 5. The agent according to claim 4, wherein the nucleic acid construct is composed of three nucleic acid molecules and includes one complementary region: nucleic acid molecule 1-1 includes a promoter region and the 5'-side region; nucleic acid molecule 1-2 comprises the 3'-side region, Nucleic acid molecule 2 comprises a complementary region comprising homologous region α and homologous region β, the homology between the homologous region α and a region including at least the 3'-end portion of the 5'-side region, which is a substrate for the homologous region α in the recombination reaction, is 40% or more but less than 100%; the homology between the homologous region β and a region including at least the 5'-end portion of the 3'-side region, which is a substrate for the homologous region β in the recombination reaction, is 40% or more but less than 100%; The agent, wherein when the 5' region, complementary region, and 3' region are arranged in this order, they encode the amino acid sequence of protein A or protein B with overlapping between each homologous region and its substrate, and protein A or protein B is expressed by the recombination reaction that occurs between each homologous region and its substrate.

12. A promoter region; a 5' region of a gene sequence encoding protein A; a 3' region of a gene sequence encoding protein A; (n-1) complementary regions, including a homologous region α having a region including at least the 3'-end portion of the 5'-side region as a substrate for a recombination reaction by single-strand annealing, and a homologous region β having a region including at least the 5'-end portion of the 3'-side region as a substrate for the recombination reaction; are arranged on (n+1) nucleic acid molecules, nucleic acid molecule 1-1 includes a promoter region and the 5'-side region; nucleic acid molecule 1-2 comprises the 3'-side region, nucleic acid molecule 2 comprises a first complementary region comprising a first homologous region and a second homologous region; the mth nucleic acid molecule m comprises an (m-1)th substrate region and an (m-1)th complementary region comprising an mth homologous region; the first homologous region is homologous region α, and the homology between the homologous region and a region of the 5′-side region that is a substrate for the homologous region in the recombination reaction and includes at least the 3′-end portion is 40% or more but less than 100%; the (m-1) homologous region uses the (m-1) substrate region as a substrate for the recombination reaction, and the homology therebetween is 40% or more but less than 100%; the nth homologous region in the nucleic acid molecule n is a homologous region β, and the homology between the homologous region and a region comprising at least the 5'-end portion of the 3'-side region, which is a substrate for the homologous region in the recombination reaction, is 40% or more but less than 100%; A nucleic acid construct in which, when arranged in this order, the 5' region, the first complementary region, the (m-1)th complementary region, and the 3' region encode the amino acid sequence of protein A or protein B having the same activity as protein A, with overlapping between each homologous region and its substrate, and the recombination reaction occurring between each homologous region and its substrate forms a nucleic acid having a base sequence encoding protein A or protein B, resulting in expression of protein A or protein B (wherein integer n is a constant satisfying 3≦n≦10, and integer m is a variable satisfying 3≦m≦n).

13. The agent according to claim 3, wherein the nucleic acid construct is composed of one circular nucleic acid molecule and includes one complementary region, the circular nucleic acid molecule comprises the 5' region, the cleavage site, and the 3' region downstream of a promoter region in this order, and a complementary region comprising homologous region α and homologous region β upstream of the promoter region and downstream of the 3' region; the homology between a region including at least the 3'-end portion of the 5'-side region and a homologous region α that uses the region as a substrate for the recombination reaction is 40% or more but less than 100%; the homology between a region including at least the 5'-end portion of the 3'-side region and a homologous region β that uses the region as a substrate for the recombination reaction is 40% or more but less than 100%; The agent, wherein when the 5' region, complementary region, and 3' region are arranged in this order, they encode the amino acid sequence of protein A or protein B with overlapping between each homologous region and its substrate, and protein A or protein B is expressed by the recombination reaction that occurs between each homologous region and its substrate.

14. A promoter region; a 5' region of a gene sequence encoding protein A; a 3' region of a gene sequence encoding protein A; (n-1) complementary regions, including a homologous region α having a region including at least the 3'-end portion of the 5'-side region as a substrate for a recombination reaction by single-strand annealing, and a homologous region β having a region including at least the 5'-end portion of the 3'-side region as a substrate for the recombination reaction; are arranged on one circular nucleic acid molecule, the circular nucleic acid molecule comprises the 5' region, the cleavage site, and the 3' region downstream of a promoter region, in this order, and (n-1) complementary regions upstream of the promoter region and downstream of the 3' region; the first complementary region comprises a first homologous region and a second homologous region; the (m-1)th complementary region comprises the (m-1)th substrate region and the mth homologous region; the first homologous region is a homologous region α that uses a region including at least the 3'-end portion of the 5'-side region as a substrate for the recombination reaction, and the homology between the substrate and the first homologous region is 40% or more but less than 100%; the (m-1) homologous region uses the (m-1) substrate region as a substrate for the recombination reaction, and the homology therebetween is 40% or more but less than 100%; the nth homologous region in the (n-1)th complementary region is a homologous region β that uses a region including at least the 5'-end portion of the 3'-side region as a substrate for the recombination reaction, and the homology between the substrate and the nth homologous region is 40% or more but less than 100%; A nucleic acid construct in which, when arranged in this order, the 5' region, the first complementary region, the (m-1)th complementary region, and the 3' region encode the amino acid sequence of protein A or protein B having the same activity as protein A, with overlapping between each homologous region and its substrate, and the recombination reaction occurring between each homologous region and its substrate forms a nucleic acid having a base sequence encoding protein A or protein B, resulting in expression of protein A or protein B (wherein integer n is a constant satisfying 3≦n≦10, and integer m is a variable satisfying 3≦m≦n).

15. The agent according to claim 13, wherein the cleavage site is a restriction enzyme recognition site.

16. The agent according to claim 3, wherein the nucleic acid construct is composed of one linear nucleic acid molecule and includes one complementary region, the linear nucleic acid molecule comprises, from upstream to downstream, the 3' region, a complementary region, a promoter region, and the 5' region, in this order; The complementary region includes a homologous region α and a homologous region β, the homology between a region including at least the 3'-end portion of the 5'-side region and a homologous region α that uses the region as a substrate for the recombination reaction is 40% or more but less than 100%; the homology between a region including at least the 5'-end portion of the 3'-side region and a homologous region β that uses the region as a substrate for the recombination reaction is 40% or more but less than 100%; The agent, wherein when the 5' region, complementary region, and 3' region are arranged in this order, they encode the amino acid sequence of protein A or protein B with overlapping between each homologous region and its substrate, and protein A or protein B is expressed by the recombination reaction that occurs between each homologous region and its substrate.

17. A promoter region; a 5' region of a gene sequence encoding protein A; a 3' region of a gene sequence encoding protein A; (n-1) complementary regions, including a homologous region α having a region including at least the 3'-end portion of the 5'-side region as a substrate for a recombination reaction by single-strand annealing, and a homologous region β having a region including at least the 5'-end portion of the 3'-side region as a substrate for the recombination reaction; are arranged on one linear nucleic acid molecule, the linear nucleic acid molecule comprises, from upstream to downstream, the 3'-side region, a promoter region, and the 5'-side region in this order, and comprises (n-1) complementary regions between the 3'-side region and the promoter region; the first complementary region comprises a first homologous region and a second homologous region; the (m-1)th complementary region comprises the (m-1)th substrate region and the mth homologous region; the first homologous region is a homologous region α that uses a region including at least the 3'-end portion of the 5'-side region as a substrate for the recombination reaction, and the homology between the substrate and the first homologous region is 40% or more but less than 100%; the (m-1) homologous region uses the (m-1) substrate region as a substrate for the recombination reaction, and the homology therebetween is 40% or more but less than 100%; the nth homologous region in the (n-1)th complementary region is a homologous region β that uses a region including at least the 5'-end portion of the 3'-side region as a substrate for the recombination reaction, and the homology between the substrate and the nth homologous region is 40% or more but less than 100%; A nucleic acid construct in which, when arranged in this order, the 5' region, the first complementary region, the (m-1)th complementary region, and the 3' region encode the amino acid sequence of protein A or protein B having the same activity as protein A, with overlapping between each homologous region and its substrate, and the recombination reaction occurring between each homologous region and its substrate forms a nucleic acid having a base sequence encoding protein A or protein B, resulting in expression of protein A or protein B (wherein integer n is a constant satisfying 3≦n≦10, and integer m is a variable satisfying 3≦m≦n).

18. The agent according to claim 1 or 83, wherein the chain length of all of the homologous regions is at least 20 bases.

19. 2. The agent according to claim 1, wherein the gene sequence encoding protein A is a gene sequence encoding a protein having the effect of reducing cell viability or a protein whose expression in cells can be detected.

20. The agent according to claim 19, wherein the gene sequence encoding protein A is the sequence of a suicide gene, a DNA damage-inducing gene, a DNA repair inhibitor gene, a luciferase gene, a fluorescent protein gene, a cell surface antigen gene, a secretory protein gene, or a membrane protein gene.

21. A therapeutic agent for mismatch repair-deficient cancer, comprising the nucleic acid construct of any one of claims 7, 12, 14 and 17, wherein the gene sequence encoding protein A is a gene sequence encoding a protein that has the effect of reducing cell viability.

22. The therapeutic agent according to claim 21, wherein the gene sequence encoding protein A is the sequence of a suicide gene, a DNA damage-inducing gene, or a DNA repair-inhibiting gene.

23. A diagnostic agent for mismatch repair-deficient cancer, comprising the nucleic acid construct of any one of claims 7, 12, 14 and 17.

24. 24. The diagnostic agent according to claim 23, wherein the gene sequence encoding protein A is a gene sequence encoding a protein whose expression in cells can be detected.

25. A companion diagnostic agent for predicting the effect of an anticancer drug on mismatch repair-deficient cancer, comprising the nucleic acid construct of any one of claims 7, 12, 14 and 17.

26. The companion diagnostic of claim 25, wherein the anticancer drug is an immune checkpoint inhibitor.

27. The companion diagnostic agent according to claim 25, wherein the gene sequence encoding protein A is a gene sequence encoding a protein whose expression in cells can be detected.

28. A method for treating mismatch repair deficient cancer, comprising administering a nucleic acid construct to a patient having the mismatch repair deficient cancer, wherein the nucleic acid construct comprises: A promoter region; a 5' region of a gene sequence encoding protein A; a 3' region of a gene sequence encoding protein A; at least one complementary region including a homologous region α having a region including at least the 3'-end portion of the 5'-side region as a substrate for a recombination reaction by single-strand annealing, and a homologous region β having a region including at least the 5'-end portion of the 3'-side region as a substrate for the recombination reaction; are located on one nucleic acid molecule or on two or more different nucleic acid molecules, The homologous region α and the homologous region β are contained in the same or different complementary regions, and in the latter case, at least one complementary region contains one or more sets of additional homologous regions and substrates that cause the recombination reaction; the homology between each homologous region and each region that serves as its substrate is 40% or more but less than 100%; The method as described above, wherein the recombination reaction results in the formation of a nucleic acid having a base sequence encoding protein A or protein B having the same activity as protein A, thereby resulting in the expression of protein A or protein B, which is a nucleic acid construct.

29. Introducing the nucleic acid construct into cancer cells of a cancer patient; and Measuring the expression of protein A or protein B A method for diagnosing mismatch repair deficient cancer, comprising: A promoter region; a 5' region of a gene sequence encoding protein A; a 3' region of a gene sequence encoding protein A; at least one complementary region including a homologous region α having a region including at least the 3'-end portion of the 5'-side region as a substrate for a recombination reaction by single-strand annealing, and a homologous region β having a region including at least the 5'-end portion of the 3'-side region as a substrate for the recombination reaction; are located on one nucleic acid molecule or on two or more different nucleic acid molecules, The homologous region α and the homologous region β are contained in the same or different complementary regions, and in the latter case, at least one complementary region contains one or more sets of additional homologous regions and substrates that cause the recombination reaction; the homology between each homologous region and each region that serves as its substrate is 40% or more but less than 100%; The method as described above, wherein the recombination reaction results in the formation of a nucleic acid having a base sequence encoding protein A or protein B having the same activity as protein A, and the nucleic acid construct expresses protein A or protein B.

30. 30. The method of claim 29, wherein the expression of protein A or protein B is measured by directly or indirectly measuring the activity of protein A or protein B.

31. 30. The method of claim 29, wherein the cancer cells are cells isolated from the cancer patient, and the introduction of the nucleic acid construct into the cancer cells is performed ex vivo.

32. The method according to claim 29, wherein protein A or protein B is a protein whose expression in cells can be detected as a signal, the introduction of the nucleic acid construct into the cancer cells is carried out by administering the nucleic acid construct to the cancer patient, and it is examined whether a signal of the protein can be detected from the cancer lesion.

33. The method according to claim 29, wherein protein A or protein B is a secreted protein whose expression in cells can be detected, the introduction of the nucleic acid construct into the cancer cells is carried out by administering the nucleic acid construct to the cancer patient, and the activity of the protein in blood isolated from the patient after administration of the nucleic acid construct is measured.

34. Introducing the nucleic acid construct into cancer cells of a cancer patient; and Measuring the expression of protein A or protein B A method for predicting the effect of an anticancer drug on a mismatch repair deficient cancer, comprising: A promoter region; a 5' region of a gene sequence encoding protein A; a 3' region of a gene sequence encoding protein A; at least one complementary region including a homologous region α having a region including at least the 3'-end portion of the 5'-side region as a substrate for a recombination reaction by single-strand annealing, and a homologous region β having a region including at least the 5'-end portion of the 3'-side region as a substrate for the recombination reaction; are located on one nucleic acid molecule or on two or more different nucleic acid molecules, The homologous region α and the homologous region β are contained in the same or different complementary regions, and in the latter case, at least one complementary region contains one or more sets of additional homologous regions and substrates that cause the recombination reaction; the homology between each homologous region and each region that serves as its substrate is 40% or more but less than 100%; The method as described above, wherein the recombination reaction results in the formation of a nucleic acid having a base sequence encoding protein A or protein B having the same activity as protein A, thereby resulting in the expression of protein A or protein B, which is a nucleic acid construct.

35. 35. The method of claim 34, wherein the anti-cancer agent is an immune checkpoint inhibitor.

36. 35. The method of claim 34, wherein the cancer cells are cells isolated from the patient and the introduction of the nucleic acid construct into the cancer cells is performed ex vivo.

37. The method according to claim 34, wherein protein A or protein B is a protein whose expression in cells can be detected as a signal, the introduction of the nucleic acid construct into the cancer cells is carried out by administering the nucleic acid construct to the patient, and it is examined whether a signal of the protein can be detected from the cancer lesion.

38. The method according to claim 34, wherein protein A or protein B is a secreted protein whose expression in cells can be detected, the introduction of the nucleic acid construct into the cancer cells is carried out by administering the nucleic acid construct to the cancer patient, and the activity of the protein in blood isolated from the patient after administration of the nucleic acid construct is measured.

39. A promoter region; two substrate regions α and β; a gene sequence encoding protein X; at least one complementary region containing two homologous regions that serve as substrates for a recombination reaction by single-strand annealing, respectively, of the two substrate regions; are located on one nucleic acid molecule or on two or more different nucleic acid molecules, The two homologous regions are contained in the same or different complementary regions, and in the latter case, at least one complementary region contains one or more sets of additional homologous regions and substrates that cause the recombination reaction; the homology between the corresponding substrate region and the homologous region is 40% or more but less than 100%; A nucleic acid construct that expresses protein X through the recombination reaction.

40. 40. The nucleic acid construct of claim 39, wherein the promoter region, substrate region α, substrate region β, and the gene sequence are located on the same nucleic acid molecule.

41. 40. The nucleic acid construct of claim 39, wherein the promoter region and substrate region α are located on one nucleic acid molecule, and the substrate region β and the gene sequence are located on another nucleic acid molecule.

42. 40. The nucleic acid construct of claim 39, wherein a poly A addition signal is operably linked downstream of the gene sequence encoding protein X.

43. 41. The nucleic acid construct of claim 40, which is comprised of two nucleic acid molecules and contains one complementary region: nucleic acid molecule 1 comprises a promoter region, a substrate region α, a substrate region β, and a gene sequence encoding protein X; nucleic acid molecule 2 comprises one complementary region comprising a first homologous region and a second homologous region; the first homologous region uses substrate region α as a substrate for the recombination reaction, and the homology between the two regions is 40% or more but less than 100%; The second homologous region uses the substrate region β as a substrate for the recombination reaction, and the homology between the two is 40% or more but less than 100%. A nucleic acid construct in which protein X is expressed by the recombination reaction occurring between each homologous region and its substrate.

44. 41. The nucleic acid construct of claim 40, which is composed of n nucleic acid molecules and contains (n-1) complementary regions, The nucleic acid molecule 1 includes a promoter region, a substrate region α, a substrate region β, and the gene sequence; nucleic acid molecule 2 comprises a first complementary region comprising a first homologous region and a second homologous region; the mth nucleic acid molecule m comprises an (m-1)th substrate region and an (m-1)th complementary region comprising an mth homologous region; the first homologous region uses substrate region α as a substrate for the recombination reaction, and the homology between the two regions is 40% or more but less than 100%; the (m-1) homologous region uses the (m-1) substrate region as a substrate for the recombination reaction, and the homology therebetween is 40% or more but less than 100%; the nth homologous region of the nucleic acid molecule n uses the substrate region β as a substrate for the recombination reaction, and the homology between the two is 40% or more but less than 100%; A nucleic acid construct in which the recombination reaction between each homologous region and its substrate results in expression of protein X (wherein integer n is a constant with 3≦n≦10, and integer m is a variable with 3≦m≦n).

45. 44. The nucleic acid construct according to claim 43, wherein nucleic acid molecule 1 is a circular nucleic acid molecule comprising, downstream of a promoter region, substrate region α, substrate region β, and a gene sequence encoding protein X, in this order, or a linear nucleic acid molecule comprising, from upstream to downstream, substrate region β, a gene sequence encoding protein X, a promoter region, and substrate region α, in this order.

46. 46. ​​The nucleic acid construct of claim 45, wherein the circular nucleic acid molecule comprises a cleavage site between substrate region α and substrate region β.

47. 47. The nucleic acid construct of claim 46, wherein the cleavage site is a restriction enzyme recognition site.

48. 42. The nucleic acid construct of claim 41, which is comprised of three nucleic acid molecules and contains one complementary region: The nucleic acid molecule 1-1 comprises a promoter region and a substrate region α, Nucleic acid molecule 1-2 comprises substrate region β and the gene sequence, the first homologous region uses substrate region α as a substrate for the recombination reaction, and the homology between the two regions is 40% or more but less than 100%; the second homologous region uses substrate region β as a substrate for the recombination reaction, and the homology between them is 40% or more but less than 100%; A nucleic acid construct in which protein X is expressed by the recombination reaction occurring between each homologous region and its substrate.

49. 42. The nucleic acid construct of claim 41, which is composed of (n+1) nucleic acid molecules and contains (n-1) complementary regions, The nucleic acid molecule 1-1 includes a promoter region and a substrate region α, Nucleic acid molecule 1-2 contains a substrate region β and the gene sequence, nucleic acid molecule 2 comprises a first complementary region comprising a first homologous region and a second homologous region; the mth nucleic acid molecule m comprises an (m-1)th substrate region and an (m-1)th complementary region comprising an mth homologous region; the first homologous region uses substrate region α as a substrate for the recombination reaction, and the homology between the two regions is 40% or more but less than 100%; the (m-1) homologous region uses the (m-1) substrate region as a substrate for the recombination reaction, and the homology therebetween is 40% or more but less than 100%; the nth homologous region of the nucleic acid molecule n uses the substrate region β as a substrate for the recombination reaction, and the homology between the two is 40% or more but less than 100%; A nucleic acid construct in which the recombination reaction between each homologous region and its substrate results in expression of protein X (wherein integer n is a constant with 3≦n≦10, and integer m is a variable with 3≦m≦n).

50. 41. The nucleic acid construct of claim 40, which is comprised of one circular nucleic acid molecule and contains one complementary region, the circular nucleic acid molecule comprises a substrate region α, a cleavage site, a substrate region β, and a gene sequence encoding protein X downstream of a promoter region, in this order, and a complementary region comprising a first homologous region and a second homologous region, located upstream of the promoter region and downstream of the gene sequence; the first homologous region uses substrate region α as a substrate for the recombination reaction, and the homology between the two regions is 40% or more but less than 100%; The second homologous region uses the substrate region β as a substrate for the recombination reaction, and the homology between the two is 40% or more but less than 100%. A nucleic acid construct in which protein X is expressed by the recombination reaction occurring between each homologous region and its substrate.

51. The nucleic acid construct according to claim 40, which is composed of one circular nucleic acid molecule and contains (n-1) complementary regions, the circular nucleic acid molecule comprises, downstream of a promoter region, a substrate region α, a cleavage site, a substrate region β, and a gene sequence encoding protein X, in this order, and (n-1) complementary regions upstream of the promoter region and downstream of the gene sequence; the first complementary region comprises a first homologous region and a second homologous region; the (m-1)th complementary region comprises the (m-1)th substrate region and the mth homologous region; the first homologous region uses substrate region α as a substrate for the recombination reaction, and the homology between the two regions is 40% or more but less than 100%; the (m-1) homologous region uses the (m-1) substrate region as a substrate for the recombination reaction, and the homology therebetween is 40% or more but less than 100%; the nth homologous region in the (n-1)th complementary region uses substrate region β as a substrate for the recombination reaction, and the homology between them is 40% or more but less than 100%; A nucleic acid construct in which the recombination reaction between each homologous region and its substrate results in expression of protein X (wherein integer n is a constant with 3≦n≦10, and integer m is a variable with 3≦m≦n).

52. 51. The nucleic acid construct of claim 50, wherein the cleavage site is a restriction enzyme recognition site.

53. 41. The nucleic acid construct of claim 40, which is comprised of one linear nucleic acid molecule and contains one complementary region, the linear nucleic acid molecule comprises, from upstream to downstream, a substrate region β, a gene sequence encoding protein X, a complementary region, a promoter region, and a substrate region α, in this order; the complementary region comprises a first homologous region and a second homologous region; the first homologous region uses substrate region α as a substrate for the recombination reaction, and the homology between the two regions is 40% or more but less than 100%; The second homologous region uses the substrate region β as a substrate for the recombination reaction, and the homology between the two is 40% or more but less than 100%. A nucleic acid construct in which protein X is expressed by the recombination reaction occurring between each homologous region and its substrate.

54. 41. The nucleic acid construct according to claim 40, which is composed of one linear nucleic acid molecule and contains (n-1) complementary regions, the linear nucleic acid molecule comprises, from upstream to downstream, a substrate region β, a gene sequence encoding protein X, a promoter region, and a substrate region α, in this order, and comprises (n-1) complementary regions between the gene sequence and the promoter region; the first complementary region comprises a first homologous region and a second homologous region; the (m-1)th complementary region comprises the (m-1)th substrate region and the mth homologous region; the first homologous region uses substrate region α as a substrate for the recombination reaction, and the homology between the two regions is 40% or more but less than 100%; the (m-1) homologous region uses the (m-1) substrate region as a substrate for the recombination reaction, and the homology therebetween is 40% or more but less than 100%; the nth homologous region in the (n-1)th complementary region uses substrate region β as a substrate for the recombination reaction, and the homology between them is 40% or more but less than 100%; A nucleic acid construct in which the recombination reaction between each homologous region and its substrate results in expression of protein X (wherein integer n is a constant with 3≦n≦10, and integer m is a variable with 3≦m≦n).

55. 40. The nucleic acid construct according to claim 39, wherein protein X is a protein that has the effect of reducing cell viability or a protein whose expression in a cell can be detected.

56. 56. The nucleic acid construct according to claim 55, wherein the gene sequence encoding protein X is a sequence of a suicide gene, a DNA damage-inducing gene, a DNA repair-inhibiting gene, a luciferase gene, a fluorescent protein gene, a cell surface antigen gene, a secreted protein gene, or a membrane protein gene.

57. 40. The nucleic acid construct according to claim 39, wherein the base sequence generated by the recombination reaction between each of the substrate regions and each of the homologous regions is a gene sequence encoding protein Y, and the recombination reaction results in the expression of protein X and protein Y.

58. 58. The nucleic acid construct of claim 57, wherein a sequence enabling polycistronic expression is located between the gene sequence encoding substrate region β and protein X.

59. 59. The nucleic acid construct of claim 58, wherein the sequence enabling polycistronic expression is an IRES sequence or a 2A peptide coding sequence.

60. 58. The nucleic acid construct according to claim 57, wherein one of protein X and protein Y is a protein that has the effect of reducing cell viability, and the other is a protein whose expression in a cell can be detected.

61. 61. The nucleic acid construct according to claim 60, wherein one of the gene sequence encoding protein X and the gene sequence encoding protein Y is the sequence of a suicide gene, a DNA damage-inducing gene, or a DNA repair-inhibiting gene, and the other is the sequence of a luciferase gene, a fluorescent protein gene, a cell surface antigen gene, a secreted protein gene, or a membrane protein gene.

62. A therapeutic agent for mismatch repair-deficient cancer, comprising the nucleic acid construct of any one of claims 39 to 61, wherein the gene sequence encoding protein X is a gene sequence encoding a protein having an effect of reducing cell viability.

63. 63. The therapeutic agent according to claim 62, wherein the gene sequence encoding protein X is a sequence of a suicide gene, a DNA damage-inducing gene, or a DNA repair-inhibiting gene.

64. An agent for detecting and treating mismatch repair-deficient cancer, comprising the nucleic acid construct of claim 60 or 61.

65. A diagnostic agent for mismatch repair-deficient cancer, comprising the nucleic acid construct according to any one of claims 39 to 61.

66. 66. The diagnostic agent according to claim 65, wherein the gene sequence encoding protein X is a gene sequence encoding a protein whose expression in a cell can be detected.

67. A companion diagnostic agent for predicting the effect of an anticancer drug on mismatch repair-deficient cancer, comprising the nucleic acid construct according to any one of claims 39 to 61.

68. The companion diagnostic of claim 67, wherein the anticancer drug is an immune checkpoint inhibitor.

69. The companion diagnostic according to claim 67, wherein the gene sequence encoding protein X is a gene sequence encoding a protein whose expression in a cell can be detected.

70. A method for treating mismatch repair deficient cancer, comprising administering the nucleic acid construct according to any one of claims 39 to 61, wherein the gene sequence encoding protein X is a gene sequence encoding a protein having the effect of reducing cell viability, to a patient having mismatch repair deficient cancer.

71. 62. A method for detecting and treating mismatch repair-deficient cancer, comprising administering to a cancer patient the nucleic acid construct of claim 60 or 61, wherein one of protein X and protein Y is a protein whose expression in cells is detectable as a signal, and examining whether a signal of said protein is detected from a cancer lesion.

72. 62. A method for detecting and treating mismatch repair-deficient cancer, comprising administering to a cancer patient the nucleic acid construct of claim 60 or 61, wherein one of protein X and protein Y is a secreted protein whose expression in cells can be detected, and measuring the activity of the protein in blood isolated from the patient after administration.

73. Introducing the nucleic acid construct according to any one of claims 39 to 61 into cancer cells of a cancer patient; and Measuring the expression of protein X A method for diagnosing mismatch repair deficient cancer, comprising:

74. 74. The method of claim 73, wherein the expression of protein X is measured by directly or indirectly measuring the activity of protein X.

75. 74. The method of claim 73, wherein the cancer cells are cells isolated from the cancer patient and the introduction of the nucleic acid construct into the cancer cells is performed ex vivo.

76. The method according to claim 73, wherein protein X is a protein whose expression in cells can be detected as a signal, the introduction of the nucleic acid construct into the cancer cells is carried out by administering the nucleic acid construct to the cancer patient, and whether a signal of the protein can be detected in a cancer lesion is examined.

77. The method according to claim 73, wherein protein X is a secreted protein whose expression in cells can be detected, the introduction of the nucleic acid construct into the cancer cells is carried out by administering the nucleic acid construct to the cancer patient, and the activity of the protein in blood isolated from the patient after administration of the nucleic acid construct is measured.

78. Introducing the nucleic acid construct according to any one of claims 39 to 61 into cancer cells of a cancer patient; and Measuring the expression of protein X A method for predicting the efficacy of an anticancer drug against mismatch repair-deficient cancer, comprising:

79. 79. The method of claim 78, wherein the anti-cancer agent is an immune checkpoint inhibitor.

80. 79. The method of claim 78, wherein the cancer cells are cells isolated from the patient and the introduction of the nucleic acid construct into the cancer cells is performed ex vivo.

81. The method according to claim 78, wherein protein X is a protein whose expression in cells can be detected as a signal, the introduction of the nucleic acid construct into the cancer cells is carried out by administering the nucleic acid construct to the patient, and whether a signal of the protein can be detected in a cancer lesion is examined.

82. 79. The method according to claim 78, wherein protein X is a secreted protein whose expression in cells can be detected, the introduction of the nucleic acid construct into the cancer cells is carried out by administering the nucleic acid construct to the cancer patient, and the activity of the protein in blood isolated from the patient after administration of the nucleic acid construct is measured.

83. A therapeutic agent for mismatch repair-deficient cancer, comprising a nucleic acid construct, A promoter region; a 5'-side region of a gene sequence encoding protein A, which is a protein that reduces cell viability; a 3'-side region of a gene sequence encoding the protein A; one complementary region including a homologous region α, which uses a region including at least the 3'-end portion of the 5'-side region as a substrate for a recombination reaction by single-strand annealing, and a homologous region β, which uses a region including at least the 5'-end portion of the 3'-side region as a substrate for the recombination reaction; are located on one nucleic acid molecule or on two or more different nucleic acid molecules, the homology between each homologous region and each region that serves as its substrate is 40% or more but less than 100%; The agent is a nucleic acid construct, which is formed by the recombination reaction with a nucleic acid having a base sequence encoding protein A or protein B having the same activity as protein A, and which expresses protein A or protein B.

84. The agent according to claim 83, wherein the gene sequence encoding protein A is the sequence of a suicide gene, a DNA damage-inducing gene, or a DNA repair-inhibiting gene.

85. The agent according to claim 1, wherein the gene sequence encoding protein A is a gene sequence encoding a protein whose expression in cells can be detected.

86. The agent according to claim 1, which is a companion diagnostic agent for predicting the effect of an anticancer drug on mismatch repair-deficient cancer, and the anticancer drug is an immune checkpoint inhibitor.