Transposable element detection method
A method for detecting transposable elements by analyzing pairs of TSD sequences from genome samples addresses the challenge of identifying active transposons, facilitating applications in agriculture and research.
Patent Information
- Application Number
- JP2024151554
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-26
- Filing Date
- 2024-09-03
- Publication Date
- 2026-02-25
- Estimated Expiration
- 2040-12-25
AI Technical Summary
Detecting active transposable elements from genome sequences is challenging due to the large amount of inactive elements and the difficulty in distinguishing new active elements from base sequence information.
A method that identifies transposable elements by detecting pairs of 5' and 3' ends with different target site duplications (TSD) sequences from two samples, without relying on a reference sequence, through a process of extracting and analyzing sequences shifted by one base at a time.
Accurately detects new transposon transpositions using next-generation sequencing, enabling applications in agriculture, healthcare, and research by identifying active transposons in various organisms, including those without a reference genome.
Smart Images

Figure 0007819961000010 
Figure 0007819961000011 
Figure 0007819961000012
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to the field of information processing of sequence information, particularly sequence information of biomolecules such as genomes. The present disclosure can be used in the fields of medical care, health care, agriculture, forestry, livestock farming, fisheries, and environmental applications and basic research by detecting transposable elements. [Background technology]
[0002] Many types of transposable elements exist in the genome in various copy numbers, and most of them are inactive due to mutations. The large amount of information on these inactive transposable elements makes it difficult to detect new active transposable elements from base sequence information. Summary of the Invention [Means for solving the problem]
[0003] It is known that when a transposable element transposes, a duplication of several bases (target site duplication, TSD) may occur at the target site. For example, in the case of a transposable element with a TSD size of five bases, the sequence of the five bases upstream of the 5' end of the transposable element is the same as the sequence of the five bases downstream of the 3' end. Focusing on both ends of the transposon and the TSD, it is thought that the base sequence of the target site of a newly transposed transposon differs with each transposition, resulting in a different TSD sequence. This disclosure has discovered that by detecting pairs of the 5' and 3' ends of transposons with different TSD sequences from the base sequences obtained from two samples, it is possible to directly detect transposition without using a reference sequence.
[0004] The present disclosure provides a method for identifying a transposable element in a sequence, the method comprising: (A) extracting a sequence of a certain length from genome sequence data by shifting the sequence by one base at a time to generate a set of excised sequences; (B) dividing each excised sequence into an n-base-long portion and the remaining portion, and selecting pairs of excised sequences in which the n-base-long portion at the 5' end and the n-base-long portion at the 3' end are identical; (C) selecting pairs of excised sequences having corresponding transposable element partial sequences from the selected pairs; and / or (D) selecting pairs of excised sequences having corresponding transposable element partial sequences from pairs of excised sequences in which the n-base-long portion differs based on the genome sequence data as transposable element matching pairs. The method may further include additional steps described herein. The present disclosure may also provide a program implementing the method, a recording medium storing the program, or a system for use therewith. Another aspect of the present disclosure relates to the identified transposable element or its use.
[0005] Examples of the present disclosure include the following: (Item 1) A method for identifying a transposable element in a sequence, comprising: (A) extracting sequences of a certain length from the first genome sequence data and the second genome sequence data, shifting the sequences by one base at a time, and generating a set of extracted sequences from the first genome sequence data and a set of extracted sequences from the second genome sequence data; (B) separating each of the excision sequences into a sequence corresponding to an n-base-long target site overlap (TSD) and a sequence other than the TSD (transposable element partial sequence), and selecting a pair of excision sequences in which the 5'-terminal TSD and the 3'-terminal TSD are identical; (C) selecting pairs of excised sequences having corresponding partial sequences of transposable elements from the selected pairs; (D) selecting, from the pairs of excision sequences having the corresponding transposable element partial sequences, pairs of excision sequences having different TSDs in the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data as transposable element corresponding pairs; A method comprising: (Item 2) The method further comprises, before step (B), a step of selecting excision sequences that differ from the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data, The method according to the preceding item, wherein the selection of a pair of excision sequences having the same TSD at the 5' end and the same TSD at the 3' end is performed from among excision sequences that differ between the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data. (Item 3) The method according to any one of the preceding items, further comprising a step of calculating the frequency of identical excision sequences in the set of excision sequences. (Item 4) The method according to any one of the preceding items, further comprising the step of excluding excision sequences having a frequency below a certain level from the set of excision sequences. (Item 5) The method according to any one of the preceding items, further comprising a step of identifying a transposable element sequence based on the transposable element match pair. (Item 6) The method according to any one of the preceding items, further comprising a step of confirming the transposition activity of the transposition element. (Item 7) The method according to any one of the preceding items, wherein the transposition activity of the transposable element is confirmed by one or more methods selected from PCR, sequencing, and hybridization. (Item 8) The method according to any one of the preceding items, further comprising repeating steps (B) to (D) while changing n. (Item 9) The method according to any one of the preceding items, wherein n is 3 to 20. (Item 10) The method according to any one of the preceding items, wherein the constant length is 17 to 50 bases long. (Item 11) The method according to any of the preceding items, wherein the corresponding transposable element partial sequence is a transposable element partial sequence having at least 90% identity. (Item 12) The method according to any of the preceding items, wherein the corresponding transposable element partial sequence is a transposable element partial sequence having at least 95% identity. (Item 13) The method according to any of the preceding items, wherein the corresponding transposable element partial sequences are identical transposable element partial sequences. (Item 14) A method according to any of the preceding items, wherein the step of generating the set of excised sequences includes excising the fixed length sequences by shifting them by one base for the complementary strands of the first genome sequence data and the second genome sequence data. (Item 15) A program for causing a computer to execute a method for identifying a transposable element in a sequence, the method comprising: (a) receiving first genome sequence data and second genome sequence data; (b) extracting sequences of a fixed length from the first genome sequence data and the second genome sequence data, shifting the sequences by one base at a time, to generate a set of extracted sequences from the first genome sequence data and a set of extracted sequences from the second genome sequence data; (c) classifying each of the excision sequences into a sequence corresponding to an n-base-long target site overlap (TSD) and a sequence other than the TSD (transposable element partial sequence), and selecting a pair of excision sequences in which the 5'-terminal TSD and the 3'-terminal TSD are identical; (d) selecting pairs of excised sequences having corresponding transposable element partial sequences from the selected pairs; (e) selecting, from the pairs of excision sequences having the corresponding transposable element partial sequences, pairs of excision sequences having different TSDs in the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data as transposable element corresponding pairs; Including, the program. (Item 15-1) The program described in the above item, further comprising one or more of the features described in the above item. (Item 16) A recording medium storing a program for causing a computer to execute a method for identifying a transposable element in a sequence, the method comprising: (a) receiving first genome sequence data and second genome sequence data; (b) extracting sequences of a fixed length from the first genome sequence data and the second genome sequence data, shifting the sequences by one base at a time, to generate a set of extracted sequences from the first genome sequence data and a set of extracted sequences from the second genome sequence data; (c) classifying each of the excision sequences into a sequence corresponding to an n-base-long target site overlap (TSD) and a sequence other than the TSD (transposable element partial sequence), and selecting a pair of excision sequences in which the 5'-terminal TSD and the 3'-terminal TSD are identical; (d) selecting pairs of excised sequences having corresponding transposable element partial sequences from the selected pairs; (e) selecting, from the pairs of excision sequences having the corresponding transposable element partial sequences, pairs of excision sequences having different TSDs in the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data as transposable element corresponding pairs; A recording medium including: (Item 16-1) A recording medium according to any one of the preceding items, further comprising the features described in any one or more of the preceding items. (Item 17) A system for identifying a transposable element in a sequence, the system comprising one or more processors, a memory, and a recording medium storing a program, the program, when executed by the one or more processors, performing the following: (a) receiving first genome sequence data and second genome sequence data; (b) extracting sequences of a fixed length from the first genome sequence data and the second genome sequence data, shifting the sequences by one base at a time, to generate a set of extracted sequences from the first genome sequence data and a set of extracted sequences from the second genome sequence data; (c) Each of the excised sequences is divided into a sequence corresponding to a target site overlap (TSD) of n bases in length. and a sequence other than the TSD (transposable element partial sequence), and selecting pairs of excision sequences in which the 5'-terminal TSD and the 3'-terminal TSD are identical; (d) selecting pairs of excised sequences having corresponding transposable element partial sequences from the selected pairs; (e) selecting, from the pairs of excision sequences having the corresponding transposable element partial sequences, pairs of excision sequences having different TSDs in the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data as transposable element corresponding pairs; A system implementing a method including: (Item 17-1) The system described in the above item further comprises the features described in any one or more of the above items. (Item 18) A system for identifying a transposable element in a sequence, the system comprising: (A) a sequence data receiving unit that receives first genome sequence data and second genome sequence data; (B) extracting sequences of a certain length from the first genome sequence data and the second genome sequence data obtained by the sequence data receiving unit, shifting the sequences by one base at a time, to generate a set of sequences extracted from the first genome sequence data and a set of sequences extracted from the second genome sequence data; Each of the excision sequences is separated into a sequence corresponding to an n-base long target site overlap (TSD) and a sequence other than the TSD (transposable element partial sequence), and pairs of excision sequences in which the 5'-terminal TSD and the 3'-terminal TSD are identical are selected; From the selected pairs, pairs of excised sequences having corresponding partial sequences of transposable elements are selected; a transposable element matching pair selection unit that executes a step of selecting, from pairs of excision sequences having corresponding transposable element partial sequences, pairs of excision sequences having different TSDs in a set of excision sequences from the first genome sequence data and a set of excision sequences from the second genome sequence data, as transposable element matching pairs; (C) a display section displaying the transposable element matching pair; Including, the system. (Item 18-1) The system described in the above item further comprises the features described in any one or more of the above items. (Item 19) A nucleic acid having at least a partial sequence of a transposable element identified by the method described in any of the preceding items. (Item 20) The nucleic acid according to the preceding item, which has a recognition sequence for a transferase in the transposable element. (Item 21) Use of the nucleic acid according to any of the preceding items for introducing mutations into a genome sequence. (Item 22) Use of the nucleic acid according to any one of the preceding items for gene disruption by transposition of a transposable element. (Item 23) Use of a nucleic acid according to any of the preceding items for the control of transcription. (Item 24) Use of the nucleic acid according to any of the preceding items for the production of an insertion mutant strain.
[0006] The present disclosure further provides: (Item A1) A method for identifying a transposable element in a sequence, comprising: (A) extracting sequences of a certain length from the first genome sequence data and the second genome sequence data, shifting the sequences by one base at a time, and generating a set of extracted sequences from the first genome sequence data and a set of extracted sequences from the second genome sequence data; (B) For each of the excised sequences, a portion containing a sequence corresponding to a target site overlap (TSD) of n bases in length is distinguished from a portion not containing a sequence corresponding to the TSD, and the same TSD is identified. Selecting a pair of excised sequences comprising: (C) selecting pairs of excised sequences having corresponding partial sequences of transposable elements from the selected pairs; (D) selecting, from the pairs of excision sequences having the corresponding transposable element partial sequences, pairs of excision sequences having different TSDs in the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data as transposable element corresponding pairs; A method comprising: (Item A2) The method according to the above item, wherein step (B) is a step of distinguishing between a sequence corresponding to the TSD and a sequence other than the TSD (transposable element partial sequence), and selecting a pair of excision sequences in which the TSD at the 5' end and the TSD at the 3' end are identical. (Item A3) A method according to any of the above items, wherein step (B) is a step of distinguishing between a portion containing a sequence corresponding to the TSD and a portion not containing a sequence corresponding to the TSD, and selecting a pair of excision sequences containing the same TSD from an excision sequence containing a sequence corresponding to the TSD at the 3' end of the first half and an excision sequence containing a sequence corresponding to the TSD at the 5' end of the second half. (Item A4) The method according to any one of the preceding items, wherein the distinction between the portion containing the sequence corresponding to the TSD and the portion not containing the sequence corresponding to the TSD comprises distinguishing between half of the excised sequence. (Item A5) The method according to any one of the above items, wherein step (B) further comprises a step of identifying the genomic locations of portions containing a sequence corresponding to the TSD and portions not containing a sequence corresponding to the TSD by mapping the excised sequence onto a reference sequence. (Item A6) The method further includes, before step (B), a step of selecting excision sequences that are different between the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data, The method according to any of the above items, wherein the selection of a pair of excision sequences having the same TSD at the 5' end and the same TSD at the 3' end is performed from among excision sequences that differ between the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data. (Item A7) The method according to any one of the above items, further comprising a step of calculating the frequency of identical excision sequences in the set of excision sequences. (Item A8) The method according to any one of the above items, further comprising the step of excluding excision sequences having a frequency below a certain level from the set of excision sequences. (Item A9) The method according to any one of the above items, further comprising a step of identifying a transposable element sequence based on the transposable element match pair. (Item A10) The method according to any one of the above items, further comprising a step of confirming the transposition activity of the transposition element. (Item A11) The method according to any one of the above items, wherein the transposition activity of the transposable element is confirmed by one or more methods selected from PCR, sequencing, and hybridization. (Item A12) A method according to any one of the above items, further comprising repeating steps (B) to (D) while changing n. (Item A13) The method according to any one of the above items, wherein n is 3 to 20. (Item A14) The method according to any one of the above items, wherein the constant length is 17 to 50 bases long. (Item A15) The method according to any one of the above items, wherein the corresponding transposable element partial sequence is a transposable element partial sequence having at least 90% identity. (Item A16) The method according to any one of the above items, wherein the corresponding transposable element partial sequence is a transposable element partial sequence having at least 95% identity. (Item A17) The method according to any one of the above items, wherein the corresponding transposable element partial sequences are identical transposable element partial sequences. (Item A18) A method according to any of the above items, wherein the step of generating the set of excised sequences includes excising the fixed length sequences by shifting them by one base for the complementary strands of the first genome sequence data and the second genome sequence data. (Item A19) The method according to any one of the above items, wherein n is 5 or 8, and the constant length is 25 bases long. (Item A20) The method according to any one of the above items, wherein n is 5 or 8, and the constant length is 40 bases long. (Item A21) A method described in any of the above items, further comprising a step of identifying the presence or absence of insertion of the transposable element partial sequence and / or the insertion position of the transposable element partial sequence by comparing the insertion position of the transposable element with a reference sequence in an individual other than the individual containing the TSD based on the transposable element correspondence pair. (Item A22) A method described in any of the above items, further comprising a step of identifying the presence or absence of insertion of the transposable element partial sequence and / or the insertion position of the transposable element partial sequence by mapping the insertion position of the transposable element onto a reference sequence based on the transposable element correspondence pair. (Item A23) A method for identifying a transposable element in a sequence, comprising: (A) A step of extracting X base sequences from short reads derived from two types of samples by shifting them by one base; (B) obtaining the frequency of each of the obtained sequences of the plurality of X bases; (C) obtaining and comparing the positions of the first and second halves of each of the obtained sequences of multiple X bases on a reference sequence; A method comprising: (Item A24) The method according to any one of the preceding items, wherein the base length of the first half and the second half is X / 2 bases. (Item A25) The method according to any one of the above items, wherein X is 20 or more bases. (Item A26) A method for detecting target site duplications (TSDs) and / or junctions of transposable elements in a sequence, comprising: (A) A step of extracting X base sequences from short reads derived from two types of samples by shifting them by one base; (B) obtaining the frequency of each of the obtained sequences of the plurality of X bases; (C) obtaining and comparing the positions of the first and second halves of each of the obtained sequences of multiple X bases on a reference sequence; A method comprising: (Item A27) The method according to any one of the preceding items, wherein the base length of the first half and the second half is X / 2 bases. (Item A28) The method according to any one of the above items, wherein X is 20 or more bases. (Item A29) A program for causing a computer to execute a method for identifying a transposable element in a sequence, the method comprising: (a) receiving first genome sequence data and second genome sequence data; (b) extracting sequences of a fixed length from the first genome sequence data and the second genome sequence data, shifting the sequences by one base at a time, to generate a set of extracted sequences from the first genome sequence data and a set of extracted sequences from the second genome sequence data; (c) classifying each of the excision sequences into a sequence corresponding to an n-base-long target site overlap (TSD) and a sequence other than the TSD (transposable element partial sequence), and selecting a pair of excision sequences in which the 5'-terminal TSD and the 3'-terminal TSD are identical; (d) selecting pairs of excised sequences having corresponding transposable element partial sequences from the selected pairs; (e) deriving the first genome sequence from the pair of excised sequences having the corresponding transposable element partial sequences; selecting pairs of excised sequences having different TSDs as transposable element matching pairs from a set of excised sequences from the sequence data and a set of excised sequences from the second genome sequence data; Including, the program. (Item A30) A recording medium storing a program for causing a computer to execute a method for identifying a transposable element in a sequence, the method comprising: (a) receiving first genome sequence data and second genome sequence data; (b) extracting sequences of a fixed length from the first genome sequence data and the second genome sequence data, shifting the sequences by one base at a time, to generate a set of extracted sequences from the first genome sequence data and a set of extracted sequences from the second genome sequence data; (c) classifying each of the excision sequences into a sequence corresponding to an n-base-long target site overlap (TSD) and a sequence other than the TSD (transposable element partial sequence), and selecting a pair of excision sequences in which the 5'-terminal TSD and the 3'-terminal TSD are identical; (d) selecting pairs of excised sequences having corresponding transposable element partial sequences from the selected pairs; (e) selecting, from the pairs of excision sequences having the corresponding transposable element partial sequences, pairs of excision sequences having different TSDs in the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data as transposable element corresponding pairs; A recording medium including: (Item A31) A system for identifying a transposable element in a sequence, the system comprising one or more processors, a memory, and a recording medium storing a program, the program, when executed by the one or more processors, performing the following: (a) receiving first genome sequence data and second genome sequence data; (b) extracting sequences of a fixed length from the first genome sequence data and the second genome sequence data, shifting the sequences by one base at a time, to generate a set of extracted sequences from the first genome sequence data and a set of extracted sequences from the second genome sequence data; (c) classifying each of the excision sequences into a sequence corresponding to an n-base-long target site overlap (TSD) and a sequence other than the TSD (transposable element partial sequence), and selecting a pair of excision sequences in which the 5'-terminal TSD and the 3'-terminal TSD are identical; (d) selecting pairs of excised sequences having corresponding transposable element partial sequences from the selected pairs; (e) selecting, from the pairs of excision sequences having the corresponding transposable element partial sequences, pairs of excision sequences having different TSDs in the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data as transposable element corresponding pairs; A system implementing a method including: (Item A32) A system for identifying a transposable element in a sequence, the system comprising: (A) a sequence data receiving unit that receives first genome sequence data and second genome sequence data; (B) extracting sequences of a certain length from the first genome sequence data and the second genome sequence data obtained by the sequence data receiving unit, shifting the sequences by one base at a time, to generate a set of sequences extracted from the first genome sequence data and a set of sequences extracted from the second genome sequence data; Each of the excision sequences is separated into a sequence corresponding to an n-base long target site overlap (TSD) and a sequence other than the TSD (transposable element partial sequence), and pairs of excision sequences in which the 5'-terminal TSD and the 3'-terminal TSD are identical are selected; From the selected pairs, pairs of excised sequences having corresponding partial sequences of transposable elements are selected; a transposable element matching pair selection unit that executes a step of selecting, from pairs of excision sequences having corresponding transposable element partial sequences, pairs of excision sequences having different TSDs in a set of excision sequences from the first genome sequence data and a set of excision sequences from the second genome sequence data, as transposable element matching pairs; (C) a display section displaying the transposable element matching pair; Including, the system. (Item A33) A nucleic acid having at least a partial sequence of a transposable element identified by the method according to any one of the above items. (Item A34) The nucleic acid according to any one of the above items, which has a recognition sequence for a transferase in the transposable element. (Item A35) Use of the nucleic acid according to any of the above items for introducing mutations into a genome sequence. (Item A36) Use of the nucleic acid according to any one of the above items for gene disruption by transposition of a transposable element. (Item A37) Use of a nucleic acid according to any of the above items for the regulation of transcription. (Item A38) Use of the nucleic acid according to any of the above items for the production of an insertional mutant line. (Item A39) A method for identifying a transposable element in a sequence according to any one of the above items, (A) extracting sequences of a fixed length of X bases + n bases from the first genome sequence data and the second genome sequence data, shifting the sequences by one base at a time, to generate a set of extracted sequences from the first genome sequence data and a set of extracted sequences from the second genome sequence data; (A1) sorting and outputting a set of excised sequences from the first genome sequence data and a set of excised sequences from the second genome sequence data together with their frequencies; (A2) removing sequences that overlap between the first genome sequence data and the second genome sequence from the set of excised sequences from the first genome sequence data and the set of excised sequences from the second genome sequence data, and selecting sequences that are specific to each of the set of excised sequences from the first genome sequence data and the set of excised sequences from the second genome sequence data; (B) separating each of the excision sequences into sequences corresponding to n-base-long target site overlaps (TSDs) and sequences other than the TSDs (candidate transposable element partial sequences), selecting, from among pairs of candidate transposable element partial sequences, the 5'-terminal TSD of a candidate transposable element partial sequence having two or more types of TSDs and the 3'-terminal TSD of the candidate transposable element partial sequence, and selecting, from the selected dataset, all pairs of excision sequences in which the 5'-terminal TSD and the 3'-terminal TSD are identical; (C) a step of selecting pairs of excision sequences having corresponding transposable element partial sequences from the selected pairs, in which the 5'-side transposable element partial sequence candidates and the 3'-side transposable element partial sequence candidates and the corresponding TSDs are output in a single line and sorted; (D) selecting, from the pairs of excision sequences having the corresponding transposable element partial sequences, pairs of excision sequences having different TSDs in the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data as transposable element corresponding pairs, selecting a pair of candidate 5'-side transposable element partial sequences and a pair of 3'-side transposable element partial sequences in which the ratio of the TSD sequence specific to the first genome sequence data or the TSD sequence specific to the second genome sequence data to the TSD of the pair of candidate 5'-side transposable element partial sequences exceeds th, where th is 0.2 or more; (E) if a reference sequence is present, mapping the adjacent sequences at the 5' and 3' ends of the detected candidate transposable element partial sequence to identify the insertion position on the genome, or if a reference sequence is not present, confirming that the insertion is not present in the reference sequence corresponding to the assumed insertion site in a sample sequence other than the sample in which the insertion was detected. A method comprising: (Item A39-1) A method according to any of the preceding items, further comprising (C1) a step of excluding sequences in which the number of A, C, G, or T bases is 1 or less from the 5'-side transposable element partial sequence candidate and the 3'-side transposable element partial sequence candidate. (Item A40) The method according to any one of the preceding items, wherein th is 0.7. (Item A41) The method according to any one of the preceding items, wherein n is 3 to 20. (Item A42) The method according to any one of the preceding items, wherein X is 17 to 50. (Item A43) (A) A step of extracting X base sequences from short reads derived from two types of samples by shifting the sequences by one base; (B) obtaining the frequency of each of the obtained sequences of the plurality of X bases; (B1) sorting and outputting each sequence of a plurality of X bases; (B2) extracting a sequence of X bases specific to either of the two types of samples; (C1) obtaining a position on a reference sequence corresponding to the first half X / 2 base sequence of each extracted X base sequence; (C2) obtaining a position on a reference sequence corresponding to the latter X / 2 base sequence of each extracted X base sequence; (C3) sorting the location data obtained in C1 and C2 by chromosome; (D) a step of outputting, from the position data sorted in C3, TSDs formed by shifting the junction positions and having a length of s1 or more and s2 or less; (E) selecting a TSD adjacent to the 3'-end of the first X / 2 bases and a TSD adjacent to the 5'-end of the second X / 2 bases, in which there are two or more types of TSDs, and selecting from the selected dataset all pairs of X base sequences in which the TSD adjacent to the 5'-end and the TSD adjacent to the 3'-end are identical; (F) selecting pairs of the X / 2 base sequences that are present in both of the two types of samples and have two or more different TSD sequences; (G) a step of outputting pairs selected in F from the position data obtained in C1 and C2; The method according to any one of the preceding items, comprising: (Item A44) The method according to any one of the preceding items, wherein X is 20 or more. (Item A45) The method according to any one of the preceding items, wherein s1 is 3 or more. (Item A46) The method according to any one of the preceding items, wherein s2 is 20 or less.
[0007] It is contemplated that the present disclosure may provide one or more of the above-described features in combinations other than those explicitly stated. Still further embodiments and advantages of the present disclosure will be recognized by those skilled in the art upon reading and understanding the following detailed description, if necessary. [Effects of the Invention]
[0008] The method of the present disclosure can detect the transposition of novel active transposons solely through sequence analysis using a next-generation sequencer, without the need for comparison with a reference sequence. While several techniques for detecting the transposition of transposons with known sequences using nucleotide sequence information are known, the method of the present disclosure can accurately detect new transposition of novel transposons whose sequences are unknown.
[0009] Because it can find active transposons even in organisms for which no reference genome sequence yet exists, it can be used to detect and breed transposons in vegetables, fruit trees, and ornamental plants, in addition to crops. Potential applications include controlling variegation in ornamental plants, and providing disease and drought resistance in crops, fruit trees, and vegetables, increasing yield, improving taste, and altering harvest times. In animal and other fields, it can be used to identify various transposon-related characteristics (e.g., conditions including diseases and disorders, or characteristics related to individuality), and for diagnostics and trait modification based on the identified characteristics. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a schematic diagram illustrating an embodiment of a system of the present disclosure. [Figure 2] FIG. 2 is a schematic diagram of a further embodiment of the system of the present disclosure. [Figure 3] FIG. 3 is a schematic diagram of a transposition sequence identified by the method of Algorithm 1 according to one embodiment of the present disclosure. [Figure 4] FIG. 4 is a schematic diagram of a transposition sequence identified by the method of Algorithm 2 according to one embodiment of the present disclosure. [Figure 5] FIG. 5 is a schematic diagram showing the LTR arrangement of the retrotransposon Tos17. DETAILED DESCRIPTION OF THE INVENTION
[0011] The present disclosure will now be described, illustrating the best mode thereof. Throughout this specification, singular expressions should be understood to include the plural concept unless otherwise specified. Thus, singular articles (e.g., "a," "an," "the," etc. in English) should be understood to include the plural concept unless otherwise specified. Furthermore, terms used in this specification should be understood to have the meaning commonly used in the art unless otherwise specified. Therefore, unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. In the event of conflict, the present specification (including definitions) will prevail.
[0012] (definition) The following provides definitions of terms particularly used in this specification and / or explains basic technical content as appropriate.
[0013] In this specification, "or" is used when "at least one or more" of the items listed in the sentence can be employed. The same applies to "alternative." In this specification, when "within a range" of "two values" is specified, the range includes the two values themselves.
[0014] As used herein, the term "sequence" refers to a plurality of variables, each of which takes on a value, and further includes information about the order of the plurality of variables. It is typically expressed as a character string. Examples of sequences include, but are not limited to, nucleic acid sequences (substantially synonymous with DNA sequences, nucleotide sequences, etc.), amino acid sequences (substantially synonymous with peptide sequences, protein sequences, etc.), etc. The determination of a sequence is referred to as sequence determination or sequencing (or sequencing), which are synonymous.
[0015] As used herein, "subject sequence" refers to any sequence in which a polymorphism is to be detected, and may also be referred to herein as "target," "target sequence," or "target." As used herein, "control sequence" refers to any sequence used as a standard for detecting differences from that sequence as polymorphisms, and may also be referred to herein as "control," "reference sequence," "comparison sequence," or "control."
[0016] As used herein, the term "reference sequence" refers to a sequence that can be treated as the full-length sequence of a target sequence and / or a control sequence. The sequence that constitutes the full-length sequence is determined appropriately depending on the sequence used as the target sequence and / or the control sequence, and is not limited to the examples given. For example, a full genome sequence, a full-length chromosome sequence, a full-length gene sequence, a full-length plasmid sequence, a full-length exon sequence, or a full-length protein sequence present in a web-based database can be used as a reference sequence. Although sequences are sometimes referred to herein as a first sequence, a second sequence, etc., using ordinal numbers, either may be referred to as the target sequence or either may be referred to as the reference sequence, it should be noted that the target sequence and the reference sequence are assigned different ordinal numbers.
[0017] As used herein, the term "excised sequence" refers to a sequence obtained by extracting (excising) a portion from a given sequence.
[0018] As used herein, "sequence data" refers to data that provides information about a sequence. Typically, the sequence itself can be referred to as sequence data, but sequence data also includes data that provides information about a portion of a sequence (e.g., analytical data obtained by sequencing a genome sequence).
[0019] As used herein, a "subsequence" of a sequence refers to any sequence contained within that sequence.
[0020] As used herein, the term "subset" refers to any subset of a set that combines a set of sequences and a set of subsequences of those sequences.
[0021] As used herein, "next-generation sequencing" refers to a sequencing technique that parallelizes the sequencing process and generates tens to hundreds of millions of sequence data in a single run. "Next-generation sequencer" refers to a device used for next-generation sequencing.
[0022] As used herein, "eliminating coincidental identity" means reducing the expected value of a sequence that appears by chance to be identical to a given sequence to less than 1.
[0023] As used herein, "coverage" refers to how many times the amount of sequence data corresponds to the full length of the sequence. It may also be referred to as "coverage rate" or "~ times the number of reads."
[0024] As used herein, the term "array structure" refers to a physically separated series of sequences within a sequence. For example, in the context of a genome sequence, each chromosome can be referred to as an array structure.
[0025] As used herein, the term "transposition" refers to the insertion of a transposable element into a sequence construct.
[0026] (transposable element) In one aspect, the present disclosure provides novel transposable elements and techniques for identifying them.
[0027] As used herein, the term "transposable element" refers to a base sequence that can transpose (transpose) its location on the genome. Transposable elements are broadly divided into DNA types, in which DNA fragments are directly transposed, and RNA types, which undergo the processes of transcription and reverse transcription. The former are transposons in the narrow sense, while the latter are also called retrotransposons (or retroposons). In this specification, when the term "transposon" is simply used, it is not intended to be limited to transposons in the narrow sense, and unless otherwise specified, it is understood to be used synonymously with transposable element.
[0028] DNA transposons require an enzyme called transposase in order to transpose, which is encoded by the transposon itself. Transposons have inverted repeat sequences at their ends, which the transposase recognizes to excise the transposon from the genome sequence and reinsert it into the appropriate genome sequence. After being transcribed, retroposons create cDNA from mRNA using the reverse transcriptase enzyme they encode, which is then reinserted into the chromosome. Both cause mutations when inserted into gene regions, and DNA transposons can sometimes delete surrounding DNA sequences during excision, inducing chromosomal abnormalities. Furthermore, incomplete transposition can leave junk sequences in the chromosome. While "RNA transposons" move in a so-called copy-and-paste manner, "DNA transposons" move in a so-called cut-and-paste manner.
[0029] RNA-type transposable elements include long terminal repeat (LTR) retrotransposons, endogenous retroviruses, long interspersed nucleotide sequences (LINEs), short interspersed nucleotide sequences (SINEs), and non-autonomous elements known as processed pseudogenes (PPs). DNA-type transposable elements include DNA transposons and small inverted repeat transposable elements (MITEs). When these elements integrate into new genomic locations, they may duplicate the sequence of their target site. This duplication is referred to herein as "target site duplication" and may be abbreviated as "TSD." The length of this target site duplication is often unique to individual transposable elements. TSDs are generally unidirectional repeat sequences of approximately 2-20 bp. Sequences other than TSDs are sometimes referred to herein as "transposable element subsequences."
[0030] Transposons are abundant in eukaryotic genomes. Most are inactive and have been considered unnecessary junk until now. However, in rare cases, some transposons can transpose, causing mutations. For example, it has been shown that red and white wine grapes turn white when transposons are inserted into anthocyanin synthase genes, resulting in their loss of activity (Science 304:982, 2004). Among rice mutants carrying the retrotransposon Tos17, strains with mutated starch biosynthesis genes have been developed, exploiting their starch gelatinization properties, resulting in new varieties such as "Akita Barari" and "Akita Sarari." While the ability to freely control transposon transposition would enable the functional modification of various genes and potentially be of great industrial value, finding active transposons has proven extremely difficult.
[0031] Methods for detecting active transposons include creating primers from consensus sequences specific to transposons, such as those in reverse transcriptase genes, and detecting transposon transfer from the transcripts (Hirochika H et al. (1996) Retrotransposons of rice involved in mutations induced by tissue culture, Proc Natl Sci U S A., 93:7783-7788), and detecting transposons that happened to be inserted into a gene of interest (Nakazaki T et al. (2003) Mobilization of a transposon in the rice genome, Nature 421, 170-172). These transposons are isolated using experimental molecular biological techniques.
[0032] (Method for identifying transposable elements) In one embodiment of the present disclosure, transposable elements can be identified, for example, as follows. Algorithm 1 (TSD method) 1. Short reads from two samples, A and B, are extracted to a size of 20 bases + the specified TSD, each shifted by one base. 2. Sort the extracted sequences and output them along with their frequencies. 3. Exclude sequences that are present in both A and B and select sequences that are specific to A or B. 4. The selected sequences are output as the TSD sequence (Head sequence) and the Tail sequence (TSD sequence), and the possible candidate sequences are output. 5. From sequences that combine Head and Tail with TSD, select a dataset of Head and Tail sequences with two or more types of TSD. 6. Create all combinations of Head and Tail sequences with the same TSD. 7. Create a sorted file by outputting Head_Tail and TSD for A and B on one line. 8. Head or tail sequences with one or less A, C, G, or T base are excluded. 9. Select the Head_Tail pair for which the proportion of A-specific or B-specific TSD sequences in the TSD for the Head_Tail pair of interest exceeds th (for example, th is 0.2, in which case the program sets th > 0.2). 10. If a reference sequence is available, map the adjacent sequences of the detected Head and Tail sequences to determine the insertion position on the genome. 11. If no reference sequence is available, verify that the sequence corresponding to the expected insertion site in a different sample sequence from the one in which the insertion was detected does not contain the insertion.
[0033] When using this method, th is typically 0.1 or greater, and may be 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, etc., with generally good results being obtained at 0.7. Of course, if the conditions are too strict and nothing can be detected at th > 0.7, the threshold can be lowered for calculation. For example, in the example of detecting P elements in Drosophila, this example shows that transposition can be detected by setting the TSD size to 8 and lowering th > 0.2.
[0034] Furthermore, if there is no reference sequence, the absence of a transposable element insertion in the sequence of another sample can be used as proof that transposition has occurred. For example, because the rice genome is diploid, in the generation immediately following transposition, the genotype of the new transposition will be heterozygous, and sequences without insertions can be detected even among the sequences of samples in which insertions have been detected. In addition, in the case of ttm2 and ttm5, because the F4 generation was selfed four times, there are also homozygous portions of the transposition site. In such cases, confirmation can be made using the sequence of an individual other than the one in which transposition was detected.
[0035] In one embodiment, when analyzing sequences of an organism for which no reference sequence exists, the candidate terminal sequences of the detected transposable element can be used as a query to search the sequences of two or more individuals to detect evidence of transposition.
[0036] In one embodiment, when a reference sequence is available, the insertion position can be identified by searching for the insertion flanking sequence against the 20-base sequence at every position of the reference sequence, which also allows identification of the insertion position in the repeat region.
[0037] When a transposable element transposes and inserts into a genomic DNA sequence, it creates a duplication of several bases at the target site (Fig. 3a). In the case of Algorithm 1, the sequence into which the transposable element has inserted is cut into pieces of a fixed length, and among these, there are sequences containing target site duplications (TSDs) adjacent to the 5' and 3' ends of the transposable element. Furthermore, since the TSD sequence differs for each insertion site, grouping by TSD sequence makes it possible to detect the 5' and 3' end sequences of the transposable element (Fig. 3b).
[0038] In one embodiment of the present disclosure, a transposable element can be identified as follows. Algorithm 2 (Junction Method) 1. Short reads from two samples, A and B, are cut into 40-base pairs, each shifted by one base. 2. Sort the extracted 40-base sequences and output them along with their frequency. 3. Extract a 40-base sequence specific to A or B. 4. Determine the genomic location of the first 20 bases of the sequence. If the sequence contains more than nine repeats of AC, AG, AT, TC, or TG, exclude them. 5. Determine the genomic location of the last 20 bases of the sequence. Exclude sequences containing 9 or more repeats of AC, AG, AT, TC, or TG. 6. After separating the first and second 20 bases of data by chromosome, sort them by position. 7. From the sorted data, the junction positions are shifted to form TSDs, and data with a TSD size of 4 to 10 is output. 8. Select a head and tail pair with multiple types of TSD. 9. Select a pair of head and tail that exists in both A and B and has two or more different TSD sequences. 10. Select and output the head and tail pairs selected from the mapped data.
[0039] In this case, the number of bases to be excised is not limited to 40 bases, but can be any length. For example, a length of 20 bases or more can be excised. In one embodiment, when the excised base length is divided into a first half and a second half in steps 4 and 5, the positions of the first half and the second half on the reference sequence can be determined by dividing the first half and the second half, and it is not necessary to divide it in half.
[0040] In one embodiment, when the extracted base length is divided into the first and second halves, the base lengths of each part are the same, so that the position data can be subdivided into parts of the same length and searched for an exact match with the reference data paired with the position data.
[0041] In the case of algorithm 2, TSDs in the first or second half of the excised sequence can be found, and transposons present in the adjacent sites can be detected (Figure 4. For each excised sequence, sequence numbers 29 to 40 are listed in order of appearance from top to bottom for individual A, and then from top to bottom for individual B).
[0042] Retrotransposons generally have duplicated sequences called long terminal repeats (LTRs) at both ends, meaning that the sequence near the 5' end of a transposon is also present in the downstream LTR, and the sequence near the 3' end is also present in the 3' end of the upstream LTR (Figure 5).
[0043] Algorithm 1 allows you to specify the length of the TSD for analysis. This means that the same analysis is repeated by changing the TSD length by one base. Algorithm 2 maps the first 20 bases of a 40-base sequence and the remaining 20 bases on the genome to detect junctions, and then performs TSD matching, so there is no need to specify the TSD length at first.
[0044] Furthermore, Algorithm 1 can detect transposition of transposable elements by comparing sequences from two samples even when there is no reference sequence. Algorithm 2 first maps to a reference sequence, and the original transposable element exists on the reference. Creating a reference sequence from an individual containing an active transposable element is difficult because inconsistencies occur in the assembly of the transposed transposable element sequence. As a result, there is a high possibility that the reference sequence will select an individual that does not normally exhibit transposable element activity. Algorithm 2 can be applied to the detection of transposable elements that are activated and transposed only under special circumstances, such as Tos17 in rice.
[0045] Thus, in one aspect of the present disclosure, a method for identifying transposable elements in a sequence (e.g., a nucleic acid sequence) is provided. Transposable elements are known to generate target site duplications (TSDs) of several bases when transposing. Therefore, it is believed that the sequences of both ends of the transposed transposable element are conserved. The method of the present disclosure can detect transposable elements in a sequence, at least in part, using this as a clue. The method of the present disclosure can detect transposable elements without requiring reference genome sequence information. For example, transposable elements can be detected by analyzing short reads from a next-generation sequencer. The method of the present disclosure can be used to detect transposable elements with unknown sequences. Furthermore, the method of the present disclosure can detect active transposable elements that are actually undergoing transposition.
[0046] In one aspect, the present disclosure provides a method for identifying a transposable element in a sequence, comprising: (A) extracting sequences of a certain length from the first genome sequence data and the second genome sequence data, shifting the sequences by one base at a time, and generating a set of extracted sequences from the first genome sequence data and a set of extracted sequences from the second genome sequence data; (B) separating each of the excision sequences into sequences corresponding to n-base-long target site overlaps (TSDs) and sequences other than the TSDs (transposable element partial sequences), and selecting pairs of excision sequences in which the 5'-terminal TSD and the 3'-terminal TSD are identical; (C) selecting pairs of excised sequences having corresponding transposable element partial sequences from the selected pairs; (D) selecting, from pairs of excision sequences having corresponding transposable element partial sequences, pairs of excision sequences having different TSDs in a set of excision sequences from the first genome sequence data and a set of excision sequences from the second genome sequence data as transposable element corresponding pairs.
[0047] In this aspect of the present disclosure, a plurality of genome sequence data are compared to detect the transposable element that transposes between them.Typically, first and second genome sequence data are compared, but more than two genome sequence data can be compared.For example, 2, 3, 4, 5, 8, 10 or more genome sequence data can be used.For example, about 10, about 15, about 20, about 50, about 100, about 200, about 500, about 1000 or more genome sequence data can be used.As one embodiment of the present disclosure, both (or either) of the first genome sequence data and the second genome sequence data can be mixed and analyzed with a plurality of genome sequence data from different origins.
[0048] The genome sequence data may be a reference genome sequence or sequencing data of a genome sequence. The genome sequence data may be recorded in a database or the like, or may be obtained by newly determining the sequence.
[0049] In one embodiment, the set of excision sequences can be generated by obtaining sequences of a fixed length while shifting the start position of each sequence in the genome sequence data by one base at a time. The fixed length can be, for example, within the range of 17 to 50 bases. An example of the fixed length is the size of 20 bases + TSD, such as 23, 24, 25, 26, 27, or 28 bases. Setting the fixed length to 17 bases or more is believed to eliminate the possibility of coincidence in the excision sequences. Furthermore, the longer the fixed length, the higher the accuracy is believed to be, but the larger the amount of data to be handled. While this is not a problem in an environment that can handle large amounts of data, it is preferable to set the upper limit to approximately 35 to 50 bases. The constant length is, for example, a base length in the range of 10 to 100, such as a base length in the range of 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, or 90 to 100. The size of the TSD can be any length, for example, 3 to 20 bases.
[0050] In one embodiment, the step of generating a set of excised sequences may include excising the fixed-length sequence from the complementary strand of the first genome sequence data and / or the second genome sequence data by shifting the sequence by one base at a time. Although not necessary if there is a sufficient amount of sequence data, including the excised sequence of the complementary strand in the analysis can improve detection sensitivity when the amount of sequence data is small.
[0051] In one embodiment, the set of excised sequences is a set of sequences of a certain length. The entire set may be used for analysis, but some sequences may be excluded from the analysis as needed. Appropriately excluding sequences from the analysis not only avoids the effects of errors during sequencing, but also reduces the amount of calculation, which may be advantageous in handling large amounts of data. In the present disclosure, when an unknown base (generally represented as N) is included in the excised sequence, the sequence may be excluded from the analysis. may be excluded from
[0052] In some embodiments of the present disclosure, the method may further include a step of calculating the frequency of identical excision sequences from a set of excision sequences of a certain length. Furthermore, the method of the present disclosure may further include a step of excluding excision sequences with a frequency below a certain level. Preferably, excision sequences with a frequency below a certain level can be excluded when using short-read genome sequences, while excision sequences with a frequency below a certain level do not need to be excluded when using assembled genome sequences. The frequency of identical excision sequences can be calculated, for example, by sorting the excision sequences lexicographically. Furthermore, sequences with a certain number of identical sequences or less can be excluded from analysis as they are likely to contain sequencer errors. The reference frequency can be determined by taking into account the balance between the amount of sequence data and the genome size of the target organism (sequence coverage). For example, if the sequence coverage is 40, sequences with 1, 2, 3, 4, or 5 or fewer identical sequences can be excluded from analysis. In rice, empirically, excluding sequences with 5 or fewer identical sequences from analysis can almost completely eliminate coincidences due to noise. In other embodiments, the sequence coverage can be any value, for example, the sequence coverage can be 20.
[0053] In one embodiment, the method of the present disclosure may include a step of selecting excision sequences that differ between a set of excision sequences from the first genome sequence data and a set of excision sequences from the second genome sequence data. Since transposons that have undergone transposition are thought to have different TSD sequences, while the transposon sequence is conserved, excision sequences containing the end of the transposon that has undergone transposition and the TSD sequence are thought to be specific to each genome sequence data. Therefore, the excision sequences selected here are thought to include a set of the transposon end and the TSD sequence. This selection can be performed before step (B) above. Different excision sequences are typically those that are present in the set of excision sequences from the first genome sequence data but not in the set of excision sequences from the second genome sequence data, or those that are present in the set of excision sequences from the second genome sequence data but not in the set of excision sequences from the first genome sequence data. However, excision sequences that are present in large numbers in the set of excision sequences from the first genome sequence data and present in small numbers in the set of excision sequences from the second genome sequence data, or that are present in large numbers in the set of excision sequences from the second genome sequence data and present in small numbers in the set of excision sequences from the first genome sequence data, can also be considered as different excision sequences.For example, when the abundance of the excision sequence is at least about 2 times, at least about 5 times, at least about 10 times, at least about 20 times, at least about 40 times, at least about 50 times, at least about 100 times, at least about 200 times, at least about 500 times, or at least about 1000 times, or more, it can be treated as different excision sequences in the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data.When the TSD is short, the possibility of coincidences occurring increases, so on the other hand, in addition to sequences that do not exist at all, it may be preferable to consider sequences with a certain or greater difference in abundance.For example, a three-base TSD has 64 possible combinations, which is thought to result in a considerable number of coincidences. To detect transposable elements more comprehensively, sequences with abundance differences of up to about five-fold may be considered.
[0054] In one embodiment, the selection of pairs of excision sequences having the same TSD at the 5' end and the same TSD at the 3' end in step (B) can be performed from among excision sequences that are different between the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data, thereby significantly reducing the enormous amount of calculation required to detect pairs.
[0055] In some embodiments, the method of the present disclosure may include a step of distinguishing each excision sequence into a sequence corresponding to an n-base-long target site overlap (TSD) and a sequence other than the TSD (transposable element partial sequence), and selecting a pair of excision sequences in which the 5'-terminal TSD and the 3'-terminal TSD are identical. The length of n bases can be set, taking into account the possible lengths of TSDs, for example, within the range of 3 to 20 (i.e., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20), for example, within the range of 3 to 10 (i.e., 3, 4, 5, 6, 7, 8, 9, 10). Because the method of the present disclosure is believed to be effective in detecting transposable elements having a TSD with a base length of n, the method may further include a step of repeating the steps of the method of the present disclosure (e.g., steps (B) to (D) described above) while changing n.
[0056] The step of distinguishing an excised sequence into a sequence corresponding to a TSD and a transposable element partial sequence can be carried out by generating, from a certain excised sequence, a set of an n-base sequence from the 5' side and the remaining sequence, and a set of an n-base sequence from the 3' side and the remaining sequence. Because a TSD is generally a unidirectional repeat, it is possible to determine whether the 5'-terminal TSD and the 3'-terminal TSD are identical by directly comparing the 5'-to-3' sequence of the n-base sequence from the 5' side with the 5'-to-3' sequence of the n-base sequence from the 3' side.
[0057] In one embodiment, it is believed that pairs of excision sequences corresponding to actual transposons can be detected by selecting pairs of sequences in which the 5'-end TSD and the 3'-end TSD are identical, and that satisfy the following two conditions between the genome sequence data: (i) the sequences of the transposon portions correspond; and (ii) the sequences of the TSD portions differ.
[0058] In certain embodiments, the method of the present disclosure may further include selecting pairs of excision sequences having corresponding transposable element sequences from the selected pairs of excision sequences having identical 5'- and 3'-terminal TSDs. Examples of corresponding transposable element sequences include transposable element sequences with at least about 90% identity, at least about 95% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, or identical transposable element sequences. Because mutations may occur in the transposable element sequences upon transposition, processing may take this into account. However, excluding mutations and considering only perfectly matching transposable element sequences may significantly reduce the amount of calculation and the time required for calculation.
[0059] In certain embodiments, the disclosed method may further include a step of selecting pairs of excision sequences having different TSDs between a set of excision sequences from the first genome sequence data and a set of excision sequences from the second genome sequence data from pairs of excision sequences having corresponding transposable element partial sequences. The pairs selected in this step may be referred to as transposable element match pairs. As used herein, the term "transposable element match pair" refers to a pair of fixed-length excision sequences believed to contain the end of a transposable element and its outer TSD. For example, each excision sequence can be located on a reference sequence, or the sequence of the region between each excision sequence can be determined from genome sequence data. Thus, the disclosed method may further include a step of identifying transposable element sequences based on the transposable element match pair.
[0060] The different TSD sequences are typically present in a pair of excised sequences from the first genome sequence data and not present in a pair of excised sequences from the second genome sequence data, or present in a pair of excised sequences from the second genome sequence data and not present in a pair of excised sequences from the first genome sequence data. A TSD sequence that is not present in a pair of excised sequences from the first genome sequence data can be considered to be a different TSD sequence. However, a TSD sequence that is present in large numbers in a pair of excised sequences from the first genome sequence data and in small numbers in a pair of excised sequences from the second genome sequence data, or a TSD sequence that is present in large numbers in a pair of excised sequences from the second genome sequence data and in small numbers in a pair of excised sequences from the first genome sequence data, can also be considered to be a different TSD sequence. For example, if the abundance of a certain TSD sequence differs by at least about 2-fold, at least about 5-fold, at least about 10-fold, at least about 20-fold, at least about 40-fold, at least about 50-fold, at least about 100-fold, at least about 200-fold, at least about 500-fold, or at least about 1000-fold or more, the pair of excised sequences from the first genome sequence data can be treated as a different TSD sequence. If the majority of TSD sequences differ between the first and second genome sequence data, the probability of a transposon is considered high. However, even if most TSDs are detected in both genomes, if only a few, for example, one, that differ are detected, this may be the result of a very low frequency of transposition, and such cases may also be detected.
[0061] In a further embodiment, once a candidate transposable element sequence has been identified, a further step of confirming the transposition activity of the transposable element can be performed, if necessary, using known methods. The transposition activity of the transposable element can be confirmed, for example, by analyzing samples in which transposition of the transposable element is suspected. For example, the transposition activity of the transposable element can be confirmed by a technique including one or more selected from PCR, sequencing, and hybridization. The transposable element can also be confirmed by using other transposable element detection techniques known in the art. Examples of other transposable element detection techniques include Transposon Insertion Finder (TIF) (Nakagome M, Solovieva E, Takahashi A, Yasue H, Hirochika H, Miyao A (2014) Transposon Insertion Finder (TIF): a novel program for detection of de novo transpositions of transposable elements. BMC Bioinformatics 5:71. doi:10.1186 / 1471-2105-15-71).
[0062] (Schematic embodiment) This section provides a schematic example of a method for detecting a transposable element. This section is provided for illustrative purposes and is not intended to limit the methods of the present disclosure. Assuming that the TSD is 5 bases, and focusing on both ends of the transposon (20 bases each) and the TSD (5 bases), the sequence derived from the transposon is as follows on the genome: [ka] In this diagram, "_" indicates the base sequence outside the transposon, while "+" and the base sequences corresponding to the Head and Tail indicate the base sequence of the transposon. If 25 bases are extracted from all short reads of the sequence data obtained by sequencing the genome, starting with all bases, this will likely include TSD-Head and Tail-TSD. Because transposon transposition is an independent event, when comparing the sequences of two individuals, the TSD sequences will differ even if the same transposon is transposed. Therefore, if a set of 25 bases is obtained from a 25-base sequence (5 bases from the TSD and 20 bases from the Head or Tail) with the same Head and Tail sequences but different TSDs between individuals, this set is considered to represent both ends of the transposon and the TSD.
[0063] (Genome sequence data) In certain embodiments, the present disclosure may use two or more genome sequence data, which may differ due to the transposition of transposable elements. Because transposable elements may transpose within the genome of a cell, genome sequence data derived from different cells of the same individual may be used. The present disclosure may also use genome sequence data derived from a single cell.
[0064] Examples of combinations of origins of the first and second genome sequence data in the present disclosure include, for example, a first cell and a second cell of an individual, a first tissue and a second tissue of an individual, a first individual and a second individual of a biological species, a first population and a second population of a biological species, an individual under first and second conditions, a cell under first and second conditions, etc.
[0065] The biological species from which the genome sequence data of the present disclosure is derived are not limited as long as they possess biological sequences. Examples of animals include vertebrates such as humans or non-human mammals (e.g., mice, rats, rabbits, sheep, pigs, cows, horses, cats, dogs, monkeys, and chimpanzees), birds, reptiles, amphibians, and fish, as well as invertebrates such as insects and nematodes. Examples of plants include rice, wheat, corn, potato, barley, sweet potato, buckwheat, Arabidopsis, Lotus japonicus, tomato, cucumber, cabbage, Chinese cabbage, eggplant, sugarcane, sorghum, apple, mandarin orange, banana, peach, poplar, pine, cedar, angiosperms, gymnosperms, ferns, mosses, and algae. Other examples include fungi, bacteria, and viruses. Genomic sequence data derived from parts of these organisms, such as tissues and cells, may be analyzed to detect transposable elements.
[0066] In one embodiment, the genome sequence data used in the method of the present disclosure is base sequence data obtained by sequencing, including Sanger sequencing, Maxam-Gilbert sequencing, single-molecule real-time sequencing (e.g., Pacific Biosciences, Menlo Park, California), ion semiconductor sequencing (e.g., Ion Torrent, South San Francisco, California), pyrosequencing (e.g., 454, Branford, Connecticut), ligation sequencing (e.g., SOLiD sequencing, Life Technologies, Carlsbad, California), synthetic and reversible terminator sequencing (e.g., Illumina, San Diego, California), nucleic acid imaging technologies such as transmission electron microscopy, and nanopore sequencing.
[0067] In one embodiment, the genome sequence data used in the method of the present disclosure can be the sequence data obtained by next-generation sequencing.Next-generation sequencing includes sequencing by synthesis, pyrosequencing, ligation sequencing, ion semiconductor sequencing, and nanopore sequencing.In detecting polymorphisms, including transposable element differences, from next-generation sequencing data, accuracy has been limited by mapping to reference or assembly, so it is believed that the method of the present disclosure can be used to great advantage.
[0068] The genome sequence data of the present disclosure may be any of the genome sequence data described above obtained from a database or the like, or may be newly sequenced and used. The genome sequence data may be a set of output (reads) from a sequencer, a set of contigs compiled from reads, a scaffold formed by further connecting contigs, the genome sequence itself, or any combination thereof.
[0069] (Programs, recording media and systems) In one embodiment, a program for causing a computer to execute a method for identifying a transposable element in a sequence, the method comprising: (a) receiving first genome sequence data and second genome sequence data; (b) extracting sequences of a fixed length from the first genome sequence data and the second genome sequence data, shifting the sequences by one base at a time, to generate a set of extracted sequences from the first genome sequence data and a set of extracted sequences from the second genome sequence data; (c) classifying each of the excision sequences into a sequence corresponding to an n-base-long target site overlap (TSD) and a sequence other than the TSD (transposable element partial sequence), and selecting a pair of excision sequences in which the 5'-terminal TSD and the 3'-terminal TSD are identical; (d) selecting pairs of excised sequences having corresponding transposable element partial sequences from the selected pairs; (e) selecting, from the pairs of excision sequences having the corresponding transposable element partial sequences, pairs of excision sequences having different TSDs in the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data as transposable element corresponding pairs; The program may be written in any language.
[0070] In another embodiment, a recording medium storing a program for causing a computer to execute a method for identifying a transposable element in a sequence, the method comprising: (a) receiving first genome sequence data and second genome sequence data; (b) extracting sequences of a fixed length from the first genome sequence data and the second genome sequence data, shifting the sequences by one base at a time, to generate a set of extracted sequences from the first genome sequence data and a set of extracted sequences from the second genome sequence data; (c) classifying each of the excision sequences into a sequence corresponding to an n-base-long target site overlap (TSD) and a sequence other than the TSD (transposable element partial sequence), and selecting a pair of excision sequences in which the 5'-terminal TSD and the 3'-terminal TSD are identical; (d) selecting pairs of excised sequences having corresponding transposable element partial sequences from the selected pairs; (e) selecting, from the pairs of excision sequences having the corresponding transposable element partial sequences, pairs of excision sequences having different TSDs in the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data as transposable element corresponding pairs; The program may be written in any language. In one embodiment, the recording medium may be an internally stored ROM or an external storage device such as a HDD, a magnetic disk, or a flash memory such as a USB memory. The recording medium may be non-transitory.
[0071] In another embodiment, a system for identifying transposable elements in a sequence comprises one or more processors, a memory, and a recording medium storing a program, the program, when executed by the one or more processors, performing the following: (a) receiving first genome sequence data and second genome sequence data; (b) extracting sequences of a fixed length from the first genome sequence data and the second genome sequence data, shifting the sequences by one base at a time, to generate a set of extracted sequences from the first genome sequence data and a set of extracted sequences from the second genome sequence data; (c) classifying each of the excision sequences into a sequence corresponding to an n-base-long target site overlap (TSD) and a sequence other than the TSD (transposable element partial sequence), and selecting a pair of excision sequences in which the 5'-terminal TSD and the 3'-terminal TSD are identical; (d) selecting pairs of excised sequences having corresponding transposable element partial sequences from the selected pairs; (e) selecting, from the pairs of excision sequences having the corresponding transposable element partial sequences, pairs of excision sequences having different TSDs in the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data as transposable element corresponding pairs; A system is provided that implements the method, including the steps of: (a) providing a program for executing the program; (b) providing a program for executing the program; (c) providing a program for executing the program; (d) providing a program for executing the program; (e) providing a program for executing the program; (f) providing a program for executing the program; (g) providing a program for executing the program; (h) providing a program for executing the program; (i) providing a program for executing the program; (ii) providing a program for executing the program; (iii) providing a program for executing the program; (iv) providing a program for executing the program; (v) providing a program for executing the program; (vi ...
[0072] Next, the configuration of the system 1 of the present disclosure will be described with reference to the functional block diagram of Figure 1. Note that while this figure shows a case where the system is realized in a single system, it will be understood that a case where the system is realized in multiple systems is also included within the scope of the present disclosure.
[0073] A system 1000 of the present disclosure is configured by connecting a CPU 1001 built into a computer system to a RAM 1003, an external storage device 1005 such as a ROM, HDD, magnetic disk, or flash memory such as a USB memory, and an input / output interface (I / F) 1025 via a system bus 1020. An input device 1009 such as a keyboard or mouse, an output device 1007 such as a display, and a communication device 1011 such as a modem are connected to the input / output I / F 1025. The external storage device 1005 includes an information database storage unit 1030 and a program storage unit 1040. Each of these is a fixed storage area secured within the external storage device 1005.
[0074] In such a hardware configuration, when various commands are input via the input device 1009 or when commands are received via the communication I / F or communication device 1011, the software program installed in this storage device 1005 is called, deployed, and executed on the RAM 1003 by the CPU 1001, thereby working in cooperation with the OS (operating system) to perform the function of a method for detecting transposons in genome sequence data. Of course, the present disclosure can also be implemented using mechanisms other than such cooperation.
[0075] In a specific embodiment, for implementing the present disclosure, genome sequence data may be input via the input device 1009, or may be input via a communication I / F or communication device 1011, or may be stored in the database storage unit 1030. Data on the identified pairs of excised sequences may be output via the output device 1007 or stored in an external storage device 1005, such as the information database storage unit 1030. Next, the excision, comparison, and / or selection of sequences may be performed by a program stored in the program storage unit 1040, or a software program installed in this external storage device 1005 by inputting various instructions (commands) via the input device 1009 or receiving commands via the communication I / F or communication device 1011, or the like. The results may be output via the output device 1007 or stored in the external storage device 1005, such as the information database storage unit 1030.
[0076] These data, calculation results, or information acquired via the communication device 1011, etc., are written and updated as needed in the database storage unit 1030. By managing information such as each sequence in each input sequence set and each gene information ID in the reference database in each master table, it becomes possible to manage information attributable to samples to be accumulated using IDs defined in each master table.
[0077] Furthermore, the computer programs stored in the program storage unit 1040 configure the computer as the above-mentioned processing system, for example, a system that performs processing such as detecting transposition elements. Each of these functions is an independent computer program, its module, or routine, and configures the computer as each system or device when executed by the CPU 1001. Note that in the examples of the present disclosure, the functions in each system cooperate to configure each system, and the programs for this processing can also be provided via an external storage device, communication device, or input device.
[0078] The method of the present disclosure may also be implemented using a computing system with a cluster structure, as shown in FIG. 2. In one embodiment, the system has a cluster configuration and is composed of a head and nodes. The nodes can use SSDs as their main storage devices to speed up searches. In one embodiment, one head can be used with multiple nodes (e.g., 12 nodes). In one embodiment, the computing system has a cluster structure, and the main computer (cluster head) is equipped with a mass storage device (HDD) to store analysis data and results. The cluster head sends divided data to each node, which performs calculations and consolidates the results back to the cluster head. Both the cluster head and the nodes are equipped with a central processing unit (CPU) and memory (RAM), and may communicate data via a communication interface (NIC). The nodes can use solid-state drives (SSDs) as their main storage devices to enable high-speed search processing. The CPU, RAM, SSD, etc. installed in each node may be shared with other nodes or may be physically separated.
[0079] In another aspect, the present disclosure provides a system for identifying transposable elements in a sequence, the system comprising: (A) a sequence data receiving unit that receives first genome sequence data and second genome sequence data; and (B) a sequence data receiving unit that extracts sequences of a certain length from the first genome sequence data and the second genome sequence data obtained by the sequence data receiving unit, shifting the sequences by one base at a time, to generate a set of extracted sequences from the first genome sequence data and a set of extracted sequences from the second genome sequence data, and classifies each of the extracted sequences into a sequence corresponding to an n-base-long target site duplication (TSD) and a sequence other than the TSD (transposable element). a transposable element matching pair selection unit that executes the steps of: distinguishing between pairs of excision sequences in which the 5'-end TSD and the 3'-end TSD are identical; selecting pairs of excision sequences having corresponding transposable element partial sequences from the selected pairs; and selecting pairs of excision sequences having different TSDs in the set of excision sequences from the first genome sequence data and the set of excision sequences from the second genome sequence data as transposable element matching pairs; and (C) a display unit that displays the transposable element matching pairs.
[0080] In this aspect, the sequence data receiving unit can be realized by an input device 1009 included in the input / output I / F 1025, and the display unit can be realized by an output device 1007 such as a display. The transposon corresponding pair selection unit can be realized by any embodiment of the implementation method of the above method. For example, a computer program stored in the program storage unit 1040 can cause a computer to function as the above processing system, thereby realizing the function of the transposon corresponding pair selection unit.
[0081] (Identified transposable elements) In a further aspect of the present disclosure, a nucleic acid is provided having at least a partial (or possibly entire) sequence of a transposable element identified by the method of the present disclosure. The nucleic acid may have a recognition sequence for a transposase in the transposable element. It is believed that by identifying an element involved in the transposition of an actually active transposable element in a genomic sequence and introducing a nucleic acid containing the element, it can be used to introduce a mutation into the genomic sequence. It is believed that by identifying an element involved in the transposition of an actually active transposable element in a genomic sequence and introducing a nucleic acid containing the element, it can be used for gene disruption by transposition of the transposable element. Alternatively, it is believed that by identifying an element involved in the transposition of an actually active transposable element in a genomic sequence and introducing a nucleic acid containing the element, it can be used for transcription control. Because transposable elements transpose randomly throughout the entire genome, they can be used to create insertional mutant lines that span the entire genome.
[0082] Furthermore, in the present disclosure, transposable elements that are activated and transposed under certain conditions can also be used. For example, Tos17 is activated and transposed only during culture, but when the cultured cells are redifferentiated into plants, Tos17 is inactivated and no longer transposed. DNA is extracted from the same individual under certain conditions and other conditions, and the analysis of the present disclosure is performed. If a transposable element is detected, the transposable element is considered to be a candidate for a transposable element that transposes under certain conditions.
[0083] An example of an application of the obtained transposable element is a koji mold transposon (e.g., https: / / www.jstage.jst.go.jp / article / jbrewsocjapan / 105 / 6 / 105_6_334 / _pdf). In this case, the transposon activity can be utilized to breed practical koji mold strains, and the technology of the present disclosure can be applied.
[0084] Retrotransposons can also be used to identify the varieties of raw materials used in processed foods, such as sweet potato. For example, as mentioned in Breeding Research 6: 169-177 (2004) (https: / / www.jstage.jst.go.jp / article / jsbbr / 6 / 4 / 6_4_169 / _pdf), retrotransposon replication sequences are scattered throughout the genome and are known to be excellent genetic markers. Therefore, using retrotransposons as multilocus probes makes it possible to identify all six Japonica and six Indica rice varieties. These technologies can be applied to identify the varieties of raw materials used in processed foods.
[0085] Furthermore, as introduced in "Silkworm Technology: Evolving from a Clothing Revolution to a Medical Revolution" (https: / / www.yakult.co.jp / healthist / 213 / img / pdf / p02_07.pdf), the transposons identified in this disclosure can be applied to efficient genetic modification of insects. Furthermore, as introduced in "Characteristics and Industrial Applications of Ion Beam Breeding Technology" (https: / / katosei.jsbba.or.jp / download_pdf.php?aid=268), it is possible to modify plants such as flowers, and the modifications can be strengthened by ion irradiation.
[0086] Furthermore, analysis of genome function can be performed by comprehensively generating mutant mice using transposons (for example, http: / / lifesciencedb.jp / houkoku / pdf / A-42_final.pdf), and this technology can be applied to the transposons of the present disclosure.
[0087] (application) The techniques provided in this disclosure for identifying transposable elements can be used to examine and identify events involving transposable elements.
[0088] For example, it is possible to diagnose and analyze cognitive functions of the brain (see, for example, the report on the discovery of a retrotransposon-derived acquired gene important for brain cognitive functions (http: / / www.tmd.ac.jp / mri / press / press29 / index.html)). Here, it has been demonstrated that the LTR retrotransposon-derived gene "SIRH11 / ZCCHC16" plays an important role in reactions such as attention and recognition in the mammalian brain, and the method of the present disclosure can be used to test and analyze such cognitive functions.
[0089] They can also be used to identify and analyze DNA mutations (see, for example, "Uncovering the mechanism by which DNA moves on DNA - A movement strategy that cleverly utilizes host factors" (http: / / www.kyoto-u.ac.jp / ja / research / research_results / 2019 / 190829_1.html); also see Miyoshi T. et al., Molecular Cell, Volume 75, ISSUE 6, P1286-1298 (https: / / doi.org / 10.1016 / j.molcel.2019.07.018)). LINEs are believed to suppress DNA mutations involved in canceration, and the present disclosure can be applied directly or indirectly.
[0090] Furthermore, with regard to brain analysis, it can also be used to analyze the formation of the nervous system (for example, The ability to create new neurons in the adult brain - National Institute of Advanced Industrial Science and Technology (https: / / www.aist.go.jp / Portals / 0 / resource_images / aist_j / aistinfo / aist_today / vol10_05 / vol10_05_p12.pdf)). As reported individually, retrotransposons are mobile genes that have explosively increased in proportion to their genomes as mammals evolved during the course of evolution, and it has been reported that when the Wnt3a signal is activated, retrotransposon sequences are also activated, and genes near retrotransposon sequences are indirectly affected by this, with their expression levels being affected, and it is believed that there are genes related to neurological diseases and important genes that regulate neurological function, and it is suggested that the present disclosure may also be applied to diagnosis and testing.
[0091] Furthermore, as reported in Chen, JM., Stenson, PD, Cooper, DN et al. Hum Genet (2005) 117: 411. (https: / / link.springer.com / article / 10.1007%2Fs00439-005-1321-0) and W' Waves Vol. 17 No. 1 2011 (http: / / www.npo-jsct.umin.jp / wwaves / WWAVES%20Vol.17_p044.pdf), it has been shown that various genetic diseases are caused by the transposition of a retrotransposon called Line-1 (Table 2), and it is understood that the method of the present disclosure can be applied to the diagnosis of these diseases (genetic diseases). Thus, the relationship of Line-1 with genetic diseases and cancer has been explained, and the method of the present disclosure can be applied to these diseases.
[0092] (General technology) The molecular biological techniques, biochemical techniques, microbiological techniques, and bioinformatics used herein are any techniques that are known, well-known, or commonly used in the art.
[0093] All references cited herein, including scientific literature, patents, patent applications, and the like, are incorporated by reference in their entirety to the same extent as if each were specifically set forth.
[0094] The present disclosure has been described above by showing preferred embodiments to facilitate understanding of the present disclosure. The present disclosure will be described below based on examples. However, the above description and the following examples are provided for illustrative purposes only and are not intended to limit the present disclosure. Therefore, the scope of the present disclosure is not limited to the embodiments or examples specifically described herein, but is limited only by the scope of the claims. [Example]
[0095] Examples are described below. The handling of organisms used in the following examples complied with the standards stipulated by the implementing organization and regulatory agency, where necessary. Perl was used as the programming language in the examples shown below, but similar results can be obtained using other programming languages.
[0096] (General Information Processing) The sequence information processing procedure performed in this example is described below. Two sets of genome sequence data (here, sequence data determined by a next-generation sequencer) were prepared for the target organism or individual, and analysis was performed using the following procedure.
[0097] 1. Extract a sequence of 20 bases + TSD size from every position of the sequence data of the next-generation sequencer. In this example, the TSD size is assumed to be 5 bases, and analysis of a 25-base sequence is illustrated. The complementary strand sequence of the base sequence data is also extracted in the same way. If N is included in the extracted 25 bases, it will not be included in subsequent analysis.
[0098] 2. Sort the extracted 25-base sequences in lexicographical order and check the number of identical sequences.
[0099] 3. Sequences with five or fewer identical sequences are likely to contain sequencer errors, so only select sequences with a frequency of five or more.
[0100] 4. Compare the target and control sequence sets and select a 25-base sequence that is target-specific or control-specific. The selected sequence should include the transposon ends and the TSD sequence set.
[0101] 5. For all target or control-specific sequences, cut them into pairs of the first 20 bases and the remaining 5 bases. The first 20 bases are candidates for the 3' end of the transposon, and the remaining 5 bases are candidates for the TSD. Also, obtain the complementary strand sequence of the same 25-base sequence and cut it into pairs of the first 5 bases and the remaining 20 bases. In this case, the first 5 bases are candidates for the TSD, and the remaining 20 bases are candidates for the 5' end of the transposon.
[0102] 6. Obtain a pair of 5'-end and 3'-end candidates with the same TSD sequence. Pairs of 5'- and 3'-end candidates with different TSD sequences may represent the end sequences of a newly transposed transposon.
[0103] 7. Among the pairs of ends with different TSD sequences, we find a pair where the majority of the TSD sequences differ between the target and control. We identify this pair as corresponding to the end sequences of the newly transposed transposon.
[0104] The following examples show the detection of transposition sites in two Tos17 individuals, where the site is known, and the detection of previously unreported transposition in a ddm1 mutant. The system disclosed here is believed to be the first to demonstrate such a phenomenon.
[0105] Example 1: Detection of known transposons ttm2 and ttm5 are mutant panel strains in which Tos17 activated by cell culture has been transposed. In this example, the genome sequence data of ttm2 and ttm5 were used for analysis. These genome sequence data were analyzed in Nakagome M, Solovieva E, Takahashi A, Yasue H, Hirochika H, Miyao A. Transposon Insertion Finder (TIF): a novel program for detection of de novo transpositions of transposable elements. BMC Bioinformatics. 2014 Mar 14;15:71. doi: 10.1186 / 1471-2105-15-71. ttm2 is registered under the accession number SRR556173, and ttm5 is registered under the accession numbers SRR556174 and SRR556175.
[0106] These genome sequence data were processed as described in (General Information Processing). The results are shown below.
[0107] The 5' end (first column) and 3' end (next column) of Tos17 were detected, and three TSDs (before |) were detected in ttm2 and eight TSDs (after |) in ttm5. The TSD sequences matched those previously analyzed and reported. Two lines are output, but the second line is the complementary strand of the first line, indicating that Tos17 is the only transposon that transposes in cell culture. [ka] (The 20 base sequences of the transposon end candidates correspond to SEQ ID NOS: 1 to 4 in order of appearance.)
[0108] As described above, the detected TSD matched the sequence detected in a previously reported experiment, demonstrating that retrotransposon transposition can be detected using this method.
[0109] (Example 2: Detection of unknown transposons) We tested whether the method of the present disclosure can detect unknown transposons. Accession numbers DRR001193 and DRR001194 are the genome sequences of Arabidopsis ddm1 mutants. ddm1 mutants are known to have reduced DNA methylation, which activates transposons and allows them to transpose.
[0110] In this example, the sequences of DRR001193 and DRR001194 were analyzed, and the genome sequence data was processed as described in (General Information Processing). The results are shown below. Analysis of the ddm1 mutant detected a potential transposon that has not yet been reported. [ka] [ka] (The 20 base sequences of the transposon end candidates correspond to SEQ ID NOS: 5 to 28 in order of appearance.)
[0111] Many of the sequences described above are likely to be terminal sequences and TSDs of unknown retrotransposons. This indicates that the method of the present disclosure is effective for detecting transposition of more than just the known transposon, Tos17. This suggests that the method of the present disclosure can be used to detect transposition of new transposons that have not yet been analyzed.
[0112] (Example 3: Modification method) The procedure was the same as that described in (General Information Processing), except for the part where the complementary strand sequence is obtained. Specifically, step 5 was modified as follows.
[0113] 5. For all target or control-specific sequences, split into pairs of the first 20 bases and the remaining 5 bases. The first 20 bases are candidates for the 3' end of the transposon, and the remaining 5 bases are candidates for the TSD. Also, split the same 25-base sequence into pairs of the first 5 bases and the remaining 20 bases. In this case, the first 5 bases are candidates for the TSD, and the remaining 20 bases are candidates for the 5' end of the transposon.
[0114] (result) The same results as in Example 1 were obtained. Specifically, the results are as follows: [ka] (The 20 base sequences of the transposon end candidates correspond to SEQ ID NOS: 1 to 4 in order of appearance.)
[0115] It was shown that similar results could be obtained even without analysis of the complementary strand.
[0116] (Example 4: Application to cancer diagnosis) DNA is extracted from normal and cancer tissues of the same cancer patient, and the base sequence is obtained, and transposons that are metastasized in the cancer tissue are detected using this method.
[0117] Transposons are detected in multiple cancer patients, and transposons that are specifically activated by cancer transformation are identified. If transposons that are activated by cancer transformation are identified, they can be used to easily determine cancer.
[0118] It is thought that cancer can be diagnosed by extracting DNA or RNA from tissues of patients suspected of having cancer and detecting an increase in transposons.
[0119] We will identify transposable elements that can cause random gene mutations in vivo, create an insertional mutation screening system, and induce tumor formation. By determining the nucleotide sequences of the genomes obtained from these tumors and identifying the sites of insertional mutations, we will identify candidate genes involved in cancer formation and / or cancer malignancy. By functionally evaluating the oncogenic potential of these candidate genes, we will identify novel oncogenes and / or colon cancer suppressor genes.
[0120] (Example 5: Use in creating mutant strains) Transposable elements are used to generate mutant line collections. For example, regarding the generation of Tos17 insertion mutant lines, see Miyao A, Tanaka K, Murata K, Sawaki H, Takeda S, Abe K, Shinozuka Y, Onosato K, Hirochika H. (2003) Target site specificity of the Tos17 retrotransposon shows a preference for insertion within genes and against insertion in retrotransposon-rich regions of the genome. Plant Cell 15(8):1771-80. Miyao A, Iwasaki Y, Kitano H, Itoh JI, Maekawa M, Murata K, Yatou O, Nagato Y, Hirochika H. (2007) A large-scale collection of phenotypic data describing an insertional mutant population to facilitate functional analysis of rice genes. Plant Mol Biol. 63(5):625-635, 2007. and Miyao A, Nakagome M, Ohnuma T, Yamagata H, Kanamori H, Katayose Y, Takahashi A, Matsumoto T, Hirochika H. (2012) Molecular spectrum of somaclonal variation in regenerated rice revealed by whole-genome sequencing. Plant Cell Physiol. 53(1):256-64. doi: 10.1093 / pcp / pcr172. Similar mutant line collections can be created using the identified transposable elements. These mutant line collections can be used for gene function analysis and for obtaining desired traits.For example, in the case of rice, it was discovered that suppressing a specific AGPase gene using a mutant line collection led to the accumulation of high levels of soluble sugar in the stem, making it possible to produce plants with high soluble sugar content ("Plant mutants, method for producing plant mutants, and method for accumulating soluble sugar", Patent No. 5623324); using a mutant line collection, a mutant in which the rice starch synthase type I (SSI) gene was knocked out was produced, which allowed the function of SSI to be elucidated and a novel starch to be produced. ("Functional Elucidation of Starch Synthase Type I and Method for Producing Novel Starch," Patent No. 4703919); and, using a mutant line collection, starch synthase type IIIa (SSIIIa) mutants were isolated from rice populations, resulting in mutants in which SSIIIa activity was completely lost compared to the wild-type, and starch was obtained that exhibited a different amylopectin chain length distribution from the wild-type and different physical properties, such as gelatinization characteristics, from existing wild-type starches ("Functional Elucidation of Starch Synthase Type IIIa and Method for Producing Novel Starch," Patent No. 4711762). Similarly, mutant line collections produced using transposable elements identified by the methods disclosed herein can be used for gene function analysis, acquisition of desired traits, and the production of novel products therefrom.
[0121] Example 6: Detection of transposable elements in Arabidopsis hypomethylated mutants Transposable elements were detected in Arabidopsis hypomethylated mutants (DRR00193, DRR00194). Some of the results are shown in the table below. [Table 1] (The transposon sequences and adjacent sequences correspond to SEQ ID NOs: 41 to 88 in order of appearance, from left to right and top to bottom in each line.) The position where the insertion occurred at a unique position is marked as unique. Insertions into repeat regions have been difficult to detect until now, but the method of the present disclosure can output possible positions from among multiple candidate positions.
[0122] Example 7: Example of detection of transposition of retrotransposon Tos17 The transposition of the retrotransposon Tos17 was detected by the method of the present invention, and some of the results are shown in the table below. [Table 2] (The head and tail sequences and adjacent sequences correspond to SEQ ID NOs: 89 to 128 in order of appearance, from left to right and top to bottom in each line.) The position where the insertion occurred at a unique position is marked as unique.
[0123] Example 8: Detection of metastasis at any TSD size To confirm whether transposition of any TSD size can be detected, we analyzed the Drosophila P elements SRR823377 and SRR823382. Some of the results are shown in the table below. . [Table 3] (The head and tail sequences and adjacent sequences correspond to SEQ ID NOs: 129 to 168 in order of appearance, from left to right and top to bottom in each line.) The position where the insertion occurred at a unique position is marked as unique.
[0124] (Example 9: Example of detection of transposition sequences using algorithm 2) The Tos17 insertion in ttm2 was detected by the method of algorithm 2. Some of the results are shown in the table below. [Table 4] (The terminal sequences and insertion site sequences correspond to SEQ ID NOs: 169 to 188 in order of appearance, from left to right and top to bottom in each line.) The original Tos17 exists in one copy each on chromosome 7 and 10. The 20-base sequence at the end of the transposable element is present at both ends of the long terminal repeat (LTR) at both ends of the transposable element, so two copies were detected on each side. The new algorithm was also able to detect the location on the genome of the original transposable element before transposition.
[0125] (Note) While the present disclosure has been illustrated using preferred embodiments thereof, it is understood that the scope of the present disclosure should be construed solely in accordance with the claims. It is understood that the patents, patent applications, and other documents cited herein are incorporated by reference in their entirety as if the contents themselves were specifically set forth herein. This application claims priority to Japanese Patent Application No. 2019-236480, filed on December 26, 2019, with the Japan Patent Office, the contents of which are incorporated herein by reference in their entirety. [Industrial Applicability]
[0126] Although much remains unknown about the dynamics of transposons within individuals, the method of the present disclosure can detect transposons that are actually transposing, and has a wide range of applications. For example, in plants, the method can be used to screen for transposons that can be used for transposon tagging. In humans, the method is likely to be useful in studying the relationship between cancer and transposons, and in studying the activation or deactivation of gene transcription due to transposon transposition. [Sequence List Free Text]
[0127] SEQ ID NOs: 1 to 4: 20-base sequences of transposon end candidates described in Examples 1 and 4 (rice) SEQ ID NOs: 5 to 28: 20 base sequences of transposon end candidates described in Example 2 (Arabidopsis) SEQ ID NOs: 29 to 40: 20 base sequences of the excised sequences shown in Figure 4 (rice) SEQ ID NOs: 41 to 88: 20 base sequences of the transposon sequence and adjacent sequences described in Example 6 (Arabidopsis) SEQ ID NOs: 89 to 128: 20 base sequences of head and tail sequences and adjacent sequences described in Example 7 (rice) SEQ ID NOs: 129 to 168: 20 base sequences of the head and tail sequences and adjacent sequences described in Example 8 (Drosophila) SEQ ID NOs: 169 to 188: Terminal sequences and insertion site sequences described in Example 9 (rice)
Claims
1. 1. A method for identifying a transposable element in a sequence, comprising: (A) extracting X base sequences from short reads derived from two types of samples by shifting the sequences by one base; (B) obtaining the frequency of each of the obtained sequences of the plurality of X bases; (C) obtaining the positions of the first and second halves of each of the obtained sequences of a plurality of X bases on a reference sequence and comparing the positions; (D) identifying transposable elements in the sequence based on the frequency and location information; A method comprising:
2. The method according to claim 1, wherein the base lengths of the first and second parts are X / 2 bases.
3. The method of claim 1 or 2, wherein X is 20 or more bases.
4. 1. A method for detecting target site duplications (TSDs) and / or junctions of transposable elements in a sequence, comprising: (A) extracting X base sequences from short reads derived from two types of samples by shifting the sequences by one base; (B) obtaining the frequency of each of the obtained sequences of the plurality of X bases; (C) obtaining the positions of the first and second halves of each of the obtained sequences of a plurality of X bases on a reference sequence and comparing the positions; (D) detecting target site duplications (TSDs) and / or junctions of transposable elements in the sequence based on the frequency and location information; A method comprising:
5. The method according to claim 4, wherein the base lengths of the first and second parts are X / 2 bases.
6. The method according to claim 4 or 5, wherein X is 20 or more bases.
7. (A) extracting X base sequences from short reads derived from two types of samples by shifting the sequences by one base; (B) obtaining the frequency of each of the obtained sequences of the plurality of X bases; (B1) sorting and outputting each sequence of a plurality of X bases; (B2) extracting a sequence of X bases specific to either of the two types of samples; (C1) obtaining a position on a reference sequence corresponding to the first X / 2 base sequence of each extracted X base sequence; (C2) obtaining a position on a reference sequence corresponding to the sequence of the latter X / 2 bases of each of the extracted X base sequences; (C3) sorting the location data obtained in C1 and C2 by chromosome; (D) a step of outputting, from the position data sorted in C3, TSDs each having a shifted junction position and a length of TSD between s1 and s2; (E) selecting a TSD adjacent to the 3'-end of the first X / 2 bases and a TSD adjacent to the 5'-end of the second X / 2 bases, when there are two or more types of TSDs, and selecting from the selected dataset all pairs of X base sequences in which the TSD adjacent to the 5'-end and the TSD adjacent to the 3'-end are identical; (F) selecting pairs of the X / 2 base sequences that are present in both of the two types of samples and have two or more different TSD sequences; (G) outputting the pairs selected in F from the position data obtained in C1 and C2; 5. The method of claim 1 or 4, comprising:
8. 8. The method of claim 7, wherein X is 20 or greater.
9. 9. The method according to claim 7 or 8, wherein s1 is 3 or more.
10. The method according to any one of claims 7 to 9, wherein s2 is 20 or less.
Citation Information
Patent Citations
Method used for identification and quantitative low frequency somatic cell mutation detection
CN108949911A
Utilization of rice transposon
JP2006238804A
Transposon gene of rice plant
JP2008187967A
Parallel functional testing of synthetic DNA segments, pathways, and genomes
JP2019501642A
Method for identifying transposons from a nucleic acid database
US20030152955A1