Identification method of sequence participating in potato centromere relocation

By using combined bioinformatics analysis methods, the centromere relocation sequence of potato was identified, which solved the problem of lack of CENH3 deposition signal, improved the stability of artificial chromosomes and the efficiency of genetic transmission, and has the potential for cross-species application.

CN121709039APending Publication Date: 2026-03-20TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511959034.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

The molecular basis of potato centromere relocation is not yet clear in the existing technology, especially the lack of direct evidence for the initial induction signal and sequence basis of CENH3 deposition, which limits the development of artificial chromosomes and precision genome engineering.

Method used

Using a combined bioinformatics analysis method, satellite DNA sequences with potential CENH3 enrichment capacity were screened and identified based on T2T-level genomic data of different potato haplotypes and CENH3 ChIP-seq data. The specific steps included predicting transposon families with EDTA software, accurately identifying repetitive sequences with RepeatMasker, and performing secondary annotation with DeepTE software, combined with Bowtie 2 alignment technology to confirm sequence function.

Benefits of technology

The identified satellite DNA sequence shows high conservation and colocalization with the CENH3 binding signal, possessing the potential to guide CENH3 deposition and initiate new centromere formation, thus improving the stability and genetic transmission efficiency of artificial chromosomes, and demonstrating cross-species compatibility and application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121709039A_ABST
    Figure CN121709039A_ABST
Patent Text Reader

Abstract

The invention discloses an identification method of a sequence participating in potato centromere relocation, and belongs to the technical field of plant heredity and molecular biology. According to the method, T2T-level genome data and CENH3 ChIP-seq data of different haplotypes of potatoes are integrated, sequence screening, annotation and function prediction are carried out by utilizing a bioinformatics analysis tool, and a satellite DNA sequence which is 2-3Kbp in length, is highly conserved in the same haplotype and can be specifically combined with CENH3 protein is identified. The sequence can induce CENH3 to deposit in a non-native centromere area, so that functional relocation of centromere is realized, and a core molecular element and a technical support are provided for artificial centromere construction, plant chromosome engineering and accurate genome operation. The method solves the problems that in the prior art, a functional sequence for clearly inducing centromere relocation is lacked, a CENH3 targeted recruitment mechanism model is not established and the like, and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of plant genetics and molecular biology, and in particular to a method for identifying sequences involved in centromere relocation in potatoes. Background Technology

[0002] Centromeres are key regions on eukaryotic chromosomes responsible for spindle attachment, accurate chromosome segregation, and maintaining genome stability. Their function is primarily mediated by the specific histone variant CENH3. The deposition of CENH3 at specific sites on chromosomes determines the position and activity of the centromere, serving as a core marker of chromosomal genetic transmission. Studies have shown that centromeres can undergo relocation during evolution, meaning that functional centromeres migrate to new genomic locations. However, the molecular basis of centromere relocation remains unclear, particularly regarding the initial induction signal and sequence basis of CENH3 deposition, which still lack direct evidence.

[0003] Centromere regions are typically enriched with specific types of satellite DNA and long terminal repeat transposons, which may provide the epigenetic background or sequence template for CENH3 loading. However, to date, no specific sequence elements have been identified that can induce CENH3 enrichment or initiate new centromere formation. In potatoes and other plants, although multi-haplotype T2T level reference genome assembly has been completed, revealing significant variations in centromere regions between different haplotypes, related studies have mainly focused on structural comparisons and sequence composition analyses, lacking validation and application studies targeting functional sequences. Summary of the Invention

[0004] The purpose of this invention is to provide a method for identifying potato centromere relocation sequences, so as to clarify the functional DNA sequences that induce CENH3 deposition and trigger centromere relocation, and provide basic elements for artificial chromosomes and precision genome engineering.

[0005] To achieve the above objectives, the present invention provides a method for identifying potato centromere relocation sequences. The method is based on a combined bioinformatics analysis of T2T-level genomic data and CENH3 ChIP-seq data of different potato haplotypes to screen and identify a satellite DNA sequence with potential CENH3 enrichment capacity; the satellite DNA sequence is shown in SEQ ID NO.1.

[0006] Preferably, the combined bioinformatics analysis method specifically includes the following steps: S1. Collect T2T genome data and CENH3 ChIP-seq data of different haplotypes of potato; S2. Using EDTA software, transposon families were predicted and annotated across the entire genome of the potato genome data from S1. Known centromere-specific sequences of potatoes were integrated into the EDTA TE reference library. RepeatMasker was used to accurately identify centromere-specific repetitive sequences. DeepTE software was used to perform secondary annotation on the initially identified UnknownLTR sequences to clarify their subfamily affiliation. S3. By evaluating the sequence characteristics and conservation of the Unknown LTR sequence obtained in S2, determine whether the Unknown LTR sequence is a sequence involved in potato centromere relocation.

[0007] Preferably, in S1, the T2T genome data includes the T2T genome data of potato haplotype DM and the T2T genome data of two haplotypes CM_Hap1 and CM_Hap2 of potato line CM, and the CENH3 ChIP-seq data are sourced from the NCBI database and the NGDC database.

[0008] Preferably, S3 specifically includes: S31. After the CENH3 ChIP-seq reads were quality controlled and filtered by FastP v0.23.4, they were aligned to the potato genome assembly results using Bowtie 2 v2.2.5. The number of aligned reads was counted in 500kb windows, and merged regions with signal values ​​greater than 20 and adjacent intervals less than 3Mb were selected as centromere relocation regions. S32. Extract the enriched region of the Unknown LTR sequence in the same way as in S31, and observe whether the centromere relocation region and the enriched region of the Unknown LTR sequence are highly correlated. If they are correlated, it is proven that the Unknown LTR sequence has the potential function of inducing centromere relocation.

[0009] On the other hand, the present invention provides a DNA construct comprising the potato centromere relocation sequence and operatively linked regulatory elements as described in any of the preceding claims.

[0010] Preferably, the DNA construct is a plant expression vector or an artificial chromosome framework vector.

[0011] On the other hand, the present invention provides the application of the above-mentioned satellite DNA sequence or DNA construct in the preparation of reagents or tools for studying plant centromere function or CENH3 deposition mechanisms.

[0012] On the other hand, the present invention provides an application of the above-mentioned satellite DNA sequence or DNA construct as a core functional module in the construction of plant artificial centromeres or artificial chromosomes.

[0013] Therefore, the method for identifying potato centromere relocation sequences of the present invention has the following beneficial effects: (1) Sequence specificity and conservation: The identified satellite DNA sequence is 2-3Kbp in length, originates from the centromere-rich region of potato, and shows high conservation and centromere-specific distribution characteristics among the same haplotype. It has significant colocalization with the CENH3 binding signal, suggesting that it has strong CENH3 recruitment potential.

[0014] (2) This sequence can be used as a core functional module in artificial centromere construction systems, embedded in artificial vectors or synthetic chromosome frameworks, to achieve targeted recruitment of CENH3 and preliminary establishment of centromere function, thereby improving the stability and genetic transmission efficiency of artificial chromosomes. Its sequence characteristics have the potential for designability and cross-species compatibility, and may be extended to a variety of plant systems, demonstrating good versatility and application prospects.

[0015] (3) This invention identifies a class of DNA sequences in plants that have the potential to induce CENH3 deposition through bioinformatics, providing new clues for elucidating the centromere relocation mechanism and providing a potential molecular tool for artificial centromere construction and chromosome engineering.

[0016] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0017] Figure 1 This section shows the distribution of the St49 sequence on chromosome 1 in different haplotypes. Part A shows the results of mapping ChIP-seq data of the potato DM line to the CM_Hap1 genome; Part B shows the results of mapping ChIP-seq data of the potato DM line to the CM_Hap2 genome; and Part C shows the enrichment of the St49 sequence in the centromere region of chromosome 1 of the potato DM line. Figure 2 The distribution of the St49 sequence on chromosome 5 in different haplotypes is shown in Part A, which shows the results of mapping ChIP-seq data of potato DM lines to the CM_Hap1 genome; Part B shows the results of mapping ChIP-seq data of potato DM lines to the CM_Hap2 genome; and Part C shows the enrichment of the St49 sequence in the centromere region of potato chromosome 5 in DM lines. Detailed Implementation

[0018] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0019] To make the objectives, technical solutions, and advantages of this application clearer, more thorough, and more complete, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings and embodiments. The following detailed descriptions are all illustrations of embodiments, intended to provide further detailed explanation of the present invention. Unless otherwise specified, all technical terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0020] In this embodiment, the T2T genome data of potato haplotype DM was obtained from publicly available data provided by Bioinformatics Lab (http: / / www.bioinformaticslab.cn / pubs / dm8), and the CenH3 ChIP-seq data was downloaded from the NCBI database (PRJNA820869). The T2T genome data and CenH3 ChIP-seq data of the two haplotypes CM, CM_Hap1 and CM_Hap2, of the potato line CM were obtained from the NGDC database (PRJCA023865).

[0021] Example A method for identifying sequences involved in potato centromere relocation includes the following steps: S1. Data Acquisition and Preprocessing: Download T2T genome data of potato haplotypes DM, CM_Hap1, and CM_Hap2, along with the corresponding CENH3 ChIP-seq data, from public databases. Use FastP v0.23.4 for quality control of the ChIP-seq data, with the following parameters: remove adapter sequences, filter bases with a Q value <20, and retain sequences ≥36 bp in length.

[0022] S2. Sequence Annotation and Screening: EDTA v2.1.0 was used to identify and annotate genome-wide repetitive sequences, with initial identification performed using default parameters. To improve the accuracy of centromere region repetitive sequence identification, known potato centromere-specific sequences (such as the St and Sv sequence families) were integrated into the EDTA TE reference library (lib file). RepeatMasker was then re-run using the integrated lib file to accurately identify centromere region-specific repetitive sequences.

[0023] Finally, a satellite DNA sequence St49 with potential CENH3 enrichment capacity was selected, as shown in SEQ ID NO.1.

[0024] S3. Sequence Identification: Sequencing reads obtained from the ChIP experiment of the DM potato CENH3 strain were quality controlled and filtered using FastP (v0.23.4) to remove low-quality sequences and ensure data reliability. Bowtie 2 (v2.2.5) was used to align the filtered reads to the genome assembly results of the CM potato strain, and the number of aligned reads was counted in 500kb windows. Regions with signal values ​​greater than 20 were screened, and if the distance between two adjacent regions was less than 3Mb, they were merged into one region, which became the candidate region for centromere relocation.

[0025] Meanwhile, the enriched regions of the identified Unknown LTR sequences were extracted in the same manner as above, and the location of the centromere relocation region was observed to be highly correlated with that of the Unknown LTR sequence enriched region. If they were correlated, it was proven that the sequence had the potential function of inducing centromere relocation.

[0026] The distribution of the St49 sequence on chromosome 1 in different haplotypes is as follows: Figure 1 As shown, Part A shows the results of mapping ChIP-seq data of the potato DM line onto the CM_Hap1 genome; Part B shows the results of mapping ChIP-seq data of the potato DM line onto the CM_Hap2 genome; and Part C shows the enrichment of the St49 sequence in the centromere region of chromosome 1 of the potato DM line.

[0027] Depend on Figure 1 It can be seen that the centromere region of chromosome 1 of the DM potato variety is enriched with the St49 sequence, and when mapped to Chr01 of CM (Hap1 & Hap2), the St49 sequence is enriched again at the CENH3 enrichment site, indicating a high correlation between the two.

[0028] The distribution of the St49 sequence on chromosome 5 in different haplotypes is as follows: Figure 2 As shown, Part A shows the results of mapping ChIP-seq data of the potato DM line onto the CM_Hap1 genome; Part B shows the results of mapping ChIP-seq data of the potato DM line onto the CM_Hap2 genome; and Part C shows the enrichment of the St49 sequence in the centromere region of chromosome 5 of the potato DM line.

[0029] Depend on Figure 2 It can be seen that the centromere region of chromosome 5 of DM potato is enriched with the St49 sequence, and when mapped to Chr05 of CM (Hap1 & Hap2), the St49 sequence is enriched again at the CENH3 enrichment site, indicating a high correlation between the two.

[0030] Based on the sequence characteristics and conservation assessment of the satellite DNA sequence, it is known that when this sequence is artificially introduced into non-centromere chromosome regions (such as the arm or proximal region), it has the potential to guide CENH3 to be deposited locally at that site and initiate the formation of a new centromere, thereby achieving functional repositioning of the centromere.

[0031] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for identifying sequences involved in potato centromere relocation, characterized in that: The method described is a bioinformatics analysis method based on the combination of T2T level genomic data of different haplotypes of potato and CENH3 ChIP-seq data, which screens and identifies a satellite DNA sequence with potential CENH3 enrichment capacity. The satellite DNA sequence is shown in SEQ ID NO.

1.

2. The method for identifying potato centromere repositioning sequences according to claim 1, characterized in that, The combined bioinformatics analysis method specifically includes the following steps: S1. Collect T2T genome data and CENH3 ChIP-seq data of different haplotypes of potato; S2. Using EDTA software, transposon families were predicted and annotated on the potato genome data from S1 across the entire genome. Known potato centromere-specific sequences were integrated into the EDTA TE reference library. RepeatMasker was used to accurately identify centromere region-specific repetitive sequences. DeepTE software was used to perform secondary annotation on the initially identified Unknown LTR sequences to clarify their subfamily affiliation. S3. By evaluating the sequence characteristics and conservation of the Unknown LTR sequence obtained in S2, determine whether the Unknown LTR sequence is a sequence involved in potato centromere relocation.

3. The method for identifying potato centromere repositioning sequences according to claim 2, characterized in that: In S1, the T2T genome data includes the T2T genome data of potato haplotype DM and the T2T genome data of two haplotypes CM_Hap1 and CM_Hap2 of potato line CM. The CENH3 ChIP-seq data are from the NCBI database and the NGDC database.

4. The method for identifying potato centromere repositioning sequences according to claim 2, characterized in that, S3 specifically includes: S31. After the CENH3 ChIP-seq reads were quality controlled and filtered by FastP v0.23.4, they were aligned to the potato genome assembly results using Bowtie 2v2.2.

5. The number of aligned reads was counted in 500kb windows, and merged regions with signal values ​​greater than 20 and adjacent intervals less than 3Mb were selected as centromere relocation regions. S32. Extract the enriched region of the Unknown LTR sequence in the same way as in S31, and observe whether the centromere relocation region and the enriched region of the Unknown LTR sequence are highly correlated. If they are correlated, it is proven that the Unknown LTR sequence has the potential function of inducing centromere relocation.

5. A DNA construct, characterized in that, It includes a regulatory element that participates in the potato centromere repositioning sequence and is operatively connected as described in any one of claims 1-4.

6. A DNA construct according to claim 5, characterized in that: The DNA construct is a plant expression vector or an artificial chromosome framework vector.

7. The use of the satellite DNA sequence as described in claim 1 or the DNA construct as described in claim 5 in the preparation of reagents or tools for studying plant centromere function or CENH3 deposition mechanisms.

8. The application of the satellite DNA sequence as described in claim 1 or the DNA construct as described in claim 5 as a core functional module in the construction of plant artificial centromeres or artificial chromosomes.