Methods for precise identification of chromosomal break and / or fusion locations and uses thereof

By combining ONT ultra-long and PacBio HiFi sequencing with HiC technology, the problems of cumbersome, unsafe, and inaccurate identification of chromosome breakage and fusion sites in existing technologies have been solved. This has enabled simple, safe, and high-precision identification of chromosome breakage and fusion sites, providing complete genetic information and supporting species evolution and breeding research.

CN116479101BActive Publication Date: 2026-07-21SOUTHWEST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTHWEST UNIV
Filing Date
2023-03-10
Publication Date
2026-07-21

Smart Images

  • Figure HDA0004119798180000011
    Figure HDA0004119798180000011
  • Figure HDA0004119798180000012
    Figure HDA0004119798180000012
  • Figure HDA0004119798180000013
    Figure HDA0004119798180000013
Patent Text Reader

Abstract

The present application relates to a method for accurately identifying the position of chromosome breakage and fusion, which comprises the following steps: performing karyotype analysis on a sample of a plant to be tested to determine the presence of chromosome breakage and fusion, taking a tender tissue sample of the plant to be tested to determine its genome size, combining ONT ultra-long sequencing and PacBio HiFi sequencing, and HIC assisted assembly technology to perform complete genome assembly and chromosome mounting, obtaining the complete genome of the plant to be tested, and accurately determining the position of chromosome breakage and fusion. The method can be used to accurately determine the position of chromosome breakage and fusion in plants, and is simple to operate and reliable in results. Moreover, the results obtained by the method have very high added value, providing convenience for analyzing the breakage / fusion mechanism and studying species evolution, and obtaining the complete blueprint of all genetic information of the species, which can accelerate the molecular genetic breeding research of the species.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to biotechnology, and more specifically to a method for accurately identifying the location of chromosome breaks and / or fusions and its applications. Background Technology

[0002] In biological research and molecular genetic breeding of a species, the first step is often a cellular study. Cellular research is both the starting point and the convergence point of life sciences. A crucial and fundamental aspect of cellular research is determining the chromosome number of the species and understanding the dynamic changes in chromosomes during meiosis and mitosis. This is essential for solving practical problems in breeding and biological research within the species. During cellular studies, such as karyotype analysis, fluctuations in the chromosome number of a species are often observed. This is due to chromosome breakage and fusion during cell division. Accurate identification of chromosome breakage and fusion sites helps elucidate the mechanisms of these mechanisms, which is significant for studying species evolution and conducting genetic breeding. For a long time, karyotype analysis has been used to locate chromosome breakage and fusion sites. However, this method is subject to numerous uncertainties and randomness. Specifically:

[0003] 1. Obtaining satisfactory karyotype results requires a significant amount of work. Karyotype analysis of plant cells typically requires taking specific tissue samples (root tips, young leaves, epidermis) from vigorous growth material at a specific stage, followed by a specific processing procedure. This procedure generally includes pretreatment, fixation, hypotonicity, enzymatic digestion, slide preparation, staining, and microscopic examination. This process exhibits several characteristics: ① Cumbersome: Beginners need repeated practice to master the technique; ② Time-consuming: From sample collection to final microscopic examination, the entire process takes nearly a day; ③ Unsafe: Some biochemical reagents used in this process are harmful to humans. For example, fixatives are often highly irritating, and liquids used in pretreatment, such as colchicine and 8-hydroxyquinoline, are harmful to human cells; ④ Low operability: For different materials, even experienced operators need to experiment with various combinations of conditions, such as sample collection time and processing temperature and time, and undergo numerous trials before they can summarize the specific procedure for that material. This process is also highly dependent on proficiency; if someone else were to perform the procedure, the desired results might not be achieved.

[0004] 2. Difficulty in pinpointing the actual location of chromosome breakage and fusion. Chromosomes are a specific form of genetic material present in cells during mitosis (or meiosis), resulting from the tight assembly of chromosomes during interphase. In the field of view for karyotype analysis, the genetic material of chromosomes where breakage and fusion occur is highly condensed. However, the fluorescence in situ hybridization (FISH) method commonly used for chromosome karyotype analysis struggles to find specific targets to pinpoint potential breakage and fusion sites. This is because, under ideal microscopic conditions, only major and minor constrictions of the chromosome might be visible, which are insufficient to pinpoint the exact location of a breakage and fusion. For example, if a species has a genome of 500 Mb containing 10 chromosomes, the average chromosome size is 50 Mb. A probe of 1 kb would require approximately 50,000 probes to cover the entire chromosome, assuming each probe is successfully used and the 1 kb resolution allows detection of the probe signal during subsequent microscopic examination. However, in reality, such an ideal scenario is practically impossible.

[0005] 3. Difficulty in elucidating the underlying mechanisms of break-fusion phenomena. Even with fluorescence in situ hybridization (FISH) and the use of probes to further narrow down the regions where break-fusion may occur, current methods lack genomic information for the species, making it impossible to obtain sequence information of the break-fusion regions or identify the components within them. This makes it impossible to elucidate the underlying mechanisms of break-fusion phenomena. Summary of the Invention

[0006] Based on this, the purpose of this invention is to provide a method for accurately identifying chromosome breakage and / or fusion sites. This method can be used to accurately locate the chromosome breakage and fusion sites in the plant under test, and it is simple, easy to implement, and reliable.

[0007] The technical solutions for achieving the above objectives include the following.

[0008] A first aspect of the present invention is to provide a method for accurately identifying the location of chromosome breaks and / or fusions, comprising the following steps:

[0009] s1 performed karyotype analysis on the plant samples to determine the presence of chromosome breakage and fusion;

[0010] s2. Take young tissue samples from the plant to be tested, perform a genome survey to determine its genome size, including the following steps:

[0011] s2.1 Extract DNA from the plant sample to be tested;

[0012] s2.2 Constructing a DNA library;

[0013] s2.3 The qualified libraries were sequenced using DNASEQ-T7, and the raw image data files obtained from sequencing were converted into raw data using base calling (base recognition technology);

[0014] s2.4 filters the raw data to obtain clean reads (data to be analyzed);

[0015] The clean reads obtained from s2.5 are extracted with a length of more than 50,000-60,000 reads and compared with public databases to determine that the sample has not been contaminated by external sources.

[0016] s2.6 Perform Kmer frequency depth distribution analysis to obtain genome size, repetition and heterozygosity information; s3 HIC library construction and sequencing, including: s3.1 Take young tissue samples of the plant to be tested, fix and cross-link them to maintain the 3D structure in the cells;

[0017] s3.2 endonuclease digestion, end repair, end repair, cyclization;

[0018] s3.3 DNA purification and capture, library construction, sequencing to obtain sequencing data, and filtering to obtain effective HIC sequencing data;

[0019] S4, combined with ONT ultra-long sequencing and PacBio HiFi sequencing, and HIC-assisted assembly technology, was used to assemble the complete genome and mount chromosomes from the sequencing data obtained in S3 and S4, yielding the complete genome of the plant under test.

[0020] s4.1 uses young mulberry leaf tissue samples to set the throughput required for ONT ultra-long and PacBio HiFi sequencing. The ONT ultra-long library preparation standard requires the selection of fragments larger than 100k; the PacBio HiFi library preparation requires fragments larger than 20k to obtain raw reads.

[0021] s4.2 filters the raw read length data, retaining read length data with an average quality score greater than 90% for subsequent assembly;

[0022] s4.3 Perform data assembly and chromosome mounting to obtain the complete genome of the plant under test (telomere-telomere genome, abbreviated as T2T);

[0023] s5. Juicebox can be used to obtain the interaction matrix between complete chromosomes in the genome of the plant sample under test at the T2T level, and accurately determine the location of chromosome breakage and / or fusion.

[0024] In some embodiments, the tender tissue sample is a newly sprouted seedling or a newly grown leaf.

[0025] In some of these embodiments, s2.3, the depth of the sequencing genome survey analysis is greater than 100x.

[0026] In some embodiments, s2.4, the filtering of the raw data includes at least the following: a) removing reads with an N base content greater than 5%; b) removing reads with a quality value less than or equal to 5 in more than 50% of the bases; c) removing reads with adapter contamination; and d) removing repetitive sequences caused by PCR amplification.

[0027] In some embodiments, s2.5, the read length data is 50000-60000.

[0028] In some embodiments, the library preparation data obtained in s3 is greater than 100x the genome, and the proportion of bases with a quality of Q20 and Q30 or higher is 95% and 90% or higher, respectively. Only after rigorous subsequent filtering according to these standards can sufficient effective HIC data be generated. The final effective HIC data volume needs to reach at least 60% to include all interaction signals; otherwise, many interaction signals will be lost.

[0029] In some embodiments, s4.1, based on the results of a genome survey, when the plant is a mulberry tree, the genome is assessed as a low heterozygosity genome, and it is set that ONT ultra-long must generate more than 100x of data and PacBio HiFi must generate more than 40x of data.

[0030] In some embodiments, s4.2, the raw read length data filtering includes: for the raw read length data of ONT ultra-long, using Filtlong to filter data segments <10kb, removing fail reads with an average quality score less than 7, and obtaining valid data for subsequent analysis; using Porechop default parameters to filter connector sequences, and continuing to use Filtlong to filter sequences <30kb, retaining read length data with mean read quality scores >90%.

[0031] In some of these embodiments, the PacBio HiFi raw data is filtered to remove low-quality subreads with sequencing cycles of less than 3 times and an SNR of less than 2.5.

[0032] In some embodiments, the data assembly and chromosome mounting described in s4.3 to obtain the complete genome of the plant to be tested includes: after obtaining the 0gap genome of the plant species to be tested, comparing the HiFi reads of not less than 10kb generated by PacBio HiFi sequencing with the 0gap genome, filtering the alignment fragments, deleting chimeric alignment fragments, and performing error correction.

[0033] In some embodiments, the karyotype analysis in s1 includes: slicing method, compression method, smear method, and wall-removing hypotonic method.

[0034] In some preferred embodiments, when the plant is a mulberry tree, the cell wall removal hypotonic method is used because, compared with the cell wall removal hypotonic method, when the sectioning method, pressing method, and smear method are used on mulberry trees, it is found that the chromosomes are difficult to disperse and are not easy to observe.

[0035] The cell wall removal hypotonic method includes enzymatic cell wall removal hypotonic method and acid hydrolysis cell wall removal hypotonic method, with enzymatic cell wall removal hypotonic method being more preferred, as it has the best effect on karyotype analysis of mulberry trees.

[0036] In some preferred embodiments, the plant is a mulberry tree.

[0037] A second aspect of the present invention is to provide the application of the above-described method in the accurate identification of chromosome breakage and / or fusion sites in plant species.

[0038] This invention develops a novel technical approach for accurately identifying chromosome breakage and / or fusion sites in plant species. It cleverly combines information on chromosome breakage and fusion phenomena and genome size in a single determination of the plant species, and then uses ONT ultra-long sequencing, PacBio HiFi sequencing, and HiC-assisted genome assembly technology to achieve telomere-to-telomere (T2T) chromosome assembly, as well as optimization of data depth and length, and optimization of strict data filtering, thereby obtaining a method for accurately determining chromosome breakage and / or fusion sites in plants.

[0039] The method described in this invention can be used to accurately determine the locations of chromosome breaks and / or fusions in plants, and it is simple to operate and yields reliable results. Furthermore, this method not only facilitates the analysis of break and / or fusion mechanisms but also provides a complete blueprint of all genetic information for the species, accelerating molecular genetic breeding research on that species.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] 1. Minimal technical workload: Based on the technical solution, only one successful routine karyotype analysis is required during the precise identification of the entire break-fusion site. Statistical analysis confirms that chromosome break-fusion occurs during cell division in this species. Simultaneously, based on information such as chromosome morphology and relative length from the karyotype analysis, the number of one or more chromosomes that have broken-fused is obtained; subsequent karyotype analysis is not required.

[0042] 2. High technical repeatability: The method described in this invention is simple, straightforward, and easy to master. The technical description provides detailed breakdowns of all steps and parameters, allowing beginners to obtain results simply by following the procedure. Unlike traditional FISH techniques, which involve numerous intermediate steps that can affect the results and require repeated practice to master, this method offers superior performance.

[0043] 3. Technical safety: The method described in this invention does not require repeated use of highly irritating or even carcinogenic biochemical reagents such as fixatives and pretreatment solutions;

[0044] 4. High technical portability: The method described in this invention is based on whole-genome sequencing and intrachromosomal interactions to solve the corresponding problems. Theoretically speaking, any living object can be studied by referring to the method of this invention to quickly locate the corresponding break and fusion sites on chromosome segments.

[0045] 5. High technical precision: This technology, while identifying chromosome break and fusion sites, obtains complete information about the corresponding species' genome, achieving precision down to the individual base level to pinpoint these sites. In species without genomic guidance, traditional FISH struggles to design effective probes. Even with probes and successful multiple trials, break and fusion sites are only visible at the chromosome scale, a resolution of tens of megabp, which is incomparable to bp-level resolution.

[0046] 6. High technological value: The value of the method described in this invention is not limited to obtaining break-fusion sites. At the same time, it obtains information about the entire genome of the species, including all sequence information within and near the break-fusion site. With clear sequence information, it can help identify elements adjacent to the site, truly facilitating the analysis of the mechanism and providing convenience for species evolution research.

[0047] 7. The technology is scalable: The complete genome obtained based on the method described in this invention (also known as telomere-to-telomere, or T2T for short) gives the technology extremely high added value. While providing convenience for elucidating the break-fusion mechanism, it also depicts a complete blueprint of all the genetic information of the species, which can accelerate the molecular genetic breeding research of the species. Attached Figure Description

[0048] Figure 1 Two types of cells with different numbers of chromosomes in mulberry trees.

[0049] Figure 2 Analysis of weak interaction signal regions between chromosomes; where A shows weak interaction signal regions present on all 6 chromosomes in 5Mb intervals; B shows the coverage of Pacbio HiFi and ONT ultra-long reads of the corresponding weak interaction regions; C shows the sequence composition analysis of the corresponding weak interaction signal regions.

[0050] Figure 3 Chromosomal interaction matrix data for the 30-32Mb region of chromosome chr5.

[0051] Figure 4 The copy number of 25S rDNA in the 30-32Mb region of chromosome chr5. Detailed Implementation

[0052] To facilitate understanding of the present invention, a more complete description will be provided below. The present invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the present invention.

[0053] Unless otherwise specified, experimental methods in the following examples were performed under standard conditions, such as those described in the fourth edition of *Molecular Cloning: A Laboratory Manual*, edited by Green and Sambrook, published in 2013, or under conditions recommended by the manufacturer. All commonly used chemical reagents used in the examples are commercially available products.

[0054] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. The term "and / or" as used in this invention includes any and all combinations of one or more of the associated listed items.

[0055] This invention provides a method for accurately identifying chromosome breakage and / or fusion sites, comprising the following core steps: 1) First, determine if chromosome fusion and breakage have occurred in the plant species: Initially, only one successful, low-threshold karyotype analysis is required; subsequent low-efficiency, uncertain, and accidental chromosome fluorescence in situ hybridization experiments are unnecessary. The initial determination of the dynamic changes in chromosome number during cell division in this species is crucial. The inventors have found that, for a given species, determining whether chromosome breakage and fusion occur is a prerequisite for this method and also provides specific chromosome numbering information for subsequent analysis.

[0056] 2) Extract DNA from young leaves or other tissues from which DNA can be easily extracted. Construct a qualified library according to the requirements of the corresponding protocol and steps, perform genome sequencing and chromosome mounting, and assemble the complete genome sequence of the species (also known as telomere-to-telomere, or T2T) using a combination of ONT ultra-long, PacBioHiFi and hic technologies.

[0057] 3. Analyze the interactions within the chromosomes of this species. By studying the chromosome interaction matrix data across the entire genome, the precise locations of chromosome breakage and fusion can be obtained.

[0058] The present invention will be further described in detail below with reference to specific embodiments.

[0059] Example 1

[0060] This embodiment uses mulberry trees as an example to illustrate in detail the method for accurately identifying chromosome breakage and / or fusion locations. The method includes the following steps:

[0061] 1. Perform a successful karyotype analysis by following the modified enzymatic hydrolysis and hypotonic method described below, including the following steps:

[0062] 1.1 Pretreatment: Young mulberry leaves were collected in the early morning and then placed in a pre-prepared 0.002M 8-hydroxychloroquine solution. The solution was pretreated at room temperature in the dark for 3 hours, during which the leaves could be gently shaken to ensure that the material was always submerged in the liquid.

[0063] 1.2 Fixation: Take out the pretreated material and place it in freshly prepared Carno fixative. Treat at 4°C for more than 2 hours. During this period, the fixative can be replaced with fresh fixative once as needed.

[0064] 1.3 Pre-treatment with low osmotic pressure: After rinsing the fixed material repeatedly with distilled water 3 times, place it in 0.075M KCl solution and treat at room temperature for 30 min;

[0065] 1.4 Enzymatic hydrolysis: After rinsing the material treated in the previous step twice with distilled water, treat it with a freshly prepared 2.5% pectinase and cellulase mixture and hydrolyze it at room temperature for 2.5 hours.

[0066] 1.5 Post-hypotonic treatment: After the previous enzymatic hydrolysis, the material was rinsed repeatedly with distilled water 3 times, and then soaked in distilled water for 30 minutes.

[0067] 1.6 Preparation of cell suspension: Gradually remove distilled water with pipette tip, grind the material into a homogenate, add freshly prepared Carno fixative, let stand for 10 min, discard the first precipitate, then let the upper cell suspension stand for another 30 min, discard the upper part of the clear liquid, and collect the remaining cell suspension.

[0068] 1.7 Droplet preparation: Use a 1000ml pipette to aspirate 1-3 drops of cell suspension and drop them onto a ground glass slide. After observing the droplets spread out rapidly, gently heat them over an alcohol lamp to dry them.

[0069] 1.8 Staining: The above slides were dried at 37°C for 30 minutes. The freshly prepared Giemsa staining solution was added to the chromosome slide specimens and incubated overnight in a humidified chamber. The next day, the slides were gently and repeatedly rinsed with distilled water to remove the surface stain. After the slides were dried at room temperature, they were covered with neutral resin and sealed with nail polish.

[0070] 1.9 Microscopic examination: Add immersion oil to the center of the coverslip, observe under a 100x objective lens of a microscope, take a picture and save the image.

[0071] The final results revealed two types of chromosome numbers: 2n=14 and 2n=12, corresponding to 7 pairs of chromosomes and 6 pairs of chromosomes, respectively. It was determined that there must be a change in chromosome number between these two morphologies in this species, involving either chromosome breakage (14→12) or fusion (12→14) (see [link to study]). Figure 1 ).

[0072] 2. Take young tissue from the material and conduct a genome survey of the species to determine information such as the genome size. The specific implementation procedure and method are as follows:

[0073] 2.1 Take fresh, young tissue from the material and extract DNA from the tissue using the classic DNA extraction CTAB method or a commercial DNA extraction kit;

[0074] 2.2 DNA purity was detected using a Nanodrop instrument and DNA was precisely quantified using a Qubit instrument. The required purity and quality must meet the requirements for library construction.

[0075] 2.3 After the sample passes the test, it is randomly broken using an ultrasonic disruptor. The entire library is prepared by following conventional methods such as end repair, A-tailing, amplification, and product purification. The library is then tested according to requirements.

[0076] 2.4 The qualified libraries were sequenced using DNASEQ-T7, with a sequencing depth of at least 100x for the genome survey, to assess genome size, redundancy, and heterozygosity. This invention found that insufficient data depth can lead to significant biases in subsequent genome information analysis and result in inappropriate sequencing strategies. Therefore, the sequencing data volume needs to reach at least 100x the expected genome size to obtain sufficient information for accurate assessment of genome size, redundancy, and heterozygosity.

[0077] 2.5 The raw image data files obtained from sequencing are converted into raw reads through base calling and stored in FASTQ format;

[0078] 2.6 The raw data is filtered to obtain clean reads (data to be analyzed), which are also stored in FASTQ format. The filtering criteria include, but are not limited to, the following: a) removing reads with an N base content greater than 5%; b) removing reads with a quality value of less than or equal to 5 bases exceeding 50%; c) removing reads contaminated with adapters; d) removing repetitive sequences caused by PCR amplification. This invention has found that it is essential to ensure that the raw reads undergo the above filtering steps to obtain clean reads that can be used for subsequent analysis; otherwise, any sequences that are not removed will lead to incorrect assembly results.

[0079] 2.7 The obtained data needs to be compared with more than 50,000-60,000 reads from the public database NT (Nucleotide Sequence Database) to determine whether the sample has been contaminated by external sources. In this embodiment, the first 50,000 reads (read length data) from all reads in the FASTQ file were extracted for analysis to determine whether the sample was uncontaminated. Research has found that if the number of data reads selected is too small, it cannot represent the contamination status of the entire dataset well, while selecting too large a number of data reads (e.g., more than 60,000) will lead to time-consuming calculations and unnecessary waste of computing resources. After exploration, the comparison results of 50,000-60,000 reads were selected to evaluate whether the sample was contaminated.

[0080] 2.8 Kmer frequency depth distribution analysis was performed using Jellyfish (v2.2.10) with default parameters, and the genome size was estimated accordingly. The estimated genome size for this material was 390 Mb. Meanwhile, the genome heterozygosity was 0.3% and the repetition rate was 48%, which determined the genome to be a low heterozygous genome.

[0081] At this point, the data generated by the second-generation sequencing of mulberry trees were also obtained, which can be used for data error correction in subsequent step 4 and for assessing the genomic consistency of the obtained T2T genome.

[0082] 3. HIC Library Construction and Sequencing. The construction of HIC libraries differs from previous library construction methods. Before library construction, after the sample has passed testing, a series of steps are required, detailed below:

[0083] 3.1 Cell Crosslinking: This invention found that this step requires fresh tissue samples (fresh is defined as young tissue, newly sprouted seedlings, or newly grown leaves, etc.). Other non-fresh tissue samples cannot obtain sufficient interaction signals, which will cause some contig sequences to fail to attach to chromosomes. Using newly grown mulberry leaves, the samples are fixed with formaldehyde to crosslink intracellular proteins with DNA and DNA with DNA, preserving their interaction relationships and maintaining the 3D structure within the cell.

[0084] 3.2 Restriction enzyme digestion: DNA is digested using restriction endonucleases to create sticky ends on both sides of the cross-links. The restriction endonuclease commonly used is DPN II.

[0085] 3.3 End Repair: Biotin-labeled bases are introduced using the end repair mechanism to facilitate subsequent DNA purification and capture;

[0086] 3.4 Circularization: The end-repaired DNA is circularized, and the DNA fragments containing interactions are circularized to ensure that the location of interacting DNA can be determined during subsequent sequencing and analysis;

[0087] 3.5 DNA purification and capture: DNA is decrosslinked, purified, and broken into 300bp-700bp fragments. DNA fragments containing interaction relationships are captured using streptavidin magnetic beads for library construction.

[0088] 3.6 Library quality control: After library construction, Qubit 2.0 and Agilent 2100 were used to detect the library concentration and insert size, respectively. Q-PCR was used to accurately quantify the effective concentration of the library to ensure library quality.

[0089] 3.7 Sequencing: After the library passes inspection, high-throughput sequencing is performed using the Illumina platform with a read length of PE150. The generated raw image data is converted into raw data and stored in FASTQ format after Base Calling.

[0090] 3.8 Using fastp (v0.21.0), the raw data (raw reads) is filtered to remove connector sequences and low-quality reads, retaining only high-quality clean data (data to be analyzed).

[0091] 3.9 Use HICUP (v0.8.0) to align clean data to the reference genome (obtained in step 4.5) and filter it. The filtering criteria are: remove reads that fail to align to the reference genome at both ends; remove invalid reads; remove repetitive sequences caused by PCR amplification; and retain only valid HIC sequencing data.

[0092] 3.10 Testing revealed that both the total amount and quality of HIC data generated must meet specific requirements to obtain sufficient interaction information. The raw data must be larger than 100x the genome size, and the proportion of bases with Q20 and Q30 or higher must be above 95% and 90%, respectively. Only after rigorous filtering according to these standards can sufficient effective HIC data be generated. The final effective HIC data volume needs to reach at least 60% to include all interaction signals; otherwise, much interaction signal will be lost.

[0093] At this point, the valid HIC sequencing data can be used for subsequent HIC-assisted assembly and mounting.

[0094] 4. By combining ONT ultra-long sequencing and PacBio HiFi sequencing with HIC-assisted assembly technology derived from Chromossome Conformation Capture-3C, a complete (also known as telomere-to-telomere, or T2T) genome sequencing and assembly of this species was performed, resulting in a high-quality genome with 0 gap. The specific operation procedure is as follows:

[0095] 4.1 Using newly grown mulberry leaves, we set the throughput required for ONT ultra-long and PacBio HiFi sequencing. For ONT ultra-long, the library preparation standard requires the selection of fragments larger than 100k; for PacBio HiFi, the library preparation requires fragments larger than 20k to obtain raw reads.

[0096] This invention reveals that library construction for ONT ultra-long and PacBio HiFi has strict requirements. The extracted DNA tissue must be sufficiently fresh, and the sample weight must be at least 20g. Testing, using ONT ultra-long as an example, shows that if the sample freshness and total weight do not meet the requirements, it will be difficult to form a qualified library. Following the DNA extraction quality testing process described in section 2, qualified libraries are constructed according to the specific requirements of the ONT and PacBio platforms, and sequencing is performed using both platforms to obtain raw reads (or raw data). Testing reveals specific requirements for the size of the library construction fragments: ONT ultra-long requires fragments larger than 100kbps for library construction; otherwise, they cannot be used to fill all gaps. PacBio HiFi requires fragments larger than 20kbps for library construction; otherwise, the length and accuracy of the long reads used for assembly cannot be guaranteed.

[0097] At the same time, there are clear requirements for the amount of data and sequencing depth that must be generated: based on the genome survey results, in this embodiment, the mulberry genome is estimated to be a low heterozygous genome, and the estimated genome size is the standard. ONTultra-long must generate more than 100x of data, and PacBio HiFi must generate more than 40x of data to ensure the length, accuracy and amount of long read data used for assembly, so as to assemble a complete genome.

[0098] 4.2 Data filtering of raw data is crucial and requires strict filtering of both ONT ultra-long and HiFi raw data according to the following standards. Specifically, for ONT ultra-long raw data, remove fail reads (useless data) with an average quality score less than 7. Use Filtlong (v0.2.4) to filter data, filtering segments <10kb to obtain valid data for subsequent analysis. Use Porechop (v0.2.4) with default parameters to filter connector sequences, and continue using Filtlong (v0.2.4) to filter sequences <30kb, retaining reads with mean read quality scores >90% for subsequent assembly.

[0099] 4.3 For HiFi raw sequencing data subreads, CCS (v6.0.0) was used to filter out low-quality subreads with sequencing cycles of less than 3 times and an SNR of less than 2.5, so as to obtain effective data for subsequent analysis.

[0100] 4.4 While assembling using ONT ultra-long and PacBio HiFi sequencing data, filtered data with an estimated genome size of 100x was generated using the second-generation sequencing data obtained in step 2 for error correction.

[0101] 4.5 Hifiasm (v0.16.1) was used to assemble the genome from the HiFi data to obtain the assembly results based on the PacBio HiFi data; nextDenovo (v2.5.0) was used to assemble the ONT ultra-long data, and Purge_haplotigs (v1.0.4) / Purge_dups (v1.2.5) was used to remove heterozygotes from the genome; the assembly results based on the ultra-long data were obtained. The sizes of the two assembled genomes were compared with the genome size obtained from the genome survey in step 2, and the genome with the closest size to the genome size assessed by the survey was used as the reference genome.

[0102] 4.6 Use BUSCO (version: 4.1.4; parameter: –evalue 1e-05) to evaluate the assembly results. The index should reach more than 90%.

[0103] 4.7 Minimapp2 (2.17-r941) was used to align mitochondria, chloroplasts, etc., to remove assembly contamination sequences; bacterial contamination was removed using the blast refseq library, and contigs with low read support were removed.

[0104] 4.8 The contigs retained in 4.7 are assembled with the help of effective data generated by HIC sequencing. Specifically, the retained contigs are clustered, ordered and oriented, and deredundant.

[0105] 4.9 After HIC assembly and adjustment, chromosome-level genome sequences were first obtained by using 100 N-patch gaps, at which point all sequences were located on 6 chromosomes;

[0106] 4.10 Extend telomere sequences on the mounted chromosomes: Use winnowmap (v1.11, parameter k=15, -MD) to align all ONT reads to the reference genome and extend the telomere sequences on both sides of the genome.

[0107] 4.11 For sequences mounted on 6 chromosomes, there are still gap sequences filled with 100 Ns. The gap regions need to be filled again. Winnowmap (v1.11, parameter k=15,-MD) is used to compare the gap filling data with the genome gap regions and fill the gaps to obtain the genome with 0 gaps in this species.

[0108] 4.12 The 0-gap genome (the genome after gap filling) that has been mounted earlier must undergo another round of error correction before it can be used as the final T2T genome. The specific operation method is as follows: Align the HiFi reads of no less than 10kb generated by PacBio HiFi sequencing with the version of the genome after gap filling. Use samtools "view" (v1.10, parameter: -F256) to filter the alignment fragments. Use the "falconc bam-filter-clipped" software to delete chimeric alignment fragments (-tF 0x104). Use racon software for error correction to obtain the assembled complete genome (T2T genome).

[0109] 4.13 The continuity, integrity, and consistency of the T2T genome were assessed. All gaps in the T2T genome were completely filled, indicating perfect genome continuity. The integrity of the T2T genome was assessed using BUSCO, with a BUSCO value of 96.1%, indicating excellent genome integrity. Furthermore, the consistency of the T2T genome was assessed based on the sequencing data obtained in step 2 using next-generation sequencing technology. 99.81% of the next-generation sequencing data could be aligned to the T2T genome in this embodiment, demonstrating excellent consistency between the pre- and post-assembly assembly methods described in this invention. Thus, the complete genome of this species was obtained. The genome information is shown in Table 1 below:

[0110] Table 1. Chromosome length information in the T2T genome of this species.

[0111] chr1 105 chr2 86 chr3 75 chr4 60 chr5 50 chr6 31

[0112] 5. After obtaining the complete T2T genome of this species (as shown in Table 1), which includes the complete sequence information of all chromosomes, Bedtools can be used to obtain the sequence information of any chromosome, any position and any interval, and also to know the specific composition of these sequences, the characteristics of GC content, etc.

[0113] 6. Using the valid HIC data obtained in step 3, the complete chromosome interaction matrix of the genome at the T2T level of this species can be obtained through Juicebox (v1.11.08). In this embodiment, regions with weakened chromosome interaction signals were found on all chromosomes, and these regions are displayed in 5Mb windows. Figure 2 A shows the interaction matrix diagram of the chr5 chromosome.

[0114] Regions where fusion breaks occur have weak interaction signals within chromosomes. The key is to screen out false positives from many regions with weak chromosomal interaction signals and retain the positive weak-force break-fusion sites. The method described in this invention establishes a clever "seeking similarities while reserving differences" differential analysis approach.

[0115] First, we need to confirm the existence of these regions. This is done using ONT ultra-long and PacBio HiFi reads from long-read sequencing technologies. Minimap2 is used to map all reads into the T2T genome to examine read coverage in these regions and confirm their authenticity. Figure 2 As can be seen from B, all these intervals are covered by Pacbio HiFi and ONT ultra-long long reads, so these intervals are reliable.

[0116] This invention employs a cleverly designed differential analysis method that "seeks similarities while reserving differences" to compare the sequence composition and differences of these weakly interacting signal regions. Step 1 has already determined that the species has 12 and 14 chromosome numbers, which translates to 6 and 7 chromosomes in a single genome. Therefore, there is only one break / fusion site between these two morphologies. For example... Figure 2 As shown in Figure C, orange indicates that these regions contain the same sequence. Comparison reveals that weak interaction signal regions on chromosomes other than chr5 all have similar sequence compositions, meaning there are no break-fusion sites in these regions. However, the weak interaction signal region of chr5 exhibits different characteristics. The same sequence composition is only present in the 28.5-30Mb region of chr5, and its abundance is significantly reduced. No similar sequence composition is found in the 30-32Mb region (the region shown by the box in Figure 2A). Therefore, the 28.5-32Mb region of chr5 is determined to be a completely different region from the weak interaction regions on other chromosomes.

[0117] Using the Straw tool to examine the interaction matrix data within the 28.5-32 Mb range of chromosome 5 (chr5), many anomalous data points with interaction values ​​of 0 were found in the 30-32 Mb range. For example... Figure 3The red highlighted area in the image shows a significant difference from other interaction matrix data within and outside this region. Combining this with the characteristic that the sequence composition in the 28.5-30Mb interval is a repetitive sequence, the precise locations of chromosome breakage and fusion were obtained.

[0118] Based on the break-fusion region indicated by the above results, the sequence and annotation information of this region were extracted from the complete T2T genome using Bedtools. It was found that this region was annotated as a tandem 25S rDNA sequence containing 185 25S rDNA copies. Figure 4 As shown.

[0119] The technique demonstrated here identified a break-fusion site in this species, consistent with previously reported results. Based on the identified region and the T2T genome assembled using this method, all sequence information for the break-fusion region and its neighboring regions has been obtained. This provides extremely complete data for the evolutionary study of this species' genome and paves the way for elucidating the break-fusion mechanism.

[0120] The results of identifying the break-fusion location are consistent with the FISH results obtained in this species' research published by the inventors in Horticulture Research in January 2022, entitled "Chromosome restructuring and number change during the evolution of Morus notabilis and Morus alba," and in Scientific Reports in 2017, entitled "FISH-based mitotic and meiotic diakinesis karyotypes of Morus notabilis reveal a chromosomal fusion-fission cycle between mitotic and meiotic phases." However, these two results only roughly indicated that the region was near the 25S rDNA, without knowing the specific internal sequence or the exact location of the break. The results obtained by the method of this invention clearly identified the break region and simultaneously obtained all sequence information of all regions inside and near the break region, paving the way for the analysis of this mechanism.

[0121] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.

Claims

1. A method for accurately identifying the location of chromosome breaks and / or fusions, characterized in that, Includes the following steps: s1 performed karyotype analysis on the plant samples to determine the presence of chromosome breakage and fusion; s2. Take young tissue samples from the plant to be tested and conduct a genome survey, including the following steps: s2.1 Extract DNA from the plant sample to be tested; s2.2 Constructing a DNA library; s2.3 Sequencing is performed on qualified libraries, and the raw image data files obtained from sequencing are converted into raw data using base recognition technology; s2.4 filters the raw data to obtain the data to be analyzed. data ; The data to be analyzed obtained from s2.5 data Extract 50,000-60,000+ reads and compare them with a public database to determine if the sample was not contaminated by external sources. s2.6 Perform Kmer frequency deep distribution analysis to obtain information on genome size, repetition, and heterozygosity of the plant under test; s3 HIC library construction and sequencing, including: s3.1 Take young tissue samples of the plant to be tested, fix and cross-link them to maintain the 3D structure within the cells; s3.2 endonuclease digestion, end repair, end repair, cyclization; s3.3 DNA purification and capture, library construction, and filtering to obtain effective HIC sequencing data; s4 combined ONT ultra-long sequencing and PacBio HiFi sequencing, along with HIC-assisted assembly technology, to obtain the complete genome of the plant being tested. s4.1 DNA was extracted from young tissue samples of the plant to be tested. The required throughput for ONT ultra-long sequencing and PacBio HiFi sequencing was set. For ONT ultra-long sequencing, fragments longer than 100k were selected for library construction. For PacBio HiFi sequencing, fragments longer than 20k were selected for library construction to obtain raw read length data. s4.2 Filter the raw read data and retain read data with an average quality score greater than 90% for subsequent assembly; s4.3 The filtered read data from s4.2 and the HIC sequencing data described in S3 are combined to perform data assembly and chromosome mounting to obtain the complete genome of the plant to be tested; s5. Obtain the interaction matrix between complete chromosomes in the complete genome of the plant to be tested, and accurately determine the location of chromosome breakage and fusion.

2. The method according to claim 1, characterized in that, The tender tissue sample is a newly sprouted seedling, a newly grown tender leaf, and / or the plant is a mulberry tree.

3. The method according to claim 1, characterized in that, In s2.3, the depth of sequencing genome survey analysis is 100x or higher.

4. The method according to claim 1, characterized in that, In s2.4, the filtering of the raw data includes at least the following: removing reads with an N base content greater than 5%; removing reads with a quality value less than or equal to 5 in more than 50% of the bases; and removing reads with connector contamination. Remove repetitive sequences caused by PCR amplification.

5. The method according to claim 1, characterized in that, In s1, the karyotype analysis of the plant samples was performed using the cell wall removal hypotonic method.

6. The method according to claim 1, characterized in that, The HIC sequencing data obtained in s3 is greater than 100x the genome, with a base quality of over 95% for Q20 and over 90% for Q30.

7. The method according to claim 1, characterized in that, In s4.1, based on the results of the genome survey, when the survey results are evaluated as a low heterozygous genome, ONT ultra-long is set to generate more than 100x the amount of data, and PacBio HiFi is set to generate more than 40x the amount of data.

8. The method according to claim 1, characterized in that, In s4.2, the raw read data filtering includes: for ONT ultra-long raw read data, using Filtlong to filter fragments <10kb, removing useless data with an average quality of less than 7, and obtaining valid data for subsequent analysis; filtering adapter sequences, filtering sequences <30kb, and retaining read data with an average read quality score >90%; and / or for PacBio HiFi raw data, performing data filtering to filter low-quality subreads with sequencing cycles less than 3 times and an SNR less than 2.

5.

9. The method according to claim 1, characterized in that, The data assembly and chromosome mounting described in s4.3 to obtain the complete genome of the plant to be tested includes: after obtaining the species 0-gap genome of the plant to be tested, aligning the HiFi reads of no less than 10kb generated by PacBioHiFi sequencing with the 0-gap genome, filtering the aligned fragments, deleting chimeric aligned fragments, and performing error correction.

10. The application of the method of claims 1-9 in the accurate identification of chromosome breakage and / or fusion sites in a plant species, wherein the plant is a mulberry tree.