A method of identifying chromosomal translocations integrating three-dimensional genomic and third-generation genomic data
By integrating three-dimensional genome and third-generation genome data and combining multi-omics analysis methods, the false positive and insufficient resolution problems of chromosome translocation detection in existing technologies are solved, and high-accuracy chromosome translocation screening and target gene identification are achieved.
Patent Information
- Application Number
- CN202411657006.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-19
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2044-11-19
AI Technical Summary
Existing chromosomal translocation detection methods have problems with false positives and inaccurate detection when detecting repetitive regions and regions with high sequence similarity, especially methods based on second-generation genome sequencing and Hi-C data.
Three-dimensional genome and third-generation genome data were integrated, and chromosomal translocations were identified using PBSV software and the deep learning tool EagleC. Combined with nucleic acid amplification and Sanger sequencing verification, promoter-enhancer chromatin loops were annotated using the deep learning tool NeoLoopFinder, and target genes with expression differences greater than a threshold were screened.
The accuracy and confidence of screening for chromosomal translocations have been improved, reaching an accuracy rate of 98%, accurately identifying potential target genes caused by chromosomal translocations.
Smart Images

Figure CN119541631B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and in particular to a method for identifying chromosome translocation by integrating three-dimensional genome and three-generation genome data. Background Art
[0002] The following statements merely provide background information related to the present disclosure and do not necessarily constitute prior art.
[0003] Structural variations (SVs) generally refer to abnormalities in DNA sequences greater than 50 base pairs within the genome. SVs primarily include deletions, insertions, duplications, inversions, and translocations, and are a major source of genomic variation.
[0004] Chromosome translocation, the repositioning of chromosome segments, is the most common chromosomal abnormality. Chromosome translocation requires DNA double-strand breaks (DSBs), which are typically caused by spontaneous DNA replication errors, exogenous stressors such as ionizing radiation or chemicals, and programmed breaks in the immune system. DSBs are generally repaired by the body's DNA repair system, but if repair goes awry, the broken ends can physically contact and interact within a certain space, resulting in incorrect linkage, leading to chromosomal translocation. In addition to DSB repair, factors such as histone modifications contribute to chromosomal translocation. The spatial folding structure and activity of chromosomes are also important factors in the formation of chromosomal translocations. Chromosome translocations can occur between different chromosomes or at different locations on the same chromosome, altering the position of genes and often leading to gene mutations and fusions.
[0005] There are two main methods for detecting chromosomal translocations: one is the classic method of identifying chromosomal translocations based on second-generation or third-generation genome sequencing, and the other is the method of identifying chromosomal translocations from Hi-C data using deep learning methods.
[0006] Chromosomal translocations are identified based on second-generation or third-generation genome sequencing technology. Due to the limitation of sequencing read length, only genomic regions within 10 to 20 kb can be detected. This leads to false positives in the detection of repetitive regions and regions with high sequence similarity on the genome.
[0007] Identifying chromosomal translocations from Hi-C (highest-throughput chromosome conformation capture) data using deep learning methods has also been a hot topic in recent years. Hi-C is a method derived from 3C (chromosome conformation capture) to study the three-dimensional structure of chromatin. It is an analytical technique that studies chromatin interactions and reveals the three-dimensional structure of the entire genome on a genome-wide scale. The basic principles of this technology are: (1) Formaldehyde cross-linking: First, use 37% formaldehyde to fix the DNA-protein or protein-protein complexes that are cross-linked together or in close spatial proximity within the cell; (2) Enzyme digestion and end labeling: Use restriction endonucleases to digest and separate chromatin, and label the fragment ends with biotin; (3) Connection of fragment ends: Use DNA ligase to connect the ends of different cuts to form a circular chimeric molecule; (4) Purification and shearing: These chimeric molecules are purified, sheared, and broken into DNA fragments; (5) Establishment of Hi-C library: Use biotin labeling precipitation technology to obtain the labeled target DNA fragments, and select DNA fragments of appropriate size to establish Hi-C library (Hi-C library); (6) Hi-C sequencing and data analysis: Use high-throughput sequencing technology to perform paired-end sequencing (paired-end sequencing) on the samples contained in the Hi-C library to generate a large number of paired-end fragments (paired-end reads) for analysis. Identifying chromosomal translocations based on Hi-C data also has two disadvantages: first, this inference based on another type of omics data has a certain degree of false positives; second, since the resolution of Hi-C is at least 5kb, the detection of chromosomal translocation breakpoints is inaccurate.
[0008] In view of this, the present invention is proposed. Summary of the Invention
[0009] The present invention aims to provide a method for identifying chromosomal translocations by integrating three-dimensional genomic and third-generation genomic data, thereby improving the ability to screen for chromosomal translocations in samples. Another object of the present invention is to provide a chromosomal translocation detection device, a computer-readable medium, and applications thereof based on this method.
[0010] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0011] In a first aspect, a method for identifying chromosomal translocations by integrating three-dimensional genome and three-generation genome data is provided, the method comprising:
[0012] (A) obtaining third-generation genome sequencing data of a sample to be tested, identifying chromosomal translocations based on the third-generation genome sequencing data, and using the obtained chromosomal translocations as candidate chromosomal translocations to form a first data set;
[0013] (B) obtaining Hi-C sequencing data or sequencing data based on a Hi-C-derived technology of the sample to be tested, identifying chromosomal translocations based on the Hi-C sequencing data or sequencing data based on a Hi-C-derived technology, and using the obtained chromosomal translocations as candidate chromosomal translocations to form a second data set;
[0014] (C) The candidate chromosomal translocations in the intersection of the first and second datasets are selected, and the verified candidate chromosomal translocations are used as the final chromosomal translocations obtained through screening.
[0015] In an optional embodiment, in step (A), PBSV software is used to identify chromosomal translocations from third-generation genome sequencing data; and / or, in step (B), the deep learning tool EagleC is used to identify chromosomal translocations from Hi-C sequencing data or sequencing data based on Hi-C derivative technology.
[0016] In an optional embodiment, in step (C), the candidate chromosomal translocation is verified by nucleic acid amplification and Sanger sequencing.
[0017] In an optional embodiment, the method further comprises using the chromosomal translocation obtained in step (C) and the Hi-C sequencing data or sequencing data based on Hi-C derivative technology obtained in step (B) as input data, and using a deep learning tool to obtain a newly generated chromatin loop due to the chromosomal translocation;
[0018] Then, the promoter-enhancer chromatin loop is annotated, and target genes with expression difference fold greater than a threshold among the promoter-enhancer chromatin loop-related genes are screened, wherein the target genes are target genes caused by chromosomal translocation.
[0019] In an optional embodiment, the deep learning tool includes NeoLoopFinder.
[0020] In an optional embodiment, the method includes obtaining ChIP-Seq or CUT&Tag data of H3K27ac of the sample to be tested, and annotating promoter-enhancer chromatin loops.
[0021] In an optional embodiment, the method includes obtaining RNA-seq data of the sample to be tested, and screening target genes with an expression difference fold greater than 2.
[0022] In a second aspect, a chromosome translocation detection device is also provided, which includes a data acquisition module and a detection module. The data acquisition module is used to obtain third-generation genome sequencing data of the sample to be tested, and Hi-C sequencing data or sequencing data based on Hi-C derivative technology; the detection module executes the method of integrating three-dimensional genome and third-generation genome data to identify chromosome translocation in the first aspect.
[0023] In a third aspect, a computer-readable medium is provided, which stores a computer program, and when the computer program is processed and executed, the method of integrating three-dimensional genome and three-generation genome data to identify chromosome translocations in the first aspect is implemented.
[0024] In a fourth aspect, the method for identifying chromosomal translocation by integrating three-dimensional genome and third-generation genome data of the first aspect, or the chromosomal translocation detection device of the second aspect, or the use of the computer-readable medium of the third aspect in screening potential target genes for diseases is also provided.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] The method for identifying chromosomal translocations provided by this invention combines the advantages of cutting-edge research methods in this field, third-generation genome sequencing and Hi-C sequencing, to identify chromosomal translocations with high confidence, with an accuracy rate of up to 98% verified by experiments. In a more preferred approach, combined multi-omics analysis can more accurately identify potential target genes causing chromosomal translocations. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] Figure 1 The framework of the chromosome translocation identification method provided in Example 1 of the present invention;
[0029] Figure 2 is the number of chromosomal translocations obtained in steps 1 and 2, respectively, for the multiple myeloma cell line sample and the multiple myeloma patient sample in Example 1 of the present invention;
[0030] Figure 3 The intersection of the chromosomal translocations obtained in each step of the multiple myeloma cell line sample and the multiple myeloma patient sample in Example 1 of the present invention;
[0031] Figure 4 The 43 chromosomal translocations in the two intersections of the multiple myeloma cell line sample and the multiple myeloma patient sample in Example 1 of the present invention are distributed on the genome;
[0032] Figure 5 The PCR verification results for t(8;14) and t(5;8) in Example 1 of the present invention are as follows;
[0033] Figure 6 The verification results of Sanger sequencing for t(8;14) and t(5;8) in Example 1 of the present invention are as follows;
[0034] Figure 7 The novel chromosomal translocation t(5;8) and its potential oncogene CPEB4 were found in a PAT1 myeloma patient using the method of Example 1 of the present invention. DETAILED DESCRIPTION
[0035] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0036] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0037] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0038] In a first aspect, a method for identifying chromosomal translocations by integrating three-dimensional genome and three-generation genome data is provided, the method comprising the following (A) to (C):
[0039] (A) Obtaining third-generation genomic sequencing data for the sample to be tested, identifying chromosomal translocations based on the third-generation genomic sequencing data, and identifying all chromosomal translocations obtained in this step as candidate chromosomal translocations. These candidate chromosomal translocations constitute a first data set. This step can be performed using third-generation sequencing technologies known in the art, including but not limited to HeliScope single molecule sequencing technology, single-molecule real-time sequencing technology (SMRT), Oxford nanopore sequencing technology, and Geno Care single molecule sequencing technology.
[0040] In an optional embodiment, in step (A), PBSV software is used to identify chromosomal translocations from third-generation genome sequencing data.
[0041] In an optional embodiment, pbmm2 is used to align the reference genome, and then pbsv is used to discover and identify genomic structural variations.
[0042] In an optional embodiment, the parameters are set as follows: use pbmm2 to align the reference genome, pbmm2align --preset HIFI --sort; then use pbsv to discover genomic structural variations, pbsv discover --hifi; finally, use pbsv to identify structural variations, pbsv call -m 50 –-hifi.
[0043] (B) Obtaining Hi-C sequencing data or sequencing data based on a Hi-C derivative technology of the sample to be tested, identifying chromosomal translocations based on the Hi-C sequencing data or sequencing data based on a Hi-C derivative technology, and using the obtained chromosomal translocations as candidate chromosomal translocations to form a second data set.
[0044] Hi-C-derived technologies include but are not limited to single-cell Hi-C, In situ Hi-C, DNase Hi-C, Micro-C, DLO Hi-C and other techniques, preferably including Micro-C. Micro-C is a Hi-C derivative technique with nucleosome-level (approximately 200 bp) resolution. Micro-C uses formaldehyde and disuccinimidyl glutarate (DSG) to fix and cross-link cells. The cell membranes are then solubilized and the cross-linked chromatin is digested with micrococcal nuclease (MNase) to nucleosome-level resolution.
[0045] In an optional embodiment, in step (B), the deep learning tool EagleC is used to identify chromosomal translocations from Hi-C sequencing data or sequencing data based on Hi-C derivative technologies.
[0046] In an optional implementation, the parameters are set as follows: predictSV -g reference –balance-type CNV –output-format NeoLoopFinder.
[0047] (C) Select candidate chromosomal translocations from the intersection of the first dataset and the second dataset, and use the verified candidate chromosomal translocations as the final chromosomal translocations obtained through screening. Verification can be performed using methods known in the art, including but not limited to nucleic acid amplification and Sanger sequencing. Nucleic acid amplification can be performed using PCR (polymerase chain reaction).
[0048] In an optional embodiment, the method for identifying chromosomal translocations by integrating 3D and third-generation genomic data further includes final screening for target genes affected by the chromosomal translocation. This involves using the chromosomal translocation obtained in step (C) and the Hi-C sequencing data or sequencing data derived from Hi-C-derived techniques obtained in step (B) as input data, and utilizing deep learning tools to identify newly generated chromatin loops (loops) resulting from the chromosomal translocation. Loops bring sites that are linearly distant into close spatial proximity. Loop sites typically have a promoter at one end and an enhancer at the other. This close proximity of promoters and enhancers modulates gene expression, resulting in changes in the expression of genes associated with such loops. Annotating promoter-enhancer loops and then screening for target genes with expression differences exceeding a threshold helps elucidate the relationship between loops and gene transcriptional regulation across samples. Genes with expression differences exceeding the threshold are considered target genes caused by the chromosomal translocation.
[0049] In an optional embodiment, the deep learning tool includes NeoLoopFinder.
[0050] In an optional implementation, NeoLoopFinder parameters are set as follows:
[0051] Assemble structural variants across breakpoints with assemble-complexSVs --balance-type CNV --protocol insitu; identify neoloops resulting from structural variants with neoloop-caller --balance-type CNV --protocol insitu --prob 0.95.
[0052] In an optional embodiment, ChIP-Seq or CUT&Tag data of H3K27ac (acetylation of histone H3 lysine 27) is obtained for the sample to be tested, and promoter-enhancer loops are annotated by analyzing DNA that interacts with histone H3K27ac.
[0053] ChIP-Seq (Chromatin immunoprecipitation followed by sequencing, ChIP-seg, chromosome immunoprecipitation sequencing) and CUT&Tag (Cleavage Under Targets and Tagmentation) are technologies for studying DNA-protein interactions. Their basic principle is to first obtain the DNA region bound by the target protein and then perform high-throughput sequencing on this DNA region.
[0054] CUT&Tag is a new method for studying protein-DNA interactions. The basic principle of CUT&Tag technology is that under the mediation of a specific antibody, the pA-Tn5 fusion protein fragments the target DNA only at the local area where the target histone modification mark, transcription factor or chromatin regulatory protein binds to chromatin, and sequencing adapters are added at the same time. Since the pA-Tn5 transposome only binds and cuts the DNA in its adjacent space, the signal-to-noise ratio of the entire experiment is greatly improved, and the experimental steps are simplified.
[0055] In an optional embodiment, RNA-seq data of the sample to be tested is obtained, and target genes with an expression difference fold greater than 2 are screened as target genes caused by chromosomal translocation.
[0056] In the second aspect, a chromosome translocation detection device is also provided, which includes a data acquisition module and a detection module. The data acquisition module is used to obtain the third-generation genome sequencing data and Hi-C sequencing data or sequencing data based on Hi-C derivative technology of the sample to be tested; the detection module executes the method of integrating three-dimensional genome and third-generation genome data to identify chromosome translocation in the first aspect.
[0057] In an optional embodiment, the data acquisition module is also used to obtain ChIP-Seq or CUT&Tag data of the sample to be tested.
[0058] In an optional embodiment, the data acquisition module is also used to obtain RNA-seq data of the sample to be tested.
[0059] In an optional embodiment, the detection module can be stored in a memory in the form of software or firmware, or embedded in the operating system (OS) of the chromosome translocation detection device. The chromosome translocation detection device may also be pre-installed with a storage module that stores data, program code, etc. required to execute some or all of the above modules.
[0060] The chromosome translocation detection device may include a memory, a processor, a bus, and a communication interface, wherein the memory, processor, and communication interface are electrically connected to each other directly or indirectly to enable data transmission or exchange. For example, these components may be electrically connected to each other via one or more buses or signal lines. The processor may process information and / or data related to target identification to perform one or more functions described in this application.
[0061] In practical applications, the chromosome translocation detection device can be a server, a cloud platform, a mobile phone, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a handheld computer, a netbook, a personal digital assistant (PDA), a wearable electronic device, a virtual reality device, and the like.
[0062] In a third aspect, a computer-readable medium is provided, the computer-readable medium storing a computer program that, when executed, implements the method for identifying chromosomal translocations by integrating three-dimensional genome and third-generation genome data according to the first aspect. Computer-readable media include, but are not limited to, various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories, random access memories, magnetic disks, or optical disks.
[0063] In a fourth aspect, the method for identifying chromosomal translocation by integrating three-dimensional genome and third-generation genome data of the first aspect, or the chromosomal translocation detection device of the second aspect, or the use of the computer-readable medium of the third aspect in screening potential target genes for diseases is also provided.
[0064] In an optional embodiment, the disease comprises a tumor.
[0065] The present invention is further described below by way of specific examples. However, it should be understood that these examples are merely provided for more detailed description and are not to be construed as limiting the present invention in any form.
[0066] Example 1
[0067] This example provides a method for identifying chromosomal translocations by integrating 3D genomic and third-generation genomic data. The samples tested were five multiple myeloma (MM) cell lines (KMS11, LP1, MM1S, RPMI8266, and U266) and three multiple myeloma patient samples (PAT1, PAT2, and PAT3). The method framework is as follows: Figure 1 As shown, the following steps are included:
[0068] 1. The above samples were sent to Berry Genomics for third-generation genome sequencing. The sequencing platform used Pacbio's Single-Molecule Sequencing in Real Time (SMRT) and the HiFi sequencing mode was selected to obtain high-quality sequencing data. The sequencing volume of each sample was 45GB. After obtaining the sequencing data, chromosomal translocations were identified from the third-generation genome sequencing data using PBSV software. The specific parameters are as follows:
[0069] 1) pbmm2 align reference.fa input.bam out.bam --preset HIFI --sort;
[0070] 2) pbsv discover –-hifi out.bam out.svsig.gz;
[0071] 3) pbsv call -m 50 –hifi out.svsig.gz out.vcf.
[0072] like Figure 2 As shown, 80 chromosomal translocations were found in 5 multiple myeloma cell line samples, including known chromosomal translocations such as t (11;14), t (14;16), t (8;14), and t (16;22). 30 chromosomal translocations were found in 3 multiple myeloma patient samples, including known chromosomal translocations such as t (8;14) and t (11;14).
[0073] 2. The above samples were sent to Fraser Genomics for Hi-C sequencing based on the Illumina platform. Five multiple myeloma cell lines underwent in situ Hi-C sequencing, and three multiple myeloma patient samples underwent Micro-C sequencing. Each sample had a sequencing capacity of 300G. After obtaining the sequencing data, EagleC, a deep learning-based tool, was used to identify chromosomal translocations from the sequencing data. The specific parameters are as follows:
[0074] 1) predictSV --hic-5k input.mcool:: / resolutions / 5000 \
[0075] --hic-10k input.mcool:: / resolutions / 10000 \
[0076] --hic-50k input.mcool:: / resolutions / 50000 \
[0077] -O out.predictsv.txt -g reference –balance-type CNV –output-formatNeoLoopFinder --prob-cutoff-5k 0.8 --prob-cutoff-10k 0.8 --prob-cutoff-50k0.99999.
[0078] like Figure 2 As shown, 71 chromosomal translocations were found in the 5 multiple myeloma cell line samples, including known chromosomal translocations t (14; 16), t (8; 14), t (16; 22). 22 chromosomal translocations were found in the 3 multiple myeloma patient samples, including known chromosomal translocation t (11; 14) and etc.
[0079] 3. As shown, the chromosomal translocations of the 5 multiple myeloma cell line samples identified in step 1 and step 2 were intersected, obtaining 35 chromosomal translocations. Figure 3
[0080] The chromosomal translocations of the 3 multiple myeloma patient samples identified in step 1 and step 2 were intersected, obtaining 8 chromosomal translocations.
[0081] The sites of the 43 chromosomal translocations of the two intersections on the genome are shown in Figure 4
[0082] The 43 chromosomal translocations of the two intersections were verified by PCR and Sanger sequencing, of which 42 chromosomal translocations were verified to be correct, with an accuracy rate of 98%, and the specific sites are shown in Table 1.
[0083] And 2 new chromosomal translocations t (8; 14) and t (5; 8) were detected in multiple myeloma patient samples, the PCR verification results are shown in Figure 5 Figure 6
[0084] 4. Take the 42 chromosomal translocations verified to be correct in step 3 and the Hi-C sequencing in step 2 as input, use the deep learning-based tool NeoLoopFinder to find the new loops generated due to chromosomal translocations. The specific parameters are as follows:
[0085] 1) assemble-complexSVs -O out -B overlap.5K_combined.txt \
[0086] --balance-type CNV --protocol insitu \
[0087] -H input.mcool::resolutions / 25000 input.mcool::resolutions / 10000
[0088] input.mcool::resolutions / 5000.
[0089] 2) neoloop-caller -O neo-loops.txt --assembly out.assemblies.txt \
[0090] --balance-type CNV --protocol insitu --prob 0.95 \
[0091] -H input.mcool::resolutions / 25000 input.mcool::resolutions / 10000
[0092] input.mcool::resolutions / 5000.
[0093] 5. The above samples were sent to Berry Genomics for ChIP-seq sequencing and RNA-seq sequencing based on the illumina platform, with a data volume of ChIP-seq 12G and RNA-seq 6G per sample. ChIP-seq sequencing data and RNA-seq sequencing data of H3K27ac of 5 multiple myeloma (MM) cell lines and 3 multiple myeloma patient samples were obtained. According to the ChIP-seq sequencing data of H3K27ac, the promoter-enhancer loop was annotated, and the target genes with a fold change > 2 were screened from the RNA-seq data, i.e. potential target genes caused by chromosomal translocation. Figure 7 An unknown chromosomal translocation t(5;8) was found in the PAT1 myeloma patient, specifically an enhancer on chromosome 8 translocated to the upstream of the CPEB4 gene on chromosome 5, i.e. reconnection occurred at chr5 site 173708142 on chromosome 5 and chr8 site 127798852 on chromosome 8; and an enhancer-promoter loop caused by the t(5;8) translocation and its corresponding highly expressed oncogene CPEB4 were found.
[0094] Table 1 Specific genomic sites of 42 translocations screened in Example 1
[0095]
[0096] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for identifying chromosomal translocations by integrating three-dimensional genome and three-generation genome data, characterized in that: include: (A) obtaining third-generation genome sequencing data of a sample to be tested, identifying chromosomal translocations based on the third-generation genome sequencing data, and using the obtained chromosomal translocations as candidate chromosomal translocations to form a first data set; (B) obtaining Hi-C sequencing data or sequencing data based on a Hi-C-derived technology of the sample to be tested, identifying chromosomal translocations based on the Hi-C sequencing data or sequencing data based on a Hi-C-derived technology, and using the obtained chromosomal translocations as candidate chromosomal translocations to form a second data set; (C) taking the candidate chromosomal translocations from the intersection of the first dataset and the second dataset, and using the verified candidate chromosomal translocations as the final chromosomal translocations obtained by screening; In step (A), PBSV software is used to identify chromosomal translocations from three-generation genome sequencing data; In step (B), the deep learning tool EagleC is used to identify chromosome translocations from Hi-C sequencing data or sequencing data based on Hi-C-derived technologies; In step (C), the candidate chromosomal translocation is verified by nucleic acid amplification and Sanger sequencing; then, the chromosomal translocation obtained in step (C) and the Hi-C sequencing data or sequencing data based on Hi-C-derived technology obtained in step (B) are used as input data, and a deep learning tool is used to obtain a newly generated chromatin loop due to the chromosomal translocation; Then, the promoter-enhancer chromatin loop is annotated, and target genes with expression difference fold greater than a threshold among the promoter-enhancer chromatin loop-related genes are screened, wherein the target genes are target genes caused by chromosomal translocation.
2. The method according to claim 1, characterized in that The deep learning tool includes NeoLoopFinder.
3. The method according to claim 1, characterized in that Obtain ChIP-Seq or CUT&Tag data for H3K27ac of the sample to be tested and annotate promoter-enhancer chromatin loops.
4. The method according to claim 1, wherein Obtain RNA-seq data of the sample to be tested and screen target genes with expression differences greater than 2.
5. A chromosome translocation detection device, characterized in that: It includes a data acquisition module and a detection module, wherein the data acquisition module is used to obtain the third-generation genome sequencing data of the sample to be tested, and Hi-C sequencing data or sequencing data based on Hi-C derivative technology; The detection module executes the method for identifying chromosomal translocation by integrating three-dimensional genome and three-generation genome data as described in any one of claims 1 to 4.
6. A computer-readable medium, characterized in that A computer program is stored, and when the computer program is processed and executed, the method for identifying chromosomal translocation by integrating three-dimensional genome and three-generation genome data according to any one of claims 1 to 4 is implemented.
7. Use of the method for identifying chromosomal translocation by integrating three-dimensional genome and third-generation genome data according to any one of claims 1 to 4, or the chromosomal translocation detection device according to claim 5, or the computer-readable medium according to claim 6 in screening potential target genes for diseases.
Citation Information
Patent Citations
Gene-assisted assembly method using Hi-C technology, chromosome level genome and application
CN116665781A
Micro chromosome assembling and identifying device and method and application of micro chromosome assembling and identifying device and method
CN118335196A