Method for analyzing DNA fusion gene sequencing data and application

CN117133361BActive Publication Date: 2026-09-11DAAN GENE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210556976.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2026-09-11
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

[0007]本申请实施例的目的在于提出一种DNA融合基因测序数据的分析方法及应用,以解决相关技术中融合基因比对分析时间长,跨越融合断点的序列长度要符合既定的长度要求以及不能稳定检出低突变频率样本的阳性结果,准确度和灵敏度低的技术问题

Benefits of technology

[0033]本申请通过从原始比对文件中获取同时存在M序列和S序列的测序序列,并记录测序序列中的M序列比对到人类参考基因组的第一比对信息;提取测序序列中的S序列,并记录S序列比对到人类参考基因组的第二比对信息;根据第一比对信息获取第一碱基序列,根据第二比对信息获取第二碱基序列;将第一碱基序列和第二碱基序列进行拼接,得到融合基因碱基序列;提取原始比对文件中所有测序序列,将提取出的测序序列与融合基因碱基序列进行比对,得到比对结果;对比对结果进行分析,得到融合基因的测序数据;本申请通过拼接第一碱基序列和第二碱基序列得到融合基因碱基序列,并将测序序列与融合基因碱基序列进行比对,可以提高比对分析结果的准确性和灵敏度,同时,优化了序列的比对算法,缩短了序列比对的分析时间,该比对方法还降低了跨越融合断点的序列长度要求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117133361B_ABST
    Figure CN117133361B_ABST
Patent Text Reader

Abstract

The embodiment of the application belongs to the technical field of biological medicine, and relates to a DNA fusion gene sequencing data analysis method and application, which comprises the following steps: obtaining a sequencing sequence with M sequences and S sequences existing simultaneously from an original alignment file, and recording first alignment information of the M sequences in the sequencing sequence aligned to a human reference genome; extracting the S sequences in the sequencing sequence, and recording second alignment information of the S sequences aligned to the human reference genome; obtaining first base sequences according to the first alignment information, and obtaining second base sequences according to the second alignment information; splicing the first base sequences and the second base sequences to obtain fusion gene base sequences; aligning the sequencing sequence extracted from the original alignment file with the fusion gene base sequences to obtain an alignment result, and analyzing the alignment result. The application can improve the accuracy and sensitivity of the alignment analysis result, and simultaneously optimizes the sequence alignment algorithm and shortens the sequence alignment analysis time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biomedical technology, and in particular to an analysis method and application of DNA fusion gene sequencing data. Background Technology

[0002] Tumor cells have chromosomal structures that differ from normal cells. This is due to chromosomal rearrangements caused by DNA (deoxyribonucleic acid) double-strand breaks and reconnections. These rearrangements include deletions, inversions, and translocations of large chromosomal segments. This rearrangement causes two genes that were far apart in the original genome to be placed in upstream or downstream positions, ultimately resulting in the rearrangement of the two original gene sequences in the genome to form a new gene.

[0003] Currently, many cancers are reported to be closely related to fusion genes, such as the presence of multiple fusion genes in non-small cell lung cancer. The US FDA (Food and Drug Administration) has approved several therapeutic drugs targeting specific fusion genes. Therefore, accurately identifying the type of fusion gene is crucial in the overall treatment of cancer.

[0004] With the development of high-throughput sequencing technology, its application in the clinical testing of tumor samples has become more widespread. DNA from tumor tissue or blood is extracted, processed, and sequenced, and finally analyzed by software to determine the actual gene fusion patterns.

[0005] Currently, many factors affect the detection of fusion genes in tumor samples. For example, clinical samples are diverse and exist in various forms; the quality of cell-free DNA in plasma depends on the quality of the extracted blood sample; hemolyzed samples or samples left in blood collection tubes for extended periods exhibit fragmented degradation of cell-free DNA. Furthermore, cell-free DNA contains only low levels of circulating tumor DNA, and patients with low tumor burden also have low levels of circulating tumor DNA in their plasma. These factors lead to a low content of DNA derived from fusion genes in the extracted total DNA, making the detection of fusion genes in tumor samples difficult. Therefore, developing an analytical method or software capable of detecting fusion genes with low mutation frequencies is a current challenge.

[0006] FACTERA (an open-source analysis software capable of detecting gene fusions) is one of the mainstream analysis software currently used for gene fusion analysis. During the analysis process, the software analyzes sequencing sequences suspected to originate from fusion genes, using third-party alignment software combined with its own algorithms to obtain sequences supporting gene fusion. However, the analysis process is time-consuming; furthermore, the software only analyzes target sequences supporting gene fusion in the alignment results file, and requires the sequence length crossing the fusion breakpoint to meet predetermined length requirements; additionally, the software cannot consistently detect positive results for samples with low mutation frequencies (less than 0.5%). Summary of the Invention

[0007] The purpose of this application is to propose an analysis method and application for DNA fusion gene sequencing data, in order to solve the technical problems in related technologies such as long fusion gene alignment analysis time, sequence length across fusion breakpoints must meet predetermined length requirements, inability to reliably detect positive results of samples with low mutation frequency, and low accuracy and sensitivity.

[0008] To address the aforementioned technical problems, this application provides a method for analyzing DNA fusion gene sequencing data, employing the following technical solution:

[0009] Obtain sequencing sequences containing both M and S sequences from the original alignment file, and record the first alignment information of the M sequence in the sequencing sequence to the human reference genome.

[0010] Extract the S sequence from the sequencing sequence and record the second alignment information of the S sequence to the human reference genome;

[0011] The first base sequence is obtained based on the first alignment information, and the second base sequence is obtained based on the second alignment information.

[0012] The first base sequence and the second base sequence are spliced ​​together to obtain the fusion gene base sequence;

[0013] Extract all sequencing sequences from the original alignment file, and align the extracted sequencing sequences with the fusion gene base sequence to obtain the alignment results;

[0014] The alignment results were analyzed to obtain the sequencing data of the fusion gene.

[0015] Furthermore, the first alignment information includes first chromosome information and first position information, and the second alignment information includes second chromosome information and second position information.

[0016] Furthermore, the first chromosome information is the first aligned chromosome of the M sequence terminal base in the human reference genome; the first position information is the alignment position of the M sequence terminal base in the first aligned chromosome, denoted as the terminal base position;

[0017] The second chromosome information is the second aligned chromosome of the S sequence start base in the human reference genome; the second position information is the alignment position of the S sequence start base in the second aligned chromosome, denoted as the start base position.

[0018] Furthermore, the step of obtaining the first base sequence based on the first alignment information includes:

[0019] Starting from the position of the last base, a base sequence of a first preset sequence length upstream of the position of the last base on the first aligned chromosome is extracted to obtain the first base sequence.

[0020] Furthermore, the step of obtaining the second base sequence based on the second alignment information includes:

[0021] Starting from the initial base position, a second base sequence of a predetermined length downstream of the initial base position on the second aligned chromosome is extracted to obtain the second base sequence.

[0022] Furthermore, the step of splicing the first base sequence and the second base sequence includes:

[0023] The last base position of the first base sequence is spliced ​​with the first base position of the second base sequence.

[0024] Furthermore, the step of comparing the extracted sequencing sequence with the fusion gene base sequence to obtain the comparison result includes:

[0025] The extracted sequencing sequence was aligned to the fusion gene base sequence using alignment software;

[0026] The matching results are obtained by determining the matching status of each extracted sequencing sequence with the base sequence of the fusion gene.

[0027] Furthermore, the step of analyzing the comparison results includes:

[0028] The number of sequences whose extracted sequencing sequences completely match the fusion gene base sequence is determined based on the matching results;

[0029] The existence of the fusion gene is determined based on the number of sequences described.

[0030] If present, the sequencing data of the fusion gene are counted.

[0031] To address the aforementioned technical problems, this application also provides an application of the DNA fusion gene sequencing data analysis method described above in the prevention and / or treatment of tumors and / or cancer.

[0032] Compared with the prior art, the embodiments of this application have the following main advantages:

[0033] This application obtains sequencing sequences containing both M and S sequences from the original alignment file and records the first alignment information of the M sequence to the human reference genome; extracts the S sequence from the sequencing sequence and records the second alignment information of the S sequence to the human reference genome; obtains the first base sequence based on the first alignment information and the second base sequence based on the second alignment information; splices the first and second base sequences to obtain the fusion gene base sequence; extracts all sequencing sequences from the original alignment file and aligns the extracted sequencing sequences with the fusion gene base sequence to obtain the alignment results; analyzes the alignment results to obtain the sequencing data of the fusion gene. This application improves the accuracy and sensitivity of the alignment analysis results by splicing the first and second base sequences to obtain the fusion gene base sequence and aligning the sequencing sequence with the fusion gene base sequence. It also optimizes the sequence alignment algorithm, shortens the sequence alignment analysis time, and reduces the sequence length requirement for crossing fusion breakpoints. Attached Figure Description

[0034] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a flowchart of an embodiment of the DNA fusion gene sequencing data analysis method according to this application;

[0036] Figure 2 This refers to the alignment information of a pair of Reads to the human reference genome in the embodiments of this application;

[0037] Figure 3 This is a schematic diagram of the alignment of a Reads portion containing both M and S sequences to the human reference genome in an embodiment of this application.

[0038] Figure 4 This is a schematic diagram of the alignment of a Read containing both M and S sequences to the human reference genome in an embodiment of this application.

[0039] Figure 5 This is a schematic diagram of the alignment of a pair of Reads containing both M and S sequences to the human reference genome in an embodiment of this application.

[0040] Figure 6 This is a schematic diagram illustrating the alignment of Reads containing both M and S sequences to the fusion gene base sequence in an embodiment of this application.

[0041] Figure 7 This is a schematic diagram illustrating the alignment of a pair of Reads containing both M and S sequences to the fusion gene base sequence in an embodiment of this application. Detailed Implementation

[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0043] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0044] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0045] This application provides a method for analyzing DNA fusion gene sequencing data, referencing... Figure 1 The flowchart of one embodiment of the method is shown, including the following steps:

[0046] Step S10: Obtain the sequencing sequence containing both M and S sequences from the original alignment file, and record the first alignment information of the M sequence in the sequencing sequence to the human reference genome.

[0047] Among them, the sequencing sequence is called Reads, which are short read sequences generated by sequencing.

[0048] Sequencing reads were aligned to the human reference genome (GRCh37) using bwa software (Burrow-Wheeler Aligner, a software for mapping DNA sequences to a reference genome). The original alignment file contains alignment information for each pair of reads within GRCh37 (Human Reference Genome Version 37), primarily considering chromosomal information and alignment location information. See [link to relevant documentation]. Figure 2 As shown.

[0049] See Figure 2 A pair of arrows represents a pair of reads that originate from the same DNA template and have a certain length. For example, each of these reads is 100 base pairs long. The bar below the arrows represents a segment of the sequence in one of the chromosomes of GRCh37. This chromosome is the alignment chromosome for this pair of reads. When this pair of reads matches this sequence, the position of this pair of reads in GRCh37 is determined, and the alignment chromosome number and actual location information are specified.

[0050] If the sample contains fused genomes, some reads will only have partial sequence alignment information during the first GRCh37 alignment, as follows: Figure 3 As shown.

[0051] Figure 3 The bars in the image represent sequences of a region on one chromosome of GRCh37, and the arrows represent individual reads. The black portion of the arrow aligns to GRCh37 and is denoted as the M sequence (the base sequence that correctly aligns to the human reference genome in the first alignment result of the entire read). The chromosome number corresponding to the last base of the M sequence in GRCh37 is denoted as chrbk1 (named for recording the chromosome number of the M sequence alignment), and the alignment position is denoted as POSend (named for recording the position information of the M sequence alignment). The white dashed portion of the arrow does not align correctly to this position and is denoted as the S sequence (the base sequence that did not correctly align to the reference genome in the entire read). This S sequence needs to be aligned to GRCh37 again to obtain the corresponding alignment information. The chromosome number corresponding to the first base of the S sequence in GRCh37 is denoted as chrbk2 (named for recording the chromosome number of the S sequence alignment), and the alignment position is denoted as POSstart (named for recording the position information of the S sequence alignment), as shown below. Figure 4 As shown.

[0052] Figure 4 The arrow on the left is divided into two parts: gray represents the M sequence and white represents the S sequence. The S sequence did not find the corresponding alignment position in the first alignment. Figure 4The white arrow on the right represents the S sequence on the left, and the gray bar indicates a different genomic location on a different chromosome or on the same chromosome than the bar on the left.

[0053] Through the second alignment, the S sequence represented by the white dashed line was able to be aligned to GRCh37, and the corresponding alignment chromosome and position information were obtained.

[0054] The presence of fusion genes in a sample can also be determined by comparing a pair of reads. In a pair of reads, the first read aligns to a location on GRCh37, while the other read aligns to a different location on the same chromosome as the first read, or to a location on a different chromosome, as shown below. Figure 5 As shown.

[0055] Figure 5 Arrows represent a pair of reads, and bars represent sequences of different regions on different chromosomes of GRCh37 or sequences of different locations on the same chromosome. Figure 5 It was found that each of the two Reads obtained a different position of GRCh37.

[0056] In this embodiment, the sequencing sequence that simultaneously contains both M and S sequences is specifically as follows: Figure 4 and Figure 5 The two types shown are (one Read or a pair of Reads).

[0057] After obtaining a sequencing sequence that contains both M and S sequences, the first alignment information of the M sequence can be further obtained from the original alignment file. The first alignment information includes the first chromosome information and the first position information.

[0058] Specifically, the first chromosome information is the first aligned chromosome of the M sequence terminal base in the human reference genome, denoted as chrbk1; the first position information is the alignment position of the M sequence terminal base in the first aligned chromosome, denoted as the terminal base position, and represented by POSend.

[0059] Step S20: Extract the S sequence from the sequencing sequence and record the second alignment information of the S sequence to the human reference genome.

[0060] The S sequence is the base sequence in the entire reads that failed to match the human reference genome correctly in the first alignment result. This part is extracted and used for the second alignment.

[0061] The S sequence was aligned to the human reference genome using alignment software, and the second alignment information was recorded.

[0062] The second alignment information includes second chromosome information and second position information. Specifically, the second chromosome information is the second alignment chromosome of the S sequence start base in the human reference genome, denoted as chrbk2; the second position information is the alignment position of the S sequence start base in the second alignment chromosome, denoted as the start base position, and represented by POSstart.

[0063] Step S30: Obtain the first base sequence based on the first alignment information, and obtain the second base sequence based on the second alignment information.

[0064] In this embodiment, the fusion gene base sequence is constructed based on the alignment information of the M and S sequences in the sequencing reads. Before constructing the fusion gene base sequence, the fusion base sequence needs to be obtained, including the following steps:

[0065] Starting from the end base position, extract the base sequence of the first preset sequence length upstream of the end base position on the first aligned chromosome to obtain the first base sequence.

[0066] Starting from the initial base position, a second base sequence of a predetermined length downstream of the initial base position on the second alignment chromosome is extracted to obtain the second base sequence.

[0067] Specifically, the position of the last base position POSend on the first aligned chromosome chrbk1 is determined, and a first predetermined sequence length upstream of POSend is extracted, for example, a sequence of 300 bases upstream; the position of the first base position POSstart on the second aligned chromosome chrbk2 is determined, and a second predetermined sequence length downstream of POSstart is extracted, for example, a sequence of 300 bases downstream.

[0068] It should be noted that the first preset sequence length and the second preset sequence length can be the same or different, depending on the actual situation.

[0069] Step S40: The first base sequence and the second base sequence are spliced ​​together to obtain the fusion gene base sequence.

[0070] In this embodiment, the splicing position of the first base sequence and the second base sequence is called the breakpoint. The first base sequence is used as the pre-base sequence of the fusion gene base sequence, and the second base sequence is used as the post-base sequence of the fusion gene base sequence. The last base position of the first base sequence and the first base position of the second base sequence are spliced ​​together to obtain the constructed fusion gene base sequence.

[0071] In some optional implementations, the sequence length of the fusion gene base sequence and the breakpoint position are recorded at the specific location of the fusion gene base sequence.

[0072] Step S50: Extract all sequencing sequences from the original alignment file, and align the extracted sequencing sequences with the fusion gene base sequence to obtain the alignment results.

[0073] In this embodiment, all sequencing sequences, i.e. all Reads, are extracted from the original alignment file. Alignment software is used to align the extracted sequencing sequences to the fusion gene base sequence, and the matching status of each extracted sequencing sequence with the fusion gene base sequence is determined to obtain the matching result. The matching result is the alignment result. The matching result is either that the currently aligned Read matches a certain region of the fusion gene base sequence or that the currently aligned Read does not match the fusion gene base sequence correctly.

[0074] Step S60: Compare and analyze the results to obtain sequencing data of the fusion gene.

[0075] Specifically, the number of sequences that completely match the base sequence of the fusion gene is determined based on the matching results; the number of sequences is used to analyze whether the fusion gene actually exists; if it exists, the sequencing data of the fusion gene is statistically analyzed.

[0076] If a sequencing sequence completely matches the fusion gene's base sequence, it indicates that the fusion gene actually exists. The same fusion gene base sequence can be matched with multiple reads. The more reads that are matched, the greater the likelihood that a fusion gene exists.

[0077] In this embodiment, the reference genome sequence for the constructed fusion gene base sequence is used instead of the normal reference genome sequence, and the fusion gene base sequence is validated. This can maximize the identification of sequences supporting the fusion gene in the sequencing sequence and has higher sensitivity.

[0078] In this embodiment, the fusion gene exists in two forms: one is a single read that perfectly matches the base sequence of the fusion gene, see [link to documentation]. Figure 6 Another type is where each read in a pair matches the fusion gene's base sequence individually; see [link to relevant documentation]. Figure 7 As shown.

[0079] exist Figure 6 In the diagram, arrows represent multiple reads, and bars represent the fusion gene's base sequence. The base sequences of these reads perfectly match the fusion gene's base sequence. Figure 7In the diagram, the arrow represents a pair of Reads, and the bar represents the fusion gene base sequence. Each Read in this pair matches the fusion gene base sequence.

[0080] In some alternative implementations, the above-mentioned methods for analyzing DNA fusion gene sequencing data all involve writing analysis scripts using computer code, installing alignment software, analyzing the actual data after alignment, and finally summarizing the analysis results.

[0081] Based on the above-described method for analyzing DNA fusion gene sequencing data, embodiments of this application also provide the application of this method in the prevention and / or treatment of tumors and / or cancer.

[0082] In one specific implementation, the experimental material selected is a positive cell sample containing the fusion of the EML4 (echinoderm microtubule-associated protein 4) gene and the ALK (anaplastic lymphoma kinase) gene. This cell sample is mixed with other cell samples that do not contain the fused genome in a certain proportion to prepare mixed cell samples with the mutation frequencies of the EML4-ALK fusion gene of 20.0%, 10.0%, 5.0%, 1.0%, 0.5%, and 0.1%, respectively.

[0083] These mixed cell samples were processed and sequenced, resulting in five batches of reads. Following the analysis method described above for DNA fusion gene sequencing data, the reads were aligned with GRCh37, and analysis scripts were used to analyze these five batches of reads. Simultaneously, FACTERA software was used to analyze these data, and the results of the two analyses were compared. The analysis results are shown in Table 1.

[0084] Table 1 Analysis Results

[0085]

[0086] As shown in the table above, after comparing the analysis method of the DNA fusion gene sequencing data of this application with the analysis results of FACTERA software, it was found that FACTERA software could not reliably detect positive results of fusion genes in mixed cell samples with a mutation frequency of less than 0.5%, while this application was able to detect positive results of fusion genes with a mutation frequency of 0.1%.

[0087] This demonstrates that the DNA fusion gene sequencing data analysis method provided in this application can identify sequences supporting fusion genes to the greatest extent possible in the sequencing sequence, exhibiting higher sensitivity.

[0088] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. A method of analyzing DNA fusion gene sequencing data for non-disease diagnosis and treatment purposes, characterized by, Includes the following steps: Sequencing sequences containing both M and S sequences are obtained from the original alignment file, and the first alignment information of the M sequence in the sequencing sequence to the human reference genome is recorded. The M sequence is the base sequence in the first alignment result that correctly aligns to the human reference genome, and the S sequence is the base sequence in the entire sequencing sequence that does not correctly align to the reference genome. The first alignment information includes first chromosome information and first position information. The first chromosome information is the first aligned chromosome of the M sequence's terminal base in the human reference genome; the first position information is the alignment position of the M sequence's terminal base in the first aligned chromosome, denoted as the terminal base position. The S sequence is extracted from the sequencing sequence, and the second alignment information of the S sequence to the human reference genome is recorded. The second alignment information includes second chromosome information and second position information. The second chromosome information is the second alignment chromosome of the S sequence start base in the human reference genome. The second position information is the alignment position of the S sequence start base in the second alignment chromosome, denoted as the start base position. A first base sequence is obtained based on the first alignment information, and a second base sequence is obtained based on the second alignment information. The step of obtaining the first base sequence based on the first alignment information includes extracting a base sequence of a first predetermined sequence length upstream of the last base position on the first aligned chromosome, starting from the last base position, to obtain the first base sequence. The step of obtaining the second base sequence based on the second alignment information includes extracting a base sequence of a second predetermined sequence length downstream of the first base position on the second aligned chromosome, starting from the first base position, to obtain the second base sequence. The first base sequence and the second base sequence are spliced ​​together to obtain the fusion gene base sequence; the step of splicing the first base sequence and the second base sequence includes: splicing the last base position of the first base sequence and the first base position of the second base sequence; Extract all sequencing sequences from the original alignment file, and align the extracted sequencing sequences with the fusion gene base sequence to obtain the alignment results; The alignment results were analyzed to obtain the sequencing data of the fusion gene.

2. The analysis method according to claim 1, characterized in that, The step of comparing the extracted sequencing sequence with the fusion gene base sequence to obtain the comparison result includes: The extracted sequencing sequence was aligned to the fusion gene base sequence using alignment software; The matching results are obtained by determining the matching status of each extracted sequencing sequence with the base sequence of the fusion gene.

3. The analysis method according to claim 2, characterized in that, The steps for analyzing the comparison results include: The number of sequences whose extracted sequencing sequences completely match the fusion gene base sequence is determined based on the matching results; Analyze the number of sequences to determine whether the fusion gene actually exists; If present, the sequencing data of the fusion gene are counted.

Citation Information

Patent Citations

  • Method and device for detection of gene fusion

    CN107480472A