Method, system, equipment and medium for searching base mutation for third-generation full-length transcript sequencing data

By using dynamic programming algorithms and gene annotation information from third-generation full-length transcript sequencing data, the problem of base mutations not being accurately mapped to transcripts in existing technologies has been solved. This enables accurate detection of base mutations and analysis of co-occurrence relationships, improving the accuracy and sensitivity of applications such as neoantigen prediction.

CN120913642APending Publication Date: 2025-11-07BEIJING VIEWSOLIDBIOTECH

Patent Information

Application Number
CN202511437943.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies cannot clearly map base mutations to specific full-length transcripts in second-generation short-read sequencing data. This makes it impossible to determine whether a mutation occurs in a specific transcript and makes it difficult to resolve the co-occurrence relationship of multiple mutations in the same transcript, affecting the accuracy and scope of applications such as neoantigen prediction.

Method used

Using third-generation full-length transcript sequencing data, sample data was acquired and compared with reference genome sequences. A dynamic programming algorithm was used to identify base level differences. Combined with gene annotation information and structural variation analysis, overlapping sequencing reads were screened to achieve accurate detection of base mutations.

Benefits of technology

It achieves precise mapping between base mutations and specific transcripts, solves the problem of co-occurrence analysis of multiple mutations on the same transcript, significantly improves the analysis accuracy and detection sensitivity in complex biological scenarios, and avoids computational redundancy and functional misjudgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913642A_ABST
    Figure CN120913642A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of bioinformatics, and discloses a method, a system, equipment and a medium for searching for base mutation aiming at third-generation full-length transcript sequencing data, the base mutation is accurately positioned to a specific transcript by directly processing the third-generation full-length transcript sequencing data, and the method and the system for searching for the base mutation aiming at the third-generation full-length transcript sequencing data are provided. The co-occurrence relation of a plurality of mutations on the same transcript is accurately analyzed, and the defects of calculation redundancy, function misjudgment and the like caused by the fact that a mutation transcript source cannot be determined due to fragmentation splicing and the co-occurrence is simulated by depending on permutation and combination in a second-generation short-read-long sequencing technology are effectively overcome; the recall rate of transcripts which are not mapped due to a high-variation region is improved through mapping correction assisted by structural variation, and false positive is effectively inhibited through a multi-dimensional filtering condition, so that the sensitivity and reliability of mutation detection are remarkably improved; particularly, mutation events with function remodeling due to reading frame change can be accurately recognized in scenes such as neoantigen prediction where mutation function consequences need to be accurately evaluated, and the method has important value in application in the fields of precision medical treatment and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics, and in particular to a method, system, device and medium for finding base mutations based on third-generation full-length transcript sequencing data. BACKGROUND

[0002] At present, the mainstream base mutation detection method is developed based on second-generation short read sequencing data, such as a variant identification process relying on a deep learning model (such as DeepVariant) or a Bayesian statistical method (such as GATK, FreeBayes). Although this method performs well in identifying whether a sample carries a mutation, it has a fundamental limitation: due to the short read length, the need for fragmentation sequencing and reassembly, it cannot explicitly map the mutation to a specific full-length transcript, so it cannot determine whether a mutation actually occurs in a specific transcript, nor can it analyze the co-occurrence relationship of multiple mutations in the same transcript. In actual applications, such as new antigen prediction scenarios, it is necessary to accurately understand the distribution of mutation combinations on real transcripts. The traditional method lacks mutation association information at the single transcript level, so it can only rely on permutation and combination to simulate possible mutation co-occurrence patterns, which not only introduces a large amount of redundant calculation, but also may lead to functional misjudgment. For example, a synonymous mutation should be filtered under normal circumstances, but if there is a frameshift mutation in the same transcript that changes the reading frame, the synonymous mutation may actually participate in the coding change, thereby affecting the generation of new antigens. The existing method cannot effectively capture the functional changes of mutations in such complex transcript backgrounds, limiting its accuracy and application breadth in full-length transcript sequencing data. SUMMARY

[0003] The present application provides a method, system, device and medium for finding base mutations based on third-generation full-length transcript sequencing data to solve the defects of the prior art.

[0004] The present application provides a method for finding base mutations based on third-generation full-length transcript sequencing data, comprising: Obtaining third-generation full-length transcript sequencing data of a sample to be tested; Mapping and aligning the third-generation full-length transcript sequencing data of the sample to be tested with a reference genome sequence, and annotating the mapping results based on gene annotation information to obtain mapping information and annotation information of each sequencing read; Based on the mapping information and annotation information of each sequencing read, screening out sequencing reads with overlapping mapping positions and reference exon regions; Using a dynamic programming algorithm, sequence aligning the screened sequencing reads with the corresponding reference exons to realize base mutation detection by identifying base level differences.

[0005] According to the method for finding base mutations for third-generation full-length transcript sequencing data provided by the application, the third-generation full-length transcript sequencing data of a sample to be detected is mapped to a reference genome sequence, and the mapping result is annotated based on gene annotation information to obtain mapping information and annotation information of each sequencing read, comprising: The sequencing read that is not successfully mapped is compared with a deletion sequence in the structural variation analysis result in terms of similarity, and if the similarity reaches a predetermined threshold, the sequencing read is determined to be a transcript fragment that is not mapped due to variation, and the mapping information and the annotation information are updated accordingly.

[0006] According to the method for finding base mutations for third-generation full-length transcript sequencing data provided by the application, the screened sequencing read is subjected to sequence alignment with the corresponding reference exon by using a dynamic programming algorithm, base mutation detection is realized by identifying base level differences, comprising: The screened sequencing read is subjected to sequence alignment with the corresponding reference exon by using a dynamic programming algorithm, base differences are identified according to a preset identification condition, and base mutation analysis results of the sample to be detected are obtained.

[0007] According to the method for finding base mutations for third-generation full-length transcript sequencing data provided by the application, the preset identification condition comprises any one or any combination of the following: If the base difference is the replacement of a single base, it is identified as a single nucleotide variation; If the base difference is the addition of one or more continuous bases in the actual sequence relative to the reference sequence, it is identified as an insertion variation; If the base difference is the loss of one or more continuous bases in the actual sequence relative to the reference sequence, it is identified as a deletion variation.

[0008] According to the method for finding base mutations for third-generation full-length transcript sequencing data provided by the application, the screened sequencing read is subjected to sequence alignment with the corresponding reference exon by using a dynamic programming algorithm, base mutation detection is realized by identifying base level differences, comprising: The base mutation analysis results of the sample to be detected are filtered according to a preset screening condition to obtain the final base mutation analysis results of the sample to be detected.

[0009] According to the method for finding base mutations for third-generation full-length transcript sequencing data provided by the application, the preset screening condition comprises any one or any combination of the following: The base mutations that exist only in a single transcript are removed; The base mutations in the exons containing more than 5 indels (indels) are excluded; The base mutations with a mutation frequency lower than 0.1 are removed; Filtering base mutations whose total TPM values of transcripts (i.e. the sum of TPM values of all transcripts containing the base mutation, and TPM value represents the number of transcripts per million) are less than 5.

[0010] According to the method for finding base mutations provided by the application, the mutation frequency is the number of transcripts containing the base mutation / the number of all transcripts passing through the site, and here, filtering is performed by considering the sequencing error, and therefore, the TPM value is not calculated.

[0011] The application further provides a system for finding base mutations based on third-generation full-length transcript sequencing data, comprising: a data receiving module, configured to receive third-generation full-length transcript sequencing data of a sample to be tested; a mapping module, configured to align and map the third-generation full-length transcript sequencing data of the sample to be tested with a reference genome sequence, and annotate the mapping result based on gene annotation information to obtain mapping information and annotation information of each sequencing read; a screening module, configured to screen sequencing reads having overlapping mapping positions with reference exon regions based on the mapping information and the annotation information of each sequencing read; a detection module, configured to perform sequence alignment between the screened sequencing reads and corresponding reference exons by using a dynamic programming algorithm, and realize base mutation detection by identifying base level differences.

[0012] The application further provides an electronic device comprising a processor and a memory storing a computer program, and the processor implements the method for finding base mutations based on third-generation full-length transcript sequencing data according to any one of the above-mentioned methods when executing the computer program.

[0013] The application further provides a non-transitory computer readable storage medium storing a computer program, and the computer program is executed by a processor to implement the method for finding base mutations based on third-generation full-length transcript sequencing data according to any one of the above-mentioned methods.

[0014] The application further provides a computer program product comprising a computer program, and the computer program can be stored on a non-transitory computer readable storage medium, and the computer program is executed by a processor to implement the method for finding base mutations based on third-generation full-length transcript sequencing data according to any one of the above-mentioned methods.

[0015] The method, system, device and medium for finding base mutations based on third-generation full-length transcript sequencing data provided by the application can at least bring the following beneficial effects: (1) Realize the accurate mapping of base mutations and specific transcripts: The present application can overcome the fundamental limitations of traditional methods based on second-generation short-read sequencing (such as GATK, DeepVariant, etc.). By directly processing third-generation full-length transcript sequencing data and associating mutation detection results with each independent sequencing read, the present application can accurately locate each base mutation (including SNV, insertion INS and deletion DEL) to the specific transcript from which it originates, thereby explicitly revealing the direct correspondence between mutations and transcript isoforms.

[0016] (2) Solve the co-occurrence analysis problem of multiple mutations on the same transcript: Traditional methods cannot determine whether multiple mutations occur on the same RNA molecule, so they can only rely on computationally redundant and inaccurate permutations and combinations for simulation. The present application can directly and accurately determine whether multiple base mutations occur on the same transcript by analyzing a single full-length transcript sequence, completely avoiding computational redundancy and functional misjudgment caused by combinatorial explosion, and providing a reliable data foundation for studying the synergistic effects of mutations.

[0017] (3) Significantly improve the analysis accuracy in complex biological scenarios: In particular, in applications such as neoantigen prediction, which highly depend on accurate interpretation of the functional consequences of mutations, the advantages of the present application are particularly prominent. For example, it can effectively identify key scenarios where a pre-existing frameshift mutation converts a previously synonymous mutation into a non-synonymous mutation, thereby avoiding missing important functional mutations. This depth of analysis of the context of the transcript is not available in existing methods.

[0018] (4) Enhance the sensitivity and reliability of detection: By introducing a mapping correction step based on structural variation (SV) analysis, the present application can effectively recall transcript fragments that were not initially mapped due to high variability (such as inversion INV or large fragment deletion), significantly improving the sensitivity of mutation detection. At the same time, by setting multi-dimensional filtering conditions (such as mutation frequency, transcript support number, TPM value, etc.), while ensuring high positive rate, false positive signals caused by sequencing errors and other factors are effectively eliminated, ensuring the reliability of the detection results.

[0019] (5) Open up new application paths for third-generation sequencing data in precise mutation analysis: The present application provides a complete and efficient base mutation detection solution designed specifically for third-generation full-length transcript data, filling a gap in the field. This method ensures that the base mutation profile of a sample can be independently, accurately and comprehensively analyzed even with only third-generation sequencing data, which has important value for promoting the in-depth application of third-generation sequencing technology in genetic disease research, precision medicine for cancer and other fields. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the present application or the prior art, the accompanying drawings required by the embodiments or the prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0021] Figure 1 A flowchart of a method for finding base mutations for third-generation full-length transcript sequencing data provided by the present application.

[0022] Figure 2 A structural diagram of a system for finding base mutations for third-generation full-length transcript sequencing data provided by the present application.

[0023] Figure 3 A structural diagram of an electronic device provided by the present application. DETAILED DESCRIPTION

[0024] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in combination with the accompanying drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments, and they should not be understood as limiting the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application. In the description of the present application, it should be understood that the terms used are only for the purpose of description, and should not be understood as indicating or implying relative importance.

[0025] In the application of full-length transcriptome sequencing (Iso-Seq) based on single molecule real-time sequencing (SMRT) technology, complete and non-fragmented mRNA sequence information can be directly obtained, which provides a powerful tool for studying gene isoform expression, alternative splicing and transcript structure variation. This technology observes the synthesis process of DNA polymerase through zero-mode waveguide holes (ZMWs), and the long-read sequences obtained are helpful to accurately identify transcript isoforms and the genetic variations carried by them, especially in the study of base mutations. As a basic type of genetic variation, base mutations are widely involved in multiple biological processes such as disease occurrence, evolutionary adaptation and biological diversity formation, and their accurate detection has key value for medical diagnosis, molecular breeding and functional genomics research.

[0026] Figure 1A flowchart of a method for finding base mutations from third-generation full-length transcript sequencing data according to the present application is shown. The execution subject of the method for finding base mutations from third-generation full-length transcript sequencing data according to the present application can be any applicable terminal-side device or network-side device, such as a device for finding base mutations from third-generation full-length transcript sequencing data, etc.

[0027] Referring to Figure 1 The method for finding base mutations from third-generation full-length transcript sequencing data according to the present application can include: S110, obtaining third-generation full-length transcript sequencing data of a sample to be tested.

[0028] S120, aligning and mapping the third-generation full-length transcript sequencing data of the sample to be tested with a reference genome sequence, and annotating the mapping results based on gene annotation information to obtain mapping information and annotation information of each sequencing read.

[0029] "Third-generation full-length transcript sequencing data" is a general term and a collection, referring to an entire FASTQ / FASTA file, containing millions of sequences. "Sequencing read" is a single member and basic unit in this collection, referring to each sequence in the file.

[0030] The mapping information contains the chromosome and position information of the measured transcript on the reference transcript; the annotation information includes the gene name to which the measured transcript belongs, whether it is a known transcript, the transcript FLC (Full-Length Transcript), the transcript TPM value, etc.

[0031] S130, performing similarity comparison between the sequencing reads that have not been successfully mapped and the deletion sequences in the structural variation analysis results, and if the similarity reaches a predetermined threshold, determining that the sequencing reads are transcript fragments that have not been mapped due to variations, and updating the mapping information and the annotation information accordingly.

[0032] S140, screening sequencing reads with overlapping mapping positions and reference exon regions based on the mapping information and the annotation information of each sequencing read.

[0033] S150, using a dynamic programming algorithm to perform sequence alignment between the screened sequencing reads and the corresponding reference exons, identifying base differences according to a preset identification condition, and obtaining base mutation analysis results of the sample to be tested, wherein the preset identification condition includes any one or any combination of the following: If the base difference is a single base substitution, it is identified as a single nucleotide variation; If the base difference is the addition of one or more consecutive bases in the measured sequence relative to the reference sequence, it is identified as an insertion variation; If the base difference is the loss of one or more consecutive bases in the actual sequence relative to the reference sequence, it is identified as a deletion mutation.

[0034] S160, filtering the base mutation analysis result of the to-be-tested sample according to a preset screening condition to obtain the final base mutation analysis result of the to-be-tested sample, wherein the preset screening condition comprises any one or any combination of the following: eliminating the base mutation existing only in a single transcript; excluding the base mutation in which the containing exon has more than 5 indels (indels); removing the base mutation with a mutation frequency lower than 0.1, the mutation frequency = the number of transcripts containing the base mutation / the number of all transcripts passing the site, and here the filtering considers the sequencing error problem, so the TPM value is not calculated; filtering the base mutation with a total TPM value (i.e. the sum of the TPM values of all transcripts containing the base mutation, and the TPM value represents the number of transcripts per million) of the containing transcript less than 5.

[0035] This embodiment uses the above-provided method for finding base mutations in three-generation full-length transcript sequencing data to find base mutations in eight groups of three-generation full-length transcripts and compares the results with existing software, and the steps include: S1, using a third-generation DNA sequencing technology platform (such as a Pacbio platform) to sequence eight to-be-tested samples to obtain three-generation full-length transcript sequencing data corresponding to the eight to-be-tested samples.

[0036] S2, renaming the three-generation full-length transcript sequencing data corresponding to the eight to-be-tested samples, with the naming form being Sample1, Sample2, …, Sample7, Sample8, and then using the TAGET software (a bioinformatics software, the core function of which is to identify and quantify transcript isoforms from raw sequencing data) to map the sequencing sequences to the reference genome sequence and annotate the transcripts to obtain the mapping information and annotation information of each sequencing read.

[0037] S3, using the SV module of the NeoTAGET software (a bioinformatics software) to compare the similarity between the unmapped fragments in the actual transcript and the missing exons in the SV, supplementing the unmapped transcripts due to base mutations or structural variation INV, and updating the mapping information and annotation information of each sequencing read accordingly.

[0038] S4, based on the updated mapping information and annotation information of each sequencing read, retaining the sequencing reads with overlapping mapping positions and reference exon regions.

[0039] S5, adopt dynamic programming algorithm to carry out sequence alignment to actual exon and corresponding reference exon, realize accurate detection of base mutation by recognizing base level difference.

[0040] S6, screening the obtained base mutation, eliminating the base mutation existing only in a single transcript, excluding the base mutation of the exon containing more than 5 insertions and deletions, removing the base mutation with a mutation frequency lower than 0.1, filtering the base mutation with a total TPM value of the transcript less than 5, and outputting a result information file.

[0041] S7, using clair3 software (an existing base mutation detection software, denoted as C in table 1) and deepvariant software (an existing base mutation detection software, denoted as D in table 1) to detect the base mutations of the eight test samples, taking the base mutation results of the second-generation RNAseq data as the standard, comparing the number of base mutations found by the embodiment (denoted as N in table 1) and the other two software, and the number of second-generation support, and the comparison results are shown in table 1.

[0042]

[0043] The above results show that the method for finding base mutations provided by the present application based on the third-generation full-length transcript sequencing data can detect new variable splicing of the third-generation full-length transcript, and has the following advantages compared with the existing software: (1) high base mutation positive rate For the eight test samples of the above embodiment, the base mutations detected by the method provided by the present application are at a high level in terms of the number of real mutations and the positive rate compared with the same type of software; (2) can determine the conjugation of multiple base mutations of the same transcript, and play a reference role for new antigen research The method provided by the present application can effectively obtain the combination of multiple base mutations in the actual sample through the third-generation full-length transcript sequencing data, and determine the corresponding reference transcript, effectively reducing the problem of whether multiple base mutations exist simultaneously in the translation of the transcript into amino acid, and playing a good reference role for the current cancer new antigen research.

[0044] The system for finding base mutations based on the third-generation full-length transcript sequencing data provided by the present application is described below, and the system for finding base mutations based on the third-generation full-length transcript sequencing data described below can be correspondingly referred to the method for finding base mutations based on the third-generation full-length transcript sequencing data described above.

[0045] Referring to Figure 2 The system for finding base mutations based on the third-generation full-length transcript sequencing data provided by the present application can include: The data receiving module is used to receive third-generation full-length transcript sequencing data of the sample to be tested; The mapping module is used to: align and map the third-generation full-length transcript sequencing data of the sample to be tested with the reference genome sequence, and annotate the mapping results based on gene annotation information to obtain the mapping information and annotation information for each sequencing read; The filtering module is used to: filter out sequencing reads whose mapping positions overlap with the reference exon regions based on the mapping and annotation information of each sequencing read; The detection module is used to: use a dynamic programming algorithm to align the selected sequencing reads with the corresponding reference exons, and detect base mutations by identifying differences in base levels.

[0046] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute the following steps: Obtain third-generation full-length transcript sequencing data of the sample to be tested; The third-generation full-length transcript sequencing data of the sample to be tested were compared and mapped with the reference genome sequence, and the mapping results were annotated based on gene annotation information to obtain the mapping information and annotation information of each sequencing read. Based on the mapping and annotation information of each sequencing read, sequencing reads whose mapping positions overlap with the reference exon region are selected. Using a dynamic programming algorithm, the selected sequencing reads are compared with the corresponding reference exons, and base mutation detection is achieved by identifying differences at the base level.

[0047] Further, the logic instructions in the memory 830 described above can be implemented in the form of software functional units and sold or used as standalone products, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0048] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to perform the following steps: obtain third-generation full-length transcript sequencing data of a sample to be tested; align and map the third-generation full-length transcript sequencing data of the sample to be tested with a reference genome sequence, and annotate the mapping results based on gene annotation information to obtain mapping information and annotation information of each sequencing read; based on the mapping information and the annotation information of each sequencing read, screening out sequencing reads with overlapping mapping positions and reference exon regions; using a dynamic programming algorithm, performing sequence alignment between the screened sequencing reads and corresponding reference exons, and realizing base mutation detection by identifying base level differences.

[0049] In another aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the following steps: obtain third-generation full-length transcript sequencing data of a sample to be tested; align and map the third-generation full-length transcript sequencing data of the sample to be tested with a reference genome sequence, and annotate the mapping results based on gene annotation information to obtain mapping information and annotation information of each sequencing read; based on the mapping information and the annotation information of each sequencing read, screening out sequencing reads with overlapping mapping positions and reference exon regions; using a dynamic programming algorithm, performing sequence alignment between the screened sequencing reads and corresponding reference exons, and realizing base mutation detection by identifying base level differences.

[0050] The apparatus embodiments described above are merely illustrative, wherein the units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0051] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0052] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for finding base mutations against third generation full-length transcript sequencing data, characterized in that, The method comprises the following steps: obtaining the third-generation full-length transcript sequencing data of the sample to be tested; aligning and mapping the third-generation full-length transcript sequencing data of the sample to be tested with the reference genome sequence, and annotating the mapping results based on gene annotation information to obtain the mapping information and annotation information of each sequencing read; based on the mapping information and annotation information of each sequencing read, screening out sequencing reads with overlapping mapping positions and reference exon regions; using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences.

2. The method for finding base mutations for third generation full-length transcript sequencing data according to claim 1, wherein, The method comprises the following steps: aligning and mapping the third-generation full-length transcript sequencing data of the sample to be tested with the reference genome sequence, and annotating the mapping results based on gene annotation information to obtain the mapping information and annotation information of each sequencing read.

3. The method for finding base mutations for third generation full-length transcript sequencing data of claim 1, wherein, The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences.

4. The method for finding base mutations for third generation full-length transcript sequencing data according to claim 3, wherein, The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences.

5. The method of finding base mutations for third generation full-length transcript sequencing data according to claim 4, wherein, The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences.

6. The method of finding base mutations for third generation full-length transcript sequencing data according to claim 5, wherein, The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps:

7. The method of finding base mutations for third generation full-length transcript sequencing data according to claim 6, wherein, using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps:

8. A system for finding base mutations against third generation full-length transcript sequencing data, characterized in that, using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and realizing base mutation detection by identifying base level differences. The method comprises the following steps: using a dynamic programming algorithm to perform sequence alignment on the screened sequencing reads and the corresponding reference exons, and The screening module is configured to screen, based on the mapping information and the annotation information of each sequencing read, sequencing reads with a mapping position overlapping a reference exon region; The detection module is configured to perform sequence alignment between the screened sequencing reads and corresponding reference exons by using a dynamic programming algorithm, and to realize base mutation detection by identifying base level differences.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the method for finding base mutations for third-generation full-length transcript sequencing data according to any one of claims 1 to 7 when executing the program.

10. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program implements the method for finding base mutations for third-generation full-length transcript sequencing data according to any one of claims 1 to 7 when executed by the processor.

Citation Information

Patent Citations

  • Method and system for detecting structural variation in third-generation whole exon sequencing transcript data

    CN117789820A

  • Gene structure variation detection method and device

    CN117935908A

  • Method and apparatus for training machine learning model for removing noise in data

    US20250239326A1

  • Method and computer program product for detecting mutation in a nucleotide sequence

    WO2014041380A1

Cited By

  • Method for searching novel variable splicing event aiming at third-generation full-length transcript sequencing data

    CN121506243A