Analysis method for HBV expression level and fragment distribution based on single-cell RNA-seq

By constructing the location information of HBV genes in the transcriptome and genome in single-cell RNA sequencing data, and comparing the mapping relationship between the transcriptome and genome for genome, the accuracy of HBV expression level and fragment distribution analysis in the prior art was solved, and more accurate quantification of HBV and precise positioning of breakpoint positions were achieved.

CN120015118AActive Publication Date: 2025-05-16WEST CHINA HOSPITAL SICHUAN UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510095005.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The prior art is difficult to accurately analyze the expression levels and fragment distribution of hepatitis B virus (HBV) in different individuals and cell types through single-cell RNA sequencing, especially after HBV is integrated into the host genome.

Method used

Using a single-cell RNA-seq analysis method, the location information of HBV genes in the transcriptome and genome is constructed, and the sequencing data is aligned to the linear transcriptome, and the mapping relationship between the transcriptome and genome is used to indirectly aligned to the HBV genome, thereby accurately analyzing the expression level and fragment distribution of HBV.

Benefits of technology

A more accurate quantification of HBV expression levels and fragment distribution is achieved, and the expression of overlapping genes and different transcripts can be distinguished. It provides accurate positioning of HBV integration breakpoints, filling the gap in HBV analysis methods based on single-cell resolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015118A_ABST
    Figure CN120015118A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biological industry, and discloses a single-cell RNA-seq-based HBV expression level and fragment distribution analysis method, which comprises the following steps: constructing position information of each HBV gene in a linear transcriptome and a circular genome to obtain a mapping relationship between the transcriptome and the genome; sequencing original data obtained by scRNAseq is compared to an HBV transcriptome reference sequence, and a human source sequence and an HBV source sequence are recognized; removing a human source sequence to obtain an HBV source transcriptome sequence; converting the position of the HBV source transcriptome sequence aligned on the transcriptome into an alignment position on the genome according to the mapping relation between the transcriptome and the genome; and calculating the number of the sequencing sequences covered on each basic group of the genome, and drawing the coverage depth of each basic group of the HBV genome to obtain the expression level and fragment distribution of the single-cell RNA-seq sequencing data on the HBV genome. The blank of the HBV analysis method based on the single cell resolution is filled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioindustry, and in particular to an analysis method for HBV expression level and fragment distribution based on single-cell RNA-seq. Background Art

[0002] Gene expression refers to the process of transcribing gene fragments from DNA into messenger RNA (mRNA) and translating RNA into proteins. This process is a key step in implementing genetic instructions in living organisms. Digital gene expression profiling (DGE) and transcriptome analysis (RNA-Seq) are new methods that use a new generation of high-throughput sequencing technology and high-performance computing analysis technology to capture sequences and accurately analyze gene expression in specific tissues and states of a species.

[0003] Single-cell RNA-seq (scRNAseq) technology is a high-throughput sequencing analysis technology for the transcriptome level of a single cell, which can detect the gene expression level in each cell. This technology mainly captures the 3' region of RNA for library construction and sequencing, and the sequenced data mainly covers the 3' or 5' region of the gene.

[0004] Single-cell RNA sequencing (scRNA-seq) technology provides unprecedented resolution and insights into biological research. Here are some of the main advantages of scRNA-seq technology:

[0005] 1. High resolution: scRNA-seq can analyze gene expression at the level of single cells, which allows researchers to identify subtle differences between cells and even discover new cell types or subpopulations.

[0006] 2. Revealing heterogeneity: Traditional large-scale cell population sequencing methods can only provide averaged data, masking the significant differences that may exist between cells. scRNA-seq can show the heterogeneity within the cell population in detail, which is particularly important for understanding complex tissue structure and function.

[0007] 3. Explore dynamic changes: By sampling the same type of cells at different time points, scRNA-seq can track the changes in gene expression, which helps to understand the dynamic changes in cell development, differentiation and response to external stimuli.

[0008] 4. Low input requirements: Even a few cells isolated from a small sample can be analyzed using scRNA-seq, making it an ideal tool for studying rare cell types or samples that are difficult to obtain.

[0009] 5. Multidimensional information: In addition to gene expression levels, some advanced scRNA-seq methods can also provide other types of molecular information, such as alternative splicing, allele-specific expression, and the function of non-coding RNA.

[0010] 6. Disease research and personalized medicine: scRNA-seq helps to deeply understand the changes at the cellular level in disease states, provide a basis for the development of new diagnostic markers and treatment strategies, and promote the development of personalized medicine.

[0011] 7. Drug response assessment: This technology can be used to evaluate the response of individual cells to drugs, help predict drug efficacy and reduce side effects, thereby optimizing treatment plans.

[0012] However, there are significant limitations in analyzing the expression level and fragment distribution of hepatitis B virus (HBV) using single-cell RNA sequencing (scRNA-seq) technology, mainly for the following reasons:

[0013] 1. Integration expression mechanism: After HBV infects cells, its genome can be integrated into the host's human genome and expressed as the host's genes are transcribed. This integration expression pattern means that the expression level of HBV and its fragment distribution may vary significantly between different individuals and different cell types. In order to better understand the mechanism of action of HBV, these differences need to be accurately measured. However, scRNA-seq technology is mainly used to detect mRNA expression in a non-integrated state, and is not good at identifying and quantifying specific changes caused by viral genome integration.

[0014] 2. Low abundance and background noise: Due to the relatively low content of HBV in infected cells and the background noise generated by the large amount of mRNA in host cells, it is difficult to effectively capture and quantify HBV RNA using scRNA-seq. This makes it difficult to accurately assess the expression level of HBV.

[0015] 3. Lack of 3' polyA tail: HBV mRNA lacks a 3' polyA tail structure, which is a key issue because most scRNA-seq library construction methods rely on the polyA tail to select mRNA for sequencing. This means that HBV mRNA is not easily captured, resulting in extremely low levels of HBV-related sequences in the final sequencing data, further limiting the accurate measurement of its expression level.

[0016] 4. Circular genome alignment problem: The HBV genome is a circular genome, and existing bioinformatics tools are mainly designed for linear genomes and used to align and analyze sequencing data. This leads to inaccurate or incomplete matches when trying to map the data obtained from HBV back to the reference genome, thus affecting the understanding of the distribution of HBV genome fragments.

[0017] In view of this, there is currently no HBV analysis method directly derived from single-cell sequencing data. Existing HBV analysis technologies such as Northern Blotting, RT-PCR, RNA-seq (bulk), Reporter Assays, etc., although they can provide important information about HBV expression levels and expression abundance in various regions of the genome, are not based on single-cell resolution data. Although traditional bulk RNA-seq and other molecular biology techniques can provide a holistic view of HBV expression, in some cases, especially when faced with heterogeneous cell populations, these methods may not be sufficient to reveal subtle differences in complex situations. Summary of the invention

[0018] The present invention aims to provide an analysis method for HBV expression level and fragment distribution based on single-cell RNA-seq to solve the current problem of lack of methods for calculating HBV expression level from single-cell sequencing data and locating expression abundance of various regions on the HBV genome, thus filling the gap in HBV analysis methods based on single-cell resolution.

[0019] In order to achieve the above object, the present invention adopts the following technical scheme:

[0020] A method for analyzing HBV expression level and fragment distribution based on single-cell RNA-seq, comprising the following steps:

[0021] Step 1: construct the location information of each HBV gene in the linear transcriptome and circular genome to obtain the mapping relationship between the transcriptome and the genome;

[0022] Step 2: Align the original sequencing data obtained by scRNAseq to the HBV transcriptome reference sequence to determine the breakpoint position when HBV integrates into the human genome and identify the human and HBV source sequences;

[0023] Step 3, removing the human source sequence, retaining the remaining HBV source sequence and the corresponding transcriptome coordinate position starting from the HBV potential breakpoint position, and obtaining the HBV source transcriptome sequence;

[0024] Step 4, using the transcriptome-genome mapping relationship constructed in step 1 to map the HBV source transcriptome sequence obtained in step 3, converting the position of the sequencing sequence alignment on the transcriptome into the alignment position on the genome;

[0025] Step 5: Calculate the number of sequencing sequences covered on each base of the genome, plot the coverage depth of each base of the HBV genome, and mark and display the gene position to obtain the expression level and fragment distribution of single-cell RNA-seq sequencing data on the HBV genome.

[0026] Preferably, in step one, the HBV genome reference sequence and the HBV transcriptome reference sequence are first obtained respectively, the annotation information of the HBV genome is obtained, the position information of each HBV gene in the transcriptome and the genome is constructed, and the mapping relationship between the transcriptome and the genome is obtained; the mapping relationship between the transcriptome and the genome is that the starting position 1376 and the ending position 1840 of the HBV genome NC_003977.2 are the X gene, which is the X protein; the starting position 1816 and the ending position 2454 of NC_003977.2 are the C gene, which is the precapsid protein; the starting position 1903 and the ending position 2454 of NC_003977.2 are the C gene, which is the capsid protein; the starting position 1903 and the ending position 2454 of NC_003977.2 are the C gene, which is the capsid protein; the starting position 1376 and the ending position 1840 of NC_003977.2 are the X gene, which is the X protein; the starting position 1816 and the ending position 2454 of NC_003977.2 are the C gene, which is the precapsid protein; the starting position 1903 and the ending position 2454 of NC_003977.2 are the C gene, which is the capsid protein; the starting position 1376 and the ending position 1840 of NC_003977.2 are the X gene, which is the precap ... 2309 and the end position 3182 are P genes, which are polymerases; NC_003977.2 starting position 1 and the end position 1625 are P genes, which are polymerases; NC_003977.2 starting position 2850 and the end position 3182 are S genes, which are large envelope proteins; NC_003977.2 starting position 1 and the end position 837 are S genes, which are large envelope proteins; NC_003977.2 starting position 3174 and the end position 3182 are S genes, which are middle membrane proteins; NC_003977.2 starting position 1 and the end position 837 are S genes, which are middle membrane proteins; NC_003977.2 starting position 157 and the end position 837 are S genes, which are small envelope proteins.

[0027] Preferably, in step 2, the FASTQ file of the sequencing raw data obtained by scRNAseq is aligned to the HBV transcriptome reference sequence using bwa mem software, and only the sequencing sequence information that can be aligned to the HBV genome is retained to generate a BAM format alignment record file aligned to the HBV genome as the HBV source sequence and the rest are human sequences.

[0028] Preferably, in step three, according to the CIGAR record information of each sequencing sequence in the BAM file, the sequencing sequence with the soft cut structure is extracted, and the soft cut part is cut off, and the remaining sequence and the corresponding transcriptome coordinate position are retained; for the sequencing sequence without the soft cut structure, the sequence and the corresponding transcriptome coordinate position are retained.

[0029] Preferably, in determining the breakpoint position when HBV is integrated into the human genome, the sequence alignment, soft-cut structure distribution and transcriptome coordinate position are comprehensively considered in the following manner:

[0030] According to the sequence comparison, the comparison results of the sequencing sequence and the HBV transcriptome reference sequence are analyzed in detail. If the base matching rate between the sequencing sequence and the reference sequence in a certain region is lower than the preset threshold, and the low matching rate is distributed continuously, and the matching pattern is significantly different from that of the normal comparison region, the region is marked as a potential breakpoint location region; the preset threshold is 70%-80%;

[0031] For the distribution of soft-cut structures, statistical analysis is performed on sequencing sequences with soft-cut structures, and the starting and ending positions of each soft-cut structure in the sequence are counted, and these positions are mapped to transcriptome coordinates; when the frequency of occurrence of the soft-cut structure in a specific transcriptome coordinate interval exceeds a preset frequency threshold, the coordinate interval is marked as a potential breakpoint position interval; the preset frequency threshold is that the number of soft-cut structures that appear in every 1,000 sequencing sequences exceeds 10 times;

[0032] Combined with the transcriptome coordinate position, the transcriptome coordinate continuity of the sequencing sequence is analyzed within the human-HBV genome chimeric region; if the transcriptome coordinates are found to be jumping, interrupted, or inconsistent with the known human genome and HBV genome transcriptome coordinate ranges, the corresponding position is marked as a potential breakpoint position region;

[0033] Finally, the potential breakpoint position regions and potential breakpoint position intervals marked by the above three methods are comprehensively considered. When multiple marked regions overlap with each other or are densely distributed in a small range, the common region or dense region is determined as the breakpoint position when HBV integrates into the human genome.

[0034] Preferably, the dense distribution within a smaller range refers to the presence of 3 or more potential breakpoint position regions marked in different ways within a range with a radius of no more than 50 bp centered on a certain position.

[0035] Preferably, when analyzing sequence alignment to determine potential breakpoint location regions, low match rates present a continuous distribution, which means that the number of continuous bases in the low match rate region exceeds 10.

[0036] Preferably, when determining the potential breakpoint position interval by statistically analyzing the soft-cut structure distribution, the length distribution of the soft-cut structure is also considered; when within a certain transcriptome coordinate interval, not only the frequency of occurrence of the soft-cut structure exceeds the preset frequency threshold, but also the average length of the soft-cut structure is at least 10 bp longer than that of the normal region, and the maximum length is at least 20 bp longer than that of the normal region.

[0037] Preferably, in step 4, the position of the sequencing sequence alignment on the transcriptome is converted into the alignment position on the genome: gene NC_003977.2, start position 1712, end position 1819, nucleotide sequence AAAGACTGTGTGTTTAATGAGTGGGAGGAGTT.

[0038] Preferably, in step 4, the position of the sequencing sequence alignment on the transcriptome is converted into the alignment position on the genome: gene NC 003977.2, starting position 1511, ending position 1633, nucleotide sequence C0GACCGACCACGGGGCGCACCTCTCTTTACG.

[0039] The advantages of the present invention are:

[0040] The present invention solves the problem of the lack of methods for calculating HBV expression levels from single-cell sequencing data and locating the expression abundance of various regions on the HBV genome, filling the gap in HBV analysis methods based on single-cell resolution.

[0041] 1. Improved accuracy: Since the HBV genome is a circular structure, and existing sequence alignment tools are mainly designed for linear genomes, this may lead to quantitative deviations when aligning S and P genes that span the start and end positions. This solution first aligns the sequencing data to the linear transcriptome, and then indirectly aligns it back to the HBV genome based on the mapping relationship between the transcriptome and the genome, so that the expression levels of these key genes can be quantified more accurately.

[0042] 2. Ability to distinguish overlapping genes and multiple transcripts: There are many overlapping genes and multiple transcripts of a gene in the HBV genome, which increases the difficulty of correctly analyzing their respective expression patterns. This solution uses the method of first aligning the data to the transcriptome and then mapping it back to the location of the genome. This strategy helps to more accurately distinguish the specific expression of different overlapping genes and different transcripts of the same gene.

[0043] 3. Providing additional information: Using the chimeric sequencing sequence of the human-HBV genome can not only evaluate the expression of HBV itself, but also be used to identify the exact breakpoint location of HBV integration into the human genome. This method is not limited to monitoring expression levels, but also provides important clues about the mechanism of viral integration, which is of great significance for a deeper understanding of HBV infection, replication process and its carcinogenic potential.

[0044] The HBV expression level and fragment distribution analysis method based on single-cell RNA-seq of the present invention has its own unique advantages, such as innovatively converting the circular HBV genome into a linear transcriptome RNA for comparison, and then using the mapping relationship between the transcriptome and the genome to achieve accurate analysis, overcoming the inaccurate comparison problem caused by the HBV circular structure in the traditional method. Because of this, it can be implemented through a software, which integrates the complete process from sequence acquisition, data comparison, sequence recognition, transcriptome sequence processing, position conversion to expression level and fragment distribution calculation and display. It uses optimization algorithms and multi-threading technology to ensure the efficiency and accuracy of data processing, and can also facilitate researchers to understand the results through built-in visualization and auxiliary interpretation functions. At the same time, the perfect system architecture guarantees data security and traceability, providing a one-stop, efficient, accurate and reliable analysis platform for HBV research, and promoting the development of related fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a schematic diagram of constructing the location information of each HBV gene in the transcriptome and genome according to an embodiment of the present invention.

[0046] Figure 2 This is a schematic diagram of extracting sequences and transcriptome alignment position information derived from HBV in scRNAseq sequencing data according to an embodiment of the present invention.

[0047] Figure 3 Schematic diagram of the alignment position of sequencing data on the genome according to an embodiment of the present invention.

[0048] Figure 4 This is a distribution diagram of the expression levels of the sequencing data of an embodiment of the present invention on the HBV genome.

[0049] Figure 5 This is a complete flowchart of the scRNA-seq-based HBV expression level and fragment distribution analysis method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0050] The following is further described in detail through specific implementation methods:

[0051] Embodiment 1

[0052] This embodiment is basically as shown in the attached Figure 1 To Attachment Figure 5 As shown: The analysis method of HBV expression level and fragment distribution based on single-cell RNA-seq in this embodiment includes the following steps:

[0053] The first step is to construct a mapping relationship between the transcriptome and the genome: obtain the HBV genome reference sequence and the HBV transcriptome reference sequence respectively, obtain the annotation information of the HBV genome, construct the position information of each HBV gene in the transcriptome and the genome, and obtain the mapping relationship between the transcriptome and the genome; this embodiment downloads the following from the National Center for Biotechnology Information (NCBI): HBV genome sequence (https: / / www.ncbi.nlm.nih.gov / nuccore / NC_003977.2?report=fasta), HBV transcriptome sequence (https: / / ftp.ncbi.nlm.nih.gov / genomes / all / GCF / 000 / 861 / 825 / GCF_000861825.2_Vira lProj15428 / GCF_000861825.2_ViralProj15428_cds_from_genomic.fna.gz), HBV genome annotation file (https: / / ncbi.nlm.nih.gov / datasets / gene / GCF_000861825.2 / ), and according to the attached Figure 1 The location information of each HBV gene in the transcriptome and genome was constructed. The X gene is located at the starting position 1376 and the ending position 1840 of the HBV genome NC_003977.2, which is the X protein; the C gene is located at the starting position 1816 and the ending position 2454 of NC_003977.2, which is the precapsid protein; the C gene is located at the starting position 1903 and the ending position 2454 of NC_003977.2, which is the capsid protein; the P gene is located at the starting position 2309 and the ending position 3182 of NC_003977.2, which is the polymerase; the P gene is located at the starting position 1 and the ending position 2454 of NC_003977.2, which is the precapsid protein; the P gene is located at the starting position 2309 and the ending position 3182 of NC_003977.2, which is the polymerase; the P gene is located at the starting position 1 and the ending position 2454 of NC_003977.2, which is the precapsid protein; the P gene is located at the starting position 2309 and the ending position 3182 of NC_003977.2, which is the precapsid protein ... Position 1625 is the P gene, which is a polymerase; NC_003977.2 starting position 2850 and ending position 3182 are the S gene, which is a large envelope protein; NC_003977.2 starting position 1 and ending position 837 are the S gene, which is a large envelope protein; NC_003977.2 starting position 3174 and ending position 3182 are the S gene, which is a middle membrane protein; NC_003977.2 starting position 1 and ending position 837 are the S gene, which is a middle membrane protein; NC_003977.2 starting position 1 and ending position 837 are the S gene, which is a small envelope protein.

[0054] The second step is to compare the sequencing data and determine the sequence source and breakpoint location: Figure 2 As shown, the original sequencing data obtained by scRNAseq (including attached Figure 2The sequences derived from humans as shown by the dotted lines and the sequences derived from HBV as shown by the solid lines) are aligned to the HBV transcriptome reference sequence, the sequences derived from humans are identified as human-derived sequences, and the sequencing data derived from HBV in the sequencing data are identified as HBV source sequences. Specifically, the FASTQ file of the sequencing raw data obtained by scRNAseq is aligned to the HBV transcriptome reference sequence using the bwa mem software (version v0.7.17), and only the sequencing sequence information that can be aligned to the HBV genome is retained to generate a BAM format alignment record file aligned to the HBV genome.

[0055] The third step is to extract and process the HBV source sequence: extract the HBV source sequence in the second step, and remove the partial sequence from the human transcriptome, that is, the human sequence that has been identified and separated, and retain the remaining HBV source sequence and the corresponding transcriptome coordinate position starting from the HBV potential breakpoint position to obtain the HBV source transcriptome sequence. Specifically, Figure 2 As shown, according to the CIGAR record information of each sequencing sequence in the BAM file, the sequencing sequence with soft clip (S) structure is extracted, and the soft clip part is cut off, and the remaining sequence and the corresponding transcriptome coordinate position are retained; for the sequencing sequence without soft clip structure, the sequence and the corresponding transcriptome coordinate position are retained.

[0056] Step 4: Convert the alignment position: The HBV source transcriptome sequence obtained in step 3 is converted into the alignment position on the genome through the transcriptome and genome mapping relationship constructed in step 1; specifically, Figure 3 As shown, the sequence and transcriptome alignment position information from HBV in the scRNAseq sequencing data are combined with the results of the first step to convert the transcriptome alignment position of the sequencing data into a genome alignment position:

[0057] Example A, gene NC_003977.2, start position 1712, end position 1819, nucleotide sequence AAAGACTGTGTGTTTAATGAGTGGGAGGAGTT, cell barcode ATGTCCCCATAACAGA, unique molecular identifier GTATGATATTCG;

[0058] Example A, gene NC 003977.2, start position 1511, end position 1633, nucleotide sequence C0GACCGACCACGGGGCGCACCTCTCTTTACG, cell barcode CTCGAGGTCTGAGGCC, unique molecular identifier TTGAAACATTTA;

[0059] Example A, gene NC 003977.2, start position 1511, end position 1606, nucleotide sequence C0GACCGACCACGGGGCGCACCTCTCTTTACG, cell barcode GATTCGATCAAACGAA, unique molecular identifier GAAGCAACCCCT;

[0060] Example A, gene NC 003977.2, start position 998, end position 1147, nucleotide sequence AATTGTGGGTCTTTTGGGTTTTGCCGCCCCTT, cell barcode ACCTACCGTCCOGGTA, unique molecular identifier TATTTTCCAAGG;

[0061] Example A, gene NC 003977.2, start position 998, end position 1147, nucleotide sequence AATIGIGGGTCTTTIGGGGTTTGCOGCCCCTT, cell barcode ACCTACCGTCCCGGTA, unique molecular identifier TCTTTTCCAAGG;

[0062] Example B, gene NC 003977.2, start position 1594, end position 1743, nucleotide sequence CACCTCIGCACGTOGCATGGAGACCACCGTGA, cell barcode CIGGACCGTAACACAT, unique molecular identifier CTCCATTAAGAT;

[0063] Example B, gene NC 003977.2, start position 1668, end position 1817, nucleotide sequence TIGGACTTTCAGCAATGTCAACGACCGACCTT, cell barcode CCCTCTCICGTTCAGA, unique molecular identifier TAGTCTAAAAGG;

[0064] Example B, gene NC 003977.2, start position 3093, end position 3182, nucleotide sequence CCTCCTGCCTCCACCAATOGGCAGTCAGGAOG, cell barcode GACGCTGGCAGAGACA, unique molecular identifier GTATTCIGAGCA;

[0065] Example B, gene NC 003977.2, start position 1657, end position 1800, nucleotide sequence TAAGAGGACTCTTGGACTTTCAGCAATGICAA, cell barcode TICAGGATCGAACTCA, unique molecular identifier ACATGCGGGCAT;

[0066] Example C, gene NC 003977.2, start position 1694, end position 1819, nucleotide sequence GACCTIGAGGCATACTTCAAAGACIGTGIGTT, cell barcode AGGGTTTICCAAGCAT, unique molecular identifier CACCGACTAGTG;

[0067] Example C, gene NC 003977.2, start position 1673, end position 1819, nucleotide sequence CTTTCAGCAATGTCAACGACCGACCTIGAGGC, cell barcode AAGCGTIICTGICGTC, unique molecular identifier TTAGGTATCAAT;

[0068] Example C, gene NC 003977.2, start position 482, end position 631, nucleotide sequence TAATTCCAGGATCATCAACAACCAGCACCGGA, cell barcode TTIGGTTTCACTCOG, unique molecular identifier ATATCATGACCT;

[0069] Example C, gene NC 003977.2, start position 482, end position 631, nucleotide sequence TAATTCCAGGATCATCAACAACCAGCACOGGA, cellular barcode TTTGGTTICACTCCGT, unique molecular identifier ATATCATGACCT.

[0070] Step 5: Calculate the expression level and fragment distribution: Calculate the number of sequencing sequences covered on each base of the genome, plot the coverage depth of each base of the HBV genome, and mark the gene position to obtain the following figure: Figure 4 The expression levels and fragment distribution of single-cell RNA-seq sequencing data on the HBV genome are shown.

[0071] When identifying human source sequences and HBV source sequences in sequencing data, the specific steps of this embodiment are as follows: First, the sequencing raw data FASTQ file obtained by scRNAseq is aligned to the HBV transcriptome reference sequence using bwa mem software (version v0.7.17). During the alignment process, the sequencing sequence can be matched with the HBV transcriptome reference sequence through the algorithm of the software. Only the sequencing sequence information that can be aligned to the HBV genome is retained, and a BAM format alignment record file aligned to the HBV genome is generated. In this BAM file, detailed alignment information for each sequencing sequence is contained.

[0072] The determination of the breakpoint position when HBV is integrated into the human genome is mainly based on the CIGAR record information of each sequencing sequence in the BAM file. The CIGAR record describes in detail the alignment of the sequencing sequence with the reference sequence, among which the softclip (S) structure is the key information. The sequencing sequence with the softclip structure is extracted. The softclip part is usually the part that is "cut" off because the sequence does not completely match the reference sequence during the alignment process. Since HBV integration into the human genome may cause such abnormal alignment of the sequencing sequence, the possible breakpoint position can be inferred by analyzing the position and sequence characteristics of the softclip part. Specifically, the softclip part is removed, and the remaining sequence and the corresponding transcriptome coordinate position are retained. For sequencing sequences without softclip structures, the sequence and the corresponding transcriptome coordinate position are retained. Through this analysis of a large number of sequencing sequences, the breakpoint position when HBV is integrated into the human genome can be accurately determined by comprehensively considering the alignment of the sequence, the distribution of the softclip structure, and the transcriptome coordinate position. For example, if multiple sequencing sequences are found to have soft-cut structures at similar positions in a certain region, and the region is associated with known human genome and HBV genome sequence features, then it can be highly suspected that the region is the breakpoint location of HBV integration, and the breakpoint location can be further confirmed through existing experimental verification and other methods.

[0073] This embodiment solves the problem of the lack of methods for calculating HBV expression levels from single-cell sequencing data and locating the expression abundance of various regions on the HBV genome, filling the gap in HBV analysis methods based on single-cell resolution. Compared with existing HBV analysis methods, it has obvious advantages:

[0074] First, the existing alignment tools are used to align sequences of linear genomes, and the HBV genome sequence (NCBI nucleic acid ID: NC_003977.2) is also recorded in the database in a linear manner. However, the HBV genome is a circular genome, and the S and P genes of HBV span the start and end positions of the HBV genome. When the existing tools align the sequencing data with the HBV genome, the expression level of the S and P genes will be quantified with deviations. This solution aligns the sequencing data with the linear transcriptome, and indirectly aligns the sequencing data with the genome based on the mapping relationship between the transcriptome and the genome, thereby achieving more accurate quantification of gene expression levels;

[0075] Secondly, there are many overlapping genes on the HBV genome, and the same gene has multiple transcripts. The method of aligning data to the transcriptome and then mapping it to the genomic location in this scheme can better distinguish the expression levels of overlapping genes and different transcripts of genes;

[0076] Finally, this example can further determine the breakpoint position when HBV integrates into the human genome through the sequencing sequence of the human-HBV genome chimera, providing more abundant research information in addition to the expression level.

[0077] This example reveals cell-to-cell heterogeneity: HBV infection may lead to different response patterns in different cell types or within the same cell type. scRNA-seq can identify which cell types are more susceptible to HBV infection and how these cells regulate their own gene expression during infection.

[0078] This embodiment overcomes the unique biological characteristics of HBV (such as the lack of polyA tail and circular genome structure), directly applies the existing scRNA-seq technology and analysis process to study the contradictory problems existing in HBV, and accurately quantifies and locates the expression activities on the HBV genome by constructing an accurate mapping relationship between the transcriptome and the genome.

[0079] In the process of developing this technical solution, we faced many challenges, among which the technical difficulties brought about by the unique biological characteristics of HBV were particularly prominent. The key to breaking the convention lies in the handling of the circular structure of the HBV genome and the gene crossing phenomenon.

[0080] Traditional alignment tools usually perform sequence alignment on linear genomes, while the HBV genome is a circular genome, and its S and P genes span the start and end positions of the HBV genome. This special structure causes deviations in the quantitative expression levels of the S and P genes when the sequencing data is aligned to the HBV genome using existing tools. This is a technical bias that has always existed in this field, that is, it is generally believed that the existing alignment methods for linear genomes can be directly applied to the analysis of the HBV genome.

[0081] To overcome this bias, this solution breaks the conventional thinking and no longer directly aligns the sequencing data to the circular genome. Instead, it aligns the sequencing data to the linear transcriptome, and indirectly aligns the sequencing data to the genome based on the mapping relationship between the constructed transcriptome and the genome. Specifically, in the first step, the HBV genome sequence, transcriptome sequence, and genome annotation files are downloaded from NCBI respectively, and the position information of each HBV gene in the transcriptome and genome is carefully constructed to obtain the mapping relationship between the transcriptome and the genome. In the subsequent steps, the raw sequencing data obtained by scRNAseq is first aligned to the HBV transcriptome reference sequence, the human sequence and the HBV source sequence are identified and separated, and then the HBV source transcriptome sequence is converted into a matching position on the genome through a mapping relationship. This innovative method effectively solves the problem of alignment bias caused by the special structure of the HBV genome and achieves more accurate quantification of gene expression levels.

[0082] Before obtaining the final solution, the R&D team tried to directly use existing alignment tools for linear genomes, such as Bowtie2, to directly align the raw sequencing data obtained by scRNAseq to the circular genome of HBV. However, due to the special structure of the HBV genome, especially the S and P genes spanning the start and end positions, the alignment results of these two genes deviated seriously from the actual situation, and the quantitative expression levels showed great deviations. After an in-depth analysis of the results, it was found that this method could not accurately distinguish between different transcripts of genes and the expression levels of overlapping genes, and could not provide reliable analysis results for the complex gene expression situation on the HBV genome.

[0083] Another abandoned solution is to try to simply modify the existing alignment tools to adapt to the circular genome of HBV. For example, try to adjust the parameters in the alignment algorithm so that it can take into account the situation of gene crossing to a certain extent. However, after a large number of experimental verifications, although this method has improved some data, the overall effect is still not ideal and cannot meet the requirements for comprehensive and accurate analysis of the HBV genome.

[0084] In terms of the accuracy of quantitative gene expression levels, this embodiment uses the S and P genes as an example. Due to the innovative method of aligning sequencing data to the transcriptome and then mapping it to the genome, compared with the traditional method of directly aligning to the circular genome, the expression levels of these two genes that span the start and end positions of the genome can be more accurately quantified. Assuming that in the traditional method, the quantitative deviation of the S gene expression level may reach 30%-50% (this is the deviation range inferred by the alignment error caused by the circular structure), after adopting this solution, through the precise construction and application of the mapping relationship between the transcriptome and the genome, this deviation can be controlled within 5%-10%, greatly improving the accuracy of quantification.

[0085] For distinguishing the expression levels of overlapping genes and different transcripts of genes, this solution can clearly distinguish the expression signals of different transcripts and overlapping genes through a unique comparison and analysis process. For example, when analyzing multiple overlapping gene regions on the HBV genome, traditional methods may not be able to accurately distinguish the expression contributions of different genes, resulting in confusing analysis results. However, this solution can accurately identify the expression of each gene transcript, increase the expression analysis resolution of overlapping gene regions by at least 2-3 times, and enable more detailed research on the expression regulation mechanism of HBV genes.

[0086] In terms of determining the breakpoints when HBV integrates into the human genome, this protocol can provide richer research information by analyzing the sequencing sequences of the human-HBV genome mosaic. Compared with traditional methods, more potential breakpoints can be detected. Assuming that traditional methods can only detect 5-8 breakpoints, this protocol can increase the number of breakpoints detected to 15-20, greatly increasing the depth and breadth of research on HBV integration mechanisms.

[0087] Embodiment 2

[0088] In this embodiment, in the process of determining the breakpoint position when HBV is integrated into the human genome, the sequence alignment, soft-cut structure distribution and transcriptome coordinate position are comprehensively considered in the following manner:

[0089] For the sequence alignment, the alignment results of the sequencing sequence and the HBV transcriptome reference sequence are analyzed in detail. If the base matching rate of the sequencing sequence and the reference sequence in a certain region is lower than the preset threshold, and the low matching rate is continuously distributed, and the matching pattern is significantly different from the normal alignment region, the region is marked as a potential breakpoint position region. The preset threshold is 70%-80%, ensuring that it can effectively distinguish between normal and abnormal alignment situations. For example, when the base matching rate of the sequencing sequence and the reference sequence in a certain region is lower than 75%, and the low matching rate is continuously distributed, that is, the number of continuous bases in the low matching rate region exceeds 10 (the length is determined according to the average fragment length of the HBV transcriptome reference sequence and the allowable range of sequencing errors), and the matching pattern is significantly different from the normal alignment region, the region is marked as a potential breakpoint position region. The purpose of this setting is to avoid misjudging the low match caused by accidental sequencing errors as a potential breakpoint position, because in the normal sequencing process, a small amount of base mismatches may occur due to various factors, but such mismatches are usually dispersed and do not present continuous low matches. When the matching rate of multiple consecutive bases is lower than 75%, it is likely that factors such as HBV integration have caused changes in the sequence structure, thus affecting the alignment results.

[0090] For the distribution of soft-cut structures, statistical analysis is performed on sequencing sequences with soft-cut structures, and the starting and ending positions of each soft-cut structure in the sequence are counted, and these positions are mapped to transcriptome coordinates; when the frequency of occurrence of the soft-cut structure exceeds the preset frequency threshold in a specific transcriptome coordinate interval, the coordinate interval is marked as a potential breakpoint position interval. The preset frequency threshold is that the number of soft-cut structures in every 1000 sequencing sequences exceeds 10 times, ensuring that it can accurately screen out the soft-cut structure aggregation area that may be caused by HBV integration. When the frequency of occurrence of the soft-cut structure exceeds this threshold in a specific transcriptome coordinate interval, for example, in a certain 100bp transcriptome coordinate interval, the number of soft-cut structures reaches more than 15 times, and the average length, maximum length and other parameters of the soft-cut structure are significantly different from other normal areas (such as the average length is more than 20% longer than the normal area), the coordinate interval is marked as a potential breakpoint position interval. This is because when HBV integrates into the human genome, it may cause abnormal soft cutting at certain positions in the sequencing sequence. By counting the frequency and related length parameters of the soft cutting structure, these regions that may be related to HBV integration can be effectively screened out.

[0091] Combined with the transcriptome coordinate position, the transcriptome coordinate continuity of the sequencing sequence is analyzed in the human-HBV genome chimeric region. If the transcriptome coordinates are found to jump, interrupt, or inconsistent with the known human genome and HBV genome transcriptome coordinate ranges, the corresponding position will be marked as a potential breakpoint position; special attention will be paid to the coordinate changes at gene boundary regions, repetitive sequence regions, and other locations where integration is prone to occur. For example, when the transcriptome coordinates suddenly jump from the human genome coordinate range to the HBV genome coordinate range at a certain point, and the sequences before and after the jump are highly similar to specific conserved sequences of the human genome and HBV genome, respectively, the jump position is very likely to be the breakpoint position of HBV integration.

[0092] Finally, the potential breakpoint position regions marked by the above three methods are comprehensively considered. When multiple marked regions overlap or are densely distributed in a small range, the common region or dense region is determined as the breakpoint position when HBV integrates into the human genome. Among them, "dense distribution in a small range" specifically refers to the presence of 3 or more potential breakpoint position regions marked by different methods within a radius of no more than 50bp centered on a certain position. When this happens, the common region or dense region is determined as the breakpoint position when HBV integrates into the human genome. This is because multiple independent judgment factors point to the possible existence of breakpoints in similar areas, which greatly increases the credibility of the region as the true breakpoint position. Through this comprehensive analysis method, the accuracy and reliability of breakpoint position determination can be effectively improved, avoiding misjudgments that may be caused by single factor judgment.

[0093] The HBV expression level and fragment distribution analysis method based on single-cell RNA-seq of the present invention significantly improves the recognition accuracy, especially in the accurate recognition of breakpoint positions. In traditional analysis, since the gene library stores linear sequences, and the HBV genome is a circular structure, some genes span the beginning and the end. When using traditional tools to compare sequencing data with the circular HBV genome, there will be inaccurate comparison problems in special areas such as spanning the beginning and the end and breakpoints due to structural differences. This is because conventional comparison algorithms are designed based on linear genome features and are difficult to handle the special sequence features and positional relationships of circular genomes.

[0094] The present invention innovatively transcribes the circular HBV genome into a linear transcriptome RNA, eliminating the comparison complexity caused by the circular structure, and its sequence is more consistent with the linear reference sequence of the gene library. By comparing the transcriptome RNA with the reference sequence in the gene library, not only the comparison deviation caused by the circular structure is avoided, but also the foundation is laid for the subsequent accurate identification of the breakpoint position. From the perspective of molecular biology and bioinformatics principles, the transcription process faithfully records the HBV genome coding information and presents it in a linear form, so that the comparison algorithm can more accurately identify the matching area and reduce mismatches or missed matches. At the same time, by constructing a mapping relationship between the transcriptome and the genome, the transcriptome comparison results are accurately mapped back to the genome, further ensuring the accuracy of the expression level and fragment distribution analysis of each region on the HBV genome, and realizing the accurate determination of the breakpoint position. Therefore, the method of the present invention effectively overcomes the limitations of traditional methods and significantly improves the ability to accurately identify HBV-related sequences and breakpoint positions.

[0095] Embodiment 3

[0096] In this embodiment, in the process of determining the breakpoint position when HBV is integrated into the human genome, the sequence alignment, soft-cut structure distribution and transcriptome coordinate position are comprehensively considered in the following manner:

[0097] For sequence alignment, the alignment results of the sequencing sequence and the HBV transcriptome reference sequence are analyzed in detail. If the base matching rate between the sequencing sequence and the reference sequence in a certain region is less than 75%, and the low matching rate is continuously distributed, that is, the number of continuous bases in the low matching rate region exceeds 15bp, and the matching pattern is significantly different from that of the normal alignment region, the region is marked as a potential breakpoint location region. The length of 15bp is determined based on the average fragment length of the HBV transcriptome reference sequence and the allowable range of sequencing errors to avoid misjudging low matches caused by accidental sequencing errors as potential breakpoint locations.

[0098] For the distribution of soft-cut structures, statistical analysis is performed on sequencing sequences with soft-cut structures, and the starting and ending positions of each soft-cut structure in the sequence are counted, and these positions are mapped to transcriptome coordinates. When the frequency of occurrence of the soft-cut structure exceeds the preset frequency threshold of 10 times in every 1000 sequencing sequences in a specific transcriptome coordinate interval, and the average length of the soft-cut structure in the interval is at least 10bp longer than the normal region (such as the average length of the normal region is 15bp, and the average length of the interval is 25bp or more), and the maximum length is at least 20bp longer than the normal region (such as the maximum length of the normal region is 20bp, and the maximum length of the interval is 40bp or more), the coordinate interval is marked as a potential breakpoint position interval.

[0099] Combined with the transcriptome coordinate position, the transcriptome coordinate continuity of the sequencing sequence is analyzed in the human-HBV genome chimeric region. If the transcriptome coordinates are found to jump, interrupt, or inconsistent with the known human genome and HBV genome transcriptome coordinate range, the corresponding position will be marked as a potential breakpoint position area. In particular, for the transcriptome coordinate jumps found, the sequence characteristics before and after the jump are further analyzed. If the sequence before the jump has a similarity of 95% or more with a conservative sequence of the human genome, and the sequence after the jump has a similarity of 90% or more with a conservative sequence of the HBV genome, the jump position will be regarded as a potential breakpoint position that is the focus of attention.

[0100] Finally, the potential breakpoint position regions and potential breakpoint position intervals marked by the above three methods were comprehensively considered. When multiple marked regions overlapped with each other or three or more potential breakpoint position regions marked by different methods appeared within a radius of no more than 50bp centered on a certain position, the common region or dense region was determined as the breakpoint position when HBV integrated into the human genome.

[0101] Among them, when analyzing the sequence alignment to determine the potential breakpoint location area, the number of continuous bases 15bp based on the continuity distribution is determined by experiments and data analysis based on the average fragment length of the HBV transcriptome reference sequence of 200bp, combined with the allowable range of sequencing errors, to avoid misjudging low matches caused by accidental sequencing errors as potential breakpoint positions, and ensure that normal alignments can be effectively distinguished from abnormal alignments that may be caused by HBV integration.

[0102] Among them, when the distribution of soft-cut structures is statistically determined to determine the potential breakpoint position interval, the preset frequency threshold is set to 10 occurrences of soft-cut structures in every 1,000 sequencing sequences. At the same time, the difference between the average length and maximum length of the soft-cut structure and the normal region is required, such as the average length is at least 10bp longer than the normal region, and the maximum length is at least 20bp longer than the normal region, so as to improve the accuracy of screening the potential breakpoint position interval.

[0103] Among them, when determining the potential breakpoint position in combination with the transcriptome coordinate position, the requirement for the similarity between the sequences before and after the jump and the corresponding conservative sequences of the genome, that is, the similarity between the sequence before the jump and a conservative sequence of the human genome reaches 95% or more, and the similarity between the sequence after the jump and a conservative sequence of the HBV genome reaches 90% or more. This is determined through comparison and analysis of a large amount of known human genome and HBV genome conservative sequence data, combined with actual verification experiments, to ensure that the transcriptome coordinate jump position caused by HBV integration can be accurately identified as a potential breakpoint position.

[0104] This embodiment determines the location of HBV integration breakpoints through multi-dimensional precise analysis, significantly improving accuracy. It not only effectively filters sequencing errors and accurately identifies potential areas, but also avoids omissions in single-dimensional analysis by comprehensively considering soft-cut structures and transcriptome coordinate characteristics. The comprehensive judgment of multiple evidences enhances credibility and fully guarantees the integrity of breakpoint determination. Based on accurate breakpoint position judgment, this embodiment enables the analysis method of HBV expression level and fragment distribution based on single-cell RNA-seq to complete all computational analysis through a software program, making the operation more accurate and faster.

[0105] The above is only an embodiment of the present invention, and the common knowledge such as the known specific technical solutions and / or characteristics in the solution is not described in detail here. It should be pointed out that for those skilled in the art, without departing from the technical solution of the present invention, several modifications and improvements can be made, which should also be regarded as the protection scope of the present invention, and these will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.

Claims

1. A method for analyzing HBV expression level and fragment distribution based on single-cell RNA-seq, characterized in that: The following steps are involved: Step 1: construct the location information of each HBV gene in the linear transcriptome and circular genome to obtain the mapping relationship between the transcriptome and the genome; Step 2: Align the original sequencing data obtained by scRNAseq to the HBV transcriptome reference sequence to determine the breakpoint position when HBV integrates into the human genome and identify the human and HBV source sequences; Step 3, removing the human source sequence, retaining the remaining HBV source sequence and the corresponding transcriptome coordinate position starting from the HBV potential breakpoint position, and obtaining the HBV source transcriptome sequence; Step 4, using the transcriptome-genome mapping relationship constructed in step 1 to map the HBV source transcriptome sequence obtained in step 3, converting the position of the sequencing sequence alignment on the transcriptome into the alignment position on the genome; Step 5: Calculate the number of sequencing sequences covered on each base of the genome, plot the coverage depth of each base of the HBV genome, and mark and display the gene position to obtain the expression level and fragment distribution of single-cell RNA-seq sequencing data on the HBV genome.

2. The method for analyzing HBV expression level and fragment distribution based on single-cell RNA-seq according to claim 1, characterized in that: In step 1, the HBV genome reference sequence and the HBV transcriptome reference sequence are first obtained respectively, the annotation information of the HBV genome is obtained, the position information of each HBV gene in the transcriptome and the genome is constructed, and the mapping relationship between the transcriptome and the genome is obtained; the mapping relationship between the transcriptome and the genome is that the starting position 1376 and the ending position 1840 of the HBV genome NC_003977.2 are the X gene, which is the X protein; the starting position 1816 and the ending position 2454 of NC_003977.2 are the C gene, which is the precapsid protein; the starting position 1903 and the ending position 2454 of NC_003977.2 are the C gene, which is the capsid protein; the starting position 23 09 The P gene at the end position 3182 is a polymerase; NC_003977.2 The P gene at the starting position 1 and the end position 1625 is a polymerase; NC_003977.2 The S gene at the starting position 2850 and the end position 3182 is a large envelope protein; NC_003977.2 The S gene at the starting position 1 and the end position 837 is a large envelope protein; NC_003977.2 The S gene at the starting position 3174 and the end position 3182 is a middle membrane protein; NC_003977.2 The S gene at the starting position 1 and the end position 837 is a middle membrane protein; NC_003977.2 The S gene at the starting position 1 and the end position 837 is a middle membrane protein; NC_003977.2 The S gene at the starting position 157 and the end position 837 is a small envelope protein.

3. The method for analyzing HBV expression level and fragment distribution based on single-cell RNA-seq according to claim 1, characterized in that: In step 2, the FASTQ file of the sequencing raw data obtained by scRNAseq was aligned to the HBV transcriptome reference sequence using bwa mem software, and only the sequencing sequence information that could be aligned to the HBV genome was retained to generate a BAM format alignment record file aligned to the HBV genome as the HBV source sequence and the rest were human sequences.

4. The method for analyzing HBV expression level and fragment distribution based on single-cell RNA-seq according to claim 3, characterized in that: In step 3, according to the CIGAR record information of each sequencing sequence in the BAM file, the sequencing sequence with soft-cut structure is extracted, and the soft-cut part is removed, retaining the remaining sequence and the corresponding transcriptome coordinate position; for the sequencing sequence without soft-cut structure, the sequence and the corresponding transcriptome coordinate position are retained.

5. The method for analyzing HBV expression level and fragment distribution based on single-cell RNA-seq according to claim 3, characterized in that: In determining the breakpoint position when HBV integrates into the human genome, the sequence alignment, soft-cut structure distribution and transcriptome coordinate position are comprehensively considered in the following ways: For sequence comparison, the comparison results of the sequencing sequence and the HBV transcriptome reference sequence are analyzed in detail. If the base matching rate between the sequencing sequence and the reference sequence in a certain area is lower than the preset threshold, and the low matching rate is distributed continuously, and the matching pattern is significantly different from that of the normal comparison area, the area will be marked as a potential breakpoint location area; The preset threshold is 70%-80%; For the distribution of soft-cut structures, statistical analysis is performed on sequencing sequences with soft-cut structures, the starting position and the ending position of each soft-cut structure in the sequence are counted, and these positions are mapped to the transcriptome coordinates; when the frequency of occurrence of the soft-cut structure exceeds the preset frequency threshold in a specific transcriptome coordinate interval, the coordinate interval is marked as a potential breakpoint position interval; The preset frequency threshold is that the soft-cut structure appears more than 10 times in every 1000 sequencing sequences; Combined with the transcriptome coordinate position, the transcriptome coordinate continuity of the sequenced sequence was analyzed within the chimeric region of the human-HBV genome; If the transcriptome coordinates are found to be jumping, interrupted, or inconsistent with the known human genome and HBV genome transcriptome coordinate ranges, the corresponding positions will be marked as potential breakpoint location regions; Finally, the potential breakpoint position regions and potential breakpoint position intervals marked by the above three methods are comprehensively considered. When multiple marked regions overlap with each other or are densely distributed in a small range, the common region or dense region is determined as the breakpoint position when HBV integrates into the human genome.

6. The method for analyzing HBV expression level and fragment distribution based on single-cell RNA-seq according to claim 5, characterized in that: The dense distribution in a smaller range refers to the presence of 3 or more potential breakpoint location regions marked in different ways within a radius of no more than 50 bp centered on a certain position.

7. The method for analyzing HBV expression level and fragment distribution based on single-cell RNA-seq according to claim 5, characterized in that: When analyzing sequence alignment to determine potential breakpoint location regions, a low match rate presents a continuous distribution, which means that the number of continuous bases in the low match rate region exceeds 10.

8. The method for analyzing HBV expression level and fragment distribution based on single-cell RNA-seq according to claim 5, characterized in that: When determining the potential breakpoint position interval by statistically analyzing the soft-cut structure distribution, the length distribution of the soft-cut structure is also considered; when within a certain transcriptome coordinate interval, not only does the frequency of soft-cut structures exceed the preset frequency threshold, but the average length of the soft-cut structure is at least 10 bp longer than that of the normal region, and the maximum length is at least 20 bp longer than that of the normal region.

9. The method for analyzing HBV expression level and fragment distribution based on single-cell RNA-seq according to claim 1, characterized in that: In step 4, the position of the sequencing sequence alignment on the transcriptome is converted into the alignment position on the genome: gene NC_003977.2, start position 1712, end position 1819, nucleotide sequence AAAGACTGTGTGTTTAATGAGTGGGAGGAGTT.

10. The method for analyzing HBV expression level and fragment distribution based on single-cell RNA-seq according to claim 1, characterized in that: In step 4, the position of the sequencing sequence alignment on the transcriptome is converted into the alignment position on the genome: gene NC 003977.2, start position 1511, end position 1633, nucleotide sequence C0GACCGACCACGGGGCGCACCTCTCTTTACG.

Citation Information

Patent Citations

  • Method and device for extracting gene fusion immune treatment neoantigen by integrating deep sequencing data of DNA and RNA

    CN111192632A

  • Methods of detecting structural variations in genomic regions

    CN112349346A

  • Method for detecting and analyzing virus expression quantity in single cell transcriptome sequencing data

    CN115512767A

  • Copy number variation detection method based on single-sample high-throughput transcriptome sequencing data

    CN117095744A

  • Single-cell multi-omics parallel sequencing method based on third-generation sequencing platform and application of single-cell multi-omics parallel sequencing method

    CN117106873A