Sequencing data processing method and apparatus, electronic device, and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BGI RESEARCH HANGZHOU
- Filing Date
- 2026-07-08
- Publication Date
- 2026-08-07
AI Technical Summary
[0004]相关技术中,缺乏对单细胞长读长测序数据的UMI进行校正的有效方法
[0018] The sequencing data processing method proposed in this application involves acquiring single-cell long-read sequencing data of a target sample, whereby single-cell long-read sequencing includes multiple sequence reads; extracting the cell identifier sequence corresponding to each sequence read, and dividing the multiple sequence reads into multiple processing units based on the cell identifier sequence, with each processing unit containing multiple reads to be binned; constructing binning features for each read to be binned within each processing unit, the binning features being generated based on the splicing site information or alignment position information of the read to be binned; binning the multiple reads to be binned based on the binning features, and correcting the molecular identifier sequence of the reads to be binned in each processing unit based on the binning results.
Smart Images

Figure CN122531473A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of bioinformatics analysis technology, specifically to a sequencing data processing method, apparatus, electronic device, and storage medium. Background Technology
[0002] Long-read sequencing, also known as third-generation sequencing or single-molecule sequencing, is a sequencing technology that can accurately identify complete long transcripts. It can precisely detect complex variations, comprehensively characterize the genome, and is suitable for diverse samples. This technology has wide applications and plays an important role in genome assembly, variation detection, haplotype analysis, epigenetics, medical research, and microbial diversity research.
[0003] Single-cell long-read sequencing is a technology that can simultaneously sequence a large number of cells and directly obtain the complete molecular sequence information within each cell, thus enabling quantitative analysis of gene expression in single cells. However, due to the presence of sequencing errors in long-read sequencing data, when these errors occur in the sequence corresponding to a unique molecular identifier (UMI), it can lead to inaccurate quantitative analysis of gene expression in single cells. Therefore, it is necessary to correct for UMIs in single-cell long-read sequencing data.
[0004] Among related technologies, there is a lack of effective methods for correcting the UMI of single-cell long-read sequencing data. Summary of the Invention
[0005] The main objective of this application is to provide a sequencing data processing method, apparatus, electronic device, and storage medium, which aims to improve the accuracy of long-read sequencing data processing.
[0006] To achieve the above objectives, a first aspect of this application provides a sequencing data processing method, the method comprising: Obtain single-cell long-read sequencing data of the target sample, wherein the single-cell long-read sequencing includes multiple sequence reads; Extract the cell identifier sequence corresponding to each sequence read segment, and divide the multiple sequence read segments into multiple processing units according to the cell identifier sequence. Each processing unit contains multiple read segments to be bucketed. In each of the processing units, a binning feature is constructed for each of the read segments to be binned. The binning feature is generated based on the cut-off point information of the read segment to be binned or the alignment position information of the read segment to be binned. The multiple read segments to be bucketed are bucketed according to the bucketing characteristics, and the molecular identifier sequence of the read segments to be bucketed in each processing unit is corrected according to the bucketing results.
[0007] To achieve the above objectives, a second aspect of this application provides a sequencing data processing apparatus, the apparatus comprising: The acquisition unit is used to acquire single-cell long-read sequencing data of the target sample, wherein the single-cell long-read sequencing includes multiple sequence reads; A partitioning unit is used to extract the cell identifier sequence corresponding to each sequence read segment, and to divide the multiple sequence read segments into multiple processing units according to the cell identifier sequence, wherein each processing unit contains multiple read segments to be binned. A construction unit is used to construct a bucketing feature for each read segment to be bucketed in each of the processing units. The bucketing feature is generated based on the cut-off site information of the read segment to be bucketed or the alignment position information of the read segment to be bucketed. The correction unit is used to divide the multiple read segments to be divided into buckets according to the bucketing characteristics, and to correct the molecular identifier sequence of the read segments to be divided into buckets in each of the processing units according to the bucketing results.
[0008] Optionally, in some embodiments, the building unit includes: The parsing subunit is used to parse each of the read segments to be bucketed and obtain the chain direction information corresponding to each read segment to be bucketed in each of the processing units. The first generation subunit is used to generate binning features based on the cell identifier sequence of the read segment to be binned, the chain direction information, and the splicing site information when the read segment to be binned contains splicing information. The second generation subunit is used to generate binning features based on the cell identifier sequence of the read segment to be binned, the chain direction information, and the alignment position information when the read segment to be binned does not contain splicing information.
[0009] Optionally, in some embodiments, the first generating subunit includes: The extraction module is used to extract the start and end coordinates of each contained sub-interval in the segment to be divided into buckets; The quantization module is used to perform dithering quantization on the start and end coordinates according to a preset base length to obtain quantized coordinates; The generation module is used to generate the bucketing features of the read segment to be bucketed based on the quantized coordinates, the cell identifier sequence of the read segment to be bucketed, and the chain direction information.
[0010] Optionally, in some embodiments, the correction unit includes: The bucketing subunit is used to bucket the multiple read segments to be bucketed according to the bucketing characteristics, so as to obtain multiple sets of read segments to be corrected. The correction subunit is used to extract the molecular identifier sequence of each read segment to be corrected from each set of read segments to be corrected, and to aggregate the molecular identifier sequences according to the similarity between the molecular identifier sequences to obtain multiple aggregate groups, so as to correct the molecular identifier sequences of multiple read segments to be corrected corresponding to each aggregate group.
[0011] Optionally, the correction subunit includes: The acquisition module is used to acquire the number of molecular identifier sequences in each of the aggregate groups; An integration module is used to calculate the distance between molecular sequence identifiers corresponding to any two aggregation groups, and to integrate the multiple aggregation groups according to the distance and the number of molecular identifier sequences.
[0012] Optionally, in some embodiments, the correction subunit includes: The construction module is used to construct a molecular identifier sequence similarity graph based on the molecular identifier sequences of the plurality of reads to be corrected. The molecular identifier sequence similarity graph includes a plurality of molecular identifier sequence nodes and a plurality of candidate edges, wherein the distance between the molecular identifier sequence nodes connected by the candidate edges is less than a preset distance. The merging module is used to calculate the merging probability of the candidate edges and merge the candidate edges based on the merging probability to update the molecular identifier sequence similarity graph. The determination module is used to identify multiple clusters based on the updated molecular representation sequence similarity map.
[0013] Optionally, in some embodiments, the correction subunit includes: The search subunit is used to perform a fast nearest neighbor search on the molecular identifier sequences of the multiple read segments to be corrected based on a preset distance, so as to obtain the similarity relationship between the molecular identifier sequences; Determine subunits for identifying multiple aggregate groups based on the similarity relationships between the molecular identifier sequences.
[0014] Optionally, in some embodiments, the sequencing data processing apparatus provided in this application further includes: The second determining subunit is used to determine multiple deduplication sequence reads corresponding to each corrected molecular identifier sequence, and to obtain the alignment quality and alignment sequence length of each deduplication read. The deduplication subunit is used to deduplicate multiple read segments corresponding to each corrected molecular identifier sequence according to the alignment quality and alignment sequence length, so as to obtain the target read segment corresponding to each corrected molecular identifier sequence.
[0015] To achieve the above objectives, a third aspect of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the sequencing data processing method described in the first aspect.
[0016] To achieve the above objectives, a fourth aspect of the present application provides a storage medium storing a computer program that, when executed by a processor, implements the sequencing data processing method described in the first aspect.
[0017] To achieve the above objectives, a fifth aspect of this application provides a computer program product comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the sequencing data processing method described in the first aspect.
[0018] The sequencing data processing method proposed in this application involves acquiring single-cell long-read sequencing data of a target sample, whereby single-cell long-read sequencing includes multiple sequence reads; extracting the cell identifier sequence corresponding to each sequence read, and dividing the multiple sequence reads into multiple processing units based on the cell identifier sequence, with each processing unit containing multiple reads to be binned; constructing binning features for each read to be binned within each processing unit, the binning features being generated based on the splicing site information or alignment position information of the read to be binned; binning the multiple reads to be binned based on the binning features, and correcting the molecular identifier sequence of the reads to be binned in each processing unit based on the binning results.
[0019] Therefore, the sequencing data processing method provided in this application, when performing UMI correction based on single-cell long-read sequencing data, can first divide the sequence reads into bucket sets corresponding to each cell based on the barcode sequence of the sequence reads. Then, for each cell's bucket set, the reads are bucketed based on the splicing features or alignment positions of the reads. Finally, the UMI of each read is corrected based on the bucketing results. This method can accurately bucket sequence reads lacking gene annotation, thereby achieving UMI correction under conditions without gene annotation and improving the accuracy of sequencing data processing. Attached Figure Description
[0020] The accompanying drawings are used to provide a further understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.
[0021] Figure 1 A flowchart illustrating the sequencing data processing method provided in this application; Figure 2 A flowchart illustrating the UMI deduplication method for the sequencing data provided in this application; Figure 3 This is a schematic diagram of the sequencing data processing apparatus provided in the embodiments of this application; Figure 4 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0023] Before providing a further detailed description of the embodiments of this application, the nouns and terms used in the embodiments of this application are explained, and the nouns and terms used in the embodiments of this application shall be interpreted as follows: Long-read sequencing: Long-read sequencing is a core feature of third-generation sequencing technology. It refers to the ability to directly read DNA or RNA fragments of tens or even hundreds of thousands of bases in a single sequencing reaction, a stark contrast to second-generation sequencing technology. Second-generation sequencing, often referred to as "short-read sequencing," requires breaking DNA into small fragments of a few hundred bases, which are then assembled using bioinformatics methods. Figure 1 These fragments are then reassembled. Compared to short-read sequencing, long-read sequencing has a higher sequencing error rate.
[0024] Gene splicing: Gene splicing is a crucial step in gene expression in eukaryotes. Simply put, it is the process by which cells "cut and splice" the original RNA transcripts of genes, removing useless parts and retaining useful parts, ultimately generating mature messenger RNA. The essence of gene splicing is the removal of introns and the preservation of exons.
[0025] Exons are DNA sequence segments that ultimately appear in mature mRNA and are used to guide protein synthesis (encoding) or function as RNA. They are the "functional components" that make up the final "product".
[0026] Introns are DNA sequence segments that are removed from mature mRNA during RNA processing. They are non-coding spacer sequences that separate the "functional parts".
[0027] Related technologies provide tools for short-read single-cell UMI correction and deduplication (such as the UMI-tools series). These tools construct a network of similar UMIs using Hamming distance in a cell × gene dimension and perform correction and deduplication using a "directional adjacency / aggregation" algorithm. However, these tools are designed for short-read sequencing and rely on reliable gene annotation. In long-read sequencing, however, sequencing errors lead to a lack of accurate gene annotation, which can easily result in inaccurate binning and consequently inaccurate UMI correction and deduplication results.
[0028] The above-described methods illustrate that in UMI correction of long-read sequences in related technologies, the lack of accurate gene annotation information leads to inaccurate UMI correction and deduplication results. To address this, this application provides a sequencing data processing method to improve the accuracy of sequencing data processing. The sequencing data processing method provided in this application will be described below.
[0029] Reference Figure 1 In some embodiments, the sequencing data processing method provided in this application includes, but is not limited to, steps S101 to S104.
[0030] Step S101: Obtain single-cell long-read sequencing data of the target sample.
[0031] Specifically, the sequencing data processing method provided in this application is a method for UMI correction of single-cell long-read sequencing data. This method addresses the problem in related technologies where the lack of accurate gene annotation information in single-cell long-read sequencing data leads to inaccurate data binning when using short-read UMI correction tools for sequencing data processing, resulting in inaccurate UMI correction results for the sequencing data.
[0032] In this embodiment, the single-cell long-read sequencing data of the target sample includes multiple reads. After obtaining the sequencing data from the single-cell long-read sequencing of the target sample to be processed, the sequencing reads can be aligned to a reference genome to obtain a binary alignment / map format (BAM) file. Each aligned read contains a cell barcode and the original UMI sequence, and some reads may also contain gene annotation information.
[0033] Thus, after obtaining multiple sequence reads of the target sample, the cell identifier sequence of each sequence read can be further extracted so that the multiple sequence reads can be binned at the cell level.
[0034] Step S102: Extract the cell identifier sequence corresponding to each sequence read, and divide the multiple sequence reads into multiple processing units based on the cell identifier sequence.
[0035] After obtaining the cell identifier sequence for each sequence read, multiple sequence reads can be bucketed based on the cell identifier sequence for each sequence read. This means dividing multiple sequence reads into multiple processing units, with each processing unit corresponding to multiple bucketed reads of a cell.
[0036] Specifically, in some embodiments, multiple sequence reads are divided into multiple processing units based on cell identifier sequences, including: Clustering multiple cell identifier sequences corresponding to multiple sequence reads yields multiple cell identifier sequence classes; Multiple sequence reads are divided into multiple processing units based on multiple cell identifier sequence classes.
[0037] Step S103: In each unit to be processed, construct the bucketing features of each segment to be bucketed.
[0038] After multiple sequence reads are bucketed into multiple processing units corresponding to multiple cells based on the barcode sequence, UMI correction can be further performed on the reads to be bucketed in each processing unit corresponding to each cell. Specifically, before performing UMI correction on the reads to be bucketed in each processing unit corresponding to each cell, these reads need to be bucketed first. In this embodiment, a corresponding bucketing feature (or bucketing key) can be constructed for each read to be bucketed. The bucketing feature is generated based on the splice site information or alignment position information of the read to be bucketed.
[0039] Specifically, in some embodiments, in each processing unit, the bucketing characteristics of each read segment to be bucketed are constructed, including: In each unit to be processed, parse each segment to be bucketed. When the segment to be binned contains splicing information, binning features are generated based on the splicing point information of the segment to be binned. When the segment to be bucketed does not contain splicing information, bucketing features are generated based on the comparison position information of the segment to be bucketed.
[0040] This application provides a method for binning sequence reads based on splicing features. Specifically, each sequence read can first be distinguished between exons and introns, and the number of introns contained in each read can be identified based on the distinction results. Specifically, the CIGAR information of each read can be parsed to extract exon and intron intervals, and then the number of introns in each read can be determined accordingly. When a read is determined to contain splicing information based on the identified intron information, binning features can be generated based on the splicing site information of that read. The splicing site information can be determined based on the start and end coordinates of the introns, and the generated binning feature can also be called a splicing signature (SJ signature). If a read is determined not to contain splicing information, it is determined that the read is not spliced. In this case, a corresponding binning feature can be constructed based on the read coverage interval; this binning feature can also be called genomic interval binning (locus-bin).
[0041] In this embodiment, due to the specificity of splice signatures and genomic regions, introducing splice signatures and genomic region binning enables accurate molecular-level differentiation and binning even in the absence of gene annotation or the presence of new genes. This avoids the problem of inaccurate sequence binning and consequently inaccurate UMI correction caused by the lack of gene annotation information when performing UMI correction on single-cell long-read sequencing data. Furthermore, this embodiment, without splicing, accurately divides sequencing reads according to genomic location ranges by introducing genomic region binning, thereby achieving accurate binning of sequencing reads and improving the accuracy of UMI correction.
[0042] In some embodiments, before parsing each segment to be bucketed in each processing unit, the method further includes: Obtain the chain direction information corresponding to each segment to be bucketed; Bucketing features are generated based on the shearing site information of the segments to be bucketed, including: Bucketing features are generated based on the cell identifier sequence, chain orientation information, and splicing site information of the read segments to be bucketed; Based on the comparison location information of the segments to be bucketed, bucketing features are generated, including: Bucketing features are generated based on the cell identifier sequence, chain direction information, and alignment position information of the read segment to be bucketed.
[0043] In this embodiment of the application, when constructing the bucketing feature, in addition to obtaining the splice signature or genomic region bucketing of the read segment to be bucketed, the chain direction information of each read segment to be bucketed can also be obtained, wherein the chain direction information can also be obtained from the CIGAR information of the read segment to be bucketed.
[0044] After obtaining the chain direction information of each read segment to be bucketed, for read segments with splicing, bucketing features can be generated based on chain direction information, cell identifier sequences, and splicing signatures; for read segments without splicing, corresponding bucketing features can be generated based on chain direction information, cell identifier sequences, and genomic region bucketing.
[0045] In some embodiments, for reads with gene annotation information, the gene annotation information corresponding to the read to be bucketed can be further used to assist in bucketing, thereby obtaining a more accurate bucketing result. In summary, in this embodiment, the reads to be bucketed in each processing unit can be divided into two parts: those with annotation information and those without. For reads with annotation information, the bucketing key can be constructed as follows: Key = (CB, GX, strand, SJ), where CB represents the cell identifier sequence, GX represents the gene annotation information, strand represents the strand direction information, and SJ represents the splice signature. For reads without annotation information, the bucketing key can be constructed as follows: Key = (CB, strand, SJ) or (CB, strand, no_sj|locus). Here, no_sj|locus indicates that the read to be bucketed corresponds to a genomic region bucket and that the read to be bucketed has no splicing. In some embodiments, the key may also include polyA / joint hit flags, anchor length grading, splice site canonicality (GT-AG / GC-AG / AT-AC) weights as parallel or secondary keys to improve the robustness of bucketing. And in some embodiments, In some embodiments, binning features are generated based on the cell identifier sequence, chain orientation information, and splicing site information of the read segment to be binned, including: Extract the start and end coordinates of each contained sub-interval in the segment to be divided into buckets; The start and end coordinates are quantized by jitter according to a preset base length to obtain quantized coordinates; The binning features of the reads to be binned are generated based on the quantized coordinates and the cell identifier sequence and chain direction information of the reads to be binned.
[0046] In long-read sequencing, alignment jitter of 1-10 bp at splice sites can cause reads of the same molecule to be incorrectly classified into different splice structures. In this embodiment, a coordinate jitter quantization method is used when generating splice signatures to ensure that reads at the same true splice site are grouped into the same bin, effectively avoiding inaccurate binning caused by splicing instability.
[0047] Specifically, when generating binning features based on the shearing site information of the read segment to be binned, the start and end coordinates corresponding to each intron interval in the read segment to be binned can be extracted first. Then, the start and end coordinates are jittered and quantized according to the preset base length (jitter) to obtain quantized coordinates. Finally, the binning features of the read segment to be binned are generated based on the quantized coordinates.
[0048] For example, the following is a schematic diagram of the original coordinates of several introns: Intron A: chr1:98-210; Intron B: chr1: 102-205; Intron C: chr1: 105-199.
[0049] Therefore, when performing jitter quantization in 10bp units, the general quantization calculation rule is as follows: Quantized coordinates = round(original coordinates / jitter) * jitter.
[0050] Thus, the coordinates of each intron after quantization are: Intron A: chr1: 100-210; Intron B: chr1: 100-210; Intron C: chr1: 110-200.
[0051] In some embodiments, the specific base length of the jitter can be adaptively adjusted based on mapping quality (MAPQ), anchoring length, and sequencing platform (e.g., 5 bp for high MAPQ and 15–20 bp for low MAPQ). In some embodiments, the size of the locus-bin can be controlled by the locus-bin parameter, for example, it can be set to 1000 bp by default. Alternatively, the bin can be dynamically set according to the local gene density / repetitive sequence density (e.g., reduced to 200–500 bp in human HLA regions, and relaxed to 2–5 kb in sparse regions).
[0052] In this embodiment of the application, by introducing coordinate jitter tolerance during splicing signature generation, the problem of inaccurate bucketing caused by splicing instability can be avoided, thereby improving the accuracy of UMI correction.
[0053] Step S104: Bucket the multiple read segments to be bucketed according to the bucketing characteristics, and correct the molecular identifier sequence of the read segments to be bucketed in each processing unit according to the bucketing results.
[0054] Once the bucketing features corresponding to each read segment to be bucketed are constructed based on the splice signature or genomic region bucketing of each read segment to be bucketed, multiple read segments to be bucketed can be bucketed based on the bucketing features corresponding to each read segment to be bucketed, and the bucketing results can be obtained.
[0055] After binning multiple reads within each cell's corresponding processing unit, the molecular identifier sequences of the reads within each processing unit can be further corrected based on the binning results. Specifically, the reads within each processing unit are binned according to their binning bonds, allowing reads with the same binning bond to be grouped into the same bin, thus obtaining the binning results.
[0056] In some embodiments, multiple read segments to be bucketed are bucketed according to bucketing characteristics, and the molecular identifier sequence of the read segments to be bucketed in each processing unit is corrected according to the bucketing results, including: Based on the binning characteristics, multiple read segments to be binned are binned to obtain multiple sets of read segments to be corrected. In each set of reads to be corrected, the molecular identifier sequence of each read to be corrected is extracted, and the molecular identifier sequences are aggregated according to the similarity between them to obtain multiple aggregate groups, so as to correct the molecular identifier sequences of multiple reads to be corrected corresponding to each aggregate group.
[0057] In this embodiment, by using the binning features of each read segment to be binned as described above, multiple read segments to be binned in the processing unit corresponding to each cell are binned, resulting in multiple read sets. These binned read sets can be called read sets to be corrected, and each read set to be corrected contains multiple reads to be UMI corrected. Next, the UMI of the sequence reads in each read set to be corrected can be corrected.
[0058] Specifically, after binning the sequencing reads in each cell's processing unit according to binning keys constructed based on splicing signatures or genomic regions to obtain multiple sets of reads to be corrected, UMI extraction and clustering can be performed on the sequence reads in each set of reads to be corrected. This results in the sequence reads contained in each set of reads to be corrected being divided into multiple clusters according to their UMIs. Each cluster corresponds to an accurate UMI, and the UMIs of other sequence reads in that cluster can be corrected based on this accurate UMI.
[0059] Specifically, UMI clustering can be performed on the sequence reads contained in each set of reads to be corrected, which can be based on UMI similarity.
[0060] In some embodiments, clustering is performed based on the similarity between molecular identifier sequences to obtain multiple molecular identifier sequence classes, including: The extracted molecular identifier sequences are clustered based on a preset clustering algorithm to obtain multiple molecular identifier sequence clusters, and the number of molecular identifier sequences corresponding to each molecular identifier sequence cluster is obtained. Calculate the Hamming distance between molecular sequence identifiers corresponding to any two molecular identifier sequence clusters, and integrate multiple molecular sequence identifier clusters based on the Hamming distance and the number of molecular identifier sequences to obtain multiple molecular identifier sequence classes.
[0061] In this embodiment of the application, in order to further improve the accuracy of UMI correction, after clustering the UMIs of multiple reads in each set of reads to be corrected to obtain multiple UMI clusters, the Hamming distance between the cluster center UMIs corresponding to the multiple UMI clusters can be further calculated, and the UMI count corresponding to each UMI cluster can be obtained.
[0062] Then, the need for UMI cluster consolidation can be determined based on the Hamming distance and the UMI count corresponding to each UMI cluster. Specifically, the smaller cluster can be consolidated into the larger cluster when the Hamming distance and UMI count between two UMI clusters satisfy the following conditions: Hamming distance ≤ ham (e.g., 1), and large cluster count ≥ small cluster count × ratio (e.g., 2.0); In short, integration is considered successful when the Hamming distance between the cluster center UMIs of two UMI clusters is less than a preset value (ham), where ham can be 1 or other preset integer values, and the number of UMIs in one UMI cluster is greater than a certain multiple (e.g., greater than 2 times) of the number of UMIs in the other UMI cluster. This condition requires that the similarity between the cluster center UMIs of the two UMI clusters be sufficiently high, and that there be a certain difference in cluster counts before integration. Specifically, since UMI sequences are typically tens of bp in length, and the sequencing error rate for long-read sequencing is around 10%, the number of erroneous bases in a UMI sequence is most likely only one base. Therefore, if the Hamming distance between the cluster center UMIs is 1, this difference is likely due to sequencing errors from long-read sequencing. Furthermore, based on the sequencing error rate, it is known that due to the difference in the number of UMI sequences with and without sequencing errors, the number of correct UMI sequences is significantly greater than the number of UMI sequences with sequencing errors.
[0063] By further integrating the multiple UMI clusters obtained by clustering using the above method, we can avoid clustering a small number of sequencing errors UMI sequences separately, thereby improving the accuracy of UMI correction.
[0064] In some embodiments, a directional adjacency algorithm can be used to construct a UMI similarity network, and then the UMIs can be clustered according to the UMI similarity network to obtain multiple UMI classes.
[0065] Specifically, aggregation is performed based on the similarity between molecular identifier sequences to obtain multiple aggregation groups, including: A molecular identifier sequence similarity graph is constructed based on the molecular identifier sequences of multiple reads to be corrected. The molecular identifier sequence similarity graph includes multiple molecular identifier sequence nodes and multiple candidate edges. The distance between the molecular identifier sequence nodes connected by the candidate edges is less than a preset distance. Calculate the merging probability of candidate edges and merge the candidate edges based on the merging probability to update the molecular identifier sequence similarity graph; Multiple clusters were identified based on the updated molecular representation sequence similarity map.
[0066] In this embodiment, the original UMIs can be constructed into a UMI similarity graph within the same cell, strand direction, splice signature, or genomic region bins. Different UMIs serve as graph nodes, and UMI pairs that satisfy preset Hamming distance or edit distance conditions are considered candidate edges. For each candidate edge, instead of solely using a fixed read count fold change threshold to determine whether to merge, the likelihood, likelihood ratio, or Bayesian posterior probability derived from errors in high-support UMIs by combining the distance between UMIs, read support count, sequencing error primacy, base quality, UMI count distribution within the bin, and local UMI density is calculated. Based on this, it is determined whether UMIs are merged and their merging direction.
[0067] Furthermore, in another optional implementation, the calculated posterior probabilities can be used to assign weights to candidate edges, and a directed weighted graph can be constructed based on node abundance differences (e.g., from low abundance to high abundance). For connected components in the graph composed of multiple similar UMIs, a graph optimization algorithm is used to determine global consistency. Specifically, algorithms such as directed adjacency propagation, maximum posterior assignment, maximum spanning tree, minimum cut, or minimum cost flow can be used to resolve the affiliation relationships between multiple candidate UMIs, thereby obtaining a corrected set of UMIs.
[0068] In this way, the traditional local greedy merging based on fixed read fold can be improved into an adaptive discrimination process based on error probability, count distribution and graph structure. This can more accurately distinguish between real independent molecules and UMI error-derived molecules, reduce the risk of excessive merging in high-expression or densely populated UMI regions, reduce the false retention of low-support error UMIs, and improve the accuracy, robustness and applicability of UMI deduplication in single-cell long-read sequencing.
[0069] In some embodiments, the process of determining candidate edges includes: Based on a preset distance, a fast nearest neighbor search is performed on the molecular identifier sequences of multiple read segments to be corrected to obtain multiple nearest neighbor molecular identifier sequence pairs; Candidate edges are determined based on the nearest neighbor molecular identifier sequence pairs.
[0070] In this embodiment, to avoid comparing all UMIs pairwise, a fast retrieval index can be established for each unique UMI within each processing unit. For fixed-length UMIs that are mainly replacement errors, a hash table combined with neighborhood enumeration can be used: when the preset distance (distance threshold) is 1, single-base substitutions at each site are enumerated sequentially; when the distance threshold is 2, joint substitutions at two sites are further enumerated, and pruning is performed using count, quality, or whitelist information to quickly obtain nearest neighbor candidates. In some embodiments, for the most common scenario where ham is 1, to further improve the efficiency of UMI correction, this application also provides a fast adjacency algorithm based on enumerated adjacency (fast_h1). By enumerating all candidate neighbors of single-base variations and performing hash queries, the computational complexity can be greatly reduced, and the efficiency of UMI sequence similarity comparison can be improved. Specifically, using the UMI adjacency algorithm in related technologies requires pairwise comparisons with O(K²) complexity within each UMI bucket obtained by bucketing based on bucketing features. However, the large number of UMIs within a UMI bucket leads to high computational complexity. The enumerated adjacency algorithm provided in this application can reduce the computational complexity to O(K·L), where K is the number of UMI sequences in the bucket and L is the sequence length of the UMI sequence, and K is much larger than L.
[0071] In some embodiments, aggregation is performed based on the similarity between molecular identifier sequences to obtain multiple aggregate groups, including: Based on a preset distance, a fast nearest neighbor search is performed on the molecular identifier sequences of multiple read segments to be corrected to obtain the similarity relationship between the molecular identifier sequences; Multiple aggregates were identified based on the similarity between molecular identifier sequences.
[0072] In this embodiment, when aggregating molecular identifier sequences based on their similarity, the aforementioned fast nearest neighbor search method can also be used to determine the similarity between molecular identifier sequences. Specifically, a distance threshold between molecular identifier sequences can be determined first (e.g., a preset distance ham=1 or ham=2, etc.). Then, it can be determined that molecular identifier sequences are similar when the distance between them is less than or equal to the preset distance, and dissimilar otherwise. Then, a fast nearest neighbor search can be performed on the molecular identifier sequences of multiple reads to be corrected based on the preset distance to obtain the similarity relationship between the molecular identifier sequences. Then, these multiple molecular identifier sequences can be further aggregated according to the similarity relationship to obtain multiple aggregate groups.
[0073] In another implementation, an inverted index can be built using single-position or double-position mask keys to quickly filter potentially similar UMIs; alternatively, UMIs can be encoded in binary form, and search efficiency can be improved using XOR, bit counting, or parallel comparison methods. For scenarios requiring support for complex errors such as insertions and missing values, a BK-tree or other tree-based index structure based on string distance can be used for nearest neighbor retrieval; for extremely large datasets, coarse screening can be performed first using locality-sensitive hashing or signature bucketing, followed by precise comparison within the candidate set. The above indexing methods can be adaptively selected based on UMI length, error type, data size, and the existence of a whitelist.
[0074] In some embodiments, before performing adjacency calculation on UMI sequences in the same bucket, the UMI sequences can be projected and mapped according to a pre-set whitelist or error correction mapping table to obtain sequences that are determined to be accurate UMIs, and then adjacency calculation and cluster integration can be performed to avoid erroneous integration.
[0075] In some embodiments, after binning multiple read segments to be binned according to binning characteristics and correcting the molecular identifier sequence of the read segments to be binned in each processing unit according to the binning results, the method further includes: Identify multiple deduplication reads corresponding to each corrected molecular identifier sequence, and obtain the alignment quality and alignment sequence length of each deduplication read; Based on the alignment quality and alignment sequence length, multiple deduplication reads corresponding to each corrected molecular identifier sequence are deduplicated to obtain the target read corresponding to each corrected molecular identifier sequence.
[0076] In this embodiment of the application, a method for deduplicating sequencing data is also provided, such as... Figure 2 The diagram shown is a flowchart of the sequencing data deduplication method provided in this application. The method includes: Step S101: Obtain single-cell long-read sequencing data of the target sample.
[0077] Step S102: Extract the cell identifier sequence corresponding to each sequence read, and divide the multiple sequence reads into multiple processing units based on the cell identifier sequence.
[0078] Step S103: In each unit to be processed, construct the bucketing features of each segment to be bucketed.
[0079] Step S104: Bucket the multiple read segments to be bucketed according to the bucketing characteristics, and correct the molecular identifier sequence of the read segments to be bucketed in each processing unit according to the bucketing results.
[0080] Steps S101 to S104 have been described in detail above and will not be repeated here.
[0081] Step S205: Determine multiple deduplication reads corresponding to each corrected molecular identifier sequence, and obtain the alignment quality and alignment sequence length of each deduplication read.
[0082] In this embodiment, to accurately quantify gene expression in the target sample, after performing UMI correction on multiple sequence reads obtained from single-cell long-read sequencing of the target sample, multiple corrected UMI sequences corresponding to each cell can be obtained. These corrected UMI sequences can be referred to as corrected molecular identifier sequences. Then, for each corrected molecular identifier sequence corresponding to each cell, multiple corresponding sequence reads can be determined, which can be referred to as multiple deduplication reads. Due to the existence of long-read sequencing errors, there are certain differences between the multiple deduplication reads corresponding to each corrected molecular identifier sequence. This embodiment can further obtain the alignment quality and alignment sequence length of each deduplication read to further deduplicate the multiple deduplication reads corresponding to each corrected molecular identifier sequence based on the alignment quality and alignment sequence length of each deduplication read.
[0083] Specifically, the alignment quality and alignment sequence length of each deduplicated read can be obtained from the BAM file obtained by aligning each single-cell long-read sequencing read to the reference genome.
[0084] Step S206: Based on the alignment quality and alignment sequence length, deduplication is performed on multiple deduplication reads corresponding to each corrected molecular identifier sequence to obtain the target read corresponding to each corrected molecular identifier sequence.
[0085] After obtaining the alignment quality and alignment sequence length of each read segment to be deduplicated, multiple read segments to be deduplicated can be sorted according to the alignment quality and alignment sequence length. Then, the best read at the top of the sorted list is retained as the target read segment, and the other read segments to be deduplicated are filtered out.
[0086] Multiple reads to be deduplicated are sorted based on alignment quality and alignment sequence length. Specifically, scores are generated for each read based on alignment quality and alignment sequence length, and then a total score is calculated by weighting these scores according to preset weight values. The reads are then sorted based on this total score. Finally, the corrected UMI is written to the UB tag and the DA tag (1 indicates retention of the representative read, 0 indicates removal of duplicate reads). The deduplicated BAM file (containing only the target read, or retaining all reads and marking them with DA in --write-all mode) and molecules.tsv are output, recording the cell, strand, SJ / no-sj, UB_corrected, count, and UR for each unique molecule; assignments.tsv is output, recording the cell, SJ / no-sj, UR_raw, UB_corrected, and whether each read is retained. The multi-slice results are then merged to generate the final BAM and TSV files.
[0087] In one specific embodiment, using a single-cell long-read transcriptome dataset as input, the sequencing data processing method provided in this application is employed to construct SJ signatures and perform locus-bin binning in an annotation-free mode. In the ham=1 scenario, the fast_h1 optimization algorithm is used to optimize the processing of sequences with more than 10 reads per bucket. 5 Even with multiple UMIs, correction can still be completed within minutes, significantly improving efficiency. The final deduplicated BAM file is approximately [size missing], and molecules.tsv and assignments.tsv are output, which can be directly used for downstream expression matrix construction and traceability analysis.
[0088] In summary, the sequencing data processing method provided in this application acquires single-cell long-read sequencing data of the target sample, which includes multiple sequence reads; extracts the cell identifier sequence corresponding to each sequence read, and divides the multiple sequence reads into multiple processing units based on the cell identifier sequence, with each processing unit containing multiple reads to be binned; in each processing unit, constructs binning features for each read to be binned, which are generated based on the splicing site information or alignment position information of the read to be binned; bins the multiple reads to be binned based on the binning features, and corrects the molecular identifier sequence of the reads to be binned in each processing unit based on the binning results.
[0089] Therefore, the sequencing data processing method provided in this application, when performing UMI correction based on single-cell long-read sequencing data, can first divide the sequence reads into bucket sets corresponding to each cell based on the barcode sequence of the sequence reads. Then, for each cell's bucket set, the reads are bucketed based on the splicing features or alignment positions of the reads. Finally, the UMI of each read is corrected based on the bucketing results. This method can accurately bucket sequence reads lacking gene annotation, thereby achieving UMI correction under conditions without gene annotation and improving the accuracy of sequencing data processing.
[0090] Reference Figure 3 In some embodiments, this application also provides a sequencing data processing apparatus 300, which includes: The acquisition unit 310 is used to acquire single-cell long-read sequencing data of the target sample, wherein the single-cell long-read sequencing includes multiple sequence reads; The partitioning unit 320 is used to extract the cell identifier sequence corresponding to each sequence read segment, and to divide the multiple sequence read segments into multiple processing units according to the cell identifier sequence, wherein each processing unit contains multiple read segments to be bucketed. The construction unit 330 is used to construct the binning feature of each read segment to be binned in each of the processing units. The binning feature is generated based on the shearing site information of the read segment to be binned or the alignment position information of the read segment to be binned. The correction unit 340 is used to divide the plurality of read segments to be divided into buckets according to the bucketing characteristics, and to correct the molecular identifier sequence of the read segments to be divided into buckets in each of the processing units according to the bucketing results.
[0091] Optionally, in some embodiments, the building unit includes: The parsing subunit is used to parse each of the read segments to be bucketed and obtain the chain direction information corresponding to each read segment to be bucketed in each of the processing units. The first generation subunit is used to generate binning features based on the cell identifier sequence of the read segment to be binned, the chain direction information, and the splicing site information when the read segment to be binned contains splicing information. The second generation subunit is used to generate binning features based on the cell identifier sequence of the read segment to be binned, the chain direction information, and the alignment position information when the read segment to be binned does not contain splicing information.
[0092] Optionally, in some embodiments, the first generating subunit includes: The extraction module is used to extract the start and end coordinates of each contained sub-interval in the segment to be divided into buckets; The quantization module is used to perform dithering quantization on the start and end coordinates according to a preset base length to obtain quantized coordinates; The generation module is used to generate the bucketing features of the read segment to be bucketed based on the quantized coordinates, the cell identifier sequence of the read segment to be bucketed, and the chain direction information.
[0093] Optionally, in some embodiments, the correction unit includes: The bucketing subunit is used to bucket the multiple read segments to be bucketed according to the bucketing characteristics, so as to obtain multiple sets of read segments to be corrected. The correction subunit is used to extract the molecular identifier sequence of each read segment to be corrected from each set of read segments to be corrected, and to aggregate the molecular identifier sequences according to the similarity between the molecular identifier sequences to obtain multiple aggregate groups, so as to correct the molecular identifier sequences of multiple read segments to be corrected corresponding to each aggregate group.
[0094] Optionally, the correction subunit includes: The acquisition module is used to acquire the number of molecular identifier sequences in each of the aggregate groups; An integration module is used to calculate the distance between molecular sequence identifiers corresponding to any two aggregation groups, and to integrate the multiple aggregation groups according to the distance and the number of molecular identifier sequences.
[0095] Optionally, in some embodiments, the correction subunit includes: The construction module is used to construct a molecular identifier sequence similarity graph based on the molecular identifier sequences of the plurality of reads to be corrected. The molecular identifier sequence similarity graph includes a plurality of molecular identifier sequence nodes and a plurality of candidate edges, wherein the distance between the molecular identifier sequence nodes connected by the candidate edges is less than a preset distance. The merging module is used to calculate the merging probability of the candidate edges and merge the candidate edges based on the merging probability to update the molecular identifier sequence similarity graph. The determination module is used to identify multiple clusters based on the updated molecular representation sequence similarity map.
[0096] Optionally, in some embodiments, the correction subunit includes: The search subunit is used to perform a fast nearest neighbor search on the molecular identifier sequences of multiple read segments to be corrected based on a preset distance, so as to obtain the similarity relationship between the molecular identifier sequences; Subunits are identified to determine multiple aggregate groups based on the similarity between molecular identifier sequences.
[0097] Optionally, in some embodiments, the sequencing data processing apparatus provided in this application further includes: The second determining subunit is used to determine multiple deduplication sequence reads corresponding to each corrected molecular identifier sequence, and to obtain the alignment quality and alignment sequence length of each deduplication read. The deduplication subunit is used to deduplicate multiple read segments corresponding to each corrected molecular identifier sequence according to the alignment quality and alignment sequence length, so as to obtain the target read segment corresponding to each corrected molecular identifier sequence.
[0098] Reference Figure 4 , Figure 4 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 401 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0099] The memory 402 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 402 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 402 and is called and executed by the processor 401 to execute the sequencing data processing method of the embodiments of this application.
[0100] Input / output interface 403 is used to implement information input and output.
[0101] The communication interface 404 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0102] Bus 405 transmits information between various components of the device, such as processor 401, memory 402, input / output interface 403, and communication interface 404.
[0103] The processor 401, memory 402, input / output interface 403 and communication interface 404 are connected to each other within the device via bus 405.
[0104] This application also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, causing the computer device to perform the sequencing data processing method described above.
[0105] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “including,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.
[0106] It should be understood that in this disclosure, "at least one item" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0107] It should be understood that in the description of the embodiments of this application, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.
[0108] In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0109] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0110] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0111] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0112] It should also be understood that the various implementation methods provided in this application can be combined arbitrarily to achieve different technical effects.
[0113] The above is a detailed description of the embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.
Claims
1. A sequencing data processing method, characterized in that, The method includes: Obtain single-cell long-read sequencing data of the target sample, wherein the single-cell long-read sequencing includes multiple sequence reads; Extract the cell identifier sequence corresponding to each sequence read segment, and divide the multiple sequence read segments into multiple processing units according to the cell identifier sequence. Each processing unit contains multiple read segments to be bucketed. In each of the processing units, a binning feature is constructed for each of the read segments to be binned. The binning feature is generated based on the cut-off point information of the read segment to be binned or the alignment position information of the read segment to be binned. The multiple read segments to be bucketed are bucketed according to the bucketing characteristics, and the molecular identifier sequence of the read segments to be bucketed in each processing unit is corrected according to the bucketing results.
2. The method according to claim 1, characterized in that, In each of the processing units, constructing the bucketing features for each segment to be bucketed includes: In each of the processing units, each read segment to be bucketed is parsed and the chain direction information corresponding to each read segment to be bucketed is obtained. When the read segment to be bucketed contains splicing information, bucketing features are generated based on the cell identifier sequence of the read segment to be bucketed, the chain direction information, and the splicing site information; When the read segment to be binned does not contain splicing information, binning features are generated based on the cell identifier sequence of the read segment to be binned, the chain direction information, and the alignment position information.
3. The method according to claim 2, characterized in that, The step of generating binning features based on the cell identifier sequence of the read segment to be binned, the chain direction information, and the splicing site information includes: Extract the start and end coordinates of each contained sub-interval in the segment to be divided into buckets; The start and end coordinates are jitter-quantized according to a preset base length to obtain quantized coordinates; The binning features of the read segment to be binned are generated based on the quantized coordinates, the cell identifier sequence of the read segment to be binned, and the chain direction information.
4. The method according to claim 1, characterized in that, The step of binning the multiple read segments to be binned according to the binning characteristics, and correcting the molecular identifier sequence of the read segments to be binned in each processing unit according to the binning results, includes: The multiple read segments to be bucketed are bucketed according to the bucketing characteristics to obtain multiple sets of read segments to be corrected. In each set of reads to be corrected, the molecular identifier sequence of each read to be corrected is extracted, and the molecular identifier sequences are aggregated according to the similarity between them to obtain multiple aggregate groups, so as to correct the molecular identifier sequences of multiple reads to be corrected corresponding to each aggregate group.
5. The method according to claim 4, characterized in that, The aggregation based on the similarity between the molecular identifier sequences yields multiple aggregation groups, including: Obtain the number of molecular identifier sequences in each of the aggregate groups; Calculate the distance between the molecular sequence identifiers corresponding to any two aggregate groups, and integrate the plurality of aggregate groups based on the distance and the number of molecular identifier sequences.
6. The method according to claim 4, characterized in that, The aggregation based on the similarity between the molecular identifier sequences yields multiple aggregation groups, including: A molecular identifier sequence similarity graph is constructed based on the molecular identifier sequences of the multiple reads to be corrected. The molecular identifier sequence similarity graph includes multiple molecular identifier sequence nodes and multiple candidate edges. The distance between the molecular identifier sequence nodes connected by the candidate edges is less than a preset distance. Calculate the merging probability of the candidate edges, and merge the candidate edges based on the merging probability to update the molecular identifier sequence similarity graph; Multiple clusters were identified based on the updated molecular representation sequence similarity map.
7. The method according to claim 4, characterized in that, The aggregation based on the similarity between the molecular identifier sequences yields multiple aggregation groups, including: Based on a preset distance, a fast nearest neighbor search is performed on the molecular identifier sequences of the multiple read segments to be corrected to obtain the similarity relationship between the molecular identifier sequences; Multiple aggregation groups are determined based on the similarity between the molecular identifier sequences.
8. The method according to claim 1, characterized in that, After binning the plurality of read segments to be binned according to the binning characteristics, and correcting the molecular identifier sequence of the read segments to be binned in each processing unit according to the binning results, the method further includes: Identify multiple deduplication reads corresponding to each corrected molecular identifier sequence, and obtain the alignment quality and alignment sequence length of each deduplication read; Based on the alignment quality and alignment sequence length, multiple deduplication reads corresponding to each corrected molecular identifier sequence are deduplicated to obtain the target read corresponding to each corrected molecular identifier sequence.
9. A sequencing data processing device, characterized in that, The device includes: The acquisition unit is used to acquire single-cell long-read sequencing data of the target sample, wherein the single-cell long-read sequencing includes multiple sequence reads; A partitioning unit is used to extract the cell identifier sequence corresponding to each sequence read segment, and to divide the multiple sequence read segments into multiple processing units according to the cell identifier sequence, wherein each processing unit contains multiple read segments to be binned. A construction unit is used to construct a bucketing feature for each read segment to be bucketed in each of the processing units. The bucketing feature is generated based on the cut-off site information of the read segment to be bucketed or the alignment position information of the read segment to be bucketed. The correction unit is used to divide the multiple read segments to be divided into buckets according to the bucketing characteristics, and to correct the molecular identifier sequence of the read segments to be divided into buckets in each of the processing units according to the bucketing results.
10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the sequencing data processing method according to any one of claims 1 to 8.
11. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the sequencing data processing method according to any one of claims 1 to 8.