A method for assembling genome sequences in single-cell microbial genomes

By automatically calculating the number of assembly wheels, grouping the similarity matrix and dynamically adjusting the similarity threshold, the problems of cumbersome processes, difficult to control the pollution, and high memory consumption in the existing single-cell microbial genome assembly methods are solved, efficient and accurate assembly results are achieved, and the flexibility of the method is improved.

CN119694397BActive Publication Date: 2025-05-06MOBIDROP (ZHEJIANG) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510221583.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-06
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The analysis process of existing single-cell microbial genome assembly methods is cumbersome, unable to effectively control the pollution in the bin, has a large memory consumption, and cannot dynamically adjust the similarity threshold to adapt to different sample types or library quality.

Method used

The method of automatically calculating the number of assembly wheels is used to simplify the process and only one task submission is required; the similarity matrix is ​​calculated by grouping to reduce memory consumption; the similarity threshold is dynamically adjusted to adapt to different sample types or library quality; only the high reliability contig is included for assembly.

Benefits of technology

It improves operational convenience and efficiency, ensures the accuracy of assembly results, reduces memory consumption, improves computing speed, and enhances the flexibility and applicability of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119694397B_ABST
    Figure CN119694397B_ABST
Patent Text Reader

Abstract

The present invention provides a method for assembling genome sequences in a single-cell microorganism genome. The method can 1) automatically calculate the number of assembly rounds, and only one task submission is required for the entire process; 2) control the contamination of bins at a low level for subsequent analysis; 3) calculate the similarity matrix in a grouping manner to reduce memory consumption when the number of SAGs is high and improve the calculation speed; 4) when calculating the similarity, only contigs with higher reliability are included; and the similarity threshold can be dynamically selected or adjusted according to the actual analysis situation to adapt to different sample types or library quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of bioinformatics, and in particular relates to a method for assembling genome sequences in a single-cell microorganism genome. Background Art

[0002] The relationship between microorganisms and humans is complex and close. Microorganisms can be pathogens that cause diseases, such as bacteria and viruses. However, most microorganisms are beneficial to humans, and they play an important role in food production (such as fermentation), medicine (such as the production of antibiotics), and environmental protection (such as sewage treatment). Different microorganisms play different biological functions through their different gene sequences. Through sequencing, microbial genes can now be detected on a large scale, so as to deeply understand their mechanism and significance.

[0003] Single-cell microbial genomics technology (microbe-seq) is a technical means of high-throughput detection of the entire gene sequence of a single microorganism using a single-cell sequencing method. It was published in Science in 2022 (Wenshan Zheng et al., High-throughput, single-microbe genomics with strain resolution, applied to a human gut microbiome. Science376, eabm1483 (2022). DOI: 10.1126 / science.abm1483). This technology uses a variety of droplet microfluidic operation technologies combined with bioinformatics analysis methods to obtain the genome information of thousands of single-cell microorganisms from complex microbial communities without cultivation. It can also efficiently distinguish similar species based on the reads of microorganisms in each droplet in the single-cell microbial genome sequencing results, distinguish strains through mutations in individual microorganisms, detect interactions between microorganisms and viruses, and discover horizontal gene transfer events in microbial populations. Compared with previous microbial genome detection technologies, single-cell microbial genome technology has significantly improved in throughput and resolution. It can not only quickly identify different strains, but also provide more detailed biological characteristics analysis based on the characteristic genes of each strain.

[0004] In the above article, a method for assembling species-level genome sequences based on single-cell microbial genome sequencing data is described. Specifically, the collection of SAG (single-amplified genome, referring to the droplets containing microorganisms and their amplified gene sequences) is called a bin, and each SAG forms a separate bin in the initial state, and then the following process is executed in a loop: use SPAde to assemble all reads of the SAG in the bin, and then use sourmash to calculate the similarity between the assembly results of the bin, and merge the bins with similarity greater than the similarity threshold of 5%; until a certain number of cycles are repeated, the species-level genome sequence without filtering contamination is obtained.

[0005] However, the above genome sequence assembly methods still have the following problems:

[0006] (1) The analysis process is cumbersome and requires manual specification of the number of assembly rounds, and each round of analysis requires manual submission of tasks;

[0007] (2) It is impossible to control the contamination level within the bin. If the contamination level is too high, it is difficult to remove the contamination from the assembly result.

[0008] (3) If the sequencing data contains a high number of SAGs, the memory consumption will be too large, slowing down the calculation speed;

[0009] (4) When calculating bin similarity, too much contamination may be included; and the similarity threshold cannot be selected dynamically. Different sample types or library qualities require manual adjustment of parameters. Summary of the invention

[0010] In order to solve the above problems, the present invention provides an optimized method for assembling genome sequences in single-cell microbial genomes, which can 1) automatically calculate the number of assembly rounds, and only one task submission is required for the whole process; 2) control the contamination of bins at a low level for subsequent analysis; 3) calculate the similarity matrix in a grouping manner to reduce memory consumption when the number of SAGs is high and improve the calculation speed; 4) only include contigs with higher reliability when calculating similarity; and the similarity threshold can be dynamically selected or adjusted according to the actual analysis situation to adapt to different sample types or library qualities.

[0011] The technical solution adopted by the present invention is: a method for assembling a genome sequence in a single-cell microorganism genome, comprising:

[0012] S1. Obtain a single-cell microbial genome sequencing library, generate a bin for each SAG, and initialize the bin list;

[0013] S2. Assemble each bin according to the k-mer set in the current assembly round, and evaluate the completeness and contamination of the assembly result of each bin to obtain assembly information;

[0014] S3. If the assembly information satisfies the assembly termination condition, the assembly result of each bin of the current assembly round is output, otherwise step S4 is executed;

[0015] S4. Group all bins, hierarchically cluster the bins in the group according to the similarity between the two bins; split the hierarchical clustering results into one or more clusters according to the dynamically adjusted similarity threshold until the number of clusters in the group or the similarity threshold meets the preset split termination condition; merge the bins in a single cluster in the group into a new bin, and then summarize the new bins of all groups to obtain the bin list for the next assembly round;

[0016] S5. If the bin list of the next assembly round has the same number of bins as the bin list of the current assembly round, output the assembly result of each bin of the current assembly round; otherwise, update the k-mer and repeat steps S2 to S4.

[0017] Preferably, in step S2, the length of the k-mer is K, the value range of K is 21-61 and K is an odd number.

[0018] Preferably, the assembly termination condition is any one of the following (1)-(2):

[0019] (1) The ratio of the number of SAGs corresponding to bins with integrity greater than 0.5 to the total number of SAGs is greater than the preset integrity threshold C;

[0020] (2) The number of assembly rounds is greater than the preset maximum number of assembly rounds.

[0021] Preferably, step S4 comprises:

[0022] S4-1. Filter the assembly results of each bin to obtain the filtered assembly results of each bin;

[0023] S4-2. Divide all bins into multiple groups, and calculate the similarity between bins in each group based on the filtered assembly results of each bin and the k-mer set in the current assembly round;

[0024] S4-3. Modify the similarity according to the completeness and contamination of the assembly result of each bin, perform hierarchical clustering on the bins in the group according to the modified similarity, and obtain the hierarchical clustering result of the group;

[0025] S4-4. Split the hierarchical clustering results of the group into one or more clusters according to the similarity threshold, and calculate the ratio R of the number of all clusters in the group to the number of all bins;

[0026] S4-5. If R is less than the ratio threshold, execute step S4-6; otherwise, lower the similarity threshold, if the lowered similarity threshold is less than the minimum similarity, execute step S4-6, otherwise repeat step S4-4;

[0027] S4-6. Merge the bins in a single cluster within a group into a new bin, then aggregate the new bins of all groups and update the bin list.

[0028] Preferably, in step S4-1, the method for filtering the assembly results of each bin includes: removing contigs in the assembly results of the bin whose length is less than a length threshold; the length threshold is one L of the length of Contig N50 in the assembly results.

[0029] Preferably, the value range of L is 5~20.

[0030] Preferably, in step S4-3, the method for modifying the similarity includes: if the integrity of a bin in the group is not less than 0.9 or the contamination is not less than 0.5, then modifying the similarity between the bin and any bin in the group to 0.

[0031] Preferably, step S4-6 also includes: if the number of bins in a cluster within the group is greater than the merging threshold, the cluster is split into multiple small clusters, and then the bins in a single small cluster are merged into a new bin; wherein the number of bins in the small cluster is not greater than the merging threshold.

[0032] Preferably, in step S5, the method for updating the k-mer includes: updating the k-mer according to the current assembly round number and the k-mer attenuation step length.

[0033] The present invention also provides a device for assembling a genome sequence in a single-cell microorganism genome, comprising a processor and a memory; the processor and the memory are connected via a communication bus; wherein the processor is used to call and execute a program stored in the memory; the memory is used to store a program, and the program is at least used to execute the method for assembling a genome sequence in a single-cell microorganism genome.

[0034] Beneficial effects of the present invention: Compared with the prior art, the present invention 1) reduces human intervention and improves the convenience and efficiency of operation by automatically calculating the number of assembly rounds and simplifying the process; 2) can effectively control the contamination of bins to ensure the accuracy of assembly results; 3) by grouping and calculating the similarity matrix, the memory consumption in the case of high SAG numbers is significantly reduced, while the calculation speed is improved, solving the problems existing in large-scale data processing; 4) can dynamically select or adjust the similarity threshold to adapt it to different sample types or library qualities, thereby enhancing the flexibility and applicability of the method. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 A schematic diagram of a simplified process of assembling a genome sequence in a single-cell microbial genome provided in an embodiment of the present invention.

[0036] Figure 2 A schematic diagram of a detailed process of a method for assembling a genome sequence in a single-cell microbial genome provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0037] The following describes the embodiments of the present invention through specific embodiments, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0038] The embodiment of the present invention provides a method for assembling a genome sequence in a single-cell microorganism genome. The method initializes the bin list after obtaining a single-cell microorganism genome sequencing library. Then, each bin is assembled according to the set k-mer and the integrity and contamination of the assembly result are evaluated. The assembly result is filtered to remove shorter contigs to control the contamination. Then, the bins are grouped and the similarity between the bins in the group is calculated. The similar bins are hierarchically clustered and the hierarchical clustering results are split into one or more clusters according to the similarity threshold. In this process, the similarity threshold is dynamically adjusted to adapt to different sample types or library quality until the number of clusters in the group or the similarity threshold meets the preset split termination condition. Finally, similar bins are merged into new bins according to clustering, and the k-mer value is updated for the next round of assembly. The above assembly process is repeated until the bin list is no longer updated or the assembly meets the preset assembly termination condition. The method reduces memory consumption and improves the efficiency and accuracy of single-cell microorganism genome sequence assembly by automating the process, controlling contamination, and dynamically adjusting the similarity threshold. The specific steps of the method are as follows.

[0039] S1. Obtain a single-cell microbial genome sequencing library and set the current assembly round counter m is 0, each SAG generates a bin, and initializes the bin list:

[0040] ;

[0041] in,

[0042] SAG i For the i No. SAG;

[0043] bin m,i For the m Assembly wheel i No.bin;

[0044] binList m For the m The bin list of the assembled wheel;

[0045] SAGCount is the total number of SAGs in the sequencing library.

[0046] S2. Acquisition bin m,i Corresponding fastq file fasta m,i , and then use SPAde, according to the current m-th assembly round set kmer m ,rightbin m,i Assemble and evaluate using checkm bin m,i The completeness of the assembly result Completeness m,i 、 Pollution Contamination m,i , to obtain assembly information. The length of k-mer is K, the value range of K is 21-61 and K is an odd number.

[0047] S3. If the assembly information meets the assembly termination condition, output the current assembly round bin m,i otherwise, execute steps S4-1 to S4-6; the assembly termination condition is any one of the following (1)-(2):

[0048] (1) The ratio of the number of SAGs corresponding to bins with integrity greater than 0.5 to the total number of SAGs is greater than the preset integrity threshold C;

[0049] (2) The number of assembly rounds m is greater than the preset maximum number of assembly rounds MaxM;

[0050] S4-1. Yes bin m,i Filter the assembly results and remove bin m,i The contigs whose length is less than one-Lth of the length of ContigN50 in the assembly result are obtained. bin m,i The filtered assembly results FilteredFasta m,i. Wherein, the value range of L is 5~20, and in a preferred embodiment, L is 10. Contig N50 is an important statistical indicator in the field of genome sequencing and assembly, which is used to evaluate the quality of genome assembly. Specifically, Contig N50 refers to arranging all contigs of the genome from large to small according to length, and then accumulating the length of these contigs until the total accumulated length reaches half of the total length of all contigs. The length of the last accumulated contig is Contig N50. For example, if the size of a genome is 1M, after the reads obtained by sequencing are spliced ​​into contigs, these contigs are arranged from long to short, and then accumulated in sequence until the accumulated length reaches 500k (i.e. 50% of 1M), then the length of the last accumulated contig is Contig N50. The larger the Contig N50 value, the better the continuity of the assembly, the more long fragments assembled, and thus closer to the real genome structure. Therefore, Contig N50 is generally considered to be an important criterion for measuring the quality of genome splicing results. Therefore, the present invention filters the bin assembly results according to the length of the contig in the bin and the length of Contig N50, discards shorter contigs with poor assembly continuity or integrity, thereby controlling the contamination of the bin at a lower level for subsequent analysis.

[0051] S4-2. Divide all bins into M groups according to the number of N bins in each group; bin m,i The filtered assembly results FilteredFasta m,i And the current mth assembly round set k-mer m , calculated using sourmash's MinHash function bin m,i The minhash feature is then calculated using the Compare function of sourmash bin m,i and bin m,ii (i.e., except for the nth group in the mth assembly round bin m,i Similarity of any bin outside the group;

[0052]

[0053] ;

[0054] in,

[0055] Gm,n The set of bins of group n in the mth assembly round;

[0056] mh m,n,i The minhash feature of bin number i of group number n in the mth assembly round;

[0057] S m,n,i,ii It is the similarity between bin No. i and bin No. ii of group No. n in the mth assembly round.

[0058] In this step, all bins are divided into M groups according to N bins in each group. For example, when N=5000 and n=2, all bins are divided into M groups according to 5000 bins in each group. Then group 2 includes SAGs 5000-9999. The similarities between the 5000 bins in the group are calculated, and then the subsequent merging steps, i.e. steps S4-3 to S4-6, are performed in groups. Compared with the method in the prior art of putting all bins together to calculate the similarity and then merging them, the present invention uses a grouping method to calculate the similarity moment, which can reduce memory consumption when the number of SAGs is high and improve the calculation speed. Moreover, after multiple rounds of co-assembly, it has no effect on the assembly results.

[0059] S4-3. If the integrity of a bin in a group is not less than 0.9 or the contamination is not less than 0.5, the similarity between the bin and any bin in the group is modified to 0; based on the modified similarity, the linkage function of scipy is used to perform hierarchical clustering on the bins in the group to obtain the hierarchical clustering results of the group;

[0060] ;

[0061] in,

[0062] C m,n,i It is the assembly result of bin No. i in group No. n in the mth assembly round.

[0063] In this step, the purpose of modifying the similarity is to exclude bins with relatively high completeness (i.e., completeness greater than or equal to 0.9) or high contamination (i.e., contamination greater than or equal to 0.5) after assembly from subsequent merging to prevent excessive contamination from being included. For bins with high completeness, continued merging will greatly increase their contamination, increasing the difficulty of subsequent decontamination. Bins with high contamination may gather polysomal SAGs or SAGs with high backgrounds, and continued merging will also increase the difficulty of subsequent decontamination. Therefore, this step modifies the similarity between bins with relatively high completeness or high contamination after assembly and any bin in the group to 0, so that the bin will not be split into the same cluster with the remaining bins during the subsequent hierarchical clustering and splitting process, and will no longer be included in the subsequent merging process. In addition, hierarchical clustering of bins within a group in small groups can effectively control memory consumption and increase computing speed.

[0064] S4-4. Set the current split round counter mm to 0, and set the similarity threshold of the current mm split round to SH mm ; Based on the similarity threshold SH mm , use scipy's fcluster function to split the group hierarchical clustering results into one or more clusters (cluster), and calculate the ratio R of the number of all clusters in the group to the number of all bins. In this step, the similarity threshold can be dynamically adjusted according to the split results. Because different sequencing libraries have different coverages, the level of coverage directly affects the similarity, so it is necessary to dynamically adjust the similarity threshold for bin merging in order to 1) adapt to libraries of different plasmids; 2) be able to merge bins of some species with fewer SAG numbers. This solves the problem in the prior art that the similarity threshold cannot be dynamically selected, and different sample types or library qualities require manual adjustment and modification of the similarity threshold.

[0065] S4-5. If R is less than the ratio threshold, execute step S4-6; otherwise, according to the preset similarity step length StepSH Lower the similarity threshold SH mm , , get the similarity threshold for the next splitting round, i.e. the mm+1 splitting round SH mm+1 , if the similarity threshold is lowered SH mm+1 Less than the minimum similarity minSH , then execute step S4-6, otherwise repeat step S4-4.

[0066] S4-6. Merge the bins in a single cluster within a group into a new bin, and then aggregate the new bins of all groups to obtain the bin list for the next assembly round (the m+1th assembly round) binList m+1 ; If the number of bins in a cluster within a group is greater than the merge threshold maxC, the cluster is split into multiple small clusters, and then the bins in a single small cluster are merged into a new bin, where the number of bins in the small cluster is not greater than the merge threshold maxC. For example, the first group in the first assembly round has 5000 bins. After hierarchical clustering, according to the similarity threshold set in the first splitting round SH 1 The split results in 3000 clusters, and the ratio of 3000 clusters to 5000 bins (R=0.6) is less than the ratio threshold. Then multiple bins in a single cluster are merged into a new bin, and the 3000 clusters in group 1 will get a total of 3000 new bins. These 3000 new bins will enter the second assembly round together with the new bins obtained from other groups.

[0067] S5. If the bin list of the next assembly round (the m+1th assembly round) has the same number of bins as the bin list of the current assembly round (the mth assembly round), then binList m+1 = binList m , then output the assembly result of each bin in the current assembly round, otherwise according to the number of current assembly rounds m and k-mer decay step length StepkMer renew kMer m , get the value for the m+1th assembly round kMer m+1 , and then repeat steps S2 to S4;

[0068] ;

[0069] in,

[0070] StepkMerLen The number of rounds required for each decay;

[0071] StepkMer For each decay step, the decay step must be an odd number.

[0072] In a specific embodiment, the above method is used to co-assemble a single-cell microbial genome sequencing library with 21914 SAGs, and some assembly results are shown in Table 1. The specific parameters are set as follows:

[0073] The length of the initial k-mer was set to 51;

[0074] The completeness threshold C (the ratio of the number of SAGs corresponding to bins with completeness greater than 0.5 to the total number of SAGs) is set to 0.75;

[0075] The maximum number of assembly rounds MaxM is set to 20;

[0076] The length threshold was set to one-tenth of the length of Contig N50 in the assembly result;

[0077] The bins for each round are divided into 8,000 bins per group;

[0078] The similarity threshold SHmm is set to 0.25;

[0079] The similarity step size StepSH is set to 0.01;

[0080] The ratio threshold (the ratio of the number of all clusters to all bins within a group) was set to 0.75;

[0081] The merge threshold maxC (the maximum number of bins in a cluster within a group) is set to 15;

[0082] The k-mer decay step length StepkMer is set to 5;

[0083] The number of rounds required for each decay, StepkMerLen, is set to 5.

[0084] Table 1. Some bins that meet high assembly quality standards in the co-assembly results of single-cell microbial genome sequencing libraries

[0085] bin_ID Completeness Contamination Genome size N50 8_3079 100 29.36782 4354587 32055 8_3100 100 11.45801 2866009 14853 8_3271 100 5.46595 2687405 25206 8_3499 100 37.22571 8003849 12305 8_3500 100 30.0627 6864230 20754 8_3145 99.91948 3.587963 3043472 28843 8_3148 99.90338 2.586188 3025084 36201 8_3098 99.9002 5.588822 2560968 32014 8_3033 99.60145 21.82574 4332449 15661

[0086] Among them, Bin_ID is the number of the bin; Completeness is the completeness of the bin; Contamination is the contamination of the bin;

[0087] N50 is the contig N50 length of the bin.

[0088] A total of 21,914 SAGs were assembled into 3,538 bins, of which 288 bins had integrity greater than 50, with an average contamination of 9.33 and a maximum contamination of 72. There were 87 bins that met the high assembly quality standards, and Table 1 shows the assembly information of some of them, such as bins 8_3271, 8_3145, 8_3148, and 8_3098, which had better assembly results.

[0089] In a specific embodiment, a device for assembling a genome sequence in a single-cell microorganism genome is also provided, comprising a processor and a memory; the processor and the memory are connected via a communication bus; wherein the processor is used to call and execute a program stored in the memory; the memory is used to store a program, and the program is at least used to execute the method for assembling a genome sequence in a single-cell microorganism genome.

[0090] The embodiments described above are merely descriptions of preferred implementations of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope of the present invention.

Claims

1. A method for assembling a genome sequence in a single-cell microorganism genome, characterized in that: include: S1. Obtain a single-cell microbial genome sequencing library, generate a bin for each SAG, and initialize the bin list; S2. Assemble each bin according to the k-mer set in the current assembly round, and evaluate the integrity and contamination of the assembly result of each bin to obtain assembly information; the length of the k-mer is K, the value range of K is 21-61 and K is an odd number; S3. If the assembly information satisfies the assembly termination condition, the assembly result of each bin in the current assembly round is output, otherwise, step S4 is executed; the assembly termination condition is any one of the following (1)-(2): (1) The ratio of the number of SAGs corresponding to bins with integrity greater than 0.5 to the total number of SAGs is greater than the preset integrity threshold; (2) The number of assembly rounds is greater than the preset maximum number of assembly rounds; S4. Group all bins, hierarchically cluster the bins within the group according to the similarity between the two bins; then split the hierarchical clustering results into one or more clusters according to the dynamically adjusted similarity threshold, until the number of clusters within the group or the similarity threshold meets the preset split termination condition; Merge the bins in a single cluster within a group into a new bin, and then aggregate the new bins of all groups to obtain the bin list for the next assembly round; S5. If the bin list of the next assembly round has the same number of bins as the bin list of the current assembly round, output the assembly result of each bin of the current assembly round, otherwise update the k-mer and repeat steps S2 to S4; Step S4 includes: S4-1. Filter the assembly results of each bin to obtain the filtered assembly results of each bin; S4-2. Divide all bins into multiple groups, and calculate the similarity between bins in each group based on the filtered assembly results of each bin and the k-mer set in the current assembly round; S4-3. Modify the similarity according to the completeness and contamination of the assembly result of each bin, perform hierarchical clustering on the bins in the group according to the modified similarity, and obtain the hierarchical clustering result of the group; S4-4. Split the hierarchical clustering results of the group into one or more clusters according to the similarity threshold, and calculate the ratio R of the number of all clusters in the group to the number of all bins; S4-5. If R is less than the ratio threshold, execute step S4-6; otherwise, lower the similarity threshold, if the lowered similarity threshold is less than the minimum similarity, execute step S4-6, otherwise repeat step S4-4; S4-6. Merge the bins in a single cluster within a group into a new bin, then aggregate the new bins of all groups and update the bin list; In step S4-1, the method for filtering the assembly results of each bin includes: removing contigs in the bin assembly results whose length is less than a length threshold; the length threshold is one L of the length of Contig N50 in the assembly results; and the value range of L is 5 to 20.

2. The method according to claim 1, characterized in that In step S4-3, the method for modifying the similarity includes: if the integrity of a bin in the group is not less than 0.9 or the contamination is not less than 0.5, then modifying the similarity between the bin and any bin in the group to 0.

3. The method according to claim 1, characterized in that Step S4-6 also includes: If the number of bins in a cluster within the group is greater than the merging threshold, the cluster is split into multiple small clusters, and then the bins in a single small cluster are merged into a new bin; wherein the number of bins in the small cluster is not greater than the merging threshold.

4. The method according to claim 1, characterized in that In step S5, the method for updating k-mer includes: Update k-mer according to the current assembly round number and k-mer decay step size.

5. A device for assembling genome sequences in a single-cell microorganism genome, characterized in that It comprises a processor and a memory; the processor and the memory are connected via a communication bus; wherein the processor is used to call and execute a program stored in the memory; the memory is used to store a program, and the program is used to execute at least the method for assembling a genome sequence in a single-cell microorganism genome as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Method and system for screening specific sequences of pathogenic species

    CN115719616A

  • Method for distinguishing strains in sequencing results of single-cell microbial genomes

    CN118737269A