Box separation method and device of metagenome assembly sequence, electronic equipment and storage medium
By acquiring and aligning the co-tag sequencing data and genome sequence collections of preset microbial communities, the correlation between genome sequences is calculated and binned, the problem of low precision and accuracy of genome sequence binning in the prior art is solved, and more efficient genome sequence grouping is achieved.
Patent Information
- Application Number
- CN202311498091.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-10
- Publication Date
- 2025-05-13
AI Technical Summary
The existing genome assembly sequence binning method has low precision and accuracy of binning genome assembly sequences in the environment, and it is impossible to effectively identify gene construction of multiple bacterial species.
By obtaining the co-tag sequencing data of the preset microbial community and the genome sequence set, sequence alignment is performed to obtain the tag sequence set, the first correlation between the genome sequences is calculated, and binning is performed based on this.
The binning fineness and accuracy of genome assembly sequences can be improved, and the gene sequences of multiple bacterial species in the environment can be more effectively grouped.
Smart Images

Figure CN119993271A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of bioinformatics, and in particular to a binning method, device, electronic device and storage medium for metagenome assembly sequences. Background Art
[0002] In the field of environmental microbial genetics, there are often multiple species of microorganisms in the environment, and the gene assembly sequences of all microorganisms are mixed together. Conventional gene assembly cannot identify the source of microorganisms in the environment. The purpose of binning gene assembly sequences is to group the gene sequences of all microorganisms in the environment so that the sequences grouped together are derived from the same species as much as possible, so as to achieve gene construction of multiple species in the environment.
[0003] The current method for binning gene assembly sequences has low precision and accuracy in binning gene assembly sequences in the environment. Summary of the invention
[0004] The embodiments of the present disclosure provide a binning method, device, electronic device and storage medium for a metagenomic assembly sequence, which can improve the binning precision and accuracy of a metagenomic assembly sequence.
[0005] According to one aspect of the present disclosure, a method for binning a metagenomic assembly sequence is provided, comprising:
[0006] Obtaining co-labeled sequencing data and a set of gene assembly sequences of a preset microbial community, wherein the co-labeled sequencing data includes a plurality of gene fragments corresponding to the preset microbial community and label data corresponding to each of the gene fragments;
[0007] Performing sequence alignment on each gene assembly sequence in the gene assembly sequence set based on the co-label sequencing data to obtain a label sequence set corresponding to each gene assembly sequence in the gene assembly sequence set;
[0008] Calculating a first degree of association between gene assembly sequences in the gene assembly sequence set according to the tag sequence set corresponding to each gene assembly sequence;
[0009] The gene assembly sequences in the gene assembly sequence set are binned based on the first association degree to obtain a binning result.
[0010] According to one aspect of the present disclosure, a binning device for a metagenomic assembly sequence is provided, comprising:
[0011] A first acquisition unit is used to acquire co-labeled sequencing data and a set of gene assembly sequences of a preset microbial community, wherein the co-labeled sequencing data includes a plurality of gene fragments corresponding to the preset biological community and label data corresponding to each of the gene fragments;
[0012] A sequence alignment unit, used to perform sequence alignment on the co-labeled sequencing data and each gene assembly sequence in the gene assembly sequence set to obtain a label sequence set corresponding to each gene assembly sequence in the gene assembly sequence set;
[0013] A first association degree calculation unit, configured to calculate a first association degree between gene assembly sequences in the gene assembly sequence set according to a tag sequence set corresponding to each gene assembly sequence;
[0014] A binning unit is used to bin the gene assembly sequences in the gene assembly sequence set based on the first association degree to obtain a binning result.
[0015] Optionally, in some embodiments, the boxing unit is specifically used for:
[0016] Based on the first association degree, dividing the plurality of gene assembly sequences in the gene assembly sequence set into a plurality of sequence groups;
[0017] Obtaining the polynucleotide frequency of each of the sequence groups;
[0018] The plurality of sequence groups are clustered based on the plurality of polynucleotide frequencies to obtain binning results.
[0019] Optionally, the binning device for the metagenome assembly sequence further comprises:
[0020] A first determining unit, configured to determine the number of bacterial species in the gene assembly sequence set based on the gene assembly sequence in the gene assembly sequence set;
[0021] The boxing unit is specifically used for:
[0022] Determining the number of cluster centers based on the number of bacterial species;
[0023] Determine the number of cluster centers based on the multiple polynucleotide frequencies;
[0024] The plurality of sequence groups are clustered according to the cluster centers to obtain binning results.
[0025] Optionally, the first determining unit is specifically configured to:
[0026] Obtaining a single copy gene set and a conserved gene set of the gene assembly sequence in the gene assembly sequence set;
[0027] Based on the single copy gene set and the conserved gene set, the number of bacterial species in the gene assembly sequence set is determined.
[0028] Optionally, the first determining unit is specifically configured to:
[0029] Obtaining the intersection of the single copy gene set and the conserved gene set;
[0030] Determine the number of times each gene in the intersection occurs in the set of gene assembly sequences;
[0031] Based on the number of times each of the genes corresponds, the number of bacterial species in the gene assembly sequence set is determined.
[0032] Optionally, the boxing unit is specifically used for:
[0033] Based on the first association degree, generating a gene assembly sequence association graph, wherein a point of the gene assembly sequence association graph indicates the gene assembly sequence, and an edge indicates the first association degree between two gene assembly sequences;
[0034] Based on the gene assembly sequence association graph, a plurality of the gene assembly sequences are divided into a plurality of sequence groups.
[0035] Optionally, the binning device for the metagenome assembly sequence further comprises:
[0036] A first segmentation unit, configured to segment each of the gene assembly sequences in the gene assembly sequence set into a plurality of first short sequences having a first predetermined length;
[0037] A second correlation calculation unit, used to determine a second correlation between two adjacent first short sequences;
[0038] A second determining unit, configured to determine a relevance threshold based on a plurality of said second relevances;
[0039] A deleting unit is used to delete the first degree of association that is less than the degree of association threshold.
[0040] Optionally, the first association degree calculation unit is specifically used to:
[0041] Determine a first set and a second set corresponding to every two of the gene assembly sequences, wherein the first set includes the same label data in the label sequence sets corresponding to the two gene assembly sequences, and the second set includes all the label data in the label sequence sets corresponding to the two gene assembly sequences;
[0042] Based on the first set and the second set, a first degree of association between two of the gene assembly sequences is determined.
[0043] Optionally, the binning device for the metagenome assembly sequence further comprises:
[0044] A second segmentation unit, used for segmenting each of the gene assembly sequences into a plurality of second short sequences having a second predetermined length;
[0045] A second acquisition unit, configured to acquire the second short sequences at both ends of each of the gene assembly sequences as target second short sequences of the gene assembly sequences;
[0046] The first association degree calculation unit is specifically used for:
[0047] A first set and a second set corresponding to the target second short sequences in every two of the gene assembly sequences are determined.
[0048] Optionally, the first association degree calculation unit is specifically used to:
[0049] Combining the target second short sequence of one of the gene assembly sequences with the target second short sequence of another of the gene assembly sequences one by one in sequence;
[0050] A first set and a second set corresponding to two target second short sequences in each combination are determined.
[0051] Optionally, the first association degree calculation unit is specifically used to:
[0052] Determine, based on the first set and the second set corresponding to each of the combinations, a candidate association degree corresponding to the combination;
[0053] Based on the candidate association degree corresponding to each of the combinations, a first association degree between two of the gene assembly sequences is determined.
[0054] Optionally, the binning device for the metagenome assembly sequence further comprises:
[0055] A third acquisition unit is used to obtain the comparison result of each of the gene assembly sequences, wherein the comparison result includes the assembly sequence length, the comparison type and the quality score;
[0056] A screening unit, configured to screen out a screened alignment result from the alignment result based on the length of the assembled sequence, the alignment type and the quality score;
[0057] The first association degree calculation unit is specifically used for:
[0058] The first degree of association between the gene assembly sequences is calculated based on the tag sequence set in each of the post-screening comparison results.
[0059] Optionally, the polynucleotide frequency is a tetranucleotide frequency;
[0060] The boxing unit is specifically used for:
[0061] Concatenating the gene assembly sequences in each of the sequence groups to generate a sequence group assembly sequence;
[0062] Slide the assembly sequence of the sequence group with a window of 4 and a step size of 1 to obtain the types of tetranucleotides contained in the sequence group and the number corresponding to each type;
[0063] Based on the quantities, the tetranucleotide frequency of each of the tetranucleotides in the sequence set is determined.
[0064] According to one aspect of the present disclosure, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the binning method of the metagenomic assembly sequence as described above is implemented.
[0065] According to one aspect of the present disclosure, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the binning method of the metagenomic assembly sequence as described above is implemented.
[0066] According to one aspect of the present disclosure, a computer program product is provided, which includes a computer program, and the computer program is read and executed by a processor of a computer device, so that the computer device executes the binning method of the metagenomic assembly sequence as described above.
[0067] The binning method of the metagenomic assembly sequence in the embodiment of the present disclosure obtains co-label sequencing data and a gene assembly sequence set of a preset microbial community, wherein the co-label sequencing data includes multiple gene fragments corresponding to the preset biological community and label data corresponding to each gene fragment; the co-label sequencing data is compared with each gene assembly sequence in the gene assembly sequence set to obtain a label sequence set corresponding to each gene assembly sequence in the gene assembly sequence set; the first association degree between the gene assembly sequences in the gene assembly sequence set is calculated according to the label sequence set corresponding to each gene assembly sequence; the gene assembly sequences in the gene assembly sequence set are binned based on the first association degree to obtain a binning result.
[0068] In this way, the binning method of the metagenomic assembly sequence of the disclosed embodiment uses a plurality of gene fragments in the preset co-label sequencing data to perform sequence comparison with each gene assembly sequence in the gene assembly sequence set to obtain a label sequence set corresponding to the gene assembly sequence. Since the label data of the short gene fragments derived from the same long DNA fragment in the co-label sequencing data are the same, the degree of association between the gene assembly sequences can be obtained using the label sequence set. The gene assembly sequences with higher association degrees are more likely to belong to the same microorganism. Therefore, the disclosed embodiment can perform more refined binning results on the gene assembly sequences that may belong to the same microorganism by calculating the correlation of the label data between the gene assembly sequences, thereby improving the binning precision and accuracy of the gene assembly sequences.
[0069] Other features and advantages of the present disclosure will be described in the following description, and partly become apparent from the description, or understood by practicing the present disclosure. The purpose and other advantages of the present disclosure can be realized and obtained by the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The accompanying drawings are used to provide further understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation on the technical solution of the present disclosure.
[0071] Figure 1 It is a system architecture diagram of the application of the binning method of the metagenomic assembly sequence according to the embodiment of the present disclosure;
[0072] Figure 2 is a schematic flow chart of a binning method for a metagenomic assembly sequence according to an embodiment of the present disclosure;
[0073] Figure 3 It is a schematic diagram for sequence alignment;
[0074] Figure 4 It is a schematic diagram of a long gene fragment covering two gene assembly sequences at the same time;
[0075] Figure 5 It is a schematic diagram of the association map of the gene assembly sequence;
[0076] Figure 6 A schematic diagram of splicing the gene assembly sequences in a sequence group;
[0077] Figure 7A-7D It is a schematic diagram for finding tetranucleotides in a sequence;
[0078] Figure 8 is another schematic flow chart of the binning method of the metagenomic assembly sequence provided by the present disclosure;
[0079] Fig. 9 is a histogram of binning effects of binning experiments using the disclosed embodiments and other tools;
[0080] Fig.10 is a histogram of binning effects of another binning experiment using the disclosed embodiment and other tools;
[0081] Fig.11 is a schematic structural diagram of a binning device for a metagenomic assembly sequence according to an embodiment of the present disclosure;
[0082] Fig.12 is a terminal structure diagram for implementing various methods according to an embodiment of the present disclosure;
[0083] Fig.13 It is a server structure diagram for implementing various methods according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0084] In order to make the purpose, technical solution and advantages of the present disclosure more clear, the present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.
[0085] Before further describing the embodiments of the present disclosure in detail, the nouns and terms involved in the embodiments of the present disclosure are described. The nouns and terms involved in the embodiments of the present disclosure are subject to the following interpretations:
[0086] Polynucleotide frequency: The frequency of all possible polynucleotides with k nucleotides in the sequence. If k = 2, the frequency of dinucleotides (i.e. AA, AT, AG, AC, ... TT) is calculated, with a total of 4^2 = 16 types; if k = 3, the frequency of trinucleotides (i.e. AAA, AAT, AAG, AAC, ... TTT) is calculated, with a total of 4^3 = 64 types; and so on.
[0087] Paired-end sequencing: A method of sequencing from both ends of the template while synthesizing. In order to sequence in two directions separately, two sequencing primers in different directions are required; in order to distinguish the reads in the two directions, a short sequence is added in front of one of the sequencing primers for labeling.
[0088] Eigenvalues and eigenvectors: For an operator A, if there is a non-zero vector X that satisfies: AX = aX, where a is a constant. Then X is called the eigenvector of the operator A, and a is the corresponding eigenvalue.
[0089] In the related art, binning is mainly achieved by clustering gene assembly sequences based on features such as the composition and abundance of gene assembly sequences. However, there are complex inter-species and intra-species repeated sequences in microbial communities, and the noise of extracting features from short gene assembly sequences is large, resulting in the inability to finely bin gene assembly sequences. In contrast, the embodiments of the present disclosure provide a binning method for metagenomic assembly sequences, in order to improve the precision and accuracy of binning of metagenomic assembly sequences.
[0090] System architecture of the application of the present disclosure
[0091] Figure 1 It is a system architecture diagram applied to the binning method of the metagenomic assembly sequence according to the embodiment of the present disclosure, which includes a terminal 140, the Internet 130, a gateway 120, a server 110, and the like.
[0092] The terminal 140 includes various forms such as desktop computers, laptop computers, PDAs (personal digital assistants), mobile phones, vehicle-mounted terminals, home theater terminals, and dedicated terminals. In addition, it can be a single device or a collection of multiple devices. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data.
[0093] The server 110 refers to a computer system that can provide certain services to the terminal 140. Compared with the ordinary terminal 140, the server 110 has higher requirements in terms of stability, security, performance, etc. The server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part of a high-performance computer (such as a virtual machine), a combination of parts of multiple high-performance computers (such as virtual machines), etc.
[0094] The gateway 120 is also called an internetwork connector or a protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. The gateway is a translator between two systems that use different communication protocols, data formats or languages, or even completely different architectures. At the same time, the gateway can also provide filtering and security functions. The message sent by the terminal 140 to the server 110 must be sent to the corresponding server 110 through the gateway 120. The message sent by the server 110 to the terminal 140 must also be sent to the corresponding terminal 140 through the gateway 120.
[0095] The binning method of the metagenomic assembly sequence of the embodiment of the present disclosure can be completely implemented in the terminal 140; can be completely implemented in the server 110; or can be implemented partially in the terminal 140 and the other part in the server 110.
[0096] General description of the disclosed embodiments
[0097] According to one embodiment of the present disclosure, a method for binning a metagenomic assembly sequence is provided.
[0098] like Figure 2 As shown, a schematic flow chart of a binning method for a metagenome assembly sequence provided by the present disclosure is provided. The method can be applied to a binning device for a metagenome assembly sequence. The binning method for a metagenome assembly sequence may include:
[0099] Step 210, obtaining co-labeled sequencing data and a gene assembly sequence set of a preset microbial community;
[0100] Step 220: perform sequence alignment between the co-label sequencing data and each gene assembly sequence in the gene assembly sequence set to obtain a label sequence set corresponding to each gene assembly sequence in the gene assembly sequence set;
[0101] Step 230, calculating a first degree of association between gene assembly sequences in the gene assembly sequence set according to the tag sequence set corresponding to each gene assembly sequence;
[0102] Step 240: bin the gene assembly sequences in the gene assembly sequence set based on the first association degree to obtain binning results.
[0103] In step 210, the preset microbial community is the microbial community where the gene assembly sequence set that needs to be binned is located, for example, the human intestinal microbial community or the oral microbial community. There are different co-label sequencing data for different microbial communities. The co-label sequencing data includes multiple gene fragments corresponding to the preset microbial community and label data corresponding to each gene fragment. The co-label sequencing data is generated by the single-tube long fragment sequencing technology developed by Shenzhen BGI Life Sciences Institute, which divides the microbial community sample from the same long gene fragment into multiple short gene fragments, and inserts the same label data into the short gene fragments derived from the same long gene fragment, and finally sequences the short gene fragments to generate co-label sequencing data for the microbial community.
[0104] The gene assembly sequence set includes multiple gene assembly sequences, which are composed of multiple gene fragments in the microbial community, and each gene assembly sequence corresponds to a bacterial species.
[0105] In step 220, the co-labeled sequencing data is sequenced with each gene assembly sequence in the gene assembly sequence set. Since the gene assembly sequence set includes multiple gene assembly sequences, in order to distinguish each gene assembly sequence, before the co-labeled sequencing data is used to perform sequence alignment on the gene assembly sequence, an index can be established for the gene assembly sequence in the gene assembly sequence set, and each gene assembly sequence corresponds to a unique index.
[0106] Since each gene fragment in the common label data corresponds to a label data, a label sequence set corresponding to the gene assembly sequence can be constructed according to the comparison results. Figure 3 As shown, Figure 3 An example of co-labeled sequencing data and gene assembly sequences is shown, where the co-labeled sequencing data shows gene fragments and corresponding label data, for example, the label data corresponding to the first gene fragment is "3-420-300". After using the co-labeled sequencing data to align the following gene assembly sequence, the first short sequence "ATAC...GTCA" in the gene assembly sequence corresponds to the label data "3-420-300" in the co-labeled sequencing data; the second short sequence "TTCC...ACCG" corresponds to the label data "3-420-387" in the co-labeled sequencing data; the third short sequence "ACCG...TTAG" corresponds to the label data "3-420-387" in the co-labeled sequencing data; the fourth short sequence "AGTT...CAGA" corresponds to the label data "3-420-311" in the co-labeled sequencing data. Therefore, the label sequence set corresponding to the gene assembly sequence is [3-420-300, 3-420-387, 3-420-387, 3-420-311]. The gene fragments in the co-labeled sequencing data can be double-end sequencing data. When performing sequence alignment, the gene fragments are aligned to the gene assembly sequence in a double-end sequencing alignment mode.
[0107] The comparison result can be input in the format of a tag sequence set, and the index of the gene assembly sequence is stored corresponding to the tag sequence set, for example, contig_0, [3-420-300, 3-420-387, 3-420-387, 3-420-311]. The comparison result can also include the comparison position information of the short sequence corresponding to each tag in the gene assembly sequence, for example, contig_0, [(3-420-300, [0, 880]), (3-420-387, [881, 1870]), (3-420-387, [1871, 3345]), (3-420-311, [3346, 5210])].
[0108] After obtaining the alignment result corresponding to each gene assembly sequence, the alignment result file corresponding to the gene assembly sequence set can be output. The alignment result file can be saved in sam format, and each line in the file represents the alignment result of a sequencing data on the gene assembly sequence.
[0109] In step 230, a first degree of association between gene assembly sequences in the gene assembly sequence set is calculated according to the tag sequence set corresponding to each gene assembly sequence.
[0110] The first correlation indicates the degree of correlation between the tag sequence sets of the gene assembly sequences. The more repeated tags there are in the tag sequence sets of two gene assembly sequences, the more short sequences in the two gene assembly sequences come from the same long gene fragment, and the greater the possibility that the two gene assembly sequences belong to the same bacterial species.
[0111] In one embodiment, calculating the first degree of association between gene assembly sequences in the gene assembly sequence set according to the tag sequence set corresponding to each gene assembly sequence includes:
[0112] Determine the first set and the second set corresponding to every two gene assembly sequences;
[0113] Based on the first set and the second set, a first degree of association between the two gene assembly sequences is determined.
[0114] The first set contains the same label data in the label sequence sets corresponding to the two gene assembly sequences, that is, the intersection of the label data corresponding to the two gene assembly sequences; the second set contains all the label data in the label sequence sets corresponding to the two gene assembly sequences, that is, the union of the label data corresponding to the two gene assembly sequences. For example, the label sequence set of gene assembly sequence A is [3-420-300, 3-420-387, 3-420-387, 3-420-311], and the label sequence set of gene assembly sequence B is [3-420-311, 3-420-300, 3-410-258, 3-415-140]. Then the first set corresponding to gene assembly sequence A and gene assembly sequence B is {3-420-300, 3-420-311}, and the second set is {3-420-300, 3-420-387, 3-420-311, 3-410-258, 3-415-140}.
[0115] In one embodiment, determining the first set and the second set corresponding to every two gene assembly sequences includes:
[0116] Obtaining a first mapping relationship between each gene assembly sequence and each tag data on the corresponding tag sequence set;
[0117] Obtaining a second mapping relationship between each tag data and a gene combination sequence corresponding to each tag sequence set;
[0118] Based on the first mapping relationship and the second mapping relationship, a first set and a second set corresponding to every two gene assembly sequences are determined.
[0119] The first mapping relationship is based on the gene assembly sequence for statistics. In the aforementioned embodiment, the comparison result of each gene assembly sequence includes the gene assembly sequence index and the corresponding label sequence set. Based on this, the first mapping relationship of each gene assembly sequence index and each corresponding label data is obtained. Tabulation can be used as a separator between the gene assembly sequence index and the label data, and each first mapping relationship can be output in a row unit. For example, the label sequence set of contig_0 is [3-420-300, 3-420-387, 3-420-387, 3-420-311], so the first mapping relationship can be obtained as shown in Table 1.
[0120] Genome assembly sequence index Label data contig_0 3-420-300 contig_0 3-420-387 contig_0 3-420-311
[0121] Table 1
[0122] In the above embodiment, the comparison result also includes the comparison position information of each tag data in the tag sequence set, so the first mapping relationship may also include the position information of the tag data. For example, the tag sequence set of contig_0 is [(3-420-300, [0, 880]), (3-420-387, [881, 1870]), (3-420-387, [1871, 3345]), (3-420-311, [3346, 5210])], and the first mapping relationship may be as shown in Table 2.
[0123] Genome assembly sequence index Label data Compare location information contig_0 3-420-300 0:880 contig_0 3-420-387 881:3345 contig_0 3-420-311 3346:5210
[0124] Table 2
[0125] The second mapping relationship is statistically analyzed for the label data, and the gene assembly sequence where the label data is located can be determined from the label sequence set of each gene assembly sequence. Tabs can also be used as separators between the label data and the gene assembly sequence index, and each second mapping relationship can be output in units of lines. For example, the label sequence set of contig_0 is [3-420-300, 3-420-387, 3-420-387, 3-420-311], and the label sequence set of contig_1 is [3-420-311, 3-420-300, 3-410-258, 3-415-140]. Then the second mapping relationship obtained based on the statistics of the label data 3-420-300 can be shown in Table 3.
[0126] Label data Genome assembly sequence index 3-420-300 contig_0 3-420-300 contig_1
[0127] Table 3
[0128] The second mapping relationship may also include the alignment position information corresponding to the label data in the gene assembly sequence, as shown in Table 4, for example.
[0129] Label data Genome assembly sequence index Compare location information 3-420-300 contig_0 0:880 3-420-300 contig_1 4130:5022
[0130] Table 4
[0131] After counting the first mapping relationship and the second mapping relationship, the first set and the second set corresponding to the two gene assembly sequences can be determined. Since the first set is the intersection of the label data contained in the two gene assembly sequences, the first set can be determined in the following way:
[0132] Traverse the first mapping relationship corresponding to a gene assembly sequence;
[0133] For each label data corresponding to the first mapping relationship, traverse the gene assembly sequence index corresponding to the label data in the second mapping relationship;
[0134] If the gene assembly sequence index corresponding to the label data contains the index of another gene assembly sequence, the label data is added to the first set.
[0135] For example, when determining the first set between contig_0 and contig_1, the first mapping relationship corresponding to contig_0 is traversed, as shown in Table 1. For label data 3-420-300, the corresponding second mapping relationship is traversed, as shown in Table 2, which includes the second mapping relationship between 3-420-300 and contig_1. Therefore, label data 3-420-300 is the common label data of contig_0 and contig_1 and needs to be added to the first set; for label data 3-420-387, there is no second mapping relationship with contig_1, so it is not the common label data of contig_0 and contig_1; for label data 3-420-311, there is a second mapping relationship with contig_1, so label data 3-420-311 is the common label data of contig_0 and contig_1 and needs to be added to the first set. Therefore, the first set is {3-420-300, 3-420-311}.
[0136] Since the second set is the union of the label data contained in the two gene assembly sequences, the second set can be determined by combining the corresponding label data in the first mapping relationship of the two gene assembly sequences to obtain the second set. For example, the first mapping relationship of contig_0 is shown in Table 1, and the first mapping relationship of contig_1 is shown in Table 5.
[0137] Genome assembly sequence index Label data contig_1 3-420-311 contig_1 3-420-300 contig_1 3-410-258 contig_1 3-415-140
[0138] Table 5
[0139] Based on Table 1 and Table 5, it can be obtained that the second set of gene assembly sequences contig_0 and contig_1 is {3-420-300, 3-420-387, 3-420-311, 3-410-258, 3-415-140}.
[0140] The advantage of determining the first set and the second set based on the first mapping relationship and the second mapping relationship in this embodiment is that the correspondence between the gene assembly sequence and the label data can be quickly obtained, thereby improving the efficiency of determining the first set and the second set.
[0141] After determining the first set and the second set, the first degree of association can be calculated by dividing the number of label data in the first set by the number of label data in the second set, as shown in Formula 1:
[0142]
[0143] In formula 1, α and β represent two gene assembly sequences; J αβ Represents the first degree of association between two gene assembly sequences; M αβ represents the first set corresponding to the two gene assembly sequences; N αβ Indicates the second set corresponding to the two gene assembly sequences. For example, the first set of contig_0 and contig_1 is {3-420-300, 3-420-311}, and the second set is {3-420-300, 3-420-387, 3-420-311, 3-410-258, 3-415-140}, then the first correlation between contig_0 and contig_1 is 2 / 5 = 0.4.
[0144] The advantage of determining the first degree of association based on the first set and the second set is that the first degree of association is determined by using the proportion of common label data to all label data, and a simple calculation is used to accurately indicate the proportion of sequences belonging to the same long gene fragment between two gene assembly sequences, thereby improving the calculation efficiency of determining the first degree of association.
[0145] Each short sequence in the gene assembly sequence is a long gene fragment from different microorganisms in the environment. Since a long gene fragment needs to cover two assembly sequences at the same time to be able to be binned based on the label data corresponding to the long gene fragment, and a long gene fragment that covers two gene assembly sequences at the same time must cover both ends of the two assembly sequences. For example Figure 4As shown, the long gene fragment covers both the gene assembly sequence A and the gene assembly sequence B, and the short sequence at the end of the gene assembly sequence A and the short sequence at the front end of the gene assembly sequence B correspond to the two short gene fragments of the long gene fragment. The sequence in the middle of the gene assembly sequence may only cover the sequence and cannot be associated with other gene assembly sequences on the label, and belongs to a noise sequence. Therefore, in one embodiment, before determining the first set and the second set corresponding to each two gene assembly sequences, the method also includes:
[0146] Dividing each gene assembly sequence into a plurality of second short sequences having a second predetermined length;
[0147] Obtaining the second short sequences at both ends of each gene assembly sequence as the target second short sequence of the gene assembly sequence;
[0148] Determining the first set and the second set corresponding to each two gene assembly sequences includes:
[0149] The first set and the second set corresponding to the target second short sequences in every two gene assembly sequences are determined.
[0150] The second predetermined length is less than the length of each gene assembly sequence in the gene assembly sequence set, and the second predetermined length can be preset to half the length of the shortest gene assembly sequence in the gene assembly sequence set. For example, if the shortest gene assembly sequence length is 5000 bp, the second predetermined length can be preset to 2500 bp.
[0151] Each gene assembly sequence is divided into a plurality of second short sequences, and the second short sequences at both ends are used as the target second short sequence. The gene assembly sequence can be divided into a plurality of second short sequences in a manner from both ends to the middle. In this way, even if the gene assembly sequence cannot be equally divided into a plurality of second short sequences, the length of the second short sequence at both ends must also be the same. For example, the length of the gene assembly sequence is 5000bp, and the second predetermined length is 2000bp. If it is divided according to the order from front to back, three second short sequences of 2000bp, 2000bp and 1000bp will be obtained successively, and the length of the target second short sequence at both ends is different like this. If it is divided according to the order from both ends to the middle, 2000bp, 1000bp and 2000bp will be obtained successively, and the length of the target second short sequence is the same at this moment. The same length of the target second short sequence can improve the accuracy of calculating the first degree of association, and the first degree of association will not be biased towards any gene assembly sequence.
[0152] After determining the target second short sequence of each gene assembly sequence, the first set and the second set of target second short sequences corresponding to every two gene assembly sequences are determined. The process of determining the first set and the second set of target second short sequences is the same as the process of determining the first set and the second set of two gene assembly sequences in the aforementioned embodiment, and will not be repeated here.
[0153] Since each gene assembly sequence has two target second short sequences and they are non-contiguous, in one embodiment, determining the first set and the second set corresponding to the target second short sequences in every two gene assembly sequences includes:
[0154] Combining the target second short sequence of one gene assembly sequence with the target second short sequence of another gene assembly sequence one by one in sequence;
[0155] A first set and a second set corresponding to two target second short sequences in each combination are determined.
[0156] In this embodiment, the target second short sequences of two gene assembly sequences are combined. For example, the target second short sequence A and the target second short sequence B of the gene assembly sequence contig_0 are combined with the target second short sequence C and the target second short sequence D of the gene assembly sequence contig_1. Four combinations can be obtained: (A, C), (A, D), (B, C), and (B, D). Based on each combination, the first set and the second set corresponding to the target second short sequence are determined.
[0157] The advantage of this embodiment is that the first correlation degree is determined by using a combination of single target second short sequences in two gene assembly sequences, which accurately indicates the coverage of long gene fragments and improves the accuracy of determining the first correlation degree.
[0158] Based on this embodiment, different combinations will calculate different associations. Therefore, in one embodiment, based on the first set and the second set, determining the first association between two gene assembly sequences includes:
[0159] Determine the candidate association degree corresponding to each combination based on the first set and the second set corresponding to each combination;
[0160] Based on the candidate association degree corresponding to each combination, a first association degree between the two gene assembly sequences is determined.
[0161] Based on the first set and the second set corresponding to each combination, the process of determining the candidate association degree corresponding to the combination is the same as that of Formula 1, and will not be repeated here.
[0162] The process of determining the first association degree of two gene assembly sequences through candidate association degrees may be: taking the maximum value among the candidate association degrees as the first association degree of the two gene assembly sequences. For example, the candidate association degrees corresponding to the combinations (A, C), (A, D), (B, C), (B, D) are: 0.6, 0.2, 0.4 and 0.3 respectively, then the first association degree of the two gene assembly sequences is 0.6.
[0163] The advantage of determining the first degree of association based on candidate degrees of association corresponding to different combinations is that the influence of label data that does not provide association on the first degree of association is reduced, thereby improving the accuracy of determining the first degree of association.
[0164] The advantage of using short sequences at both ends of the gene assembly sequence to determine the first degree of association is that it reduces noise sequences and improves the accuracy of determining the first degree of association; at the same time, it reduces the sequence length and improves the efficiency of determining the first degree of association.
[0165] Each sequencing data will correspond to a comparison result after being compared to the gene assembly sequence. However, not all comparison results can convey accurate comparison information. For example, if the gene assembly sequence is short, there will be more noise in the sequence, and the comparison result may be inaccurate. Therefore, before calculating the first correlation between the gene assembly sequences, the comparison results can be screened to obtain accurate gene assembly sequence comparison results.
[0166] In one embodiment, after performing sequence alignment between the co-labeled sequencing data and each gene assembly sequence in the gene assembly sequence set to obtain a label sequence set corresponding to each gene assembly sequence in the gene assembly sequence set, the method further comprises:
[0167] Obtain the alignment results of each gene assembly sequence, which include the assembly sequence length, alignment type and quality score;
[0168] Based on the length of the assembled sequence, the alignment type and the quality score, the filtered alignment results are selected from the alignment results.
[0169] The assembly sequence length indicates the sequence length of the gene assembly sequence corresponding to the alignment result. The assembly sequence length can be used to filter gene assembly sequences with large noise. The alignment type can be used to filter gene assembly sequences whose alignment method does not conform to the alignment type. The quality score can be used to filter out gene assembly sequences with a high probability of alignment errors.
[0170] The comparison results after screening by using the length of the assembled sequence can be deleted by setting a length threshold, and the comparison results whose assembled sequence length is less than the length threshold can be deleted. For example, the length threshold can be set to 5000bp, so the assembled sequence length less than 5000bp in the comparison results will be deleted.
[0171] The comparison results after screening can be deleted by presetting the comparison type. For example, if the comparison type is set to: the double-end sequencing data must be accurately aligned to the assembly sequence and it is the main comparison, then the secondary comparison and supplementary comparison in the comparison result that do not meet the comparison type requirements will be deleted.
[0172] The quality score can be used to filter out the comparison results after screening. By setting a minimum score threshold, the comparison results with a quality score less than the minimum score threshold can be deleted. For example, if the minimum score threshold is set to 24, the comparison results with a quality score less than 24 will be deleted.
[0173] It should be noted that the assembly sequence length, alignment type and quality score can be used alone to screen the alignment results, or can be combined in any way to filter the alignment results.
[0174] Based on the above steps, calculating the first association degree between the gene assembly sequences in the gene assembly sequence set according to the tag sequence set corresponding to each gene assembly sequence includes:
[0175] The first degree of association between the gene assembly sequences is calculated based on the tag sequence set in each post-screening alignment result.
[0176] The process of calculating the first degree of association between gene assembly sequences based on the label sequence set of the filtered comparison results is the same as the process of calculating the first degree of association between gene assembly sequences based on the label sequence set corresponding to each gene assembly sequence in the aforementioned embodiment, and will not be repeated here.
[0177] The advantage of screening the alignment results based on the length of the assembled sequence, the alignment type and the quality score is that the alignment results that lead to inaccurate sequence alignment can be deleted from multiple aspects, the screened alignment results are retained, and the accuracy of calculating the first degree of association is improved.
[0178] Since the number of gene assembly sequences is relatively large, the data volume of the comparison result file composed of the comparison results will also be very large. Therefore, in order to improve the screening efficiency of the comparison results, the comparison result file can be divided into multiple sub-files, and multi-threaded screening can be performed simultaneously. In the comparison result file, the comparison results are arranged in rows, so the comparison result file can be evenly divided into multiple sub-files according to the number of rows. The number of rows of each sub-file can be determined by the total number of rows and the total number of threads of the comparison result file. For example, the total number of rows of the comparison result file is 1200 rows, and the total number of threads is 6, then the number of rows of each sub-file is 1200 / 6=200 rows. The sub-files from the same comparison result file can be set with the same prefix to facilitate the management of the sub-files.
[0179] For each subfile, a thread screens each row of comparison results in the subfile, and the screening process is the same as the above embodiment, which will not be repeated here. After the screening is completed, the screened subfiles output by each thread are merged into a screened file, and the first correlation between the gene assembly sequences is determined based on the screening results in the screened file.
[0180] The advantage of dividing the comparison result file into multiple sub-files is that it improves the efficiency of screening the comparison results, thereby improving the efficiency of binning.
[0181] In step 240, the gene assembly sequences in the gene assembly sequence set are binned based on the first association degree to obtain binning results.
[0182] In one embodiment, the gene assembly sequences in the gene assembly sequence set are binned based on the first association degree to obtain binning results, including:
[0183] Based on the first association degree, multiple gene assembly sequences in the gene assembly sequence set are divided into multiple sequence groups;
[0184] Obtaining the polynucleotide frequencies for each sequence group;
[0185] Multiple sequence groups are clustered based on multiple polynucleotide frequencies to obtain binning results.
[0186] In this embodiment, the gene assembly sequences are first preliminarily grouped according to the first degree of association, and then further clustered according to the polynucleotide frequencies to obtain binning results.
[0187] In one embodiment, based on the first association degree, multiple gene assembly sequences in the gene assembly sequence set are divided into multiple sequence groups, including:
[0188] Based on the first association degree, generating a gene assembly sequence association graph;
[0189] Based on the gene assembly sequence association graph, multiple gene assembly sequences are divided into multiple sequence groups.
[0190] In this embodiment, the gene assembly sequence association graph may be an undirected connected graph, in which a node indicates a gene assembly sequence and an edge indicates a first degree of association between two gene assembly sequences. Figure 5 As shown, every two gene assembly sequences are connected and show the first degree of association. For example, the first degree of association between the assembly sequences contig_0 and contig_1 is 0.4.
[0191] In one embodiment, based on the gene assembly sequence association, multiple gene assembly sequences are divided into multiple sequence groups, including:
[0192] Determine the maximum eigenvalue and eigenvector corresponding to the gene assembly sequence association graph;
[0193] Calculate the residence probability and transition probability of each node in the gene assembly sequence association graph;
[0194] Based on the hierarchical coding model, multiple gene assembly sequences are divided into multiple sequence groups.
[0195] The largest eigenvalue and eigenvector can be determined using the power method.
[0196] The residence probability is the probability of staying at each node. The residence probability can be calculated by formula 2.
[0197]
[0198] In formula 2, P α is the residence probability of node α; V α is the value corresponding to node α in the eigenvector of the gene assembly sequence association graph; |V| is the modulus of the eigenvector.
[0199] The transition probability is the probability that a node changes from its current state to the next state.
[0200] Formula 3 is used to calculate.
[0201]
[0202] In formula 3, P αβ is the transition probability between nodes α and β; J αβ is the first degree of association between the gene assembly sequences α and β, and λ is the maximum eigenvalue.
[0203] Dividing multiple gene assembly sequences into multiple sequence groups based on the hierarchical coding model is to achieve the most accurate grouping by minimizing the coding length. The calculation process of the coding length is shown in Formula 4.
[0204]
[0205] In Formula 4, M is the coding scheme that divides all nodes into m regions, and length(M) is the coding length corresponding to the coding scheme; represents the probability that a single random walk code belongs to the i-th zone among m zones, which is equal to the sum of the transition probabilities of each node; H(Q) represents the Shannon information entropy corresponding to all cross-zone codes; The probability of a single random behavior residing in the i-th zone among m zones is equal to the sum of the residence probabilities of all nodes in the i-th zone and The sum of H(P i) represents the Shannon information entropy corresponding to the random walk code in the i-th partition. By minimizing the value of Formula 4, the coding scheme M is obtained, and the nodes belonging to the same partition in the coding scheme M are clustered into the same sequence group.
[0206] When dividing the sequence groups based on the first correlation, if the first correlation is very low, it may mean that the corresponding two gene assembly sequences are definitely not from the same bacterial species. Therefore, in one embodiment, before generating a gene assembly sequence association graph based on the first correlation, the method includes:
[0207] Splitting each gene assembly sequence in the gene assembly sequence set into a plurality of first short sequences having a first predetermined length;
[0208] determining a second correlation degree between two adjacent first short sequences;
[0209] Determining a relevance threshold based on the plurality of second relevances;
[0210] The first relevance that is less than the relevance threshold is deleted.
[0211] In this embodiment, the first degree of association is screened by calculating the degree of association threshold, and the second degree of association between multiple adjacent first short sequences on each gene assembly sequence can be used to determine the degree of association threshold. The length of each first short sequence is the first predetermined length. The calculation process of the second degree of association is similar to the calculation process of the first degree of association in the aforementioned embodiment, and will not be repeated here. After obtaining the second degree of association between all adjacent first short sequences, the frequencies of all second degrees of association can be counted and fitted to calculate the degree of association threshold, and the specific process can be shown in Formula 5.
[0212]
[0213] In Formula 5, K refers to the correlation threshold; refers to the mean of the second correlation; δ refers to the standard deviation of the second correlation; n refers to the number of second correlations; z α Refers to confidence. After the correlation threshold is calculated based on Formula 5, the first correlations that are less than the correlation threshold are deleted, and the remaining first correlations are used to generate a gene assembly sequence correlation graph.
[0214] The advantage of this implementation is that the first correlation between gene assembly sequences with low correlation and which are unlikely to be grouped into the same group is deleted, thereby improving the grouping efficiency.
[0215] The advantage of dividing multiple gene assembly sequences into multiple sequence groups based on the gene assembly sequence association graph is that the accuracy of grouping is improved.
[0216] After the plurality of gene assembly sequences are divided into a plurality of sequence groups, the polynucleotide frequency of each sequence group is obtained. In one embodiment, the polynucleotide frequency is a tetranucleotide frequency. There are 256 tetranucleotides, and the tetranucleotide frequency is the frequency of each tetranucleotide in the sequence group.
[0217] Get the polynucleotide frequencies for each sequence group, including:
[0218] The gene assembly sequences in each sequence group are concatenated to generate the sequence group assembly sequence;
[0219] The assembled sequence of the sequence group is slid with a window of 4 and a step size of 1 to obtain the types of tetranucleotides contained in the sequence group and the corresponding number of each type;
[0220] Based on the quantity, the tetranucleotide frequency of each tetranucleotide in the sequence group was determined.
[0221] In this embodiment, each gene assembly sequence in each sequence group is spliced together to form a sequence group assembly sequence. Since direct splicing will cause the end-to-end parts of the two gene assembly sequences to also form multiple tetranucleotides, the calculation result is inaccurate. Therefore, a special sequence can be used to splice each gene assembly sequence together, such as "NNNN". Figure 6 As shown, multiple gene assembly sequences are spliced together using the sequence "NNNN".
[0222] For the sequence group assembly sequence, 4 is used as a window to indicate the range of tetranucleotides, and the step length is 1 to slide on the sequence group assembly sequence to obtain the types and frequencies of tetranucleotides contained in the sequence group assembly sequence. If the gene assembly sequence is spliced using a special sequence, then when the window appears with characters contained in the special sequence during sliding, the sequence is not collected, and the sliding continues until the window does not contain the characters in the special sequence. Figure 7A-7D As shown, the window slides from the beginning and end of the sequence group assembly sequence, first collects "ATAC" and records the corresponding quantity as 1; after sliding one position backward, collects "TACC" and records the corresponding quantity as 1; after sliding one position backward, collects "ACCG" and records the corresponding quantity as 1; and so on, when "ACCG" is collected again, the corresponding quantity is increased by 1. If sequences containing "N" such as "TCAN", "CANN", "ANNN", "NNNN" are identified during the window sliding process, it is not necessary to collect and continue to slide backward. According to the above process, when the window slides to the last position of the sequence group assembly sequence, stop sliding, and the types and corresponding quantities of tetranucleotides contained in the sequence group can be obtained at this time.
[0223] Determining the tetranucleotide frequency of each tetranucleotide in a sequence group based on quantity can be performed by dividing the quantity of each tetranucleotide by the total number of tetranucleotides to obtain the corresponding tetranucleotide frequency.
[0224] After obtaining the frequency of each tetranucleotide in each sequence group, the tetranucleotide frequencies may be stored in a matrix, wherein the number of rows in the matrix may be the number of sequence groups plus one, and the number of columns may be 257. The matrix is output for subsequent clustering.
[0225] The advantage of obtaining tetranucleotide frequencies is that it can accurately express the gene composition information in the sequence group and improve clustering accuracy.
[0226] The sequence groups are clustered based on the polynucleotide frequencies corresponding to each sequence group to obtain binning results. In order to improve clustering efficiency, the number of cluster centers can also be determined before clustering. Therefore, in one embodiment, before clustering multiple sequence groups based on multiple polynucleotide frequencies and obtaining binning results, the method further includes:
[0227] Based on the gene assembly sequences in the gene assembly sequence set, the number of bacterial species in the gene assembly sequence set is determined.
[0228] The number of bacterial species can be used as the number of cluster centers. In one embodiment, based on the gene assembly sequence in the gene assembly sequence set, determining the number of bacterial species in the gene assembly sequence set includes:
[0229] Obtaining a single copy gene set and a conserved gene set of the gene assembly sequence in the gene assembly sequence set;
[0230] Based on the single-copy gene set and the conserved gene set, the number of bacterial species in the gene assembly sequence set was determined.
[0231] In this embodiment, a single copy gene appears only once in the gene assembly sequence of a single bacterial species; a conserved gene is a gene that can appear in the gene assembly sequence of most bacterial species. Therefore, the number of both single copy genes and conserved genes is approximately the same as the number of bacterial species. To obtain the single copy gene set and the conserved gene set, a gene annotation tool can be used to compare the genes on the gene assembly sequence with the gene sequences in the known single copy gene set and the conserved gene set to determine the single copy gene set and the conserved gene set present in the gene assembly sequence set.
[0232] In one embodiment, determining the number of bacterial species in the gene assembly sequence set based on the single copy gene set and the conserved gene set includes:
[0233] Obtain the intersection of the single-copy gene set and the conserved gene set;
[0234] Determine the number of times each gene in the intersection occurs in the set of gene assembly sequences;
[0235] Based on the number of occurrences of each gene, the number of bacterial species in the gene assembly sequence set was determined.
[0236] The intersection of the single copy gene set and the conserved gene set includes genes that are both single copy genes and conserved genes in the gene assembly sequence set. The number of occurrences of each gene in the intersection is determined in the gene assembly sequence set. According to the number of times each gene corresponds to, the number of bacterial species in the gene assembly sequence set can be determined by calculating the average number of times each gene corresponds to, and the average number is used as the number of bacterial species.
[0237] The advantage of this embodiment is that the number of bacterial species is determined by using the number of times the genes in the intersection of the single-copy gene set and the conserved gene set appear in the gene assembly sequence, which has high search efficiency and improves the efficiency of determining the number of bacterial species based on the single-copy gene set and the conserved gene set.
[0238] If a gene only appears once in each bacterial species and appears in most bacterial species, then the number of times this gene appears can represent the number of bacterial species. Therefore, this embodiment determines the number of bacterial species based on the combination of single copy genes and conservative gene sets, which can improve the accuracy of determining the number of bacterial species.
[0239] After determining the number of bacterial species, it is equivalent to determining the number of categories for clustering the sequence groups. Therefore, multiple sequence groups are clustered based on multiple polynucleotide frequencies to obtain binning results, including:
[0240] Determine the number of cluster centers based on the number of bacterial species;
[0241] Determining a number of cluster centers based on a plurality of polynucleotide frequencies;
[0242] According to the cluster center, multiple sequence groups are clustered to obtain binning results.
[0243] The number of bacterial species is the number of cluster centers. Cluster centers are determined in multiple sequence groups based on the polynucleotide frequencies corresponding to each sequence group. The polynucleotide frequency corresponding to each sequence group is a vector mapped in a high-dimensional space. The dimension of the high-dimensional space is determined by the number of polynucleotide species. For example, there are 256 tetranucleotides, so the tetranucleotide frequency corresponding to the sequence group is a vector mapped in a 256-dimensional space, where the frequency of the i-th tetranucleotide is defined as the component of the corresponding coordinate axis.
[0244] The clustering process can be achieved by calculating the Euclidean distance of different sequence groups in high-dimensional space, and clustering points with similar distances into the same cluster. In one embodiment, the clustering process includes:
[0245] One is randomly selected from multiple sequence groups as an initial cluster center, and the initial cluster center is added to the cluster center set.
[0246] Calculate the probability of the remaining sequence group being selected as the next cluster center. The calculation process is shown in Formula 6:
[0247]
[0248] In formula 6, represents the probability that one of the remaining sequence groups is selected as the next cluster center; Represents the minimum value of the Euclidean distance between the sequence group and the cluster center in the cluster center set. According to each corresponding Randomly select a sequence group as a new cluster center and add it to the cluster center set. Repeat this process until the number of sequence groups in the cluster center set reaches the number of cluster centers.
[0249] The distance between each cluster center and each of the remaining sequence groups is calculated one by one. The calculation process can use the Euclidean distance formula. The remaining sequence groups are each assigned to the cluster center closest to it. At this point, the preliminary clustering is completed and multiple clusters are formed.
[0250] After completing the preliminary clustering, calculate the centroid of each cluster in the high-dimensional space and use the centroid as the new cluster center of each cluster. If the new cluster center is different from the previous cluster center, it means that the clustering has not achieved convergence, and it is necessary to repeat the previous step and recalculate the distance between each new cluster center and each of the remaining sequence groups to form a new cluster. If the new cluster center does not change, it means that convergence has been achieved, and the sequence groups in the same cluster belong to the same bin.
[0251] Put the assembled sequences belonging to the same cluster into the same file to obtain the binning results.
[0252] Through the above process, the present embodiment achieves clustering of sequence groups with similar polynucleotide frequency distances to obtain binning results, thereby ensuring that assembled sequences belonging to the same bin have similar polynucleotide features, thereby improving clustering accuracy.
[0253] After obtaining the first correlation degree, preliminary grouping is first performed based on the first correlation degree, and then clustering is performed according to the polynucleotide frequency to obtain the binning result. The advantage is that the co-label correlation degree and polynucleotide characteristics between the gene assembly sequences are comprehensively considered for binning, thereby improving the accuracy of binning.
[0254] The embodiments of the present disclosure are described in detail in combination with specific application scenarios.
[0255] like Figure 8As shown, it is another flow chart of the binning method of the metagenomic assembly sequence provided by the present disclosure. In the embodiment of the present disclosure, the binning method of the metagenomic assembly sequence provided by the present disclosure will be introduced exemplarily. Take the binning of the gene assembly sequence in the artificial microbial community as an example. The community contains 8 bacterial species and 2 fungal species, and the gene sequences of different microorganisms are mixed in equal proportions. The method specifically includes the following steps:
[0256] Step 810: align the co-labeled sequencing data to each gene assembly sequence in the gene assembly sequence set.
[0257] The co-labeled sequencing data consisted of double-end sequencing data with a read length of 100 bp, with a total of approximately 59.54 GB of qualified data. The genome assembly sequence set contained 1,280 assembly sequences with a length greater than 5,000 bp, with a total length of 68,276,771 bp.
[0258] The co-labeled sequencing data can be directly aligned to the gene assembly sequence set using an alignment tool, such as the minimap2 tool. Use the -d parameter in minimap2 to index the gene assembly sequence, and then align the co-labeled data to the gene assembly sequence in double-end sequencing alignment mode. When aligning, select the -sr parameter to adapt to short fragment alignment. Finally, save the alignment results in a sam format file, where each line in the file represents the alignment result of a short sequence (or sequencing data) on the gene assembly sequence.
[0259] Step 820: Calculate the co-label association between the gene assembly sequences using the comparison results.
[0260] Take the sam format file as input and divide the file into several sub-files evenly according to the number of lines. Divide the total number of lines in the file by the number of threads to get the number of lines in each sub-file, and set the file name prefix of the sub-files to the same value.
[0261] For each subfile, filter the alignment results in the subfile according to the length of the assembly sequence, alignment type and quality score, and extract the key alignment information. Among them, the minimum length threshold of the assembly sequence length is set to 5000bp; the alignment type is set to meet the requirement that the double-end sequencing data is accurately aligned to the assembly sequence and is the main alignment; the quality score is set to the minimum quality score threshold of 24. The extracted key alignment information includes: label name, assembly sequence name and alignment position. Each line outputs the key alignment information corresponding to a alignment result with a tab as the delimiter. Splice the key alignment information output by all subfiles into one file and sort them according to the assembly sequence name as the first keyword, the label name as the second keyword, and the alignment position as the third keyword.
[0262] The key alignment information is summarized with the assembly sequence name as the key. Finally, the assembly sequence name, label name and alignment position are output in lines, with tabs as separators. For example, contig_0\t3_420_387\t3810:4200, where 3810:4200 means that the alignment position is from 3810 to 4200 in the gene assembly sequence.
[0263] The tag name is used as the key to summarize the key alignment information. Finally, the tag name, assembly sequence name and alignment position are output in lines, with tabs as separators. For example, 3_420_387\tcontig_0\t3810:4200.
[0264] Calculate the co-label association between any two gene assembly sequences with a length greater than 5000 bp. Split the gene assembly sequence into short sequences of equal length of 2000 bp from both ends to the middle, retain the first and last short sequences, record the start and end positions of each interval, and save them in array format.
[0265] Extract all the tag types and corresponding tag numbers in each short sequence and output them in dictionary format. The first short sequence of each gene assembly sequence is output to the dic_1.txt file. Each line of this file contains the assembly sequence name, tag name and tag number, and uses tabs as delimiters, such as
[0266] contig_0\t3_420_387\t2. The last short sequence is output to the dic_2.txt file in the same format as above.
[0267] The association degree between the short sequences in any two gene assembly sequences is calculated using formula 1 in the above embodiment, and the maximum value is taken as the co-label association degree between the two gene assembly sequences, and the assembly sequence names of the two gene assembly sequences and the co-label association degree are stored in the association degree file. Each line of the file contains two assembly sequence names and co-label association degrees, and a tab character is used as a separator, such as: contig_0\tcontig_5\t0.23.
[0268] Step 830: Divide the gene assembly sequence into multiple sequence groups using the co-label association degree.
[0269] Determine the co-label association interval. Split each assembled sequence longer than 20,000 bp into multiple short sequences at 5,000 bp intervals, calculate the association between adjacent short sequences in each assembled sequence, statistically analyze the association frequency distribution, and fit it with a normal distribution curve. Calculate the one-sided confidence interval based on the fitted curve, and determine the lower limit of the confidence interval with a 95% confidence level. If the lower limit is 0.039, then the co-label association interval is [0.039, +∞], which means that the co-label association information in the association file that is not in this interval is deleted.
[0270] Construct a co-label association undirected graph based on the co-label association in the association file. Use the assembly sequence name as the key and the digital number as the value to construct a dictionary structure in which the assembly sequence name and the digital number correspond one to one. Use the digital number corresponding to the assembly sequence as the node name and the co-label association as the edge weight to construct an undirected graph, and store the undirected graph in the data structure of the adjacency list.
[0271] Based on the co-labeled association undirected graph, the gene assembly sequences are divided into multiple sequence groups. Specifically, the infomap algorithm can be used to implement grouping, and the algorithm parameters silent are set to TRUE and num_trials is set to 20. The grouping results are output to the grouping file, and each line contains a single gene assembly sequence and the corresponding sequence group number, such as contig_0\t13.
[0272] Step 840: Calculate the tetranucleotide frequency of each sequence group.
[0273] According to the grouping result of step 830, the gene assembly sequences belonging to the same sequence group are connected, and every two gene assembly sequences are separated by "NNNN", and the sequence group number is used as the name of the merged sequence. Furthermore, if a gene assembly sequence does not belong to any sequence group and its length is greater than 10,000 bp, it is also regarded as a merged sequence to calculate the tetranucleotide frequency.
[0274] Take the merged sequence as input and use the tetranucleotide frequency calculation tool to get the tetranucleotide frequency corresponding to each sequence group. The tetranucleotide frequency calculation tool can be CheckM, which can calculate the tetranucleotide frequencies of different merged sequences through the parameter --tetra, and output the matrix in the format of (n+1) rows and (256+1) columns, where n is the number of merged sequences and 256 is the number of tetranucleotide types.
[0275] Step 850: Estimate the number of bacterial species contained in the collection based on the gene assembly sequence collection.
[0276] The number of bacterial species can be estimated by using the CheckM model to calculate the number of occurrences of single-copy conservative genes, and the number of occurrences of single-copy conservative genes is taken as the number of bacterial species. When there are multiple single-copy conservative genes, the average number of occurrences of multiple single-copy genes is taken as the number of bacterial species.
[0277] Step 860: cluster multiple sequence groups using tetranucleotide frequencies.
[0278] Based on the number of bacterial species, multiple sequence groups are clustered with tetranucleotide frequencies as features. The number of cluster centers is obtained by rounding up 1.5 times the number of bacterial species. The K-means clustering algorithm can be used for the clustering process. K-means uses the k++ algorithm to select cluster centers, and the Euclidean distance is used to calculate the distance. The rest of the parameters are used by default. After clustering is completed, the name of each assembled sequence and the corresponding bin number are output, for example, contig_0\t12. The gene assembly sequences belonging to the same bin are stored in the same fasta file.
[0279] Through the above process, the gene assembly sequence was binned. In order to test the binning effect, other tools were used to bin the same gene assembly sequence, including MetaBAT2, Maxbin2 and VAMB. The above three models all use default parameters. CheckM was used to evaluate the binning results of different tools. The high-quality genome given by CheckM is defined as a completeness greater than 90% and a contamination less than 5%. The binning results obtained by the evaluation are as follows: Fig. 9 As shown, the binning method of the metagenomic assembly sequence of the embodiment of the present disclosure obtains the largest number of high-quality genomes, which is 8, corresponding to 8 bacterial species in the microbial colony. The number of high-quality genomes obtained by the above three tools is 6, 4, and 6, respectively, and they are all bacterial species. For fungal species, CheckM was used to evaluate the fungal reference genome, and the fungal reference genome had low integrity and high contamination. This result is similar to the evaluation results of the other two genomes with larger total sequence lengths obtained by clustering the embodiment of the present disclosure. It was found that the accuracy and recall rates of these two non-high-quality genomes were relatively high when compared to the reference genomes of two fungi, one of which had an accuracy and recall rate of 0.991 and 0.973, and an F1-score of 0.985; the other had an accuracy and recall rate of 0.979 and 0.949, and an F1-score of 0.960. This shows that the binning method of the embodiment of the present disclosure is also relatively good for fungal binning, and can achieve high integrity and low contamination. The above results confirm that the embodiments of the present disclosure can achieve higher precision binning, and the accuracy of the binning results is high.
[0280] In addition to evaluating the binning results, the binning efficiency of the disclosed embodiment is also evaluated. Table 6 shows the total time consumed for binning using different tools.
[0281] Tools (6 threads) Total time MetaBAT2 7h17min Maxbin2 3h59min VAMB 2h44min Embodiments of the present disclosure 3h40min
[0282] Table 6
[0283] It can be seen from Table 6 that the total time consumption of the binning method of the embodiment of the present disclosure is less than MetaBAT2 and Maxbin2, so the binning efficiency of the embodiment of the present disclosure is also higher.
[0284] Using one data set for verification may be accidental. In order to further verify the precision and accuracy of the binning of the disclosed embodiment, the co-labeled data of human intestinal fecal microorganisms are used to bin the gene assembly sequences through the above steps 810-860. At the same time, the tools MetaBAT2, Maxbin2 and VAMB are used for binning. CheckM is used to evaluate the binning results, and a high-quality genome is defined as a genome with a completeness greater than 95% and a contamination degree less than 5%. A medium-quality genome is defined as a genome with a completeness greater than 70%, a contamination degree less than 10% and that does not fall within the definition of high quality. The binning results of each tool are shown below. Fig.10 As shown, the embodiment of the present disclosure obtained the most medium and high-quality genomes, with a value of 45; the number of high-quality genomes was the largest, with a value of 25; the number of medium-quality genomes was second, with a value of 20. MetaBAT2 obtained a total of 41 medium and high-quality genomes, second only to the embodiment of the present disclosure, but the number of medium-quality genomes it obtained accounted for the majority, and its number of high-quality genomes was the least among the four tools. VAMB obtained 21 high-quality genomes, second only to the embodiment of the present disclosure, and the number of high-quality genomes was also less than that of the present invention. In addition to the binning results, the classification levels to which different high-quality genomes belonged were further counted, as shown in Table 7.
[0285] tool Head division MetaBAT2 11 1 Maxbin2 13 4 VAMB 11 6 Embodiments of the present disclosure 17 5
[0286] Table 7
[0287] It can be seen from Table 7 that the number of categories (5) in the binning results of the embodiment of the present disclosure is slightly less than that of VAMB (6), but the number of purposes (17) is much greater than that of the other three. Therefore, the embodiment of the present disclosure has better binning precision.
[0288] The total time consumed for binning using different tools in this experiment is shown in Table 8.
[0289] Tools (6 threads) Total time MetaBAT2 6h7min Maxbin2 4h33min VAMB 4h1min Embodiments of the present disclosure 6h4min
[0290] Table 8
[0291] It can be seen from Table 8 that the binning method of the embodiment of the present disclosure is at the same order of magnitude as other tools in terms of time consumption. Therefore, the embodiment of the present disclosure improves binning precision and accuracy while maintaining computational efficiency.
[0292] Description of the apparatus and device of the present disclosure
[0293] It is to be understood that, although the steps in the above-mentioned flowcharts are sequentially displayed according to the characterization of arrows, these steps are not necessarily executed in sequence according to the order of arrow characterization. Unless there is a clear description in the present embodiment, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the above-mentioned flowcharts can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of the steps or stages in other steps.
[0294] It should be noted that in each specific embodiment of the present disclosure, when it comes to the need to perform relevant processing based on data related to the characteristics of the target object such as the target object attribute information or attribute information set, the permission or consent of the target object will be obtained first, and the collection, use and processing of these data will comply with the relevant laws, regulations and standards of the relevant region. In addition, when the embodiment of the present application needs to obtain the attribute information of the target object, the separate permission or separate consent of the target object will be obtained through a pop-up window or jump to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary target object-related data used to enable the normal operation of the embodiment of the present application will be obtained.
[0295] Fig.11 A schematic diagram of the structure of a binning device 1100 for a metagenome assembly sequence provided in an embodiment of the present disclosure. The binning device 1100 for a metagenome assembly sequence includes:
[0296] A first acquisition unit 1110 is used to acquire co-labeled sequencing data and a set of gene assembly sequences of a preset microbial community, wherein the co-labeled sequencing data includes a plurality of gene fragments corresponding to the preset biological community and label data corresponding to each gene fragment;
[0297] A sequence alignment unit 1120 is used to align the co-labeled sequencing data with each gene assembly sequence in the gene assembly sequence set to obtain a label sequence set corresponding to each gene assembly sequence in the gene assembly sequence set;
[0298] A first association degree calculation unit 1130 is used to calculate a first association degree between gene assembly sequences in the gene assembly sequence set according to a tag sequence set corresponding to each gene assembly sequence;
[0299] The binning unit 1140 is used to bin the gene assembly sequences in the gene assembly sequence set based on the first association degree to obtain binning results.
[0300] Optionally, in some embodiments, the boxing unit 1140 is specifically used for:
[0301] Based on the first association degree, multiple gene assembly sequences in the gene assembly sequence set are divided into multiple sequence groups;
[0302] Obtaining the polynucleotide frequencies for each sequence group;
[0303] Multiple sequence groups are clustered based on multiple polynucleotide frequencies to obtain binning results.
[0304] Optionally, the binning device 1100 for the metagenomic assembly sequence further includes:
[0305] A first determining unit (not shown) is used to determine the number of bacterial species in the gene assembly sequence set based on the gene assembly sequence in the gene assembly sequence set;
[0306] The boxing unit 1140 is specifically used for:
[0307] Determine the number of cluster centers based on the number of bacterial species;
[0308] Determining a number of cluster centers based on a plurality of polynucleotide frequencies;
[0309] According to the cluster center, multiple sequence groups are clustered to obtain binning results.
[0310] Optionally, the first determining unit (not shown) is specifically configured to:
[0311] Obtaining a single copy gene set and a conserved gene set of the gene assembly sequence in the gene assembly sequence set;
[0312] Based on the single-copy gene set and the conserved gene set, the number of bacterial species in the gene assembly sequence set was determined.
[0313] Optionally, the first determining unit (not shown) is specifically configured to:
[0314] Obtain the intersection of the single-copy gene set and the conserved gene set;
[0315] Determine the number of times each gene in the intersection occurs in the set of gene assembly sequences;
[0316] Based on the number of occurrences of each gene, the number of bacterial species in the gene assembly sequence set was determined.
[0317] Optionally, the boxing unit 1140 is specifically used for:
[0318] Based on the first association degree, generating a gene assembly sequence association graph, wherein a point of the gene assembly sequence association graph indicates a gene assembly sequence, and an edge indicates a first association degree between two gene assembly sequences;
[0319] Based on the gene assembly sequence association graph, multiple gene assembly sequences are divided into multiple sequence groups.
[0320] Optionally, the binning device 1100 for the metagenomic assembly sequence further includes:
[0321] A first segmentation unit (not shown), configured to segment each gene assembly sequence in the gene assembly sequence set into a plurality of first short sequences having a first predetermined length;
[0322] A second correlation degree calculation unit (not shown), used to determine a second correlation degree between two adjacent first short sequences;
[0323] A second determining unit (not shown), configured to determine a relevance threshold based on the plurality of second relevances;
[0324] A deleting unit (not shown) is used to delete the first degree of association that is less than a threshold value of the degree of association.
[0325] Optionally, the first association degree calculation unit 1130 is specifically configured to:
[0326] Determine a first set and a second set corresponding to every two gene assembly sequences, wherein the first set includes the same label data in the label sequence sets corresponding to the two gene assembly sequences, and the second set includes all label data in the label sequence sets corresponding to the two gene assembly sequences;
[0327] Based on the first set and the second set, a first degree of association between the two gene assembly sequences is determined.
[0328] Optionally, the binning device 1100 for the metagenomic assembly sequence further includes:
[0329] A second segmentation unit (not shown), configured to segment each gene assembly sequence into a plurality of second short sequences having a second predetermined length;
[0330] A second acquisition unit (not shown), used to acquire second short sequences at both ends of each gene assembly sequence as target second short sequences of the gene assembly sequence;
[0331] The first correlation calculation unit 1130 is specifically used for:
[0332] The first set and the second set corresponding to the target second short sequences in every two gene assembly sequences are determined.
[0333] Optionally, the first association degree calculation unit 1130 is specifically configured to:
[0334] Combining the target second short sequence of one gene assembly sequence with the target second short sequence of another gene assembly sequence one by one in sequence;
[0335] A first set and a second set corresponding to two target second short sequences in each combination are determined.
[0336] Optionally, the first association degree calculation unit 1130 is specifically configured to:
[0337] Determine the candidate association degree corresponding to each combination based on the first set and the second set corresponding to each combination;
[0338] Based on the candidate association degree corresponding to each combination, a first association degree between the two gene assembly sequences is determined.
[0339] Optionally, the binning device 1100 for the metagenomic assembly sequence further includes:
[0340] A third acquisition unit (not shown) is used to obtain the alignment result of each gene assembly sequence, wherein the alignment result includes the assembly sequence length, alignment type and quality score;
[0341] A screening unit (not shown), for screening the screened alignment results from the alignment results based on the length of the assembled sequence, the alignment type and the quality score;
[0342] The first correlation calculation unit 1130 is specifically used for:
[0343] The first degree of association between the gene assembly sequences is calculated based on the tag sequence set in each post-screening alignment result.
[0344] Optionally, the polynucleotide frequency is a tetranucleotide frequency;
[0345] The boxing unit 1140 is specifically used for:
[0346] The gene assembly sequences in each sequence group are concatenated to generate the sequence group assembly sequence;
[0347] The assembled sequence of the sequence group is slid with a window of 4 and a step size of 1 to obtain the types of tetranucleotides contained in the sequence group and the corresponding number of each type;
[0348] Based on the quantity, the tetranucleotide frequency of each tetranucleotide in the sequence group was determined.
[0349] Reference Fig.12 , Fig.12 The structural block diagram of the terminal 140 for implementing the binning method of the metagenome assembly sequence of the embodiment of the present disclosure includes: a radio frequency (RF) circuit 1210, a memory 1215, an input unit 1230, a display unit 1240, a sensor 1250, an audio circuit 1260, a wireless fidelity (WiFi) module 1270, a processor 1280, and a power supply 1290. Those skilled in the art can understand that Fig.12 The structure of the terminal 140 shown does not constitute a limitation on a mobile phone or a computer, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.
[0350] The RF circuit 1210 may be used for receiving and sending signals during information transmission or communication. In particular, after receiving downlink information from the base station, the information is sent to the processor 1280 for processing. In addition, the uplink data is sent to the base station.
[0351] The memory 1215 may be used to store software programs and modules. The processor 1280 executes various functional applications of the terminal and identifies lane change points by running the software programs and modules stored in the memory 1215 .
[0352] The input unit 1230 may be used to receive input digital or character information and generate key signal input related to the terminal's settings and function control. Specifically, the input unit 1230 may include a touch panel 1231 and other input devices 1232 .
[0353] The display unit 1240 may be used to display input information or provided information and various menus of the terminal. The display unit 1240 may include a display panel 1241 .
[0354] The audio circuit 1260 , the speaker 1261 , and the microphone 1262 may provide an audio interface.
[0355] In this embodiment, the processor 1280 included in the terminal 140 can execute the binning method of the metagenomic assembly sequence of the previous embodiment.
[0356] The terminal 140 of the embodiment of the present disclosure includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc. The embodiment of the present invention can be applied to various scenarios, including but not limited to gene sequencing, etc.
[0357] Fig.13A block diagram of the structure of a portion of a server 110 for implementing the binning method of a metagenome assembly sequence of an embodiment of the present disclosure. The server 110 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPUs) 1322 (e.g., one or more processors) and a storage device 1332, and one or more storage media 1330 (e.g., one or more massive storage devices) storing application programs 1342 or data 1344. Among them, the storage device 1332 and the storage medium 1330 may be short-term storage or permanent storage. The program stored in the storage medium 1330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 110. Furthermore, the central processing unit 1322 may be configured to communicate with the storage medium 1330, and execute a series of instruction operations in the storage medium 1330 on the server 110.
[0358] The server 110 may also include one or more power supplies 1326, one or more wired or wireless network interfaces 1350, one or more input and output interfaces 1358, and / or one or more operating systems 1341, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0359] The central processor 1322 in the server 110 can be used to execute the binning method of the metagenomic assembly sequence of the embodiment of the present disclosure.
[0360] The embodiments of the present disclosure also provide a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the binning method of the metagenomic assembly sequence of each of the aforementioned embodiments.
[0361] The present disclosure also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, so that the computer device executes the binning method for implementing the above-mentioned metagenomic assembly sequence.
[0362] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present disclosure described herein can, for example, be implemented in an order other than those illustrated or described herein. In addition, the terms "comprises" and "comprising" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0363] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0364] It should be understood that in the description of the embodiments of the present disclosure, the meaning of multiple (or multiple items) is more than two, greater than, less than, exceed, etc. are understood to not include the number, and above, below, within, etc. are understood to include the number.
[0365] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0366] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0367] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0368] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store program codes.
[0369] It should also be understood that the various implementations provided in the embodiments of the present disclosure can be combined arbitrarily to achieve different technical effects.
[0370] The above is a specific description of the implementation methods of the present disclosure, but the present disclosure is not limited to the above implementation methods. Technical personnel familiar with the art can also make various equivalent modifications or substitutions without violating the spirit of the present disclosure. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present disclosure.
Claims
1. A binning method for a metagenomic assembly sequence, characterized in that: include: Obtaining co-labeled sequencing data and a set of gene assembly sequences of a preset microbial community, wherein the co-labeled sequencing data includes a plurality of gene fragments corresponding to the preset microbial community and label data corresponding to each of the gene fragments; Perform sequence alignment of the co-labeled sequencing data with each gene assembly sequence in the gene assembly sequence set to obtain a label sequence set corresponding to each gene assembly sequence in the gene assembly sequence set; Calculating a first degree of association between gene assembly sequences in the gene assembly sequence set according to the tag sequence set corresponding to each gene assembly sequence; The gene assembly sequences in the gene assembly sequence set are binned based on the first association degree to obtain a binning result.
2. The method according to claim 1, characterized in that: The step of binning the gene assembly sequences in the gene assembly sequence set based on the first association degree to obtain binning results includes: Based on the first association degree, dividing the plurality of gene assembly sequences in the gene assembly sequence set into a plurality of sequence groups; Obtaining the polynucleotide frequency of each of the sequence groups; The plurality of sequence groups are clustered based on the plurality of polynucleotide frequencies to obtain binning results.
3. The method according to claim 2, characterized in that Before clustering the plurality of sequence groups based on the plurality of polynucleotide frequencies to obtain binning results, the method further comprises: Based on the gene assembly sequence in the gene assembly sequence set, determining the number of bacterial species in the gene assembly sequence set; The clustering of the plurality of sequence groups based on the plurality of polynucleotide frequencies to obtain binning results comprises: Determining the number of cluster centers based on the number of bacterial species; Determine the number of cluster centers based on the multiple polynucleotide frequencies; The plurality of sequence groups are clustered according to the cluster centers to obtain binning results.
4. The method according to claim 3, characterized in that The step of determining the number of bacterial species in the gene assembly sequence set based on the gene assembly sequence in the gene assembly sequence set comprises: Obtaining a single copy gene set and a conserved gene set of the gene assembly sequence in the gene assembly sequence set; Based on the single copy gene set and the conserved gene set, the number of bacterial species in the gene assembly sequence set is determined.
5. The method according to claim 4, characterized in that The step of determining the number of bacterial species in the gene assembly sequence set based on the single copy gene set and the conserved gene set includes: Obtaining the intersection of the single copy gene set and the conserved gene set; Determine the number of times each gene in the intersection occurs in the set of gene assembly sequences; Based on the number of times each of the genes corresponds, the number of bacterial species in the gene assembly sequence set is determined.
6. The method according to claim 2, characterized in that The dividing the plurality of gene assembly sequences in the gene assembly sequence set into a plurality of sequence groups based on the first association degree comprises: Based on the first association degree, generating a gene assembly sequence association graph, wherein a point of the gene assembly sequence association graph indicates the gene assembly sequence, and an edge indicates the first association degree between two gene assembly sequences; Based on the gene assembly sequence association graph, a plurality of the gene assembly sequences are divided into a plurality of sequence groups.
7. The method according to claim 6, characterized in that Before generating a gene assembly sequence association graph based on the first association degree, the method includes: Splitting each of the gene assembly sequences in the gene assembly sequence set into a plurality of first short sequences having a first predetermined length; determining a second correlation degree between two adjacent first short sequences; Determining a relevance threshold based on a plurality of said second relevances; The first degree of association that is less than the degree of association threshold is deleted.
8. The method according to claim 1, characterized in that The calculating the first degree of association between the gene assembly sequences in the gene assembly sequence set according to the tag sequence set corresponding to each gene assembly sequence includes: Determine a first set and a second set corresponding to every two of the gene assembly sequences, wherein the first set includes the same label data in the label sequence sets corresponding to the two gene assembly sequences, and the second set includes all the label data in the label sequence sets corresponding to the two gene assembly sequences; Based on the first set and the second set, a first degree of association between two of the gene assembly sequences is determined.
9. The method according to claim 8, characterized in that Before determining the first set and the second set corresponding to every two gene assembly sequences, the method further comprises: Dividing each of the gene assembly sequences into a plurality of second short sequences having a second predetermined length; Acquire the second short sequences at both ends of each of the gene assembly sequences as target second short sequences of the gene assembly sequences; The determining of the first set and the second set corresponding to every two gene assembly sequences comprises: A first set and a second set corresponding to the target second short sequences in every two of the gene assembly sequences are determined.
10. The method according to claim 9, characterized in that The determining of the first set and the second set corresponding to the target second short sequences in every two of the gene assembly sequences comprises: Combining the target second short sequence of one of the gene assembly sequences with the target second short sequence of another of the gene assembly sequences one by one in sequence; A first set and a second set corresponding to two target second short sequences in each combination are determined.
11. The method according to claim 10, characterized in that The determining, based on the first set and the second set, a first degree of association between two of the gene assembly sequences comprises: Determine, based on the first set and the second set corresponding to each of the combinations, a candidate association degree corresponding to the combination; Based on the candidate association degree corresponding to each of the combinations, a first association degree between two of the gene assembly sequences is determined.
12. The method according to claim 8, characterized in that The determining of the first set and the second set corresponding to every two gene assembly sequences comprises: Obtaining a first mapping relationship between each of the gene assembly sequences and each of the label data on the corresponding label sequence set; Obtaining a second mapping relationship between each of the label data and the gene combination sequence corresponding to each of the label sequence sets; Based on the first mapping relationship and the second mapping relationship, a first set and a second set corresponding to every two gene assembly sequences are determined.
13. The method according to claim 1, characterized in that After performing sequence alignment between the co-labeled sequencing data and each gene assembly sequence in the gene assembly sequence set to obtain a label sequence set corresponding to each gene assembly sequence in the gene assembly sequence set, the method further comprises: Obtaining an alignment result of each of the gene assembly sequences, wherein the alignment result includes the assembly sequence length, alignment type, and quality score; Based on the length of the assembled sequence, the alignment type and the quality score, screening out a screened alignment result from the alignment results; The calculating the first degree of association between the gene assembly sequences in the gene assembly sequence set according to the tag sequence set corresponding to each gene assembly sequence includes: The first degree of association between the gene assembly sequences is calculated based on the tag sequence set in each of the post-screening comparison results.
14. The method according to claim 2, characterized in that The polynucleotide frequency is a tetranucleotide frequency; The obtaining of the polynucleotide frequency of each sequence group comprises: Concatenating the gene assembly sequences in each of the sequence groups to generate a sequence group assembly sequence; Slide the assembly sequence of the sequence group with a window of 4 and a step size of 1 to obtain the types of tetranucleotides contained in the sequence group and the number corresponding to each type; Based on the quantities, the tetranucleotide frequency of each of the tetranucleotides in the sequence set is determined.
15. A binning device for a metagenomic assembly sequence, characterized in that: include: A first acquisition unit is used to acquire co-labeled sequencing data and a set of gene assembly sequences of a preset microbial community, wherein the co-labeled sequencing data includes a plurality of gene fragments corresponding to the preset biological community and label data corresponding to each of the gene fragments; A sequence alignment unit, used to perform sequence alignment on the co-labeled sequencing data and each gene assembly sequence in the gene assembly sequence set to obtain a label sequence set corresponding to each gene assembly sequence in the gene assembly sequence set; A first association degree calculation unit, configured to calculate a first association degree between gene assembly sequences in the gene assembly sequence set according to a tag sequence set corresponding to each gene assembly sequence; A binning unit is used to bin the gene assembly sequences in the gene assembly sequence set based on the first association degree to obtain a binning result.
16. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for binning the metagenomic assembly sequence according to any one of claims 1 to 14 is implemented.
17. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the binning method for the metagenomic assembly sequence according to any one of claims 1 to 14 is implemented.
18. A computer program product, comprising a computer program, wherein the computer program is read and executed by a processor of a computer device, so that the computer device executes the binning method for the metagenomic assembly sequence according to any one of claims 1 to 14.