A method, system, device, and medium for splitting multi-sample amplicon hybrid library sequencing data

By combining tag pairs and primer pairs, the problems of inaccurate barcode recognition and difficulty in handling degenerate bases in existing technologies are solved, enabling efficient splitting of multi-sample amplicon sequencing data and improving the accuracy and throughput of data processing.

CN122436009APending Publication Date: 2026-07-21GUANGDONG MEIGE GENE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610902878.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing software has inaccurate barcode recognition when processing multi-sample amplicon sequencing data, cannot distinguish between barcode and primer sequences, and cannot handle degenerate bases, resulting in low data splitting efficiency.

Method used

A dual mechanism of tag pairs and primer pairs is employed to split samples by matching the 5' end sequence of amplicon sequencing data. The combination of barcode and primer pairs is used to split samples when barcode is insufficient, dynamically skipping ligation sequences and allowing a certain degree of mismatch to improve the recognition rate.

Benefits of technology

It improves the efficiency and throughput of sequencing data splitting, can handle degenerate bases and hypoxanthine, and is suitable for high-throughput sequencing where the sample size exceeds the available barcodes, thus improving the accuracy and efficiency of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122436009A_ABST
    Figure CN122436009A_ABST
Patent Text Reader

Abstract

The application discloses a multi-sample amplicon mixed library sequencing data splitting method, system, device and medium, and belongs to the technical field of sequencing data preprocessing. The method, system, device or medium of the application is used for splitting sequencing data, which not only considers barcode matching conditions, but also considers primer pair alignment conditions. In the case that the sample quantity exceeds the available barcode, the splitting can be further based on the primer pair, so as to reversely improve the sequencing throughput.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of sequencing data preprocessing technology, and in particular to a method, system, device and medium for splitting sequencing data from multi-sample amplicon mixed library construction. Background Technology

[0002] Amplicon sequencing technology refers to the use of second- or third-generation sequencing platforms to sequence specific regions of PCR products, such as 16S rRNA, 18S rRNA, ITS, and functional genes, thereby obtaining information such as the microbial community structure, evolutionary relationships, and the correlation between microorganisms and the environment in environmental samples. From the development of second-generation sequencing technology to third-generation sequencing technology, amplicon sequencing has been widely used in many research fields, especially in the study of microbial diversity and community structure.

[0003] With advancements in sequencing technology (especially platforms like the Illumina NovaSeq series), the amount of data produced in a single run has increased exponentially, making it possible to perform amplicon sequencing on thousands of samples simultaneously. However, existing software has its own limitations when dealing with splitting large volumes of amplicon data, for example: 1. For a given barcode sequence, the check for a complete match does not start from the 5' end of the sequence, resulting in inaccurate barcode recognition; 2. It cannot handle degenerate bases, but treats them all as N bases, or considers them as mismatches; 3. It cannot distinguish between barcode sequences and primer sequences, and cannot match the correspondence between barcodes and primers; it can only match according to fixed sequences. 4. It has limited functionality and can only extract sequences according to a given pattern sequence. Summary of the Invention

[0004] To solve at least one of the above-mentioned technical problems, the technical solution adopted in this application is as follows.

[0005] The first aspect of this application provides a method for splitting multi-sample amplicon mixed library preparation and sequencing data, wherein the multi-sample amplicon mixed library preparation and sequencing data is obtained through the following steps: Amplicon is obtained for each sample. The 5' end of the amplicon includes the sequence of the forward tag FB and the sequence of the forward primer FP. The 3' end of the amplicon includes the complementary sequence of the reverse primer RP. The complementary sequence of the reverse primer RP may or may not include the complementary sequence of the reverse tag RB. FB and RB constitute a tag pair. Each sample corresponds to one tag pair. The number of samples exceeds the number of tag pairs. FP and RP constitute a primer pair. For any two samples with the same tag pair, the corresponding primer pairs are different. Amplicon samples from multiple samples were mixed to construct a library and then subjected to paired-end sequencing. Each sequencing data included two reads, R1 and R2. The splitting method includes the following steps: Obtain the label pair corresponding to each sample, as well as the primer pair set corresponding to each sample; For any sequencing data, if the 5' ends of R1 and R2 completely match a tag pair, the sequencing data is classified as identified sequences, and sequences in which the 5' ends of R1 and R2 completely match the forward tag sequence or the reverse tag sequence are removed. For an identified sequence, if the 5' ends of R1 and R2 can be matched with a primer pair, the sequencing data is classified according to the sample information corresponding to the primer pair.

[0006] Those skilled in the art will understand that the number of available bars is limited due to factors such as barcode sequence length and specificity. When performing high-throughput sequencing on a large number of samples, there may be a shortage of bars. In this application, a dual mechanism of bars and primer pairs is employed to address this situation. For samples with different amplification regions, or samples with the same amplification region but different primer sequences, the same barcode can be used. When barcodes alone cannot achieve the desired separation, further separation based on primer sequences can be achieved.

[0007] In some embodiments of this application, where the tag pair contains only FB, the statement "the 5' ends of R1 and R2 completely match a tag pair" means that FB completely matches the 5' end of either R1 or R2. Conversely, if FB does not completely match the 5' end of either R1 or R2, then the sequencing sequence does not completely match the tag pair.

[0008] In some other embodiments of this application, the tag pair consists of FB and RB. The statement "the 5' ends of R1 and R2 perfectly match the 5' end of a tag pair" means that the FB of the tag pair perfectly matches the 5' end of one of R1 and R2, and the RB of the tag pair perfectly matches the 5' end of the other of R1 and R2. In one case, the FB perfectly matches the 5' end of R1, and the RB perfectly matches the 5' end of R2. In another case, the FB perfectly matches the 5' end of R2, and the RB perfectly matches the 5' end of R1. If the FB of the tag pair does not perfectly match the 5' end of either R1 or R2, or if the FB of the tag pair perfectly matches the 5' end of one of R1 and R2 but the RB of the tag pair does not perfectly match the 5' end of the other of R1 and R2, then the sequencing sequence does not perfectly match the tag pair.

[0009] Furthermore, in order to distinguish the sequencing data, if the 5' ends of R1 and R2 completely match a tag pair, the sequencing data is marked as an identified sequence; if the 5' ends of R1 and R2 do not completely match any tag pairs, the sequencing data is marked as an unidentified sequence.

[0010] In this application, "the 5' ends of R1 and R2 can align with a primer pair" means that the FP of the primer pair can align with the 5' end of one of R1 and R2, and the RP of the primer pair can align with the 5' end of the other of R1 and R2. In one case, the FP can align with the 5' end of R1, and the RP can align with the 5' end of R2. In another case, the FP can align with the 5' end of R2, and the RP can align with the 5' end of R1. To further segment the identified sequences, if the 5' ends of R1 and R2 can align with a primer pair, the identified sequence is marked as a matched sequence. For matched sequences, the sequencing data is classified according to the sample information corresponding to the primer pairs. If the 5' ends of R1 and R2 cannot align with any primer pairs, the identified sequence is marked as an unmatched sequence.

[0011] In some embodiments of this application, the amplicon for each sample is obtained using the following steps: Obtain the DNA sample. Amplification is performed using at least one primer pair, wherein the upstream primer 5'-3' of the primer pair comprises a forward tag FB sequence and a forward primer FP sequence in sequence, and the downstream primer 5'-3' of the primer pair comprises a reverse primer RP sequence, and the 5' end of the downstream primer may or may not include a reverse tag RB sequence; Obtain the amplicon of the sample.

[0012] In some embodiments of this application, the phrase "different corresponding primer pairs" includes at least the following situations: When two primer pairs amplify different regions, their sequences are generally completely different. For example, one primer pair is used to amplify region A, and the other primer pair is used to amplify region B. Since regions A and B are different, for specificity reasons, the sequences of the two primers will not be the same when designing the primers. The two primer pairs have the same amplification region, but due to different designed binding sites, at least one of the upstream and downstream primers has a different sequence.

[0013] In some embodiments of this application, if FP and / or RP include degenerate bases or hypoxanthine, then the primer pair set includes all primer pairs in which the degenerate bases and / or hypoxanthine are replaced with specific bases.

[0014] For example, the FP sequence is ATCGRACCGAATTG (SEQ ID No. 1), and the RP sequence is CCGAYGCACGGAAT (SEQ ID No. 2). Where R is A or G, and Y is C or T, then the primer pairs formed by FP and RP include: FP1: ATCGAACCGAATTG (SEQ ID No. 3); RP1: CCGACGCACGGAAT (SEQ ID No. 4).

[0015] FP2: ATCGAACCGAATTG (SEQ ID No. 5); RP2: CCGAGGCACGGAAT (SEQ ID No. 6).

[0016] FP3: ATCGTACCGAATTG (SEQ ID No. 7); RP1: CCGACGCACGGAAT (SEQ ID No. 8).

[0017] FP4: ATCGTACCGAATTG (SEQ ID No. 9); RP1: CCGAGGCACGGAAT (SEQ ID No. 10).

[0018] In some embodiments of this application, different primer pairs may have completely identical FP or RP.

[0019] In some embodiments of this application, the amplicon for each sample is obtained using the following steps: Obtain the DNA sample. The DNA sample is amplified using the first primer pair to obtain the first amplification product. The upstream primer 5'-3' of the first primer pair includes a linker sequence and a forward primer FP sequence in sequence, and the downstream primer 5'-3' of the first primer pair includes a linker sequence and a reverse primer RP sequence in sequence. The first amplification product is amplified using a second primer pair to obtain the amplicon of the sample. The upstream primers 5'-3' of the second primer pair sequentially include the positive tag FB sequence and the complementary sequence of the linker sequence, and the downstream primers 5'-3' of the second primer pair include the complementary sequence of the linker sequence.

[0020] In this case, the 5' end of the resulting amplicon includes the forward tag FB sequence, the complementary sequence of the linker sequence, and the forward primer FR sequence in sequence, while the 3' end includes the complementary sequence of the linker sequence and the complementary sequence of the reverse primer RP in sequence.

[0021] In other embodiments of this application, the amplicon for each sample is obtained using the following steps: Obtain the DNA sample. The DNA sample is amplified using the first primer pair to obtain the first amplification product. The upstream primer 5'-3' of the first primer pair includes a linker sequence and a forward primer FP sequence in sequence, and the downstream primer 5'-3' of the first primer pair includes a linker sequence and a reverse primer RP sequence in sequence. The first amplification product is amplified using a second primer pair to obtain the amplicon of the sample. The upstream primers 5'-3' of the second primer pair sequentially include the positive tag FB sequence and the complementary sequence of the linker sequence, and the downstream primers 5'-3' of the second primer pair sequentially include the reverse tag FB sequence and the complementary sequence of the linker sequence.

[0022] In this case, the 5' end of the resulting amplicon includes, in sequence, the positive tag FB sequence, the complementary sequence of the linker sequence, and the positive primer FR sequence, while the 3' end includes, in sequence, the complementary sequence of the reverse tag RB, the complementary sequence of the linker sequence, and the complementary sequence of the reverse primer RP.

[0023] In some embodiments of this application, the 5' ends of the upstream and downstream primers of the second primer pair also include sequencing adapters.

[0024] In some embodiments of this application, the linker sequence consists of N bases, where N = 3 to 5. Therefore, the alignment of the 5' ends of R1 and R2 with a primer pair means that the sequence following the n1 bases from the 5' end of R1 and the sequence following the n2 bases from the 5' end of R2 can align with a primer pair, where 0 ≤ n1 ≤ N, 0 ≤ n2 ≤ N, and n1 and n2 are both natural numbers.

[0025] In some specific embodiments of this application, "the 5' ends of R1 and R2 can align with a primer pair" means that FP can align with the sequence following n1 bases from the 5' end of R1, and RP can align with the sequence following n2 bases from the 5' end of R2. In other specific embodiments of this application, "the 5' ends of R1 and R2 can align with a primer pair" means that RP can align with the sequence following n1 bases from the 5' end of R1, and FP can align with the sequence following n2 bases from the 5' end of R2.

[0026] In some embodiments of this application, the number of mismatched bases in the aligned region does not exceed M. In some specific embodiments of this application, M is determined according to the primer length range, generally not exceeding 20% ​​of the primer length, and generally 10-20%. In some specific embodiments of this application, the primer length is 20, and M is set to 2.

[0027] In some embodiments of this application, if for any two samples, regardless of whether the corresponding label pairs are the same, the corresponding primer pairs are different, then: For sequencing data where the 5' ends of R1 and R2 do not perfectly match any tag pairs (unidentified sequences), split them according to the following steps: Obtain representative values ​​L1 and L2 for the lengths of all samples FB and RB. Remove the sequence with a 5' end length of L1 from one of R1 and R2, and remove the sequence with a 5' end length of L2 from the other. Obtain the set of primer pairs for all samples; If the 5' ends of R1 and R2 after sequence excision can be matched with a primer pair, the sequencing data is classified according to the sample information corresponding to the primer pair.

[0028] In some embodiments of this application, the representative value is selected from one of the maximum value, the mode, the upper quartile, or determined according to any one of them, for example, the representative value is selected to be greater than the maximum value.

[0029] In some embodiments of this application, the 5' end of R1 with a length of L1 is first removed, and the 5' end of R2 with a length of L2 is removed. The removed R1 and R2 are then compared with the set of primer pairs of all samples. If a match is found, the removed sequence is used as a candidate barcode. If a match is not found, the 5' end of R1 with a length of L2 is removed, and the 5' end of R2 with a length of L1 is removed. The removed R1 and R2 are then compared with the set of primer pairs of all samples. Similarly, if a match is found, the removed sequence is used as a candidate barcode. If a match is not found, the next unidentified sequence is matched.

[0030] In some embodiments of this application, "the 5' ends of R1 and R2 can align with a primer pair" means that the FP of the primer pair can align with the 5' end of one of R1 and R2, and the RP of the primer pair can align with the 5' end of the other of R1 and R2. In one case, FP can align with the 5' end of R1, and RP can align with the 5' end of R2. In another case, FP can align with the 5' end of R2, and RP can align with the 5' end of R1. Similarly, in some embodiments of this application, there is a linker sequence between FB and FP, and between RB and RP, wherein the linker sequence consists of N bases, where N = 3 to 5. In some specific embodiments of this application, "the 5' ends of R1 and R2 can be aligned with a primer pair" means that FP can be aligned with the sequence following n1 bases from the 5' end of R1, and RP can be aligned with the sequence following n2 bases from the 5' end of R2. In other specific embodiments of this application, "the 5' ends of R1 and R2 can be aligned with a primer pair" means that RP can be aligned with the sequence following n1 bases from the 5' end of R1, and FP can be aligned with the sequence following n2 bases from the 5' end of R2, where 0 ≤ n1 ≤ N, 0 ≤ n2 ≤ N, and n1 and n2 are both natural numbers.

[0031] It should be noted that, during primer pair matching, the 5' ends of R1 and R2 that completely match the forward tag FB or the reverse tag RB have been removed.

[0032] A second aspect of this application provides a system for splitting multi-sample amplicon mixed library construction and sequencing data, comprising the following modules: The data input module is used to input the multi-sample amplicon pooled library construction and sequencing data, which is obtained through the following steps: Amplicon for each sample is obtained. The 5' end of the amplicon includes the sequence of the forward tag FB and the sequence of the forward primer FP. The 3' end of the amplicon includes the complementary sequence of the reverse primer RP. The complementary sequence of the reverse primer RP may or may not include the complementary sequence of the reverse tag RB. FB, alone or together with RB, constitutes a tag pair. Each sample corresponds to one tag pair. The number of samples exceeds the number of tag pairs. Furthermore, for any two samples with the same tag pair, the corresponding primer pairs are different. Amplicon samples from multiple samples were mixed to construct a library and then subjected to paired-end sequencing. Each sequencing data included two reads, R1 and R2. The data input module is also used to input the sample information from which the multi-sample amplicon mixed library construction and sequencing data comes from, as well as the tag pairs and primer pairs corresponding to each sample; The tag-based splitting module is connected to the data input module. The data input module is also used to input the sample information of the source of the multi-sample amplicon mixed library construction and sequencing data and the tag pairs corresponding to each sample, for splitting based on the following steps: For any sequencing data, if the 5' ends of R1 and R2 completely match a tag pair, the sequencing data is classified as identified sequences, and sequences in which the 5' ends of R1 and R2 completely match the forward tag sequence or the reverse tag sequence are removed. The primer pair splitting module is connected to both the data input module and the tag pair splitting module. The data input module is also used to obtain the primer pairs corresponding to the samples. This module is used to further split the identified sequences based on the following steps: If the 5' ends of R1 and R2 can be matched with a primer pair, the sequencing data is classified according to the sample information corresponding to the primer pair.

[0033] In some embodiments of this application, the splitting system further includes: The unidentified sequence splitting module, together with the data input module and the label-based splitting module, performs splitting based on the input information from the data input module if, for any two samples, regardless of whether the corresponding label pairs are the same, the corresponding primer pairs are different: For sequencing data where the 5' ends of R1 and R2 do not perfectly match any of the tag pairs, split the data as follows: Obtain representative values ​​L1 and L2 for the lengths of all samples FB and RB. Remove the sequence with a 5' end length of L1 from one of R1 and R2, and remove the sequence with a 5' end length of L2 from the other. Obtain the set of primer pairs for all samples; If the 5' ends of R1 and R2 after sequence excision can be matched with a primer pair, the sequencing data is classified according to the sample information corresponding to the primer pair.

[0034] A third aspect of this application provides a computer device, comprising: a memory for storing a computer program; and a processor for executing the steps of any of the splitting methods described in the first aspect of this application when executing the computer program.

[0035] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the splitting methods described in the first aspect of this application.

[0036] Compared with the prior art, this application has the following advantages: Using the methods, systems, devices, or media of this application to split sequencing data can support the identification of degenerate bases and hypoxanthine (I).

[0037] Using the methods, systems, devices, or media of this application to split sequencing data not only considers barcode matching but also primer pair alignment. When the sample size exceeds the available barcodes, it can further split the data based on primer pairs, thereby increasing sequencing throughput.

[0038] Using the methods, systems, devices, or media of this application to split sequencing data can dynamically skip the linker sequences between barcodes and primers, making it particularly suitable for situations where barcodes are introduced through two rounds of amplification, and allowing for some primer mismatches, thus offering greater tolerance.

[0039] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0040] The above and other objects, features, and advantages of exemplary embodiments of this application will become readily apparent from the following detailed description taken in conjunction with the accompanying drawings. Several embodiments of this application are illustrated in the drawings by way of example and not limitation, in which: Figure 1 This paper illustrates the amplicon structure and sequencing data coverage diagram in Embodiment 1 of this application; Figure 2 This diagram illustrates data splitting based on tag pairs in Embodiment 1 of this application; Figure 3 This diagram illustrates data splitting based on primer pairs in Embodiment 1 of this application; Figure 4 A schematic diagram of the amplicon structure and sequencing data coverage in Embodiment 2 of this application is shown. Detailed Implementation

[0041] Unless otherwise stated, implied from the context, or as is customary in the art, all parts and percentages in this application are based on weight, and all testing and characterization methods used are concurrent with the filing date of this application. Where applicable, any patent, patent application, or disclosure relating to this application is incorporated herein by reference in its entirety, and its equivalent patent families are also incorporated herein by reference, particularly the definitions of relevant terms in the art disclosed in such documents. If any definition of a specific term disclosed in the prior art is inconsistent with any definition provided in this application, the definition provided in this application shall prevail.

[0042] To make the technical problems, technical solutions and beneficial effects solved by this application clearer, the following detailed description is provided in conjunction with embodiments.

[0043] The following examples are used to illustrate preferred embodiments of this application. Those skilled in the art will understand that the techniques disclosed in the examples represent technologies discovered by the inventors that can be used to implement this application, and therefore can be considered preferred embodiments of this application. However, those skilled in the art should understand from this specification that many modifications can be made to the specific embodiments disclosed herein, still yielding the same or similar results, without departing from the spirit or scope of this application.

[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains, and all materials cited herein and referenced by them are incorporated herein by reference.

[0045] Those skilled in the art will recognize, or can learn through routine experimentation, many equivalents of specific embodiments of the invention described herein. These equivalents will be included in the claims.

[0046] Unless otherwise specified, the experimental methods used in the following examples are conventional methods. Unless otherwise specified, the instruments and equipment used in the following examples are all conventional laboratory instruments and equipment; unless otherwise specified, the experimental materials used in the following examples were all purchased from conventional biochemical reagent stores.

[0047] Example 1 1. Obtain sample information Sample information for obtaining sequencing data includes: (1) Sample name.

[0048] (2) Forward barcode (FB) and reverse barcode (RB) information, where RB may not exist, i.e. only the forward barcode exists.

[0049] (3) Information on the front primer (FP) and the back primer (RP), and an enumeration method is used to list all combinations of degenerate bases and hypoxanthine, replacing degenerate bases or hypoxanthine with actual bases, with the following replacement rules: 1) Replace R with A or G; 2) Replace Y with C or T; 3) Replace M with A or C; 4) Replace K with G or T; 5) Replace S with C or G; 6) Replace W with A or T; 7) Replace H with A, C, or T; 8) Replace B with C, G, or T; 9) Replace V with A, C, or G; 10) Replace D with A, G, or T; 11) Replace N with A, C, G, or T; 12) Replace I with A, C, G, T.

[0050] For example, the amplified region is bacterial V3V4, and the FP and RP sequences are as follows.

[0051] FP: 5'-ACTCCTACGGGAGGCAGCA-3' (SEQ ID No. 11); RP: 5'-GGACTACHVGGGTWTCTAAT-3' (SEQ ID No. 12), The primer pairs described above contain degenerate bases H, V, and W. After substitution, there are 3 × 3 × 2 = 18 actual primer pairs, which are as follows: 5'-GGACTACAAGGGTATCTAAT-3' (SEQ ID No. 13) 5'-GGACTACCAGGGTATCTAAT-3' (SEQ ID No. 14) 5'-GGACTACTAGGGTATCTAAT-3' (SEQ ID No. 15) 5'-GGACTACACGGGTATCTAAT-3' (SEQ ID No. 16) 5'-GGACTACCCGGGTATCTAAT-3' (SEQ ID No. 17) 5'-GGACTACTCGGGTATCTAAT-3' (SEQ ID No. 18) 5'-GGACTACAGGGGTATCTAAT-3' (SEQ ID No. 19) 5'-GGACTACCGGGGTATCTAAT-3' (SEQ ID No. 20) 5'-GGACTACTGGGGTATCTAAT-3' (SEQ ID No. 21) 5'-GGACTACAAGGGTTTCTAAT-3' (SEQ ID No. 22) 5'-GGACTACCAGGGTTTCTAAT-3' (SEQ ID No. 23) 5'-GGACTACTAGGGTTTCTAAT-3' (SEQ ID No. 24) 5'-GGACTACACGGGTTTCTAAT-3' (SEQ ID No. 25) 5'-GGACTACCCGGGTTTCTAAT-3' (SEQ ID No. 26) 5'-GGACTACTCGGGTTTCTAAT-3' (SEQ ID No. 27) 5'-GGACTACAGGGGTTTCTAAT-3' (SEQ ID No. 28) 5'-GGACTACCGGGGTTTCTAAT-3' (SEQ ID No. 29) 5'-GGACTACTGGGGTTTCTAAT-3' (SEQ ID No. 30) For example, the FP and RP sequences of the amplified fungus ITS2 are as follows.

[0052] FP: 5'-GCATCGATGAAGAACGCAGC-3' (SEQ ID No. 31); RP: 5'-TCCTCCGCTTATTGATATGC-3' (SEQ ID No. 32), This primer pair does not contain degenerate bases and does not require further processing.

[0053] (4) The amount of sequencing data required.

[0054] Table 1 shows some sample information and primer information. In this embodiment, the tag pairs do not include reverse tags (RB).

[0055] Table 1: Partial Sample Information and Primer Information .

[0056] Table 1 shows that samples A-05 and B-01 used the same barcode. The samples were amplified using the aforementioned bacterial V3V4 region primers or fungal ITS2 primers. After adding the barcode sequence, mixed library construction and sequencing (paired ends) were performed to obtain sequencing data. A total of 1,597,498 raw sequencing reads were generated from the 12 samples. Each sequencing read includes R1 and R2, as shown below. Figure 1 As shown.

[0057] 2. Data Splitting Read the sequencing files and split the data. For each sequencing data point (paired-end data, R1 and R2 respectively), perform the following processing: (1) Splitting based on barcode refer to Figure 2 Follow these steps to match: a. Determine if the 5' ends of FB and R1 are completely matched (i.e., no mismatch is allowed): i. If FB and R1's 5' ends are completely matched and RB exists, determine whether RB and R2's 5' ends are completely matched. If RB and R2's 5' ends are completely matched, the sequencing data is classified as identified sequences. At the same time, the sequence whose R1 end is completely matched with FB is removed, and the sequence whose R2 end is completely matched with RB is removed. If RB and R2's 5' ends are not completely matched, proceed to step b. ii. If the FB and R1 5' ends are perfectly matched but the RB is not present, the sequencing data is classified as identified sequences, and the sequences whose R1 ends are perfectly matched with the FB are removed. iii. If FB does not fully match the 5' end of R1, proceed to step b.

[0058] b. Determine if the 5' ends of FB and R2 are fully matched: i. If the 5' end of FB and R2 are completely matched and RB exists, determine whether the 5' end of RB and R1 are completely matched. If the 5' end of RB and R1 are completely matched, the sequencing data is classified as an identified sequence. At the same time, the sequence whose R2 end is completely matched with FB is removed, and the sequence whose R1 end is completely matched with RB is removed. If the 5' end of RB and R1 are not completely matched, proceed to the next sample judgment (i.e., change a barcode and continue the above judgment). ii. If the FB and R2 5' ends are perfectly matched but the RB is not present, the sequencing data is classified as identified sequences, and the sequences whose R2 ends are perfectly matched with the FB are removed. iii. If FB does not perfectly match the 5' end of R2, proceed to the next sample judgment.

[0059] If all barcode sequences fail to match, the sequencing data is classified as unidentified sequences.

[0060] (2) Splitting based on primer sequences Since a single pool typically contains approximately 2000-3000 samples, far exceeding the number of available barcodes, in practice, samples with the same barcode but different amplified fragments are often mixed together. For example, in this embodiment, samples A-05 and B-01 use the same barcode. Therefore, for identified sequences (sequences whose 5' ends have been completely removed and perfectly match the barcode), further matching based on primer sequences is required.

[0061] Compared to a perfect match of the barcode sequence, primer matching allows for a maximum of M (approximately 10-20% of the primer length; in this example, the primer length is approximately 20 bp, and M is 3) base mismatches. Furthermore, since there are N (e.g., 2) linking bases between the barcode and primer sequences, these also need to be considered during relevant matching.

[0062] refer to Figure 3 Follow these steps to match: a. Calculate the number of mismatched bases in the sequence following FP and R1 at position n1 (a natural number from 0 to N): i. If, for any value of n1, the number of mismatched bases between a given FP and R1 is no greater than M (i.e., FP and R1 match), then calculate the number of mismatched bases in the sequence following RP and R2 at position n2 (a natural number from 0 to N): 1) If, for any n² value, the number of mismatched bases between RP and R2 is no greater than M (RP and R2 match), then the sequencing sequence is classified as a matched sequence. Simultaneously, the first n¹ bases of R1 and the first n² bases of R2 are removed. 2) If for all n2 values, the number of mismatched bases between RP and R2 is greater than M (RP and R2 do not match), then proceed to step b; ii. If for all n1 values, the number of mismatched bases between all FP and R1 is greater than M (FP and R1 do not match), then proceed to step b.

[0063] b. Calculate the number of mismatched bases in the sequence following FP and R2 at position n1: i. If for any value of n1, the number of mismatched bases between FP and R2 is no greater than M, then calculate the number of mismatched bases in the sequence following RP and R1 at position n2: 1) If, for any n² value, the number of mismatched bases between RP and R1 is no greater than M, then the sequencing sequence is classified as a matched sequence. Simultaneously, the first n¹ bases of R2 are removed, and the first n² bases of R1 are removed. 2) If for all n2 values, the number of mismatched bases between RP and R1 is greater than M (PR and R1 do not match), then the current primer pair fails to match, and the next primer pair is checked. ii. If for all n1 values, the number of mismatched bases between FP and R1 is greater than M, then the current primer pair fails to match, and the next primer pair is determined.

[0064] If the sequence fails to match any primer pair, it is classified as an unmatched sequence.

[0065] For sequences that fail to match the barcode (unidentified sequences), use the following steps to perform a match: First, obtain the lengths of all FBs and all RBs, and take the maximum values ​​to obtain L1 and L2. In the sequencing data of this embodiment, since there are no RBs, L2=0.

[0066] Next, the 5' end of R1 with a length of L1 is cut off, and the 5' end of R2 with a length of L2 is cut off. The remaining sequence after the cut-off is then split according to "(2) based on primer sequence". The primer pairs used are the primer pairs of all samples. If the match is successful, the cut-off sequence is used as a candidate barcode. If the match fails, the 5' end of R1 with a length of L2 is cut off, and the 5' end of R2 with a length of L1 is cut off. The remaining sequence after the cut-off is then split according to "(2) based on primer sequence". The primer pairs used are the primer pairs of all samples. Similarly, if the match is successful, the cut-off sequence is used as a candidate barcode. If the match fails, the next unidentified sequence is matched.

[0067] 3. Data Statistics (1) Count the number of identified sequences, the number of matched sequences and the Q30 index for each sample. Determine whether the data volume meets the requirements based on the matched sequences. The results are shown in Table 2.

[0068] Table 2: Results of Sequencing Data Splitting .

[0069] The formula for calculating the efficiency e of barcode usage is as follows:

[0070] Where M is the number of matched sequences in the identified sequences, and C is the number of identified sequences.

[0071] As shown in Table 2, a total of 1,559,333 identified sequences were obtained. For some samples, the actual number of sequences was greater than the required amount of sequencing data. For example, for sample A-01, the actual amount of data exceeded 12W, which is more than twice the required amount of sequencing data (6W). The inventors took the first X sequences for flattening, where X is a random number between 1.4 and 1.5 times the required amount of sequencing data. In this embodiment, it was taken as 1.49 times.

[0072] In addition, 38,165 unidentified sequences were obtained from the sequencing data, and 10,843 matched sequences were obtained from the unidentified sequences. Eight candidate barcodes were obtained. All candidate barcodes and the number of sequences they contain are shown in Table 3.

[0073] Table 3: Candidate barcodes and the number of sequences they contain .

[0074] Example 2 The difference between this embodiment and Example 1 is that the amplified region is the bacterial V4V5 region, with barcode sequences at both ends. The FP and RP sequences are as follows.

[0075] FP: 5'-GTGYCAGCMGCCGCGGTAA-3' (SEQ ID No. 52); RP: 5'-CCGYCAATTYMTTTRAGTTT-3' (SEQ ID No. 53), The degenerate bases in the primer sequences were treated in the same way as in Example 1.

[0076] Some sample information and primer information are shown in Table 4.

[0077] Table 4: Partial Sample Information and Primer Information .

[0078] The samples were amplified using the aforementioned bacterial V4V5 region primers. After adding barcode sequences, mixed library construction and sequencing (paired ends) were performed to obtain sequencing data. A total of 1,898,286 raw sequencing reads were generated from the 9 samples. Each sequencing read includes R1 and R2, as shown below. Figure 4 As shown.

[0079] The sequencing data splitting process is the same as in Example 1. The number of identified sequences, the number of matched sequences, and the Q30 index for each sample are counted, as shown in Table 5.

[0080] Table 5: Results of Sequencing Data Splitting .

[0081] As shown in Table 5, a total of 1,898,286 identified sequences were obtained. In addition, 86,869 unidentified sequences were obtained from the sequencing data, and 74,934 matched sequences were obtained from the unidentified sequences. Five candidate barcodes were obtained. All candidate barcodes and the number of sequences they contain are shown in Table 6.

[0082] Table 6: Candidate barcodes and the number of sequences they contain .

[0083] Furthermore, it should be understood that after reading the foregoing contents of this application, those skilled in the art can make various alterations or modifications to this application, and these equivalent forms also fall within the scope defined by the appended claims.

Claims

1. A method for splitting multi-sample amplicon mixed library construction and sequencing data, characterized in that, The multi-sample amplicon pooled library construction and sequencing data were obtained through the following steps: Amplicon is obtained for each sample. The 5' end of the amplicon includes the sequence of the forward tag FB and the sequence of the forward primer FP. The 3' end of the amplicon includes the complementary sequence of the reverse primer RP. The complementary sequence of the reverse primer RP may or may not include the complementary sequence of the reverse tag RB. FB alone or together with RB constitutes a tag pair. Each sample corresponds to one tag pair. The number of samples exceeds the number of tag pairs. FP and RP constitute a primer pair. For any two samples with the same tag pair, the corresponding primer pairs are different. Amplicon samples from multiple samples were mixed to construct a library and then subjected to paired-end sequencing. Each sequencing data included two reads, R1 and R2. The splitting method includes the following steps: Obtain the label pair corresponding to each sample, as well as the primer pair set corresponding to each sample; For any sequencing data, if the 5' ends of R1 and R2 completely match a tag pair, the sequencing data is classified as identified sequences, and the sequences in which the 5' ends of R1 and / or R2 completely match the forward tag sequence or the reverse tag sequence are removed. For an identified sequence, if the 5' ends of R1 and R2 can be matched with a primer pair, the sequencing data is classified according to the sample information corresponding to the primer pair.

2. The splitting method according to claim 1, characterized in that, If FP and / or RP include degenerate primers and / or hypoxanthine, then the primer pair set includes primer pairs consisting of all specific primers corresponding to the degenerate primers and / or hypoxanthine.

3. The splitting method according to claim 1, characterized in that, For an identified sequence, after deleting the sequence whose 5' end of R1 and / or R2 completely matches the forward or reverse tag sequence, the 5' end of R1 and / or R2 still includes a linker sequence of length N, where N = 3~5. Then, the alignment of the 5' end of R1 and R2 with a primer pair means that the sequence after n1 bases from the 5' end of R1 and the sequence after n2 bases from the 5' end of R2 can align with a primer pair, where 0 ≤ n1 ≤ N, 0 ≤ n2 ≤ N, and n1 and n2 are both natural numbers.

4. The splitting method according to any one of claims 1 to 3, characterized in that, In the region of the alignment, the number of mismatched bases does not exceed M, where M = 0-2.

5. The splitting method according to any one of claims 1 to 3, characterized in that, If for any two samples, regardless of whether the corresponding label pairs are the same, the corresponding primer pairs are different, then: For sequencing data where the 5' ends of R1 and R2 do not perfectly match any of the tag pairs, split the data as follows: Obtain representative values ​​L1 and L2 for the lengths of all samples FB and RB. Remove the sequence with a 5' end length of L1 from one of R1 and R2, and remove the sequence with a 5' end length of L2 from the other. Obtain the set of primer pairs for all samples; If the 5' ends of R1 and R2 after sequence excision can be matched with a primer pair, the sequencing data is classified according to the sample information corresponding to the primer pair.

6. The splitting method according to claim 5, characterized in that, The representative value is selected from one of the maximum value, mode, and upper quartile, or determined based on any one of them.

7. A system for splitting multi-sample amplicon mixed library construction and sequencing data, comprising the following modules: The data input module is used to input the multi-sample amplicon pooled library construction and sequencing data, which is obtained through the following steps: Amplicones were obtained for each sample. The 5' end of each amplicon sequentially included the sequence of the forward tag FB and the sequence of the forward primer FP. The 3' end of each amplicon included the complementary sequence of the reverse primer RP. The complementary sequence of the reverse primer RP may or may not include the complementary sequence of the reverse tag RB. FB (First Implantable Filter) can be used alone or together with RB (Second Implantable Filter) to form a tag pair. Each sample corresponds to one tag pair, and the number of samples exceeds the number of tag pairs. FP (First Implantable Filter) and RP (Second Implantable Filter) form a primer pair. For any two samples with the same tag pair, the corresponding primer pairs are different. Amplicon samples from multiple samples were mixed to construct a library and then subjected to paired-end sequencing. Each sequencing data included two reads, R1 and R2. The data input module is also used to input the sample information from which the multi-sample amplicon mixed library construction and sequencing data comes from, as well as the tag pairs and primer pairs corresponding to each sample; The tag-based splitting module is connected to the data input module. The data input module is also used to input the sample information of the source of the multi-sample amplicon mixed library construction and sequencing data and the tag pairs corresponding to each sample, for splitting based on the following steps: For any sequencing data, if the 5' ends of R1 and R2 completely match a tag pair, the sequencing data is classified as identified sequences, and sequences in which the 5' ends of R1 and R2 completely match the forward tag sequence or the reverse tag sequence are removed. The primer pair splitting module is connected to both the data input module and the tag pair splitting module. The data input module is also used to obtain the primer pairs corresponding to the samples. This module is used to further split the identified sequences based on the following steps: If the 5' ends of R1 and R2 can be matched with a primer pair, the sequencing data is classified according to the sample information corresponding to the primer pair.

8. The splitting system according to claim 7, characterized in that, Also includes: The unidentified sequence splitting module, together with the data input module and the label-based splitting module, performs splitting based on the input information from the data input module if, for any two samples, regardless of whether the corresponding label pairs are the same, the corresponding primer pairs are different: For sequencing data where the 5' ends of R1 and R2 do not perfectly match any of the tag pairs, split the data as follows: Obtain representative values ​​L1 and L2 for the lengths of all samples FB and RB. Remove the sequence with a 5' end length of L1 from one of R1 and R2, and remove the sequence with a 5' end length of L2 from the other. Obtain the set of primer pairs for all samples; If the 5' ends of R1 and R2 after sequence excision can be matched with a primer pair, the sequencing data is classified according to the sample information corresponding to the primer pair.

9. A computer device, characterized in that, include: Memory, used to store computer programs; A processor, configured to execute the steps of the splitting method according to any one of claims 1 to 6 when executing the computer program.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the splitting method according to any one of claims 1 to 6.