Method and apparatus for pooled sequencing, electronic device and storage medium
By directly sequencing different batches of libraries to be sequenced and performing sequence assembly and evaluation in the same sequencing chip, the high cost and time waste caused by chip cleaning in Nanopore sequencing technology are solved, thus improving sequencing efficiency.
Patent Information
- Application Number
- CN202310462349.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-21
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2043-04-21
AI Technical Summary
Nanopore sequencing technology requires cleaning the sequencing chip when sequencing adjacent batches, resulting in high reagent and labor costs, and the cleaning time is long, which affects sequencing efficiency.
In the same sequencing chip, different batches of libraries to be sequenced are directly sequenced. Through sequence assembly and evaluation, the assembly sequence of each is determined by similarity, reducing the cleaning steps.
This reduces the time spent cleaning sequencing chips and the amount of reagents used, thus improving sequencing efficiency.
Smart Images

Figure CN116504309B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of data processing, and particularly relates to a method and device for mixed sample sequencing, an electronic device and a storage medium. BACKGROUND
[0002] Compared with previous sequencing technologies, the Nanopore sequencing technology does not need to perform polymerase chain reaction (PCR) amplification in the sequencing process, realizes individual sequencing of each DeoxyriboNucleic Acid (DNA) molecule, reduces the process of constructing a library, and improves the efficiency of sequencing. The Nanopore sequencing technology has the characteristics of longer read length, fast running speed, and real-time acquisition of sequence information. In the same sequencing chip, a single barcode can support mixed sample library construction of 96 samples at most. The Nanopore sequencing chip supports clean reuse, that is, when sequencing adjacent two batches of samples, after the sequencing of the first batch is completed, the sequencing chip is cleaned, and then the sequencing of the second batch can be continued based on the cleaned chip. However, the reagent, labor and related material costs used for cleaning each time are high, and the cleaning takes a long time, which affects the efficiency of sequencing. SUMMARY
[0003] The present disclosure provides a method and device for mixed sample sequencing, an electronic device and a storage medium. The main purpose is to realize sequencing between different batches without cleaning the sequencing chip for direct mixed sample sequencing.
[0004] According to a first aspect of the present disclosure, a method for mixed sample sequencing is provided, which comprises the following steps.
[0005] In the same sequencing chip, after at least one first test sequence set is obtained by performing sequencing processing on a first to-be-sequenced library, at least one second test sequence set is obtained by performing sequencing processing on a second to-be-sequenced library; wherein the first to-be-sequenced library comprises at least one first to-be-sequenced sample, and the second to-be-sequenced library comprises at least one second to-be-sequenced sample;
[0006] Each test sequence in each of the first test sequence sets is subjected to sequence assembly to obtain a first assembly sequence set corresponding to each of the first to-be-sequenced samples; and each test sequence in each of the second test sequence sets is subjected to sequence assembly to obtain a second assembly sequence set corresponding to each of the second to-be-sequenced samples;
[0007] evaluating each of the first assembly sequence set to obtain a first assembly sequence corresponding to each of the first sequencing sample to be sequenced respectively;
[0008] determining the assembly sequence corresponding to each of the at least one first sequencing sample to be sequenced, the assembly sequence corresponding to each of the at least one second sequencing sample to be sequenced respectively based on the similarity between the at least one first assembly sequence and the at least one second assembly sequence, the at least one third assembly sequence respectively.
[0009] Optionally, after obtaining the at least one second test sequence set by sequencing the second sequencing library, the method further comprises:
[0010] continuing to sequence the third sequencing library in the same sequencing chip to obtain at least one third test sequence set; wherein the third sequencing library comprises at least one third sequencing sample to be sequenced;
[0011] performing sequence assembly on each of the test sequence in the third test sequence set to obtain a third assembly sequence set corresponding to each of the third sequencing sample to be sequenced;
[0012] evaluating each of the third assembly sequence set to obtain a fourth assembly sequence, a fifth assembly sequence, a sixth assembly sequence corresponding to each of the third sequencing sample to be sequenced respectively;
[0013] determining the assembly sequence corresponding to each of the at least one third sequencing sample to be sequenced based on the similarity between the assembly sequence corresponding to each of the at least one first sequencing sample to be sequenced, the assembly sequence corresponding to each of the at least one second sequencing sample to be sequenced and the fourth assembly sequence, the fifth assembly sequence, the sixth assembly sequence respectively.
[0014] Optionally, the performing sequence assembly on each of the test sequence in the first test sequence set to obtain a first assembly sequence set corresponding to each of the first sequencing sample to be sequenced; the performing sequence assembly on each of the test sequence in the second test sequence set to obtain a second assembly sequence set corresponding to each of the second sequencing sample to be sequenced comprises:
[0015] performing sequence assembly on each of the test sequence in the first test sequence set to obtain a first assembly sequence set corresponding to each of the first sequencing sample to be sequenced respectively;
[0016] performing sequence assembly on each of the test sequence in the second test sequence set to obtain a second assembly sequence set corresponding to each of the second sequencing sample to be sequenced respectively.
[0017] Optionally, the sequence assembling is performed on each test sequence in the first test sequence set to obtain a first assembled sequence set corresponding to each of the first to-be-sequenced samples; and the sequence assembling is performed on each test sequence in the second test sequence set to obtain a second assembled sequence set corresponding to each of the second to-be-sequenced samples.
[0018] The preset assembly algorithm is called to perform sequence assembling on each test sequence in the first test sequence set to obtain a first sequence set corresponding to each of the first to-be-sequenced samples.
[0019] The preset assembly algorithm is called to perform sequence assembling on each test sequence in the second test sequence set to obtain a second sequence set corresponding to each of the second to-be-sequenced samples.
[0020] Optionally, after the preset assembly algorithm is called to perform sequence assembling on each test sequence in the second test sequence set to obtain a second sequence set corresponding to each of the second to-be-sequenced samples, the method further comprises:
[0021] It is determined whether the number of sequences in the first sequence set and / or the second sequence set is greater than one.
[0022] In a case where the number of sequences in the first sequence set and / or the second sequence set is greater than one, a de-redundancy processing is performed on each of the first sequence set and / or the second sequence set to obtain each of the first assembled sequence set and the second assembled sequence set.
[0023] Optionally, the de-redundancy processing performed on each of the first sequence set and / or the second sequence set comprises:
[0024] Based on a first preset similarity threshold, assembled sequences in the first sequence set are compared with each other to remove the assembled sequences with a short sequence length.
[0025] Based on the first preset similarity threshold, assembled sequences in the second sequence set are compared with each other to remove the assembled sequences with a short sequence length.
[0026] Optionally, the evaluation performed on each of the first assembled sequence set to obtain a first assembled sequence corresponding to each of the first to-be-sequenced samples comprises:
[0027] In a case where the first assembled sequence set includes one assembled sequence, the assembled sequence is taken as the first assembled sequence corresponding to the first assembled sequence set.
[0028] In a case where it is determined that the first assembly sequence set comprises at least two assembly sequences, performing abundance evaluation on all assembly sequences in the first assembly sequence set, and determining an assembly sequence ranked first in the abundance evaluation as a first assembly sequence corresponding to the first assembly sequence set.
[0029] Optionally, the evaluating each of the second assembly sequence set obtains a second assembly sequence and a third assembly sequence corresponding to each of the second sequencing sample respectively, comprises:
[0030] In a case where it is determined that the second assembly sequence set comprises one assembly sequence, aligning the second assembly sequence set with the first assembly sequence in the first assembly sequence set;
[0031] In a case where it is determined that the second assembly sequence set comprises two assembly sequences, taking the two assembly sequences as the second assembly sequence and the third assembly sequence corresponding to the second assembly sequence set respectively;
[0032] In a case where it is determined that the second assembly sequence set comprises at least three assembly sequences, performing abundance evaluation on assembly sequences in the second assembly sequence set, and taking assembly sequences ranked first and second in the abundance evaluation as the second assembly sequence and the third assembly sequence corresponding to the second assembly sequence set respectively.
[0033] Optionally, the aligning the second assembly sequence set with the first assembly sequence in the first assembly sequence set comprises:
[0034] Calculating a third similarity between the assembly sequence in the second assembly sequence set and the first assembly sequence with the same marker sequence;
[0035] In a case where it is determined that the third similarity is greater than or equal to the second preset similarity threshold, resequencing the second sequencing sample corresponding to the marker sequence;
[0036] In a case where it is determined that the third similarity is less than the second preset similarity threshold, taking the assembly sequence in the second assembly sequence set as an assembly sequence corresponding to the second sequencing library.
[0037] Optionally, the determining the assembly sequence corresponding to each of the first sequencing sample, the assembly sequence corresponding to each of the second sequencing sample based on the similarity between at least one of the first assembly sequence and at least one of the second assembly sequence and at least one of the third assembly sequence, comprises:
[0038] determining the first assembly sequence corresponding to each of the first to-be-sequenced samples as the first assembly sequence;
[0039] performing similarity comparison between the second assembly sequence with the same identification sequence and the first assembly sequence respectively to obtain a first similarity corresponding to each of the identification sequences;
[0040] performing similarity comparison between the third assembly sequence with the same identification sequence and the first assembly sequence respectively to obtain a second similarity corresponding to each of the identification sequences;
[0041] determining the first assembly sequence corresponding to each of the second to-be-sequenced samples as the first assembly sequence based on the first similarity and the second similarity; wherein the first assembly sequence corresponding to the first similarity is the same as the first assembly sequence corresponding to the second similarity.
[0042] Optionally, the determining the first assembly sequence corresponding to each of the second to-be-sequenced samples as the first assembly sequence based on the first similarity and the second similarity comprises:
[0043] in a case where it is determined that the first similarity is greater than or equal to a second preset similarity threshold and the second similarity is less than the second preset similarity threshold, or there is no result, determining the third assembly sequence as the assembly sequence corresponding to the second to-be-sequenced sample;
[0044] in a case where it is determined that the second similarity is greater than or equal to the second preset similarity threshold and the first similarity is less than the second preset similarity threshold, or there is no result, determining the second assembly sequence as the assembly sequence corresponding to the second to-be-sequenced sample;
[0045] in a case where it is determined that the first similarity and the second similarity are both greater than or both less than the second preset similarity threshold, re-sequencing the second to-be-sequenced sample.
[0046] According to a second aspect of the present disclosure, a device for mixed sample sequencing is provided, comprising:
[0047] a first sequencing unit, configured to, in a same sequencing chip, after obtaining at least one first test sequence set by performing sequencing processing on a first to-be-sequenced library, obtain at least one second test sequence set by performing sequencing processing on a second to-be-sequenced library; wherein the first to-be-sequenced library comprises at least one first to-be-sequenced sample, and the second to-be-sequenced library comprises at least one second to-be-sequenced sample;
[0048] a first assembling unit, configured to perform sequence assembly on each test sequence in the first test sequence set to obtain a first assembled sequence set corresponding to each of the first to-be-sequenced samples, and perform sequence assembly on each test sequence in the second test sequence set to obtain a second assembled sequence set corresponding to each of the second to-be-sequenced samples;
[0049] a first evaluating unit, configured to evaluate each of the first assembled sequence set to obtain a first assembled sequence corresponding to each of the first to-be-sequenced samples, and evaluate each of the second assembled sequence set to obtain a second assembled sequence and a third assembled sequence corresponding to each of the second to-be-sequenced samples;
[0050] a first determining unit, configured to determine, based on similarity between at least one of the first assembled sequence and at least one of the second assembled sequence and at least one of the third assembled sequence, an assembled sequence corresponding to each of the at least one first to-be-sequenced sample and an assembled sequence corresponding to each of the at least one second to-be-sequenced sample.
[0051] Optionally, the apparatus further includes:
[0052] a second sequencing unit, configured to, after obtaining the at least one second test sequence set by performing sequencing processing on the second to-be-sequenced library, continue to perform sequencing processing on a third to-be-sequenced library in the same sequencing chip to obtain at least one third test sequence set, wherein the third to-be-sequenced library includes at least one third to-be-sequenced sample;
[0053] a second assembling unit, configured to perform sequence assembly on each test sequence in the third test sequence set to obtain a third assembled sequence set corresponding to each of the third to-be-sequenced samples;
[0054] a second evaluating unit, configured to evaluate each of the third assembled sequence set to obtain a fourth assembled sequence, a fifth assembled sequence and a sixth assembled sequence corresponding to each of the third to-be-sequenced samples;
[0055] a second determining unit, configured to determine, based on similarity between the assembled sequence corresponding to each of the at least one first to-be-sequenced sample and the assembled sequence corresponding to each of the at least one second to-be-sequenced sample and the fourth assembled sequence, the fifth assembled sequence and the sixth assembled sequence, an assembled sequence corresponding to each of the at least one third to-be-sequenced sample.
[0056] Optionally, the first assembling unit includes:
[0057] a first assembling module, configured to perform sequence assembly on each test sequence in the first test sequence set to obtain a first assembled sequence set corresponding to each of the first to-be-sequenced samples;
[0058] a second assembling module, configured to perform sequence assembly on each test sequence in the second test sequence set respectively to obtain a second assembled sequence set corresponding to each of the second to-be-sequenced samples.
[0059] Optionally, the first assembling unit further includes:
[0060] a third assembling module, configured to call a preset assembly algorithm to perform sequence assembly on each test sequence in the first test sequence set respectively to obtain a first sequence set corresponding to each of the first to-be-sequenced samples;
[0061] a fourth assembling module, configured to call the preset assembly algorithm to perform sequence assembly on each test sequence in the second test sequence set respectively to obtain a second sequence set corresponding to each of the second to-be-sequenced samples.
[0062] Optionally, the apparatus further includes:
[0063] a third determining unit, configured to determine whether the number of sequences in the first sequence set and / or the second sequence set is greater than one after the preset assembly algorithm is called to perform sequence assembly on each test sequence in the second test sequence set respectively to obtain the second sequence set corresponding to each of the second to-be-sequenced samples;
[0064] a processing unit, configured to perform de-redundancy processing on each of the first sequence set and / or each of the second sequence set respectively to obtain each of the first assembled sequence set and each of the second assembled sequence set in the case that the number of sequences in the first sequence set and / or the second sequence set is greater than one.
[0065] Optionally, the processing unit includes:
[0066] a first comparing module, configured to compare the assembled sequences in the first sequence set with each other based on a first preset similarity threshold to remove the assembled sequences with short sequence length;
[0067] a second comparing module, configured to compare the assembled sequences in the second sequence set with each other based on the first preset similarity threshold to remove the assembled sequences with short sequence length.
[0068] Optionally, the first evaluating unit includes:
[0069] a first determining module, configured to take an assembled sequence as the first assembled sequence corresponding to the first assembled sequence set in the case that the first assembled sequence set includes the assembled sequence;
[0070] The second determining module is configured to, in a case where the first assembly sequence set comprises at least two assembly sequences, evaluate the abundance of all assembly sequences in the first assembly sequence set, and determine an assembly sequence with the highest abundance evaluation as the first assembly sequence corresponding to the first assembly sequence set.
[0071] Optionally, the first evaluation unit further comprises:
[0072] The first comparison module is configured to, in a case where the second assembly sequence set comprises one assembly sequence, compare the second assembly sequence set with the same identifier sequence with the first assembly sequence in the first assembly sequence set;
[0073] The third determining module is configured to, in a case where the second assembly sequence set comprises two assembly sequences, determine the two assembly sequences as the second assembly sequence and the third assembly sequence corresponding to the second assembly sequence set, respectively.
[0074] The fourth determining module is configured to, in a case where the second assembly sequence set comprises at least three assembly sequences, evaluate the abundance of the assembly sequences in the second assembly sequence set, and determine the assembly sequences with the first and second highest abundance evaluations as the second assembly sequence and the third assembly sequence corresponding to the second assembly sequence set, respectively.
[0075] Optionally, the first comparison module is further configured to:
[0076] Calculate a third similarity between the assembly sequence in the second assembly sequence set and the first assembly sequence with the same identifier sequence;
[0077] In a case where the third similarity is greater than or equal to the second preset similarity threshold, re-sequencing the second sample to be sequenced corresponding to the identifier sequence;
[0078] In a case where the third similarity is less than the second preset similarity threshold, determining the assembly sequence in the second assembly sequence set as the assembly sequence corresponding to the second library to be sequenced.
[0079] Optionally, the first determining unit comprises:
[0080] The fifth determining module is configured to determine the first assembly sequence as the assembly sequence corresponding to each of the first samples to be sequenced, respectively.
[0081] The second comparison module is configured to compare the second assembly sequence with the same identifier sequence with the first assembly sequence for similarity, to obtain a first similarity corresponding to each identifier sequence.
[0082] a third comparison module, configured to perform similarity comparison between the third assembled sequences with the same identifier sequence and the first assembled sequences respectively, to obtain a second similarity corresponding to each of the identifier sequences;
[0083] a sixth determination module, configured to determine an assembled sequence corresponding to each of the second to-be-sequenced samples based on the first similarity and the second similarity, wherein the first similarity corresponds to the first assembled sequence, and the second similarity corresponds to the first assembled sequence.
[0084] Optionally, the sixth determination module is further configured to:
[0085] determine the third assembled sequence as the assembled sequence corresponding to the second to-be-sequenced sample when it is determined that the first similarity is greater than or equal to a second preset similarity threshold and the second similarity is less than the second preset similarity threshold, or there is no result;
[0086] determine the second assembled sequence as the assembled sequence corresponding to the second to-be-sequenced sample when it is determined that the second similarity is greater than or equal to the second preset similarity threshold and the first similarity is less than the second preset similarity threshold, or there is no result;
[0087] re-sequence the second to-be-sequenced sample when it is determined that the first similarity and the second similarity are both greater than or both less than the second preset similarity threshold.
[0088] According to a third aspect of the present disclosure, an electronic device is provided, comprising:
[0089] at least one processor; and
[0090] a memory connected to the at least one processor in communication; wherein
[0091] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of the first aspect.
[0092] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method of the first aspect.
[0093] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method of the first aspect.
[0094] The method and device, electronic device and storage medium provided by the present disclosure can reduce the time spent on cleaning the sequencing chip during adjacent batch sequencing, improve the efficiency of sequencing, and reduce the use of cleaning reagents during cleaning of the sequencing chip.
[0095] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0096] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:
[0097] Figure 1 A flowchart of a mixed sample sequencing method provided by an embodiment of the present disclosure is shown in the figure;
[0098] Figure 2 A flowchart of another mixed sample sequencing method provided by an embodiment of the present disclosure is shown in the figure;
[0099] Figure 3 A flowchart of another mixed sample sequencing method provided by an embodiment of the present disclosure is shown in the figure;
[0100] Figure 4A schematic diagram of a method for contig de-duplication;
[0101] Figure 5 A schematic diagram of another method for contig de-duplication;
[0102] Figure 6 A schematic diagram of a method for another method of pooled sequencing provided by an embodiment of the present disclosure;
[0103] Figure 7 A schematic diagram of a method for another method of pooled sequencing provided by an embodiment of the present disclosure;
[0104] Figure 8 A schematic diagram of a method for another method of pooled sequencing provided by an embodiment of the present disclosure;
[0105] Figure 9 A schematic diagram of a device for pooled sequencing provided by an embodiment of the present disclosure;
[0106] Figure 10 A schematic diagram of a device for pooled sequencing provided by an embodiment of the present disclosure;
[0107] Figure 11 A schematic block diagram of an example electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0108] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, which are provided for the purpose of illustration only. The details of the embodiments of the present disclosure are described herein for the purpose of providing what is believed to be the presently understood best mode of the disclosure, and, thus, the description is not intended to limit the scope of the disclosure. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, it will be apparent to one skilled in the art that the embodiments of the present disclosure can be practiced without the specific details. In other instances, well-known methods have not been described in detail in order to avoid obscuring aspects of the embodiments of the present disclosure.
[0109] In existing Nanopore sequencing, two libraries of adjacent batches are sequenced according to the following steps:
[0110] Step 1. Add the prepared library 1 to the Nanopore sequencer.
[0111] Step 2. Perform Nanopore sequencing on library 1.
[0112] Step 3. Pause the sequencing when enough data for analysis is reached.
[0113] Step 4. Remove the remaining sample on the sequencing chip with a cleaning reagent.
[0114] Step 5. Add library 2 to the Nanopore sequencer.
[0115] Step 6, Nanopore sequencing of library 2.
[0116] The sequencing chip can be fully utilized by repeating steps 1-6 for six steps. After the sequencing of library 1 and library 2 is completed, the sequences obtained by sequencing steps 2 and 6 are assembled to obtain the assembled sequencing results of library 1 and library 2, respectively. Using this method, different plasmid sample data can be effectively distinguished. This method is also suitable for plasmid Nanopore sequencing using multi-sample barcode pooling (centralized resources), that is, the plasmid 1 library and the plasmid 2 library can be a multi-sample barcode pooling library or a single-sample library. If it is a multi-barcode, each barcode sequence is assembled separately to obtain the assembly result of the plasmid sample corresponding to the barcode. However, the disadvantage of this method is that the cost of reagents, labor and related materials for each cleaning is several hundred yuan, which is high in cost, and the cleaning takes a long time, about 1 hour.
[0117] The method and device for mixed sample sequencing, the electronic device and the storage medium of the embodiments of the present disclosure are described below with reference to the accompanying drawings.
[0118] Figure 1 A flowchart of a method for mixed sample sequencing provided by an embodiment of the present disclosure.
[0119] As Figure 1 shown, the method comprises the following steps:
[0120] Step 101, after obtaining at least one first test sequence set by sequencing processing of a first library to be sequenced in the same sequencing chip, obtaining at least one second test sequence set by sequencing processing of a second library to be sequenced; wherein the first library to be sequenced comprises at least one first sample to be sequenced, and the second library to be sequenced comprises at least one second sample to be sequenced.
[0121] In the embodiment of the present disclosure, the first library to be sequenced is added to the Nanopore sequencer for sequencing, and the sequencing is paused after reaching the preset sequencing data amount (the preset sequencing data amount generally reaches more than five times the length of the sequencing sequence, but in order to obtain better quality of the subsequent sequence assembly result, the sequencing is paused after reaching fifty times), and the second library to be sequenced is added to the same sequencing chip for continuous sequencing. After the sequencing is completed, at least one first test sequence set and at least one second test sequence set are obtained. It should be noted that the foregoing description is exemplary and does not constitute a limitation on the present disclosure. After the sequencing of the first library to be sequenced is completed, the sequencing chip does not need to be cleaned, and the second library to be sequenced is directly added to the same sequencing chip for sequencing; which can reduce the use of cleaning reagents, and can also reduce the time spent on cleaning.
[0122] The first sequencing library is a first batch of sequencing library, and the second sequencing library is a second batch of sequencing library. The first sequencing library and the second sequencing library can be a multi-sample barcode pooling library or a single-sample library without barcode. That is, the first sequencing library includes at least one first sequencing sample, and the second sequencing library includes at least one second sequencing sample, and the present disclosure does not constitute a limitation in this regard.
[0123] In step 102, sequence assembly is performed on the test sequences in each of the first test sequence set to obtain a first assembly sequence set corresponding to each of the first sequencing sample; and sequence assembly is performed on the test sequences in each of the second test sequence set to obtain a second assembly sequence set corresponding to each of the second sequencing sample.
[0124] In the embodiments of the present disclosure, since the DNA sequence needs to be split into a short sequence with a length within the read length of the sequencer for sequencing, the sequencing result of each sequencing sample is the result of the read of the DNA sequence with different lengths, that is, the test sequence set. After the first sequencing library and the second sequencing library are sequenced, the obtained sequencing sequences need to be subjected to sequence assembly to obtain relatively complete DNA sequences.
[0125] For the convenience of understanding the embodiments of the present disclosure, the first sequencing library and the second sequencing library can be understood as single-sample libraries. Due to contamination of the test chip and other reasons, the first test sequence set can include not only the collection of reads of the first sequencing sample, but also reads of the pollution source; after the first sequencing library is sequenced, the sequencing chip is not cleaned and is directly used for sequencing the second sequencing library; in this way, the second test sequence set can also include the collection of reads of the first sequencing sample and the collection of reads of the second sequencing sample. Sequence assembly is performed on the first test sequence set and the second test sequence set respectively to obtain a first assembly sequence set and a second assembly sequence set. The first assembly sequence set and the second assembly sequence set can include more than one DNA assembly sequence.
[0126] For example, after the plasmid 1 library and the plasmid 2 library are sequenced according to step 101, the first assembly sequence set corresponding to the plasmid 1 and the second assembly sequence set corresponding to the plasmid 2 are obtained by assembling the sequencing results. The first assembly sequence set corresponding to the plasmid 1 includes the DNA sequence corresponding to the plasmid 1, and other DNA sequences can exist if there is contamination. Since the sequencing chip is not cleaned after the plasmid 1 is sequenced and is directly used for sequencing the plasmid 2, the second assembly sequence set corresponding to the plasmid 2 includes the DNA sequence corresponding to the plasmid 2 and can also include the DNA sequence corresponding to the plasmid 1.
[0127] By analogy, the first sequencing library and the second sequencing library are the results of sequencing, assembling of multi-sample pooling library. The corresponding assembly results are obtained by assembling the fastq sequences of each barcode separately. For example, plasmids 1 and 2 are used for library construction using barcode 1 and barcode 2 respectively to obtain the first sequencing library, and the corresponding sequences a1.fastq and a2.fastq are obtained after sequencing, and the assembly results a1 and a2 are obtained by assembling. Plasmids 3 and 4 are used for library construction using barcode 1 and barcode 2 respectively to obtain the second sequencing library, and the corresponding sequences b1.fastq and b2.fastq are obtained after sequencing, and the assembly results b1 and b2 are obtained by assembling. b1.fastq corresponds to the same barcode 1 as a1.fastq, wherein a1 contains sequencing data of plasmid 1 (without considering the case of being contaminated), and b1 contains sequencing data of plasmid 1 and plasmid 3. The foregoing description is exemplary and does not constitute a limitation on the disclosure.
[0128] In step 103, each of the first assembly sequence sets is evaluated to obtain a corresponding first assembly sequence of each of the first sequencing samples; and each of the second assembly sequence sets is evaluated to obtain a corresponding second assembly sequence and a third assembly sequence of each of the second sequencing samples.
[0129] In the embodiments of the disclosure, in order to facilitate understanding of the embodiments of the disclosure, the first sequencing library and the second sequencing library can be understood as single-sample libraries. By evaluating the assembly sequences in the assembly sequence sets, assembly sequences closer to the true sequences of the sequencing samples are obtained.
[0130] Due to contamination of the test chip and other reasons, there can be multiple assembly sequences in the corresponding first assembly sequence set obtained from the first sequencing sample; it is necessary to evaluate the assembly sequences in the first assembly sequence set to obtain the first assembly sequence corresponding to the first sequencing sample. Since the second sequencing sample is directly added to the sequencer for sequencing without cleaning the sequencing chip, the first sequencing sample and the second sequencing sample are present in the second assembly sequence set. By evaluating the assembly sequences in the second assembly sequence set, the second assembly sequence and the third assembly sequence can be obtained; however, it is necessary to further align to determine which assembly sequence is the assembly sequence of the second sequencing sample. It should be noted that the disclosure does not limit the assembly sequence of which sample to correspond to the second assembly sequence or the third assembly sequence.
[0131] For the case of multi-sample pooling library construction of the first library and the second library, the embodiments of the present disclosure will not be repeated.
[0132] Step 104, based on the similarity between at least one of the first assembled sequences and at least one of the second assembled sequences and at least one of the third assembled sequences, respectively, determine the respective assembled sequences corresponding to the at least one first sequencing sample, the respective assembled sequences corresponding to the at least one second sequencing sample.
[0133] In the embodiments of the present disclosure, the first library and the second library are also understood as single-sample library. The similarity between the second assembled sequence and the first assembled sequence, and the similarity between the third assembled sequence and the first assembled sequence are calculated respectively; the assembled sequence corresponding to the second sequencing sample is determined by comparing the similarities. The first assembled sequence obtained by evaluation is the assembled sequence corresponding to the first sequencing sample. By the method of the present disclosure, the first sequencing sample and the second sequencing sample are sequenced on the same sequencing chip in turn, which can reduce the reagent consumption and time waste caused by cleaning the sequencing chip between different batches. By evaluating and selecting the assembled sequence that is more consistent with the actual sequence of the sequencing sample, the similarity calculation is performed to determine the assembled sequence of the second sequencing sample in the second assembled sequence. Although a certain amount of computing resources are required, compared with the method in the prior art, the method disclosed in the present disclosure can effectively improve the sequencing efficiency.
[0134] This disclosure provides a method for mixed-sample sequencing. In the same sequencing chip, after sequencing a first library to be sequenced to obtain at least one first test sequence set, sequencing a second library to be sequenced to obtain at least one second test sequence set. The first library to be sequenced includes at least one first sample to be sequenced, and the second library to be sequenced includes at least one second sample to be sequenced. Sequences are assembled from the test sequences in each of the first test sequence sets to obtain a first assembled sequence set corresponding to each of the first samples to be sequenced. Sequences are also assembled from the test sequences in each of the second test sequence sets to obtain a second assembled sequence set corresponding to each of the second samples to be sequenced. Each of the first assembled sequence sets is evaluated to obtain a first assembled sequence corresponding to each of the first samples to be sequenced. Each of the second assembled sequence sets is evaluated to obtain a second assembled sequence and a third assembled sequence corresponding to each of the second samples to be sequenced. Based on the similarity between at least one first assembled sequence and at least one second assembled sequence and at least one third assembled sequence, the assembled sequences corresponding to each of the at least one first sample to be sequenced and the assembled sequences corresponding to each of the at least one second sample to be sequenced are determined. Compared with related technologies, the embodiments of this disclosure, by sequentially sequencing the first and second sequencing libraries in the same sequencing chip, assembling and evaluating the sequenced sequences, obtain the assembled sequences corresponding to the first and second sequencing libraries; this disclosure can reduce the time spent cleaning the sequencing chip during adjacent sequencing batches and improve sequencing efficiency.
[0135] To clearly illustrate the embodiments of this disclosure, a flowchart of another method for pooled sequencing is provided.
[0136] like Figure 2 As shown, the method includes the following steps:
[0137] Step 201: In the same sequencing chip, continue sequencing the third library to be sequenced to obtain at least one third test sequence set; wherein the third library to be sequenced includes at least one third sample to be sequenced.
[0138] Specifically, in this embodiment of the disclosure, the third sequencing library is a third batch of sequencing libraries. The third sequencing library can be a multi-sample barcode pooling library or a single-sample library without barcodes.
[0139] After sequencing the second batch of libraries, the sequencing chip is not cleaned; the second library is directly added to the same sequencing chip for sequencing to obtain the third test sequence set.
[0140] Step 202, sequence assembly is performed on each test sequence in each of the third test sequence set, to obtain a third assembly sequence set corresponding to each of the third to-be-sequenced sample.
[0141] Specifically, the third test sequence set may include sequencing sequences of the first two batches of sequencing. After assembly of the third test sequence set, a third assembly sequence set is obtained.
[0142] Step 203, evaluation is performed on each of the third assembly sequence set, to obtain a fourth assembly sequence, a fifth assembly sequence, and a sixth assembly sequence corresponding to each of the third to-be-sequenced sample.
[0143] Specifically, the fourth assembly sequence, the fifth assembly sequence, and the sixth assembly sequence corresponding to each of the third to-be-sequenced sample are obtained by evaluating the obtained third assembly sequence set. However, which assembly sequence is the assembly sequence of the third to-be-sequenced sample needs to be determined through further comparison. It should be noted that the present disclosure does not limit which assembly sequence of the sample corresponds to the fourth assembly sequence, the fifth assembly sequence, and the sixth assembly sequence.
[0144] Step 204, based on the similarity between the assembly sequence corresponding to each of the at least one first to-be-sequenced sample, the assembly sequence corresponding to each of the at least one second to-be-sequenced sample, and the fourth assembly sequence, the fifth assembly sequence, and the sixth assembly sequence, respectively, the assembly sequence corresponding to each of the at least one third to-be-sequenced sample is determined.
[0145] The similarity between the assembly sequence corresponding to each of the at least one first to-be-sequenced sample, the assembly sequence corresponding to each of the at least one second to-be-sequenced sample, and the fourth assembly sequence, the fifth assembly sequence, and the sixth assembly sequence, respectively, is calculated. The assembly sequence corresponding to the third to-be-sequenced sample is determined by comparing the similarity.
[0146] The present embodiment of the present disclosure describes the case of mixed sample sequencing of the third batch of to-be-sequenced libraries. The mixed sample sequencing method provided by the present disclosure can also be applied to more batches of mixed sample sequencing. However, in the case of mixed sample sequencing of multiple batches, the redundancy of the to-be-sequenced sample leads to a decrease in the accuracy of sequencing. Therefore, when performing more batches of mixed sample sequencing, the size of the to-be-sequenced sample and the sequencing throughput at each time of sequencing need to be controlled as much as possible. The amount of data to be evaluated and the amount of similarity data to be calculated in more batches of mixed sample sequencing will be more. It should be noted that the present embodiment of the present disclosure does not limit the maximum number of batches of sequencing supported by the mixed sample sequencing.
[0147] In order to clearly illustrate the present embodiment of the present disclosure, the present embodiment of the present disclosure provides a flowchart of another method of mixed sample sequencing.
[0148] AsFigure 3 As shown, the method comprises the following steps:
[0149] Step 301, any first sequencing sample in the first sequencing library and any second sequencing sample in the second sequencing library are used to construct a library using the same identification sequence.
[0150] In particular, in the embodiments of the present disclosure, in order to have a better understanding of the present disclosure, the embodiments of the present disclosure are described as a multi-sample pooling library of the first sequencing library and the second sequencing library. The identification sequence is a barcode used when constructing a multi-sample pooling library. The first sequencing library corresponds to the first batch of samples, and the second sequencing library corresponds to the second batch of samples. A total of 12 samples in two batches are sequenced by Nanopore mixed sample sequencing. As shown in Table 1, the first batch of six samples is plasmid 1, 2, 3, 4, 5, and 6, which corresponds to the use of barcodes 1, 2, 3, 4, 5, and 6. The second batch of six samples is plasmid 7, 8, 9, 10, 11, and 12, which corresponds to the use of barcodes 1, 2, 3, 4, 5, and 6.
[0151] Table 1 Two batches of 12 sample tables
[0152] barcode1 barcode2 barcode3 barcode4 barcode5 barcode6 Batch 1 Plasmid 1 Plasmid 2 Plasmid 3 Plasmid 4 Plasmid 5 Plasmid 6 Batch 2 Plasmid 7 Plasmid 8 Plasmid 9 Plasmid 10 Plasmid 11 Plasmid 12
[0153] Currently, Nanopore uses a single barcode to support up to 96 samples for mixed sample library construction. The embodiments of the present disclosure take 12 samples of the first sequencing library and the second sequencing library (the first batch and the second batch) as an example for description, which does not constitute a limitation on the present disclosure.
[0154] Step 302, after obtaining at least one first test sequence set by sequencing processing of the first sequencing library in the same sequencing chip, at least one second test sequence set is obtained by sequencing processing of the second sequencing library; wherein the first sequencing library comprises at least one first sequencing sample, and the second sequencing library comprises at least one second sequencing sample.
[0155] Specifically in the embodiments of the present disclosure, the first batch of samples to be sequenced (first library to be sequenced) is added to a sequencer for sequencing, and after a preset amount of sequencing data is reached, the second batch of samples to be sequenced (second library to be sequenced) is added to the same sequencing chip for sequencing. The fastq sequence obtained is automatically split according to each barcode and stored in a folder named by the barcode; for example, barcode 1 corresponds to a1.fastq and b1.fastq; a1.fastq is the sequencing result of plasmid 1 in the first batch, and b1.fastq is the sequencing result of plasmid 1 (may contain) and plasmid 2 in the second batch. In this way, the fastq data of plasmids 1-12 can be obtained.
[0156] For details of step 302, please refer to the above embodiments, and the embodiments of the present disclosure will not be repeated here.
[0157] In step 303, a preset assembly algorithm is called to perform sequence assembly on each test sequence in the first test sequence set to obtain a first sequence set corresponding to each first sample to be sequenced.
[0158] Specifically in the embodiments of the present disclosure, the reads in the first test sequence set corresponding to each first sample to be sequenced obtained in step 302 are subjected to sequence assembly to obtain corresponding contigs (first sequence set). The fastq sequence of each barcode is separately assembled to obtain the corresponding assembly result. For example, test sequences a1.fastq, a2.fastq, a3.fastq, a4.fastq, a5.fastq, and a6.fastq corresponding to plasmids 1-6 are assembled to obtain assembly results a1, a2, a3, a4, a5, and a6.
[0159] In step 304, a preset assembly algorithm is called to perform sequence assembly on each test sequence in the second test sequence set to obtain a second sequence set corresponding to each second sample to be sequenced.
[0160] Specifically in the embodiments of the present disclosure, the reads in the second test sequence set corresponding to each second sample to be sequenced obtained in step 302 are subjected to sequence assembly to obtain corresponding contigs (second sequence set). The fastq sequence of each barcode is separately assembled to obtain the corresponding assembly result. For example, test sequences b1.fastq, b2.fastq, b3.fastq, b4.fastq, b5.fastq, and b6.fastq corresponding to plasmids 7-12 are assembled to obtain assembly results b1, b2, b3, b4, b5, and b6.
[0161] Step 305, respectively, de-redundant processing is performed on each of the first sequence set and / or each of the second sequence set, to obtain each of the first assembled sequence set and each of the second assembled sequence set.
[0162] As another implementable manner of the embodiments of the present disclosure, the de-redundant processing respectively performed on each of the first sequence set and each of the second sequence set comprises: based on a first preset similarity threshold, comparing the assembled sequences in the first sequence set with each other, and removing the assembled sequences with short sequence length; based on the first preset similarity threshold, comparing the assembled sequences in the second sequence set with each other, and removing the assembled sequences with short sequence length.
[0163] As a refinement of the embodiments of the present disclosure, it is determined whether the number of sequences in the first sequence set and / or the second sequence set is greater than one; in the case that the number of sequences in the first sequence set and / or the second sequence set is greater than one, respectively, de-redundant processing is performed on each of the first sequence set and / or each of the second sequence set, to obtain each of the first assembled sequence set and each of the second assembled sequence set.
[0164] Specifically in the embodiments of the present disclosure, after respectively assembling the first and second test sequences corresponding to each of the first and second to-be-tested sequence samples, the first and second sequence sets are obtained, and the first sequence set and the second sequence set include a large number of contigs (assembled sequences) with similar sequences; it is possible that the sequence of one contig is all on the sequence of another contig, which will make the data volume large and not conducive to the later evaluation processing, and thus it is necessary to perform de-redundant processing on the assembled contigs according to a first preset similarity threshold. When performing sequence assembly, canu software can be used but is not limited thereto. It should be noted that the embodiments of the present disclosure do not limit when de-redundant processing is performed on each of the first sequence set and each of the second sequence set.
[0165] Specifically, if the head-to-middle part of one contig is aligned with the middle-to-tail part of another contig, and the similarity of the aligned region is greater than 98%, the contig with short length is removed, and finally the assembly result is obtained. As shown in FIG. 6, this is the case. If one contig is aligned on another contig from head to tail, and the similarity is greater than 98%, the contig with short length is removed, and finally the assembly result is obtained, as shown in FIG. 7, which is the case. Figure 4 Figure 5 If three contigs are included in the first sequence set or the second sequence set, after alignment, there is no Figure 4 Figure 5 In the case shown, and when the similarity is less than 98%, three contigs are taken as the results after de-redundancy. The number of contigs after assembly and de-redundancy of 12 samples is shown in Table 2.
[0166] Table 2 Number of contigs after assembly and de-redundancy of 12 samples
[0167]
[0168] It should be noted that the first preset similarity threshold is 98% in the embodiments of the present disclosure, which does not constitute a limitation of the present disclosure. The embodiments of the present disclosure can extract accurate assembly sequences and reduce the amount of data for subsequent processing by assembling and de-redundancy processing of the test sequences obtained by sequencing.
[0169] In step 306, each of the first assembly sequence sets is evaluated to obtain a first assembly sequence corresponding to each of the first samples to be sequenced, and each of the second assembly sequence sets is evaluated to obtain a second assembly sequence and a third assembly sequence corresponding to each of the second samples to be sequenced.
[0170] For the description of step 306, please refer to the above embodiments, and the embodiments of the present disclosure will not be repeated here.
[0171] In step 307, based on the similarity between at least one of the first assembly sequences and at least one of the second assembly sequences and at least one of the third assembly sequences, the assembly sequence corresponding to each of the at least one first sample to be sequenced and the assembly sequence corresponding to each of the at least one second sample to be sequenced are determined.
[0172] For the description of step 307, please refer to the above embodiments, and the embodiments of the present disclosure will not be repeated here.
[0173] In the embodiments of the present disclosure, the first assembly sequence is determined as the assembly sequence corresponding to the first sample to be sequenced, the similarity between the second assembly sequence and the first assembly sequence and the similarity between the third assembly sequence and the first assembly sequence are calculated respectively, and the assembly sequence corresponding to the second sample to be sequenced is determined by comparing the similarities.
[0174] In order to clearly illustrate the embodiments of the present disclosure, the present disclosure provides a flowchart of another method of mixed sample sequencing.
[0175] As Figure 6 shown, the method comprises the following steps:
[0176] In step 401, any of the first samples to be sequenced in the first library to be sequenced and any of the second samples to be sequenced in the second library to be sequenced are constructed using the same identification sequence.
[0177] Step 402, after obtaining at least one first test sequence set by sequencing processing on a first to-be-sequenced library in the same sequencing chip, obtaining at least one second test sequence set by sequencing processing on a second to-be-sequenced library; wherein the first to-be-sequenced library comprises at least one first to-be-sequenced sample, and the second to-be-sequenced library comprises at least one second to-be-sequenced sample.
[0178] Step 403, performing sequence assembly on the test sequences in each of the first test sequence set to obtain a first assembled sequence set corresponding to each of the first to-be-sequenced sample; and performing sequence assembly on the test sequences in each of the second test sequence set to obtain a second assembled sequence set corresponding to each of the second to-be-sequenced sample.
[0179] For the description of steps 401 to 403, please refer to the above-mentioned embodiments, and the embodiments of the present disclosure will not be described one by one.
[0180] Step 404, in the case where the first assembled sequence set comprises one assembled sequence, taking the assembled sequence as the first assembled sequence corresponding to the first assembled sequence set.
[0181] Specifically in the embodiments of the present disclosure, the assembled sequence in the first assembled sequence set that is closest to the sequence of the first to-be-sequenced sample is obtained by evaluating the assembled sequence in each first assembled sequence set. For example, referring to Table 2, after removing redundancy, plasmids 1, 2, 3, 5, and 6 each have only one contig, so there is no need to perform abundance evaluation. Taking the contigs in plasmids 1, 2, 3, 5, and 6 as the assembled sequences of assembled sequence sets a1, a2, a3, a5, and a6 respectively, c1, c2, c3, c5, and c6 are obtained.
[0182] Step 405, in the case where the first assembled sequence set comprises at least two assembled sequences, performing abundance evaluation on all the assembled sequences in the first assembled sequence set, and determining the assembled sequence with the highest abundance evaluation ranking as the first assembled sequence corresponding to the first assembled sequence set.
[0183] Specifically in the embodiments of the present disclosure, when the number of assembled sequences in the first assembled sequence set is greater than or equal to two, all the assembled sequences are subjected to alignment evaluation to obtain an abundance ranking. For example, referring to Table 2, the assembled sequence set corresponding to plasmid 4 contains two contigs, and the contig with the highest abundance is selected as the assembled result c4 of a4.
[0184] Step 406, in the case where the second assembled sequence set comprises one assembled sequence, aligning the second assembled sequence set with the same identifier sequence with the first assembled sequence in the first assembled sequence set.
[0185] As a refinement of an embodiment of the present disclosure, the aligning the second set of assembled sequences with the same marker sequence as the assembled sequences in the first set of assembled sequences comprises: calculating a third similarity between the assembled sequences in the second set of assembled sequences and the assembled sequences with the same marker sequence; in a case where it is determined that the third similarity is greater than or equal to the second preset similarity threshold, resequencing the second sample to be sequenced corresponding to the marker sequence; in a case where it is determined that the third similarity is less than the second preset similarity threshold, taking the assembled sequences in the second set of assembled sequences as the assembled sequences corresponding to the second sample to be sequenced.
[0186] Specifically in the embodiment of the present disclosure, since the fastq obtained by sequencing the second sample to be sequenced in the second batch contains the test sequence data of the first sample to be sequenced with the same barcode, it is necessary to evaluate the assembled sequences in the second set of assembled sequences to obtain the second assembled sequence and the third assembled sequence; the second assembled sequence can be the assembled sequence of the first sample to be sequenced, or can be the assembled sequence of the second sample to be sequenced.
[0187] For example, if b1 (the second set of assembled sequences) includes a contig (assembled sequence), it is necessary to confirm whether the contig is the contig corresponding to plasmid 7; by aligning with c1, if it is determined that the contig is the contig of plasmid 1, plasmid 7 is resequenced; if it is determined that the contig is not the contig of plasmid 1, the contig in b1 is taken as the sequencing result of plasmid 7.
[0188] Step 407, in a case where it is determined that the second set of assembled sequences includes two assembled sequences, taking the two assembled sequences as the second assembled sequence and the third assembled sequence corresponding to the second set of assembled sequences, respectively.
[0189] Specifically in the embodiment of the present disclosure, please refer to Table 2, only two contigs in b1, b2, b3, b5, b6, the two contigs are taken as the second assembled sequence and the third assembled sequence, and are named as d1, d2, d3, d4, d5, d6, d9, d10, d11, d12, respectively; wherein d1 and d2 are the second assembled sequence and the third assembled sequence corresponding to b1, and the second assembled sequence and the third assembled sequence corresponding to b2, b3, b5, b6 are obtained in the same manner.
[0190] Step 408, in the case of determining that the second assembly sequence set includes at least three assembly sequences, performing abundance evaluation on the assembly sequences in the second assembly sequence set, and taking the assembly sequences with the first and second abundance evaluation rankings as the second assembly sequence and the third assembly sequence corresponding to the second assembly sequence set, respectively.
[0191] Specifically in the embodiments of the present disclosure, referring to Table 2, b4 has three contigs, which can be evaluated by, but not limited to, the alignment of the blasr software to obtain the abundance rankings. Finally, the two contigs with the highest abundance are selected as the second assembly sequence (d7) and the third assembly sequence (d8) corresponding to b4.
[0192] Step 409, determining the assembly sequence corresponding to each of the at least one first test sequence sample and the assembly sequence corresponding to each of the at least one second test sequence sample based on the similarity between at least one of the first assembly sequences and at least one of the second assembly sequences and at least one of the third assembly sequences.
[0193] For the description of step 409, please refer to the above embodiments, and the embodiments of the present disclosure will not be described one by one.
[0194] In order to clearly illustrate the embodiments of the present disclosure, the embodiments of the present disclosure provide a flowchart of another method for mixed sample sequencing.
[0195] As shown in the method, the method comprises the following steps: Figure 7
[0196] Step 501, using the same identification sequence to construct a library for any of the first test sequence samples in the first test sequence library and any of the second test sequence samples in the second test sequence library.
[0197] Step 502, after obtaining at least one first test sequence set by sequencing the first test sequence library in the same sequencing chip, obtaining at least one second test sequence set by sequencing the second test sequence library; wherein the first test sequence library includes at least one first test sequence sample, and the second test sequence library includes at least one second test sequence sample.
[0198] Step 503, performing sequence assembly on the test sequences in each of the first test sequence sets to obtain a first assembly sequence set corresponding to each of the first test sequence samples; and performing sequence assembly on the test sequences in each of the second test sequence sets to obtain a second assembly sequence set corresponding to each of the second test sequence samples.
[0199] Step 504, evaluating each of the first assembly sequence set to obtain the first assembly sequence corresponding to each of the first sequencing sample; evaluating each of the second assembly sequence set to obtain the second assembly sequence and the third assembly sequence corresponding to each of the second sequencing sample.
[0200] For the description of steps 501-504, please refer to the above embodiments, and the embodiments of the present disclosure will not be repeated here.
[0201] Step 505, determining the first assembly sequence as the assembly sequence corresponding to each of the first sequencing sample.
[0202] Specifically in the embodiments of the present disclosure, please refer to steps 404-405, the evaluation results c1, c2, c3, c4, c5, c6 of the first batch of plasmids 1, 2, 3, 4, 5 and 6.
[0203] Step 506, respectively performing similarity comparison between the second assembly sequence with the same identification sequence and the first assembly sequence to obtain the first similarity corresponding to each of the identification sequence.
[0204] Specifically in the embodiments of the present disclosure, respectively performing similarity comparison between d1 and c1, d3 and c2, d5 and c3, d7 and c4, d9 and c5, d11 and c6 to obtain the first similarity D1, D3, D5, D7, D9, D11.
[0205] Step 507, respectively performing similarity comparison between the third assembly sequence with the same identification sequence and the first assembly sequence to obtain the second similarity corresponding to each of the identification sequence.
[0206] Specifically in the embodiments of the present disclosure, respectively performing similarity comparison between d2 and c1, d4 and c2, d6 and c3, d8 and c4, d10 and c5, d12 and c6 to obtain the second similarity D2, D4, D6, D8, D10, D12.
[0207] Step 508, determining the assembly sequence corresponding to each of the second sequencing sample based on the first similarity and the second similarity; wherein the first assembly sequence corresponding to the first similarity is the same as the first assembly sequence corresponding to the second similarity.
[0208] As a refinement of an embodiment of the present disclosure, the determining of the respective assembly sequence corresponding to each of the second sequencing samples based on the first similarity and the second similarity comprises: in a case where it is determined that the first similarity is greater than or equal to a second preset similarity threshold and the second similarity is less than the second preset similarity threshold, or there is no result, determining the third assembly sequence as the assembly sequence corresponding to the second sequencing sample; in a case where it is determined that the second similarity is greater than or equal to the second preset similarity threshold and the first similarity is less than the second preset similarity threshold, or there is no result, determining the second assembly sequence as the assembly sequence corresponding to the second sequencing sample; and in a case where it is determined that the first similarity and the second similarity are both greater than or both less than the second preset similarity threshold, resequencing the second sequencing sample.
[0209] In particular, in the embodiment of the present disclosure, the assembly sequence corresponding to d1 or d2 is determined to be the assembly sequence corresponding to plasmid 7 (the second sequencing sample) by comparing the sizes of D1 and D2, the second preset similarity threshold. For example, if D1 is greater than 97% and D2 is less than 97% or there is no result, it is determined that the assembly sequence corresponding to d2 is the assembly sequence corresponding to plasmid 7 (the second sequencing sample). If the similarity D1 between d1 and c1 is greater than the preset threshold, it can be considered that d1 is the sample of plasmid 1 remaining in the second batch of sequencing. If D2 is greater than 97% and D1 is less than 97% or there is no result, it is determined that the assembly sequence corresponding to d1 is the assembly sequence corresponding to plasmid 7 (the second sequencing sample). If D1 and D2 are both greater than or both less than, it cannot be determined that the assembly sequence corresponding to d1 or d2 is the assembly sequence corresponding to plasmid 7 (the second sequencing sample). Therefore, it is necessary to resequence plasmid 7.
[0210] The sanger sequences obtained by sequencing plasmids 1-12 using the sanger method are used to verify the sequencing results of plasmids 1-12 obtained by using Nanopore sequencing in the above embodiment, and the alignment statistics table is shown in Tables 3 and 4.
[0211] Table 3 Alignment statistics table of sanger sequence and Nanopore sequencing result
[0212] Plasmid 1 Plasmid 2 Plasmid 3 Plasmid 4 Plasmid 5 Plasmid 6 Similarity 99.9656 100 100 99.8249 100 100
[0213] Table 4 Alignment statistics table of sanger sequence and Nanopore sequencing result
[0214] Plasmid 7 Plasmid 8 Plasmid 9 Plasmid 10 Plasmid 11 Plasmid 12 Similarity 99.8379 99.7964 99.9418 99.7229 99.8321 99.8249
[0215] The similarity of the alignment of the sanger sequence and the Nanopore sequencing result is more than 98%; therefore, it is proved that the mixed sample sequencing method of the present disclosure is accurate and reliable.
[0216] To clearly illustrate the embodiments of this disclosure, a flowchart of another method for pooled sequencing is provided.
[0217] like Figure 8 As shown, the method includes the following steps:
[0218] Step 601: Construct a library by combining any of the first sequencing samples from the first sequencing library with any of the second sequencing samples from the second sequencing library.
[0219] Specifically, in this embodiment, the first and second sequencing libraries are described as single-sample libraries. After constructing the libraries using plasmid 1 and plasmid 2 according to Nanopore's standard library construction method, the plasmid 1 library (the first sequencing library) is added to Nanopore for sequencing. The plasmid 2 library (the second sequencing library) is then added to Nanopore and mixed with the remaining plasmid 1 library to form a mixed library.
[0220] Step 602: In the same sequencing chip, after sequencing the first library to be sequenced to obtain at least one first test sequence set, sequencing the second library to be sequenced to obtain at least one second test sequence set; wherein, the first library to be sequenced includes at least one first sample to be sequenced, and the second library to be sequenced includes at least one second sample to be sequenced.
[0221] Specifically, in this embodiment, plasmid 1 library is added to Nanopore for sequencing. Sequencing is paused after the sequencing data volume reaches 10MB. Sequencing data "A.fastq" is obtained (sequencing data "A.fastq" is the first test sequence set). Plasmid 2 library is added to Nanopore to mix with the remaining plasmid 1 library, forming a mixed library. Sequencing is paused after the mixed sample sequencing data volume reaches 20MB. Sequencing data "B.fastq" is obtained (sequencing data "B.fastq" is the second test sequence set).
[0222] Step 603: Assemble the test sequences in each of the first test sequence sets to obtain a first assembled sequence set corresponding to each of the first sequencing samples; assemble the test sequences in each of the second test sequence sets to obtain a second assembled sequence set corresponding to each of the second sequencing samples.
[0223] Specifically in the embodiments of the present disclosure, canu assembly software can be used, but is not limited to, to assemble "A.fastq" and "B.fastq" respectively, and assembly results "A.contigs.fasta" and "B.contigs.fasta" (the assembly results "A.contigs.fasta" and "B.contigs.fasta" are respectively the first assembly sequence set and the second assembly sequence set) are obtained. Viewing "A.contigs.fasta" determines that there is only one contig, and no redundancy and abundance evaluation is needed. For the convenience of subsequent description, the contig ID command is "Contig_A". Viewing "B.contigs.fasta" has 3 contigs, and redundancy is removed. The 3 contigs are compared with each other by blasr, and there is no Figure 4 and Figure 5 case, the similarity is also less than 98%, and the original assembly result is retained.
[0224] Step 604, evaluating each of the first assembly sequence set to obtain the first assembly sequence corresponding to each of the first sample to be sequenced respectively; evaluating each of the second assembly sequence set to obtain the second assembly sequence and the third assembly sequence corresponding to each of the second sample to be sequenced respectively.
[0225] Specifically in the embodiments of the present disclosure, the abundance of "B.contigs.fasta" is evaluated. The "B.fastq" can be compared with "B.contigs.fasta" by using blasr software, and the comparison command is "blasr B.fastq B.contigs.fasta -out B.m4 -m 4 -bestn 1". The second column of B.m4 is the contig ID of each compared contig. The total number N of each contig in the second column is counted, and N is divided by the length of the corresponding contig to obtain the abundance ranking. Finally, the two contigs "tig00000001.fasta" and "tig00000002.fasta" with the highest abundance (corresponding to the second assembly sequence and the third assembly sequence respectively) are selected.
[0226] Step 605, determining the assembly sequence corresponding to each of the at least one first sample to be sequenced and the assembly sequence corresponding to each of the at least one second sample to be sequenced based on the similarity between at least one of the first assembly sequence and at least one of the second assembly sequence and at least one of the third assembly sequence.
[0227] Specifically in the embodiments of the present disclosure, the two contigs with the highest abundance, "tig00000001.fasta" and "tig00000002.fasta", are aligned to "A.contigs.fasta" by blasr to obtain alignment results tig1.m4 and tig2.m4. The similarity statistics are shown in Table 5. The similarity of "tig00000002.fasta" is less than 98%, and it is determined as the sequence of plasmid 2.
[0228] Table 5 Similarity statistics table
[0229] Plasmid 1 Contig ID Plasmid 2 Contig ID Similarity Contig_A tig00000001 99.0641% Contig_A tig00000002 85.2081%
[0230] The sanger sequence "Sanger1.fa" of plasmid 1 is aligned with the assembly result "A.contigs.fasta", and the similarity between the two is 99.8968%; the sanger sequence "Sanger2.fa" of plasmid 2 is aligned with the assembly result "tig00000002.fasta", and the similarity between the two is 99.9344%. Through plasmid sanger sequence verification, the similarity of both is more than 98%, proving that the method of the present application is reliable.
[0231] It should be noted that the embodiments of the present disclosure can include a plurality of steps, which are numbered for the convenience of description, but these numbers are not a limitation on the execution time slot, execution order between steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not limit this.
[0232] Corresponding to the above-mentioned mixed sample sequencing method, the present application also proposes a mixed sample sequencing device. Since the device embodiments of the present application correspond to the above-mentioned method embodiments, for the details not disclosed in the device embodiments, please refer to the above-mentioned method embodiments, which will not be described in detail in the present application.
[0233] Figure 9 A structure diagram of a mixed sample sequencing device provided by the embodiments of the present disclosure is shown in Figure 9 as shown, comprising:
[0234] The first sequencing unit 71 is used to perform sequencing processing on a second library to be sequenced to obtain at least one second test sequence set after performing sequencing processing on a first library to be sequenced to obtain at least one first test sequence set in the same sequencing chip; wherein the first library to be sequenced includes at least one first sample to be sequenced, and the second library to be sequenced includes at least one second sample to be sequenced;
[0235] The first assembly unit 72 is configured to perform sequence assembly on the test sequences in each of the first test sequence set to obtain a first assembly sequence set corresponding to each of the first to-be-sequenced samples; and perform sequence assembly on the test sequences in each of the second test sequence set to obtain a second assembly sequence set corresponding to each of the second to-be-sequenced samples.
[0236] The first evaluation unit 73 is configured to evaluate each of the first assembly sequence set to obtain a first assembly sequence corresponding to each of the first to-be-sequenced samples; and evaluate each of the second assembly sequence set to obtain a second assembly sequence and a third assembly sequence corresponding to each of the second to-be-sequenced samples.
[0237] The first determination unit 74 is configured to determine, based on a similarity between at least one of the first assembly sequence and at least one of the second assembly sequence and at least one of the third assembly sequence, an assembly sequence corresponding to each of the at least one first to-be-sequenced sample and an assembly sequence corresponding to each of the at least one second to-be-sequenced sample.
[0238] The present disclosure provides a device for mixed sample sequencing. After at least one first test sequence set is obtained by performing sequencing on a first to-be-sequenced library, at least one second test sequence set is obtained by performing sequencing on a second to-be-sequenced library in the same sequencing chip. The first to-be-sequenced library includes at least one first to-be-sequenced sample, and the second to-be-sequenced library includes at least one second to-be-sequenced sample. Sequence assembly is performed on the test sequences in each of the first test sequence set to obtain a first assembly sequence set corresponding to each of the first to-be-sequenced samples. Sequence assembly is performed on the test sequences in each of the second test sequence set to obtain a second assembly sequence set corresponding to each of the second to-be-sequenced samples. Evaluation is performed on each of the first assembly sequence set to obtain a first assembly sequence corresponding to each of the first to-be-sequenced samples. Evaluation is performed on each of the second assembly sequence set to obtain a second assembly sequence and a third assembly sequence corresponding to each of the second to-be-sequenced samples. Based on a similarity between at least one of the first assembly sequence and at least one of the second assembly sequence and at least one of the third assembly sequence, an assembly sequence corresponding to each of the at least one first to-be-sequenced sample and an assembly sequence corresponding to each of the at least one second to-be-sequenced sample are determined. Compared with the related art, the present disclosure performs sequence assembly and evaluation on the sequences obtained by performing sequencing on the first to-be-sequenced library and the second to-be-sequenced library in the same sequencing chip to obtain an assembly sequence corresponding to the first to-be-sequenced library and an assembly sequence corresponding to the second to-be-sequenced library. The present disclosure can reduce the time spent on cleaning the sequencing chip when two adjacent batches of sequencing are performed, and improve the efficiency of sequencing.
[0239] Further, in a possible implementation manner of the present embodiment, as shown in FIG. 2, the device for mixed sample sequencing includes a first sequencing unit 71, a first assembly unit 72, a first evaluation unit 73, a first determination unit 74, a second sequencing unit 75, a second assembly unit 76, a second evaluation unit 77, and a second determination unit 78.Figure 10 As shown in the figure, the device further comprises:
[0240] a second sequencing unit 75, configured to, after obtaining at least one second test sequence set by performing sequencing processing on a second to-be-sequenced library, continue to perform sequencing processing on a third to-be-sequenced library in the same sequencing chip to obtain at least one third test sequence set, wherein the third to-be-sequenced library comprises at least one third to-be-sequenced sample;
[0241] a second assembly unit 76, configured to perform sequence assembly on test sequences in each of the third test sequence set to obtain a third assembly sequence set corresponding to each of the third to-be-sequenced sample;
[0242] a second evaluation unit 77, configured to evaluate each of the third assembly sequence set to obtain a fourth assembly sequence, a fifth assembly sequence and a sixth assembly sequence corresponding to each of the third to-be-sequenced sample respectively;
[0243] a second determination unit 78, configured to determine an assembly sequence corresponding to each of the third to-be-sequenced sample based on similarities between the assembly sequence corresponding to each of the at least one first to-be-sequenced sample, the assembly sequence corresponding to each of the at least one second to-be-sequenced sample and the fourth assembly sequence, the fifth assembly sequence and the sixth assembly sequence respectively.
[0244] Further, in a possible implementation manner of the embodiment, as shown in the figure, Figure 10 the first assembly unit 72 comprises:
[0245] a first assembly module 721, configured to perform sequence assembly on test sequences in each of the first test sequence set respectively to obtain a first assembly sequence set corresponding to each of the first to-be-sequenced sample;
[0246] a second assembly module 722, configured to perform sequence assembly on test sequences in each of the second test sequence set respectively to obtain a second assembly sequence set corresponding to each of the second to-be-sequenced sample.
[0247] Further, in a possible implementation manner of the embodiment, as shown in the figure, Figure 10 the first assembly unit 72 further comprises:
[0248] a third assembly module 723, configured to call a preset assembly algorithm to perform sequence assembly on test sequences in each of the first test sequence set respectively to obtain a first sequence set corresponding to each of the first to-be-sequenced sample;
[0249] The fourth assembling module 724 is configured to call the preset assembling algorithm to perform sequence assembly on each test sequence in the second test sequence set respectively to obtain a second sequence set corresponding to each second sequencing sample.
[0250] Further, in a possible implementation of the embodiment, as shown in Figure 10 The device further includes:
[0251] The third determining unit 79 is configured to determine whether the number of sequences in the first sequence set and / or the second sequence set is greater than one after calling the preset assembling algorithm to perform sequence assembly on each test sequence in the second test sequence set respectively to obtain a second sequence set corresponding to each second sequencing sample.
[0252] The processing unit 710 is configured to perform de-redundancy processing on each first sequence set and / or each second sequence set respectively to obtain each first assembled sequence set and each second assembled sequence set in the case where the number of sequences in the first sequence set and / or the second sequence set is greater than one.
[0253] Further, in a possible implementation of the embodiment, as shown in Figure 10 The processing unit 710 includes:
[0254] The first comparison module 7101 is configured to compare the assembled sequences in the first sequence set with each other based on a first preset similarity threshold, and remove the assembled sequences with a short sequence length.
[0255] The second comparison module 7102 is configured to compare the assembled sequences in the second sequence set with each other based on the first preset similarity threshold, and remove the assembled sequences with a short sequence length.
[0256] Further, in a possible implementation of the embodiment, as shown in Figure 10 The first evaluation unit 73 includes:
[0257] The first determining module 731 is configured to, in the case where the first assembled sequence set includes one assembled sequence, determine the assembled sequence as the first assembled sequence corresponding to the first assembled sequence set.
[0258] The second determining module 732 is configured to, in the case where the first assembled sequence set includes at least two assembled sequences, perform abundance evaluation on all the assembled sequences in the first assembled sequence set, and determine the assembled sequence with the first ranking in the abundance evaluation as the first assembled sequence corresponding to the first assembled sequence set.
[0259] Further, in a possible implementation of the embodiment, as shown in Figure 10 the first evaluation unit 73 further includes:
[0260] a first comparison module 733, configured to, in a case where the second assembled sequence set includes one assembled sequence, compare the second assembled sequence set with the first assembled sequence in the first assembled sequence set, the second assembled sequence set and the first assembled sequence being identical in sequence;
[0261] a third determination module 734, configured to, in a case where the second assembled sequence set includes two assembled sequences, take the two assembled sequences as the second assembled sequence and the third assembled sequence corresponding to the second assembled sequence set respectively;
[0262] a fourth determination module 735, configured to, in a case where the second assembled sequence set includes at least three assembled sequences, perform abundance evaluation on the assembled sequences in the second assembled sequence set, and take the assembled sequences ranked first and second in abundance evaluation as the second assembled sequence and the third assembled sequence corresponding to the second assembled sequence set respectively.
[0263] Further, in a possible implementation of the embodiment, the first comparison module 733 is further configured to:
[0264] calculate a third similarity between the assembled sequences in the second assembled sequence set and the first assembled sequence identical in sequence to the marker sequence;
[0265] in a case where the third similarity is greater than or equal to the second preset similarity threshold, re-sequencing the second sample to be sequenced corresponding to the marker sequence;
[0266] in a case where the third similarity is less than the second preset similarity threshold, taking the assembled sequences in the second assembled sequence set as the assembled sequences corresponding to the second library to be sequenced.
[0267] Further, in a possible implementation of the embodiment, as shown in Figure 10 the first determination unit 74 includes:
[0268] a fifth determination module 741, configured to determine the first assembled sequence as the assembled sequence corresponding to each of the first samples to be sequenced respectively;
[0269] a second comparison module 742, configured to perform similarity comparison on the second assembled sequence identical in sequence and the first assembled sequence respectively, to obtain a first similarity corresponding to each of the marker sequences;
[0270] The third comparison module 743 is configured to perform similarity comparison between the third assembly sequence and the first assembly sequence respectively, to obtain a second similarity corresponding to each of the identification sequences.
[0271] The sixth determination module 744 is configured to determine an assembly sequence corresponding to each of the second to-be-sequenced samples based on the first similarity and the second similarity, wherein the first assembly sequence corresponding to the first similarity is the same as the first assembly sequence corresponding to the second similarity.
[0272] Further, in a possible implementation of the embodiment, the sixth determination module 744 is further configured to:
[0273] In a case where it is determined that the first similarity is greater than or equal to a second preset similarity threshold and the second similarity is less than the second preset similarity threshold, or there is no result, the third assembly sequence is determined as the assembly sequence corresponding to the second to-be-sequenced sample.
[0274] In a case where it is determined that the second similarity is greater than or equal to the second preset similarity threshold and the first similarity is less than the second preset similarity threshold, or there is no result, the second assembly sequence is determined as the assembly sequence corresponding to the second to-be-sequenced sample.
[0275] In a case where it is determined that the first similarity and the second similarity are both greater than or both less than the second preset similarity threshold, the second to-be-sequenced sample is re-sequenced.
[0276] It should be noted that the foregoing explanation of the method embodiment is also applicable to the apparatus of the embodiment, and the principle is the same, which is not limited in the embodiment.
[0277] According to embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0278] Figure 11 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.
[0279] As Figure 11As shown, the device 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a ROM (Read-Only Memory) 802 or a computer program loaded into a RAM (Random Access Memory) 803 from the storage unit 808. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An I / O (Input / Output) interface 805 is also connected to the bus 804.
[0280] A plurality of components in the device 800 are connected to the I / O interface 805, including an input unit 806 such as a keyboard, a mouse, and the like, an output unit 807 such as various types of displays, a speaker, and the like, a storage unit 808 such as a magnetic disk, an optical disk, and the like, and a communication unit 809 such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0281] The computing unit 801 can be various general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, and the like. The computing unit 801 performs various methods and processes described above, such as the methods of pooled sequencing. For example, in some embodiments, the methods of pooled sequencing can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the aforementioned methods of pooled sequencing by any other appropriate means, such as by means of firmware.
[0282] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a Field Programmable Gate Array (FPGA), an Application-Specific Integrated Circuit (ASIC), an Application Specific Standard Product (ASSP), a System on a Chip (SOC), a Complex Programmable Logic Device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0283] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general or special purpose computer, such that the program code, when executed by the processor or controller, causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be implemented in a wholly in machine language, in partially in machine language, in partially in a high level language, and other combinations thereof. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server.
[0284] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a linearly-programmed electronic storage, a portable computer diskette, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory), or flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0285] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0286] The systems and techniques described here can be implemented in a computing system that includes a back-end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front-end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.
[0287] The computer system can include clients and servers. This relationship can be between a client and a server that are typically remote from each other and typically interact through a communication network. The relationship between client and server exists by virtue of computer programs running on the respective computer systems and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS (Virtual Private Server, or VPS for short) services. The server can also be a server of a distributed system, or a server combined with a blockchain.
[0288] It should be noted that artificial intelligence is a discipline that studies enabling computers to simulate some thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), both hardware and software technologies. Artificial intelligence hardware technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing, etc.; artificial intelligence software technology mainly includes computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, knowledge graph technology, etc. several major directions.
[0289] It should be understood that the various forms of the flow shown above can be used to reorder, add or delete steps. For example, each step described in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, which is not limited herein.
[0290] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method of sequencing by hybridization, characterized in that, The method comprises the following steps: sequencing a second library to obtain at least one second test sequence set after the first library is sequenced to obtain at least one first test sequence set; wherein the first library comprises at least one first sample to be sequenced, and the second library comprises at least one second sample to be sequenced; performing sequence assembly on each test sequence in the first test sequence set to obtain a first assembly sequence set corresponding to each first sample to be sequenced; and performing sequence assembly on each test sequence in the second test sequence set to obtain a second assembly sequence set corresponding to each second sample to be sequenced; evaluating each first assembly sequence set to obtain a first assembly sequence corresponding to each first sample to be sequenced; and evaluating each second assembly sequence set to obtain a second assembly sequence and a third assembly sequence corresponding to each second sample to be sequenced; determining an assembly sequence corresponding to each first sample to be sequenced and an assembly sequence corresponding to each second sample to be sequenced based on the similarity between at least one first assembly sequence and at least one second assembly sequence and at least one third assembly sequence.
2. The method of claim 1, wherein, After the second library is sequenced to obtain at least one second test sequence set, the method further comprises: sequencing a third library in the same sequencing chip to obtain at least one third test sequence set; wherein the third library comprises at least one third sample to be sequenced; performing sequence assembly on each test sequence in the third test sequence set to obtain a third assembly sequence set corresponding to each third sample to be sequenced; evaluating each third assembly sequence set to obtain a fourth assembly sequence, a fifth assembly sequence, and a sixth assembly sequence corresponding to each third sample to be sequenced; determining an assembly sequence corresponding to each third sample to be sequenced based on the similarity between the assembly sequence corresponding to each first sample to be sequenced, the assembly sequence corresponding to each second sample to be sequenced, and the fourth assembly sequence, the fifth assembly sequence, and the sixth assembly sequence.
3. The method of claim 1, wherein, The method of performing sequence assembly on each test sequence in the first test sequence set to obtain a first assembly sequence set corresponding to each first sample to be sequenced and performing sequence assembly on each test sequence in the second test sequence set to obtain a second assembly sequence set corresponding to each second sample to be sequenced comprises: performing sequence assembly on each test sequence in the first test sequence set to obtain a first assembly sequence set corresponding to each first sample to be sequenced; performing sequence assembly on each test sequence in the second test sequence set to obtain a second assembly sequence set corresponding to each second sample to be sequenced.
4. The method of claim 1, wherein, The sequence assembling is performed on each test sequence in the first test sequence set to obtain a first assembled sequence set corresponding to each of the first to-be-sequenced samples, and the sequence assembling is performed on each test sequence in the second test sequence set to obtain a second assembled sequence set corresponding to each of the second to-be-sequenced samples, and the method further comprises: calling a preset assembly algorithm to perform sequence assembling on each test sequence in the first test sequence set to obtain a first sequence set corresponding to each of the first to-be-sequenced samples; calling the preset assembly algorithm to perform sequence assembling on each test sequence in the second test sequence set to obtain a second sequence set corresponding to each of the second to-be-sequenced samples.
5. The method of claim 4, wherein, After calling the preset assembly algorithm to perform sequence assembling on each test sequence in the second test sequence set to obtain a second sequence set corresponding to each of the second to-be-sequenced samples, the method further comprises: determining whether the number of sequences in the first sequence set and / or the second sequence set is greater than one; in the case that the number of sequences in the first sequence set and / or the second sequence set is greater than one, performing de-redundancy processing on each of the first sequence set and / or the second sequence set to obtain each of the first assembled sequence set and the second assembled sequence set.
6. The method of claim 5, wherein, The de-redundancy processing on each of the first sequence set and / or the second sequence set comprises: comparing the assembled sequences in the first sequence set with each other based on a first preset similarity threshold to remove the assembled sequences with a short sequence length; comparing the assembled sequences in the second sequence set with each other based on the first preset similarity threshold to remove the assembled sequences with a short sequence length.
7. The method of claim 1, wherein, The evaluation on each of the first assembled sequence set to obtain a first assembled sequence corresponding to each of the first to-be-sequenced samples comprises: in the case that the first assembled sequence set comprises one assembled sequence, taking the assembled sequence as the first assembled sequence corresponding to the first assembled sequence set; in the case that the first assembled sequence set comprises at least two assembled sequences, performing abundance evaluation on all the assembled sequences in the first assembled sequence set, and determining the assembled sequence with the first ranking in the abundance evaluation as the first assembled sequence corresponding to the first assembled sequence set.
8. The method of claim 7, wherein, The evaluation on each of the second assembled sequence set to obtain a second assembled sequence and a third assembled sequence corresponding to each of the second to-be-sequenced samples comprises: in the case that the second assembled sequence set comprises one assembled sequence, comparing the second assembled sequence set with the same identification sequence with the first assembled sequence in the first assembled sequence set; in the case that the second assembled sequence set comprises two assembled sequences, taking the two assembled sequences as the second assembled sequence and the third assembled sequence corresponding to the second assembled sequence set, respectively; and in the case that the second assembled sequence set comprises two assembled sequences, taking the two assembled sequences as the second assembled sequence and the third assembled sequence corresponding to the second assembled sequence set, respectively. In a case where it is determined that the second assembly sequence set comprises at least three assembly sequences, the assembly sequences in the second assembly sequence set are subjected to abundance evaluation, and the assembly sequences ranked first and second in the abundance evaluation are respectively taken as the second assembly sequence and the third assembly sequence corresponding to the second assembly sequence set.
9. The method of claim 8, wherein, The aligning of the second assembly sequence set with the first assembly sequence in the first assembly sequence set comprises: calculating a third similarity between the assembly sequences in the second assembly sequence set and the first assembly sequence with the same marker sequence; In a case where it is determined that the third similarity is greater than or equal to the second preset similarity threshold, the second sample to be sequenced corresponding to the marker sequence is subjected to re-sequencing; In a case where it is determined that the third similarity is less than the second preset similarity threshold, the assembly sequences in the second assembly sequence set are taken as the assembly sequences corresponding to the second library to be sequenced.
10. The method of claim 8, wherein, The determining of the assembly sequence corresponding to each of the first sample to be sequenced, the assembly sequence corresponding to each of the second sample to be sequenced based on the similarity between at least one first assembly sequence and at least one second assembly sequence, at least one third assembly sequence respectively comprises: determining the first assembly sequence as the assembly sequence corresponding to each of the first sample to be sequenced; aligning the second assembly sequence with the same identifier sequence and the first assembly sequence respectively to obtain a first similarity corresponding to each of the identifier sequences; aligning the third assembly sequence with the same identifier sequence and the first assembly sequence respectively to obtain a second similarity corresponding to each of the identifier sequences; determining the assembly sequence corresponding to each of the second sample to be sequenced based on the first similarity and the second similarity; wherein the first assembly sequence corresponding to the first similarity is the same as the first assembly sequence corresponding to the second similarity.
11. The method of claim 10, wherein, The determining of the assembly sequence corresponding to each of the second sample to be sequenced based on the first similarity and the second similarity comprises: in a case where it is determined that the first similarity is greater than or equal to a second preset similarity threshold and the second similarity is less than the second preset similarity threshold, or there is no result, determining that the third assembly sequence is the assembly sequence corresponding to the second sample to be sequenced; in a case where it is determined that the second similarity is greater than or equal to the second preset similarity threshold and the first similarity is less than the second preset similarity threshold, or there is no result, determining that the second assembly sequence is the assembly sequence corresponding to the second sample to be sequenced; in a case where it is determined that the first similarity and the second similarity are both greater than or both less than the second preset similarity threshold, re-sequencing the second sample to be sequenced.
12. An apparatus for sequencing by hybridization, characterized in that comprises: a first sequencing unit, configured to, in a same sequencing chip, after obtaining at least one first test sequence set by sequencing processing on a first library to be sequenced, obtain at least one second test sequence set by sequencing processing on a second library to be sequenced. a second test sequence set; wherein the first test sequence library comprises at least one first test sequence sample, and the second test sequence library comprises at least one second test sequence sample; a first assembly unit, configured to perform sequence assembly on each test sequence in the first test sequence set to obtain a first assembly sequence set corresponding to each first test sequence sample, and perform sequence assembly on each test sequence in the second test sequence set to obtain a second assembly sequence set corresponding to each second test sequence sample; a first evaluation unit, configured to evaluate each first assembly sequence set to obtain a first assembly sequence corresponding to each first test sequence sample respectively, and evaluate each second assembly sequence set to obtain a second assembly sequence and a third assembly sequence corresponding to each second test sequence sample respectively; a first determination unit, configured to determine an assembly sequence corresponding to each first test sequence sample and an assembly sequence corresponding to each second test sequence sample based on a similarity between at least one first assembly sequence and at least one second assembly sequence and at least one third assembly sequence respectively.
13. An electronic device, comprising: comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-11.
14. A non-transitory computer-readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-11.
15. A computer program product, characterised in that, comprising a computer program, which, when executed by a processor, implements the method of any one of claims 1-11.
Citation Information
Patent Citations
Large-scale genetic typing method based on SLAF-seq (Specific-Locus Amplified Fragment Sequencing) technology
CN103088120A
Method of predicting and then producing a mix of microbiota samples
EP4086337A1